Cordium documentation · Latest

How Cordium Works

Cordium is a free and open source, self-hosted, horizontally scalable, identity-based sandbox platform built on Kubernetes and Octelium. A single Cordium system is called a Cluster and it is defined and addressed by its Octelium Cluster domain (e.g. example.com, octelium.example.com, etc...). In other words, Cordium is installed as a package on top of an existing Octelium Cluster and it reuses all of its infrastructure.

Kubernetes provides the container orchestration layer. Every Workspace run is a Kubernetes pod and every Workspace's storage, Volume and snapshot is backed by Kubernetes persistent volumes and CSI volume snapshots. The Cluster can run on a single-node Kubernetes cluster installed on top of a single cheap VM (see the quick installation guide here) as well as on scalable on-prem or managed Kubernetes clusters. You don't need prior Kubernetes experience to operate Cordium.

Octelium provides the identity and access control infrastructure. Every Cordium User is an Octelium User, every Workspace run is an Octelium Session, and every Cordium component that is reachable from outside the Cluster (i.e. the API, the web portal and the SSH endpoint) is an Octelium Service protected by Octelium's identity-aware proxies and Policies. Secretless access, tunneling, authentication, authorization and auditing are all provided by Octelium. Cordium does not implement its own identity system.

Components

A Cordium installation consists of the following components deployed in every Region of the Cluster:

  • API Server exposes the full Cordium API over gRPC as the Octelium GRPC Service default-cordium.octelium-api (i.e. via octelium-api.<DOMAIN>, the same endpoint as the Octelium API). It implements three gRPC services: MainService for resource management (Spaces, Templates, Workspaces, Volumes, snapshots, Secrets, etc...), WorkspaceService for the operations inside a running Workspace (i.e. command execution, terminals and initialization logs), and ManagementService for the Cluster-wide ClusterConfig. Every call is authenticated and authorized by Octelium before it reaches the API server and it is stateless and horizontally scalable.

  • Portal is the web console that is served as the Octelium WEB Service default.cordium at https://cordium.<DOMAIN>. Besides serving the web application, it multiplexes the browser terminals over a single WebSocket connection, reverse-proxies the Workspace applications at https://<WORKSPACE>.cordium.<DOMAIN> and handles the OAuth2 authorization flows of GitProviders (read more here).

  • SSH Proxy is an Octelium SSH Service named default-ssh.cordium that is implemented by a Cordium-specific identity-aware proxy. It authorizes every SSH connection against the requesting Session's User, the target Workspace's owner and its Space configuration, and then proxies the connection to the embedded SSH server running inside the Workspace. This is what the cordium ssh and cordium cp commands use (read more here).

  • Nocturne is the Workspace controller. It watches Workspace state changes and drives the Kubernetes-level operations: creating and deleting pods and persistent volume claims, taking and restoring volume snapshots, provisioning Volumes, enforcing the inactivity timeouts, and cleaning up after stopped, failed or deleted Workspaces.

  • Workspace Supervisor runs as the single container of every Workspace pod. It bootstraps the sandbox, pulls or builds the Workspace image, enforces the network policy, launches the sandbox container, runs the Octelium authentication proxy and the SSH agent for the sandbox, and reports the state of the Workspace back to the Cluster.

  • Workspace Agent runs inside the sandbox container itself. It sets up the Workspace user, clones the repositories, installs the devcontainer features and dotfiles, configures the git credential helper, runs the lifecycle tasks, manages the terminals and the command executions, and runs the octelium connect process that gives the Workspace its secretless access to Octelium Services.

Workspace Isolation Model

Each Workspace pod uses a layered isolation model:

  • The supervisor container is the only container of the pod. It owns the pod's network namespace, the Workspace's persistent volume and the Kubernetes-level resources limits. It runs the sandbox via rootless Podman (i.e. without any container daemon) under a dedicated, unprivileged user.

  • The sandbox container is the user-facing Workspace. It is launched by the supervisor with the crun runtime, a custom seccomp profile and its own user, PID, mount, IPC, cgroup and network namespaces. The root user inside the sandbox is mapped to an unprivileged user outside of it. The sandbox has no Kubernetes service account token, it cannot see the processes, mounts or cgroups of the supervisor, and its CPU, memory and process counts are enforced by its own nested cgroup. The Octelium and Cordium binaries (i.e. octelium, octeliumctl and cordium) are mounted into it read-only.

  • The network policy is enforced by the supervisor outside of the sandbox using nftables. The supervisor (including its loopback interface), the pod's own addresses, its default gateway and link-local addresses such as the cloud metadata endpoint (i.e. 169.254.169.254) are always unreachable from the sandbox regardless of any configuration. Private networks (e.g. 10.0.0.0/8 and 192.168.0.0/16, which typically include the Kubernetes cluster's pod and service networks and the Kubernetes API) are denied by default while the public internet is allowed by default, and every Workspace can tighten that further with its own egress rules (read more here).

From the User's perspective, however, the sandbox behaves like a full Linux machine: the Workspace user can escalate to root via sudo, packages can be installed, nested containers can be run via rootless Podman, NET_ADMIN operations inside the sandbox's own network namespace are permitted and services can be run as daemons. Root inside the sandbox cannot, however, change any host-wide kernel settings.

The Lifecycle of a Workspace Run

When a User starts a Workspace (e.g. via cordium run, cordium start, the portal or the SDKs), the following happens:

  1. The API server validates the request, resolves the effective resource limits, chooses the Region (i.e. the Region of its storage or mounted Volumes, the requested Region or the User's preferred Region), creates a dedicated Octelium Session for the run, and moves the Workspace to INIT_REQUEST.

  2. Nocturne creates the Workspace pod and its persistent volume claim. A fresh run of a Template that has a ready pre-build, or of a Workspace created from a WorkspaceSnapshot, has its volume restored from the corresponding snapshot. A persistent Workspace that has already run reuses its existing volume.

  3. The supervisor pulls or builds the image (PULLING_IMAGE, BUILDING_IMAGE), applies the network policy and starts the sandbox (STARTING_RUNTIME).

  4. The agent prepares the Workspace (PREPARING): it clones the repositories, merges any repository-defined configuration, installs the devcontainer features and the dotfiles, configures git, starts octelium connect and runs the ON_CREATE tasks (on fresh runs only) followed by the POST_START tasks. Background tasks are started without being waited for.

  5. Once all the foreground tasks complete, the Workspace is RUNNING. With autoStop enabled, the Workspace immediately stops itself at this point.

  6. Upon a stop request, an auto-stop or an inactivity timeout, the PRE_STOP tasks are run, the pod is removed, the Session is deleted and the storage is either kept (persistent) or discarded (ephemeral).

You can follow every transition with cordium logs, the web portal, the WatchWorkspace API or the SDKs' wait helpers.

Identity and Secretless Access

Every Workspace run is an Octelium Session that belongs to the Workspace's owner. That Session carries the identity of the Workspace, its Space and its Template, which is what lets Octelium Policies distinguish a User's Workspace from the same User's laptop, a CI Workspace from a development one, or one Space from another (read more here).

The Session's credentials never enter the sandbox. Instead, the supervisor runs an authentication proxy that is mounted into the sandbox as a Unix socket and exposed via the OCTELIUM_AUTH_PROXY_SOCKET environment variable. The octelium connect process, the cordium and octelium CLIs and the Cordium Agent running inside the Workspace all transparently use that socket. In other words, a process inside a Workspace can call the Cordium API (e.g. to create, start or exec into other Workspaces of the same User) and access Octelium Services as the Workspace's owner, but it can never read or exfiltrate a token.

Secretless access to a protected resource (e.g. a PostgreSQL database, an SSH server, a Kubernetes cluster, an HTTP API or an LLM provider) then works as follows:

process in the Workspace (e.g. psql, curl, ssh, kubectl, claude) │ plain request to the Service's hostname (e.g. pg-staging.payments) ▼ octelium connect (inside the Workspace, authenticated by the Workspace Session) │ WireGuard/QUIC tunnel ▼ Octelium Vigil (the Service's identity-aware proxy) │ 1. identifies the Workspace Session, User, Space and Template │ 2. evaluates the Octelium Policies against the full L7 request │ 3. injects the upstream credential (password, API key, private key, kubeconfig, mTLS certificate) │ 4. emits an OpenTelemetry access log ▼ upstream resource

The credential is stored in the Octelium Cluster as an Octelium Secret and it is only used by Vigil at the moment of the request. It never exists inside the Workspace, so a compromised or manipulated process inside it, including an AI agent, has nothing to leak (read more about secretless access here).

Storage

Every Workspace has its own persistent volume claim that stores its home directory, /workspace and its container layer, which means that packages installed via apt or npm -g persist across the runs of a persistent Workspace. Storage is provisioned by the Kubernetes StorageClass selected by the ClusterConfig rules and it can be restored from CSI volume snapshots, which is what powers both Template pre-builds and WorkspaceSnapshots. Volumes are separate persistent volume claims that are mounted into the sandbox at the requested paths (read more here).

Activity and Timeouts

A running Workspace is stopped automatically once it has been inactive for longer than its inactivity timeout, which is configured in the ClusterConfig and defaults to 30 hours. Terminal input, command executions, SSH connections and requests to the Workspace's applications all count as activity. A Workspace can opt out of the inactivity timeout via its runtime's timeout mode, but only if the ClusterConfig explicitly allows it (read more here).