Cordium documentation · Latest

Parallel Agents

Coding agents are non-deterministic: the same task given to different agents, different models, or even the same model twice often produces solutions of very different quality. This example fans one task out to several agents running in parallel, each in its own Workspace forked from the exact same state, and compares their results by running the test suite. It is the basis of "best-of-N" workflows and of evaluating agents and models on your own codebase.

WorkspaceSnapshots make the fan-out cheap and fair: a developer prepares a Workspace once (e.g. the repository at the right commit, its dependencies downloaded, its build cache warm, and a failing test that reproduces the bug), snapshots it, and every agent starts from an identical copy of it within seconds (read more about snapshots here).

The Template

The following Template of the payments Space installs both Claude Code and Codex, and configures both of them to use the anthropic and openai Octelium LLM Services, so that no API key is present in any Workspace (read more here):

spec: image: registry: url: golang:1.26-bookworm repository: url: https://github.com/acme-corp/payments-api authentication: http: username: x-access-token password: fromSecret: github-read-token.payments.cordium runtime: envVars: - key: ANTHROPIC_BASE_URL value: http://anthropic - key: ANTHROPIC_API_KEY value: injected-by-octelium - key: OPENAI_API_KEY value: injected-by-octelium tasks: - name: install-agents type: ON_CREATE runAsRoot: true onFailure: ON_FAILURE_ABORT run: | apt-get update apt-get install -y --no-install-recommends nodejs npm npm install -g @openai/codex - name: install-claude-code type: ON_CREATE workingDir: /workspace onFailure: ON_FAILURE_ABORT run: | curl -fsSL https://claude.ai/install.sh | bash mkdir -p "$HOME/.codex" cat > "$HOME/.codex/config.toml" <<'EOF' model_provider = "octelium" [model_providers.octelium] name = "OpenAI via Octelium" base_url = "http://openai/v1" env_key = "OPENAI_API_KEY" wire_api = "responses" EOF - name: go-mod-download type: ON_CREATE workingDir: /workspace/repo run: go mod download && go build ./... limit: cpu: millicores: 4000 memory: megabytes: 8192 storage: megabytes: 20000
cordium create template agents.payments.cordium --file agents.yaml cordium build agents.payments.cordium

Preparing the Baseline

A developer creates a persistent Workspace from the Template, checks out the commit to work on, and reproduces the bug with a failing test:

cordium run --template agents.payments.cordium
cd /workspace/repo git checkout -b fix/1842 origin/main # Add TestWebhookRetryUsesClock to internal/webhooks/retry_test.go go test ./internal/webhooks/ -run TestWebhookRetryUsesClock
--- FAIL: TestWebhookRetryUsesClock (0.00s) retry_test.go:88: backoff computed from the wall clock: got 2026-10-09T09:14:21Z FAIL

Then, they snapshot the Workspace. Stopping it first gives a clean snapshot, while a running Workspace can be snapshotted too:

cordium stop <WORKSPACE> cordium create snapshot issue-1842 --workspace <WORKSPACE>

The Orchestrator

The following Go program forks the issue-1842 snapshot into one ephemeral Workspace per agent, runs the agents in parallel with the same prompt, runs the whole test suite in each Workspace, saves each agent's diff locally, and prints a comparison:

package main import ( "context" "fmt" "os" "sync" "time" cordium "github.com/octelium/cordium/cordium-go" ) const prompt = `The test TestWebhookRetryUsesClock in internal/webhooks fails because the retry backoff reads the wall clock. Fix the implementation so that it uses the injected clock, without changing the test. Run "go test ./..." before finishing. Do not commit.` type agent struct { name string command string } var agents = []agent{ {"claude-opus", `$HOME/.local/bin/claude -p "$PROMPT" --model claude-opus-5-5 --dangerously-skip-permissions < /dev/null`}, {"claude-sonnet", `$HOME/.local/bin/claude -p "$PROMPT" --model claude-sonnet-5-5 --dangerously-skip-permissions < /dev/null`}, {"codex", `codex exec --sandbox danger-full-access "$PROMPT" < /dev/null`}, } func main() { ctx, cancel := context.WithTimeout(context.Background(), 90*time.Minute) defer cancel() c, err := cordium.New(ctx) if err != nil { fmt.Fprintln(os.Stderr, err) os.Exit(1) } defer c.Close() if _, err := c.Snapshots().WaitUntilReady(ctx, "issue-1842"); err != nil { fmt.Fprintln(os.Stderr, err) os.Exit(1) } results := make([]string, len(agents)) var wg sync.WaitGroup for i, a := range agents { wg.Add(1) go func() { defer wg.Done() results[i] = attempt(ctx, c, a) }() } wg.Wait() for _, line := range results { fmt.Println(line) } } func attempt(ctx context.Context, c *cordium.Client, a agent) string { ws, err := c.Workspaces().Run(ctx, cordium.FromSnapshot("issue-1842"), cordium.Ephemeral(), cordium.WithDisplayName("issue-1842: "+a.name), ) if ws != nil { defer ws.Delete(context.WithoutCancel(ctx)) } if err != nil { return fmt.Sprintf("%-14s could not start: %v", a.name, err) } start := time.Now() run, err := ws.Exec(ctx, a.command, cordium.WithWorkingDir("/workspace/repo"), cordium.WithExecEnv("PROMPT", prompt), cordium.WithExecTimeout(45*time.Minute), ) if err != nil { return fmt.Sprintf("%-14s could not run: %v", a.name, err) } elapsed := time.Since(start).Round(time.Second) tests, err := ws.Exec(ctx, "go test ./...", cordium.WithWorkingDir("/workspace/repo"), cordium.WithExecTimeout(20*time.Minute), ) if err != nil { return fmt.Sprintf("%-14s could not test: %v", a.name, err) } diff, err := ws.Exec(ctx, "git add -A && git diff --cached", cordium.WithWorkingDir("/workspace/repo"), ) if err != nil { return fmt.Sprintf("%-14s could not diff: %v", a.name, err) } if err := os.WriteFile(a.name+".patch", diff.Stdout, 0o644); err != nil { return fmt.Sprintf("%-14s could not save the diff: %v", a.name, err) } return fmt.Sprintf("%-14s agent exit %d in %s, tests passed: %t, diff: %d bytes", a.name, run.ExitCode, elapsed, tests.Success(), len(diff.Stdout)) }
export OCTELIUM_DOMAIN=<DOMAIN> go run .
claude-opus agent exit 0 in 3m41s, tests passed: true, diff: 2817 bytes claude-sonnet agent exit 0 in 2m12s, tests passed: true, diff: 1904 bytes codex agent exit 0 in 4m05s, tests passed: false, diff: 3360 bytes

The developer then reviews the patches of the successful attempts and applies the best one to their own Workspace via git apply. Here are a few notes about this program:

  • The prompt is passed as an environment variable of each exec session, so it never needs to be quoted for the shell.

  • The agents' standard input is redirected from /dev/null, since exec sessions do not signal the end of the standard input.

  • Every Workspace is ephemeral and deleted after its attempt, while the snapshot remains available for more attempts. Delete it via cordium delete snapshot issue-1842 once you are done.

  • Each Workspace run is a separate Octelium Session, which means that the LLM token usage of every attempt is attributed and can be budgeted separately (read more here).