Cordium documentation · Latest

Secretless LLM Access

AI agents are the workloads that are the most likely to leak their own credentials: they read their environment, write logs, generate code, and can be manipulated by prompt injection. This example shows how to give every agent running in a Cordium Workspace access to LLM providers without any provider API key inside the Workspace, while controlling which Workspaces can use which models, giving every agent run its own token budget, and auditing every inference request.

The building block is the Octelium LLM Service: an identity-aware AI gateway that understands the OpenAI, Anthropic and Gemini APIs, injects the provider credentials, and authorizes every request against the identity of the Workspace that sends it (read more here).

The LLM Services

A Cluster administrator stores the provider API keys as Octelium Secrets and creates one LLM Service per provider:

octeliumctl create secret anthropic-api-key octeliumctl create secret openai-api-key
kind: Service metadata: name: anthropic spec: mode: LLM config: upstream: url: https://api.anthropic.com llm: protocol: ANTHROPIC auth: custom: header: x-api-key value: fromSecret: anthropic-api-key --- kind: Service metadata: name: openai spec: mode: LLM config: upstream: url: https://api.openai.com llm: protocol: OPENAI auth: bearer: fromSecret: openai-api-key

From inside any authorized Workspace, the providers are now reachable at http://anthropic and http://openai.

Configuring Clients

Clients only need to point their base URL to the Service. Most clients refuse to start without an API key, which is why they are given a placeholder: Octelium never forwards the client's own credentials to the provider. Here is how to configure the most common clients:

ClientConfiguration
Claude CodeANTHROPIC_BASE_URL=http://anthropic and ANTHROPIC_API_KEY=<PLACEHOLDER> (read more here).
CodexA custom model provider with base_url = "http://openai/v1" in ~/.codex/config.toml (read more here).
OpenCodebaseURL in the provider options of opencode.json (read more here).
Anthropic SDKsbase_url="http://anthropic" (Python) or baseURL: "http://anthropic" (TypeScript).
OpenAI SDKsbase_url="http://openai/v1" (Python) or baseURL: "http://openai/v1" (TypeScript).

Here is an example via the Anthropic Python SDK:

import anthropic client = anthropic.Anthropic(base_url="http://anthropic", api_key="injected-by-octelium") message = client.messages.create( model="claude-opus-5-5", max_tokens=1024, messages=[{"role": "user", "content": "Summarize the CHANGELOG.md of this repository."}], ) print(message.content[0].text)

And here is a quick test via curl:

curl -s http://anthropic/v1/messages \ -H "content-type: application/json" \ -H "anthropic-version: 2023-06-01" \ -d '{"model": "claude-haiku-5-5", "max_tokens": 64, "messages": [{"role": "user", "content": "Say hi"}]}'

Authorizing Workspaces

Like any other Octelium Service, an LLM Service is only accessible to the Users allowed by a Policy. Since every Workspace run is an Octelium Session that carries the identity of its Workspace, Space and Template (read more here), you can allow the LLM Services only from Workspaces, and only from the Spaces that need them. Here is an example of a Policy that lets the members of the engineering Group use both providers from the Workspaces of the payments and storefront Spaces, but not from their laptops:

kind: Policy metadata: name: llm-from-workspaces spec: rules: - effect: ALLOW condition: all: of: - match: ctx.service.metadata.name in ["anthropic.default", "openai.default"] - match: '"ext" in ctx.session.status && "cordium" in ctx.session.status.ext' - match: ctx.session.status.ext.cordium.spaceRef.name in ["payments.cordium", "storefront.cordium"] --- kind: Group metadata: name: engineering spec: authorization: policies: ["cordium-users", "llm-from-workspaces"]

Restricting Models

LLM Services expose the requested model to Policies via ctx.request.llm.model. Here is an example that only allows the Claude Haiku and Sonnet models to the Workspaces of the claude-agent Template, which run unattended and at scale, while every other Workspace can use any model:

kind: Policy metadata: name: agents-models spec: rules: - effect: DENY condition: all: of: - match: ctx.service.metadata.name == "anthropic.default" - match: '"ext" in ctx.session.status && "cordium" in ctx.session.status.ext' - match: ctx.session.status.ext.cordium.templateRef.name == "claude-agent.payments.cordium" - match: '!(ctx.request.llm.model in ["claude-haiku-5-5", "claude-sonnet-5-5"])'

Attach it to the Users or Groups that run the agents, or make it global via the Octelium ClusterConfig (read more here).

Token Budgets per Run

Every Workspace run has its own Session, which makes Octelium's per-Session token rate limits a natural per-run budget for agents: a runaway agent loop stops consuming tokens as soon as its run exhausts its budget, without affecting any other run. Here is the anthropic Service with a budget of 2 million tokens per Workspace run and per day, in addition to a per-User hourly budget (read more about LLM plugins here):

kind: Service metadata: name: anthropic spec: mode: LLM config: upstream: url: https://api.anthropic.com llm: protocol: ANTHROPIC auth: custom: header: x-api-key value: fromSecret: anthropic-api-key plugins: - name: workspace-run-budget condition: all: of: - match: ctx.request.llm.operation == "GENERATE" - match: '"ext" in ctx.session.status && "cordium" in ctx.session.status.ext' tokenRateLimit: scope: TOTAL key: perSession: true limit: 2000000 window: days: 1 defaultOutputTokens: 8192 denyMessage: This Workspace run has exhausted its token budget - name: user-hourly-budget condition: match: ctx.request.llm.operation == "GENERATE" tokenRateLimit: scope: TOTAL key: perUser: true limit: 5000000 window: hours: 1 defaultOutputTokens: 8192

Auditing

Every inference request is logged by Octelium with its model, token usage and outcome, along with the identity of the User and the Session of the Workspace that sent it (read more here). Since the Session of a Workspace run identifies the Workspace, its Space and its Template, you can attribute the LLM usage and cost of every agent run, Template and team precisely.