Cordium documentation · Latest
Secretless LLM Access
AI agents are the workloads that are the most likely to leak their own credentials: they read their environment, write logs, generate code, and can be manipulated by prompt injection. This example shows how to give every agent running in a Cordium Workspace access to LLM providers without any provider API key inside the Workspace, while controlling which Workspaces can use which models, giving every agent run its own token budget, and auditing every inference request.
The building block is the Octelium LLM Service: an identity-aware AI gateway that understands the OpenAI, Anthropic and Gemini APIs, injects the provider credentials, and authorizes every request against the identity of the Workspace that sends it (read more here).
The LLM Services
A Cluster administrator stores the provider API keys as Octelium Secrets and creates one LLM Service per provider:
From inside any authorized Workspace, the providers are now reachable at http://anthropic and http://openai.
Configuring Clients
Clients only need to point their base URL to the Service. Most clients refuse to start without an API key, which is why they are given a placeholder: Octelium never forwards the client's own credentials to the provider. Here is how to configure the most common clients:
| Client | Configuration |
| Claude Code | ANTHROPIC_BASE_URL=http://anthropic and ANTHROPIC_API_KEY=<PLACEHOLDER> (read more here). |
| Codex | A custom model provider with base_url = "http://openai/v1" in ~/.codex/config.toml (read more here). |
| OpenCode | baseURL in the provider options of opencode.json (read more here). |
| Anthropic SDKs | base_url="http://anthropic" (Python) or baseURL: "http://anthropic" (TypeScript). |
| OpenAI SDKs | base_url="http://openai/v1" (Python) or baseURL: "http://openai/v1" (TypeScript). |
Here is an example via the Anthropic Python SDK:
And here is a quick test via curl:
Authorizing Workspaces
Like any other Octelium Service, an LLM Service is only accessible to the Users allowed by a Policy. Since every Workspace run is an Octelium Session that carries the identity of its Workspace, Space and Template (read more here), you can allow the LLM Services only from Workspaces, and only from the Spaces that need them. Here is an example of a Policy that lets the members of the engineering Group use both providers from the Workspaces of the payments and storefront Spaces, but not from their laptops:
Restricting Models
LLM Services expose the requested model to Policies via ctx.request.llm.model. Here is an example that only allows the Claude Haiku and Sonnet models to the Workspaces of the claude-agent Template, which run unattended and at scale, while every other Workspace can use any model:
Attach it to the Users or Groups that run the agents, or make it global via the Octelium ClusterConfig (read more here).
Token Budgets per Run
Every Workspace run has its own Session, which makes Octelium's per-Session token rate limits a natural per-run budget for agents: a runaway agent loop stops consuming tokens as soon as its run exhausts its budget, without affecting any other run. Here is the anthropic Service with a budget of 2 million tokens per Workspace run and per day, in addition to a per-User hourly budget (read more about LLM plugins here):
Auditing
Every inference request is logged by Octelium with its model, token usage and outcome, along with the identity of the User and the Session of the Workspace that sent it (read more here). Since the Session of a Workspace run identifies the Workspace, its Space and its Template, you can attribute the LLM usage and cost of every agent run, Template and team precisely.