# Secretless LLM Access

> Cordium documentation. Canonical page: <https://octelium.com/docs/cordium/latest/examples/ai/secretless-llm>.

AI agents are the workloads that are the most likely to leak their own credentials: they read their environment, write logs, generate code, and can be manipulated by prompt injection. This example shows how to give every agent running in a Cordium *Workspace* access to LLM providers without any provider API key inside the *Workspace*, while controlling which *Workspaces* can use which models, giving every agent run its own token budget, and auditing every inference request.

The building block is the Octelium `LLM` *Service*: an identity-aware AI gateway that understands the OpenAI, Anthropic and Gemini APIs, injects the provider credentials, and authorizes every request against the identity of the *Workspace* that sends it (read more [here](https://octelium.com/docs/octelium/latest/management/core/service/llm/overview.md)).

## The LLM Services

A *Cluster* administrator stores the provider API keys as Octelium *Secrets* and creates one `LLM` *Service* per provider:

```bash
octeliumctl create secret anthropic-api-key
octeliumctl create secret openai-api-key
```

```yaml
kind: Service
metadata:
  name: anthropic
spec:
  mode: LLM
  config:
    upstream:
      url: https://api.anthropic.com
    llm:
      protocol: ANTHROPIC
      auth:
        custom:
          header: x-api-key
          value:
            fromSecret: anthropic-api-key
---
kind: Service
metadata:
  name: openai
spec:
  mode: LLM
  config:
    upstream:
      url: https://api.openai.com
    llm:
      protocol: OPENAI
      auth:
        bearer:
          fromSecret: openai-api-key
```

From inside any authorized *Workspace*, the providers are now reachable at `http://anthropic` and `http://openai`.

## Configuring Clients

Clients only need to point their base URL to the *Service*. Most clients refuse to start without an API key, which is why they are given a placeholder: Octelium never forwards the client's own credentials to the provider. Here is how to configure the most common clients:

| Client         | Configuration                                                                                                                                                             |
| -------------- | ------------------------------------------------------------------------------------------------------------------------------------------------------------------------- |
| Claude Code    | `ANTHROPIC_BASE_URL=http://anthropic` and `ANTHROPIC_API_KEY=<PLACEHOLDER>` (read more [here](https://octelium.com/docs/cordium/latest/examples/ai/claude-code.md)).      |
| Codex          | A custom model provider with `base_url = "http://openai/v1"` in `~/.codex/config.toml` (read more [here](https://octelium.com/docs/cordium/latest/examples/ai/codex.md)). |
| OpenCode       | `baseURL` in the provider options of `opencode.json` (read more [here](https://octelium.com/docs/cordium/latest/examples/ai/opencode.md)).                                |
| Anthropic SDKs | `base_url="http://anthropic"` (Python) or `baseURL: "http://anthropic"` (TypeScript).                                                                                     |
| OpenAI SDKs    | `base_url="http://openai/v1"` (Python) or `baseURL: "http://openai/v1"` (TypeScript).                                                                                     |

Here is an example via the Anthropic Python SDK:

```python
import anthropic

client = anthropic.Anthropic(base_url="http://anthropic", api_key="injected-by-octelium")

message = client.messages.create(
    model="claude-opus-5-5",
    max_tokens=1024,
    messages=[{"role": "user", "content": "Summarize the CHANGELOG.md of this repository."}],
)
print(message.content[0].text)
```

And here is a quick test via `curl`:

```bash
curl -s http://anthropic/v1/messages \
  -H "content-type: application/json" \
  -H "anthropic-version: 2023-06-01" \
  -d '{"model": "claude-haiku-5-5", "max_tokens": 64, "messages": [{"role": "user", "content": "Say hi"}]}'
```

## Authorizing Workspaces

Like any other Octelium *Service*, an `LLM` *Service* is only accessible to the *Users* allowed by a *Policy*. Since every *Workspace* run is an Octelium *Session* that carries the identity of its *Workspace*, *Space* and *Template* (read more [here](https://octelium.com/docs/cordium/latest/workspaces/secretless.md#policies-for-workspaces)), you can allow the LLM *Services* only from *Workspaces*, and only from the *Spaces* that need them. Here is an example of a *Policy* that lets the members of the `engineering` *Group* use both providers from the *Workspaces* of the `payments` and `storefront` *Spaces*, but not from their laptops:

```yaml
kind: Policy
metadata:
  name: llm-from-workspaces
spec:
  rules:
    - effect: ALLOW
      condition:
        all:
          of:
            - match: ctx.service.metadata.name in ["anthropic.default", "openai.default"]
            - match: '"ext" in ctx.session.status && "cordium" in ctx.session.status.ext'
            - match: ctx.session.status.ext.cordium.spaceRef.name in ["payments.cordium", "storefront.cordium"]
---
kind: Group
metadata:
  name: engineering
spec:
  authorization:
    policies: ["cordium-users", "llm-from-workspaces"]
```

## Restricting Models

LLM *Services* expose the requested model to *Policies* via `ctx.request.llm.model`. Here is an example that only allows the Claude Haiku and Sonnet models to the *Workspaces* of the `claude-agent` *Template*, which run unattended and at scale, while every other *Workspace* can use any model:

```yaml
kind: Policy
metadata:
  name: agents-models
spec:
  rules:
    - effect: DENY
      condition:
        all:
          of:
            - match: ctx.service.metadata.name == "anthropic.default"
            - match: '"ext" in ctx.session.status && "cordium" in ctx.session.status.ext'
            - match: ctx.session.status.ext.cordium.templateRef.name == "claude-agent.payments.cordium"
            - match: '!(ctx.request.llm.model in ["claude-haiku-5-5", "claude-sonnet-5-5"])'
```

Attach it to the *Users* or *Groups* that run the agents, or make it global via the Octelium `ClusterConfig` (read more [here](https://octelium.com/docs/octelium/latest/management/core/policy.md#global-policies)).

## Token Budgets per Run

Every *Workspace* run has its own *Session*, which makes Octelium's per-*Session* token rate limits a natural per-run budget for agents: a runaway agent loop stops consuming tokens as soon as its run exhausts its budget, without affecting any other run. Here is the `anthropic` *Service* with a budget of 2 million tokens per *Workspace* run and per day, in addition to a per-*User* hourly budget (read more about LLM plugins [here](https://octelium.com/docs/octelium/latest/management/core/service/llm/plugins.md)):

```yaml
kind: Service
metadata:
  name: anthropic
spec:
  mode: LLM
  config:
    upstream:
      url: https://api.anthropic.com
    llm:
      protocol: ANTHROPIC
      auth:
        custom:
          header: x-api-key
          value:
            fromSecret: anthropic-api-key
      plugins:
        - name: workspace-run-budget
          condition:
            all:
              of:
                - match: ctx.request.llm.operation == "GENERATE"
                - match: '"ext" in ctx.session.status && "cordium" in ctx.session.status.ext'
          tokenRateLimit:
            scope: TOTAL
            key:
              perSession: true
            limit: 2000000
            window:
              days: 1
            defaultOutputTokens: 8192
            denyMessage: This Workspace run has exhausted its token budget
        - name: user-hourly-budget
          condition:
            match: ctx.request.llm.operation == "GENERATE"
          tokenRateLimit:
            scope: TOTAL
            key:
              perUser: true
            limit: 5000000
            window:
              hours: 1
            defaultOutputTokens: 8192
```

## Auditing

Every inference request is logged by Octelium with its model, token usage and outcome, along with the identity of the *User* and the *Session* of the *Workspace* that sent it (read more [here](https://octelium.com/docs/octelium/latest/management/core/service/llm/visibility.md)). Since the *Session* of a *Workspace* run identifies the *Workspace*, its *Space* and its *Template*, you can attribute the LLM usage and cost of every agent run, *Template* and team precisely.
