# Hermes Agent on Codex Pooler

Hermes Agent is an AI assistant from Nous Research that can remember past conversations and build reusable skills. Use it from the terminal or connected messaging channels, with Codex Pooler supplying the models for its conversations. This guide also covers image generation and voice-message transcription through the same Pooler instance.

![Codex Pooler Hermes Agent integration](/codex-pooler-hermes.webp)

## Before you start

- Install [Hermes Agent](https://hermes-agent.nousresearch.com/docs/getting-started/installation) using the official instructions for your operating system.
- Have a Codex Pooler URL reachable from the client.
- Create a [Pool API key](/getting-started/quick-start/) and choose a model available to that Pool.

<a id="recommended-openai-api-provider"></a>

## Configure the connection

Use Hermes' `openai-api` provider with the `/v1` base URL for this setup. Model requests and optional image generation use the Pool API key; operator MCP uses a separate token.

### Config file paths

**API keys and addresses**

| OS | Default config file |
| --- | --- |
| macOS | `~/.hermes/.env` |
| Linux | `~/.hermes/.env` |
| Windows | `%LOCALAPPDATA%\hermes\.env` |

**Model and feature settings**

| OS | Default config file |
| --- | --- |
| macOS | `~/.hermes/config.yaml` |
| Linux | `~/.hermes/config.yaml` |
| Windows | `%LOCALAPPDATA%\hermes\config.yaml` |

**Credentials for the alternate provider setup**

| OS | Default config file |
| --- | --- |
| macOS | `~/.hermes/auth.json` |
| Linux | `~/.hermes/auth.json` |
| Windows | `%LOCALAPPDATA%\hermes\auth.json` |

The native Windows installer uses `%LOCALAPPDATA%\hermes` and sets `HERMES_HOME` to that folder. If `HERMES_HOME` points elsewhere, use that folder for all three files. A WSL installation has its own `~/.hermes` folder.

These are the default locations. On Windows, paste the `%USERPROFILE%`, `%APPDATA%` or `%LOCALAPPDATA%` path into File Explorer's address bar. For a client installed inside WSL, use the Linux paths and commands inside WSL. Keep any custom configuration folder or profile you already use.

Store the Pool API key and base URL in the `.env` file for your installation:

```dotenv title=".env" frame="code"
OPENAI_API_KEY=<pool-api-key>
OPENAI_BASE_URL=https://codex-pooler.example.com/v1
```

Then point `config.yaml` at the same `/v1` base URL and use the Responses transport mode:

```yaml title="config.yaml" frame="code"
model:
  default: gpt-6-sol
  provider: openai-api
  base_url: https://codex-pooler.example.com/v1
  api_mode: codex_responses
  context_length: 828400
  supports_vision: true

agent:
  image_input_mode: native
  api_max_retries: 2
  auto_recovery_cycles: 1

compression:
  threshold: 0.95

auxiliary:
  compression:
    timeout: 900
```

For local setup, use `http://localhost:4000/v1` in both the environment and model configuration.

## Choose a model

Set `model.default` to a model available to your Pool.

Current Codex Pooler releases expose an SDK-readable `context_length` on `/v1/models`, flattened from the selected native raw context window and its effective percentage. Treat that per-model endpoint value as authoritative because provider accounts can temporarily report different catalog ceilings. The `828400` value in this example is a long-profile value for a Pool whose selected model source reports 872000 raw tokens; a short 272000-token profile reports `258400` instead. Match it to the Pool's `/v1/models` result rather than configuring the raw ceiling. Set `model.context_length` explicitly rather than relying on detection: without it Hermes probes `/v1/models` itself, and that probe carries an `Authorization` header only when the provider's API key resolves to a plain string, so a key supplied through a command or another indirect source probes anonymously, is refused, and is remembered as a failure for five minutes. Hermes then logs `Could not detect context length … (probe-down)` on every resolution and silently falls back to its own bundled catalog, which for these models is larger than the window your Pool actually advertises — so compression starts too late and a turn can overflow upstream. An explicit `context_length` removes the probe entirely.

Hermes compresses the conversation when it reaches the lower of two limits: `compression.threshold` times the context window, and `compression.threshold_tokens`, which defaults to 256000 tokens. With the long-profile example, 0.95 of the window is 786980 tokens, so the 256000-token cap applies. With a 258400-token short profile, 0.95 of the window is lower than the cap and applies instead. Keep the default cap: a smaller per-turn context uses less quota. Set `compression.threshold_tokens: null` only if you want Hermes to compress by the ratio alone. Hermes context compression uses its own auxiliary request timeout. Keep `auxiliary.compression.timeout: 900` so large retained contexts can finish instead of cycling through the older 120-second compression budget. This is independent from the optional MCP server `timeout` and from an application output cap.

Hermes generates session titles with a separate auxiliary call. Current Hermes releases send that call with reasoning disabled, which overrides `auxiliary.title_generation.reasoning_effort`, so the key has no effect on titles and you can leave it out; the model's default effort applies.

Hermes fast mode (`agent.service_tier: fast`, `/fast`) sends `service_tier: priority` only when the provider is `openai` on `api.openai.com` or `openai-codex` on `chatgpt.com`. Since the Hermes 0.21 releases, an `openai-api` provider pointed at Codex Pooler never sends it, even though `/fast` and its tip still report fast mode as on. Codex Pooler accepts `service_tier: priority` (and `fast` as an alias) whenever a client sends it, so this is a Hermes routing rule, not a Pooler limitation. In that setup the request log shows `tier default` with no requested tier, because nothing was requested. To get priority processing, use one of these:

- In Hermes, declare a named provider under `providers:` with the same `/v1` base URL, `api_mode: codex_responses`, your Pool API key and `extra_body: {service_tier: priority}`, then point `model.provider` at it. The model's `context_length` and the other settings above still apply.

  ```yaml title="config.yaml" frame="code"
  model:
    default: gpt-6-sol
    provider: codex-pooler
    base_url: https://codex-pooler.example.com/v1
    api_mode: codex_responses
    context_length: 828400
    supports_vision: true

  providers:
    codex-pooler:
      base_url: https://codex-pooler.example.com/v1
      api_mode: codex_responses
      key_env: OPENAI_API_KEY
      extra_body:
        service_tier: priority
  ```

  Hermes adds `extra_body` to the main agent turns only. Its auxiliary calls, such as session titles and context compression, use their own `auxiliary.<task>.extra_body` and request no tier.
- In Codex Pooler, set the API key's [enforced service tier](/operators/api-keys/) to `priority`. This needs no client change, but it only routes when the Pool's model declares that tier.

Either way the request log then shows `tier default` next to `priority requested`, since the Codex backend reports `default` even for priority requests (see [service tiers](/operators/request-logs/#service-tiers)). Priority processing uses included quota faster.

## Retries and fallback

When every account in a Pool has used up its quota, Codex Pooler answers `429` with `usage_limit_reached`, the reset time in `resets_at`, and `Retry-After`. Hermes treats this as a rate limit: before each retry it waits for `Retry-After`, up to 600 seconds, and then ends the turn with the reset time. With the default of three attempts, a turn can wait twice while the Pool cannot recover. `agent.api_max_retries: 2` keeps it to one wait.

A retryable `503`, for example while an account's circuit is open, also carries `Retry-After`. Once a turn's own retries are spent, Hermes retries server errors through a recovery ladder of up to five cycles by default. `agent.auto_recovery_cycles: 1` keeps that to one extra cycle.

If you have a second Pool API key, or another Codex Pooler instance, you can list it under `fallback_providers`. Hermes switches to it once the retries on the primary provider are spent. A fallback key in the same Pool does not help when that Pool is out of quota, because it uses the same accounts.

```yaml title="config.yaml" frame="code"
providers:
  second-pool:
    base_url: https://codex-pooler.example.com/v1
    api_mode: codex_responses
    key_env: SECOND_POOL_API_KEY

fallback_providers:
  - provider: second-pool
    model: gpt-6-luna
```

## Verify the connection

Test the model path with a one-shot prompt:

```bash
hermes -z 'Reply with exactly: hermes openai api ok' --ignore-rules
```

That test should create normal `/v1` traffic using the Pool API key. The MCP entry is separate and should authenticate only with an operator-owned MCP token.

In Codex Pooler's request logs, match the request time, API key, model, and final status to your test. A reply alone does not confirm that the client used your Pooler instance.

## Advanced configuration

### Image generation and speech-to-text

Add these optional blocks to `config.yaml`:

```yaml title="config.yaml" frame="code"
image_gen:
  provider: openai
  model: gpt-image-2.5-flare-medium

stt:
  enabled: true
  provider: openai
  openai:
    model: gpt-4o-transcribe
```

For speech-to-text, also add this value to `.env`:

```dotenv title=".env" frame="code"
STT_OPENAI_BASE_URL=https://codex-pooler.example.com/v1
```

For local setup, use `http://localhost:4000/v1`.

Speech-to-text uses `POST /v1/audio/transcriptions`. Hermes' audio client needs
`STT_OPENAI_BASE_URL`; it does not inherit `model.base_url` or `OPENAI_BASE_URL`.
It uses `OPENAI_API_KEY` unless `VOICE_TOOLS_OPENAI_KEY` overrides it. Keep these
variables visible to the running Hermes process. Set `stt.provider: openai`
explicitly to select this route instead of automatic local or cloud provider
selection, and use `gpt-4o-transcribe` because Hermes' default `whisper-1` is
not supported by Codex Pooler. The `gpt-transcribe` alias is also accepted.
This enables transcription only; text-to-speech and Realtime audio are not supported.

Use a Hermes version whose OpenAI image provider lists GPT Image 2.5. Its
`gpt-image-2.5-flare-medium` preset sends `gpt-image-2.5-flare` to the API
with `quality: medium`; use `gpt-image-2.5-sunburst-medium` for Sunburst.
The `-low`, `-high`, `-xhigh`, and `-max` presets are also accepted and forwarded.
The Codex backend can return a different quality or size than requested; these
presets do not guarantee measured output quality. Older Hermes versions hardcode `gpt-image-2` in this
provider; update Hermes before applying the new preset. See the
[Hermes OpenAI image provider](https://github.com/NousResearch/hermes-agent/blob/b1f003e18633298d549668b8e186af84cca45b76/plugins/image_gen/openai/__init__.py)
for the model-to-quality mapping.

This provider uses the OpenAI SDK environment, so `OPENAI_API_KEY` and
`OPENAI_BASE_URL` must be visible to the running Hermes process, not only to
the shell where you edited the config. `model.base_url` configures Hermes'
text/model provider path; the OpenAI image provider still needs the SDK
environment so image requests go through Codex Pooler's `/v1` surface instead
of OpenAI directly.

### Alternate `openai-codex` credential-pool path

Use the `openai-api` setup above. Hermes can also reach Codex Pooler through its `openai-codex` provider pointed at the Pooler's `/v1` surface, but treat this as an advanced option with limits:

- Always set `model.context_length` from your Pool's `/v1/models` result. The Codex picker expects a native catalog and cannot read that endpoint's OpenAI-compatible format; do not use it to select your Pool's models or context limits.
- With Hermes 0.21.5, the unconfigured context probe, `/model` picker and `openai-codex` image plugin can send the Pool API key to `chatgpt.com`. Keep the context limit explicit, avoid that picker and use `image_gen.provider: openai`. If the key was sent outside your Pooler instance, rotate it.
- Web search switches to the provider's hosted `web_search` tool.
- Hermes never sends `service_tier` on this provider, and the named-provider option above does not apply to it.
- Image generation and speech-to-text still need the `OPENAI_*` variables pointed at the Pooler's `/v1` surface.
- Hermes sends a `session_id` header here. An ingress proxy that drops headers containing underscores removes it.

This provider supplies a session header that Codex Pooler's websocket bridge can use when owner forwarding is enabled. A body `prompt_cache_key` alone does not enable the bridge. Keep the base URL on `/v1`: the recommended setup already supports prompt-cache locality, so switching providers or routes is unnecessary for that purpose.

Keep the endpoint value in the environment:

```dotenv title=".env" frame="code"
HERMES_CODEX_BASE_URL=https://codex-pooler.example.com/v1
```

Set the provider in `config.yaml`. `context_length` is required here:

```yaml title="config.yaml" frame="code"
model:
  default: gpt-6-sol
  provider: openai-codex
  base_url: https://codex-pooler.example.com/v1
  context_length: 828400
  supports_vision: true

agent:
  image_input_mode: native
  api_max_retries: 2
  auto_recovery_cycles: 1

compression:
  threshold: 0.95

auxiliary:
  compression:
    timeout: 900
```

Hermes treats `openai-codex` as an OAuth provider by default, so add a Pool API key credential ahead of any device-code credential, and keep the credential's `base_url` on `/v1`. This example shows only placeholders. Don't paste a real key into public docs or shared files.

```json title="auth.json" frame="code"
{
  "active_provider": "openai-codex",
  "credential_pool": {
    "openai-codex": [
      {
        "label": "codex-pooler",
        "auth_type": "api_key",
        "priority": -10,
        "source": "manual",
        "access_token": "<pool-api-key>",
        "base_url": "https://codex-pooler.example.com/v1"
      }
    ]
  }
}
```

## Operator MCP (optional)

Add a separate operator token to `.env`:

```dotenv title=".env" frame="code"
CODEX_POOLER_MCP_KEY=<operator-mcp-token>
```

Add the following top-level block to `config.yaml` with either provider setup:

```yaml title="config.yaml" frame="code"
# Optional operator-only MCP metadata add-on. Omit for model/runtime use.
mcp_servers:
  codex_pooler:
    url: https://codex-pooler.example.com/mcp
    headers:
      Authorization: "Bearer ${CODEX_POOLER_MCP_KEY}"
    enabled: true
    timeout: 120
    connect_timeout: 15
```

For local MCP setup, use `http://localhost:4000/mcp`.

Remote HTTP MCP servers require Hermes' `mcp` extra. If `hermes mcp test codex_pooler` reports `mcp.client.streamable_http is not available`, install MCP support into the Hermes environment, following the [Hermes MCP Integration docs](https://hermes-agent.nousresearch.com/docs/user-guide/features/mcp), and rerun the test.

## Troubleshooting

If text requests work but image generation fails with `invalid_api_key`, check
the environment of the long-running Hermes process or gateway service first.
It may not have loaded `OPENAI_BASE_URL`, so the OpenAI SDK image client may be
using OpenAI's default endpoint instead of Codex Pooler.

## Compatibility notes

Hermes model requests use Codex Pooler's narrow OpenAI-compatible `/v1` support for selected SDK routes. Codex Pooler doesn't provide full OpenAI API parity.

`GET /v1/responses` is narrow Responses websocket compatibility, not `/v1/realtime` support. `/v1/realtime` and OpenAI Realtime SDK websocket or session routes are unsupported.

The operator MCP endpoint is rooted at `/mcp`. It uses an operator-owned MCP token, not a Pool API key.