Skip to content

Hermes Agent on Codex Pooler

Hermes Agent is an AI assistant from Nous Research that can remember past conversations and build reusable skills. Use it from the terminal or connected messaging channels, with Codex Pooler supplying the models for its conversations. This guide also covers image generation and voice-message transcription through the same Pooler instance.

Codex Pooler Hermes Agent integration

  • Install Hermes Agent using the official instructions for your operating system.
  • Have a Codex Pooler URL reachable from the client.
  • Create a Pool API key and choose a model available to that Pool.

Use Hermes’ openai-api provider with the /v1 base URL for this setup. Model requests and optional image generation use the Pool API key; operator MCP uses a separate token.

API keys and addresses

OS Default config file
macOS ~/.hermes/.env
Linux ~/.hermes/.env
Windows %LOCALAPPDATA%\hermes\.env

Model and feature settings

OS Default config file
macOS ~/.hermes/config.yaml
Linux ~/.hermes/config.yaml
Windows %LOCALAPPDATA%\hermes\config.yaml

Credentials for the alternate provider setup

OS Default config file
macOS ~/.hermes/auth.json
Linux ~/.hermes/auth.json
Windows %LOCALAPPDATA%\hermes\auth.json

The native Windows installer uses %LOCALAPPDATA%\hermes and sets HERMES_HOME to that folder. If HERMES_HOME points elsewhere, use that folder for all three files. A WSL installation has its own ~/.hermes folder.

These are the default locations. On Windows, paste the %USERPROFILE%, %APPDATA% or %LOCALAPPDATA% path into File Explorer’s address bar. For a client installed inside WSL, use the Linux paths and commands inside WSL. Keep any custom configuration folder or profile you already use.

Store the Pool API key and base URL in the .env file for your installation:

.env
OPENAI_API_KEY=<pool-api-key>
OPENAI_BASE_URL=https://codex-pooler.example.com/v1

Then point config.yaml at the same /v1 base URL and use the Responses transport mode:

config.yaml
model:
default: gpt-6-sol
provider: openai-api
base_url: https://codex-pooler.example.com/v1
api_mode: codex_responses
context_length: 828400
supports_vision: true
agent:
image_input_mode: native
api_max_retries: 2
auto_recovery_cycles: 1
compression:
threshold: 0.95
auxiliary:
compression:
timeout: 900

For local setup, use http://localhost:4000/v1 in both the environment and model configuration.

Set model.default to a model available to your Pool.

Current Codex Pooler releases expose an SDK-readable context_length on /v1/models, flattened from the selected native raw context window and its effective percentage. Treat that per-model endpoint value as authoritative because provider accounts can temporarily report different catalog ceilings. The 828400 value in this example is a long-profile value for a Pool whose selected model source reports 872000 raw tokens; a short 272000-token profile reports 258400 instead. Match it to the Pool’s /v1/models result rather than configuring the raw ceiling. Set model.context_length explicitly rather than relying on detection: without it Hermes probes /v1/models itself, and that probe carries an Authorization header only when the provider’s API key resolves to a plain string, so a key supplied through a command or another indirect source probes anonymously, is refused, and is remembered as a failure for five minutes. Hermes then logs Could not detect context length … (probe-down) on every resolution and silently falls back to its own bundled catalog, which for these models is larger than the window your Pool actually advertises — so compression starts too late and a turn can overflow upstream. An explicit context_length removes the probe entirely.

Hermes compresses the conversation when it reaches the lower of two limits: compression.threshold times the context window, and compression.threshold_tokens, which defaults to 256000 tokens. With the long-profile example, 0.95 of the window is 786980 tokens, so the 256000-token cap applies. With a 258400-token short profile, 0.95 of the window is lower than the cap and applies instead. Keep the default cap: a smaller per-turn context uses less quota. Set compression.threshold_tokens: null only if you want Hermes to compress by the ratio alone. Hermes context compression uses its own auxiliary request timeout. Keep auxiliary.compression.timeout: 900 so large retained contexts can finish instead of cycling through the older 120-second compression budget. This is independent from the optional MCP server timeout and from an application output cap.

Hermes generates session titles with a separate auxiliary call. Current Hermes releases send that call with reasoning disabled, which overrides auxiliary.title_generation.reasoning_effort, so the key has no effect on titles and you can leave it out; the model’s default effort applies.

Hermes fast mode (agent.service_tier: fast, /fast) sends service_tier: priority only when the provider is openai on api.openai.com or openai-codex on chatgpt.com. Since the Hermes 0.21 releases, an openai-api provider pointed at Codex Pooler never sends it, even though /fast and its tip still report fast mode as on. Codex Pooler accepts service_tier: priority (and fast as an alias) whenever a client sends it, so this is a Hermes routing rule, not a Pooler limitation. In that setup the request log shows tier default with no requested tier, because nothing was requested. To get priority processing, use one of these:

  • In Hermes, declare a named provider under providers: with the same /v1 base URL, api_mode: codex_responses, your Pool API key and extra_body: {service_tier: priority}, then point model.provider at it. The model’s context_length and the other settings above still apply.

    config.yaml
    model:
    default: gpt-6-sol
    provider: codex-pooler
    base_url: https://codex-pooler.example.com/v1
    api_mode: codex_responses
    context_length: 828400
    supports_vision: true
    providers:
    codex-pooler:
    base_url: https://codex-pooler.example.com/v1
    api_mode: codex_responses
    key_env: OPENAI_API_KEY
    extra_body:
    service_tier: priority

    Hermes adds extra_body to the main agent turns only. Its auxiliary calls, such as session titles and context compression, use their own auxiliary.<task>.extra_body and request no tier.

  • In Codex Pooler, set the API key’s enforced service tier to priority. This needs no client change, but it only routes when the Pool’s model declares that tier.

Either way the request log then shows tier default next to priority requested, since the Codex backend reports default even for priority requests (see service tiers). Priority processing uses included quota faster.

When every account in a Pool has used up its quota, Codex Pooler answers 429 with usage_limit_reached, the reset time in resets_at, and Retry-After. Hermes treats this as a rate limit: before each retry it waits for Retry-After, up to 600 seconds, and then ends the turn with the reset time. With the default of three attempts, a turn can wait twice while the Pool cannot recover. agent.api_max_retries: 2 keeps it to one wait.

A retryable 503, for example while an account’s circuit is open, also carries Retry-After. Once a turn’s own retries are spent, Hermes retries server errors through a recovery ladder of up to five cycles by default. agent.auto_recovery_cycles: 1 keeps that to one extra cycle.

If you have a second Pool API key, or another Codex Pooler instance, you can list it under fallback_providers. Hermes switches to it once the retries on the primary provider are spent. A fallback key in the same Pool does not help when that Pool is out of quota, because it uses the same accounts.

config.yaml
providers:
second-pool:
base_url: https://codex-pooler.example.com/v1
api_mode: codex_responses
key_env: SECOND_POOL_API_KEY
fallback_providers:
- provider: second-pool
model: gpt-6-luna

Test the model path with a one-shot prompt:

Terminal window
hermes -z 'Reply with exactly: hermes openai api ok' --ignore-rules

That test should create normal /v1 traffic using the Pool API key. The MCP entry is separate and should authenticate only with an operator-owned MCP token.

In Codex Pooler’s request logs, match the request time, API key, model, and final status to your test. A reply alone does not confirm that the client used your Pooler instance.

Add these optional blocks to config.yaml:

config.yaml
image_gen:
provider: openai
model: gpt-image-2.5-flare-medium
stt:
enabled: true
provider: openai
openai:
model: gpt-4o-transcribe

For speech-to-text, also add this value to .env:

.env
STT_OPENAI_BASE_URL=https://codex-pooler.example.com/v1

For local setup, use http://localhost:4000/v1.

Speech-to-text uses POST /v1/audio/transcriptions. Hermes’ audio client needs STT_OPENAI_BASE_URL; it does not inherit model.base_url or OPENAI_BASE_URL. It uses OPENAI_API_KEY unless VOICE_TOOLS_OPENAI_KEY overrides it. Keep these variables visible to the running Hermes process. Set stt.provider: openai explicitly to select this route instead of automatic local or cloud provider selection, and use gpt-4o-transcribe because Hermes’ default whisper-1 is not supported by Codex Pooler. The gpt-transcribe alias is also accepted. This enables transcription only; text-to-speech and Realtime audio are not supported.

Use a Hermes version whose OpenAI image provider lists GPT Image 2.5. Its gpt-image-2.5-flare-medium preset sends gpt-image-2.5-flare to the API with quality: medium; use gpt-image-2.5-sunburst-medium for Sunburst. The -low, -high, -xhigh, and -max presets are also accepted and forwarded. The Codex backend can return a different quality or size than requested; these presets do not guarantee measured output quality. Older Hermes versions hardcode gpt-image-2 in this provider; update Hermes before applying the new preset. See the Hermes OpenAI image provider for the model-to-quality mapping.

This provider uses the OpenAI SDK environment, so OPENAI_API_KEY and OPENAI_BASE_URL must be visible to the running Hermes process, not only to the shell where you edited the config. model.base_url configures Hermes’ text/model provider path; the OpenAI image provider still needs the SDK environment so image requests go through Codex Pooler’s /v1 surface instead of OpenAI directly.

Alternate openai-codex credential-pool path

Section titled “Alternate openai-codex credential-pool path”

Use the openai-api setup above. Hermes can also reach Codex Pooler through its openai-codex provider pointed at the Pooler’s /v1 surface, but treat this as an advanced option with limits:

  • Always set model.context_length from your Pool’s /v1/models result. The Codex picker expects a native catalog and cannot read that endpoint’s OpenAI-compatible format; do not use it to select your Pool’s models or context limits.
  • With Hermes 0.21.5, the unconfigured context probe, /model picker and openai-codex image plugin can send the Pool API key to chatgpt.com. Keep the context limit explicit, avoid that picker and use image_gen.provider: openai. If the key was sent outside your Pooler instance, rotate it.
  • Web search switches to the provider’s hosted web_search tool.
  • Hermes never sends service_tier on this provider, and the named-provider option above does not apply to it.
  • Image generation and speech-to-text still need the OPENAI_* variables pointed at the Pooler’s /v1 surface.
  • Hermes sends a session_id header here. An ingress proxy that drops headers containing underscores removes it.

This provider supplies a session header that Codex Pooler’s websocket bridge can use when owner forwarding is enabled. A body prompt_cache_key alone does not enable the bridge. Keep the base URL on /v1: the recommended setup already supports prompt-cache locality, so switching providers or routes is unnecessary for that purpose.

Keep the endpoint value in the environment:

.env
HERMES_CODEX_BASE_URL=https://codex-pooler.example.com/v1

Set the provider in config.yaml. context_length is required here:

config.yaml
model:
default: gpt-6-sol
provider: openai-codex
base_url: https://codex-pooler.example.com/v1
context_length: 828400
supports_vision: true
agent:
image_input_mode: native
api_max_retries: 2
auto_recovery_cycles: 1
compression:
threshold: 0.95
auxiliary:
compression:
timeout: 900

Hermes treats openai-codex as an OAuth provider by default, so add a Pool API key credential ahead of any device-code credential, and keep the credential’s base_url on /v1. This example shows only placeholders. Don’t paste a real key into public docs or shared files.

auth.json
{
"active_provider": "openai-codex",
"credential_pool": {
"openai-codex": [
{
"label": "codex-pooler",
"auth_type": "api_key",
"priority": -10,
"source": "manual",
"access_token": "<pool-api-key>",
"base_url": "https://codex-pooler.example.com/v1"
}
]
}
}

Add a separate operator token to .env:

.env
CODEX_POOLER_MCP_KEY=<operator-mcp-token>

Add the following top-level block to config.yaml with either provider setup:

config.yaml
# Optional operator-only MCP metadata add-on. Omit for model/runtime use.
mcp_servers:
codex_pooler:
url: https://codex-pooler.example.com/mcp
headers:
Authorization: "Bearer ${CODEX_POOLER_MCP_KEY}"
enabled: true
timeout: 120
connect_timeout: 15

For local MCP setup, use http://localhost:4000/mcp.

Remote HTTP MCP servers require Hermes’ mcp extra. If hermes mcp test codex_pooler reports mcp.client.streamable_http is not available, install MCP support into the Hermes environment, following the Hermes MCP Integration docs, and rerun the test.

If text requests work but image generation fails with invalid_api_key, check the environment of the long-running Hermes process or gateway service first. It may not have loaded OPENAI_BASE_URL, so the OpenAI SDK image client may be using OpenAI’s default endpoint instead of Codex Pooler.

Hermes model requests use Codex Pooler’s narrow OpenAI-compatible /v1 support for selected SDK routes. Codex Pooler doesn’t provide full OpenAI API parity.

GET /v1/responses is narrow Responses websocket compatibility, not /v1/realtime support. /v1/realtime and OpenAI Realtime SDK websocket or session routes are unsupported.

The operator MCP endpoint is rooted at /mcp. It uses an operator-owned MCP token, not a Pool API key.