Hermes Agent on Codex Pooler
Hermes Agent is an AI assistant from Nous Research that can remember past conversations and build reusable skills. Use it from the terminal or connected messaging channels, with Codex Pooler supplying the models for its conversations. This guide also covers image generation and voice-message transcription through the same Pooler instance.

Before you start
Section titled “Before you start”- Install Hermes Agent using the official instructions for your operating system.
- Have a Codex Pooler URL reachable from the client.
- Create a Pool API key and choose a model available to that Pool.
Configure the connection
Section titled “Configure the connection”Use Hermes’ openai-api provider with the /v1 base URL for this setup. Model requests and optional image generation use the Pool API key; operator MCP uses a separate token.
Config file paths
Section titled “Config file paths”API keys and addresses
| OS | Default config file |
|---|---|
| macOS | ~/.hermes/.env |
| Linux | ~/.hermes/.env |
| Windows | %LOCALAPPDATA%\hermes\.env |
Model and feature settings
| OS | Default config file |
|---|---|
| macOS | ~/.hermes/config.yaml |
| Linux | ~/.hermes/config.yaml |
| Windows | %LOCALAPPDATA%\hermes\config.yaml |
Credentials for the alternate provider setup
| OS | Default config file |
|---|---|
| macOS | ~/.hermes/auth.json |
| Linux | ~/.hermes/auth.json |
| Windows | %LOCALAPPDATA%\hermes\auth.json |
The native Windows installer uses %LOCALAPPDATA%\hermes and sets HERMES_HOME to that folder. If HERMES_HOME points elsewhere, use that folder for all three files. A WSL installation has its own ~/.hermes folder.
These are the default locations. On Windows, paste the %USERPROFILE%, %APPDATA% or %LOCALAPPDATA% path into File Explorer’s address bar. For a client installed inside WSL, use the Linux paths and commands inside WSL. Keep any custom configuration folder or profile you already use.
Store the Pool API key and base URL in the .env file for your installation:
OPENAI_API_KEY=<pool-api-key>OPENAI_BASE_URL=https://codex-pooler.example.com/v1Then point config.yaml at the same /v1 base URL and use the Responses transport mode:
model: default: gpt-6-sol provider: openai-api base_url: https://codex-pooler.example.com/v1 api_mode: codex_responses context_length: 828400 supports_vision: true
agent: image_input_mode: native api_max_retries: 2 auto_recovery_cycles: 1
compression: threshold: 0.95
auxiliary: compression: timeout: 900For local setup, use http://localhost:4000/v1 in both the environment and model configuration.
Choose a model
Section titled “Choose a model”Set model.default to a model available to your Pool.
Current Codex Pooler releases expose an SDK-readable context_length on /v1/models, flattened from the selected native raw context window and its effective percentage. Treat that per-model endpoint value as authoritative because provider accounts can temporarily report different catalog ceilings. The 828400 value in this example is a long-profile value for a Pool whose selected model source reports 872000 raw tokens; a short 272000-token profile reports 258400 instead. Match it to the Pool’s /v1/models result rather than configuring the raw ceiling. Set model.context_length explicitly rather than relying on detection: without it Hermes probes /v1/models itself, and that probe carries an Authorization header only when the provider’s API key resolves to a plain string, so a key supplied through a command or another indirect source probes anonymously, is refused, and is remembered as a failure for five minutes. Hermes then logs Could not detect context length … (probe-down) on every resolution and silently falls back to its own bundled catalog, which for these models is larger than the window your Pool actually advertises — so compression starts too late and a turn can overflow upstream. An explicit context_length removes the probe entirely.
Hermes compresses the conversation when it reaches the lower of two limits: compression.threshold times the context window, and compression.threshold_tokens, which defaults to 256000 tokens. With the long-profile example, 0.95 of the window is 786980 tokens, so the 256000-token cap applies. With a 258400-token short profile, 0.95 of the window is lower than the cap and applies instead. Keep the default cap: a smaller per-turn context uses less quota. Set compression.threshold_tokens: null only if you want Hermes to compress by the ratio alone. Hermes context compression uses its own auxiliary request timeout. Keep auxiliary.compression.timeout: 900 so large retained contexts can finish instead of cycling through the older 120-second compression budget. This is independent from the optional MCP server timeout and from an application output cap.
Hermes generates session titles with a separate auxiliary call. Current Hermes releases send that call with reasoning disabled, which overrides auxiliary.title_generation.reasoning_effort, so the key has no effect on titles and you can leave it out; the model’s default effort applies.
Hermes fast mode (agent.service_tier: fast, /fast) sends service_tier: priority only when the provider is openai on api.openai.com or openai-codex on chatgpt.com. Since the Hermes 0.21 releases, an openai-api provider pointed at Codex Pooler never sends it, even though /fast and its tip still report fast mode as on. Codex Pooler accepts service_tier: priority (and fast as an alias) whenever a client sends it, so this is a Hermes routing rule, not a Pooler limitation. In that setup the request log shows tier default with no requested tier, because nothing was requested. To get priority processing, use one of these:
-
In Hermes, declare a named provider under
providers:with the same/v1base URL,api_mode: codex_responses, your Pool API key andextra_body: {service_tier: priority}, then pointmodel.providerat it. The model’scontext_lengthand the other settings above still apply.config.yaml model:default: gpt-6-solprovider: codex-poolerbase_url: https://codex-pooler.example.com/v1api_mode: codex_responsescontext_length: 828400supports_vision: trueproviders:codex-pooler:base_url: https://codex-pooler.example.com/v1api_mode: codex_responseskey_env: OPENAI_API_KEYextra_body:service_tier: priorityHermes adds
extra_bodyto the main agent turns only. Its auxiliary calls, such as session titles and context compression, use their ownauxiliary.<task>.extra_bodyand request no tier. -
In Codex Pooler, set the API key’s enforced service tier to
priority. This needs no client change, but it only routes when the Pool’s model declares that tier.
Either way the request log then shows tier default next to priority requested, since the Codex backend reports default even for priority requests (see service tiers). Priority processing uses included quota faster.
Retries and fallback
Section titled “Retries and fallback”When every account in a Pool has used up its quota, Codex Pooler answers 429 with usage_limit_reached, the reset time in resets_at, and Retry-After. Hermes treats this as a rate limit: before each retry it waits for Retry-After, up to 600 seconds, and then ends the turn with the reset time. With the default of three attempts, a turn can wait twice while the Pool cannot recover. agent.api_max_retries: 2 keeps it to one wait.
A retryable 503, for example while an account’s circuit is open, also carries Retry-After. Once a turn’s own retries are spent, Hermes retries server errors through a recovery ladder of up to five cycles by default. agent.auto_recovery_cycles: 1 keeps that to one extra cycle.
If you have a second Pool API key, or another Codex Pooler instance, you can list it under fallback_providers. Hermes switches to it once the retries on the primary provider are spent. A fallback key in the same Pool does not help when that Pool is out of quota, because it uses the same accounts.
providers: second-pool: base_url: https://codex-pooler.example.com/v1 api_mode: codex_responses key_env: SECOND_POOL_API_KEY
fallback_providers: - provider: second-pool model: gpt-6-lunaVerify the connection
Section titled “Verify the connection”Test the model path with a one-shot prompt:
hermes -z 'Reply with exactly: hermes openai api ok' --ignore-rulesThat test should create normal /v1 traffic using the Pool API key. The MCP entry is separate and should authenticate only with an operator-owned MCP token.
In Codex Pooler’s request logs, match the request time, API key, model, and final status to your test. A reply alone does not confirm that the client used your Pooler instance.
Advanced configuration
Section titled “Advanced configuration”Image generation and speech-to-text
Section titled “Image generation and speech-to-text”Add these optional blocks to config.yaml:
image_gen: provider: openai model: gpt-image-2.5-flare-medium
stt: enabled: true provider: openai openai: model: gpt-4o-transcribeFor speech-to-text, also add this value to .env:
STT_OPENAI_BASE_URL=https://codex-pooler.example.com/v1For local setup, use http://localhost:4000/v1.
Speech-to-text uses POST /v1/audio/transcriptions. Hermes’ audio client needs
STT_OPENAI_BASE_URL; it does not inherit model.base_url or OPENAI_BASE_URL.
It uses OPENAI_API_KEY unless VOICE_TOOLS_OPENAI_KEY overrides it. Keep these
variables visible to the running Hermes process. Set stt.provider: openai
explicitly to select this route instead of automatic local or cloud provider
selection, and use gpt-4o-transcribe because Hermes’ default whisper-1 is
not supported by Codex Pooler. The gpt-transcribe alias is also accepted.
This enables transcription only; text-to-speech and Realtime audio are not supported.
Use a Hermes version whose OpenAI image provider lists GPT Image 2.5. Its
gpt-image-2.5-flare-medium preset sends gpt-image-2.5-flare to the API
with quality: medium; use gpt-image-2.5-sunburst-medium for Sunburst.
The -low, -high, -xhigh, and -max presets are also accepted and forwarded.
The Codex backend can return a different quality or size than requested; these
presets do not guarantee measured output quality. Older Hermes versions hardcode gpt-image-2 in this
provider; update Hermes before applying the new preset. See the
Hermes OpenAI image provider
for the model-to-quality mapping.
This provider uses the OpenAI SDK environment, so OPENAI_API_KEY and
OPENAI_BASE_URL must be visible to the running Hermes process, not only to
the shell where you edited the config. model.base_url configures Hermes’
text/model provider path; the OpenAI image provider still needs the SDK
environment so image requests go through Codex Pooler’s /v1 surface instead
of OpenAI directly.
Alternate openai-codex credential-pool path
Section titled “Alternate openai-codex credential-pool path”Use the openai-api setup above. Hermes can also reach Codex Pooler through its openai-codex provider pointed at the Pooler’s /v1 surface, but treat this as an advanced option with limits:
- Always set
model.context_lengthfrom your Pool’s/v1/modelsresult. The Codex picker expects a native catalog and cannot read that endpoint’s OpenAI-compatible format; do not use it to select your Pool’s models or context limits. - With Hermes 0.21.5, the unconfigured context probe,
/modelpicker andopenai-codeximage plugin can send the Pool API key tochatgpt.com. Keep the context limit explicit, avoid that picker and useimage_gen.provider: openai. If the key was sent outside your Pooler instance, rotate it. - Web search switches to the provider’s hosted
web_searchtool. - Hermes never sends
service_tieron this provider, and the named-provider option above does not apply to it. - Image generation and speech-to-text still need the
OPENAI_*variables pointed at the Pooler’s/v1surface. - Hermes sends a
session_idheader here. An ingress proxy that drops headers containing underscores removes it.
This provider supplies a session header that Codex Pooler’s websocket bridge can use when owner forwarding is enabled. A body prompt_cache_key alone does not enable the bridge. Keep the base URL on /v1: the recommended setup already supports prompt-cache locality, so switching providers or routes is unnecessary for that purpose.
Keep the endpoint value in the environment:
HERMES_CODEX_BASE_URL=https://codex-pooler.example.com/v1Set the provider in config.yaml. context_length is required here:
model: default: gpt-6-sol provider: openai-codex base_url: https://codex-pooler.example.com/v1 context_length: 828400 supports_vision: true
agent: image_input_mode: native api_max_retries: 2 auto_recovery_cycles: 1
compression: threshold: 0.95
auxiliary: compression: timeout: 900Hermes treats openai-codex as an OAuth provider by default, so add a Pool API key credential ahead of any device-code credential, and keep the credential’s base_url on /v1. This example shows only placeholders. Don’t paste a real key into public docs or shared files.
{ "active_provider": "openai-codex", "credential_pool": { "openai-codex": [ { "label": "codex-pooler", "auth_type": "api_key", "priority": -10, "source": "manual", "access_token": "<pool-api-key>", "base_url": "https://codex-pooler.example.com/v1" } ] }}Operator MCP (optional)
Section titled “Operator MCP (optional)”Add a separate operator token to .env:
CODEX_POOLER_MCP_KEY=<operator-mcp-token>Add the following top-level block to config.yaml with either provider setup:
# Optional operator-only MCP metadata add-on. Omit for model/runtime use.mcp_servers: codex_pooler: url: https://codex-pooler.example.com/mcp headers: Authorization: "Bearer ${CODEX_POOLER_MCP_KEY}" enabled: true timeout: 120 connect_timeout: 15For local MCP setup, use http://localhost:4000/mcp.
Remote HTTP MCP servers require Hermes’ mcp extra. If hermes mcp test codex_pooler reports mcp.client.streamable_http is not available, install MCP support into the Hermes environment, following the Hermes MCP Integration docs, and rerun the test.
Troubleshooting
Section titled “Troubleshooting”If text requests work but image generation fails with invalid_api_key, check
the environment of the long-running Hermes process or gateway service first.
It may not have loaded OPENAI_BASE_URL, so the OpenAI SDK image client may be
using OpenAI’s default endpoint instead of Codex Pooler.
Compatibility notes
Section titled “Compatibility notes”Hermes model requests use Codex Pooler’s narrow OpenAI-compatible /v1 support for selected SDK routes. Codex Pooler doesn’t provide full OpenAI API parity.
GET /v1/responses is narrow Responses websocket compatibility, not /v1/realtime support. /v1/realtime and OpenAI Realtime SDK websocket or session routes are unsupported.
The operator MCP endpoint is rooted at /mcp. It uses an operator-owned MCP token, not a Pool API key.