Configuration
Codex Pooler keeps boot-time release settings in environment variables and day-to-day operator settings in the database-backed admin UI.
Local Defaults
Section titled “Local Defaults”For Docker Compose, scripts/self-host/generate-env.sh writes a local .env file with generated secrets and local defaults. Keep that file private. Don’t reuse generated values across public installs.
The local web URL is http://localhost:4000.
The generator writes CODEX_POOLER_IMAGE_TAG=latest unless you set a tag first.
For self-hosted installs, set it to a tagged stable release from
GitHub Releases before
running the generator. The latest tag follows verified stable releases and does not move backward when an older version is published again. A version tag keeps the selected installation reproducible:
export CODEX_POOLER_IMAGE_TAG=<release-tag>scripts/self-host/generate-env.shIf port 4000 is already in use, set CODEX_POOLER_HTTP_PORT before
generation or edit the generated .env before starting the stack. Use the
Compose URL above even if the Phoenix startup banner prints an endpoint URL such
as https://localhost; the port mapping is the local URL to open.
Release Environment Variables
Section titled “Release Environment Variables”Set release environment variables for values the app needs before it can read database settings.
Set http_proxy and https_proxy to http:// proxy URLs to route outbound HTTP,
HTTPS, and WebSocket traffic by target scheme. Credentials embedded in a proxy
URL are sent as HTTP Basic proxy authentication. Use the comma-separated
no_proxy list for exact hosts, leading-dot or wildcard domain suffixes, IP
addresses, CIDR ranges, optional ports, or *. Lowercase variables take
precedence over HTTP_PROXY, HTTPS_PROXY, and NO_PROXY; an explicitly empty
lowercase value disables its uppercase fallback.
Proxy and no_proxy selection is evaluated again when an allowed Req caller follows a redirect, so the redirected destination cannot inherit the first URL’s proxy route.
| Variable | Purpose |
|---|---|
CODEX_POOLER_IMAGE | Release image name used by the Compose flow |
CODEX_POOLER_IMAGE_TAG | Release image tag used by the Compose flow |
CODEX_POOLER_HTTP_PORT | Local host port, default 4000 |
http_proxy / HTTP_PROXY | Optional HTTP proxy for outbound HTTP and WS destinations; lowercase takes precedence |
https_proxy / HTTPS_PROXY | Optional HTTP CONNECT proxy for outbound HTTPS and WSS destinations; lowercase takes precedence |
no_proxy / NO_PROXY | Optional comma-separated destination bypass list; lowercase takes precedence |
DATABASE_URL | Postgres connection URL |
SECRET_KEY_BASE | Phoenix signing and encryption secret |
PHX_HOST | Public host used by the HTTP endpoint |
PORT | HTTP port inside the release container |
PHX_SERVER | Set to true so the release serves HTTP |
OBAN_MODE | Release role, such as web, worker, scheduler, or all |
OBAN_JOBS_QUEUE_LIMIT | Job queue limit for the release role |
DNS_CLUSTER_QUERY | DNS clustering query when clustering is enabled |
CODEX_POOLER_TOTP_ENCRYPTION_KEY | TOTP encryption root key |
CODEX_POOLER_TOTP_KEY_VERSION | TOTP key version |
CODEX_POOLER_UPSTREAM_SECRET_KEY | Upstream secret encryption root key |
CODEX_POOLER_UPSTREAM_SECRET_KEY_VERSION | Upstream secret key version |
The upstream secret key must decode to 32 raw bytes. The TOTP and upstream key versions let operators rotate encrypted secret roots deliberately.
Every app, worker, and scheduler process of one installation must run with the same SECRET_KEY_BASE. Besides signing browser sessions, it keys the duplicate-turn and reconnect-replay protection of native Codex websocket turns. Changing it signs out every admin and API Key Observatory browser session, and for turns that are in flight while the change rolls out, a client resend of a turn started under the old value is not recognized as a duplicate and can be served again, and a reconnect replay of such a turn is refused. Change it during low traffic.
Official release images include the OS IANA timezone database used for operator timezone display. Custom runtime images or hosts must provide zoneinfo files at /usr/share/zoneinfo or set TZDIR.
Admin-Managed Settings
Section titled “Admin-Managed Settings”After the app boots, manage operational settings from the admin UI instead of editing code. Instance settings cover areas such as file limits, ingress trust, gateway diagnostics, route-class admission, circuit thresholds, metrics auth, operator email, model metadata, upstream timeouts and connection reuse, pricing catalog settings, and SMTP delivery.
OpenAI incident polling makes an outbound HTTPS request to status.openai.com every five minutes by default. On the System page’s Gateway tab, owners can turn off OpenAI status polling (operator.openai_status_polling_enabled, default true). The database-backed setting applies after the settings cache reloads, without a restart. Disabled scheduled checks make no outbound request and retain the last known incident history. Expected network or provider failures remain visible as feed status and cancel that check until the next scheduled poll; internal sync failures remain job failures.
Secret settings stay write-only in the UI. For example, metrics bearer material is stored as a keyed digest and SMTP credentials are stored encrypted for mail delivery and credential checks.
Forwarded client IP policy
Section titled “Forwarded client IP policy”The runtime firewall applies only to runtime compatibility routes and /mcp,
not /metrics. An empty client allowlist disables it. When the allowlist is
set, use exact client IPs or CIDRs. trusted_proxies is a prerequisite for
forwarding: list only the immediate proxy IPs or CIDRs allowed to supply a
client address. If the directly connected peer is not trusted, every forwarded
field is ignored and the peer address is evaluated instead.
The policy has three sources and one default:
forwarded_client_ip_source |
Allowed forwarded_proxy_depth |
Resolution |
|---|---|---|
peer |
0 |
Ignore forwarding headers and use the directly connected peer |
x_forwarded_for |
0..16 |
Use the selected XFF rule below, but only after the trusted-peer check |
x_real_ip |
0 |
Ignore XFF and require exactly one X-Real-IP field after the trusted-peer check |
The default is x_forwarded_for with depth 0 (default x_forwarded_for with
depth 0). At depth 0, the resolver walks XFF from right to left and selects
the first nontrusted address. At depth 1..16, it selects the positional entry
instead: Depth 2 selects the second XFF entry from the right. The directly
connected peer counts toward the configured proxy depth, but is not an XFF entry.
Duplicate XFF field occurrences are combined in wire order before this selection.
There is no mixed-header fallback: the selected source is authoritative. The
result is one bounded client identity for that request.
Forwarded addresses accept IPv4, IPv6, bracketed IPv6, and optional decimal ports
from 1 through 65535. Values are bounded to 64 printable ASCII bytes after
ASCII space or tab trimming, and XFF traversal is bounded to 32 entries. IP and
CIDR rules use one strict grammar: only outer ASCII space or tab is accepted,
prefixes are canonical unsigned decimal values, and IPv4-mapped IPv6 values
normalize to IPv4. A stored allowlist or trusted-proxy list with one invalid rule
fails closed rather than partially applying the rest.
If settings are unavailable on a cold start, runtime and MCP requests return
503; a warm node continues enforcing its last known good snapshot. Existing
websocket clients are rechecked after a locally applied firewall update. New work
is refused, already admitted work can finish, and the socket then closes with
code 1008.
Firewall denials emit codex_pooler_ingress_firewall_denied_count with only
scope and reason labels. They occur before authenticated request accounting.
Metrics uses a separate bearer boundary: it is open with no configured
metrics bearer, bearer-protected when configured, and unavailable
fail-closed when its settings cannot be read. The runtime firewall does not
protect /metrics.
The System page shows only the signed-in operator’s current active, unexpired browser session IP. It is not an inventory of other sessions; use Settings to review those sessions.
Paths are decoded once for classification. Unsafe decoded candidates on runtime and MCP routes receive fixed, sanitized invalid-path responses before authentication, parsing, or dispatch. Non-runtime routes retain their ordinary router behavior.
Malformed, overlong, unresolved, or over-bound forwarding input fails closed on runtime and MCP routes. Other routes retain the immediate peer and continue with their normal behavior. Each request uses one settings snapshot and one resolved client address, so a settings update cannot change the decision halfway through that request.
Compressed JSON ingress is enabled for gzip, deflate, and zstd when runtime
support is available. The defaults are 32 MiB compressed, 64 MiB decompressed, a
200:1 expansion ratio, and a 10-second decode timeout. Tune these four values and
the algorithm list together. Multipart uploads are not decompressed here; file,
audio, and image routes keep their own route-specific limits.
Pool image permission is checked after Pool API-key authentication and before
body parsing for exactly these four actions: POST /backend-api/codex/images/generations,
POST /backend-api/codex/images/edits, POST /v1/images/generations, and
POST /v1/images/edits. The gateway remains the final authority when dispatching
the request.
Gateway websocket upgrades use two separate admin-managed idle settings:
websocket_idle_timeout_mscontrols downstream idle close behavior for new backend Codex websocket upgrades and the narrow public/v1/responseswebsocket route. It also bounds owner-forwarded full-turn submit waits for new websocket requests. It is separate from upstream receive timeouts.websocket_owner_idle_timeout_mscontrols how long an owner and its upstream websocket may remain available after the final downstream detaches and no turn is active. Reattaching before expiry cancels that owner’s retention timer.
Both settings default to 1_800_000 ms and accept 60_000..3_600_000 ms. Each new or recovered owner captures websocket_owner_idle_timeout_ms from the node where it starts. An existing owner retains its captured value when the setting changes; later owners capture the new value.
A longer owner-retention window can improve the chance of reusing an upstream connection, but keeps owner processes and connections allocated longer after clients detach. A shorter window releases those resources sooner, with less opportunity for reuse. To roll back a change, restore the previous value in the admin UI. New and recovered owners use the restored value; replace existing owners as part of the rollback only when immediate convergence is required.
To roll back an ingress policy change, restore the previous allowlist and trusted-proxy settings in the admin UI. New requests use the restored snapshot; a websocket already revoked by the prior policy must reconnect.
Owner lease TTL
Section titled “Owner lease TTL”HTTP requests and websocket connections that continue a Codex session hold a short owner lease on that session. The lease is renewed while the owner is alive. bridge_owner_lease_ttl_seconds is how long a lease stays valid without renewal, and bridge_owner_lease_renewal_seconds is how often an active owner renews it. Change both on the System page, Gateway tab, in the Continuity group. The TTL defaults to 45 s and accepts 24 s or more. The renewal interval defaults to 15 s and accepts at most a third of the TTL, so a live owner gets at least two renewal attempts before its lease would expire; with the minimum TTL of 24 s that is 8 s. A save that sets either value checks the pair.
The minimum follows from the gateway’s own budgets, not from any installation’s latency. An HTTP request renews its lease once more just before it reserves capacity and dispatches upstream. If the lease expired during the work before that point, and it is still the one the request was given (no other replica renewed or replaced it), the request takes it over under a new lease, and the old lease stops being valid everywhere; this is also what happens when the previous turn of the session ran on another replica and its lease ran out a moment after the new request picked it up. If another replica holds the session by then, the request is refused with 409 stale_owner, and a resend attaches to the current owner. Every such refusal appears in the request log as a rejected request with the refusal code and no usage. The TTL must still outlast one database statement at its full 15-second budget plus that final renewal.
A value saved before this minimum existed is not rewritten. The gateway uses the minimum in its place and logs instance setting clamped at read setting=bridge_owner_lease_ttl_seconds once per replica. Save a value of 24 or more to clear it. A saved renewal interval above a third of the TTL in effect is handled the same way: the gateway renews every third of the TTL instead and logs instance setting clamped at read setting=bridge_owner_lease_renewal_seconds once per replica. Work between that final renewal and dispatch that runs longer than the TTL still ends in 503 owner_unavailable. Clients that retry server errors recover: the retry starts a new session under the same session key.
Upstream connection idle bound
Section titled “Upstream connection idle bound”Upstream HTTP requests reuse pooled keep-alive connections to each provider origin. upstream_conn_max_idle_time_ms is the longest time a pooled connection may sit idle before it is replaced instead of reused. It applies to every outbound HTTP request Codex Pooler sends: gateway traffic, usage and quota probes, token refresh and device sign-in, saved-reset redemption, model catalog sync, the pricing catalog import, the OpenAI status feed, alert webhook deliveries, and file uploads to the provider’s storage URL. Background requests such as the hourly pricing import run far less often than gateway traffic, so their connections are the most likely to sit idle past an egress timeout; cloud NAT gateways commonly drop idle connections after a few minutes. Change it on the System page, Gateway tab, in the Upstream timing group, next to the connect, pool, and receive timeouts. The default is 45000 ms and the accepted range is 1000..3600000 ms. A saved value applies immediately to new upstream requests without a restart; the first requests after a change open new connections. Treat this setting and the upstream connect timeout as low-churn: each distinct combination of their saved values keeps its own connection pools until the next restart, so change them deliberately and do not drive them from automation. Changing the pool or receive timeout does not create new pools.
The bound is checked only when a connection is taken for a new request, so it never interrupts an in-flight or streaming request, however long that request runs.
A NAT gateway, load balancer, or proxy on the egress path can silently drop a connection that has been idle longer than its own timeout. Reusing such a connection fails the request within milliseconds as an upstream network error with a closed transport failure in the request phase, and the gateway does not retry it on the same account, because the request may already have reached the provider. Alert webhook deliveries and file uploads are not retried in place either; scheduled jobs and webhook deliveries try again on their normal retry schedule. If those failures appear, lower the bound below the shortest idle timeout on your egress path. A longer bound means fewer reconnects but more exposure to dropped connections; a shorter bound means more connection setup but fewer stale reuses.
Runtime URLs
Section titled “Runtime URLs”Use only the surface that matches your client.
| Client type | Local URL | Deployed example |
|---|---|---|
| Codex backend client | http://localhost:4000/backend-api/codex | https://codex-pooler.example.com/backend-api/codex |
| Narrow OpenAI-compatible SDK client | http://localhost:4000/v1 | https://codex-pooler.example.com/v1 |
| Operator MCP host | http://localhost:4000/mcp | https://codex-pooler.example.com/mcp |
Use Authorization: Bearer <pool-api-key> for /backend-api and /v1. Use Authorization: Bearer <operator-mcp-token> for /mcp.
The operator MCP endpoint is metadata-only and read-only. It doesn’t accept Pool API keys, browser sessions, cookies, invite tokens, upstream tokens, or query tokens.
Health And Readiness
Section titled “Health And Readiness”Use these checks for a local instance:
curl -fsS http://localhost:4000/healthzcurl -fsS http://localhost:4000/readyzFor a deployed instance, use the same paths on https://codex-pooler.example.com.
Support And Responsible Operation
Section titled “Support And Responsible Operation”The support boundary is intentionally narrow. Public docs can help you start the service, configure supported route surfaces, and understand metadata-only behavior. They can’t decide whether your organization may use a provider account, share a Pool, or connect a given MCP host.
For responsible operation:
- Operate only accounts you are allowed to use
- Keep Pool API keys separate from operator MCP tokens
- Rotate exposed credentials immediately
- Connect MCP only to hosts you trust with operator metadata
- Don’t send raw prompts, request bodies, response bodies, file bytes, audio, images, cookies, bearer tokens, Pool API keys, MCP tokens, upstream secrets, or
auth.jsonmaterial in support messages
Compatibility Notes
Section titled “Compatibility Notes”The Codex backend compatibility route is rooted at /backend-api/codex/*. The /v1 surface provides narrow OpenAI-compatible support for selected SDK routes, then routes supported requests through the same Pool policy and accounting path.
It does not provide full OpenAI API parity. OpenAI Realtime SDK websocket and session routes are not supported. GET /v1/responses is narrow Responses websocket compatibility, not /v1/realtime support.