# Configuration

Codex Pooler keeps boot-time release settings in environment variables and day-to-day operator settings in the database-backed admin UI.

## Local Defaults

For Docker Compose, `scripts/self-host/generate-env.sh` writes a local `.env` file with generated secrets and local defaults. Keep that file private. Don't reuse generated values across public installs.

The local web URL is `http://localhost:4000`.

The generator writes `CODEX_POOLER_IMAGE_TAG=latest` unless you set a tag first.
For self-hosted installs, set it to a tagged stable release from
[GitHub Releases](https://github.com/icoretech/codex-pooler/releases) before
running the generator. The `latest` tag follows verified stable releases and does not move backward when an older version is published again. A version tag keeps the selected installation reproducible:

```bash
export CODEX_POOLER_IMAGE_TAG=<release-tag>
scripts/self-host/generate-env.sh
```

If port `4000` is already in use, set `CODEX_POOLER_HTTP_PORT` before
generation or edit the generated `.env` before starting the stack. Use the
Compose URL above even if the Phoenix startup banner prints an endpoint URL such
as `https://localhost`; the port mapping is the local URL to open.

## Release Environment Variables

Set release environment variables for values the app needs before it can read database settings.

Set `http_proxy` and `https_proxy` to `http://` proxy URLs to route outbound HTTP,
HTTPS, and WebSocket traffic by target scheme. Credentials embedded in a proxy
URL are sent as HTTP Basic proxy authentication. Use the comma-separated
`no_proxy` list for exact hosts, leading-dot or wildcard domain suffixes, IP
addresses, CIDR ranges, optional ports, or `*`. Lowercase variables take
precedence over `HTTP_PROXY`, `HTTPS_PROXY`, and `NO_PROXY`; an explicitly empty
lowercase value disables its uppercase fallback.
Proxy and `no_proxy` selection is evaluated again when an allowed Req caller follows a redirect, so the redirected destination cannot inherit the first URL's proxy route.

<table>
  <thead>
    <tr>
      <th>Variable</th>
      <th>Purpose</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td><code>CODEX_POOLER_IMAGE</code></td>
      <td>Release image name used by the Compose flow</td>
    </tr>
    <tr>
      <td><code>CODEX_POOLER_IMAGE_TAG</code></td>
      <td>Release image tag used by the Compose flow</td>
    </tr>
    <tr>
      <td><code>CODEX_POOLER_HTTP_PORT</code></td>
      <td>Local host port, default <code>4000</code></td>
    </tr>
    <tr>
      <td><code>http_proxy</code> / <code>HTTP_PROXY</code></td>
      <td>Optional HTTP proxy for outbound HTTP and WS destinations; lowercase takes precedence</td>
    </tr>
    <tr>
      <td><code>https_proxy</code> / <code>HTTPS_PROXY</code></td>
      <td>Optional HTTP CONNECT proxy for outbound HTTPS and WSS destinations; lowercase takes precedence</td>
    </tr>
    <tr>
      <td><code>no_proxy</code> / <code>NO_PROXY</code></td>
      <td>Optional comma-separated destination bypass list; lowercase takes precedence</td>
    </tr>
    <tr>
      <td><code>DATABASE_URL</code></td>
      <td>Postgres connection URL</td>
    </tr>
    <tr>
      <td><code>SECRET_KEY_BASE</code></td>
      <td>Phoenix signing and encryption secret</td>
    </tr>
    <tr>
      <td><code>PHX_HOST</code></td>
      <td>Public host used by the HTTP endpoint</td>
    </tr>
    <tr>
      <td><code>PORT</code></td>
      <td>HTTP port inside the release container</td>
    </tr>
    <tr>
      <td><code>PHX_SERVER</code></td>
      <td>Set to <code>true</code> so the release serves HTTP</td>
    </tr>
    <tr>
      <td><code>OBAN_MODE</code></td>
      <td>Release role, such as <code>web</code>, <code>worker</code>, <code>scheduler</code>, or <code>all</code></td>
    </tr>
    <tr>
      <td><code>OBAN_JOBS_QUEUE_LIMIT</code></td>
      <td>Job queue limit for the release role</td>
    </tr>
    <tr>
      <td><code>DNS_CLUSTER_QUERY</code></td>
      <td>DNS clustering query when clustering is enabled</td>
    </tr>
    <tr>
      <td><code>CODEX_POOLER_TOTP_ENCRYPTION_KEY</code></td>
      <td>TOTP encryption root key</td>
    </tr>
    <tr>
      <td><code>CODEX_POOLER_TOTP_KEY_VERSION</code></td>
      <td>TOTP key version</td>
    </tr>
    <tr>
      <td><code>CODEX_POOLER_UPSTREAM_SECRET_KEY</code></td>
      <td>Upstream secret encryption root key</td>
    </tr>
    <tr>
      <td><code>CODEX_POOLER_UPSTREAM_SECRET_KEY_VERSION</code></td>
      <td>Upstream secret key version</td>
    </tr>
  </tbody>
</table>

The upstream secret key must decode to 32 raw bytes. The TOTP and upstream key versions let operators rotate encrypted secret roots deliberately.

Every app, worker, and scheduler process of one installation must run with the same `SECRET_KEY_BASE`. Besides signing browser sessions, it keys the duplicate-turn and reconnect-replay protection of native Codex websocket turns. Changing it signs out every admin and API Key Observatory browser session, and for turns that are in flight while the change rolls out, a client resend of a turn started under the old value is not recognized as a duplicate and can be served again, and a reconnect replay of such a turn is refused. Change it during low traffic.

Official release images include the OS IANA timezone database used for operator timezone display. Custom runtime images or hosts must provide zoneinfo files at `/usr/share/zoneinfo` or set `TZDIR`.

## Admin-Managed Settings

After the app boots, manage operational settings from the admin UI instead of editing code. Instance settings cover areas such as file limits, ingress trust, gateway diagnostics, route-class admission, circuit thresholds, metrics auth, operator email, model metadata, upstream timeouts and connection reuse, pricing catalog settings, and SMTP delivery.

OpenAI incident polling makes an outbound HTTPS request to `status.openai.com` every five minutes by default. On the System page's Gateway tab, owners can turn off **OpenAI status polling** (`operator.openai_status_polling_enabled`, default `true`). The database-backed setting applies after the settings cache reloads, without a restart. Disabled scheduled checks make no outbound request and retain the last known incident history. Expected network or provider failures remain visible as feed status and cancel that check until the next scheduled poll; internal sync failures remain job failures.

Secret settings stay write-only in the UI. For example, metrics bearer material is stored as a keyed digest and SMTP credentials are stored encrypted for mail delivery and credential checks.

### Forwarded client IP policy

The runtime firewall applies only to runtime compatibility routes and `/mcp`,
not `/metrics`. An empty client allowlist disables it. When the allowlist is
set, use exact client IPs or CIDRs. `trusted_proxies` is a prerequisite for
forwarding: list only the immediate proxy IPs or CIDRs allowed to supply a
client address. If the directly connected peer is not trusted, every forwarded
field is ignored and the peer address is evaluated instead.

The policy has three sources and one default:

| `forwarded_client_ip_source` | Allowed `forwarded_proxy_depth` | Resolution |
| --- | --- | --- |
| `peer` | `0` | Ignore forwarding headers and use the directly connected peer |
| `x_forwarded_for` | `0..16` | Use the selected XFF rule below, but only after the trusted-peer check |
| `x_real_ip` | `0` | Ignore XFF and require exactly one X-Real-IP field after the trusted-peer check |

The default is `x_forwarded_for` with depth `0` (default `x_forwarded_for` with
depth `0`). At depth `0`, the resolver walks XFF from right to left and selects
the first nontrusted address. At depth `1..16`, it selects the positional entry
instead: Depth `2` selects the second XFF entry from the right. The directly
connected peer counts toward the configured proxy depth, but is not an XFF entry.
Duplicate XFF field occurrences are combined in wire order before this selection.
There is no mixed-header fallback: the selected source is authoritative. The
result is one bounded client identity for that request.

Forwarded addresses accept IPv4, IPv6, bracketed IPv6, and optional decimal ports
from `1` through `65535`. Values are bounded to 64 printable ASCII bytes after
ASCII space or tab trimming, and XFF traversal is bounded to 32 entries. IP and
CIDR rules use one strict grammar: only outer ASCII space or tab is accepted,
prefixes are canonical unsigned decimal values, and IPv4-mapped IPv6 values
normalize to IPv4. A stored allowlist or trusted-proxy list with one invalid rule
fails closed rather than partially applying the rest.

If settings are unavailable on a cold start, runtime and MCP requests return
`503`; a warm node continues enforcing its last known good snapshot. Existing
websocket clients are rechecked after a locally applied firewall update. New work
is refused, already admitted work can finish, and the socket then closes with
code `1008`.

Firewall denials emit `codex_pooler_ingress_firewall_denied_count` with only
`scope` and `reason` labels. They occur before authenticated request accounting.

Metrics uses a separate bearer boundary: it is **open** with no configured
metrics bearer, **bearer-protected** when configured, and **unavailable**
fail-closed when its settings cannot be read. The runtime firewall does not
protect `/metrics`.

The System page shows only the signed-in operator's current active, unexpired
browser session IP. It is not an inventory of other sessions; use Settings to
review those sessions.

Paths are decoded once for classification. Unsafe decoded candidates on runtime
and MCP routes receive fixed, sanitized invalid-path responses before
authentication, parsing, or dispatch. Non-runtime routes retain their ordinary
router behavior.

Malformed, overlong, unresolved, or over-bound forwarding input fails closed on
runtime and MCP routes. Other routes retain the immediate peer and continue with
their normal behavior. Each request uses one settings snapshot and one resolved
client address, so a settings update cannot change the decision halfway through
that request.

Compressed JSON ingress is enabled for `gzip`, `deflate`, and `zstd` when runtime
support is available. The defaults are 32 MiB compressed, 64 MiB decompressed, a
200:1 expansion ratio, and a 10-second decode timeout. Tune these four values and
the algorithm list together. Multipart uploads are not decompressed here; file,
audio, and image routes keep their own route-specific limits.

Pool image permission is checked after Pool API-key authentication and before
body parsing for exactly these four actions: `POST /backend-api/codex/images/generations`,
`POST /backend-api/codex/images/edits`, `POST /v1/images/generations`, and
`POST /v1/images/edits`. The gateway remains the final authority when dispatching
the request.

Gateway websocket upgrades use two separate admin-managed idle settings:

- `websocket_idle_timeout_ms` controls downstream idle close behavior for new backend Codex websocket upgrades and the narrow public `/v1/responses` websocket route. It also bounds owner-forwarded full-turn submit waits for new websocket requests. It is separate from upstream receive timeouts.
- `websocket_owner_idle_timeout_ms` controls how long an owner and its upstream websocket may remain available after the final downstream detaches and no turn is active. Reattaching before expiry cancels that owner's retention timer.

Both settings default to `1_800_000` ms and accept `60_000..3_600_000` ms. Each new or recovered owner captures `websocket_owner_idle_timeout_ms` from the node where it starts. An existing owner retains its captured value when the setting changes; later owners capture the new value.

A longer owner-retention window can improve the chance of reusing an upstream connection, but keeps owner processes and connections allocated longer after clients detach. A shorter window releases those resources sooner, with less opportunity for reuse. To roll back a change, restore the previous value in the admin UI. New and recovered owners use the restored value; replace existing owners as part of the rollback only when immediate convergence is required.

To roll back an ingress policy change, restore the previous allowlist and
trusted-proxy settings in the admin UI. New requests use the restored snapshot;
a websocket already revoked by the prior policy must reconnect.

### Owner lease TTL

HTTP requests and websocket connections that continue a Codex session hold a short owner lease on that session. The lease is renewed while the owner is alive. `bridge_owner_lease_ttl_seconds` is how long a lease stays valid without renewal, and `bridge_owner_lease_renewal_seconds` is how often an active owner renews it. Change both on the System page, Gateway tab, in the Continuity group. The TTL defaults to `45` s and accepts `24` s or more. The renewal interval defaults to `15` s and accepts at most a third of the TTL, so a live owner gets at least two renewal attempts before its lease would expire; with the minimum TTL of `24` s that is `8` s. A save that sets either value checks the pair.

The minimum follows from the gateway's own budgets, not from any installation's latency. An HTTP request renews its lease once more just before it reserves capacity and dispatches upstream. If the lease expired during the work before that point, and it is still the one the request was given (no other replica renewed or replaced it), the request takes it over under a new lease, and the old lease stops being valid everywhere; this is also what happens when the previous turn of the session ran on another replica and its lease ran out a moment after the new request picked it up. If another replica holds the session by then, the request is refused with `409` `stale_owner`, and a resend attaches to the current owner. Every such refusal appears in the request log as a rejected request with the refusal code and no usage. The TTL must still outlast one database statement at its full 15-second budget plus that final renewal.

A value saved before this minimum existed is not rewritten. The gateway uses the minimum in its place and logs `instance setting clamped at read setting=bridge_owner_lease_ttl_seconds` once per replica. Save a value of `24` or more to clear it. A saved renewal interval above a third of the TTL in effect is handled the same way: the gateway renews every third of the TTL instead and logs `instance setting clamped at read setting=bridge_owner_lease_renewal_seconds` once per replica. Work between that final renewal and dispatch that runs longer than the TTL still ends in `503` `owner_unavailable`. Clients that retry server errors recover: the retry starts a new session under the same session key.

### Upstream connection idle bound

Upstream HTTP requests reuse pooled keep-alive connections to each provider origin. `upstream_conn_max_idle_time_ms` is the longest time a pooled connection may sit idle before it is replaced instead of reused. It applies to every outbound HTTP request Codex Pooler sends: gateway traffic, usage and quota probes, token refresh and device sign-in, saved-reset redemption, model catalog sync, the pricing catalog import, the OpenAI status feed, alert webhook deliveries, and file uploads to the provider's storage URL. Background requests such as the hourly pricing import run far less often than gateway traffic, so their connections are the most likely to sit idle past an egress timeout; cloud NAT gateways commonly drop idle connections after a few minutes. Change it on the System page, Gateway tab, in the Upstream timing group, next to the connect, pool, and receive timeouts. The default is `45000` ms and the accepted range is `1000..3600000` ms. A saved value applies immediately to new upstream requests without a restart; the first requests after a change open new connections. Treat this setting and the upstream connect timeout as low-churn: each distinct combination of their saved values keeps its own connection pools until the next restart, so change them deliberately and do not drive them from automation. Changing the pool or receive timeout does not create new pools.

The bound is checked only when a connection is taken for a new request, so it never interrupts an in-flight or streaming request, however long that request runs.

A NAT gateway, load balancer, or proxy on the egress path can silently drop a connection that has been idle longer than its own timeout. Reusing such a connection fails the request within milliseconds as an upstream network error with a `closed` transport failure in the request phase, and the gateway does not retry it on the same account, because the request may already have reached the provider. Alert webhook deliveries and file uploads are not retried in place either; scheduled jobs and webhook deliveries try again on their normal retry schedule. If those failures appear, lower the bound below the shortest idle timeout on your egress path. A longer bound means fewer reconnects but more exposure to dropped connections; a shorter bound means more connection setup but fewer stale reuses.

## Runtime URLs

Use only the surface that matches your client.

<table>
  <thead>
    <tr>
      <th>Client type</th>
      <th>Local URL</th>
      <th>Deployed example</th>
    </tr>
  </thead>
  <tbody>
    <tr>
      <td>Codex backend client</td>
      <td><code>http://localhost:4000/backend-api/codex</code></td>
      <td><code>https://codex-pooler.example.com/backend-api/codex</code></td>
    </tr>
    <tr>
      <td>Narrow OpenAI-compatible SDK client</td>
      <td><code>http://localhost:4000/v1</code></td>
      <td><code>https://codex-pooler.example.com/v1</code></td>
    </tr>
    <tr>
      <td>Operator MCP host</td>
      <td><code>http://localhost:4000/mcp</code></td>
      <td><code>https://codex-pooler.example.com/mcp</code></td>
    </tr>
  </tbody>
</table>

Use `Authorization: Bearer <pool-api-key>` for `/backend-api` and `/v1`. Use `Authorization: Bearer <operator-mcp-token>` for `/mcp`.

The operator MCP endpoint is metadata-only and read-only. It doesn't accept Pool API keys, browser sessions, cookies, invite tokens, upstream tokens, or query tokens.

## Health And Readiness

Use these checks for a local instance:

```bash
curl -fsS http://localhost:4000/healthz
curl -fsS http://localhost:4000/readyz
```

For a deployed instance, use the same paths on `https://codex-pooler.example.com`.

## Support And Responsible Operation

The support boundary is intentionally narrow. Public docs can help you start the service, configure supported route surfaces, and understand metadata-only behavior. They can't decide whether your organization may use a provider account, share a Pool, or connect a given MCP host.

For responsible operation:

- Operate only accounts you are allowed to use
- Keep Pool API keys separate from operator MCP tokens
- Rotate exposed credentials immediately
- Connect MCP only to hosts you trust with operator metadata
- Don't send raw prompts, request bodies, response bodies, file bytes, audio, images, cookies, bearer tokens, Pool API keys, MCP tokens, upstream secrets, or `auth.json` material in support messages

## Compatibility Notes

The Codex backend compatibility route is rooted at `/backend-api/codex/*`. The `/v1` surface provides narrow OpenAI-compatible support for selected SDK routes, then routes supported requests through the same Pool policy and accounting path.

It does not provide full OpenAI API parity. OpenAI Realtime SDK websocket and session routes are not supported. `GET /v1/responses` is narrow Responses websocket compatibility, not `/v1/realtime` support.