Runtime Routes
Codex Pooler exposes a deliberately narrow runtime surface. It is not a wildcard proxy and it does not claim full OpenAI API parity. Each route below is either a Codex backend compatibility route, a backend bridge route, a translated /v1 compatibility route, an operator metadata route, or an explicit unsupported boundary.
Use this page when you need to answer three questions before wiring a client:
- Which endpoint should the client call?
- Which upstream route does Codex Pooler call after admission and routing?
- Is the request passed through as Codex backend traffic, translated into Codex work, partially supported, or blocked?
All runtime endpoints require Pool API key bearer auth unless noted otherwise. A Pool API key represents a Pool, not one upstream account. Codex Pooler still applies Pool policy, model support, limits, account health, route class admission, session continuity, and accounting before dispatching supported work upstream.
Which Runtime Route Should I Use?
Section titled “Which Runtime Route Should I Use?”Use /backend-api/codex for Codex backend-compatible clients, /v1 for selected OpenAI SDK-compatible clients, and /mcp only for operator metadata tools. Runtime work on /backend-api and /v1 uses Pool API key bearer auth. MCP uses operator-owned MCP tokens and does not accept Pool API keys.
Quick map
Section titled “Quick map”| Route family | Use it for | Auth boundary |
|---|---|---|
/backend-api/codex | Codex backend compatibility route for Codex-compatible clients | Pool API key bearer auth |
/backend-api | Backend file bridge, audio transcription, and backend usage routes | Pool API key bearer auth |
/v1 | Narrow OpenAI-compatible /v1 support for selected SDK routes | Pool API key bearer auth |
/mcp | Read-only operator MCP endpoint for metadata lookup | Operator-owned MCP bearer token |
| Usage routes | Runtime usage checks exposed on compatibility paths | Pool API key bearer auth |
Status labels
Section titled “Status labels”| Status | Meaning |
|---|---|
| Supported | The route is an active public route and enters the normal authenticated gateway path. |
| Translated | The client calls an OpenAI-shaped /v1 route, and Codex Pooler converts the request to Codex-compatible work before dispatch. |
| Partial | The route exists, but only a narrower behavior is supported than the name may imply. |
| Unsupported | The route is blocked with a deterministic unsupported response or intentionally absent from the route surface. |
Codex backend compatibility
Section titled “Codex backend compatibility”Use /backend-api/codex for Codex-compatible clients:
https://codex-pooler.example.com/backend-api/codexThese endpoints keep Codex backend semantics and do not translate through the public OpenAI SDK adapter.
| Exposed endpoint | Upstream destination | Translation | Support | Notes |
|---|---|---|---|---|
GET /backend-api/codex/models | Model metadata served from Pool/catalog state | No Codex request translation | Supported | Lists models visible to the authenticated Pool, preserving the selected upstream’s default context_window, supported max_context_window, compaction metadata (including an absent or null threshold), effective_context_window_percent, and optional comp_hash. Pricing categories do not change the context or catalog ETag. Explicit operator context overrides still apply. Provider accounts can report different limits for the same model; Pooler selects one canonical source cohort for the Pool. Codex applies the percentage once, while /v1/models publishes the selected default’s effective value as context_length. |
POST /backend-api/codex/responses | /backend-api/codex/responses | No; backend payload is proxied through gateway normalization | Supported | Primary Codex Responses route for JSON and SSE work, including supported backend compaction. |
GET /backend-api/codex/responses | Persistent upstream Codex websocket response session | No HTTP adapter translation | Supported | Narrow backend websocket response-stream compatibility for Codex-compatible clients, including native backend compaction support. |
POST /backend-api/codex/responses/compact | /backend-api/codex/responses/compact | No | Supported | Backend compact compatibility route. |
POST /backend-api/codex/images/generations | /backend-api/codex/images/generations | No public /v1 image translation | Supported | Explicit authenticated native image proxy route. Pool image permission applies before body parsing. On either native image route, a policy-authorized effective image model that is genuinely absent from the Pool catalog may use eligible visible host capacity while its effective identifier is preserved exactly. A catalog-present but invisible target remains invalid. Image prompt and source fields stay image-specific. |
POST /backend-api/codex/images/edits | /backend-api/codex/images/edits | No public /v1 image translation | Supported | Explicit authenticated native image edit proxy route. Pool image permission applies before body parsing. The same policy-authorized, catalog-absent native image routing applies: eligible visible host capacity anchors the request while the effective image identifier is preserved exactly. A catalog-present but invisible target remains invalid. |
Backend /v1 aliases
Section titled “Backend /v1 aliases”Some Codex clients use /backend-api/codex/v1 as their base URL. Codex Pooler exposes exact aliases for that shape. These are still backend routes, not the public OpenAI-compatible /v1 surface.
| Exposed endpoint | Canonical gateway target | Translation | Support | Notes |
|---|---|---|---|---|
GET /backend-api/codex/v1/models | /backend-api/codex/models | No | Supported | Alias for backend model listing, including backend-only fields such as optional comp_hash. |
POST /backend-api/codex/v1/responses | /backend-api/codex/responses | No | Supported | Alias for backend Responses, including supported backend compaction. Prompt-cache routing locality can apply here. |
GET /backend-api/codex/v1/responses | Backend websocket response session | No | Supported | Alias for the backend websocket response-stream compatibility route, including native backend compaction support. |
POST /backend-api/codex/v1/responses/compact | /backend-api/codex/responses/compact | No | Supported | Alias for backend compact. Prompt-cache routing locality is excluded. |
POST /backend-api/codex/v1/chat/completions | /backend-api/codex/responses | Chat payload is coerced to backend Responses work | Supported | Backend alias for chat-completion-shaped Codex clients. Prompt-cache routing locality can apply here. |
Backend app-server helper routes
Section titled “Backend app-server helper routes”Codex Pooler is a model-provider runtime boundary. It does not proxy Codex
account helpers, analytics posting, thread-goal helpers, memory summaries,
search helpers, realtime helper calls, safety helper calls, identity JWKS, or
reset-credit consume operations. Configure Codex clients by pointing
model_providers.*.base_url at /backend-api/codex.
Supported Responses stream event metadata is relayed on model-provider streams as upstream sends it. Codex Pooler does not synthesize app-server notifications such as turn/safetyBuffering/updated.
On HTTP SSE, Codex Pooler holds a response’s lifecycle events (response.created, response.in_progress, response.metadata) until the attempt streams output or a terminal event, so a retry before the first output can move to another account without mixing two responses. While nothing else arrives the stream carries : keepalive comments at the SSE keepalive interval (System page, Gateway tab, 10 seconds by default). On the native /backend-api/codex/responses route, the first keepalive after a held lifecycle event is a data event instead, event: keepalive with {"type":"keepalive"}: the Codex client’s stream idle timeout (300 seconds by default) counts only events, not comments, so without it a provider that stays in its lifecycle phase longer than that made the client give up and resend while the provider was still working. Silence after that event stays comments, so the client’s idle timer still measures the provider as it would on a direct connection, and nothing but comments follows a terminal event. Codex ignores the event; a client with its own SSE parser on this route should ignore unknown event types. The /v1 routes keep comments only, and the native websocket route holds nothing back: it relays these events as they arrive.
Backend bridge and usage routes
Section titled “Backend bridge and usage routes”These routes live under /backend-api or legacy usage paths. They are still Pool API key routes. The backend file bridge stores metadata only; file bytes stay upstream-backed.
| Exposed endpoint | Upstream destination | Translation | Support | Notes |
|---|---|---|---|---|
POST /backend-api/files | Upstream file create/upload-url flow | No OpenAI multipart translation | Supported | Accepts JSON metadata only and returns upstream file metadata plus upload URL. Codex Pooler stores metadata only. |
POST /backend-api/files/:file_id/uploaded | Upstream file finalization flow | No | Supported | Marks an upstream-backed file upload as complete. |
POST /backend-api/transcribe | /backend-api/transcribe | Separate public /v1 audio transcription translation | Supported | Backend multipart transcription route. It forces the backend transcription model and preserves safe multipart fields. |
GET /api/codex/usage | Codex usage resolver | No | Supported | Compatibility usage route. |
GET /wham/usage | Codex usage resolver | No | Supported | Compatibility usage route. |
GET /backend-api/wham/usage | Codex usage resolver | No | Supported | Backend usage alias. |
The usage routes report the account quota evidence Codex Pooler keeps. A window whose reset time has passed is still reported, marked by its reset time, for 30 days after that reset. After that it is no longer evidence, and an account with nothing newer answers 404 with error code no_upstream_usage, the same answer as for an account that has never reported usage. The next successful usage read reports the account again.
The app-server JSON-RPC reset-credit method
account/rateLimitResetCredit/consume remains unsupported. Codex Pooler does
not add a wildcard backend app-server proxy for reset-credit methods.
OpenAI-compatible /v1 routes
Section titled “OpenAI-compatible /v1 routes”Use /v1 only for clients that require an OpenAI-shaped base URL:
https://codex-pooler.example.com/v1Supported work is translated into Codex-compatible requests and then routed through the same Pool policy, limit, account-selection, accounting, and session-continuity machinery as backend traffic. Codex Pooler does not provide full OpenAI API parity.
| Exposed endpoint | Gateway/upstream destination | Translated to Codex? | Support | Notes |
|---|---|---|---|---|
GET /v1/models | OpenAI-shaped model metadata from Pool/catalog state | No upstream Codex request | Supported | Returns an OpenAI-shaped model list for the authenticated Pool. Includes effective context_length when Codex context metadata is available; backend-only fields such as context_window, max_context_window, auto_compact_token_limit, and comp_hash are omitted. |
POST /v1/responses | /backend-api/codex/responses | Yes | Translated | OpenAI Responses payloads are coerced to Codex-compatible work. Input audio uses the public {data, format} shape and accepts WAV, MP3, M4A, WebM, and OGG. The decoded audio ceiling is 50 MiB (52,428,800 bytes), while the configured request envelope may reject a request earlier. Malformed, unsupported, or oversized audio returns a sanitized invalid_request before upstream dispatch. System/developer input-message text is lifted into top-level instructions. Exactly one final compaction_trigger after visible input requests an explicit compaction turn; non-streaming JSON and public SSE return the normalized compact item. Encrypted type: “compaction” replay items from prior remote compaction turns are forwarded. Remote MCP tool definitions are rejected before dispatch. Streaming responses use the public Responses stream adapter, and terminal-only failed or incomplete responses receive a response.created lifecycle prefix without implying success. If visible public output ends without a terminal Responses event, public HTTP SSE returns a sequence-valid sanitized type: “error” / nested server_error terminal. Native backend streams and websocket surfaces retain their route-native behavior. |
GET /v1/responses | Backend Codex websocket response session | Partial | Partial | Narrow Responses websocket compatibility only. It is not OpenAI Realtime SDK support. Exactly one final compaction_trigger after visible input returns the same normalized compaction item as HTTP through response.created, response.output_item.added, response.output_item.done and response.completed. Encrypted compaction replay follows the same field and ordering contract as POST /v1/responses. It uses the same downstream idle and frame guardrails as backend Codex websockets. Malformed JSON and JSON non-object provider frames are dropped without consuming a public sequence number or creating a local terminal. |
POST /v1/chat/completions | /backend-api/codex/responses | Yes | Translated | Chat Completions payloads are coerced to Codex Responses work and normalized back to Chat Completions shape. Official nested custom-tool definitions and named custom choices are translated to the supported flat Responses subset; completed and streamed custom calls return the Chat custom shape, and streamed free-form input is never parsed as function JSON. Input audio uses the public {data, format} shape and accepts WAV, MP3, M4A, WebM, and OGG. The decoded audio ceiling is 50 MiB (52,428,800 bytes), while the configured request envelope may reject a request earlier. Malformed, unsupported, or oversized audio returns a sanitized invalid_request before upstream dispatch. Early streaming terminal errors emit one public error chunk before any assistant role chunk. After visible output, a stream that ends without an upstream terminal emits one nested data: {“error”:{…}} chunk with no following [DONE]; the translated POST /backend-api/codex/v1/chat/completions alias has the same OpenAI-compatible HTTP SSE behavior. |
GET /v1/usage | Codex usage resolver | No upstream Codex work request | Supported | OpenAI-compatible usage read surface for the authenticated Pool. |
GET /v1/files | Codex Pooler file metadata | No | Supported | Lists metadata for upstream-backed files visible to the Pool. |
POST /v1/files | Upstream-backed file metadata create flow | Partial | Partial | Uploads to Codex-compatible upstream file storage. Transient storage failures retry the same upload, with at most five PUT attempts within five minutes. Backend /backend-api/files remains JSON-only; this is the OpenAI-compatible file-create surface. |
GET /v1/files/:file_id | Codex Pooler file metadata | No | Supported | Retrieves metadata only. |
GET /v1/files/:file_id/content | No upstream content read | No | Partial | Checks ownership, records an unsupported operation, then returns an OpenAI-shaped unsupported endpoint response. File bytes are not served by Codex Pooler. |
DELETE /v1/files/:file_id | No upstream delete | No | Partial | Checks ownership, records an unsupported operation, then returns an OpenAI-shaped unsupported endpoint response. |
POST /v1/audio/transcriptions | /backend-api/transcribe | Yes | Translated | OpenAI-style multipart transcription request dispatched through the backend transcription path. gpt-transcribe is a caller alias for canonical gpt-4o-transcribe. Output is JSON by default, or plain UTF-8 text with response_format: “text”. SRT, VTT and timestamped or diarized JSON are unsupported. Nonempty keywords and languages lists retain order and duplicates, while empty lists are omitted. Responses omit language-detection fields. Audio support is speech-to-text only: speech synthesis, audio translation and Realtime voice sessions are unsupported. This is request-shape compatibility, not a transcription-quality, model-discovery, or general Audio API coverage claim. |
POST /v1/images/generations | /backend-api/codex/images/generations or /backend-api/codex/responses | Yes | Translated | GPT Image 2/2.5 uses native Images generation; legacy image models use Responses translation. Pool image permission applies before body parsing. |
POST /v1/images/edits | /backend-api/codex/images/edits or /backend-api/codex/responses | Yes | Translated | Unmasked GPT Image 2/2.5 edits use native Images. Masked edits require an eligible Full Responses image-input host and explicit image tool selection; without one, mask is rejected before dispatch or reservation. Legacy models use Responses translation. Pool image permission applies before body parsing. |
For explicit remote compaction on the narrow public surface, the request must contain visible input followed by exactly one final {"type":"compaction_trigger"} item. Invalid placement returns 400 invalid_request on input before upstream dispatch. Successful HTTP JSON returns one completed Responses object; public SSE emits response.created with an empty output, response.output_item.added and response.output_item.done at output_index 0, response.completed, and [DONE]; narrow Responses websocket completion emits the same four Responses events without [DONE].
The returned replay item has type: "compaction", nonblank opaque encrypted_content, and a string id: the upstream id when it sent a nonblank one, otherwise a derived cmp_ id that Codex Pooler removes again when the item is replayed. Start the next request as a new chain without previous_response_id, put the compaction item first, and append new visible input after it. When an accepted narrow Responses websocket compaction continuation carries an explicit opaque response anchor, Codex Pooler preserves it only on the current live upstream websocket connection with the same generation and effective Full/Lite mode. That incremental compact collects provider output before validation, settlement, and adaptation; it never reconnects, retries, switches assignment, or falls back to HTTP. If that connection or mode is unavailable, the request returns the existing previous_response_not_found recovery result and the client can submit a separate no-anchor full-history request. When no anchor is present, Codex Pooler sends the supplied full history through the existing HTTP compact path. Unknown replay fields and malformed values reject before upstream dispatch. Malformed upstream compact JSON or missing encrypted content returns sanitized 502 invalid_compaction_response. Direct POST /v1/responses/compact remains unsupported.
Pool image permission
Section titled “Pool image permission”One Pool setting, allow_image_generation, controls exactly four authenticated
POST routes: /backend-api/codex/images/generations,
/backend-api/codex/images/edits, /v1/images/generations, and
/v1/images/edits. It defaults to true for new Pools and existing routing
rows without a stored value. When it is disabled, runtime ingress authenticates
the Pool API key, then returns 403 with code image_generation_disabled
before decompression, request parsing, compatibility coercion, admission, or
upstream dispatch.
Pool audio permission
Section titled “Pool audio permission”The Pool setting allow_audio_transcription controls POST /backend-api/transcribe and POST /v1/audio/transcriptions. It defaults to true, including for existing Pools. A disabled Pool returns 403 audio_transcription_disabled after Pool API-key authentication and before decompression, multipart parsing, admission, accounting reservation, or upstream dispatch. The gateway also reads the current shared setting before transcription execution. Image permission and unsupported audio routes retain their separate behavior.
POST /v1/responses accepts reasoning.context only for auto,
current_turn, and all_turns after trimming and lowercasing. Unknown, empty,
or non-string context values fail before dispatch with
param: "reasoning.context".
Reasoning effort values can come from the client request or API-key policy. The policy is derived from its configured fields:
- Unrestricted preserves the current route behavior. Omission stays absent, and any currently accepted explicit effort, including a custom value, passes through this policy.
- Allow up to permits only known efforts at or below the selected ceiling:
none,minimal,low,medium,high,xhigh,max, andultra. The Pool model’s known levels (the union across the Pool’s assignments, as/v1/modelslists them) narrow the permitted set; admission happens before routing, so it cannot know which assignment will serve. An omitted effort resolves to the permitted model default, or the highest permitted known effort. - Always use applies the legacy exact configured effort. It remains compatible even when the configured effort is absent from model metadata.
Allow up to never clamps or downgrades a request. An above-ceiling, unknown, or
custom effort, or an omission with no permitted effort, returns 400 with code
reasoning_effort_not_allowed and message
reasoning effort is not available for this API key before reservation or
upstream dispatch. Responses, backend Responses, and compact routes return
param: "reasoning.effort"; Chat Completions returns
param: "reasoning_effort". API-key model denial remains first and returns
400 model_not_allowed with param: "model".
The same policy runs after websocket upgrade for every response.create frame.
A forbidden frame uses the existing websocket error frame with the same status,
code, message, and route-native parameter. The connection is not rejected at
upgrade time.
Authenticated backend Codex model metadata does not change with the policy:
the Pool catalog serves each model’s selected upstream entry with all of its
advertised reasoning levels and its default, whatever the key’s policy, and the
policy applies when a request arrives. A Codex client can therefore offer an
effort the key does not allow; that request is refused with
400 reasoning_effort_not_allowed. Models stay visible, and public
/v1/models remains unchanged.
minimal and ultra remain distinct for policy evaluation. Backend Codex
compatibility rewrites minimal to low and ultra to max before upstream
dispatch. When the serving assignment’s own source model lists reasoning
levels without max, ultra becomes the highest level that assignment
lists (the Pool-wide union only when the assignment has no levels of its
own), and it never becomes none or minimal. none and every other explicit effort are forwarded
unchanged, even when the model catalog does not list them, so the upstream
decides whether that model accepts the value. Routing prefers a candidate whose
own source model lists the effort the turn will send, so an effort the Pool-wide
union advertises is served by an assignment that advertises it whenever one is
eligible; when no candidate lists it, every candidate stays and the upstream
refusal stands. Safe reasoning summaries retain
requested, applied, and effective values and name the rewrite as
minimal_to_low or ultra_to_<level>.
Native backend turns apply the same rule within their canonical cap. When otherwise-equivalent source partitions differ only in reasoning levels, the models catalog advertises the union from the quota-routable capability family and all reasoning variants in that family remain allowed through canonical filtering. After quota and circuit eligibility, an explicit known effort prefers an eligible assignment that lists it, preserving hard continuation pins and healthy fallbacks. This never crosses a partition that differs in another capability such as context window or serving mode. If no eligible assignment lists the effort, every effective candidate stays and the upstream decides. The resulting x-models-etag remains byte-identical to the authenticated backend models ETag because both snapshots use the same stable family union.
Non-strict function tool schemas are lowered before local validation and
upstream dispatch for backend Responses HTTP, backend Responses websocket
response.create, and public /v1/responses compatibility paths. Lowering is
limited to function tools, including nested function tools inside accepted
namespace tools. Strict function tools and strict structured-output schemas stay
on the strict validation path and are not made looser.
Direct public Responses requests also accept exact executable custom tools on
POST /v1/responses and websocket response.create. Exact top-level custom
definitions and nested custom definitions are accepted in an already-valid
namespace. functions is the
canonical Codex namespace example, not a restriction: any nonblank valid
namespace accepts the same flat function and exact custom children. A custom
tool requires type: "custom" and a nonblank name. It may include a
description, boolean defer_loading, nullable allowed_callers using direct
or programmatic, and omitted, unconstrained text, or a lark/regex grammar
format. Hosted, MCP, tool-search, nested-namespace, malformed, and duplicate
executable-name shapes are rejected; executable names must be globally unique
across top-level and namespace children.
A typed custom tool_choice resolves only a declared same-kind custom tool with
the same exact name, whether it is top-level or an accepted namespace child.
Full mode preserves that typed choice. Lite mode rejects every map-shaped
tool_choice before upstream dispatch with unsupported_parameter and
param: "tool_choice"; use automatic or explicit Full mode when a client needs
forced typed selection. That rejection follows the model’s serving mode rather
than the endpoint, so any map-shaped tool_choice is rejected on a Lite-served
model, including a Chat Completions named-function choice. See
Responses Lite and Full for what the two
serving modes change on the outgoing request. String choices such
as auto are accepted in both modes. This is separate from accepted
custom_tool_call replay input.
Translated Chat Completions accepts the official nested custom definition
shape, with type: "custom" and a nested custom object containing a nonblank
name plus optional description and format. Its named custom choice uses the
same outer wrapper and a nested custom name. Codex Pooler flattens those values
into the supported Responses request shape, then projects completed and streamed
custom_tool_call output back into the Chat custom object. Custom input stays
free-form text across split SSE deltas and is never parsed as function JSON.
Direct-Responses-only fields such as defer_loading and allowed_callers are
not accepted inside the Chat wrapper. Malformed wrappers and unrelated tool
families still fail before dispatch.
Actual execution availability still depends on the selected model and upstream account. Smoke verification retains metadata only. This narrow contract does not imply backend or broad OpenAI tool parity.
When an accepted namespace custom declaration has one exact executable name,
public Responses HTTP, SSE, and websocket output
restores that namespace on a custom_tool_call only if the provider omits it or
returns null. An explicit provider namespace is preserved. Flat, unknown, and
non-unique names remain unchanged rather than guessed.
For direct public Responses only, a strict flat function tool whose parameters
already have an object root may receive a missing nested object or array
type when the surrounding schema provides complete, unambiguous structural
evidence. This covers top-level flat functions and flat function children of an
accepted namespace, and the same repair applies to websocket response.create.
It does not repair the parameters root, explicit type values, refs, definition
tables, combinators or their descendants, annotations, unknown keywords,
ambiguous or incomplete evidence, structured outputs, Chat Completions, the
older nested function wrapper shape, or backend routes. Public Responses and
Chat reject malformed, duplicate, and unsupported explicit type values. This
strict compatibility repair is not non-strict schema lowering.
Strict structured output and function parameters on the narrow public
OpenAI-compatible surface require a direct concrete object root. A root
$ref or root anyOf is rejected. Supported nested constructs, including
local refs, remain valid below that root. Non-strict requests and native backend
Responses behavior are unchanged. POST /backend-api/codex/v1/chat/completions
uses translated Chat semantics, so it follows this public strict-root contract.
Invalid strict structured output returns HTTP 400 with
invalid_json_schema at its root schema parameter, such as
text.format.schema or response_format.json_schema.schema. Invalid strict
function parameters return HTTP 400 with invalid_function_parameters at
the applicable root parameter family: tools.<index>.parameters,
tools.<index>.function.parameters, or
tools.<namespace_index>.tools.<tool_index>.parameters. These public
rejections occur before upstream dispatch and durable accounting.
Catalog revision and final Responses envelope
Section titled “Catalog revision and final Responses envelope”The authenticated backend model routes return the same effective catalog body
for the same Pool and catalog snapshot. Both
GET /backend-api/codex/models and GET /backend-api/codex/v1/models attach a
deterministic weak ETag derived from that policy-visible body.
Successful backend Responses streams expose the same token as
X-Models-Etag. It appears on HTTP SSE response headers for the canonical and
backend-alias POST routes, and on websocket upgrade headers for the matching
GET routes. The token is produced by Codex Pooler, not relayed from upstream.
It is not exposed by compact, public /v1, usage, or unauthenticated routes.
Catalog convergence across replicas is eventual, so clients should compare a
successful Responses token with a later authenticated backend models token.
Every non-compact request that reaches the backend Responses destination has a
reasoning object and exactly one reasoning.encrypted_content entry in
include after final normalization. This applies to canonical backend HTTP and
websocket traffic, backend /v1 Responses and Chat Completions aliases, and
translated POST /v1/responses, GET /v1/responses, and
POST /v1/chat/completions traffic. Compact dispatch is excluded and keeps its
narrow compact request shape.
When upstream returns a valid parameter path with an error, failed attempt
detail may expose it as upstream_error_param. The value is limited to a
bounded field or numeric-index path. Invalid values and successful attempts
omit the field, and raw upstream error messages or rejected values are never
projected through it.
Upstream parameter validation rejections
Section titled “Upstream parameter validation rejections”When upstream rejects an ordinary Responses or Chat Completions HTTP request
with status 400, an error object of type invalid_request_error, and one of
the parameter-validation codes unsupported_value, invalid_value,
unsupported_parameter, missing_required_parameter, invalid_type, or
string_above_max_length, the client receives that rejection instead of an
empty or generic error:
{ "error": { "type": "invalid_request_error", "code": "unsupported_value", "param": "reasoning.effort", "message": "upstream rejected parameter reasoning.effort (unsupported_value); supported values: low, medium, high" }}param is a bounded field path such as reasoning.effort or input[0].content,
or null when upstream supplies no valid path. An input[N] index names the
position of the item in the input the client sent, also when Lite adds the tool
manifest and the instructions message in front of it; when an index points at an
item Codex Pooler added, or cannot be mapped back, it is left out
(input[].id). Chat Completions requests
receive the Chat field they sent where Codex Pooler renamed it, such as
reasoning_effort, max_tokens or max_completion_tokens, verbosity,
response_format, and nested tools[N].function fields. A path into the
Responses input that Codex Pooler built from messages is reported as
messages, because those input items do not correspond one to one with the
messages the client sent; other paths stay as upstream reported them. The message is written by Codex Pooler from the code
and param; the provider message is never forwarded because validation messages
can quote submitted or gateway-normalized values. For unsupported_value and
invalid_value, the message may list up to 12 supported values when upstream
reports them as a plain list of simple identifiers; the rejected value is never
included, and the list is omitted when it cannot be read safely.
Auto, Lite, Full relay the same bounded supported-values list from persisted attempt metadata.
The Codex backend answers some unsupported parameters with a {"detail": "Unsupported parameter: <name>"} body instead of an error object, notably previous_response_id on every HTTP request, because it resolves that anchor only on the websocket connection that produced the response. A public /v1/responses request anchored on previous_response_id never reaches the provider over HTTP: Codex Pooler answers it with previous_response_not_found first (see OpenAI-compatible clients), so this relay concerns the native route. A detail that is exactly that text followed by a bounded field path is relayed as unsupported_parameter with that path as param, for example upstream rejected parameter previous_response_id (unsupported_parameter), so a client can send the complete input again without the parameter. Any other detail text is not relayed.
A streaming
POST /backend-api/codex/responses request receives this JSON body with
content-type: application/json and no SSE stream, and a non-streaming native
request receives the same body instead of the upstream one.
A native /backend-api/codex/responses websocket turn receives the same error
object in one wrapped error event, {"type":"error","status":400,"error":{...}},
so the Codex client reports it as an invalid request instead of retrying it, as
it does for the HTTP answer. The public GET /v1/responses websocket sends the
same error object in its error event. Both websocket events leave an
input[N] index out, because the socket does not know how the turn’s input was
rewritten.
Any other 400 refusal of a native websocket turn, including the usual one
that carries no code, also reaches the client as one wrapped error event, so the
Codex client stops instead of resending the turn and then retrying it over
HTTPS. Codex Pooler writes its error object from the refusal’s sanitized type,
code and param only: code is the upstream code, or invalid_request when
upstream sent none, param follows the websocket rule above, and the message
reads upstream rejected the request (<code>) or names the param. The provider
message is not forwarded. A refusal naming a code the Codex client classifies
itself keeps its response.failed event, so the client still compacts on
context_length_exceeded, stops on the quota codes (insufficient_quota,
credit_balance_exhausted, organization_spend_limit_exceeded,
project_spend_limit_exceeded) and usage_not_included, and reports
invalid_prompt, cyber_policy, bio_policy and
misalignment_policy_violation as it would from the provider; overload and
rate-limit codes stay retryable.
A native websocket refusal with another final 4xx status, such as 404,
409, 413 or 422, reaches the client the same way, as one wrapped error
event with status 400, because the Codex client treats a wrapped event of any
other status as a retryable unexpected status: before, it resent the turn over
the websocket, then fell back to HTTPS and retried there, and every HTTPS
attempt reached the provider again. The message names the provider status, for
example upstream rejected the request (invalid_request); upstream status 404.
The same code exceptions apply. A 401, a 408 and a 429 keep their
response.failed event, and so does a 403 whose code makes Codex Pooler
demote the account (a code it knows as an account failure, such as a
credential code): the Codex client’s HTTPS fallback is then routed to another
account of the Pool first. When the Pool has no other routable account for the
model, that 403 is final like the others, since a retry could only reach the
demoted account again. A 403 that demotes nothing, with no code or an
unknown one, would only reach the same account again and is final like the
others. The kept 401 and demoting 403 are about the Pool’s upstream
account, not the client’s request, so their message is the Codex
Pooler-written one naming the status, for example
upstream rejected the request (unauthorized); upstream status 403; a 408
and a 429 keep the provider’s message, whose retry or limit detail the
client acts on.
A websocket refusal with any 4xx status other than 401 and 429 and a code
Codex Pooler does not know leaves route health alone, as the HTTP answer of the
same refusal does: it does not demote the account or count a circuit failure.
A code it knows keeps its own classification, so a quota or credential code
still demotes the account.
A native HTTP POST /backend-api/codex/responses answers a 400 refusal
outside the relayed validation codes (no code, an unknown code, another error
type, or any other {"detail": ...} body) with the same Codex Pooler-authored error
object as the websocket event, as JSON, streaming or
not, keeping the 400 status; an input[N] param keeps the index the client
sent. A streaming request used to receive the 400 with an empty body, which
the Codex client showed as an error with no text, and a non-streaming one the
upstream body. The Codex client reads a cyber_policy or bio_policy code in
that body as the policy refusal it names, which an empty body did not allow.
A native HTTP refusal with another final 4xx status, such as 403, 404,
409, 413 or 422, is answered the same way with status 400, whatever the
serving mode, and the message names the provider status, for example
upstream rejected the request (invalid_request); upstream status 404. The
Codex client retries every HTTP status but 400 as an unexpected status, and
each retry used to reach the provider again: a single refused turn cost six
provider requests. The request and its attempt keep the provider status.
Over HTTP, every other upstream failure keeps its existing behavior: a 401
and a 403 with a credential code (refreshed, then a retryable 503), 408,
429, 5xx, misalignment and model-unavailability refusals, and compact
routes. The request remains a non-retried failure recorded as upstream_status,
and a 4xx rejection other than 401 and 429 does not demote the upstream or
affect its circuit.
On /v1, such a failure keeps the redacted error (upstream request failed,
code upstream_status or the upstream code, no param) on HTTP and in the
GET /v1/responses websocket error event. Its type follows the status it is
answered with: invalid_request_error for a refused 4xx such as 400 or
422, rate_limit_error for a 429 (OpenAI’s type for a throttle; SDKs retry
it on the status alone), and server_error for an upstream 401 or 403 (the
upstream account’s credentials or standing, not the request), a 5xx, and an
upstream 404, which /v1 answers as 502. A 429 or a 5xx the upstream
websocket sends reaches a GET /v1/responses websocket client as the masked
response.failed event rather than an error event; its error is typed the
same way, rate_limit_error for the 429 and server_error for the 5xx.
Unsupported /v1 Responses request shapes
Section titled “Unsupported /v1 Responses request shapes”These POST /v1/responses request shapes return OpenAI-shaped invalid_request before gateway dispatch, not unsupported_endpoint:
- top-level
tools[].type = "mcp" - nested
input[].type = "additional_tools"withtools[].type = "mcp"
Routed but unsupported /v1 endpoints
Section titled “Routed but unsupported /v1 endpoints”These routes are deliberately present so clients receive deterministic OpenAI-shaped unsupported endpoint errors before gateway admission or upstream dispatch.
| Exposed endpoint | Upstream destination | Translated to Codex? | Support | Notes |
|---|---|---|---|---|
POST /v1/responses/compact | None | No | Unsupported | The backend compact route exists under /backend-api/codex; the public /v1 compact route returns unsupported_endpoint. |
POST /v1/images/variations | None | No | Unsupported | Image variations are not implemented. |
POST /v1/content_provenance_checks | None | No | Unsupported | Content provenance checks require a Platform API surface that Codex Pooler does not expose. |
POST /v1/embeddings | None | No | Unsupported | Embeddings are outside the Codex Pooler route surface. |
POST /v1/batches | None | No | Unsupported | Batch jobs are not dispatched by Codex Pooler. |
POST /v1/moderations | None | No | Unsupported | Moderations are not implemented. |
POST /v1/fine_tuning/jobs | None | No | Unsupported | Fine-tuning jobs are not implemented. |
GET /v1/responses/:response_id | None | No | Unsupported | Response retrieval by id is not part of the public compatibility surface. |
POST /v1/responses/:response_id/cancel | None | No | Unsupported | Response cancellation by id is not part of the public compatibility surface. |
DELETE /v1/responses/:response_id | None | No | Unsupported | Response deletion by id is not part of the public compatibility surface. |
Intentionally absent /v1 route families
Section titled “Intentionally absent /v1 route families”/v1/realtime and OpenAI Realtime SDK websocket/session routes are intentionally outside the public route surface. A client calling those paths should treat Codex Pooler as not supporting OpenAI Realtime. Use GET /v1/responses only for the narrow Responses websocket compatibility route documented above.
MCP endpoint
Section titled “MCP endpoint”The operator MCP endpoint is rooted at /mcp, not under /backend-api or /v1.
| Route | Meaning |
|---|---|
POST /mcp | JSON-RPC Streamable HTTP endpoint |
GET /mcp | Routed endpoint, but stateless SSE is unavailable today |
OPTIONS /mcp | Allowed MCP methods response |
MCP uses operator-owned bearer MCP tokens. It doesn’t accept Pool API keys, browser sessions, cookies, query tokens, invite tokens, upstream tokens, or custom headers as authentication.
MCP output is metadata-only and scoped by the operator’s owner or assigned-Pool visibility. It is an operator inspection endpoint, not a runtime client endpoint.
The request-log and audit-log list tools accept offset from 0 through 10,000 and reject larger values as invalid_arguments. Refine the filters to inspect another slice of history. Counts are bounded; totalExact indicates whether the total is exact, and nextOffset becomes null at the traversal cap. Audit outcome accepts only success or failure.
The root /mcp operator endpoint is not a bridge for /v1/responses remote MCP tools.
Responses access programs
Section titled “Responses access programs”The narrow /v1/responses surface accepts access_programs on HTTP JSON, HTTP SSE, and websocket response.create requests in both Full and Lite modes. The object may contain only the optional cyber selection: standard, daybreak_blue, or daybreak_red. Codex Pooler validates the shape and forwards the selection unchanged; an omitted field or empty object leaves selection to the provider. Invalid shapes, unknown keys, and unsupported selections return 400 invalid_request before upstream dispatch.
Forwarding a selection does not grant access. The upstream account and model must be eligible for the requested program, and provider refusals follow the normal upstream error contract.
Prompt-cache and continuity boundaries
Section titled “Prompt-cache and continuity boundaries”Prompt-cache routing locality is a local routing hint. It can apply on POST /backend-api/codex/responses, POST /backend-api/codex/v1/responses, POST /backend-api/codex/v1/chat/completions, POST /v1/responses, and POST /v1/chat/completions. It is excluded from websocket, compact, file, audio, image, usage, and app-server helper routes. Locality is always a heuristic and never guarantees a provider cache hit or cached-token accounting.
Websocket Guardrails
Section titled “Websocket Guardrails”Backend Codex websocket routes and the narrow public GET /v1/responses websocket route use bounded downstream guardrails. websocket_idle_timeout_ms controls the downstream websocket idle close window for new upgrades. Its default is 1_800_000 ms, and accepted values are 60_000..3_600_000 ms.
Inbound websocket frames are also bounded by the configured gateway request body limit. Operators should treat reason_class=max_frame_size_exceeded as an oversized client frame, reason_class=timeout as downstream websocket idle close, and upstream receive timeout errors as separate upstream-side failures. These classifications are metadata-only and must not include raw websocket frames or request bodies.
Tool-output preservation
Section titled “Tool-output preservation”Accepted tool-output text is preserved through the existing protocol adapters on backend Responses and their aliases, public Responses, translated Chat, native compact, and supported Responses websocket dispatches. JSON whitespace, numeric text, duplicate keys, logs, search results, diffs, source text and Unicode are not minified or summarized by the gateway. When an adapter splits a text value into ordered text parts, their concatenation preserves the original value; the complete HTTP envelope can still change during normalization.
Full and Lite retain their distinct request contracts, including tool manifests and schema normalization. HTTP gzip, deflate and zstd decoding, provider compaction and client-owned context management remain separate features. Public POST /v1/responses/compact remains unsupported.
The retired compression setting and savings display are unavailable. Existing database columns and historical metadata can remain inert during upgrades; new requests do not produce compression metadata. Old serving nodes can still apply their former behavior until replaced. Removing a rewrite changes the upstream serialized prefix on upgrade and may change prompt-cache reuse or token usage; content preservation does not guarantee the same cache-hit ratio.
Continuity headers are also local routing inputs. Codex Pooler chooses them in this order:
x-codex-window-idx-codex-session-idsession-idx-session-idx-session-affinitysession_idx-codex-conversation-id
The local session those headers select belongs to the calling API key within its Pool. Two API keys of one Pool that send the same header, for example one Codex thread resumed on two machines configured with different keys, each get their own session: neither joins, renews, or closes the other’s. A rotated API key resumes its own sessions, because rotation replaces the secret and keeps the key.
x-session-id and x-session-affinity are never forwarded upstream. The Codex client’s session-id, thread-id, and x-client-request-id are forwarded only on the native /backend-api/codex/responses and /backend-api/codex/responses/compact routes, on HTTP requests and on the websocket handshake, when each value is a short ASCII identifier; /v1 routes never forward them. On /v1/responses and /v1/chat/completions, a prompt_cache_key in the body produces a Codex Pooler-derived upstream session-id for provider sticky routing, scoped to the calling API key and its Pool so different API keys or Pools that send the same key never share one; see Routing strategies. The native /backend-api/codex/responses and /backend-api/codex/responses/compact HTTP routes send the same derived session-id when the client sent no usable session-id of its own, for example a client that names its conversation only with session_id or x-session-id. When such a request has no prompt_cache_key either, the session-id is derived the same way from that conversation header instead, under a separate namespace and with the same API key and Pool scoping; the header itself is never forwarded. A client session-id is always forwarded unchanged, and the native websocket handshake is not affected. If a pinned continuation points at an upstream account that now requires reauthentication, /v1/responses HTTP and websocket requests fail closed with a recovery hint to restart with full context and remove stale continuation anchors.
A resend of a native Codex turn that has already been sent is refused with 409 duplicate_turn on both the HTTP and the websocket form of /backend-api/codex/responses, before any upstream dispatch, accounting reservation, or request row. That refusal is scoped to the thread identity the client carries in its own turn metadata, not to the local session the continuity headers above select, so it survives the window rotation a Codex client performs after a remote compaction: a client that resends the first turn of the new window is refused instead of buying a second provider turn for history the provider already holds. Local routing, affinity and session ownership keep following the window headers above, a genuinely new turn of a rotated window is ordinary work, and a request that carries no thread identity is scoped to its local session as before. A resend that reaches a turn still running upstream rejoins that turn and is served its output rather than refused, whatever window it carries; the refusal is for a resend whose turn has already settled. A turn that settled as a failure the Codex client retries, such as a retryable provider failure, a stream cut before any completed output, or a websocket client that disconnected before any output reached it (the response lifecycle events response.created, response.in_progress and response.queued are not output), is the exception: once that request has settled, its byte-identical resend is served as a new request linked to the failed one. Byte-identical means the same request, not the same bytes on the wire: the Codex client sends every turn after the first on a websocket as an increment anchored with previous_response_id, and after a reconnect it resends that turn as full history without the anchor. That full-history resend is the same request when it ends with exactly the items the anchored request carried and every other field matches; a resend that changes the anchor, a trailing item, or any other field is a different request. A websocket turn whose anchored request was refused because its connection cannot resolve previous_response_id, either by the provider (whose Invalid previous_response_id refusal reaches the client as previous_response_not_found) or by Codex Pooler before sending an anchor that connection cannot serve, recovers the same way: the Codex client sends the turn again as full history without the anchor, and that resend is served as a new request linked to the refused one, with or without owner forwarding. With owner forwarding, a websocket that closes before the owner has accepted any turn of it, as a client cut within milliseconds of its request does, is detached at once and its turn never reaches the provider; the client’s resend is then served as that turn, on its first retry. A turn the client sends the moment its previous response completes, as a tool continuation, recovers the same way as one sent later: with owner forwarding, a disconnect before any of its output keeps it as one request that the resend completes, instead of a failed request followed by a linked one. On the websocket form, and on the HTTP form when the client falls back to it, the resend of a turn whose provider refusal went out as a final 400 is answered with that same error rather than 409 duplicate_turn, and is not sent to the provider again; the same holds for the HTTP resend of a turn first sent over HTTP and refused there with a validation error the provider repeats for the same request, while any other refusal of an HTTP turn is sent to the provider again; this is what a Codex client’s in-band compaction, which resends a refused compaction request several times, then shows. A websocket turn the provider completed after its client had already disconnected, and of which nothing was sent to that client, is served again as a new request when the client resends it; the provider is asked twice and each request is recorded and billed once. The same holds for a websocket turn cut after its client was shown only the response lifecycle events, the opening of an output item or of a content or reasoning-summary part, and output deltas: the Codex client discards that partial output and resends the request unchanged, as it would to the provider directly, and the resend is served as a new request linked to the cut one, whether the cut turn was stopped or the provider completed it afterwards. A turn whose client was shown one or more completed output items, and not the end of the response, is resent differently: the Codex client keeps those completed items and resends the request with exactly those items appended, as it would to the provider directly. That resend is served as a new request linked to the cut one when it is the cut request plus exactly the completed items its client was shown, in the order it was shown them; each item is recognised by its content, ignoring only the fields the Codex client does not keep, such as the item status and text annotations. An appended item the client was not shown, an extra item, or any other change keeps the 409 duplicate_turn refusal, and so does a turn whose client was shown the end of the response. What a client was shown is what Codex Pooler managed to write to its connection: when a write fails, as it does for a client that stopped reading once its connection is dropped or its writes time out, nothing from that write on counts as shown, so a turn whose connection broke before the end of the response was written is treated as cut at the last output written before the failure, while a turn whose end was written keeps the refusal even if its connection then dropped. Such a resend is served only within 30 seconds of the cut request’s settlement or, when a write to its connection failed, within 30 seconds of that failure: a client that stopped reading without disconnecting is noticed only when a write times out, 30 seconds later by default, and its resend is still served when it arrives within 30 seconds of that timeout. Without owner forwarding, the closing connection stops a turn cut in either of these ways instead of letting it keep generating beside its resend. The resend the Codex client sends over HTTPS after its websocket retries have failed is the same request as its websocket form, so it is admitted, linked or answered by the same rules. A later request of a turn is not a resend of it: the Codex client sends user input typed while a turn is running as another request of that same turn, ending with the new message, including right after a compaction the turn made. On the HTTP form that request is served as a later request of the turn when it is further along than the turn’s first HTTP request: more user input after the same compaction (or with no compaction), or a compaction the first request did not end on. Its own resend is refused like any other, and so is a resend of the first request that carries less user input or has lost its compaction, since that is the first request again with part of its history trimmed. On the websocket form it arrives on the same connection anchored on the response the turn just completed there, and is served as a later request of the turn, after that response has settled; its identical resend is refused. When that connection closed before the client sent it, the client sends it as full history on a new websocket connection, and when the session has fallen back to HTTPS it sends it over HTTP; either is served as a later request of the turn when it is further along than the turn’s first request in the same sense, whichever transport that first request used, and the same request in any of these forms, sent again, is refused. A turn whose first request was recorded by an earlier release, or went over a websocket connection that could not know the history its anchor stood for, keeps refusing such a request. A remote compaction the Codex client resends over HTTP because it never received the reply, which it does up to twice with the same request, is served as a new request linked to the previous attempt, and each attempt is billed once.
Native HTTP streaming requests have one additional recovery case: if the stream ends during a single incomplete client tool call, Codex Pooler can admit one identical client retry within 30 seconds. This requires complete stream evidence that no output item or terminal response was completed. Finishing the tool’s input alone does not execute it; completing the output item does. Requests with completed output, missing or malformed evidence, changed authorization or changed request contents retain their existing duplicate protection. This recovery applies to ordinary native turns and tool continuations in Full and Lite modes.
Related references
Section titled “Related references”- Routing Strategies explains Pool routing policy, account selection, and prompt-cache locality.
Frequently Asked Questions
Section titled “Frequently Asked Questions”Is /v1 the same as the OpenAI API?
Section titled “Is /v1 the same as the OpenAI API?”No. /v1 is narrow OpenAI-compatible support for selected SDK routes. Supported requests are translated into Codex-compatible work and routed through Pool policy. Unsupported routes are either intentionally absent or return deterministic OpenAI-shaped unsupported endpoint errors when explicitly routed.
Is GET /v1/responses OpenAI Realtime support?
Section titled “Is GET /v1/responses OpenAI Realtime support?”No. GET /v1/responses is narrow Responses websocket compatibility. It is not /v1/realtime support, and OpenAI Realtime SDK websocket or session routes are outside the Codex Pooler route surface.
Can a Pool API key call /mcp?
Section titled “Can a Pool API key call /mcp?”No. Pool API keys authenticate runtime clients for /backend-api and /v1. The root /mcp endpoint is for operator metadata only and requires an operator-owned MCP bearer token.
Does Codex Pooler store uploaded file bytes?
Section titled “Does Codex Pooler store uploaded file bytes?”No. The backend file bridge stores file metadata and uses upstream-backed upload or download URLs. Raw file bytes, upload URLs, prompts, response bodies, media bodies, credentials, and websocket frames are not stored or exposed as public docs evidence.
For multipart POST /v1/files, the storage upload URL must use HTTPS and resolve exclusively to public addresses. Codex Pooler checks both address families before creating local file metadata and pins one validated address for every upload attempt. A configured HTTPS proxy receives that literal address in CONNECT, while TLS verification and the storage Host header retain the original hostname. DNS failure, private or reserved answers, and mixed public/private answers are rejected; storage redirects are not followed.