Skip to content

Runtime Routes

Codex Pooler exposes a deliberately narrow runtime surface. It is not a wildcard proxy and it does not claim full OpenAI API parity. Each route below is either a Codex backend compatibility route, a backend bridge route, a translated /v1 compatibility route, an operator metadata route, or an explicit unsupported boundary.

Use this page when you need to answer three questions before wiring a client:

  • Which endpoint should the client call?
  • Which upstream route does Codex Pooler call after admission and routing?
  • Is the request passed through as Codex backend traffic, translated into Codex work, partially supported, or blocked?

All runtime endpoints require Pool API key bearer auth unless noted otherwise. A Pool API key represents a Pool, not one upstream account. Codex Pooler still applies Pool policy, model support, limits, account health, route class admission, session continuity, and accounting before dispatching supported work upstream.

Use /backend-api/codex for Codex backend-compatible clients, /v1 for selected OpenAI SDK-compatible clients, and /mcp only for operator metadata tools. Runtime work on /backend-api and /v1 uses Pool API key bearer auth. MCP uses operator-owned MCP tokens and does not accept Pool API keys.

Route familyUse it forAuth boundary
/backend-api/codexCodex backend compatibility route for Codex-compatible clientsPool API key bearer auth
/backend-apiBackend file bridge, audio transcription, and backend usage routesPool API key bearer auth
/v1Narrow OpenAI-compatible /v1 support for selected SDK routesPool API key bearer auth
/mcpRead-only operator MCP endpoint for metadata lookupOperator-owned MCP bearer token
Usage routesRuntime usage checks exposed on compatibility pathsPool API key bearer auth
StatusMeaning
SupportedThe route is an active public route and enters the normal authenticated gateway path.
TranslatedThe client calls an OpenAI-shaped /v1 route, and Codex Pooler converts the request to Codex-compatible work before dispatch.
PartialThe route exists, but only a narrower behavior is supported than the name may imply.
UnsupportedThe route is blocked with a deterministic unsupported response or intentionally absent from the route surface.

Use /backend-api/codex for Codex-compatible clients:

https://codex-pooler.example.com/backend-api/codex

These endpoints keep Codex backend semantics and do not translate through the public OpenAI SDK adapter.

Exposed endpointUpstream destinationTranslationSupportNotes
GET /backend-api/codex/modelsModel metadata served from Pool/catalog stateNo Codex request translationSupportedLists models visible to the authenticated Pool, preserving the selected upstream’s default context_window, supported max_context_window, compaction metadata (including an absent or null threshold), effective_context_window_percent, and optional comp_hash. Pricing categories do not change the context or catalog ETag. Explicit operator context overrides still apply. Provider accounts can report different limits for the same model; Pooler selects one canonical source cohort for the Pool. Codex applies the percentage once, while /v1/models publishes the selected default’s effective value as context_length.
POST /backend-api/codex/responses/backend-api/codex/responsesNo; backend payload is proxied through gateway normalizationSupportedPrimary Codex Responses route for JSON and SSE work, including supported backend compaction.
GET /backend-api/codex/responsesPersistent upstream Codex websocket response sessionNo HTTP adapter translationSupportedNarrow backend websocket response-stream compatibility for Codex-compatible clients, including native backend compaction support.
POST /backend-api/codex/responses/compact/backend-api/codex/responses/compactNoSupportedBackend compact compatibility route.
POST /backend-api/codex/images/generations/backend-api/codex/images/generationsNo public /v1 image translationSupportedExplicit authenticated native image proxy route. Pool image permission applies before body parsing. On either native image route, a policy-authorized effective image model that is genuinely absent from the Pool catalog may use eligible visible host capacity while its effective identifier is preserved exactly. A catalog-present but invisible target remains invalid. Image prompt and source fields stay image-specific.
POST /backend-api/codex/images/edits/backend-api/codex/images/editsNo public /v1 image translationSupportedExplicit authenticated native image edit proxy route. Pool image permission applies before body parsing. The same policy-authorized, catalog-absent native image routing applies: eligible visible host capacity anchors the request while the effective image identifier is preserved exactly. A catalog-present but invisible target remains invalid.

Some Codex clients use /backend-api/codex/v1 as their base URL. Codex Pooler exposes exact aliases for that shape. These are still backend routes, not the public OpenAI-compatible /v1 surface.

Exposed endpointCanonical gateway targetTranslationSupportNotes
GET /backend-api/codex/v1/models/backend-api/codex/modelsNoSupportedAlias for backend model listing, including backend-only fields such as optional comp_hash.
POST /backend-api/codex/v1/responses/backend-api/codex/responsesNoSupportedAlias for backend Responses, including supported backend compaction. Prompt-cache routing locality can apply here.
GET /backend-api/codex/v1/responsesBackend websocket response sessionNoSupportedAlias for the backend websocket response-stream compatibility route, including native backend compaction support.
POST /backend-api/codex/v1/responses/compact/backend-api/codex/responses/compactNoSupportedAlias for backend compact. Prompt-cache routing locality is excluded.
POST /backend-api/codex/v1/chat/completions/backend-api/codex/responsesChat payload is coerced to backend Responses workSupportedBackend alias for chat-completion-shaped Codex clients. Prompt-cache routing locality can apply here.

Codex Pooler is a model-provider runtime boundary. It does not proxy Codex account helpers, analytics posting, thread-goal helpers, memory summaries, search helpers, realtime helper calls, safety helper calls, identity JWKS, or reset-credit consume operations. Configure Codex clients by pointing model_providers.*.base_url at /backend-api/codex. Supported Responses stream event metadata is relayed on model-provider streams as upstream sends it. Codex Pooler does not synthesize app-server notifications such as turn/safetyBuffering/updated.

On HTTP SSE, Codex Pooler holds a response’s lifecycle events (response.created, response.in_progress, response.metadata) until the attempt streams output or a terminal event, so a retry before the first output can move to another account without mixing two responses. While nothing else arrives the stream carries : keepalive comments at the SSE keepalive interval (System page, Gateway tab, 10 seconds by default). On the native /backend-api/codex/responses route, the first keepalive after a held lifecycle event is a data event instead, event: keepalive with {"type":"keepalive"}: the Codex client’s stream idle timeout (300 seconds by default) counts only events, not comments, so without it a provider that stays in its lifecycle phase longer than that made the client give up and resend while the provider was still working. Silence after that event stays comments, so the client’s idle timer still measures the provider as it would on a direct connection, and nothing but comments follows a terminal event. Codex ignores the event; a client with its own SSE parser on this route should ignore unknown event types. The /v1 routes keep comments only, and the native websocket route holds nothing back: it relays these events as they arrive.

These routes live under /backend-api or legacy usage paths. They are still Pool API key routes. The backend file bridge stores metadata only; file bytes stay upstream-backed.

Exposed endpointUpstream destinationTranslationSupportNotes
POST /backend-api/filesUpstream file create/upload-url flowNo OpenAI multipart translationSupportedAccepts JSON metadata only and returns upstream file metadata plus upload URL. Codex Pooler stores metadata only.
POST /backend-api/files/:file_id/uploadedUpstream file finalization flowNoSupportedMarks an upstream-backed file upload as complete.
POST /backend-api/transcribe/backend-api/transcribeSeparate public /v1 audio transcription translationSupportedBackend multipart transcription route. It forces the backend transcription model and preserves safe multipart fields.
GET /api/codex/usageCodex usage resolverNoSupportedCompatibility usage route.
GET /wham/usageCodex usage resolverNoSupportedCompatibility usage route.
GET /backend-api/wham/usageCodex usage resolverNoSupportedBackend usage alias.

The usage routes report the account quota evidence Codex Pooler keeps. A window whose reset time has passed is still reported, marked by its reset time, for 30 days after that reset. After that it is no longer evidence, and an account with nothing newer answers 404 with error code no_upstream_usage, the same answer as for an account that has never reported usage. The next successful usage read reports the account again.

The app-server JSON-RPC reset-credit method account/rateLimitResetCredit/consume remains unsupported. Codex Pooler does not add a wildcard backend app-server proxy for reset-credit methods.

Use /v1 only for clients that require an OpenAI-shaped base URL:

https://codex-pooler.example.com/v1

Supported work is translated into Codex-compatible requests and then routed through the same Pool policy, limit, account-selection, accounting, and session-continuity machinery as backend traffic. Codex Pooler does not provide full OpenAI API parity.

Exposed endpointGateway/upstream destinationTranslated to Codex?SupportNotes
GET /v1/modelsOpenAI-shaped model metadata from Pool/catalog stateNo upstream Codex requestSupportedReturns an OpenAI-shaped model list for the authenticated Pool. Includes effective context_length when Codex context metadata is available; backend-only fields such as context_window, max_context_window, auto_compact_token_limit, and comp_hash are omitted.
POST /v1/responses/backend-api/codex/responsesYesTranslatedOpenAI Responses payloads are coerced to Codex-compatible work. Input audio uses the public {data, format} shape and accepts WAV, MP3, M4A, WebM, and OGG. The decoded audio ceiling is 50 MiB (52,428,800 bytes), while the configured request envelope may reject a request earlier. Malformed, unsupported, or oversized audio returns a sanitized invalid_request before upstream dispatch. System/developer input-message text is lifted into top-level instructions. Exactly one final compaction_trigger after visible input requests an explicit compaction turn; non-streaming JSON and public SSE return the normalized compact item. Encrypted type: “compaction” replay items from prior remote compaction turns are forwarded. Remote MCP tool definitions are rejected before dispatch. Streaming responses use the public Responses stream adapter, and terminal-only failed or incomplete responses receive a response.created lifecycle prefix without implying success. If visible public output ends without a terminal Responses event, public HTTP SSE returns a sequence-valid sanitized type: “error” / nested server_error terminal. Native backend streams and websocket surfaces retain their route-native behavior.
GET /v1/responsesBackend Codex websocket response sessionPartialPartialNarrow Responses websocket compatibility only. It is not OpenAI Realtime SDK support. Exactly one final compaction_trigger after visible input returns the same normalized compaction item as HTTP through response.created, response.output_item.added, response.output_item.done and response.completed. Encrypted compaction replay follows the same field and ordering contract as POST /v1/responses. It uses the same downstream idle and frame guardrails as backend Codex websockets. Malformed JSON and JSON non-object provider frames are dropped without consuming a public sequence number or creating a local terminal.
POST /v1/chat/completions/backend-api/codex/responsesYesTranslatedChat Completions payloads are coerced to Codex Responses work and normalized back to Chat Completions shape. Official nested custom-tool definitions and named custom choices are translated to the supported flat Responses subset; completed and streamed custom calls return the Chat custom shape, and streamed free-form input is never parsed as function JSON. Input audio uses the public {data, format} shape and accepts WAV, MP3, M4A, WebM, and OGG. The decoded audio ceiling is 50 MiB (52,428,800 bytes), while the configured request envelope may reject a request earlier. Malformed, unsupported, or oversized audio returns a sanitized invalid_request before upstream dispatch. Early streaming terminal errors emit one public error chunk before any assistant role chunk. After visible output, a stream that ends without an upstream terminal emits one nested data: {“error”:{…}} chunk with no following [DONE]; the translated POST /backend-api/codex/v1/chat/completions alias has the same OpenAI-compatible HTTP SSE behavior.
GET /v1/usageCodex usage resolverNo upstream Codex work requestSupportedOpenAI-compatible usage read surface for the authenticated Pool.
GET /v1/filesCodex Pooler file metadataNoSupportedLists metadata for upstream-backed files visible to the Pool.
POST /v1/filesUpstream-backed file metadata create flowPartialPartialUploads to Codex-compatible upstream file storage. Transient storage failures retry the same upload, with at most five PUT attempts within five minutes. Backend /backend-api/files remains JSON-only; this is the OpenAI-compatible file-create surface.
GET /v1/files/:file_idCodex Pooler file metadataNoSupportedRetrieves metadata only.
GET /v1/files/:file_id/contentNo upstream content readNoPartialChecks ownership, records an unsupported operation, then returns an OpenAI-shaped unsupported endpoint response. File bytes are not served by Codex Pooler.
DELETE /v1/files/:file_idNo upstream deleteNoPartialChecks ownership, records an unsupported operation, then returns an OpenAI-shaped unsupported endpoint response.
POST /v1/audio/transcriptions/backend-api/transcribeYesTranslatedOpenAI-style multipart transcription request dispatched through the backend transcription path. gpt-transcribe is a caller alias for canonical gpt-4o-transcribe. Output is JSON by default, or plain UTF-8 text with response_format: “text”. SRT, VTT and timestamped or diarized JSON are unsupported. Nonempty keywords and languages lists retain order and duplicates, while empty lists are omitted. Responses omit language-detection fields. Audio support is speech-to-text only: speech synthesis, audio translation and Realtime voice sessions are unsupported. This is request-shape compatibility, not a transcription-quality, model-discovery, or general Audio API coverage claim.
POST /v1/images/generations/backend-api/codex/images/generations or /backend-api/codex/responsesYesTranslatedGPT Image 2/2.5 uses native Images generation; legacy image models use Responses translation. Pool image permission applies before body parsing.
POST /v1/images/edits/backend-api/codex/images/edits or /backend-api/codex/responsesYesTranslatedUnmasked GPT Image 2/2.5 edits use native Images. Masked edits require an eligible Full Responses image-input host and explicit image tool selection; without one, mask is rejected before dispatch or reservation. Legacy models use Responses translation. Pool image permission applies before body parsing.

For explicit remote compaction on the narrow public surface, the request must contain visible input followed by exactly one final {"type":"compaction_trigger"} item. Invalid placement returns 400 invalid_request on input before upstream dispatch. Successful HTTP JSON returns one completed Responses object; public SSE emits response.created with an empty output, response.output_item.added and response.output_item.done at output_index 0, response.completed, and [DONE]; narrow Responses websocket completion emits the same four Responses events without [DONE].

The returned replay item has type: "compaction", nonblank opaque encrypted_content, and a string id: the upstream id when it sent a nonblank one, otherwise a derived cmp_ id that Codex Pooler removes again when the item is replayed. Start the next request as a new chain without previous_response_id, put the compaction item first, and append new visible input after it. When an accepted narrow Responses websocket compaction continuation carries an explicit opaque response anchor, Codex Pooler preserves it only on the current live upstream websocket connection with the same generation and effective Full/Lite mode. That incremental compact collects provider output before validation, settlement, and adaptation; it never reconnects, retries, switches assignment, or falls back to HTTP. If that connection or mode is unavailable, the request returns the existing previous_response_not_found recovery result and the client can submit a separate no-anchor full-history request. When no anchor is present, Codex Pooler sends the supplied full history through the existing HTTP compact path. Unknown replay fields and malformed values reject before upstream dispatch. Malformed upstream compact JSON or missing encrypted content returns sanitized 502 invalid_compaction_response. Direct POST /v1/responses/compact remains unsupported.

One Pool setting, allow_image_generation, controls exactly four authenticated POST routes: /backend-api/codex/images/generations, /backend-api/codex/images/edits, /v1/images/generations, and /v1/images/edits. It defaults to true for new Pools and existing routing rows without a stored value. When it is disabled, runtime ingress authenticates the Pool API key, then returns 403 with code image_generation_disabled before decompression, request parsing, compatibility coercion, admission, or upstream dispatch.

The Pool setting allow_audio_transcription controls POST /backend-api/transcribe and POST /v1/audio/transcriptions. It defaults to true, including for existing Pools. A disabled Pool returns 403 audio_transcription_disabled after Pool API-key authentication and before decompression, multipart parsing, admission, accounting reservation, or upstream dispatch. The gateway also reads the current shared setting before transcription execution. Image permission and unsupported audio routes retain their separate behavior.

POST /v1/responses accepts reasoning.context only for auto, current_turn, and all_turns after trimming and lowercasing. Unknown, empty, or non-string context values fail before dispatch with param: "reasoning.context".

Reasoning effort values can come from the client request or API-key policy. The policy is derived from its configured fields:

  1. Unrestricted preserves the current route behavior. Omission stays absent, and any currently accepted explicit effort, including a custom value, passes through this policy.
  2. Allow up to permits only known efforts at or below the selected ceiling: none, minimal, low, medium, high, xhigh, max, and ultra. The Pool model’s known levels (the union across the Pool’s assignments, as /v1/models lists them) narrow the permitted set; admission happens before routing, so it cannot know which assignment will serve. An omitted effort resolves to the permitted model default, or the highest permitted known effort.
  3. Always use applies the legacy exact configured effort. It remains compatible even when the configured effort is absent from model metadata.

Allow up to never clamps or downgrades a request. An above-ceiling, unknown, or custom effort, or an omission with no permitted effort, returns 400 with code reasoning_effort_not_allowed and message reasoning effort is not available for this API key before reservation or upstream dispatch. Responses, backend Responses, and compact routes return param: "reasoning.effort"; Chat Completions returns param: "reasoning_effort". API-key model denial remains first and returns 400 model_not_allowed with param: "model".

The same policy runs after websocket upgrade for every response.create frame. A forbidden frame uses the existing websocket error frame with the same status, code, message, and route-native parameter. The connection is not rejected at upgrade time.

Authenticated backend Codex model metadata does not change with the policy: the Pool catalog serves each model’s selected upstream entry with all of its advertised reasoning levels and its default, whatever the key’s policy, and the policy applies when a request arrives. A Codex client can therefore offer an effort the key does not allow; that request is refused with 400 reasoning_effort_not_allowed. Models stay visible, and public /v1/models remains unchanged.

minimal and ultra remain distinct for policy evaluation. Backend Codex compatibility rewrites minimal to low and ultra to max before upstream dispatch. When the serving assignment’s own source model lists reasoning levels without max, ultra becomes the highest level that assignment lists (the Pool-wide union only when the assignment has no levels of its own), and it never becomes none or minimal. none and every other explicit effort are forwarded unchanged, even when the model catalog does not list them, so the upstream decides whether that model accepts the value. Routing prefers a candidate whose own source model lists the effort the turn will send, so an effort the Pool-wide union advertises is served by an assignment that advertises it whenever one is eligible; when no candidate lists it, every candidate stays and the upstream refusal stands. Safe reasoning summaries retain requested, applied, and effective values and name the rewrite as minimal_to_low or ultra_to_<level>.

Native backend turns apply the same rule within their canonical cap. When otherwise-equivalent source partitions differ only in reasoning levels, the models catalog advertises the union from the quota-routable capability family and all reasoning variants in that family remain allowed through canonical filtering. After quota and circuit eligibility, an explicit known effort prefers an eligible assignment that lists it, preserving hard continuation pins and healthy fallbacks. This never crosses a partition that differs in another capability such as context window or serving mode. If no eligible assignment lists the effort, every effective candidate stays and the upstream decides. The resulting x-models-etag remains byte-identical to the authenticated backend models ETag because both snapshots use the same stable family union.

Non-strict function tool schemas are lowered before local validation and upstream dispatch for backend Responses HTTP, backend Responses websocket response.create, and public /v1/responses compatibility paths. Lowering is limited to function tools, including nested function tools inside accepted namespace tools. Strict function tools and strict structured-output schemas stay on the strict validation path and are not made looser.

Direct public Responses requests also accept exact executable custom tools on POST /v1/responses and websocket response.create. Exact top-level custom definitions and nested custom definitions are accepted in an already-valid namespace. functions is the canonical Codex namespace example, not a restriction: any nonblank valid namespace accepts the same flat function and exact custom children. A custom tool requires type: "custom" and a nonblank name. It may include a description, boolean defer_loading, nullable allowed_callers using direct or programmatic, and omitted, unconstrained text, or a lark/regex grammar format. Hosted, MCP, tool-search, nested-namespace, malformed, and duplicate executable-name shapes are rejected; executable names must be globally unique across top-level and namespace children.

A typed custom tool_choice resolves only a declared same-kind custom tool with the same exact name, whether it is top-level or an accepted namespace child. Full mode preserves that typed choice. Lite mode rejects every map-shaped tool_choice before upstream dispatch with unsupported_parameter and param: "tool_choice"; use automatic or explicit Full mode when a client needs forced typed selection. That rejection follows the model’s serving mode rather than the endpoint, so any map-shaped tool_choice is rejected on a Lite-served model, including a Chat Completions named-function choice. See Responses Lite and Full for what the two serving modes change on the outgoing request. String choices such as auto are accepted in both modes. This is separate from accepted custom_tool_call replay input.

Translated Chat Completions accepts the official nested custom definition shape, with type: "custom" and a nested custom object containing a nonblank name plus optional description and format. Its named custom choice uses the same outer wrapper and a nested custom name. Codex Pooler flattens those values into the supported Responses request shape, then projects completed and streamed custom_tool_call output back into the Chat custom object. Custom input stays free-form text across split SSE deltas and is never parsed as function JSON. Direct-Responses-only fields such as defer_loading and allowed_callers are not accepted inside the Chat wrapper. Malformed wrappers and unrelated tool families still fail before dispatch.

Actual execution availability still depends on the selected model and upstream account. Smoke verification retains metadata only. This narrow contract does not imply backend or broad OpenAI tool parity.

When an accepted namespace custom declaration has one exact executable name, public Responses HTTP, SSE, and websocket output restores that namespace on a custom_tool_call only if the provider omits it or returns null. An explicit provider namespace is preserved. Flat, unknown, and non-unique names remain unchanged rather than guessed.

For direct public Responses only, a strict flat function tool whose parameters already have an object root may receive a missing nested object or array type when the surrounding schema provides complete, unambiguous structural evidence. This covers top-level flat functions and flat function children of an accepted namespace, and the same repair applies to websocket response.create. It does not repair the parameters root, explicit type values, refs, definition tables, combinators or their descendants, annotations, unknown keywords, ambiguous or incomplete evidence, structured outputs, Chat Completions, the older nested function wrapper shape, or backend routes. Public Responses and Chat reject malformed, duplicate, and unsupported explicit type values. This strict compatibility repair is not non-strict schema lowering.

Strict structured output and function parameters on the narrow public OpenAI-compatible surface require a direct concrete object root. A root $ref or root anyOf is rejected. Supported nested constructs, including local refs, remain valid below that root. Non-strict requests and native backend Responses behavior are unchanged. POST /backend-api/codex/v1/chat/completions uses translated Chat semantics, so it follows this public strict-root contract. Invalid strict structured output returns HTTP 400 with invalid_json_schema at its root schema parameter, such as text.format.schema or response_format.json_schema.schema. Invalid strict function parameters return HTTP 400 with invalid_function_parameters at the applicable root parameter family: tools.<index>.parameters, tools.<index>.function.parameters, or tools.<namespace_index>.tools.<tool_index>.parameters. These public rejections occur before upstream dispatch and durable accounting.

Catalog revision and final Responses envelope

Section titled “Catalog revision and final Responses envelope”

The authenticated backend model routes return the same effective catalog body for the same Pool and catalog snapshot. Both GET /backend-api/codex/models and GET /backend-api/codex/v1/models attach a deterministic weak ETag derived from that policy-visible body.

Successful backend Responses streams expose the same token as X-Models-Etag. It appears on HTTP SSE response headers for the canonical and backend-alias POST routes, and on websocket upgrade headers for the matching GET routes. The token is produced by Codex Pooler, not relayed from upstream. It is not exposed by compact, public /v1, usage, or unauthenticated routes. Catalog convergence across replicas is eventual, so clients should compare a successful Responses token with a later authenticated backend models token.

Every non-compact request that reaches the backend Responses destination has a reasoning object and exactly one reasoning.encrypted_content entry in include after final normalization. This applies to canonical backend HTTP and websocket traffic, backend /v1 Responses and Chat Completions aliases, and translated POST /v1/responses, GET /v1/responses, and POST /v1/chat/completions traffic. Compact dispatch is excluded and keeps its narrow compact request shape.

When upstream returns a valid parameter path with an error, failed attempt detail may expose it as upstream_error_param. The value is limited to a bounded field or numeric-index path. Invalid values and successful attempts omit the field, and raw upstream error messages or rejected values are never projected through it.

When upstream rejects an ordinary Responses or Chat Completions HTTP request with status 400, an error object of type invalid_request_error, and one of the parameter-validation codes unsupported_value, invalid_value, unsupported_parameter, missing_required_parameter, invalid_type, or string_above_max_length, the client receives that rejection instead of an empty or generic error:

{
"error": {
"type": "invalid_request_error",
"code": "unsupported_value",
"param": "reasoning.effort",
"message": "upstream rejected parameter reasoning.effort (unsupported_value); supported values: low, medium, high"
}
}

param is a bounded field path such as reasoning.effort or input[0].content, or null when upstream supplies no valid path. An input[N] index names the position of the item in the input the client sent, also when Lite adds the tool manifest and the instructions message in front of it; when an index points at an item Codex Pooler added, or cannot be mapped back, it is left out (input[].id). Chat Completions requests receive the Chat field they sent where Codex Pooler renamed it, such as reasoning_effort, max_tokens or max_completion_tokens, verbosity, response_format, and nested tools[N].function fields. A path into the Responses input that Codex Pooler built from messages is reported as messages, because those input items do not correspond one to one with the messages the client sent; other paths stay as upstream reported them. The message is written by Codex Pooler from the code and param; the provider message is never forwarded because validation messages can quote submitted or gateway-normalized values. For unsupported_value and invalid_value, the message may list up to 12 supported values when upstream reports them as a plain list of simple identifiers; the rejected value is never included, and the list is omitted when it cannot be read safely.

Auto, Lite, Full relay the same bounded supported-values list from persisted attempt metadata.

The Codex backend answers some unsupported parameters with a {"detail": "Unsupported parameter: <name>"} body instead of an error object, notably previous_response_id on every HTTP request, because it resolves that anchor only on the websocket connection that produced the response. A public /v1/responses request anchored on previous_response_id never reaches the provider over HTTP: Codex Pooler answers it with previous_response_not_found first (see OpenAI-compatible clients), so this relay concerns the native route. A detail that is exactly that text followed by a bounded field path is relayed as unsupported_parameter with that path as param, for example upstream rejected parameter previous_response_id (unsupported_parameter), so a client can send the complete input again without the parameter. Any other detail text is not relayed.

A streaming POST /backend-api/codex/responses request receives this JSON body with content-type: application/json and no SSE stream, and a non-streaming native request receives the same body instead of the upstream one.

A native /backend-api/codex/responses websocket turn receives the same error object in one wrapped error event, {"type":"error","status":400,"error":{...}}, so the Codex client reports it as an invalid request instead of retrying it, as it does for the HTTP answer. The public GET /v1/responses websocket sends the same error object in its error event. Both websocket events leave an input[N] index out, because the socket does not know how the turn’s input was rewritten.

Any other 400 refusal of a native websocket turn, including the usual one that carries no code, also reaches the client as one wrapped error event, so the Codex client stops instead of resending the turn and then retrying it over HTTPS. Codex Pooler writes its error object from the refusal’s sanitized type, code and param only: code is the upstream code, or invalid_request when upstream sent none, param follows the websocket rule above, and the message reads upstream rejected the request (<code>) or names the param. The provider message is not forwarded. A refusal naming a code the Codex client classifies itself keeps its response.failed event, so the client still compacts on context_length_exceeded, stops on the quota codes (insufficient_quota, credit_balance_exhausted, organization_spend_limit_exceeded, project_spend_limit_exceeded) and usage_not_included, and reports invalid_prompt, cyber_policy, bio_policy and misalignment_policy_violation as it would from the provider; overload and rate-limit codes stay retryable.

A native websocket refusal with another final 4xx status, such as 404, 409, 413 or 422, reaches the client the same way, as one wrapped error event with status 400, because the Codex client treats a wrapped event of any other status as a retryable unexpected status: before, it resent the turn over the websocket, then fell back to HTTPS and retried there, and every HTTPS attempt reached the provider again. The message names the provider status, for example upstream rejected the request (invalid_request); upstream status 404. The same code exceptions apply. A 401, a 408 and a 429 keep their response.failed event, and so does a 403 whose code makes Codex Pooler demote the account (a code it knows as an account failure, such as a credential code): the Codex client’s HTTPS fallback is then routed to another account of the Pool first. When the Pool has no other routable account for the model, that 403 is final like the others, since a retry could only reach the demoted account again. A 403 that demotes nothing, with no code or an unknown one, would only reach the same account again and is final like the others. The kept 401 and demoting 403 are about the Pool’s upstream account, not the client’s request, so their message is the Codex Pooler-written one naming the status, for example upstream rejected the request (unauthorized); upstream status 403; a 408 and a 429 keep the provider’s message, whose retry or limit detail the client acts on.

A websocket refusal with any 4xx status other than 401 and 429 and a code Codex Pooler does not know leaves route health alone, as the HTTP answer of the same refusal does: it does not demote the account or count a circuit failure. A code it knows keeps its own classification, so a quota or credential code still demotes the account.

A native HTTP POST /backend-api/codex/responses answers a 400 refusal outside the relayed validation codes (no code, an unknown code, another error type, or any other {"detail": ...} body) with the same Codex Pooler-authored error object as the websocket event, as JSON, streaming or not, keeping the 400 status; an input[N] param keeps the index the client sent. A streaming request used to receive the 400 with an empty body, which the Codex client showed as an error with no text, and a non-streaming one the upstream body. The Codex client reads a cyber_policy or bio_policy code in that body as the policy refusal it names, which an empty body did not allow.

A native HTTP refusal with another final 4xx status, such as 403, 404, 409, 413 or 422, is answered the same way with status 400, whatever the serving mode, and the message names the provider status, for example upstream rejected the request (invalid_request); upstream status 404. The Codex client retries every HTTP status but 400 as an unexpected status, and each retry used to reach the provider again: a single refused turn cost six provider requests. The request and its attempt keep the provider status.

Over HTTP, every other upstream failure keeps its existing behavior: a 401 and a 403 with a credential code (refreshed, then a retryable 503), 408, 429, 5xx, misalignment and model-unavailability refusals, and compact routes. The request remains a non-retried failure recorded as upstream_status, and a 4xx rejection other than 401 and 429 does not demote the upstream or affect its circuit.

On /v1, such a failure keeps the redacted error (upstream request failed, code upstream_status or the upstream code, no param) on HTTP and in the GET /v1/responses websocket error event. Its type follows the status it is answered with: invalid_request_error for a refused 4xx such as 400 or 422, rate_limit_error for a 429 (OpenAI’s type for a throttle; SDKs retry it on the status alone), and server_error for an upstream 401 or 403 (the upstream account’s credentials or standing, not the request), a 5xx, and an upstream 404, which /v1 answers as 502. A 429 or a 5xx the upstream websocket sends reaches a GET /v1/responses websocket client as the masked response.failed event rather than an error event; its error is typed the same way, rate_limit_error for the 429 and server_error for the 5xx.

These POST /v1/responses request shapes return OpenAI-shaped invalid_request before gateway dispatch, not unsupported_endpoint:

  • top-level tools[].type = "mcp"
  • nested input[].type = "additional_tools" with tools[].type = "mcp"

These routes are deliberately present so clients receive deterministic OpenAI-shaped unsupported endpoint errors before gateway admission or upstream dispatch.

Exposed endpointUpstream destinationTranslated to Codex?SupportNotes
POST /v1/responses/compactNoneNoUnsupportedThe backend compact route exists under /backend-api/codex; the public /v1 compact route returns unsupported_endpoint.
POST /v1/images/variationsNoneNoUnsupportedImage variations are not implemented.
POST /v1/content_provenance_checksNoneNoUnsupportedContent provenance checks require a Platform API surface that Codex Pooler does not expose.
POST /v1/embeddingsNoneNoUnsupportedEmbeddings are outside the Codex Pooler route surface.
POST /v1/batchesNoneNoUnsupportedBatch jobs are not dispatched by Codex Pooler.
POST /v1/moderationsNoneNoUnsupportedModerations are not implemented.
POST /v1/fine_tuning/jobsNoneNoUnsupportedFine-tuning jobs are not implemented.
GET /v1/responses/:response_idNoneNoUnsupportedResponse retrieval by id is not part of the public compatibility surface.
POST /v1/responses/:response_id/cancelNoneNoUnsupportedResponse cancellation by id is not part of the public compatibility surface.
DELETE /v1/responses/:response_idNoneNoUnsupportedResponse deletion by id is not part of the public compatibility surface.

/v1/realtime and OpenAI Realtime SDK websocket/session routes are intentionally outside the public route surface. A client calling those paths should treat Codex Pooler as not supporting OpenAI Realtime. Use GET /v1/responses only for the narrow Responses websocket compatibility route documented above.

The operator MCP endpoint is rooted at /mcp, not under /backend-api or /v1.

RouteMeaning
POST /mcpJSON-RPC Streamable HTTP endpoint
GET /mcpRouted endpoint, but stateless SSE is unavailable today
OPTIONS /mcpAllowed MCP methods response

MCP uses operator-owned bearer MCP tokens. It doesn’t accept Pool API keys, browser sessions, cookies, query tokens, invite tokens, upstream tokens, or custom headers as authentication.

MCP output is metadata-only and scoped by the operator’s owner or assigned-Pool visibility. It is an operator inspection endpoint, not a runtime client endpoint.

The request-log and audit-log list tools accept offset from 0 through 10,000 and reject larger values as invalid_arguments. Refine the filters to inspect another slice of history. Counts are bounded; totalExact indicates whether the total is exact, and nextOffset becomes null at the traversal cap. Audit outcome accepts only success or failure.

The root /mcp operator endpoint is not a bridge for /v1/responses remote MCP tools.

The narrow /v1/responses surface accepts access_programs on HTTP JSON, HTTP SSE, and websocket response.create requests in both Full and Lite modes. The object may contain only the optional cyber selection: standard, daybreak_blue, or daybreak_red. Codex Pooler validates the shape and forwards the selection unchanged; an omitted field or empty object leaves selection to the provider. Invalid shapes, unknown keys, and unsupported selections return 400 invalid_request before upstream dispatch.

Forwarding a selection does not grant access. The upstream account and model must be eligible for the requested program, and provider refusals follow the normal upstream error contract.

Prompt-cache routing locality is a local routing hint. It can apply on POST /backend-api/codex/responses, POST /backend-api/codex/v1/responses, POST /backend-api/codex/v1/chat/completions, POST /v1/responses, and POST /v1/chat/completions. It is excluded from websocket, compact, file, audio, image, usage, and app-server helper routes. Locality is always a heuristic and never guarantees a provider cache hit or cached-token accounting.

Backend Codex websocket routes and the narrow public GET /v1/responses websocket route use bounded downstream guardrails. websocket_idle_timeout_ms controls the downstream websocket idle close window for new upgrades. Its default is 1_800_000 ms, and accepted values are 60_000..3_600_000 ms.

Inbound websocket frames are also bounded by the configured gateway request body limit. Operators should treat reason_class=max_frame_size_exceeded as an oversized client frame, reason_class=timeout as downstream websocket idle close, and upstream receive timeout errors as separate upstream-side failures. These classifications are metadata-only and must not include raw websocket frames or request bodies.

Accepted tool-output text is preserved through the existing protocol adapters on backend Responses and their aliases, public Responses, translated Chat, native compact, and supported Responses websocket dispatches. JSON whitespace, numeric text, duplicate keys, logs, search results, diffs, source text and Unicode are not minified or summarized by the gateway. When an adapter splits a text value into ordered text parts, their concatenation preserves the original value; the complete HTTP envelope can still change during normalization.

Full and Lite retain their distinct request contracts, including tool manifests and schema normalization. HTTP gzip, deflate and zstd decoding, provider compaction and client-owned context management remain separate features. Public POST /v1/responses/compact remains unsupported.

The retired compression setting and savings display are unavailable. Existing database columns and historical metadata can remain inert during upgrades; new requests do not produce compression metadata. Old serving nodes can still apply their former behavior until replaced. Removing a rewrite changes the upstream serialized prefix on upgrade and may change prompt-cache reuse or token usage; content preservation does not guarantee the same cache-hit ratio.

Continuity headers are also local routing inputs. Codex Pooler chooses them in this order:

  1. x-codex-window-id
  2. x-codex-session-id
  3. session-id
  4. x-session-id
  5. x-session-affinity
  6. session_id
  7. x-codex-conversation-id

The local session those headers select belongs to the calling API key within its Pool. Two API keys of one Pool that send the same header, for example one Codex thread resumed on two machines configured with different keys, each get their own session: neither joins, renews, or closes the other’s. A rotated API key resumes its own sessions, because rotation replaces the secret and keeps the key.

x-session-id and x-session-affinity are never forwarded upstream. The Codex client’s session-id, thread-id, and x-client-request-id are forwarded only on the native /backend-api/codex/responses and /backend-api/codex/responses/compact routes, on HTTP requests and on the websocket handshake, when each value is a short ASCII identifier; /v1 routes never forward them. On /v1/responses and /v1/chat/completions, a prompt_cache_key in the body produces a Codex Pooler-derived upstream session-id for provider sticky routing, scoped to the calling API key and its Pool so different API keys or Pools that send the same key never share one; see Routing strategies. The native /backend-api/codex/responses and /backend-api/codex/responses/compact HTTP routes send the same derived session-id when the client sent no usable session-id of its own, for example a client that names its conversation only with session_id or x-session-id. When such a request has no prompt_cache_key either, the session-id is derived the same way from that conversation header instead, under a separate namespace and with the same API key and Pool scoping; the header itself is never forwarded. A client session-id is always forwarded unchanged, and the native websocket handshake is not affected. If a pinned continuation points at an upstream account that now requires reauthentication, /v1/responses HTTP and websocket requests fail closed with a recovery hint to restart with full context and remove stale continuation anchors.

A resend of a native Codex turn that has already been sent is refused with 409 duplicate_turn on both the HTTP and the websocket form of /backend-api/codex/responses, before any upstream dispatch, accounting reservation, or request row. That refusal is scoped to the thread identity the client carries in its own turn metadata, not to the local session the continuity headers above select, so it survives the window rotation a Codex client performs after a remote compaction: a client that resends the first turn of the new window is refused instead of buying a second provider turn for history the provider already holds. Local routing, affinity and session ownership keep following the window headers above, a genuinely new turn of a rotated window is ordinary work, and a request that carries no thread identity is scoped to its local session as before. A resend that reaches a turn still running upstream rejoins that turn and is served its output rather than refused, whatever window it carries; the refusal is for a resend whose turn has already settled. A turn that settled as a failure the Codex client retries, such as a retryable provider failure, a stream cut before any completed output, or a websocket client that disconnected before any output reached it (the response lifecycle events response.created, response.in_progress and response.queued are not output), is the exception: once that request has settled, its byte-identical resend is served as a new request linked to the failed one. Byte-identical means the same request, not the same bytes on the wire: the Codex client sends every turn after the first on a websocket as an increment anchored with previous_response_id, and after a reconnect it resends that turn as full history without the anchor. That full-history resend is the same request when it ends with exactly the items the anchored request carried and every other field matches; a resend that changes the anchor, a trailing item, or any other field is a different request. A websocket turn whose anchored request was refused because its connection cannot resolve previous_response_id, either by the provider (whose Invalid previous_response_id refusal reaches the client as previous_response_not_found) or by Codex Pooler before sending an anchor that connection cannot serve, recovers the same way: the Codex client sends the turn again as full history without the anchor, and that resend is served as a new request linked to the refused one, with or without owner forwarding. With owner forwarding, a websocket that closes before the owner has accepted any turn of it, as a client cut within milliseconds of its request does, is detached at once and its turn never reaches the provider; the client’s resend is then served as that turn, on its first retry. A turn the client sends the moment its previous response completes, as a tool continuation, recovers the same way as one sent later: with owner forwarding, a disconnect before any of its output keeps it as one request that the resend completes, instead of a failed request followed by a linked one. On the websocket form, and on the HTTP form when the client falls back to it, the resend of a turn whose provider refusal went out as a final 400 is answered with that same error rather than 409 duplicate_turn, and is not sent to the provider again; the same holds for the HTTP resend of a turn first sent over HTTP and refused there with a validation error the provider repeats for the same request, while any other refusal of an HTTP turn is sent to the provider again; this is what a Codex client’s in-band compaction, which resends a refused compaction request several times, then shows. A websocket turn the provider completed after its client had already disconnected, and of which nothing was sent to that client, is served again as a new request when the client resends it; the provider is asked twice and each request is recorded and billed once. The same holds for a websocket turn cut after its client was shown only the response lifecycle events, the opening of an output item or of a content or reasoning-summary part, and output deltas: the Codex client discards that partial output and resends the request unchanged, as it would to the provider directly, and the resend is served as a new request linked to the cut one, whether the cut turn was stopped or the provider completed it afterwards. A turn whose client was shown one or more completed output items, and not the end of the response, is resent differently: the Codex client keeps those completed items and resends the request with exactly those items appended, as it would to the provider directly. That resend is served as a new request linked to the cut one when it is the cut request plus exactly the completed items its client was shown, in the order it was shown them; each item is recognised by its content, ignoring only the fields the Codex client does not keep, such as the item status and text annotations. An appended item the client was not shown, an extra item, or any other change keeps the 409 duplicate_turn refusal, and so does a turn whose client was shown the end of the response. What a client was shown is what Codex Pooler managed to write to its connection: when a write fails, as it does for a client that stopped reading once its connection is dropped or its writes time out, nothing from that write on counts as shown, so a turn whose connection broke before the end of the response was written is treated as cut at the last output written before the failure, while a turn whose end was written keeps the refusal even if its connection then dropped. Such a resend is served only within 30 seconds of the cut request’s settlement or, when a write to its connection failed, within 30 seconds of that failure: a client that stopped reading without disconnecting is noticed only when a write times out, 30 seconds later by default, and its resend is still served when it arrives within 30 seconds of that timeout. Without owner forwarding, the closing connection stops a turn cut in either of these ways instead of letting it keep generating beside its resend. The resend the Codex client sends over HTTPS after its websocket retries have failed is the same request as its websocket form, so it is admitted, linked or answered by the same rules. A later request of a turn is not a resend of it: the Codex client sends user input typed while a turn is running as another request of that same turn, ending with the new message, including right after a compaction the turn made. On the HTTP form that request is served as a later request of the turn when it is further along than the turn’s first HTTP request: more user input after the same compaction (or with no compaction), or a compaction the first request did not end on. Its own resend is refused like any other, and so is a resend of the first request that carries less user input or has lost its compaction, since that is the first request again with part of its history trimmed. On the websocket form it arrives on the same connection anchored on the response the turn just completed there, and is served as a later request of the turn, after that response has settled; its identical resend is refused. When that connection closed before the client sent it, the client sends it as full history on a new websocket connection, and when the session has fallen back to HTTPS it sends it over HTTP; either is served as a later request of the turn when it is further along than the turn’s first request in the same sense, whichever transport that first request used, and the same request in any of these forms, sent again, is refused. A turn whose first request was recorded by an earlier release, or went over a websocket connection that could not know the history its anchor stood for, keeps refusing such a request. A remote compaction the Codex client resends over HTTP because it never received the reply, which it does up to twice with the same request, is served as a new request linked to the previous attempt, and each attempt is billed once.

Native HTTP streaming requests have one additional recovery case: if the stream ends during a single incomplete client tool call, Codex Pooler can admit one identical client retry within 30 seconds. This requires complete stream evidence that no output item or terminal response was completed. Finishing the tool’s input alone does not execute it; completing the output item does. Requests with completed output, missing or malformed evidence, changed authorization or changed request contents retain their existing duplicate protection. This recovery applies to ordinary native turns and tool continuations in Full and Lite modes.

  • Routing Strategies explains Pool routing policy, account selection, and prompt-cache locality.

No. /v1 is narrow OpenAI-compatible support for selected SDK routes. Supported requests are translated into Codex-compatible work and routed through Pool policy. Unsupported routes are either intentionally absent or return deterministic OpenAI-shaped unsupported endpoint errors when explicitly routed.

Is GET /v1/responses OpenAI Realtime support?

Section titled “Is GET /v1/responses OpenAI Realtime support?”

No. GET /v1/responses is narrow Responses websocket compatibility. It is not /v1/realtime support, and OpenAI Realtime SDK websocket or session routes are outside the Codex Pooler route surface.

No. Pool API keys authenticate runtime clients for /backend-api and /v1. The root /mcp endpoint is for operator metadata only and requires an operator-owned MCP bearer token.

Does Codex Pooler store uploaded file bytes?

Section titled “Does Codex Pooler store uploaded file bytes?”

No. The backend file bridge stores file metadata and uses upstream-backed upload or download URLs. Raw file bytes, upload URLs, prompts, response bodies, media bodies, credentials, and websocket frames are not stored or exposed as public docs evidence.

For multipart POST /v1/files, the storage upload URL must use HTTPS and resolve exclusively to public addresses. Codex Pooler checks both address families before creating local file metadata and pins one validated address for every upload attempt. A configured HTTPS proxy receives that literal address in CONNECT, while TLS verification and the storage Host header retain the original hostname. DNS failure, private or reserved answers, and mixed public/private answers are rejected; storage redirects are not followed.