OpenAI-Compatible SDKs
Use Codex Pooler in your own applications with the OpenAI Python and Node SDKs or the Vercel AI SDK. This guide covers the supported ways to generate text, use tools, work with images and transcribe audio with a Pool API key, along with examples you can adapt to your application.

For a short citable summary, see OpenAI-compatible Codex gateway.
What OpenAI-Compatible Means Here
Section titled “What OpenAI-Compatible Means Here”OpenAI-compatible means selected SDK request shapes can use a /v1 base URL and a Pool API key. Codex Pooler still routes the request through Codex account Pools, not through a separate OpenAI engine. Unsupported OpenAI API areas remain unsupported, including embeddings, batches, fine-tuning, moderation, response retrieve/cancel/delete, image variations, OpenAI Responses remote MCP tool definitions, and OpenAI Realtime SDK routes.
Base URL and authentication
Section titled “Base URL and authentication”Use the /v1 base URL and a Pool API key:
Base URL:https://codex-pooler.example.com/v1
Authorization:Bearer <pool-api-key>For local setup, use http://localhost:4000/v1.
Set the key in the terminal used to run the SDK examples:
macOS / Linux / WSL
export CODEX_POOLER_API_KEY="<pool-api-key>"Windows PowerShell
$env:CODEX_POOLER_API_KEY = "<pool-api-key>"These commands apply to the current terminal. If the application runs as a service, configure the same variable in that service’s environment. The Python, JavaScript and TypeScript examples below use the same code on all supported operating systems.
Service tier compatibility
Section titled “Service tier compatibility”Use service_tier: "priority" for priority processing. fast is an accepted
equivalent request spelling, but priority is the canonical spelling for new
configuration. On /v1, Codex Pooler translates supported OpenAI request and
response shapes while preserving any projected provider service_tier value in
its literal provider vocabulary. Codex backend relay routes under
/backend-api/codex preserve provider bytes, frames, and service-tier
vocabulary unchanged. The ChatGPT Codex backend reports
service_tier: "default" on completed responses even when the request asked
for priority, and Codex Pooler relays that value as reported.
ultrafast is a separate tier, not an alias for fast or priority. Use
service_tier: "ultrafast" only with direct /v1/responses requests when the
selected model metadata advertises it:
response = client.responses.create( model="your-model-id", input="Your request input.", service_tier="ultrafast",)JSON, SSE, and narrow Responses WebSocket requests preserve a returned
ultrafast tier literally. POST /v1/chat/completions rejects ultrafast.
Provider availability, access, and price control whether the tier can be used;
model metadata advertisement is not an entitlement or price promise.
Tool-specific setup pages
Section titled “Tool-specific setup pages”Use the dedicated setup pages when configuring an agent or editor that has its own provider shape:
- Aider
- Cline
- Continue
- DeepSeek Harness
- Goose
- Hermes Agent
- Kilo Code
- OpenCode
- OpenClaw
- OpenHands Agent Canvas
- OMP
- Pi
- Trae
- Windmill AI
Python SDK
Section titled “Python SDK”import os
from openai import OpenAI
client = OpenAI( api_key=os.environ["CODEX_POOLER_API_KEY"], base_url="https://codex-pooler.example.com/v1",)
response = client.responses.create( model="gpt-6-sol", input="Write a short setup confirmation.",)
print(response.output_text)Node SDK
Section titled “Node SDK”import OpenAI from "openai";
const client = new OpenAI({ apiKey: process.env.CODEX_POOLER_API_KEY, baseURL: "https://codex-pooler.example.com/v1",});
const response = await client.responses.create({ model: "gpt-6-sol", input: "Write a short setup confirmation.",});
console.log(response.output_text);Vercel AI SDK
Section titled “Vercel AI SDK”import { createOpenAI } from "@ai-sdk/openai";import { generateText } from "ai";
const pooler = createOpenAI({ apiKey: process.env.CODEX_POOLER_API_KEY, baseURL: "https://codex-pooler.example.com/v1",});
const { text } = await generateText({ model: pooler.responses("gpt-6-sol"), providerOptions: { openai: { promptCacheOptions: { mode: "explicit", ttl: "30m" }, }, }, messages: [ { role: "system", content: "Keep setup guidance concise.", providerOptions: { openai: { promptCacheBreakpoint: { mode: "explicit" } }, }, }, { role: "user", content: "Write a short setup confirmation." }, ],});
console.log(text);Codex Pooler accepts these OpenAI cache controls as public input. For
account-backed requests, it adapts the explicit controls automatically before
upstream dispatch while preserving prompt_cache_key independently from Pool
routing affinity. GPT-6 model ids are passed through to the Pool catalog and
assignment policy; naming one here does not guarantee that every Pool has an
eligible account for that model.
Programmatic tools and stateless replay
Section titled “Programmatic tools and stateless replay”For Vercel AI SDK programmatic tools, set providerOptions.openai.store: false
on every generation that participates in the tool loop:
const result = await generateText({ model: pooler.responses("gpt-6-sol"), providerOptions: { openai: { store: false }, }, // Your programmatic tool and tool-choice configuration goes here. prompt: appPrompt,});Vercel chooses complete stateless replay or a stored item_reference before
the request reaches Codex Pooler. store: false is therefore required for the
complete stateless replay path. Independently, Codex Pooler forces upstream
stream: true and store: false.
This is narrow, closed-world compatibility, not general OpenAI programmatic-tool
or Responses parity. Supported shapes are program and program_output,
caller-enabled function_call and function_call_output, the type-only
programmatic_tool_calling tool and tool choice, plus allowed_callers and
map-only output_schema. Reference-only and ordinary stored-response
continuations are not supported. Codex Pooler does not locally persist raw
program code, results, schemas, or identifiers. A request accepted by the
gateway can still be rejected when its selected upstream model or account lacks
the required capability.
Responses function_call_output replay has two closed forms over HTTP and the
narrow Responses websocket bridge. A paired item requires a nonblank call_id;
its existing output form and paired-only legacy result form are unchanged. A
named standalone item requires a nonblank name and an output field, permits
call_id only when omitted or null, and permits namespace only when omitted,
null, or a nonblank string. Blank or non-string values and standalone result
reject before dispatch. Classification and debug summaries remain metadata-only.
This does not expand general Responses parity or prove provider-live acceptance.
Hosted shell history replay
Section titled “Hosted shell history replay”OpenAI Responses history may include completed shell_call and
shell_call_output items from an earlier response. Codex Pooler accepts this
closed-key replay subset, including direct or program callers, local or container
environment metadata, in_progress, completed, or incomplete status, and
empty command or output arrays where the shape permits them. It forwards accepted
history for stateless replay and semantic tool-output continuations without
requiring call/output pairing or a particular item order.
This is history forwarding only. Codex Pooler does not execute shell commands or
accept a top-level shell tool declaration. local_shell_call history and remote
MCP tool definitions remain unsupported. It also does not implement SDK-local
command-index accumulation or claim general hosted-tool parity.
Responses SSE and the narrow Responses websocket preserve the five hosted-shell
relay event types: command added, command delta, command done, output-content
delta, and output-content done. Existing public sequence normalization and the
websocket stream_id behavior are the only permitted relay adaptations.
Commands and shell output remain transient request data. They are not persisted or rendered in request metadata, logs, or operator views.
For a previous_response_id tool continuation, a native reasoning replay item
may include content with reasoning_text parts. Codex Pooler preserves that
accepted reasoning content for the upstream continuation. Stateless replay drops
reasoning items before dispatch instead.
Assistant-history replay may include output_text URL citations. Codex Pooler
accepts only ordered url_citation annotations with exactly type,
start_index, end_index, url, and title; it preserves accepted values,
explicit empty annotation lists, and omission exactly. Malformed or unsupported
annotation shapes are rejected before dispatch rather than filtered.
GET /v1/models may include context_length for clients that probe OpenAI-compatible model lists, such as Hermes. That probe is an ordinary authenticated request: a client that issues it without the Pool API key is refused like any other route, and a client that caches the refusal will keep using whatever context window it bundles for the model instead. Configure the context window explicitly wherever the client allows it. The official OpenAI SDK request APIs and Vercel AI SDK generation APIs do not expose Codex model-catalog context controls. Their output-budget fields are max_output_tokens in OpenAI Responses, max_completion_tokens in Chat Completions, and maxOutputTokens at the Vercel AI SDK layer; availability at the SDK layer does not establish upstream support or a hard generation limit. Codex Pooler’s public /v1/responses rejects context_management, and direct POST /v1/responses/compact remains unsupported. Explicit compaction is available on the normal /v1/responses route.
Continuing with previous_response_id
Section titled “Continuing with previous_response_id”A nonblank previous_response_id is accepted only for a tool-output continuation. If the input has no tool-result shape, Codex Pooler returns 400 invalid_request with param: previous_response_id before checking the upstream connection. This validation rule applies to HTTP and the Responses websocket.
Codex Pooler forces store: false upstream, and the upstream provider resolves previous_response_id only on the websocket connection that produced the response. Over HTTP it rejects the parameter, and on any other websocket connection, a new one included, it rejects the anchor. On /v1/responses an anchored request therefore reaches the earlier context in two cases: on the Responses websocket route, or as a stream: true HTTP request on a deployment with websocket owner forwarding (see the websocket replica caveat) when every request of the chain, the first one included, sends the same session header, such as session-id. Codex Pooler then sends those requests over its session’s upstream websocket.
After input validation, an unavailable producing connection returns 400 previous_response_not_found. Codex Pooler rejects locally detectable failures before dispatch and maps the provider’s matching stale-anchor refusal to the same error:
{ "error": { "type": "invalid_request_error", "code": "previous_response_not_found", "param": "previous_response_id", "message": "Previous response not found on this request's upstream connection: ..." }}This covers a non-streaming request, every HTTP request on a deployment without owner forwarding, and a streaming request whose session has no upstream connection that produced the anchor, such as a chain without a session header or a session whose connection was closed. The message states the requirement above. When you receive this error, send the complete input again without previous_response_id; clients that recognize the previous_response_not_found code as a stale chain do this on their own.
Output budgets and local usage
Section titled “Output budgets and local usage”Codex Pooler checks token policy against admission estimates. The output reservation floor is 512 tokens, or 2,048 when context is opaque; higher accepted request or policy values can increase the estimate. Continuation references and other opaque input prevent a complete input count from the visible payload. These are local accounting estimates in both Full and Lite serving modes, not provider hard bounds. The selected upstream may refuse max_output_tokens; forwarding a field or switching serving modes does not guarantee output enforcement. Actual measured usage is recorded without clipping, so a single response can exceed its reservation or remaining token budget. There is no absolute per-response spend cap.
GET /v1/usage separates measured reporting from local enforcement. Existing total_tokens and priced cost totals remain measured-usage reporting; a price snapshot produces an estimate, not an invoice or a provider-confirmed charge. budget_usage.daily and budget_usage.weekly each expose known_total_tokens, provisional_total_tokens, pending_total_tokens, effective_total_tokens, and admission_count. Daily means midnight UTC onward; weekly means the trailing seven days. These components cover only the authenticated key across models and routes, never other keys.
Effective pressure is known plus provisional plus pending. Attempted work with unknown terminal usage keeps its original estimate provisionally until a known correction or expiry of the original terminal window. A correction replaces that estimate once at the original terminal time. Pre-attempt failures and requests the Pooler refuses before anything is sent upstream create no provisional consumption; pending reservations remain pressure across a window boundary. RPM counts each admitted reservation once, without refunds or movement on correction.
The local limits list describes active default-policy thresholds; it is not a model-override simulator. A matching active model override replaces the default binding for admission, including its blank fields, while comparing its thresholds with whole-key pressure. Provider quota windows remain separate in upstream_limits; Codex-compatible usage endpoints retain their upstream quota authority when available. See API key limits for policy and concurrency configuration.
An optional key-wide max_active_requests limit spans the fleet and all models and routes. A trusted local saturation denial is 429 api_key_concurrency_limit_exceeded, typed rate_limit_error, with HTTP Retry-After: 1. Back off before retrying; the interval does not promise availability. Native and narrow public Responses websockets receive the equivalent turn error and can retry on the same connection when a slot becomes available. A provider error that happens to use the same code is not trusted as a local concurrency denial and does not acquire its retry header or type. Authentication 401 and policy-budget 403 remain distinct.
Request an explicit compaction turn
Section titled “Request an explicit compaction turn”Provider clients that expose an OpenAI compaction option can use the normal /v1/responses route. Vercel AI SDK serializes the required trigger with providerOptions.openai.compactionTrigger: true in @ai-sdk/openai 4.0.42 or later:
import { createOpenAI, type OpenAILanguageModelResponsesOptions,} from "@ai-sdk/openai";import { generateText } from "ai";
const pooler = createOpenAI({ apiKey: process.env.CODEX_POOLER_API_KEY, baseURL: "https://codex-pooler.example.com/v1",});
const result = await generateText({ model: pooler.responses("gpt-6-sol"), providerOptions: { openai: { store: false, compactionTrigger: true, } satisfies OpenAILanguageModelResponsesOptions, }, prompt: "Compact the visible conversation context.",});
const compaction = result.content.find( (part) => part.type === "custom" && part.kind === "openai.compaction",);The serialized request must contain visible input followed by exactly one final {"type":"compaction_trigger"} item. Non-terminal, duplicate, trigger-only, hidden-only, or otherwise malformed placement returns an OpenAI-shaped 400 invalid_request on input before upstream dispatch.
Successful non-streaming HTTP returns a completed Responses JSON object containing the normalized compaction item. Public SSE follows the Responses streaming grammar the official SDK stream helpers expect: response.created with an empty output, response.output_item.added and response.output_item.done for the compaction item at output_index 0, then response.completed and [DONE]. Narrow Responses websocket completion emits the same four Responses events without the HTTP [DONE] sentinel. Direct POST /v1/responses/compact remains unsupported.
If upstream compact output is malformed JSON or does not contain nonblank encrypted compaction content, Codex Pooler returns a sanitized 502 invalid_compaction_response. Other provider failures follow the public error rules documented below.
Replay a remote-compaction item
Section titled “Replay a remote-compaction item”When a compaction turn returns an encrypted type: "compaction" output item, submit that item unchanged at the start of the next POST /v1/responses input, followed by the new user input. Start a new chain by omitting previous_response_id; the compact response envelope id is not a replay anchor.
The stable public replay fields are type, a nonblank opaque encrypted_content string, and a string id. When the upstream item carries no usable id, Codex Pooler returns a derived cmp_ id and removes that derived id again when the item is replayed, so the upstream receives the item as it produced it; a provider id is forwarded unchanged. JSON, SSE, and narrow Responses websocket surfaces return the same normalized item. Native-only metadata and unknown fields are removed from output; unknown replay fields or malformed values are rejected before upstream dispatch. Treat encrypted_content as opaque and do not log or persist it in application telemetry.
Supported or translated route support
Section titled “Supported or translated route support”The OpenAI-compatible /v1 surface supports or translates selected routes only:
GET /v1/modelsPOST /v1/responsesGET /v1/responses, narrow Responses websocket compatibility onlyPOST /v1/chat/completionsGET /v1/usageGET /v1/filesPOST /v1/filesGET /v1/files/:file_idPOST /v1/audio/transcriptionsPOST /v1/images/generationsPOST /v1/images/edits
The /v1 surface is compatibility over Codex routing, not a separate OpenAI engine. Supported requests still require a Pool API key and a Pool with eligible upstream capacity for the requested model.
Chat Completions request compatibility
Section titled “Chat Completions request compatibility”POST /v1/chat/completions accepts normal messages requests. It also accepts
Responses-shaped input with reasoning, text, and include when messages
is absent or empty. Combining these fallback fields with nonempty messages
returns invalid_request, except that a string reasoning value is accepted
as an alias for Chat’s reasoning_effort. The alias follows the same validation
and API-key reasoning policy. If both effort fields are present, their normalized
values must match; conflicting values return invalid_request. A Responses
reasoning object remains invalid alongside nonempty messages.
Both shapes return Chat Completions JSON or SSE. stream_options.include_usage
controls the terminal streamed usage chunk and is not forwarded upstream.
Successful function or custom-tool turns return finish_reason: "tool_calls",
including streams whose terminal event omits the output list.
Streamed function arguments and custom-tool input can be completed from the upstream’s final argument, item, or completed-response snapshot. Pooler appends only the missing suffix when the tool identity and already-delivered bytes agree. Conflicting snapshots end the stream with a sanitized error rather than reporting a successful tool call.
Chat function tools keep Chat’s non-strict default: omitted or null strict
is forwarded as strict: false. Explicit strict: true remains strict.
Direct Responses requests and Responses-shaped fallback requests retain their
own strict-mode defaults.
For ordinary Chat messages, text fields larger than 10 MiB are divided into
smaller UTF-8 content parts before dispatch, preserving all text and its order.
This includes message history and textual tool outputs. Oversized combined
system/developer instructions keep their normalized content in one leading
developer message with multiple parts. Nothing is truncated or summarized.
Request-size and model context limits still apply; opaque fields such as tool
arguments and encoded media are not split. Direct Responses and input
fallback requests retain their existing representation.
Replayed Chat tool-call identifiers longer than 64 UTF-8 bytes are mapped to
stable, bounded identifiers before dispatch. The matching call and result use
the same identifier, including across repeated history requests. Short
identifiers, tool names, arguments and outputs stay unchanged. If a mapped
identifier would collide with another identifier in the request, Pooler rejects
the request instead of merging distinct calls. Direct Responses and input
fallback requests keep their original identifiers.
The optional user field accepts a string or null and is discarded before
dispatch. It does not establish ownership, permissions, or routing identity.
A tool message may carry image_url parts next to its text, as Hermes sends a
screenshot tool result when it talks Chat Completions. Codex Pooler forwards them
as images of the tool output, with image_url.detail handled like a user image.
A file part in a tool message returns invalid_request.
Image generation and edits
Section titled “Image generation and edits”Use gpt-image-2.5-flare for current image examples, or select gpt-image-2.5-sunburst explicitly. Both Flare and Sunburst are accepted; gpt-image-2, gpt-image-1.5, gpt-image-1, and gpt-image-1-mini remain available for compatibility. Media identifiers do not have to appear in the Pool’s model catalog for these translated routes to work.
const image = await client.images.generate({ model: "gpt-image-2.5-flare", prompt: "A simple watercolor landscape", size: "1024x1024", quality: "medium",});For GPT Image 2.5, the translated Images routes accept auto or custom WIDTHxHEIGHT sizes: both edges must be multiples of 16, neither may exceed 3840 pixels, the aspect ratio must stay between 1:3 and 3:1, and the total area must be between 655,360 and 8,294,400 pixels. Resolutions above 2560x1440 are experimental upstream. Quality accepts auto, low, medium, high, xhigh, and max. Older models retain the three standard sizes plus auto, with quality up to high. Requests still produce one image. The dated identifiers gpt-image-2.5-flare-2026-09-08 and gpt-image-2.5-sunburst-2026-09-08 use the same adapter options as their aliases.
Standard gpt-image-2.5-flare, gpt-image-2.5-sunburst, and gpt-image-2 generation and edit requests use the native Codex image service. Uploaded edit images are forwarded transiently as image data URLs. Accepted size and quality values are forwarded unchanged. The backend accepts these options but has returned quality: low and 1254x1254 even when higher quality and custom dimensions were requested. Acceptance does not establish output adherence; inspect the returned size and quality fields and decoded image dimensions rather than treating request values as measured output. The native Codex service is distinct from the public OpenAI Images API.
Edits with mask use a Full Responses host and explicitly request the image generation tool. Codex Pooler chooses an eligible Full host without changing Pool serving-mode overrides. If no eligible Full host is available under the configured model and serving policy, the request fails before upstream dispatch or accounting reservation with 400 unsupported_parameter and param: "mask". Unmasked GPT Image 2/2.5 edits remain on the native image route.
A mask identifies editable regions through transparent pixels while opaque pixels indicate areas to preserve. Supply a mask with the same dimensions and format as the source image, and describe the intended edit clearly in the prompt. Image editing is generative: inspect the output rather than assuming exact pixel preservation. Codex Pooler forwards the original mask as input_image_mask; it does not invert alpha, append it as an ordinary reference image, or rewrite the prompt.
input_fidelity is accepted only for gpt-image-1 and gpt-image-1.5. Omit it for gpt-image-2, which always processes image inputs at high fidelity, and for gpt-image-1-mini, which does not support the option. Codex Pooler also rejects configurable input_fidelity for both GPT Image 2.5 models; this is an adapter limitation, not a claim about their public OpenAI API capabilities. Unsupported model/option combinations are rejected before dispatch.
The bundled pricing catalog includes both GPT Image 2.5 models. Their separate text/image token prices are not converted into generic token pricing snapshots; catalog presence does not guarantee a calculated image cost.
Image responses contain base64 image data. response_format may be omitted or set to b64_json; url and other values are rejected. The optional user hint accepts a string or null and is discarded before dispatch; it does not establish ownership or routing identity. Unsupported options, including image streaming and output compression, are rejected rather than silently ignored.
Audio transcriptions
Section titled “Audio transcriptions”The Pool’s Allow Audio Transcription setting must be enabled (the default). Disabling it rejects transcription requests with 403 audio_transcription_disabled before the upload is parsed or sent upstream.
Audio support covers speech-to-text transcription. Text-to-speech (POST /v1/audio/speech), audio translation (POST /v1/audio/translations), and
Realtime voice sessions are not supported. Use a provider that supports those
endpoints when a client needs spoken output or live voice conversations.
POST /v1/audio/transcriptions accepts gpt-transcribe as a caller alias. The
gateway uses the fixed canonical backend identity gpt-4o-transcribe. The alias
is not a model-list entry or a model-discovery guarantee.
The transcription adapter accepts file, model, an optional string prompt, decoded keywords and languages arrays, and response_format: "json" or "text". Omitting the format returns JSON. With "text", Pooler returns the backend transcript as an unquoted UTF-8 string with Content-Type: text/plain; charset=utf-8; whitespace and line breaks are preserved, and errors still use the JSON error envelope.
Subtitle formats srt and vtt, verbose_json, and diarized_json are unsupported. The Codex transcription backend returns text without timed segments or speaker annotations, so Pooler does not generate subtitle timestamps or infer them from the audio length. A model or format accepted by the OpenAI SDK is not automatically available through this backend. Scalar language and temperature are also rejected. Use the explicitly supported languages array only when its backend language hints fit your client; it is not an automatic translation of the OpenAI language parameter.
import fs from "node:fs";import OpenAI from "openai";
const client = new OpenAI({ apiKey: process.env.CODEX_POOLER_API_KEY, baseURL: "http://localhost:4000/v1",});
const transcription = await client.audio.transcriptions.create({ file: fs.createReadStream("audio.mp3"), model: "gpt-transcribe", keywords: ["example", "example"], languages: ["en", "it"],});
console.log(transcription.text);import os
from openai import OpenAI
client = OpenAI( api_key=os.environ["CODEX_POOLER_API_KEY"], base_url="http://localhost:4000/v1",)
with open("audio.mp3", "rb") as audio_file: transcription = client.audio.transcriptions.create( file=audio_file, model="gpt-transcribe", keywords=["example", "example"], languages=["en", "it"], )
print(transcription.text)keywords and languages are optional nonempty string lists. Empty lists are
omitted. Nonempty lists retain caller order and duplicate entries. Responses omit
language-detection fields. This documents the accepted request shape only, not
transcription quality, model discovery, or general Audio API coverage.
A successful upstream transcription response must contain a string text field. Pooler validates this JSON before returning the requested JSON or plain-text format. An empty string is valid for silence; missing or invalid text and error objects are reported as upstream failures rather than successful transcriptions.
POST /v1/responses lifts system and developer input-message text into top-level instructions before dispatching to Codex-compatible work. When the provider supplies only a failed, incomplete, or error terminal, the public SSE adapter adds a response.created lifecycle prefix so SDK accumulators can initialize; this prefix does not imply success. Chat Completions streams still return early upstream errors without a synthetic assistant prefix. Non-streaming failures remain OpenAI-shaped JSON errors.
Relayed validation errors name a client input index only when its original position remains known. If coercion removed or expanded items, the public parameter uses input[] rather than an incorrect index. Tool-output image detail is validated for both list outputs and object outputs containing a content list.
Streaming /v1/responses clients must also handle a terminal type: "error" SSE event whose flat fields and nested error object both carry the sanitized server_error code.
Streaming /v1/chat/completions clients must separately handle a terminal data: {"error":{...}} chunk with nested server_error fields and no top-level type; no [DONE] chunk follows that terminal.
For public SSE, an absent, blank, or whitespace-only event: label is treated as absent before the label is compared with the JSON type. A nonblank mismatch is still rejected. Ordinary incomplete POST /v1/responses SSE blocks are capped at 8 MiB so single large provider events, such as reasoning items carrying encrypted content, can finish decoding; structurally recognizable terminal candidates may retain up to 64 MiB so a large terminal split across upstream chunks can finish decoding. Crossing the applicable cap sends one bounded sanitized type: "error" event, relays none of that source block, and drops later source frames.
On narrow GET /v1/responses websocket compatibility, malformed JSON and JSON values that are not objects are ignored without consuming a sequence number or creating a local terminal. Native backend websocket behavior is unchanged.
Responses WebSocket stream IDs
Section titled “Responses WebSocket stream IDs”Narrow GET /v1/responses websocket response.create accepts an optional
stream_id only when it is a 1 through 256 byte string matching
^[A-Za-z0-9_.-]+$. A valid accepted ID is echoed on every attributable Open
Responses server event and is stripped before upstream dispatch. It is transient
socket-turn state and is not stored in request metadata, accounting, logs, or
telemetry.
Requests with the same ID are FIFO. Different valid IDs are accepted and echoed,
but Codex Pooler conservatively serializes public turns per connection; it does
not promise cross-ID concurrency or fairness. previous_response_id controls
conversation lineage independently. REST POST /v1/responses, native backend
WebSockets, Chat, compact, and batches do not accept this field.
Error types on errors Codex Pooler authors
Section titled “Error types on errors Codex Pooler authors”Errors that Codex Pooler itself authors carry an error.type that agrees with
the HTTP status about whether the failure is worth retrying. The same
classification applies to HTTP error bodies on /v1/* and
/backend-api/codex/* and to error frames on GET /v1/responses, so one code
has one type on every surface.
server_error: the failure is on the server side, including failures raised while the server hands a turn between connections. The same request can succeed on a retry.rate_limit_error: the request was throttled with HTTP429. Retry after backing off.usage_limit_reached: HTTP429when every upstream account is exhausted until a known reset (see below). Retry after the reset it names. The samequota_exhaustedcode on a retryable503keepsserver_error.invalid_request_error: the client must change the request before retrying, or the connection was superseded by its own newer replacement.
No released client branches on error.type for retries. openai-node,
openai-python and the Vercel AI SDK retry on the HTTP status (408, 409,
429, 5xx) and on connection errors, and the first two honour an
x-should-retry response header that overrides the status in both directions.
Codex’s HTTP client retries 5xx and transport errors, not 429, with a
default budget of four retries on a 200 ms doubling backoff; its sampling loop
then retries any retryable turn error, including an unexpected status such as
403 or 409 and a stream-level rate limit, up to five more times from its own
error variant, surfacing each as a stream disconnect and falling back from
websocket to HTTPS once that budget is spent. A 400 response or frame and a
429 status end the turn at once, so a Pooler policy denial that no resend can
pass (model_not_allowed, a per-request token cap) answers 400. None of that
reads the type. Keep that behaviour and use the type and error.code for diagnosis,
treating error.code as an open string as described below; on the
GET /v1/responses frame surface there is no HTTP status, and the type is what
a reader sees. The status and the type never disagree: a 5xx response or frame
is never typed invalid_request_error. Errors relayed from the Codex backend
follow the rules in the next section.
Policy denials that Codex Pooler authors (api_key_missing,
api_key_disabled, api_key_policy_malformed, model_not_allowed,
image_generation_disabled) keep their
own code and message on /v1/*, exactly as on /backend-api/codex/*; they are
never rewritten to server_error or upstream request failed, because no
upstream was involved.
When every upstream account is exhausted
Section titled “When every upstream account is exhausted”When every account that can serve the requested model is excluded because its quota is exhausted, and Codex Pooler knows when each of them resets, the request is answered before any upstream work with the same terminal shape the Codex backend uses for an exhausted account, on every route and transport:
- HTTP status
429(and"status": 429on a websocket error frame); error.typeusage_limit_reached,error.codequota_exhaustedand Codex Pooler’s own message, also on/v1/*;resets_at(Unix seconds) andresets_in_secondsfor the earliest reset among those accounts;- a
Retry-Afterheader with the same number of seconds (in the websocket error frame’sheaders), plusx-should-retry: falseon HTTP when the wait is longer than 60 seconds.
The Codex client ends the turn and names the reset time; the OpenAI SDKs do not
retry it. The time is advice rather than a promise: an automatically redeemed
saved reset can bring an account back sooner. When the reset of any of those
accounts is unknown (missing or stale quota evidence), or an account was taken
out by an open circuit rather than by quota, the answer stays a retryable 503
(quota_evidence_unavailable or quota_exhausted) with no reset fields. When
an open circuit took an account out, that 503 (and the no_eligible_backend
503 when open circuits took out every account) carries a Retry-After of the
seconds until the earliest of those circuits lets a request through again,
between 1 and 60, in the websocket error frame’s headers too, and never
x-should-retry: false, because the request can succeed then. The same
applies when a circuit opens between routing and dispatch, and to file
uploads. When
at least one account is still eligible, the request is routed to it as usual,
and a provider usage-limit refusal that arrives before any output moves the
same request to another eligible account. That includes an account Codex
Pooler kept out of the request because it advertises the model with a
different catalog shape: when the last account of the chosen shape refuses
with a usage limit, the request moves to such an account once.
When no other account is left for the request and the Codex backend’s
usage-limit refusal names the account’s reset, that refusal gets the same
answer: 429, usage_limit_reached, quota_exhausted and Codex Pooler’s
message, and resets_at and resets_in_seconds for the soonest of the
backend’s reset and the resets of the other accounts, which must all be
exhausted with a known reset (otherwise the refusal stays retryable),
Retry-After, and x-should-retry: false when the wait is longer than 60
seconds. The backend’s own message and other fields are not relayed. Over
the GET /v1/responses websocket the same refusal is the error event with
"status": 429, those fields and headers.retry-after. A streaming
/v1/responses request that Codex Pooler carries over the backend’s
websocket gets the same HTTP answer, or moves to another eligible account,
when the refusal arrives before any output. A usage-limit
refusal that names no reset still ahead, or one answered while another
account’s return is not known, keeps the redacted 429 rate_limit_error
described below in both serving modes, with Retry-After (and, over the
websocket, an error event with headers.retry-after) when an account that
an open circuit took out bounds the wait.
Relayed terminal errors
Section titled “Relayed terminal errors”When the Codex backend returns a genuine terminal error, Codex Pooler preserves its trimmed error.code only when
the value is at most 80 bytes and matches ^[A-Za-z0-9_.-]+$. Other code values are redacted to upstream_error.
A surviving code may be a Codex-backend value that is not part of the OpenAI platform ResponseError enum.
Upstream error messages are replaced with upstream request failed, and upstream error types are replaced with
server_error, except that an upstream refusal answered with a 4xx status other than 401, 403, 404 and
429 (for example 400 or 422) is typed invalid_request_error, the class its status names, and an upstream
429 is typed rate_limit_error, on HTTP and in the response.failed event a /v1/responses websocket receives
for it. SDKs retry a 429 on the status alone, so the type changes what a reader sees, not whether the request is
retried.
Clients must therefore treat error.code as an open string: handle the values they understand and
keep an unknown-code fallback rather than validating against the platform enum.
Public relays construct each response.failed envelope from a named-field projection. Unknown event, response,
error, and usage siblings are excluded. The response identifier is retained only when it is a valid resp_
identifier; otherwise it becomes resp_failed. Usage is either null or a bounded projection of the declared token
counters. Content-bearing response fields are empty or null: output and tools are empty arrays, output_text
is an empty string, and instructions, metadata, temperature, and top_p are null.
Top-level and nested errors are sanitized independently and are never copied between locations. A safe code remains
unchanged at its original location, while every invalid code becomes upstream_error; message and type remain the
fixed upstream request failed and server_error values.
One narrow exception applies to misalignment_policy_violation. An eligible
direct 400 or 403, or an exact terminal SSE or websocket failure, is
health-neutral and non-retryable. Public clients receive the exact code,
invalid_request_error, and the nonblank provider message when available, or a
fixed safe fallback when it is blank. The error never includes a provider
param, body, or sibling fields. Durable accounting and logs retain only the
exact code, fixed accounting text, and bounded facts. Every other provider error
continues to use the generic redaction above.
A second narrow exception covers upstream parameter validation. When upstream
rejects a POST /v1/responses or POST /v1/chat/completions request before
any stream starts with HTTP 400, type invalid_request_error, and one of the
codes unsupported_value, invalid_value, unsupported_parameter,
missing_required_parameter, invalid_type, or string_above_max_length,
streaming and non-streaming clients receive a 400 JSON error with type
invalid_request_error, that code, a bounded param field path or null, and
a message written by Codex Pooler from the code and param, for example
upstream rejected parameter reasoning.effort (unsupported_value); supported values: low, medium, high.
SDKs expose these as BadRequestError with code and param. The provider
message is never forwarded; for unsupported_value and invalid_value the
message may list up to 12 supported values when upstream reports them as simple
identifiers, never the rejected value. Chat Completions requests receive the
Chat field they sent where Codex Pooler renamed it, for example
reasoning_effort, max_tokens, verbosity, or response_format, and the
message names that same Chat field; other paths name the upstream Responses
field. Other statuses, types, and codes keep
the generic redaction.
Schema-strict clients that enforce the platform enum cannot accept every relayed Codex-backend code. In
openai-python, leave _strict_response_validation at its default false when using this compatibility surface.
Ordinary OpenAI Responses incomplete terminals are preserved. If an upstream returns response.incomplete with status: "incomplete" for output limits or content filtering and no embedded error, /v1/responses streaming and websocket clients receive that incomplete terminal instead of a synthetic failure. Error-coded incomplete terminals, such as context overflow or stale continuation anchors, are still returned through the sanitized failure path.
Accepted POST /v1/responses tool definitions are narrow. OpenAI Responses remote MCP tool definitions are rejected before dispatch.
Declaration-backed allowed tools
Section titled “Declaration-backed allowed tools”In Full serving mode, direct POST /v1/responses and narrow Responses websocket
response.create accept an allowed_tools choice only when every member is
backed by a direct top-level declaration. Named function and custom members
must match the declared kind and name. The only type-only built-ins are
programmatic_tool_calling, web_search_preview, web_search, and
image_generation, each declared at the top level. Caller order and repeated
members are preserved.
This is a narrow declaration-backed contract, not broad OpenAI tool parity. It
doesn’t cover Chat, backend Responses, namespaces, additional tools, deferred
tools, aliases, Realtime, or remote MCP. Invalid or undeclared Full choices fail
before dispatch. Lite rejects any map-shaped allowed_tools choice with the
existing unsupported_parameter error on tool_choice.
Remote MCP remains unsupported in Responses. A top-level MCP declaration is a
tools error, while an MCP member inside allowed_tools is a tool_choice
error. The separate /mcp endpoint is operator metadata access, not a Responses
remote MCP bridge.
Web search domain filters
Section titled “Web search domain filters”web_search.filters accepts only allowed_domains and blocked_domains. Each
supplied field is a list of 1 through 100 nonblank strings without leading-
whitespace, case-insensitive HTTP(S) schemes. Both fields can appear together.
Codex Pooler forwards accepted values unchanged, including their order, case,
duplicates, and bytes.
external_web_access is optional. Codex Pooler guarantees local validation and
forwarding of this accepted shape only. Web search availability and any
allow-list or block-list enforcement can vary by selected upstream model and
account.
Tool-output preservation
Section titled “Tool-output preservation”Accepted tool-output text on POST /v1/responses, POST /v1/chat/completions and the supported Responses websocket route reaches the upstream without gateway minification or summarization. Protocol adapters may change the surrounding request or split text into ordered parts; the output text itself is preserved. Full/Lite tool handling and strict-schema normalization remain in effect.
No Pool compression switch or savings display remains. Raw outputs are not stored. Public /v1/responses/compact remains unsupported, while provider and client compaction keep their separate contracts.
Routed but unsupported boundaries
Section titled “Routed but unsupported boundaries”These /v1 routes are unsupported and may return deterministic OpenAI-shaped unsupported endpoint errors when explicitly routed:
POST /v1/responses/compactGET /v1/files/:file_id/content, after ownership checksDELETE /v1/files/:file_id, after ownership checksPOST /v1/images/variationsPOST /v1/content_provenance_checks, deliberately routed to deterministic OpenAI-shapedunsupported_endpointPOST /v1/embeddingsPOST /v1/batchesPOST /v1/moderationsPOST /v1/fine_tuning/jobsGET /v1/responses/:response_idPOST /v1/responses/:response_id/cancelDELETE /v1/responses/:response_id/v1/realtimeand OpenAI Realtime SDK websocket or session routes
GET /v1/responses is narrow Responses websocket compatibility, not /v1/realtime support. OpenAI Realtime SDK websocket and session routes are not supported.
Within POST /v1/responses, OpenAI Responses remote MCP tool definitions are unsupported. A top-level tools entry with type: "mcp", or an input item with type: "additional_tools" whose tools list contains type: "mcp", is rejected before upstream dispatch with an OpenAI-shaped invalid_request error.
MCP stays separate
Section titled “MCP stays separate”The operator MCP endpoint is rooted at /mcp, not under /v1. It uses operator-owned MCP bearer tokens and returns metadata only.
The root /mcp endpoint is not a /v1/responses remote MCP bridge and is not invoked by Responses tools[type=mcp] definitions.
Request-log and audit-log tools accept offsets from 0 through 10,000. Refine the filters to inspect more distant history. Their totalExact field distinguishes an exact total from a bounded lower estimate; nextOffset becomes null at the traversal cap. Audit outcomes are success or failure. Clients that cache tool schemas should refresh the tool list after upgrading.
MCP URL:https://codex-pooler.example.com/mcp
Authorization:Bearer <operator-mcp-token>Don’t use Pool API keys, browser sessions, cookies, query tokens, invite tokens, upstream tokens, or custom headers as MCP authentication.
Frequently Asked Questions
Section titled “Frequently Asked Questions”Can I use the official OpenAI SDK with Codex Pooler?
Section titled “Can I use the official OpenAI SDK with Codex Pooler?”Yes, for selected SDK routes. Set the SDK base_url or baseURL to https://codex-pooler.example.com/v1, use a Pool API key as the bearer credential, and keep requests on supported route shapes such as responses, chat completions, models, usage, files, audio transcription, and image generation or edits.
Should Codex backend clients use /v1?
Section titled “Should Codex backend clients use /v1?”No. Codex backend-compatible clients should use /backend-api/codex. The /v1 surface is for selected OpenAI SDK-compatible clients and translates supported requests into Codex-compatible work.
Does /v1 support OpenAI Realtime?
Section titled “Does /v1 support OpenAI Realtime?”No. /v1/realtime and OpenAI Realtime SDK websocket or session routes are unsupported. GET /v1/responses is narrow Responses websocket compatibility only, not OpenAI Realtime support.
Does /v1/responses/compact work?
Section titled “Does /v1/responses/compact work?”No. POST /v1/responses/compact returns a deterministic OpenAI-shaped unsupported_endpoint error. Explicit compaction is supported on normal /v1/responses requests when the client appends exactly one final compaction_trigger after visible input. Backend compact compatibility remains available under /backend-api/codex.
Can I use the same key for /v1 and /mcp?
Section titled “Can I use the same key for /v1 and /mcp?”No. /v1 uses Pool API keys for runtime work. /mcp uses operator-owned MCP bearer tokens for metadata-only lookup. Keep those credentials separate in client configuration and storage.