# OpenAI-Compatible SDKs

Use Codex Pooler in your own applications with the OpenAI Python and Node SDKs or the Vercel AI SDK. This guide covers the supported ways to generate text, use tools, work with images and transcribe audio with a Pool API key, along with examples you can adapt to your application.

![Codex Pooler OpenAI-compatible SDK integrations](/codex-pooler-sdks.webp)

For a short citable summary, see [OpenAI-compatible Codex gateway](/discovery/openai-compatible-codex-gateway/).

## What OpenAI-Compatible Means Here

OpenAI-compatible means selected SDK request shapes can use a `/v1` base URL and a Pool API key. Codex Pooler still routes the request through Codex account Pools, not through a separate OpenAI engine. Unsupported OpenAI API areas remain unsupported, including embeddings, batches, fine-tuning, moderation, response retrieve/cancel/delete, image variations, OpenAI Responses remote MCP tool definitions, and OpenAI Realtime SDK routes.

## Base URL and authentication

Use the `/v1` base URL and a Pool API key:

```text
Base URL:
https://codex-pooler.example.com/v1

Authorization:
Bearer <pool-api-key>
```

For local setup, use `http://localhost:4000/v1`.

Set the key in the terminal used to run the SDK examples:

**macOS / Linux / WSL**

```bash
export CODEX_POOLER_API_KEY="<pool-api-key>"
```

**Windows PowerShell**

```powershell
$env:CODEX_POOLER_API_KEY = "<pool-api-key>"
```

These commands apply to the current terminal. If the application runs as a service, configure the same variable in that service's environment. The Python, JavaScript and TypeScript examples below use the same code on all supported operating systems.

## Service tier compatibility

Use `service_tier: "priority"` for priority processing. `fast` is an accepted
equivalent request spelling, but `priority` is the canonical spelling for new
configuration. On `/v1`, Codex Pooler translates supported OpenAI request and
response shapes while preserving any projected provider `service_tier` value in
its literal provider vocabulary. Codex backend relay routes under
`/backend-api/codex` preserve provider bytes, frames, and service-tier
vocabulary unchanged. The ChatGPT Codex backend reports
`service_tier: "default"` on completed responses even when the request asked
for `priority`, and Codex Pooler relays that value as reported.

`ultrafast` is a separate tier, not an alias for `fast` or `priority`. Use
`service_tier: "ultrafast"` only with direct `/v1/responses` requests when the
selected model metadata advertises it:

```python
response = client.responses.create(
    model="your-model-id",
    input="Your request input.",
    service_tier="ultrafast",
)
```

JSON, SSE, and narrow Responses WebSocket requests preserve a returned
`ultrafast` tier literally. `POST /v1/chat/completions` rejects `ultrafast`.
Provider availability, access, and price control whether the tier can be used;
model metadata advertisement is not an entitlement or price promise.

## Tool-specific setup pages

Use the dedicated setup pages when configuring an agent or editor that has its own provider shape:

- [Aider](/clients/aider/)
- [Cline](/clients/cline/)
- [Continue](/clients/continue/)
- [DeepSeek Harness](/clients/deepseek-harness/)
- [Goose](/clients/goose/)
- [Hermes Agent](/clients/hermes/)
- [Kilo Code](/clients/kilo-code/)
- [OpenCode](/clients/opencode/)
- [OpenClaw](/clients/openclaw/)
- [OpenHands Agent Canvas](/clients/openhands/)
- [OMP](/clients/omp/)
- [Pi](/clients/pi/)
- [Trae](/clients/trae/)
- [Windmill AI](/clients/windmill/)

## Python SDK

```python
import os

from openai import OpenAI

client = OpenAI(
    api_key=os.environ["CODEX_POOLER_API_KEY"],
    base_url="https://codex-pooler.example.com/v1",
)

response = client.responses.create(
    model="gpt-6-sol",
    input="Write a short setup confirmation.",
)

print(response.output_text)
```

## Node SDK

```js
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.CODEX_POOLER_API_KEY,
  baseURL: "https://codex-pooler.example.com/v1",
});

const response = await client.responses.create({
  model: "gpt-6-sol",
  input: "Write a short setup confirmation.",
});

console.log(response.output_text);
```

## Vercel AI SDK

```ts
import { createOpenAI } from "@ai-sdk/openai";
import { generateText } from "ai";

const pooler = createOpenAI({
  apiKey: process.env.CODEX_POOLER_API_KEY,
  baseURL: "https://codex-pooler.example.com/v1",
});

const { text } = await generateText({
  model: pooler.responses("gpt-6-sol"),
  providerOptions: {
    openai: {
      promptCacheOptions: { mode: "explicit", ttl: "30m" },
    },
  },
  messages: [
    {
      role: "system",
      content: "Keep setup guidance concise.",
      providerOptions: {
        openai: { promptCacheBreakpoint: { mode: "explicit" } },
      },
    },
    { role: "user", content: "Write a short setup confirmation." },
  ],
});

console.log(text);
```

Codex Pooler accepts these OpenAI cache controls as public input. For
account-backed requests, it adapts the explicit controls automatically before
upstream dispatch while preserving `prompt_cache_key` independently from Pool
routing affinity. GPT-6 model ids are passed through to the Pool catalog and
assignment policy; naming one here does not guarantee that every Pool has an
eligible account for that model.

### Programmatic tools and stateless replay

For Vercel AI SDK programmatic tools, set `providerOptions.openai.store: false`
on every generation that participates in the tool loop:

```ts
const result = await generateText({
  model: pooler.responses("gpt-6-sol"),
  providerOptions: {
    openai: { store: false },
  },
  // Your programmatic tool and tool-choice configuration goes here.
  prompt: appPrompt,
});
```

Vercel chooses complete stateless replay or a stored `item_reference` before
the request reaches Codex Pooler. `store: false` is therefore required for the
complete stateless replay path. Independently, Codex Pooler forces upstream
`stream: true` and `store: false`.

This is narrow, closed-world compatibility, not general OpenAI programmatic-tool
or Responses parity. Supported shapes are `program` and `program_output`,
caller-enabled `function_call` and `function_call_output`, the type-only
`programmatic_tool_calling` tool and tool choice, plus `allowed_callers` and
map-only `output_schema`. Reference-only and ordinary stored-response
continuations are not supported. Codex Pooler does not locally persist raw
program code, results, schemas, or identifiers. A request accepted by the
gateway can still be rejected when its selected upstream model or account lacks
the required capability.

Responses `function_call_output` replay has two closed forms over HTTP and the
narrow Responses websocket bridge. A paired item requires a nonblank `call_id`;
its existing `output` form and paired-only legacy `result` form are unchanged. A
named standalone item requires a nonblank `name` and an `output` field, permits
`call_id` only when omitted or `null`, and permits `namespace` only when omitted,
`null`, or a nonblank string. Blank or non-string values and standalone `result`
reject before dispatch. Classification and debug summaries remain metadata-only.
This does not expand general Responses parity or prove provider-live acceptance.

### Hosted shell history replay

OpenAI Responses history may include completed `shell_call` and
`shell_call_output` items from an earlier response. Codex Pooler accepts this
closed-key replay subset, including direct or program callers, local or container
environment metadata, `in_progress`, `completed`, or `incomplete` status, and
empty command or output arrays where the shape permits them. It forwards accepted
history for stateless replay and semantic tool-output continuations without
requiring call/output pairing or a particular item order.

This is history forwarding only. Codex Pooler does not execute shell commands or
accept a top-level shell tool declaration. `local_shell_call` history and remote
MCP tool definitions remain unsupported. It also does not implement SDK-local
command-index accumulation or claim general hosted-tool parity.

Responses SSE and the narrow Responses websocket preserve the five hosted-shell
relay event types: command added, command delta, command done, output-content
delta, and output-content done. Existing public sequence normalization and the
websocket `stream_id` behavior are the only permitted relay adaptations.

Commands and shell output remain transient request data. They are not persisted
or rendered in request metadata, logs, or operator views.

For a `previous_response_id` tool continuation, a native `reasoning` replay item
may include `content` with `reasoning_text` parts. Codex Pooler preserves that
accepted reasoning content for the upstream continuation. Stateless replay drops
reasoning items before dispatch instead.

Assistant-history replay may include `output_text` URL citations. Codex Pooler
accepts only ordered `url_citation` annotations with exactly `type`,
`start_index`, `end_index`, `url`, and `title`; it preserves accepted values,
explicit empty annotation lists, and omission exactly. Malformed or unsupported
annotation shapes are rejected before dispatch rather than filtered.

`GET /v1/models` may include `context_length` for clients that probe OpenAI-compatible model lists, such as Hermes. That probe is an ordinary authenticated request: a client that issues it without the Pool API key is refused like any other route, and a client that caches the refusal will keep using whatever context window it bundles for the model instead. Configure the context window explicitly wherever the client allows it. The official OpenAI SDK request APIs and Vercel AI SDK generation APIs do not expose Codex model-catalog context controls. Their output-budget fields are `max_output_tokens` in OpenAI Responses, `max_completion_tokens` in Chat Completions, and `maxOutputTokens` at the Vercel AI SDK layer; availability at the SDK layer does not establish upstream support or a hard generation limit. Codex Pooler's public `/v1/responses` rejects `context_management`, and direct `POST /v1/responses/compact` remains unsupported. Explicit compaction is available on the normal `/v1/responses` route.

### Continuing with `previous_response_id`

A nonblank `previous_response_id` is accepted only for a tool-output continuation. If the input has no tool-result shape, Codex Pooler returns `400 invalid_request` with `param: previous_response_id` before checking the upstream connection. This validation rule applies to HTTP and the Responses websocket.

Codex Pooler forces `store: false` upstream, and the upstream provider resolves `previous_response_id` only on the websocket connection that produced the response. Over HTTP it rejects the parameter, and on any other websocket connection, a new one included, it rejects the anchor. On `/v1/responses` an anchored request therefore reaches the earlier context in two cases: on the Responses websocket route, or as a `stream: true` HTTP request on a deployment with websocket owner forwarding (see the [websocket replica caveat](/deployment/helm/#websocket-replica-caveat)) when every request of the chain, the first one included, sends the same session header, such as `session-id`. Codex Pooler then sends those requests over its session's upstream websocket.

After input validation, an unavailable producing connection returns `400 previous_response_not_found`. Codex Pooler rejects locally detectable failures before dispatch and maps the provider’s matching stale-anchor refusal to the same error:

```json
{
  "error": {
    "type": "invalid_request_error",
    "code": "previous_response_not_found",
    "param": "previous_response_id",
    "message": "Previous response not found on this request's upstream connection: ..."
  }
}
```

This covers a non-streaming request, every HTTP request on a deployment without owner forwarding, and a streaming request whose session has no upstream connection that produced the anchor, such as a chain without a session header or a session whose connection was closed. The message states the requirement above. When you receive this error, send the complete input again without `previous_response_id`; clients that recognize the `previous_response_not_found` code as a stale chain do this on their own.

### Output budgets and local usage

Codex Pooler checks token policy against admission estimates. The output reservation floor is 512 tokens, or 2,048 when context is opaque; higher accepted request or policy values can increase the estimate. Continuation references and other opaque input prevent a complete input count from the visible payload. These are local accounting estimates in both Full and Lite serving modes, not provider hard bounds. The selected upstream may refuse `max_output_tokens`; forwarding a field or switching serving modes does not guarantee output enforcement. Actual measured usage is recorded without clipping, so a single response can exceed its reservation or remaining token budget. There is no absolute per-response spend cap.

`GET /v1/usage` separates measured reporting from local enforcement. Existing `total_tokens` and priced cost totals remain measured-usage reporting; a price snapshot produces an estimate, not an invoice or a provider-confirmed charge. `budget_usage.daily` and `budget_usage.weekly` each expose `known_total_tokens`, `provisional_total_tokens`, `pending_total_tokens`, `effective_total_tokens`, and `admission_count`. Daily means midnight UTC onward; weekly means the trailing seven days. These components cover only the authenticated key across models and routes, never other keys.

Effective pressure is known plus provisional plus pending. Attempted work with unknown terminal usage keeps its original estimate provisionally until a known correction or expiry of the original terminal window. A correction replaces that estimate once at the original terminal time. Pre-attempt failures and requests the Pooler refuses before anything is sent upstream create no provisional consumption; pending reservations remain pressure across a window boundary. RPM counts each admitted reservation once, without refunds or movement on correction.

The local `limits` list describes active default-policy thresholds; it is not a model-override simulator. A matching active model override replaces the default binding for admission, including its blank fields, while comparing its thresholds with whole-key pressure. Provider quota windows remain separate in `upstream_limits`; Codex-compatible usage endpoints retain their upstream quota authority when available. See [API key limits](/operators/api-keys/#limits) for policy and concurrency configuration.

An optional key-wide `max_active_requests` limit spans the fleet and all models and routes. A trusted local saturation denial is `429 api_key_concurrency_limit_exceeded`, typed `rate_limit_error`, with HTTP `Retry-After: 1`. Back off before retrying; the interval does not promise availability. Native and narrow public Responses websockets receive the equivalent turn error and can retry on the same connection when a slot becomes available. A provider error that happens to use the same code is not trusted as a local concurrency denial and does not acquire its retry header or type. Authentication `401` and policy-budget `403` remain distinct.

### Request an explicit compaction turn

Provider clients that expose an OpenAI compaction option can use the normal `/v1/responses` route. Vercel AI SDK serializes the required trigger with `providerOptions.openai.compactionTrigger: true` in `@ai-sdk/openai` 4.0.42 or later:

```ts
import {
  createOpenAI,
  type OpenAILanguageModelResponsesOptions,
} from "@ai-sdk/openai";
import { generateText } from "ai";

const pooler = createOpenAI({
  apiKey: process.env.CODEX_POOLER_API_KEY,
  baseURL: "https://codex-pooler.example.com/v1",
});

const result = await generateText({
  model: pooler.responses("gpt-6-sol"),
  providerOptions: {
    openai: {
      store: false,
      compactionTrigger: true,
    } satisfies OpenAILanguageModelResponsesOptions,
  },
  prompt: "Compact the visible conversation context.",
});

const compaction = result.content.find(
  (part) => part.type === "custom" && part.kind === "openai.compaction",
);
```

The serialized request must contain visible input followed by exactly one final `{"type":"compaction_trigger"}` item. Non-terminal, duplicate, trigger-only, hidden-only, or otherwise malformed placement returns an OpenAI-shaped `400 invalid_request` on `input` before upstream dispatch.

Successful non-streaming HTTP returns a completed Responses JSON object containing the normalized compaction item. Public SSE follows the Responses streaming grammar the official SDK stream helpers expect: `response.created` with an empty output, `response.output_item.added` and `response.output_item.done` for the compaction item at `output_index` 0, then `response.completed` and `[DONE]`. Narrow Responses websocket completion emits the same four Responses events without the HTTP `[DONE]` sentinel. Direct `POST /v1/responses/compact` remains unsupported.

If upstream compact output is malformed JSON or does not contain nonblank encrypted compaction content, Codex Pooler returns a sanitized `502 invalid_compaction_response`. Other provider failures follow the public error rules documented below.

### Replay a remote-compaction item

When a compaction turn returns an encrypted `type: "compaction"` output item, submit that item unchanged at the start of the next `POST /v1/responses` input, followed by the new user input. Start a new chain by omitting `previous_response_id`; the compact response envelope id is not a replay anchor.

The stable public replay fields are `type`, a nonblank opaque `encrypted_content` string, and a string `id`. When the upstream item carries no usable id, Codex Pooler returns a derived `cmp_` id and removes that derived id again when the item is replayed, so the upstream receives the item as it produced it; a provider id is forwarded unchanged. JSON, SSE, and narrow Responses websocket surfaces return the same normalized item. Native-only metadata and unknown fields are removed from output; unknown replay fields or malformed values are rejected before upstream dispatch. Treat `encrypted_content` as opaque and do not log or persist it in application telemetry.

## Supported or translated route support

The OpenAI-compatible `/v1` surface supports or translates selected routes only:

- `GET /v1/models`
- `POST /v1/responses`
- `GET /v1/responses`, narrow Responses websocket compatibility only
- `POST /v1/chat/completions`
- `GET /v1/usage`
- `GET /v1/files`
- `POST /v1/files`
- `GET /v1/files/:file_id`
- `POST /v1/audio/transcriptions`
- `POST /v1/images/generations`
- `POST /v1/images/edits`

The `/v1` surface is compatibility over Codex routing, not a separate OpenAI engine. Supported requests still require a Pool API key and a Pool with eligible upstream capacity for the requested model.

### Chat Completions request compatibility

`POST /v1/chat/completions` accepts normal `messages` requests. It also accepts
Responses-shaped `input` with `reasoning`, `text`, and `include` when `messages`
is absent or empty. Combining these fallback fields with nonempty `messages`
returns `invalid_request`, except that a string `reasoning` value is accepted
as an alias for Chat's `reasoning_effort`. The alias follows the same validation
and API-key reasoning policy. If both effort fields are present, their normalized
values must match; conflicting values return `invalid_request`. A Responses
`reasoning` object remains invalid alongside nonempty `messages`.

Both shapes return Chat Completions JSON or SSE. `stream_options.include_usage`
controls the terminal streamed usage chunk and is not forwarded upstream.
Successful function or custom-tool turns return `finish_reason: "tool_calls"`,
including streams whose terminal event omits the output list.

Streamed function arguments and custom-tool input can be completed from the
upstream's final argument, item, or completed-response snapshot. Pooler appends
only the missing suffix when the tool identity and already-delivered bytes
agree. Conflicting snapshots end the stream with a sanitized error rather than
reporting a successful tool call.

Chat function tools keep Chat's non-strict default: omitted or null `strict`
is forwarded as `strict: false`. Explicit `strict: true` remains strict.
Direct Responses requests and Responses-shaped fallback requests retain their
own strict-mode defaults.

For ordinary Chat `messages`, text fields larger than 10 MiB are divided into
smaller UTF-8 content parts before dispatch, preserving all text and its order.
This includes message history and textual tool outputs. Oversized combined
system/developer instructions keep their normalized content in one leading
developer message with multiple parts. Nothing is truncated or summarized.
Request-size and model context limits still apply; opaque fields such as tool
arguments and encoded media are not split. Direct Responses and `input`
fallback requests retain their existing representation.

Replayed Chat tool-call identifiers longer than 64 UTF-8 bytes are mapped to
stable, bounded identifiers before dispatch. The matching call and result use
the same identifier, including across repeated history requests. Short
identifiers, tool names, arguments and outputs stay unchanged. If a mapped
identifier would collide with another identifier in the request, Pooler rejects
the request instead of merging distinct calls. Direct Responses and `input`
fallback requests keep their original identifiers.

The optional `user` field accepts a string or null and is discarded before
dispatch. It does not establish ownership, permissions, or routing identity.

A `tool` message may carry `image_url` parts next to its text, as Hermes sends a
screenshot tool result when it talks Chat Completions. Codex Pooler forwards them
as images of the tool output, with `image_url.detail` handled like a user image.
A `file` part in a `tool` message returns `invalid_request`.

### Image generation and edits

Use `gpt-image-2.5-flare` for current image examples, or select `gpt-image-2.5-sunburst` explicitly. Both [Flare](https://developers.openai.com/api/docs/models/gpt-image-2.5-flare) and [Sunburst](https://developers.openai.com/api/docs/models/gpt-image-2.5-sunburst) are accepted; `gpt-image-2`, `gpt-image-1.5`, `gpt-image-1`, and `gpt-image-1-mini` remain available for compatibility. Media identifiers do not have to appear in the Pool's model catalog for these translated routes to work.

```js
const image = await client.images.generate({
  model: "gpt-image-2.5-flare",
  prompt: "A simple watercolor landscape",
  size: "1024x1024",
  quality: "medium",
});
```

For GPT Image 2.5, the translated Images routes accept `auto` or custom `WIDTHxHEIGHT` sizes: both edges must be multiples of 16, neither may exceed 3840 pixels, the aspect ratio must stay between 1:3 and 3:1, and the total area must be between 655,360 and 8,294,400 pixels. Resolutions above `2560x1440` are experimental upstream. Quality accepts `auto`, `low`, `medium`, `high`, `xhigh`, and `max`. Older models retain the three standard sizes plus `auto`, with quality up to `high`. Requests still produce one image. The dated identifiers `gpt-image-2.5-flare-2026-09-08` and `gpt-image-2.5-sunburst-2026-09-08` use the same adapter options as their aliases.

Standard `gpt-image-2.5-flare`, `gpt-image-2.5-sunburst`, and `gpt-image-2` generation and edit requests use the native Codex image service. Uploaded edit images are forwarded transiently as image data URLs. Accepted `size` and `quality` values are forwarded unchanged. The backend accepts these options but has returned `quality: low` and `1254x1254` even when higher quality and custom dimensions were requested. Acceptance does not establish output adherence; inspect the returned `size` and `quality` fields and decoded image dimensions rather than treating request values as measured output. The native Codex service is distinct from the public OpenAI Images API.

Edits with `mask` use a Full Responses host and explicitly request the image generation tool. Codex Pooler chooses an eligible Full host without changing Pool serving-mode overrides. If no eligible Full host is available under the configured model and serving policy, the request fails before upstream dispatch or accounting reservation with `400 unsupported_parameter` and `param: "mask"`. Unmasked GPT Image 2/2.5 edits remain on the native image route.

A mask identifies editable regions through transparent pixels while opaque pixels indicate areas to preserve. Supply a mask with the same dimensions and format as the source image, and describe the intended edit clearly in the prompt. Image editing is generative: inspect the output rather than assuming exact pixel preservation. Codex Pooler forwards the original mask as `input_image_mask`; it does not invert alpha, append it as an ordinary reference image, or rewrite the prompt.

`input_fidelity` is accepted only for `gpt-image-1` and `gpt-image-1.5`. Omit it for `gpt-image-2`, which [always processes image inputs at high fidelity](https://developers.openai.com/api/docs/guides/image-generation#image-input-fidelity), and for `gpt-image-1-mini`, which does not support the option. Codex Pooler also rejects configurable `input_fidelity` for both GPT Image 2.5 models; this is an adapter limitation, not a claim about their public OpenAI API capabilities. Unsupported model/option combinations are rejected before dispatch.

The bundled pricing catalog includes both GPT Image 2.5 models. Their separate text/image token prices are not converted into generic token pricing snapshots; catalog presence does not guarantee a calculated image cost.

Image responses contain base64 image data. `response_format` may be omitted or set to `b64_json`; `url` and other values are rejected. The optional `user` hint accepts a string or null and is discarded before dispatch; it does not establish ownership or routing identity. Unsupported options, including image streaming and output compression, are rejected rather than silently ignored.

### Audio transcriptions

The Pool's `Allow Audio Transcription` setting must be enabled (the default). Disabling it rejects transcription requests with `403 audio_transcription_disabled` before the upload is parsed or sent upstream.

Audio support covers speech-to-text transcription. Text-to-speech (`POST /v1/audio/speech`), audio translation (`POST /v1/audio/translations`), and
Realtime voice sessions are not supported. Use a provider that supports those
endpoints when a client needs spoken output or live voice conversations.

`POST /v1/audio/transcriptions` accepts `gpt-transcribe` as a caller alias. The
gateway uses the fixed canonical backend identity `gpt-4o-transcribe`. The alias
is not a model-list entry or a model-discovery guarantee.

The transcription adapter accepts `file`, `model`, an optional string `prompt`, decoded `keywords` and `languages` arrays, and `response_format: "json"` or `"text"`. Omitting the format returns JSON. With `"text"`, Pooler returns the backend transcript as an unquoted UTF-8 string with `Content-Type: text/plain; charset=utf-8`; whitespace and line breaks are preserved, and errors still use the JSON error envelope.

Subtitle formats `srt` and `vtt`, `verbose_json`, and `diarized_json` are unsupported. The Codex transcription backend returns text without timed segments or speaker annotations, so Pooler does not generate subtitle timestamps or infer them from the audio length. A model or format accepted by the OpenAI SDK is not automatically available through this backend. Scalar `language` and `temperature` are also rejected. Use the explicitly supported `languages` array only when its backend language hints fit your client; it is not an automatic translation of the OpenAI `language` parameter.

```js
import fs from "node:fs";
import OpenAI from "openai";

const client = new OpenAI({
  apiKey: process.env.CODEX_POOLER_API_KEY,
  baseURL: "http://localhost:4000/v1",
});

const transcription = await client.audio.transcriptions.create({
  file: fs.createReadStream("audio.mp3"),
  model: "gpt-transcribe",
  keywords: ["example", "example"],
  languages: ["en", "it"],
});

console.log(transcription.text);
```

```python
import os

from openai import OpenAI

client = OpenAI(
    api_key=os.environ["CODEX_POOLER_API_KEY"],
    base_url="http://localhost:4000/v1",
)

with open("audio.mp3", "rb") as audio_file:
    transcription = client.audio.transcriptions.create(
        file=audio_file,
        model="gpt-transcribe",
        keywords=["example", "example"],
        languages=["en", "it"],
    )

print(transcription.text)
```

`keywords` and `languages` are optional nonempty string lists. Empty lists are
omitted. Nonempty lists retain caller order and duplicate entries. Responses omit
language-detection fields. This documents the accepted request shape only, not
transcription quality, model discovery, or general Audio API coverage.

A successful upstream transcription response must contain a string `text` field. Pooler validates this JSON before returning the requested JSON or plain-text format. An empty string is valid for silence; missing or invalid text and error objects are reported as upstream failures rather than successful transcriptions.

`POST /v1/responses` lifts system and developer input-message text into top-level `instructions` before dispatching to Codex-compatible work. When the provider supplies only a failed, incomplete, or error terminal, the public SSE adapter adds a `response.created` lifecycle prefix so SDK accumulators can initialize; this prefix does not imply success. Chat Completions streams still return early upstream errors without a synthetic assistant prefix. Non-streaming failures remain OpenAI-shaped JSON errors.

Relayed validation errors name a client input index only when its original position remains known. If coercion removed or expanded items, the public parameter uses `input[]` rather than an incorrect index. Tool-output image `detail` is validated for both list outputs and object outputs containing a `content` list.

Streaming `/v1/responses` clients must also handle a terminal `type: "error"` SSE event whose flat fields and nested `error` object both carry the sanitized `server_error` code.

Streaming `/v1/chat/completions` clients must separately handle a terminal `data: {"error":{...}}` chunk with nested `server_error` fields and no top-level `type`; no `[DONE]` chunk follows that terminal.

For public SSE, an absent, blank, or whitespace-only `event:` label is treated as absent before the label is compared with the JSON `type`. A nonblank mismatch is still rejected. Ordinary incomplete `POST /v1/responses` SSE blocks are capped at 8 MiB so single large provider events, such as reasoning items carrying encrypted content, can finish decoding; structurally recognizable terminal candidates may retain up to 64 MiB so a large terminal split across upstream chunks can finish decoding. Crossing the applicable cap sends one bounded sanitized `type: "error"` event, relays none of that source block, and drops later source frames.

On narrow `GET /v1/responses` websocket compatibility, malformed JSON and JSON values that are not objects are ignored without consuming a sequence number or creating a local terminal. Native backend websocket behavior is unchanged.

### Responses WebSocket stream IDs

Narrow `GET /v1/responses` websocket `response.create` accepts an optional
`stream_id` only when it is a 1 through 256 byte string matching
`^[A-Za-z0-9_.-]+$`. A valid accepted ID is echoed on every attributable Open
Responses server event and is stripped before upstream dispatch. It is transient
socket-turn state and is not stored in request metadata, accounting, logs, or
telemetry.

Requests with the same ID are FIFO. Different valid IDs are accepted and echoed,
but Codex Pooler conservatively serializes public turns per connection; it does
not promise cross-ID concurrency or fairness. `previous_response_id` controls
conversation lineage independently. REST `POST /v1/responses`, native backend
WebSockets, Chat, compact, and batches do not accept this field.

### Error types on errors Codex Pooler authors

Errors that Codex Pooler itself authors carry an `error.type` that agrees with
the HTTP status about whether the failure is worth retrying. The same
classification applies to HTTP error bodies on `/v1/*` and
`/backend-api/codex/*` and to error frames on `GET /v1/responses`, so one code
has one type on every surface.

- `server_error`: the failure is on the server side, including failures raised
  while the server hands a turn between connections. The same request can
  succeed on a retry.
- `rate_limit_error`: the request was throttled with HTTP `429`. Retry after
  backing off.
- `usage_limit_reached`: HTTP `429` when every upstream account is exhausted
  until a known reset (see below). Retry after the reset it names. The same
  `quota_exhausted` code on a retryable `503` keeps `server_error`.
- `invalid_request_error`: the client must change the request before retrying,
  or the connection was superseded by its own newer replacement.

No released client branches on `error.type` for retries. openai-node,
openai-python and the Vercel AI SDK retry on the HTTP status (`408`, `409`,
`429`, `5xx`) and on connection errors, and the first two honour an
`x-should-retry` response header that overrides the status in both directions.
Codex's HTTP client retries `5xx` and transport errors, not `429`, with a
default budget of four retries on a 200 ms doubling backoff; its sampling loop
then retries any retryable turn error, including an unexpected status such as
`403` or `409` and a stream-level rate limit, up to five more times from its own
error variant, surfacing each as a stream disconnect and falling back from
websocket to HTTPS once that budget is spent. A `400` response or frame and a
`429` status end the turn at once, so a Pooler policy denial that no resend can
pass (`model_not_allowed`, a per-request token cap) answers `400`. None of that
reads the type. Keep that behaviour and use the type and `error.code` for diagnosis,
treating `error.code` as an open string as described below; on the
`GET /v1/responses` frame surface there is no HTTP status, and the type is what
a reader sees. The status and the type never disagree: a `5xx` response or frame
is never typed `invalid_request_error`. Errors relayed from the Codex backend
follow the rules in the next section.

Policy denials that Codex Pooler authors (`api_key_missing`,
`api_key_disabled`, `api_key_policy_malformed`, `model_not_allowed`,
`image_generation_disabled`) keep their
own code and message on `/v1/*`, exactly as on `/backend-api/codex/*`; they are
never rewritten to `server_error` or `upstream request failed`, because no
upstream was involved.

### When every upstream account is exhausted

When every account that can serve the requested model is excluded because its
quota is exhausted, and Codex Pooler knows when each of them resets, the
request is answered before any upstream work with the same terminal shape the
Codex backend uses for an exhausted account, on every route and transport:

- HTTP status `429` (and `"status": 429` on a websocket error frame);
- `error.type` `usage_limit_reached`, `error.code` `quota_exhausted` and Codex
  Pooler's own message, also on `/v1/*`;
- `resets_at` (Unix seconds) and `resets_in_seconds` for the earliest reset
  among those accounts;
- a `Retry-After` header with the same number of seconds (in the websocket
  error frame's `headers`), plus `x-should-retry: false` on HTTP when the wait
  is longer than 60 seconds.

The Codex client ends the turn and names the reset time; the OpenAI SDKs do not
retry it. The time is advice rather than a promise: an automatically redeemed
saved reset can bring an account back sooner. When the reset of any of those
accounts is unknown (missing or stale quota evidence), or an account was taken
out by an open circuit rather than by quota, the answer stays a retryable `503`
(`quota_evidence_unavailable` or `quota_exhausted`) with no reset fields. When
an open circuit took an account out, that `503` (and the `no_eligible_backend`
`503` when open circuits took out every account) carries a `Retry-After` of the
seconds until the earliest of those circuits lets a request through again,
between 1 and 60, in the websocket error frame's `headers` too, and never
`x-should-retry: false`, because the request can succeed then. The same
applies when a circuit opens between routing and dispatch, and to file
uploads. When
at least one account is still eligible, the request is routed to it as usual,
and a provider usage-limit refusal that arrives before any output moves the
same request to another eligible account. That includes an account Codex
Pooler kept out of the request because it advertises the model with a
different catalog shape: when the last account of the chosen shape refuses
with a usage limit, the request moves to such an account once.

When no other account is left for the request and the Codex backend's
usage-limit refusal names the account's reset, that refusal gets the same
answer: `429`, `usage_limit_reached`, `quota_exhausted` and Codex Pooler's
message, and `resets_at` and `resets_in_seconds` for the soonest of the
backend's reset and the resets of the other accounts, which must all be
exhausted with a known reset (otherwise the refusal stays retryable),
`Retry-After`, and `x-should-retry: false` when the wait is longer than 60
seconds. The backend's own message and other fields are not relayed. Over
the `GET /v1/responses` websocket the same refusal is the `error` event with
`"status": 429`, those fields and `headers.retry-after`. A streaming
`/v1/responses` request that Codex Pooler carries over the backend's
websocket gets the same HTTP answer, or moves to another eligible account,
when the refusal arrives before any output. A usage-limit
refusal that names no reset still ahead, or one answered while another
account's return is not known, keeps the redacted `429` `rate_limit_error`
described below in both serving modes, with `Retry-After` (and, over the
websocket, an `error` event with `headers.retry-after`) when an account that
an open circuit took out bounds the wait.

### Relayed terminal errors

When the Codex backend returns a genuine terminal error, Codex Pooler preserves its trimmed `error.code` only when
the value is at most 80 bytes and matches `^[A-Za-z0-9_.-]+$`. Other code values are redacted to `upstream_error`.
A surviving code may be a Codex-backend value that is not part of the OpenAI platform `ResponseError` enum.

Upstream error messages are replaced with `upstream request failed`, and upstream error types are replaced with
`server_error`, except that an upstream refusal answered with a `4xx` status other than `401`, `403`, `404` and
`429` (for example `400` or `422`) is typed `invalid_request_error`, the class its status names, and an upstream
`429` is typed `rate_limit_error`, on HTTP and in the `response.failed` event a `/v1/responses` websocket receives
for it. SDKs retry a `429` on the status alone, so the type changes what a reader sees, not whether the request is
retried.
Clients must therefore treat `error.code` as an open string: handle the values they understand and
keep an unknown-code fallback rather than validating against the platform enum.

Public relays construct each `response.failed` envelope from a named-field projection. Unknown event, response,
error, and usage siblings are excluded. The response identifier is retained only when it is a valid `resp_`
identifier; otherwise it becomes `resp_failed`. Usage is either `null` or a bounded projection of the declared token
counters. Content-bearing response fields are empty or `null`: `output` and `tools` are empty arrays, `output_text`
is an empty string, and `instructions`, `metadata`, `temperature`, and `top_p` are `null`.

Top-level and nested errors are sanitized independently and are never copied between locations. A safe code remains
unchanged at its original location, while every invalid code becomes `upstream_error`; message and type remain the
fixed `upstream request failed` and `server_error` values.

One narrow exception applies to `misalignment_policy_violation`. An eligible
direct `400` or `403`, or an exact terminal SSE or websocket failure, is
health-neutral and non-retryable. Public clients receive the exact code,
`invalid_request_error`, and the nonblank provider message when available, or a
fixed safe fallback when it is blank. The error never includes a provider
`param`, body, or sibling fields. Durable accounting and logs retain only the
exact code, fixed accounting text, and bounded facts. Every other provider error
continues to use the generic redaction above.

A second narrow exception covers upstream parameter validation. When upstream
rejects a `POST /v1/responses` or `POST /v1/chat/completions` request before
any stream starts with HTTP `400`, type `invalid_request_error`, and one of the
codes `unsupported_value`, `invalid_value`, `unsupported_parameter`,
`missing_required_parameter`, `invalid_type`, or `string_above_max_length`,
streaming and non-streaming clients receive a `400` JSON error with type
`invalid_request_error`, that code, a bounded `param` field path or `null`, and
a message written by Codex Pooler from the code and param, for example
`upstream rejected parameter reasoning.effort (unsupported_value); supported values: low, medium, high`.
SDKs expose these as `BadRequestError` with `code` and `param`. The provider
message is never forwarded; for `unsupported_value` and `invalid_value` the
message may list up to 12 supported values when upstream reports them as simple
identifiers, never the rejected value. Chat Completions requests receive the
Chat field they sent where Codex Pooler renamed it, for example
`reasoning_effort`, `max_tokens`, `verbosity`, or `response_format`, and the
message names that same Chat field; other paths name the upstream Responses
field. Other statuses, types, and codes keep
the generic redaction.

Schema-strict clients that enforce the platform enum cannot accept every relayed Codex-backend code. In
`openai-python`, leave `_strict_response_validation` at its default `false` when using this compatibility surface.

Ordinary OpenAI Responses incomplete terminals are preserved. If an upstream returns `response.incomplete` with `status: "incomplete"` for output limits or content filtering and no embedded error, `/v1/responses` streaming and websocket clients receive that incomplete terminal instead of a synthetic failure. Error-coded incomplete terminals, such as context overflow or stale continuation anchors, are still returned through the sanitized failure path.

Accepted `POST /v1/responses` tool definitions are narrow. OpenAI Responses remote MCP tool definitions are rejected before dispatch.

### Declaration-backed allowed tools

In Full serving mode, direct `POST /v1/responses` and narrow Responses websocket
`response.create` accept an `allowed_tools` choice only when every member is
backed by a direct top-level declaration. Named `function` and `custom` members
must match the declared kind and name. The only type-only built-ins are
`programmatic_tool_calling`, `web_search_preview`, `web_search`, and
`image_generation`, each declared at the top level. Caller order and repeated
members are preserved.

This is a narrow declaration-backed contract, not broad OpenAI tool parity. It
doesn't cover Chat, backend Responses, namespaces, additional tools, deferred
tools, aliases, Realtime, or remote MCP. Invalid or undeclared Full choices fail
before dispatch. Lite rejects any map-shaped `allowed_tools` choice with the
existing `unsupported_parameter` error on `tool_choice`.

Remote MCP remains unsupported in Responses. A top-level MCP declaration is a
`tools` error, while an MCP member inside `allowed_tools` is a `tool_choice`
error. The separate `/mcp` endpoint is operator metadata access, not a Responses
remote MCP bridge.

### Web search domain filters

`web_search.filters` accepts only `allowed_domains` and `blocked_domains`. Each
supplied field is a list of 1 through 100 nonblank strings without leading-
whitespace, case-insensitive HTTP(S) schemes. Both fields can appear together.
Codex Pooler forwards accepted values unchanged, including their order, case,
duplicates, and bytes.

`external_web_access` is optional. Codex Pooler guarantees local validation and
forwarding of this accepted shape only. Web search availability and any
allow-list or block-list enforcement can vary by selected upstream model and
account.

## Tool-output preservation

Accepted tool-output text on `POST /v1/responses`, `POST /v1/chat/completions` and the supported Responses websocket route reaches the upstream without gateway minification or summarization. Protocol adapters may change the surrounding request or split text into ordered parts; the output text itself is preserved. Full/Lite tool handling and strict-schema normalization remain in effect.

No Pool compression switch or savings display remains. Raw outputs are not stored. Public `/v1/responses/compact` remains unsupported, while provider and client compaction keep their separate contracts.

## Routed but unsupported boundaries

These `/v1` routes are unsupported and may return deterministic OpenAI-shaped unsupported endpoint errors when explicitly routed:

- `POST /v1/responses/compact`
- `GET /v1/files/:file_id/content`, after ownership checks
- `DELETE /v1/files/:file_id`, after ownership checks
- `POST /v1/images/variations`
- `POST /v1/content_provenance_checks`, deliberately routed to deterministic OpenAI-shaped `unsupported_endpoint`
- `POST /v1/embeddings`
- `POST /v1/batches`
- `POST /v1/moderations`
- `POST /v1/fine_tuning/jobs`
- `GET /v1/responses/:response_id`
- `POST /v1/responses/:response_id/cancel`
- `DELETE /v1/responses/:response_id`
- `/v1/realtime` and OpenAI Realtime SDK websocket or session routes

`GET /v1/responses` is narrow Responses websocket compatibility, not `/v1/realtime` support. OpenAI Realtime SDK websocket and session routes are not supported.

Within `POST /v1/responses`, OpenAI Responses remote MCP tool definitions are unsupported. A top-level `tools` entry with `type: "mcp"`, or an `input` item with `type: "additional_tools"` whose `tools` list contains `type: "mcp"`, is rejected before upstream dispatch with an OpenAI-shaped `invalid_request` error.

## MCP stays separate

The operator MCP endpoint is rooted at `/mcp`, not under `/v1`. It uses operator-owned MCP bearer tokens and returns metadata only.

The root `/mcp` endpoint is not a `/v1/responses` remote MCP bridge and is not invoked by Responses `tools[type=mcp]` definitions.

Request-log and audit-log tools accept offsets from 0 through 10,000. Refine the filters to inspect more distant history. Their `totalExact` field distinguishes an exact total from a bounded lower estimate; `nextOffset` becomes null at the traversal cap. Audit outcomes are `success` or `failure`. Clients that cache tool schemas should refresh the tool list after upgrading.

```text
MCP URL:
https://codex-pooler.example.com/mcp

Authorization:
Bearer <operator-mcp-token>
```

Don't use Pool API keys, browser sessions, cookies, query tokens, invite tokens, upstream tokens, or custom headers as MCP authentication.

## Frequently Asked Questions

### Can I use the official OpenAI SDK with Codex Pooler?

Yes, for selected SDK routes. Set the SDK `base_url` or `baseURL` to `https://codex-pooler.example.com/v1`, use a Pool API key as the bearer credential, and keep requests on supported route shapes such as responses, chat completions, models, usage, files, audio transcription, and image generation or edits.

### Should Codex backend clients use `/v1`?

No. Codex backend-compatible clients should use `/backend-api/codex`. The `/v1` surface is for selected OpenAI SDK-compatible clients and translates supported requests into Codex-compatible work.

### Does `/v1` support OpenAI Realtime?

No. `/v1/realtime` and OpenAI Realtime SDK websocket or session routes are unsupported. `GET /v1/responses` is narrow Responses websocket compatibility only, not OpenAI Realtime support.

### Does `/v1/responses/compact` work?

No. `POST /v1/responses/compact` returns a deterministic OpenAI-shaped `unsupported_endpoint` error. Explicit compaction is supported on normal `/v1/responses` requests when the client appends exactly one final `compaction_trigger` after visible input. Backend compact compatibility remains available under `/backend-api/codex`.

### Can I use the same key for `/v1` and `/mcp`?

No. `/v1` uses Pool API keys for runtime work. `/mcp` uses operator-owned MCP bearer tokens for metadata-only lookup. Keep those credentials separate in client configuration and storage.