Skip to content

Responses Lite And Full

Codex backend Responses has two request dialects. The ordinary dialect, which Codex Pooler calls Full, sends instructions and tools as top-level request fields. The Lite dialect carries the same information inside the request input array instead, and sets a marker so the provider knows which dialect it is reading.

Both dialects describe the same turn, reach the same upstream route, and return the same response shape. They differ only in how the outgoing request is assembled.

Codex Pooler decides the dialect per Pool and per model, then guarantees it on the way upstream. Clients keep one Pool API key, one base URL, and one model id in either mode.

Use this page when you need to know what Lite actually changes, why a request shape differs from what your client sent, or what a serving mode does and does not affect. For the operator workflow that sets the mode, see Pools. For where these requests are admitted and dispatched, see Runtime Routes.

Some Codex models are served by a backend that expects the Lite request shape. Its model catalog entry advertises this with a use_responses_lite boolean. A Codex-compatible client reads that flag and builds the matching request; a client that sends the ordinary shape to a Lite model, or the Lite shape to an ordinary one, can be rejected by the provider.

Because a Pool routes one model id across several upstream accounts, Codex Pooler cannot leave that decision to the client. It resolves the dialect itself, advertises the result in its own catalog, and applies the matching request shape at dispatch.

Serving mode never changes the endpoint, the model id, or the credentials.

PropertyLite and Full
Client endpointIdentical. The same /backend-api/codex or /v1 route in both modes
Upstream routeIdentical. Ordinary Responses work and backend compact work reach their matching backend Responses route in both modes
Client-visible model idIdentical
Pool API key and base URLIdentical
Response shapeIdentical. Codex Pooler applies no Lite-specific handling to the response or the event stream
Transport choiceIdentical. HTTP, SSE, and websocket eligibility never read the serving mode

The dialect marker itself travels differently per transport:

  • Over HTTP and SSE it is the upstream request header x-openai-internal-codex-responses-lite: true. It is absent, not false, in Full.
  • Over websockets it is a request-body entry, client_metadata.ws_request_header_x_openai_internal_codex_responses_lite, whose value is the string "true". It is set per turn, so one websocket connection can legitimately carry Lite and Full turns.

There is no Lite marker in the HTTP request body, and no marker of either kind in any response Codex Pooler returns to a client.

On the ordinary Responses lanes — HTTP, SSE, and the websocket response.create turn — Lite rewrites the outgoing request body. Backend compact routes receive the same Lite rewrite before their compact-specific projection runs. This is a real translation, not a flag.

Request fieldFullLite
toolsSent as a top-level arrayRemoved from the top level. The tool list becomes the first input item, an additional_tools item with role developer
instructionsSent as a top-level stringRemoved from the top level. A non-blank value becomes a developer message item placed immediately after the tools item
parallel_tool_callsThe client’s value, preserved exactly, including absent when the client omitted itForced to the boolean false
reasoning.contextLeft as sent, so the provider default appliesForced to the string all_turns
input_image detailPreserved in message content and tool outputs, on native routes and /v1/responses alike; /v1/chat/completions carries image_url.detail into the same field, and so do a chat-style image_url part in a /v1/responses tool message and an image in a Chat tool message or tool result. A null detail is dropped on both /v1 routes. They refuse a value other than low, high, auto or original in either mode with 400 invalid_value and the field path the client sent as param, before anything is sent upstreamThe detail field is removed from input_image entries in typed message content and in function_call_output and custom_tool_call_output outputs
Dialect markerRemoved if a client supplied oneAdded by Codex Pooler

A request that opens a context, one without previous_response_id, always carries the additional_tools item in Lite, even when the request declared no tools. In that case its tools array is empty.

A request anchored on previous_response_id continues a context the provider already holds, and that context already carries the tools item and the instructions message of the request that opened it. Lite therefore sends an anchored request’s input as the client sent it: it adds neither item, and still removes top-level tools and instructions. This is how a Lite-aware Codex client builds its own anchored requests, since it only anchors a request whose tools and instructions equal those of the request before it. A client that changes its tools or instructions has to send a request without previous_response_id for the change to reach the provider in Lite. An anchor whose response was served in Full, after the model’s mode changed, is the exception described in Stability Within A Request.

A client sends this ordinary Responses body:

{
"model": "example-model",
"instructions": "Example system instructions.",
"tools": [
{ "type": "function", "name": "lookup", "parameters": { "type": "object", "properties": {} } }
],
"parallel_tool_calls": true,
"input": [
{ "type": "message", "role": "user", "content": [{ "type": "input_text", "text": "Example user turn." }] }
]
}

Served in Full, the tools and instructions stay where they are. Served in Lite, the same turn is restructured before dispatch:

{
"model": "example-model",
"parallel_tool_calls": false,
"reasoning": { "context": "all_turns" },
"include": ["reasoning.encrypted_content"],
"input": [
{
"type": "additional_tools",
"role": "developer",
"tools": [
{ "type": "function", "name": "lookup", "parameters": { "type": "object", "properties": {} } }
]
},
{ "type": "message", "role": "developer", "content": [{ "type": "input_text", "text": "Example system instructions." }] },
{ "type": "message", "role": "user", "content": [{ "type": "input_text", "text": "Example user turn." }] }
]
}

Top-level instructions and tools are gone, the tool list leads the input, and the instructions follow it as a developer message.

These blocks isolate the serving-mode difference. Codex Pooler applies other normalization to every ordinary backend Responses request in both modes — it maps the model id to the selected upstream account’s identifier, ensures include carries reasoning.encrypted_content, and lowers non-strict tool schemas — so a captured upstream body carries those changes too, in Lite and in Full alike.

Codex Pooler decides, and it closes the loop at both ends.

  1. The Pool’s serving mode for that model resolves to Lite or Full before dispatch.
  2. GET /backend-api/codex/models advertises use_responses_lite as that effective serving mode, not as the raw value an upstream account reported. This is what a Codex-compatible client reads to decide which dialect to build.
  3. A Lite-aware client therefore builds the Lite shape itself, and Codex Pooler applies the same translation again at egress.

Applying the translation twice is safe because it is idempotent. An additional_tools developer item the request already carries is reused rather than wrapped a second time, and a request that already omits top-level instructions produces no duplicate developer message. Normalizing an already-normalized request produces an identical request.

additional_tools is also an ordinary public Responses item type. When the request declared no top-level tools, an additional_tools item with role developer, a tools array and either no id or a non-blank id is the request’s manifest wherever it sits in input, and Codex Pooler forwards it as sent instead of adding another one; a Lite-aware Codex client puts it first and gives it an id. When the request declared top-level tools, the Lite tools item is built from those tools alone, and any additional_tools item in input stays ordinary input.

Clients cannot set the dialect. Codex Pooler removes any client-supplied x-openai-internal-codex-responses-lite header and any client-supplied ws_request_header_x_openai_internal_codex_responses_lite entry from client_metadata on every request, then adds its own only when the effective mode is Lite. Codex Pooler never reads a Lite preference out of a request body or header.

The public GET /v1/models response carries no serving-mode field.

A serving mode belongs to one Pool and one canonical exposed model id. The same model can use different modes in different Pools, and the choice applies to every Pool API key in that Pool.

Configured modeEffective modeRecorded source
auto (the default, stored as no override)Lite when any routable catalog source for that Pool-model pair reports literal true for Lite support, otherwise Fullcatalog
liteLiteoverride
fullFulloverride

Two details of Auto matter in practice:

  • Only a literal true selects Lite. false, a missing key, the string "true", the number 1, and any other malformed value all select Full. When a model entry has no per-source map at all, only a literal true top-level value selects Lite.
  • “Routable” means health-level routable. The assignment is active, eligible, has an active health status, is past any cooldown, and belongs to a visible upstream account. It does not mean that account has quota available, or that this turn will select it. A single Lite-advertising account is enough to put the whole Pool-model pair in Lite.

A model with no routable source at all is not resolved to a serving mode. The request fails eligibility with a 503 and no upstream dispatch, rather than defaulting to a dialect.

Serving-mode choices are stored independently of the model catalog. They persist while a model is stale, retired, or suppressed, remain editable, and re-apply when the model returns under the same exposed model id.

This section exists because the mode is easy to over-attribute. Lite and Full primarily control the outgoing Responses request shape. Features requiring typed tool selection can also constrain the compatible host, as with masked image edits.

ConcernEffect of serving mode
Context windowNone. A model advertises the same context window in Lite and in Full. On native Codex routes, Codex Pooler derives the raw advertised window from the operator context-window override and the model’s pricing bucket, while preserving the effective context-window percentage for Codex to apply once. On /v1/models, Pooler flattens that percentage into context_length. Serving mode is not an input, and the Lite boolean is applied after context policy has already run
Token budgets and auto-compaction limitsNone
Pricing and accountingNone. Reservation is derived from the request payload and the API key’s output-token policy. The mode appears in request and attempt metadata only as diagnostic labels
Routing and eligibilityOrdinary model routing keeps its existing assignment eligibility and ordering. Masked Images edits additionally require an eligible Full image-input host; their adapter selects a compatible host while respecting persisted model overrides
Catalog partition selectionNone. Partitions are computed from the pristine upstream model entries, before the effective mode is applied
Plan familyNone. Lite support is a model catalog property. Codex Pooler has no rule that ties Lite to Free, Go, Plus, Pro, Team, Business, or Enterprise
Response payloadsNone

In the Codex model catalog, the Lite flag and the context-window family are independent fields. A model can advertise Lite with a large context window, or the ordinary dialect with a small one. Neither implies the other.

Codex Pooler applies the Lite translation server-side, at egress, on every lane that dispatches ordinary or backend compact Responses work. A client that knows nothing about Lite can send an ordinary Full-shaped body and Codex Pooler produces the Lite shape for it. Compact dispatch then keeps its own narrow projection and validation rules: compact-only controls are still removed, while Lite tools and instructions are carried in the rewritten input array.

Most OpenAI-compatible clients are unaware of Lite and need no configuration for it. For those clients, a Lite Pool has these observable effects:

  • parallel_tool_calls is overridden to false, so parallel tool calling is off regardless of what was requested.
  • reasoning.context is forced to all_turns.
  • detail pins on input images are not forwarded. A detail of high or low has no effect in Lite.
  • The response, including the event stream, is unchanged.

On native, non-compact Codex backend Responses routes, a present input or tools value must be an array. A non-list value is rejected before dispatch with 400 invalid_request in both Lite and Full, so it is never silently dropped or forwarded with changed meaning. The narrow /v1 surface continues to normalize a string input into an item array before serving-mode handling, so existing /v1 string-input behavior is unaffected.

There is one hard limit. On a Lite-served model, a map-shaped tool_choice is rejected before any upstream request, with HTTP 400, code unsupported_parameter, and param: "tool_choice". This includes the declaration-backed allowed_tools map accepted only by direct public Responses in Full mode. The rejection follows the model’s serving mode rather than the endpoint, so it applies to a Chat Completions named-function choice too. Scalar choices such as "auto", "none", and "required" are accepted in both modes. If a client needs forced typed tool selection, that model must be served in Full.

Clients that implement Lite themselves stay compatible, because the translation is idempotent: a Lite-aware client talking to a Lite Pool produces the same upstream request whether the client or Codex Pooler built the shape.

The conversion runs in one direction only. Codex Pooler translates an ordinary request into the Lite dialect when the mode is Lite; it never translates a Lite-shaped request back into the ordinary dialect. A client that forces Lite on its own side while the Pool serves Full will have its Lite-shaped body forwarded as sent, without the Lite marker. Leave client-side Lite settings at their default and let the Pool decide, which is what a client does when it reads use_responses_lite from the Pool’s own catalog.

Unmasked GPT Image 2/2.5 generation and edits use the native Codex Images service. Its catalog host is an account-capacity carrier; a Full or Lite label on that request does not change the native image payload or establish a quality level.

Masked POST /v1/images/edits requests instead require a Full Responses host with image-input support. Codex Pooler preserves the original mask in input_image_mask and explicitly selects the image generation tool. An eligible Full host is selected under the Pool’s configured model and serving policy; an explicit Lite override on an exact requested model is not bypassed. The resolved mode is checked again before reservation.

If no eligible Full host exists under that policy, the request returns 400 unsupported_parameter with param: "mask", without an upstream request or accounting reservation. Pooler does not discard the mask, silently override Lite, or fall back to an unmasked edit. Describe the intended edit clearly and inspect the generated result: mask forwarding does not guarantee exact pixel preservation. See image generation and edits.

One HTTP request, or one websocket response.create turn, keeps the serving mode it started with. The resolved mode is captured once, before dispatch, and is then immutable for that request. It survives:

  • same-assignment retry,
  • cross-assignment failover to a different upstream account,
  • cross-node owner forwarding.

Saving a new mode does not mutate a request or turn that is already in flight. A change is picked up by the next HTTP request, or by the next response.create turn on an already-open websocket connection.

A context the provider holds keeps the dialect of the requests that built it: a context opened under Full holds no tools item and no instructions message, and one opened under Lite holds both as input items. So on a backend websocket, a Lite response.create turn anchored on previous_response_id whose previous response on that connection was served in Full is not sent, because neither the request nor the context would carry the tools item and the instructions message. It receives the same previous_response_not_found error a Codex client gets for an anchor its connection no longer holds, and a Codex client answers it by sending the full request again without previous_response_id, which opens a Lite context with both items. The reverse change is sent as usual: a Full anchored turn carries its tools and instructions at top level.

The provider resolves previous_response_id only on the upstream websocket connection that produced the response, and over HTTP it rejects the parameter with a 400 (Unsupported parameter: previous_response_id) in either mode. Outside the backend websocket, an anchor therefore reaches a context only on the public /v1/responses websocket, or on a streaming /v1/responses HTTP request that Codex Pooler sends over its session’s upstream websocket, as described in OpenAI-compatible clients. An OpenAI SDK does not send a request again after previous_response_not_found, so these anchors are not refused after a mode change. Instead, Codex Pooler records on each response of a session the mode it was served in, and a Lite request anchored on a response served in Full carries the tools and instructions it declares at top level as the tools item and the instructions message, which gives the context what it lacks. A request that declares neither, such as the anchored turn of a client that builds Lite-shaped requests itself, gets nothing added. Every other anchored request receives a 400, whatever its mode: a tool-output continuation on the native HTTP route goes upstream over HTTP and the provider refuses it, and every other anchored /v1/responses request, a non-streaming one or a continuation whose session has no connection that produced the anchor, is answered with previous_response_not_found before it reaches the provider.

Because the effective mode is part of the backend catalog body, changing it changes the GET /backend-api/codex/models ETag. A backend websocket keeps its upgrade ETag for backward compatibility, but every accepted response.create turn also receives a codex.response.metadata event whose x-models-etag value is authoritative for that turn. An already-open connection can therefore observe the new ETag on its next turn; an in-flight retry keeps the turn snapshot it started with.

Codex Pooler records the mode on each request and attempt as three metadata labels: the configured mode, the effective mode, and the source. They are bounded values only — auto, lite, or full for the configured mode; lite or full for the effective mode; catalog or override for the source. No request body, prompt, tool payload, header, or credential is retained with them.

The mode a Pool is using is visible in two places:

  • the Pool’s edit dialog, Models step, which shows the configured choice and the resolved effective mode per model, and
  • the authenticated GET /backend-api/codex/models response, whose use_responses_lite boolean is the effective mode for that Pool.

Full is an advanced, provider-dependent override, and Codex Pooler never silently downgrades it. When an ordinary Responses HTTP request under an explicit Full override receives a terminal non-rate-limit 4xx, the request log classifies it as upstream_status — the same code the same rejection receives under Auto or Lite, because the serving mode is recorded in the routing metadata rather than in the error code — and the client receives the upstream status with a server-owned message. The error body relays the same bounded type, code, and param that the attempt already records as sanitized metadata, so a terminal invalid_request_error is not presented as a retryable server_error. When the provider supplies a type but no code, the client-visible code is derived from the type: invalid_request_error becomes invalid_request, the same code Codex Pooler uses for its own pre-dispatch rejections, and any other type is reused as the code. The server-owned message names the relayed param and code (upstream rejected parameter tools.defer_loading (invalid_request), or upstream rejected the request (<code>) when the rejection carries no param), built by the same constructor the Auto and Lite parameter-validation relay uses, so serving mode does not decide how much a client is told. Auto, Lite, Full relay the same bounded supported-values list from persisted attempt metadata. For an allowlisted HTTP 400 unsupported_value or invalid_value rejection, the suffix includes at most 12 identifier-shaped alternatives parsed from a strictly shaped trailing provider list, excluding every value quoted earlier in the message. An absent or unsafe list produces no suffix. The provider message and body are never forwarded. A rejection with no sanitized type keeps the fixed server_error body. A 429 is classified upstream_rate_limited and an ordinary 5xx keeps upstream_status; neither relays rejection fields, because neither records them, and both are unrelated to the serving mode. No serving mode produces an error code of its own.

For example, an explicit Full request with a qualifying validation rejection receives:

{
"error": {
"type": "invalid_request_error",
"code": "unsupported_value",
"param": "reasoning.effort",
"message": "upstream rejected parameter reasoning.effort (unsupported_value); supported values: low, medium, high"
}
}

Serving modes are database-backed Pool configuration. There is no environment variable or Helm value that sets them.

  • Pools for the operator workflow that selects Auto, Lite, or Full
  • Runtime Routes for the routes these requests are admitted on and the custom tool contract that interacts with tool_choice
  • Routing Strategies for ordinary upstream account selection and eligibility
  • Request logs for per-request metadata and error classification