Skip to content

Routing Strategies

Routing is the path from a client request to one upstream account. Codex Pooler uses a Pool API key to identify the Pool, checks whether the request can run there, then chooses an eligible upstream account for that turn.

A Pool API key represents a Pool. It doesn’t represent one upstream account. That is the main difference from direct account credentials: clients stay configured with one stable key, while Codex Pooler can choose among the Pool’s active upstream assignments for each supported request.

Codex Pooler first removes upstream accounts that fail hard eligibility checks, then orders the remaining accounts with the Pool’s routing strategy. Eligibility covers assignment state, upstream lifecycle, model support, capability support, quota evidence, route health, file affinity, and session continuity. Strategy only ranks accounts that are already allowed.

Use this page with Runtime Routes and Operator admin UI when you need to understand why a request did, or didn’t, choose a specific account.

Every runtime request follows the same broad path:

  1. Codex Pooler authenticates the bearer token as a Pool API key.
  2. It admits the request into the route family and route class, such as normal HTTP, stream, websocket, compact, file upload, or audio transcription.
  3. It classifies the requested model and route surface, either the Codex backend compatibility route or the narrow OpenAI-compatible /v1 surface.
  4. It finds the authenticated Pool and starts from that Pool’s configured upstream assignments.
  5. It applies hard eligibility checks for assignment status, upstream status, model support, requested capabilities, quota, health, route-class circuit state, file affinity, and session continuity.
  6. It orders the remaining eligible upstream accounts using the Pool’s routing strategy and routing preferences.
  7. It tries the ordered shortlist. If one upstream fails in a retryable way, Codex Pooler can move to the next eligible account in that shortlist.
  8. It records metadata-only request, route, attempt, quota, and audit evidence. It doesn’t store raw prompts, completions, files, audio, images, credentials, or websocket frames.

The important split is this: eligibility decides what is allowed. Strategy decides which allowed account is preferred first.

Hard eligibility checks remove upstream accounts from consideration before strategy ordering begins. A routing strategy can’t select an account that failed these checks.

CheckWhat it means for users
Pool accessThe Pool API key must be active and tied to the Pool that should receive the request
Pool assignmentsThe upstream account must be assigned to that Pool and active for routing
Upstream lifecyclePaused, disabled, deleted, or reauth-required upstream accounts aren't eligible for new ordinary work
Model availabilityThe requested model must be exposed to the Pool and mapped to the assignment
Capability supportThe assignment must support the requested shape, such as streaming, tools, reasoning, image input, audio transcription, service tier, or compact responses when that support is known and required
QuotaFresh reset-bearing windows are the normal authority for the account, model, and limit family. A narrow lower-priority fallback can use fresh provider-attested availability when no account windows exist; explicit blocked, unknown, stale, credential-mismatched, wrong-model, malformed, or untrusted evidence still removes the account
HealthRoute-class circuit state can temporarily remove an account for a route after repeated backend failures
Session continuityExisting Codex sessions, websocket sessions, previous response links, and file affinity can pin the request to the upstream assignment that owns that state

If every account fails eligibility, the request is rejected before upstream dispatch. Common user-facing outcomes include no eligible backend, no compatible backend, quota exhaustion, quota evidence unavailable, session assignment unavailable, or file assignment conflict.

Provider-attested availability without windows

Section titled “Provider-attested availability without windows”

Fresh reset-bearing quota windows remain the normal routing authority. When the provider returns no account windows, Codex Pooler can use a distinct lower-priority availability observation only when the provider explicitly reports that requests are allowed and the limit is not reached, or explicitly reports usable or unlimited credits. The observation must be fresh, belong to the current credential generation, and have no retained account-window evidence or applicable model or additional-limit blocker.

This fallback is fail closed. An explicit provider block, contradictory or unknown status, stale observation, credential-generation mismatch, malformed status, exhausted applicable model limit, or blocked additional limit keeps the account out of routing. A neutral credit report with neither available nor unlimited credits is not positive capacity by itself.

Codex Pooler doesn’t create a quota window, reset timestamp, duration, percentage, or remaining-capacity value for this state. It doesn’t depend on an operator setting, plan badge, plan name, or SKU. Those fields can describe an account, but they don’t grant routing.

The availability flags belong to Codex Pooler’s direct provider usage response. Current Codex client quota projections omit those raw flags, so a client-side quota snapshot is not evidence that the same no-window contract was observed.

After eligibility, Codex Pooler orders the remaining accounts. The Pool’s strategy affects preference, not permission.

StrategyUser-facing behavior
Bridge ringDefault strategy. It spreads requests with stable scoring, then uses continuity, prompt-cache locality, and temporary demotions to keep useful stickiness without pinning every stateless request
Deterministic rotationRotates the eligible set from a stable request seed so repeated independent requests don't always start with the same account
Least recent successPrefers accounts that haven't recently completed successful work, which can help spread successful turns across the Pool
Quota firstPrefers eligible accounts with more usable remaining quota evidence for the requested model. A fresh reset-bearing report at 100% used counts as zero capacity, while eligibility confidence remains a separate check

The ring size controls how many ordered eligible accounts are tried for a request. A ring size of 3 means Codex Pooler prepares up to three eligible upstream accounts in order. If the first account fails in a retryable way, it can try the next one. The ring doesn’t include accounts that failed eligibility.

Temporary demotion also affects order. When an upstream attempt fails with a retryable backend or network reason, Codex Pooler can demote that assignment briefly so other eligible accounts in the same quota tier are tried first. A later success clears that demotion.

Demotion doesn’t change the quota tier. Accounts that route through provider-attested availability without windows stay behind every account with fresh reset-bearing quota evidence, even when that account is demoted. Among the provider-attested accounts, a demoted one moves behind the others. The ring size applies after this ordering, so a small ring can leave out demoted accounts or provider-attested accounts entirely.

Session Continuity Versus Stateless Routing

Section titled “Session Continuity Versus Stateless Routing”

Session continuity is stronger than ordinary stateless routing.

Stateless requests can be routed to any eligible upstream account in the Pool. Strategy, remaining quota, prompt-cache locality, and recent failures shape the order.

An existing Codex session prefers its current upstream account. A native websocket request carrying portable full history can use another eligible account while keeping the client’s websocket open, including history that starts with the provider’s compaction checkpoint or contains recognized encrypted messages between Codex agents. Previous-response anchors, uploaded file references, item references, and a compaction still being collected on the upstream connection require the account that owns that state. Additional file references or opaque state inside an agent message retain those restrictions.

If an upstream rejects a portable full-history websocket request with a quota-limit terminal before output and without reported token usage, Codex Pooler can try the next eligible account in the configured ring. Successful recovery updates the session’s account for later turns. Encrypted reasoning items and compaction checkpoints can remain in that history; item references retain their continuity restrictions. Work that has already produced output or reported token usage is not silently repeated on another account.

Continuity headers are local routing inputs. Codex Pooler reads them in this order:

  1. x-codex-window-id
  2. x-codex-session-id
  3. session-id
  4. x-session-id
  5. x-session-affinity
  6. session_id
  7. x-codex-conversation-id

x-session-id and x-session-affinity aren’t forwarded upstream. They help Codex Pooler find the local routing state only. On the native /backend-api/codex/responses and /backend-api/codex/responses/compact routes, the Codex client’s own session-id, thread-id, and x-client-request-id headers are forwarded verbatim when they are short ASCII identifiers, because the provider uses them for sticky routing and a full-history HTTP turn without them misses the prompt cache the previous turn warmed. The native /backend-api/codex/responses websocket forwards the same three headers, under the same limits, on the upstream websocket handshake, so a fresh or reopened upstream connection reaches the same provider session. An upstream websocket connection is reused only for turns that carry the same values; a different or missing value opens a new upstream connection. /v1 routes keep every client continuity header local, including on /v1/responses websocket and bridged turns. When a /v1/responses or /v1/chat/completions request carries a prompt_cache_key, Codex Pooler sends the provider a session-id derived from that key together with the calling API key and its Pool (a UUID v5 over a fixed Codex Pooler namespace), so consecutive HTTP turns of one conversation reach the replica that holds the warm prompt cache. The derived value is stable across nodes and restarts, is never stored or logged, and is not sent for keys that are empty, longer than 512 bytes, or not strings. Because the API key and Pool are part of the derivation, two API keys or two Pools that send the same prompt_cache_key (for example a generic default) never share a provider session-id, even when their requests reach the same upstream account; the model is not part of it, so switching models inside a conversation keeps the same session-id. The native /backend-api/codex/responses and /backend-api/codex/responses/compact HTTP routes derive the same value, from the same key, API key and Pool, when the client sent no usable session-id of its own, for example a third-party client that names its conversation only with session_id or x-session-id. Such an alias still selects the local continuity session but is never forwarded, because it carries no tenant scope. When the request carries no prompt_cache_key either, as with clients that send only a conversation header, Codex Pooler derives the provider session-id from that header with the same API key and Pool scoping, under a separate namespace, so consecutive turns of one conversation still share one provider session-id. The order is: the client’s own session-id, then prompt_cache_key, then the conversation header. A client session-id is always forwarded unchanged, and the native websocket handshake still carries only the session headers of the upgrade. The derived session-id only steers provider-side sticky routing: Codex Pooler creates no local session from it, so bridge eligibility and continuity still require a real continuity header. Within one API key the derivation assumes a conversation-scoped key. A client that reuses one prompt_cache_key across unrelated conversations on the same API key (for example one key per end user or per application) sends all of them to the provider under the same session-id, which concentrates them on the same provider replica; use a per-conversation key when that is not the intent.

If a request still requires an upstream-owned anchor or file and that account is no longer eligible, Codex Pooler cannot silently move that request. A client-authored full-history request can recover when it contains the portable context needed by another eligible account. Check the upstream’s lifecycle state, reauth status, Pool assignment, model support, quota evidence, and route health when no eligible recovery target remains.

Prompt-cache locality is a preference for stateless requests that share a prompt_cache_key.

When prompt-cache affinity is enabled for the Pool and there is no stronger continuity rule, Codex Pooler uses the transient prompt_cache_key as a routing seed. That makes repeat stateless requests prefer the same eligible upstream account, which can improve provider-side cache locality.

This doesn’t store prompts or responses locally. The raw prompt-cache key isn’t stored as request evidence. It is used as a routing input, then request and routing logs remain metadata-only.

Prompt-cache locality can be skipped when:

ConditionResult
No prompt_cache_key is presentNormal strategy ordering applies
Prompt-cache affinity is disabled on the PoolNormal strategy ordering applies
Only one upstream account is eligibleThere is nothing to choose among
A Codex session, idempotency key, or durable affinity already appliesThe stronger continuity rule wins

Fallback is bounded by eligibility and ring size.

If the first chosen upstream account returns a retryable backend status, rate limit, network error, authorization failure, or server error, Codex Pooler can try the next eligible account in the route plan. It records safe attempt metadata and may briefly demote the failing assignment for the same Pool and model.

Fallback doesn’t bypass hard checks. It won’t use a paused account, an account without model support, an account with unusable quota evidence, or an account that conflicts with session or file affinity. For session-bound work, fallback is intentionally limited because moving the request to another account could break the conversation state.

Saved reset redemption is quota recovery, not a routing strategy. Saved-reset auto redemption cannot make duplicate or stale turns, circuit-open assignments, session-conflicting assignments, unsupported models, paused or deleted identities, or unhealthy assignments eligible.

Auto redemption is default-disabled per upstream account because saved resets are limited provider capacity. After required quota filtering reaches quota_exhausted, the gateway may attempt one redemption on a candidate whose own exclusions are limited to its account long window: weekly or 30-day monthly. The window is identified by account scope and reported duration, not the plan name. Other candidates may have different quota blockers without preventing that candidate’s recovery; their exclusions do not authorize spending a reset on them. The selected candidate still requires corroborated exhaustion, a reported saved-reset count greater than its resets-to-keep policy, no blocking redemption already claimed, and a natural reset far enough away to satisfy the Natural reset buffer. The target’s other applicable windows must remain usable, including when a provider-level blocked status hides their individual reasons; this is checked again before consumption. A reset never recovers a 5-hour, model-specific or additional-meter blocker through this policy. Conflicting long-window descriptors fail closed. Existing continuity, circuit, credential and consumption safeguards remain in force, and reset candidates retain their existing order rather than being ranked by credit expiration.

Early request-driven redemption can also run before hard exhaustion when the operator selects near-limit mode and every eligible candidate account has fresh long-window quota evidence at or above the configured threshold. The Natural reset buffer applies to that pool-wide threshold path, so near-limit auto redemption waits when the natural reset is close. Long-window exhaustion and configured threshold pressure are the only request-time recovery triggers; request traffic does not own expiration rescue.

Separately, a known saved reset inside the next 24 hours can enter traffic-independent scheduled observation and consideration after canonical account reconciliation persists current quota and saved-reset evidence. Entering that horizon does not consume a reset by itself. Scheduled redemption requires B1 long-window exhaustion, B2 the configured threshold for that identity, or B3 the final inclusive 90-minute last call. Scheduled threshold rescue is evaluated per identity; request-driven threshold pressure remains pool-wide. B1’s ordinary exhausted path and B2 require the configured Natural reset buffer. B1 may bypass it when a fresh expiration observation shows expiration before the natural long-window reset (E < R). B3 requires positive long-window usage, fresh expiration evidence, and E < R instead of the full configured buffer.

This timing policy is a value-preservation heuristic, not a guarantee that the saved capacity will be realized. Reconciliation downtime during the final margin can miss the last-call opportunity, and the expiration that justified a scheduled decision is timing evidence rather than an exact target. Expiration at decision time does not identify the credit the provider ultimately consumes. The ChatGPT consume path still sends the first usable credit identifier in provider-returned order, while the Codex consume path omits credit_id and lets the provider select the credit.

When a redemption succeeds, the provider consuming a reset credit is an authoritative side effect. The credit is spent even if the immediate usage refresh is partial or temporarily omits the account’s refreshed quota window. Codex Pooler records that consumed credit truthfully and treats the account as pending confirmation, rather than assuming it is usable or that it is still exhausted. It never fabricates a quota value.

Only redemption triggered by an in-flight request can let that triggering request act as one guarded probe against the redeemed account. Codex Pooler generates one internal lease token and persists its exact scope: Pool assignment, upstream identity, effective model, and route class. This is pooler-owned correlation material, not a client or provider credential. It is never returned, rendered, logged, or included in request, audit, or routing metadata. Scheduled expiry rescue has no request context: it creates neither a request probe nor an accounting request. If its immediate confirmation is incomplete, it remains fail-closed until immediate or later fresh reconciliation evidence confirms recovery or reblocks the account.

A PostgreSQL-coordinated claim guarantees that at most one request across all nodes probes. The claim is irreversible and stays bound to that exact scope. If the probe fails, is cancelled, times out, or finishes after its confirmation deadline, Codex Pooler ignores the success result and never unlocks, replaces, reroutes, or broadens the claim to another request, account, model, or route class. A successful probe only confirms before its deadline and makes the account temporarily routeable for a bounded confirmation window. Real quota evidence observed at or after the credit was consumed then confirms recovery or reblocks the account. If no such evidence arrives within the window, the account fails closed. Reconciliation applies the same evidence-only convergence, so an account left pending recovers automatically once the provider reports a real quota window again.

If no saved reset can be applied, or the redemption result does not change quota evidence, the original quota decision stands. Manual Redeem saved reset from the upstream cockpit queues one account-level recovery attempt through the same metadata-only redemption boundary; it is an operator action, not an in-flight routing fallback.

Saved-reset observations and redemption attempt state are stored on upstream_identities.metadata; routing and redemption metadata stores bounded status, result codes, timestamps, HTTP status, and before/after counts only. It does not store raw provider payloads, credit identifiers, redemption request identifiers, client or provider credentials, prompts, request bodies, or response bodies. The internal one-shot lease token is the sole non-secret token in the redemption lifecycle, and it remains non-renderable and absent from public metadata.

Saved reset counts do not boost ranking. quota_first still orders only already-eligible candidates by usable remaining quota; saved reset metadata only determines whether an opt-in quota recovery attempt is allowed.

Operators configure routing from the admin UI. The exact visibility depends on role, but the user-facing controls are:

Admin areaWhat users can configure or inspect
PoolsPool lifecycle, Pool assignments, Pool API key assignments, routing strategy, ring size, sticky websocket sessions, HTTP affinity, prompt-cache affinity, and /v1 compatibility
UpstreamsUpstream account readiness, lifecycle state, Pool assignment, model support, quota evidence freshness, saved reset count, saved reset redemption policy, and reauth state
Pool API keysWhich client credential represents the Pool, plus key status and policy metadata
System settingsGateway admission defaults, route-class limits, diagnostics, circuit thresholds, model metadata, pricing catalog settings, and runtime policy controls
Request logsMetadata-only evidence about route family, Pool, model, selected upstream metadata, status, retries, duration, quota state, and safe error codes

The admin UI doesn’t expose raw prompts, completions, request bodies, response bodies, file bytes, audio bytes, image bytes, websocket frames, bearer tokens, raw Pool API keys after creation, upstream tokens, or Codex auth.json contents.

Can a routing strategy bypass eligibility checks?

Section titled “Can a routing strategy bypass eligibility checks?”

No. Routing strategies only order upstream accounts that already passed hard eligibility checks. A strategy cannot choose a paused upstream, an unassigned upstream, an account without the requested model, an account with unusable quota evidence, or an account that conflicts with session or file continuity.

Why does session continuity override stateless routing?

Section titled “Why does session continuity override stateless routing?”

Session continuity protects resumable Codex work. If a request belongs to an existing Codex session, websocket session, previous response chain, or file affinity, Codex Pooler keeps it with the upstream assignment that owns that state instead of treating it as a new stateless request.

What happens when the preferred upstream fails?

Section titled “What happens when the preferred upstream fails?”

For stateless work, Codex Pooler can try the next eligible upstream in the prepared ring when the first attempt fails in a retryable way. Fallback does not bypass hard checks, and session-bound work is intentionally stricter because moving state to another upstream can break the conversation.

No. Prompt-cache locality uses the transient prompt_cache_key as a routing input when Pool policy allows it and no stronger continuity rule applies. The raw prompt-cache key is not stored as request evidence, and raw prompts or responses are not persisted.