Skip to content

Magpie on Codex Pooler

Magpie is a local gateway and model switcher for coding agents. It edits an agent’s model settings, listens on localhost, and forwards each request to the provider you picked, rebuilding the request in another API format when that provider needs one. Connect it to Codex Pooler to keep Magpie’s model picker and gateway in front of Codex while your Pool serves the requests.

Codex CLI and Codex Desktop can also use Codex Pooler directly, and Codex CLI / Desktop on Codex Pooler remains the primary setup for them. This guide covers the Magpie recipe that keeps Codex’s requests closest to what Codex sends, and the limits that come with Magpie in the path.

  • Install Magpie and Codex CLI or Codex Desktop using their official instructions for your operating system.
  • Have a Codex Pooler URL reachable from the machine that runs Magpie. Magpie, not Codex, calls the Pooler.
  • Create a Pool API key and choose a model available to that Pool.
  • Keep Magpie signed out of any ChatGPT account that is also a Codex Pooler upstream.

Run these commands in a terminal on the machine that runs Magpie. They use no shell-specific syntax, so they read the same in macOS and Linux shells, WSL and Windows PowerShell.

Terminal window
magpie provider add "Codex Pooler" id=pooler responses=https://codex-pooler.example.com/backend-api/codex/v1 key=<pool-api-key> models=gpt-6-luna search=yes

For a local Pooler, use responses=http://localhost:4000/backend-api/codex/v1.

  • responses= is the Codex backend alias. Magpie appends /responses and sends Codex’s Responses request to POST /backend-api/codex/v1/responses. A base of https://codex-pooler.example.com/backend-api/codex sends to POST /backend-api/codex/responses instead, which works as well.
  • Set responses= alone. A provider that declares responses= gets Codex’s requests on the Responses route whether or not it also has url=; only a provider with url= and no responses= is spoken to in Chat Completions, see Chat Completions recipe.
  • models= lists the model ids your Pool serves, separated by commas. The alias lists models in Codex’s own catalog format, which Magpie does not read: without models=, Magpie reports no models listed and exposes no model. To follow the Pool’s list instead, drop models= and add models.url=https://codex-pooler.example.com/v1/models. Magpie then exposes every model the key can see, including models the Codex catalog hides.
  • search=yes tells Magpie that the provider answers Codex’s hosted web search itself. Magpie then relays Codex’s request instead of rebuilding it; see What reaches Codex Pooler.

Magpie takes the key as an argument and stores it in plain text in its own configuration directory (on macOS, a file readable only by your user). Treat that file like the key itself, and keep the command out of shared shell history.

Terminal window
magpie codex pooler/gpt-6-luna

magpie codex edits your user-level Codex configuration in place and keeps your other settings:

OS Default config file
macOS ~/.codex/config.toml
Linux ~/.codex/config.toml
Windows %USERPROFILE%\.codex\config.toml

Magpie edits the file under your home directory and does not read CODEX_HOME. With CODEX_HOME pointing elsewhere, it still writes ~/.codex/config.toml and leaves the CODEX_HOME file untouched, so Codex does not see the change. Keep CODEX_HOME at its default while you use Magpie, or make the same edits in your CODEX_HOME file by hand. For a client installed inside WSL, run Magpie and the commands inside WSL as well.

After the command, the file contains the following, in addition to what was there before:

CODEX_HOME/config.toml
model = "pooler/gpt-6-luna"
model_provider = "magpie"
model_catalog_json = "<home>/.codex/magpie-models.json"
openai_base_url = "http://127.0.0.1:3425/backend-api/codex"
[model_providers.magpie]
name = "magpie"
base_url = "http://127.0.0.1:3425/v1"
wire_api = "responses"

Magpie also writes a placeholder credential and an x-openai-actor-authorization header into the provider table; its gateway accepts any value on localhost. openai_base_url applies to Codex’s built-in openai provider, which the selected magpie provider does not use. Restart Codex Desktop and any open Codex sessions afterwards, because Codex builds its model list at start-up.

magpie codex default takes Magpie out again: the previous model returns and model_provider, model_catalog_json and openai_base_url are removed. It leaves the unused [model_providers.magpie] table in the file. Codex lists resumable conversations per model_provider, so sessions started through Magpie are separate from your others; see Existing session migration.

Codex sends its requests to Magpie, so Magpie must be running: open the Magpie app, or run magpie serve in a terminal. The gateway listens on 127.0.0.1:3425 unless MAGPIE_ADDR names another address, and magpie codex writes the address it is configured for into the Codex config. MAGPIE_DEBUG=1 magpie serve prints one line per request, for example pooler/gpt-6-luna pooler → pooler (responses→responses) 200 2116ms.

Select a model with magpie codex pooler/<model-id>, where pooler is the provider id from the command above. Magpie sends the model id without the prefix to the Pool (gpt-6-luna in the example). magpie models lists what Magpie exposes.

Magpie builds the Codex catalog entry for each model from its own model data. For a model it has no data for, the entry declares text input only, no reasoning levels, no service tiers and no Responses Lite mode. With such an entry:

  • Codex sends reasoning.summary but no reasoning.effort, so the Pool’s default effort for the model applies. To choose an effort in Codex, give Magpie the levels, for example magpie model efforts pooler/gpt-6-luna low,medium,high, and restart Codex. The selected effort is then forwarded.
  • Fast mode is not offered. With a service tier configured, Codex reports that the tier is not advertised as supported for the model and omits it from requests.
  • Codex builds Full-shaped requests, even for a model the Pool serves as Lite, and the Pool accepts them. See Responses Lite versus Full for how the Pool decides the mode.

Run a short prompt from a disposable folder, with Magpie’s gateway running:

Terminal window
codex exec --skip-git-repo-check "Reply with exactly: magpie ok"

Confirm that the reply arrives and, if you started the gateway with MAGPIE_DEBUG=1, that its line for the request ends in 200. In Codex Pooler’s request logs, match the request time, API key, model and final status. A reply alone does not confirm that the request went through your Pooler.

To check tool use and continuation, ask Codex to run echo magpie-probe in the same folder. The turn makes two requests, and the second one normally reports cached input tokens in its usage.

Magpie decides per request whether to relay Codex’s Responses body or to rebuild it. The table shows what reaches the Pooler for each provider setup and which turns it serves.

Magpie provider Codex request What reaches Codex Pooler Serves
responses= alias, search=yes Full-shaped, from Magpie’s catalog entry Codex’s body; tool_search becomes a plain function; web_search, apply_patch and sealed reasoning items stay Plain and tool turns
responses= alias Lite-shaped, from Codex’s own entry for a Lite model The additional_tools manifest moved into tools with the functions namespace opened; namespaces such as collaboration and the x-openai-internal-codex-responses-lite header kept Plain turns and the lead turns of a multi-agent run
responses= alias, without search=yes Full-shaped, from Magpie’s catalog entry A rebuilt request: developer messages folded into instructions, every tool a plain function (namespaces flattened to namespace__name), hosted web_search removed, sealed reasoning not replayed Plain and tool turns, without hosted search or sealed reasoning
responses= at /v1 Full-shaped A rebuilt request as for the alias without search=yes, sent to POST /v1/responses Plain and tool turns, without hosted search or sealed reasoning
responses= at /v1 Lite-shaped Codex’s body with the manifest lifted Plain turns and replayed multi-agent history
url= without responses= Full-shaped A rebuilt Chat Completions request on POST /v1/chat/completions Plain and tool turns, without reasoning continuity or hosted search

The alias recipes answer on POST /backend-api/codex/v1/responses and the /v1 recipe on POST /v1/responses, for Full-shaped and Lite-shaped requests alike. The Chat Completions recipe answers on POST /v1/chat/completions.

Lite-shaped requests appear when Codex uses its own catalog entry for a Lite model, for example the bundled gpt-6-luna entry, instead of Magpie’s entry. They carry no hosted tools, so Magpie relays them with or without search=yes. Codex’s identification and continuity headers (originator, session-id, thread-id, x-client-request-id, x-codex-turn-metadata, x-codex-window-id and x-codex-beta-features) reach the Pooler on every route above.

Use search=yes: it keeps Codex’s request, including the sealed reasoning items and the freeform apply_patch tool, as Codex sent it. Without it Magpie rebuilds every default Full request, because Codex’s default request carries the hosted web_search tool.

With search=yes, the request that reaches the Pooler lists the hosted web_search tool as Codex sent it, and the Pooler forwards it upstream. Without search=yes, Magpie removes the tool from the request, or substitutes a search API you configured in Magpie.

A Magpie provider that has url= and no responses= has no Responses route, so Magpie rebuilds a default Full request as POST /v1/chat/completions. The rebuilt Chat request has no reasoning continuity and no hosted search, and Codex’s freeform apply_patch arrives as a function. Prefer the responses= alias recipe above.

A responses=https://codex-pooler.example.com/v1 provider sends Codex’s requests to the narrow OpenAI-compatible surface instead of the Codex backend alias. Codex backend clients should use the alias; see Should Codex backend clients use /v1?. The /v1 surface accepts the agent_message items of a multi-agent history, plaintext or sealed (see Multi-agent mailbox history replay), the hosted web_search options Codex sends (see Web search tool options), and the web_search_call items a hosted search leaves in a history (see Hosted web search history replay). It drops reasoning items from stateless requests by design, as described in Continuing with previous_response_id.

no models listed while adding the provider means Magpie could not read the model list the alias returned. Add models= or models.url= as described above; magpie provider set pooler models=gpt-6-luna adds models to an existing provider.

400 sealed subagent task comes from Magpie, not from Codex Pooler. It appears when MultiAgentV2 (features.multi_agent_v2) is enabled and a lead model spawns a subagent: the provider seals the spawn_agent task, Codex carries it into the subagent’s first request as an agent_message with encrypted_content, and Magpie refuses to send it to a provider that has no ChatGPT account of its own. The lead’s other turns are answered normally; the subagent’s request never reaches the Pooler, and the lead reports a subagent error. Run multi-agent work with Codex pointed directly at Codex Pooler.

input item shape is not translatable from Codex Pooler means the request reached /v1 and carried an input item the narrow surface does not admit. Point the provider at the alias.

Configured service tier ... is not advertised as supported is Codex declining a service tier that Magpie’s catalog entry does not list. It is not an error.

If Codex shows the old model list after magpie codex, restart Codex Desktop and the open sessions.

  • Transport. Under Magpie’s wiring Codex sends every model request as an HTTP POST answered as server-sent events and does not open the backend websocket routes, so websocket continuation and previous-response anchors are not used.
  • Compaction. Codex compacts the conversation itself with an ordinary tool-less Responses request carrying a summarization prompt and does not call POST /backend-api/codex/responses/compact, so Pooler-side backend compaction is not used.
  • Response headers. Magpie does not pass Codex Pooler’s response headers such as x-codex-turn-state and x-models-etag back to Codex. Codex receives the request id and Magpie’s own x-magpie-model and x-magpie-provider headers.
  • Reasoning continuity. Rebuilt requests (no search=yes) discard the sealed reasoning items a provider returns and do not replay them. Relayed requests replay them.
  • Prompt cache key. After a 400 that does not name a parameter, Magpie retries without prompt_cache_key and, when that succeeds, stops sending the key until Magpie restarts. Codex Pooler uses the key as a routing hint for cache locality, so restart Magpie after a rejected request if cache reuse drops.
  • Retries. A failing turn is retried by Codex and again by Magpie. When every attempt fails, one Codex turn can produce dozens of requests to Magpie and about three times as many from Magpie to the Pool within two minutes, all with the same session and request identifiers.
  • Subagents. See 400 sealed subagent task above. Codex’s other multi-agent tools reach the Pooler unchanged on the relayed routes.
  • Scope. Codex Pooler supports Codex model-provider traffic only. Do not point Magpie’s Codex sign-in or any account helper at Codex Pooler. Operator MCP is not part of this setup.