Skip to content

Runtime triage

Start from the symptom and its time window. Compare the affected pods, client/provider transports and deployment changes before attributing an error to a client, Codex Pooler or the provider.

PromQL recipes contains the matching queries. Logs and memory explains what to collect when metrics cannot identify the cause.

codex_pooler_gateway_websocket_control_failure_count records observed callback failures with bounded phase and reason labels. Phases are init, serve and terminate.

These failures can occur before request reservation, so an empty Request logs error list does not rule them out. cleanup_deferred means cleanup exceeded its short synchronous budget and continues under supervision; it does not, by itself, prove that cleanup failed.

When the rate rises, compare database checkout pressure, owner-process logs and rollout timing. This counter does not measure TCP resets or prove that a close frame reached the client.

codex_pooler_gateway_duplicate_turn_refused_count counts client-visible 409 duplicate_turn refusals. The refused duplicate does not get its own request row, so database-derived request/error totals do not include it.

A few refusals can be normal after a disconnect: a client resends while an earlier execution is still reserved or has already produced output.

Use the stage and transport labels to find where the refusal happened, then search the matching replay rejection application log. A rise after a rollout or concentration on one pod deserves investigation.

owner_busy and owner_unavailable are different outcomes. A new turn arriving while the owner is handling another turn is not automatically a duplicate, and a preflight refusal without a matched turn is not necessarily counted here.

Database connection failures, timeouts or cancelled statements during authentication and request preparation can return 503 service_unavailable before upstream dispatch. A WebSocket turn can receive an error event while the connection stays open.

Look for these log messages:

  • runtime request refused before admission stage=authentication
  • runtime request refused before dispatch stage=…

The bounded stage and reason_class help locate the boundary. The refused request may have no request row because persistence was unavailable.

Use HTTP status metrics for HTTP refusals; an error inside an already-upgraded WebSocket is not a new HTTP 503 response. For that case, inspect the application logs and WebSocket diagnostics.

A burst across replicas points to a shared dependency or resource problem. Check PostgreSQL availability, connection-pool pressure, query cancellation and deployment changes before concluding that storage or the database server is the cause.

codex_pooler_gateway_routing_affinity_stale_write_count records routing-affinity updates rejected because their event timestamp is older than the stored value.

The bounded labels are operation (success_upsert or miss_update) and affinity_kind (codex_session, idempotency_key or request_correlation); invalid values become unknown.

Some reordering is expected when overlapping work finishes on different nodes. The refused update loses an ordering hint; it does not itself fail the completed turn or change its accounting.

If the signal is sustained or concentrated on one replica, compare node clocks and overlapping request timing. The pod label identifies the writer whose update was refused, not necessarily the node that caused the timestamp mismatch. Clock skew is a hypothesis to verify, not a conclusion from this metric alone.

codex_pooler_gateway_native_compaction_admission_clear_count records the end of an admission for native in-band compaction. A clear is not automatically an error.

Label combination Interpretation
reason="connection_closed", stage="armed" With topology="direct", a connection ended while a possible compaction was armed; with topology="forwarded", the provider closed the upstream connection (a closing socket there ends the admission as downstream_detached)
reason="connection_closed" while compacting or finalizing The upstream connection closed under the admission: a compaction already collected is still delivered and its final runs as an ordinary turn
Other clears while compacting or finalizing Investigate an incomplete compaction using its reason and logs

A completed in-band compaction ends no admission: the final request’s own success arms the next one, so no clear counts completed compactions. final_success belongs to the reason vocabulary, but no serving path clears with it.

topology distinguishes direct from forwarded. Reasons such as compact_failure and final_failure identify unsuccessful exchanges; owner lifecycle, client departure and explicit request rejection have their own reasons.

These events come from serving application nodes and have no worker/scheduler reporter gap. A clear when no admission existed is not counted.

codex_pooler_accounting_reservation_pre_attempt_release_count counts reservations released before an attempt row exists. Read phase before treating a release as a failure:

Phase What to check
routing_rejected A deliberate refusal before dispatch; inspect the routing reason
turn_interrupted Client departure or owner lifecycle; inspect the release reason
task_exception A process raised before creating an attempt; investigate the pre-dispatch path
stale_sweep Cleanup found a reservation abandoned beyond the six-hour threshold
unrecorded The releasing path did not declare its boundary; a rising count needs investigation

The cleanup job runs every 15 minutes; six hours is the staleness threshold, not its schedule.

On split OBAN_MODE=worker/scheduler deployments, the job-side share can reach the reporter as via="job_relay". The starter dashboard selects only via="in_process". A flat direct-series graph does not exclude stale-sweep activity.

The release ledger’s pre_attempt_phase detail is the durable evidence. A release after a dispatched or terminal attempt belongs to a different lifecycle even if no later retry started.

codex_pooler_saved_reset_convergence_count records observed committed outcomes after a saved reset is consumed: confirmed_by_quota, reblocked, expired or unknown.

Three histograms describe different parts of the observed timeline:

Histogram suffix Interval
applied_to_canonical_seconds Reset application to accepted quota evidence
canonical_to_lifecycle_seconds Accepted evidence to lifecycle completion
applied_to_lifecycle_seconds Reset application to lifecycle completion

These are observations, not a timer for database replication or proof that every transition was measured. Missing or invalid timestamps produce no latency sample.

Under split OBAN_MODE, some transitions are emitted by worker jobs and reach the app reporter only through the best-effort relay. The direct panels can therefore undercount. Current account lifecycle metadata describes the latest state; it is overwritten as the lifecycle changes and cannot reconstruct every historical transition.

Use the metric to locate a time window, then inspect the account’s saved-reset and quota state in Upstreams. Do not infer duplicate reset consumption from an empty or delayed panel.

Record the UTC window, affected role/pod, application release, relevant error or reason codes, and whether the client used HTTP/SSE or WebSocket. Include known request/correlation IDs when available.

Keep client symptoms, edge responses, gateway events and provider results separate. Share only metadata; do not attach raw conversation bodies, credentials or frames.