PromQL recipes
Use these queries after metric collection is working. Each code block contains one expression that you can paste into Prometheus or a Grafana panel.
Replace namespace="example-app" and job="codex-pooler-app" with your target labels. Examples assume one cluster; if your data source combines clusters, select one cluster consistently before joining or aggregating pod-level series.
Application metrics come from Codex Pooler. Container metrics require kubelet/cAdvisor, and kube_* metrics require kube-state-metrics.
Traffic and latency
Section titled “Traffic and latency”HTTP request rate
Section titled “HTTP request rate”Requests per second by method and HTTP status class. This describes HTTP responses, not error frames within an already-open WebSocket.
sum by (method, status_class) ( rate(codex_pooler_http_request_count{namespace="example-app", job="codex-pooler-app"}[5m]))Endpoint latency p95
Section titled “Endpoint latency p95”Seconds at the 95th percentile for application HTTP requests. Long-lived or streamed requests can have different duration patterns from short endpoints.
histogram_quantile(0.95, sum by (le) ( rate(phoenix_endpoint_stop_duration_seconds_bucket{namespace="example-app", job="codex-pooler-app"}[5m]) ))Stream outcomes
Section titled “Stream outcomes”The direct request-path share, split by outcome and both transport directions. Under split OBAN_MODE, job-recovered interruptions are not included in this direct series; see coverage.
sum by (outcome, downstream_transport, upstream_transport) ( rate(codex_pooler_gateway_stream_outcome_count{namespace="example-app", job="codex-pooler-app", via="in_process"}[5m]))Gateway pressure
Section titled “Gateway pressure”Running requests
Section titled “Running requests”Current admitted work per pod and route class. This is a gauge, so read the value directly rather than applying rate().
codex_pooler_gateway_admission_running{namespace="example-app", job="codex-pooler-app"}Queued requests
Section titled “Queued requests”Current local queue depth. Compare pods individually before interpreting a cluster aggregate.
codex_pooler_gateway_admission_queued{namespace="example-app", job="codex-pooler-app"}Newly queued requests
Section titled “Newly queued requests”Requests per second entering the local admission queue.
sum by (route_class, transport) ( rate(codex_pooler_gateway_admission_enqueued_count{namespace="example-app", job="codex-pooler-app"}[5m]))Admission wait p95
Section titled “Admission wait p95”Seconds spent waiting by requests that were eventually dequeued. Timed-out requests have a separate histogram and are not part of this percentile.
histogram_quantile(0.95, sum by (le, route_class, transport) ( rate(codex_pooler_gateway_admission_dequeued_time_seconds_bucket{namespace="example-app", job="codex-pooler-app"}[5m]) ))Websocket control failures
Section titled “Websocket control failures”A nonzero increase is a starting point for investigation, not proof of a lost connection. For an alert, consider a one-minute hold and tune it to your workload; cleanup_deferred includes supervised work continuing after the synchronous budget.
sum by (phase, reason) ( increase(codex_pooler_gateway_websocket_control_failure_count{namespace="example-app", job="codex-pooler-app"}[5m])) > 0Bridge fallbacks
Section titled “Bridge fallbacks”Fallback events by bounded reason. Interpret these with client and upstream transport evidence.
sum by (reason) ( rate(codex_pooler_gateway_websocket_bridge_fallback_count{namespace="example-app", job="codex-pooler-app"}[5m]))Bridge precommit overflow
Section titled “Bridge precommit overflow”A separate counter for overflowing the bridge’s precommit buffer.
sum( rate(codex_pooler_gateway_websocket_bridge_precommit_overflow_count{namespace="example-app", job="codex-pooler-app"}[5m]))Stream buffer releases
Section titled “Stream buffer releases”Oversized incomplete buffers released without retaining them, grouped by transport and route class.
sum by (transport, route_class) ( rate(codex_pooler_gateway_stream_buffer_oversized_count{namespace="example-app", job="codex-pooler-app"}[5m]))Stream buffer truncations
Section titled “Stream buffer truncations”Retained stream bodies reduced to a bounded suffix. This is distinct from the oversized-buffer counter.
sum by (transport, route_class) ( rate(codex_pooler_gateway_stream_buffer_truncated_count{namespace="example-app", job="codex-pooler-app"}[5m]))Database timings
Section titled “Database timings”Repository query rate
Section titled “Repository query rate”Queries per second from each scraped application pod. Worker/scheduler queries are not exported by those roles.
sum by (pod) ( rate(codex_pooler_repo_query_count{namespace="example-app", job="codex-pooler-app"}[5m]))Most active query sources
Section titled “Most active query sources”The ten busiest bounded query-source and SQL-command labels. These labels do not contain raw SQL.
topk(10, sum by (source, command) ( rate(codex_pooler_repo_query_count{namespace="example-app", job="codex-pooler-app"}[5m]) ))Unclassified query sources
Section titled “Unclassified query sources”Queries whose source was classified as unknown. A rise means less attribution, not necessarily a failed query.
sum by (pod, command) ( rate(codex_pooler_repo_query_count{namespace="example-app", job="codex-pooler-app", source="unknown"}[5m]))Total repository time p95
Section titled “Total repository time p95”Total query time by source and command, in seconds. Compare this with connection checkout wait and database execution timing.
histogram_quantile(0.95, sum by (le, source, command) ( rate(codex_pooler_repo_query_total_time_seconds_bucket{namespace="example-app", job="codex-pooler-app"}[5m]) ))Connection checkout wait p95
Section titled “Connection checkout wait p95”Seconds waiting for a database connection, per pod. This can rise even when SQL execution remains fast.
histogram_quantile(0.95, sum by (le, pod) ( rate(codex_pooler_repo_query_queue_time_seconds_bucket{namespace="example-app", job="codex-pooler-app"}[5m]) ))Memory and restarts
Section titled “Memory and restarts”Container working set
Section titled “Container working set”Application-container working set in bytes. Replace the container selector when your deployment uses a different name.
max by (namespace, pod) ( container_memory_working_set_bytes{namespace="example-app", container="app", image!=""})BEAM total memory
Section titled “BEAM total memory”Memory reported by the Erlang runtime, in bytes.
vm_memory_total_bytes{namespace="example-app", job="codex-pooler-app"}Use the same selector with these gauges for more detail:
| Metric | Meaning |
|---|---|
vm_memory_binary_bytes |
Binary data |
vm_memory_processes_bytes |
Process memory |
vm_memory_ets_bytes |
ETS tables |
vm_memory_system_bytes |
BEAM system memory |
vm_system_counts_process_count |
Process count |
vm_system_counts_port_count |
Port count |
vm_total_run_queue_lengths_total |
Total scheduler run queue |
Container memory beyond BEAM total
Section titled “Container memory beyond BEAM total”An approximate comparison of two memory measurements, not an exact leak size. Select the same cluster, namespace and application pod on both sides.
clamp_min( max by (namespace, pod) ( container_memory_working_set_bytes{namespace="example-app", container="app", image!=""} ) - on (namespace, pod) max by (namespace, pod) ( vm_memory_total_bytes{namespace="example-app", job="codex-pooler-app"} ), 0)Recent restarts
Section titled “Recent restarts”Restart increments over 15 minutes. This example follows the dashboard’s kube-state-metrics label convention; use container instead of exported_container if that is how your collector exposes the workload container.
increase(kube_pod_container_status_restarts_total{ namespace="example-app", exported_container="app"}[15m])Routing and lifecycle signals
Section titled “Routing and lifecycle signals”See runtime triage before deciding whether these events require action.
Routing circuit changes
Section titled “Routing circuit changes”Transitions by route class and bounded reason, useful around retry or availability changes.
sum by (transition, route_class, reason_class) ( rate(codex_pooler_gateway_routing_circuit_transition_count{namespace="example-app", job="codex-pooler-app"}[5m]))Rejected stale affinity writes
Section titled “Rejected stale affinity writes”The pod label identifies the writer whose update was refused. Check timing and clocks before attributing a sustained pattern to one node.
sum by (pod, operation, affinity_kind) ( rate(codex_pooler_gateway_routing_affinity_stale_write_count{namespace="example-app", job="codex-pooler-app"}[5m]))Duplicate-turn refusals
Section titled “Duplicate-turn refusals”Refusals that do not create a separate request row for the duplicate.
sum by (stage, transport) ( rate(codex_pooler_gateway_duplicate_turn_refused_count{namespace="example-app", job="codex-pooler-app"}[5m]))Native compaction admission clears
Section titled “Native compaction admission clears”Count normal completion and interruption separately using reason, stage and topology.
sum by (reason, stage, topology) ( rate(codex_pooler_gateway_native_compaction_admission_clear_count{namespace="example-app", job="codex-pooler-app"}[5m]))Pre-attempt releases
Section titled “Pre-attempt releases”Direct observations only. Under split OBAN_MODE, stale-sweep job activity can be absent here; the release ledger is the durable evidence.
sum by (phase, transport) ( rate(codex_pooler_accounting_reservation_pre_attempt_release_count{namespace="example-app", job="codex-pooler-app", via="in_process"}[5m]))Saved-reset outcomes
Section titled “Saved-reset outcomes”Direct observations only. Under split OBAN_MODE, job-side transitions are partial and relayed separately.
sum by (source, outcome) ( rate(codex_pooler_saved_reset_convergence_count{namespace="example-app", job="codex-pooler-app", via="in_process"}[5m]))Reset application to accepted evidence p95
Section titled “Reset application to accepted evidence p95”Seconds until canonical quota evidence was observed. Under split OBAN_MODE this query includes only the direct share. Missing timing observations do not mean zero delay.
histogram_quantile(0.95, sum by (le, source, outcome) ( rate(codex_pooler_saved_reset_convergence_applied_to_canonical_seconds_bucket{namespace="example-app", job="codex-pooler-app", via="in_process"}[5m]) ))Accepted evidence to lifecycle completion p95
Section titled “Accepted evidence to lifecycle completion p95”Seconds from canonical evidence to the observed lifecycle finish. Under split OBAN_MODE this query includes only the direct share. Missing timing observations do not mean zero delay.
histogram_quantile(0.95, sum by (le, source, outcome) ( rate(codex_pooler_saved_reset_convergence_canonical_to_lifecycle_seconds_bucket{namespace="example-app", job="codex-pooler-app", via="in_process"}[5m]) ))Reset application to lifecycle completion p95
Section titled “Reset application to lifecycle completion p95”Seconds across the complete observed interval. Under split OBAN_MODE this query includes only the direct share. Missing timing observations do not mean zero delay.
histogram_quantile(0.95, sum by (le, source, outcome) ( rate(codex_pooler_saved_reset_convergence_applied_to_lifecycle_seconds_bucket{namespace="example-app", job="codex-pooler-app", via="in_process"}[5m]) ))Telemetry relay
Section titled “Telemetry relay”These are shared database observations repeated by application reporters. Use max across observers; summing would multiply a shared value by the number of reporting pods.
Unclaimed relay backlog
Section titled “Unclaimed relay backlog”Unclaimed samples, including expired backlog until cleanup removes it. A sustained rise can indicate that consumers are not keeping up.
max(codex_pooler_telemetry_relay_backlog_samples{namespace="example-app", job="codex-pooler-app"})Fresh relay consumers
Section titled “Fresh relay consumers”Consumers reporting within the freshness window. Zero warrants investigation; it does not stop producers from inserting events.
max(codex_pooler_telemetry_relay_consumers_fresh{namespace="example-app", job="codex-pooler-app"})Known lost relay samples
Section titled “Known lost relay samples”Durable cumulative loss by reason. These are gauges of shared totals, not per-pod counters. Hard-kill and post-claim/pre-scrape losses are not fully measured.
max by (reason) ( codex_pooler_telemetry_relay_loss_samples{namespace="example-app", job="codex-pooler-app"})Reading the result
Section titled “Reading the result”A five-minute rate needs enough samples in that window. Missing event series and empty quantiles are not measured zeros. Keep the le label when aggregating classic histogram buckets before histogram_quantile.
Do not combine direct and relayed shares accidentally: keep the explicit via="in_process" selector for the partial families above, or select via="job_relay" deliberately for a separate diagnostic view. Neither a graph nor its absence replaces request/accounting evidence.