Guide

Route Claude Code and Codex through the proxy

Point Claude Code, Codex, or another coding agent at gobstopper proxy, choose its settings, and fix connection, routing, and service problems.

gobstopper proxy sits on 127.0.0.1 between a coding agent and its model provider. When a request passes the threshold, it sends the head verbatim, one mechanical summary of the older turns, and the last three turns verbatim; a positive --keep-tail-percent can keep older whole turns too. The provider then reports the smaller size back to the client, which can delay the client’s own auto-compaction trigger. Session files are not changed. The summary rule is CliffCompaction's (Nguyen, Cho, Chen and Dettmers, arXiv:2609.26779); the port's MIT notice is in THIRD_PARTY_NOTICES.md, and How it works describes where the proxy departs from it.

The proxy understands three API dialects. Any agent that lets you set a custom provider address can use it:

DialectEndpointAgents
Anthropic Messages.../messagesClaude Code, opencode, Crush
OpenAI Responses.../responsesCodex
OpenAI Chat Completions.../chat/completionsopencode, Crush, Aider, Goose, other OpenAI-compatible clients

The proxy requires system curl 8.3 or later. Check it with curl --version.

What you get

  • A smaller request each turn. Past the threshold, older file reads and command output are summarized. Selected original observations can survive in a size-limited evidence carry; see context retention. The system prompt, the first task, and the newest turns stay word for word, so the agent keeps what it was just working on.
  • No model call to compact. The summary is built by a fixed rule, so a compaction costs no extra request and adds no model-written text.
  • No summary of a summary. Each compaction starts from the original history the client resends, so detail is lost once, not again at every compaction.
  • A stable prompt cache. Between compactions, every request carries the same compacted prefix byte for byte while the context policy and calibration stay the same, so the provider's prompt cache can keep matching.
  • Fewer native compaction triggers. The provider reports the compacted size, which can keep the client below its trigger. Client limits still apply, and its transcript keeps the full history.
  • Optional compaction can be skipped. If a rewrite fails or times out, the proxy can send the client's original bytes. Explicit context limits, strict sizing and scoped policy still apply; see Limits.

CliffCompaction's authors report up to 50% lower cost with a capped context, with Terminal-Bench 2.0 scores held or improved, on the Kimi K2.6 and GLM 5.1 models they tested. In one run through Claude Code, their proxy scored above Claude Code's own auto-compaction. Those are their measurements of their proxy, computed with a model of perfect prompt caching; they report that the benefit depends on the agent and the task and matters only for medium-to-long tasks.

Gobstopper's own run, on September 27 and 28, 2026, put the 89 tasks of Terminal-Bench 2.1 through Claude Code 2.1.283 with GLM 5.3 Flash via Vercel AI Gateway, one trial per arm, at a 45,000-token threshold (the default is 128,000). At tail 0, the proxy solved 61 tasks against 60 with no proxy, within single-trial noise, and sent 29% fewer provider-reported input tokens (84.3 million against 118.6 million). Its provider-reported cost for that model was 16% lower, with a 95% interval from 32% lower to 2% higher, so one run does not establish a saving. The benchmarks page has the full study; replays and estimates are in the README.

Try it on one session

gobstopper proxy run -- claude

run starts the proxy on a free port, sets ANTHROPIC_BASE_URL and OPENAI_BASE_URL for that command only, and stops when the command exits. Codex takes its model provider from ~/.codex/config.toml or -c overrides, so route it with the provider block in Codex.

Run it in the background

gobstopper proxy serve        # http://127.0.0.1:8260; --port changes it

proxy install creates an owned user service on macOS, Linux or Windows:

gobstopper proxy install
gobstopper proxy status
gobstopper proxy doctor
gobstopper proxy repair

The service manager restarts failed processes. Installation verifies service identity and readiness, and replacement preserves active inference. --print previews the definition. Use proxy install --replace to update owned settings; proxy repair diagnoses and reconciles the installed service. Existing manually written definitions are not overwritten without an exact supported migration. See startup, recovery, and sleep behavior for platform prerequisites, legacy Mac migration, logs and rollback.

For a difficult phase, use temporary context budgets. For local usage, throughput, tool activity and portable exports, see session data.

Claude Code

Add this to your shell profile, then start Claude Code from a new terminal:

export ANTHROPIC_BASE_URL=http://127.0.0.1:8260

Claude Code sends your claude.ai sign-in or API key through the proxy unchanged. If ANTHROPIC_API_KEY is set in your environment, Claude Code uses that key instead of your claude.ai sign-in, with or without the proxy. For this shell-only setup, env -u ANTHROPIC_BASE_URL claude runs one session without the proxy. A URL saved in Claude settings can reenable the proxy. Use gobstopper proxy launch --client claude to check readiness before starting a session and choose direct fallback when the configuration permits it; see launch requirements.

Codex

Codex reads its provider from ~/.codex/config.toml. With ChatGPT sign-in:

model_provider = "gobstopper"    # top-level key: above the first [table]

[model_providers.gobstopper]
name = "OpenAI via gobstopper"
base_url = "http://127.0.0.1:8260/backend-api/codex"
wire_api = "responses"
requires_openai_auth = true

With an OpenAI API key, use base_url = "http://127.0.0.1:8260/v1" and env_key = "OPENAI_API_KEY" instead of requires_openai_auth. To run one session without the proxy, use codex -c model_provider=openai.

opencode

opencode providers take an options.baseURL. For an Anthropic-shaped provider, point it at the proxy root in opencode.json:

{
  "provider": {
    "gobstopper-anthropic": {
      "npm": "@ai-sdk/anthropic",
      "options": { "baseURL": "http://127.0.0.1:8260/v1" },
      "models": { "claude-sonnet-4-5": {} }
    }
  }
}

For an OpenAI-compatible provider, use the Chat Completions dialect:

{
  "provider": {
    "gobstopper-openai": {
      "npm": "@ai-sdk/openai-compatible",
      "options": { "baseURL": "http://127.0.0.1:8260/v1" },
      "models": { "gpt-5": {} }
    }
  }
}

API keys come from opencode's own provider credentials; the proxy forwards them unchanged.

Crush

Crush providers accept a base_url and a type. In crush.json:

{
  "providers": {
    "gobstopper": {
      "type": "openai-compat",
      "base_url": "http://127.0.0.1:8260/v1",
      "models": [{ "id": "gpt-5", "name": "GPT-5" }]
    }
  }
}

A "type": "anthropic" provider with base_url pointed at the proxy uses the Messages dialect instead.

Aider

Aider routes OpenAI-shaped models through a configurable base URL:

aider --openai-api-base http://127.0.0.1:8260/v1 --model openai/<model>
# or: OPENAI_API_BASE=http://127.0.0.1:8260/v1

Goose

Goose's OpenAI provider reads OPENAI_HOST (and OPENAI_BASE_PATH):

OPENAI_HOST=http://127.0.0.1:8260 goose

A declarative custom provider with engine: openai pointed at http://127.0.0.1:8260/v1 works the same way.

Any OpenAI-compatible client that posts to {base}/chat/completions works: give it a base URL of http://127.0.0.1:8260/v1. The bare path /chat/completions (no /v1) also routes to the OpenAI upstream, but most providers expect the /v1 prefix, so configure the base URL with it.

Preview on a recorded session

gobstopper proxy replay <session>
gobstopper proxy replay <session> --threshold 100000 --json

replay rebuilds the requests a Claude Code or Codex session sent, runs each through the engine, and reports the peak request with and without the proxy, the number of compactions and reused prefixes, the context sent across all requests, and any history left with an unpaired tool call. It calls no provider. Transcripts do not record the system prompt and tool definitions, so --fixed-tokens (20,000 by default) stands in for them. It also reads a Claude Code subagent's own transcript. --json adds the estimated cache reads and writes (est_cache_read_tokens, est_cache_write_tokens), repeated reads and how many the kept history still held (repeated_reads, repeated_reads_covered), and the spacing of compactions (back_to_back_compactions, min_compaction_gap), and each compaction lists the characters its summary carried from the turns earlier compactions summarized (carry_chars). From a Claude Code transcript, replay also reads the provider-reported input and calibrates as the proxy does; see Estimate calibration. From a Claude Code transcript, replay rebuilds only the user and assistant records, so messages typed while the agent worked are missing from its carried text. replay has no --threshold-1m; to model a session that declares a 1M-token window, pass --threshold 256000.

Settings

These settings apply to proxy serve, proxy run, and proxy install, except where noted.

FlagDefaultMeaning
--threshold128000Compact when the estimated outgoing request exceeds this many tokens. Applies to every OpenAI-dialect request and to Anthropic requests that do not declare a 1M-token window; those that do use --threshold-1m. Keep it below the client's own auto-compaction point.
--threshold-1m256000, or --threshold if higherThe threshold for Anthropic Messages requests whose anthropic-beta header lists a token starting with context-1m. It can't be lower than --threshold; an equal value applies one threshold to every request. Keep it below the client's own auto-compaction point, including any claude --autocompact value.
--context-windowunsetDeclare the context capacity supported by the upstream route. Output headroom reduces the input capacity; original-body fallbacks and retries must also fit. This does not increase the provider's supported window.
--client-context-windowunsetrun scope only. Declare the client's capacity when it is lower than the provider's. See context budgets.
--output-reserve32000Tokens reserved for output when creating a run scope. A request asking for more output reserves more room.
--adaptive-contextoffrun scope only. Allow a temporary context increase after repeated reads of unchanged evidence the proxy removed, within configured capacity.
--transform-timeout-ms5000Deadline for optional compaction work, from 1 to 60000 milliseconds. A timeout sends the original request only when the context and strict-sizing policies permit it.
--keep-recent3Newest assistant steps kept verbatim. A request still over the threshold after one pass keeps one.
--keep-tail-percent0Share of the room under the threshold, after the fixed request fields and the head, that the summary and the kept steps may fill. Older whole steps are kept while they fit. From 0 to 60; 0 keeps exactly --keep-recent steps, as CliffCompaction does. Use a positive value to keep additional older turns.
--result-max-chars500Older tool results longer than this are dropped from the summary; shorter ones stay verbatim.
--carry-max-chars24000Characters of the human's words and the assistant's visible replies that each summary carries forward from the turns earlier compactions summarized. The oldest text drops out first, and the carried text takes at most a quarter of the room under the threshold after the fixed request fields and the head. 0 turns carrying off.
--evidence-max-bytes262144Serialized UTF-8 bytes of selected original tool evidence kept across summaries. 0 disables evidence retention. Also a replay flag.
--evidence-max-chars32000Billable characters of retained evidence, including image cost, also subject to byte and context limits. Also a replay flag. See evidence retention.
--drop-thinkingoffLeave thinking and reasoning text out of summaries.
--no-calibrateoffCompare the plain four-characters-per-token estimate with the threshold. By default the threshold is divided by the ratio of provider-reported to estimated input, learned per upstream and model (see Estimate calibration). Also a replay flag.
--shadowoffLog what would change and forward requests unchanged. Explicit capacity and scoped policy still apply.
--strictoffRefuse (HTTP 400) a request still over the threshold after every step, instead of sending it. Outside shadow mode, sizing must succeed before forwarding.
--no-keep-awakeoffDisable idle-sleep prevention during inference. Display sleep is always allowed; see sleep behavior.
--no-session-dataoffDisable the local session observation journal. This is separate from the legacy JSONL statistics ledger.
--anthropic-upstreamhttps://api.anthropic.comWhere Anthropic requests go.
--openai-upstreamhttps://api.openai.comWhere OpenAI API requests (/v1/..., .../chat/completions) go. Point it at any OpenAI-compatible provider.
--chatgpt-upstreamhttps://chatgpt.comWhere ChatGPT-signed-in Codex requests (/backend-api/...) go.

How it works

Claude Code → local Gobstopper proxy → model provider.

The proxy changes outgoing history. Optional rewrite failures can send the original bytes when policy permits. A provider HTTP 400 rejection can trigger another trim for a length error, or an original-body retry for another error when the original fits the configured capacity.

Diagram: a full-width bar labelled original request, and below it a shorter bar of six parts: head, summary, carry, and the last three turns.

Inside a rewritten request: the head, one summary, the carried words, and the last three turns. In the logged part of the tail-0 Terminal-Bench arm (about 68 of the 89 trials), compacted requests had a median of 31.5K estimated tokens, against a median of 55K before compaction. The example lines are illustrative.

  • The proxy compacts Anthropic Messages (.../messages), OpenAI Responses (.../responses), and OpenAI Chat Completions (.../chat/completions) requests. Token counts, provider-side compaction endpoints, and every other path pass through unchanged.

  • The head is everything before the first model turn: for Claude Code, the first user message; for Codex, the environment and instruction messages and the first prompt; for Chat Completions, the system, developer, and user messages that precede the first assistant turn. It is always sent verbatim, as are the system prompt, tool definitions, and other request fields.

  • A turn starts at a model message and includes any model messages right after it and the tool results that answer them: tool_result blocks for Anthropic, function_call_output items for Responses, and the run of tool messages answering a tool_calls turn for Chat Completions. The summary keeps human and assistant text (and readable thinking), keeps tool results of at most 500 characters, reduces each tool call to its name and up to 150 characters of arguments. A size-limited evidence carry also retains selected original tool results and supported images, with invocation provenance and explicit excerpt labels; see context retention. Calls are never separated from their results: the kept tail is whole turns, so a tool message can never outlive the call it answers. CliffCompaction starts a turn at every Anthropic or Chat Completions assistant message; the proxy keeps a run of them together, because Claude Code can record one step as two assistant messages (the tool calls, then the text).

  • The kept tail starts with the newest --keep-recent turns and grows one older whole turn at a time while the summary and the tail fit in --keep-tail-percent of the room under the threshold, the threshold minus the system prompt, tool definitions, and head. The kept turns hold the files and command output the agent read most recently, which the summary omits once they pass 500 characters unless evidence carry selects them. At the default, 0, the tail is exactly the newest --keep-recent turns. At 40%, a compacted request leaves 60% of that room for new turns before the next compaction. A higher floor leaves less room, so we expect the proxy to compact more often and every request between compactions to be larger. Replay agrees in direction: over 24 recorded sessions, tail 40 compacted 369 times against 343 at 32,000 tokens and 50 against 38 at 128,000 (estimates). In the Terminal-Bench run at a 45,000-token threshold, tail 40 sent 118.5 million input tokens against 84.3 million and cost 39% more in total, with solved counts within noise.

    Diagram: two stacks after a rewrite. Tail 0 keeps your task, a summary and the last three turns. Tail 40 also keeps older turns, so it sits closer to the threshold and is rewritten again sooner.

    Keeping more old turns meant more rewrites and more tokens. In the benchmark, tail 40 cost 39% more than tail 0 in provider-reported terms (95% interval 2% to 87% more), at a 45,000-token threshold. v0.7.3 makes tail 0 the default. Terminal-Bench 2.1 · 89 tasks · one trial per arm · Gobstopper v0.7.2 · September 27–28, 2026 · 21 of 89 tail-0 trials may have run an earlier build

    If the summary, with its carried text, and the newest --keep-recent turns need more, the proxy keeps those turns anyway, and the next request can compact again.

  • Each compaction discards the previous summary, as CliffCompaction does, but the conversation's words carry forward. The proxy keeps the human's messages, including those typed while the agent works, after an interrupt, or when rejecting a tool call, and the assistant's visible replies from every summarized turn, and each later summary opens with them, oldest first. Each carried part, one message's text or one queued message, is capped at 4,000 characters. When the carried text passes --carry-max-chars (24,000 by default) or a quarter of the room under the threshold, the oldest parts drop out first. Tool calls, other tool results, thinking, skill instructions, shell and local-command output, system reminders, and task notifications are never carried. A message queued while Claude Code works is read from the system message Claude Code sends it in, or from a human turn's text block that is wholly that message, never from a tool result, which can quote the same words. In the Responses and Chat Completions dialects every user-role message is carried, except the context items Codex sends again (roadmap). The first threshold crossing has nothing to carry. A proxy that starts on a long history, after a restart or when its cache dropped the entry, replays every crossing and rebuilds the carry, so its first logged compaction can show a nonzero carry N chars. --carry-max-chars 0 restores the reference rule for every compaction. When the fixed request fields and the head take more than half the threshold, the carried text brings some compactions one request sooner, which added 2% to the estimated cost of the one replayed session of that kind (design).

  • Anthropic Messages requests whose anthropic-beta header lists a context-1m token use --threshold-1m. Claude Code sends that token for a model such as opus[1m]. The proxy never reads the window from the model name, so a request that does not declare the window uses --threshold, even on a model whose default window is larger. gobstopper proxy status counts the Anthropic requests that declared a 1M-token window (requests_1m), and each compaction log line names the window it applied (window=1m or window=base).

  • Prefix reuse. Clients resend their original history on every request. The proxy keys each compaction by a hash of the original prefix and substitutes it into later requests with the same context policy and calibration, so the compacted prefix stays byte-stable until the next compaction and the provider's prompt cache can match it. A policy or calibration change rebuilds the prefix from the original history. The cache lives in memory; after a restart, the proxy replays the threshold crossings over the full history and reaches the same result. The carried text is the exception in two cases, until newer words fill it again: after the proxy shortened a summary to fit, which starts the carry over, and when the carry's quarter-of-the-room bound grew, which can happen only when --carry-max-chars is at least half the threshold, rounded down (at the default, at a threshold of 48,001 tokens or less).

  • Retry ladder. If one pass leaves a request over the threshold, or the provider rejects the rewrite for length, the proxy tries harsher settings in order:

    1. after a length rejection of a request not yet compacted, a compaction at tail 0, the newest --keep-recent turns and no tail extension;
    2. one kept turn;
    3. assistant text capped at 300 characters and thinking dropped;
    4. after a length rejection, or with --strict, the summary shortened to its newest parts.

    Without --strict, a request still over the threshold is then sent if it fits any configured hard input capacity; with strict mode, the proxy refuses it (HTTP 400). If the provider rejects the rewritten request with HTTP 400 for a reason other than length, the proxy resends the client's original bytes only when they fit the configured capacity.

  • The threshold is calibrated to the provider's count; see Estimate calibration.

  • If the verbatim head alone approaches the threshold, as it can after the client compacted a session itself or when a subagent starts with a long prompt, the threshold for that session rises to the head plus half the request's threshold. A request over the configured threshold but under the raised one is sent unchanged, and the log says so (sent unchanged ... under the threshold raised to ~Nk by a large verbatim head). A request over the threshold with nothing to compact, such as one with too few turns, is also sent unchanged and logged (... with nothing to compact). An explicit hard context capacity limits this allowance. If the preserved head and fixed fields already exceed it, the proxy refuses the request.

Where it departs from CliffCompaction

The proxy ports CliffCompaction's summary rule, prefix reuse, and retry on a length rejection. The options below control its departures from that rule.

#DepartureDefaultRestore the reference
1Keep older whole turns beyond the last three within a tail budgetoff (tail 0)--keep-tail-percent 0, the default
2Count a run of assistant messages as one turn in every dialectonnone; keeps tool calls paired with their results
3Separate threshold for Anthropic requests that declare a 1M-token windowon--threshold-1m equal to --threshold
4Resend the original after an HTTP 400 rejection for a reason other than length, when it fits configured capacityonnone
5Raise the threshold when the verbatim head alone approaches it, within configured capacityonnone
6Carry human words and assistant replies from summarized turns, up to 24,000 characterson--carry-max-chars 0
7Calibrate the threshold from provider-reported input, 1.0 to 2.0on--no-calibrate

Choosing a threshold

The default threshold, 128,000 estimated tokens, sits below the point where a 200,000-token client compacts on its own. A lower threshold compacts sooner and cuts more: replaying 24 recorded sessions (12 Claude Code, 12 Codex) at tail 0 cut cumulative estimated input by 78% at 32K, 73% at 64K, 61% at 128K, and 38% at 256K. Three large sessions dominate those pooled figures; a typical session's cut at 32K is about 46%, and at 128K most Claude Code sessions never cross the threshold. The Terminal-Bench run used 45,000. Try gobstopper proxy replay <session> --threshold N on your own sessions before you lower it.

Keep the threshold below the point where the client compacts on its own. Claude Code's environment variables can move that point or turn it off, as Anthropic's reference described them on October 4, 2026. CLAUDE_AUTOCOMPACT_PCT_OVERRIDE sets “the percentage (1-100) of the auto-compact window at which auto-compaction triggers”, and “the variable can't raise the threshold”, so after you set it, keep --threshold below the earlier point. DISABLE_AUTO_COMPACT=1 disables “automatic compaction when approaching the context limit”, and “the manual /compact command remains available.”

Line chart: estimated input cut across 24 recorded sessions. Tail 0 cuts 78% at 32K, 73% at 64K, 61% at 128K and 38% at 256K. Tail 40 is a little lower at every threshold.

Lower thresholds cut more. Estimates, not billed · 24 recorded sessions (12 Claude Code, 12 Codex), 665M tokens · main fdeb099 · September 26, 2026

Estimate calibration

The proxy estimates a request at four characters per token. Claude models count more: the ratio depends on the model and the request. Calibration uses reported usage to adjust later compaction thresholds when the byte estimate runs low.

  • After relaying a response, the proxy reads its usage: for Anthropic Messages, input_tokens plus cache_creation_input_tokens and cache_read_input_tokens, from a JSON body or a stream's message_start event; for OpenAI Responses, input_tokens from a JSON body or a stream's last event carrying response.usage (normally response.completed); for OpenAI Chat Completions, prompt_tokens from a JSON body or a stream's last chunk carrying usage (sent when the client asks for usage in the stream). The OpenAI counts already include cached input. It reads a copy of the body after each chunk reached the client, so the response is neither changed nor delayed. A missing, malformed, or oversized usage record (a JSON body over 4 MiB, no message_start in the first 64 KiB of an Anthropic stream, or no usage event in the last 64 KiB of an OpenAI stream) is skipped.
  • Each response adds one sample: reported input divided by the proxy's estimate of the request it forwarded. Requests estimated under 1,000 tokens and samples outside 0.25 to 4.0 are skipped. The proxy keeps a running ratio per upstream and model (each sample moves it one eighth of the way), for up to 64 pairs, in memory only.
  • After five samples, a request's threshold is divided by that ratio, limited to 1.0 to 2.0: at a ratio of 1.25, a 128,000-token threshold compacts at 102,400 estimated tokens. The limit means calibration can only compact earlier, never later, and never below half the threshold. Stored compactions include the applied ratio and context policy in their identity. A changed ratio or policy rebuilds the prefix from the original history.
  • gobstopper proxy status shows the ratio applied and measured, and the sample count, for each upstream and model (calibrate, calibrations). Compaction log lines name a ratio other than 1.0 after the window (window=base, ratio 1.25), the ledger records ratio_permille (1250 for 1.25), and the proxy logs a line when the applied ratio first departs from 1.0 or one sample moves it by 0.05 or more.
  • --no-calibrate restores the uncalibrated threshold: every request is sized and compacted exactly as before calibration existed.

gobstopper proxy replay calibrates the same way from the usage a Claude Code transcript records, unless given --no-calibrate. It assumes the recorded session did not run behind the proxy: behind the proxy, the recorded usage describes the compacted request the proxy sent, not the history the transcript holds. The ratio also depends on --fixed-tokens. --json adds usage_requests, calibration_samples, last_ratio_permille, the lowest, median, and highest reported ratio (reported_ratio_min_permille, reported_ratio_median_permille, reported_ratio_max_permille), the largest request sent in reported tokens (peak_reported_tokens_out), the requests over the threshold in reported tokens (reported_over_threshold), and each compaction's reported_tokens_before.

Privacy and security

  • The proxy binds 127.0.0.1 and refuses requests addressed to any host name other than 127.0.0.1, localhost, or ::1.
  • It forwards your request headers, including API keys and sign-in tokens, to curl through curl's environment, not its command line.
  • Logs contain paths, sizes, counts, and error summaries, never request or response text. The status page names each calibrated upstream and model. Carried text stays in the proxy's memory and in the requests it forwards; log lines and the ledger below record only its size.
  • Prepared requests queue JSONL statistics records (timestamp, dialect, path, estimated tokens in and out, the estimated head, summary, and tail sizes, the carried characters, the window, the threshold applied to that request, the calibration ratio, and flags) to ~/.local/share/gobstopper/proxy-stats.jsonl. A background worker appends them, and gobstopper proxy status reports estimated-token totals for this run and loaded history. Queue overflow, write failures, or incomplete history can leave totals partial; inspect stats_persistence and see files and logs. GOBSTOPPER_STATS_FILE overrides the path; set it to off to disable the ledger.
  • Unparseable or compressed request bodies can pass unchanged when no explicit capacity or strict-sizing requirement needs to verify them.

Limits

  • Explicit context capacity and strict sizing can prevent original-body fallback. If required sizing times out or its workers are occupied, the proxy returns HTTP 503 with Retry-After; a payload it cannot size returns HTTP 400. Unavailable scoped context storage also returns a retryable 503 before inference is sent; an invalid scope returns 400. See recovery and fallback for launch-time choices and existing-session limits.
  • Sizes are estimates at four characters per token, with images priced by their dimensions; the provider's count can differ. Calibration corrects the threshold only after five responses, so the first requests of each upstream and model use the plain estimate, as does any response that reports no usage.
  • A summary drops the details of long tool results. The agent can read the file or rerun the command, but nothing makes it notice that a detail is missing.
  • Agents that send requests through a vendor service with no configurable model address cannot use the proxy.
  • Live use through the proxy has been checked for routing with Claude Code 2.1.282 and Codex 0.156.1. Chat Completions coverage is tested against synthetic histories and recorded contracts, not a live opencode, Crush, Aider, or Goose session; provider acceptance of that dialect is unqualified.
  • Task results under the proxy come from one Terminal-Bench 2.1 run: one trial per arm, one model (GLM 5.3 Flash), one host, and a 45,000-token threshold. It did not test an Anthropic model, other agents, or the default threshold, and its dollar figures are provider-reported prices for that model through Vercel AI Gateway, not a general bill.

Troubleshooting

Claude Code and Codex behavior in this table follows their documentation as checked on October 4, 2026: Claude Code's environment variables, gateway connection and gateway rollout guides, and Codex's advanced configuration. For the service itself, see diagnose, repair, and remove.

SymptomWhat to check
Claude Code shows no output for several minutes, then reports a connection errorNothing is answering at ANTHROPIC_BASE_URL, and Claude Code retries an unreachable base URL with backoff before it reports the error. gobstopper proxy status prints No proxy is running on 127.0.0.1:8260 when the proxy is stopped. Start it with gobstopper proxy serve, or check an installed service with gobstopper proxy doctor --json and repair it with gobstopper proxy repair. env -u ANTHROPIC_BASE_URL claude runs one session without the proxy.
gobstopper proxy run -- claude does not start Claude CodeCheck that claude runs directly in the same terminal.
Claude Code still uses the proxy after you remove the exportAn ANTHROPIC_BASE_URL in the env block of a Claude Code settings file applies over the shell. Run /status in Claude Code to see the base URL in use, then remove the entry from ~/.claude/settings.json or the project's settings file.
Claude Code uses an API key instead of your claude.ai sign-inANTHROPIC_API_KEY is set. Claude Code uses it instead of a subscription sign-in, with or without the proxy. Run unset ANTHROPIC_API_KEY to use the sign-in.
Codex requests do not reach the proxyPut model_provider = "gobstopper" at the top level of ~/.codex/config.toml, above the first table, and check that base_url ends in /backend-api/codex with ChatGPT sign-in or /v1 with an API key. Codex ignores model_provider and model_providers in a project's .codex/config.toml.
Codex cannot reach a stopped proxyCodex has no direct fallback. codex -c model_provider=openai runs one session without the proxy; otherwise repair the proxy. gobstopper proxy launch --client codex refuses to start Codex until the proxy is healthy.
gobstopper proxy launch reports Custom client upstream is configuredAn environment variable or a client settings file sets a base URL that is neither the provider's nor this proxy's. The launcher reads the user settings, managed settings, and the project settings in the working directory and each of its parents. For Codex, match --codex-auth to the provider block: chatgpt for a base_url ending in /backend-api/codex, api-key for one ending in /v1.
Claude Code compacts its own history while the proxy is runningThe proxy threshold must stay below Claude Code's trigger. Lower --threshold, especially after setting CLAUDE_AUTOCOMPACT_PCT_OVERRIDE; see Choosing a threshold.
gobstopper proxy status shows no compacted requestsCompaction starts only when a supported request crosses the threshold, 128,000 estimated tokens by default, so a short session can report zero. Preview a recorded session with gobstopper proxy replay <session> --threshold N.
A request fails with HTTP 503 and Retry-AfterRequired sizing timed out or its workers were busy, scoped context storage was unavailable, or a service change is pausing new requests. Retry after the stated delay; gobstopper proxy doctor --json reports the drain phase.
A request fails with HTTP 400 from the proxyThe proxy could not size the payload, the scope was invalid, or --strict refused a request still over the threshold. See Limits and Settings.

Diagnose, repair, and remove

gobstopper proxy doctor --json
gobstopper proxy repair --print
gobstopper proxy repair
gobstopper proxy uninstall

The doctor reports configuration ownership, registration with the service manager, the executable and port, any pending operation, live health, sleep-inhibition state, and the drain phase. Waiting and committed drains are reported as unhealthy for ordinary use. The public status response exposes the proxy process identity, drain protocol, epoch, phase, and remaining wait; it does not expose the controller token. A matching HTTP response alone doesn't establish that startup is configured correctly.

proxy status uses the installed service's recorded port. An explicit --port checks that address even when the saved service configuration is damaged. Its network check has a two-second total deadline and a response size limit, so a listening socket that never responds cannot hang the command indefinitely. If a background transform holds the prefix-store lock, status returns store_details_available: false and null store sizes instead of waiting for it.

GET /gobstopper/ready is a small local check of the listener and of whether the proxy accepts new requests. It returns the process and service identity, the active inference count, and whether new requests are accepted. It returns 503 while draining or while that state is busy. It does not wait for statistics, calibration, context storage, or power-status collection. This checks whether Gobstopper can accept work, not provider credentials or upstream availability.

Normal forwarding has its own concurrency limit. Extra capacity remains for health and service-control requests. Request headers and bodies have total read deadlines, including clients that keep sending small amounts of data. Early rejections send their response before draining unread input for at most 300 milliseconds and 64 KiB; stalled uploads cannot hold that worker indefinitely. At the absolute socket limit, new sockets close immediately so the accept thread continues handling connections. Existing inference is not restarted or replayed to recover capacity.

Repair recreates missing managed definitions and restarts an absent owned job. Running repair or an identical installation from the service's recorded executable also restarts a proxy when its running version differs from the installed version, using the same waiting lease on capable proxies. It resumes a legacy idle pause and reconciles recorded lease operations before starting a service. Installations record an immutable configuration snapshot and a journal before replacement. Repair reconciles an interrupted operation only when its files and loaded job match those recorded identities. It restores the predecessor configuration when available, or completes a first installation. It leaves externally edited files and unrelated jobs intact, and reports the mismatch.

A damaged manifest, a missing executable, unavailable service manager, or changed job requires investigation before repair can proceed. On Windows, a process interruption between task registration and recording the queried task definition can require manual reconciliation; repair won't overwrite a task whose ownership it cannot establish.

Uninstall uses the same drain procedure, then removes the owned startup definition. It retains observation data, logs, and backups. Older proxies require idle inference; unresolved operations require recovery first.

Launch with a direct fallback

gobstopper proxy launch --client claude --print
gobstopper proxy launch --client claude
gobstopper proxy launch --client codex --codex-auth chatgpt
gobstopper proxy launch --client codex --codex-auth api-key

The launcher checks the configured proxy with a two-second network deadline, then starts one client. Claude can use its official provider directly when the proxy is unavailable and the recorded route permits it. Codex requires a healthy proxy and an existing explicitly selected custom provider that already points to it. Codex direct fallback is unavailable: cloud configuration can contain provider fields the launcher cannot completely verify. If the check fails, run Codex normally with its existing settings or repair the proxy.

--print reports proxy readiness and the Claude route, or that Codex will use its existing settings, without starting a client. It does not verify Codex's effective cloud or project route. Saved settings and authentication are retained; pass ordinary client arguments after --. For Codex, --codex-auth identifies the existing ChatGPT or API-key route to check; it does not sign in or obtain a key.

Claude direct fallback refuses custom upstreams, uncertain provider or authentication settings, scoped context reservations, and configured capacity constraints. This includes X-Gobstopper-Scope inside Claude's ANTHROPIC_CUSTOM_HEADERS in environment or settings. Healthy launches preserve scoped headers. Codex retains its entire existing provider table, including literal and environment-backed headers: the launcher does not override its provider selection, URL, headers or environment. Built-in Codex providers and existing direct routes are refused.

The launcher also refuses Codex system and managed configuration layers it cannot inspect. Checking effective macOS managed preferences has a two-second deadline; unavailable inspection refuses launch. Pass short client options separately, since combined flags can conceal a change of working directory or settings.

This choice happens before the client starts. It does not change existing sessions, retry a failed client, or replay inference. Sessions already using a fixed proxy URL still need that listener to remain available.