A coding agent often feels slow long before the model itself is the only bottleneck. Each turn has to reintroduce a working scene: base instructions, tool schemas, project rules, permission state, prior messages, tool outputs, and the new user request. In a long task, the cost of rebuilding that scene becomes visible.
The official OpenAI prompt caching guide requires a matching reusable prefix and recommends stable content before dynamic input. Eligibility, billing, TTL, and key semantics depend on the model. GPT-5.6 and later route automatically and may use prompt_cache_key for separate customer or user cache accounting; earlier models use a stable key to optimize routing. Responses reports reads under usage.input_tokens_details.cached_tokens, rather than the Chat Completions usage path.
prompt_cache_key is not a cache-hit switch. It can guide routing on earlier models or separate cache accounting on newer models; neither use repairs a changed prefix. The client controls the stability of ordered input, tool definitions, and context updates.
Source scope.
OpenAI docs define the server-side prompt caching conditions:
exact prefixes, eligibility threshold, routing, retention,
prompt_cache_key, and cached_tokens.
Public Codex source shows request construction, history
persistence, compaction, and token usage accounting. This
article does not infer private server serialization, machine
routing, or KV-cache placement from client code.
This part follows six questions:
- Which view is actually cacheable?
- Where does Codex get
prompt_cache_key, and how do sessions and internal request families determine it? - How do
instructions,tools, andinputform a stable prefix? - Why should settings, environment context, and tool output land after that prefix?
- Why does compaction reduce pressure while changing future cache shape?
- Which metrics connect cache behavior to perceived latency?
1. Cache the Model View, Not the UI Transcript
Asking whether a turn “hit the cache” is too coarse. First ask: what was sent to the provider, in what order, and through which fields? The visible UI history, the rollout persisted on disk, and the request assembled for the model are related representations, but they are not identical.
| Layer | Implementation | Role | Cache relevance |
|---|---|---|---|
| Visible history | Client event conversion | Shows turns, tools, hooks, events, and final answers. | Useful for reading; not proof of the next model input. |
| Durable record | Rollout / thread store | Supports resume, replay, fork, and audit. | Seeds reconstruction, but is not the provider cache. |
| Model view | Prompt / Responses request |
Contains this turn's instructions, tools, input, and controls. |
The provider can only reuse this ordered request view. |
Codex makes this separation visible in Prompt. It
carries conversation input, model-visible
tools, parallel_tool_calls,
base_instructions and output schema.
ModelClient::build_responses_request turns that into
a Responses payload: instructions, input,
tools, reasoning, text, and
prompt_cache_key.
Shape-level request:
Responses request
instructions: stable base instructions
tools: model-visible tool schemas
input: prior context + dynamic turn tail
text: output schema / verbosity
prompt_cache_key: session/request-family cache key
2. Cache Keys Follow Sessions and Request Families
ModelClient is session-scoped and retains authentication, provider, and transport state across turns. That does not make every key equal to the thread id. The durable thread identity, the current runtime session, and an internal helper request family must be distinguished.
ModelClient::prompt_cache_key() returns an override first. For an internal session with a parent thread id, it returns source:parent_thread_id. Otherwise it uses responses_metadata.session_id. The request builder serializes that value, so helper requests can share a request-family key while ordinary requests follow the current session; the default is no longer simply thread_id.
The key neither chooses a message breakpoint nor guarantees a hit. Public API behavior must also be separated from the ChatGPT client adapter: responses_session_id() preserves the actual session id for non-root agents and otherwise uses the selected key for request affinity. This visible header choice does not prove one server routing algorithm across all model families.
Callers such as Guardian review sessions can supply an override through with_session_options. That groups related requests; a new user message does not need a newly invented key. Observe usage separately to establish actual reuse.
3. Stable Prefix: Instructions, Tools, Context Base
OpenAI's best practice is to put stable content first. Codex has three obvious candidates: base instructions, tool schemas, and the initial or diffed context baseline.
3.1 Instructions Stay Stable Across Ordinary Turns
For ordinary Responses, build_responses_request reads
prompt.base_instructions.text into
instructions. The prompt caching test suite submits
two ordinary turns and checks that instructions
remains unchanged. When a model path needs stable apply-patch
guidance because the tool itself is absent, the test also checks
that the appended instruction block remains identical.
3.2 Tools Are Prefix Material
OpenAI docs say tool definitions can be cached and count toward
the eligibility threshold. Ordinary Responses converts prompt.tools into the top-level tools field; Responses Lite places AdditionalTools in input. Both use the tool router captured by the current StepContext. Parts IV and VII
already showed why that list can be large: core tools, MCP tools,
deferred tool search, dynamic tools, and plugin or skill supplied
capabilities all feed the model-visible tool list.
A large tool list has a cost. But a stable tool list is also excellent prefix material. Unnecessary per-turn tool churn hurts cacheability before the model even looks at the new user input.
3.3 Context Base: Inject Once, Diff Later
The context-management part introduced
reference_context_item. It is the baseline for later
context diffs. When the baseline is still valid, Codex can append
an update rather than reinjecting the whole context bundle.
The test named
prefixes_context_and_instructions_once_and_consistently_across_requests
demonstrates the effect: the second request preserves the first
request's input prefix and appends the next user message.
3.4 Request shape and WebSocket deltas are separate choices
The builder has two request shapes. Ordinary Responses sends top-level instructions and tools. Responses Lite prepends AdditionalTools and base-instruction messages to input, leaving top-level instructions empty and tools absent. Prefix item ids are stable hashes within a thread namespace. Tool eligibility and the minimum cacheable length still depend on the model.
WebSocket continuation compares the complete logical request. When non-input properties are compatible and previous input plus returned output remains a prefix, it sends only new input with previous_response_id. Compaction, configuration changes, or a prefix mismatch can require full input. A full network payload does not imply a complete cache miss.
shape-level, fields omitted:
Ordinary Responses: instructions + tools + input
Responses Lite: input = AdditionalTools + base instructions + history
Logical input: S + prior output + new tool result
Reusable WebSocket continuation: previous_response_id + new tool result
No continuation: full input
Cache result: observe usage separately from transmitted bytes
4. Dynamic Tail: Change Later, Not Earlier
Dynamic material is inevitable. Users switch settings, change cwd, change approval policy, request a different model, trigger tools, receive tool outputs, or let hooks add context. Codex does not try to remove all change. It tries to keep change behind the already-stable prefix.
Two tests make that concrete. One applies thread settings
overrides and checks that prompt_cache_key stays
constant while updated permissions and environment messages are
appended after the existing prefix. Another applies per-turn
overrides and checks the same invariant while adding a model
switch and new environment context.
| Change source | Codex handling | Protected property |
|---|---|---|
| New user input | Append a new user message. | Prior prefix is not reformatted. |
| Thread settings override | Append settings and environment updates. | Stable context remains before the change. |
| Tool output | Record history, then normalize with for_prompt(). |
Call/output invariants stay intact. |
| Hook additional context | Write a contextual fragment into model context. | Policy additions become visible input changes. |
| Compaction | Install summary or replacement history as a new baseline. | Future turns get a smaller but reconstructible prefix. |
5. Compaction Relieves Pressure and Changes Shape
Prompt cache does not replace the context window. Once history
grows too large, Codex still has to compact. ContextManager
stores token usage, estimates counts, and bumps
history_version when history is rewritten. The
get_context_remaining tool reports the smaller remaining allowance across the auto-compact scope and effective full window, using the shared context-window status.
The current remote compact v2 attempt clones history, trims tool outputs to fit the window, retains history metadata, and appends CompactionTrigger. It uses the shared ModelClientSession streaming path. Collection requires response.completed and exactly one compaction output before replacement history can be installed. Sending a request, receiving an output item, and installing a checkpoint are distinct states.
Compaction has two performance effects: it reduces future history pressure, and it establishes a new prefix shape. The old prefix is no longer the mainline model view; the summary, retained outputs, and reinjected context become the next cache candidate.
6. Metrics: Cached Input, Non-Cached Input, First Token
The provider reports cache behavior through usage. Codex tracks
input_tokens, cached_input_tokens,
output_tokens, reasoning_output_tokens,
and total_tokens. At turn completion,
on_task_finished compares the token snapshot from
turn start with current total usage to compute this turn's input,
cached input, non-cached input, output, and total tokens.
The same path records tracing fields, session telemetry
histograms, analytics events, time_to_first_token_ms,
and total turn duration. That matters because a high cached-token
count and a fast-feeling turn are related but not equivalent.
Tool loops, approvals, network retries, compaction, and output
length all shape user-visible latency.
| Metric | Meaning | Do not read it as |
|---|---|---|
cached_input_tokens |
Input tokens served from provider cache. | Tokens the model did not see. |
non_cached_input() |
Input tokens that were not cache hits. | All non-cached tokens are waste. |
time_to_first_token_ms |
Delay until the first visible model output. | A value controlled only by prompt cache. |
total_tokens |
The turn's total token count. | A guarantee of semantic completeness. |
7. Common Misreadings
The API fields look small, but the runtime discipline behind them is large. These are the misreadings that tend to make performance debugging noisy.
| Misreading | Better reading | Source check |
|---|---|---|
prompt_cache_key decides the hit. |
The key helps routing; exact prefix still decides eligibility. | Key construction and request shape are separate. |
| If the chat shows it, the model sees it. | The model sees for_prompt() output. |
History normalization owns the request input. |
| Fewer tools are always faster. | Stable tool schemas are costly but cacheable prefix material. | The tools field comes from prompt.tools. |
| Compaction automatically improves cache. | Compaction changes the future prefix and must remain reconstructible. | Remote compact reuses the key while rewriting history. |
| A cache hit guarantees a fast turn. | Tool loops, approvals, compaction, output length, and network still matter. | Codex records token usage, TTFT, and duration separately. |
8. Why Recovery Comes Next
Prompt cache pushes performance analysis onto model view shape. But that shape has to be reconstructed from durable history: rollout items, context updates, tool outputs, compaction replacement, rollback markers, and fork state. If that durable record cannot rebuild the same history, the next request shape drifts and cache analysis loses its anchor.
The next part returns to persistence and recovery: how Codex writes
a turn into rollout, reconstructs history from RolloutItem,
handles historical rollback markers, fork, compact transitions, and token usage. That
connects what the user saw, what disk retained, and
what the model sees on the next turn.
Source References
- OpenAI Prompt caching guide
- OpenAI Prompt Caching 201
- OpenAI latest model guide: reasoning models and prompt caching
- openai/codex pinned source snapshot
Promptandget_formatted_input_for_request()ModelClientand turn-scopedModelClientSessionwith_session_options()and session / internal-family keysbuild_responses_request()ResponsesApiRequestand websocket request mapping- session model client initialization and Guardian override
run_turnrequest input constructionContextManager,reference_context_item, andfor_prompt()- token usage update and local estimation
TokenUsageInfoandTokenCountEvent- turn start token usage snapshot
- turn completion cached / non-cached token usage
- turn duration and time-to-first-token event
get_context_remaining- remote compaction v2 prompt construction
- compact attempt and the shared request path
- prompt caching tests: stable tools and instructions
- prompt caching tests: reused contextual prefix
- prompt caching tests: thread settings overrides
- prompt caching tests: per-turn overrides
- remote compact v2 shared sampling path