Start with a normal Codex session. You ask it to fix a bug in a
repository. It reads AGENTS.md, searches source files,
runs a test, inspects the failure, and edits the patch. Halfway
through, you add a constraint: do not touch generated files. Later
the thread approaches the context limit and compacts. The next
morning, you resume the thread and expect the agent to continue
from the right place.
A plain "append every message to a chat history" model cannot explain this workflow. Repository instructions, working directory, sandbox policy, skill metadata, tool schemas, tool outputs, images, reasoning tokens, pending user input, compaction summaries, and resume baselines are stored and used by different components. Some must be durable. Some are only model-visible for one request. Some may be truncated. Some must be re-injected from scratch if the baseline becomes unreliable.
The official Codex docs describe the product behavior: a thread is a session made from prompts, model outputs, and tool calls; as the agent works, it gathers context from files, tool output, and its running record; everything must fit in the model's context window, and long threads may be compacted. The source answers the next question: where does that information live, when does it become model-visible, and what gets installed after compaction?
OpenAI's
agent-loop article
makes the external constraints concrete: a model request is made from
prompt, tools, input, and related request fields; long-running work also
has to deal with prompt cache, the context window, and compaction.
The source-level question is which of those constraints is carried by
ContextManager, TurnContext, and replacement history,
and which state remains durable versus turn-local.
The easy-to-miss step is what happens after a tool call. The next
model request does not start from a blank page. The previous
reasoning, the tool call, and the returned
function_call_output are appended to the next request's
input in order. The second request therefore keeps the stable prefix
from the first request and adds the new tool facts at the tail.
first inference:
instructions
tools
input:
- user task
-> reasoning
-> function_call(call_id=A)
tool execution:
command returns:
function_call_output(
call_id=A
)
next inference:
instructions
tools
input:
- user task
- reasoning
- function_call(call_id=A)
- function_call_output(
call_id=A
)
Complete logical input does not imply a full network payload every time. The client assembles its complete input before transport; WebSocket continuation can reuse previous_response_id and send only new input when non-input properties remain compatible and previous input plus returned items remains a prefix. Without a usable response id or matching prefix, it sends full input. Client-side reconstructability remains true, while a blanket claim that Codex never uses previous_response_id does not. Cache reuse and reduced network bytes are separate questions.
Source scope. This article uses official documentation for product behavior and the public openai/codex source for implementation details. Source links point to one fixed public snapshot. Where the source exposes only the remote-compaction selection conditions, this article describes the visible lifecycle and does not infer the server-side summarization algorithm.
The article follows six questions:
- In one agent loop, how do model output, tool calls, and tool results return to the next request?
- Why is durable history not the same thing as the prompt?
- When does Codex inject full initial context, and when does it emit only a diff?
- What does
for_prompt()normalize or filter? - During compaction, which history is replaced and what recovery record remains?
- Why do legacy rollback replay, resume, and model switches affect the next context injection?
1. Separate the four kinds of data
The first step is to stop looking for one memory object. Codex has at least four relevant kinds of data: session-scoped state, per-turn configuration, durable history, and the model request. They are not interchangeable.
TurnContext keeps turn identity, initial settings, and control state. Each sampling request captures a StepContext containing one version of model settings, environments, MCP connections, and the tool router. Changes during tool execution can affect a later step without mixing the advertised capabilities and execution settings of an already captured request.
| Data | Stored or generated by | Model-visible? | What must remain true |
|---|---|---|---|
| Session state | SessionState |
Not as a whole | Keep history, permission baselines, compaction windows, and recovery state across turns. |
| Turn and sampling snapshots | TurnContext + StepContext |
Partly, through context messages and request fields | Keep each sampling request and its tool execution on the same captured configuration. |
| Durable history | ContextManager.items |
Only after filtering and normalization | Retain recoverable API messages, tool calls, outputs, and compaction records. |
| Model-visible prompt | for_prompt() + Prompt |
Yes | Send a legal request view adapted to the current model and tools. |
The ContextManager
comment calls it a transcript of thread history, but the struct
also carries history_version, token_info,
and reference_context_item. That last field is the
baseline used to generate model-visible settings updates. If it is
missing, the next regular turn falls back to full context
reinjection instead of diffing against stale state.
2. A Turn Does Not Jump Straight From User Input To Model
In run_turn,
the first step is not to record the new user message. Codex first
runs pre-sampling compaction. A nearby source note says this happens
before context updates and the new user message are recorded, and
that pending incoming items should eventually be estimated before
deciding whether the thread crosses the compaction threshold.
After the pre-sampling compaction check, run_turn captures the first StepContext, records environment updates, and prepares skills and user input. Before sampling it clones history and applies for_prompt() using that step’s model input modalities. Later steps can capture refreshed settings and capabilities; a turn-wide object does not freeze the tool catalog for every request forever.
turn starts
-> maybe compact old history
-> capture StepContext for this sampling request
-> inject full context or WorldState changes
-> load skills/plugins chosen for this turn
-> record user input and additional context
-> clone history
-> for_prompt(input_modalities)
-> build Prompt with tools and base instructions
-> stream model response
The turn loop also treats pending input carefully. A
source comment
says the turn-scoped ModelClientSession is reused
across retries, while pending input is deliberately not drained at
the very start of a turn or immediately after auto-compaction.
That ordering preserves continuation semantics. If a steer is
inserted in the wrong place, the model may resume the wrong work.
2.1 Initial Context Is Synthesized Runtime State
Full initial context comes from
build_initial_context_with_world_state.
It collects model-switch guidance, permission instructions,
developer instructions, collaboration mode, realtime state,
personality, apps instructions, available skills, plugins,
extension fragments, token budget context, environment context,
and user instructions. These become developer or contextual user
messages. They are not the user chat message.
The structured baseline is stored as a
TurnContextItem.
That lets Codex later decide whether a new turn can emit only a
settings diff or must re-inject the full context bundle.
2.2 Compare WorldState against the established baseline
Suppose the bug-fix task changes from workspace-write to read-only. The runtime must distinguish the old environment, the current environment, and what it has already told the model. It builds a WorldState and keeps a world_state_baseline in history. The separate reference_context_item remains the durable turn-settings baseline.
record_context_updates_and_set_reference_context_item builds WorldState for the captured step. With no reference baseline it emits full initial context plus a full WorldState record. Otherwise update_world_state produces changed fragments, with turn contributions added when turn settings change.
shape-level:
step = capture(settings + environments + MCP + tools)
world = build_world_state_for_step(step)
if reference_context_item is None:
history += full_initial_context(step, world)
rollout += full_world_state
else:
history += visible_world_state_changes
rollout += world_state_changes_if_any
persist TurnContextItem when required
advance the reference baseline
The order matters: record model-visible messages first, persist the WorldState that produced them next, then write the required TurnContextItem. A snapshot-only change with no visible fragment or turn-setting change does not require another TurnContext record. Recovery must not restore a baseline that claims the model received an update whose message was never recorded.
Both full reinjection and incremental updates remain. The current diff mechanism is WorldState-based, rather than the former fixed settings-field list. Missing baselines require complete context; trusted baselines let new facts join the history tail.
3. for_prompt() filters and normalizes history
Before sampling, Codex does not send raw
ContextManager.items. It calls
for_prompt(),
which normalizes history and returns the items suitable for the
current model request. Durable history stays recoverable, while the generated request must satisfy the current model's capabilities.
normalize_history supplies missing call outputs, removes orphaned paired outputs, and strips unsupported image and audio content. Function/custom outputs need their matching calls; named external outputs may stand alone. for_prompt() then unwraps ResponseItemEnvelope into API ResponseItem values. History-only provenance and truncation-budget metadata stay local instead of becoming the API envelope.
durable history, shape-level:
user text
function_call(call_id=A)
function_call_output(call_id=A)
orphan function_call_output(call_id=old)
user image + audio
model-visible view for a text-only model:
user text
function_call(call_id=A)
function_call_output(call_id=A)
user text with unsupported image/audio stripped
The same layer also truncates large tool outputs. The
record_annotated_items
path applies a truncation policy to function output payloads.
Then build_prompt
adds model-visible tool specs, parallel tool call support, base
instructions and output schema. Tools are not
stored as normal chat messages, but their schemas still enter the
model-visible request.
3.1 Before compaction, combine server counts and local estimates
for_prompt() explains what the model sees. The next question is when the
stored history is large enough to compact. Token accounting is not just the
server's last usage value.
get_total_token_usage
adds the latest server usage, local items appended after the last
model-generated item, and, when the server has not included them,
estimated non-last reasoning tokens. The local estimate is coarse:
the estimator is byte-based, not tokenizer-accurate.
context_window_token_status
uses that active context count together with the configured
auto-compact scope, a scope limit, and the full context window
limit. In BodyAfterPrefix mode, it subtracts a
prefill baseline held by
AutoCompactWindow.
Server-observed input tokens replace an estimated baseline when
they become available.
4. Compaction Installs Replacement History
Compaction is easy to under-describe as "make a summary." In Codex, the visible mechanism is a recoverable history rewrite: generate a handoff summary, preserve recent real user messages, install replacement history, advance the auto-compact window, and decide how initial context should be represented after the rewrite.
4.1 Pre-Turn And Mid-Turn Compaction Differ
The pre-turn path is
run_pre_sampling_compact.
It checks whether compaction is required before the new user turn
is sampled. There are also pre-turn compactions for compaction
compatibility hash changes and model downshifts to smaller context
windows.
The mid-turn path happens after sampling when the model still
needs a follow-up or pending input exists. The post-sampling state
collection checks token status and, if the context limit has been
reached while follow-up work is needed, runs auto-compaction with
InitialContextInjection::BeforeLastUserMessage.
The enum comment in
InitialContextInjection
explains the split. Pre-turn and manual compactions use
DoNotInject, replace history with a summary, and
clear reference_context_item; the next regular turn
will fully reinject initial context. Mid-turn compaction injects
initial context into replacement history just above the last real
user message, preserving the model's expected continuation shape.
4.2 The summary prompt defines what a future turn must retain
Local inline compaction starts in
run_inline_auto_compact_task.
It turns turn_context.compact_prompt() into user
input. The default
compact prompt
asks the model to create a handoff summary for another LLM that
will resume the task: current progress, key decisions, constraints,
remaining steps, and critical references.
In
run_compact_task_inner_impl,
Codex clones history, appends the compact prompt to that compact
request, runs the model to completion, extracts the last assistant
message as the summary suffix, collects real user messages, builds
replacement history, advances the auto-compact window id, installs
the compacted history, and recomputes token usage.
old history:
initial context
user A
model/tool work
user B
model/tool work
replacement history, shape-level:
recent real user messages, capped
user message: SUMMARY_PREFIX + handoff summary
build_compacted_history
retains recent real user messages, capped by a 20,000-token
approximate budget, then appends the summary as a user message.
When mid-turn compaction needs initial context, the
insertion helper
places canonical initial context before the last real user message
when possible, or before the summary or compaction item as a
fallback so the compacted item remains last.
4.3 Remote Compact V2 Has an Explicit Completion Boundary
The handoff summary above describes local compaction. A remote v2 attempt also clones history, but uses for_prompt_annotated() to retain metadata needed by this specialized request and appends CompactionTrigger. It streams through the shared ModelClientSession and receives a compaction item, rather than assuming a readable assistant summary.
Collection requires response.completed and exactly one compaction output before installing replacement history. Dispatching a request, receiving an item, and installing a checkpoint are distinct states. Failures follow bounded retries and conditional fallback; an unfinished stream is not a successful compact. How the server generates the content is outside the public client implementation.
4.4 Replacement Is Persisted To Rollout
First place rollout: it is the replayable durable record layer,
not the current in-memory history. Compaction changes durable state too.
replace_compacted_history
replaces in-memory history, persists RolloutItem::Compacted,
persists a full WorldState when supplied and a RolloutItem::TurnContext when a new reference baseline was established. A settings persistence lock and a current-settings event prevent an old turn snapshot from overwriting accepted updates. It then queues a session-start hook for
the compacted state. That makes compaction part of resume and fork
semantics, not just a temporary token-saving step.
5. Recovery and New Windows Invalidate Baselines
The difficult part is not appending a new turn. It is keeping the baseline valid after legacy rollback replay, resume, fork, and compaction. Codex does not keep diffing against a baseline it can no longer reconstruct. When the surviving history no longer contains the bundle that established the baseline, Codex clears the baseline.
drop_last_n_user_turns now exists to replay rollback markers in historical rollouts, not as an active rollback operation in the current submission protocol. It still removes context updates adjacent to the cut. If this trims a mixed initial-context developer bundle, the trim logic clears reference_context_item, forcing complete reinjection on the next real turn.
Resume reconstruction treats history, previous settings, reference
context, and window id as a single recovery group.
apply_rollout_reconstruction
installs the reconstructed history and baseline. In
BodyAfterPrefix mode, it also estimates prefix tokens
for the recovered compaction window.
start_new_context_window advances the window number and ids, constructs initial context from the captured step and WorldState, and installs an empty-message CompactedItem with a fresh baseline. When the client-developer-message retention feature is enabled, it also preserves those messages within a budget. This path creates a fresh prefix without generating a semantic summary of previous work.
6. Common Misreadings
The mechanism becomes easier to reuse once the common shortcuts are made explicit.
| Misreading | What the source shows | What breaks if implemented literally |
|---|---|---|
| History is the prompt | ContextManager stores durable items; for_prompt() derives a request view. |
Orphan outputs, unsupported images, or bad call/output pairs can poison the request. |
| Initial context is repeated every turn | A valid baseline allows settings diffs; missing or unreliable baselines trigger full reinjection. | Either waste tokens or, worse, diff against stale runtime state. |
| Compaction is just a summary | Compaction creates a handoff summary and installs replacement history in rollout. | Resume cannot know which history was replaced or how to restore the baseline. |
| Token usage is only server usage | Codex combines server usage, local appended items, reasoning estimates, and window baselines. | Large local additions can push a request over the window without being noticed early. |
| Rollback just deletes user messages | Rollback also trims nearby context updates and may clear the reference baseline. | The next turn can send diffs based on an initial context bundle that no longer survives. |
7. Runtime Rules To Carry Forward
The Codex implementation suggests a small set of general rules for coding-agent runtimes.
| State type | Where it belongs | Before model visibility | What must remain true |
|---|---|---|---|
| Conversation and tool evidence | Durable history / rollout | Filter by model capability and preserve valid call/output pairs. | Recovery and later inspection retain complete records. |
| Runtime environment and permissions | Turn context plus reference baseline | Inject fully first, then emit trusted diffs. | The model sees current settings without stale configuration. |
| Tool capability | Tool router / prompt request | Expose as model-visible specs, not as normal chat text. | Tool schemas match the tools the runtime can execute. |
| Long-thread compaction | Replacement history plus compacted rollout item | Install a handoff summary and reinsert initial context when needed. | Long work can continue while acknowledging compaction is lossy. |
| Context-window accounting | token_info plus auto_compact_window |
Combine server usage with local estimates. | Neither the server count nor the local estimate is trusted alone. |
That is why context management is a natural first mechanism article after the overview. Tool calling, permissions, multi-agent work, skills, and plugins all come back to the same discipline: decide where each piece of state is stored, then decide which parts belong in the next model request.
Sources
- openai/codex fixed source snapshot
- Codex docs overview
- OpenAI: Unrolling the Codex agent loop
- AGENTS.md docs
- Codex skills docs
- SessionState fields
- TurnContext definition
- ContextManager fields
- run_turn order
- build_initial_context_with_world_state
- record_context_updates_and_set_reference_context_item
- normalize_history
- context_window_token_status
- run_compact_task_inner_impl
- Default compact prompt
- replace_compacted_history
- drop_last_n_user_turns