Codex Source Notes · Part X

Rebuilding sessions and forks from rollout records

Part IX placed performance on request shape. A thread also needs to survive process exit, resume from an old session, fork into a new branch, and interpret historical rollback records without resurrecting stale tool output in the next prompt. Those capabilities all depend on the same thing: Codex must persist a turn as records it can interpret again.

Project: openai/codex Topic: rollout / recovery Scope: public source
Codex rollout recovery showing append-only JSONL, reverse scan, replay suffix, model history, context baseline, and resume fork output
Select a surviving checkpoint in reverse, then replay the suffix forward. Resume and fork consume the rebuilt context. Legacy rollback markers are compatibility records, not a current active rollback API.

Imagine a long coding task. The agent has read files, run tools, compacted history, changed model settings, and accumulated enough context that a simple transcript is no longer the whole story. If the process stops, the next runtime must recover more than visible messages. It needs the model-visible history, the current context baseline, session identity, lineage, and the event facts that a client will project back to the user.

That is the job of rollout. Codex writes an append-only JSONL stream of typed RolloutItems, keeps write ordering behind a recorder, and reconstructs the runtime state by replaying those items through session logic. The recovery path is not a convenience wrapper around chat history; it is part of runtime execution.

Rollout is an append-only stream of typed JSONL records. It persists enough information to rebuild model history, context baselines, token usage, state after legacy rollback markers, and fork starting points.

Source scope. This article describes public openai/codex source: rollout schema, recorder behavior, session initialization, resume, fork, legacy rollback compatibility, reconstruction, and tests. It does not infer private backend storage behavior or treat local rollout paths as a public interface.

This part follows six questions:

  1. Why is continuing a thread harder than showing old messages?
  2. What does each RolloutItem variant own?
  3. How does the writer queue, persist, flush, and recover from failed writes?
  4. Why do resume and fork first become InitialHistory?
  5. Why does reconstruction scan backward before replaying forward?
  6. How do legacy rollback records, compaction, and prompt caching depend on these records?

1. Recovery Needs Four Kinds of State, Not Just Messages

A client needs visible facts. The next model request needs ordered input. The runtime also needs a context baseline: working directory, sandbox, permission profile, model, collaboration mode, and the latest full-context checkpoint. Those requirements overlap, but they cannot collapse into one saved string.

Recovered state Visible layer Runtime requirement Rollout role
Client state User messages, tool progress, warnings, and legacy rollback events. Event order, turn start/end, resume/fork source. Persist EventMsg for client event replay.
Model history Messages and tool results the model should see. Compacted replacement history and surviving suffix. Persist ResponseItem and CompactedItem.
Context baseline cwd, sandbox, permissions, model, collaboration mode. The last durable full-context baseline. Persist TurnContextItem and WorldState.
Thread identity Thread id, source, parent and fork lineage. Resume keeps an id; fork creates a new id with lineage. Persist SessionMeta.

This is why rollout connects earlier parts of the series. Client event conversion, multi-agent spawning, context management, and prompt caching all rely on the same durable records.

2. Rollout Types Preserve More Than Messages

RolloutItem now lives in the history crate and has more than five variants. Alongside SessionMeta, ResponseItem, Compacted, TurnContext, and EventMsg, it includes WorldState, TokenUsageRecord, RetainedContext, inter-agent communication, and model-invisible realtime presentation facts. Each consumer interprets the records it needs; persistence does not make every record model-visible.

This shape-level fragment omits required fields to focus on the envelope. RolloutLine keeps timestamp and optional ordinal at the top level and flattens type/payload. The ResponseItem wire representation can store metadata beside the payload. Reconstruction dispatches on top-level type, not a nested item.type.

{"timestamp":"...","ordinal":1,"type":"session_meta","payload":{"id":"t1"}}
{"timestamp":"...","ordinal":2,"type":"world_state","payload":{"full":true,"state":{}}}
{"timestamp":"...","ordinal":3,"type":"response_item","payload":{"type":"function_call","call_id":"A","name":"shell","arguments":"{}"},"metadata":{"client_authored":false}}
{"timestamp":"...","ordinal":4,"type":"compacted","payload":{"message":"...","window_number":1,"replacement_history":[]}}

TurnContextItem preserves recoverable settings, while WorldState stores full environment snapshots or changes. The write path records model-visible context before WorldState and the required TurnContextItem. A bare settings object does not prove its context reached history; recovery also looks for a real input boundary or a valid full WorldState after compaction.

CompactedItem can hold replacement history, retained context, window numbers and ids, and the latest token usage record. The replacement gives the model-history base; the remaining records let recovery restore settings, authorization-relevant facts, and token totals without indefinitely scanning older history. Fields are optional, so older records still need compatibility handling.

3. Writes Queue Before File Writes Complete

New and resumed sessions initialize the recorder differently. A new session precomputes rollout path and session metadata, but it can defer file creation until persist(). A resumed session opens the existing rollout for append immediately.

Actual writes go through a background writer task. The recorder sends AddItems, Persist, Flush, and Shutdown commands. Items live in pending_items until they are written successfully. If I/O fails, the writer drops the file handle, keeps the unwritten suffix, and retries after reopening at the next persist or flush command.

writer discipline:
record_canonical_items(items) -> queue AddItems
persist()                   -> materialize file + write pending
flush()                     -> wait for preceding file writes (not fsync)
shutdown()                  -> final drain before exit

Preparing a fork from an active source thread flushes before reading so queued records are not omitted. The completion boundary is file writes and flush(): the writer does not call fsync/sync_all here. Queue acceptance, completed writes, and power-loss-safe stable storage are different guarantees.

4. Resume and Fork Enter Through InitialHistory

Session startup normalizes all starting points into InitialHistory: New, Cleared, Resumed(ResumedHistory), and Forked(Vec<RolloutItem>). New and cleared sessions defer initial context insertion until the first real turn. Resumed and forked sessions create a default turn context and call apply_rollout_reconstruction.

Start Thread id History source Startup behavior
New / Cleared Fresh id No prior rollout. Defer initial context to the first real turn.
Resumed Existing id Recoverable history loaded through the thread store. Rebuild history, settings, and token usage.
Forked Fresh id Snapshot items from the source thread. Rebuild history; persist a copy or a parent-history reference.

Resume also warns when the last recorded model differs from the current model. Reconstruction can restore the shape of history, but changing models can still change context windows, cache behavior, and runtime performance.

5. Reconstruction Scans Backward, Then Replays Forward

reconstruct_history_from_rollout does not replay the whole file from the first line. It first scans newest-to-oldest. The scan looks for the newest surviving replacement-history checkpoint, previous turn settings, reference context item, and window id. Once those are known, older items cannot affect the rebuilt state.

Historical rollback markers become a counter during this reverse pass. ThreadRolledBack means “drop the newest N real user turns.” In reverse, that means skipping the next N finalized turn segments that actually contain a user message. Tests cover the distinction: standalone task turns should not consume rollback skips.

After selecting the newest surviving checkpoint, forward replay installs its replacement history and replays the suffix: annotated response items, inter-agent messages, retained context, and legacy compaction/rollback records. A newer replacement checkpoint in that suffix may belong to a rolled-back turn. It must not blindly overwrite the selected base; original records preserve the boundaries needed by compatibility replay.

reconstruction:
reverse scan
  find newest surviving replacement_history
  recover previous_turn_settings
  recover reference_context_item
  account for ThreadRolledBack markers

forward replay
  seed ContextManager from replacement_history
  append surviving ResponseItem suffix
  apply rollback markers to the rebuilt history

The shape matters. The reverse pass avoids interpreting old log prefixes after a surviving checkpoint. The forward pass preserves the exact ordering semantics of the live tail.

6. Legacy Rollback Records Limit Bounded Reads

The current submission protocol no longer exposes the former ThreadRollback operation. ThreadRolledBack remains in persistence and recovery for historical records. The drop_last_n_user_turns comment explicitly limits it to that replay role. Compatibility code does not establish a current callable rollback API.

ModelContextScan supports bounded reverse reads of the latest model context. It needs both a checkpoint containing replacement_history and window_number and a turn establishing a valid settings baseline. A raw role=user item is insufficient: contextual fragments can use that role too. Legacy rollback markers or compacted records missing either required field force a scan back to the beginning.

shape-level read decision:
Complete checkpoint + settings baseline -> bounded scan may stop
Legacy rollback marker                  -> scan to the beginning
Legacy compact missing required fields  -> scan to the beginning
Contextual user fragment alone          -> no proven user-turn boundary

The local thread store actually feeds rollout segments through this scanner. Core reconstruction then interprets the supplied item sequence. Storage selection and semantic replay are separate responsibilities; a reverse algorithm does not prove that every recovery entry avoids reading a complete file.

7. Fork Copies Replayable History

Fork shares the recovery machinery with resume, but lineage is different. Resume keeps the old thread id. Fork creates a new id while recording forked_from_thread_id. The manager reads rollout-backed history, shapes it with fork_history_from_snapshot, and starts a new thread.

Fork persistence has two strategies. Ordinary forks use ForkPersistence::Copied; prepared forks may use Referenced with a history_base and inherited item count. The child still receives reconstructable model context without necessarily rewriting the parent’s entire history into its own rollout. Forking is therefore not universally a full JSONL copy.

8. The Prompt Cache Connection

Part IX argued that prompt cache hits depend on stable request shape. Recovery preserves the precondition for that stability. TurnContextItem tells the runtime what context has already been established. CompactedItem.replacement_history gives history a new base. Rollback markers trim the tail consistently.

If recovery were based only on visible transcript text, Codex could re-inject full context as repeated diffs or carry rolled-back tool output into the next model view. That would be a semantic bug and a prompt-cache-shape bug. The compact/resume/fork tests assert that compacted input prefixes remain predictable across recovery.

9. What to Carry Forward

Rollout/recovery gives a useful agent-design rule: a long-running agent should persist replayable facts, not only final state snapshots.

Rule Codex mechanism Result
Append facts RolloutLine timestamp plus typed item. Recovery and audit read the same records, including legacy rollback compatibility markers.
Separate record types History, context, session meta, and events are distinct. Client state and model view can rebuild differently.
Checkpoint history replacement_history after compaction. Long threads can resume from a newer base.
Use semantic markers ThreadRolledBack as an appended event. Old facts remain; replay computes the current world.
Flush before snapshots Fork flushes the active source before reading; shutdown drains pending writes. Snapshots are based on persisted facts.

The internal execution path is now almost complete: requests enter the runtime, context forms the model view, tools create side effects, events go to clients, rollout persists the records, and recovery turns those records back into a usable next turn. The next part should look at the public edge: how SDK and app-server callers enter the same runtime.

Sources