Reading contract: Follow “fix this failing test” from queued input to a model request, checked tool actions, and recoverable records. Distinguish when work is queued, executed, recorded, and verified.

The implementation evidence is the public mirror's March 31, 2026 snapshot. This is not an official Anthropic open-source implementation or evidence of the current release. The official product overview, memory documentation, and tool-use contract describe product and API behavior separately. Missing feature-gated modules and provider internals remain outside the visible implementation.

Imagine the user types: fix this failing test. Claude Code does not pass that sentence straight to the model. The request is first separated from early-return CLI modes, enters the interactive app, gets ordered with slash commands and background notifications, is wrapped with system and user context, then finally reaches query().

That route matters because the codebase is layered. The CLI decides whether a normal agent session should even start. The command queue decides which piece of work gets the next turn. The REPL prepares memory, repository state, tool context, and permission callbacks. The query loop turns selected runtime state into a provider request. The tool layer turns tool_use into local side effects only after validation and permission checks. The transcript and resume paths keep enough state for a future session to continue.

The main invariant is this: Claude Code is not one big function that asks a model and runs whatever comes back. It is a sequence of stages, and each stage changes or checks the input before handing it onward.

type NormalTurnRoute = {
  cli: "early-return flags" | "interactive REPL"
  queue: "user prompt" | "slash command" | "permission response"
  repl: { systemPrompt; memory; tools; permissionContext }
  queryLoop: { modelVisibleMessages; apiStream; toolUseBlocks }
  toolGate: { validate; permission; execute; toolResult }
  transcript: { uiMessages; durableRows; resumeState }
}
Source shape: the overview keeps the responsible modules in view before later chapters zoom into one stage at a time.

Five questions guide the route: which CLI modes return early; how competing inputs are ordered; what the REPL assembles; how tool proposals become checked actions; and which records a resumed session needs.

1. The seven stages on the normal route

A useful first map has seven checkpoints: CLI routing, command queue, REPL turn assembly, query(), model stream, tool execution, and persistence. The names are not decorative. They describe where the code changes shape.

StageWhat changes thereSource to read first
CLI routingVersion, MCP host, remote-control, resume, and normal interaction are separated.entrypoints/cli.tsx, main.tsx
QueueUser input, slash commands, task notifications, and orphaned permission requests become ordered commands.commands.ts, messageQueueManager.ts
REPL assemblySystem prompt, memory, repository state, tools, MCP clients, and permission functions are prepared.screens/REPL.tsx, context.ts, Tool.ts
Query loopThe model-visible view is rebuilt before each provider call.query.ts, utils/api.ts
Model streamText, thinking, tool_use, stop reasons, and API errors become runtime events.query.ts, services/api/claude.ts
Tool checks and executionTool names, schemas, permissions, concurrency, progress, and results are handled.toolOrchestration.ts, toolExecution.ts
Record and resumeUI messages, transcript rows, content replacement, compact state, and resume metadata stay separate.sessionStorage.ts, conversationRecovery.ts

2. CLI is a router; the REPL is the main road

The CLI entry has several paths that return before the interactive agent starts. The version path loads no further modules; other early paths, such as prompt dumping or remote control, have their own feature or runtime conditions. Some flags print information, some paths start MCP host behavior, and resume-related paths load older state. The normal coding conversation continues through main.tsx, where session config, commands, tools, MCP clients, and turn-complete callbacks are assembled. Continue and resume load older state first; the fresh interactive path starts a new REPL session.

The handoff to replLauncher.tsx is where a terminal command becomes an interactive runtime. From that point on, Claude Code can render progress, accept shortcuts, show permission prompts, and keep a local message model that is richer than one text input.

3. The queue orders pending inputs

A slash command, a natural-language prompt, a background notification, and a permission response are not the same thing, but they compete for the same turn. The queue manager gives them now, next, and later priority classes with FIFO order inside each class. Ordinary enqueue defaults to next; pending task notifications default to later. A dequeue filter can leave commands for another agent untouched. This orders eligible pending inputs, rather than promising every kind of background work runs deterministically. Without that layer, a coding agent would feel random as soon as background work and user input overlap.

4. The REPL assembles the turn

Before the failing-test request reaches the loop, the REPL creates a tool context, reads available tools and MCP clients, resolves system and user context, and constructs the effective system prompt. It then consumes query() as an async generator. Assembly matters because the tool executor needs the same session's permissions, messages, and tools when a result arrives later.

getSystemContext() and getUserContext() are memoized. System context can include an initial git-status snapshot, but remote mode or disabled git instructions skip it. User context can include discovered memory files and the date. Calling these helpers again usually reuses cached results; it does not imply disk reads on every turn. At the request boundary, appendSystemContext() and prependUserContext() place the data in system text and a meta user message respectively.

ToolUseContext also carries commands, agent definitions, access to application state, content-replacement state, and rendered prompt information. These are runtime capabilities and records, not all fields that the provider sees directly.

5. query() is the real agent loop

The outer query() delegates to queryLoop() and marks consumed commands complete during cleanup. The inner loop owns repeated model calls and tool follow-up.

Query-loop branches: tool_use runs through tool scheduling and returns results to context; a response without tool_use enters stop checks before continuation or turn end
A focus map of the normal loop. Error and budget branches are omitted; ending a turn does not prove the task has passed verification.

5.1 Select the next request's context

The REPL does not simply forward raw history. Before each model call, queryLoop() takes messages after the most recent compact marker, applies tool-result budgeting, then conditionally considers snip, microcompact, context collapse, and automatic compaction. Some feature-gated implementation modules are absent from this snapshot: visible call sites establish ordering, not their hidden selection strategies. The result is the message view sent to the next model call. What is saved locally and what is sent to the model are related, but they are not identical.

5.2 Build a provider request

The API call contains more than messages. It carries the effective system prompt, thinking options, tools, MCP tools, agents, query source, task budget, cache policy, and permission context access. That is why a later article on prompt cache or tools must come back to this same loop.

Message conversion handles user and assistant payloads, cache control, media, thinking, and tool results; queryModel() starts translating the remaining runtime options into API behavior. A local transcript row is therefore not automatically a valid provider request block.

5.3 Follow actual tool blocks and stop checks

During the stream, the loop collects actual tool_use blocks and sets its follow-up state. It does not rely solely on stop_reason. With streaming tool execution enabled, complete tool calls enter the executor while the model stream continues, and completed results can be yielded immediately. Without tool follow-up, output-limit recovery, stop hooks, and token-budget handling can still continue or end the turn.

6. A tool call is a proposal, not an execution

When the model emits tool_use, the runtime collects those blocks and routes them into the tool layer. There are two alternative scheduling routes: the streaming executor collects its remaining results, or runTools() runs a batch after streaming. The latter groups only adjacent calls whose parsed arguments are concurrency-safe; it preserves serial barriers rather than gathering all reads ahead of writes. Bash delegates this decision to its read-only check, so shell commands are not categorically serial. Concurrency eligibility does not grant execution permission. runToolUse() finds the tool and handles unknown-tool and abort cases. Its execution wrapper parses the schema and applies tool-specific validation; the permission callback routes allow, deny, and ask decisions. Progress and tool_result are separate outputs of that process.

This is the most important mental model for the rest of the series: the model proposes an action; the runtime decides whether that action exists, whether the arguments are valid, whether policy allows it, and how the result should return to the next turn.

read A, read B → one batch if both calls are concurrency-safe
write A        → wait for the previous batch; execute alone
read A again   → begin after the write finishes
A scheduling example of contiguous batches. Eligibility depends on each invocation's parsed input.

7. Records are not one thing

The events shown on screen, the rows written to transcript storage, the content-replacement state, and the resume metadata serve different readers. The UI needs responsive rendering. The transcript stores filtered records for recovery; visible UI output alone is not proof of a completed disk write. Resume needs enough runtime state to reconstruct a workable session. Treating all of that as chat history hides the most interesting engineering work.

onQueryEvent updates REPL messages and handles display-specific compact markers or progress removal. Separately, recordTranscript() cleans messages, skips UUIDs already recorded, and inserts a new message chain. Replacement and compaction records have additional persistence paths.

For continuation, loadConversationForResume() reads old messages, file-history snapshots, replacements, collapse commits, and metadata. processResumedConversation() then restores applicable session, cost, worktree, replacement, and collapse state. Loading records and installing runtime state are separate steps.

Return to the failing-test example. Enqueueing means input is waiting; tool_use means the model proposed an action; a tool_result allows the runtime to inspect its outcome. Even a loop return reason of completed requires checking the test exit status and resulting change before calling the user's goal achieved. Stop handling defines turn control flow, not a proof of correctness.

8. Follow the same route into the other chapters

Reading route for the Claude Code source series with overview, memory, context, tools, permissions, capabilities, subagents, hooks, resume, and prompt cache
The rest of the series zooms into long-term information, runtime checks, recovery, and performance, but keeps returning to this same task route.
ChapterQuestionReturn point
OverviewHow does one task traverse the runtime?The complete route.
MemoryHow do durable instructions enter later requests?Memory files and user context.
ContextWhat happens before a full summary becomes necessary?The next request's selected messages.
ToolsHow are calls validated, ordered, and paired with results?Streaming and batch execution.
PermissionsWhen does a request allow, deny, or ask?Permission context and tool checks.
Commands, Skills, MCPHow do extensions become callable?Commands, MCP clients, and schemas.
Subagents and forkWhat is isolated and what prefix can be reused?AgentTool and worker context.
HooksHow do lifecycle callbacks change the running turn?Tool, stop, and compact events.
Transcript and resumeWhich records rebuild runnable state?Transcript, loader, and state restoration.
Prompt cacheHow do request shape and fork tails affect reuse?Cache markers, breakpoints, and microcompact.

9. What to remember before going deeper

First, every user-visible feature has a specific place in the runtime. Second, every model request contains a selected view, not a dump of all local state. Third, tool schemas and permission policy guard local side effects. Fourth, resume reconstructs runtime state instead of merely pasting old chat. With those facts in place, the remaining articles become much easier to read.

InvariantWhy it matters
Entry is not the loopCLI routing may return before an agent session starts.
Eligible input is orderedPriorities and FIFO order govern competing pending commands.
The REPL assembles the turnContext, tools, and callbacks are prepared before query.
Query connects the mechanismsContext selection, provider calls, streams, and follow-up share one loop.
Tool calls are proposalsExecution requires validation, permission, scheduling, and result handling.
Records serve different readersUI rendering, persistent rows, and resume state cannot be treated as one history.

Continue with the durable instruction layer, then return to the selected request view in the context-management chapter.

Sources

The source-code claims in this article are based on the public mirror and the linked official documentation. For server-side behavior and private feature switches, the article states only what the client sends or receives.