After this article: you should be able to tell the story of one run, explain why every observation must change the next decision, and design clear state, exit, budget, and retry rules.
Source note: the vendor-neutral payment-test example is compared with a fixed OpenAI Codex source snapshot verified on September 20, 2026. Figure states, done checks, and per-round checkpoints describe a recommended controller. Codex turn completion is not business-task acceptance.
1. Understand the loop through one task
1.1 Watch one task advance through four rounds
An Agent receives “fix the failing payment test.” It cannot responsibly edit the first plausible line and stop. A useful run might unfold like this:
- Round 1: read the failing log and discover that the error involves refunds for expired orders.
- Round 2: read the relevant function and callers, confirming that an expired-state check is missing.
- Round 3: make the smallest edit and run payment tests; they still fail, revealing another boundary case.
- Round 4: adjust the code and rerun tests. In this example, a configured completion check marks the task complete only after the target tests and relevant regressions pass.
Each round uses what the previous round discovered. If the Agent keeps issuing the same command after the same error, it is not progressing; it is only repeating.
1.2 Give this repeated process a name
An Agent Loop is the repeated process inside one run: prepare the current information, ask the model for the next step, execute an allowed action through the harness, observe the result, and decide whether to continue or stop.
Three roles are enough to understand the basic loop. The model proposes the next step. The harness turns an allowed proposal into a real action and observation. The loop controller carries state forward and decides whether another round is needed.
For a beginner, the loop first needs only three broad states: running, waiting, and stopped. Later we will split them into completed, blocked, failed, cancelled, and other precise states.
2. Make every round move toward completion
2.1 Every round must change the next decision
In plain language, each round must bring back a change for the next one: reading a file produces a new fact, running a test produces a new error, and editing code produces a new diff. The outer program organizes that result, updates progress, and prepares the next packet of material. Continuing is useful only when known facts, the plan, risk, or the done judgment has changed.
The table below only gives engineering names to the actions you already saw. Preparing material is Assemble, model judgment is Infer, action is Execute, organizing the result is Observe, and choosing whether to continue is Decide. It is not a second process.

| State | Input | New fact it must produce |
|---|---|---|
| Assemble | Goal, control state, recent observation | Next model input and budgets |
| Infer | Model input | A proposal that respects the limits, or a final candidate |
| Execute | Authorized tool call | Typed result, artifact, environment delta |
| Observe | Result and world delta | Success, failure, or unknown with a recorded source |
| Decide | Done criteria, risk, budgets | Continue or an explicit exit reason |
2.2 Stopping has more than two meanings
“The model returned final” is only a signal; it does not prove success. A recommended controller distinguishes verified completion, waiting for user input, blocked by permission, retryable failure, terminal failure, exhausted budget, or cancellation. Each result needs different saved information and a different next handler.

| Exit | What must be saved | Who handles it next |
|---|---|---|
| Completed | Done check passed plus artifact or diff | Evaluator or delivery flow |
| Waiting user | Missing decision, options, default impact | User |
| Blocked | Required capability or permission and current scope | Harness permission component or approver |
| Retryable failure | Error class, attempt, backoff condition | Current loop or outer loop |
| Terminal failure | Unrecoverable reason, checked in-flight effects, failure artifact | Outer loop, evaluator, or person |
| Budget exhausted | Completed, remaining, checkpoint | Outer loop or person |
| Cancelled | Cancel source, in-flight action, cleanup | Harness cleanup |
Codex’s actual stopping decision is more specific than receiving final text. run_turn combines the need for further model sampling with pending input: tool results may need to go back to the model, or queued input may require another request. Reaching a context threshold can first trigger context management. Even without another tool call, end_turn: false in the server’s ResponseEvent::Completed requests further sampling. With no follow-up input, a Stop hook can still supply instructions to continue. This decides whether to call the model again; it does not automatically establish that payment tests and regressions passed.
Read terminal-event fields as well as names. The task runner’s TurnComplete may carry an error; aborts use TurnAborted. A finished turn, absence of a terminal error, and an accepted business task therefore need separate judgments. If the last observation only says that the test process is still running, a subsequent end to the turn supplies no evidence that the test passed.
2.3 Define “done” before the first action
Without explicit completion criteria, an agent substitutes fluent prose for completion. A bug fix may require the target test, relevant regressions, a scoped diff, and no unexplained effects. Research may require source coverage, citations for every conclusion, and marked conflicts. In a controller you design, store those criteria before the loop starts and check them after observations. A general-purpose Agent client cannot automatically know every repository’s acceptance rules.
The model may propose “I believe this is done.” Deterministic checks should run in the harness; subjective quality can go to an independent grader or person. Reflection helps, but one component should not be contestant and sole judge.
Put the three checks on different time scales and their jobs become clearer. An Agent Loop with a verifier asks whether the available evidence meets this run’s acceptance criteria. The Outer Loop reads test, CI, or review results and asks, “may this long-lived work item end?” Evals repeat many tasks and ask, “did this system version become reliably better?” They check one run, one work item, and one system version respectively.
2.4 Why an Agent cannot run forever
A loop design should manage task-appropriate time, model-call, tool-call, token, external-API, concurrency, and risk budgets. A maximum iteration count treats a cheap read and an expensive deployment equally. Better controllers estimate value and cost before each action and reserve capacity for validation.
- Near a limit, narrow search, reduce parallelism, or request a choice instead of stopping abruptly.
- Reserve an independent verification budget so editing cannot consume the ability to test.
- Risky writes require a separate permission check; more tokens cannot compensate for missing permission.
- Budget exhaustion produces a checkpoint and exit reason, never a disguised completion.
2.5 A retry must introduce a change
Retry is meaningful only if input, environment, strategy, or time changes. A transient connection failure can justify retrying the same request after a delay; a deterministic failing test assertion needs new diagnosis or a code change. These are different reasons to try again. Classify transient, invalid input, permission, not found, conflict, invariant violation, and unknown errors to choose recovery.

| Error class | Useful change | Do not |
|---|---|---|
| Transient | Backoff, jitter, bounded retry | Replay at full speed |
| Invalid input | Read schema and repair arguments | Retry identical arguments |
| Permission | Request explicit permission or use a read-only path | Bypass the permission check |
| Conflict / stale | Refresh state and replan | Overwrite new truth |
| Invariant failure | Rollback, shrink the change, escalate | Stack more changes |
| Unknown | Save artifacts, stop, or probe in isolation | Explain forever |
Codex puts model-connection recovery in handle_response_stream_error. The error determines whether retry is possible and how long to wait. The normal path follows the provider’s stream-retry configuration and may switch transport. A separate feature-gated connection-retry branch, when its conditions hold, is not bounded by that retry count, so bounded retries are not a guarantee across all configurations. run_sampling_request rereads current history before retrying; this does not instruct the system to blindly repeat a payment test or external refund that already ran. Model reconnection, tool retries, and reconciliation of business effects need separate designs.
3. Advanced: expand three broad states into a recoverable state machine
By this point, a beginner can review one run through evidence changes, done criteria, exit reasons, budgets, and meaningful retries. The following turns those judgments into a recommended durable controller; it does not reproduce Codex’s internal state structure.
while (toolCalls.length) expresses a syntax condition, not an engineering state. Depending on product needs, distinguish running, waiting approval, waiting user, blocked, completed, failed, cancelled, and budget exhausted. State determines who may wake the run, which resources remain valid, and whether an outer loop can retry safely.
Before naming fields, follow one payment-task round in this recommended design:
- The runtime begins by reading the goal, the previous test failure, and remaining budgets.
- The assembler puts those facts into the current model request, and the model proposes one next action with a call id.
- The harness validates and executes the tool, producing a new observation; the proposal itself does not directly change run state.
- The controller writes the observation into state and updates progress, budgets, repeated-failure counts, and any candidate exit reason.
- Only after checkpointing does it start another round. A passed done check, required approval, or lack of meaningful change selects a different exit instead.
{
"run_id": "run_18",
"state": "running",
"iteration": 7,
"goal": { "id": "fix-payment-test", "done_check": "test://payment" },
"last_observation": { "kind": "test_failure", "ref": "log://9f2" },
"budgets": { "tool_calls_left": 12, "time_s_left": 480 },
"progress": { "changed_files": 2, "same_failure_count": 1 }
}while run.state == "running":
view = assemble_context(run)
proposal = model.infer(view)
observation = harness.dispatch(proposal, run)
run = transition(run, observation)
if run.state != "running":
checkpoint(run)
break
if done_check(run):
run.state = "completed"
elif repeated_without_change(run):
run.state = "blocked"
checkpoint(run)policy.ask moves running to waiting_approval; approval_granted wakes it back to running; only a test observation that satisfies the done check enters completed. In this design, model final text alone cannot complete the task. A failed checkpoint stops progress and reports an error; it does not make external effects and local state writes atomic.4. Review the health of one run
| Review point | Healthy signal | Danger signal |
|---|---|---|
| Progress | New facts or a state delta every iteration | Same calls and errors repeat |
| Done | Defined before, verified after | Final-answer tone |
| Exit | Machine-readable reason and checkpoint | Success/failure boolean only |
| Budget | Multiple budgets, dynamic, verification reserved | Only max iterations |
| Recovery | Strategy changes by error class | Retry every error |
| Human | Intervenes for judgment or extra permission | Approves everything or never appears |
The agent loop makes one run approach completion through real results. Next we lengthen the horizon: when work requires triggers, queues, checkpoints, and handoffs across runs, design the outer loop.
Official sources
- run_turn: continuation and stopping
- TurnComplete: terminal turn event
- handle_response_stream_error: connection recovery
- OpenAI: Unrolling the Codex agent loop
- OpenAI: A practical guide to building agents
- Anthropic: Building effective agents
- Anthropic: Managed agents