After this article: you should be able to tell the story of one run, explain why every observation must change the next decision, and design clear state, exit, budget, and retry rules.

Source note: the vendor-neutral payment-test example is compared with a fixed OpenAI Codex source snapshot verified on September 20, 2026. Figure states, done checks, and per-round checkpoints describe a recommended controller. Codex turn completion is not business-task acceptance.

1. Understand the loop through one task

1.1 Watch one task advance through four rounds

An Agent receives “fix the failing payment test.” It cannot responsibly edit the first plausible line and stop. A useful run might unfold like this:

  1. Round 1: read the failing log and discover that the error involves refunds for expired orders.
  2. Round 2: read the relevant function and callers, confirming that an expired-state check is missing.
  3. Round 3: make the smallest edit and run payment tests; they still fail, revealing another boundary case.
  4. Round 4: adjust the code and rerun tests. In this example, a configured completion check marks the task complete only after the target tests and relevant regressions pass.

Each round uses what the previous round discovered. If the Agent keeps issuing the same command after the same error, it is not progressing; it is only repeating.

1.2 Give this repeated process a name

An Agent Loop is the repeated process inside one run: prepare the current information, ask the model for the next step, execute an allowed action through the harness, observe the result, and decide whether to continue or stop.

Three roles are enough to understand the basic loop. The model proposes the next step. The harness turns an allowed proposal into a real action and observation. The loop controller carries state forward and decides whether another round is needed.

For a beginner, the loop first needs only three broad states: running, waiting, and stopped. Later we will split them into completed, blocked, failed, cancelled, and other precise states.

2. Make every round move toward completion

2.1 Every round must change the next decision

In plain language, each round must bring back a change for the next one: reading a file produces a new fact, running a test produces a new error, and editing code produces a new diff. The outer program organizes that result, updates progress, and prepares the next packet of material. Continuing is useful only when known facts, the plan, risk, or the done judgment has changed.

The table below only gives engineering names to the actions you already saw. Preparing material is Assemble, model judgment is Infer, action is Execute, organizing the result is Observe, and choosing whether to continue is Decide. It is not a second process.

Five Agent Loop steps with continue and stop branches; stopping does not prove task acceptance
An observation need not change the outside world, but it must update known facts, the plan, risk, or the done judgment. Otherwise the next iteration is probably expensive repetition.
StateInputNew fact it must produce
AssembleGoal, control state, recent observationNext model input and budgets
InferModel inputA proposal that respects the limits, or a final candidate
ExecuteAuthorized tool callTyped result, artifact, environment delta
ObserveResult and world deltaSuccess, failure, or unknown with a recorded source
DecideDone criteria, risk, budgetsContinue or an explicit exit reason

2.2 Stopping has more than two meanings

“The model returned final” is only a signal; it does not prove success. A recommended controller distinguishes verified completion, waiting for user input, blocked by permission, retryable failure, terminal failure, exhausted budget, or cancellation. Each result needs different saved information and a different next handler.

Seven recommended exit outcomes: verified complete, waiting for user, permission blocked, retryable failure, terminal failure, budget exhausted, and cancelled
Recommended task outcomes for the outer loop, UI, audit, and recovery to share; these labels do not map one-to-one to Codex event variants.
ExitWhat must be savedWho handles it next
CompletedDone check passed plus artifact or diffEvaluator or delivery flow
Waiting userMissing decision, options, default impactUser
BlockedRequired capability or permission and current scopeHarness permission component or approver
Retryable failureError class, attempt, backoff conditionCurrent loop or outer loop
Terminal failureUnrecoverable reason, checked in-flight effects, failure artifactOuter loop, evaluator, or person
Budget exhaustedCompleted, remaining, checkpointOuter loop or person
CancelledCancel source, in-flight action, cleanupHarness cleanup

Codex’s actual stopping decision is more specific than receiving final text. run_turn combines the need for further model sampling with pending input: tool results may need to go back to the model, or queued input may require another request. Reaching a context threshold can first trigger context management. Even without another tool call, end_turn: false in the server’s ResponseEvent::Completed requests further sampling. With no follow-up input, a Stop hook can still supply instructions to continue. This decides whether to call the model again; it does not automatically establish that payment tests and regressions passed.

Read terminal-event fields as well as names. The task runner’s TurnComplete may carry an error; aborts use TurnAborted. A finished turn, absence of a terminal error, and an accepted business task therefore need separate judgments. If the last observation only says that the test process is still running, a subsequent end to the turn supplies no evidence that the test passed.

2.3 Define “done” before the first action

Without explicit completion criteria, an agent substitutes fluent prose for completion. A bug fix may require the target test, relevant regressions, a scoped diff, and no unexplained effects. Research may require source coverage, citations for every conclusion, and marked conflicts. In a controller you design, store those criteria before the loop starts and check them after observations. A general-purpose Agent client cannot automatically know every repository’s acceptance rules.

The model may propose “I believe this is done.” Deterministic checks should run in the harness; subjective quality can go to an independent grader or person. Reflection helps, but one component should not be contestant and sole judge.

Put the three checks on different time scales and their jobs become clearer. An Agent Loop with a verifier asks whether the available evidence meets this run’s acceptance criteria. The Outer Loop reads test, CI, or review results and asks, “may this long-lived work item end?” Evals repeat many tasks and ask, “did this system version become reliably better?” They check one run, one work item, and one system version respectively.

2.4 Why an Agent cannot run forever

A loop design should manage task-appropriate time, model-call, tool-call, token, external-API, concurrency, and risk budgets. A maximum iteration count treats a cheap read and an expensive deployment equally. Better controllers estimate value and cost before each action and reserve capacity for validation.

  • Near a limit, narrow search, reduce parallelism, or request a choice instead of stopping abruptly.
  • Reserve an independent verification budget so editing cannot consume the ability to test.
  • Risky writes require a separate permission check; more tokens cannot compensate for missing permission.
  • Budget exhaustion produces a checkpoint and exit reason, never a disguised completion.

2.5 A retry must introduce a change

Retry is meaningful only if input, environment, strategy, or time changes. A transient connection failure can justify retrying the same request after a delay; a deterministic failing test assertion needs new diagnosis or a code change. These are different reasons to try again. Classify transient, invalid input, permission, not found, conflict, invariant violation, and unknown errors to choose recovery.

Recovery manual for transient errors, invalid input, permission limits, state conflicts, broken invariants, and unknown state
Recovery strategy belongs to an error class, not to one undifferentiated “failure” bucket. These are suggested actions; inspect the result after retrying, because success is not guaranteed.
Error classUseful changeDo not
TransientBackoff, jitter, bounded retryReplay at full speed
Invalid inputRead schema and repair argumentsRetry identical arguments
PermissionRequest explicit permission or use a read-only pathBypass the permission check
Conflict / staleRefresh state and replanOverwrite new truth
Invariant failureRollback, shrink the change, escalateStack more changes
UnknownSave artifacts, stop, or probe in isolationExplain forever

Codex puts model-connection recovery in handle_response_stream_error. The error determines whether retry is possible and how long to wait. The normal path follows the provider’s stream-retry configuration and may switch transport. A separate feature-gated connection-retry branch, when its conditions hold, is not bounded by that retry count, so bounded retries are not a guarantee across all configurations. run_sampling_request rereads current history before retrying; this does not instruct the system to blindly repeat a payment test or external refund that already ran. Model reconnection, tool retries, and reconciliation of business effects need separate designs.

3. Advanced: expand three broad states into a recoverable state machine

By this point, a beginner can review one run through evidence changes, done criteria, exit reasons, budgets, and meaningful retries. The following turns those judgments into a recommended durable controller; it does not reproduce Codex’s internal state structure.

while (toolCalls.length) expresses a syntax condition, not an engineering state. Depending on product needs, distinguish running, waiting approval, waiting user, blocked, completed, failed, cancelled, and budget exhausted. State determines who may wake the run, which resources remain valid, and whether an outer loop can retry safely.

Before naming fields, follow one payment-task round in this recommended design:

  1. The runtime begins by reading the goal, the previous test failure, and remaining budgets.
  2. The assembler puts those facts into the current model request, and the model proposes one next action with a call id.
  3. The harness validates and executes the tool, producing a new observation; the proposal itself does not directly change run state.
  4. The controller writes the observation into state and updates progress, budgets, repeated-failure counts, and any candidate exit reason.
  5. Only after checkpointing does it start another round. A passed done check, required approval, or lack of meaningful change selects a different exit instead.
{
  "run_id": "run_18",
  "state": "running",
  "iteration": 7,
  "goal": { "id": "fix-payment-test", "done_check": "test://payment" },
  "last_observation": { "kind": "test_failure", "ref": "log://9f2" },
  "budgets": { "tool_calls_left": 12, "time_s_left": 480 },
  "progress": { "changed_files": 2, "same_failure_count": 1 }
}
Shape-level example: loop state is not the message array. It is deterministic control state that the runtime can inspect and recover.
while run.state == "running":
    view = assemble_context(run)
    proposal = model.infer(view)
    observation = harness.dispatch(proposal, run)
    run = transition(run, observation)

    if run.state != "running":
        checkpoint(run)
        break

    if done_check(run):
        run.state = "completed"
    elif repeated_without_change(run):
        run.state = "blocked"
    checkpoint(run)
Teaching controller: the application implements this acceptance and persistence policy; policy.ask moves running to waiting_approval; approval_granted wakes it back to running; only a test observation that satisfies the done check enters completed. In this design, model final text alone cannot complete the task. A failed checkpoint stops progress and reports an error; it does not make external effects and local state writes atomic.

4. Review the health of one run

Review pointHealthy signalDanger signal
ProgressNew facts or a state delta every iterationSame calls and errors repeat
DoneDefined before, verified afterFinal-answer tone
ExitMachine-readable reason and checkpointSuccess/failure boolean only
BudgetMultiple budgets, dynamic, verification reservedOnly max iterations
RecoveryStrategy changes by error classRetry every error
HumanIntervenes for judgment or extra permissionApproves everything or never appears

The agent loop makes one run approach completion through real results. Next we lengthen the horizon: when work requires triggers, queues, checkpoints, and handoffs across runs, design the outer loop.

Official sources