In this series, an agent is a work system made from a model, an outer program, and tools. The model chooses a next step. The program prepares information, executes real actions, and records results. Tools let the system read files, edit code, and run tests.
After this part, you should be able to answer: what must the model know, who executes actions, why the system repeats, who carries work after a run ends, who confirms completion, and how a team proves that a layer change actually improved the system?
Source note: this article combines official OpenAI and Anthropic material with a pinned Codex source snapshot. The six responsibilities are a teaching and design framework, not six separate modules guaranteed in every product. Diagrams and teaching JSON describe recommended designs; concrete source behavior is identified separately. Ending a run, passing checks, and accepting a task are distinct decisions.
1. Build the complete picture from one task
1.1 Follow one agent task from start to finish
Suppose a payment test fails. The task is: “Find the cause, make the smallest safe change, run the payment tests, and leave evidence for review.” A target workflow with independent acceptance looks like this:
- The system gives the model the task requirements.
- It prepares the failure log, relevant code, and project rules so the model sees the current facts.
- The model proposes which file to read or which test command to run.
- The outer program checks that the action exists, its arguments are valid, and permission allows it before execution.
- The file content or test result returns to the model as new evidence.
- The model keeps reading, editing, and verifying until it finishes or can clearly explain why it cannot continue.
- Tests, static checks, or a reviewer independently confirm the result instead of trusting “fixed” on its own.

1.2 Name the six problems along that path
AI Engineering begins with six ordinary questions, not six terms to memorize.
1.2.1 How do we make the requirements clear?
Writing the goal, limits, completion checks, and delivery format as an explicit work instruction is prompt engineering.
1.2.2 What should the model see right now?
Investigation needs the failure log; editing needs relevant code; verification needs the current diff and latest test output. Selecting that material is context engineering.
1.2.3 How does a proposed action really happen?
“Run the tests” is only a proposal. The software that checks arguments and permission, isolates execution, runs the tool, and stores results is the runtime harness.
1.2.4 How can the agent act, observe, and act again?
The outer program repeatedly asks the model for a step, executes it, returns the result, and decides whether to continue. That within-run cycle is the agent loop.
1.2.5 How does work continue after one run ends?
A run may stop for time, budget, or a human decision. The cross-run process that saves progress, schedules another run, limits retries, and prevents duplicate work is the outer loop.
1.2.6 Who proves the system really succeeded?
Repeatable tests, rules, model judges, and human review form the system's exam. Those mechanisms are evals.
If you only need the beginner’s map, you already have the main story. The second half restates it in engineering language for readers who need to review APIs, responsibilities, and the source of a failure.
2. Compress the story into an engineering map
2.1 Put the six responsibilities on their time scales
Step back from the story and the responsibilities last for different lengths of time. The model judges one next step at a time. Reading, editing, and testing repeat within one run. The same task card may survive several runs. Evals compare many tasks and system versions.
Engineering gives those scales names. One model call is an inference. One “model decision–tool action–returned result” exchange is a tool cycle. One Agent execution from start to completion or interruption is a run. A task card that survives several runs is a work item. The table below only compresses the six responsibilities you have already learned onto those scales.

| System part | What it does | Primary time scale | Typical failure |
|---|---|---|---|
| Prompt | States the model's role, goal, limits, instruction precedence, output format, and tool semantics. | Behavior for an inference | Ambiguous intent, conflicting rules, output drift |
| Context | The working set actually visible to and usable by the model in one call. | Reassembled per inference | Missing, noisy, stale, or poorly placed evidence |
| Runtime harness | Assembles requests, checks and executes tools, and records one run outside the model. | One run | Uncontrolled effects, broken state, no recovery |
| Agent loop | Repeat reasoning, action, and observation until an exit condition is met. | One run, potentially many inferences | Premature stop, spinning, exhausted budget |
| Outer loop | Trigger, select, verify, persist, retry, and escalate across runs. | Multiple runs | Duplicate work, lost state, retry storms |
| Evals | Repeatable outcome truth and feedback for system change. | System lifecycle | Treating “the agent said done” as done |
These are not mutually exclusive job titles. One change can touch prompt, context policy, and harness. A review still needs to name each changed part; otherwise the team cannot tell what improved or where a regression began.
2.2 Ask who decides each step and where the result is stored
Ask “who has the final say?” at each point in the seven-step story. The user decides the goal. The runtime prepares material and checks tool requests. The repository and test process hold real results. Durable storage keeps progress across runs. An evaluator accepts or rejects the checks. A task passes through several responsible components, so the model does not decide its overall state alone.
2.2.1 Turn a goal into verifiable done
“Fix the test” gives direction but leaves the allowed work unspecified. May production code change? Which commands must pass? Is network access allowed? What results should a reviewer receive? The prompt states those requirements. Evals turn the verifiable ones into machine or human checks. They serve the same goal but do different jobs.
2.2.2 Select stored material for one model call
OpenAI’s official prompt engineering guide discusses high-level instructions, message roles, tools, and input as parts of request design. Anthropic’s context engineering article emphasizes that an agent chooses what belongs in the next inference at every step. A useful conceptual request shape is:
{
"instructions": "Fix the failing payment test; do not change public APIs.",
"tools": ["read_file", "search", "apply_patch", "run_tests"],
"input": [
{ "type": "task", "text": "..." },
{ "type": "repo_facts", "ref": "selected working set" }
]
}A repository may store ten thousand facts, but only selected and serialized material in the current request is model-visible context. Instructions also consume the context window. Prompt is a semantic layer cut by responsibility; context is a physical working set cut by visibility. They overlap, but they are not substitutes.
2.2.3 A tool call is a proposal; the harness executes the effect
When a model emits “run the tests” or a structured tool call, no process has to start automatically. A harness validates arguments and permission for the specific tool, selects an execution environment, handles timeouts and cancellation, bounds results, records events, and returns an observation. Approval and isolation are separate controls: permission to execute does not remove a sandbox, and the absence of an approval prompt does not prove that execution is sandboxed. The tool and configuration determine which checks apply; part four follows that path in source.
Official usage varies in scope. OpenAI’s Harness engineering discusses the Codex harness as the core agent loop plus its tool and instruction environment. Anthropic’s Building and evaluating trustworthy agents stresses the instructions and guardrails that form a harness, while Managed agents separates session, harness, and sandbox. For sharper reviews, this series uses a narrower definition: the runtime harness is the execution and control plane outside the model and inside one run; the loop is the control flow that advances state within it.
2.3 Separate one run from a long-lived task
OpenAI’s public walk-through of the Codex agent loop shows a single-run path: prepare input, call the model, receive events, execute tools, return results, and continue until completion or interruption. Anthropic’s Building effective agents likewise describes agents as systems where models dynamically direct their own processes and tool use. This is the inner loop.
Codex app-server makes this distinction concrete. A successful turn/start response contains status: "inProgress": it acknowledges the run, not passing tests. run_turn then checks whether the model needs another step or input is pending. When neither applies, configurable Stop hooks can still affect termination. These conditions govern a turn's exit; they do not independently prove that payment tests and regressions passed.
turn/start succeeds → inProgress: request accepted
Tool result returns → model decides again: still running
turn/completed → inspect final status / error: this run ended
Tests, regressions, and diff-scope checks pass → delivery accepts the resultA production system has another set of questions. Who triggers the next run? Which work item is selected? What recoverable state survives? After verification fails, do we retry, change strategy, or escalate to a human? This cross-run control is the outer loop in this series.
| Question | Inner / agent loop | Outer loop |
|---|---|---|
| Where input comes from | Current task and observations | Events, plans, queues, or people |
| How long state lasts | Messages and tool results during execution; session records can also persist | Work items, scheduling decisions, and recoverable progress across runs |
| When it stops | Completion, error, budget, or cancellation | Goal reached, retry cap, time window, or escalation |
| Main risk | Infinite tool use, bad observations, premature final | Duplicate triggers, state drift, failures that no component handles |
Collapsing both into one “loop” creates predictable mistakes. The inner loop writes a polished completion message while the outer loop has no independent verification. Or the outer loop keeps replaying the same failure without changing evidence, strategy, or authority. A retry is repetition, not learning.
3. Use the map to diagnose real systems
3.1 Route a failure to the layer that can control it
A symptom may pass through several components, but the fix should land near the source of the problem. If the model repeatedly ignores a stable constraint, inspect prompt precedence. If it lacks a recently changed API, inspect context selection. If it runs a dangerous command, inspect the harness permission check. If it finds the fix but never tests it, inspect loop termination. If it claims tests passed when they did not, inspect which results the evaluator reads.

3.1.1 Change only the component that can control the problem
- Put universal rules in stable instructions rather than retrieving them afresh.
- Let context policy select changing facts instead of freezing them into the prompt.
- Control side effects with permissions, sandboxes, and deterministic checks—not a plea to “be careful.”
- When done is executable, make the loop or evaluator execute the check instead of trusting self-report.
- When failure accumulates across runs, give checkpoints, retry budgets, and escalation to the outer loop.
3.2 Return to real coding agents
The map is not detached from products. Parts two through five inspect instructions, context, tool controls, and the loop in Codex. Part six examines cross-run scheduling in Symphony; part seven examines samples, grading, and records in open-source Evals. These are slices of different projects, not one assembled product. Other source articles on this site offer deeper reading:
| Engineering question | Continue with | What to observe |
|---|---|---|
| How does a task enter a model–tool loop? | Codex runtime, Claude Code runtime | Request assembly, event streams, tool-result feedback |
| What does the model see each turn? | Codex context, Claude Code context | History selection, serialization, compaction, and persistence |
| Where are effects constrained? | Codex permissions and sandbox, Claude Code permissions | Proposal, validation, approval, execution, observation |
| How does long-running work recover? | Transcript and resume, Hermes runtime loop | Durable fact versus the next-run working set |
Source code grounds the abstraction but does not replace system definitions. The rest of this series follows one source rule: official documentation states the product's public behavior; when a chapter enters a concrete implementation, public code shows how the program produces that behavior; our diagrams explain the mechanism between them.
3.3 After a change, rerun the same task and checks
Finding the cause is only the beginning. Suppose the payment task failed because a fresh test result never entered the next model request, so the team changes context selection. A higher retrieval score does not prove success. Run the same task from the same starting point and confirm that the new result reaches the model, changes the Agent’s next decision, and does not regress neighboring tasks.
- Execute first: hold the task’s starting state, model, tools, permissions, and budgets constant; record the old outcome and trace.
- Localize next: use logs, tool results, and eval output to identify the responsible component instead of editing the prompt by reflex.
- Change one primary component: for example, the context-selection policy, with an observable expected result.
- Rerun comparably: check the target failure, critical regressions, cost, and latency under the same conditions.
- Adopt last: a validated candidate may be released, while the delivery process still decides rollout and rollback.
The complete AI Engineering move is not “see a problem, add a rule.” It is execute, observe, find the responsible component, change it, rerun, and retain only changes that pass the checks. The final chapter turns this process into repeatable evals.
4. Review the full system with six questions
| If you are deciding… | Ask first… | Primary component |
|---|---|---|
| How the model should behave | Are the goal, limits, instruction precedence, and output format explicit? | Prompt |
| Which facts belong in this call | What is relevant, current, trustworthy, and worth the budget? | Context |
| Whether an action may occur | Who validates, authorizes, executes, and recovers? | Harness |
| How one run converges | Are observations reliable, and what are the stop and budget rules? | Agent loop |
| How runs hand work off | Which component handles triggers, checkpoints, verification, retries, and escalation? | Outer loop |
| Why we believe the system improved | Where are outcome truth, failure classes, and regressions? | Evals |
The point is not that prompt engineering is obsolete. It is that prompt engineering now has a precise place in a larger system. Next, we turn the prompt from one-off wording into a versioned behavioral interface.
Official sources
- Codex: runtime loop source
- Codex: turn request handling
- OpenAI: Prompt engineering
- Anthropic: Effective context engineering for AI agents
- OpenAI: Harness engineering
- OpenAI: Unrolling the Codex agent loop
- Anthropic: Building effective agents
- Anthropic: Building and evaluating trustworthy agents
- Anthropic: Managed agents
