After this article: you should be able to trace a proposed tool call into a controlled action and explain the roles of tools, permission, sandbox, state, recovery, and observation.

Terminology note: vendors use “harness” at slightly different widths. This series uses it for the execution environment and controls outside the model but inside one run. The repeated control flow is introduced separately as the Agent Loop. The payment-test path is an engineering design example, compared with a fixed OpenAI Codex source snapshot verified on September 20, 2026. Seven checks, artifacts, idempotency, and recovery states are recommendations, not a protocol implemented by every tool.

1. See how a proposal becomes a real action

1.1 First understand what a harness does

A harness is the software environment that lets a model use tools under explicit rules. It prepares the request, exposes available tools, checks arguments and permissions, runs actions in a limited environment, records what happened, and returns observations to the model.

Think of a workshop. The model is the worker deciding what to try. The harness supplies labeled tools, a protected workbench, safety rules, a logbook, and a way to resume after interruption. A more capable worker does not remove the need for the workshop.

1.2 Follow “run the payment tests” from words to evidence

When the model proposes a test command, the design should answer seven questions; real code may combine steps or skip an approval request according to policy:

  1. Find the real test tool and its current interface.
  2. Check that the arguments are complete and valid.
  3. Decide whether this Agent may use it for this repository.
  4. Limit which files, processes, network locations, and credentials it can reach.
  5. Run the command with timeout and cancellation controls.
  6. Store large logs and return a useful result.
  7. Record state for later reconciliation; a log alone does not guarantee recovery of external effects.

If step 2 is missing, malformed input reaches the executor. If step 3 is missing, a prompt warning is the only protection against unauthorized actions. If step 6 is missing, a huge log floods the next model request. Harness engineering is the design of this whole path, not merely a wrapper around an API call.

Put this Tool Call back into a complete run and a repeating process appears: the harness prepares a request → the model proposes an action → the harness checks and executes it → the result becomes an observation → the observation enters the next request → the model decides again. The harness checks and executes each action; the Agent Loop decides whether another round is needed.

2. Separate the parts of the action environment

2.1 Gather the seven steps into one runtime environment

Models are good at proposing plans under uncertainty. The runtime must handle those uncertain proposals with deterministic code. Which tools are visible to the current role, whether arguments satisfy a schema, which paths are writable, which network destinations are reachable, and how results are serialized must be defined by policy and code—not improvised by the model.

A minimal harness only needs to prepare a request, expose tools, validate input and permission, execute within limits, and return results. A mature system adds a versioned tool registry, event records, progress observation, artifacts, and checkpoints for reproducibility and recovery. The table below puts both layers together for engineering review.

Harness duties: the model exchanges results with controlled tools while state records support recovery reconciliation
The model is a decision component, not the whole agent. The harness lets decisions touch the world safely.
ComponentWhat it records or decidesWhy the model cannot decide it
Request assemblerModel, instructions, tools, current context, budgetsRequests must be reproducible and observable
Tool registryName, schema, capability, version, result typeThe model cannot invent real interfaces
Policy / permissionAllow, deny, ask, and accessible scopePermission decisions must be repeatable
Sandbox / executorFiles, processes, network, and resource limitsPrompt text cannot isolate effects
State / transcriptRun state, events, result references, checkpointsRecovery and audit need durable records
Observer / verifierProgress, metrics, tests, and execution resultsSelf-report cannot be the only truth

The harness records and provides tool results, log references, and test output. The agent loop decides whether this run should continue; evals compare quality across tasks and versions. A verifier may run tests inside the harness, but one test does not represent every product requirement.

2.2 Pass one tool call through seven checks and actions

A tool call is structured model output: a proposal. A robust harness does not map it directly to a function. It separates interpretation, authority, execution, and observation.

A manual of seven tool-call duties: parse, resolve, validate, authorize, isolate by configuration, execute, and normalize
Design goal: correlate important checks and state changes with events. These seven entries are a responsibility checklist, not seven implementation stages that every Codex call must traverse.
  1. Parse: confirm the call is complete and belongs to the current request.
  2. Resolve: find the exact tool version in the current visible registry.
  3. Validate: check types, ranges, exclusions, and required fields.
  4. Authorize: combine actor, resource, action, and environment into allow / deny / ask.
  5. Isolate: limit filesystem, process, network, and credential access.
  6. Execute: manage timeout, cancellation, concurrency, retries, and idempotency keys.
  7. Normalize: turn stdout, structured results, errors, and artifacts into a stable observation.
async function dispatch(call, run) {
  const tool = registry.resolve(call.name, run.toolsetVersion)
  if (!tool) return observe("unknown_tool", call.id)

  const args = tool.schema.parse(call.arguments)
  const decision = policy.decide(run.actor, tool.capability, args)
  if (decision === "deny") return observe("denied", call.id)
  if (decision === "ask") return pauseForApproval(run, call)

  const box = await sandbox.open(tool.limits)
  events.append("tool.started", { runId: run.id, callId: call.id })
  try {
    const raw = await executor.run(tool, args, { sandbox: box })
    return normalize(raw, artifactStore)
  } finally {
    await box.close()
  }
}
Teaching pseudocode: tool-specific validation, asynchronous completion, and cancellation branches are omitted; this handles one tool proposal; it does not own a while loop. Calling the model again belongs to the agent loop in Part 5.
{
  "call_id": "call_42",
  "tool": "run_tests@3",
  "arguments": { "target": "./payment", "mode": "read-write" },
  "policy": { "decision": "allow", "rule": "repo-test" },
  "sandbox": { "fs": "workspace", "network": "off", "timeout_s": 120 },
  "result": { "status": "failed", "exit_code": 1, "artifact_ref": "log://9f2" }
}
Shape-level example: bind decision, execution environment, and result evidence to one call id so an effect remains traceable.

In this recommended design, the same call_id should associate three records for different uses. The artifact store keeps the full log. An event log keeps states such as proposed, allowed, started, and completed. The next model request receives only the observation needed for the next decision. During recovery, the runtime reads the event log, queries the real environment, and opens the artifact by reference when needed; it does not infer whether an effect happened from chat text.

Codex also separates a returned tool result from a terminated process. Its command-result header distinguishes Process running with session ID … from Process exited with code …. A payment test returning a session identifier is still running; the controller must read subsequent output and obtain the exit code and test result before deciding that it passed. Approval, process start, and test completion are separate facts.

Multiple calls do not always traverse the path serially. ToolCallRuntime reads each tool’s parallel capability: parallel calls acquire a shared read lock, while other calls acquire an exclusive write lock before entering the handler. This gate controls dispatch within that runtime; it does not prove business independence or cover every process already left running in the background. Two independent reads can run concurrently, but testing the payment code still depends on completing its edit first.

2.3 Tool design shapes what the Agent can do

Part 2 covered the tool description the model sees and how it should choose. This part covers the registry, version, schema validator, authorization, dispatcher, and artifacts that decide whether execution can actually happen. More tools are not automatically better: overlapping write tools, a universal shell, and vague free-text arguments expand the action space and force the model to guess.

2.3.1 A good tool serves the model and the runtime

DimensionModel interfaceRuntime interface
SemanticsWhen to call and when not toCapability id and version
InputFew, clear fieldsStrict schema and normalization
OutputEnough to choose the next stepTyped result, artifact, and error class
EffectWhat will changePermission, idempotency, compensation, audit
CostWhen the call is worth makingTimeout, rate, token, and resource budgets

Do not dump large output back into context. Store full logs as artifacts and return a summary, critical errors, and a reference that can be paged later. This improves both context budget and recovery.

2.4 Permission to act is not permission to reach everything

Permission answers whether this actor may perform this action. The sandbox answers what the allowed action can actually reach. Permission to run tests does not imply access to the home directory. Permission to read a repository does not imply permission to transmit it to arbitrary network destinations. Both layers are required.

Approval permits an action, sandbox configuration limits execution, and audit records outcomes; approval is policy-dependent
Approval, isolation, and recording have separate roles: policy triggers approval, execution configuration sets sandbox scope, and audit retains evidence without granting or revoking permission.
  • Capability scope: do not expose irrelevant tools to the model.
  • Policy decision: match actor, action, resource, and environment.
  • Approval: follow an explicit policy and show the action, target, and relevant diff; configured automatic review may also handle approval.
  • Sandbox: limit worst-case impact with filesystem, network, process, and credential isolation.
  • Audit: record who proposed the effect, which rule allowed it, and what actually occurred.

Codex’s ToolOrchestrator::run classifies approval requirements as skip, forbidden, or needs approval; strict automatic review can additionally review an action that would normally skip approval. Sandbox selection follows. The approval and sandbox selection show that approval does not automatically remove isolation, while configuration can permit execution without a process sandbox. Retrying a sandbox denial with escalation depends on tool support, approval policy, network decisions, and allowed scope; failure is not a general instruction to run again without isolation. Define the directories and network access needed by the payment test, then inspect the effective policy to determine the path.

3. Advanced: make real actions recoverable and traceable

3.1 After interruption, chat history alone cannot recover the work

Recovery design needs to distinguish UI conversation, model-visible messages, append-only events, external artifacts, checkpoints, and environment snapshots. They have different lifetimes and recovery roles.

Restarting by sending old chat back to a model is unsafe. The runtime must know which tool calls actually executed, which were in flight, how the workspace changed, whether approvals remain valid, and which external objects already exist. Reconcile execution records with queries to the external system rather than model guesses. Resuming a conversation is not transactional recovery of a command or payment.

tool.proposed → tool.validated → policy.allowed → tool.started
→ runtime.crashed → run.recovered → effect.reconciled
→ observation.ready
Recommended recovery trace: if call_42 has started but not completed, the runtime first queries the environment. A read-only call may be replayed under limits; an idempotent write is reconciled first; a non-idempotent write with unknown state must block or escalate rather than execute again.

Persistence itself has stages. In Codex’s RolloutRecorder, record_canonical_items enqueues records; persist and flush wait for the writer’s acknowledgment and return failures to their caller. Accepting a record and confirming its write are different events. This stores conversation records without turning arbitrary tool effects and record writes into one transaction. Recovering payment tests still requires checking files and processes; real refunds additionally need business-level idempotency and reconciliation.

3.2 Both the user and the system need to see what happened

A final response and total token count are not enough to debug a harness. A run trace needs request assembly, model events, tool proposal, policy decision, sandbox execution, observation, checkpoint, and exit reason. Sensitive bodies can be redacted; control events cannot disappear.

  • Stable run, turn, and call ids join every event.
  • Record tool latency, queue time, retry, cancellation, and artifact size.
  • Separate model error, tool error, policy denial, environment failure, and user cancellation.
  • Retain the exit reason and unfinished work so the outer loop can decide safely.

4. Review the harness with six questions

Review questionIf unansweredRisk
Which capabilities can the model see now?No versioned registryCapability drift and wrong selection
Where are tool calls validated deterministically?The model self-checksInvalid arguments reach real execution
Which component authorizes, and which one isolates?Only prompt warnings existA command can reach resources the task does not require
How are large results and secrets handled?Everything returns to contextDisclosure and window pollution
Which facts recover an interrupted process?Only chat is replayedDuplicate effects and false state
Does every exit have a machine-readable reason?Only final prose existsThe outer loop cannot take over safely

The harness turns “the model can think of it” into “the system can do it safely.” Next we enter the control flow that keeps advancing inside it: how an agent loop makes one run converge correctly.

Official sources

Further reading