1. Why “I ran the tests” still does not mean “ready to deliver”

The previous two chapters followed one example. The generated-code-testing Skill used to tell the Agent to regenerate code after a schema change. An offline candidate added one missing action: inspect the generated diff. A human adopts the revised Skill. On the next real task, that Skill guides the Agent to edit generated-code checking logic in the target application repository, not Hermes core. Every project root, command, and evidence row below belongs to that target repository. This is the moment when an apparently natural sentence appears: “The fix is done and the tests pass.”

That sentence omits every fact that determines whether it is trustworthy. Was one targeted test run, or the full repository suite? Did the run happen before or after the last edit? Did it prove only a Python function, or did the application build and start? Was the command an accepted project entry point? Without those answers, “tests passed” is narrative, not delivery evidence.

Think of verification evidence as a receipt with a time and a boundary. It states which command ran in which workspace, what scope it covered, what exit status and output it produced, and whether the run happened after the newest change. Delivery becomes credible only when this order holds:

change code
-> invalidate old evidence
-> find a project-approved verification entry point
-> execute it and record scope plus result
-> obtain a fresh pass after the latest edit
-> only then claim “verified”

Carry four questions through this chapter. How does Hermes know what to run? How does it distinguish one-file evidence from repository-wide evidence? What happens when the model tries to finish without a fresh pass? And why are local verification, offline improvement, and human delivery still three separate decisions?

2. Before execution: how Hermes learns what this project expects

Suppose the Agent just changed generated-code checking logic. The simplest strategy is to guess from the ecosystem: try pytest for Python or npm test for Node. That sometimes works in a small repository, but real projects often wrap tests, require a particular directory, or separate lint, type checking, build, and runtime checks. Hermes therefore begins by collecting stable project facts, not by firing a guessed command.

2.1 Project facts are a run sheet, not a new configuration authority

detect_project_facts() reads signals the repository already owns: manifests, package managers, scripts/run_tests.sh, test/check/lint/build/typecheck scripts in package.json, pytest configuration, and Makefile targets. It places accepted entries in ProjectFacts.verify_commands. The same facts can be rendered into model context and returned structurally to the evidence classifier.

@dataclass
class ProjectFacts:
    manifests: list[str]
    package_managers: list[str]
    verify_commands: list[str]
    context_files: list[str]

ProjectFacts itself does not store the root. Its caller, project_facts_for(), first resolves a git or marker root and then returns that root beside the four fact groups. Hermes is not inventing a new project standard here. It is respecting the repository's existing base entry points. In our example, project facts may discover only pytest. The Agent appends tests/generated_code at execution time, and only then does the classifier label that run targeted. A base entry point plus target arguments becomes one concrete check. See ProjectFacts, detection, and structured output.

2.2 One code change passes through five ordinary moments

Ignore the internal type names for a moment. An ordinary delivery passes through the following five moments. The evidence ledger, verification recipe, and stop gate merely protect different transitions in this sequence.

MomentWhat happenedStrongest honest claim
1. Before editA previous full run passedThe old version passed; the new one is not covered
2. After editFiles were written and old evidence detachedThe implementation changed and is now unverified
3. Fast checkThe targeted generated-code test passedThe targeted scope passes, not automatically the repository
4. Complete checkThe full suite or recipe passedThis workspace passed within the recorded scope
5. Delivery decisionA human reviews the diff, evidence, risk, and policyThe change may be adopted, merged, or released

This separates two responsibilities early: a verification command cannot make a release decision, and a review opinion cannot substitute for execution. Hermes now turns moments two through four into queryable state.

3. The evidence ledger stores a concrete check, not a sentence

If conversation history contains only “tests passed,” later code cannot tell which command it means or whether a newer edit invalidated it. Hermes therefore uses a passive ledger. Terminal tools or hermes verify submit facts after execution; the ledger classifies, stores, and queries them. It never runs commands, blocks delivery, or decides readiness by itself.

Project facts and terminal results are classified by kind, scope, and status, then written to verification_events and verification_state; a new edit clears last_event_id and a new passing event restores passed
An event records what happened; the state row answers which evidence is usable for this workspace now.

3.1 One event must carry command, scope, and result together

VerificationEvidence stores more than an exit code. It includes the original and canonical command, verification kind, scope, status, working directory, project root, session, and an output summary. Only then can the same exit_code = 0 mean “this particular check passed in this particular scope.”

class VerificationEvidence:
    command: str
    canonical_command: str
    kind: str       # test / lint / typecheck / build / ...
    scope: str      # targeted / full
    status: str     # passed / failed
    exit_code: int
    cwd: str
    root: str
    session_id: str
    output_summary: str

The classifier first compares a command with canonical entries in project facts. It then examines trailing file, directory, or test-selection arguments to determine scope. pytest -q tests/generated_code/test_diff.py becomes targeted; the project's complete pytest -q may become full. The implementation is in command kind, scope, and temporary-script rules and terminal and verify recording entry points.

targeted is not weak evidence. It is often the right fast diagnostic and may support a narrow claim. It simply cannot promote itself into “repo green.” Scope labels keep a conclusion honest without forcing every iteration to start with the slowest command.

3.2 Why events and current state live in separate tables

SQLite table verification_events is append-only history: each check retains its command, scope, status, and summary. verification_state is keyed by session_id + root and stores a current event pointer, last edit time, and changed paths. The first supports auditing; the second quickly answers, “does this session have usable evidence for this repository now?”

When a tool writes code, mark_workspace_edited() does not invent a failing test. It clears state.last_event_id and records the edit. The normal current edit path therefore reports unverified: the new version has not failed, but it has not proved itself yet. verification_status() also contains a stale branch for event/edit ordering, but it would be inaccurate to claim that every edit currently produces that label. See event insertion, edit invalidation, and status queries.

old full pass remains in events for audit
+ latest edit
-> state.last_event_id = null
-> current status: unverified
+ new targeted pass
-> current status: passed, scope remains targeted
+ complete canonical run passes
-> current status: passed, scope becomes full

Retention is intentionally bounded: events default to 30 days, 100 entries per session and root, plus a cap on unreferenced rows. This is operational evidence for local delivery, not an unlimited compliance archive.

4. hermes verify: executing the test, build, and start phases a Recipe declares

A targeted test can prove a function satisfies its assertions while the application still fails to build, boot, or bind a port. A web repository may be green at the unit layer but broken by packaging, environment, or start-command changes. hermes verify organizes these checks into a repeatable recipe.

4.1 A Recipe turns environment knowledge into ordered phases

A Recipe can contain bootstrap, build, test, start, port, readiness URL, and evidence notes. Hermes first loads a saved .hermes/environment.json; otherwise it statically detects Node, Python, Go, Rust, Java, Make, Docker, and related project shapes. That detection is cheap because it reads only a few project files; a discovered pytest command or wrapper script may still be expensive to run. The CLI also merges canonical commands from project facts into a detected recipe, so learning how to start the app does not hide its own lint or test entry points.

hermes verify reads a saved environment or static project detection, then runs bootstrap, build, test, background start, readiness polling, process-group teardown, and returns VerifyResult
Running the current Recipe completely records full; when a Recipe has no start, full still does not mean startup or readiness ran. Selecting one phase or skipping start is downgraded to targeted.
result = run_verify(recipe, phase=phase, skip_start=skip_start)

for name in ("bootstrap", "build", "test"):
    run_phase(name)
    if stop_on_failure and failed:
        return result

if not failed and recipe.start and not skip_start:
    start_in_background()
    poll_readiness()
    terminate_process_group()

The runner executes bootstrap, build, and test in order. In stop-on-failure mode, any failed phase returns immediately. Only when prior phases pass and a start command exists does it launch a background process, poll readiness, and tear down the process group afterward. Follow the path in run_verify, readiness, and teardown; the recipe contract is in Recipe ownership and fields.

4.2 Readiness proves the polled address answered, not that business logic is correct

Hermes currently treats any HTTP response as ready, including 4xx and 5xx. This proves that the configured URL returned an HTTP response during polling, but not strictly that the response came from the process Hermes just launched; an existing service on that port could also satisfy the check. It does not prove login, code generation, or data writes. Those behaviors still require targeted tests or a real smoke check.

Likewise, hermes verify --phase test and --skip-start are useful fast paths, but their evidence is downgraded to targeted. Without either partial option, the run is recorded as full. Here full means “all phases this Recipe actually declares.” If the Recipe contains test but no start, full does not invent a boot check. The CLI and scope logic are visible in run_verify_command and _record_evidence.

5. Who checks freshness when the model tries to finish

Project facts tell the model what to run, the ledger records what actually ran, and hermes verify adds runtime coverage. The model can still edit a file and immediately draft a final answer without invoking any of them. A prompt that says “remember to test” is behavioral guidance, not an enforceable gate.

Hermes therefore applies verify-on-stop before a genuine turn end. The policy does not execute tests. It asks whether files changed this turn, whether the changes are docs or Skill text only, and whether the ledger contains a fresh passing event. When code needs verification but no evidence exists, the policy refuses this finish attempt and gives the model a precise continuation instruction.

After a real final answer candidate, the gate checks for file edits, docs-only changes, and fresh passing evidence; without evidence it persists the candidate, appends a synthetic nudge, and silently continues for at most two attempts
The gate returns control to the model; command choice, execution, repair, and recording still belong to the model and its tools.

5.1 What exactly is a nudge?

In this source, a nudge is a lightweight runtime prompt. It is not a new user request, and it does not secretly run a test. It is a synthetic user message saying, in effect: “Code changed, but no fresh evidence exists. Run relevant verification, repair failures, or explain the blocker; do not claim this work is verified.”

build_verify_on_stop_nudge() prefers canonical commands from project facts. If a runnable Recipe exists, it suggests hermes verify --json. Only when neither exists may it recommend a specially prefixed temporary verification script. This ordering prevents the model from making the ledger green with an irrelevant true or an arbitrary command. See recipe eligibility and nudge construction.

The current gate checks freshness, not scope. As soon as the ledger's current status is passed, even for fresh targeted evidence, the gate allows the turn to finish; it does not require a full run. The final answer must therefore state scope, and reviewers must not read “the gate allowed finish” as “the whole repository is green.”

5.2 Why docs skip the gate and code gets at most two reminders

Markdown, ordinary text, README files, and Skill prose usually lack an executable verification path, so docs-only edits finish normally. “Local coding and programmatic surfaces” means the CLI, TUI, desktop app, and direct programmatic callers; they default on. Telegram, Discord, and similar messaging channels default off so internal continuation does not become chat noise. Environment and configuration can override either default. The filtering and surface policy are in verification_stop.py.

The continuation is capped at two attempts. Without a bound, missing dependencies, broken tests, or an unfixable model could loop forever between finish and reminder. Reaching the cap does not make verification pass. It only guarantees that control flow can terminate; the answer must still report failure or a blocker honestly.

The observed edit set also has a current implementation boundary. The conversation loop passes _turn_file_mutation_paths, and Hermes currently adds paths only for landed write_file and patch tool results. Files changed through a terminal command or an external program do not automatically enter that set; if no project facts can be resolved from the observed paths, no nudge is produced. The gate therefore protects edits it observed through those tools, not every filesystem mutation on the operating system. See the conversation-loop handoff and FILE_MUTATING_TOOL_NAMES.

5.3 Why the real answer candidate is shown and persisted first

By the time the gate intervenes, the model has already produced real answer content. Hermes keeps it: the candidate is emitted as interim and persisted, then a synthetic message marked _verification_stop_synthetic is appended before the model loop continues silently. A successful continuation produces a revised final answer. If the continuation consumes the remaining budget, the runtime can recover the real candidate instead of ending on an internal reminder.

The handoff is implemented in the genuine turn-end branch of the conversation loop. It explains why a user may see an interim answer followed by a verified update: this is one delivery continuing at an evidence boundary, not a second user turn.

6. Durable history must still look like a real conversation

The synthetic nudge is useful runtime scaffolding, but it is not a real user message. Persisting it forever would make resumed models believe that the user explicitly requested verification, and could leave the visible transcript ending on an internal command instead of an assistant answer.

turn_finalizer removes marked synthetic nudges before durable persistence while retaining real assistant candidates. If verification continuation exhausts the remaining budget, it uses the pending real candidate as a fallback. If the model later produces an updated answer, that answer replaces the provisional candidate in replay. The core behavior is in scaffolding removal and budget fallback and final persistence cleanup.

user-visible, resumable history:
user task
assistant candidate / tool work
assistant verified final response

runtime-only scaffolding:
synthetic verify nudge
verification continuation flags

The invariant is conversational truth. A gate may alter control flow, but it must not invent user intent. A failed verification may change the conclusion, but it must not erase the real work that preceded it.

7. Local pass, offline improvement, and human adoption are three gates

Put the previous chapter's offline evolution beside this chapter's local verification. Both use evidence, but they answer different questions. Local verification asks whether the latest code in this workspace passes declared checks. Offline evaluation asks whether a Skill candidate improves over a frozen baseline across representative tasks. Human adoption asks whether the organization accepts the diff, risks, cost, and release policy.

To keep this chapter self-contained, translate the dataset names once more: training tasks expose failures and provide rewrite signals; validation tasks help the search choose candidates; holdout tasks stay sealed until the final independent comparison; and the frozen baseline is the old Skill that remains unchanged throughout the experiment. Reusing one split for all three jobs lets the optimizer memorize its measuring stick.

GateInputWhat it provesWhat it cannot replace
Local verificationCurrent code, canonical command, exit status, outputThe latest workspace passes within recorded scopeGeneral improvement across tasks or release judgment
Offline evaluationFrozen baseline, candidate, train / validation / holdoutThe candidate is more reliable on a controlled task setThe current integration builds and starts
Human adoptionDiff, both evidence types, cost, risk, policyWhether to merge, enable, release, or roll backActual execution or a fair comparison

7.1 Replay the generated-code example end to end

  1. A runtime task exposes the missing “inspect generated diff” behavior. The experience enters a candidate area instead of rewriting the active Skill mid-turn.
  2. An offline run produces candidates on train, navigates with validation, and compares the frozen baseline on holdout. A human adopts the new instruction.
  3. On a later application task, the new Skill guides the Agent to change generated-code checking in the target repository. Edits landed through write_file or patch invalidate old local evidence.
  4. Project facts expose canonical base entries such as pytest. The Agent adds target arguments for fast diagnosis, then runs the complete entry or hermes verify.
  5. The ledger records command, scope, status, and output. Without a fresh pass after the newest edit, verify-on-stop continues the turn.
  6. The final answer states the verified scope honestly. A human then combines offline results, source diff, local evidence, and risk into the merge or release decision.

No single green mark proves everything. passed + targeted is not repository-wide evidence; readiness is not business correctness; holdout improvement is not proof that integration boots; and human approval cannot retroactively prove that tests ran. Trustworthy delivery comes from connecting these signals while preserving each boundary.

8. Conclusion: an Agent finishes only after its evidence

“The fix is done and the tests pass” becomes meaningful only when command, scope, result, workspace, and time are traceable. Hermes does not pretend one universal test can settle every engineering question. It protects the sequence most likely to be lost at delivery time: edit, invalidate, rerun, record, then finish.

ProjectFacts names project-approved entry points;
VerificationEvidence records one concrete check;
hermes verify executes Recipe phases, including build, startup, and readiness when declared;
verify-on-stop blocks unsupported finish attempts;
turn_finalizer removes internal scaffolding and preserves real conversation;
humans combine local verification, offline evaluation, and risk into delivery.

The seven chapters now complete Hermes Agent's two clocks. The runtime clock handles turns, tools, memory, Skills, and background curation. The offline clock turns experience into comparable candidates. The verification loop asks one final plain question before either clock's output reaches a person: did the evidence happen after the latest change you are claiming is better?

Source references