Reading contract: Track turn-finalization scheduling, review write permissions, and Skill lifecycle. Keep pending proposals, applied writes, and later visibility distinct. Claims follow the linked public source snapshot; examples explain mechanics rather than prove improvement on independent tasks.

1. Why review follows foreground answer generation

After turns, state stores, and tool permissions are clear, "self-evolving" becomes concrete. It should not mean that the model reflects while answering, or that every turn is summarized into long-term knowledge. Both designs are risky: the foreground task is delayed, and temporary trial paths can harden into future rules.

What is a nudge? The ordinary word means a gentle push. In Hermes it is closer to a due reminder: runtime counters track elapsed user turns or tool-call iterations, and reaching a configured interval only sets the corresponding review flag to true. A nudge does not inject text into the user message, call a model immediately, or mean that a memory or skill has already been created. It tells the finalizer, "once this turn has a final answer and has not been interrupted, a review may be worth scheduling."

The opening figure separates foreground finalization, turn review, and periodic maintenance. Finalization here means the answer has been generated, not that a channel has delivered it to the user. Eligible review may start immediately or wait for the managed local runtime to become idle. A proposed write can then remain pending approval; only applied state enters later reads. The independent maintenance clock checks Skill lifecycle, with semantic consolidation as an optional phase rather than the next step after every turn review.

Use the coding-session example again. The agent finds a fixed project command, learns your commit preference, and also tries several paths that turn out to be wrong. The test log belongs in SessionDB. A repeatable project command may become a Skill. “Do not edit generated files automatically” belongs in the relevant Skill or project Context; “the user consistently prefers Simplified Chinese” is a cross-task preference suitable for Memory. The failed detours should usually remain transcript only.

Hermes implements self-evolution as a controlled write scheduled during turn finalization. The foreground turn produces the final answer. The finalizer checks whether review is allowed. A background review fork may look at the completed turn snapshot. That fork uses dispatch-limited tools, with additional write guards. Successful writes land in their stores, and later turns use them through the normal Memory and Skills paths. The already-generated answer is not rewritten.

Step What happens Why this shape matters
1. Generate the answer The foreground turn finishes tool calls and generates a final answer. The user is waiting for the task result, not for self-summary.
2. Check after The finalizer checks final response, interrupt status, the skip flag, and review reminders; later scheduling also checks configuration and origin. Interrupted turns or turns without a final answer are skipped; these gates do not prove a candidate lesson is correct.
3. Fork review Background review receives the completed messages snapshot and forks a restricted AIAgent. It can inspect user messages, assistant output, tool calls, and tool results without taking over the foreground turn.
4. Narrow writes The default whitelist includes Skills and read-only file tools; Memory also requires its review trigger. External provider ingestion is skipped. Default writes go through Memory / Skills management tools; explicit extra_tools may admit existing parent tools.
5. Affect future turns Successful writes are read later through normal Memory or Skills paths. Learning changes future work, not the answer already generated.

2. Finalizer schedules review only after the answer

Expand the final steps of finalization. Hermes already has the final answer, runs post hooks, synchronizes external memory, and only then checks the background-learning conditions. With error handling removed, the order is visible directly in the source:

agent._sync_external_memory_for_turn(...)

if final_response and not interrupted and not skip_background_review and (
    _should_review_memory or _should_review_skills
):
    agent._spawn_background_review(
        messages_snapshot=list(messages),
        review_memory=_should_review_memory,
        review_skills=_should_review_skills,
    )

The reduced finalizer condition also requires skip_background_review to be off. Memory counts user turns; Skills count tool iterations, with the relevant tool available. Reaching the interval resets that counter and marks review due. See Memory counting, Skill counting, and the finalizer gate. The scheduling entry additionally skips automatic reviews from delegated children and checks auxiliary.background_review.enabled. A due reminder alone does not authorize a fork from every source.

Suppose the Memory interval is three user turns. The first two turns only advance the counter; the third sets "review is due." If that third turn is interrupted, the finalizer still does not launch review. If it completes, review may still find no durable lesson and write nothing. Conversely, if the foreground agent already uses memory or skill_manage, the corresponding counter resets in pre-dispatch counter reset. Review is therefore not a per-turn tax. It is best-effort work after a safe ending: failure cannot revoke the generated answer, and success is still allowed to write nothing.

Scheduling is not the same as starting a model request. With the default defer: auto, reviews targeting Hermes's managed local llama-server wait in an in-memory queue. One slot per session keeps the newest snapshot; dispatch normally waits for 15 seconds without a live turn and idle server slots. After the default 30-minute maximum age, dispatch no longer requires idleness. A process exit loses pending work, and dispatch checks enabled again. Explicit /refine does not defer. See the idle review queue.

3. Review forks preserve evidence while isolating conversation and writes

The completed conversation is structurally copied, including nested tool calls and content, so fork sanitizers cannot modify the foreground history. Both failed and successful steps remain visible:

user: Fix this failing test
assistant -> terminal: Run pytest directly
tool -> assistant: Failed; generated files are stale
assistant -> terminal: Run the generator, then pytest
tool -> assistant: Tests passed
assistant: Fixed and verified

This lets review extract “run the generator first” rather than recommend the failed attempt. A same-model fork inherits the parent's system prompt, tool schemas, reasoning configuration, and resolved cache scope to preserve its first warm request prefix. A different auxiliary model receives a shorter history digest. Advertised schemas describe what the model sees; a separate dispatch whitelist decides what it may execute. See fork construction and history selection and execution.

3.1 Isolate persistence while allowing independent compaction

Sharing the parent session id would let normal persistence write the review harness into the real transcript. Hermes therefore sets _persist_disabled=true, clears _session_db, skips external memory-provider ingestion, and prevents the fork from closing its parent's session. Fully disabling compression, however, lets a long replay grow across tool calls. The current approach first detaches the compressor's database/session binding, then permits in-memory compaction of the fork alone:

review_agent._persist_disabled = True
review_agent._session_db = None
compressor.bind_session_state(session_db=None, session_id="")
review_agent.compression_in_place = True
review_agent.compression_enabled = True  # only after successful detachment
review_agent._review_defer_compaction_before_first_response = True

This simplified shape shows successful detachment. Compaction waits until the first response to preserve the first warm cache read; later compaction changes only the fork. If a plugin compressor cannot detach, compression remains disabled. Review permits at most 16 iterations by default. Its aggregate input budget defaults to 75% of the resolved context window, capped at 600,000 tokens, or 120,000 when the window is unknown; max_input_tokens can override it. Compaction bounds individual context size, while the aggregate budget bounds the whole review's input cost. See budget definitions and compression isolation.

3.2 Read permissions and write permissions are separate checks

The default whitelist includes Skills tools plus read_file / search_files. It includes memory only when this review was triggered for Memory and memory or user-profile storage is enabled. auxiliary.background_review.extra_tools can admit named tools already present in the parent's schemas; it cannot invent a schema. Terminal, ordinary file writes, and delegation remain denied by default. See the dispatch whitelist.

Permission to call skill_manage still does not authorize editing every Skill. Autonomous writes require Curator-managed records and reject pinned, user-owned, Hub, bundled, and external Skills. The fork must also have read the exact target file during this review before changing it. read_file records that read, so a read-then-patch sequence can pass. See background Skill write guards.

3.3 Pending proposals are not applied lessons

An unattended Memory add follows normal write rules, but replace / remove, including those inside a batch, stage a proposal instead of deleting prior memory. A staging failure denies the operation. With skills.write_approval enabled, Skill mutations also stage and replay only after approval. Both result shapes below can carry success: true, but they describe different states:

[
  {"success": true, "staged": true, "pending_id": "proposal-id"},
  {"success": true, "message": "Skill 'generated-code-testing' created."}
]

The first accepts a proposal without changing the durable Skill; the second reports an executed creation. Review summaries distinguish pending proposals from applied changes and exclude old tool results inherited in the snapshot. A notification alone does not prove application. See Memory proposal rules, Skill approval entry, and review summaries.

When a new user turn arrives before review finishes, the foreground requests cancellation and waits at most two seconds for request-exit acknowledgement. It logs a warning and proceeds after that deadline, so cancellation does not guarantee the fork has stopped or that resource overlap is impossible. A preempted managed-local review in deferred mode can requeue at most three times. Applied writes are not automatically rolled back by cancellation. See the foreground entry, cancellation handshake, and bounded requeue.

4. Curator is long-term maintenance, not per-turn reflection

Background review asks whether one completed turn revealed something worth saving. Months later, the problem changes shape. A library may contain "generate code before testing," "refresh generated files before testing," and "rerun tests after a schema change" as three separate skills. Every one may come from a real success, but keeping all three forever may make retrieval worse. Another rarely used skill may still be important and should not disappear merely because its task has not appeared recently. Curator handles maintenance pressure that becomes visible only across many turns.

4.1 It wakes on periodic entry-point checks, not after every turn

Curator does not sit on the finalizer-to-background-review path above. The CLI checks at new-session startup, while a long-running Gateway asks from its housekeeping loop once an hour. Whether a pass actually starts is then decided by maybe_run_curator and should_run_now:

CLI session starts / Gateway polls hourly
  -> is curator.enabled true?
  -> is the curator paused?
  -> has interval_hours elapsed since last_run_at?
  -> if supplied, is idle time at least min_idle_hours?
  -> yes: start one Curator run
  -> no: stop this check without reading or changing skills

The defaults are a seven-day interval and a two-hour idle requirement. The current CLI startup and Gateway poll both pass infinite idle time, so this is not evidence that those entry points measure a two-hour user pause. On the first observation of an install with no Curator history, Hermes only seeds last_run_at and waits a full interval. It does not reorganize the library immediately after an update. A manual hermes curator run bypasses this scheduled check. See the CLI startup hook, Gateway poll, and should_run_now.

4.2 Phase one applies time thresholds without a model

Hermes keeps operational telemetry out of SKILL.md. A sidecar at ~/.hermes/skills/.usage.json records how often a skill was actually loaded or referenced, viewed with skill_view, and patched, plus the latest time for each event. last_activity_at is the newest of last_used_at, last_viewed_at, and last_patched_at:

{
  "generated-code-testing": {
    "use_count": 12,
    "view_count": 4,
    "patch_count": 2,
    "last_used_at": "2026-04-10T09:00:00Z",
    "last_viewed_at": "2026-04-12T08:00:00Z",
    "last_patched_at": "2026-03-28T18:00:00Z",
    "state": "active",
    "pinned": false
  }
}

This record's latest activity is April 12. Counter updates and timestamp derivation live in bump_view, bump_use, and bump_patch and activity derivation. The lifecycle is not four mutually exclusive states. It has three states plus one orthogonal protection flag:

State or flag Discoverable by later work? How it enters How it leaves
active Yes Creation, restoration, or renewed activity after stale Becomes stale after 14 inactive days by default
stale Yes. Its directory remains in the normal skill tree Crosses stale_after_days Returns active after activity and the next pass, or ages into archive
archived No. Its directory moves under skills/.archive/ Crosses 30 inactive days by default, or is manually archived hermes curator restore <name> restores it as active
pinned Yes; pin does not replace active or stale, and is not a fourth state The user explicitly pins it The user unpins it; automatic transitions skip it while set

Phase one walks these records and applies one configured threshold policy. Reduced to its decision shape, the source reads like this:

anchor = newest(last_used_at, last_viewed_at, last_patched_at)
anchor = anchor or created_at

if skill.pinned or skill.referenced_by_cron:
    continue
if now - anchor >= archive_after:
    archive(skill)                 # move to .archive; recoverable
elif now - anchor >= stale_after and skill.state == "active":
    skill.state = "stale"
elif now - anchor < stale_after and skill.state == "stale":
    skill.state = "active"

For generated-code-testing, use on day 10 refreshes the activity anchor. A Curator pass on day 25 sees 15 inactive days and marks it stale while leaving it discoverable. Viewing or using it on day 28 lets the next pass reactivate it. With no later activity, it becomes stale again on day 43 and eligible for archive on day 59. This uses the current 14 / 30-day defaults; movement happens during a Curator pass, not automatically at the exact deadline. See default thresholds and transition and grace rules.

The candidate set is not every skill on disk. Skills autonomously created and marked by background review qualify; users may also opt in with hermes curator adopt <name>. With curator.prune_builtins: true, bundled built-ins also receive time-based maintenance, but their first observation only seeds a clock. Hub-installed, external-directory, and protected built-in skills remain outside the set. Skills referenced by cron jobs and pinned skills skip automatic transitions. See candidate enumeration.

4.3 This is neither LRU nor LFU

"Recently used," "frequency," and "archive" sound like cache eviction, but Curator has no rule saying a full library must evict one entry. It does not rank every skill and sacrifice the last one. Phase one is better understood as an independent lifecycle clock for each skill:

Question LRU / LFU cache Hermes Curator phase one
When removal happens Usually under capacity or memory pressure When a periodic pass finds one skill beyond its time threshold
Primary signal Global recency order or global frequency That skill's latest use, view, or patch timestamp
Global ranking required? Yes, to choose a victim No; many skills may remain active, become stale, or archive together
Is removal destructive? A cache entry is usually dropped and rebuilt from its source The directory moves to .archive and can be restored or rolled back

Status and the semantic candidate list expose use_count, view_count, and patch_count, but automatic stale/archive decisions do not sort by frequency. use_count == 0 only adds a grace rule: a newly created skill whose trigger has not appeared must survive at least the stale window. Zero frequency is absence of evidence, not an immediate LFU eviction signal.

4.4 Phase two is optional semantic consolidation

Timestamps can say how long a skill has been quiet; they cannot say whether three testing skills belong in one reusable procedure. Phase two therefore renders each candidate's state, pin and cron flags, counters, and latest activity for a restricted auxiliary agent to inspect. In the current source, consolidate defaults to false. A normal automatic pass only runs phase one. The model pass costs tokens only after setting curator.consolidate: true or manually running hermes curator run --consolidate.

counts = apply_automatic_transitions(now=start)  # phase one always runs

if not consolidate:
    write_report("llm: skipped (consolidation off)")
    return

candidate_list = _render_candidate_list()
if candidate_list:
    llm_meta = _run_llm_review(candidate_list)   # phase two is optional

Return to the three testing skills. The auxiliary agent must inspect complete skill packages rather than merge on name similarity. If they implement one class of workflow, it may choose or create a generated-code-testing umbrella, preserve unique details in its body, references/, templates/, or scripts/, and then archive the absorbed source directories. It can patch a drifting skill, consolidate overlapping skills, keep an already broad umbrella, and archive absorbed packages. A background delete must declare absorbed_into pointing to an existing umbrella; a bare deletion is refused. Obsolescence-based pruning belongs to phase one. See the consolidation archive guard.

before
  test-after-codegen
  refresh-generated-files
  pytest-schema-change

semantic decision
  all three serve generated-code validation
  each still contains a distinct tool, failure signal, or command

after
  generated-code-testing/
    SKILL.md
    references/schema-change.md
    scripts/verify-generated-files.sh
  old directories -> skills/.archive/ (absorbed_into recorded; recoverable)

A pin is a user-owned safety fence, not an excellence grade the model may award; phase two must skip pinned skills. Bundled built-ins may only be archived when pruning is enabled, not rewritten into umbrellas. Hub and external skills remain out of bounds. The consolidation condition and ordering are in run_curator_review, and the rendered fields are in _render_candidate_list.

4.5 What later work observes

Only a consolidation-enabled pass attempts a best-effort whole-tree snapshot of ~/.hermes/skills/. Time-only pruning avoids copying the whole tree: archived directories provide single-Skill recovery, and old snapshots are pruned. State, transitions, consolidation destinations, tool calls, and recovery information are written under ~/.hermes/logs/curator/<timestamp>/. Active and stale directories remain discoverable; archived directories leave that path, and umbrellas become available through later scans. See snapshot conditions and reporting.

hermes curator run --dry-run is the safety preview. It skips automatic state mutations, instructs the model to report without write tools, and does not advance the scheduled last_run_at. After a real pass, one archived skill can be restored; if the pre-run snapshot was successfully created, it can restore the pre-consolidation tree. The two phases now have separate questions: phase one asks where a skill sits in its lifecycle; phase two asks how the knowledge that remains worth keeping should be organized.

5. Conclusion: three clocks improve one agent without becoming one chain

The foreground turn, background review, and curator all improve the agent, but they run on different clocks. The foreground solves the current task. Review extracts a preference or procedure after one turn. Curator maintains the skill library after many turns. Their trigger, input snapshot, decision method, and write permissions are deliberately different.

This closes runtime learning, not offline optimization. Background review asks what this turn is worth preserving, and the curator asks how to maintain the skill library. Neither proves that a particular rewrite performs better on an independent set of tasks. That requires comparison on an independent task set. Before starting that offline evaluation, the next chapter covers the user's active-learning entry point: how one /learn request reads named material, creates a Skill, and exposes relationships through a Learning Graph that is a view rather than a hidden learning engine.

Sources