Source scope. This article reads one fixed public source snapshot. Claims linked to a file or function are verified source; conclusions assembled from visible constraints are labeled engineering inferences. Performance numbers are project-reported from the README and technical report; I did not rerun the benchmarks here.
Reading goal. We will carry one shape-level example throughout: several SpreadsheetBench agents damage formulas or choose the wrong sheet when writing a workbook back. Those failure buckets are named by the optimization prompt, but no unpublished sample count is invented. By the end, you should be able to state who triggers optimization, which files it reads, what it may edit, which score retains a change, and when the result is ready to deploy.
Model weights stay frozen; changes land in skill.md or an allowlisted harness file.
Trajectories become files, recurring patterns are mined, and passed traces try to falsify the diagnosis.
The optimizer produces a candidate; validation scores and thresholds decide retention or rollback.
The loop is clear, while scripts, split definitions, and HarnessOpt commands still break at HEAD.
1. Translate “self-evolution” into one repeatable revision
The first mistake is to read SkillOpt-Lite as online fine-tuning. It computes no parameter gradient and changes no model weights. A more precise description is: the natural-language skill is the program under revision, the coding agent generates a patch, and evaluator scores decide whether it remains. During one skill experiment, model M and harness H remain fixed; the search variable is the text artifact skill.md.
The report writes the objective as reward over H(M, z, s), where s is the skill. Text is discrete, and the model–harness–environment composition is non-differentiable, so the system does not approximate a numerical gradient. It exploits information already present in rollouts: errors, tool calls, generated code, and environment outcomes. Optimization becomes less like blind perturbation and more like debugging a program, then compiling a better instruction artifact. That is the useful substance behind the report's phrase language-mediated program compilation.
1.1 Five participants: what each reads, stores, and changes
| Participant | What it reads or stores | What it may do | What it cannot do |
|---|---|---|---|
Target model M | Current request and generated trajectory | Attempt the task under the current skill | It does not rewrite its own weights |
Harness H | Tools, loop, timeouts, and environment execution | Turn a task into a scored result | It is fixed in a SkillOpt-Lite round |
skill.md | Reusable domain rules and procedures | Receive small patches, snapshots, and rollback | It does not prove its own quality |
| Coding agent | Current skill, passed/failed traces, and source tools | Diagnose patterns and produce a candidate | It cannot declare that candidate best |
| Evaluator / gate | Splits, scores, and current/best records | Retain, reject, restore, and report | It does not author the patch |
1.2 The trainer is a prompt file
On the Lite path, the optimization loop is not a newly deployed optimizer service. A coding agent reads and executes skillopt-loop.prompt.md. The README likewise says to open the repository in a host that can read prompt files and issue the command in the agent chat, not a plain shell. The filesystem therefore stores both optimization data and the optimizer's control program.
This is powerful because the loop policy is readable, diffable, and benchmark-specific. It is also fragile in a very ordinary way: working-directory assumptions, host tools, long-command continuation, and a single path typo can stop the “algorithm” before evaluation begins.
2. What one optimization round actually does
Foreground product requests do not silently trigger this evolution. A user explicitly invokes the slash command, starting an offline or maintenance workflow. It sees a fixed current skill and the current train rollout; it should not read validation or test items to author the patch.
2.1 From snapshot to the next sample batch
| Stage | Input snapshot | Writes | Advance condition |
|---|---|---|---|
| Baseline | Initial skill.md and full validation | Round-0 best snapshot | Baseline score is recorded |
| Train rollout | Current skill, train slice, and seed | results.jsonl plus per-item sample files | Readable trajectories exist |
| Improve | Current skill and latest train samples | __before.md plus candidate skill | A cross-sample, generalizable pattern exists |
| Validation gate | Candidate and full validation split | Current/best/rollback state | The score clears a threshold or becomes flat/reject |
| Next round | The post-gate skill on disk | New-seed train samples | The next improve sees traces matching the active artifact |
setup
-> run baseline val; snapshot initial best
-> run train(seed = 1); write failed/*.md + passed/*.md
round r
-> snapshot __before.md
-> diagnose latest train samples
-> patch skill.md
-> run full val
-> accept | flat | reject
-> run train(seed = r + 1); replace sample files
final
-> restore best snapshot
-> run test exactly once
2.2 Read/write permissions, storage, and a valid no-op
The improve prompt narrows authority deliberately. It must read workspace/skill.md in full, then sample samples/failed/ and samples/passed/. It may edit the workspace skill, not the baseline. Single-task edge cases are skipped; a change needs support from at least two failures. An empty sample directory stops the process, and a round with no stable common pattern may leave the artifact unchanged. Evolution is allowed to do nothing.
Successful work becomes visible through three stores: the sample directories preserve observations, .skillopt/history/ preserves rollback snapshots, and workspace/skill.md is the artifact used by the next rollout. Later rounds do not retrieve a mystical internal memory. They read files and fresh trajectories produced by the currently active skill.
3. Trajectories as files: why a coding agent can absorb optimizer layers
The key is not merely “ask an LLM to reflect.” It is to turn the evidence required for reflection into the medium coding agents already handle well. row_to_md converts every results.jsonl row into a Markdown file. Frontmatter carries id, status, score, environment, and tags; the body carries Input, Expected, Agent output, Trace, and Notes. Failed and passed items go to separate directories.

3.1 One sample first answers “what did the agent actually see?”
For the spreadsheet example, a sample is not just success=0. It collocates the instruction, expected output, agent output, tail of the ReAct trace, and fail_reason. For an execution failure, the prompt asks the optimizer to inspect the generated solution.py directly. For an agent-phase failure, it reads the final turns where the run repeatedly became stuck.
---
id: task_x
status: failed
score: 0
tags: [sheet_level, exec]
---
## Input
## Expected
## Agent output
## Trace
## Notes
fail_reason: ...
This is a shape-level example, not a published project sample. The information changes form: an episode buried in evaluator output becomes an independently addressable file. The coding agent can inspect the directory distribution, then open a small high-signal subset instead of flooding its context with every long trace.
3.2 Consensus mining means cross-task invariants, not majority vote
The reading budget buckets by task_type, then caps failed/passed samples and total trace text. Diagnosis follows four rules:
- Failure-first: recurring failure repairs win when passed and failed traces suggest conflicting edits, while regressions against successful traces are still checked.
- At least two supports: a one-task edge case is skipped; sheet names, columns, and row counts must not be hardcoded.
- Passed-trace falsification: inspect at least one success of the same task type to see whether the current skill already handles the alleged rule.
- Smallest effective patch: SpreadsheetBench permits at most four edits, favoring concrete code patterns over vague reminders.
“Several tasks left formulas behind” should not compile into “never use formulas.” A useful patch names its conditions and actions: when openpyxl is needed to preserve workbook structure, how to iterate over actual rows, and why a truncated preview cannot define the dataset. At this point a candidate has been produced; it has not yet been shown effective.
4. Validation decides whether the patch stays or rolls back
If the same agent can diagnose a failure and declare its repair successful, the process collapses into self-persuasion. SkillOpt-Lite separates writing patches from selecting versions: train produces modification signals, aggregate validation scores decide retention or rollback, and the prompt reserves test for one run after all selection is over.
4.1 Train, validation, and test have different permitted uses
| Stage | What the prompt allows the optimizer to read | What it may change | Question answered |
|---|---|---|---|
train | Task content, passed/failed traces, and generated code | May drive a skill.md patch | “What should we try next?” |
val | Use comparable aggregate scores only | May retain, flatten, or roll back | “Is this candidate worth carrying forward?” |
test | Opened only after the loop | May no longer affect candidate selection | “How does the final best do on untouched samples?” |

4.2 The Lite dead band adds a state beyond pure evaluate_gate
The repository's pure evaluate_gate has three outcomes: beat current and best for accept_new_best; beat current only for accept; otherwise reject. The SpreadsheetBench Lite prompt refers to this comparison policy, but asks the coding agent to make the decision rather than calling the pure function. It separately specifies a ±0.02 dead band and a flat state:
Δ ≥ +0.02is treated as improvement;Δ ≤ -0.02rolls back;|Δ| < 0.02leaves the candidate on disk for the next train batch but does not updatecurrent_accorbest;cand_soft − current_soft ≥ 0.05may promote a hard-flat result to accept.
Flat is therefore neither ordinary acceptance nor a no-op. It can alter the behavioral distribution observed by the next round, but it cannot become the final restore target unless a later gate makes it best. Reading only gate.py misses state semantics owned by the loop prompt.
4.3 Validation reading rules are not enforced access isolation
The project calls validation held out and instructs the agent not to read validation items, using only aggregate scores for gating. This is a prompt-level reading rule, not access isolation enforced by files or tools. The validation step still produces per-item samples in the shared samples/ directory; the next train batch clears and replaces them. Keeping validation content out of patch generation therefore depends on the host agent following that instruction.
Even when the instruction is followed, the same validation set selects candidates across many rounds. Over the entire optimization process it is better understood as a search set, not an untouched final estimate. As an engineering inference, round count, dead bands, and stopping rules can adapt indirectly to it. The one-shot test, reserved until search ends, estimates the final best on samples that did not participate in selection.
| State | What has happened | What is still missing |
|---|---|---|
| Output produced | The coding agent wrote a candidate skill.md | Comparable independent validation |
| Output validated | The candidate passed validation; final best ran one test | Deployment-specific cost, latency, safety, and rollback checks |
| Output adopted | An authorized release process installed the best artifact | Ongoing monitoring and review remain necessary |
5. Full SkillOpt versus Lite: fewer optimizer steps, the same scored validation
The same repository retains the Full SkillOpt Python trainer. The header of trainer.py names six stages: Rollout, Reflect, Aggregate, Select, Update, and Evaluate. Lite keeps rollout and evaluation. It lets the coding agent and filesystem absorb much of reflection pooling, patch merging, ranking, and update orchestration.

| Question | Full SkillOpt | SkillOpt-Lite |
|---|---|---|
| How are trajectories read? | Minibatch reflection and structured patches | Per-item Markdown plus coding-agent exploration |
| How are suggestions combined? | Hierarchical merge_patches | Cross-task pattern mining in one agent context |
| How is update size bounded? | Edit budget and scheduler | Minimal-change prompt rules and edit cap |
| How is failed history retained? | Rejection/step buffers | History snapshots, diffs, and fresh next-round samples |
| Who accepts? | Selection/evaluation gate | Full validation gate, best snapshot, and one-shot test |
Do not borrow guarantees across the two paths. Full trainer defaults enable use_slow_update: true. Its implementation unconditionally injects slow-update guidance into current and best, marking it force_accept. That is a Full SkillOpt behavior, not the Lite prompt's strict validation-gated retention. Similar names in one repository do not mean the two paths use the same code to retain or roll back changes.
6. HarnessOpt: widening the target from instructions to execution
If the failure comes from a tiny workbook preview, brittle code extraction, an unsuitable timeout policy, or a missing tool, more wording in skill.md merely wraps a harness defect in prompt text. HarnessOpt therefore lets the agent search Python harness code. But every additional writable file must increase approval, validation cost, and rollback discipline with it.

6.1 The allowlist names exactly six writable files
The HarnessOpt prompt allowlists rollout.py, react_agent.py, codegen_agent.py, executor.py, recalc_harness.py, and adapter.py. Skill files, evaluator, dataloader, configs, and prompts are denylisted. Skill-content failures are routed back to a separate /skillopt-loop. This is more precise than the README table's broader “skill.md and agent code,” so the prompt's allowlist and denylist determine what the current path can actually edit.
6.2 Round 0 makes architectural decisions; later rounds become surgical
Round 0 runs baseline validation and full train, scans the failure distribution, then considers memory, tools, prompt context, loop policy, and codegen/executor shape. A new tool, memory system, or execution shape may expand the data and actions available to the code, so the prompt prints a brief and waits for user approve before applying the patch. Round 1 onward narrows to automatic surgical edits, diff guard, smoke, full validation, and rollback.
This protects a useful invariant: architectural authority expansion receives human approval; only bounded repairs enter automatic gating. The shipped Round 0 still has a gap. It commits and tags the bootstrap as round-0-best, then runs only a six-item debug batch. The prompt explicitly says no full validation gate runs there; Round 1 is expected to validate it. If no later round creates a new best, final restore may select a bootstrap that passed smoke but never a full gate.
6.3 The checkpoint shows what harness optimization can produce
harnessopt_ckpt contains concrete code, not just an architectural claim. In codegen_agent.py, enabling SPREADSHEETBENCH_HARNESS_NANO=1 expands workbook preview from 5×20 to 15×30 and activates gold-free output introspection that detects formula strings left in answer cells.
A separate flag, SPREADSHEETBENCH_TIMEOUT_FALLBACK=1, enables reasoning-effort fallback: after a turn’s normal model-request retries fail in run_multi, it can make another attempt at a lower reasoning effort when the remaining time and current effort allow it. This applies to the Chat / Responses backend, not the exec backend. Both flags default off, allowing independent rollback and A/B comparison.
7. Reading HEAD: a clear method with breaks in the public execution path
We can now separate “is the design coherent?” from “does a fresh clone follow the documented route end to end?” The design has strong source support. The current public snapshot still has reproducible discontinuities. For the audit below, a capability is in the main path only when a real entry point reaches its side effect; a present but unconnected file is disconnected; a documented route whose call chain cannot complete is target route not demonstrated.

7.1 Five discontinuities to resolve before reproduction
| Observation | Source evidence | Status | Practical consequence |
|---|---|---|---|
| The loop instructions are readable | .github/prompts/*-loop.prompt.md defines snapshot, improve, gate, rollback, and test | Main workflow described | The intended steps and permissions can be reconstructed accurately |
HEAD has no scripts/ | run.sh calls scripts/eval_only.py, and pyproject.toml registers script entry points; a later commit removed the directory | Target route not demonstrated | The documented fresh-clone path fails before evaluator entry |
| LiveMath split definitions disagree | The config names 2-1-7_seed42; download and prompt name 2-2-6_seed42; the prompt also alternates between validation size 35 and 18 | Dataset definitions disagree | Different entry paths may evaluate different items, making scores incomparable |
| Round-0 best lacks a full gate | Tag, smoke, and best initialization order | Bootstrap gap | Round 0 should pass full validation before entering the best slot |
| HarnessOpt repeats its path under the new cwd | The prompt first changes into harness_example/spreadsheetbench, then passes the same path to git status/add | Command path breaks | Diff guard can see an empty set and git add fails its pathspec |
| Host coverage remains roadmap work | The README names several hosts while marking Codex CLI and Claude Code plugins TODO | Adapter not delivered | VS Code prompt files are not an installable extension for every coding-agent host |
Why the pathspec fails
cwd = harness_example/spreadsheetbench/
pathspec = harness_example/spreadsheetbench/
resolved intent:
harness_example/spreadsheetbench/
+ harness_example/spreadsheetbench/
result:
no tracked files for that doubled relative path
This is deterministic path resolution, not a guess about model behavior. Run -- . from that cwd, or remain at repository root and keep the full pathspec. The prompt also uses git reset --hard for rollback. In a worktree containing unrelated edits, its blast radius is wider than one harness directory; a production implementation should isolate the run in its own worktree/branch or restore an exact snapshot.
7.2 What the project reports—and what this article does not establish
| Project setup | Reported result | Evidence level here |
|---|---|---|
| LiveMath · GPT-5.4-nano + SkillOpt-Lite | 25.4 points above Full SkillOpt | Project-reported, not reproduced |
| LiveMath · GPT-5.5 + SkillOpt-Lite | 8.8 points above Full SkillOpt | Project-reported, not reproduced |
| ALFWorld · GPT-5.4-nano | 81.3, +9.5 over SkillOpt | Project-reported, not reproduced |
| Spreadsheet average | +12.6 points | Project-reported, not reproduced |
| HarnessOpt · SpreadsheetBench | Nano 0.7758 versus GPT-5.5 standard harness + SkillOpt 0.7620 | Project-reported, not reproduced; the current entry path is incomplete |
These results make the project worth reproducing; they do not establish “fewer modules always wins.” The defensible claim is narrower: a strong coding agent may replace several specialized text-optimization stages with direct file inspection and revision. Whether it is better still depends on the target model, benchmark, repeated validation use, token cost, and host tooling.
8. Where this design transfers well
The portable part is not the slash command. It is the division of responsibility: make the mutable artifact a file, make each observation an addressable record, separate patch generation from retention authority, and keep test sealed until search ends. The implementation can be a coding agent, a CI job, or a smaller dedicated service.
A prompt, skill, policy, or harness file can be precisely allowlisted; every change is diffable and reversible.
Inputs are stable, scoring is deterministic or calibrated, splits are large enough, and rollout cost is affordable.
Without an independent environment outcome, reflection remains a hypothesis; more rounds amplify a coherent story, not evidence.
If a patch can alter permissions, data, billing, or deployment, it needs isolation and a human adoption gate first.
8.1 A minimal transferable rule set
- Fix comparable conditions first: pin target model, harness, data version, seed policy, and scorer; optimize one artifact type at a time.
- Give every trajectory an address: input, output, trace, score, failure reason, and generated artifact must be independently inspectable.
- Repair cross-task patterns only: require multiple independent failures and try to falsify the rule with successful traces.
- Limit the patch: use an allowlist, edit budget, feature flags, snapshots, and rollback.
- Let validation decide retention and run test once: repeated test inspection simply turns test into a new validation set.
- Add a separate adoption check: verify intended outcome, critical regressions, cost/latency, recovery, permissions, and the person responsible for release; benchmark best is not a release decision.
8.2 Final assessment
SkillOpt-Lite's core self-evolution mechanism is real and more disciplined than “ask the agent to reflect.” It has an explicit trigger, fixed snapshot, prompt-level read/write constraints, a valid no-op, durable history, comparable reruns, and restoration of best before the final test. It evolves an external behavioral artifact, not model weights.
The current repository is best understood as a research/engineering snapshot with results, checkpoints, and clear loop instructions, but gaps in its public reproduction steps. Before adopting it in an existing agent, restore the evaluator entry point, unify the split definitions, fix the HarnessOpt pathspec, validate Round 0 before naming it best, and isolate destructive Git rollback. That moves the method from candidate design to verifiable implementation. It still becomes adopted only after business-specific safety, cost, and rollback checks.
