Agent Self-Evolution

How agents improve from feedback

Start with SkillOpt-Lite, then compare what SkillOpt, GEPA, TextGrad, and EvoAgentX change, how they obtain feedback, and who decides what to keep.

Questions for reading self-evolution projects: feedback produces candidate edits, evaluation compares baseline and candidate, and a decision retains or rolls back the change.
A conceptual reading framework; verify which steps each project actually implements.

Reading route. First replay one revision with SkillOpt-Lite, then compare fuller skill optimization, textual feedback, and workflow search. A better score, a passed independent test, and production adoption are three separate decisions.

Published chapters

Begin with one concrete revision

Part 1 establishes the shared questions: who triggers a change, which files may change, what data evaluates it, and how the result is retained or rolled back.

Open-source reading

From skill edits to textual feedback and workflow search

These projects cover different editable artifacts and feedback mechanisms. The links below open official repositories and resources; they do not imply that a corresponding source-reading chapter has been published here.

Microsoft’s SkillOpt and SkillOpt-Lite live in separate repositories. Lite’s official description says it is built on SkillOpt and hands the optimization loop to a coding agent. HarnessOpt extends the scope to agent code in the Lite repository. The Full trainer bundled with Lite is not identical to the current Microsoft implementation.

SkillOpt-Lite / HarnessOpt

Revise skills or harness code from trajectory files

Lite lets a coding agent read trajectories and revise skill.md. A separate HarnessOpt path permits changes to an allowlist of agent Python files.

Check the boundary: Validation and rollback depend on prompt execution discipline; Part 1 explains the gaps in the public reproduction path.

SkillOpt

Optimize a natural-language skill

Maintain a skill outside the model, turn scored trajectories into reflection, aggregation, candidate selection, and bounded edits, then use validation scores to decide acceptance.

Check the boundary: A validation set repeatedly used for selection is not an independent final test; acceptance rules still need checking for each extension mode.

GEPA

Search for candidates with complementary strengths

Use execution traces to propose textual changes, then score candidates and retain complementary strengths through Pareto-aware selection. The adapter and candidate definition set the editable scope.

Check the boundary: When valset is omitted, GEPA reuses trainset; callers must explicitly provide independent validation data.

TextGrad

Propagate textual feedback through a computation graph

Send language-model feedback back to text variables such as prompts, answers, or code, then rewrite those variables with an optimizer. The “gradient” is textual feedback.

Check the boundary: optimizer.step() directly replaces variables. Validation rollback is optional in the official prompt example, rather than a built-in optimizer guarantee.

EvoAgentX

Optimize multi-agent workflows

The framework hosts several optimizers. SEW, for example, evaluates prompt or workflow-structure changes on development data, records snapshots, and selects higher-scoring workflows.

Check the boundary: Rules differ across optimizers. Callers must supply appropriate development and test sets; SEW is not a framework-wide guarantee.