These projects cover different editable artifacts and feedback mechanisms. The links below open official repositories and resources; they do not imply that a corresponding source-reading chapter has been published here.
Revise skills or harness code from trajectory files
Lite lets a coding agent read trajectories and revise skill.md. A separate HarnessOpt path permits changes to an allowlist of agent Python files.
Check the boundary: Validation and rollback depend on prompt execution discipline; Part 1 explains the gaps in the public reproduction path.
Official documentation and source
Optimize a natural-language skill
Maintain a skill outside the model, turn scored trajectories into reflection, aggregation, candidate selection, and bounded edits, then use validation scores to decide acceptance.
Check the boundary: A validation set repeatedly used for selection is not an independent final test; acceptance rules still need checking for each extension mode.
Official documentation and source
Search for candidates with complementary strengths
Use execution traces to propose textual changes, then score candidates and retain complementary strengths through Pareto-aware selection. The adapter and candidate definition set the editable scope.
Check the boundary: When valset is omitted, GEPA reuses trainset; callers must explicitly provide independent validation data.
Official documentation and source
Propagate textual feedback through a computation graph
Send language-model feedback back to text variables such as prompts, answers, or code, then rewrite those variables with an optimizer. The “gradient” is textual feedback.
Check the boundary: optimizer.step() directly replaces variables. Validation rollback is optional in the official prompt example, rather than a built-in optimizer guarantee.
Official documentation and source · Validation example
Optimize multi-agent workflows
The framework hosts several optimizers. SEW, for example, evaluates prompt or workflow-structure changes on development data, records snapshots, and selects higher-scoring workflows.
Check the boundary: Rules differ across optimizers. Callers must supply appropriate development and test sets; SEW is not a framework-wide guarantee.
Official documentation and source