A skill library should not only grow. Mutation children enter as trial, seed skills that survive pre-retirement start as active or stable, and anything the improved policy outgrows is demoted, rewritten or retired.
Every method in this family keeps the library append-only. A skill that was correct at step 20 encodes a procedure the model has outgrown by step 120 — and it is still being retrieved into the context, still phrased as an instruction.
Each skill carries a fitness fi = Ci / Ui and a usage count Ui. Active is the hub — nothing leaves the library without passing through it. Hover a state or an arrow.
One parent, two children: the variant that adds a verification step stabilizes, the one that imposes a fixed ordering is retired at step 65. The lifecycle keeps the first and drops the second.
ALFWorld, WebShop and Search-Augmented QA. SkillForge beats the strongest memory-augmented RL baseline on all three, and the closed-source GPT-4o / Gemini-2.5-Pro references by a wide margin on the agentic environments.
All four series of the paper's Fig. 3(b). SkillForge separates from SkillRL after step 75 and holds near 90% while SkillRL oscillates in the 80–88% band; both stay well clear of the two GRPO variants.
Seven datasets over the same 51,713-sample union. SkillForge is best or second-best on every axis; the largest gains over the strongest baseline are on Bamboogle (+3.4) and HotpotQA (+0.7).
Mutation explores both directions from the same parent. The surviving child often has a lower aggregate fitness, because it targets a narrower niche — the lifecycle scores within-niche accuracy, not popularity.
@misc{ge2026skillforgecoevolvingskillsagents,
title={SkillForge: Co-Evolving Skills and Agents via Dynamic Skill Lifecycles},
author={Yuyao Ge and Yiwei Wang and Yuchen He and Baolong Bi and
Lingrui Mei and Jiayu Yao and Lizhe Chen and Shenghua Liu},
year={2026},
eprint={2610.09832},
archivePrefix={arXiv},
primaryClass={cs.AI},
url={https://arxiv.org/abs/2610.09832},
}