Researchers introduce LongWoF-Bench, a benchmark of 778 machine-verifiable tasks spanning code generation, agent synthesis, mathematical reasoning, and rule-following, alongside EvoMap, a system that consolidates verifier-confirmed execution trajectories into structured “Genes” to enable knowledge reuse across models. Testing on 252 tasks with verified trajectories shows evolved genes outperform baseline skills by 8.7-15.5 percentage points across seven models, complete 39 additional tasks, and reduce token consumption by 9.9% for Claude Opus.