CogEvol is a family of models that automatically generate learning environments by converting course briefs into structured JSON slides or interactive HTML pages in a single pass, using a production-grounded data pipeline that created 53,687 verified training samples from real failures and hybrid rule-plus-VLM rewards with GRPO-based reinforcement learning. CogEvol-27B scores 83.7 on slide quality and 63.7 on a 500-case interactive-HTML benchmark with 26.9 times fewer parameters than flagship coding models, and the authors release a smaller 4B variant open-source.
