EvoGenUI-Bench evaluates LLMs’ ability to maintain and evolve interactive web interfaces across multiple turns, comprising 150 five-turn tasks across three scenarios with evaluation via browser execution, screenshots, and DOM analysis. Even the strongest models tested achieve only 74.9% per-turn accuracy and 37.3% success on complete five-turn episodes, highlighting the difficulty of maintaining consistency across sequential interface modifications.