Researchers developed a full-stack framework comprising GameUI-Taxonomy and G2WEngine that automatically extracts reusable UI assets from real gameplay footage and synthesizes temporally coherent UI overlays, backed by a dataset of 96,000 synthetic paired videos and 1,079 in-the-wild clips spanning 303 games. The team also introduced GameCleaner, a mask-free model using multimodal semantic understanding and video editing to identify and remove HUD elements while preserving scene integrity. GameCleaner reached 95.36 average accuracy on synthetic footage and 80.05 on real gameplay, and world models trained on the cleaned footage showed a 6.83% improvement in VideoReward scores.
