H3-World adapts the MiniMax-H3 video generator into an interactive world model by converting character and camera actions into compositional natural-language instructions aligned with video latent intervals. It uses temporal attention routing for precise temporal control and lightweight LoRA adaptation (0.199% of parameters) to learn action-conditioned visual dynamics from just 8,000 gameplay samples. The system generalizes to unseen action combinations and visually distinct scenarios beyond its training distribution.