Researchers introduce WorldToken, a robotic imitation learning approach that organizes multimodal inputs through a time-first architecture, fusing multiview images, proprioception, and task conditioning into one world token per policy timestep before a causal temporal Transformer processes the sequence and a diffusion head generates action outputs. Testing on RoboCasa tasks shows an 85.3-million-parameter model achieving 59.45% mean closed-loop success, with experiments revealing that visible temporal history significantly affects performance across all tested policies.