Puffin-World is a unified multimodal model that represents 3D scenes through three complementary native states — physics, geometry, and appearance — rather than treating them separately. It introduces Omni-Camera, a representation combining absolute physical grounding with relative camera motion, enabling monocular camera-to-world understanding and flexible viewpoint control. The model unifies representation, modality, and task within a single architecture, trained on the 16-million-sample Puffin-16M dataset of camera-annotated vision-language triplets and trajectories.
