Researchers developed EchoWM, an omnimodal world model that generates synchronized 720p video, environmental sound, music, and speech while responding to continuous navigation inputs. The system organizes interaction through camera intent, mapping discrete commands and continuous poses to a shared metric-scale 6-DoF trajectory with dataset-level calibration. A complementary data engine alongside progressive training and autoregressive post-training enables long-horizon generation, with evaluations showing strong trajectory following, high visual quality, and sustained audio-visual synchronization across first- and third-person perspectives.