EchoWM: Open and Enterable Omnimodal World Models
JD.com
World model for enterable generative media that responds to continuous navigation while producing 720p video with synchronized environmental sound, music, and speech. Interaction is organized around camera intent: first-person scenes use the camera to define observer motion; third-person scenes learn camera-character dynamics from data. Both discrete commands and continuous poses are mapped to a shared metric-scale relative 6-DoF trajectory with dataset-level motion calibration. A complementary data engine and progressive training scheme plus autoregressive post-training support joint learning of audio-visual generation and trajectory control. Authors report strong trajectory following and visual quality on public world-model benchmarks with both first- and third-person interaction.
Why it matters
HF Daily 66 upvotes on Aug 25. The omnimodal (video + audio + speech) enterable formulation under continuous 6-DoF control is a step beyond prior world models that focused on a single camera motion type or visual-only output; calibration across heterogeneous data sources is the practical lever.
Importance: 3/5
HF Daily 66 upvotes