EchoWM: Open and Enterable Omnimodal World Models

JD.com

Research official + media 2 src. ~1 min

World model for enterable generative media that responds to continuous navigation while producing 720p video with synchronized environmental sound, music, and speech. Interaction is organized around camera intent: first-person scenes use the camera to define observer motion; third-person scenes learn camera-character dynamics from data. Both discrete commands and continuous poses are mapped to a shared metric-scale relative 6-DoF trajectory with dataset-level motion calibration. A complementary data engine and progressive training scheme plus autoregressive post-training support joint learning of audio-visual generation and trajectory control. Authors report strong trajectory following and visual quality on public world-model benchmarks with both first- and third-person interaction.

Why it matters

HF Daily 66 upvotes on Aug 25. The omnimodal (video + audio + speech) enterable formulation under continuous 6-DoF control is a step beyond prior world models that focused on a single camera motion type or visual-only output; calibration across heterogeneous data sources is the practical lever.

Importance: 3/5

HF Daily 66 upvotes

Sources

official arXiv abstract