EchoWM generates navigable worlds with synced video, sound and speech

Researchers describe EchoWM, an omnimodal world model built for enterable generative media: a system that lets someone navigate a scene while it generates 720p video, environmental sound, music and speech at the same time. The model organizes interaction around camera intent rather than treating first- and third-person views the same way. In first-person scenes, the user's input directly specifies the observer's motion. In third-person scenes, the relationship between the camera and the character is instead learned from data, without a separate view-specific controller for each case. To make navigation input usable across both modes, discrete commands and continuous poses are mapped onto a shared metric-scale relative 6-DoF trajectory, with calibration applied at the dataset level so that motion magnitude stays consistent across data drawn from different sources. Training audio-visual generation and trajectory control together relies on a purpose-built, complementary data engine, followed by progressive training and then autoregressive post-training aimed at long-horizon generation, where a scene has to stay coherent well beyond a short clip. The authors report that EchoWM achieves strong trajectory following and high visual quality on public world-model benchmarks, works across both first- and third-person interaction for varied subjects, and keeps environmental sound and speech synchronized with the visuals over long-horizon generation. The source text does not name which benchmarks were used, does not give numeric scores or comparisons against other world models, and does not list the paper's authors, institutions or venue beyond the submitting name attached to the listing. Despite the word "Open" in the title, no model size, training compute, license terms, release date or code repository are mentioned.
Key facts
- EchoWM jointly generates 720p video, environmental sound, music and speech while responding to continuous navigation.
- First-person scenes let user input directly specify observer motion; third-person camera-character dynamics are learned from data instead, with no separate view-specific controller.
- Discrete commands and continuous poses are mapped to a shared metric-scale relative 6-DoF trajectory, with dataset-level calibration preserving motion magnitude across heterogeneous data.
- Training combines a complementary data engine with progressive training followed by autoregressive post-training for long-horizon generation.
- The authors report strong trajectory following and high visual quality on public world-model benchmarks and synchronized sound and speech over long-horizon generation, but the text names no benchmarks and gives no numeric scores.
Why it matters
World models that let a viewer navigate a scene usually generate video alone. EchoWM's pitch is joint generation across video, environmental sound, music and speech, plus a single system that handles both first-person navigation (where the user's motion is the camera) and third-person navigation (where a character moves through a scene the camera tracks), without needing a separate controller built for each case. Unifying those two interaction modes under one shared 6-DoF trajectory representation, and keeping audio and speech in sync as the generated scene stretches over a long horizon, is the specific technical claim being made.
Who it affects
The audience is researchers working on generative world models, video generation, and interactive or navigable media, rather than end users. The text describes a research system evaluated on benchmarks, not a shipped product, app or API.
How to use it
Nothing in the text describes how to access EchoWM. Despite the word "Open" in the paper's title, no code repository, model weights, license terms or release date are given, and no pricing or availability information appears anywhere in the source.
How solid is it
The claims come from the paper's own description of its results: strong trajectory following and high visual quality on unnamed public world-model benchmarks, plus synchronized sound and speech over long-horizon generation. The source text gives no benchmark names, no numeric scores and no comparison against other world models, so the strength of the results cannot be checked from what is available here.
Risks and caveats
The source text does not list the paper's authors, their institutions or a publication venue, and does not state model size, training compute or license terms. That leaves the "Open" framing in the title unsubstantiated by anything in the text itself, and the benchmark claims unverifiable without the underlying scores or baselines.