ACE-Data-0 packs 17M frames of home robot data across 75,000 episodes

ACE-Data-0 packs 17M frames of home robot data across 75,000 episodes

Researchers led by Yukang Cao describe the Ambient Capture Engine (ACE), a system that converts ordinary home environments into spatially calibrated, temporally synchronized recording setups for capturing how humans perceive and act in the world. The motivation is a data bottleneck in embodied intelligence: models need to learn how first-person perception, whole-body motion, dexterous hand manipulation, object state, sound, and touch evolve together as a person pursues a goal, but existing datasets split this experience across separate viewpoints, modalities, or spatial scales, so the full perception-action loop is only ever partially captured.

ACE runs at two complementary scales. A table-scale configuration is built to resolve fine hand-object manipulation. A room-scale configuration captures whole-body motion, locomotion, and interactions across a furnished home. Across both, ACE records egocentric (first-person) and multi-view exocentric (third-person) video, full-body and articulated hand motion, object geometry with 6-DoF (six degrees of freedom) trajectories, audio, and tactile signals, all as one unified multisensory stream rather than as separate, misaligned recordings.

Using this engine, the team built ACE-Data-0: 150 hours and 17 million video frames, spanning 200 task categories, performed by 50 participants across 2 environments, for a total of 75,000 interaction episodes. The dataset covers atomic manipulation tasks, long-horizon chains of household activities, and general human-scene interaction. Participants were given goal-level instructions rather than step-by-step scripts, which the authors say preserves natural variation in how people actually perform tasks.

Alongside the dataset, the authors introduce a hierarchical benchmark that progresses from raw signals to scene components and then to full interactions. Evaluating current state-of-the-art methods on this benchmark exposes substantial gaps under contact, occlusion, egomotion (the camera's own motion as the wearer moves), and long time horizons. The authors position ACE-Data-0 as a scalable foundation for imitation learning, world models, vision-language-action systems, and embodied AI more broadly, since it provides synchronized human demonstrations with aligned perceptual, kinematic, and contact supervision.

Key facts

  • ACE-Data-0 contains 150 hours and 17 million video frames across 200 task categories, collected from 50 participants in 2 environments.
  • The dataset totals 75,000 interaction episodes, covering atomic manipulation, long-horizon household activity chains, and human-scene interaction.
  • The Ambient Capture Engine (ACE) runs at table scale (hand-object manipulation) and room scale (whole-body motion and locomotion), recording egocentric and multi-view exocentric video, body and hand motion, 6-DoF object trajectories, audio, and touch as one synchronized stream.
  • Tasks were directed with goal-level instructions rather than step-by-step scripts, to preserve natural variation in how people perform them.
  • A hierarchical benchmark built on the dataset shows current state-of-the-art methods have substantial gaps under contact, occlusion, egomotion, and long temporal horizons.

Why it matters

Embodied AI models such as robot policies and world models need to learn how perception and action unfold together over time, but until now datasets have fragmented that experience across separate viewpoints, sensors, or spatial scales. ACE is built specifically to solve the synchronization and coverage problem by capturing egocentric video, exocentric video, motion, object trajectories, audio, and touch in one aligned stream, and to do it inside real homes rather than staged labs.

Who it affects

Researchers and teams building imitation learning systems, world models, and vision-language-action systems for robotics stand to gain a large, richly annotated dataset. The benchmark also gives anyone evaluating embodied AI methods a structured way to test performance from raw signals up through full interactions.

How to use it

The paper describes ACE-Data-0 as a dataset and accompanying benchmark rather than a product; the text does not state a license, price, or access process, so use is limited to what the paper itself documents: the data engine's design, the dataset's scale, and the benchmark's structure.

How solid is it

The dataset scale is concretely specified: 150 hours, 17 million frames, 200 task categories, 50 participants, 2 environments, 75,000 episodes. The authors also ran a hierarchical benchmark against state-of-the-art methods and report the gaps found, which supports the dataset's value as a genuine stress test rather than an easy benchmark.

Risks and caveats

The evaluation itself surfaces the caveat: current state-of-the-art methods show substantial gaps under contact, occlusion, egomotion, and long temporal horizons on this benchmark, meaning the dataset reveals unsolved problems in embodied AI rather than easy wins. The paper covers only 2 environments and 50 participants, so how well the data generalizes beyond those homes and people is not addressed in the text.