Survey of 700+ practitioners finds data work beats model scale

Survey of 700+ practitioners finds data work beats model scale

IEEE Spectrum and Wiley published a white paper, sponsored by Voxel51, presenting the results of a 2026 survey of more than 700 professionals working on visual and physical AI, systems built on video, LiDAR point clouds, sensor streams and other high-dimensional data that let machines perceive, reason and act in physical space.

The report frames this as a shift in where AI progress is happening: the last decade, it says, was built on text, but the frontier has moved to data from the physical world.

On adoption, 78% of the surveyed teams say they already see measurable value from visual and physical AI, while 74% still consider the field underinvested relative to its opportunity. On what actually makes production succeed, the report finds that data problems cause the majority of model failures, and that curating data matters more than chasing larger architectures. Teams that ship successfully invest nearly 3 times more time in data work than teams that struggle.

Annotation is called out as a specific source of waste: teams commonly label everything and then discard much of it before it ever reaches production. The report's headline conclusion is that data work, not data collection, is what separates teams that ship from teams that stall. It also reports that 92% of practitioners hold a view on where the field is heading next, without detailing what that direction is.

Key facts

  • The white paper is based on a 2026 survey of more than 700 professionals building visual and physical AI, published by IEEE Spectrum and Wiley and sponsored by Voxel51.
  • 78% of surveyed teams already see measurable value from visual and physical AI, while 74% still consider the field underinvested relative to its opportunity.
  • Teams that ship successfully invest nearly 3 times more time in data work than teams that struggle.
  • Data problems cause the majority of model failures, and curating data matters more than chasing larger architectures, according to the report.
  • Annotation remains costly and wasteful because teams often label everything and then discard much of it before production.

Why it matters

The report positions visual and physical AI, systems trained on video, LiDAR point clouds and sensor streams rather than text, as the next frontier after a decade of text-driven AI progress. Its central claim cuts against the assumption that bigger models are the main lever for production success: across the surveyed teams, data problems caused the majority of model failures, and curating data mattered more than chasing larger architectures. For an industry that has spent years chasing parameter counts, that redirects the question of where investment should actually go.

Who it affects

The findings speak most directly to teams building physical AI, robotics, autonomous systems and other applications that reason over video, LiDAR and sensor data, and to the vendors, like survey sponsor Voxel51, that sell tools for managing and curating that data. The gap the survey highlights, 78% of teams already seeing measurable value against 74% who still call the field underinvested, points to a market still working out how much to commit and where.

How to use it

The white paper is a free download, gated behind a 'Look Inside' click on the IEEE Spectrum and Wiley page sponsored by Voxel51. It functions as a benchmark: the report suggests teams aiming to ship should expect to spend nearly 3 times as much time on data work as teams that are struggling, and should treat annotation quality, not annotation volume, as the priority, since surveyed teams often label everything and then discard much of it before production.

How solid is it

The findings rest on self-reported answers from more than 700 practitioners surveyed in 2026; no specific month or date, sample selection method, margin of error, or question wording is disclosed. The paper is sponsored by Voxel51, a company that sells data curation and annotation tooling for visual and physical AI, giving it a direct commercial interest in a finding that data curation matters more than model scale. No breakdown of respondents by role, company size, industry or geography is given, and no specific examples of which model failures or projects the data problems caused are cited.

Risks and caveats

As presented, the material is promotional: a sponsored white paper distributed as a lead-generation download rather than an independently reviewed study, and the reported 92% of practitioners with a view on where the field is heading next is cited without saying what that view actually is. Readers should treat the headline figures, 78%, 74%, nearly 3 times, 92%, as marketing statistics from the sponsor's own survey rather than independently verified research, and should not assume they generalize beyond the population Voxel51 and its partners reached.