Micro1 reaches $500M gross run rate amid AI training-data boom

Micro1, a four-year-old startup that supplies AI training data, has expanded its gross annual run rate from $100 million to $500 million over the past eight months, according to a person familiar with the company. Like other companies in the space, Micro1 hires domain experts such as doctors, lawyers and scientists on a contract basis to label and generate data; it keeps roughly 60% to 70% of that gross figure as net revenue, putting its net annual run rate at $150 million to $200 million.
Micro1 still trails larger rivals: Mercor hit $2 billion in gross annualized revenue this summer, and Handshake reached $1 billion earlier this year. But Micro1’s own growth, TechCrunch notes, shows that demand for AI training data is large enough to support multiple competing suppliers rather than a single winner. Some researchers have hypothesized that future AI spending on data could eventually rival spending on compute.
Part of the growth comes from margin, not just volume. Micro1 is increasingly generating synthetic data without direct human involvement, for instance by producing automated descriptions of video content. Some of this data can also be resold to multiple customers as “off-the-shelf” datasets, which a person familiar with Micro1’s finances said carries gross margins as high as 80% to 90%. Contract sizes are growing at an accelerated pace, and Micro1 expects its overall margins to keep expanding.
Reselling the same datasets to multiple clients has also drawn criticism: critics argue that supplying off-the-shelf training data to Chinese AI developers helps their models match the capability of leading US systems. Micro1 founder Ali Ansari addressed the issue in a post on X last month, saying that, unlike some competitors, Micro1 does not sell its data to Chinese model makers, and citing Kimi K3 as an example of what he said results when rival data companies do.
Micro1 began as an AI recruiting startup. Ansari pivoted the company into data labeling after noticing that clients were using his recruiting platform to vet and hire engineers specifically for annotation work. He has previously told TechCrunch that, beyond having its experts evaluate model outputs, a practice known as reinforcement learning gyms, Micro1 is building a robotics pre-training dataset by having hundreds of generalists record everyday object interactions inside their own homes.
Micro1 raised its Series A at a $500 million valuation last September, a figure that happens to match its current gross run rate but is a separate, older metric describing the company’s worth rather than its revenue. TechCrunch reports that it understands Micro1 may have since raised another round at a significantly higher valuation, though that has not been confirmed or quantified. Micro1 did not respond to a request for comment for the story.
Key facts
- Micro1’s gross annual run rate grew from $100 million to $500 million over the past eight months; after keeping 60% to 70% of that as net revenue, its net annual run rate is $150 million to $200 million.
- Rivals are still bigger: Mercor hit $2 billion in gross annualized revenue this summer and Handshake reached $1 billion earlier this year; even so, Micro1’s growth suggests the AI training-data market can support several large suppliers at once.
- Reusable “off-the-shelf” datasets, sold to multiple customers and increasingly generated synthetically (such as automated video descriptions), carry gross margins of 80% to 90%.
- Founder Ali Ansari says Micro1, unlike some competitors, does not sell data to Chinese model makers, and cited Kimi K3 as an example of what happens when rival data companies do.
- Micro1 raised its Series A at a $500 million valuation last September and may have since raised another round at a significantly higher valuation, though TechCrunch says that is unconfirmed.
Why it matters
The AI industry’s demand for training data is now large enough to support several billion-dollar-scale suppliers rather than a single dominant one. Micro1’s gross run rate growing five-fold in eight months, alongside Mercor’s $2 billion and Handshake’s $1 billion in annualized revenue, indicates that data has become a significant AI spending category in its own right, alongside compute. Some researchers have even hypothesized that future AI spending on data could eventually rival spending on compute, a shift that would reshape where AI money flows.
Who it affects
AI labs and large corporations buying training data are the direct customers; domain experts such as doctors, lawyers and scientists who work as contractors for Micro1 and similar firms supply it. Competing data-labeling startups, including Mercor and Handshake, are affected by how large the addressable market turns out to be. The report also touches the broader dispute over reselling training data to Chinese AI developers: unnamed critics quoted in the piece argue the practice helps Chinese models catch up to leading US systems, while Micro1’s founder disputes that his company is among those doing it.
How to use it
There is no consumer product here; the news is a set of financial figures about Micro1’s business. For anyone assessing the AI data-labeling market, the useful numbers are the gross-to-net split (Micro1 keeps 60% to 70% of gross revenue as net) and the margin structure: standard contract labeling carries lower margins, while reusable “off-the-shelf” data resold to several customers reaches 80% to 90% gross margins. That split is a useful lens for comparing Micro1 to Mercor and Handshake, whose reported figures in the source are not broken out the same way.
How solid is it
The core figures (the $500 million gross run rate, the 60% to 70% retention rate, and the 80% to 90% margin on off-the-shelf data) all come from two unnamed sources, described only as “a person familiar with the company” and “a person familiar with the startup’s finances,” not from Micro1 itself; the company did not respond to TechCrunch’s request for comment. Every date in the story is relative to publication only, phrases like “last month,” “last September,” “this summer” and “earlier this year,” rather than an absolute calendar date. The claim that Micro1 may have raised a new round at a higher valuation is explicitly presented by TechCrunch as something it merely understands may have happened, not a confirmed, quantified event.
Risks and caveats
All of the headline financial figures rest on unnamed sourcing rather than Micro1’s own disclosure, so they read as reported estimates rather than audited numbers. The report does not say what share of Micro1’s current revenue is synthetic versus human-labeled, only that synthetic generation is “increasingly” used, so the underlying data mix is unclear. The dispute over data sold to Chinese developers rests on Ansari’s own denial and unnamed critics’ argument, with no independent verification of either side, and the report does not explain what Kimi K3 is beyond that passing mention.
“Some human data companies work with foreign adversaries. [A]nd the results show today in Kimi K3. We believe it’s shameful to claim American AI dominance desires while selling millions worth of data to countries that we are in adversarial competition with.”
— Ali Ansari, Micro1 founder, in a post on X