Vals raises $40M Series A led by Andreessen Horowitz for AI benchmarking

Vals, an AI-benchmarking startup founded in 2024, raised a $40 million Series A led by Andreessen Horowitz last month, following an earlier seed round led by 8VC and Bloomberg Beta. Co-founder Rayan Krishnan, 25, previously interned at Palantir and, as a Stanford undergraduate, worked for both Microsoft and Stanford's AI lab. He says Vals was born from watching benchmarking fall behind the pace of model releases: "We were seeing a bunch of new, very capable models come to market quickly, and the academic benchmarks [were] not keeping up with that frontier advance."
Vals positions itself against legacy benchmarking systems that publish their test materials, which Krishnan says lets companies train their models against the tests and effectively cheat on the exam. Vals keeps its specific test materials private instead. Rather than testing general knowledge, it evaluates models on complex, industry-specific work in fields like law, finance and coding, checking not only for good outcomes but for negative ones too, modeling what could go wrong "if these models ran wild in the world." Its benchmark coverage now extends into recursive self-improvement, mental health, cybersecurity, biosecurity, and how models apply the law of armed conflict, including the Geneva Convention.
Companies pay Vals to test their own models, a model Krishnan compares to a student paying the College Board to take the SAT: an uncomfortable purchase that nonetheless helps a company find and fix weaknesses. Vals also recently launched a program evaluating models for federal agencies. The startup's revenue is now eight times what it was a year ago. Headcount has grown from eight people at the start of the year to a team of 25, and Krishnan plans to add another 10 to 15 people and move Vals out of its current two-floor San Francisco office into a bigger one. He frames the company's evaluations as increasingly central to how AI companies operate as they head toward public markets, citing SpaceX's IPO, Anthropic's planned listing later this year and his expectation that OpenAI will go public soon.
Key facts
- Vals raised a $40 million Series A led by Andreessen Horowitz last month, after an earlier seed round led by 8VC and Bloomberg Beta.
- Co-founder Rayan Krishnan, 25, previously interned at Palantir and worked for Microsoft and Stanford's AI lab as a Stanford undergraduate.
- Vals keeps its test materials private, unlike benchmarks that publish theirs and risk being trained against, and evaluates models on complex, industry-specific tasks in law, finance and coding, plus areas like biosecurity and the law of armed conflict.
- The startup's revenue is now eight times what it was a year ago, and headcount has grown from eight people at the start of the year to a team of 25, with plans to add 10 to 15 more.
- Vals has launched a program providing model evaluations to federal agencies.
Why it matters
Benchmarks double as marketing: a good score becomes a company's PR line, and Krishnan's pitch is that many of the benchmarks producing those scores are old, publicly disclosed, and gameable by training against them. Vals is betting that as AI models get folded into more of the economy, and as AI companies themselves move toward public markets, there will be sustained demand for evaluations that measure whether a model can actually do the specialized work it is sold as doing, not just answer exam-style questions.
Who it affects
Directly: AI model developers who pay Vals to evaluate their own models, plus federal agencies now covered by its new evaluations program. Indirectly: any industry Vals builds domain benchmarks for, including law, finance, coding, mental health, cybersecurity, biosecurity, and questions of how models handle the law of armed conflict.
How to use it
Vals sells evaluations directly to AI companies, who pay to have their models tested, and now also runs a program evaluating models for federal agencies. The source gives no pricing or contract terms, so none should be inferred.
How solid is it
The account rests on a single TechCrunch interview with co-founder Rayan Krishnan and an office visit; the funding, revenue and headcount figures are the company's own disclosures rather than independently verified numbers. No valuation, exact closing date for the Series A, or dollar figure for the earlier seed round is given.
Risks and caveats
Vals's growth story, revenue multiple and headcount figures are self-reported by the startup itself, with no third-party confirmation in the source. Keeping test materials undisclosed helps prevent models from being trained against them, but it also means outside parties cannot independently audit what Vals is actually measuring.
“What we're doing is actually looking at what are the real impacts of the models. Can they do work that produces a product of the same quality as a human within every domain?”
— Rayan Krishnan, Vals co-founder