Vals raises $40M to fix AI benchmarking that companies keep gaming
The startup says legacy benchmarks can't keep up with modern AI. Its solution: private tests that measure real-world impact, not just abstract intelligence.
The startup says legacy benchmarks can't keep up with modern AI. Its solution: private tests that measure real-world impact, not just abstract intelligence.
The AI industry has a measurement problem. Companies release new models faster than academics can benchmark them, and when the metrics favor a particular system, suddenly that company’s marketing team has gold. Good benchmarks equal good PR, which means there’s enormous incentive to find ways around them.
Enter Vals, a 2024 startup that just raised $40 million in Series A funding led by Andreessen Horowitz. The company is betting that the benchmarking world needs a complete overhaul.
Rayan Krishnan, the 25-year-old co-founder, watched this problem unfold firsthand. After interning at Palantir and working at Microsoft and Stanford’s AI lab, he saw capable models hitting the market at breakneck speed while academic benchmarks couldn’t keep pace. “We were seeing a bunch of new, very capable models come to market quickly, and the academic benchmarks were not keeping up with that frontier advance,” Krishnan explains.
The deeper issue: most benchmarking systems measure abstract intelligence. Can a model pass a bar exam? Does it know enough trivia? These questions miss what actually matters in the real world.
Vals takes a different approach. Instead of publicly available tests that companies can potentially train against (gaming the system), Vals keeps its test materials private. More importantly, it evaluates models on their ability to complete complex, domain-specific tasks in law, finance, and coding. The question shifts from “does the model know stuff” to “can it actually do the work.”
“What we’re doing is actually looking at what are the real impacts of the models,” Krishnan says. “Can they do work that produces a product of the same quality as a human within every domain?”
The startup doesn’t just check for positive outcomes either. Vals examines potential harms: what happens if these models are deployed without guardrails? This darker assessment includes benchmarks on recursive self-improvement, mental health applications, cybersecurity, biosecurity, and even how models understand the Geneva Convention.
Companies actually pay Vals to test their models, which might sound counterintuitive. Why pay to learn you’re failing? But effective measurement drives improvement. Think of it like the SAT: students pay College Board to identify gaps and track progress.
More importantly, these evaluations are becoming decision-making tools for companies acquiring new AI models. That’s real market value.
The company’s trajectory tells you everything. Revenue is eight times what it was last year. The team started 2024 with eight people and has already tripled to 25. Krishnan plans to add another 10 to 15 employees and relocate to a much larger San Francisco office.
Vals also recently launched a program providing model evaluations to federal agencies, signaling that benchmarking has moved beyond internal corporate interests into regulatory territory.
Krishnan believes this is just the beginning. As AI companies go public, benchmarking will become central to how they communicate with investors and regulators. “AI companies are starting to go public. Anthropic is slated for later this year. I suspect OpenAI will be public soon,” he notes. “As AI models become a core part of the economy, the types of benchmarks and evaluations that we do are going to drive their usage and be a central part of how these companies submit public filings.”
The startup is essentially building the infrastructure that will determine which AI systems the world actually trusts. In an industry obsessed with benchmarks as marketing tools, someone had to create benchmarks that actually measure what matters.
Source: TechCrunch