a16z-backed Vals raises $40M to fix AI benchmarking
Original: Vals, backed by Andreessen Horowitz, is looking to become the gold standard for AI benchmarking
Why This Matters
As AI claims flood enterprise sales pitches, independent task-based evaluation is gaining commercial urgency.
Vals, a San Francisco-based AI benchmarking startup founded in 2024, raised $40 million in a Series A led by Andreessen Horowitz. The company targets flaws in legacy benchmarks by testing models on real-world industry tasks — in law, finance, and coding — rather than abstract knowledge exams, and keeps its test materials private to prevent training-data gaming.
AI benchmarking has long served double duty: validating model capabilities internally while functioning as marketing ammunition externally. The problem is that legacy benchmarks, many designed for an earlier era of AI, haven't kept pace with frontier models — and companies have learned to game them by training against publicly available test sets.
Vals wants to close that gap. Founded in 2024 by 25-year-old Stanford alum Rayan Krishnan — who previously interned at Palantir and worked with Microsoft and Stanford's AI lab — the startup argues that benchmarks should answer one concrete question: can a model actually do the work a human does, at the same quality, within a specific domain?
To prevent cheating, Vals does not publish its test materials. Instead of measuring broad knowledge (think bar-exam-style trivia), it evaluates models on task completion across law, finance, and coding, checking for both successful outputs and failure modes.
The company secured an earlier seed round from 8VC and Bloomberg Beta before closing its $40 million Series A last month, led by Andreessen Horowitz. Krishnan told TechCrunch: "We were seeing a bunch of new, very capable models come to market quickly, and the academic benchmarks [were] not keeping up with that frontier advance." The team operates out of a converted 19th-century brewery on San Francisco's Folsom Street.