Tech
EN AZ
Vals, backed by Andreessen Horowitz, is looking to become the gold standard for AI benchmarking

Vals, backed by Andreessen Horowitz, is looking to become the gold standard for AI benchmarking

techcrunch.com 19.09.2026 15:00 4 views
Vals AI is hoping to make AI benchmarking a more neutral and trustworthy resource in a world increasingly inundated by AI models.

Benchmarking has become the industry norm for how AI companies validate their models’ capabilities and, when the metrics swing in their favor, stand out from competitors and advertise their superiority. In other words, good benchmarks pretty much always mean good PR. Unfortunately, companies have also figured out how to outwit legacy benchmarking systems — many of which are older, and not built to measure the capabilities of modern models.

Vals, a startup formed in 2024, says that it is on a mission to fix this very imperfect system. In the span of less than two years, the company has established itself as a notable presence in the tech industry and, last year it managed to secure a seed round led by 8VC and Bloomberg Beta. Then, last month, after a period of rapid growth, it raised $40 million in a series A led by Andreessen Horowitz.

Rayan Krishnan, the company’s 25-year-old co-founder, previously interned at Palantir, and, as an undergraduate at Stanford, worked for Microsoft and the school’s much lauded artificial intelligence lab. Krishnan says Vals was born from his own observations about how benchmarking was falling behind the advances of the industry it was designed to measure. With AI being integrated into every part of society, benchmarks should really exist to verify that models can do what companies advertise they can do, Krishnan said.

Last week, the young founder showed me around his company’s two-floor office on San Francisco’s Folsom Street — an old brick building that, a century ago, served as the site of a large brewery. Instead of an industrial output of beer, the historical structure is now home to a number of different startups looking to ship the future of the tech industry. While many benchmarking systems offer tests that are publicly available (this can allow a company to train its model against those tests, thus arguably cheating on their exam), Vals doesn’t publicly disclose its specific test materials.

Instead of measuring an AI model’s general knowledge, Vals also evaluates models on their ability to complete complex tasks associated with specific industries like law, finance, and coding. The hope is to analyze how, “if these models ran wild in the world, what the negative implications would be.” The capabilities that Vals is measuring are growing. In additional to more traditional industries, the startup continues to push into more unique terrain.

We’re doing some work in mental health, cybersecurity, biosecurity, and even law of armed conflict to models to understand how to apply the Geneva Convention,” Krishnan shares. Companies pay Vals to test their models, which can be an odd concept to wrap your head around. Why would a company pay to learn its model isn’t performing well?

Extract — continue reading at the source.

Read full story