Vals AI Wants to Become the Benchmark Layer for Models

Every AI company says its latest model is better. Vals AI wants to become the company that decides whether that claim actually holds up.
The San Francisco startup is building independent benchmarks designed to evaluate how AI models perform on real-world tasks rather than relying only on widely used academic tests.
Founded in 2024, Vals raised a $40 million Series A led by Andreessen Horowitz last month after previously securing seed funding from investors including 8VC and Bloomberg Beta.
The company is entering a market that barely existed a few years ago.
AI evaluation is becoming infrastructure.
Traditional benchmarks are getting easier to game
AI benchmarks historically worked a little like exams.
Give a model a collection of questions.
Measure how many answers are correct.
Compare the score with competing systems.
There is one obvious problem.
If the exam becomes public, model builders can train around it.
A benchmark that once measured intelligence can slowly become something closer to a test-preparation exercise.
Vals tries to avoid that by keeping much of its test material private.
That makes the evaluation harder to optimize against artificially.
The company cares about work, not trivia
Vals is also shifting the focus away from abstract intelligence tests.
Instead, it evaluates whether models can perform tasks associated with industries such as:
law,
finance,
coding,
cybersecurity,
biosecurity,
and other specialized domains.
The question isn't simply:
Does this model know the answer?
It's closer to:
Could this model produce work comparable to what a capable human professional would deliver?
That distinction becomes more important as businesses begin selecting models for production systems.
AI buyers need neutral measurements
Enterprises now face a confusing purchasing decision.
OpenAI may claim one model is strongest.
Anthropic publishes its own evaluations.
Google has its benchmark results.
Open-model providers highlight different tests.
Every company naturally prefers metrics that make its product look strong.
That creates room for an independent evaluator.
A law firm choosing between several AI systems may care far more about contract analysis than a generalized reasoning leaderboard.
A bank may care about financial workflows.
A government may care about cyber resilience or policy interpretation.
Vals wants to supply those domain-specific measurements.
Revenue has grown eightfold
The company says its revenue is currently eight times higher than a year ago.
Its workforce has also grown from eight employees at the beginning of 2026 to around 25, with further hiring planned.
Those numbers remain small compared with frontier AI labs.
But the growth suggests model evaluation itself is becoming a business.
Companies pay Vals to discover where their systems fail and how they compare with alternatives.
It is somewhat counterintuitive.
A company is paying someone to expose weaknesses in its product.
But reliable evaluation helps identify problems before customers do.
Benchmarks could become financial infrastructure
Vals co-founder Rayan Krishnan sees an even larger opportunity.
As AI companies mature and potentially enter public markets, model evaluations could begin influencing investment decisions, corporate disclosures and purchasing.
That would make benchmarks function less like academic research and more like standardized ratings.
The analogy might eventually look closer to credit ratings, safety certifications or standardized testing.
A neutral score becomes useful because many parties need a common point of comparison.
Safety could become another major market
Vals isn't only measuring productivity.
It has also developed evaluations around areas such as recursive self-improvement, mental health, cybersecurity, biosecurity and the laws of armed conflict.
This matters as regulators and enterprises increasingly ask whether an AI system is safe enough for a particular environment.
A model capable of performing well on coding tasks may also need to demonstrate that it does not behave dangerously when connected to sensitive systems.
Capability and safety evaluations are therefore starting to merge.
The benchmarking market could become crowded
Vals is not alone.
AI labs run their own internal evaluation teams.
Research organizations develop independent tests.
Safety startups are building certification systems.
Governments are increasingly interested in standardized AI evaluations.
The opportunity is large.
So is the competition.
Vals' advantage may depend on becoming a trusted neutral brand before evaluation standards consolidate.
What happens next?
AI models are improving so quickly that benchmark scores can become outdated within months.
That means evaluation companies need to move almost as fast as the labs themselves.
If Vals succeeds, the most important number attached to a future AI model may not come directly from the company that built it.
It could come from an independent evaluator telling customers what the model can actually do.
AI companies built the intelligence race.
Vals wants to build the scoreboard everyone trusts.
