Snorkel AI’s $350M Round Puts Training Data Back on Top

AI models keep getting smarter. That is making the data needed to improve them harder — and more valuable — to produce.

Snorkel AI has raised a $350 million Series E at a $3.5 billion valuation, nearly tripling its valuation from the $1.3 billion attached to its previous round 17 months ago.

Insight Partners and S32 co-led the financing, with investors including Addition, Greylock, Lightspeed, GV and Wells Fargo also participating.

The funding comes amid explosive demand for something the AI industry once treated as relatively unglamorous:

training data.

Snorkel says revenue has grown 18-fold

The startup says its annualized revenue run rate has reached roughly $375 million, an 18-fold increase over the previous 12 months. Reuters separately reported a figure around $350 million and said CEO Alex Ratner expects the company to reach profitability this year.

That growth reflects an important change in model development.

The early large-language-model era relied heavily on massive quantities of internet data.

The next generation increasingly needs something more difficult.

Expert reasoning.

Specialized examples.

Synthetic environments.

And data designed specifically to teach models complicated tasks.

The easy internet data has already been harvested

AI labs have already trained on enormous portions of the easily accessible web.

More data still exists.

But simply adding another billion generic webpages does not necessarily teach a model how to become a better financial analyst, software engineer or scientific researcher.

Frontier models increasingly require carefully constructed examples representing the behavior developers want them to learn.

That can mean working with domain experts.

Creating simulated environments.

Generating synthetic examples.

Testing the model.

Finding where it fails.

Then producing new data targeted at those weaknesses.

Training data is becoming less like mining the internet and more like engineering a curriculum.

Snorkel changed its business to follow that market

Snorkel originally built software designed to help customers automate data labeling.

Last year it moved deeper into what it calls data-as-a-service, delivering completed datasets and reinforcement-learning environments rather than only giving customers tools to build them.

That shift mirrors a wider change across AI data companies.

Customers increasingly do not want another platform.

They want the finished training resource.

This creates an interesting new category of AI company: part software platform, part research lab and part expert workforce.

AI data is becoming a crowded gold rush

Snorkel is not alone.

TechCrunch reported extraordinary annualized revenue figures across companies working around AI training data, including Mercor, Handshake and Micro1. Some of these figures represent gross revenue from businesses that pass a large share of sales directly to human experts, so comparisons require caution.

Still, the direction is obvious.

AI labs are spending aggressively on high-quality data.

And investors increasingly believe the businesses supplying that data can become multibillion-dollar companies themselves.

Synthetic data is part of the answer

Human experts are expensive.

They are also difficult to scale.

Synthetic data can help.

Models generate examples.

Software filters them.

Experts verify or improve important cases.

The combination can generate substantially more training material than relying entirely on manual annotation.

Snorkel's approach blends software, synthetic generation and domain expertise.

That matters because the future data market probably will not be a simple choice between human-created and machine-created information.

It will increasingly be hybrid.

Better models actually create more demand for data

There is a paradox here.

It would be reasonable to assume smarter AI eventually reduces the need for data suppliers.

So far, the opposite appears to be happening.

A more capable model creates demand for harder tasks.

Harder tasks require richer evaluation environments.

Those environments reveal new weaknesses.

Developers then need additional data to address them.

Model intelligence therefore raises the quality bar for the data ecosystem underneath it.

What happens next?

Snorkel's $3.5 billion valuation is a bet that AI development will continue requiring large amounts of specialized training infrastructure.

If that thesis holds, one of the biggest businesses in the frontier-AI economy may not build a frontier model at all.

It may build the questions, environments and examples that make the next model smarter.


Our latest news