Vals AI has raised $40 million in a Series A round led by Andreessen Horowitz to build independent evaluation benchmarks for AI models, targeting the growing gap between how language models perform on standardized tests and how they actually behave when doing economically valuable work.

The funding marks a significant leap for a company that was described as bootstrapped with an estimated annual recurring revenue of $1.3 million as recently as late 2025.

What Vals AI actually does

Instead of testing whether an AI can solve abstract logic puzzles or complete sentences from Wikipedia, the company evaluates frontier LLMs on tasks that businesses actually pay for: financial analysis, coding, legal research, and web search.

The company’s flagship product, the Vals Index, aggregates performance data across these real-world categories and ranks models accordingly. The most recent update to the index, from August 2026, showed Claude Fable 5 sitting at the top with a score of 75.14%.