Most AI research tools point a language model at the web and trust the output. Saarth Shah built Sixtyfour around the opposite instinct: grade everything, ship only what improves the score.
Saarth Shah keeps a scoreboard. Every build of Sixtyfour’s research agents is graded against questions a team of experts assembles by hand and checks against real-world cases, vertical by vertical, and the grade is the only thing that decides what ships. The discipline behind that scoreboard is why a payments company will let his software, rather than a human analyst, decide whether a stranger is real.
Most AI research tools start from the same shortcut: point a language model at the open web, ask it to look something up, and let it write. The output reads well. What Shah fixed on early was whether it was right, and how anyone would know.
For a while, the industry blamed hallucination, the models’ habit of inventing facts and sources. That problem is real, and it is slowly getting better. Frontier models hallucinate less than they did a year ago, and pairing them with live web search has narrowed the gap further. But the limit Shah cares about is not the model’s imagination. It is the model’s reach.









