Public LLM benchmarks are just vibes. So @sergical built his own.
In this video, he instruments a Next.js app with Sentry to capture every real conversation (inputs, outputs, cost, tool calls), and then feeds it into @braintrust to eval models on quality, speed, and cost.
The








