No synthetic puzzles, no contaminated suites: 25 tasks harvested from PRs merged in 2026 across five languages, graded by each project's own held-out tests. Four agents at stock settings. octomind with an open model solved 24/25 - ahead of Claude Code with Opus - while the same model in another harness solved 19 at double the cost. Here's how we built the benchmark and what it taught us.
We wanted one number we could actually trust: if you hand a coding agent the kind of task a maintainer faces on a normal Tuesday - a real bug, a real feature request, in a real codebase - how often does it deliver a fix the project's own test suite accepts?
None of the public benchmarks could give us that number, so we built octobench. This post is the story of how, what broke along the way, and what the scoreboard says.
Why not an existing benchmark
Two problems kept biting us with the popular suites.








