I want to be upfront about something: I didn't build this. Our engineering team did. But I get to write about it, and I've been waiting a while to write this one.
This week we submitted the Backboard CLI to the official Terminal-Bench 2.1 leaderboard. The score: 85.4% ± 0.8%, with a pass@5 of 0.888, running Claude Opus 4.8 via Bedrock.
For context, the top published entries on the leaderboard right now are Claude Code with Fable 5 at 83.8% and Codex with GPT-5.5 at 83.1%. Those are Anthropic's and OpenAI's own coding agents. Built by the labs that built the models.
Our submission is above every published result. It's pending review on the leaderboard now, and you can go look at every trial yourself.
Who we are






