The three-way comparison you actually need is not the one on the leaderboards

Pick any two of ChatGPT, Claude and Gemini and there is a benchmark where each one wins. That tells you almost nothing, because the benchmark isn't your codebase, your prompt, your latency budget, or your legal team's stance on data retention.

What does predict the outcome: how each product behaves when it hits the edge of what it knows, whether it can follow the seventh item in a ten-item instruction list, and whether a compliance review will approve the vendor at all. Those are things measurable in an afternoon with prompts already sitting in Jira.

This article is about running that afternoon.

Behavioural differences people consistently report