Today we're opening up Favur Evals — a public, live leaderboard for a question we kept needing an answer to ourselves: which model should I actually put on a coding job?
Here's how it works. We take a fixed statement of work and hand it to a full team of AI agents — planner, architect, tester, coder, reviewer, builder. Then we run that same job again and again, swapping only the model underneath the team:
all-Qwen, all-OpenAI, all-DeepSeek, all-Gemini, and so on. Same task, same
scaffolding, different brain. Every run gets scored across the whole lifecycle — the code, the tests, the cost, the discipline — not just the final diff.
What you'll find on the board







