Disclosure: I maintain Lynkr, the router being benchmarked, so read everything here with that in mind. The mitigation: every number in this post comes from RouterArena's automated evaluation pipeline, run by their CI on their infrastructure, not by me — including the numbers that make us look mediocre.

Every LLM router's README — ours included — makes the same claim: it sends easy queries to cheap models and hard queries to good ones, and saves you money without hurting quality. Almost none of them attach evidence. The numbers that do exist are self-reported, measured on datasets the vendor picked, with baselines the vendor chose.

So when RouterArena showed up — an open, standardized benchmark for LLM routers out of the RouteWorks group (paper, leaderboard) — we submitted Lynkr to it. This post is what we learned, including the metrics where we did badly, because a benchmark result you only quote selectively is just marketing with extra steps.

What RouterArena actually measures

RouterArena evaluates a router over 8,400 queries spanning 9 domains and 44 categories at three difficulty levels, and scores it on five axes: