Same DeepSeek V4 Flash. Different runtime. Very different long-task outcomes.
My local sample is bounded: Codex + Flash completed a long, cross-file, repeatedly verified deck task; Claude Code + Flash launched multiple reviews, but their quality was not independently verified. This supports different pairings, not a universal ranking.
The useful unit is not a model ID. It is a complete runtime: model × protocol × tools × context × recovery × acceptance.
Four layers where the result changes
Protocol is a trajectory interface. It defines how goals, tool results, intermediate state, and continuation are represented. A compatibility layer can connect successfully and still lose long-horizon affordances.









