Same DeepSeek V4 Flash. Different runtime. Very different long-task outcomes.

My local sample is bounded: Codex + Flash completed a long, cross-file, repeatedly verified deck task; Claude Code + Flash launched multiple reviews, but their quality was not independently verified. This supports different pairings, not a universal ranking.

The useful unit is not a model ID. It is a complete runtime: model × protocol × tools × context × recovery × acceptance.

Four layers where the result changes

Protocol is a trajectory interface. It defines how goals, tool results, intermediate state, and continuation are represented. A compatibility layer can connect successfully and still lose long-horizon affordances.