I started with a simple assumption: if one coding agent was useful, several

coding agents working in parallel would be even better.

Like most, I figured out that works until a point where you are just copying and pasting outputs between sessions, and sessions would get worse on delivering as time and context grew.

Codex, Claude Code, and GitHub Copilot could each write, research, review, or test a bounded piece of work. The problem was everything between those pieces.

Agents duplicated effort, missed repository-specific instructions, reviewed stale changes, and returned individually plausible results that did not form a coherent whole.