I had a dumb little experiment I wanted to try.

And, as dumb little experiments sometimes do, it got way more interesting than I expected.

I gave three coding agents the exact same task:

Claude (Opus 5)

Codex (GPT 5.6 Sol)