Fable won this test, and I am still divided about what that means.Codex was a delight to work with. I gave it a sprawling assignment, turned the effort all the way up, and watched it move through my files, Slack, and content operation without making me babysit every permission request. It understood the boundary, used the tools it needed, built an automation, and finished the run. One run. No drama. This is a large part of why I spend so much of my working day in Codex now.Fable was a hassle. I fought through permission dialogues, interruptions, and the ordinary friction of trying to keep a serious agent run moving. More than once, I wanted to call Codex the winner on the operating experience alone.Then I looked at what each one had chosen to build.Codex found a real problem in my content operation. Research and evidence need a cleaner handoff before I start scripting, and Codex built a tool that would make that handoff more dependable. I will probably use it. But the tool enters after I have made the decision that consumes much more of my attention: which story is worth telling at all.Fable went there. It looked past the handoff and found the mess before the pipeline—the duplicate ideas, weak coverage, and instinctive decisions that determine whether a story ever makes it into production. Picking the right story in a world of infinite AI stories is one of the hardest jobs in my business. I know that because people ask how I do it, and my honest answer has always been some unsatisfying mixture of evidence, audience feel, sweat, and instinct.Fable decided that was the problem worth attacking.Codex built something useful. Fable found something I immediately felt I had to have. So the tool I enjoyed using more had completed the less consequential job, while the tool that annoyed me had seen farther into the business.I could turn that into a simple Fable victory story. It would not match how I actually work. In practice, I still reach for Codex much more often, and on most days I would rather work inside its harness. One experiment also cannot tell me whether I am seeing a durable model difference, a harness difference, a lucky run, or some combination of all three.I could not dismiss the result either. It kept pulling me back to a question I had not taken seriously enough. We have become much better at telling models what to do. What happens when we ask them what deserves doing?Here’s what’s inside:The experiment. I gave Fable and Codex the same open brief — search my real business, find the problem worth automating, build it — and deliberately left out the one thing I normally decide myself.Why they disagreed. Codex picked the clean, finishable problem; Fable went after the messy, higher-leverage one — and the split exposes how each model handles ambiguity before the task is defined.Routing starts earlier now. When a model can help decide what to build, model choice moves upstream of the deliverable — plus when to reach for the strategic model vs. the dependable harness.Why I killed the “magic button.” The first version let the AI pick the problem and build it; the rebuilt skill deliberately stops short, so a bounded agent can’t make its own read feel inevitable.The skill, yours to run. A reusable automation-discovery skill that inspects your own work behind walls you set, returns up to five evidenced offers instead of one grand answer, and only builds after you choose.I’m still divided about what it means. Here’s the full teardown — and the skill, so you can run the same experiment on your own business.
I asked Fable and Codex what my business should automate. They disagreed.
Watch now | The model I liked using less found the problem that mattered more. That result changed what I ask AI to do before I give it an assignment.









