Ploy builds production marketing websites with an AI agent. They've been benchmarking every frontier release for months. Nothing beat Claude Opus until GPT-5.6 Sol — 2.2× faster, 27% cheaper, better visual scores.

Then they actually tried to ship it. That's where it got interesting.

"We use Vercel's AI SDK, but switching from Claude Opus 4.8 to GPT-5.6 Sol still exposed provider-specific assumptions throughout our stack."

What the numbers look like

Claude Opus 4.8