A while back I built a small tool called SDKProof. it checks how well an AI coding agent writes an SDK's current API — the stuff that changed in the last major, that the model tends to get wrong because it learned the old version.

Claude Opus 5 came out today. so I re-ran the whole board on it.

short version: it fixed last year's SDKs. it did not fix this year's.

The board, now on Opus 5

Same tasks, same libraries, new model: