A while back I built a small tool called SDKProof. it checks how well an AI coding agent writes an SDK's current API — the stuff that changed in the last major, that the model tends to get wrong because it learned the old version.
Claude Opus 5 came out today. so I re-ran the whole board on it.
short version: it fixed last year's SDKs. it did not fix this year's.
The board, now on Opus 5
Same tasks, same libraries, new model:









