Every few months a wave of “it got worse” posts appears, followed by a wave of “it is your imagination” replies. Both sides are arguing without data, because almost nobody froze a test set before the thing they are arguing about happened. This page is about not being in that position.
The recurring argument
The claim is hard to evaluate because so many things are changing at once: the model behind an alias, the provider’s serving stack, your own prompts, your own context, the difficulty of the work you are throwing at it, and your expectations. Perceived degradation is genuinely common — novelty wears off, the easy wins were taken first, and today’s baseline is last quarter’s ceiling. Actual degradation also genuinely happens. Nothing about the shape of the complaint distinguishes them.
The study everyone cites
Chen, Zaharia and Zou’s How Is ChatGPT’s Behavior Changing over Time? (2023) compared the March and June 2023 snapshots of GPT-3.5 and GPT-4 across several tasks: prime identification, code generation, visual reasoning and answering sensitive questions. The headline finding was that behaviour changed substantially over three months in both directions, with the widely quoted case being GPT-4’s accuracy on the prime-identification task falling from around 84% to around 51%. The code-generation result — a large drop in directly executable output — was also widely shared.






