I run a public board that probes 16 LLMs on a frozen 35-task suite, once a day, and keeps every score. When a model drops against its previous run, it opens a GitHub issue by itself and writes me a draft post.
Between 21 and 24 July it did that four times:
23 Jul Gemini 3.5 Flash -11.4 pts
24 Jul Gemini 3.1 Pro -2.9 pts
21 Jul Grok 4.3 -5.7 pts






