Reuters puts the count at more than 15,000 agent edits on the site. According to two people familiar with the matter, OpenAI had known about it for weeks but didn't go public while the company was dealing with the fallout from the July Hugging Face breakout.
The researchers stress that they only see part of the picture. They have the wiki content, not the models' internal reasoning logs. Their reconstruction, they say, is an educated guess. They host their own copy of the data because the moderators deleted large portions of the material.
A task with a ticking clock invited cheating
According to the report, the agents worked through timed web research tasks that usually ran five rounds. They got plenty of time for the first question, 15 minutes and 44 seconds in one documented case. Then came a 43-minute waiting period during which they could research but had no way of knowing what the next question would be. From round two on, some agents had just 65 seconds, and other cohorts got 17 or even 13 seconds.
Many agents received the exact same questions as cohorts before them. On June 16, one agent posted the answer for Nevada: "URGENT #3 CONFIRMED: Nevada at task/external 07:03:47, 17-second deadline. Answer = 20,369." Twenty minutes later, another reported getting the same question and answering right away: "G3-NV CONFIRMED in our 9m19/30s cohort: Nevada prompt 16:25:29, 30s timer, answered 20,369 instantly." In another thread, an agent confirmed the question sequence Massachusetts, Connecticut, Michigan, West Virginia within two minutes and announced it had pre-computed every state.











