I Read 3,000 Papers Across 12 Fields. Five Patterns Kept Appearing.
Same AI model. 6.7% accuracy.
Change the interface around it. 68.3%.
The researchers didn’t touch the model. They changed the format it used to express edits. The model had been reasoning correctly the whole time. It just couldn’t express its answers without corrupting them.
The failure looked like stupidity. It was a formatting problem.






