If you’ve ever watched someone try to figure out the rules of a board game mid-play, you have a rough idea of what Epoch AI is now asking artificial intelligence to do. The nonprofit research institute has rolled out a pair of game-based benchmarks, Mystery Game Puzzles and Chess Puzzles, designed to stress-test the reasoning abilities that AI companies love to brag about.
The early results are humbling. The top score on the Mystery Game Puzzles sits at 59%, and open-weight models max out at just 38%.
What the benchmarks actually measure
Each benchmark consists of 100 programmatically generated puzzles. The Chess Puzzles are relatively straightforward in concept: given a board position, find the best move. The Mystery Game Puzzles are something different entirely.
In the mystery variant, the AI doesn’t even know which game it’s playing. The game’s identity is deliberately obscured, which means the model can’t fall back on memorized patterns or training data shortcuts.








