AI handles two of three types of reasoning

To pinpoint where the gap lies, Zahavy draws on a classic distinction from philosopher Charles Sanders Peirce, who categorized all reasoning by how it connects rules, cases, and results.

Deduction derives guaranteed conclusions from fixed rules, like running a program that produces a provably correct output. Induction spots patterns in data: observe a thousand white swans, and you generalize that all swans are white. Abduction is the creative leap. It invents a cause to explain a surprising phenomenon.

This third form is where Zahavy sees the critical bottleneck, and he draws a line between two levels of it. Ordinary abduction picks the most plausible explanation from a set of known candidates, the way a doctor matches symptoms to a disease. Language models can do this, he concedes. The harder version is what he calls "manipulative abduction": inventing a cause for which no linguistic template exists yet. That, he argues, is the real bottleneck of scientific invention, and machines can't do it.

Induction and deduction, the paper argues, are well within reach. Language models already excel at statistical pattern recognition, and they're rapidly conquering formal derivation too. Systems like AlphaProof, Gemini, and GPT-5 now achieve gold-level scores on International Mathematical Olympiad problems. Zahavy even concedes that a language model could probably derive general relativity if given Einstein's assumptions as a starting point. But formulating those assumptions in the first place, making the manipulative leap to reach them, remains the bottleneck.