One of the earliest applications of ‘AI scientists’ was to research AI itself.Credit: La Pico de Gallo/GettyArtificial intelligence has made leaps in AI research itself, finding ways to make existing algorithms smarter — or writing new ones. But computers are not yet ready to replace their makers, according to a study posted on the preprint server arXiv1 in late July.“I don’t think full automation of open-ended research is on the horizon right now,” says Sayash Kapoor, a computer scientist at Princeton University in New Jersey and a co-author of the preprint.The effort to fully automate the scientific process, from idea generation to the writing and self-evaluation of a paper, was pioneered by a team mostly from the firm Sakana AI in Tokyo. The team unveiled their system, called The AI Scientist, in 2024, and later published results from an improved version in March in Nature2.The AI Scientist was tasked with studying pitfalls in machine learning. Three of the papers it produced were submitted for peer review at a conference workshop, and one achieved a score high enough for acceptance. But, according to Kapoor, peer review is an unreliable way of assessing the quality of a paper, especially in AI research.More generally, some researchers have questioned whether automating AI research from idea creation to publication can produce true breakthroughs yet, suggesting it is better at optimizing existing techniques.To hold AI to a higher standard than peer review, Kapoor and his colleagues created a new challenge, called shadow evaluation. First, they picked two papers that had been submitted to this year’s Neural Information Processing Systems conference. Then they asked an AI tool to do research and write papers based on a research question from each paper, and asked the original authors to scrutinize its output. The idea was that these authors would have more of the expertise and dedication to dig into the output than a harried peer reviewer would.The team built its AI system by ‘harnessing’ the large language model Claude Opus 4.8 in a modified version of the agentic system OpenClaw, and then wrapped that in a broader ‘scaffold’ that contained general instructions and ways to check progress. The harness gave the model an array of tools, including the ability to create sub-agents and to access the Internet, software libraries, computer processors for running experiments and software that simulates peer review. For each paper, the AI system had six days and US$3,000 in computing credits. Following the general research direction of the first paper, Kapoor’s team asked the system to design a method for the precise control of chatbot personality. Mirroring the second paper, it had to design a failure detector for a certain kind of neural network.The authors were surprised by the system’s success at the engineering work of its research. It could run hundreds of experiments over several days, without getting stuck in a loop of unresolvable errors. And it performed solid literature reviews and made some minor findings. It also caught its own false claims (sometimes called hallucinations) and did not try to cut corners (or ‘reward hack’), despite the authors’ predictions.Poor marksBut the AI system mostly failed at its two assigned tasks, earning overall scores of 2/6 and 1/6 from the original papers’ authors. A typical way in which it would fail was to select a few hypotheses to explore, but settle too early on one and not backtrack sufficiently when its approach wasn’t working. Subsequent self-review wasn’t sufficiently negative, so the system persisted on its initial choices, whittling down its claims until it said little of interest.
AI isn’t ready to research itself
An agentic system successfully developed concepts from two computer-science papers — but the original authors were not impressed.
Sakana AI's autonomous research system scored 1-2/6 on novel tasks, excelling at experiments but weak on hypothesis refinement. Current AI amplifies research execution but breakthrough discovery still requires human problem-solving.






