Among the many predictions about the future of artificial intelligence is that models will one day be able to conduct scientific research on their own, leaving humans out of the equation. Already, they can write code, run experiments and search scientific literature, but carrying out open-ended research would require a significant leap in ability.

In a paper posted on the arXiv preprint server, researchers tested AI's ability to conduct open-ended research and found that it came up short.

Putting AI to the test

The study authors gave frontier agents (cutting-edge, state-of-the-art AI tools designed to carry out complex, multi-step tasks autonomously) six days to conduct research and write papers based on two then-unpublished AI conference submissions. This ensured they couldn't just find the answers online.

The agents had full access to the internet, dedicated computing power and approximately $3,000 in model-use credits, meaning they had a budget to conduct open-ended exploration and run experiments. The topics they had to research and write about were the structure and controllability of language-model personas and designing a detector for distribution shifts in tabular foundation models.