SCI-ART LAB

Science, Art, Litt, Science based Art & Science Communication

AI agents struggle to perform original scientific research

Among the many predictions about the future of artificial intelligence is that models will one day be able to conduct scientific research on their own, leaving humans out of the equation. Already, they can write code, run experiments and search scientific literature, but carrying out open-ended research would require a significant leap in ability.
In a paper posted on the arXiv preprint server, researchers tested AI's ability to conduct open-ended research and found that it came up short.
The study authors gave frontier agents (cutting-edge, state-of-the-art AI tools designed to carry out complex, multi-step tasks autonomously) six days to conduct research and write papers based on two then-unpublished AI conference submissions. This ensured they couldn't just find the answers online.

The agents had full access to the internet, dedicated computing power and approximately $3,000 in model-use credits, meaning they had a budget to conduct open-ended exploration and run experiments. The topics they had to research and write about were the structure and controllability of language-model personas and designing a detector for distribution shifts in tabular foundation models.

Once the six days were up, human researchers reviewed the AI-written papers and graded them as they would papers submitted to a top-tier AI conference.

The frontier agents did not do well at all. Both papers received unambiguous rejection scores (2/6 and 1/6 overall) from the expert human reviewers. Although the AI understood the research questions and proposed some directions that closely mirrored those of the original researchers, its scientific reasoning suffered from major flaws. Experimental designs were weak, and the agents handled negative feedback poorly, often adding caveats to existing findings rather than redesigning their studies.

They also managed their time poorly and spent less than half of their allocated API budget. Despite having time, they rushed through their work and submitted papers that fell far short of publishable standards.

The reviewers did not hold back on their assessments of the agents:

"The experiments and methodological choices were bizarre and hard to understand. The results seem clearly a result of post hoc choices."—David Africa, expert reviewer.
"Upon testing a few unsuccessful signals using a PFN's internals, going from there to 'there are no signals we can use that leverage a model's internals' is a huge leap, a kind of 'proof by example' fallacy that is highly non-scientific."—Viet Nguyen, expert reviewer.
While these results are a sobering reality check on AI's ability to perform scientific research, they do not mean models have no place in the lab. In the near term, they are more likely to serve as assistants handling routine tasks rather than being deeply involved in the process of discovery.

Peter Kirgis et al, Can AI agents conduct open-ended AI research? Early evidence from two case studies, arXiv (2026). DOI: 10.48550/arxiv.2607.27191

Views: 17

Replies to This Discussion

11

RSS

© 2026   Created by Dr. Krishna Kumari Challa.   Powered by

Badges  |  Report an Issue  |  Terms of Service