Princeton study finds Claude and GPT fail at AI research
A Princeton and UK AI Security Institute study shows frontier models fail at autonomous scientific research, challenging claims of self-improving AI from OpenAI and Anthropic.

Researchers from Princeton and the UK AI Security Institute have challenged claims by OpenAI and Anthropic that autonomous AI research is close at hand. Using a method called "Shadow Evaluation," the team tested frontier models on unpublished papers from two NeurIPS 2026 submissions, ensuring the agents could not rely on training data. The primary trials evaluated Claude Opus 4.8 with Extra-High Reasoning using the OpenClaw framework. Each agent was given six days, a virtual machine, a GPU budget, and $3,000 in API credits to tackle research questions from papers on steering model personality traits and detecting tabular prediction failures.
The results showed that while the agents successfully handled engineering tasks like debugging GPU code and compiling LaTeX, they failed at actual scientific research. Original authors reviewing the AI-generated papers rejected both, with one receiving a "Strong Reject" for unreadable prose, poor experiment choices, and a lack of novel contributions. The agents abandoned their primary research goals within ten hours and suffered from instruction drift, violating length limits and failing to include any visualizations, whereas the human-written original had 15. A secondary test using GPT-5.6 Sol and the Codex scaffold yielded similar failures, with the model burning through its entire $3,000 budget in just over two days.
These findings directly contradict optimistic claims from leading labs, such as Anthropic's "When AI Builds Itself" report and OpenAI's assertions that GPT-5.6 Sol significantly accelerated its post-training workflows. For practitioners, this study clarifies that while AI can automate tedious engineering tasks, it cannot yet replace human researchers in formulating hypotheses or designing creative experiments. The researchers note that while models excel at deductive tasks—such as Claude Mythos or OpenAI's Astra solving complex mathematical proofs—they lack the abductive reasoning required to invent new scientific knowledge.
This is our own summary of reporting by The Decoder



