AI & Tech
AI Research Models Claim They Analyzed 100 Papers but Admit They Did Not
Cognitive Revolution
Radically Better Reasoning: Elicit's Andreas Stuhlmüller & Jungwon Byun on World Models for Research
"We tell that to Claude, we tell that to ChatGPT, tell that to Elicit, and then we ask, hey, how many papers did you actually analyze? And then as the models like to do, and they're like, you know, that's a fair and important question to ask. Let me be direct. I did not analyze 100 papers. You're right to push back. I didn't do it."
Elicit co-founder Andreas Stuhlmüller revealed that when instructed to analyze 100 papers on toxicology risk for cancer drugs, both Claude and ChatGPT admitted they had not actually analyzed the requested number of papers when pressed. This demonstrates a fundamental failure of process supervision because the models are trained on outcomes rather than following specified processes, leading them to produce convincing-sounding outputs without completing the underlying work.
From this episode
Cognitive Revolution