Take Caution in Using LLMs as Human Surrogates: Scylla Ex Machina
Yuan Gao; Dokyun Lee; Gordon Burtch; Sina Fazelpour · 2024 · arXiv
WASTE classifies this as Negative / Null Result Report · AI classification, approximate
The study found no significant effect — useful as a negative control or null benchmark for your own design.
Abstract (excerpt)
Recent studies suggest large language models (LLMs) can exhibit human-like reasoning, aligning with human behavior in economic experiments, surveys, and political discourse. This has led many to propose that LLMs can be used as surrogates or simulations for humans in social science research. However, LLMs differ fundamentally from humans, relying on probabilistic patterns, absent the embodied experiences or survival objectives that shape human cognition. We assess the reasoning depth of LLMs using the 11-20 money request game. Nearly all advanced approaches fail to replicate human behavior dis
Excerpt shown for reference under fair use — read the full paper at the publisher.
About to run something similar?
Run an AI Precheck on your own design to catch failure modes like this one before you spend the time. Your first desk check is free.
Related failures
The Oregon Experiment — Effects of Medicaid on Clinical Outcomes
Negative / Null Result ReportMicrocredit in Theory and Practice: Using Randomized Credit Scoring for Impact Evaluation
Negative / Null Result ReportThe Cost of Carbon: Capital Market Effects of the Proposed Emission Trading Scheme (ETS)
Negative / Null Result ReportPushing on a string: US monetary policy is less powerful in recessions ∗
Negative / Null Result ReportThe Evidence on Globalisation
Negative / Null Result ReportStatistical tests for power-law cross-correlated processes
WASTE indexes this work — it does not host or republish it. Failure-type classification is automated and approximate.
Metadata source: arXiv
