Researchers keep using LLMs as fake humans without agreeing what “human” means
A new review divides language-model proxies into four roles and argues that each needs a different test before its behavior can support a scientific claim.
OddBrief EditorialAI-assisted, human-reviewed
AIKey facts
- Publication
- Nature Computational Science review
- Framework
- Four roles for LLM human proxies
- Roles
- Believable agents, task agents, experimental subjects and silicon samples
- Claim
- Human similarity requires role-specific validity tests
- Caveat
- The paper reviews existing work rather than introducing a new benchmark
Large language models are increasingly being asked to stand in for people: as survey respondents, experimental participants, customers, voters and characters in simulations. A new review in Nature Computational Science argues that researchers are treating all of those substitutions as one idea when they actually support very different claims.
The authors organize LLM-based human proxies into four roles: believable agents, task agents, experimental subjects and silicon samples. Each role imitates a different slice of human behavior and therefore needs a different kind of validation.
One model, four scientific jobs
A believable agent is judged by whether people experience its behavior as coherent and humanlike. A task agent must perform a role well enough to accomplish an objective. Those standards can overlap, but neither proves that the model predicts how a population would behave.
Experimental subjects are used in place of people to test interventions or theories. Silicon samples go further, attempting to reproduce the distribution of responses across demographic or social groups. In those settings, surface plausibility is a weak test. A model can sound convincing while systematically missing the causes and variation behind real decisions.
The review emphasizes that similarity to humans is not a single property. Matching language style, average survey answers or one task score does not establish validity for another use.
Proxies can repeat the data they were trained on
LLMs learn patterns from large collections of human-produced text. That makes them useful approximations in some contexts, but it also creates a circularity risk. A model may reproduce familiar findings because those findings were present in its training data, not because it independently simulates the mechanism under study.
Models also compress differences. A prompt describing age, income or political identity may not reconstruct the lived experience associated with those labels. When a simulated sample appears cleaner than human data, the missing noise may be the most important result.
The paper is a review of research across disciplines rather than a new benchmark. Its framework is intended to help researchers connect each use to the evidence required.
Cheap participants can make expensive mistakes
Synthetic subjects are attractive because they are fast, reproducible and inexpensive. They can help explore a study design before recruiting people or generate hypotheses for later testing.
The danger begins when convenience becomes evidence. Policy, product and social-science decisions can look empirically grounded even when the “participants” are instances of the same model responding to slightly different prompts.
The review does not argue that LLM proxies are useless. It asks researchers to state exactly what the proxy represents, what human reference it was tested against and where the comparison fails.
A machine can imitate a person for a conversation. Turning that performance into a claim about people requires a much harder experiment.
Sources
- Large language models as human proxiesNature Computational Scienceprimary source


