Over one year ago, we got a lot of response for our preprint "Large language models do not simulate human psychology". While psychologists seemed to be somewhat puzzled that such claims are made in the first place, some colleagues from computer science or application reacted quite skeptically to our claim - we even got some comments that our paper would never make it through peer review.
@bpaassen congratulations! 🎉
Our core argument is that LLMs should not be regarded as reliable simulators of human participants because subtle changes to the input text can lead to very different reactions from humans and LLMs. This is also shown in our key figure below. The x axis corresponds to different moral scenarios, the y axis to the reaction if we subtly change the text of the scenario. You can see that the reactions of LLMs and humans are clearly very different for most scenarios (exact correlations in the paper).
I want to thank all authors for their tireless work on this - but also our reviewers for some helpful discussion and feedback. Some highlights of what we changed since the preprint:
1. We have now evaluated more LLMs, including more recent ones. Our core effect is still maintained: under rewording, the correlation between human responses and LLM responses significantly reduces. But we also find differences between LLMs: Older and smaller LLMs tend to just ignore the rewordings, whereas newer LLMs do react to them - but not in the same way humans do.
Well, it turns out, it has. The journal "Technology, Mind, and Behavior" just published our paper under the new title "Exploring the Limits of Large Language Models as a Reflection of Human Cognition: An Illustration in the Context of Moral Judgment".
https://doi.org/10.1037/tmb0000210
Exploring the limits of large language models as a reflection of human cognition: An illustration in the context of moral judgment.