Mind attribution, or ascribing mental capacities such as consciousness, emotion, and agency to an entity, shapes human moral judgment. As LLMs increasingly mediate consequential decisions, the folk psychology that they inherit from their training data could become consequential if, like humans, it also informs their behaviors and moral judgments. This project characterizes the structure of mind attribution in instruction-tuned LLMs, its internal representational basis, and its causal relationship to moral-status judgment. Adapting Gray, Gray, and Wegner's (2007) mind-perception survey to six instruction-tuned models, we measure per-entity Experience and Agency attributions across thirteen entities, probe internal activations for a corresponding representational geometry using representational similarity analysis, and test whether behavioral mind attribution predicts the model's own moral-status judgments. We then move from correlation to intervention, ablating the mind attribution direction and measuring the consequences for both mind and moral-status judgments. We ask which entities models over- and under-attribute relative to the human reference, and how a model characterizes itself within that space, and compare to public perceptions of AI minds. Together these results characterize mental state attributions in LLMs, establish whether models use information about mind to assign moral standing, and identify where those attributions diverge from human folk psychology.
Google Research & University of London, USA
Google Research & University of London, UK
I’m currently studying perception of mind and human-AI interaction at Princeton University, where I use cognitive neuroscience techniques to investigate brain mechanisms supporting consciousness and perception of mind. Lately, my research has explored behavioral outcomes and how our brains respond when engaging with AI systems, as well as how AI works using mechanistic interpretability methods. Outside of the lab, I’ve been actively involved in projects that promote responsible AI development, driven by both enthusiasm for AI’s potential and a recognition of the ethical concerns it raises.
"I'm most excited by the opportunity to collaborate across philosophy, cognitive science, and machine learning, at a time where questions surrounding AI sentience are becoming increasingly urgent."
AI Sentience Scholars Dezhi Luo and Mariana Amendoeira Duarte, AI Sentience Mentor Geoff Keeling, and AI Sentience Scholar Rachel Metzgar (left to right) meet up at the Eleos Conference on AI Consciousness and Welfare in Berkeley, California. Learn more.