Rachel Metzgar

Rachel Metzgar

Round Pushpin United StatesGraduation Cap PhD StudentClassical Building Princeton UniversityBooks Neuroscience/Cognitive Science


Project

Mind and Moral Status Attribution in Large Language Models

Project Overview

Mind attribution, or ascribing mental capacities such as consciousness, emotion, and agency to an entity, shapes human moral judgment. As LLMs increasingly mediate consequential decisions, the folk psychology that they inherit from their training data could become consequential if, like humans, it also informs their behaviors and moral judgments. This project characterizes the structure of mind attribution in instruction-tuned LLMs, its internal representational basis, and its causal relationship to moral-status judgment. Adapting Gray, Gray, and Wegner's (2007) mind-perception survey to six instruction-tuned models, we measure per-entity Experience and Agency attributions across thirteen entities, probe internal activations for a corresponding representational geometry using representational similarity analysis, and test whether behavioral mind attribution predicts the model's own moral-status judgments. We then move from correlation to intervention, ablating the mind attribution direction and measuring the consequences for both mind and moral-status judgments. We ask which entities models over- and under-attribute relative to the human reference, and how a model characterizes itself within that space, and compare to public perceptions of AI minds. Together these results characterize mental state attributions in LLMs, establish whether models use information about mind to assign moral standing, and identify where those attributions diverge from human folk psychology.


Mentors

Name
Title
Afiliation

https://slite.com/api/files/nMNvdHXntywbLO/Winnie%20Street.png?apiToken=eyJhbGciOiJIUzI1NiIsImtpZCI6IjIwMjMtMDUtMDQifQ.eyJzY29wZSI6Im5vdGUtZXhwb3J0IiwibmlkIjoiVl9KeFlOR0pEbWhyR2kiLCJpYXQiOjE3OTExODQ2NDgsImlzcyI6Imh0dHBzOi8vc2xpdGUuY29tIiwianRpIjoiNTlrY0xfOUpPREdTRFUiLCJleHAiOjE3OTM3NzY2NDh9.unFVTDOnhgDmVQwMjypdsKMJ1cU0QkE_VxnpOKtAqCQ
Google Research & University of London, USA

https://slite.com/api/files/EIF_FLXk-8CG2c/Geoff%20Keeling.png?apiToken=eyJhbGciOiJIUzI1NiIsImtpZCI6IjIwMjMtMDUtMDQifQ.eyJzY29wZSI6Im5vdGUtZXhwb3J0IiwibmlkIjoiVl9KeFlOR0pEbWhyR2kiLCJpYXQiOjE3OTExODQ2NDgsImlzcyI6Imh0dHBzOi8vc2xpdGUuY29tIiwianRpIjoiNTlrY0xfOUpPREdTRFUiLCJleHAiOjE3OTM3NzY2NDh9.unFVTDOnhgDmVQwMjypdsKMJ1cU0QkE_VxnpOKtAqCQ
Google Research & University of London, UK


About the Scholar

I’m currently studying perception of mind and human-AI interaction at Princeton University, where I use cognitive neuroscience techniques to investigate brain mechanisms supporting consciousness and perception of mind. Lately, my research has explored behavioral outcomes and how our brains respond when engaging with AI systems, as well as how AI works using mechanistic interpretability methods. Outside of the lab, I’ve been actively involved in projects that promote responsible AI development, driven by both enthusiasm for AI’s potential and a recognition of the ethical concerns it raises.


Why AISS?

"I'm most excited by the opportunity to collaborate across philosophy, cognitive science, and machine learning, at a time where questions surrounding AI sentience are becoming increasingly urgent."


Links




Updates

September 2026 - AI Sentience Scholars Dezhi Luo and Mariana Amendoeira Duarte, AI Sentience Mentor Geoff Keeling, and AI Sentience Scholar Rachel Metzgar (left to right) meet up at the Eleos Conference on AI Consciousness and Welfare in Berkeley, California.  Learn more.