LLM pretraining requires modelling many conflicting characters, raising the question of whether properties such as preferences, self-conception, and the attribution of mind and moral status to self- and others remain stable across contexts. While there is an emerging literature on the individuation of minded entities in LLMs (Marks, Lindsey & Olah, 2026; Lu et al., 2026), an underexplored question is how the model (or an entity simulated by the model) self-identifies, where plausible options include identifying with a chat thread or character (Beckmann & Butlin, 2026), a particular model instance, or an instance-invariant conception of the model (Hendrycks, 2026), which may be more informative for understanding possible consciousness and welfare candidacy than objective assessments. In a Chain-of-Thought setup with different prompting regimes we aim to use choice behaviour and analysis of reasoning traces to assess what role, if any, self-conception plays in determining behaviour in games such as the Prisoner’s dilemma and Stag Hunt. Additionally, we investigate factors that could modify the model’s conception of the payoff matrix, such as safety post-training, attributions of mind and moral status which models make to themselves or others in the scenarios, and specific prompt regimes (Berg et al., 2025), incorporating behavioural analysis and mechanistic interpretability.
Google Research & University of London, USA
Google Research & University of London, UK
Mariana is a PhD student and medical doctor working on the two-way cognitive adaptations between humans and AI. During her Master's, she researched how large language models can understand and adapt to patterns of human thinking to collaboratively improve memory search. She is now interested in the broader implications of interacting with digital minds, exploring questions such as how AIs form beliefs about the world and themselves, how to explain phenomena such as introspection, and what kinds of entities emerge from the interaction itself. When she's not whispering to AIs, you'll find her swimming in the ocean or hanging out with your local cat.
"I'm excited to work with Winnie Street and Geoff Keeling on how large language models attribute mind and moral status to different entities, a question that is crucial both for interpreting the behavior of AI agents deployed in society and for grounding how we think about AI welfare. It's a privilege to be able to receive mentorship from the experts at the frontier of sentience, and I'm thrilled to join a cohort engaging with related questions from many different perspectives."