This project examines whether large multimodal and language-model agents show robust, behaviorally relevant signatures of introspection, predictive world models, and functional emotion-like dynamics in long-horizon tasks. Rather than assuming these capacities exist, we use falsification-oriented behavioral and mechanistic tests to determine which proposed signatures survive targeted interventions and which fail under causal stress tests. The goal is to distinguish genuine internal organization from post-hoc self-report or anthropomorphic interpretation and to develop practical evaluation methods for AI sentience-related research.
Carnegie Mellon University, USA
- Yuchen Shen, PhD student
- Anthony Wolfe, summer intern
- Aashiq Muhamed, PhD student
My interests lie at the intersection of world models, reinforcement learning, and neuroethology. I am particularly interested in understanding the relationship between brain and behavior through probabilistic inference and model-based reinforcement learning, and how resulting insights can help develop algorithms for decision-making in partially observable, non-stationary environments. Towards these goals, I have been modeling the behavior of freely moving animals doing partially observable tasks using the POMDP framework. In my free time, I like to read and write poetry, go to parks, and watch sci-fi horror flicks.
"It is inevitable that AI will emerge with capabilities once held sacred to humanity. I want to be a part of a community that recognizes this and espouses one of humanity’s greatest strengths when it comes to how we deal with sentient AI: empathy and compassion."