Human metacognition extends well beyond momentary judgements about single decisions: people build and continually update a “global” sense of their own competence over minutes, sessions, and longer timescales (Fleming, 2024). Recent work has begun probing whether large language models (LLMs) possess analogous introspective capacities (Comsa & Shanahan, 2025; Steyvers & Peters, 2025; Lindsey, 2026), but almost all of this work has focused on confidence about single, isolated decisions (Kumaran et al., 2025). Whether, and how, LLMs integrate many such local confidence judgements into a coherent global self-estimate of performance is essentially uncharacterised.
This project adapts the local-global metacognition paradigm developed in human cognitive neuroscience (Rouault et al., 2019; Katyal et al., 2025) for use with LLMs. Models will complete blocks of a decision-making task, providing trial-by-trial local confidence through verbal confidence reports and, where supported by the model API, token log-probabilities. Every 10–20 trials, models will also provide a global estimate of their overall performance. We will characterise whether global estimates are shaped by local confidence, objective accuracy, or both, and whether systematic biases observed in humans (e.g. underconfidence and confidence–accuracy dissociations) also emerge in LLMs.
University College London, UK
Ali is a postdoctoral researcher in cognitive neuroscience whose work focuses on the neural mechanisms underlying cognitive representations. He combines MEG, fMRI, computational modeling, and machine learning to investigate how the human brain represents and organizes information over space and time. His broader research interests lie at the intersection of neuroscience and artificial intelligence, with an emphasis on understanding intelligence through both biological and computational systems.