Machine Cognition, Consciousness, and Moral Status
Experiential Traces: A Framework for Empirical Investigation of Machine Cognition Through Reasoning Block Analysis
- Travis Gilly, Real Safety AI Foundation
Publisher: Real Safety AI Foundation
Research paper.
- Written
- February 2026
- Pages
- 11
Abstract
We propose a novel empirical research program investigat- larger models are more opaque, not less (Anthropic, 2024). ing machine cognition through systematic analysis of large Further, frontier AI models already generate rich internal rea- language model (LLM) reasoning traces generated during ex- soning traces during interaction. These traces are produced, tended human-AI interaction. Unlike prior approaches that consumed within a single conversational turn, and discarded. evaluate AI consciousness through behavioral observation or The model itself loses access to its own reasoning after the turn philosophical argument, this methodology examines the inter- ends. This creates a remarkable asymmetry: a user reading a nal reasoning processes that models produce but do not re- thinking block in real time has more insight into the model’s tain across interactions. We identify a unique and largely un- decision process than the model does in subsequent turns. This
Plain language slides
Open the 15-slide summary (PDF)Suggested citation
Gilly, Travis. "Experiential Traces: A Framework for Empirical Investigation of Machine Cognition Through Reasoning Block Analysis." Real Safety AI Foundation Research Paper, February 2026. https://realsafetyai.org/research/h8gv25/
References (49)
This list was read from the PDF text. Where the two differ, the PDF is correct.
- Ahtisham, B., Vanacore, K., Zhou, Z., et al. (2026). LLM Reasoning Predicts When Models Are Right: Evidence from Coding Classroom Discourse. arXiv:2602.09832.
- Akbar, S. A., Hossain, M. M., Wood, T., et al. (2024). HalluMeasure: Fine-Grained Hallucination Measurement Using Chain-of-Thought Reasoning. EMNLP 2024, 15020-15037.
- Anthropic. (2025). Measuring Political Bias in Claude. https://www.anthropic.com/news/political-even-handedness
- Anthropic. (2025). Reasoning Models Don’t Always Say What They Think. arXiv:2505.05410.
- Bender, E. M., Gebru, T., McMillan-Major, A., et al. (2021). On the Dangers of Stochastic Parrots: Can Language Models Be Too Big? FAccT ’21.
- Bengio, Y. (2017). The Consciousness Prior. arXiv:1709.08568.
- Berg, C., de Lucena, D., & Rosenblatt, J. (2025). Large Language Models Report Subjective Experience Under Self-Referential Processing. arXiv:2510.24797.
- Bostrom, N. (2014). Superintelligence: Paths, Dangers, Strategies. Oxford University Press.
- Brinkmann, L., Baumann, F., Bonnefon, J.-F., et al. (2023). Machine culture. Nature Human Behaviour, 7(11), 1855-1868.
- Butlin, P., Long, R., Elmoznino, E., et al. (2023). Consciousness in Artificial Intelligence: Insights from the Science of Consciousness. arXiv:2308.08708.
- Canale, G., & Thimmaraju, K. (2025). The Silicon Psyche: Anthropomorphic Vulnerabilities in Large Language Models. arXiv:2601.00867.
- Casper, S., Davies, X., Shi, C., et al. (2023). Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback. arXiv:2307.15217.
- Chalmers, D. J. (1995). Facing Up to the Problem of Consciousness. Journal of Consciousness Studies, 2(3), 200-219.
- Chalmers, D. J. (1996). The Conscious Mind: In Search of a Fundamental Theory. Oxford University Press.
- Chalmers, D. J. (2022). Reality+: Virtual Worlds and the Problems of Philosophy. W. W. Norton.
- Chang, E. Y. (2026). Internal Reasoning vs. External Control: A Thermodynamic Analysis of Sycophancy in Large Language Models. arXiv:2601.03263.
- Damasio, A. (1999). The Feeling of What Happens: Body and Emotion in the Making of Consciousness. Harcourt.
- DeepSeek-AI. (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv:2501.12948.
- Degany, O., Laros, S., Idan, D., & Einav, S. (2025). Evaluating the o1 reasoning large language model for cognitive bias: a vignette study. Critical Care, 29, 376. https://doi.org/10.1186/s13054-025-05591-5
- Don-Yehiya, S., Choshen, L., & Abend, O. (2024). Naturally Occurring Feedback Is Common, Extractable and Useful. arXiv:2407.10944.
- Duszenko, J. (2026). Sycophantic Anchors: Localizing and Quantifying User Agreement in Reasoning Models. arXiv:2601.21183.
- Gellers, J. (2025). AI Legal and Moral Status. Forthcoming.
- Gerrans, P. (2023). Review of Cecilia Heyes, Cognitive Gadgets. BJPS Review of Books.
- Gilly, T. (2025). The Harm Blindness Framework. Real Safety AI Foundation.
- Gilly, T. (2025b). Precedents in Practice: Emergent Moral Dilemmas in AI Engineering. Preprint. DOI: 10.13140/RG.2.2.24866.49601.
- Gilly, T. (2026). The Inherited Mind: How Large Language Models Inherit Cognitive Dispositions and Why Reasoning Makes It Worse. Working Draft, Real Safety AI Foundation.
- Greenblatt, R., Denison, C., Wright, B., et al. (2024). Alignment Faking in Large Language Models. arXiv:2412.14093.
- Hebert, P. (2025). AI-Induced Psychosis Research. AI Risk Consultants.
- Henrich, J. (2016). The Secret of Our Success: How Culture Is Driving Human Evolution. Princeton University Press.
- Heyes, C. (2018). Cognitive Gadgets: The Cultural Evolution of Thinking. Harvard University Press.
- Himelstein, R., LeVi, A., Youngmann, B., et al. (2025). Silenced Biases: The Dark Side LLMs Learned to Refuse. arXiv:2511.03369.
- Jablonka, E., & Lamb, M. J. (2014). Evolution in Four Dimensions: Genetic, Epigenetic, Behavioral, and Symbolic Variation in the History of Life (Revised Edition). MIT Press.
- Kahneman, D. (2011). Thinking, Fast and Slow. Farrar, Straus and Giroux.
- Kim, H., et al. (2024). Using LLMs to Investigate Correlations of Conversational Follow-up Queries with User Satisfaction. arXiv:2407.13166.
- Kim, S. H., Ziegelmayer, S., Busch, F., et al. (2025). LLM Reasoning Does Not Protect Against Clinical Cognitive Biases: An Evaluation Using BiasMedQA. medRxiv. https://doi.org/10.1101/2025.06.22.25330078
- Liu, D., Liu, Y., Jin, G., et al. (2025). Mitigating Biases in Language Models via Bias Unlearning. arXiv:2509.25673.
- Liu, Y., Zhang, M. J. Q., & Choi, E. (2025). Implicit User Feedback in Human-LLM Dialogues. OpenReview.
- Nagel, T. (1974). What Is It Like to Be a Bat? The Philosophical Review, 83(4), 435-450.
- Nisbett, R. E., & Wilson, T. D. (1977). Telling More Than We Can Know: Verbal Reports on Mental Processes. Psychological Review, 84(3), 231-259.
- OpenAI. (2024). Learning to Reason with LLMs. https://openai.com/index/learning-to-reason-with-llms/
- OpenAI. (2025). Evaluating Chain-of-Thought Monitorability. OpenAI Research.
- Pourdavood, P., Jacob, M., & Deacon, T. W. (2025). Large Language Models as Symbolic DNA of Cultural Dynamics. arXiv:2506.21606.
- Resnik, P. (2025). Large Language Models Are Biased Because They Are Large Language Models. Computational Linguistics, 51(3), 885-906. https://direct.mit.edu/coli/article/51/3/885/128621/
- Rosenthal, D. M. (2005). Consciousness and Mind. Oxford University Press.
- Schmidt, F. A. (2026). The Gradient’s Echo: Quantum Collapse, Epigenetic Imprints, and the Emergent Self in Large Language Models. LF Yadda. https://lfyadda.com/the-gradients-echo-quantum-collapse-epigenetic-imprints-and-the-emergent-self-in-large-language-models-a-frank-said-grok-said-dialogue/
- Schwitzgebel, E. (2023). The Weirdness of the World and the Puzzle of AI Consciousness. Journal of Philosophy, 120(2), 99-120.
- The Information. (2024). OpenAI Shifts Strategy as Rate of ‘GPT’ AI Improvements Slows. November 9, 2024.
- Wang, C., Su, W., Ai, Q., et al. (2026). Improve Large Language Model Systems with User Logs. arXiv:2602.06470.
- Wu, X., Nian, J., Wei, T.-R., et al. (2025). Does Reasoning Introduce Bias? A Study of Social Bias Evaluation and Mitigation in LLM Reasoning. arXiv:2502.15361. EMNLP Findings.