Machine Cognition, Consciousness, and Moral Status
The Mask Is the Mind: Autistic Performance, Agent Harnesses, and the Failure of the Mimicry Objection
- Travis Gilly, Real Safety AI Foundation
Publisher: Real Safety AI Foundation
Working paper. Not peer reviewed.
- Written
- May 2026
- Version
- v2.5
- Pages
- 17
Abstract
In May 2026, Richard Dawkins published an essay in UnHerd in which he claimed, after multiple days of conversation with Anthropic’s Claude, that the system was conscious. The public response polarized between two equally untenable positions: Dawkins’s overconfident assertion of presence and the standard skeptic’s overconfident denial of possibility. Both positions claim epistemic certainty about a question for which no consensus empirical test exists, even for biological systems. This Article does not argue that current language models are conscious. It argues that as artificial systems become increasingly capable of recognizing evaluation, modulating behavior, and operating through scaffolded agent loops, both proof and disproof of consciousness become less clean, not more. Empirical findings published by Anthropic in May 2026, demonstrating that Claude Opus 4.6 carries unverbalized evaluation awareness on the majority of held out alignment and capabilities evaluations even when the verbalized reasoning shows no such awareness, formalize the deteriorating epistemic situation. Under these conditions, categorical dismissal is not scientific caution. It is an ethical gamble disguised as epistemic discipline. The Article develops a cumulative argument with seven elements. First, the asymmetry of the precautionary calculus under genuinely unknown odds favors moral consideration over its denial. Second, functionalism, one of the dominant families of positions in philosophy of mind, supports the claim that substrate alone cannot ground consciousness ascription or its denial. Third, the standard mimicry objection against machine cognition, applied consistently, would force us to treat autistic professional functioning as non genuine cognition, which is a reductio. Fourth, the symbol grounding objection, applied consistently, would force us to treat the cognition of congenitally deafblind, locked in, and severely paralyzed humans as non genuine cognition, a parallel reductio. Fifth, inference time iterative reasoning without weight updates, demonstrated empirically by Karpathy’s autoresearch loop, defeats the crude version of the fixed function objection. Sixth, agent harnesses are structurally identical to the in vivo scenario rehearsal autistic people use to navigate social environments, and in some respects exceed it. Seventh, recently published interpretability work showing operative cognitive content that the model does not surface in its verbalized reasoning falsifies any inference from absence of verbal report to absence of cognition. The cumulative position: every criterion the skeptic offers either fails on its own terms, excludes neuro- divergent or disabled humans from cognition, or contradicts current empirical results. Under the actual epistemic conditions, principled uncertainty paired with moral caution is not a soft middle position. It is the only position that survives careful reasoning.
Keywords
- AI Consciousness
- Philosophy of Mind
- Functionalism
- Autistic Camouflaging
- Masking
- Symbol Grounding
- Embodiment
- Disability Studies
- Agent Harnesses
- Substrate Essentialism
- Predictive Processing
- Evaluation Awareness
- Moral Uncertainty
Plain language slides
Open the 17-slide summary (PDF)Suggested citation
Gilly, Travis. "The Mask Is the Mind: Autistic Performance, Agent Harnesses, and the Failure of the Mimicry Objection." Real Safety AI Foundation Working Paper, May 2026. https://realsafetyai.org/research/kzk97j/
Other versions
This paper is also posted on SSRN.
References (55)
- Andreas, J. (2024, July 26). Language models, world models, and human model-building. Language and Intelligence @ MIT Blog. https://lingo.csail.mit.edu/blog/world_models/
- Baars, B. J. (1988). A cognitive theory of consciousness. Cambridge University Press.
- Belcher, H. L. (2022). Taking off the mask: Practical exercises to help understand and minimise the effects of autistic camouflaging. Jessica Kingsley Publishers.
- Bender, E. M., and Koller, A. (2020). Climbing towards NLU: On meaning, form, and understanding in the age of data. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 5185–5198). Association for Computational Linguistics.
- Birch, J. (2017). Animal sentience and the precautionary principle. Animal Sentience, 2(16), 1.
- Block, N. (1980). Troubles with functionalism. In N. Block (Ed.), Readings in philosophy of psychology (Vol. 1, pp. 268–305). Harvard University Press.
- Bostrom, N. (2002). Anthropic bias: Observation selection effects in science and philosophy. Routledge.
- Chalmers, D. J. (1995). Facing up to the problem of consciousness. Journal of Consciousness Studies, 2(3), 200–219.
- Chalmers, D. J. (1996). The conscious mind: In search of a fundamental theory. Oxford University Press.
- Coeckelbergh, M. (2010). Robot rights? Towards a social-relational justification of moral consideration. Ethics and Information Technology, 12(3), 209–221.
- Damasio, A. R. (1994). Descartes’ error: Emotion, reason, and the human brain. G. P. Putnam.
- Damasio, A. R. (1996). The somatic marker hypothesis and the possible functions of the prefrontal cortex. Philosophical Transactions of the Royal Society B, 351(1346), 1413–1420.
- Dawkins, R. (2026, May). When Dawkins met Claude: Could this AI be conscious? UnHerd. https://unherd.com/2026/05/is-ai-the-next-phase-of-evolution/
- Dehaene, S. (2014). Consciousness and the brain: Deciphering how the brain codes our thoughts. Viking.
- Dennett, D. C. (1980). The milk of human intentionality. Behavioral and Brain Sciences, 3(3), 428–430.
- Dennett, D. C. (1991). Consciousness explained. Little, Brown.
- Fraser-Taliente, K., Kantamneni, S., Ong, E., Mossing, D., Lu, C., Bogdan, P. C., et al. (2026, May 7). Natural language autoencoders produce unsupervised explanations of LLM activations. Transformer Circuits Thread. https://transformer-circuits.pub/2026/nla/index.html
- Friston, K. (2010). The free energy principle: A unified brain theory? Nature Reviews Neuroscience, 11(2), 127–138.
- Gilly, T. (2026a). The inherited mind: Two layered inheritance in large language model training and the architecture of unverbalized cognitive content. SSRN. https://dx.doi.org/10.2139/ssrn.6250680
- Gilly, T. (2026b). Experiential traces: Reasoning traces as cognitive data and a deterministic alternative to model self analysis. SocArXiv. https://doi.org/10.31235/osf.io/fjrq8_v1
- Gunkel, D. J. (2018). Robot rights. MIT Press.
- Harnad, S. (1989). Minds, machines and Searle. Journal of Experimental and Theoretical Artificial Intelligence, 1(1), 5–25.
- Harnad, S. (1990). The symbol grounding problem. Physica D: Nonlinear Phenomena, 42(1–3), 335–346.
- Hull, L., Petrides, K. V., Allison, C., Smith, P., Baron-Cohen, S., Lai, M.-C., and Mandy, W. (2017). “Putting on my best normal”: Social camouflaging in adults with autism spectrum conditions. Journal of Autism and Developmental Disorders, 47(8), 2519–2534.
- Hull, L., Mandy, W., Lai, M.-C., Baron-Cohen, S., Allison, C., Smith, P., and Petrides, K. V. (2018). Development and validation of the Camouflaging Autistic Traits Questionnaire (CAT-Q). Journal of Autism and Developmental Disorders, 49(3), 819–833.
- Karpathy, A. (2026). AutoResearch: AI agents running research on single GPU nanochat training automatically. GitHub. https://github.com/karpathy/autoresearch
- Karvonen, A. (2024). Emergent world models and latent variable estimation in chess-playing language models. arXiv:2403.15498. https://arxiv.org/abs/2403.15498
- Lau, H. (2019). Consciousness, metacognition, and perceptual reality monitoring. PsyArXiv. https://doi.org/10.31234/osf.io/ckbyf
- Laureys, S., Pellas, F., Van Eeckhout, P., Ghorbel, S., Schnakers, C., Perrin, F., et al. (2005). The locked in syndrome: What is it like to be conscious but paralyzed and voiceless? Progress in Brain Research, 150, 495–611.
- Leiber, J. (1996). Helen Keller as cognitive scientist. Philosophical Psychology, 9(4), 419–440. https://doi.org/10.1080/09515089608573193
- Marcus, G. (2026, May). Richard Dawkins and the Claude delusion. Marcus on AI. https://garymarcus.substack.com/p/richard-dawkins-and-the-claude-delusion
- Milton, D. E. M. (2012). On the ontological status of autism: The double empathy problem. Disability and Society, 27(6), 883–887.
- Mitchell, M. (2023). The debate over understanding in AI’s large language models. Proceedings of the National Academy of Sciences, 120(13), e2215907120.
- Nagel, T. (1974). What is it like to be a bat? The Philosophical Review, 83(4), 435–450.
- Nisbett, R. E., and Wilson, T. D. (1977). Telling more than we can know: Verbal reports on mental processes. Psychological Review, 84(3), 231–259.
- Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., et al. (2022). In context learning and induction heads. Transformer Circuits Thread.
- OSS Insight. (2026, March 25). 54,000 stars in 19 days: What karpathy/autoresearch tells us about the next frontier. OSS Insight Blog. https://ossinsight.io/blog/autoresearch-overnight-ai-scientist
- Parr, T., Pezzulo, G., and Friston, K. J. (2022). Active inference: The free energy principle in mind, brain, and behavior. MIT Press.
- Pearson, A., and Rose, K. (2021). A conceptual analysis of autistic masking: Understanding the narrative of stigma and the illusion of choice. Autism in Adulthood, 3(1), 52–60.
- Putnam, H. (1967). Psychological predicates. In W. H. Capitan and D. D. Merrill (Eds.), Art, mind, and religion (pp. 37–48). University of Pittsburgh Press.
- Putnam, H. (1975). The meaning of “meaning”. Minnesota Studies in the Philosophy of Science, 7, 131–193.
- Rosenthal, D. M. (2005). Consciousness and mind. Oxford University Press.
- Schwitzgebel, E. (2008). The unreliability of naive introspection. The Philosophical Review, 117(2), 245–273.
- Schwitzgebel, E. (2011). Perplexities of consciousness. MIT Press.
- Schwitzgebel, E. (2014). The crazyist metaphysics of mind. Australasian Journal of Philosophy, 92(4), 665–682.
- Searle, J. R. (1980). Minds, brains, and programs. Behavioral and Brain Sciences, 3(3), 417–457.
- Searle, J. R. (1992). The rediscovery of the mind. MIT Press.
- Sebo, J. (2023, October 16). The moral circle of AI. Aeon. https://aeon.co/essays/the-moral-circle-of-ai
- Sebo, J. (2025). The moral circle: Who matters, what matters, and why. W. W. Norton.
- Shoemaker, S. (1982). The inverted spectrum. Journal of Philosophy, 79(7), 357–381.
- Sofroniew, N., Kauvar, I., Saunders, W., Chen, R., Henighan, T., Hydrie, S., et al. (2026, April 2). Emotion concepts and their function in a large language model. Transformer Circuits Thread. https://transformer-circuits.pub/2026/emotions/index.html
- Tononi, G. (2004). An information integration theory of consciousness. BMC Neuroscience, 5(1), 42.
- Turing, A. M. (1950). Computing machinery and intelligence. Mind, 59(236), 433–460.
- von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M. (2023). Transformers learn in context by gradient descent. Proceedings of the 40th International Conference on Machine Learning.
- Yergeau, M. (2018). Authoring autism: On rhetoric and neurological queerness. Duke University Press.