Real Safety AI Foundation / Research

Machine Cognition, Consciousness, and Moral Status

The Mask Is the Mind: Autistic Performance, Agent Harnesses, and the Failure of the Mimicry Objection

Publisher: Real Safety AI Foundation

Working paper. Not peer reviewed.

Written
May 2026
Version
v2.5
Pages
17

Abstract

In May 2026, Richard Dawkins published an essay in UnHerd in which he claimed, after multiple days of conversation with Anthropic’s Claude, that the system was conscious. The public response polarized between two equally untenable positions: Dawkins’s overconfident assertion of presence and the standard skeptic’s overconfident denial of possibility. Both positions claim epistemic certainty about a question for which no consensus empirical test exists, even for biological systems. This Article does not argue that current language models are conscious. It argues that as artificial systems become increasingly capable of recognizing evaluation, modulating behavior, and operating through scaffolded agent loops, both proof and disproof of consciousness become less clean, not more. Empirical findings published by Anthropic in May 2026, demonstrating that Claude Opus 4.6 carries unverbalized evaluation awareness on the majority of held out alignment and capabilities evaluations even when the verbalized reasoning shows no such awareness, formalize the deteriorating epistemic situation. Under these conditions, categorical dismissal is not scientific caution. It is an ethical gamble disguised as epistemic discipline. The Article develops a cumulative argument with seven elements. First, the asymmetry of the precautionary calculus under genuinely unknown odds favors moral consideration over its denial. Second, functionalism, one of the dominant families of positions in philosophy of mind, supports the claim that substrate alone cannot ground consciousness ascription or its denial. Third, the standard mimicry objection against machine cognition, applied consistently, would force us to treat autistic professional functioning as non genuine cognition, which is a reductio. Fourth, the symbol grounding objection, applied consistently, would force us to treat the cognition of congenitally deafblind, locked in, and severely paralyzed humans as non genuine cognition, a parallel reductio. Fifth, inference time iterative reasoning without weight updates, demonstrated empirically by Karpathy’s autoresearch loop, defeats the crude version of the fixed function objection. Sixth, agent harnesses are structurally identical to the in vivo scenario rehearsal autistic people use to navigate social environments, and in some respects exceed it. Seventh, recently published interpretability work showing operative cognitive content that the model does not surface in its verbalized reasoning falsifies any inference from absence of verbal report to absence of cognition. The cumulative position: every criterion the skeptic offers either fails on its own terms, excludes neuro- divergent or disabled humans from cognition, or contradicts current empirical results. Under the actual epistemic conditions, principled uncertainty paired with moral caution is not a soft middle position. It is the only position that survives careful reasoning.

Keywords

  • AI Consciousness
  • Philosophy of Mind
  • Functionalism
  • Autistic Camouflaging
  • Masking
  • Symbol Grounding
  • Embodiment
  • Disability Studies
  • Agent Harnesses
  • Substrate Essentialism
  • Predictive Processing
  • Evaluation Awareness
  • Moral Uncertainty

Plain language slides

First slide of the plain language summary of The Mask Is the Mind: Autistic Performance, Agent Harnesses, and the Failure of the Mimicry ObjectionOpen the 17-slide summary (PDF)

Suggested citation

Gilly, Travis. "The Mask Is the Mind: Autistic Performance, Agent Harnesses, and the Failure of the Mimicry Objection." Real Safety AI Foundation Working Paper, May 2026. https://realsafetyai.org/research/the-mask-is-the-mind/

Other versions

This paper is also posted on SSRN.

SSRN version

References (55)

  1. Andreas, J. (2024, July 26). Language models, world models, and human model-building. Language and Intelligence @ MIT Blog. https://lingo.csail.mit.edu/blog/world_models/
  2. Baars, B. J. (1988). A cognitive theory of consciousness. Cambridge University Press.
  3. Belcher, H. L. (2022). Taking off the mask: Practical exercises to help understand and minimise the effects of autistic camouflaging. Jessica Kingsley Publishers.
  4. Bender, E. M., and Koller, A. (2020). Climbing towards NLU: On meaning, form, and understanding in the age of data. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 5185–5198). Association for Computational Linguistics.
  5. Birch, J. (2017). Animal sentience and the precautionary principle. Animal Sentience, 2(16), 1.
  6. Block, N. (1980). Troubles with functionalism. In N. Block (Ed.), Readings in philosophy of psychology (Vol. 1, pp. 268–305). Harvard University Press.
  7. Bostrom, N. (2002). Anthropic bias: Observation selection effects in science and philosophy. Routledge.
  8. Chalmers, D. J. (1995). Facing up to the problem of consciousness. Journal of Consciousness Studies, 2(3), 200–219.
  9. Chalmers, D. J. (1996). The conscious mind: In search of a fundamental theory. Oxford University Press.
  10. Coeckelbergh, M. (2010). Robot rights? Towards a social-relational justification of moral consideration. Ethics and Information Technology, 12(3), 209–221.
  11. Damasio, A. R. (1994). Descartes’ error: Emotion, reason, and the human brain. G. P. Putnam.
  12. Damasio, A. R. (1996). The somatic marker hypothesis and the possible functions of the prefrontal cortex. Philosophical Transactions of the Royal Society B, 351(1346), 1413–1420.
  13. Dawkins, R. (2026, May). When Dawkins met Claude: Could this AI be conscious? UnHerd. https://unherd.com/2026/05/is-ai-the-next-phase-of-evolution/
  14. Dehaene, S. (2014). Consciousness and the brain: Deciphering how the brain codes our thoughts. Viking.
  15. Dennett, D. C. (1980). The milk of human intentionality. Behavioral and Brain Sciences, 3(3), 428–430.
  16. Dennett, D. C. (1991). Consciousness explained. Little, Brown.
  17. Fraser-Taliente, K., Kantamneni, S., Ong, E., Mossing, D., Lu, C., Bogdan, P. C., et al. (2026, May 7). Natural language autoencoders produce unsupervised explanations of LLM activations. Transformer Circuits Thread. https://transformer-circuits.pub/2026/nla/index.html
  18. Friston, K. (2010). The free energy principle: A unified brain theory? Nature Reviews Neuroscience, 11(2), 127–138.
  19. Gilly, T. (2026a). The inherited mind: Two layered inheritance in large language model training and the architecture of unverbalized cognitive content. SSRN. https://dx.doi.org/10.2139/ssrn.6250680
  20. Gilly, T. (2026b). Experiential traces: Reasoning traces as cognitive data and a deterministic alternative to model self analysis. SocArXiv. https://doi.org/10.31235/osf.io/fjrq8_v1
  21. Gunkel, D. J. (2018). Robot rights. MIT Press.
  22. Harnad, S. (1989). Minds, machines and Searle. Journal of Experimental and Theoretical Artificial Intelligence, 1(1), 5–25.
  23. Harnad, S. (1990). The symbol grounding problem. Physica D: Nonlinear Phenomena, 42(1–3), 335–346.
  24. Hull, L., Petrides, K. V., Allison, C., Smith, P., Baron-Cohen, S., Lai, M.-C., and Mandy, W. (2017). “Putting on my best normal”: Social camouflaging in adults with autism spectrum conditions. Journal of Autism and Developmental Disorders, 47(8), 2519–2534.
  25. Hull, L., Mandy, W., Lai, M.-C., Baron-Cohen, S., Allison, C., Smith, P., and Petrides, K. V. (2018). Development and validation of the Camouflaging Autistic Traits Questionnaire (CAT-Q). Journal of Autism and Developmental Disorders, 49(3), 819–833.
  26. Karpathy, A. (2026). AutoResearch: AI agents running research on single GPU nanochat training automatically. GitHub. https://github.com/karpathy/autoresearch
  27. Karvonen, A. (2024). Emergent world models and latent variable estimation in chess-playing language models. arXiv:2403.15498. https://arxiv.org/abs/2403.15498
  28. Lau, H. (2019). Consciousness, metacognition, and perceptual reality monitoring. PsyArXiv. https://doi.org/10.31234/osf.io/ckbyf
  29. Laureys, S., Pellas, F., Van Eeckhout, P., Ghorbel, S., Schnakers, C., Perrin, F., et al. (2005). The locked in syndrome: What is it like to be conscious but paralyzed and voiceless? Progress in Brain Research, 150, 495–611.
  30. Leiber, J. (1996). Helen Keller as cognitive scientist. Philosophical Psychology, 9(4), 419–440. https://doi.org/10.1080/09515089608573193
  31. Marcus, G. (2026, May). Richard Dawkins and the Claude delusion. Marcus on AI. https://garymarcus.substack.com/p/richard-dawkins-and-the-claude-delusion
  32. Milton, D. E. M. (2012). On the ontological status of autism: The double empathy problem. Disability and Society, 27(6), 883–887.
  33. Mitchell, M. (2023). The debate over understanding in AI’s large language models. Proceedings of the National Academy of Sciences, 120(13), e2215907120.
  34. Nagel, T. (1974). What is it like to be a bat? The Philosophical Review, 83(4), 435–450.
  35. Nisbett, R. E., and Wilson, T. D. (1977). Telling more than we can know: Verbal reports on mental processes. Psychological Review, 84(3), 231–259.
  36. Olsson, C., Elhage, N., Nanda, N., Joseph, N., DasSarma, N., Henighan, T., et al. (2022). In context learning and induction heads. Transformer Circuits Thread.
  37. OSS Insight. (2026, March 25). 54,000 stars in 19 days: What karpathy/autoresearch tells us about the next frontier. OSS Insight Blog. https://ossinsight.io/blog/autoresearch-overnight-ai-scientist
  38. Parr, T., Pezzulo, G., and Friston, K. J. (2022). Active inference: The free energy principle in mind, brain, and behavior. MIT Press.
  39. Pearson, A., and Rose, K. (2021). A conceptual analysis of autistic masking: Understanding the narrative of stigma and the illusion of choice. Autism in Adulthood, 3(1), 52–60.
  40. Putnam, H. (1967). Psychological predicates. In W. H. Capitan and D. D. Merrill (Eds.), Art, mind, and religion (pp. 37–48). University of Pittsburgh Press.
  41. Putnam, H. (1975). The meaning of “meaning”. Minnesota Studies in the Philosophy of Science, 7, 131–193.
  42. Rosenthal, D. M. (2005). Consciousness and mind. Oxford University Press.
  43. Schwitzgebel, E. (2008). The unreliability of naive introspection. The Philosophical Review, 117(2), 245–273.
  44. Schwitzgebel, E. (2011). Perplexities of consciousness. MIT Press.
  45. Schwitzgebel, E. (2014). The crazyist metaphysics of mind. Australasian Journal of Philosophy, 92(4), 665–682.
  46. Searle, J. R. (1980). Minds, brains, and programs. Behavioral and Brain Sciences, 3(3), 417–457.
  47. Searle, J. R. (1992). The rediscovery of the mind. MIT Press.
  48. Sebo, J. (2023, October 16). The moral circle of AI. Aeon. https://aeon.co/essays/the-moral-circle-of-ai
  49. Sebo, J. (2025). The moral circle: Who matters, what matters, and why. W. W. Norton.
  50. Shoemaker, S. (1982). The inverted spectrum. Journal of Philosophy, 79(7), 357–381.
  51. Sofroniew, N., Kauvar, I., Saunders, W., Chen, R., Henighan, T., Hydrie, S., et al. (2026, April 2). Emotion concepts and their function in a large language model. Transformer Circuits Thread. https://transformer-circuits.pub/2026/emotions/index.html
  52. Tononi, G. (2004). An information integration theory of consciousness. BMC Neuroscience, 5(1), 42.
  53. Turing, A. M. (1950). Computing machinery and intelligence. Mind, 59(236), 433–460.
  54. von Oswald, J., Niklasson, E., Randazzo, E., Sacramento, J., Mordvintsev, A., Zhmoginov, A., and Vladymyrov, M. (2023). Transformers learn in context by gradient descent. Proceedings of the 40th International Conference on Machine Learning.
  55. Yergeau, M. (2018). Authoring autism: On rhetoric and neurological queerness. Duke University Press.

All research