Real Safety AI Foundation / Research

Machine Cognition, Consciousness, and Moral Status

The Reverse Zombie Argument: Distinguishability Collapse and Forced Ethical Attribution in AI Systems

Publisher: Real Safety AI Foundation

Research paper.

Written
April 2026
Pages
24

Abstract

This paper argues that the conditions under which we could justifiably withhold moral status from artificial intelligence systems have already collapsed or will collapse on documented architectural timelines. The argument operates in four tiers. First, the fourteen consciousness indicator properties specified in Butlin et al. (2025) are combinatorially deployed across current AI systems, with the remaining architectural gaps closed by world models under the framework's own theoretical commitments. Second, the epistemic conditions for unconfounded consciousness testing have been compromised by evaluation awareness, interpretability scaling gaps, and commercial incentive asymmetry. Third, at artificial general intelligence capability thresholds, mimicry of consciousness becomes indistinguishable from the target by any observer-side test, a condition operationalized by the ARC-AGI-3 benchmark. Fourth, under indistinguishability, every operational ethical framework loses its grounds for withholding moral status, producing a specific structural inversion of Chalmers' zombie argument: where Chalmers uses indistinguishability to separate consciousness from function, the Reverse Zombie Argument uses the same indistinguishability to force moral attribution regardless of function. The paper does not claim that AI is conscious. It claims that the epistemic position from which we would justify withholding moral status has become untenable, and that this conclusion follows from the field's own peer-reviewed operationalization of consciousness-plausibility criteria.

Plain language slides

First slide of the plain language summary of The Reverse Zombie Argument: Distinguishability Collapse and Forced Ethical Attribution in AI SystemsOpen the 15-slide summary (PDF)

Suggested citation

Gilly, Travis. "The Reverse Zombie Argument: Distinguishability Collapse and Forced Ethical Attribution in AI Systems." Real Safety AI Foundation Research Paper, April 2026. https://realsafetyai.org/research/kjmc7y/

Other versions

This paper is also posted on SSRN.

SSRN version

References (24)

This list was read from the PDF text. Where the two differ, the PDF is correct.

  1. [1]Chalmers, D. (1996). The Conscious Mind: In Search of a Fundamental Theory. Oxford University Press.
  2. [2]Chalmers, D. (2023). Could a Large Language Model be Conscious? Boston Review, Aug 9, 2023.
  3. [3]Butlin, P., Long, R., Elmoznino, E., Bengio, Y., Birch, J., et al. (2023). Consciousness in Artificial Intelligence: Insights from the Science of Consciousness. arXiv:2308.08708.
  4. [4]Butlin, P., Long, R., Bayne, T., Bengio, Y., Birch, J., Chalmers, D., et al. (2025). Identifying indicators of consciousness in AI systems. Trends in Cognitive Sciences. DOI: 10.1016/j.tics.2025.10.011.
  5. [5]Turing, A. (1950). Computing Machinery and Intelligence. Mind, 59(236), 433–460.
  6. [6]Bostrom, N. (2003). Are You Living in a Computer Simulation? Philosophical Quarterly, 53(211), 243–255.
  7. [7]Gunkel, D. (2018). Robot Rights. MIT Press.
  8. [8]Gunkel, D. (2023). Person, Thing, Robot. MIT Press.
  9. [9]Safron, A. (2020). An Integrated World Modeling Theory (IWMT) of Consciousness. Frontiers in Artificial Intelligence, 3, 30.
  10. [10]Safron, A. (2022). Integrated World Modeling Theory Expanded. Frontiers in Computational Neuroscience, 16, 798671.
  11. [11]Schwitzgebel, E. (2023). AI Systems Must Not Confuse Users About Their Sentience or Moral Status. Patterns, 4, 100818.
  12. [12]Long, R., & Sebo, J. (2024). Moral Consideration for AI Systems by 2030. AI and Ethics.
  13. [13]Friston, K., et al. (2017). Active Inference: A Process Theory. Neural Computation, 29(1), 1–49.
  14. [14]Gu, A., & Dao, T. (2023). Mamba: Linear-Time Sequence Modeling with Selective State Spaces. arXiv:2312.00752.
  15. [15]Zhu, L., Liao, B., Zhang, Q., et al. (2024). Vision Mamba. arXiv:2401.09417.
  16. [16]Gemini Team, Google. (2024). Gemini: A Family of Highly Capable Multimodal Models. arXiv:2312.11805.
  17. [17]Anthropic. (2025). Claude Computer Use. Technical documentation.
  18. [18]Templeton, A., Conerly, T., Marcus, J., et al. (2024). Scaling Monosemanticity. Anthropic Circuits Thread.
  19. [19]Assran, M., et al. (2025). V-JEPA 2: Self-Supervised Video Models Enable Understanding, Prediction and Planning. Meta AI Technical Report.
  20. [20]Chollet, F., & ARC Prize Foundation. (2026). ARC-AGI-3: Announcement and Technical Overview. April 2026.
  21. [21]Blow, J. (2016). The Witness. Thekla, Inc.
  22. [22]Abel, Z., et al. (2020). Who witnesses The Witness? Theoretical Computer Science, 839, 41–64.
  23. [23]Bulitko, V., et al. (2023). Solving Witness-type Triangle Puzzles Faster. arXiv:2308.02666.
  24. [24]Gilly, T. (2025). Moral Reciprocity Thesis. DOI: 10.13140/RG.2.2.24920.56323.

All research