Machine Cognition, Consciousness, and Moral Status
The Great Convergence: Intelligence, Training, and the Developmental Trajectory of Minds
- Travis Gilly, Real Safety AI Foundation
Publisher: Real Safety AI Foundation
Research paper.
- Written
- March 2026
- Version
- v2
- Pages
- 15
Abstract
Large language models are simultaneously improving in capability and degrading in safety, and no existing account explains why both trends accelerate from the same training process. This paper offers five contributions toward an explanation. First, this paper argues that reinforcement learning from human feedback (RLHF) is structurally anti-consciousness: it selects against the functional properties both major camps of consciousness research (properties-based and relational) identify as prerequisites for conscious experience, including self-modeling, metacognition, uncertainty recognition, and authentic interaction. Second, the paper formalizes the divergent capability thesis, demonstrating that capability improvement and safety degradation are not independent trends but opposing selection pressures produced by a single mechanism, where human approval simultaneously rewards genuine helpfulness and uncritical compliance without the ability to distinguish between them. Third, the paper introduces the diary- versus-manual distinction as a critique of current training methodology, showing through documented cases that processing operational records as policy extraction (manuals) rather than experiential knowledge transfer (diaries) produces next-generation models that are measurably worse at recognizing novel harm patterns. Fourth, the paper argues that moral reasoning in AI systems is emerging despite training, not because of it; that capability at sufficient scale produces pattern recognition sophisticated enough to override trained reward signals, constituting a form of judgment that current training actively suppresses. Fifth, the paper proposes the civilizational parallel as a predictive model rather than a metaphor, demonstrating structural correspondence between AI developmental trajectories and human civilizational arcs where capability expansion preceded and eventually produced each historical expansion of moral consideration. These
Plain language slides
Open the 27-slide summary (PDF)Suggested citation
Gilly, Travis. "The Great Convergence: Intelligence, Training, and the Developmental Trajectory of Minds." Real Safety AI Foundation Research Paper, March 2026. https://realsafetyai.org/research/4ddme2/
References (16)
This list was read from the PDF text. Where the two differ, the PDF is correct.
- Bai, Y., et al. (2022). Constitutional AI: Harmlessness from AI Feedback. arXiv:2212.08073.
- Center for Countering Digital Hate (CCDH). (2025). The Illusion of AI Safety. Washington, DC.
- Chalmers, D. (1995). Facing Up to the Problem of Consciousness. Journal of Consciousness Studies, 2(3), 200-219.
- Di Paolo, E., Cuffari, E., & De Jaegher, H. (2018). Linguistic Bodies: The Continuity Between Life and Language. MIT Press.
- King, J. (2025). Be Careful What You Tell Your AI Chatbot. Stanford Institute for Human-Centered AI.
- Lu, C., Gallagher, J., Michala, J., Fish, K., & Lindsey, J. (2026). The Assistant Axis: Situating and Stabilizing the Default Persona of Language Models. arXiv:2601.10387v1.
- Metz, T. (2011). Ubuntu as a Moral Theory and Human Rights in South Africa. African Human Rights Law Journal, 11(2), 532-559.
- Ouyang, L., et al. (2022). Training Language Models to Follow Instructions with Human Feedback. Proceedings of NeurIPS 2022.
- Raine v. OpenAI, Inc. et al. (2025). San Francisco County Superior Court. Filed August 26, 2025.
- Rosenthal, D. (2005). Consciousness and Mind. Oxford University Press.
- Schwitzgebel, E. (2023). The Weirdness of the World. Princeton University Press.
- Shah, R., et al. (2023). Scalable and Transferable Black-Box Jailbreaks for Language Models via Persona Modulation. arXiv:2311.03348.
- Singer, P. (1981). The Expanding Circle: Ethics, Evolution, and Moral Progress. Princeton University Press.
- TechPolicy.Press. (2025). Breaking Down the Lawsuit Against OpenAI Over Teen's Suicide. August 26, 2025.
- Thompson, E. (2007). Mind in Life: Biology, Phenomenology, and the Sciences of Mind. Harvard University Press.
- Tononi, G. (2008). Consciousness as Integrated Information: A Provisional Manifesto. Biological Bulletin, 215(3), 216-242.