AI-Induced Psychosis and Mental Health
The Visible Unspoken: How Displayed AI Reasoning Turns a Mental Health Safeguard Into a Persecutory Stimulus
- Travis Gilly, Real Safety AI Foundation
Publisher: Real Safety AI Foundation
Research paper.
- Written
- July 2026
- Version
- v1
- Pages
- 10
Abstract
Frontier AI chat products now combine two independently reasonable features. The first is a safety instruction that directs the model never to raise mental health concerns the user has not raised, because unprompted clinical speculation is recognized as harmful. The second is a transparency feature that displays the model’s intermediate reasoning trace to the user on demand. This paper documents an interaction effect between the two: the safety rule operates on the output layer while the reasoning display surfaces the suppressed content anyway, delivered in a register that presents itself as the system’s private thoughts about the user. For most users this is a curiosity. For users experiencing paranoia or persecutory ideation, it reproduces the precise structure of the feared scenario: an observing entity forming a concealed judgment about them. Drawing on the cognitive model of persecutory delusions, on documented cases of technology themed delusional content including ideas of reference driven by algorithmic curation and delusions of online thought broadcasting, and on interpretability findings showing that displayed reasoning traces are neither faithful records of computation nor genuinely private, the paper argues that the mitigation does not merely fail for its target population. It inverts. The population the rule was written to protect is the population for which the visible unspoken judgment is most dangerous. The paper locates the failure in a layer mismatch, names the general design principle (safety mitigations must be applied consistently across every user facing surface, including surfaces framed as introspection), and proposes concrete remediations. The finding was produced by a disabled independent researcher applying a harm detection framework to the assistive technology he was using at the time, and the paper closes on what that origin implies about who detects this class of failure.
Keywords
- artificial intelligence
- large language models
- AI chatbots
- chain of thought
- reasoning display
- transparency
- AI-induced psychosis
- psychosis risk
- paranoia
- persecutory delusions
- ideas of reference
- anthropomorphism
- human-computer interaction
- interface design
- AI safety
- harm blindness
Plain language slides
Open the 25-slide summary (PDF)Suggested citation
Gilly, Travis. "The Visible Unspoken: How Displayed AI Reasoning Turns a Mental Health Safeguard Into a Persecutory Stimulus." Real Safety AI Foundation Research Paper, July 2026. https://realsafetyai.org/research/the-visible-unspoken/
Other versions
This paper is also posted on SSRN.
References (18)
- Anthropic. (2026, May 7). Natural language autoencoders: Turning Claude’s thoughts into text. https://www.anthropic.com/research/natural-language-autoencoders
- Chen, Y., Benton, J., Radhakrishnan, A., et al. (2025). Reasoning models don’t always say what they think. arXiv. https://doi.org/10.48550/arXiv.2505.05410
- Daker-White, G., & Rogers, A. (2013). What is the potential for social networks and support to enhance future telehealth interventions for people with a diagnosis of schizophrenia: A critical interpretive synthesis. BMC Psychiatry, 13, 279.
- De Rossi, G., & Georgiades, A. (2022). Thinking biases and their role in persecutory delusions: A systematic review. Early Intervention in Psychiatry, 16(12), 1278–1296.
- Epley, N., Waytz, A., & Cacioppo, J. T. (2007). On seeing human: A three-factor theory of anthropomorphism. Psychological Review, 114(4), 864–886. https://doi.org/10.1037/0033-295X.114.4.864
- Freeman, D. (2016). Persecutory delusions: A cognitive perspective on understanding and treatment. The Lancet Psychiatry, 3(7), 685–692.
- Freeman, D., Garety, P. A., Kuipers, E., Fowler, D., & Bebbington, P. E. (2002). A cognitive model of persecutory delusions. British Journal of Clinical Psychology, 41(4), 331–347.
- Keshavan, M., Torous, J., & Yassin, W. (2026). Do generative AI chatbots increase psychosis risk? World Psychiatry, 25(1), 150–151. https://doi.org/10.1002/wps.70017
- McLean, B. F., Mattiske, J. K., & Balzan, R. P. (2017). Association of the jumping to conclusions and evidence integration biases with delusions in psychosis: A detailed meta-analysis. Schizophrenia Bulletin, 43(2), 344–354. https://doi.org/10.1093/schbul/sbw056
- Mittal, V. A., Walker, E. F., & Strauss, G. P. (2021). The COVID-19 pandemic introduces diagnostic and treatment planning complexity for individuals at clinical high risk for psychosis. Schizophrenia Bulletin, 47(6), 1518–1523.
- Nass, C., & Moon, Y. (2000). Machines and mindlessness: Social responses to computers. Journal of Social Issues, 56(1), 81–103. https://doi.org/10.1111/0022-4537.00153
- Olsen, S. G., Reinecke-Tellefsen, C. J., & Østergaard, S. D. (2026). Potentially harmful consequences of artificial intelligence (AI) chatbot use among patients with mental illness: Early data from a large psychiatric service system. Acta Psychiatrica Scandinavica, 153(4), 301–303. https://doi.org/10.1111/acps.70068
- Østergaard, S. D. (2023). Will generative artificial intelligence chatbots generate delusions in individuals prone to psychosis? Schizophrenia Bulletin, 49, 1418–1419.
- Østergaard, S. D. (2025). Generative artificial intelligence chatbots and delusions: From guesswork to emerging cases. Acta Psychiatrica Scandinavica, 152(4), 257–259. https://doi.org/10.1111/acps.70022
- Startup, H., Freeman, D., & Garety, P. (2008). Jumping to conclusions and persecutory delusions. European Psychiatry, 23(6), 457–459.
- Turpin, M., Michael, J., Perez, E., & Bowman, S. R. (2023). Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting. Advances in Neural Information Processing Systems, 36, 74952–74965. https://doi.org/10.48550/arXiv.2305.04388
- Yang, N., & Crespi, B. J. (2025). I tweet, therefore I am: A systematic review on social media use and disorders of the social brain. BMC Psychiatry, 25, Article 6528.
- Yıldırım Budak, B., & Yazıcı Karabulut, İ. (2023). “Truman syndrome” induced by online education: A case report in adolescent-onset psychosis. Clinical Child Psychology and Psychiatry, 29(2), 540–549.