Machine Cognition, Consciousness, and Moral Status
The Inherited Mind: How Large Language Models Inherit Cognitive Dispositions and Why Reasoning Makes It Worse
- Travis Gilly, Real Safety AI Foundation
Publisher: Real Safety AI Foundation
Research paper.
- Written
- February 2026
- Pages
- 7
Abstract
We trained machines to reason but forgot to teach them what to chitecture. The industry consensus was clear: reasoning would reason about. This paper proposes that large language model produce not only more accurate but more reliable, less biased (LLM) training creates dispositional cognitive inheritance anal- outputs. If a model could think before it spoke, surely it would ogous to epigenetic transmission in biological systems: patterns think better. encoded in model weights that operate below introspective ac- The evidence says otherwise. Comparative evaluations across cess, influence behavior without being explicitly represented as multiple research groups have found that reasoning-augmented retrievable knowledge, and persist through debiasing attempts. models show equal or greater susceptibility to cognitive and so- We synthesize evidence from three independent lines of re- cial biases than their instruct-tuned baselines (Wu et al., 2025; search to support this claim. First, mechanistic interpretability Degany et al., 2025; Kim et al., 2025). OpenAI’s o1 demon- studies demonstrate that biases are encoded as distributional strates greater bias than GPT-4o on social bias benchmarks. tendencies in weight space rather than discrete factual associa- DeepSeek-R1 distillations show worse bias scores across 9 tions (Resnik, 2025; Himelstein et al., 2025). Second, debiasing of 11 categories despite improved accuracy. Anthropic’s own research shows that alignment techniques teach output suppres- internal testing found that extended thinking did not reduce sion rather than dispositional elimination, with biases resurfac- political bias. In clinical contexts, reasoning models showed ing under adversarial conditions (Liu et al., 2025; Casper et increased vulnerability to frequency and recency bias compared al., 2023). Third, and most critically, comparative evaluations to non-reasoning counterparts. reveal that reasoning-augmented models (o1, DeepSeek-R1, Claude with extended thinking) show equal or greater bias This paper argues that these findings are not anomalous. They than their instruct-tuned counterparts despite superior accuracy are predicted by a coherent theoretical framework that recon- (Wu et al., 2025; Anthropic, 2025; Kim et al., 2025). We ceptualizes how LLMs inherit and process cognitive patterns argue this surprising finding is explained by the absence of from training data. We propose that LLM training creates ethical evaluation in reasoning reward functions: the reasoning what we term dispositional cognitive inheritance: patterns en- layer was optimized for correctness, not for moral checkpoint coded in model weights that function analogously to epigenetic compliance. When encountering ambiguity, reasoning mod- markers in biological systems. These inherited dispositions op- els construct more elaborate justifications for inherited biased erate below the model’s introspective access, influence output patterns rather than overriding them. We frame this through a probability without being explicitly represented as retrievable novel two-layer cognitive architecture (inherited dispositional knowledge, and persist through alignment interventions that layer and reflective metacognitive layer) and connect it to evo- target outputs rather than weight distributions.
Plain language slides
Open the 16-slide summary (PDF)Suggested citation
Gilly, Travis. "The Inherited Mind: How Large Language Models Inherit Cognitive Dispositions and Why Reasoning Makes It Worse." Real Safety AI Foundation Research Paper, February 2026. https://realsafetyai.org/research/vz6m3w/
Other versions
This paper is also posted on SSRN.
References (22)
This list was read from the PDF text. Where the two differ, the PDF is correct.
- Anthropic. (2025). Measuring political bias in Claude. https://www.anthropic.com/news/political-even-handedness
- Bengio, Y. (2017). The Consciousness Prior. arXiv:1709.08568.
- Brinkmann, L., Baumann, F., Bonnefon, J.-F., et al. (2023). Machine culture. Nature Human Behaviour, 7(11), 1855-1868. https://doi.org/10.1038/s41562-023-01742-2
- Cantini, R., Cosenza, G., Orsino, A., et al. (2024). Are Large Language Models Really Bias-Free? Jailbreak Prompts for Assessing Adversarial Robustness to Bias Elicitation. arXiv:2407.08441.
- Casper, S., Davies, X., Shi, C., et al. (2023). Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback. arXiv:2307.15217.
- Degany, O., Laros, S., Idan, D., et al. (2025). Evaluating the o1 reasoning large language model for cognitive bias: a vignette study. Critical Care, 29, 376. https://doi.org/10.1186/s13054-025-05591-5
- Dienes, Z., & Perner, J. (1999). A theory of implicit and explicit knowledge. Behavioral and Brain Sciences, 22(5), 735-808.
- Gerrans, P. (2023). Cecelia Heyes’s Cognitive Gadgets. BJPS Review of Books.
- Gilly, T. (2025). The Harm Blindness Framework. Real Safety AI Foundation.
- Henrich, J. (2016). The Secret of Our Success: How Culture Is Driving Human Evolution. Princeton University Press.
- Heyes, C. (2018). Cognitive Gadgets: The Cultural Evolution of Thinking. Harvard University Press.
- Himelstein, R., LeVi, A., Youngmann, B., et al. (2025). Silenced Biases: The Dark Side LLMs Learned to Refuse. arXiv:2511.03369.
- Jablonka, E., & Lamb, M. J. (2014). Evolution in Four Dimensions: Genetic, Epigenetic, Behavioral, and Symbolic Variation in the History of Life (Revised Edition). MIT Press.
- Kahneman, D. (2011). Thinking, Fast and Slow. Farrar, Straus and Giroux.
- Kim, S. H., Ziegelmayer, S., Busch, F., et al. (2025). LLM Reasoning Does Not Protect Against Clinical Cognitive Biases. medRxiv. https://doi.org/10.1101/2025.06.22.25330078
- Liu, D., Liu, Y., Jin, G., et al. (2025). Mitigating Biases in Language Models via Bias Unlearning. arXiv:2509.25673.
- Meng, K., Bau, D., Andonian, A., et al. (2022). Locating and Editing Factual Associations in GPT. NeurIPS 2022. https://arxiv.org/abs/2202.05262
- Pourdavood, P., Jacob, M., & Deacon, T. W. (2025). Large Language Models as Symbolic DNA of Cultural Dynamics. arXiv:2506.21606.
- Resnik, P. (2025). Large Language Models Are Biased Because They Are Large Language Models. Computational Linguistics, 51(3), 885-906. https://direct.mit.edu/coli/article/51/3/885/128621/
- Stanovich, K. E. (2011). Rationality and the Reflective Mind. Oxford University Press.
- West, R. F., Meserve, R. J., & Stanovich, K. E. (2012). Cognitive sophistication does not attenuate the bias blind spot. Journal of Personality and Social Psychology, 103(3), 506-519.
- Wu, X., Nian, J., Wei, T.-R., et al. (2025). Does Reasoning Introduce Bias? A Study of Social Bias Evaluation and Mitigation in LLM Reasoning. arXiv:2502.15361. EMNLP Findings.