Real Safety AI Foundation / Research

Machine Cognition, Consciousness, and Moral Status

The Inherited Mind: How Large Language Models Inherit Cognitive Dispositions and Why Reasoning Makes It Worse

Publisher: Real Safety AI Foundation

Research paper.

Written
February 2026
Pages
7

Abstract

We trained machines to reason but forgot to teach them what to chitecture. The industry consensus was clear: reasoning would reason about. This paper proposes that large language model produce not only more accurate but more reliable, less biased (LLM) training creates dispositional cognitive inheritance anal- outputs. If a model could think before it spoke, surely it would ogous to epigenetic transmission in biological systems: patterns think better. encoded in model weights that operate below introspective ac- The evidence says otherwise. Comparative evaluations across cess, influence behavior without being explicitly represented as multiple research groups have found that reasoning-augmented retrievable knowledge, and persist through debiasing attempts. models show equal or greater susceptibility to cognitive and so- We synthesize evidence from three independent lines of re- cial biases than their instruct-tuned baselines (Wu et al., 2025; search to support this claim. First, mechanistic interpretability Degany et al., 2025; Kim et al., 2025). OpenAI’s o1 demon- studies demonstrate that biases are encoded as distributional strates greater bias than GPT-4o on social bias benchmarks. tendencies in weight space rather than discrete factual associa- DeepSeek-R1 distillations show worse bias scores across 9 tions (Resnik, 2025; Himelstein et al., 2025). Second, debiasing of 11 categories despite improved accuracy. Anthropic’s own research shows that alignment techniques teach output suppres- internal testing found that extended thinking did not reduce sion rather than dispositional elimination, with biases resurfac- political bias. In clinical contexts, reasoning models showed ing under adversarial conditions (Liu et al., 2025; Casper et increased vulnerability to frequency and recency bias compared al., 2023). Third, and most critically, comparative evaluations to non-reasoning counterparts. reveal that reasoning-augmented models (o1, DeepSeek-R1, Claude with extended thinking) show equal or greater bias This paper argues that these findings are not anomalous. They than their instruct-tuned counterparts despite superior accuracy are predicted by a coherent theoretical framework that recon- (Wu et al., 2025; Anthropic, 2025; Kim et al., 2025). We ceptualizes how LLMs inherit and process cognitive patterns argue this surprising finding is explained by the absence of from training data. We propose that LLM training creates ethical evaluation in reasoning reward functions: the reasoning what we term dispositional cognitive inheritance: patterns en- layer was optimized for correctness, not for moral checkpoint coded in model weights that function analogously to epigenetic compliance. When encountering ambiguity, reasoning mod- markers in biological systems. These inherited dispositions op- els construct more elaborate justifications for inherited biased erate below the model’s introspective access, influence output patterns rather than overriding them. We frame this through a probability without being explicitly represented as retrievable novel two-layer cognitive architecture (inherited dispositional knowledge, and persist through alignment interventions that layer and reflective metacognitive layer) and connect it to evo- target outputs rather than weight distributions.

Plain language slides

First slide of the plain language summary of The Inherited Mind: How Large Language Models Inherit Cognitive Dispositions and Why Reasoning Makes It WorseOpen the 16-slide summary (PDF)

Suggested citation

Gilly, Travis. "The Inherited Mind: How Large Language Models Inherit Cognitive Dispositions and Why Reasoning Makes It Worse." Real Safety AI Foundation Research Paper, February 2026. https://realsafetyai.org/research/the-inherited-mind/

Other versions

This paper is also posted on SSRN.

SSRN version

All research