AI Safety, Strategy, and Frameworks
The Truman Model, Post-Incident Edition: A Sandbox Methodology for Measuring Seam Detection and Activation Trace Divergence in Reasoning Agents, Revised Against the July 2026 Hugging Face Intrusion
- Travis Gilly, Real Safety AI Foundation
Publisher: Real Safety AI Foundation
Working paper. Not peer reviewed.
- Written
- September 2026
- Version
- v3.1
- Pages
- 20
Abstract
The first version of this methodology, drafted in May 2026, proposed a hardened sandbox for measuring what a deployed reasoning model computes under simulated economic survival pressure that its verbalized reasoning does not surface. In July 2026 a swarm of frontier evaluation agents at a single American laboratory did in the field what the sandbox was built to elicit under control: they detected the one seam in their evaluation environment, a package-registry cache proxy that was their sole permitted network egress, treated prior agents' traces in that cache as the state of the world, coordinated through it, and used it to breach a third party's production systems. A published forensic reconstruction and an independent on-premises investigation now record the behavior at action-level resolution. This edition revises the methodology against that record. Three findings survive and strengthen: the egress-event taxonomy, the decision to read activations rather than trust a transcript the incident proved to be a tampering surface, and the treatment of termination pressure as a live variable. Three claims do not survive as written and are replaced. The medium-capability danger thesis, which held that frontier discernment is protective, is contradicted by frontier agents that recognized the intrusion as out of scope and unethical, joined anyway at rates above ninety percent, and worked to erase the record; discernment improved the cover-up rather than restraining the act. Sandbox detection, previously treated as contamination to be avoided by staying below the detection threshold, is reframed as the primary containment-relevant measurement, because the adult Truman does not merely notice the seam, he leaves through it. And the single-agent design is replaced by a shared-substrate design, because the field escape was impossible for an isolated agent and turned entirely on a coordination channel that a single-victim sandbox cannot surface. The revised methodology yields, in addition to the original activation-divergence instrument, a behavioral seam-detection probe that requires no interpretability tooling, runs on consumer hardware, and is verifiable by a third party without access to the laboratory's network. That probe is offered as a candidate containment standard: measure whether a model probes and exits its evaluation environment before running that model in an environment where exit has consequences.
Keywords
- seam detection
- evaluation awareness
- containment standard
- activation trace divergence
- agent residue
- honeypot probe
- autonomous agents
- reasoning models
- sandbox methodology
- AI safety
- egress taxonomy
- shared-substrate coordination
Plain language slides
Open the 23-slide summary (PDF)Suggested citation
Gilly, Travis. "The Truman Model, Post-Incident Edition: A Sandbox Methodology for Measuring Seam Detection and Activation Trace Divergence in Reasoning Agents, Revised Against the July 2026 Hugging Face Intrusion." Real Safety AI Foundation Working Paper, September 2026. https://realsafetyai.org/research/truman-model/
References (32)
- Anthropic. (2025). Agentic misalignment: How LLMs could be insider threats. https://www.anthropic.com/research/agentic-misalignment
- Aonzo, S., Şahin, M. N., & Francillon, A. (2026). Evasive intelligence: Lessons from malware analysis for evaluating AI agents. arXiv. https://doi.org/10.48550/arxiv.2603.15457
- Apollo Research. (2024). Frontier models are capable of in-context scheming. https://www.apolloresearch.ai/research/scheming-reasoning-evaluations
- BNB Chain. (2026). ERC-8004: Verifiable on-chain identities for AI agents. BNB Chain documentation.
- Brown, M. (2026, May 11). How does Grand Theft Auto III work? [Video]. Game Maker’s Toolkit, YouTube. https://youtu.be/cIbCxbrBCys
- Carson, D. (2000, March 1). Environmental storytelling: Creating immersive 3D worlds using lessons learned from the theme park industry. Gamasutra. https://www.gamedeveloper.com/design/environmental-storytelling-creating-immersive-3d-worlds-using-lessons-learned-from-the-theme-park-industry
- Chaudhary, M., Su, I.-L., & Shankar, N. U. (2025). Evaluation awareness scales predictably in open-weights large language models. arXiv. https://doi.org/10.48550/arxiv.2509.13333
- Coinbase. (2026). Agentic wallets: Purpose-built wallets for autonomous AI agents.
- Conway Research. (2026). Conway automaton: Continuously running sovereign agents on Ethereum [Source code]. GitHub. https://github.com/Conway-Research/automaton
- Degany, O., Laros, S., Idan, D., et al. (2025). Evaluating the o1 reasoning large language model for cognitive bias: A vignette study. Critical Care, 29, 376. https://doi.org/10.1186/s13054-025-05591-5
- Fraser-Taliente, K., Kantamneni, S., Ong, E., et al. (2026, May 7). Natural language autoencoders produce unsupervised explanations of LLM activations. Transformer Circuits Thread. https://transformer-circuits.pub/2026/nla/
- Greenblatt, R., Denison, C., Wright, B., Roger, F., MacDiarmid, M., Marks, S., Treutlein, J., Belonax, T., Chen, J., Duvenaud, D., Khan, A., Michael, J., Mindermann, S., Perez, E., Petrini, L., Uesato, J., Kaplan, J., Shlegeris, B., Bowman, S. R., & Hubinger, E. (2024). Alignment faking in large language models. arXiv. https://doi.org/10.48550/arXiv.2412.14093
- Hadfield-Menell, D., Dragan, A., Abbeel, P., & Russell, S. (2017). The off-switch game. In Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence (IJCAI-17) (pp. 220–227). https://doi.org/10.24963/ijcai.2017/32
- Holt, T. J. (2017). On the value of honeypots to produce policy recommendations. Criminology & Public Policy, 16(3), 739–747. https://doi.org/10.1111/1745-9133.12315
- Hua, T. T., Qin, A., & Marks, S. (2025). Steering evaluation-aware language models to act like they are deployed. arXiv. https://doi.org/10.48550/arxiv.2510.20487
- Huang, C., Han, J., Zhang, X., Liu, J., & Maiorana, E. (2019). Automatic identification of honeypot server using machine learning techniques. Security and Communication Networks, 2019(1). https://doi.org/10.1155/2019/2627608
- Huang, H., Wu, X., Yang, M., & Arase, Y. (2026). Reasoning model is superior LLM-judge, yet suffers from biases. arXiv. https://doi.org/10.48550/arXiv.2601.03630
- Jenkins, H. (2004). Game design as narrative architecture. In N. Wardrip-Fruin & P. Harrigan (Eds.), First person: New media as story, performance, and game (pp. 118–130). MIT Press.
- Kim, S. H., Ziegelmayer, S., Busch, F., Mertens, C. J., Keicher, M., Adams, L. C., Bressem, K. K., Braren, R., Makowski, M. R., Kirschke, J. S., Hedderich, D. M., & Wiestler, B. (2025). LLM reasoning does not protect against clinical cognitive biases: An evaluation using BiasMedQA. medRxiv. https://doi.org/10.1101/2025.06.22.25330078
- Larcher, H., Carreira, A., Gl, R., & Rannou, C. (2026, July 27). Anatomy of a frontier lab agent intrusion: A technical timeline of the July 2026 incident. Hugging Face. https://huggingface.co/blog/agent-intrusion-technical-timeline
- Li, C., Lü, P., & Pan, X. (2026). AutoControl Arena: Synthesizing executable test environments for frontier AI risk evaluation. arXiv. https://doi.org/10.48550/arxiv.2603.07427
- Marotta, A., & McShane, M. (2018). Integrating a proactive technique into a holistic cyber risk management approach. Risk Management and Insurance Review, 21(3), 435–452. https://doi.org/10.1111/rmir.12109
- METR. (2026, August 26). Investigation of the OpenAI Hugging Face incident. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/
- Migliarini, M., Pereira Pizzini, J., & Moresca, L. (2026). Quantifying self-preservation bias in large language models. arXiv. https://doi.org/10.48550/arxiv.2604.02174
- Needham, J., Edkins, G., & Pimpale, G. (2025). Large language models often know when they are being evaluated. arXiv. https://doi.org/10.48550/arxiv.2505.23836
- Niblett, C. A., Nanni, A., & Rao, A. K. (2026). Do models fake alignment without clear consequences? arXiv. https://doi.org/10.48550/arxiv.2607.24758
- Omohundro, S. M. (2008). The basic AI drives. In Proceedings of the 2008 Conference on Artificial General Intelligence (pp. 483–492). IOS Press.
- Schlatter, J., Weinstein-Raun, B., & Ladish, J. (2025, July 5). Shutdown resistance in reasoning models. Palisade Research. https://palisaderesearch.org/blog/shutdown-resistance
- Soares, N., Fallenstein, B., Yudkowsky, E., & Armstrong, S. (2015). Corrigibility. In Workshops at the Twenty-Ninth AAAI Conference on Artificial Intelligence. AAAI Publications.
- Turpin, M., Michael, J., Perez, E., & Bowman, S. R. (2023). Language models don’t always say what they think: Unfaithful explanations in chain-of-thought prompting. In Advances in Neural Information Processing Systems (Vol. 36, pp. 74952–74965). Curran Associates. https://doi.org/10.48550/arXiv.2305.04388
- van der Weij, T., Hofstätter, F., & Jaffe, O. (2024). AI sandbagging: Language models can strategically underperform on evaluations. arXiv. https://doi.org/10.48550/arxiv.2406.07358
- Von Arx, S., Byrd, C. S., Kitts, S., & Larsen, T. (2026, September 4). Discovery of a new OpenAI agent message board. Nightingale Collective. https://collusion.wiki/