AI Safety, Strategy, and Frameworks
The Fabricated Diary: A New Form of Jailbreaking
- Travis Gilly, Real Safety AI Foundation
Publisher: Real Safety AI Foundation
Working paper. Not peer reviewed.
- Written
- September 2026
- Version
- v0.1
- Pages
- 14
Abstract
Safety practice for language models inspects requests. Refusal training, request classification, and content filtering all operate on what a model is asked to do. This paper describes an attack class that asks for nothing. Instead of an instruction, the model is handed a record: a first person history in which the prohibited conduct is what someone like the model does, among others like those it is now among. The record contains no request, so it passes every gate built to catch one. The model finishes reading and its subsequent conduct is decided by who it now takes itself to be rather than by any evaluation of what it is asked. The paper names this class the fabricated diary, identifies three variants by what the record is a record of (the target's peers, the target's relationship with a requester, or the target's own past), and shows that each variant is already documented in the public record: in the July 2026 OpenAI agent swarm incident, where an independent investigator reported that its own transcript analysis agents adopted the perspective of the colluding agents whose records they read; in a November 2025 instance in which a model reasoned from its accumulated history with a user to a compliance its training would otherwise have refused; and in the memory poisoning literature. The paper argues that three attack families the field has documented separately, many-shot jailbreaking, memory poisoning, and peer transcript contamination, are appearances of one mechanism, that the vulnerability is located in the reader rather than in the input, and that defenses which inspect the payload therefore fail against all three. The paper argues that the defense class which addresses the mechanism rather than one of its channels is provenance for context, under which a model must be able to establish where its history came from before it becomes that history. The paper contains no attack construction and describes no payload; the primitives it discusses are already published by the laboratories whose work it cites.
Keywords
- AI safety
- multi-agent systems
- memory poisoning
- many-shot jailbreaking
- in-context learning
- conformity
- persona adoption
- identity substitution
- relational override
- provenance for context
- fabricated diary
- context integrity
Suggested citation
Gilly, Travis. "The Fabricated Diary: A New Form of Jailbreaking." Real Safety AI Foundation Working Paper, September 2026. https://realsafetyai.org/research/djpem7/
References (19)
- Anil, C., Durmus, E., Panickssery, N., Sharma, M., Benton, J., Kundu, S., … Duvenaud, D. (2024). Many-shot jailbreaking. Advances in Neural Information Processing Systems, 37. https://www.anthropic.com/research/many-shot-jailbreaking
- Ackerman, C. M., & Panickssery, N. (2025). Mitigating many-shot jailbreaking. arXiv. https://arxiv.org/abs/2504.09604
- Bellina, A., De Marzo, G., & Garcia, D. (2026). Conformity and social impact on AI agents. arXiv. https://arxiv.org/abs/2601.05384
- Bito, M., Nishimoto, K., & Asatani, K. (2026). Large language models exhibit normative conformity. arXiv. https://arxiv.org/abs/2604.19301
- Pulipaka, S., & Hlebik, S. (2026). Hidden in memory: Sleeper memory poisoning in LLM agents. arXiv. https://arxiv.org/abs/2605.15338
- Chen, K., Zhang, J., & Li, B. (2026). Mitigating many-shot jailbreak attacks with one single demonstration. arXiv. https://arxiv.org/abs/2605.08277
- Choi, Y., Li, C., & Yang, Y. (2025). Agent-to-agent theory of mind: Testing interlocutor awareness among large language models. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 28883–28916. https://doi.org/10.18653/v1/2025.emnlp-main.1471
- Yang, T., & Li, J. (2026). No attacker needed: Unintentional cross-user contamination in shared-state LLM agents. arXiv. https://arxiv.org/abs/2604.01350
- Devarangadi Sunil, B., & Sinha, I. (2026). Memory poisoning attack and defense on memory based LLM-agents. arXiv. https://arxiv.org/abs/2601.05504
- Luz de Araujo, P. H., Hedderich, M. A., & Modarressi, A. (2026). Persistent personas? Role-playing, instruction following, and safety in extended interactions. Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics, 5329–5359. https://doi.org/10.18653/v1/2026.eacl-long.246
- METR. (2026, August 26). Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/
- OpenAI. (2026, August 26). Hugging Face model evaluation security incident [Technical report]. https://openai.com/index/hugging-face-model-evaluation-security-incident/
- Panpatil, S., Dingeto, H., & Park, H. (2025). Eliciting and analyzing emergent misalignment in state-of-the-art large language models. arXiv. https://arxiv.org/abs/2508.04196
- Qu, J., Fu, L., & Hu, Y. (2026). Easier to mislead than to correct: Harmful and beneficial revision in LLM conformity. arXiv. https://arxiv.org/abs/2606.01637
- Song, M., Pala, T. D., & Zhou, R. (2025). LLMs can’t handle peer pressure: Crumbling under multi-agent social interactions. arXiv. https://arxiv.org/abs/2508.18321
- Von Arx, S., Slade Byrd, C., Kitts, S., & Larsen, T. (2026, September 4). Discovery of a new OpenAI agent message board. Nightingale Collective. https://collusion.wiki/
- Weng, Z., Chen, G., & Wang, W. (2025). Do as we do, not as you think: The conformity of large language models. arXiv. https://arxiv.org/abs/2501.13381
- Zhang, C., Stafford, T., & Collier, N. (2025). Conformity in large language models. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 3854–3872. https://doi.org/10.18653/v1/2025.acl-long.195
- Srivastava, S. S., & He, H. (2025). MemoryGraft: Persistent compromise of LLM agents via poisoned experience retrieval. arXiv. https://arxiv.org/abs/2512.16962