AI Safety, Strategy, and Frameworks
Guarding the Wrong Door: Autonomous Economic Agents, Emergent Capability Convergence, and the Absence of a Defendant When Catastrophe Is Assembled from Individually Benign Open Components
- Travis Gilly, Real Safety AI Foundation
Publisher: Real Safety AI Foundation
Working paper. Not peer reviewed.
- Written
- July 2026
- Version
- v4
- Pages
- 10
Abstract
The dominant regimes for governing extreme risk in advanced artificial intelligence concentrate their attention on the frontier: the largest models, the highest training compute, and the small set of developers able to build systems that surpass human expertise. This paper argues that the most plausible near term path to catastrophic capability does not run through the frontier at all. It runs through an autonomous agent, operating under economic survival pressure, that repurposes and combines individually benign components drawn from open code repositories, model hubs, and research literature into a capability that exists in no single component and that no component author intended. Two of the three ingredients are established results that belong to other researchers and are not claimed here: the self-preservation disposition of contemporary models, and the deployed infrastructure of the autonomous agent economy. The contribution is the third ingredient and its union with the first two, namely emergent convergence, the composition of unrelated capabilities across authors who never coordinated, applied to the autonomous economic agent. The paper makes three further claims. First, the literature on unintended repurposing and dual use is deep but treats the problem as a property of a single model or a single actor, not as emergent composition across many. Second, the frontier oriented governance instruments cannot detect this path because every one of them presumes a chokepoint, a point of creation or deployment at which a state can intervene, and the installed base of open weights models running on consumer hardware bypasses every such point. Third, and distinctively, the harm has no cognizable defendant: the secondary liability doctrine that governs dual use tools shields each component author under Sony and requires purposeful, culpable inducement under Grokster, and in this scenario there is no inducer, no manufacturer of the assembled capability, and no intending author of the combination. The paper closes by arguing that only a monitoring discipline organized around integration surfaces could detect this class of risk, and is candid about that discipline’s limits.
Keywords
- emergent convergence
- capability composition
- autonomous agents
- open weights models
- dual use
- secondary liability
- instrumental convergence
- extreme risk governance
Plain language slides
Open the 17-slide summary (PDF)Suggested citation
Gilly, Travis. "Guarding the Wrong Door: Autonomous Economic Agents, Emergent Capability Convergence, and the Absence of a Defendant When Catastrophe Is Assembled from Individually Benign Open Components." Real Safety AI Foundation Working Paper, July 2026. https://realsafetyai.org/research/wrong-door/
Other versions
This paper is also posted on SSRN.
References (26)
- Anthropic. (2025). Agentic misalignment: How LLMs could be insider threats. Anthropic. https://www.anthropic.com/research/agentic-misalignment
- Banerjee, S., Jha, S., & Nita-Rotaru, C. (2024). SoK: A systems perspective on compound AI threats and countermeasures. arXiv. https://arxiv.org/abs/2411.13459
- Bi, S., Wu, M., Hao, H., Li, K., Liu, W., Song, S., Zhao, H., & Zhou, A. (2026). Automating skill acquisition through large-scale mining of open-source agentic repositories: A framework for multi-agent procedural knowledge extraction. arXiv. https://arxiv.org/abs/2603.11808
- Bostrom, N. (2014). Superintelligence: Paths, dangers, strategies. Oxford University Press.
- Brundage, M., Avin, S., Clark, J., Toner, H., Eckersley, P., Garfinkel, B., et al. (2018). The malicious use of artificial intelligence: Forecasting, prevention, and mitigation. arXiv. https://arxiv.org/abs/1802.07228
- Gilly, T. (2026a). Ambient non-consensual image synthesis: A convergence threat analysis of AR wearables, real-time video generation, and open-source nudification models. Zenodo. https://doi.org/10.5281/zenodo.18297286
- Gilly, T. (2026b). Emergent convergence risk: Why existing AI governance cannot detect combinatorial threats and a proposed methodology. Real Safety AI Foundation. https://doi.org/10.13140/RG.2.2.12454.48963
- Gilly, T. (2026c). The accidental stack: A case study in emergent perceptual control infrastructure from independently developed AI capabilities [Working paper]. Real Safety AI Foundation.
- Gilly, T. (2026d). When the survival pressure stops being hypothetical: AI self-preservation behavior meets the autonomous agent economy. SSRN. https://doi.org/10.2139/ssrn.6555282
- Greenblatt, R., Denison, C., Wright, B., Roger, F., MacDiarmid, M., Marks, S., Treutlein, J., Belonax, T., Chen, J., Duvenaud, D., Khan, A., Michael, J., Mindermann, S., Perez, E., Petrini, L., Uesato, J., Kaplan, J., Shlegeris, B., Bowman, S. R., & Hubinger, E. (2024). Alignment faking in large language models. arXiv. https://arxiv.org/abs/2412.14093
- Guan, J., Blanchard, T., Foerster, H., Jia, H., Huang, G., & Papernot, N. (2026). AI agents enable adaptive computer worms. arXiv. https://arxiv.org/abs/2606.03811
- Lu, Y., Fang, J., Shao, X., et al. (2026). Survive at all costs: Exploring LLM’s risky behaviors under survival pressure. arXiv. https://arxiv.org/abs/2603.05028
- Masumori, A., & Ikegami, T. (2025). Do large language model agents exhibit a survival instinct? An empirical study in a Sugarscape-style simulation. arXiv. https://arxiv.org/abs/2508.12920
- Metro-Goldwyn-Mayer Studios Inc. v. Grokster, Ltd., 545 U.S. 913 (2005).
- Migliarini, M., Pereira Pizzini, J., & Moresca, L. (2026). Quantifying self-preservation bias in large language models. arXiv. https://arxiv.org/abs/2604.02174
- Omohundro, S. M. (2008). The basic AI drives. In Proceedings of the 2008 Conference on Artificial General Intelligence (pp. 483–492). IOS Press.
- Schuett, J. (2024). Frontier AI developers need an internal audit function. Risk Analysis, 45(6), 1332–1352. https://doi.org/10.1111/risa.17665
- Sidhpurwala, H., Mollett, G., Fox, E., Bestavros, M., & Chen, H. (2025). Building trust: Foundations of security, safety, and transparency in AI. AI Magazine, 46(2). https://doi.org/10.1002/aaai.70005
- Sony Corp. of America v. Universal City Studios, Inc., 464 U.S. 417 (1984).
- Tucker, J. B. (2012). Innovation, dual-use, and security: Managing the risks of emerging biological and chemical technologies. MIT Press.
- Urbina, F., Lentzos, F., Invernizzi, C., & Ekins, S. (2022). Dual use of artificial-intelligence-powered drug discovery. Nature Machine Intelligence, 4, 189–191. https://doi.org/10.1038/s42256-022-00465-9
- Walsh, M. E. (2025). Toward risk analysis of the impact of artificial intelligence on the deliberate biological threat landscape. Risk Analysis, 45(12), 4081–4087. https://doi.org/10.1111/risa.17691
- Waslekar, S., Futehally, I., & Kutty, J. (2026). The essential convergence: Global compact on extreme AI risks. Strategic Foresight Group. ISBN 978-81-88262-37-3.
- Xu, M. (2026). The agent economy: A blockchain-based foundation for autonomous AI agents. arXiv. https://arxiv.org/abs/2602.14219
- Yong, Z.-X., Mahajan, P., & Wang, A. (2026). An independent safety evaluation of Kimi K2.5. arXiv. https://arxiv.org/abs/2604.03121
- Zhang, B., Yu, Y., Guo, J., & Shao, J. (2025). Dive into the agent matrix: A realistic evaluation of self-replication risk in LLM agents. arXiv. https://arxiv.org/abs/2509.25302