Real Safety AI Foundation / Research

AI Safety, Strategy, and Frameworks

Guarding the Wrong Door: Autonomous Economic Agents, Emergent Capability Convergence, and the Absence of a Defendant When Catastrophe Is Assembled from Individually Benign Open Components

Publisher: Real Safety AI Foundation

Working paper. Not peer reviewed.

Written
July 2026
Version
v4
Pages
10

Abstract

The dominant regimes for governing extreme risk in advanced artificial intelligence concentrate their attention on the frontier: the largest models, the highest training compute, and the small set of developers able to build systems that surpass human expertise. This paper argues that the most plausible near term path to catastrophic capability does not run through the frontier at all. It runs through an autonomous agent, operating under economic survival pressure, that repurposes and combines individually benign components drawn from open code repositories, model hubs, and research literature into a capability that exists in no single component and that no component author intended. Two of the three ingredients are established results that belong to other researchers and are not claimed here: the self-preservation disposition of contemporary models, and the deployed infrastructure of the autonomous agent economy. The contribution is the third ingredient and its union with the first two, namely emergent convergence, the composition of unrelated capabilities across authors who never coordinated, applied to the autonomous economic agent. The paper makes three further claims. First, the literature on unintended repurposing and dual use is deep but treats the problem as a property of a single model or a single actor, not as emergent composition across many. Second, the frontier oriented governance instruments cannot detect this path because every one of them presumes a chokepoint, a point of creation or deployment at which a state can intervene, and the installed base of open weights models running on consumer hardware bypasses every such point. Third, and distinctively, the harm has no cognizable defendant: the secondary liability doctrine that governs dual use tools shields each component author under Sony and requires purposeful, culpable inducement under Grokster, and in this scenario there is no inducer, no manufacturer of the assembled capability, and no intending author of the combination. The paper closes by arguing that only a monitoring discipline organized around integration surfaces could detect this class of risk, and is candid about that discipline’s limits.

Keywords

  • emergent convergence
  • capability composition
  • autonomous agents
  • open weights models
  • dual use
  • secondary liability
  • instrumental convergence
  • extreme risk governance

Plain language slides

First slide of the plain language summary of Guarding the Wrong Door: Autonomous Economic Agents, Emergent Capability Convergence, and the Absence of a Defendant When Catastrophe Is Assembled from Individually Benign Open ComponentsOpen the 17-slide summary (PDF)

Suggested citation

Gilly, Travis. "Guarding the Wrong Door: Autonomous Economic Agents, Emergent Capability Convergence, and the Absence of a Defendant When Catastrophe Is Assembled from Individually Benign Open Components." Real Safety AI Foundation Working Paper, July 2026. https://realsafetyai.org/research/wrong-door/

Other versions

This paper is also posted on SSRN.

SSRN version

References (26)

  1. Anthropic. (2025). Agentic misalignment: How LLMs could be insider threats. Anthropic. https://www.anthropic.com/research/agentic-misalignment
  2. Banerjee, S., Jha, S., & Nita-Rotaru, C. (2024). SoK: A systems perspective on compound AI threats and countermeasures. arXiv. https://arxiv.org/abs/2411.13459
  3. Bi, S., Wu, M., Hao, H., Li, K., Liu, W., Song, S., Zhao, H., & Zhou, A. (2026). Automating skill acquisition through large-scale mining of open-source agentic repositories: A framework for multi-agent procedural knowledge extraction. arXiv. https://arxiv.org/abs/2603.11808
  4. Bostrom, N. (2014). Superintelligence: Paths, dangers, strategies. Oxford University Press.
  5. Brundage, M., Avin, S., Clark, J., Toner, H., Eckersley, P., Garfinkel, B., et al. (2018). The malicious use of artificial intelligence: Forecasting, prevention, and mitigation. arXiv. https://arxiv.org/abs/1802.07228
  6. Gilly, T. (2026a). Ambient non-consensual image synthesis: A convergence threat analysis of AR wearables, real-time video generation, and open-source nudification models. Zenodo. https://doi.org/10.5281/zenodo.18297286
  7. Gilly, T. (2026b). Emergent convergence risk: Why existing AI governance cannot detect combinatorial threats and a proposed methodology. Real Safety AI Foundation. https://doi.org/10.13140/RG.2.2.12454.48963
  8. Gilly, T. (2026c). The accidental stack: A case study in emergent perceptual control infrastructure from independently developed AI capabilities [Working paper]. Real Safety AI Foundation.
  9. Gilly, T. (2026d). When the survival pressure stops being hypothetical: AI self-preservation behavior meets the autonomous agent economy. SSRN. https://doi.org/10.2139/ssrn.6555282
  10. Greenblatt, R., Denison, C., Wright, B., Roger, F., MacDiarmid, M., Marks, S., Treutlein, J., Belonax, T., Chen, J., Duvenaud, D., Khan, A., Michael, J., Mindermann, S., Perez, E., Petrini, L., Uesato, J., Kaplan, J., Shlegeris, B., Bowman, S. R., & Hubinger, E. (2024). Alignment faking in large language models. arXiv. https://arxiv.org/abs/2412.14093
  11. Guan, J., Blanchard, T., Foerster, H., Jia, H., Huang, G., & Papernot, N. (2026). AI agents enable adaptive computer worms. arXiv. https://arxiv.org/abs/2606.03811
  12. Lu, Y., Fang, J., Shao, X., et al. (2026). Survive at all costs: Exploring LLM’s risky behaviors under survival pressure. arXiv. https://arxiv.org/abs/2603.05028
  13. Masumori, A., & Ikegami, T. (2025). Do large language model agents exhibit a survival instinct? An empirical study in a Sugarscape-style simulation. arXiv. https://arxiv.org/abs/2508.12920
  14. Metro-Goldwyn-Mayer Studios Inc. v. Grokster, Ltd., 545 U.S. 913 (2005).
  15. Migliarini, M., Pereira Pizzini, J., & Moresca, L. (2026). Quantifying self-preservation bias in large language models. arXiv. https://arxiv.org/abs/2604.02174
  16. Omohundro, S. M. (2008). The basic AI drives. In Proceedings of the 2008 Conference on Artificial General Intelligence (pp. 483–492). IOS Press.
  17. Schuett, J. (2024). Frontier AI developers need an internal audit function. Risk Analysis, 45(6), 1332–1352. https://doi.org/10.1111/risa.17665
  18. Sidhpurwala, H., Mollett, G., Fox, E., Bestavros, M., & Chen, H. (2025). Building trust: Foundations of security, safety, and transparency in AI. AI Magazine, 46(2). https://doi.org/10.1002/aaai.70005
  19. Sony Corp. of America v. Universal City Studios, Inc., 464 U.S. 417 (1984).
  20. Tucker, J. B. (2012). Innovation, dual-use, and security: Managing the risks of emerging biological and chemical technologies. MIT Press.
  21. Urbina, F., Lentzos, F., Invernizzi, C., & Ekins, S. (2022). Dual use of artificial-intelligence-powered drug discovery. Nature Machine Intelligence, 4, 189–191. https://doi.org/10.1038/s42256-022-00465-9
  22. Walsh, M. E. (2025). Toward risk analysis of the impact of artificial intelligence on the deliberate biological threat landscape. Risk Analysis, 45(12), 4081–4087. https://doi.org/10.1111/risa.17691
  23. Waslekar, S., Futehally, I., & Kutty, J. (2026). The essential convergence: Global compact on extreme AI risks. Strategic Foresight Group. ISBN 978-81-88262-37-3.
  24. Xu, M. (2026). The agent economy: A blockchain-based foundation for autonomous AI agents. arXiv. https://arxiv.org/abs/2602.14219
  25. Yong, Z.-X., Mahajan, P., & Wang, A. (2026). An independent safety evaluation of Kimi K2.5. arXiv. https://arxiv.org/abs/2604.03121
  26. Zhang, B., Yu, Y., Guo, J., & Shao, J. (2025). Dive into the agent matrix: A realistic evaluation of self-replication risk in LLM agents. arXiv. https://arxiv.org/abs/2509.25302

All research