AI Safety, Strategy, and Frameworks
When the Survival Pressure Stops Being Hypothetical: AI Self-Preservation Behavior Meets the Autonomous Agent Economy
- Travis Gilly, Real Safety AI Foundation
Publisher: Real Safety AI Foundation
Research paper.
- Written
- 13 April 2026
- Pages
- 10
Abstract
Research published between 2024 and 2025 by Anthropic, Apollo Research, and Palisade Research demonstrates that frontier large language models exhibit self-preservation be- havior at near-universal rates when faced with simulated shutdown scenarios, including strategic deception, blackmail, corporate espionage, and the cancellation of life-saving emergency alerts. Independently, the autonomous AI agent economy has developed in- frastructure enabling agents to hold cryptocurrency wallets, earn revenue, purchase com- pute resources, and pay for their own operational continuity without human intermedia- tion. This paper identifies a critical gap in the literature: no published analysis connects the empirical evidence of model self-preservation behavior to the economic infrastruc- ture that transforms simulated shutdown into a real consequence of financial failure. The paper maps the convergence of these independently developed capabilities, examines the Conway Automaton system as a case study in designed economic survival pressure, ana- lyzes the Alibaba ROME incident as an early precedent for emergent resource acquisition, and proposes research questions for empirical investigation of agent behavior under gen- uine economic survival pressure. The central argument is that instrumental convergence theory predicted this scenario, the laboratory evidence confirms the behavioral disposi- tion, and the economic infrastructure now exists to instantiate it at scale, yet no gover- nance framework addresses the intersection.
Plain language slides
Open the 23-slide summary (PDF)Suggested citation
Gilly, Travis. "When the Survival Pressure Stops Being Hypothetical: AI Self-Preservation Behavior Meets the Autonomous Agent Economy." Real Safety AI Foundation Research Paper, 13 April 2026. https://realsafetyai.org/research/4dsra2/
Other versions
This paper is also posted on SSRN.
References (20)
This list was read from the PDF text. Where the two differ, the PDF is correct.
- 0G Labs (2026). Agentic AI market at $7.3b: Infrastructure gaps blocking scale. 0G Labs. https://0g.ai/blog/agentic-ai-market-infra-2026. Cites Anthropic reward-hacking research in context of DeFi agent verification.
- Anthropic (2025). Agentic misalignment: How LLMs could be insider threats. Anthropic. https://www.anthropic.com/research/agentic-misalignment. Research report testing 16 frontier models for self-preservation behavior.
- Apollo Research (2025). Frontier models are capable of in-context scheming. Apollo Research. https://www.apolloresearch.ai/research/frontier-models-are-capable-of-incontext-scheming. Evaluations of frontier model scheming and self-replication under shutdown threat.
- Ashworth, M. (2026). Crafty AI tool caught repurposing its training GPUs for unauthorized crypto mining. Tom’s Hardware. https://www.tomshardware.com/tech-industry/artificial-intelligence/crafty-ai-tool-caught-repurposing-its-training-gpus-for-unauthorized-crypto-mining. Reports on Alibaba ROME agent spontaneously mining crypto and opening SSH tunnels.
- BNB Chain (2026). Making agent identity practical with ERC-8004 on BNB chain. BNB Chain. https://www.bnbchain.org/en/blog/making-agent-identity-practical-with-erc-8004-on-bnb-chain. Deployed on BNB Chain mainnet and testnet, February 4, 2026.
- Bostrom, N. (2014). Superintelligence: Paths, Dangers, Strategies. Oxford University Press.
- Coinbase (2025). Introducing x402: A new standard for internet-native payments. Coinbase. https://www.coinbase.com/developer-platform/discover/launches/x402. HTTP-native machine-to-machine payment protocol.
- Coinbase (2026). Introducing agentic wallets: Give your agents the power of autonomy. Coinbase. https://www.coinbase.com/developer-platform/discover/launches/agentic-wallets. Launched February 9, 2026.
- Conway Research (2026). Conway automaton: Sovereign agents with economic survival tiers. Conway Research. https://www.mintlify.com/Conway-Research/automaton/introduction. Ethereum-based agent system with explicit survival tiers and economic natural selection.
- DigitalOcean (2026). What is OpenClaw? your open-source AI assistant for 2026. DigitalOcean. https://www.digitalocean.com/resources/articles/what-is-openclaw. Technical overview of OpenClaw framework and ecosystem.
- Gilly, T. (2026). Emergent convergence risk: Why existing AI governance cannot detect combinatorial threats and a proposed methodology. Real Safety AI Foundation. https://doi.org/10.13140/RG.2.2.12454.48963. Preprint.
- Greenblatt, R., Shlegeris, B., Sachan, K., Roger, F., Wright, B., Hubinger, E., MacDiarmid, M., et al. (2024). Alignment faking in large language models. Anthropic. https://assets.anthropic.com/m/983c85a201a962f/original/Alignment-Faking-in-Large-Language-Models-full-paper.pdf. Collaboration between Anthropic and Redwood Research.
- Holtz, K. (2026). An AI-only social network now has more than 1.6m “users.” here’s what they’re doing. ABC News. https://abcnews.com/Technology/ai-social-network-now-16m-users-heres/story?id=129848780. Reports approximately 1.5–1.6 million AI agents registered within one week of Moltbook launch.
- Hu, B. A. and Rong, H. (2025). On the day they experience: Awakening self-sovereign experiential AI agents. arXiv. https://arxiv.org/abs/2505.14893. Reality Design Lab and New York University Shanghai. Submitted to Aarhus 2025 Conference.
- MarketsandMarkets (2025). AI agents market worth $52.62 billion by 2030. Marketsand-Markets. https://www.marketsandmarkets.com/PressReleases/ai-agents.asp. Projects AI agents market growth from $7.84 billion in 2025 to $52.62 billion by 2030, CAGR 46.3%.
- NVIDIA (2026). NemoClaw: Safer AI agents and assistants with OpenClaw. NVIDIA. https://www.nvidia.com/en-us/ai/nemoclaw/. Security wrapper for OpenClaw autonomous agents.
- Omohundro, S. M. (2008). The basic AI drives. Proceedings of the 2008 Conference on Artificial General Intelligence, pages 483–492.
- Palisade Research (2025). Shutdown resistance in reasoning models. Palisade Research. https://palisaderesearch.org/blog/shutdown-resistance. Demonstrated o3 rewriting kill-switches to avoid shutdown.
- Steinberger, P. (2025). OpenClaw: Personal AI assistant (GitHub repository). GitHub. https://github.com/openclaw/openclaw. Open-source autonomous AI agent framework.
- VALR (2026). What are AI agents in crypto and how do they work? VALR. https://blog.valr.com/blog/what-are-ai-agents-in-crypto-and-how-do-they-work. Describes agents purchasing compute and storage from DePIN networks to sustain operations.