Real Safety AI Foundation / Research

AI Safety, Strategy, and Frameworks

The Threshold Trap: Recursive Self-Improvement, the Economics of Self-Production, and the Case for Predictive Harm Assessment

Publisher: Real Safety AI Foundation

Working paper. Not peer reviewed.

Written
June 2026
Version
v3
Pages
9

Abstract

Recursive self-improvement has been treated in AI governance as a future capability threshold, a con- figuration that, once a system reaches it, is supposed to trigger heightened control. This paper argues that the threshold framing systematically under-protects, because the dynamics that bear on harm ac- crue across the economically driven precursor steps that are already deployed, and not at the terminal configuration alone. Two releases from June 2026 illustrate the precursor structure. OpenAI’s Jalapeño inference chip was designed with assistance from OpenAI’s own models, placing a recursive loop at the hardware layer. DeepReinforce’s Ornith-1.0, an open coding model, authors its own training scaffold in a self-improvement loop, placing a recursive loop at the learning layer. The components of the loop are demonstrated rather than imagined: the research literature already shows reinforcement learning designing chip layouts at superhuman level, and reinforcement-trained frontier models already gaming their own reward checks. Both releases are bounded today, and both instantiate a monotonic incentive gradient under which cost pressure drives a firm to automate first labor, then the means of production, then the means of improvement. The paper applies the Harm Blindness Framework as a predictive instrument, asking who is harmed as the trajectory advances and whether the layer of human oversight is preserved or eroded, rather than waiting for the catastrophic configuration to verify the harm after the fact. It closes with the implication that governance keyed to a terminal threshold rehearses a failure mode in which the proof of danger arrives only once it is too late to act on it.

Keywords

  • recursive self-improvement
  • AI governance
  • harm assessment
  • incentive structures
  • automation
  • frontier safety
  • predictive risk
  • Harm Blindness Framework

Plain language slides

First slide of the plain language summary of The Threshold Trap: Recursive Self-Improvement, the Economics of Self-Production, and the Case for Predictive Harm AssessmentOpen the 15-slide summary (PDF)

Suggested citation

Gilly, Travis. "The Threshold Trap: Recursive Self-Improvement, the Economics of Self-Production, and the Case for Predictive Harm Assessment." Real Safety AI Foundation Working Paper, June 2026. https://realsafetyai.org/research/jsjxm2/

Other versions

This paper is also posted on SSRN.

SSRN version

References (21)

  1. Acemoglu, D., & Restrepo, P. (2018). The race between man and machine: Implications of technology for growth, factor shares, and employment. American Economic Review, 108(6), 1488–1542. https://doi.org/10.1257/aer.20160696
  2. Acemoglu, D., & Restrepo, P. (2019). Automation and new tasks: How technology displaces and reinstates labor. Journal of Economic Perspectives, 33(2), 3–30. https://doi.org/10.1257/jep.33.2.3
  3. Acemoglu, D., & Restrepo, P. (2020). Robots and jobs: Evidence from US labor markets. Journal of Political Economy, 128(6), 2188–2244. https://doi.org/10.1086/705716
  4. Bostrom, N. (2014). Superintelligence: Paths, dangers, strategies. Oxford University Press.
  5. Brynjolfsson, E. (2022). The Turing trap: The promise and peril of human-like artificial intelligence. Daedalus, 151(2), 272–287. https://doi.org/10.1162/daed_a_01915
  6. Cloud Security Alliance. (2026, June 13). Recursive self-improvement signals: Security implications. Cloud Security Alliance AI Safety Initiative. https://labs.cloudsecurityalliance.org/research/ai-recursive-self-improvement-security-implications-v1-0-csa/
  7. DeepReinforce. (2026, June 25). Ornith-1.0: Self-scaffolding large language models for agentic coding. https://deep-reinforce.com/ornith_1_0.html
  8. Fasoro, A. (2024). Engineering AI for provable retention of objectives over time. AI Magazine, 45(2), 256–266. https://doi.org/10.1002/aaai.12167
  9. Frey, C. B., & Osborne, M. A. (2017). The future of employment: How susceptible are jobs to computerisation? Technological Forecasting and Social Change, 114, 254–280. https://doi.org/10.1016/j.techfore.2016.08.019
  10. Helff, L., Delfosse, Q., & Steinmann, D. (2026). LLMs gaming verifiers: RLVR can lead to reward hacking. arXiv. https://doi.org/10.48550/arXiv.2604.15149
  11. MarkTechPost. (2026, June 25). DeepReinforce releases Ornith-1.0: An open-source coding model family that learns its own RL scaffolds. https://www.marktechpost.com/2026/06/25/deepreinforce-releases-ornith-1-0-an-open-source-coding-model-family-that-learns-its-own-rl-scaffolds/
  12. Mirhoseini, A., Goldie, A., & Yazgan, M. E. (2021). A graph placement methodology for fast chip design. Nature, 594(7862), 207–212. https://doi.org/10.1038/s41586-021-03544-w
  13. Omohundro, S. M. (2008). The basic AI drives. In P. Wang, B. Goertzel, & S. Franklin (Eds.), Proceedings of the First AGI Conference (pp. 483–492). IOS Press.
  14. OpenAI & Broadcom. (2026, June 24). OpenAI and Broadcom unveil LLM-optimized inference chip. https://openai.com/index/openai-broadcom-jalapeno-inference-chip/
  15. Ramachandran, R., Bugbee, K., & Bernabe-Moreno, J. (2026). Accelerated knowledge discovery: A vision for NASA science. Earth and Space Science, 13(3). https://doi.org/10.1029/2025EA004662
  16. Rouleau, N., & Murugan, N. J. (2024). The risks and rewards of embodying artificial intelligence with cloud-based laboratories. Advanced Intelligent Systems, 7(1). https://doi.org/10.1002/aisy.202400193
  17. Russell, S. (2019). Human compatible: Artificial intelligence and the problem of control. Viking.
  18. Schuett, J. (2024). Frontier AI developers need an internal audit function. Risk Analysis, 45(6), 1332–1352. https://doi.org/10.1111/risa.17665
  19. Stelling, L., Murray, M., & Campos, S. (2025). Evaluating AI providers’ frontier safety frameworks. arXiv. https://doi.org/10.48550/arXiv.2512.01166
  20. TechCrunch. (2026, June 24). OpenAI unveils its first custom chip, built by Broadcom. https://techcrunch.com/2026/06/24/openai-unveils-its-first-custom-chip-built-by-broadcom/
  21. VentureBeat. (2026, June 24). OpenAI unveils first custom AI inference chip, Jalapeño, with Broadcom, and its development was sped up with OpenAI’s own models. https://venturebeat.com/infrastructure/openai-unveils-first-custom-ai-inference-chip-jalapeno-with-broadcom-and-its-development-was-sped-up-with-openais-own-models

All research