Convergence Risk and Perceptual Modification
The Accidental Stack: Emergent Perceptual Control Infrastructure from Independently Developed AI Capabilities
- Travis Gilly, Real Safety AI Foundation
Publisher: Real Safety AI Foundation
Research paper.
- Pages
- 9
Abstract
We present a case study in emergent convergence risk, documenting how a single week of independent AI capability releases (March 2026) produced the component stack for a population-scale perceptual control system that no individual developer intended, designed, or recognized. We inventory eight open-source tools released or updated within a seven-day period: a person segmentation model (MatAnyone2), a real-time 3D world generator (InSpatio-WorldFM), a mobile 3D renderer (MobileGS), a video generation accelerator (Diagonal Distillation), a voice cloner (TaDa), a tagged speech synthesizer (Fish Audio S2), a visual effects transfer system (Effect Maker), and a multi-shot cinematic video generator (Shotverse). Each tool was reviewed in isolation by a popular technology channel and evaluated as a standalone capability with legitimate applications. None were identified as combinable into a unified system. We apply integration surface analysis, extending the Harm Blindness Framework from product-level checkpoints to capability-combination checkpoints, to demonstrate how individually benign tools become infrastructure for perceptual modification when composed. We argue this case study illustrates a structural gap in existing AI governance: no current methodology, including the EU AI Act risk classification, NIST AI RMF, or MIT AI Risk Repository, systematically scans for emergent harm at integration surfaces between independently developed capabilities. This gap is not an oversight in implementation; it is a limitation of the methodological paradigm.
Plain language slides
Open the 18-slide summary (PDF)Suggested citation
Gilly, Travis. "The Accidental Stack: Emergent Perceptual Control Infrastructure from Independently Developed AI Capabilities." Real Safety AI Foundation Research Paper, n.d.. https://realsafetyai.org/research/the-accidental-stack/
Other versions
This paper is also posted on SSRN.
References (14)
This list was read from the PDF text. Where the two differ, the PDF is correct.
- AI Search. (2026, March). AI maps, realtime 3D worlds, multi-shot videos, new TTS, new anime model: AI NEWS [Video]. YouTube.
- CD Projekt Red. (2020). Cyberpunk 2077 [Video game]. CD Projekt.
- Compulsion Games. (2018). We Happy Few [Video game]. Gearbox Publishing.
- Daniel, M. (2026, March 12). How we're reimagining Maps with Gemini. Google Blog. https://blog.google/products-and-platforms/products/maps/ask-maps-immersive-navigation/
- Du, X., Jiang, Y., Ye, T., Han, S., & Liu, S. (2026). Mobile-GS: Real-time Gaussian splatting for mobile devices. In Proceedings of the International Conference on Learning Representations (ICLR). arXiv:2603.11531. https://xiaobiaodu.github.io/mobile-gs-project/
- European Parliament & Council. (2024). Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (AI Act). Official Journal of the European Union. EUR-Lex: 32024R1689.
- Gilly, T. (2025-2026). The Harm Blindness Framework: A practical application methodology for stakeholder harm prevention in technology development. Real Safety AI Foundation. https://realsafetyai.org/framework
- Gilly, T. (2026b). We Happy Few: Therapeutic perceptual modification, regulatory pathway precedent, and the governance gap in cloud-rendered reality. Forthcoming.
- Liu, J., Liu, X., Mei, K., Wen, Y., Yang, M.-H., & Liu, W. (2026). Streaming autoregressive video generation via diagonal distillation. In Proceedings of the International Conference on Learning Representations (ICLR). arXiv:2603.09488. https://github.com/Sphere-AI-Lab/diagdistill
- MIT FutureTech. (2024). AI Risk Repository. Massachusetts Institute of Technology. https://airisk.mit.edu
- National Institute of Standards and Technology. (2023). Artificial Intelligence Risk Management Framework (AI RMF 1.0) (NIST AI 100-1). U.S. Department of Commerce. https://doi.org/10.6028/NIST.AI.100-1
- Wang, J., et al. (2026). Geometry-guided reinforcement learning for multi-view consistent 3D scene editing. arXiv preprint arXiv:2603.03143. https://github.com/AMAP-ML/RL3DEdit
- Yang, P., Zhou, S., Hao, K., & Tao, Q. (2026a). MatAnyone 2: Scaling video matting via a learned quality evaluator. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). arXiv:2512.11782. https://github.com/pq-yang/MatAnyone2
- Zhang, X., et al. (2026). InSpatio-WorldFM: An open-source real-time generative frame model for spatial intelligence. arXiv preprint arXiv:2603.11911. https://inspatio.github.io/worldfm/