AI Safety, Strategy, and Frameworks
BIOS: Bootstrap Instruction for Operational Safety
- Travis Gilly, Real Safety AI Foundation
Publisher: Real Safety AI Foundation
Working paper. Not peer reviewed.
- Written
- December 25, 2025
- Pages
- 16
Abstract
Current LLM safety architectures deploy sophisticated guardrail systems (NeMo Guardrails, RoboGuard, AGrail, LlamaFirewall) that define WHAT to check and HOW to verify safety. However, these systems share a critical vulnerability: they assume their own operational instructions persist in context. When context windows overflow, not only do safety rules degrade; the meta-instruction to invoke the guardrail itself can be pushed out, leaving the LLM operating without active safety enforcement. This paper introduces BIOS (Bootstrap Instruction for Operational Safety), a meta-layer architecture that ensures guardrail systems continue to execute regardless of context window state. Rather than injecting safety rules per-turn (expensive, easily defeated), BIOS injects lightweight meta-instructions that remind the system to invoke its guardrails; analogous to how a PC's BIOS ensures the operating system loads rather than running applications directly.
Plain language slides
Open the 20-slide summary (PDF)Suggested citation
Gilly, Travis. "BIOS: Bootstrap Instruction for Operational Safety." Real Safety AI Foundation Working Paper, December 25, 2025. https://realsafetyai.org/research/bios-bootstrap-instruction-for-operational-safety/
References (15)
This list was read from the PDF text. Where the two differ, the PDF is correct.
- AgentSpec Consortium. (2025). AgentSpec: Per-step execution enforcement for AI agents. arXiv. https://doi.org/10.48550/arXiv.2503.18666
- Asimov, I. (1942). Runaround. Astounding Science Fiction, 29(1), 94-103.
- AWS Security. (2024, July 7). Context window overflow: Breaking the barrier. AWS Security Blog. https://aws.amazon.com/blogs/security/context-window-overflow-breaking-the-barrier/
- Chen, Y., Aniketh, K., Xie, C., Zhao, D., Lee, I., & Matni, N. (2025). RoboGuard: A robotic agent safeguard framework via safety-aware task planning. arXiv. https://doi.org/10.48550/arXiv.2503.15538
- Goertzel, B. (2024). Metagoals: Endowing self-modifying AGI systems with goal stability or moderated goal evolution. arXiv. https://doi.org/10.48550/arXiv.2412.16559
- Liu, X., Xu, Z., Liu, Y., Wang, X., & Chen, D. (2024). Cognitive overload attack: Prompt injection for long context. arXiv. https://doi.org/10.48550/arXiv.2410.11272
- LongSafety Team. (2024). LongSafety: Enhance safety for long-context LLMs. arXiv. https://doi.org/10.48550/arXiv.2411.06899
- Mei, L., Liu, S., Wang, Y., Bi, B., Mao, J., & Cheng, X. (2025). BABYBLUE: Benchmark for reliability and jailbreak hallucination evaluation. arXiv. https://doi.org/10.48550/arXiv.2406.11668
- Meta AI. (2025). LlamaFirewall: An open-source guardrail system for building secure AI agents. https://ai.meta.com/research/publications/llamafirewall-an-open-source-guardrail-system-for-building-secure-ai-agents/
- Multi-Agent Safety Consortium. (2025). Trading off security and collaboration capabilities in multi-agent systems. arXiv. https://doi.org/10.48550/arXiv.2502.19145
- Sequent Safety Team. (2025). Sequent safety for machine-checkable control in multi-agent systems. arXiv. https://doi.org/10.48550/arXiv.2512.16279
- Rebedea, T., Dinu, R., Sreedhar, M., Parisien, C., & Cohen, J. (2023). NeMo Guardrails: A toolkit for controllable and safe LLM applications with programmable rails. arXiv. https://doi.org/10.48550/arXiv.2310.10501
- Unit 42. (2025, February 12). Indirect prompt injection poisons AI long-term memory. Palo Alto Networks. https://unit42.paloaltonetworks.com/indirect-prompt-injection-poisons-ai-long-term-memory/
- Young, R. J. (2025). Evaluating the robustness of large language model safety guardrails against adversarial attacks. arXiv. https://doi.org/10.48550/arXiv.2511.22047
- Zhang, W., Yang, Y., Xue, H., Xu, D., & Xia, L. (2025). AGrail: A lifelong agent guardrail with effective and adaptive safety detection. arXiv. https://doi.org/10.48550/arXiv.2502.11448