Real Safety AI Foundation / Research

AI Safety, Strategy, and Frameworks

BIOS: Bootstrap Instruction for Operational Safety

Publisher: Real Safety AI Foundation

Working paper. Not peer reviewed.

Written
December 25, 2025
Pages
16

Abstract

Current LLM safety architectures deploy sophisticated guardrail systems (NeMo Guardrails, RoboGuard, AGrail, LlamaFirewall) that define WHAT to check and HOW to verify safety. However, these systems share a critical vulnerability: they assume their own operational instructions persist in context. When context windows overflow, not only do safety rules degrade; the meta-instruction to invoke the guardrail itself can be pushed out, leaving the LLM operating without active safety enforcement. This paper introduces BIOS (Bootstrap Instruction for Operational Safety), a meta-layer architecture that ensures guardrail systems continue to execute regardless of context window state. Rather than injecting safety rules per-turn (expensive, easily defeated), BIOS injects lightweight meta-instructions that remind the system to invoke its guardrails; analogous to how a PC's BIOS ensures the operating system loads rather than running applications directly.

Plain language slides

First slide of the plain language summary of BIOS: Bootstrap Instruction for Operational SafetyOpen the 20-slide summary (PDF)

Suggested citation

Gilly, Travis. "BIOS: Bootstrap Instruction for Operational Safety." Real Safety AI Foundation Working Paper, December 25, 2025. https://realsafetyai.org/research/95brmp/

References (15)

This list was read from the PDF text. Where the two differ, the PDF is correct.

  1. AgentSpec Consortium. (2025). AgentSpec: Per-step execution enforcement for AI agents. arXiv. https://doi.org/10.48550/arXiv.2503.18666
  2. Asimov, I. (1942). Runaround. Astounding Science Fiction, 29(1), 94-103.
  3. AWS Security. (2024, July 7). Context window overflow: Breaking the barrier. AWS Security Blog. https://aws.amazon.com/blogs/security/context-window-overflow-breaking-the-barrier/
  4. Chen, Y., Aniketh, K., Xie, C., Zhao, D., Lee, I., & Matni, N. (2025). RoboGuard: A robotic agent safeguard framework via safety-aware task planning. arXiv. https://doi.org/10.48550/arXiv.2503.15538
  5. Goertzel, B. (2024). Metagoals: Endowing self-modifying AGI systems with goal stability or moderated goal evolution. arXiv. https://doi.org/10.48550/arXiv.2412.16559
  6. Liu, X., Xu, Z., Liu, Y., Wang, X., & Chen, D. (2024). Cognitive overload attack: Prompt injection for long context. arXiv. https://doi.org/10.48550/arXiv.2410.11272
  7. LongSafety Team. (2024). LongSafety: Enhance safety for long-context LLMs. arXiv. https://doi.org/10.48550/arXiv.2411.06899
  8. Mei, L., Liu, S., Wang, Y., Bi, B., Mao, J., & Cheng, X. (2025). BABYBLUE: Benchmark for reliability and jailbreak hallucination evaluation. arXiv. https://doi.org/10.48550/arXiv.2406.11668
  9. Meta AI. (2025). LlamaFirewall: An open-source guardrail system for building secure AI agents. https://ai.meta.com/research/publications/llamafirewall-an-open-source-guardrail-system-for-building-secure-ai-agents/
  10. Multi-Agent Safety Consortium. (2025). Trading off security and collaboration capabilities in multi-agent systems. arXiv. https://doi.org/10.48550/arXiv.2502.19145
  11. Sequent Safety Team. (2025). Sequent safety for machine-checkable control in multi-agent systems. arXiv. https://doi.org/10.48550/arXiv.2512.16279
  12. Rebedea, T., Dinu, R., Sreedhar, M., Parisien, C., & Cohen, J. (2023). NeMo Guardrails: A toolkit for controllable and safe LLM applications with programmable rails. arXiv. https://doi.org/10.48550/arXiv.2310.10501
  13. Unit 42. (2025, February 12). Indirect prompt injection poisons AI long-term memory. Palo Alto Networks. https://unit42.paloaltonetworks.com/indirect-prompt-injection-poisons-ai-long-term-memory/
  14. Young, R. J. (2025). Evaluating the robustness of large language model safety guardrails against adversarial attacks. arXiv. https://doi.org/10.48550/arXiv.2511.22047
  15. Zhang, W., Yang, Y., Xue, H., Xu, D., & Xia, L. (2025). AGrail: A lifelong agent guardrail with effective and adaptive safety detection. arXiv. https://doi.org/10.48550/arXiv.2502.11448

All research