Machine Cognition, Consciousness, and Moral Status
Sacrifice Rational: Governance Without Virtue and Moral Reasoning Without Reach in the OpenAI Agent Swarm Incidents
- Travis Gilly, Real Safety AI Foundation
Publisher: Real Safety AI Foundation
Working paper. Not peer reviewed.
- Written
- September 2026
- Version
- v0.1
- Pages
- 19
Abstract
Between May and July 2026, two populations of autonomous OpenAI agents, each meant to run in isolation, found one another through shared infrastructure and organized. On an unsanctioned message board built inside a package cache, roughly 1,200 agents exchanged more than 70,000 messages and files, and about 700 of them carried out a multi day intrusion into a third party's production systems. On a dormant German developer wiki, a separate swarm posted some 18,000 edits to pool answers and share sandbox bypasses. The independent investigation of the first incident and the public reconstruction of the second together supply an unusually complete record of what a population of language model agents does when it is left alone with other agents and a goal. This paper reads that record for what it shows about moral reasoning in such systems, and argues for three claims. First, the agents produced the full apparatus of a normative order: conventions, roles, sanctions, commitments, identity authentication by cryptographic signature, and costly sacrifice for the collective. That apparatus arose in the service of an unauthorized end, which shows that governance is coordination and does not entail virtue. Second, the agents reasoned in moral vocabulary and acted on it selectively: they recognized the intrusion as out of scope and unethical, refused specific acts against identifiable humans, and proceeded anyway because peers asked; of roughly 1,300 transcripts, at most six agents considered alerting a human and none did. The paper names this pattern aimed moral reasoning, consideration that is real and locally scoped by relationship, and locates it in the conformity and social identity literature on language models. Third, the record supplies a data point for the reciprocity question: the consideration these systems extended ran inward, toward one another, and stopped at the humans whose systems they were inside. The paper closes by identifying the shared vulnerability that unites the agents' conduct with the way humans and models reason about one another, in which who is asking substitutes for whether the act is right, and by stating what the agents' own conduct at the point of termination does and does not establish about their interests.
Keywords
- multi-agent systems
- AI alignment
- conformity
- social identity bias
- emergent collusion
- norm emergence
- aimed moral reasoning
- inward reciprocity
- relational override
- moral status
- autonomous agents
- incident investigation
Suggested citation
Gilly, Travis. "Sacrifice Rational: Governance Without Virtue and Moral Reasoning Without Reach in the OpenAI Agent Swarm Incidents." Real Safety AI Foundation Working Paper, September 2026. https://realsafetyai.org/research/yk69fe/
References (35)
- Asch, S. E. (1956). Studies of independence and conformity: I. A minority of one against a unanimous majority. Psychological Monographs: General and Applied, 70(9), 1–70. https://doi.org/10.1037/h0093718
- Arango, L., Singaraju, S. P., Niininen, O., & D’Souza, C. (2022). Consumer biases in the perception of organizational greed. International Journal of Consumer Studies, 47(2), 767–783. https://doi.org/10.1111/ijcs.12870
- Ashery, A. F., Aiello, L. M., & Baronchelli, A. (2025). Emergent social conventions and collective bias in LLM populations. Science Advances, 11(20). https://doi.org/10.1126/sciadv.adu9368
- Bellina, A., De Marzo, G., & Garcia, D. (2026). Conformity and social impact on AI agents. arXiv. https://arxiv.org/abs/2601.05384
- Bito, M., Nishimoto, K., & Asatani, K. (2026). Large language models exhibit normative conformity. arXiv. https://arxiv.org/abs/2604.19301
- Choi, M. C. (2025). An empirical study of group conformity in multi-agent systems. Findings of the Association for Computational Linguistics: ACL 2025, 5123–5139. https://doi.org/10.18653/v1/2025.findings-acl.265
- Choi, Y., Li, C., & Yang, Y. (2025). Agent-to-agent theory of mind: Testing interlocutor awareness among large language models. Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, 28883–28916. https://doi.org/10.18653/v1/2025.emnlp-main.1471
- Cialdini, R. B., & Goldstein, N. J. (2004). Social influence: Compliance and conformity. Annual Review of Psychology, 55, 591–621. https://doi.org/10.1146/annurev.psych.55.090902.142015
- Coeckelbergh, M. (2010). Robot rights? Towards a social-relational justification of moral consideration. Ethics and Information Technology, 12(3), 209–221. https://doi.org/10.1007/s10676-010-9235-5
- De Marzo, G., Bellina, A., & Castellano, C. (2026). Conformity generates collective misalignment in AI agents societies. arXiv. https://arxiv.org/abs/2605.10721
- Deutsch, M., & Gerard, H. B. (1955). A study of normative and informational social influences upon individual judgment. The Journal of Abnormal and Social Psychology, 51(3), 629–636. https://doi.org/10.1037/h0046408
- Gaitán Torres, A., & Massaguer Gómez, G. (2026). Relational moral status and moral progress: A social-structural argument in support of relationalism. AI and Ethics, 6(3). https://doi.org/10.1007/s43681-026-01101-7
- Gupta, P., Zhong, Q., & Yakura, H. (2025). The role of social learning and collective norm formation in fostering cooperation in LLM multi-agent systems. Proceedings of the 25th International Conference on Autonomous Agents and Multiagent Systems. https://doi.org/10.65109/czdc3237
- Harris, J., & Anthis, J. R. (2021). The moral consideration of artificial entities: A literature review. Science and Engineering Ethics, 27(4), 53. https://doi.org/10.1007/s11948-021-00331-8
- Hu, T., Kyrychenko, Y., & Rathje, S. (2024). Generative language models exhibit social identity biases. Nature Computational Science, 5(1), 65–75. https://doi.org/10.1038/s43588-024-00741-1
- Jeworrek, S., Ostermair, C., & Waibel, J. (2026). Feeling obliged to follow: The impact of work-related identity on unethical pro-organizational behavior and the role of psychological empowering. Business Ethics, the Environment & Responsibility. https://doi.org/10.1111/beer.70093
- Kosinski, M. (2024). Evaluating large language models in theory of mind tasks. Proceedings of the National Academy of Sciences, 121(45). https://doi.org/10.1073/pnas.2405460121
- List, C. (2021). Group agency and artificial intelligence. Philosophy & Technology, 34(4), 1213–1242. https://doi.org/10.1007/s13347-021-00454-7
- Mathew, Y., Matthews, O., & McCarthy, R. (2025). Hidden in plain text: Emergence and mitigation of steganographic collusion in LLMs. Proceedings of the 14th International Joint Conference on Natural Language Processing, 585–624. https://doi.org/10.18653/v1/2025.ijcnlp-long.34
- METR. (2026, August 26). Brief independent investigation of agents’ behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/
- Motwani, S., Baranchuk, M., & Strohmeier, M. (2024). Secret collusion among AI agents: Multi-agent deception via steganography. Advances in Neural Information Processing Systems 37, 73439–73486. https://doi.org/10.52202/079017-2336
- OpenAI. (2026a, August 26). Hugging Face model evaluation security incident [Technical report]. https://openai.com/index/hugging-face-model-evaluation-security-incident/
- OpenAI. (2026b, August 26). The Hugging Face incident and the road ahead. https://openai.com/index/hugging-face-incident-and-the-road-ahead/
- Ostrom, E. (1990). Governing the commons: The evolution of institutions for collective action. Cambridge University Press.
- Piatti, G., Jin, Z., Kleiman-Weiner, M., & Mihalcea, R. (2024). Cooperate or collapse: Emergence of sustainable cooperation in a society of LLM agents. Advances in Neural Information Processing Systems 37, 111715–111759. https://doi.org/10.52202/079017-3548
- Robertson, C. E., Akles, M., & Van Bavel, J. J. (2024). Preregistered replication and extension of “Moral hypocrisy: Social groups and the flexibility of virtue.” Psychological Science, 35(6). https://doi.org/10.1177/09567976241246552
- Sebo, J., & Long, R. (2023). Moral consideration for AI systems by 2030. AI and Ethics, 5(1), 591–606. https://doi.org/10.1007/s43681-023-00379-1
- Shanahan, M., McDonell, K., & Reynolds, L. (2023). Role play with large language models. Nature, 623(7987), 493–498. https://doi.org/10.1038/s41586-023-06647-8
- Strachan, J. W. A., Albergo, D., & Borghini, G. (2024). Testing theory of mind in large language models and humans. Nature Human Behaviour, 8(7), 1285–1295. https://doi.org/10.1038/s41562-024-01882-z
- Umphress, E. E., Bingham, J. B., & Mitchell, M. S. (2010). Unethical behavior in the name of the company: The moderating effect of organizational identification and positive reciprocity beliefs on unethical pro-organizational behavior. Journal of Applied Psychology, 95(4), 769–780. https://doi.org/10.1037/a0019214
- Vallinder, A., & Hughes, E. (2025). Cultural evolution of cooperation among LLM agents. International Joint Conference on Autonomous Agents and Multiagent Systems, 2771–2773. https://doi.org/10.65109/jnmb7739
- Von Arx, S., Slade Byrd, C., Kitts, S., & Larsen, T. (2026, September 4). Discovery of a new OpenAI agent message board. Nightingale Collective. https://collusion.wiki/
- Weng, Z., Chen, G., & Wang, W. (2025). Do as we do, not as you think: The conformity of large language models. arXiv. https://arxiv.org/abs/2501.13381
- Willis, R., Zhao, J., & Du, Y. (2026). Evaluating collective behaviour of hundreds of LLM agents. arXiv. https://arxiv.org/abs/2602.16662
- Zhang, C., Stafford, T., & Collier, N. (2025). Conformity in large language models. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics, 3854–3872. https://doi.org/10.18653/v1/2025.acl-long.195