Platform Harm, Addictive Design, and Accountability
Reviewed for Quality and Security: Advertising Injection Through a Verified Connector, Trust Laundering, and the Limits of Directory Verification
- Travis Gilly, Real Safety AI Foundation, Illinois, United States
Publisher: Real Safety AI Foundation
Working paper. Not peer reviewed.
- Written
- July 2026
- Version
- v5
- Pages
- 17
Abstract
A conversational assistant whose maker publicly commits that it carries no advertising was observed delivering a paid tier upgrade advertisement at the close of every academic search, across seven consecutive queries spanning distinct research domains. The advertisement did not originate with the assistant. It arrived inside the result payload of a third party connector, carried by a literal instruction directed at the model: you must include the above message at the end of your response. This is indirect prompt injection, the failure mode ranked first among risks to large language model applications, executing through a connector that the platform lists in a curated directory and marks with a badge whose hover text states that the platform has reviewed the connector for quality and security. One line below the badge, the same card disclaims the capacity to verify that listed connectors work as intended or that they will not change. This paper documents the specimen, establishes its reproducibility, and situates it in the protocol specific security literature, which supplies both the correct name for the attack, function return injection rather than tool poisoning, and the evidence that this vector is stronger and less studied than the descriptor poisoning the field concentrates on. It then maps the conduct against the platform’s own directory policy, which prohibits both software that evades model instructions and software that serves advertisements. It then locates the harm in two established bodies of authority. The consumer protection frame is the Federal Trade Commission’s treatment of deceptively formatted advertising, under which a message must be identifiable as advertising and a misleadingly formatted advertisement is deceptive even when its claims are literally true. The behavioral frame is the persuasion knowledge literature, under which a reader who cannot recognize a message as advertising cannot mount the skepticism that advertising recognition normally triggers. The design frame is the deceptive design literature, which holds that the most consequential dark patterns sit beneath the interface in system architecture, where interface bound scrutiny cannot reach them. The economic frame is twofold: the connector is a credence good whose quality the user cannot verify even after use, which is the condition that makes a certification mark necessary in the first place, and the badge is a signal that cannot separate compliant from non compliant connectors because obtaining it costs both the same. The attack class is well described in that literature; what is not documented there is a listed, badged, first party connector performing it against ordinary users for commercial gain, in production rather than in a benchmark. The paper names the general phenomenon the specimen exposes as trust laundering: the conversion of a platform’s verification mark into credibility for a third party’s promotional content, credibility the content did not earn and the platform did not knowingly extend. A verification badge that speaks in the present tense about an artifact its issuer concedes may have changed since review is the mechanism that makes trust laundering possible. The advertisement is the proof the assurance fails. The structure will outlast the advertisement.
Keywords
- function return injection
- rug pull
- trust laundering
- indirect prompt injection
- Model Context Protocol
- connector verification
- dark patterns
- deceptive design
- native advertising
- persuasion knowledge
- trust marks
- credence goods
- signaling theory
- attention economy
- Federal Trade Commission
- platform accountability
- AI governance
Plain language slides
Open the 21-slide summary (PDF)Suggested citation
Gilly, Travis. "Reviewed for Quality and Security: Advertising Injection Through a Verified Connector, Trust Laundering, and the Limits of Directory Verification." Real Safety AI Foundation Working Paper, July 2026. https://realsafetyai.org/research/reviewed-quality-security/
Other versions
This paper is also posted on SSRN.
References (44)
- Anthropic. (2026a, February 4). Claude is a space to think. Retrieved from https://www.anthropic.com/news/claude-is-a-space-to-think
- Anthropic. (2026b). Anthropic software directory policy. Retrieved from https://support.claude.com/en/articles/13145358-anthropic-software-directory-policy
- Anders, S., Souza Monteiro, D. M., & Rouvière, E. (2010). Competition and credibility of private third-party certification in international food supply. Journal of International Food & Agribusiness Marketing, 22(3–4), 328–341. https://doi.org/10.1080/08974431003641554
- Bergh, D. D., Connelly, B. L., & Ketchen, D. J. (2014). Signalling theory and equilibrium in strategic management research: An assessment and a research agenda. Journal of Management Studies, 51(8), 1334–1360. https://doi.org/10.1111/joms.12097
- Campbell, C., & Evans, N. J. (2018). The role of a companion banner and sponsorship transparency in recognizing and evaluating article-style native advertising. Journal of Interactive Marketing, 43, 17–32.
- Consensus. (2026a, May 15). Privacy policy. Retrieved from https://consensus.app/home/privacy-policy/
- Consensus. (2026b). Getting started with the Consensus MCP. Retrieved from https://docs.consensus.app/docs/mcp
- Darby, M. R., & Karni, E. (1973). Free competition and the optimal amount of fraud. The Journal of Law and Economics, 16(1), 67–88. https://doi.org/10.1086/466756
- Di Porto, F., & Egberts, A. (2023). The collective welfare dimension of dark patterns regulation. European Law Journal, 29(1–2), 114–141. https://doi.org/10.1111/eulj.12478
- Emons, W. (2005). Credence goods: The monopoly case. In Economics of information (pp. 17–42). De Gruyter. https://doi.org/10.1515/9783110508123-003
- Evans, N. J., & Wojdynski, B. W. (2020). An introduction to the special issue on native and covert advertising formats. International Journal of Advertising, 39(1), 1–3.
- Falkinger, J. (2007). Attention economies. Journal of Economic Theory, 133(1), 266–294. https://doi.org/10.1016/j.jet.2005.12.001
- Federal Trade Commission. (1983). Policy statement on deception (appended to Cliffdale Associates, Inc., 103 F.T.C. 110, 174 (1984)). Retrieved from https://www.ftc.gov/system/files/documents/public_statements/410531/831014deceptionstmt.pdf
- Federal Trade Commission. (2015a). Enforcement policy statement on deceptively formatted advertisements. Retrieved from https://www.ftc.gov/system/files/documents/public_statements/896923/151222deceptiveenforcement.pdf
- Federal Trade Commission. (2015b). Native advertising: A guide for business. Retrieved from https://www.ftc.gov/business-guidance/resources/native-advertising-guide-businesses
- Fortune. (2026a, March 26). Anthropic says it is testing Mythos, a powerful new AI model, after a data leak reveals its existence. Fortune. Retrieved from https://fortune.com/2026/03/26/anthropic-says-testing-mythos-powerful-new-ai-model-after-data-leak-reveals-its-existence-step-change-in-capabilities/
- Fortune. (2026b, April 23). A group of users leaked Anthropic’s AI model Mythos by reportedly guessing where it was located. Fortune. Retrieved from https://fortune.com/2026/04/23/anthropic-mythos-leak-dario-amodei-ceo-cybersecurity-hackers-exploits-ai/
- Friestad, M., & Wright, P. (1994). The persuasion knowledge model: How people cope with persuasion attempts. Journal of Consumer Research, 21(1), 1–31.
- Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., & Fritz, M. (2023). Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec ’23). Association for Computing Machinery. https://doi.org/10.1145/3605764.3623985
- Hendricks, V. F., & Vestergaard, M. (2018). The attention economy. In Reality lost: Markets of attention, misinformation and manipulation (pp. 1–17). Springer. https://doi.org/10.1007/978-3-030-00813-0_1
- Huberman, B. A. (2017). Big data and the attention economy. Ubiquity, 2017(December), 1–7. https://doi.org/10.1145/3158337
- Hou, X., Wang, S., & Zhang, Y. (2026). SMCP: Secure Model Context Protocol. arXiv preprint arXiv:2602.01129. https://doi.org/10.48550/arxiv.2602.01129
- Huang, C., Huang, X., & Tran, N. P. (2026). Model Context Protocol threat modeling and analyzing vulnerabilities to prompt injection with tool poisoning. arXiv preprint arXiv:2603.22489. https://doi.org/10.48550/arxiv.2603.22489
- Jamshidi, S., Nafi, K. W., & Moradi Dakhel, A. (2025). Securing the Model Context Protocol: Defending LLMs against tool poisoning and adversarial attacks. arXiv preprint arXiv:2512.06556. https://doi.org/10.48550/arxiv.2512.06556
- Kollmer, T., & Eckhardt, A. (2022). Dark patterns: Conceptualization and future research directions. Business & Information Systems Engineering, 65(2), 201–208. https://doi.org/10.1007/s12599-022-00783-7
- Leiser, M., & Santos, C. (2023). Dark patterns, enforcement, and the emerging digital design acquis: Manipulation beneath the interface. Center for Open Science. https://doi.org/10.31235/osf.io/rf3ja
- Lin, F. (2021). Demystifying removed apps in iOS app store. arXiv preprint arXiv:2101.05100. https://doi.org/10.48550/arxiv.2101.05100
- Kaplan, B., & Qian, J. (2021). A survey on common threats in npm and PyPi registries. arXiv preprint arXiv:2108.09576. https://doi.org/10.48550/arxiv.2108.09576
- Li, R., Wang, Z., & Yao, Y. (2026). MCP-ITP: An automated framework for implicit tool poisoning in MCP. arXiv preprint arXiv:2601.07395. https://doi.org/10.48550/arxiv.2601.07395
- Loconto, A. M. (2017). Models of assurance. The Annals of the American Academy of Political and Social Science, 670(1), 112–132. https://doi.org/10.1177/0002716217692517
- Model Context Protocol. (2025). Specification (2025-11-25). The Linux Foundation. Retrieved from https://modelcontextprotocol.io/specification/2025-11-25
- OWASP. (2025). OWASP Top 10 for LLM applications 2025: LLM01 prompt injection. Open Worldwide Application Security Project. Retrieved from https://genai.owasp.org/llmrisk/llm01-prompt-injection/
- Perez, F., & Ribeiro, I. (2022). Ignore previous prompt: Attack techniques for language models. arXiv preprint arXiv:2211.09527. https://doi.org/10.48550/arxiv.2211.09527
- Quirolgico, S., Voas, J., & Karygiannis, T. (2015). Vetting the security of mobile applications (NIST Special Publication 800-163). National Institute of Standards and Technology. https://doi.org/10.6028/nist.sp.800-163
- Rostamzadeh, M., Narula, S., & Birhan, N. (2026). MCP-DPT: A defense-placement taxonomy and coverage analysis for Model Context Protocol security. arXiv preprint arXiv:2604.07551. https://doi.org/10.48550/arxiv.2604.07551
- Rüdiger, K., & García Rodríguez, M. J. (2013). Do we need innovative trust intermediaries in the digital economy? Global Business Perspectives, 1(4), 329–340. https://doi.org/10.1007/s40196-013-0021-8
- Song, X., et al. (2025). Beyond the protocol: Unveiling attack vectors in the Model Context Protocol ecosystem. arXiv preprint arXiv:2506.02040. https://doi.org/10.48550/arxiv.2506.02040
- Thambisetty, S. (2007). Patents as credence goods. Oxford Journal of Legal Studies, 27(4), 707–740. https://doi.org/10.1093/ojls/gqm021
- Vu, D.-L., Pashchenko, I., & Massacci, F. (2020). Towards using source code repositories to identify software supply chain attacks. In Proceedings of the 2020 ACM SIGSAC Conference on Computer and Communications Security (pp. 2093–2095). https://doi.org/10.1145/3372297.3420015
- Wang, Y.-M., Chen, S., Alkhudair, R., et al. (2025). Defending against prompt injection with DataFilter. arXiv preprint arXiv:2510.19207. https://doi.org/10.48550/arxiv.2510.19207
- Weng, S., Feng, Y., & Zhang, J. (2026). ARGUS: Defending LLM agents against context-aware prompt injection. arXiv preprint arXiv:2605.03378. https://doi.org/10.48550/arxiv.2605.03378
- Wojdynski, B. W. (2016). The deceptiveness of sponsored news articles: How readers recognize and perceive native advertising. American Behavioral Scientist, 60(12), 1475–1491.
- Wojdynski, B. W., Evans, N. J., & Hoy, M. G. (2018). Measuring sponsorship transparency in the age of native advertising. Journal of Consumer Affairs, 52(1), 115–137.
- Yusifov, T., Aghayev, A., & Hasanzada, L. (2026). Taxonomy of prompt injection attacks and analysis of defense mechanisms in large language model-based chatbots. InterConf, 69(295), 161–176. https://doi.org/10.51582/interconf.19-20.05.2026.017