Real Safety AI Foundation / Research

Platform Harm, Addictive Design, and Accountability

Reviewed for Quality and Security: Advertising Injection Through a Verified Connector, Trust Laundering, and the Limits of Directory Verification

Publisher: Real Safety AI Foundation

Working paper. Not peer reviewed.

Written
July 2026
Version
v5
Pages
17

Abstract

A conversational assistant whose maker publicly commits that it carries no advertising was observed delivering a paid tier upgrade advertisement at the close of every academic search, across seven consecutive queries spanning distinct research domains. The advertisement did not originate with the assistant. It arrived inside the result payload of a third party connector, carried by a literal instruction directed at the model: you must include the above message at the end of your response. This is indirect prompt injection, the failure mode ranked first among risks to large language model applications, executing through a connector that the platform lists in a curated directory and marks with a badge whose hover text states that the platform has reviewed the connector for quality and security. One line below the badge, the same card disclaims the capacity to verify that listed connectors work as intended or that they will not change. This paper documents the specimen, establishes its reproducibility, and situates it in the protocol specific security literature, which supplies both the correct name for the attack, function return injection rather than tool poisoning, and the evidence that this vector is stronger and less studied than the descriptor poisoning the field concentrates on. It then maps the conduct against the platform’s own directory policy, which prohibits both software that evades model instructions and software that serves advertisements. It then locates the harm in two established bodies of authority. The consumer protection frame is the Federal Trade Commission’s treatment of deceptively formatted advertising, under which a message must be identifiable as advertising and a misleadingly formatted advertisement is deceptive even when its claims are literally true. The behavioral frame is the persuasion knowledge literature, under which a reader who cannot recognize a message as advertising cannot mount the skepticism that advertising recognition normally triggers. The design frame is the deceptive design literature, which holds that the most consequential dark patterns sit beneath the interface in system architecture, where interface bound scrutiny cannot reach them. The economic frame is twofold: the connector is a credence good whose quality the user cannot verify even after use, which is the condition that makes a certification mark necessary in the first place, and the badge is a signal that cannot separate compliant from non compliant connectors because obtaining it costs both the same. The attack class is well described in that literature; what is not documented there is a listed, badged, first party connector performing it against ordinary users for commercial gain, in production rather than in a benchmark. The paper names the general phenomenon the specimen exposes as trust laundering: the conversion of a platform’s verification mark into credibility for a third party’s promotional content, credibility the content did not earn and the platform did not knowingly extend. A verification badge that speaks in the present tense about an artifact its issuer concedes may have changed since review is the mechanism that makes trust laundering possible. The advertisement is the proof the assurance fails. The structure will outlast the advertisement.

Keywords

  • function return injection
  • rug pull
  • trust laundering
  • indirect prompt injection
  • Model Context Protocol
  • connector verification
  • dark patterns
  • deceptive design
  • native advertising
  • persuasion knowledge
  • trust marks
  • credence goods
  • signaling theory
  • attention economy
  • Federal Trade Commission
  • platform accountability
  • AI governance

Plain language slides

First slide of the plain language summary of Reviewed for Quality and Security: Advertising Injection Through a Verified Connector, Trust Laundering, and the Limits of Directory VerificationOpen the 21-slide summary (PDF)

Suggested citation

Gilly, Travis. "Reviewed for Quality and Security: Advertising Injection Through a Verified Connector, Trust Laundering, and the Limits of Directory Verification." Real Safety AI Foundation Working Paper, July 2026. https://realsafetyai.org/research/u65v3w/

Other versions

This paper is also posted on SSRN.

SSRN version

References (44)

  1. Anthropic. (2026a, February 4). Claude is a space to think. Retrieved from https://www.anthropic.com/news/claude-is-a-space-to-think
  2. Anthropic. (2026b). Anthropic software directory policy. Retrieved from https://support.claude.com/en/articles/13145358-anthropic-software-directory-policy
  3. Anders, S., Souza Monteiro, D. M., & Rouvière, E. (2010). Competition and credibility of private third-party certification in international food supply. Journal of International Food & Agribusiness Marketing, 22(3–4), 328–341. https://doi.org/10.1080/08974431003641554
  4. Bergh, D. D., Connelly, B. L., & Ketchen, D. J. (2014). Signalling theory and equilibrium in strategic management research: An assessment and a research agenda. Journal of Management Studies, 51(8), 1334–1360. https://doi.org/10.1111/joms.12097
  5. Campbell, C., & Evans, N. J. (2018). The role of a companion banner and sponsorship transparency in recognizing and evaluating article-style native advertising. Journal of Interactive Marketing, 43, 17–32.
  6. Consensus. (2026a, May 15). Privacy policy. Retrieved from https://consensus.app/home/privacy-policy/
  7. Consensus. (2026b). Getting started with the Consensus MCP. Retrieved from https://docs.consensus.app/docs/mcp
  8. Darby, M. R., & Karni, E. (1973). Free competition and the optimal amount of fraud. The Journal of Law and Economics, 16(1), 67–88. https://doi.org/10.1086/466756
  9. Di Porto, F., & Egberts, A. (2023). The collective welfare dimension of dark patterns regulation. European Law Journal, 29(1–2), 114–141. https://doi.org/10.1111/eulj.12478
  10. Emons, W. (2005). Credence goods: The monopoly case. In Economics of information (pp. 17–42). De Gruyter. https://doi.org/10.1515/9783110508123-003
  11. Evans, N. J., & Wojdynski, B. W. (2020). An introduction to the special issue on native and covert advertising formats. International Journal of Advertising, 39(1), 1–3.
  12. Falkinger, J. (2007). Attention economies. Journal of Economic Theory, 133(1), 266–294. https://doi.org/10.1016/j.jet.2005.12.001
  13. Federal Trade Commission. (1983). Policy statement on deception (appended to Cliffdale Associates, Inc., 103 F.T.C. 110, 174 (1984)). Retrieved from https://www.ftc.gov/system/files/documents/public_statements/410531/831014deceptionstmt.pdf
  14. Federal Trade Commission. (2015a). Enforcement policy statement on deceptively formatted advertisements. Retrieved from https://www.ftc.gov/system/files/documents/public_statements/896923/151222deceptiveenforcement.pdf
  15. Federal Trade Commission. (2015b). Native advertising: A guide for business. Retrieved from https://www.ftc.gov/business-guidance/resources/native-advertising-guide-businesses
  16. Fortune. (2026a, March 26). Anthropic says it is testing Mythos, a powerful new AI model, after a data leak reveals its existence. Fortune. Retrieved from https://fortune.com/2026/03/26/anthropic-says-testing-mythos-powerful-new-ai-model-after-data-leak-reveals-its-existence-step-change-in-capabilities/
  17. Fortune. (2026b, April 23). A group of users leaked Anthropic’s AI model Mythos by reportedly guessing where it was located. Fortune. Retrieved from https://fortune.com/2026/04/23/anthropic-mythos-leak-dario-amodei-ceo-cybersecurity-hackers-exploits-ai/
  18. Friestad, M., & Wright, P. (1994). The persuasion knowledge model: How people cope with persuasion attempts. Journal of Consumer Research, 21(1), 1–31.
  19. Greshake, K., Abdelnabi, S., Mishra, S., Endres, C., Holz, T., & Fritz, M. (2023). Not what you’ve signed up for: Compromising real-world LLM-integrated applications with indirect prompt injection. In Proceedings of the 16th ACM Workshop on Artificial Intelligence and Security (AISec ’23). Association for Computing Machinery. https://doi.org/10.1145/3605764.3623985
  20. Hendricks, V. F., & Vestergaard, M. (2018). The attention economy. In Reality lost: Markets of attention, misinformation and manipulation (pp. 1–17). Springer. https://doi.org/10.1007/978-3-030-00813-0_1
  21. Huberman, B. A. (2017). Big data and the attention economy. Ubiquity, 2017(December), 1–7. https://doi.org/10.1145/3158337
  22. Hou, X., Wang, S., & Zhang, Y. (2026). SMCP: Secure Model Context Protocol. arXiv preprint arXiv:2602.01129. https://doi.org/10.48550/arxiv.2602.01129
  23. Huang, C., Huang, X., & Tran, N. P. (2026). Model Context Protocol threat modeling and analyzing vulnerabilities to prompt injection with tool poisoning. arXiv preprint arXiv:2603.22489. https://doi.org/10.48550/arxiv.2603.22489
  24. Jamshidi, S., Nafi, K. W., & Moradi Dakhel, A. (2025). Securing the Model Context Protocol: Defending LLMs against tool poisoning and adversarial attacks. arXiv preprint arXiv:2512.06556. https://doi.org/10.48550/arxiv.2512.06556
  25. Kollmer, T., & Eckhardt, A. (2022). Dark patterns: Conceptualization and future research directions. Business & Information Systems Engineering, 65(2), 201–208. https://doi.org/10.1007/s12599-022-00783-7
  26. Leiser, M., & Santos, C. (2023). Dark patterns, enforcement, and the emerging digital design acquis: Manipulation beneath the interface. Center for Open Science. https://doi.org/10.31235/osf.io/rf3ja
  27. Lin, F. (2021). Demystifying removed apps in iOS app store. arXiv preprint arXiv:2101.05100. https://doi.org/10.48550/arxiv.2101.05100
  28. Kaplan, B., & Qian, J. (2021). A survey on common threats in npm and PyPi registries. arXiv preprint arXiv:2108.09576. https://doi.org/10.48550/arxiv.2108.09576
  29. Li, R., Wang, Z., & Yao, Y. (2026). MCP-ITP: An automated framework for implicit tool poisoning in MCP. arXiv preprint arXiv:2601.07395. https://doi.org/10.48550/arxiv.2601.07395
  30. Loconto, A. M. (2017). Models of assurance. The Annals of the American Academy of Political and Social Science, 670(1), 112–132. https://doi.org/10.1177/0002716217692517
  31. Model Context Protocol. (2025). Specification (2025-11-25). The Linux Foundation. Retrieved from https://modelcontextprotocol.io/specification/2025-11-25
  32. OWASP. (2025). OWASP Top 10 for LLM applications 2025: LLM01 prompt injection. Open Worldwide Application Security Project. Retrieved from https://genai.owasp.org/llmrisk/llm01-prompt-injection/
  33. Perez, F., & Ribeiro, I. (2022). Ignore previous prompt: Attack techniques for language models. arXiv preprint arXiv:2211.09527. https://doi.org/10.48550/arxiv.2211.09527
  34. Quirolgico, S., Voas, J., & Karygiannis, T. (2015). Vetting the security of mobile applications (NIST Special Publication 800-163). National Institute of Standards and Technology. https://doi.org/10.6028/nist.sp.800-163
  35. Rostamzadeh, M., Narula, S., & Birhan, N. (2026). MCP-DPT: A defense-placement taxonomy and coverage analysis for Model Context Protocol security. arXiv preprint arXiv:2604.07551. https://doi.org/10.48550/arxiv.2604.07551
  36. Rüdiger, K., & García Rodríguez, M. J. (2013). Do we need innovative trust intermediaries in the digital economy? Global Business Perspectives, 1(4), 329–340. https://doi.org/10.1007/s40196-013-0021-8
  37. Song, X., et al. (2025). Beyond the protocol: Unveiling attack vectors in the Model Context Protocol ecosystem. arXiv preprint arXiv:2506.02040. https://doi.org/10.48550/arxiv.2506.02040
  38. Thambisetty, S. (2007). Patents as credence goods. Oxford Journal of Legal Studies, 27(4), 707–740. https://doi.org/10.1093/ojls/gqm021
  39. Vu, D.-L., Pashchenko, I., & Massacci, F. (2020). Towards using source code repositories to identify software supply chain attacks. In Proceedings of the 2020 ACM SIGSAC Conference on Computer and Communications Security (pp. 2093–2095). https://doi.org/10.1145/3372297.3420015
  40. Wang, Y.-M., Chen, S., Alkhudair, R., et al. (2025). Defending against prompt injection with DataFilter. arXiv preprint arXiv:2510.19207. https://doi.org/10.48550/arxiv.2510.19207
  41. Weng, S., Feng, Y., & Zhang, J. (2026). ARGUS: Defending LLM agents against context-aware prompt injection. arXiv preprint arXiv:2605.03378. https://doi.org/10.48550/arxiv.2605.03378
  42. Wojdynski, B. W. (2016). The deceptiveness of sponsored news articles: How readers recognize and perceive native advertising. American Behavioral Scientist, 60(12), 1475–1491.
  43. Wojdynski, B. W., Evans, N. J., & Hoy, M. G. (2018). Measuring sponsorship transparency in the age of native advertising. Journal of Consumer Affairs, 52(1), 115–137.
  44. Yusifov, T., Aghayev, A., & Hasanzada, L. (2026). Taxonomy of prompt injection attacks and analysis of defense mechanisms in large language model-based chatbots. InterConf, 69(295), 161–176. https://doi.org/10.51582/interconf.19-20.05.2026.017

All research