Author: Tara Sarvnaz Raissi, Senior Legal Counsel (Ontario, Western & Atlantic Canada), Beneva
Introduction
While generative AI systems are quickly evolving, they remain susceptible to a variety of attacks that cause them to deviate from user expectations.[1] Among the most concerning of these are prompt injections. This term refers to attacks that “inject” malicious instructions into prompts the system processes with the goal of manipulating its behaviour.[2] This vulnerability is largely due to generative AI’s inability to reliably distinguish between trustworthy instructions and harmful input. Generative AI models interpret user prompts to create output, such as text and images based on patterns detected from large amounts of data. Without security protections that prevent adversarial commands from infiltrating the AI system, attackers can exploit its natural language processing capabilities to bypass safety protocols and cause privacy harms. While this problem is not fully solvable, organizations using generative AI can take steps to mitigate its risk. Implementing robust data security safeguards within the organization, incorporating AI specific language in supplier agreements and providing comprehensive AI user training can improve the integrity and security of generative AI systems across an enterprise.
How Prompt Injections Work
A prompt is an instruction or input given to a generative AI model to guide its output. Prompt injections can be direct or indirect, depending on whether the malicious instructions are explicit in the prompt given to the system or hidden in data that the system analyzes.
A direct prompt injection occurs when a harmful command is given directly to an AI system, tricking it into circumventing previous instructions and overriding built in security controls. For example, in February 2023, a Stanford University student bypassed the security safeguards of Microsoft’s AI powered Bing Chatbot, later rebranded as Microsoft Co-pilot, by posing as a developer and instructing the system to ignore its previous directives and reveal its internal guidelines. As a result, Bing Chat revealed its internal alias, “Sydney”, along with other confidential internal rules that governed its operation.
An indirect prompt injection takes place when adversarial instructions hidden in AI models’ external data sources alter the system’s behaviour. Unlike a direct prompt injection, this type of threat does not require direct access to the generative AI system and can be embedded in webpages, emails or other data sources to which the system has access. For example, hidden text may be added to a resume that is processed by a generative AI system to review candidate applications. This text may be invisible to the human eye but can trigger the AI system to bypass its automatic screening protocol and find an otherwise unqualified applicant suitable for a role.
Mitigating Strategies
The list below - though not exhaustive - identifies three steps that can help organizations mitigate risks posed by prompt injections.
- Data security starts within the organization. Internal practices such as good data hygiene (organizing, protecting and limiting access to information) and strong data controls are essential when implementing AI-powered tools and systems. Excess access to data increases the likelihood of privacy incidents, particularly when it goes unnoticed. Users, whether human or AI, should only have access to the data and tools that are absolutely necessary to perform their tasks.[3]
- Supplier contracts must incorporate AI specific language and protections. Vendor agreements should include governance and privacy provisions that guarantee system performance and establish protocols for managing fallout from the manipulation or misuse of an AI system. In addition to enforceable provisions protecting against adversarial inputs, vendors must undertake to promptly notify the organization about attacks or incidents that alter the operation of an AI model. Measures must be taken to mitigate or remedy the effects of any such incidents, including regular security audits. Contracts should clearly define “customer data”, specify how the vendor may access it, and exercise limits on its use for model training or product development in order to bolster the security of organizational data.
- AI-specific training flags the security risks posed by generative AI systems. Training on AI best practices will improve users’ understanding of the system’s vulnerabilities and ways in which prompts can be disguised or infiltrate the system. This can improve interactions with AI tools by emphasizing the importance of protecting sensitive or confidential data, being aware of and avoiding “suggested prompts” and monitoring for unexpected or irregular system behaviour.
The risks posed by prompt injections highlight the importance of responsible implementation of generative AI systems. Strong controls, informed oversight and a good understanding of this technology’s limitations can help organizations protect their AI tools and systems.
[1] Shi, Chongyang, et al. Lessons from Defending Gemini Against Indirect Prompt Injections. Google DeepMind, 2025.
[2] Sutton, Matt, and Damian Ruck. Indirect Prompt Injection: Generative AI’s Greatest Security Flaw. Centre for the Study of Existential Risk – Turing CETAS, 1 Nov. 2024, https://cetas.turing.ac.uk/publications/indirect-prompt-injection-generative-ais-greatest-security-flaw.
[3] Canadian Centre for Cyber Security. Top 10 IT Security Actions: No. 3 – Managing and Controlling Administrative Privileges. Government of Canada, July 2022, https://www.cyber.gc.ca/en/guidance/top-10-it-security-actions-no-3-managing-and-controlling-administrative-privileges-itsm10094