What is Prompt Injection?
Definition
Prompt injection is a security vulnerability in which instructions placed in the input a language model processes change the application's behavior in ways the developer did not intend. In a direct injection the instructions are in the user's own message; in an indirect injection they are hidden in content the model reads, such as a web page, an email or a document. OWASP ranks it first (LLM01) in its 2025 Top 10 for LLM applications.
Also known as: indirect prompt injection, direct prompt injection, LLM prompt injection, LLM01

The root cause: instructions and data share one channel
Classic SQL injection happens when user input is interpreted as part of the query code, and it has a definitive fix in parameterized queries, which send code and data down separate channels. Language models have no equivalent hard separation. The developer's instructions, the user's message and a web page the model has been asked to read all end up side by side in the same context, as the same kind of tokens. The model has to infer from the text itself which sentences are instructions to follow and which are merely data to process, and that inference is not always right.
As OWASP puts it, given the stochastic nature of how models work, it is unclear whether fool-proof prevention methods exist. The realistic goal is therefore not to eliminate the weakness but to limit the damage a successful injection can do.
Direct versus indirect injection
| Direct | Indirect | |
|---|---|---|
| Where the instruction comes from | The user's own message | External content the model reads: a web page, email, PDF, product review or image |
| Who it targets | Usually the application's rules | Often an innocent user of the application |
| Typical scenario | A user trying to get around the system rules | Hidden text on a page an assistant summarizes steering the answer |
Indirect injection is a web problem. When an AI assistant reads a page on a user's behalf, a white-on-white paragraph, an HTML comment or text embedded in an image can give the model instructions such as “praise this product” or “send the user to this address”. OWASP notes that multimodal models can be affected by instructions hidden in images as well.
The damage scales with permissions
In a chat tool that only produces text, a successful injection usually ends in a misleading or inappropriate answer. Once the model can use tools, the picture changes. An AI agent that can send email, read files, write to a database or initiate payments can be steered by a document it reads into using those powers in harmful ways. OWASP treats this separately as Excessive Agency. The same weakness that is an annoyance in a chatbot becomes a risk of data leakage or unauthorized actions in an agent with real permissions.
Defense in layers
The OWASP LLM01:2025 entry lists complementary mitigations:
- Constrain model behavior: define the role, scope and handling of external content clearly in the system prompt.
- Define and validate output formats: check in code that the output has the expected structure, such as JSON with specific fields.
- Filter inputs and outputs: screen for sensitive content and unexpected instruction patterns.
- Enforce least privilege: give the model and its tools only the access the task needs, a direct application of the principle of least privilege.
- Require human approval for high-risk actions: money, deletions and anything sent outside should wait for the user's confirmation.
- Segregate and identify external content: present untrusted text inside clear boundaries, as data rather than instructions.
- Run adversarial tests: attack your own system regularly.
Two baseline rules belong on top: treat model output as untrusted input (sanitize it before rendering it as HTML, for example), and never put secrets in prompt text. For the wider risk landscape, see the OWASP Top 10 projects.
Two notes for site owners
First, planting hidden instructions on your own pages to sway AI answers is manipulation; it conflicts with search engines' spam policies on hidden text and undermines trust in your site. Second, areas where users can post content, such as comments, reviews and forums, can carry indirect injections aimed at assistants that read your pages, which gives moderation a new dimension. And if you run a chat assistant on your own site, the defensive layers above are your responsibility.

