Description
Indirect prompt injection occurs when an attacker places instructions in an external resource, such as a web page or document, that an application adds to the model's context. If the model fails to distinguish data from instructions, the inserted text may change the task or influence later actions.
The impact depends largely on the secrets, tools, and permissions available to the model and how its output is used. Consider who can change external content together with the model's authorization boundaries.
Potential impact
- Manipulated summaries, classifications, or recommendations
- Data leakage or unauthorized actions when the model can access secrets or external tools
- Greater integrity or availability risks if model output drives commands, code, approvals, or security decisions
Remediation
No general-purpose string sanitizer reliably removes prompt injection in every context. Delimiters, role separation, encoding, RAG, fine-tuning, or instructions to ignore external directions are not complete defenses.
- Maintain trust boundaries. Track sources and content types, and limit unnecessary sources, sizes, and formats. A domain on an allow-list does not make its response trustworthy.
- Minimize access. Provide only necessary tools and data. Isolate the model from secrets and actions with side effects.
- Enforce authorization outside the model. Validate tool arguments and destinations with strict schemas and allow-lists in deterministic code. Require user confirmation for high-risk actions.
- Restrict and validate output. Prefer structured output or finite choices where possible, and validate them before further processing. A schema checks structure, not semantic safety; check values and authorization separately. Do not execute commands or make security decisions solely from model output.
- Evaluate attacks repeatedly. Reassess indirect injection scenarios when prompts or models change, and monitor inputs, outputs, and tool calls.
Examples
Before
const express = require("express");
const { fetch } = require("undici");
const { streamText } = require("ai");
const { openai } = require("@ai-sdk/openai");
const app = express();
app.post("/bad1", async (_req, res) => {
const retrieved = await (
await fetch("https://example.com/manual")
).text();
await streamText({
model: openai("gpt-5.6"),
prompt: "Summarize this document:\n" + retrieved
});
res.send("ok");
});
After
const express = require("express");
const { streamText } = require("ai");
const { openai } = require("@ai-sdk/openai");
const app = express();
app.post("/good", async (_req, res) => {
await streamText({
model: openai("gpt-5.6"),
prompt: "Summarize this internal FAQ: Support is available Monday through Friday."
});
res.send("ok");
});
Explanation:
- Before: The remote document enters the model's input. If an attacker can change it, hidden instructions may influence the summary.
- After: Only fixed, server-authored FAQ content is supplied. If remote content is required, use the layered controls above to limit its impact. Adding string delimiters alone does not provide equivalent protection.
References
- OWASP Top 10 for LLM Applications 2026
- OWASP LLM01:2025 Prompt Injection (risk explanation)
- CWE-1427: Improper Neutralization of Input Used for LLM Prompting
- NIST AI 100-2 E2025: Adversarial Machine Learning
- OpenAI: Understanding prompt injections
- OpenAI: Designing AI agents to resist prompt injection
- Vercel AI SDK: Generating and Streaming Text
- OpenAI Node SDK
- Anthropic TypeScript SDK