Indirect prompt injection through remote content

Indirect prompt injection through remote content

Description

Indirect prompt injection occurs when an attacker places instructions in an external resource, such as a web page or document, that an application adds to the model's context. If the model fails to distinguish data from instructions, the inserted text may change the task or influence later actions.

The impact depends largely on the secrets, tools, and permissions available to the model and how its output is used. Consider who can change external content together with the model's authorization boundaries.

Potential impact

  • Manipulated summaries, classifications, or recommendations
  • Data leakage or unauthorized actions when the model can access secrets or external tools
  • Greater integrity or availability risks if model output drives commands, code, approvals, or security decisions

Remediation

No general-purpose string sanitizer reliably removes prompt injection in every context. Delimiters, role separation, encoding, RAG, fine-tuning, or instructions to ignore external directions are not complete defenses.

  1. Maintain trust boundaries. Track sources and content types, and limit unnecessary sources, sizes, and formats. A domain on an allow-list does not make its response trustworthy.
  2. Minimize access. Provide only necessary tools and data. Isolate the model from secrets and actions with side effects.
  3. Enforce authorization outside the model. Validate tool arguments and destinations with strict schemas and allow-lists in deterministic code. Require user confirmation for high-risk actions.
  4. Restrict and validate output. Prefer structured output or finite choices where possible, and validate them before further processing. A schema checks structure, not semantic safety; check values and authorization separately. Do not execute commands or make security decisions solely from model output.
  5. Evaluate attacks repeatedly. Reassess indirect injection scenarios when prompts or models change, and monitor inputs, outputs, and tool calls.

Examples

Before

javascript
const express = require("express");
const { fetch } = require("undici");
const { streamText } = require("ai");
const { openai } = require("@ai-sdk/openai");
const app = express();

app.post("/bad1", async (_req, res) => {
  const retrieved = await (
    await fetch("https://example.com/manual")
  ).text();
  await streamText({
    model: openai("gpt-5.6"),
    prompt: "Summarize this document:\n" + retrieved
  });
  res.send("ok");
});

After

javascript
const express = require("express");
const { streamText } = require("ai");
const { openai } = require("@ai-sdk/openai");
const app = express();

app.post("/good", async (_req, res) => {
  await streamText({
    model: openai("gpt-5.6"),
    prompt: "Summarize this internal FAQ: Support is available Monday through Friday."
  });
  res.send("ok");
});

Explanation:

  • Before: The remote document enters the model's input. If an attacker can change it, hidden instructions may influence the summary.
  • After: Only fixed, server-authored FAQ content is supplied. If remote content is required, use the layered controls above to limit its impact. Adding string delimiters alone does not provide equivalent protection.

References