User-controlled system prompts

User-controlled system prompts

Description

Placing user input or untrusted external data directly in high-priority prompt fields such as system, developer, or instructions may let an attacker redirect the model's behavior. This can undermine safety policies, expose sensitive information, or lead to misuse of privileged tools.

Potential impact

  • User input may override the model's intended policies or role.
  • Safeguards may be bypassed, internal instructions exposed, or sensitive operations misused.
  • Later tool calls or workflows may follow attacker-controlled instructions.

Remediation

  • Keep system prompts and developer instructions fixed. Pass user input only as ordinary messages or structured data.
  • Select behavior from a server-owned map of fixed prompts instead of constructing policies from user input.
  • Treat external documents and user content as untrusted data. Do not assume string sanitization or lower-priority message placement prevents prompt injection.
  • Authorize tool calls outside the model, with argument schemas, allow-lists, and least privilege.

Examples

Before

javascript
const express = require("express");
const OpenAI = require("openai");
const app = express();
app.use(express.json());
const client = new OpenAI();

app.post("/bad1", async (req, res) => {
  await client.chat.completions.create({
    model: "gpt-4o-mini",
    messages: [
      { role: "system", content: req.body.systemPrompt },
      { role: "user", content: "hello" }
    ]
  });
  res.send("ok");
});

After

javascript
const express = require("express");
const OpenAI = require("openai");
const app = express();
app.use(express.json());
const client = new OpenAI();

app.post("/good", async (req, res) => {
  if (typeof req.body.message !== "string") {
    return res.status(400).send("bad message");
  }
  await client.chat.completions.create({
    model: "gpt-4o-mini",
    messages: [
      { role: "system", content: "You are a careful support assistant." },
      { role: "user", content: req.body.message }
    ]
  });
  res.send("ok");
});

app.post("/good-mode", async (req, res) => {
  if (typeof req.body.message !== "string") {
    return res.status(400).send("bad message");
  }
  const SYSTEM_PROMPTS = {
    support: "You are a careful support assistant.",
    sales: "You are a careful sales assistant.",
  };
  const mode = String(req.body.mode || "support");
  if (!Object.hasOwn(SYSTEM_PROMPTS, mode)) {
    return res.status(400).send("bad mode");
  }

  await client.responses.create({
    model: "gpt-4o-mini",
    instructions: SYSTEM_PROMPTS[mode],
    input: req.body.message
  });
  res.send("ok");
});

Explanation:

  • Before: User input goes directly into the system message and may redefine the model's policy.
  • After: System prompts are fixed strings or selected from a server-owned map. Validate that user input is a string, then pass it as an ordinary user message or text input to Responses. Reject message arrays that could specify privileged roles.

References