Incomplete URL substring validation

Incomplete URL substring validation

Description

Checking only whether a URL contains an allowed domain with includes, indexOf or a broad endsWith comparison can be bypassed. An attacker can place the expected text in user information, a path, a query or a hostname with a different domain. Examples include https://example.com.attacker.io, http://example.com@attacker.io and https://attacker.io/?next=https://example.com. Depending on the operation, this can enable an open redirect or SSRF.

Potential impact

  • Open redirects may support phishing or bypass OAuth redirect-URI checks, exposing sessions or tokens.
  • Server-side requests may reach internal APIs or metadata endpoints such as 169.254.169.254, enabling information disclosure or port scanning.
  • A server may access sensitive resources and return their contents externally.
  • Redirect or proxy policies intended to restrict domains may be bypassed.

Remediation

  • Parse the URL and compare new URL(url).hostname with an explicit allow-list.
  • To allow subdomains, check host === 'example.com' || host.endsWith('.example.com'). Permit that scope only if every subdomain is trusted. A bare endsWith('example.com') also accepts domains such as evil-example.com.
  • Validate the scheme, such as https:, and restrict ports where required.
  • Prefer selection from server-defined internal paths. Reject network-path references such as //external-host.
  • For outbound server requests, restrict the actual IPv4 and IPv6 destination addresses, including loopback, private, link-local and metadata ranges. Prevent bypasses through DNS changes or redirects, and bound request time and response size.
  • Do not base authorization on substring checks or incomplete regular expressions.

Examples

The example checks the HTTPS scheme and domain boundary for redirects. It assumes all subdomains of example.com are trusted. Apply port policy and DNS/IP controls for server requests separately.

Before

javascript
const express = require('express');
const app = express();

// BAD: Decide whether to allow a URL using a substring
app.get('/go', (req, res) => {
  const dest = req.query.dest || '/';
  // Allow any string containing "example.com" (unsafe)
  if (dest.includes('example.com')) {
    return res.redirect(dest);
  }
  res.redirect('/');
});

app.listen(3000);

After

javascript
const express = require('express');
const app = express();

// Parse the URL, check hostname and label boundaries, and validate the scheme
const ALLOWED_HOST = 'example.com';

app.get('/go', (req, res) => {
  const dest = req.query.dest;
  if (!dest) return res.redirect('/');

  let parsed;
  try {
    parsed = new URL(dest);
  } catch (e) {
    return res.redirect('/'); // Reject invalid URLs
  }

  const host = parsed.hostname.toLowerCase();
  const protocolOk = parsed.protocol === 'https:'; // Require HTTPS here
  const hostOk = host === ALLOWED_HOST || host.endsWith('.' + ALLOWED_HOST);

  if (protocolOk && hostOk) {
    return res.redirect(parsed.toString());
  }
  return res.redirect('/');
});

app.listen(3000);

Explanation:

  • Before: The expected text can appear inside an attacker-controlled hostname, user information, path or query. This route may redirect externally; the same validation in a server-request flow may permit SSRF.
  • After: The parsed hostname must be the allowed domain or a subdomain at a label boundary, and the scheme must be https:. Other inputs fall back to the fixed internal path.

References