Defending SSE Endpoints Against Connection Exhaustion Permalink to this section

Part of Security Headers for Event Streams, under SSE Protocol Fundamentals & Architecture.

A Server-Sent Events endpoint is, by design, an invitation to hold a connection open indefinitely. That makes it an efficient target for resource exhaustion: an attacker does not need bandwidth or request volume, only many idle connections, each of which costs the server a file descriptor, a socket buffer and some memory. A few thousand held connections can exhaust a small service; tens of thousands can exhaust a large one. This guide layers defences so that holding connections is expensive for the attacker and cheap to refuse for you.

Symptom & Developer Intent Permalink to this section

  • The service stops accepting connections with “too many open files”, while request rates look normal.
  • A handful of IP addresses hold thousands of streams each.
  • Anonymous clients can open streams on endpoints that were meant for signed-in users.
  • Legitimate users cannot connect during an incident, because the connection budget is full of junk.
  • Rejecting abusive connections is itself expensive, because authentication runs a database query first.

The intent is a service where no single client, address or account can hold more than its share of connections, unauthenticated connections are refused before they cost anything, and legitimate traffic keeps a guaranteed budget.

Root Cause Analysis Permalink to this section

The cost of an SSE connection to the server is almost entirely in holding it, not in serving it. Defences that look at request rate miss the attack, because an attacker can open connections slowly and simply never close them. The effective controls count concurrent connections, at several scopes, and reject early.

Layers of defence against connection holding Layers from the network edge through the reverse proxy, authentication, per-identity limits and global admission, each rejecting a different class of abusive connection. Layers of defence against connection holding Edge / CDN / WAF per-IP concurrency, bot rules Reverse proxy limit_conn per IP, total Authentication before any stream work Per-identity caps per user, per tenant Global admission node budget, soft shedding
Each layer rejects cheaply what it can see. The edge sees addresses, the application sees identities, and admission control protects the total.

Step-by-Step Resolution Permalink to this section

Step 1 — Require authentication before streaming, and make it cheap Permalink to this section

app.get('/api/stream', (req, res) => {
  const claims = verifySignedToken(req.cookies.stream_token);    // local HMAC/JWT check, no I/O
  if (!claims) return res.status(401).end();                     // EventSource stops on 401
  // …only now allocate subscriptions, buffers, timers…
});

A local signature check costs microseconds. A session lookup in a database on every connection attempt turns a connection flood into a database flood; if one is needed, cache it.

Step 2 — Cap concurrent connections per IP at the proxy Permalink to this section

limit_conn_zone $binary_remote_addr zone=sse_per_ip:10m;
limit_conn_zone $server_name        zone=sse_total:1m;

location /api/stream {
    limit_conn sse_per_ip 20;          # generous for offices behind NAT, fatal for single-host floods
    limit_conn sse_total 50000;        # global ceiling for this proxy
    limit_conn_status 429;
    proxy_pass http://sse_nodes;
    proxy_buffering off;
    proxy_read_timeout 1h;
}

Set the per-IP value from real data: large offices and mobile carriers put many users behind one address. Twenty to a hundred is typical. Where a CDN or WAF fronts the service, apply the per-IP limit there instead, so floods never reach your proxies.

Step 3 — Cap concurrent streams per user and per tenant Permalink to this section

const perUser = new Map();
function admitUser(userId, max = 10) {
  const n = perUser.get(userId) ?? 0;
  if (n >= max) return false;
  perUser.set(userId, n + 1);
  return true;
}
// On rejection: 429 for abuse-level counts; for a user who merely opened many tabs,
// close their oldest stream instead of refusing the newest.

In a multi-node deployment, keep counts in Redis with expiring keys, refreshed by each stream’s heartbeat, so a crashed node’s counts expire. Limiting SSE connections per user has the distributed implementation.

Server cost per rejected connection attempt Bar chart comparing CPU time spent per rejected connection attempt when rejecting at the edge, at the proxy, after a local token check, and after a database session lookup. Server cost per rejected connection attempt Edge / WAF ~0 (not your servers) Proxy limit_conn ~15 µs Local token check ~40 µs Database session lookup ~2,500 µs microseconds of server CPU per rejected attempt (log-like spread)
Reject as early and as cheaply as possible. A flood of attempts that each trigger a database query is an attack on the database.

Step 4 — Reclaim idle and dead connections Permalink to this section

Every stream should have a heartbeat, and every hop should have a finite idle timeout the heartbeat beats. Heartbeats expose clients that vanished, so their connections are reclaimed. For attackers who read the heartbeats and hold the connection, a maximum stream age (with a jittered reconnect) forces them to re-authenticate and re-pass every check periodically.

Also watch for clients that never read: a connection whose socket buffer stays full is costing memory. Close streams whose writes have been blocked beyond a deadline.

Step 5 — Reserve capacity for legitimate traffic Permalink to this section

Global admission control protects the node: when open streams approach the node’s budget, shed new connections softly (a 200 with a long, jittered retry: and an immediate end, as in using retry hints for load shedding). Prioritise by identity: authenticated paying tenants first, free-tier next, anonymous public streams last.

function admitGlobal(priority) {
  const used = openStreams.size / NODE_BUDGET;
  if (priority === 'anon' && used > 0.70) return false;
  if (priority === 'free' && used > 0.85) return false;
  return used < 0.98;
}
A connection flood absorbed by layered limits Timeline of ten minutes during a connection flood, showing the edge limiting per-IP concurrency, the proxy capping totals, and legitimate users remaining connected throughout. A connection flood absorbed by layered limits Attack traffic Edge limits Legit streams flood of held connections per-IP caps active unaffected 0 2 4 6 8 10 minutes flood starts flood ends
Legitimate streams never lose their reserved share. Attack connections are refused at the first layer that can identify them.

Finally, raise the operating-system ceilings deliberately rather than leaving them as the de facto limit. When a process runs out of file descriptors, it fails in unpredictable places — log files, database connections, health checks — rather than cleanly refusing new streams. Set the application’s own admission budget comfortably below the descriptor limit, so the application sheds load while it still has room to do so gracefully.

Validation & Monitoring Permalink to this section

# Per-IP cap: the 21st concurrent stream from one address should get 429.
for i in $(seq 1 21); do curl -s -o /dev/null -w '%{http_code}\n' -N --max-time 3 https://app.example.com/api/stream -b auth.txt & done; wait

# Unauthenticated attempts must be refused before any stream work.
curl -si https://app.example.com/api/stream | head -1        # HTTP/2 401

Dashboards should show open streams by scope — per node, top IPs, top users, top tenants — and rejections by layer and reason. Alert on open streams approaching the node budget, and on any single identity holding an unusual share.

Production Checklist Permalink to this section

Frequently Asked Questions Permalink to this section

Are request-rate limits enough to protect an SSE endpoint?

No. An attacker can open connections slowly and hold them. Limits must count concurrent connections, not requests per second.

Should rejected clients receive 429 or 503?

For abusive or over-quota clients, 429 is accurate and makes EventSource stop. For legitimate clients during overload, a 200 with a long retry hint keeps them reconnecting on their own.

How many streams should one user be allowed?

Enough for normal multi-tab and multi-device use — typically five to ten — with the oldest stream closed rather than the newest refused when a real user exceeds it.

Do HTTP/2 connections change the picture?

Each HTTP/2 connection can carry many streams, so limit concurrent streams per connection as well as connections per IP; servers advertise the former with SETTINGS_MAX_CONCURRENT_STREAMS.