Using Retry Hints for Load Shedding Permalink to this section
Part of Event ID & Retry Mechanism Design, under SSE Protocol Fundamentals & Architecture.
When an SSE service is overloaded — after an outage, during a deploy, when a popular event starts — the instinct from request/response systems is to return 503 Service Unavailable and let clients back off. With EventSource that instinct backfires: a 503 does not make the browser retry later, it makes it stop forever. The retry: field is the tool the protocol provides for pacing reconnects, and used deliberately it turns a reconnect storm into a smooth ramp. This guide covers the browser’s actual behaviour and a shedding strategy built on it.
Symptom & Developer Intent Permalink to this section
- After an outage, the service comes back and immediately falls over again under the reconnect wave.
- The team added 503 responses under load, and users now have to reload the page to get live updates back.
- Every client reconnects at the same moment because they all received the same
retry:value. - A deploy causes a load spike on the remaining nodes as every drained client arrives at once.
The intent is to reduce connection pressure when overloaded while guaranteeing that every client eventually reconnects on its own, with arrivals spread over time.
Root Cause Analysis Permalink to this section
EventSource reconnects in exactly two situations: a network error, and a successful (200 with text/event-stream) response that ends. Any other outcome — a 4xx or 5xx status, a wrong content type — “fails the connection”: readyState becomes CLOSED, an error event fires, and the browser never tries again unless application code creates a new EventSource.
The reconnect delay is the last retry: value the stream sent, or a browser default of a few seconds if none was sent. Browsers may add their own backoff, but you cannot rely on it. If every client holds the same value and loses its connection at the same moment, they return in the same instant.
Step-by-Step Resolution Permalink to this section
Step 1 — Send a jittered retry value on every stream Permalink to this section
function retryMs(base = 3000, spread = 7000) {
return base + Math.floor(Math.random() * spread); // 3–10 s, different per client
}
res.write(`retry: ${retryMs()}\n\n`);
Randomising per connection spreads any mass reconnect — an outage, a network blip at a large office — over the whole interval instead of a single instant.
Step 2 — Shed new connections softly: 200, a long retry, then close Permalink to this section
When a node is over its connection budget, accept the request just enough to tell the client when to come back:
app.get('/api/stream', (req, res) => {
if (openStreams.size >= MAX_STREAMS || overloaded()) {
res.writeHead(200, { 'Content-Type': 'text/event-stream', 'Cache-Control': 'no-cache' });
res.end(`retry: ${retryMs(20_000, 40_000)}\n: shedding load\n\n`); // come back in 20–60 s
metrics.shed.inc();
return;
}
// …normal stream…
});
This is a successful, ended stream, so the browser schedules a reconnect after the new retry value. The client stays live-capable, and arrivals are spread over forty seconds.
Step 3 — Handle genuine errors in application code Permalink to this section
For cases where a non-200 is correct — 401 when a session expired, 403 when access was revoked — handle the permanent failure explicitly:
es.onerror = () => {
if (es.readyState !== EventSource.CLOSED) return; // browser is retrying by itself
// Permanent failure: find out why, then decide whether to reopen.
fetch('/api/stream/status', { credentials: 'include' }).then((r) => {
if (r.status === 401) return refreshSessionThen(openStream);
setTimeout(openStream, backoff.next()); // e.g. a proxy returned 502
});
};
This also protects against 502s and 504s from proxies, which your server did not send but which fail the connection all the same. The full client pattern is in distinguishing fatal from transient SSE errors.
Step 4 — Budget reconnects per node Permalink to this section
A node recovering from a restart receives a burst of reconnects, each of which may trigger an expensive replay. Rate-limit accepted connections per second and shed the excess softly:
const bucket = tokenBucket({ ratePerSec: 500, burst: 1000 });
if (!bucket.take()) return softReject(res, retryMs(5_000, 25_000));
With 20,000 clients and a budget of 500 per second, the node fills over about 40 seconds instead of receiving 20,000 connections in two.
Pair the admission budget with cheap connection setup. The first moments of a stream are usually its most expensive — authentication, a replay query, a snapshot — so a reconnect wave concentrates exactly that work. Cache snapshots for a second or two so simultaneous reconnects share one computation, bound replay size per connection, and make authentication a local signature check rather than a database lookup where possible. Each of these lowers the cost per admitted connection and lets the budget be set higher.
Step 5 — Raise retry values during incidents Permalink to this section
A central setting can raise the base retry value for all streams during an incident, sent in a comment-plus-retry frame to existing connections:
function broadcastRetry(ms) {
for (const res of openStreams) res.write(`retry: ${ms + Math.floor(Math.random() * ms)}\n\n`);
}
Existing streams are unaffected until they drop; if they do, they come back slowly.
Validation & Monitoring Permalink to this section
# Soft rejection returns 200 with a long retry and ends immediately.
curl -si http://localhost:3000/api/stream | head -8
# The browser must reconnect later on its own: watch the Network panel for a second request
# after roughly the retry interval, with no page reload.
Track accepted and shed connections per second, the retry values being sent, and the time for connection counts to recover after an incident. An incident review should be able to show that no client needed a reload to get live updates back.
Production Checklist Permalink to this section
Frequently Asked Questions Permalink to this section
Does EventSource retry after a 503?
No. Any non-200 status fails the connection permanently. The browser only reconnects after network errors or after a successful stream ends.
What is the default retry interval?
It is browser-defined, typically a few seconds. Always send your own value so behaviour is predictable across browsers.
Can the server tell a client to stop reconnecting?
Yes: respond with 204 No Content, which the specification defines as a signal to stop. Use it for finished streams, not for load shedding.
Should the retry value change over the life of a stream?
It can. A new retry field takes effect for the next reconnect, which makes it a lightweight control channel for raising delays during incidents.