Distinguishing Fatal from Transient SSE Errors Permalink to this section
Part of Error Handling & Reconnection UX, under Frontend Consumption & Client Patterns.
EventSource reports every problem the same way: an error event with no details. A Wi-Fi blip that the browser will fix in three seconds, an expired session that needs a login, a deleted resource that will never come back and an overloaded server that needs a minute to recover all look identical to the onerror handler. Treating them the same is how applications end up either retrying forever against a 401 or giving up after a harmless network drop. This guide classifies SSE errors, shows how to get the information EventSource hides, and maps each class to the right reaction.
Symptom & Developer Intent Permalink to this section
- After a session expires, the live view silently stops updating and never recovers.
- A “Connection lost” banner flashes on every brief network hiccup.
- A deleted document keeps reconnecting in a tight loop in the background.
- During an outage, every client retries at once when the service returns.
- Error telemetry shows thousands of
errorevents with no way to tell what they mean.
The intent is a client that recovers silently from transient problems, stops or re-authenticates on permanent ones, backs off on overload, and reports each class distinctly.
Root Cause Analysis Permalink to this section
The browser’s behaviour depends on what went wrong, and that behaviour is the only signal available:
- Transient. A network error or a stream that ended normally. The browser sets
readyStatetoCONNECTINGand reconnects after the retry interval. The application need do nothing. - Permanent, as far as the browser knows. A non-200 status or a wrong content type. The browser sets
readyStatetoCLOSEDand never retries. The status is not exposed.
So the first question is always the readyState inside the error handler. For CLOSED, the application must find out why, because the right reaction depends on the status: 401 means re-authenticate, 403/404/410 mean stop, 429/5xx or a proxy’s 502 mean back off and reopen.
Step-by-Step Resolution Permalink to this section
Step 1 — Classify with readyState and a debounce Permalink to this section
es.onerror = () => {
if (es.readyState === EventSource.CONNECTING) {
scheduleStatus('reconnecting', 2000); // show only if it lasts more than 2 s
return; // the browser is handling it
}
onPermanentFailure(); // CLOSED: find out why
};
es.onopen = () => { cancelScheduledStatus(); setStatus('live'); };
Debouncing the visible status avoids flashing warnings for the many sub-second reconnects that users never notice. The status UX itself is covered in showing connection status in the UI.
Step 2 — Probe the status the browser hid Permalink to this section
async function onPermanentFailure() {
let status = 0;
try {
// A cheap HEAD or a dedicated endpoint with the same auth and authorisation checks as the stream.
const r = await fetch(streamUrl, { method: 'HEAD', credentials: 'include', cache: 'no-store' });
status = r.status;
} catch { status = 0; } // network down: treat as transient
react(status);
}
Make sure the stream route answers HEAD (or add a small /stream/status endpoint) and applies exactly the same authentication and authorisation, so the probe’s status is the status the stream would return.
Step 3 — React per class Permalink to this section
function react(status) {
switch (true) {
case status === 401:
telemetry.count('sse_error', { cls: 'auth' });
return auth.refresh().then(reopen, () => setStatus('signed-out'));
case status === 403 || status === 404 || status === 410:
telemetry.count('sse_error', { cls: 'fatal', status });
return setStatus('unavailable'); // stop: retrying cannot help
case status === 429 || status >= 500 || status === 0:
telemetry.count('sse_error', { cls: 'overload', status });
return setTimeout(reopen, backoff.next()); // jittered exponential backoff
default:
telemetry.count('sse_error', { cls: 'unknown', status });
return setTimeout(reopen, backoff.next());
}
}
Reset the backoff after a stream has been open and healthy for a while, not merely on open, so a server that accepts connections and immediately drops them does not reset the delay each time. Exponential backoff with jitter for SSE reconnects covers the backoff itself.
Step 4 — Let the server help the classification Permalink to this section
Servers can make classification easier: return accurate statuses before streaming; shed load with a 200 plus a long retry: rather than a 503, so overload becomes a transient case the browser handles itself (see using retry hints for load shedding); and send an explicit event before ending a stream for a reason the client should know, such as event: access-revoked.
The words matter as much as the logic. For each class, decide in advance what the interface says: nothing at all for transient blips under a couple of seconds; a quiet “Reconnecting…” for longer ones and for server overload; a specific, actionable message for authentication (“Your session expired — sign in to resume live updates”); and a plain statement for permanent loss (“This document was deleted”). Generic “Connection error” text for every class trains users to ignore it, which defeats the purpose on the rare occasions it matters.
Step 5 — Or use a fetch-based client that sees statuses directly Permalink to this section
If probing feels indirect, a fetch-based SSE client receives the real status of every connection attempt and can classify without a second request, at the cost of implementing reconnection itself.
Validation & Monitoring Permalink to this section
Simulate each class against a test server: drop the connection (transient), return 401, 404 and 503, and put a proxy in front that returns 502. Assert the client’s final state and the number of reconnect attempts for each. In production, graph sse_error by class; a spike in auth at a round interval points at token lifetimes, a spike in fatal after a deploy points at a routing change, and overload should correlate with server incidents.
Production Checklist Permalink to this section
Frequently Asked Questions Permalink to this section
Why does EventSource not expose the HTTP status?
The API was designed to be simple and to avoid leaking cross-origin information. The only signal it gives is readyState, which distinguishes retrying from given up.
Is a 204 response an error?
No. 204 No Content tells EventSource to stop reconnecting cleanly. Servers use it to end finished streams; the client should treat it as completion, not failure.
Should the client ever give up on transient errors?
Not automatically while the page is open; networks come back. Cap the backoff delay, and pause attempts while the page is hidden or offline, resuming on visibility and online events.
How do I test proxy errors locally?
Put nginx or a fault-injection proxy in front of the test server and return 502 or 504 from it. EventSource treats them like any other non-200 status.