Character Encoding and BOM Handling in SSE Permalink to this section
Part of Understanding the Event Stream Format, under SSE Protocol Fundamentals & Architecture.
Server-Sent Events settle the encoding question once: the stream is UTF-8, always. That removes a whole class of charset negotiation bugs, and introduces a few of its own. Servers that emit Latin-1 or Windows-1252 text produce garbled characters. A byte order mark at the wrong place corrupts the first field name. Custom parsers that decode each network chunk separately mangle any multi-byte character that happens to straddle a chunk boundary. This guide explains the rules and fixes each failure.
Symptom & Developer Intent Permalink to this section
- Accented letters appear as
éor�in the browser. - The very first event of a stream is never dispatched, or its first field is ignored.
- A custom fetch-based parser occasionally shows
�in the middle of otherwise correct text, at random positions. - Emoji or CJK text displays correctly in
EventSourcebut not in a Node.js or Python consumer. - Setting
charset=iso-8859-1in the Content-Type has no effect.
The intent is a stream that carries any Unicode text correctly to every client — browsers and custom parsers alike.
Root Cause Analysis Permalink to this section
The specification tells user agents to decode the stream with UTF-8 decode, which ignores any charset parameter on the Content-Type and strips a single leading byte order mark (U+FEFF, bytes EF BB BF) at the very start of the stream. Consequences:
- Text encoded in anything other than UTF-8 is misinterpreted. Latin-1
é(0xE9) is not valid UTF-8 on its own and becomes U+FFFD; double-encoded UTF-8 shows asé. - A BOM anywhere other than the start of the stream is just a character. If a server writes a BOM at the start of each event, the second and later events begin with U+FEFF, so their first field name is
datarather thandata, and the field is ignored. - A custom parser that decodes each chunk with a fresh, non-streaming decoder splits a multi-byte character when the chunk boundary falls inside it, producing replacement characters.
Step-by-Step Resolution Permalink to this section
Step 1 — Emit UTF-8, and say so Permalink to this section
res.writeHead(200, { 'Content-Type': 'text/event-stream; charset=utf-8' });
res.write(`data: ${JSON.stringify({ name: 'Zoë', city: '東京' })}\n\n`); // Node writes strings as UTF-8
The charset=utf-8 parameter is ignored by EventSource but helps proxies, debugging tools and non-browser clients that do look at it. In languages where the default encoding is platform-dependent, encode explicitly:
// Java servlet: never rely on the platform default.
response.setContentType("text/event-stream");
response.setCharacterEncoding("UTF-8");
# Python: yield str from an ASGI generator (encoded as UTF-8 by the framework), or bytes you encoded yourself.
yield f"data: {json.dumps(obj, ensure_ascii=False)}\n\n".encode("utf-8")
ensure_ascii=False sends characters as UTF-8 rather than \uXXXX escapes. Both decode correctly; escapes are larger but immune to encoding mistakes further down the pipeline.
Step 2 — Never emit a BOM except, optionally, once Permalink to this section
Most servers never produce a BOM. If a template, file or library prepends one, make sure it occurs at most once, at byte zero. The safest rule is never:
const BOM = '';
function frame(s) {
if (s.startsWith(BOM)) s = s.slice(1); // strip accidental BOMs from templated fragments
return s;
}
Step 3 — Decode with a streaming decoder in custom parsers Permalink to this section
// Correct: one TextDecoder in streaming mode for the whole response.
const decoder = new TextDecoder('utf-8');
let buffer = '';
for await (const chunk of response.body) {
buffer += decoder.decode(chunk, { stream: true }); // holds partial characters until complete
buffer = parseCompleteLines(buffer);
}
buffer += decoder.decode(); // flush at the end
{ stream: true } tells the decoder to keep an incomplete trailing byte sequence and prepend it to the next chunk. response.body.pipeThrough(new TextDecoderStream()) does the same. The equivalent in Python is an incremental decoder:
import codecs
dec = codecs.getincrementaldecoder("utf-8")(errors="replace")
async for chunk in response.aiter_bytes():
text = dec.decode(chunk) # partial characters carried over
Custom parsers must also strip a leading BOM once, to match browser behaviour.
Step 4 — Treat ids and event names as UTF-8 strings too Permalink to this section
Event ids and names may contain any Unicode characters except line breaks (and NUL, for ids). They are compared as strings by addEventListener and echoed back byte-for-byte in Last-Event-ID. Non-ASCII ids work, but an HTTP header carrying non-ASCII text can be mangled by intermediaries that assume Latin-1. Keep ids ASCII — numbers, or base64url of anything richer — and names ASCII identifiers.
Step 5 — Watch for double encoding in the pipeline Permalink to this section
Double encoding happens when text that is already UTF-8 bytes is treated as Latin-1 characters and encoded to UTF-8 again. It typically creeps in at boundaries: a database connection configured with the wrong client encoding, a message broker client that decodes bytes with a platform default, or a template engine that receives bytes where it expected a string. The telltale sign is that every non-ASCII character becomes two or more characters starting with à or Â. Fix it at the boundary where bytes became text with the wrong decoder, not by adding a compensating decode at the end, which only moves the bug. Set explicit encodings on every connection string and client (client_encoding=UTF8 for PostgreSQL, charset=utf8mb4 for MySQL — the older utf8 there cannot store emoji), and treat any payload from a broker as bytes to be decoded exactly once, as UTF-8.
Validation & Monitoring Permalink to this section
# Check the bytes on the wire: é must be C3 A9, not E9; no EF BB BF except possibly at offset 0.
curl -sN http://localhost:3000/events | head -c 400 | xxd | head -20
Add a fixture to the parser’s tests that contains multi-byte characters and is split at every byte offset, asserting identical output — the chunk-boundary test from parsing SSE from a fetch ReadableStream.
Production Checklist Permalink to this section
Frequently Asked Questions Permalink to this section
Can an SSE stream use UTF-16 or Latin-1?
No. Browsers decode event streams as UTF-8 regardless of the declared charset. Anything else will be misread.
Should JSON in data fields escape non-ASCII characters?
Either works. Raw UTF-8 is smaller; \u escapes survive any encoding mistake in the pipeline. Raw UTF-8 is fine when the whole server path is known to be UTF-8.
Why does EventSource handle split characters but my parser does not?
The browser decodes the stream incrementally, carrying partial characters between chunks. A parser that decodes each chunk independently loses that state; use a streaming decoder.
Is a BOM at the start of the stream an error?
No. The specification says to ignore one leading BOM, so browsers accept it. A BOM anywhere else becomes part of the text and breaks the field it prefixes.