Unit Testing Event Stream Serialization Permalink to this section
Part of Testing & Load Testing SSE Endpoints, under Backend Stream Generation & Connection Management.
Every SSE server has a function that turns an event into bytes, usually a template string written in five minutes: `id: ${id}\nevent: ${type}\ndata: ${data}\n\n`. It works for every payload the author thought of and silently corrupts the stream for the ones they did not — a newline in a message body, a carriage return from a Windows client, an event name containing a newline supplied by user input. Because the encoder is a pure function, it is the cheapest part of the system to test exhaustively. This guide writes the encoder properly and tests it with examples and with round-trip property tests against a spec-compliant parser.
Symptom & Developer Intent Permalink to this section
Encoder bugs look like data bugs downstream:
- A message containing a line break arrives in the browser truncated at the break.
- Some events never reach their listener; others arrive with the wrong event type.
- A user-controlled field makes the client receive events that the server never sent.
- Occasionally the whole stream stops dispatching until the next reconnect.
- Non-ASCII text arrives garbled on some clients.
The intent is an encoder whose output, parsed by the specification’s algorithm, always yields exactly the event that was encoded — for every possible input — and a test suite that proves it.
Root Cause Analysis Permalink to this section
The event stream format is line-oriented. A line ends at LF, CR, or CRLF; a blank line dispatches the event. Any of those characters inside a field value therefore ends the field early. In data the specification provides the escape hatch: split on line breaks and emit one data: line per piece, and the parser rejoins them with LF. In id and event there is no escape — the value must not contain line breaks at all.
Injection is the security form of the same bug. If an event name or id comes from user input and contains \n\ndata: {"admin":true}\n\n, the client parses an extra event the server never meant to send. Stream-stopping bugs come from \r alone: a CR is a line terminator, so a payload ending in CR followed by the encoder’s own LF produces an unexpected blank line, and data can be dispatched early.
Step-by-Step Resolution Permalink to this section
Step 1 — Write a strict encoder Permalink to this section
// sse-format.js
const LINE_BREAK = /\r\n|\r|\n/;
export function formatEvent({ id, event, data, retry, comment } = {}) {
let out = '';
if (comment !== undefined) {
for (const line of String(comment).split(LINE_BREAK)) out += `: ${line}\n`;
}
if (id !== undefined) {
const v = String(id);
if (LINE_BREAK.test(v) || v.includes('\0')) throw new TypeError('SSE id must not contain CR, LF or NUL');
out += `id: ${v}\n`;
}
if (event !== undefined) {
const v = String(event);
if (LINE_BREAK.test(v)) throw new TypeError('SSE event name must not contain CR or LF');
out += `event: ${v}\n`;
}
if (retry !== undefined) {
if (!Number.isInteger(retry) || retry < 0) throw new TypeError('SSE retry must be a non-negative integer');
out += `retry: ${retry}\n`;
}
if (data !== undefined) {
for (const line of String(data).split(LINE_BREAK)) out += `data: ${line}\n`;
}
return out + '\n';
}
The NUL check on ids follows the specification, which tells parsers to ignore an id field containing NUL. Throwing on invalid ids and names turns an injection into a loud error in your code instead of a silent one in every client.
Step 2 — Test the rules with examples Permalink to this section
import { test, expect } from 'vitest';
import { formatEvent } from './sse-format.js';
test('single-line data', () => {
expect(formatEvent({ data: 'hi' })).toBe('data: hi\n\n');
});
test('each line break in data becomes a data line', () => {
expect(formatEvent({ data: 'a\nb\r\nc\rd' })).toBe('data: a\ndata: b\ndata: c\ndata: d\n\n');
});
test('empty data still dispatches', () => {
expect(formatEvent({ data: '' })).toBe('data: \n\n');
});
test('ids and event names cannot smuggle lines', () => {
expect(() => formatEvent({ id: '1\n\ndata: x', data: 'ok' })).toThrow();
expect(() => formatEvent({ event: 'a\rb', data: 'ok' })).toThrow();
});
test('retry must be an integer', () => {
expect(() => formatEvent({ retry: 1.5 })).toThrow();
expect(formatEvent({ retry: 3000 })).toBe('retry: 3000\n\n');
});
test('UTF-8 passes through unchanged', () => {
expect(formatEvent({ data: 'naïve — 日本 ✓' })).toBe('data: naïve — 日本 ✓\n\n');
});
Step 3 — Round-trip with a spec-compliant parser Permalink to this section
Examples only cover the cases you think of. A property test generates thousands of arbitrary events, encodes them, parses the bytes with a parser that implements the specification (such as eventsource-parser), and checks that the result equals the input.
import fc from 'fast-check';
import { createParser } from 'eventsource-parser';
const noBreaks = fc.string().filter((s) => !/[\r\n\0]/.test(s));
test('encode → parse is the identity', () => {
fc.assert(fc.property(
fc.record({
id: fc.option(noBreaks, { nil: undefined }),
event: fc.option(noBreaks.filter((s) => s.length > 0), { nil: undefined }),
data: fc.string(),
}),
(evt) => {
const got = [];
const p = createParser({ onEvent: (e) => got.push(e) });
p.feed(formatEvent(evt));
expect(got).toHaveLength(1);
// The parser normalises CR and CRLF inside data to LF.
expect(got[0].data).toBe(evt.data.replace(/\r\n|\r/g, '\n'));
expect(got[0].event ?? undefined).toBe(evt.event);
if (evt.id !== undefined) expect(got[0].id).toBe(evt.id);
},
), { numRuns: 5000 });
});
Property tests find surprising cases quickly: leading spaces in data (the parser strips exactly one space after the colon, which the encoder’s ": " accounts for), lone CRs, and empty event names.
Step 4 — Test chunk boundaries on the parsing side Permalink to this section
Streams arrive in arbitrary chunks. If you also maintain a client-side parser, feed the same encoded stream split at every possible byte offset and assert that the parsed events are identical — including splits in the middle of a multi-byte UTF-8 character and between the CR and LF of a CRLF.
test('parsing is independent of chunk boundaries', () => {
const stream = formatEvent({ id: '1', data: 'héllo\r\nworld' }) + formatEvent({ data: '✓' });
const bytes = new TextEncoder().encode(stream);
const whole = parseAll([bytes]);
for (let i = 1; i < bytes.length; i++) {
expect(parseAll([bytes.slice(0, i), bytes.slice(i)])).toEqual(whole);
}
});
parseAll must decode with a streaming TextDecoder ({ stream: true }) so split characters are reassembled; the fetch-based stream parsing guide shows the pattern.
Validation & Monitoring Permalink to this section
Run the property test in CI with a fixed seed for reproducibility and a nightly job with random seeds and more runs. Make any failing input a permanent example test. In production, count encoder exceptions: a non-zero rate means some code path is trying to put user input into an id or event name, which is worth investigating as a potential injection attempt rather than just a bug.
npx vitest run sse-format # example + property tests, seconds
FC_NUM_RUNS=100000 npx vitest run sse-format --reporter=verbose # nightly
Production Checklist Permalink to this section
Frequently Asked Questions Permalink to this section
Is JSON.stringify output always safe as data?
Yes, as long as indentation is off: JSON escapes newlines inside strings, so the output is a single line. Pretty-printed JSON contains line breaks and must go through the multi-line data rule.
Should the encoder escape instead of throwing for bad ids?
There is no escape syntax for id or event fields. Throwing is safer than silently altering an id, because a changed id breaks resume in ways that are hard to debug.
Do I need to handle the byte order mark?
Only when parsing. A parser must ignore a leading UTF-8 BOM at the start of the stream; encoders should simply never emit one.
Which parser should property tests use?
One that implements the WHATWG algorithm faithfully, such as eventsource-parser in JavaScript. Using your own parser for both sides only proves the two agree with each other.