Load Testing SSE with k6 Permalink to this section
Part of Testing & Load Testing SSE Endpoints, under Backend Stream Generation & Connection Management.
k6’s built-in HTTP client waits for each response to complete, which an event stream never does, so a plain k6 script against an SSE endpoint either hangs or reports one very slow request per virtual user. The xk6-sse extension adds a streaming client that fires a callback per event. With it, k6 becomes a good SSE load generator: each virtual user holds one stream, parses events, and records custom metrics such as delivery latency. This guide builds the extension, writes a script that measures what matters, and interprets the results.
Symptom & Developer Intent Permalink to this section
Teams usually arrive here because an existing load test does not reflect production:
- The HTTP load test reports excellent throughput, yet production struggles at a fraction of that traffic.
- k6 scripts time out at 60 seconds with
request timeouton every virtual user. - Nobody knows how many concurrent streams one node can hold before latency degrades.
- A deploy under load has never been rehearsed.
- The load generator itself runs out of file descriptors or ephemeral ports before the server is stressed.
The intent is a repeatable test that holds a target number of concurrent streams, measures publish-to-receive latency per event, includes slow and churning clients, and finds the concurrency at which latency breaks the objective.
Root Cause Analysis Permalink to this section
A request-rate test measures the wrong resource. SSE capacity is bounded by concurrent connections (memory, file descriptors, per-connection bookkeeping) and by fan-out work per event (serialisation and one write per subscriber). Neither shows up when requests complete quickly.
The timeout symptom comes from k6’s default 60-second request timeout applying to the stream. The generator limits come from operating system defaults: 1,024 open files per process and roughly 28,000 ephemeral ports per source IP and destination.
Step-by-Step Resolution Permalink to this section
Step 1 — Build k6 with the SSE extension Permalink to this section
go install go.k6.io/xk6/cmd/xk6@latest
xk6 build --with github.com/phymbert/xk6-sse
./k6 version # should list the sse extension
Step 2 — Stamp publish time into events Permalink to this section
Latency can only be measured if each event carries its publish timestamp. Add it at the source for the test environment (or permanently — it is useful in production too):
// Publisher side, test mode: include the wall-clock publish time in milliseconds.
hub.publish({ id: String(seq++), event: 'tick', data: JSON.stringify({ ts: Date.now(), n: seq }) });
Run the load generator and the publisher on hosts with synchronised clocks (NTP or chrony, sub-millisecond on the same network), or measure on the same host for small tests.
Step 3 — Write the virtual-user script Permalink to this section
// sse-load.js
import sse from 'k6/x/sse';
import { check } from 'k6';
import { Trend, Counter } from 'k6/metrics';
const latency = new Trend('sse_event_latency_ms', true);
const events = new Counter('sse_events_received');
const resyncs = new Counter('sse_resyncs');
export const options = {
scenarios: {
streams: {
executor: 'ramping-vus',
startVUs: 0,
stages: [
{ duration: '5m', target: 10000 }, // ramp slowly: this is not an accept-queue test
{ duration: '10m', target: 10000 }, // hold
{ duration: '5m', target: 20000 },
{ duration: '10m', target: 20000 },
],
gracefulRampDown: '30s',
},
},
thresholds: {
sse_event_latency_ms: ['p(99)<250'], // the objective
},
};
export default function () {
const holdMs = 60000 + Math.random() * 240000; // realistic churn: 1–5 minute sessions
const res = sse.open(`${__ENV.BASE}/api/stream`, { headers: { Authorization: `Bearer ${__ENV.TOKEN}` } }, (client) => {
const started = Date.now();
client.on('event', (e) => {
if (e.name === 'resync') { resyncs.add(1); return; }
const body = JSON.parse(e.data);
latency.add(Date.now() - body.ts);
events.add(1);
if (Date.now() - started > holdMs) client.close(); // end the session, VU will reconnect
});
client.on('error', (err) => console.error(err.error()));
});
check(res, { 'status 200': (r) => r && r.status === 200 });
}
Each virtual user opens a stream, records latency for every event, and closes after a randomised session length; the next iteration reconnects, which exercises connection setup and replay continuously.
Step 4 — Add a slow-client cohort Permalink to this section
Real audiences include slow networks. Run a second scenario whose clients sleep in their handlers, so their sockets back up:
// Added to options.scenarios:
slow: {
executor: 'constant-vus', vus: 500, duration: '30m', exec: 'slow',
},
export function slow() {
sse.open(`${__ENV.BASE}/api/stream`, {}, (client) => {
client.on('event', () => { const until = Date.now() + 200; while (Date.now() < until) {} });
});
}
The pass condition is that the fast cohort’s p99 latency is unchanged by the slow cohort. If it rises, the server’s fan-out is letting slow sockets delay others — see handling slow consumers with SSE backpressure.
Step 5 — Prepare the load generator Permalink to this section
# On each load generator host
ulimit -n 1048576
sudo sysctl -w net.ipv4.ip_local_port_range="1024 65000"
sudo sysctl -w net.ipv4.tcp_tw_reuse=1
# Distributed: split VUs across machines (or use k6's distributed operator on Kubernetes).
BASE=https://staging.example.com TOKEN=$T ./k6 run --out json=results.json sse-load.js
A single generator comfortably holds 20,000–40,000 streams; beyond that, use several machines or several source IPs so no single IP runs out of ephemeral ports to one destination.
Validation & Monitoring Permalink to this section
Read the results against server-side metrics captured at the same time: open connections, memory, CPU, event-loop or scheduler latency, and file descriptors.
Look at the restart window separately. When one node restarts at minute 24, its clients reconnect to the survivors; the useful numbers are how long it takes until the open-stream count recovers, how high p99 latency spikes on the surviving nodes while they absorb the reconnect wave and serve replays, and whether any client received a resync because its position fell outside the replay window. A healthy system recovers within the retry interval plus a few seconds, with a latency spike that stays inside the objective.
Sanity-check the test itself: sse_events_received divided by test duration and virtual users should match the publish rate. If clients receive fewer events than published, events were lost — a correctness bug, not a performance one.
Production Checklist Permalink to this section
Frequently Asked Questions Permalink to this section
Can plain k6 test SSE without an extension?
Not meaningfully. The built-in HTTP client waits for the full response, so it cannot observe individual events. It can only test connection establishment with a short timeout, which says nothing about streaming.
How many virtual users can one machine run?
Tens of thousands of mostly idle streams on a modern machine with raised limits. The constraint is usually memory in the k6 process and ephemeral ports per destination; spread larger tests across machines.
Should load tests go through the CDN or load balancer?
Run both. Direct-to-node tests measure the application's capacity; through-the-edge tests find timeouts, buffering and connection limits in the infrastructure, which are often lower.
Are there alternatives to k6 for SSE load?
Yes: small custom clients in Go are common, Artillery has SSE plugins, and Gatling supports SSE natively. Whatever the tool, it must hold streams open and measure per-event latency.