Load Testing SSE with k6 Permalink to this section

Part of Testing & Load Testing SSE Endpoints, under Backend Stream Generation & Connection Management.

k6’s built-in HTTP client waits for each response to complete, which an event stream never does, so a plain k6 script against an SSE endpoint either hangs or reports one very slow request per virtual user. The xk6-sse extension adds a streaming client that fires a callback per event. With it, k6 becomes a good SSE load generator: each virtual user holds one stream, parses events, and records custom metrics such as delivery latency. This guide builds the extension, writes a script that measures what matters, and interprets the results.

Symptom & Developer Intent Permalink to this section

Teams usually arrive here because an existing load test does not reflect production:

  • The HTTP load test reports excellent throughput, yet production struggles at a fraction of that traffic.
  • k6 scripts time out at 60 seconds with request timeout on every virtual user.
  • Nobody knows how many concurrent streams one node can hold before latency degrades.
  • A deploy under load has never been rehearsed.
  • The load generator itself runs out of file descriptors or ephemeral ports before the server is stressed.

The intent is a repeatable test that holds a target number of concurrent streams, measures publish-to-receive latency per event, includes slow and churning clients, and finds the concurrency at which latency breaks the objective.

Root Cause Analysis Permalink to this section

A request-rate test measures the wrong resource. SSE capacity is bounded by concurrent connections (memory, file descriptors, per-connection bookkeeping) and by fan-out work per event (serialisation and one write per subscriber). Neither shows up when requests complete quickly.

Request-rate metrics versus streaming metrics Two panels contrasting the metrics a request-rate load test reports with the metrics a streaming load test must report. Request-rate metrics versus streaming metrics Request-rate test requests per second response time to completion error rate per request one connection, many requests Streaming test concurrent open streams publish-to-receive latency reconnects and resyncs events lost per client
The right-hand column is what users experience on a streaming product. A test that cannot produce it is not testing the stream.

The timeout symptom comes from k6’s default 60-second request timeout applying to the stream. The generator limits come from operating system defaults: 1,024 open files per process and roughly 28,000 ephemeral ports per source IP and destination.

Step-by-Step Resolution Permalink to this section

Step 1 — Build k6 with the SSE extension Permalink to this section

go install go.k6.io/xk6/cmd/xk6@latest
xk6 build --with github.com/phymbert/xk6-sse
./k6 version     # should list the sse extension

Step 2 — Stamp publish time into events Permalink to this section

Latency can only be measured if each event carries its publish timestamp. Add it at the source for the test environment (or permanently — it is useful in production too):

// Publisher side, test mode: include the wall-clock publish time in milliseconds.
hub.publish({ id: String(seq++), event: 'tick', data: JSON.stringify({ ts: Date.now(), n: seq }) });

Run the load generator and the publisher on hosts with synchronised clocks (NTP or chrony, sub-millisecond on the same network), or measure on the same host for small tests.

Step 3 — Write the virtual-user script Permalink to this section

// sse-load.js
import sse from 'k6/x/sse';
import { check } from 'k6';
import { Trend, Counter } from 'k6/metrics';

const latency = new Trend('sse_event_latency_ms', true);
const events = new Counter('sse_events_received');
const resyncs = new Counter('sse_resyncs');

export const options = {
  scenarios: {
    streams: {
      executor: 'ramping-vus',
      startVUs: 0,
      stages: [
        { duration: '5m', target: 10000 },     // ramp slowly: this is not an accept-queue test
        { duration: '10m', target: 10000 },    // hold
        { duration: '5m', target: 20000 },
        { duration: '10m', target: 20000 },
      ],
      gracefulRampDown: '30s',
    },
  },
  thresholds: {
    sse_event_latency_ms: ['p(99)<250'],       // the objective
  },
};

export default function () {
  const holdMs = 60000 + Math.random() * 240000;   // realistic churn: 1–5 minute sessions
  const res = sse.open(`${__ENV.BASE}/api/stream`, { headers: { Authorization: `Bearer ${__ENV.TOKEN}` } }, (client) => {
    const started = Date.now();
    client.on('event', (e) => {
      if (e.name === 'resync') { resyncs.add(1); return; }
      const body = JSON.parse(e.data);
      latency.add(Date.now() - body.ts);
      events.add(1);
      if (Date.now() - started > holdMs) client.close();   // end the session, VU will reconnect
    });
    client.on('error', (err) => console.error(err.error()));
  });
  check(res, { 'status 200': (r) => r && r.status === 200 });
}

Each virtual user opens a stream, records latency for every event, and closes after a randomised session length; the next iteration reconnects, which exercises connection setup and replay continuously.

Step 4 — Add a slow-client cohort Permalink to this section

Real audiences include slow networks. Run a second scenario whose clients sleep in their handlers, so their sockets back up:

// Added to options.scenarios:
slow: {
  executor: 'constant-vus', vus: 500, duration: '30m', exec: 'slow',
},

export function slow() {
  sse.open(`${__ENV.BASE}/api/stream`, {}, (client) => {
    client.on('event', () => { const until = Date.now() + 200; while (Date.now() < until) {} });
  });
}

The pass condition is that the fast cohort’s p99 latency is unchanged by the slow cohort. If it rises, the server’s fan-out is letting slow sockets delay others — see handling slow consumers with SSE backpressure.

A thirty-minute SSE load test plan Timeline of a load test with a slow ramp to ten thousand streams, a hold, a second ramp to twenty thousand, a hold with a node restart in the middle, and a constant slow-client cohort throughout. A thirty-minute SSE load test plan Fast cohort Slow cohort ramp hold 10k ramp hold 20k 500 slow readers 0 6 12 18 24 30 minutes restart one node
Holding at each level is what reveals memory growth and latency drift. The restart at minute 24 measures the reconnect wave under load.

Step 5 — Prepare the load generator Permalink to this section

# On each load generator host
ulimit -n 1048576
sudo sysctl -w net.ipv4.ip_local_port_range="1024 65000"
sudo sysctl -w net.ipv4.tcp_tw_reuse=1

# Distributed: split VUs across machines (or use k6's distributed operator on Kubernetes).
BASE=https://staging.example.com TOKEN=$T ./k6 run --out json=results.json sse-load.js

A single generator comfortably holds 20,000–40,000 streams; beyond that, use several machines or several source IPs so no single IP runs out of ephemeral ports to one destination.

Validation & Monitoring Permalink to this section

Read the results against server-side metrics captured at the same time: open connections, memory, CPU, event-loop or scheduler latency, and file descriptors.

Example result: latency against concurrent streams Line chart of p99 event latency and server memory against concurrent streams from a k6 test, with the latency objective marked. Example result: latency against concurrent streams p99 latency memory (MB ÷ 10) 0 150 300 450 600 0 4 8 12 16 20 objective concurrent streams (thousands) p99 latency (ms)
Memory grows linearly and harmlessly; latency is flat until about 17,000 streams and then climbs. The knee, not the memory curve, sets the per-node capacity.

Look at the restart window separately. When one node restarts at minute 24, its clients reconnect to the survivors; the useful numbers are how long it takes until the open-stream count recovers, how high p99 latency spikes on the surviving nodes while they absorb the reconnect wave and serve replays, and whether any client received a resync because its position fell outside the replay window. A healthy system recovers within the retry interval plus a few seconds, with a latency spike that stays inside the objective.

Sanity-check the test itself: sse_events_received divided by test duration and virtual users should match the publish rate. If clients receive fewer events than published, events were lost — a correctness bug, not a performance one.

Production Checklist Permalink to this section

Frequently Asked Questions Permalink to this section

Can plain k6 test SSE without an extension?

Not meaningfully. The built-in HTTP client waits for the full response, so it cannot observe individual events. It can only test connection establishment with a short timeout, which says nothing about streaming.

How many virtual users can one machine run?

Tens of thousands of mostly idle streams on a modern machine with raised limits. The constraint is usually memory in the k6 process and ephemeral ports per destination; spread larger tests across machines.

Should load tests go through the CDN or load balancer?

Run both. Direct-to-node tests measure the application's capacity; through-the-edge tests find timeouts, buffering and connection limits in the infrastructure, which are often lower.

Are there alternatives to k6 for SSE load?

Yes: small custom clients in Go are common, Artillery has SSE plugins, and Gatling supports SSE natively. Whatever the tool, it must hold streams open and measure per-event latency.