Chaos Testing SSE Through Proxies Permalink to this section

Part of Testing & Load Testing SSE Endpoints, under Backend Stream Generation & Connection Management.

The majority of Server-Sent Events incidents are caused not by the application but by the path to it: a reverse proxy that buffers, a load balancer that closes idle connections, a network that resets connections during a deploy. Tests that talk directly to the application cannot see any of that. This guide builds a small harness — the production nginx configuration plus Toxiproxy for fault injection — and a set of tests that prove the stream survives each fault.

Symptom & Developer Intent Permalink to this section

  • Every change to proxy configuration is a leap of faith, verified only in production.
  • A new ingress controller or CDN setting broke streaming last quarter and nobody noticed for days.
  • Heartbeat intervals were chosen by guesswork, not against the real idle timeout.
  • The client’s reconnect logic has never been exercised against a real reset.
  • Slow-client behaviour is theoretical.

The intent is an automated suite, runnable locally and in CI, that fails if buffering is enabled, if heartbeats do not beat the idle timeout, if a reset loses events, or if a slow client harms others.

Root Cause Analysis Permalink to this section

Proxies and networks change the stream in ways the application never sees. nginx buffers responses by default. Idle timeouts close connections that carry no bytes for a configured period. Resets happen during deploys and network changes. Bandwidth-limited clients fill their socket buffers. Each is a property of infrastructure configuration, so it must be tested with that configuration in place.

The chaos harness Flow from the test runner through the production nginx configuration and a Toxiproxy fault injector to the application under test with an in-memory hub. The chaos harness Test runner curl, node, k6 HTTP nginx production config upstream Toxiproxy inject faults proxied App under test in-memory hub
nginx runs the exact configuration from the deployment repository. Toxiproxy sits behind it so faults hit the proxy-to-app hop, and a second Toxiproxy listener can sit in front to simulate client-side networks.

Step-by-Step Resolution Permalink to this section

Step 1 — Compose the harness Permalink to this section

# docker-compose.chaos.yml
services:
  app:
    build: .
    environment: { NODE_ENV: test }
  toxiproxy:
    image: ghcr.io/shopify/toxiproxy:2.9.0
    ports: ["8474:8474"]                 # control API
  nginx:
    image: nginx:1.27
    volumes: ["./deploy/nginx/sse.conf:/etc/nginx/conf.d/default.conf:ro"]   # the real config
    ports: ["8080:80"]
    depends_on: [toxiproxy]
docker compose -f docker-compose.chaos.yml up -d
# Route nginx's upstream through Toxiproxy: nginx → toxiproxy:9000 → app:3000
curl -s -X POST localhost:8474/proxies -d '{"name":"app","listen":"0.0.0.0:9000","upstream":"app:3000"}'

The deployment’s nginx config should read its upstream from a variable or an include so the harness can point it at toxiproxy:9000 without editing the real file.

Step 2 — Test that nothing buffers Permalink to this section

# Publish one event per second for 5 s; each must arrive within ~200 ms of publication.
node test/chaos/assert-unbuffered.js http://localhost:8080/api/stream
// assert-unbuffered.js — fails if events arrive in bursts.
const gaps = [];
let last = null;
for await (const evt of streamEvents(process.argv[2], { count: 5 })) {
  const lag = Date.now() - JSON.parse(evt.data).ts;
  if (lag > 200) throw new Error(`event delayed ${lag} ms — buffering in the path`);
  if (last) gaps.push(Date.now() - last);
  last = Date.now();
}

Remove proxy_buffering off from the config and this test must fail — run that mutation once to prove the test works.

Step 3 — Test heartbeats against the idle timeout Permalink to this section

Set Toxiproxy’s timeout toxic, which closes a connection that sees no data for a given period, to your load balancer’s real idle timeout. Publish nothing and hold the stream open for several multiples of that period:

curl -s -X POST localhost:8474/proxies/app/toxics \
  -d '{"type":"timeout","attributes":{"timeout":60000},"stream":"downstream"}'
timeout 200 curl -sN http://localhost:8080/api/stream | grep -c '^:'   # heartbeats seen
# The connection must still be open at 200 s and heartbeats must be ≥ 10.
Heartbeats against a 60-second idle timeout Timeline of 200 seconds comparing a stream with 90-second heartbeats, which is closed at 60 seconds, and a stream with 15-second heartbeats, which stays open. Heartbeats against a 60-second idle timeout 90 s heartbeat 15 s heartbeat open closed, reconnect loop open throughout 0 40 80 120 160 200 seconds with no application events idle timeout
The test holds the stream idle for more than three timeout periods. Only the stream whose heartbeat beats the timeout survives.

Step 4 — Test resets with resume Permalink to this section

curl -s -X POST localhost:8474/proxies/app/toxics \
  -d '{"type":"reset_peer","attributes":{"timeout":5000},"stream":"downstream"}'

A client test publishes events continuously, lets the reset fire, and asserts that the ids it received are contiguous — that the reconnect with Last-Event-ID replayed everything published during the drop. Remove the toxic after the test (DELETE /proxies/app/toxics/reset_peer_downstream).

Step 5 — Test slow-client isolation Permalink to this section

Create a second Toxiproxy listener in front of nginx with a bandwidth toxic of 1 KB/s, and connect a slow client through it while fast clients connect directly:

curl -s -X POST localhost:8474/proxies -d '{"name":"slowclient","listen":"0.0.0.0:9100","upstream":"nginx:80"}'
curl -s -X POST localhost:8474/proxies/slowclient/toxics \
  -d '{"type":"bandwidth","attributes":{"rate":1},"stream":"downstream"}'

The assertion is on the fast clients: their latency must not change while the slow client is connected, and the application’s memory must stay bounded. If either fails, the server needs per-client buffering limits, as in handling slow consumers with SSE backpressure.

Step 6 — Test latency spikes against the client watchdog Permalink to this section

A latency toxic of 3,000 ms with jitter should not make the client’s silence watchdog reconnect if the watchdog is set to twice the heartbeat interval. This catches watchdogs tuned too aggressively for real networks.

Validation & Monitoring Permalink to this section

The chaos suite and what a failure means Matrix of the five chaos tests, the toxic or configuration used, and what a failure of each indicates. The chaos suite and what a failure means Test Fault A failure means Unbuffered real nginx config buffering or compression on Idle survival timeout toxic 60 s heartbeat too slow Reset + resume reset_peer replay or id bug Slow isolation bandwidth 1 KB/s shared write path Latency spike latency 3 s watchdog too tight
Each failing test points at a specific configuration value or code path. None of them can fail in a direct-to-app test.

Record the results as numbers, not just pass or fail: the worst event delay in the buffering test, the number of heartbeats observed in the idle test, the reconnect time and number of replayed events in the reset test, and the fast cohort’s p99 latency in the slow-client test. Trends in those numbers reveal regressions long before a threshold is crossed — a reconnect that took 400 ms last month and 2.5 s today points at a change in retry hints or authentication cost even though the test still passes.

Run the suite in CI whenever proxy configuration, ingress manifests or connection-handling code change, and nightly otherwise. Keep each test’s toxics scoped to the test and removed in teardown, or one test’s fault leaks into the next.

Production Checklist Permalink to this section

Frequently Asked Questions Permalink to this section

Why Toxiproxy rather than tc/netem?

Toxiproxy is controlled per connection through an HTTP API, runs in a container, and needs no privileges, which suits CI. netem is more powerful for host-level network simulation and works well on dedicated test machines.

Can I test cloud load balancers this way?

Not the managed service itself, but you can reproduce its documented idle timeout and header behaviour with the timeout toxic and nginx settings, then verify against a staging environment behind the real balancer.

How long do these tests take?

Most run in seconds; the idle-timeout test needs several timeout periods, typically three to four minutes. Run that one in a separate, less frequent job if CI time matters.

What about HTTP/2 between client and proxy?

Add a TLS listener to nginx with HTTP/2 enabled and repeat the suite; resets appear as RST_STREAM frames rather than connection closes, which exercises different client code paths.