Configuring Kubernetes Ingress for SSE Permalink to this section
Part of Proxy & CDN Configuration for SSE, under SSE Protocol Fundamentals & Architecture.
In Kubernetes, the reverse proxy in front of a Server-Sent Events service is usually an ingress controller or a Gateway API implementation, configured through annotations or route resources rather than a proxy config file. The underlying requirements are the same as anywhere — no buffering, long enough timeouts, graceful draining — but they are expressed differently and are easy to miss because the defaults are tuned for short requests. This guide covers ingress-nginx annotations, which many existing clusters still run, the Gateway API equivalents for newer clusters, and the pod-lifecycle settings that make deploys invisible to connected clients.
Symptom & Developer Intent Permalink to this section
- Streams close after 60 seconds of quiet, matching the controller’s default proxy read timeout.
- Events arrive in bursts through the ingress but promptly when port-forwarding to the pod.
- Every rollout disconnects all clients at once, and some see errors instead of a clean reconnect.
- Setting a timeout annotation on one Ingress had no effect because another Ingress for the same host overrides it.
- Moving from Ingress to Gateway API brought back the 60-second disconnects.
The intent is a declarative configuration, committed alongside the service, that makes streams survive quiet periods, flow without buffering, and hand over cleanly during rollouts.
Root Cause Analysis Permalink to this section
ingress-nginx generates an nginx configuration from Ingress resources. Its defaults include response buffering and a 60-second proxy_read_timeout — exactly the two settings that break SSE. Annotations override them per Ingress. Gateway API implementations have their own defaults, and HTTPRoute timeouts are expressed differently.
Rollouts disconnect clients abruptly when pods receive SIGTERM and exit immediately, or when the terminationGracePeriod expires before streams end. Without a drain step, clients see reset connections rather than a clean end with a retry hint.
The ingress-nginx project has been retired upstream, so new clusters generally use a Gateway API implementation instead. The same requirements apply; only the syntax changes.
Step-by-Step Resolution Permalink to this section
Step 1 — Annotate the SSE Ingress (ingress-nginx) Permalink to this section
apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
name: sse
annotations:
nginx.ingress.kubernetes.io/proxy-buffering: "off"
nginx.ingress.kubernetes.io/proxy-read-timeout: "3600"
nginx.ingress.kubernetes.io/proxy-send-timeout: "3600"
nginx.ingress.kubernetes.io/proxy-http-version: "1.1"
spec:
ingressClassName: nginx
rules:
- host: app.example.com
http:
paths:
- path: /api/stream
pathType: Prefix
backend: { service: { name: sse, port: { number: 8080 } } }
Put the stream path in its own Ingress so its annotations do not affect the rest of the host, and so the rest of the host’s annotations do not affect it. When several Ingress resources share a host, ingress-nginx merges them per path, but conflicting server-level settings are resolved by resource order — another reason to keep stream settings on location-level annotations only.
Also send X-Accel-Buffering: no from the application. It disables buffering for that response even if an annotation is lost in a later refactor.
Step 2 — Configure the same with Gateway API Permalink to this section
apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
name: sse
spec:
parentRefs: [{ name: public-gateway }]
hostnames: ["app.example.com"]
rules:
- matches: [{ path: { type: PathPrefix, value: /api/stream } }]
timeouts:
request: 0s # zero disables the request timeout for this route
backendRefs: [{ name: sse, port: 8080 }]
Response buffering and idle timeouts are implementation-specific in Gateway API; check your implementation’s policy resources for streaming (for example, backend traffic or timeout policies) and verify behaviour with a quiet stream. Envoy and Istio timeouts for SSE covers Envoy-based gateways in detail.
Step 3 — Drain pods before they stop Permalink to this section
spec:
template:
spec:
terminationGracePeriodSeconds: 60
containers:
- name: sse
lifecycle:
preStop:
exec:
command: ["/bin/sh", "-c", "curl -s -X POST localhost:8080/admin/drain; sleep 40"]
readinessProbe:
httpGet: { path: /ready, port: 8080 }
periodSeconds: 2
The drain endpoint marks the pod not-ready (so the controller stops routing new streams to it), then ends existing streams gradually with a short, jittered retry: hint. Clients reconnect to other pods and resume from Last-Event-ID. The sleep keeps the container alive while that happens; the grace period must exceed it. The same drain logic applies outside Kubernetes, as in rebalancing SSE connections after a deploy.
Two details make draining reliable. First, readiness must flip before streams are ended, and the controller needs a moment to notice — ingress controllers and gateways update their endpoint lists asynchronously, typically within a few seconds. Ending streams immediately risks clients reconnecting to the very pod that is shutting down. A short pause after marking not-ready, then gradual closure, avoids it. Second, the controller itself holds the client-side connection: when the pod ends a stream cleanly, the controller ends the client response cleanly too, and EventSource reconnects as intended. When the pod is killed instead, the controller may return a 502 or reset the stream, which fails the connection permanently in the browser.
Step 4 — Align the cloud load balancer Permalink to this section
Horizontal pod autoscaling interacts with long-lived streams too. New pods receive only new connections, so scaling out on CPU does little until existing streams recycle; recycle streams on a jittered maximum age so new pods fill within minutes. Scale on open-stream count per pod rather than CPU, since an SSE pod is often limited by connections and memory long before it is busy.
If a cloud load balancer sits in front of the ingress, its idle timeout also applies. Raise it through the Service annotations your provider supports, or rely on heartbeats shorter than its default. Measure the whole path, not just the ingress.
Validation & Monitoring Permalink to this section
# Compare direct-to-pod and through-ingress timing.
kubectl port-forward svc/sse 8080:8080 &
curl -sN localhost:8080/api/stream | head -3 # baseline
curl -sN https://app.example.com/api/stream | while read l; do echo "$(date +%T) $l"; done
# Inspect the generated nginx config for the stream location.
kubectl exec -n ingress-nginx deploy/ingress-nginx-controller -- \
cat /etc/nginx/nginx.conf | grep -A12 'location /api/stream'
During a test rollout, watch reconnects per second and client-side errors; with draining, reconnects form low plateaus and errors stay near zero.
Production Checklist Permalink to this section
Frequently Asked Questions Permalink to this section
Why do streams work with kubectl port-forward but not through the ingress?
Port-forward bypasses the ingress controller, so its buffering and timeouts do not apply. The difference isolates the problem to ingress configuration.
Is a 3600-second read timeout safe?
It is safe for the stream route alone, especially with heartbeats. Apply it only to that route so ordinary endpoints keep short timeouts.
Do I need sticky sessions for SSE in Kubernetes?
Not if streams are resumable from shared storage. Any pod can serve any reconnect, which also makes draining and scaling simpler.
What should replace ingress-nginx for new clusters?
A Gateway API implementation is the usual path. Re-verify buffering, timeouts and draining after the move, since defaults differ between implementations.