Configuring Kubernetes Ingress for SSE Permalink to this section

Part of Proxy & CDN Configuration for SSE, under SSE Protocol Fundamentals & Architecture.

In Kubernetes, the reverse proxy in front of a Server-Sent Events service is usually an ingress controller or a Gateway API implementation, configured through annotations or route resources rather than a proxy config file. The underlying requirements are the same as anywhere — no buffering, long enough timeouts, graceful draining — but they are expressed differently and are easy to miss because the defaults are tuned for short requests. This guide covers ingress-nginx annotations, which many existing clusters still run, the Gateway API equivalents for newer clusters, and the pod-lifecycle settings that make deploys invisible to connected clients.

Symptom & Developer Intent Permalink to this section

  • Streams close after 60 seconds of quiet, matching the controller’s default proxy read timeout.
  • Events arrive in bursts through the ingress but promptly when port-forwarding to the pod.
  • Every rollout disconnects all clients at once, and some see errors instead of a clean reconnect.
  • Setting a timeout annotation on one Ingress had no effect because another Ingress for the same host overrides it.
  • Moving from Ingress to Gateway API brought back the 60-second disconnects.

The intent is a declarative configuration, committed alongside the service, that makes streams survive quiet periods, flow without buffering, and hand over cleanly during rollouts.

Root Cause Analysis Permalink to this section

ingress-nginx generates an nginx configuration from Ingress resources. Its defaults include response buffering and a 60-second proxy_read_timeout — exactly the two settings that break SSE. Annotations override them per Ingress. Gateway API implementations have their own defaults, and HTTPRoute timeouts are expressed differently.

The layers between a browser and an SSE pod Stack of the cloud load balancer, the ingress controller or gateway, the Service and the pod, each with the setting that matters for SSE. The layers between a browser and an SSE pod Cloud load balancer idle timeout often 60 s by default Ingress / Gateway buffering, read timeout defaults tuned for short requests Service (kube-proxy) conntrack usually transparent Pod preStop + grace period drain before SIGKILL
Each layer has its own idle timeout. The heartbeat must beat the smallest; the ingress also needs buffering off.

Rollouts disconnect clients abruptly when pods receive SIGTERM and exit immediately, or when the terminationGracePeriod expires before streams end. Without a drain step, clients see reset connections rather than a clean end with a retry hint.

The ingress-nginx project has been retired upstream, so new clusters generally use a Gateway API implementation instead. The same requirements apply; only the syntax changes.

Step-by-Step Resolution Permalink to this section

Step 1 — Annotate the SSE Ingress (ingress-nginx) Permalink to this section

apiVersion: networking.k8s.io/v1
kind: Ingress
metadata:
  name: sse
  annotations:
    nginx.ingress.kubernetes.io/proxy-buffering: "off"
    nginx.ingress.kubernetes.io/proxy-read-timeout: "3600"
    nginx.ingress.kubernetes.io/proxy-send-timeout: "3600"
    nginx.ingress.kubernetes.io/proxy-http-version: "1.1"
spec:
  ingressClassName: nginx
  rules:
    - host: app.example.com
      http:
        paths:
          - path: /api/stream
            pathType: Prefix
            backend: { service: { name: sse, port: { number: 8080 } } }

Put the stream path in its own Ingress so its annotations do not affect the rest of the host, and so the rest of the host’s annotations do not affect it. When several Ingress resources share a host, ingress-nginx merges them per path, but conflicting server-level settings are resolved by resource order — another reason to keep stream settings on location-level annotations only.

Also send X-Accel-Buffering: no from the application. It disables buffering for that response even if an annotation is lost in a later refactor.

Step 2 — Configure the same with Gateway API Permalink to this section

apiVersion: gateway.networking.k8s.io/v1
kind: HTTPRoute
metadata:
  name: sse
spec:
  parentRefs: [{ name: public-gateway }]
  hostnames: ["app.example.com"]
  rules:
    - matches: [{ path: { type: PathPrefix, value: /api/stream } }]
      timeouts:
        request: 0s              # zero disables the request timeout for this route
      backendRefs: [{ name: sse, port: 8080 }]

Response buffering and idle timeouts are implementation-specific in Gateway API; check your implementation’s policy resources for streaming (for example, backend traffic or timeout policies) and verify behaviour with a quiet stream. Envoy and Istio timeouts for SSE covers Envoy-based gateways in detail.

The same SSE requirements in Ingress and Gateway API Matrix mapping four SSE requirements to their ingress-nginx annotation and their Gateway API expression. The same SSE requirements in Ingress and Gateway API Requirement ingress-nginx Gateway API No response buffering proxy-buffering: off implementation policy Long read timeout proxy-read-timeout timeouts.request: 0s HTTP/1.1 upstream proxy-http-version implementation default Separate stream route own Ingress own HTTPRoute rule
The requirements do not change between APIs. Check each one explicitly after any migration.

Step 3 — Drain pods before they stop Permalink to this section

spec:
  template:
    spec:
      terminationGracePeriodSeconds: 60
      containers:
        - name: sse
          lifecycle:
            preStop:
              exec:
                command: ["/bin/sh", "-c", "curl -s -X POST localhost:8080/admin/drain; sleep 40"]
          readinessProbe:
            httpGet: { path: /ready, port: 8080 }
            periodSeconds: 2

The drain endpoint marks the pod not-ready (so the controller stops routing new streams to it), then ends existing streams gradually with a short, jittered retry: hint. Clients reconnect to other pods and resume from Last-Event-ID. The sleep keeps the container alive while that happens; the grace period must exceed it. The same drain logic applies outside Kubernetes, as in rebalancing SSE connections after a deploy.

Two details make draining reliable. First, readiness must flip before streams are ended, and the controller needs a moment to notice — ingress controllers and gateways update their endpoint lists asynchronously, typically within a few seconds. Ending streams immediately risks clients reconnecting to the very pod that is shutting down. A short pause after marking not-ready, then gradual closure, avoids it. Second, the controller itself holds the client-side connection: when the pod ends a stream cleanly, the controller ends the client response cleanly too, and EventSource reconnects as intended. When the pod is killed instead, the controller may return a 502 or reset the stream, which fails the connection permanently in the browser.

Step 4 — Align the cloud load balancer Permalink to this section

Horizontal pod autoscaling interacts with long-lived streams too. New pods receive only new connections, so scaling out on CPU does little until existing streams recycle; recycle streams on a jittered maximum age so new pods fill within minutes. Scale on open-stream count per pod rather than CPU, since an SSE pod is often limited by connections and memory long before it is busy.

If a cloud load balancer sits in front of the ingress, its idle timeout also applies. Raise it through the Service annotations your provider supports, or rely on heartbeats shorter than its default. Measure the whole path, not just the ingress.

Validation & Monitoring Permalink to this section

# Compare direct-to-pod and through-ingress timing.
kubectl port-forward svc/sse 8080:8080 &
curl -sN localhost:8080/api/stream | head -3                       # baseline
curl -sN https://app.example.com/api/stream | while read l; do echo "$(date +%T) $l"; done

# Inspect the generated nginx config for the stream location.
kubectl exec -n ingress-nginx deploy/ingress-nginx-controller -- \
  cat /etc/nginx/nginx.conf | grep -A12 'location /api/stream'
A rolling update with drain versus without Timeline of a three-minute rollout comparing abrupt pod termination, which disconnects clients with errors, against preStop draining, which ends streams gradually. A rolling update with drain versus without Without drain With drain all cut gradual gradual gradual 0 36 72 108 144 180 seconds
With draining, each pod's clients leave over thirty seconds with a retry hint; without it, they are cut off together.

During a test rollout, watch reconnects per second and client-side errors; with draining, reconnects form low plateaus and errors stay near zero.

Production Checklist Permalink to this section

Frequently Asked Questions Permalink to this section

Why do streams work with kubectl port-forward but not through the ingress?

Port-forward bypasses the ingress controller, so its buffering and timeouts do not apply. The difference isolates the problem to ingress configuration.

Is a 3600-second read timeout safe?

It is safe for the stream route alone, especially with heartbeats. Apply it only to that route so ordinary endpoints keep short timeouts.

Do I need sticky sessions for SSE in Kubernetes?

Not if streams are resumable from shared storage. Any pod can serve any reconnect, which also makes draining and scaling simpler.

What should replace ingress-nginx for new clusters?

A Gateway API implementation is the usual path. Re-verify buffering, timeouts and draining after the move, since defaults differ between implementations.