Envoy and Istio Timeouts for SSE Permalink to this section
Part of Proxy & CDN Configuration for SSE, under SSE Protocol Fundamentals & Architecture.
Envoy streams responses without buffering by default, which makes it a good proxy for Server-Sent Events — once its timeouts are understood. Envoy has several, at different levels, and the one that most often ends SSE streams is not an idle timeout at all but a route timeout that limits how long the entire response may take. In a service mesh such as Istio, sidecars add a second Envoy on each side of the call, each with its own settings. This guide maps the timeouts, shows the configuration for plain Envoy and for Istio, and covers the filters and lifecycle settings that also matter.
Symptom & Developer Intent Permalink to this section
- Streams end after exactly 15 seconds, or after exactly 5 minutes, regardless of traffic.
- Streams survive when heartbeats are frequent but still end at a fixed duration.
- A service works directly but not through the mesh, or works through the ingress gateway but not between services.
- Enabling a compression filter made events arrive in batches.
- Pod restarts in the mesh reset streams with HTTP/2 errors rather than ending them cleanly.
The intent is to know which Envoy timeout ends a stream, set it deliberately for the stream route only, and keep sidecars from breaking stream lifecycle.
Root Cause Analysis Permalink to this section
Envoy distinguishes timeouts that measure inactivity from timeouts that measure total duration. Heartbeats defeat the first kind and do nothing against the second.
The route timeout (timeout on a route action) covers the time until the complete upstream response has been received. Its default is 15 seconds. For an infinite response that means every stream ends 15 seconds after it starts, whatever it carries. The HTTP connection manager’s stream_idle_timeout defaults to 5 minutes of inactivity, which heartbeats easily beat.
Istio configures these through its own resources. A VirtualService route’s timeout maps to the Envoy route timeout. Current Istio versions leave it disabled by default, but any timeout: set on a route applies to streams too.
Step-by-Step Resolution Permalink to this section
Step 1 — Disable the route timeout for stream routes (plain Envoy) Permalink to this section
route_config:
virtual_hosts:
- name: app
domains: ["app.example.com"]
routes:
- match: { prefix: "/api/stream" }
route:
cluster: sse
timeout: 0s # no total-duration limit for the stream
idle_timeout: 120s # inactivity limit; heartbeats every 15 s beat it
- match: { prefix: "/" }
route: { cluster: api, timeout: 15s }
Keep a finite idle_timeout on the stream route. It is how Envoy frees streams whose clients vanished without closing; heartbeats keep healthy streams under it.
Step 2 — Check connection-manager-level limits Permalink to this section
http_connection_manager:
stream_idle_timeout: 300s # default; fine with heartbeats
# max_stream_duration: unset # a total-duration cap would end streams regardless
common_http_protocol_options:
idle_timeout: 3600s # connection idle; applies when no streams are active
If max_stream_duration is set globally for other reasons, streams will end at that limit; either raise it or design the client for planned reconnects.
Step 3 — Configure the same in Istio Permalink to this section
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata:
name: sse
spec:
hosts: ["app.example.com"]
gateways: ["istio-system/public"]
http:
- match: [{ uri: { prefix: /api/stream } }]
route: [{ destination: { host: sse.default.svc.cluster.local, port: { number: 8080 } } }]
# no timeout here: Istio leaves the route timeout disabled unless set
- route: [{ destination: { host: api.default.svc.cluster.local } }]
timeout: 15s
Mesh-internal calls (service to service) go through both the client’s sidecar and the server’s sidecar. A VirtualService bound to the mesh (gateways: [mesh]) governs the client side; make sure no mesh-wide default applies a route timeout to the streaming service.
Retries are the other mesh default worth checking. Istio and many Envoy configurations retry failed requests automatically. For an ordinary request that is harmless; for a stream that fails after an hour, a proxy-level retry replays the original request without the Last-Event-ID the browser would have sent, which can reset the client’s position or duplicate events. Disable proxy retries on the stream route and let EventSource handle reconnection, since only the browser knows the last id it received:
- match: [{ uri: { prefix: /api/stream } }]
retries: { attempts: 0 }
route: [{ destination: { host: sse.default.svc.cluster.local } }]
Step 4 — Keep compression filters off event streams Permalink to this section
Envoy’s compressor filter can be configured with content types; exclude text/event-stream or leave it out of the list:
http_filters:
- name: envoy.filters.http.compressor
typed_config:
"@type": type.googleapis.com/envoy.extensions.filters.http.compressor.v3.Compressor
response_direction_config:
common_config:
content_type: [application/json, text/html, application/javascript]
Buffer filters and any filter that needs the complete body (some auth or transformation filters) must also be disabled on the stream route.
Step 5 — Drain sidecars in the right order Permalink to this section
On pod shutdown, the application should end streams cleanly before its sidecar stops, or clients see resets. In Istio, configure the proxy to wait for the application:
metadata:
annotations:
proxy.istio.io/config: |
terminationDrainDuration: 45s
spec:
terminationGracePeriodSeconds: 60
Combine with an application preStop drain that marks the pod not-ready and ends streams gradually with a retry hint, as in configuring Kubernetes ingress for SSE.
Validation & Monitoring Permalink to this section
# Does a quiet stream survive longer than 15 s and 5 min through the gateway?
timeout 400 curl -sN https://app.example.com/api/stream | while read l; do echo "$(date +%T) $l"; done
# Inspect the route Envoy actually has (Istio).
istioctl proxy-config route deploy/frontend -o json | jq '.. | objects | select(.match?.prefix? == "/api/stream")'
When a stream still ends early, change one Envoy at a time. Test directly against the server sidecar’s inbound port from inside the pod network, then through the client sidecar, then through the ingress gateway; the first hop that cuts the stream owns the offending setting.
Envoy’s stats show why streams end: downstream_rq_idle_timeout, upstream_rq_timeout and downstream_cx_max_duration_reached counters on the relevant listeners and clusters. A rising upstream_rq_timeout on the SSE cluster means a route timeout is still in effect somewhere.
Production Checklist Permalink to this section
Frequently Asked Questions Permalink to this section
Why do my streams end after exactly 15 seconds behind Envoy?
That is the default route timeout, which limits total response time. Set timeout: 0s on the stream route; heartbeats cannot help against a total-duration limit.
Does Envoy buffer SSE responses?
Not by default. Buffering appears only with filters that need the whole body, such as buffer or some compression and transformation filters.
Is it safe to disable the route timeout?
For stream routes, yes, provided an idle timeout remains to reclaim abandoned streams. Keep finite route timeouts on ordinary request routes.
Do other Envoy-based gateways behave the same way?
Yes. Contour, Emissary, Envoy Gateway and cloud gateways built on Envoy expose the same route and idle timeouts under their own resource names.