Streaming SSE from Flask Permalink to this section

Part of Python & FastAPI SSE Implementation Guide, under Backend Stream Generation & Connection Management.

Flask is a WSGI framework, and WSGI was designed for responses that end. A Flask view can still stream Server-Sent Events — return a Response wrapping a generator — but each open stream holds whatever the WSGI server uses to run a request: a process, a thread or a greenlet. Choosing that unit well is the whole game. This guide builds a correct Flask stream, then deploys it with a worker model that can hold hundreds or thousands of connections.

Symptom & Developer Intent Permalink to this section

  • The app works with one browser tab and hangs as soon as a few users open the live page.
  • Gunicorn logs WORKER TIMEOUT and kills workers serving streams every 30 seconds.
  • Generators keep running after clients disconnect, holding database connections.
  • RuntimeError: Working outside of request context appears inside the generator.
  • Events arrive in batches behind nginx.

The intent is a Flask endpoint that streams events immediately, releases resources when clients leave, and runs under a worker model sized for the number of concurrent viewers you expect.

Root Cause Analysis Permalink to this section

With Gunicorn’s default sync workers, one worker process handles one request at a time. A stream never finishes, so each viewer takes a worker until they leave; with four workers, the fifth viewer — and every other request to the site — waits. Gunicorn’s timeout also kills sync workers that have not responded to the arbiter within 30 seconds, which a long-running stream triggers.

Gunicorn worker classes for streaming Three panels comparing sync workers, threaded workers and gevent workers on streams per worker and caveats. Gunicorn worker classes for streaming sync 1 stream per process killed by worker timeout unsuitable for SSE gthread 1 stream per thread --threads caps streams fine for dozens to hundreds gevent 1 greenlet per stream thousands per process needs cooperative libraries
The worker class sets the unit a stream consumes. Greenlets are cheap enough to hold thousands; processes are not.

The request-context error comes from the generator running after the view has returned: by the time Flask iterates it, the request context is gone. stream_with_context keeps it alive for the generator’s lifetime. Continuing after disconnect happens because WSGI gives the application no disconnect callback; the server learns of it only when a write fails and then closes the generator, raising GeneratorExit inside it.

Step-by-Step Resolution Permalink to this section

Step 1 — A generator that heartbeats and cleans up Permalink to this section

# app.py
import json, queue
from flask import Flask, Response, request, stream_with_context

app = Flask(__name__)

@app.get("/events")
def events():
    last_id = int(request.headers.get("Last-Event-ID", 0) or 0)
    user_id = current_user_id()                      # read request state before streaming

    @stream_with_context
    def generate():
        q = hub.subscribe(user_id)
        try:
            yield "retry: 5000\n\n"
            for n in replay_after(user_id, last_id):
                yield frame(n)
            while True:
                try:
                    n = q.get(timeout=15)            # blocks this greenlet/thread only
                    yield frame(n)
                except queue.Empty:
                    yield ": hb\n\n"                 # a write that exposes dead clients
        finally:
            hub.unsubscribe(user_id, q)              # runs on GeneratorExit (client gone)

    return Response(generate(), mimetype="text/event-stream",
                    headers={"Cache-Control": "no-cache", "X-Accel-Buffering": "no"})

def frame(n):
    return f"id: {n['id']}\nevent: notification\ndata: {json.dumps(n)}\n\n"

The heartbeat does double duty: it keeps proxies from closing idle streams, and it produces a write every 15 seconds, so a vanished client causes an error and a GeneratorExit within that window rather than never.

Step 2 — Run under gevent (or threads) Permalink to this section

pip install gunicorn gevent
gunicorn app:app -k gevent --worker-connections 2000 -w 4 --timeout 0 --bind 0.0.0.0:8000

--timeout 0 disables the arbiter’s worker timeout, which is not meaningful for async workers holding long requests. With gevent, queue.Queue.get(timeout=…) must be the gevent-patched version: Gunicorn’s gevent worker monkey-patches the standard library at startup, which covers queue, socket and time. Database drivers must be cooperative as well (psycopg 3 with gevent support, or psycogreen for psycopg2), or a blocking query stalls every stream in the process.

For smaller deployments, threads are simpler and need no patching:

gunicorn app:app -k gthread -w 2 --threads 100 --timeout 0

That caps concurrent streams at 200 across the two processes.

One subscription per process, a queue per stream A Redis subscriber greenlet in each Flask worker process receives events and puts them into the queues of local streams, each served by its own greenlet. One subscription per process, a queue per stream Redis pub/sub notify channel Subscriber greenlet one per process message Stream greenlet 1 queue.get Stream greenlet 2 queue.get Stream greenlet N queue.get put
The broker sees one subscription per process. Each stream only waits on its own in-memory queue.

Step 3 — Build the hub around one background subscriber Permalink to this section

import threading, redis

class Hub:
    def __init__(self):
        self.lock = threading.Lock()
        self.queues = {}
        self.started = False

    def subscribe(self, uid):
        q = queue.Queue(maxsize=256)
        with self.lock:
            self.queues.setdefault(uid, set()).add(q)
            if not self.started:
                threading.Thread(target=self._pump, daemon=True).start()   # a greenlet under gevent
                self.started = True
        return q

    def unsubscribe(self, uid, q):
        with self.lock:
            self.queues.get(uid, set()).discard(q)

    def _pump(self):
        ps = redis.Redis.from_url(REDIS_URL).pubsub()
        ps.psubscribe("notify:*")
        for m in ps.listen():
            if m["type"] != "pmessage":
                continue
            uid = int(m["channel"].split(b":")[1])
            with self.lock:
                targets = list(self.queues.get(uid, ()))
            for q in targets:
                try: q.put_nowait(json.loads(m["data"]))
                except queue.Full: pass          # slow client; its reconnect replays

hub = Hub()

Start the subscriber lazily inside the worker process, not at import time: Gunicorn forks workers after importing the app, and threads started before the fork do not survive it.

Two details of the pump deserve care. Redis pub/sub connections drop occasionally; wrap the listen() loop in a retry with backoff, and after reconnecting, let each stream catch up from its own cursor rather than trusting that nothing was missed. And the queue’s maxsize is the per-stream buffer: a client too slow to drain 256 events will start losing them, so for feeds where loss matters, close that stream instead so the browser reconnects and replays.

Step 4 — Configure the proxy Permalink to this section

Put Flask behind nginx with buffering off for the stream location, and make sure no WSGI middleware compresses or buffers the response. The per-proxy settings are in proxy and CDN configuration for SSE.

Validation & Monitoring Permalink to this section

# 500 concurrent streams on 4 gevent workers, then check the site still answers.
for i in $(seq 1 500); do curl -sN http://localhost:8000/events > /dev/null & done
curl -s -o /dev/null -w '%{time_total}\n' http://localhost:8000/health

# Disconnect cleanup: kill the curls, wait one heartbeat, check the hub.
kill $(jobs -p); sleep 20; curl -s http://localhost:8000/debug/hub   # expect 0 queues
Concurrent streams before the site stops responding Bar chart comparing concurrent SSE streams sustainable with four sync workers, two gthread workers with 100 threads, and four gevent workers. Concurrent streams before the site stops responding 4 sync workers 4 2 × 100 gthread 200 4 gevent workers ~4,000 concurrent streams with /health still under 100 ms
The worker class decides capacity. With sync workers, the fifth viewer blocks the whole site.

If Flask’s limits become the bottleneck, the async alternatives are close at hand: Quart offers a Flask-compatible async API, and the FastAPI guide shows the ASGI equivalent.

Production Checklist Permalink to this section

Frequently Asked Questions Permalink to this section

Why does my Flask stream work with the development server but not Gunicorn?

The development server runs threaded by default, so each stream gets a thread. Gunicorn's default sync workers handle one request per process, so a single stream blocks a whole worker.

How does Flask know the client disconnected?

It does not, directly. The WSGI server notices when writing to the socket fails and closes the generator, which raises GeneratorExit inside it. Heartbeats make sure that write happens regularly.

Is gevent safe to use with Flask?

Yes, provided every blocking library is patched or cooperative. An unpatched C-extension database driver will block the whole process during queries, stalling every stream on it.

Should I migrate to an async framework?

If streaming is a core feature with thousands of concurrent viewers, an ASGI framework such as Quart or FastAPI is a better fit. For a few hundred viewers, Flask with gevent is entirely adequate.