SSE on Vercel and Netlify Functions Permalink to this section

Part of Edge & Serverless SSE Deployment, under Backend Stream Generation & Connection Management.

Frontend platforms such as Vercel and Netlify make it easy to add an API route next to a web app, and both support streaming responses from their functions. That is enough to serve Server-Sent Events. What they do not offer is an unbounded, long-lived process: every function invocation has a maximum duration, instances come and go, and you pay for the time a function holds a connection. This guide shows how to build SSE that works within those rules — streams that hand over before the limit, events delivered through an external pub/sub service, and a cost model you can check before launch.

Symptom & Developer Intent Permalink to this section

  • Streams die after a fixed number of seconds that matches the plan’s function timeout.
  • Events published by one request never reach streams served by other instances.
  • The monthly bill rose sharply after a live feature launched.
  • Local development works indefinitely; production cuts off.
  • Some platform regions or runtimes buffer the response until the function ends.

The intent is a streaming endpoint on the platform that reconnects invisibly at each limit, delivers every event regardless of instance, and has a known cost per connected user.

Root Cause Analysis Permalink to this section

Serverless functions are billed and bounded by duration. A stream is a function invocation that runs for as long as the client stays connected, so it hits the platform’s maximum duration and is terminated. Each invocation may run on a different instance with its own memory, so an in-process event emitter cannot fan out across users.

What changes when SSE runs on a serverless platform Two panels contrasting a long-lived server process with a serverless function for serving SSE. What changes when SSE runs on a serverless platform Long-lived server stream lasts hours in-memory fan-out works cost is per server reconnects are rare Serverless function stream ends at max duration fan-out needs external pub/sub cost is per stream-second reconnect at every limit
Nothing about the protocol changes. Duration limits and instance isolation change the architecture around it.

Limits vary by platform, runtime (Node.js serverless versus edge) and plan, and they change over time; check current documentation for your account. The design below does not depend on the exact numbers.

Step-by-Step Resolution Permalink to this section

Step 1 — Stream a Web-standard response Permalink to this section

Both platforms accept a Response with a ReadableStream body from their modern function APIs:

// Vercel: app/api/stream/route.ts (Next.js) or api/stream.ts
// Netlify: netlify/functions/stream.mts with the same handler signature
export const config = { path: '/api/stream' };          // Netlify routing; ignored by Vercel

export default async function handler(req: Request): Promise<Response> {
  const enc = new TextEncoder();
  const lastId = req.headers.get('last-event-id');
  let stop = () => {};

  const body = new ReadableStream({
    async start(controller) {
      const send = (s: string) => controller.enqueue(enc.encode(s));
      send('retry: 1000\n\n');
      for (const e of await replayAfter(lastId)) send(frame(e));
      stop = await subscribe((e) => send(frame(e)));          // external pub/sub (step 3)
      const hb = setInterval(() => send(': hb\n\n'), 15_000);
      const handover = setTimeout(() => {                       // step 2
        send('retry: 250\n: handover\n\n');
        clearInterval(hb); stop(); controller.close();
      }, HANDOVER_MS);
      req.signal.addEventListener('abort', () => { clearInterval(hb); clearTimeout(handover); stop(); });
    },
    cancel() { stop(); },
  });

  return new Response(body, {
    headers: { 'Content-Type': 'text/event-stream', 'Cache-Control': 'no-cache, no-transform' },
  });
}

Step 2 — Hand over before the platform’s limit Permalink to this section

Set HANDOVER_MS a few seconds below the configured maximum duration for the function. At the handover the stream sends a short retry: value and ends; the browser reconnects almost immediately with Last-Event-ID, and the new invocation replays anything published during the gap. Where the platform lets you raise the maximum duration per function, raise it for the stream route to reduce handovers.

Stream lifecycle on a function with a 60-second limit Timeline of three minutes showing a sequence of function invocations each ending a few seconds before a 60-second limit, with reconnects and replay in between. Stream lifecycle on a function with a 60-second limit Invocations stream 1 stream 2 stream 3 0 36 72 108 144 180 seconds handover handover handover
The user sees a continuous feed. Each invocation is short-lived by design, and the replay closes each quarter-second gap.

Step 3 — Use an external pub/sub service for fan-out Permalink to this section

Each invocation must receive events published anywhere. Options that work from serverless and edge runtimes include managed Redis with an HTTP or WebSocket-based client, managed pub/sub services with HTTP streaming or WebSocket subscriptions, and database change feeds. Requirements:

  • The client library must work in the function’s runtime (edge runtimes lack raw TCP sockets in many cases).
  • Subscribing must be fast, since it happens on every connect and every handover.
  • A replay API is strongly preferred, so the stream can serve Last-Event-ID without a separate store.
async function subscribe(onEvent: (e: Evt) => void) {
  const sub = await pubsub.subscribe(`user:${userId}`, onEvent);   // provider SDK
  return () => sub.unsubscribe();
}

Step 4 — Estimate cost before launch Permalink to this section

Duration-billed platforms charge for wall-clock time a function is running, including idle time spent waiting for events (some platforms bill active CPU rather than wall time for certain runtimes; check which applies). For a stream, that makes the cost proportional to connected user-hours:

monthly stream cost ≈ concurrent streams × hours per month × cost per function-hour
                      + invocations (one per handover) × cost per invocation
Relative monthly cost for 5,000 concurrent streams Bar chart comparing the relative monthly cost of holding five thousand concurrent SSE streams on wall-time-billed functions, CPU-time-billed edge functions and a small always-on server pool. Relative monthly cost for 5,000 concurrent streams Wall-time billed functions ~14× CPU-time billed edge ~2.5× Small always-on pool 1× relative monthly cost (always-on pool = 1)
Illustrative ratios only — pricing differs by plan and changes over time. The shape is typical: wall-time billing is expensive for idle streams.

If the estimate is uncomfortable, keep the web app on the platform and run the streaming endpoint on a small long-lived service or an edge platform designed for connections, such as the approach in fanning out SSE with Durable Objects.

Validation & Monitoring Permalink to this section

# Confirm streaming is not buffered on the deployed URL, and time the handover.
curl -sN https://myapp.example.com/api/stream | while IFS= read -r l; do echo "$(date +%T) $l"; done

Deploy to a preview environment and watch: events should arrive one by one, a : handover comment should appear just before the limit, and the next connection should begin within a second. In the platform’s logs, invocation duration for the stream route should cluster just under the handover time; invocations ending exactly at the maximum mean the handover is not firing and the platform is killing the function.

Production Checklist Permalink to this section

Frequently Asked Questions Permalink to this section

Can I avoid the duration limit entirely?

Not on serverless functions; the limit is intrinsic to the model. Planned handovers make it invisible to users. For truly long-lived connections without reconnects, use a long-running service or a connection-oriented edge platform.

Should I use edge functions or Node.js functions?

Edge functions typically start faster and may bill differently for idle time, which suits streaming. Node.js functions support more libraries, including TCP-based broker clients. Choose based on the pub/sub client you need and the cost model.

Why do events not reach streams on other instances?

Each invocation has its own memory. An in-process emitter only reaches streams in the same instance. Publish and subscribe through an external service so every invocation sees every event.

Is the reconnect at each handover visible to users?

With a short retry value and replay from Last-Event-ID, the gap is typically a fraction of a second and no events are lost. Connection indicators should not flicker for planned handovers; debounce them by a second or two.