Keeping Last-Event-ID Valid Across Deploys Permalink to this section

Part of Event ID & Retry Mechanism Design, under SSE Protocol Fundamentals & Architecture.

Last-Event-ID is a promise: the client hands back the last id it saw and trusts the server to know where that is. Deploys are where the promise is most often broken. A new process starts its counter from zero, a replay buffer held in memory disappears, an id format changes between versions, or a reconnect lands on a node that never saw the old ids. The client either misses events or is sent everything again. This guide makes ids durable across every kind of restart and gives the server a safe answer when it genuinely cannot honour one.

Symptom & Developer Intent Permalink to this section

  • After every deploy, some users miss the events published during the rollout.
  • After a restart, clients receive a flood of duplicate events.
  • Event ids jump backwards after a deploy (from 58,211 to 3), and clients ignore new events as “already seen”.
  • A reconnect to a different node returns nothing, because that node’s ids do not match.
  • A new version changed the id format and older clients break on reconnect.

The intent is that an id issued by any node, before or after any deploy, is either resolved correctly by any node afterwards or answered with an explicit resync.

Root Cause Analysis Permalink to this section

Three things must survive a deploy for resumption to work: the id sequence (new ids must be greater than old ones), the history (events after the client’s id must still be retrievable), and the format (new code must parse old ids).

What each restart can destroy Three panels listing how in-memory counters, in-memory replay buffers and changing id formats each break Last-Event-ID across a deploy. What each restart can destroy In-memory counter restarts from zero new ids < old ids clients drop new events In-memory buffer history gone on restart replay returns nothing events silently missed Changed id format old ids unparseable server errors or resets needs a version prefix
All three must live outside the process, or be versioned, for ids to outlive the code that issued them.

Per-node state has the same flaw at a smaller scale: a client reconnecting to another node presents an id that node’s memory has never seen.

Step-by-Step Resolution Permalink to this section

Step 1 — Take ids from a durable, shared source Permalink to this section

// Options, from simplest to most scalable:
// 1. Database sequence / identity column (the event row's primary key).
// 2. Redis INCR on a per-stream key, or Redis Stream entry ids.
// 3. Broker positions (Kafka offsets, NATS JetStream sequences).
const id = await redis.incr(`seq:${stream}`);         // survives restarts, shared by all nodes

Never derive ids from process uptime, a per-process counter or an in-memory array index. Timestamps are tempting but collide within a millisecond and go backwards when clocks are corrected; if used, combine with a sequence (ms-seq, as Redis Streams do).

Step 2 — Store history outside the process Permalink to this section

The replay window must live where every node can read it: the database table the events came from, a Redis Stream, or the broker’s retained log. A reconnect then works from any node and after any deploy. Node-local memory can still serve as a cache in front of the shared store, but never as the only copy. Replaying missed events with Redis Streams is a compact implementation.

Step 3 — Version the id format Permalink to this section

Prefix ids with a format version so that future changes remain parseable:

// v2 ids: "2.<stream-seq>"; v1 ids were bare integers.
function parseId(raw) {
  if (!raw) return null;
  if (/^\d+$/.test(raw)) return { v: 1, seq: Number(raw) };        // legacy
  const m = raw.match(/^2\.(\d+)$/);
  if (m) return { v: 2, seq: Number(m[1]) };
  return { v: 0 };                                                  // unknown: resync
}

Keep parsers for old versions for at least as long as the replay window, so a client that reconnects just after the deploy still resumes.

A reconnect during a rolling deploy Sequence diagram of a client connected to an old-version node that restarts, reconnecting to a new-version node with a v1 id, which parses it, replays from the shared store and continues with v2 ids. A reconnect during a rolling deploy Browser Old node v1 New node v2 Shared store id: 58211 node restarts (deploy) reconnect Last-Event-ID: 58211 events after seq 58211 58212 … 58219 id: 2.58212 … 2.58219
The new code reads the old id, finds the history in the shared store, and continues the sequence. The client notices nothing.

Step 4 — Answer unknown or expired ids with a resync Permalink to this section

Whatever the server cannot resolve — an unparseable id, an id older than retention, an id from a future the server does not know (a clock or sequence reset) — must produce an explicit signal, not silence:

const parsed = parseId(req.get('Last-Event-ID'));
if (parsed && (parsed.v === 0 || parsed.seq < await oldestRetainedSeq(stream) || parsed.seq > await currentSeq(stream))) {
  res.write('event: resync\ndata: {"reason":"unknown-cursor"}\n\n');   // client refetches state
} else if (parsed) {
  for (const e of await store.after(stream, parsed.seq)) res.write(frame(e));
}

The “greater than current” check catches a sequence that was reset — for example, a Redis key lost in a failover without persistence. Without it, the server would wait for the sequence to climb past the client’s id while the client drops every event as old.

Step 5 — Drain before shutdown so reconnects are orderly Permalink to this section

Send a short retry: hint and end each stream during shutdown, so clients reconnect promptly to healthy nodes instead of waiting for TCP timeouts, and so the reconnect wave is spread out. Rebalancing SSE connections after a deploy covers draining in detail.

Validation & Monitoring Permalink to this section

# Restart test: note the last id, restart the service, reconnect with it.
last=$(curl -sN https://app.example.com/api/stream | grep -m5 '^id:' | tail -1 | cut -d' ' -f2)
kubectl rollout restart deploy/sse && kubectl rollout status deploy/sse
curl -sN -H "Last-Event-ID: $last" https://app.example.com/api/stream | grep -m3 -E '^(id|event):'
# Expect ids greater than $last, or an explicit resync event — never silence.
Events missed per client during a rolling deploy Bar chart comparing average events missed per connected client during a deploy for in-memory ids and buffers, durable ids with in-memory buffers, and durable ids with a shared replay store. Events missed per client during a rolling deploy In-memory ids + buffer 23 Durable ids, local buffer 9 Durable ids + shared store 0 average events missed per client during a 4-minute rollout
Durable ids stop the duplicates and the drops caused by resets; only a shared replay store closes the gap entirely.

Monitor resync events by reason during and after deploys. A burst of unknown-cursor resyncs at a deploy means an id format or sequence source changed incompatibly.

Production Checklist Permalink to this section

Frequently Asked Questions Permalink to this section

Can I use UUIDs as event ids?

They are unique but not ordered, so the server cannot find "events after this id" without an index from UUID to position. Use ordered ids, or keep such an index in the replay store.

What happens if the id sequence resets?

New ids are smaller than those clients hold, so clients that compare ids ignore new events, and the server finds nothing after the client's id. Detect ids greater than the current sequence and answer with resync.

Do I need to keep old id parsers forever?

Only as long as a client could still present an old id and have it be within retention — the replay window plus a margin. After that, unknown ids fall through to resync, which is always safe.

Is Last-Event-ID sent by every client?

EventSource sends it automatically on reconnect when an id was received. A fresh page load starts a new EventSource and does not send it; pass a cursor as a query parameter if a reload should resume.