Resuming an Interrupted AI Stream Permalink to this section
Part of Streaming AI Responses in the Browser, under Frontend Consumption & Client Patterns.
A long AI answer can take a minute or more to generate. Over that minute, a phone changes networks, a laptop sleeps, a user reloads the page or switches tabs. In most implementations, any of those ends the answer: the stream breaks, the generation is lost or continues unseen, and the user has to ask again — paying twice. The fix is architectural: the generation must outlive the HTTP request that started it, and the stream must be a view onto the generation’s output that can be reopened from any point. This guide builds that.
Symptom & Developer Intent Permalink to this section
- When the network blips mid-answer, the text stops and a “network error” appears.
- Reloading the page during an answer loses it entirely, although the provider charged for it.
- Retrying after a drop produces a different answer, which confuses users who read half of the first.
- Opening the conversation in a second tab shows nothing until the answer finishes.
- Server logs show generations completing after their client disconnected, with the output discarded.
The intent is that a started answer is generated exactly once, and any client — the same tab after a drop, a reloaded page, another device — can attach to it and see the whole answer, live.
Root Cause Analysis Permalink to this section
In the naive design, the generation runs inside the request handler and writes tokens straight to the response. The request’s lifetime is the generation’s lifetime: when the connection drops, either the generation is aborted (losing work) or continues with nowhere to write (losing output). Neither allows reattachment, because the tokens exist nowhere but in the socket.
Decoupling separates three things: starting a generation (a POST that returns an id), running it (a background job that appends each chunk to a log with a sequence number), and viewing it (a GET stream that replays the log from a position and then follows it live).
Step-by-Step Resolution Permalink to this section
Step 1 — Start the generation as a job and return its id Permalink to this section
app.post('/api/conversations/:cid/messages', requireUser, async (req, res) => {
const gen = await generations.create({ user: req.user.id, conversation: req.params.cid,
prompt: req.body.text, idempotencyKey: req.get('Idempotency-Key') });
runGeneration(gen).catch((err) => log.error({ err, gen: gen.id }, 'generation failed')); // not awaited
res.status(202).json({ generationId: gen.id, streamUrl: `/api/generations/${gen.id}/stream` });
});
The idempotency key ensures a retried POST returns the same generation rather than starting another, as covered in sending POST requests that return SSE.
Step 2 — Append every chunk to a durable token log Permalink to this section
async function runGeneration(gen) {
const key = `gen:${gen.id}`;
let seq = 0;
const ac = registerCancellation(gen.id); // explicit cancel can still stop it
for await (const chunk of model.stream({ prompt: gen.prompt, signal: ac.signal })) {
await redis.xadd(key, 'MAXLEN', '~', 20000, `${Date.now()}-${++seq}`, 'type', 'token', 'd', chunk.text);
}
await redis.xadd(key, '*', 'type', 'done', 'd', '');
await redis.expire(key, 24 * 3600); // keep for a day, then the stored message suffices
await messages.saveFinal(gen); // persist the complete answer
}
A Redis Stream is a convenient log: entries are ordered, ids are seekable, and readers can block for new entries. A database table keyed by (generation_id, seq) works too.
Step 3 — Serve a resumable stream from the log Permalink to this section
app.get('/api/generations/:id/stream', requireUser, async (req, res) => {
const gen = await generations.getForUser(req.params.id, req.user.id);
if (!gen) return res.sendStatus(404);
if (req.get('Last-Event-ID') === 'done') return res.sendStatus(204); // client already has the end
openStream(res, { retryMs: 1000 });
let cursor = req.get('Last-Event-ID') || '0';
const reader = redis.duplicate();
req.on('close', () => reader.disconnect());
for (;;) {
const out = await reader.xread('BLOCK', 15000, 'STREAMS', `gen:${gen.id}`, cursor);
if (!out) { res.write(': hb\n\n'); continue; }
for (const [id, fields] of out[0][1]) {
const f = Object.fromEntries(pairs(fields));
cursor = id;
if (f.type === 'done') { res.end('id: done\nevent: done\ndata: {}\n\n'); return; }
res.write(`id: ${id}\nevent: token\ndata: ${JSON.stringify({ t: f.d })}\n\n`);
}
}
});
Starting from cursor 0 replays the entire answer so far; starting from the browser’s Last-Event-ID resumes exactly after the last token it received.
Step 4 — Reattach after reload and on other devices Permalink to this section
The client stores the active generation id with the conversation (in the URL or in the conversation’s server-side state). On page load, for any message whose generation is still running, open its stream from 0:
for (const m of conversation.messages) {
if (m.status === 'generating') attach(m.generationId, { from: '0' }); // replays so far, then live
}
Because the stream is a view onto a log, a second tab or a phone opened mid-answer does exactly the same and sees the answer build in real time.
Step 5 — Keep cancellation explicit Permalink to this section
Decoupling means a closed tab no longer stops generation. That is usually desirable — the user returns to a finished answer — but the stop button must still stop it. Keep the explicit cancel endpoint from cancelling an in-flight AI stream, wired to the job’s abort controller.
Validation & Monitoring Permalink to this section
- Start a long answer; set DevTools to Offline for ten seconds; restore. The answer continues without gaps or repeated text.
- Reload mid-answer; the partial answer reappears and continues.
- Open the conversation on a second device mid-answer; both show the same text as it grows.
- Check provider usage: one generation per user message.
Monitor resumes per generation, replayed tokens per resume and generations still running with no attached viewer; the last is normal briefly, but a large backlog suggests viewers are failing to reattach.
Production Checklist Permalink to this section
Frequently Asked Questions Permalink to this section
Does this double the storage or cost of every answer?
The token log is small and short-lived compared with the model cost it protects. Most answers are a few kilobytes, kept for hours, then replaced by the stored final message.
Can I resume from the model provider instead?
Some providers offer background or resumable responses, which can play the role of the token log. The browser-facing design is the same: a stable id and a resumable stream from your backend.
What if the job itself crashes?
Mark the generation failed from a supervisor or heartbeat check and emit a terminal error event to the log, so attached viewers end cleanly and can offer to retry.
Should a closed tab cancel the generation?
That is a product choice. Continuing lets users return to a finished answer; cancelling saves cost for answers nobody will read. Make it explicit either way.