Sending POST Requests That Return SSE Permalink to this section
Part of Fetch-Based SSE Clients, under Frontend Consumption & Client Patterns.
AI chat, search-as-you-type, report generation and many other features send a request body — a prompt, a query, a set of options — and receive a stream of results. EventSource cannot send a body, so these endpoints are consumed with fetch. That is straightforward until the connection drops: a naive retry sends the POST again and starts the work twice, doubling cost and producing two diverging answers. This guide builds POST streaming that is safe to retry, resumable and cancellable.
Symptom & Developer Intent Permalink to this section
- After a network blip, the chat shows two different answers interleaved, or the provider bill shows duplicate generations.
- A long answer stops half-way when the phone switches networks, and the only option is to ask again.
- Pressing stop closes the stream, but the server keeps generating.
- Putting the prompt in a query string for
EventSourcehits URL length limits and leaks prompts into logs.
The intent is one generation per user action regardless of retries, a stream that resumes after drops, and a stop that actually stops the work.
Root Cause Analysis Permalink to this section
A POST is not idempotent by default: each request is a new instruction to do work. When a streamed POST fails part-way, the client cannot tell whether the server received it, started work, or finished. Retrying the POST blindly risks duplicate work; not retrying loses the result.
The fix separates starting work from reading its output. The POST carries a client-generated idempotency key and returns an identifier for the work; reading the output is a GET that can be retried freely and resumed from the last event id.
Step-by-Step Resolution Permalink to this section
Step 1 — Send the POST with an idempotency key Permalink to this section
async function startGeneration(prompt, { signal }) {
const key = crypto.randomUUID(); // one per user action, reused on retry
const res = await fetch('/api/generations', {
method: 'POST',
headers: {
'Content-Type': 'application/json',
'Accept': 'text/event-stream',
'Idempotency-Key': key,
},
body: JSON.stringify({ prompt }),
credentials: 'include',
signal,
});
return { res, key };
}
Step 2 — Make the server honour the key and announce the generation id Permalink to this section
app.post('/api/generations', requireUser, express.json(), async (req, res) => {
const key = req.get('Idempotency-Key');
if (!key) return res.status(400).json({ error: 'Idempotency-Key required' });
// Same user + same key → same generation, never a second one.
const gen = await generations.findOrCreate(req.user.id, key, () => ({ prompt: req.body.prompt }));
res.writeHead(200, { 'Content-Type': 'text/event-stream', 'Cache-Control': 'no-cache' });
res.write(`event: meta\ndata: ${JSON.stringify({ id: gen.id })}\n\n`);
await pipeGeneration(gen, req, res, Number(req.get('Last-Event-ID')?.split(':')[1] ?? 0));
});
findOrCreate must be atomic — a unique constraint on (user_id, idempotency_key) in the database does it. The generation runs as background work owned by the server, not by the HTTP request, so a dropped connection does not kill it.
Step 3 — Resume with GET, not POST Permalink to this section
app.get('/api/generations/:id/stream', requireUser, async (req, res) => {
const gen = await generations.get(req.params.id, req.user.id);
if (!gen) return res.sendStatus(404);
res.writeHead(200, { 'Content-Type': 'text/event-stream', 'Cache-Control': 'no-cache' });
const after = Number(req.get('Last-Event-ID')?.split(':')[1] ?? 0);
await pipeGeneration(gen, req, res, after); // replay stored tokens after `after`, then follow live
});
// Client: POST once, then read; on any drop, resume with GET from the last id.
async function runGeneration(prompt, onToken, signal) {
let genId = null, lastId = '';
const onEvent = (e) => {
if (e.id) lastId = e.id;
if (e.type === 'meta') genId = JSON.parse(e.data).id;
if (e.type === 'token') onToken(JSON.parse(e.data).t);
};
try {
const { res } = await startGeneration(prompt, { signal });
await parseStream(res.body, onEvent);
} catch (err) { if (signal.aborted) return; }
while (genId && !signal.aborted && !finished) {
try {
const res = await fetch(`/api/generations/${genId}/stream`, {
headers: { 'Last-Event-ID': lastId }, credentials: 'include', signal,
});
await parseStream(res.body, onEvent);
} catch { await sleep(backoff.next(), signal); }
}
}
If the POST itself fails before the meta event arrives, retry the POST with the same idempotency key: the server returns the existing generation instead of creating a new one.
Storage makes the resume possible. The generation must write its output — tokens, partial results, the final state — somewhere the GET handler can read, with a position per event: a Redis Stream per generation, a row per chunk, or an append-only blob with byte offsets. Keep that storage for as long as a user might reasonably reconnect or reload, typically hours, then let it expire. The same storage lets a second tab, or the same user on another device, open the generation and see it complete live, which is often a welcome side effect.
Key retention needs thought too. The server must remember (user, idempotency key) → generation at least as long as a client might retry the original POST — minutes are enough for network retries, but a day is a safer default because it also covers a client that retries after being offline. After that, the same key may create new work, which is harmless because clients generate a new key for every new user action.
Step 4 — Cancel the work, not just the stream Permalink to this section
Aborting the fetch stops reading; the server should stop generating too. Send an explicit cancel, keyed by the generation id, and have the server abort the model call:
async function stop(genId, controller) {
controller.abort(); // stop reading now
await fetch(`/api/generations/${genId}/cancel`, { method: 'POST', keepalive: true, credentials: 'include' });
}
The details, including cancelling on page unload, are in cancelling an in-flight AI stream.
Validation & Monitoring Permalink to this section
# Idempotency: the same key twice yields the same generation id.
K=$(uuidgen)
for i in 1 2; do
curl -sN -X POST -H "Idempotency-Key: $K" -H 'Content-Type: application/json' \
-d '{"prompt":"hi"}' https://app.example.com/api/generations | grep -m1 -A1 'event: meta'
done
Monitor generations per user action (should be exactly one), resumes per generation, and cancels that did not stop the model within a second.
Production Checklist Permalink to this section
Frequently Asked Questions Permalink to this section
Why not put the prompt in the query string and use EventSource?
URLs have length limits, are logged by proxies and servers, and end up in browser history. Request bodies avoid all three, at the cost of using fetch.
Is the Idempotency-Key header a standard?
It is a widely used convention with an IETF draft behind it. What matters is that the server treats the key as the identity of the operation for a reasonable retention period.
What if the POST succeeds but the meta event is lost?
Retry the POST with the same key. The server finds the existing generation, returns its id in a new meta event and streams from the start or from the provided Last-Event-ID.
Should the generation continue if the user closes the tab?
That is a product decision. For chat, continuing lets the answer appear when the user returns; for expensive work nobody will read, cancel on pagehide with a beacon.