Sending POST Requests That Return SSE Permalink to this section

Part of Fetch-Based SSE Clients, under Frontend Consumption & Client Patterns.

AI chat, search-as-you-type, report generation and many other features send a request body — a prompt, a query, a set of options — and receive a stream of results. EventSource cannot send a body, so these endpoints are consumed with fetch. That is straightforward until the connection drops: a naive retry sends the POST again and starts the work twice, doubling cost and producing two diverging answers. This guide builds POST streaming that is safe to retry, resumable and cancellable.

Symptom & Developer Intent Permalink to this section

  • After a network blip, the chat shows two different answers interleaved, or the provider bill shows duplicate generations.
  • A long answer stops half-way when the phone switches networks, and the only option is to ask again.
  • Pressing stop closes the stream, but the server keeps generating.
  • Putting the prompt in a query string for EventSource hits URL length limits and leaks prompts into logs.

The intent is one generation per user action regardless of retries, a stream that resumes after drops, and a stop that actually stops the work.

Root Cause Analysis Permalink to this section

A POST is not idempotent by default: each request is a new instruction to do work. When a streamed POST fails part-way, the client cannot tell whether the server received it, started work, or finished. Retrying the POST blindly risks duplicate work; not retrying loses the result.

How a blind POST retry duplicates work Sequence diagram of a client posting a prompt, the server starting a generation, the connection dropping, and a blind retry starting a second, independent generation. How a blind POST retry duplicates work Browser API Model POST /chat {prompt} start generation A connection drops mid-stream POST /chat {prompt} (retry) start generation B generation A continues unobserved; the user sees B
Two generations, two bills, two different answers. The server had no way to know the second POST was a retry of the first.

The fix separates starting work from reading its output. The POST carries a client-generated idempotency key and returns an identifier for the work; reading the output is a GET that can be retried freely and resumed from the last event id.

Step-by-Step Resolution Permalink to this section

Step 1 — Send the POST with an idempotency key Permalink to this section

async function startGeneration(prompt, { signal }) {
  const key = crypto.randomUUID();                      // one per user action, reused on retry
  const res = await fetch('/api/generations', {
    method: 'POST',
    headers: {
      'Content-Type': 'application/json',
      'Accept': 'text/event-stream',
      'Idempotency-Key': key,
    },
    body: JSON.stringify({ prompt }),
    credentials: 'include',
    signal,
  });
  return { res, key };
}

Step 2 — Make the server honour the key and announce the generation id Permalink to this section

app.post('/api/generations', requireUser, express.json(), async (req, res) => {
  const key = req.get('Idempotency-Key');
  if (!key) return res.status(400).json({ error: 'Idempotency-Key required' });

  // Same user + same key → same generation, never a second one.
  const gen = await generations.findOrCreate(req.user.id, key, () => ({ prompt: req.body.prompt }));
  res.writeHead(200, { 'Content-Type': 'text/event-stream', 'Cache-Control': 'no-cache' });
  res.write(`event: meta\ndata: ${JSON.stringify({ id: gen.id })}\n\n`);
  await pipeGeneration(gen, req, res, Number(req.get('Last-Event-ID')?.split(':')[1] ?? 0));
});

findOrCreate must be atomic — a unique constraint on (user_id, idempotency_key) in the database does it. The generation runs as background work owned by the server, not by the HTTP request, so a dropped connection does not kill it.

Step 3 — Resume with GET, not POST Permalink to this section

app.get('/api/generations/:id/stream', requireUser, async (req, res) => {
  const gen = await generations.get(req.params.id, req.user.id);
  if (!gen) return res.sendStatus(404);
  res.writeHead(200, { 'Content-Type': 'text/event-stream', 'Cache-Control': 'no-cache' });
  const after = Number(req.get('Last-Event-ID')?.split(':')[1] ?? 0);
  await pipeGeneration(gen, req, res, after);           // replay stored tokens after `after`, then follow live
});
// Client: POST once, then read; on any drop, resume with GET from the last id.
async function runGeneration(prompt, onToken, signal) {
  let genId = null, lastId = '';
  const onEvent = (e) => {
    if (e.id) lastId = e.id;
    if (e.type === 'meta') genId = JSON.parse(e.data).id;
    if (e.type === 'token') onToken(JSON.parse(e.data).t);
  };
  try {
    const { res } = await startGeneration(prompt, { signal });
    await parseStream(res.body, onEvent);
  } catch (err) { if (signal.aborted) return; }
  while (genId && !signal.aborted && !finished) {
    try {
      const res = await fetch(`/api/generations/${genId}/stream`, {
        headers: { 'Last-Event-ID': lastId }, credentials: 'include', signal,
      });
      await parseStream(res.body, onEvent);
    } catch { await sleep(backoff.next(), signal); }
  }
}

If the POST itself fails before the meta event arrives, retry the POST with the same idempotency key: the server returns the existing generation instead of creating a new one.

Start once, read many times Flow from a user action creating an idempotency key, through a POST that finds or creates a generation, to resumable GET streams that read it from the last event id. Start once, read many times User action new key once POST + key findOrCreate returns meta: gen id first event on drop GET stream Last-Event-ID finished done generation complete
The only non-idempotent operation is protected by the key. Everything that may be retried is a GET.

Storage makes the resume possible. The generation must write its output — tokens, partial results, the final state — somewhere the GET handler can read, with a position per event: a Redis Stream per generation, a row per chunk, or an append-only blob with byte offsets. Keep that storage for as long as a user might reasonably reconnect or reload, typically hours, then let it expire. The same storage lets a second tab, or the same user on another device, open the generation and see it complete live, which is often a welcome side effect.

Key retention needs thought too. The server must remember (user, idempotency key) → generation at least as long as a client might retry the original POST — minutes are enough for network retries, but a day is a safer default because it also covers a client that retries after being offline. After that, the same key may create new work, which is harmless because clients generate a new key for every new user action.

Step 4 — Cancel the work, not just the stream Permalink to this section

Aborting the fetch stops reading; the server should stop generating too. Send an explicit cancel, keyed by the generation id, and have the server abort the model call:

async function stop(genId, controller) {
  controller.abort();                                   // stop reading now
  await fetch(`/api/generations/${genId}/cancel`, { method: 'POST', keepalive: true, credentials: 'include' });
}

The details, including cancelling on page unload, are in cancelling an in-flight AI stream.

Validation & Monitoring Permalink to this section

# Idempotency: the same key twice yields the same generation id.
K=$(uuidgen)
for i in 1 2; do
  curl -sN -X POST -H "Idempotency-Key: $K" -H 'Content-Type: application/json' \
    -d '{"prompt":"hi"}' https://app.example.com/api/generations | grep -m1 -A1 'event: meta'
done
Duplicate generations per 1,000 user actions on mobile networks Bar chart comparing duplicate generations per thousand user actions with blind POST retries, no retries, and idempotent POST plus GET resume. Duplicate generations per 1,000 user actions on mobile networks Blind POST retry: duplicates 31 No retry: lost answers 29 Key + GET resume: either 0 per 1,000 user actions on a mobile cohort
No retries avoids duplicates by losing answers. Idempotency plus GET resume avoids both.

Monitor generations per user action (should be exactly one), resumes per generation, and cancels that did not stop the model within a second.

Production Checklist Permalink to this section

Frequently Asked Questions Permalink to this section

Why not put the prompt in the query string and use EventSource?

URLs have length limits, are logged by proxies and servers, and end up in browser history. Request bodies avoid all three, at the cost of using fetch.

Is the Idempotency-Key header a standard?

It is a widely used convention with an IETF draft behind it. What matters is that the server treats the key as the identity of the operation for a reasonable retention period.

What if the POST succeeds but the meta event is lost?

Retry the POST with the same key. The server finds the existing generation, returns its id in a new meta event and streams from the start or from the provided Last-Event-ID.

Should the generation continue if the user closes the tab?

That is a product decision. For chat, continuing lets the answer appear when the user returns; for expensive work nobody will read, cancel on pagehide with a beacon.