Keep-Alive and Disconnects in Axum SSE Permalink to this section

Part of Rust SSE with Axum and Tokio, under Backend Stream Generation & Connection Management.

In axum, a client disconnect has no callback. The response body is dropped, your stream is dropped with it, and anything the stream owns is released by Rust’s ordinary Drop. That is elegant when it works and baffling when it does not: per-connection state that lives outside the stream is never cleaned up, and a client that vanished without closing its connection is never dropped at all until something tries to write to it. This guide wires keep-alive, cleanup and shutdown together so every connection’s resources are released promptly, however the client leaves.

Symptom & Developer Intent Permalink to this section

  • A gauge of connected users rises all day and only resets on deploy.
  • Presence shows users online long after they left.
  • tokio-console shows thousands of tasks idle for hours.
  • Graceful shutdown hangs until the orchestrator kills the process.
  • Behind a load balancer, streams die every 60 seconds when nothing is happening.

The intent is that per-connection state is released within about a minute of any departure, idle streams survive intermediaries, and deploys finish in seconds.

Root Cause Analysis Permalink to this section

Hyper learns that a peer is gone in two ways: the client closes the connection (a FIN or RST arrives), or a write fails. A laptop that sleeps or a phone that loses signal sends nothing, so on an idle stream neither happens, and the body — and your stream — lives on.

What gets dropped when Hyper notices a disconnect Stack of the connection, the Sse response body, the event stream, and a guard value owned by the stream, showing that only state owned by the stream is released automatically. What gets dropped when Hyper notices a disconnect TCP connection closed or write fails detected by Hyper Sse body dropped automatic Your Stream dropped automatic ConnGuard Drop runs releases external state
Anything reachable from the stream is released by Drop. Anything registered elsewhere — a map entry, a Redis lease — needs a guard owned by the stream to release it.

Keep-alive comments solve detection: they create regular writes, and writes to a dead peer eventually fail. They also solve the idle timeouts of proxies and load balancers, which close connections that carry no bytes for 60 seconds or so. Graceful shutdown hangs for the opposite reason: with_graceful_shutdown waits for in-flight responses to finish, and an infinite stream never finishes unless something ends it.

Step-by-Step Resolution Permalink to this section

Step 1 — Enable keep-alive with a deliberate interval Permalink to this section

use axum::response::sse::{KeepAlive, Sse};
use std::time::Duration;

Sse::new(stream).keep_alive(
    KeepAlive::new()
        .interval(Duration::from_secs(15))   // under every common idle timeout
        .text("hb"),                          // appears as ": hb" on the wire
)

Axum only sends the comment when the stream itself has been quiet for the interval, so busy streams pay nothing. Fifteen seconds keeps most proxies happy; see choosing heartbeat intervals for the trade-offs.

Step 2 — Own external state with a guard Permalink to this section

struct ConnGuard {
    user: String,
    registry: Registry,
}

impl ConnGuard {
    fn new(user: String, registry: Registry) -> Self {
        registry.add(&user);
        metrics::gauge!("sse_open_streams").increment(1.0);
        Self { user, registry }
    }
}

impl Drop for ConnGuard {
    fn drop(&mut self) {
        self.registry.remove(&self.user);            // synchronous, cheap
        metrics::gauge!("sse_open_streams").decrement(1.0);
    }
}

Move the guard into the stream so the two share a lifetime:

async fn stream(State(st): State<AppState>, user: AuthUser) -> Sse<impl Stream<Item = Result<Event, Infallible>>> {
    let guard = ConnGuard::new(user.id.clone(), st.registry.clone());
    let rx = st.topics.subscribe(&format!("user:{}", user.id));
    let s = BroadcastStream::new(rx).filter_map(move |m| {
        let _keep = &guard;                          // the closure owns the guard
        async move { m.ok().map(|f| Ok(to_event(&f))) }
    });
    Sse::new(s).keep_alive(KeepAlive::default())
}

Drop cannot be async. If cleanup needs I/O — deleting a Redis presence key, publishing a “left” event — spawn it from drop, or better, rely on expiring leases that the stream renews, as described in showing who is online with SSE presence.

impl Drop for PresenceGuard {
    fn drop(&mut self) {
        let (redis, key) = (self.redis.clone(), self.key.clone());
        tokio::spawn(async move { let _ = redis.del::<_, ()>(key).await; });   // best effort
    }
}
A vanished client is dropped after keep-alive writes fail Sequence diagram of axum sending keep-alive comments to a client that has disappeared, the writes eventually failing, Hyper dropping the body, and the guard's Drop releasing state. A vanished client is dropped after keep-alive writes fail Axum stream Hyper Vanished client Guard : hb no ACK, retransmits : hb write error → drop body Drop::drop
Without keep-alive there is no write to fail, and the drop in the last step never happens.

Step 3 — End streams on shutdown Permalink to this section

let (shutdown_tx, shutdown_rx) = tokio::sync::watch::channel(false);

// In each handler: end the stream when shutdown is signalled.
let mut rx = st.shutdown.clone();
let s = s.take_until(async move { let _ = rx.changed().await; });

// In main:
axum::serve(listener, app)
    .with_graceful_shutdown(async move {
        tokio::signal::ctrl_c().await.ok();
        let _ = shutdown_tx.send(true);              // every stream ends now…
    })
    .await?;                                         // …so this returns promptly

Before ending, a stream may emit a final retry: hint so clients spread their reconnects. With the stream finished, the response completes, the guard drops, and graceful shutdown can return.

Step 4 — Tune TCP keepalive as a backstop Permalink to this section

Application keep-alive handles almost everything. As an extra safety net for connections that stop acknowledging, enable TCP keepalive on accepted sockets with socket2 or a custom listener; it catches peers that vanished while the stream was mid-write and the kernel is still retransmitting.

On Linux, a write to a peer that stopped acknowledging does not fail immediately: the kernel keeps retransmitting for up to tcp_retries2 attempts, which by default adds up to roughly fifteen minutes. The TCP_USER_TIMEOUT socket option caps that period, so unacknowledged data fails the connection after, say, 30 seconds. Combined with a 15-second keep-alive comment, that bounds how long a vanished client can hold its resources:

use socket2::{SockRef, TcpKeepalive};

let listener = tokio::net::TcpListener::bind("0.0.0.0:8080").await?;
loop {
    let (sock, _) = listener.accept().await?;
    let s = SockRef::from(&sock);
    s.set_tcp_keepalive(&TcpKeepalive::new().with_time(Duration::from_secs(60)))?;
    #[cfg(target_os = "linux")]
    s.set_tcp_user_timeout(Some(Duration::from_secs(30)))?;   // fail unacked writes after 30 s
    // hand `sock` to hyper-util's connection builder with the axum service…
}

When a load balancer terminates client connections, these settings govern the hop from the balancer to your process instead; the balancer’s own idle and health settings then decide how quickly a vanished browser is noticed, which is another reason the application-level comment must be frequent enough to keep the balancer’s view fresh.

Validation & Monitoring Permalink to this section

# Idle survival: the stream stays open past a 60 s proxy timeout.
curl -sN http://localhost:8080/api/stream | grep --line-buffered '^:' | while read l; do date +%T; done

# Vanished client: drop outbound packets to the test client, then watch the gauge fall.
sudo iptables -A OUTPUT -p tcp --sport 8080 -d 10.0.0.50 -j DROP
curl -s localhost:9000/metrics | grep sse_open_streams

Use tokio-console during the test: the task for the vanished connection should disappear shortly after the writes start failing. Remove the iptables rule afterwards.

Time for a deploy's graceful shutdown to complete Bar chart comparing shutdown duration for an axum server with 5,000 open streams, with and without ending streams on the shutdown signal. Time for a deploy's graceful shutdown to complete No stream termination 30 s (killed) take_until(shutdown) 0.4 s seconds from SIGTERM to process exit, 5,000 open streams
Graceful shutdown waits for every response. Ending streams on the signal turns a forced kill into a clean, sub-second exit.

Production Checklist Permalink to this section

Frequently Asked Questions Permalink to this section

Is there an on-disconnect callback in axum?

No. Disconnection drops the response body and your stream. Put cleanup in Drop implementations of values the stream owns.

Why does my stream stay alive after the client disappeared?

Because nothing was written to it. Without writes, Hyper cannot learn the peer is gone. Enable keep-alive so there is a periodic write that eventually fails.

Can Drop run async cleanup?

Not directly. Spawn a task from drop for best-effort async work, and design shared state such as presence to expire on its own so a missed cleanup cannot leave it wrong forever.

How do I know which connections are idle right now?

Record the time of the last real event in the guard and export a histogram of idle durations, or inspect tasks with tokio-console. Long-idle streams are normal for quiet feeds; long-idle streams that never receive keep-alive writes indicate keep-alive is not configured on that route.

Does the keep-alive comment reach EventSource listeners?

No. Comment lines are consumed by the parser and never dispatched, so clients are unaffected apart from the connection staying open.