Sharded Pub/Sub in Redis Cluster for SSE Permalink to this section

Part of Redis Pub/Sub Fan-Out for SSE, under Backend Stream Generation & Connection Management.

Classic Redis pub/sub on a cluster has a scaling property that surprises teams: every PUBLISH is broadcast to every node in the cluster, so adding shards adds load instead of spreading it. Redis 7 introduced sharded pub/sub — SPUBLISH and SSUBSCRIBE — where each channel lives on the shard that owns its hash slot, just like a key. For an SSE fan-out tier with many channels, that turns the cluster’s pub/sub capacity from “one node’s worth, copied everywhere” into “the sum of the shards”. This guide migrates an SSE service to sharded pub/sub and handles the operational details that come with it.

Symptom & Developer Intent Permalink to this section

  • Redis Cluster CPU is high on every node, even though each node owns only part of the data.
  • Adding shards to the cluster does not increase pub/sub throughput.
  • The cluster bus carries far more traffic than the application publishes.
  • After a failover, some SSE nodes stop receiving events for certain channels until restarted.
  • Subscribing to thousands of per-user channels through one connection overloads a single Redis node.

The intent is pub/sub throughput that scales with shard count, subscriptions spread across shards, and fan-out that recovers automatically when a shard fails over.

Root Cause Analysis Permalink to this section

In Redis Cluster, classic PUBLISH does not know which node a subscriber is connected to, so the message is propagated across the cluster bus to every node, and each node delivers it to its local subscribers. With N nodes, every publish costs N deliveries internally.

Classic versus sharded pub/sub on a cluster Two panels comparing classic PUBLISH, which is broadcast to every cluster node, with SPUBLISH, which goes only to the shard owning the channel's slot. Classic versus sharded pub/sub on a cluster Classic PUBLISH sent to every node cluster bus traffic × N subscribe on any node no gain from more shards Sharded SPUBLISH sent to one shard only slot = CRC16(channel) subscribe on owning shard scales with shard count
Sharded pub/sub treats a channel like a key. Throughput now scales with shards, at the cost of subscribing on the right node.

With sharded pub/sub, a channel hashes to a slot exactly like a key, and messages stay on the shard owning that slot (and its replicas). Subscribers must connect to that shard. Client libraries with cluster support route SSUBSCRIBE to the right node automatically, but they must also re-subscribe when slots move — during resharding or failover — which is where the “stops receiving after failover” symptom comes from with older or misconfigured clients.

Step-by-Step Resolution Permalink to this section

Per-user channels spread naturally across slots: sse:{user:42} and sse:{user:43} hash differently. Hash tags (the part in braces) control placement. Use them to keep a user’s channels on one shard, not to force everything onto one slot:

// Good: one channel per user, spread across the cluster by the hash tag.
const channel = (userId) => `sse:{u${userId}}`;
// Bad: a single hash tag for everything puts all traffic on one shard.
const bad = (userId) => `{sse}:user:${userId}`;

Step 2 — Publish with SPUBLISH Permalink to this section

import { Cluster } from 'ioredis';
const cluster = new Cluster([{ host: 'redis-0', port: 6379 }], { scaleReads: 'slave' });

await cluster.spublish(channel(userId), JSON.stringify(evt));   // routed to the owning shard

Step 3 — Subscribe with SSUBSCRIBE on each SSE node Permalink to this section

const sub = new Cluster([{ host: 'redis-0', port: 6379 }], {
  shardedSubscribers: true,          // one connection per shard for sharded subscriptions (ioredis 5.4+)
});

sub.on('smessage', (ch, msg) => {
  const userId = ch.slice(5, -1).replace(/^u/, '');   // "sse:{u42}" → "42"
  hub.deliver(userId, msg);
});

export async function onUserConnected(userId) {
  if (hub.localCount(userId) === 1) await sub.ssubscribe(channel(userId));   // first stream here
}
export async function onUserDisconnected(userId) {
  if (hub.localCount(userId) === 0) await sub.sunsubscribe(channel(userId));
}

Each SSE node subscribes only to channels of users connected to it, and only once per user regardless of how many tabs they have open. Check your client library’s documentation for sharded subscription support and its reconnection behaviour; the option name above is ioredis-specific.

Sharded channels spread across the cluster A publisher sends SPUBLISH to channels that hash to three different shards; each SSE node holds sharded subscriptions only on the shards owning its users' channels. Sharded channels spread across the cluster Publisher SPUBLISH sse:{u42} Slot routing CRC16 → shard route Shard A slots 0–5460 Shard B slots 5461–10922 Shard C slots 10923–16383 owning shard
Each shard handles only its channels. An SSE node talks to a shard only if one of its users' channels lives there.

Step 4 — Handle slot migration and failover Permalink to this section

When a shard fails over or slots migrate, the owning node for a channel changes. Redis notifies subscribers by sending an sunsubscribe for the affected channels; a correct client re-issues SSUBSCRIBE against the new owner. Make the application robust even if the client does not:

sub.on('sunsubscribe', async (ch) => {
  const userId = ch.slice(5, -1).replace(/^u/, '');
  if (hub.localCount(userId) > 0) {
    await sub.ssubscribe(ch).catch((e) => log.warn({ e, ch }, 'resubscribe failed'));
  }
});

Events published during the switch can be missed by the live path. That is the same situation as any pub/sub gap, and the same remedy applies: clients reconnect or receive a nudge, and replay from a durable store such as a Redis Stream.

Step 5 — Keep broadcast channels few Permalink to this section

Some events go to everyone — maintenance notices, global announcements. A single sharded channel for them concentrates that load on one shard, which is fine for low-rate broadcasts. For high-rate global feeds, publish to K channels (sse:{global:0} … sse:{global:K-1}) and have each SSE node subscribe to one chosen by its own id, spreading the fan-in across shards.

Validation & Monitoring Permalink to this section

# Where does a channel live?
redis-cli -c CLUSTER KEYSLOT 'sse:{u42}'
redis-cli -c CLUSTER NODES | grep master

# Sharded subscription counts per shard (run against each master).
redis-cli -h redis-0 PUBSUB SHARDNUMSUB 'sse:{u42}'
redis-cli -h redis-0 PUBSUB SHARDCHANNELS 'sse:*' | wc -l
Per-shard CPU for the same publish load on a six-shard cluster Bar chart comparing per-shard CPU for 50,000 publishes per second using classic PUBLISH versus sharded SPUBLISH across six shards. Per-shard CPU for the same publish load on a six-shard cluster Classic PUBLISH ~78 % Sharded SPUBLISH ~14 % approximate CPU per shard at 50,000 publishes per second
Classic pub/sub makes every shard handle every message. Sharded pub/sub divides the work.

Run a failover drill in staging (CLUSTER FAILOVER on a replica) while a test client streams, and confirm events resume within a second or two and that the client’s replay covers anything published during the switch.

Production Checklist Permalink to this section

Frequently Asked Questions Permalink to this section

Do I need sharded pub/sub if I run a single Redis primary?

No. Sharded pub/sub only changes behaviour on Redis Cluster. On a standalone primary with replicas, classic pub/sub is equivalent.

Can I mix PUBLISH and SPUBLISH?

They are separate namespaces: SPUBLISH messages reach only SSUBSCRIBE subscribers, and PUBLISH only SUBSCRIBE and PSUBSCRIBE subscribers. Migrate publishers and subscribers of a channel together, or publish to both during a transition.

Is there a pattern subscription for sharded channels?

No. There is no sharded equivalent of PSUBSCRIBE, because a pattern could match channels on every shard. Subscribe to concrete channels, which suits per-user SSE routing anyway.

Do replicas deliver sharded messages?

Yes. Sharded messages are propagated to the shard's replicas, so subscribers may connect to replicas to spread load, as long as the client routes them to the right shard.