WebSockets Cheatsheet
Scaling
Use this WebSockets reference while you build software engineering projects, review code for technical interview prep, or polish examples for a software engineer resume.
The Scaling Problem
WebSocket connections are stateful and sticky — a client connected to Server A cannot receive a message published by Server B unless the servers share state. HTTP is stateless and trivially load-balanced; WebSocket is not.
Client 1 ──── Server A ──┐ ├── need a shared message bus Client 2 ──── Server B ──┘
Redis Pub/Sub Fan-Out
The standard pattern: each server subscribes to a Redis channel and re-publishes incoming messages so all servers can broadcast to their local clients.
npm install ioredis
import { WebSocketServer, WebSocket } from "ws"; import Redis from "ioredis"; const pub = new Redis(process.env.REDIS_URL); const sub = new Redis(process.env.REDIS_URL); // separate connection for subscribe const wss = new WebSocketServer({ port: 4000 }); // Subscribe to the shared channel sub.subscribe("chat", (err) => { if (err) console.error("subscribe error:", err); }); // When Redis delivers a message, broadcast to all LOCAL clients sub.on("message", (channel, message) => { for (const client of wss.clients) { if (client.readyState === WebSocket.OPEN) { client.send(message); } } }); wss.on("connection", (ws) => { ws.on("message", (data) => { // Publish to Redis so ALL server instances get it pub.publish("chat", data.toString()); }); });
Each server instance runs this same code. A message from any client flows: client → local ws → Redis
PUBLISH→ Redis delivers to allSUBSCRIBErs → each serversend()s to its local clients.
Redis Pub/Sub with Rooms
// Room-scoped channels function getRoomChannel(roomId) { return `room:${roomId}`; } function joinRoom(ws, roomId) { ws.rooms ??= new Set(); ws.rooms.add(roomId); sub.subscribe(getRoomChannel(roomId)); } function leaveRoom(ws, roomId) { ws.rooms?.delete(roomId); // Only unsubscribe if no local clients are in this room const hasLocal = [...wss.clients].some( (c) => c !== ws && c.rooms?.has(roomId) ); if (!hasLocal) sub.unsubscribe(getRoomChannel(roomId)); } sub.on("message", (channel, message) => { const roomId = channel.replace(/^room:/, ""); for (const client of wss.clients) { if (client.rooms?.has(roomId) && client.readyState === WebSocket.OPEN) { client.send(message); } } });
Sticky Sessions (Load Balancer Config)
If you need room state to stay in-memory, force all connections from the same user/room to land on the same server instance.
nginx — ip_hash (stickiness by client IP):
upstream ws_backend {
ip_hash; # hash by client IP — same IP hits the same backend
server 10.0.0.1:4000;
server 10.0.0.2:4000;
}nginx — stickiness by cookie (open source: hash an existing session cookie; sticky cookie is nginx Plus only):
upstream ws_backend {
hash $cookie_sessionid consistent; # route by your app's session cookie
server 10.0.0.1:4000;
server 10.0.0.2:4000;
}
ip_hashbreaks behind a shared corporate NAT or CDN (many users, one IP) — prefer cookie/session hashing there.
AWS ALB — sticky sessions:
Target Group → Attributes → Stickiness → Enable Duration: 86400 (1 day) Cookie name: AWSALB
Sticky sessions are a workaround, not a solution — if a server dies, sticky clients lose state. Combine with Redis for resilience.
Horizontal Scaling Options
| Approach | Complexity | Fault tolerance | Notes |
|---|---|---|---|
| Single server | Low | None | Fine up to ~50 k connections |
| Sticky sessions only | Low | Poor (server death = lost state) | Simple, fragile |
| Redis pub/sub | Medium | Good | Standard choice |
| Redis Streams | Medium | Good + persistence | Replay missed messages |
| NATS / Kafka | High | Excellent | Large-scale, low-latency |
| Socket.IO Adapter | Low (config) | Good | Adapter abstracts the bus |
Redis Streams — Persistent Message Log
import Redis from "ioredis"; const redis = new Redis(process.env.REDIS_URL); // Publish with stream async function publish(streamKey, data) { await redis.xadd(streamKey, "*", "data", JSON.stringify(data)); } // Consumer group — each server reads its own slice async function consume(streamKey, group, consumer) { await redis.xgroup("CREATE", streamKey, group, "$", "MKSTREAM").catch(() => {}); while (true) { const results = await redis.xreadgroup( "GROUP", group, consumer, "COUNT", 100, "BLOCK", 1000, "STREAMS", streamKey, ">" ); if (!results) continue; for (const [, messages] of results) { for (const [id, fields] of messages) { const data = JSON.parse(fields[1]); broadcastLocally(data); await redis.xack(streamKey, group, id); } } } }
Connection Limits
| Environment | Typical limit | Tune via |
|---|---|---|
| Node.js | ~100 k (RAM-bound) | --max-old-space-size |
| Linux file descriptors | 1024 (default) | ulimit -n 65535 |
| nginx worker_connections | 1024 (default) | worker_connections 65535 |
| AWS ALB | No fixed per-target connection quota (LB scales) | Raise idle timeout — default 60 s, max 4000 s |
# Increase OS file descriptor limit (add to /etc/security/limits.conf) * soft nofile 65535 * hard nofile 65535 # Or per-process ulimit -n 65535
Health Check Endpoint
Load balancers need an HTTP health check alongside WebSocket.
app.get("/health", (req, res) => { res.json({ status: "ok", connections: wss.clients.size, uptime: process.uptime(), }); });
Worker Threads / Cluster
import cluster from "cluster"; import os from "os"; if (cluster.isPrimary) { const cpus = os.cpus().length; for (let i = 0; i < cpus; i++) cluster.fork(); cluster.on("exit", (worker) => { console.log(`Worker ${worker.process.pid} died, restarting...`); cluster.fork(); }); } else { // Each worker runs its own WebSocketServer + Redis pub/sub const { createServer } = await import("http"); const { startApp } = await import("./app.js"); startApp(createServer()); }
With cluster, each worker process has its own
wss.clients. Redis pub/sub is required for cross-worker messaging.
Metrics to Monitor
| Metric | What it tells you |
|---|---|
wss.clients.size | Current open connections per instance |
| Redis pub/sub latency | Cross-server message delay |
process.memoryUsage().heapUsed | Memory per connection |
| Node event loop lag | Throughput saturation |
Close code 1006 rate | Abnormal drops (network/proxy issues) |
Gotchas
- Redis
subscribeandpublishmust use separate client instances — a subscribed Redis connection cannot send commands. - Redis pub/sub has no persistence — if a subscriber is down when a message is published, it is lost. Use Redis Streams for durability.
- Each Node.js connection is just an event-loop socket, but it still costs heap (per-socket buffers, your per-client state) plus kernel send/receive buffers — and TLS handshake bursts are CPU-heavy. Profile memory and event-loop lag at your expected concurrency before choosing instance size.
- On AWS, Application Load Balancer (ALB) has a default idle timeout of 60 s — set it to 3600+ and configure heartbeats to prevent silent drops.
- When using cluster,
shutdownsignals must be forwarded to workers; unhandledSIGTERMin a worker will abruptly close all its connections with code1006.