Backend · Distributed Systems
SQS, RabbitMQ, Kafka — when to pick which
Three different jobs, three different tools. The decision tree, the pitfalls, and what to do when 'just use Postgres' is the right answer.
The first message queue choice every team makes is wrong. Not because the team picked badly, but because the team usually doesn't know yet what shape their problem is. Six months in they realize "oh, I needed a stream not a queue" or "actually, I needed Postgres, not Kafka."
Here's how to skip that round.
What "queue" actually means in 2026#
Three categories, often confused:
Job queues: producer enqueues a task, exactly one worker consumes and processes it. Failed jobs retry. Usage: "send this email," "generate this thumbnail," "process this payment." → SQS, RabbitMQ, Postgres-backed (SELECT FOR UPDATE SKIP LOCKED), Sidekiq.
Pub/sub: producer publishes an event, multiple consumers each get a copy. Usage: "user signed up — send to email service AND analytics service AND CRM." → RabbitMQ topics, Redis pub/sub, NATS, Kafka.
Event streams: append-only log; consumers read at their own pace, can rewind. Usage: "every event in the system" — auditing, replay, building event-sourced systems. → Kafka, Redpanda, Pulsar.
Different job → different tool. The mistake is treating them all as "a queue."
The decision tree#
A simplified version of how I make the call:
Is the volume < 100 msg/sec AND already running Postgres?
→ Postgres-backed queue. Stop reading.
Need fan-out (multiple consumers per message)?
→ Pub/sub. Probably RabbitMQ unless you have streaming requirements.
Need to replay historical events / event sourcing / log all events?
→ Kafka (or Redpanda).
On AWS, simple job queue, fire-and-forget?
→ SQS.
Self-hosted, complex routing (topics, headers, TTLs)?
→ RabbitMQ.
High volume (>100K msg/sec), many consumers, durable history?
→ Kafka.SQS: the boring choice#
SQS is bad at almost everything except the one thing it's good at: being completely forgettable.
Pros:
- Zero ops. AWS runs it. You won't think about scaling, patching, or failover.
- Long visibility timeout lets workers take their time. Good for slow jobs.
- DLQ built in. Failed messages move to a dead-letter queue automatically.
- Cheap-ish. $0.40 per million messages. For most workloads, lost-in-the-noise.
Cons:
- No fan-out. One queue → one consumer group. Use SNS+SQS for fan-out.
- No ordering in standard queues. FIFO queues are slow.
- Polling overhead. Long-poll is fine but not free.
- At-least-once. Plan for dedup.
Use it for: AWS-hosted services, simple async jobs, anywhere "just put it in a queue and forget" is the goal.
RabbitMQ: the swiss army knife#
RabbitMQ is the most flexible of the three. It does pub/sub, work queues, RPC, headers/topic-based routing, priority queues, deferred messages.
Pros:
- Routing is rich. Direct, topic, fanout, headers exchanges.
- Mature ops story. Years of production lore.
- Plugins for everything: HTTP API, MQTT, Federation, Shovel.
Cons:
- Throughput per node tops out around 50K msg/sec. Scale-out is via clustering or sharding, both with operational tax.
- Message size limit (default 128MB but practically <1MB for performance).
- No replay. Once consumed and acked, gone.
Use it for: complex routing, on-prem deployments, mid-volume work that needs more than SQS's bare-bones model.
Kafka: the long-haul truck#
Kafka is a durable log, not a queue. The mental model: every message goes to a topic; topics are partitioned; consumers read at their own pace and remember their offset.
Pros:
- Stupendous throughput. Single broker handles 100K+ msg/sec. Horizontal scale to millions/sec.
- Replay. Consumers rewind to any offset. Build new analytics by reprocessing history.
- Durability. Replicated across brokers; tunable consistency.
- Strong ordering within partitions.
Cons:
- Operational complexity. Brokers, ZooKeeper (or KRaft), schemas, partitions, consumer groups. Not casual.
- Latency. End-to-end >50ms typical. Not for microsecond-sensitive work.
- Storage cost. You're keeping all messages for retention period.
Use it for: high-volume event streams, audit logs, any system where replay is a requirement, change-data-capture, real-time analytics.
- Postgres queue5
- RabbitMQ50
- SQS (FIFO)3
- SQS (standard)300
- Kafka500
- Pulsar400
"Just use Postgres" — when it's right#
For up to ~1000 jobs/sec, Postgres beats everything in operational simplicity. The pattern:
CREATE TABLE jobs (
id BIGSERIAL PRIMARY KEY,
payload JSONB NOT NULL,
status TEXT NOT NULL DEFAULT 'pending', -- pending, running, done, failed
attempts INT NOT NULL DEFAULT 0,
available_at TIMESTAMP NOT NULL DEFAULT NOW(),
locked_at TIMESTAMP,
locked_by TEXT
);
CREATE INDEX ON jobs (status, available_at) WHERE status = 'pending';Worker:
WITH next_job AS (
SELECT id FROM jobs
WHERE status = 'pending' AND available_at <= NOW()
ORDER BY available_at
FOR UPDATE SKIP LOCKED
LIMIT 1
)
UPDATE jobs SET status = 'running', locked_at = NOW(), locked_by = $1
WHERE id IN (SELECT id FROM next_job)
RETURNING *;SKIP LOCKED is the magic. Workers don't block each other; whichever gets the row wins.
You get:
- Transactions! Enqueue + business write in the same TXN. Impossible with external queues.
- Easy introspection (just SQL).
- One less system to operate.
- Free retries (UPDATE the row, increment attempts).
Caveats:
- VACUUM tuning matters (high-churn tables).
- Throughput tops at ~5K jobs/sec on a beefy box.
- Polling has overhead; LISTEN/NOTIFY can replace it.
For most teams' background work: Postgres queue. For the 5% that genuinely outgrow it: SQS or Kafka.
Operational pitfalls (in any queue)#
Idempotent consumers. Always. At-least-once delivery means duplicates are routine. Every consumer must handle "process this message twice without bad effects."
Dead-letter queues. Always. Failed messages go to DLQ; alert on DLQ depth.
Backpressure. If consumers can't keep up, the queue grows unbounded. Either size for peak, autoscale workers, or shed load (drop low-priority messages).
Schema evolution. Producers and consumers deploy at different times. Old messages might be in the queue when new code rolls out. Plan for backward-compatible payload changes (versioned event schemas, additive fields only).
Poison messages. A message that crashes every consumer will retry forever. Set max retries; route to DLQ; alert.
A concrete recommendation#
If you're starting from scratch and don't know yet:
- Start with Postgres queue for jobs. (
pg-boss, Sidekiq with PG, your own.) - When you outgrow it (rare): SQS if AWS, RabbitMQ if on-prem.
- When you need replay or event sourcing: Kafka, but only with operational headcount to support it.
- Don't deploy Kafka because it's "the modern choice." Kafka punishes teams that aren't ready.
Most teams discover Kafka was the wrong call about 18 months in. Most teams that started with Postgres queue are still happy three years in.
Further reading#
- SQS docs — short, accurate.
- RabbitMQ tutorials — a useful pattern primer.
- "Designing Data-Intensive Applications" by Kleppmann — chapter 11 on streams is the classic.
- "Why Postgres for queues" — the canonical defense.