Backend · Databases
Redis: cache, queue, lock, mistake — pick one
Redis is fantastic at four things and dangerous when you mix them. A pragmatic field guide for production use.
Redis is one of those tools that works so well in early prototypes that teams reach for it for everything. Then production happens and they find out that "cache" Redis and "queue" Redis and "primary store" Redis have very different operational shapes.
This article disentangles them.
Pattern 1: Redis as a cache#
The original use case. The mental model: a lossy key-value store with TTL. If Redis goes down, your app slows down but stays correct.
The non-negotiable invariant: never read from Redis for data that doesn't exist somewhere else. Cache misses must hit the source of truth.
def get_user(user_id: int) -> User:
cached = redis.get(f"user:{user_id}")
if cached:
return User.from_json(cached)
user = db.query(User).get(user_id) # source of truth
redis.setex(f"user:{user_id}", 300, user.to_json()) # 5 min TTL
return userThe pitfalls:
- Cache stampede. When the cached entry expires, 1000 concurrent requests all miss simultaneously and hammer the DB. Mitigations: jittered TTL,
SETNXlock so only one request refreshes, or a "stale-while-revalidate" pattern where the expired value is still served while a refresh runs in the background. - Stale reads after writes. Update a user → forget to invalidate → old data lives until TTL expires. Always
redis.delete()(orredis.set()to overwrite) on write paths. - Memory pressure. Redis is in-memory. Set
maxmemoryandmaxmemory-policy allkeys-lru. Without that policy, a runaway application can fill Redis until OOM.
- GET80
- SET80
- LPUSH85
- RPOP85
- SADD80
- ZADD120
Pattern 2: Redis as a job queue#
This is where things get sketchy. Redis-as-queue is fine for fire-and-forget background work where occasional message loss is OK (e.g. "send a welcome email"). It is not OK for "process this order" or "charge this card."
Use Redis Streams if you must, not LPUSH/RPOP lists. Streams give you consumer groups, ack semantics, and a backlog you can replay:
# Producer
redis.xadd("orders", {"order_id": "42", "user_id": "100"})
# Consumer (in a worker)
while True:
msgs = redis.xreadgroup(
groupname="order-workers",
consumername="worker-1",
streams={"orders": ">"},
block=5000,
)
for stream, entries in msgs:
for msg_id, data in entries:
try:
process_order(data)
redis.xack("orders", "order-workers", msg_id)
except Exception:
# Message stays unacked, will be redelivered after timeout
passFor anything where loss matters: use a real queue. SQS, RabbitMQ, Postgres-backed queue (LISTEN/NOTIFY, or SELECT FOR UPDATE SKIP LOCKED). Redis sits in a weird middle where you pay for queue semantics without getting the durability guarantees.
Pattern 3: Redis as a distributed lock#
The famous "Redlock" pattern. Used everywhere; rarely correctly.
Simple version:
import secrets
token = secrets.token_hex(16)
acquired = redis.set("lock:resource:42", token, nx=True, ex=30)
if not acquired:
raise LockBusy()
try:
do_critical_section()
finally:
# Lua script: only delete if value still matches our token
redis.eval(
"if redis.call('get', KEYS[1]) == ARGV[1] then "
" return redis.call('del', KEYS[1]) "
"else return 0 end",
1, "lock:resource:42", token,
)The pitfalls:
- Clock skew between Redis and clients. Your lock's TTL is enforced by Redis's clock; your "did the work finish in time" check is your client's clock. They drift.
- Network partition + failover. If the master holding your lock fails over to a replica that hadn't received the write, two clients can hold "the same" lock.
- Long pauses. GC pauses, container suspensions, debugger breakpoints — all can extend "I think I have the lock" past the TTL while another client takes it.
For real correctness you need either a single-node Postgres advisory lock, or a proper consensus system (etcd, Consul). Redlock is "fine for cache invalidation, not fine for money."
Pattern 4: Redis as a primary store#
Don't.
Caveat: there is a class of workloads where Redis is the right primary store — extremely fast leaderboards, real-time presence (who's online), session storage where you've separately built the durability story. Discord, Twitter's spike caches, etc. These teams have invested heavily in operations.
For most of us: Redis is not a database. It does not have transactions you can trust the way Postgres has. Its replication is asynchronous and has a non-trivial failure mode. It silently truncates writes when memory is tight (with an LRU policy).
If you find yourself building "primary store" use cases on Redis, take a hard look at whether you actually need <1ms reads or you've just memorized that Redis is fast.
Operational checklist#
Whatever your use case:
maxmemoryset explicitly. Pick a number; don't let Redis grow until OOM.- Eviction policy chosen on purpose.
allkeys-lrufor cache,noevictionfor queue (so a full Redis fails the producer rather than dropping work). - Persistence configured. RDB snapshots are fine for cache. AOF (append-only file) for anything you'd be unhappy to lose. Both for "I want Redis as a backing store but I expect to use cache patterns."
- Replicas with failover. Redis Sentinel is the minimum. Redis Cluster if you need horizontal scale. Don't run a single Redis instance in production for anything that matters.
- Monitoring.
INFO, slow log, key-eviction rate, replication lag. Track them like you track DB metrics.
A concrete decision tree#
When someone asks "should I use Redis for this?":
- Is the data already in another store, and Redis is a speed-up cache? → Yes, with TTL. Test the cold-cache path.
- Is the data ephemeral by design (session, rate-limit counters, presence)? → Yes, fine.
- Is it a queue for non-critical work? → OK, but use Streams, with consumer groups, and accept ~99.9% delivery.
- Is it a lock for cross-process coordination? → Use Redis only if eventual incorrectness is recoverable. Otherwise, Postgres advisory locks.
- Is it your only copy of important data? → No. Use a real database.
Stick to that and Redis is one of the best tools in the kit. Drift from it and Redis is one of the easiest tools to be wrong about.
Further reading#
- Redis docs on persistence — required reading before any "Redis as primary" decision.
- Martin Kleppmann, "How to do distributed locking" — the definitive critique of naive Redlock.
- Redis Streams docs — the right primitive for queue-shaped problems.
- "The little Redis book" — small, free, surprisingly current.