Backend · Operations
Caching: where to put it, how it bites you back
Caching is the easiest way to make a slow system fast and the easiest way to make a correct system wrong. A field guide to the layers.
A senior engineer told me once: "caching is just performance with a really good incident story." He was right. Every cache layer is a perf win in the steady state and a fascinating bug under specific conditions. Here's the map of those layers and the bugs.
The cache layers, from outside in#
Browser
└─ CDN (CloudFront, Fastly, Cloudflare)
└─ Reverse proxy (Nginx, Caddy)
└─ App-level cache (Redis, in-process)
└─ Query cache (Postgres `shared_buffers`)
└─ DiskEach layer has different rules. Each can save you from the layer below. Each has its own ways to lie.
Layer 1: CDN#
The fastest cache: edge servers near the user. A CDN-cached static asset (CSS, JS, images, public API responses) returns in <50ms globally.
Configuration is a Cache-Control header. Get it right or get burned:
# Static asset, immutable (renamed on every build)
Cache-Control: public, max-age=31536000, immutable
# HTML page, mostly static but might change
Cache-Control: public, max-age=300, s-maxage=3600, stale-while-revalidate=86400
# API response, personalized
Cache-Control: private, no-cache
# Don't cache (logout, dynamic dashboard)
Cache-Control: no-stores-maxage = how long shared caches (CDN) can hold it. max-age = how long the browser can. They diverge intentionally; CDN can hold a "stale" response while browsers refresh more often.
stale-while-revalidate is the underrated directive: the cache can serve a slightly stale response while fetching a fresh one in the background. User sees instant load; cache stays current. Use it.
The bug to fear: cache-poisoning. A request with a header your origin echoes back becomes a cached response served to other users. Set Vary correctly (Vary: Authorization), or use a CDN that auto-handles it.
Layer 2: Reverse proxy#
Less common as a cache layer in 2026 — most teams let the CDN handle it. But Nginx/Caddy/Varnish can cache at the edge of your origin, which helps when:
- You don't have a CDN (yet).
- Specific endpoints need cache rules different from the CDN.
- You want to cache responses from upstream services in your own data center.
Caddy makes this easy:
example.com {
cache @api {
path /api/products/*
max_age 5m
}
reverse_proxy /api/* upstream:8080
}The bugs: stale responses after a deploy if cache TTLs are long; routing failures cached as 5xx (avoid this — never cache errors).
Layer 3: App-level cache (Redis or in-process)#
The most actively-managed layer. The two flavors:
Redis (or Memcached). Shared across instances. Survives restarts (with persistence). The default for "we need a cache."
In-process. A dict in your application. Per-instance. Lost on restart. Useful for things that change rarely and are small (config, feature flags, small reference data).
Both face the same correctness questions:
- What's the TTL? Too short = wasted hits. Too long = stale data.
- Who invalidates? Time-based (TTL only) or event-based (write-path deletes the key).
- How do you handle stampedes? When 1000 requests hit a missing key simultaneously.
Stampede mitigations:
# Single-flight pattern: only one request refreshes; others wait.
def get_product(id):
cached = redis.get(f"product:{id}")
if cached:
return Product.from_json(cached)
# Try to acquire a refresh lock
lock_key = f"lock:product:{id}"
have_lock = redis.set(lock_key, "1", nx=True, ex=10)
if have_lock:
try:
product = db.query(Product).get(id)
redis.setex(f"product:{id}", 300, product.to_json())
return product
finally:
redis.delete(lock_key)
else:
# Another worker is refreshing. Wait briefly, retry.
time.sleep(0.05)
return get_product(id)Layer 4: Query cache#
Postgres's shared_buffers is itself a cache. So is the OS page cache. Tuning these is a separate article (linked at the end), but the principle: enough RAM that your working set fits.
If your working set fits in shared_buffers: most queries serve from RAM regardless of any app-level caching. Adding Redis on top is then mostly redundant for steady-state perf — though still useful for limiting connections and reducing CPU.
Cache-aside vs. write-through vs. write-behind#
Three patterns, picking right is most of the discipline:
Cache-aside (lazy loading). App reads cache, falls through to DB on miss. Most common. Simple to reason about. Bug: cache and DB can diverge after a write that didn't invalidate.
def write(item):
db.save(item)
redis.delete(item_key(item.id)) # or set with new valueWrite-through. Every write goes through the cache, which writes to the DB synchronously. Cache is always fresh. Slow writes.
Write-behind. Writes hit cache only; cache asynchronously persists to DB. Fast writes. Lossy on cache failure.
For 95% of products: cache-aside. Don't over-engineer.
The bugs that actually happen#
Stale reads after writes. Customer updates email; next read returns the old one. Fix: invalidate on write. Always.
Negative cache. Caching "this user does not exist" is good for unfound IDs. But if a real user is created later, the cache says no. Fix: short TTL on negative cache (60s typical).
Expired-stampede. TTL hits zero across many keys at once (because they were warmed up at the same time). Fix: jitter the TTL (e.g. 300–360s).
Inconsistent invalidation across instances. Instance A invalidates Redis. Instance B's in-process cache still has the old value. Fix: prefer Redis over in-process for anything that mutates.
Cache poisoning via input. A malicious header gets cached as part of the response key. Fix: explicit Vary headers; never include user-controlled values in cache keys.
Cache as a database. "It's been working in Redis for two years" — until Redis's persistence config silently failed and you lost everything. Fix: caches are never the source of truth.
What to cache (and what not to)#
Cache:
- Read-heavy data with low write rate (product catalog, user profiles, config)
- Computed results (aggregations, derived stats, rendered HTML)
- External API responses (rate-limited, slow)
- Session data (HttpOnly cookies + Redis is the standard pattern)
Don't cache:
- Real-time stock prices, payment statuses, anything where staleness causes a customer-visible bug
- Per-request, per-user data that won't be reused (cache hit rate near zero)
- Tiny computations (cache lookup costs more than re-computing)
Observability for caches#
Every cache layer should expose:
- Hit rate. Below 80% means it's barely earning its keep.
- Eviction rate. High eviction = under-provisioned cache.
- Latency. The cache itself should be <1ms; if not, it's a worse cache than no cache.
- Memory utilization. Watch the trajectory.
A dashboard showing these per-layer makes the difference between "we think caching is helping" and "caching saved us 70% of DB load and we have the chart."
Further reading#
- "Caching is hard" — Shopify's writeup on cache invalidation strategies.
- Cloudflare's caching docs — best CDN-level reference.
- "Distributed caching at Stack Overflow" — Nick Craver's deep dive. Read this once per career.
- See also: my Postgres tuning article for the layer-4 settings.