myibrahim.cloud

Backend · Operations

Caching: where to put it, how it bites you back

Caching is the easiest way to make a slow system fast and the easiest way to make a correct system wrong. A field guide to the layers.

A senior engineer told me once: "caching is just performance with a really good incident story." He was right. Every cache layer is a perf win in the steady state and a fascinating bug under specific conditions. Here's the map of those layers and the bugs.

The cache layers, from outside in#

Browser
  └─ CDN (CloudFront, Fastly, Cloudflare)
       └─ Reverse proxy (Nginx, Caddy)
            └─ App-level cache (Redis, in-process)
                 └─ Query cache (Postgres `shared_buffers`)
                      └─ Disk

Each layer has different rules. Each can save you from the layer below. Each has its own ways to lie.

Layer 1: CDN#

The fastest cache: edge servers near the user. A CDN-cached static asset (CSS, JS, images, public API responses) returns in <50ms globally.

Configuration is a Cache-Control header. Get it right or get burned:

# Static asset, immutable (renamed on every build)
Cache-Control: public, max-age=31536000, immutable

# HTML page, mostly static but might change
Cache-Control: public, max-age=300, s-maxage=3600, stale-while-revalidate=86400

# API response, personalized
Cache-Control: private, no-cache

# Don't cache (logout, dynamic dashboard)
Cache-Control: no-store

s-maxage = how long shared caches (CDN) can hold it. max-age = how long the browser can. They diverge intentionally; CDN can hold a "stale" response while browsers refresh more often.

stale-while-revalidate is the underrated directive: the cache can serve a slightly stale response while fetching a fresh one in the background. User sees instant load; cache stays current. Use it.

The bug to fear: cache-poisoning. A request with a header your origin echoes back becomes a cached response served to other users. Set Vary correctly (Vary: Authorization), or use a CDN that auto-handles it.

Layer 2: Reverse proxy#

Less common as a cache layer in 2026 — most teams let the CDN handle it. But Nginx/Caddy/Varnish can cache at the edge of your origin, which helps when:

  • You don't have a CDN (yet).
  • Specific endpoints need cache rules different from the CDN.
  • You want to cache responses from upstream services in your own data center.

Caddy makes this easy:

example.com {
    cache @api {
        path /api/products/*
        max_age 5m
    }
    reverse_proxy /api/* upstream:8080
}

The bugs: stale responses after a deploy if cache TTLs are long; routing failures cached as 5xx (avoid this — never cache errors).

Layer 3: App-level cache (Redis or in-process)#

The most actively-managed layer. The two flavors:

Redis (or Memcached). Shared across instances. Survives restarts (with persistence). The default for "we need a cache."

In-process. A dict in your application. Per-instance. Lost on restart. Useful for things that change rarely and are small (config, feature flags, small reference data).

Both face the same correctness questions:

  • What's the TTL? Too short = wasted hits. Too long = stale data.
  • Who invalidates? Time-based (TTL only) or event-based (write-path deletes the key).
  • How do you handle stampedes? When 1000 requests hit a missing key simultaneously.

Stampede mitigations:

# Single-flight pattern: only one request refreshes; others wait.
def get_product(id):
    cached = redis.get(f"product:{id}")
    if cached:
        return Product.from_json(cached)

    # Try to acquire a refresh lock
    lock_key = f"lock:product:{id}"
    have_lock = redis.set(lock_key, "1", nx=True, ex=10)
    if have_lock:
        try:
            product = db.query(Product).get(id)
            redis.setex(f"product:{id}", 300, product.to_json())
            return product
        finally:
            redis.delete(lock_key)
    else:
        # Another worker is refreshing. Wait briefly, retry.
        time.sleep(0.05)
        return get_product(id)

Layer 4: Query cache#

Postgres's shared_buffers is itself a cache. So is the OS page cache. Tuning these is a separate article (linked at the end), but the principle: enough RAM that your working set fits.

If your working set fits in shared_buffers: most queries serve from RAM regardless of any app-level caching. Adding Redis on top is then mostly redundant for steady-state perf — though still useful for limiting connections and reducing CPU.

Cache-aside vs. write-through vs. write-behind#

Three patterns, picking right is most of the discipline:

Cache-aside (lazy loading). App reads cache, falls through to DB on miss. Most common. Simple to reason about. Bug: cache and DB can diverge after a write that didn't invalidate.

def write(item):
    db.save(item)
    redis.delete(item_key(item.id))  # or set with new value

Write-through. Every write goes through the cache, which writes to the DB synchronously. Cache is always fresh. Slow writes.

Write-behind. Writes hit cache only; cache asynchronously persists to DB. Fast writes. Lossy on cache failure.

For 95% of products: cache-aside. Don't over-engineer.

The bugs that actually happen#

Stale reads after writes. Customer updates email; next read returns the old one. Fix: invalidate on write. Always.

Negative cache. Caching "this user does not exist" is good for unfound IDs. But if a real user is created later, the cache says no. Fix: short TTL on negative cache (60s typical).

Expired-stampede. TTL hits zero across many keys at once (because they were warmed up at the same time). Fix: jitter the TTL (e.g. 300–360s).

Inconsistent invalidation across instances. Instance A invalidates Redis. Instance B's in-process cache still has the old value. Fix: prefer Redis over in-process for anything that mutates.

Cache poisoning via input. A malicious header gets cached as part of the response key. Fix: explicit Vary headers; never include user-controlled values in cache keys.

Cache as a database. "It's been working in Redis for two years" — until Redis's persistence config silently failed and you lost everything. Fix: caches are never the source of truth.

What to cache (and what not to)#

Cache:

  • Read-heavy data with low write rate (product catalog, user profiles, config)
  • Computed results (aggregations, derived stats, rendered HTML)
  • External API responses (rate-limited, slow)
  • Session data (HttpOnly cookies + Redis is the standard pattern)

Don't cache:

  • Real-time stock prices, payment statuses, anything where staleness causes a customer-visible bug
  • Per-request, per-user data that won't be reused (cache hit rate near zero)
  • Tiny computations (cache lookup costs more than re-computing)

Observability for caches#

Every cache layer should expose:

  • Hit rate. Below 80% means it's barely earning its keep.
  • Eviction rate. High eviction = under-provisioned cache.
  • Latency. The cache itself should be <1ms; if not, it's a worse cache than no cache.
  • Memory utilization. Watch the trajectory.

A dashboard showing these per-layer makes the difference between "we think caching is helping" and "caching saved us 70% of DB load and we have the chart."

Further reading#

  • caching
  • performance
  • redis
  • cdn
  • production
  • backend
Need this built? I build full-stack web app or saas projects for clients worldwide. Tell me about yours.