Backend · Distributed Systems
Distributed systems for backend devs
Idempotency, retries, and the eight fallacies that bite you when one service becomes ten.
The first time a backend dev's monolith becomes two services, they relearn computing.
Inside one process: function calls always succeed (modulo bugs). Across the network: every call has three outcomes — success, failure, and "we have no idea." The third is the one that breaks systems.
This is a pragmatic tour of the patterns that take you from "I have two services" to "I have a system that doesn't wake me up at 3 AM."
The eight fallacies, restated#
Peter Deutsch's classic list, with my one-liner annotations from years of getting bitten:
- The network is reliable — packet loss happens; assume retries.
- Latency is zero — N+1 over the network is N+1 seconds, not milliseconds.
- Bandwidth is infinite — JSON over the wire is bigger than you think.
- The network is secure — every internal call needs auth.
- Topology doesn't change — services move, IPs change, DNS lies.
- There is one administrator — config drift is real.
- Transport cost is zero — every hop adds CPU somewhere.
- The network is homogeneous — protocols and clients vary.
These aren't trivia. They are the design space. Every distributed system mistake I've made traces back to assuming one of these is false.
Idempotency is the primitive#
If a request can be retried safely, every other failure mode shrinks. So the first design question for any cross-service operation is: can the receiver process this twice without breaking?
The simplest implementation: an idempotency key — a unique ID supplied by the caller, stored by the receiver, used to dedupe.
def process_payment(idempotency_key, amount, customer_id):
# Atomic insert-if-not-exists
existing = PaymentAttempt.objects.filter(key=idempotency_key).first()
if existing:
return existing.result # Already processed; return cached result
# Reserve the key BEFORE doing the work
PaymentAttempt.objects.create(key=idempotency_key, status='pending')
try:
result = charge_card(customer_id, amount)
PaymentAttempt.objects.filter(key=idempotency_key).update(
status='complete', result=result
)
return result
except CardError as e:
PaymentAttempt.objects.filter(key=idempotency_key).update(
status='failed', error=str(e)
)
raiseThe caller retries with the same key. The first call does the work; subsequent calls return the cached result. No double-charge.
Two implementation traps to avoid:
- Reserve the key before doing the work. If you check-then-reserve, two concurrent retries can race and both do the work.
- Persist the result. Just storing "we saw this key" isn't enough — the retry needs to return the same response, including any IDs the original call generated.
Stripe's API is the gold-standard reference here; their idempotency-key docs are worth reading even if you'll never use Stripe.
Timeouts: the ones you'll regret not setting#
Default timeouts in most HTTP clients are too long. Default retry counts are usually wrong. Defaults have killed more services than bugs in my experience.
import httpx
# Per-request: connect / read / write / pool — all explicit
TIMEOUT = httpx.Timeout(connect=2.0, read=5.0, write=5.0, pool=2.0)
client = httpx.Client(timeout=TIMEOUT)A working set of defaults I use unless I have a reason to deviate:
- Connect timeout: 1–2s. If the SYN-ACK didn't come back fast, the network is sad; fail fast.
- Read timeout: 5–10s for synchronous user-facing calls; 30s+ for background jobs talking to slower services.
- Total per-request budget: < 1s for anything in a request path. Each microservice you add should consume a fraction of the user-visible latency budget.
The discipline: write down your latency budget, then sum the per-call budgets and confirm they fit. If they don't, you can't fix it with more retries.
Retries: do them, but bounded#
Retries are how you recover from transient failures. They're also how you turn a flap into a thundering herd. Three rules:
- Only retry idempotent operations. Otherwise you risk duplicate side-effects.
- Exponential backoff with jitter. Not just exponential — jitter spreads the retries across time so 1000 retrying clients don't all retry at the same instant.
- Cap retries. 3–5 is usually right. Past that, the server is unhealthy and your retries are making it worse.
import random, time
def retry_with_backoff(fn, max_attempts=4, base=0.2, cap=5.0):
for attempt in range(max_attempts):
try:
return fn()
except (TransientError, TimeoutError):
if attempt == max_attempts - 1:
raise
sleep = min(cap, base * (2 ** attempt))
sleep = sleep * (0.5 + random.random()) # full jitter
time.sleep(sleep)For anything serious, use a battle-tested library — tenacity for Python, backoff, or framework-specific clients with retry policies. Don't roll your own past the demonstration above.
Circuit breakers: stop hammering a dying service#
When a downstream service is genuinely down, retrying just delays the inevitable while consuming your worker pool. Circuit breakers detect the failure pattern and stop sending requests for a window.
States:
- Closed (normal): requests flow through; track failure rate
- Open (downstream is sad): fail fast without sending; periodically...
- Half-open: send a single probe request; if it succeeds, go closed
The Hystrix paper from Netflix is dated but still the canonical reference. Modern equivalents: `resilience4j\