🚀 The Simple Version (ELI5)
Imagine ordering pizza from a kitchen that sometimes freezes. If you keep calling the same kitchen over and over while it’s frozen, you’ll just waste time and eventually give up. A circuit breaker is like a smart waiter who says, “I’m not calling this kitchen right now. Try again later.” Retries are like politely asking the waiter again after a short break. Together they keep your pizza order service working smoothly, even if the kitchen is down.
Why Fault Tolerance Matters
Modern APIs rarely run in isolation. They depend on other services, databases, or external providers that can fail, slow, or become unreachable. Without protection, a single failure can cascade, causing slow responses, timeouts, or crashes for end users. Fault‑tolerant design mitigates these risks by:
- Preventing repeated failures from overwhelming a service (circuit breaker).
- Automatically retrying transient errors (retry logic).
- Providing graceful degradation or fallback responses.
- Reducing operational noise and alert fatigue.
Circuit Breaker Pattern
States & Transitions
- CLOSED: Normal operation. Calls pass through.
- OPEN: Too many failures. Calls are short‑circuited and fail fast.
- HALF‑OPEN: After a timeout, a few test calls are allowed to see if the dependency has recovered.
Key Parameters
- Failure Threshold – number of consecutive failures before opening.
- Success Threshold – number of consecutive successes before closing.
- Timeout – how long to stay OPEN before trying HALF‑OPEN.
- Reset Timeout – optional back‑off when re‑opening.
Implementation Example (Node.js)
class CircuitBreaker {
constructor({ failureThreshold = 5, successThreshold = 2, timeout = 10000 }) {
this.state = 'CLOSED';
this.failureCount = 0;
this.successCount = 0;
this.failureThreshold = failureThreshold;
this.successThreshold = successThreshold;
this.timeout = timeout;
this.nextAttempt = Date.now();
}
async call(fn) {
if (this.state === 'OPEN' && Date.now() < this.nextAttempt) {
throw new Error('Circuit is open');
}
try {
const result = await fn();
this.recordSuccess();
return result;
} catch (err) {
this.recordFailure();
throw err;
}
}
recordSuccess() {
if (this.state === 'HALF-OPEN') {
this.successCount += 1;
if (this.successCount >= this.successThreshold) {
this.state = 'CLOSED';
this.failureCount = 0;
this.successCount = 0;
}
}
}
recordFailure() {
this.failureCount += 1;
if (this.failureCount >= this.failureThreshold) {
this.state = 'OPEN';
this.nextAttempt = Date.now() + this.timeout;
}
}
}
Retry Strategy
When to Retry
- Transient network errors (ECONNRESET, ETIMEDOUT).
- HTTP 5xx responses from downstream services.
- Temporary rate‑limit or quota hits (HTTP 429).
Back‑off & Jitter
Use exponential back‑off with random jitter to avoid thundering herd problems. Example algorithm:
async function retry(fn, attempts = 3, baseDelay = 200) {
for (let i = 0; i < attempts; i++) {
try {
return await fn();
} catch (err) {
if (i === attempts - 1) throw err;
const jitter = Math.random() * baseDelay;
const delay = Math.pow(2, i) * baseDelay + jitter;
await new Promise(r => setTimeout(r, delay));
}
}
}
Idempotency Matters
Retries should be safe. For non‑idempotent operations (e.g., POST that creates a record), include an Idempotency-Key header or design the backend to detect duplicates.
Best Practices
- Separate concerns: keep circuit breaker logic in a wrapper, not in business code.
- Configure thresholds per service based on SLA and traffic patterns.
- Use a library or framework support (e.g., Polly for .NET, Resilience4j for Java, opossum for Node).
- Instrument state changes (OPEN, HALF‑OPEN, CLOSED) and metrics.
- Set a reasonable retry limit; unlimited retries can back‑fire.
- Gracefully degrade: return cached data or a friendly error message when a circuit is open.
Common Pitfalls
- Too many retries causing request amplification.
- Not accounting for side effects on non‑idempotent calls.
- Using a single timeout for all calls – leads to race conditions.
- Ignoring the latency cost of opening/closing a circuit.
- Failing to reset the circuit after a successful call.
Monitoring & Observability
Track:
- Failure rate per service.
- Circuit state transitions.
- Retry counts and back‑off times.
- Latency distribution of wrapped calls.
- Alert thresholds for high failure rates.
Integrate with Prometheus, Grafana, or vendor dashboards. Use distributed tracing (Jaeger, Zipkin) to see where calls are failing.
Case Study: E‑Commerce Checkout API
During a flash sale, the payment gateway experienced intermittent outages. By wrapping the payment call in a circuit breaker with a 3‑failure threshold and a 30‑second timeout, the checkout API stayed responsive. Retries with exponential back‑off and jitter handled transient network hiccups. The result: 99.8% uptime during peak traffic and a 15% reduction in error logs compared to the previous monolithic approach.
Conclusion
Fault‑tolerant APIs are built on two pillars: circuit breakers to stop cascading failures and retries to handle transient glitches. When combined with proper configuration, observability, and graceful degradation, they turn brittle services into resilient ones that deliver a better user experience even in the face of chaos.