Reliability & Resilience
Engineering patterns that keep systems working correctly under adverse conditions — from transient failures to cascading outages and message delivery edge cases.
Retry Pattern
Reattempting transient failures when the failure is likely temporary and safe to repeat.
Exponential Backoff
Progressively spacing retries with jitter to avoid thundering herd under concurrent load.
Circuit Breaker
Stop calling a failing dependency temporarily to prevent cascading failures across services.
Bulkhead Pattern
Isolate resource pools so one failing workload cannot starve or destroy another.
Fallback Pattern
Provide a degraded alternative when the primary path fails — partial service over total outage.
Graceful Degradation
Reduce functionality under pressure while preserving core service availability.
Load Shedding
Reject or drop low-priority work under overload to preserve core service health.
Dead Letter Queues
A quarantine lane for failed messages that need inspection and later recovery.
Poison Message Handling
Isolating messages that repeatedly fail to prevent blocking queue processing.
Exactly-Once Delivery
Ensuring each message is processed exactly once despite network failures and retries.
At-Least-Once Delivery
Guaranteeing message delivery at the cost of potential duplicates — and how to handle them.
Chaos Engineering
Deliberately injecting failures in production to discover and fix reliability weaknesses.
Self-Healing Systems
Designing systems that automatically detect and recover from failures without human intervention.
Failure Domains
Containing failures within bounded zones to limit the blast radius of any single incident.