Section 5 · 14 Articles

Reliability & Resilience

Engineering patterns that keep systems working correctly under adverse conditions — from transient failures to cascading outages and message delivery edge cases.

Retry Pattern

Reattempting transient failures when the failure is likely temporary and safe to repeat.

Reliability Essential

Exponential Backoff

Progressively spacing retries with jitter to avoid thundering herd under concurrent load.

Reliability Essential

Circuit Breaker

Stop calling a failing dependency temporarily to prevent cascading failures across services.

Reliability Essential

Bulkhead Pattern

Isolate resource pools so one failing workload cannot starve or destroy another.

Reliability

Fallback Pattern

Provide a degraded alternative when the primary path fails — partial service over total outage.

Reliability

Graceful Degradation

Reduce functionality under pressure while preserving core service availability.

Reliability

Load Shedding

Reject or drop low-priority work under overload to preserve core service health.

Reliability Scalability

Dead Letter Queues

A quarantine lane for failed messages that need inspection and later recovery.

Messaging Reliability

Poison Message Handling

Isolating messages that repeatedly fail to prevent blocking queue processing.

Messaging Reliability

Exactly-Once Delivery

Ensuring each message is processed exactly once despite network failures and retries.

Messaging Advanced

At-Least-Once Delivery

Guaranteeing message delivery at the cost of potential duplicates — and how to handle them.

Messaging

Chaos Engineering

Deliberately injecting failures in production to discover and fix reliability weaknesses.

Reliability Advanced

Self-Healing Systems

Designing systems that automatically detect and recover from failures without human intervention.

Reliability Automation

Failure Domains

Containing failures within bounded zones to limit the blast radius of any single incident.

Reliability Architecture