Home›Telecom›Non-Functional Architecture›Resilience & High Availability← NFR overview
Non-functional architecture
Reference Architecture

Resilience & High Availability

Resilience, High Availability & Idempotency: Circuit Breaker · Bulkhead · Retry · Idempotency · HA · Active-Active · Redundancy.

Resilience, High Availability & Idempotency
Circuit Breaker · Bulkhead · Retry · Idempotency · HA · Active-Active · Redundancy
99.99%
Target Availability
Minutes downtime per year
3
Min Availability Zones
Active-active across AZs
2+
Replicas per Service
Minimum in production
<30s
Auto-Recovery Time
Pod restart + health check
Circuit Breaker & Bulkhead
Fail-fast · Timeout · Fallback · Isolation
Circuit BreakerHalf-open → Open → Closed state machine. Opens when error rate >50% in a configurable measurement window (min 20 requests). Auto-recovers after a configurable probe interval. Configured via service mesh policy — no code change required.
TimeoutAll inter-service calls have explicit timeouts: P99 latency × 3. No unbounded waits. Timeout is non-negotiable — never rely on upstream timeout alone.
Retry PolicyMax 3 retries with exponential backoff (base-delay → 2× → 4× (with jitter)) + jitter (configurable jitter). Retry only on idempotent operations and 5xx / connection errors. Never retry on 4xx or mutation endpoints without idempotency key.
BulkheadSeparate thread/connection pool per downstream dependency. Billing calls to payment cannot exhaust threads used by billing calls to party management. Pool sizes tuned per SLA.
FallbackGraceful degradation: if notification service is unavailable, order proceeds and notification is queued for retry. Service defines explicit fallback behaviour — not just error.
Circuit Breaker Bulkhead Exponential Backoff Graceful Degradation
Idempotency Design
Idempotency Keys · Exactly-Once Semantics · Outbox Pattern
Every state-changing API operation must be idempotent — executing the same request multiple times produces the same result with no duplicate side effects. Mandatory for all payment, order and provisioning endpoints.
Idempotency KeyClient generates a Unique Idempotency Key idempotency key per request. Server stores the response keyed by this ID for 24 h. Duplicate requests within TTL receive the cached response — no re-execution.
ScopeMandatory on: POST /orders, POST /payments, POST /activations, POST /subscriptions. All provisioning adapter calls. All financial mutations.
Event DeduplicationEvent consumers maintain a deduplication store (event ID → processed timestamp). Duplicate events within a 1 h window are detected and silently discarded.
Outbox PatternDatabase writes and event publications are atomic via the transactional outbox pattern. State change is written to DB and outbox table in one transaction. Outbox poller publishes to event bus. Eliminates dual-write inconsistency.
At-Least-OnceEvent streaming delivers at-least-once. Idempotent consumers handle duplicates. Exactly-once semantics achieved via dedup + idempotent processing — not by trusting the broker.
Idempotency Key Outbox Pattern Event Deduplication Exactly-Once Logic
High Availability & Redundancy
Active-Active · Multi-AZ · Auto-Failover · PDB
TopologyActive-active across 3 availability zones within a region. No active-passive for tier-1 BSS services — single standby introduces failover latency. All zones serve live traffic.
Database HAPrimary + 2 synchronous replicas across AZs. Automatic leader election (Database HA Manager/Database HA Manager). Read replicas for analytics and reporting workloads. No single point of failure in data layer.
Pod Disruption BudgetPDB enforced: minimum 2 replicas always running during voluntary disruption (node drain, rolling upgrade). Tier-1 services: min 60% replicas available during upgrades.
Queue RedundancyEvent streaming: 3 broker replicas, replication factor 3, min-insync-replicas 2. A single broker failure does not cause message loss or producer/consumer unavailability.
Cache RedundancyCache Cluster (Sentinel mode) (3 nodes) or Cache Cluster (Sharded mode) (6 nodes) for HA cache. Cache miss falls back to database — no thundering herd protection via cache stampede mitigation (probabilistic TTL).
Active-Active Multi-AZ PDB Leader Election
Non-Functional Requirements — Summary Reference
NFR CategoryRequirementTarget / StandardStatus
AvailabilityPlatform availability SLO99.99% (≈ minutes per year)Mandatory
LatencyAPI P99 latency (tier-1 services)< Sub-second P99Mandatory
ThroughputCDR processing rateVery High — stream processing sustainedMandatory
RTORevenue-critical service recovery<Minutes (Tier 0 — near-zero)Mandatory
RPOMaximum data loss windowNear-zero (Tier 0 — synchronous replication)Mandatory
SecurityEncryption in transitTLS 1.3+ on external traffic, mTLS on all inter-service communicationMandatory
SecurityEncryption at restAES-256-GCM all data storesMandatory
IdentityAuthentication protocolOAuth 2.0 / OIDC + Workload Identity Standard workload identityMandatory
PII ComplianceGDPR — Right to Erasure SLAConfirmed deletion within the regulatory SLA windowMandatory
ResilienceCircuit breaker on all external callsError rate >50% in 10 s window triggers openMandatory
IdempotencyAll financial mutation endpointsIdempotency key + 24 h dedup storeMandatory
ScalabilityHPA on all stateless servicesCPU target 60%, scale within a short cool-down windowMandatory
Anti-AffinityPod distributionRequired: no two replicas on same node; Preferred: spread across AZsMandatory
Technical DebtSprint debt remediation allocationMinimum 20% per sprintRecommended
ObservabilityDistributed tracing coverage100% of requests carry traceId; 10% sampledMandatory
CI/CD SecuritySAST + SCA in every pipelineCritical/High CVEs block mergeMandatory
Non-functional architecturePrevious: Network & Infrastructure→Non-functional architectureNext: Disaster Recovery & Continuity→