Home›Telecom›Non-Functional Architecture›Disaster Recovery & Continuity← NFR overview
Non-functional architecture
Reference Architecture

Disaster Recovery & Continuity

Disaster Recovery & Business Continuity: RTO · RPO · Multi-Region · Backup Strategy · DR Runbooks.

Disaster Recovery & Business Continuity
RTO · RPO · Multi-Region · Backup Strategy · DR Runbooks
4hr
RTO Target
Revenue-critical services
15min
RPO Target
Maximum data loss window
2
Geographic Regions
Primary + DR site
24hr
DR Test Frequency
Chaos engineering cadence
Service TierExamplesRTORPODR StrategyBackup Freq
Tier 0 — Critical RevenueBilling, Rating, Payment15 min0 min (sync)Active-Active multi-region; real-time sync replicationContinuous (WAL shipping)
Tier 1 — Customer-FacingParty Mgmt, CRM, Orders1 hr5 minActive-Passive warm standby; async replication <5 min lagEvery 5 min (incremental)
Tier 2 — ImportantCatalog, Notification, IAM4 hr15 minActive-Passive cold standby; restore from snapshotHourly snapshots
Tier 3 — Non-CriticalReporting, DMS, Monitoring24 hr1 hrBackup and restore; rebuild from event replay if neededDaily backup
Backup Strategy
3-2-1 Rule · Immutable Backups · Tested Restores
3-2-1 Rule3 copies of data. 2 different storage media. 1 offsite (different region). Applied to all Tier 0 and Tier 1 databases.
Immutable BackupsBackups written to Immutable (Write-Once Read-Many) storage. Cannot be deleted or modified for the retention period. Ransomware protection. Retention: 90 days operational, 7 years financial.
Restore TestingAutomated restore test on a daily cadence for Tier 0. Every 72 h for Tier 1. A backup that has not been successfully restored is not a backup — untested restores are not accepted.
Event Replay DREvent streaming acts as a secondary DR mechanism. Any service can be rebuilt by replaying its domain events from the topic. Event retention: configurable hot and cold retention periods.
Chaos Engineering & DR Testing
Failure Injection · GameDay · DR Runbooks
Chaos EngineeringScheduled failure injection in pre-production: pod kill, network partition, disk full, CPU spike, dependency latency injection. Tools: Chaos Engineering Framework / Chaos Engineering Framework.
GameDay ExercisesQuarterly full-team DR simulation: trigger a region-level failure scenario in pre-production. Measure actual RTO vs target. Document gaps and remediation actions.
DR RunbooksAutomated runbooks for every DR scenario. Runbooks version-controlled in Git. Executed in minimal steps to minimise operator error under stress. Tested in GameDays.
Region FailoverAutomated DNS failover triggers within seconds (configurable) of primary region health check failure (GSLB). Traffic routed to DR region. Manual runbook confirms data consistency before declaring DR active.
Non-Functional Requirements — Summary Reference
NFR CategoryRequirementTarget / StandardStatus
AvailabilityPlatform availability SLO99.99% (≈ minutes per year)Mandatory
LatencyAPI P99 latency (tier-1 services)< Sub-second P99Mandatory
ThroughputCDR processing rateVery High — stream processing sustainedMandatory
RTORevenue-critical service recovery<Minutes (Tier 0 — near-zero)Mandatory
RPOMaximum data loss windowNear-zero (Tier 0 — synchronous replication)Mandatory
SecurityEncryption in transitTLS 1.3+ on external traffic, mTLS on all inter-service communicationMandatory
SecurityEncryption at restAES-256-GCM all data storesMandatory
IdentityAuthentication protocolOAuth 2.0 / OIDC + Workload Identity Standard workload identityMandatory
PII ComplianceGDPR — Right to Erasure SLAConfirmed deletion within the regulatory SLA windowMandatory
ResilienceCircuit breaker on all external callsError rate >50% in 10 s window triggers openMandatory
IdempotencyAll financial mutation endpointsIdempotency key + 24 h dedup storeMandatory
ScalabilityHPA on all stateless servicesCPU target 60%, scale within a short cool-down windowMandatory
Anti-AffinityPod distributionRequired: no two replicas on same node; Preferred: spread across AZsMandatory
Technical DebtSprint debt remediation allocationMinimum 20% per sprintRecommended
ObservabilityDistributed tracing coverage100% of requests carry traceId; 10% sampledMandatory
CI/CD SecuritySAST + SCA in every pipelineCritical/High CVEs block mergeMandatory
Non-functional architecturePrevious: Resilience & High Availability→Non-functional architectureNext: Performance & Scalability→