Home›Telecom›Architecture & guides›Non-Functional Architecture← All modules
Telecom reference
Reference Architecture

Non-Functional Architecture
Security · Resilience · Scalability

A comprehensive non-functional requirements reference — covering security zones, identity, encryption, resilience patterns, high availability, disaster recovery, observability, compliance, performance and cloud-native design principles for enterprise microservices platforms.

Security & Identity Network & Infrastructure Resilience & HA Performance Observability Compliance & Governance
Security & Identity Architecture
Defense-in-depth · Zero Trust · IAM/RBAC · Encryption · DMZ · Firewalls
Network Security Zones (Defense-in-Depth / DMZ Architecture)
Public Zone (Internet)
CDN / Edge Cache DDoS Protection WAF (Web Application Firewall) DNS / GSLB TLS 1.3 Termination
▼ Perimeter Firewall (L4/L7) ▼
DMZ — Demilitarised Zone
API Gateway Reverse Proxy / Load Balancer Identity Provider (IdP) Rate Limiter Bot Protection Certificate Management
▼ Internal Firewall + Network Policy ▼
Application Zone (Service Mesh — mTLS enforced)
Microservices (K8s Pods) Service Mesh (mTLS) Event Streaming Bus Cache Layer Internal Load Balancer
▼ Data Firewall + Network Segment Isolation ▼
Data Zone (Encrypted at Rest — AES-256)
Relational Databases NoSQL Stores Object Storage Secret Secrets Management Platform (Hardware-Backed) Key Management Service Audit Log Store
IAM — Identity & Access Management
Authentication · Authorisation · RBAC · ABAC
AuthN ProtocolOAuth 2.0 / OIDC for all human and service identities. Enterprise Federation Protocol (SAML/WS-Fed) for enterprise federation.
Token TypeJWT (Asymmetric Signing Algorithm signed). Access token TTL: configurable. Refresh token TTL: 8 h. Revocation via short-lived JTI allowlist.
RBAC ModelRole → Permission → Resource. Roles assigned to user groups. Least-privilege by default. No wildcard permissions in production.
ABAC ExtensionAttribute-based policies for fine-grained data access (e.g. agent can only view subscribers in their assigned region).
Service IdentityEvery microservice has a workload identity (Workload Identity Framework). mTLS certificates issued per pod — no shared service accounts.
MFATime-based OTP / Hardware Key / Passkey mandatory for all administrative and privileged access. Not optional for production console access.
Session MgmtStateless — no server-side sessions. Token validation at API gateway on every request. Zero-trust: no implicit trust from network location.
Zero-Trust RBAC Workload Identity Standard mTLS Identity
Encryption — At Rest, In Transit & In Use
AES-256 · TLS 1.3 · mTLS · KMS · HSM
At RestAES-256-GCM for all data at rest. Database-level encryption + application-level field encryption for PII and financial data. Keys stored in Hardware-Backed Key Management Service — never in application config.
In TransitTLS 1.3 minimum on all external traffic. TLS 1.2 deprecated. Cipher suites restricted to ECDHE + AES-GCM / CHACHA20. Certificate pinning on mobile clients.
mTLS (Service-to-Service)Mutual TLS enforced for all inter-service communication via service mesh. Certificates auto-rotated on a daily cadence. Short-lived — compromise window minimised.
Key ManagementEnvelope encryption: data encrypted with Data Encryption Key (DEK); DEK encrypted with Key Encryption Key (KEK) in KMS. Key rotation every 90 days. Automatic.
SecretsNo secrets in code, config files or environment variables. All secrets injected at runtime from a secrets vault with short-lived dynamic credentials.
In UseSensitive computations in isolated secure enclaves (TEE/Secure Enclave) for payment card processing and KYC biometric matching where applicable.
AES-256-GCM TLS 1.3 mTLS Hardware-Backed Key Management Service Envelope Encryption
Firewalls, WAF & DDoS Protection
L3/L4/L7 · WAF Rules · Rate Limiting · Anti-Bot
Perimeter FWStateful L4 firewall. Default-deny ingress. Whitelist-only egress. All rules version-controlled in IaC — no manual changes in production.
WAF (L7)OWASP Top 10 ruleset enforced. Custom rules for SQL injection, XSS, path traversal, XXE. Managed rule groups updated weekly. Block mode, not detect mode.
DDoSNetwork-layer volumetric DDoS protection at CDN edge. Application-layer protection at API gateway: per-IP and per-API-key rate limiting with adaptive throttling.
Internal FWKubernetes NetworkPolicy for pod-to-pod traffic. Application zone to data zone traffic restricted to approved microservice-to-database pairs only. No unrestricted lateral movement.
Egress ControlAll outbound traffic from application zone routed through egress proxy. Domain allowlist. Unknown destinations blocked. Egress logs to SIEM.
OWASP Top 10 Default-Deny NetworkPolicy IaC-managed
Network & Infrastructure Architecture
CDN · Load Balancers · TLS/mTLS · API Gateway · Service Mesh · Anti-Affinity
CDN & Global Traffic Management
Edge Caching · GSLB · Anycast · Geo-routing
CDNStatic assets, API responses (GET, cacheable) and portal pages served from CDN edge. TTL configured per content type. Stale-while-revalidate for catalog API responses.
Cache TTL PolicyStatic assets: 1 year (versioned). Product catalog API: 5 min. Subscriber data: 0 (no cache). Invoice PDFs: 24 h (authenticated).
GSLBGlobal Server Load Balancing routes traffic to nearest healthy data centre. Health checks at a configurable interval. Automatic failover to secondary region on health check failure.
Anycast DNSAnycast IP routing delivers requests to the geographically nearest PoP. Low-latency DNS resolution at the edge.
Edge Cache GSLB Anycast Cache-Control
Load Balancers — L4 & L7
External · Internal · Health Checks · Algorithms
External LB (L7)HTTP/2 + gRPC aware. Path-based routing to API gateway. TLS termination. Connection draining on pod scale-down. Sticky sessions via consistent hashing where required.
Internal LB (L4)Cluster Network Proxy for east-west service-to-service traffic. Round-robin by default. Least-connections for stateful services (billing, payment).
Health ChecksLiveness: container alive? Readiness: can accept traffic? Startup: slow-starting containers. All three probes mandatory. Unhealthy pods removed from LB pool within seconds.
AlgorithmRound-robin for stateless services. Consistent hash on subscriber ID for session-sensitive flows. Least outstanding requests for high-variance latency services.
L7 Routing Health Probes Connection Draining Consistent Hash
Stateless Services & Anti-Affinity
Stateless Design · Pod Anti-Affinity · Zone Spreading
Stateless microservices have no local in-memory state between requests. All session state is externalised to In-Memory Cache or a distributed cache. Any pod replica can serve any request — enabling true horizontal scalability and zero-downtime rolling deployments.
Stateless RuleNo pod may store session state, file system state or in-memory cache that is not shared. Configuration injected via environment / ConfigMap — not baked into image.
Pod Anti-AffinityRequired anti-affinity rule: no two replicas of the same service may run on the same node. Preferred anti-affinity: spread replicas across availability zones.
Zone SpreadingTopologySpreadConstraints enforce even distribution across AZs. Minimum 2 replicas per zone for tier-1 services (billing, order management, IAM).
Node AffinityRevenue-critical services (billing, rating) scheduled on dedicated node pools — isolated from dev/test workloads and batch jobs to prevent noisy-neighbour CPU contention.
Stateless Anti-Affinity Zone Spread Node Isolation
Resilience, High Availability & Idempotency
Circuit Breaker · Bulkhead · Retry · Idempotency · HA · Active-Active · Redundancy
99.99%
Target Availability
Minutes downtime per year
3
Min Availability Zones
Active-active across AZs
2+
Replicas per Service
Minimum in production
<30s
Auto-Recovery Time
Pod restart + health check
Circuit Breaker & Bulkhead
Fail-fast · Timeout · Fallback · Isolation
Circuit BreakerHalf-open → Open → Closed state machine. Opens when error rate >50% in a configurable measurement window (min 20 requests). Auto-recovers after a configurable probe interval. Configured via service mesh policy — no code change required.
TimeoutAll inter-service calls have explicit timeouts: P99 latency × 3. No unbounded waits. Timeout is non-negotiable — never rely on upstream timeout alone.
Retry PolicyMax 3 retries with exponential backoff (base-delay → 2× → 4× (with jitter)) + jitter (configurable jitter). Retry only on idempotent operations and 5xx / connection errors. Never retry on 4xx or mutation endpoints without idempotency key.
BulkheadSeparate thread/connection pool per downstream dependency. Billing calls to payment cannot exhaust threads used by billing calls to party management. Pool sizes tuned per SLA.
FallbackGraceful degradation: if notification service is unavailable, order proceeds and notification is queued for retry. Service defines explicit fallback behaviour — not just error.
Circuit Breaker Bulkhead Exponential Backoff Graceful Degradation
Idempotency Design
Idempotency Keys · Exactly-Once Semantics · Outbox Pattern
Every state-changing API operation must be idempotent — executing the same request multiple times produces the same result with no duplicate side effects. Mandatory for all payment, order and provisioning endpoints.
Idempotency KeyClient generates a Unique Idempotency Key idempotency key per request. Server stores the response keyed by this ID for 24 h. Duplicate requests within TTL receive the cached response — no re-execution.
ScopeMandatory on: POST /orders, POST /payments, POST /activations, POST /subscriptions. All provisioning adapter calls. All financial mutations.
Event DeduplicationEvent consumers maintain a deduplication store (event ID → processed timestamp). Duplicate events within a 1 h window are detected and silently discarded.
Outbox PatternDatabase writes and event publications are atomic via the transactional outbox pattern. State change is written to DB and outbox table in one transaction. Outbox poller publishes to event bus. Eliminates dual-write inconsistency.
At-Least-OnceEvent streaming delivers at-least-once. Idempotent consumers handle duplicates. Exactly-once semantics achieved via dedup + idempotent processing — not by trusting the broker.
Idempotency Key Outbox Pattern Event Deduplication Exactly-Once Logic
High Availability & Redundancy
Active-Active · Multi-AZ · Auto-Failover · PDB
TopologyActive-active across 3 availability zones within a region. No active-passive for tier-1 BSS services — single standby introduces failover latency. All zones serve live traffic.
Database HAPrimary + 2 synchronous replicas across AZs. Automatic leader election (Database HA Manager/Database HA Manager). Read replicas for analytics and reporting workloads. No single point of failure in data layer.
Pod Disruption BudgetPDB enforced: minimum 2 replicas always running during voluntary disruption (node drain, rolling upgrade). Tier-1 services: min 60% replicas available during upgrades.
Queue RedundancyEvent streaming: 3 broker replicas, replication factor 3, min-insync-replicas 2. A single broker failure does not cause message loss or producer/consumer unavailability.
Cache RedundancyCache Cluster (Sentinel mode) (3 nodes) or Cache Cluster (Sharded mode) (6 nodes) for HA cache. Cache miss falls back to database — no thundering herd protection via cache stampede mitigation (probabilistic TTL).
Active-Active Multi-AZ PDB Leader Election
Disaster Recovery & Business Continuity
RTO · RPO · Multi-Region · Backup Strategy · DR Runbooks
4hr
RTO Target
Revenue-critical services
15min
RPO Target
Maximum data loss window
2
Geographic Regions
Primary + DR site
24hr
DR Test Frequency
Chaos engineering cadence
Service TierExamplesRTORPODR StrategyBackup Freq
Tier 0 — Critical RevenueBilling, Rating, Payment15 min0 min (sync)Active-Active multi-region; real-time sync replicationContinuous (WAL shipping)
Tier 1 — Customer-FacingParty Mgmt, CRM, Orders1 hr5 minActive-Passive warm standby; async replication <5 min lagEvery 5 min (incremental)
Tier 2 — ImportantCatalog, Notification, IAM4 hr15 minActive-Passive cold standby; restore from snapshotHourly snapshots
Tier 3 — Non-CriticalReporting, DMS, Monitoring24 hr1 hrBackup and restore; rebuild from event replay if neededDaily backup
Backup Strategy
3-2-1 Rule · Immutable Backups · Tested Restores
3-2-1 Rule3 copies of data. 2 different storage media. 1 offsite (different region). Applied to all Tier 0 and Tier 1 databases.
Immutable BackupsBackups written to Immutable (Write-Once Read-Many) storage. Cannot be deleted or modified for the retention period. Ransomware protection. Retention: 90 days operational, 7 years financial.
Restore TestingAutomated restore test on a daily cadence for Tier 0. Every 72 h for Tier 1. A backup that has not been successfully restored is not a backup — untested restores are not accepted.
Event Replay DREvent streaming acts as a secondary DR mechanism. Any service can be rebuilt by replaying its domain events from the topic. Event retention: configurable hot and cold retention periods.
Chaos Engineering & DR Testing
Failure Injection · GameDay · DR Runbooks
Chaos EngineeringScheduled failure injection in pre-production: pod kill, network partition, disk full, CPU spike, dependency latency injection. Tools: Chaos Engineering Framework / Chaos Engineering Framework.
GameDay ExercisesQuarterly full-team DR simulation: trigger a region-level failure scenario in pre-production. Measure actual RTO vs target. Document gaps and remediation actions.
DR RunbooksAutomated runbooks for every DR scenario. Runbooks version-controlled in Git. Executed in minimal steps to minimise operator error under stress. Tested in GameDays.
Region FailoverAutomated DNS failover triggers within seconds (configurable) of primary region health check failure (GSLB). Traffic routed to DR region. Manual runbook confirms data consistency before declaring DR active.
Performance, Scalability & Throughput
Latency SLOs · Throughput · HPA/VPA · Auto-scaling · Caching · Technical Debt
API / OperationP50 SLOP95 SLOP99 SLOThroughput ScaleDesign Notes
Product Catalog QueryVery LowLowSub-100msHigh — scale horizontallyIn-Memory Cache-cached; cache TTL (configurable)
Subscriber Profile APIVery LowLowSub-200msHigh — scale horizontallyDB read replica + In-Memory Cache L1 cache
Order PlacementSub-500msSub-second<1 secModerate — async capableAsync provisioning; sync validation only
Real-time Rating (CDR)Very LowVery LowLowVery High — stream processingIn-memory rate table; no DB on critical path
Invoice Generation<2 sec<5 sec<10 sececBatch — parallel per bill runAsync PDF generation; batch processing
Authentication (Token)Very LowVery LowLowHigh — scale horizontallyJWT validation cached; no DB lookup
Notification DispatchSub-second<2 sec<5 secHigh — scale horizontallyAsync event-driven; SLA is delivery time not dispatch
Auto-Scaling — HPA & VPA
Horizontal · Vertical · Event-driven · Predictive
HPAHorizontal Pod Autoscaler scales replica count based on CPU (target 60%), memory (target 70%), and custom metrics (queue depth, RPS). Scale-out: seconds (configurable). Scale-in: cool-down window (configurable) to prevent thrashing.
VPAVertical Pod Autoscaler recommends optimal resource requests/limits based on observed usage. Applied in "recommend" mode — actual requests set by platform team per sprint. Prevents right-sizing drift.
Event-driven autoscalingAutoscaling driven by event backlog for event-consumer services. Scale to zero when no events. Scale from zero rapidly on first event arrival. Mediation and notification services scale on event-stream consumer lag.
Predictive ScalingScheduled HPA rules pre-scale billing pods 30 min before month-end bill run. CDR ingestion pods pre-scale before known high-traffic windows (month-end, campaigns).
HPA VPA Event-driven autoscaling Scale-to-Zero
Technical Debt Management
Debt Registry · Fitness Functions · Code Quality Platform · DORA
Technical debt is a first-class engineering concern — not a backlog item to be deferred indefinitely. Unmanaged debt compounds: the cost of deferral increases non-linearly as systems scale.
Debt RegistryEvery known technical debt item is recorded with: description, impact assessment, estimated remediation effort, and owner. Registry reviewed in monthly architecture forum.
Fitness FunctionsAutomated architectural tests run on every CI build: no cross-service database calls, no shared libraries with circular dependencies, no deprecated API versions in use, test coverage ≥80%.
Code Quality GatesCode Quality Platform quality gate: 0 new critical bugs, 0 new security hotspots, <5% code duplication on new code, coverage ≥80% on new code. Pipeline fails if gate not met.
DORA MetricsDeploy frequency, Lead time for changes, MTTR and Change failure rate tracked per service. Shared dashboard visible to engineering and leadership. Targets reviewed quarterly.
Remediation BudgetMinimum 20% of each sprint allocated to technical debt remediation. This is non-negotiable — not traded away for feature velocity. Enforced at planning.
Debt Registry Fitness Functions Code Quality Platform DORA Metrics
Observability & CI/CD
Metrics · Logs · Traces · Alerting · SLOs · Pipelines · SAST/DAST · GitOps
Three Pillars of Observability
Metrics · Structured Logs · Distributed Traces
MetricsRED method per service: Rate (req/s), Errors (error %), Duration (latency histogram). USE method per node: Utilisation, Saturation, Errors. Metrics Collection Platform scrape interval: configurable. Retention: configurable hot/cold retention tiers.
Structured LogsJSON-structured logs only — no free-text log lines. Mandatory fields: timestamp, traceId, spanId, serviceId, level, message. PII must never appear in logs — masked at emission point.
Distributed TracingW3C TraceContext propagation across all services. Every request carries a traceId. configurable sampling rate in production; 100% for errors and slow requests and slow requests (>P99). Span retention: configurable.
SLOsEach tier-1 service defines: Availability SLO (99.9%), Latency SLO (P99 < 500ms), Error Rate SLO (<0.1%). Error budget burn rate alerts at 2× and 5× burn rate. SLO dashboards public to engineering.
AlertingAlert on symptoms (SLO breach, error budget burn), not causes (CPU%). Multi-window multi-burn-rate alerting (1h/6h windows). Runbook URL mandatory in every alert. On-call rotation with defined escalation path.
RED Method SLO/SLA W3C TraceContext Error Budget No PII in Logs
CI/CD Pipeline — Security & Quality Gates
SAST · DAST · SCA · Container Scan · GitOps
Pipeline Stages — Every Service
Build → Unit Tests → SAST Scan → SCA → Image Scan
Contract Tests → Integration Tests → DAST Scan → Deploy (Canary)
SASTStatic analysis on every commit. Blocks merge on critical/high severity findings. OWASP Top 10, injection flaws, insecure deserialization, hardcoded secrets.
SCASoftware Composition Analysis: all third-party dependencies scanned for CVEs. CVSS ≥7.0 blocks pipeline. Licence compliance checked — GPL dependencies flagged.
Container ScanningBase image and application layer scanned for CVEs before push to registry. Only images from approved base registries accepted. Distroless images preferred.
DASTDynamic Application Security Testing runs against deployed service in integration environment. DAST Scanner / DAST Scanner. Automated attack scenarios including SQLi, XSS, IDOR.
GitOpsKubernetes manifests in Git are the single source of truth for cluster state. GitOps Controller/GitOps Controller continuously reconciles. No manual kubectl apply in production. All changes via PR with review.
SAST DAST SCA GitOps Distroless Canary Deploy
Compliance & Data Governance
PII Compliance · GDPR · Data Classification · Audit Trail · Retention Policies
PII & Data Protection
GDPR · Data Minimisation · Pseudonymisation · Right to Erasure
Data Classification4-tier: Public · Internal · Confidential · Restricted (PII/PCI). Classification applied at field level — not document level. Drives encryption, access control and retention rules.
PII HandlingPII fields (name, address, MSISDN, IMEI, NIN) encrypted at field level with service-specific encryption keys. PII never logged. PII masked in non-production environments. PII map maintained per service.
Data MinimisationServices collect only the PII fields they are authorised to process. No data hoarding — PII not replicated to services that do not require it. Enforced via API contract reviews.
Right to ErasureErasure request triggers a propagated deletion event across all services holding subscriber PII. Services must confirm deletion within SLA (72 h). Audit trail of deletion is retained (without the PII).
PseudonymisationAnalytics and reporting pipelines use pseudonymised subscriber IDs — not real MSISDN or account IDs. Mapping table held in restricted access vault accessible only to authorised data teams.
Audit TrailImmutable audit log of every read and write on PII fields. Log entries: who accessed, what data, from which service, when. Audit logs stored in tamper-evident write-once storage. Retained 7 years.
GDPR Field Encryption Data Minimisation Right to Erasure Pseudonymisation Immutable Audit Log
Data Retention & Lifecycle
Retention Policies · Automated Purge · Archival · Legal Hold
Data TypeHot RetentionArchivePurge
Subscriber PIIActive + 6 months post-churn6–13 months cold13 months (unless legal hold)
Call Records (CDR)6 months6–24 months24 months (regulatory minimum)
Invoice & Billing2 years2–7 years cold7 years (financial regulation)
Order History2 years2–5 years5 years
Audit Logs90 days90 days – 7 years7 years (never auto-purge)
Application Logs30 days30–90 days90 days
Non-Functional Requirements — Summary Reference
NFR CategoryRequirementTarget / StandardStatus
AvailabilityPlatform availability SLO99.99% (≈ minutes per year)Mandatory
LatencyAPI P99 latency (tier-1 services)< Sub-second P99Mandatory
ThroughputCDR processing rateVery High — stream processing sustainedMandatory
RTORevenue-critical service recovery<Minutes (Tier 0 — near-zero)Mandatory
RPOMaximum data loss windowNear-zero (Tier 0 — synchronous replication)Mandatory
SecurityEncryption in transitTLS 1.3+ on external traffic, mTLS on all inter-service communicationMandatory
SecurityEncryption at restAES-256-GCM all data storesMandatory
IdentityAuthentication protocolOAuth 2.0 / OIDC + Workload Identity Standard workload identityMandatory
PII ComplianceGDPR — Right to Erasure SLAConfirmed deletion within the regulatory SLA windowMandatory
ResilienceCircuit breaker on all external callsError rate >50% in 10 s window triggers openMandatory
IdempotencyAll financial mutation endpointsIdempotency key + 24 h dedup storeMandatory
ScalabilityHPA on all stateless servicesCPU target 60%, scale within a short cool-down windowMandatory
Anti-AffinityPod distributionRequired: no two replicas on same node; Preferred: spread across AZsMandatory
Technical DebtSprint debt remediation allocationMinimum 20% per sprintRecommended
ObservabilityDistributed tracing coverage100% of requests carry traceId; 10% sampledMandatory
CI/CD SecuritySAST + SCA in every pipelineCritical/High CVEs block mergeMandatory