At a glance
Watches platform health with metrics, logs, traces and alerts.
Key data
Receives from
- Every service and platform
Sends to
- Operations teams
- Ticketing
- Reporting
How it works
The main steps, end to end.
- 1Collect metrics, logs and traces
- 2Visualise health on dashboards
- 3Detect thresholds and anomalies
- 4Alert on-call teams
- 5Track incidents to resolution
Core Capabilities
What Monitoring & Observability provides
Service Health
- Real-time health status per microservice: healthy, degraded, down
- Liveness and readiness probes per service instance
- Dependency health: data store, event bus, external integrations
- Service health dashboard aggregates status across all services
- Automated alerting on health state transitions
Metrics Collection
- Collects latency, throughput, error rate and resource utilisation per service
- Tracks key business metrics: orders per minute, activations, payments
- Time-series storage with configurable retention periods
- Custom metric registration per service for domain-specific KPIs
- Pre-built dashboards per module: billing, order flow, mediation
Alerting
- Threshold-based alerts: p95 latency, error rate, queue depth
- Anomaly detection for unusual traffic patterns or metric spikes
- Alert routing to on-call engineer via notification channel
- Alert severity levels: Info, Warning, Critical, Page
- Escalation policy if alert not acknowledged within SLA
Log Aggregation
- Aggregates structured logs from all microservices into central store
- Log levels: DEBUG, INFO, WARN, ERROR per service
- Full-text search across all service logs in a single interface
- Log correlation using trace IDs for end-to-end request tracing
- Log retention policy: hot tier (recent), cold tier (archive)
Distributed Tracing
- Traces a single request across multiple microservices
- Span-level visibility: each service call with duration and status
- Identifies bottlenecks in multi-service transaction chains
- Trace sampling configurable per environment and criticality
- Integrates with log aggregation via shared trace ID
Infrastructure Monitoring
- Monitors cluster CPU, memory, network and storage per node
- Container and pod resource utilisation across namespaces
- Storage utilisation for data store and search index clusters
- Network throughput and latency between service namespaces
- Capacity planning dashboards for infra sizing decisions
Storage & Persistence
How and where Monitoring stores its data
Time-series Store
- Per-service metrics: latency, throughput, error rate, resource utilisation
- Custom business metrics per module
- Configurable retention: hot (recent) / cold (archive) tiers
Log Store
- Centralised log aggregation from all microservices
- Structured JSON logs with trace ID correlation
- Full-text search across all service logs
Tracing Store
- Span-level trace data per request across services
- Trace-to-log correlation via shared trace ID
- Configurable sampling rate per environment
Design Principles
Key architectural decisions behind Monitoring & Observability
Unified Observability Stack
Metrics, logs and traces share a common trace ID. An alert triggered by a metric spike links directly to the relevant logs and traces for that time window — reducing mean time to diagnosis from minutes to seconds.
Alert before Impact
Monitoring tracks leading indicators (queue depth, p95 latency) alongside lagging indicators (error rate). Alerts on leading indicators fire before subscribers experience degradation, enabling proactive intervention before an issue becomes an incident.
Per-module Dashboards
Each business module has a dedicated operations dashboard with the metrics most relevant to its function. Billing teams see invoice generation rates. Order teams see order throughput. Mediation teams see CDR processing rates. Shared infrastructure dashboards sit below the business layer.
Learn normal behaviour, detect anomalies and suggest root causes. See the AIOps for BSS page.
Group related alerts into a single incident.
More in 30 AI & ML use cases and AIOps for BSS.