Home›Telecom›Operations›Monitoring← All modules
Operations · BSS module

Monitoring & Observability

The Monitoring module provides end-to-end observability across all platform services — aggregating metrics, logs and traces into unified dashboards, enabling real-time health visibility, proactive alerting and rapid incident diagnosis.

Service Health Metrics Alerting Log Aggregation Distributed Tracing Infrastructure
On this pageAt a glanceHow it worksCore CapabilitiesStorage & PersistenceDesign PrinciplesWhere AI helpsStandards

At a glance

Watches platform health with metrics, logs, traces and alerts.

Key data

MetricLogTraceAlert

Receives from

  • Every service and platform
Monitoring

Sends to

  • Operations teams
  • Ticketing
  • Reporting

How it works

The main steps, end to end.

  1. 1
    Collect metrics, logs and traces
  2. 2
    Visualise health on dashboards
  3. 3
    Detect thresholds and anomalies
  4. 4
    Alert on-call teams
  5. 5
    Track incidents to resolution

Core Capabilities

What Monitoring & Observability provides

Service Health

Health Checks
  • Real-time health status per microservice: healthy, degraded, down
  • Liveness and readiness probes per service instance
  • Dependency health: data store, event bus, external integrations
  • Service health dashboard aggregates status across all services
  • Automated alerting on health state transitions
All ServicesAlerting

Metrics Collection

Observe
  • Collects latency, throughput, error rate and resource utilisation per service
  • Tracks key business metrics: orders per minute, activations, payments
  • Time-series storage with configurable retention periods
  • Custom metric registration per service for domain-specific KPIs
  • Pre-built dashboards per module: billing, order flow, mediation
Time-series StoreAll Services

Alerting

Notify On-call
  • Threshold-based alerts: p95 latency, error rate, queue depth
  • Anomaly detection for unusual traffic patterns or metric spikes
  • Alert routing to on-call engineer via notification channel
  • Alert severity levels: Info, Warning, Critical, Page
  • Escalation policy if alert not acknowledged within SLA
On-call SystemNotification Gateway

Log Aggregation

Centralise Logs
  • Aggregates structured logs from all microservices into central store
  • Log levels: DEBUG, INFO, WARN, ERROR per service
  • Full-text search across all service logs in a single interface
  • Log correlation using trace IDs for end-to-end request tracing
  • Log retention policy: hot tier (recent), cold tier (archive)
Log StoreObservability Platform

Distributed Tracing

End-to-End Traces
  • Traces a single request across multiple microservices
  • Span-level visibility: each service call with duration and status
  • Identifies bottlenecks in multi-service transaction chains
  • Trace sampling configurable per environment and criticality
  • Integrates with log aggregation via shared trace ID
Tracing Platform

Infrastructure Monitoring

Platform Health
  • Monitors cluster CPU, memory, network and storage per node
  • Container and pod resource utilisation across namespaces
  • Storage utilisation for data store and search index clusters
  • Network throughput and latency between service namespaces
  • Capacity planning dashboards for infra sizing decisions
InfrastructureOperations

Storage & Persistence

How and where Monitoring stores its data

Time-series Store

Metrics
  • Per-service metrics: latency, throughput, error rate, resource utilisation
  • Custom business metrics per module
  • Configurable retention: hot (recent) / cold (archive) tiers

Log Store

Structured logs
  • Centralised log aggregation from all microservices
  • Structured JSON logs with trace ID correlation
  • Full-text search across all service logs

Tracing Store

Distributed traces
  • Span-level trace data per request across services
  • Trace-to-log correlation via shared trace ID
  • Configurable sampling rate per environment

Design Principles

Key architectural decisions behind Monitoring & Observability

Unified Observability Stack

Metrics, logs and traces share a common trace ID. An alert triggered by a metric spike links directly to the relevant logs and traces for that time window — reducing mean time to diagnosis from minutes to seconds.

Alert before Impact

Monitoring tracks leading indicators (queue depth, p95 latency) alongside lagging indicators (error rate). Alerts on leading indicators fire before subscribers experience degradation, enabling proactive intervention before an issue becomes an incident.

Per-module Dashboards

Each business module has a dedicated operations dashboard with the metrics most relevant to its function. Billing teams see invoice generation rates. Order teams see order throughput. Mediation teams see CDR processing rates. Shared infrastructure dashboards sit below the business layer.

AI opportunity
Where AI helps
AIOps

Learn normal behaviour, detect anomalies and suggest root causes. See the AIOps for BSS page.

Alert noise reduction

Group related alerts into a single incident.

More in 30 AI & ML use cases and AIOps for BSS.

Standards & references
OpenTelemetryMetrics, logs and traces standard
TMF642Alarm Management API