SRE First Microservices Monitoring: 3–5 SLIs, OpenTelemetry, One Sprint
SRE First Microservices Monitoring: 3–5 SLIs, OpenTelemetry, One Sprint

Effective microservices monitoring means tracking user-facing reliability through Service Level Objectives, correlated across metrics, logs, and traces rather than isolated dashboards per service. The immediate next step for most teams is to pick three to five Service Level Indicators and map the four golden signals to your most critical services this sprint. OpenTelemetry has become the standard for exporting and correlating that telemetry across polyglot service fleets.
TL;DR:
- Monitoring a microservices fleet requires collecting and correlating metrics, logs, and traces across service boundaries, not just individual service dashboards.
- Using OpenTelemetry with a focus on core signals and minimal custom instrumentation ensures scalable, vendor-neutral telemetry collection that supports effective correlation.
- Implementing SLOs based on key SLIs like latency, error rate, and business metrics helps prevent alert fatigue and prioritize systemic issues that impact user experience.
- Combining anomaly detection models with alerting on error budgets provides earlier incident signals, especially for slow degradations that fixed thresholds might miss.
- A unified operational layer that integrates telemetry sources reduces incident investigation time by providing a consolidated, read-only interface for fast root cause analysis.
Table of Contents
- What is microservices monitoring?
- Why monitoring distributed systems demands a different approach
- Key metrics: golden signals plus business context
- Using metrics, logs, and traces together
- Distributed tracing and OpenTelemetry in Kubernetes
- Choosing monitoring tools that fit a microservices fleet
- A step-by-step monitoring workflow you can run this sprint
- Common challenges at scale and how to mitigate them
- Security considerations for microservices monitoring
- Machine learning and anomaly detection in microservices monitoring
- Bringing it together with a unified operational layer
- Sources
- FAQ
What is microservices monitoring?
Monitoring a microservices architecture means tracking the health of dozens or hundreds of independently deployed services, plus the network calls, message queues, and infrastructure that connect them. A monolith gives you one process to watch. A microservices fleet gives you ephemeral containers, rolling deployments, and inter-service calls that can fail in combinations no single dashboard captures.
Comprehensive telemetry collection spans several layers:
- Service-level metrics: request rate, error rate, and latency per endpoint.
- Infrastructure metrics: CPU, memory, and disk pressure on the nodes running your workloads.
- Network and mesh metrics: retries, timeouts, and connection pool exhaustion between services.
- Queue and messaging metrics: consumer lag, throughput, and dead-letter volume for asynchronous systems.
Success is not “all dashboards green.” It is defined in user-facing terms: SLIs that reflect what customers experience, and SLOs that set the bar for acceptable reliability, an approach the Kubernetes observability documentation frames around the three pillars of metrics, logs, and traces rather than raw resource utilization alone.
Why monitoring distributed systems demands a different approach
Monolith monitoring answers one question: is the process healthy? Microservices monitoring has to answer that question for every service, plus a harder one: is the request healthy across every service it touched? Ephemeral pods rotate every few minutes under autoscaling, so any monitoring approach built around static hosts breaks quickly.
This shift changes how SRE and DevOps teams work day to day:
- Dynamic topology increases correlation demand: services scale, restart, and reschedule constantly, so metrics must be tagged with identity that survives pod churn.
- The unit of health becomes the request, not the service: a single checkout flow might touch a dozen services, and a slowdown in one buried dependency can look like a symptom everywhere else.
- Ownership and runbooks fragment by design: each team owns its service, so incident response depends on shared tracing and consistent alerting conventions rather than one team knowing the whole system.
Without correlation across services, teams miss systemic issues like cascading failures that start in one dependency and spread outward before anyone notices the root cause.
Key metrics: golden signals plus business context
Start with the four golden signals defined in the Google SRE book: latency, traffic, errors, and saturation. Each one maps cleanly to microservice behavior.
- Latency: time per request, measured in percentiles (p50, p95, p99) rather than averages, since a slow tail often hides behind a healthy mean.
- Traffic: request volume per service or endpoint, useful for spotting unexpected drops or spikes tied to upstream changes.
- Errors: the rate of failed requests, broken down by status code or exception type so you can tell a client error from a dependency failure.
- Saturation: how close a resource, CPU, memory, connection pool, or queue depth, is to its limit.
Relying on the four golden signals alone often produces alert fatigue. Google’s SRE guidance recommends correlating them with SLOs and watching tail-latency percentiles specifically, since averages routinely mask the requests that actually hurt users.
Beyond the golden signals, add business or custom metrics that reflect real user experience: checkout completion rate, search result relevance, or login success rate. These metrics matter because they catch failures that look technically fine but are broken for the customer.
To convert this into an operating model, pick a small number of SLIs, set an SLO target for each with a high reliability goal, and calculate an error budget. The budget becomes your shared language for how much unreliability is acceptable before you stop shipping features and fix reliability instead.
Using metrics, logs, and traces together
Each observability pillar answers a different question, and none of them alone gives you the full picture. The Kubernetes documentation on observability describes metrics as aggregated system performance, logs as granular event history, and traces as the end-to-end record of a request’s journey.
- Metrics tell you something is wrong and roughly how bad it is, fast enough to alert on.
- Logs tell you the specific event or error message tied to a failure.
- Traces tell you where in the request path the failure happened and how it propagated across services.
The correlation pattern that makes this work: a trace ID attached to both logs and metrics lets you jump from an alert straight to the specific request and the exact log lines that explain it, instead of guessing which service to check first.
SLO-driven alerting keeps this practical. Paging thresholds should trigger only when error budget burn threatens the SLO, not on every metric blip. Everything else, capacity trends, minor latency drift, belongs on a dashboard or an automated ticket, not in someone’s pager at 2 a.m.
Pro Tip: Alert on symptoms that burn error budget, and route everything else to dashboards or automated tickets instead of pages.
Distributed tracing and OpenTelemetry in Kubernetes
OpenTelemetry Protocol, OTLP, is the recommended export format for traces, metrics, and logs across a microservices stack, and it is what most modern collectors and backends expect natively. Kubernetes components themselves can export traces via OTLP, and clusters commonly default to port 4317 for that traffic, according to Kubernetes system traces documentation.
Practical setup decisions matter more than tooling brand names:
- Collector placement: run a sidecar collector for per-pod enrichment or a centralized collector deployment for simpler fleet-wide configuration, depending on how much per-service context you need.
- Sampling strategy: kube-apiserver and kubelet tracing configs support sampling controls like
samplingRatePerMillion, letting you keep full fidelity on error paths while sampling healthy traffic down to control volume. - Trace IDs in logs: every log line should carry the trace ID of the request that produced it, otherwise correlation across the three pillars breaks down exactly when you need it most.
- Semantic conventions: adopting the OpenTelemetry semantic conventions for HTTP, messaging, and database spans keeps naming consistent across teams writing services in different languages.
Full tracing on every request adds measurable CPU and network overhead. Sampling on healthy paths while keeping errors and slow requests at full fidelity is the standard trade-off.
Pro Tip: Instrument one high-traffic customer flow end-to-end before rolling tracing out fleet-wide. It surfaces integration gaps early.
Choosing monitoring tools that fit a microservices fleet
Tool selection should start with OpenTelemetry compatibility, not brand reputation. Native OTLP ingestion means new services get correlated telemetry from day one instead of a custom exporter for every backend.
Beyond OTEL support, Redis’s guidance on choosing a microservice monitoring tool points to a practical checklist:
- Scale and retention: can the backend handle your cardinality and keep enough history for meaningful SLO trend analysis.
- Cost model: does pricing scale predictably with data volume, or does it punish you for instrumenting more services.
- Alerting and incident integration: does it plug into your paging and ticketing tools without custom glue code.
- Runbook automation: can alerts trigger documented response steps automatically, rather than relying on tribal knowledge.
- Avoiding lock-in: does the vendor let you export your own telemetry, or does switching later mean re-instrumenting everything.
Prioritizing tools that natively integrate OpenTelemetry avoids the vendor lock-in that comes from proprietary agents, and it keeps cross-signal correlation intact as your stack evolves. Platform teams juggling a growing number of point tools often reach a stage where reducing tool sprawl becomes as important as any single tool’s feature set.
A step-by-step monitoring workflow you can run this sprint
Rolling out SLO-driven monitoring does not require a quarter-long project. A focused sprint on one customer flow proves the model before you scale it.
- Pick one customer-facing flow: checkout, login, or search are common starting points because they map directly to revenue or retention.
- Define three to five SLIs: latency, error rate, and one business metric specific to that flow.
- Instrument end-to-end with OpenTelemetry: add OTLP export at every service in the flow, and confirm trace IDs propagate through logs.
- Deploy collectors: choose sidecar or centralized placement, and configure sampling so error paths stay at full fidelity.
- Build one dashboard per flow: golden signals plus your custom SLI, not a wall of infrastructure charts.
- Set SLO targets and error budgets: agree on the number with the team that owns the service, not just the SRE on call.
- Configure SLO-driven alerts: page only when burn rate threatens the budget within a defined window.
- Write the runbook before the first incident: link the alert directly to remediation steps and the relevant dashboard.
- Review after each incident: update the SLO, the runbook, or the sampling rate based on what the postmortem reveals.
| Step | Primary output | Owner |
|---|---|---|
| Select SLIs | Three to five metrics tied to user experience | Service team |
| Instrument | OTLP traces, metrics, and logs with shared trace IDs | Engineering |
| Deploy collectors | Sidecar or centralized OTLP collector configuration | Platform team |
| Set SLOs and alerts | Error budget and burn-rate paging thresholds | SRE |
| Runbook and review | Documented response steps, updated after each incident | On-call rotation |
Teams applying this to platform-wide rollouts often extend the same pattern across every service a platform engineering team owns, rather than repeating the exercise from scratch per flow.
Common challenges at scale and how to mitigate them
As the fleet grows, three failure modes show up repeatedly: alert fatigue, cascading failures, and runaway telemetry cost.
- Alert fatigue: too many low-value pages train on-call engineers to ignore alerts. SLO-driven alerting, paging only on error budget burn, cuts this at the source rather than tuning thresholds forever.
- Cascading failures: a single slow dependency can degrade a dozen services downstream. Dependency mapping and automated impact analysis let you see which services actually depend on the failing one, instead of guessing during an incident.
- Telemetry cost growth: full-fidelity tracing and unlimited log retention get expensive fast. Sampling on healthy traffic and tiered retention, full detail for recent data, aggregated summaries for older data, keeps cost predictable.
- Ownership drift: without clear service ownership, alerts get ignored or escalated to the wrong team. Regular SLO reviews keep ownership, thresholds, and runbooks current as services change hands.
None of these problems are solved by adding more dashboards. They are solved by correlating signals across services so the root cause surfaces automatically instead of requiring a manual hunt through a dozen tools.
Security considerations for microservices monitoring
Monitoring infrastructure itself becomes an attack surface in a microservices architecture, since collectors, dashboards, and alerting pipelines often have broad read access across the fleet. A few principles reduce that exposure without slowing down observability work.
Telemetry pipelines should carry only the access they need. Collectors that scrape metrics or receive traces do not need to write access to production systems, and keeping that boundary read-only limits the damage from a compromised agent. Trace and log data frequently contains sensitive fields, request payloads, user identifiers, authentication tokens, so redaction or scrubbing at the collector level matters before that data lands in a long-lived backend.

Access to dashboards and alert configurations should follow the same role-based controls as production systems, since a monitoring stack that reveals internal service topology and error patterns is useful reconnaissance if it falls into the wrong hands. Audit trails on who changed an alert threshold or silenced a page matter during incident review, since a silenced alert can hide an active attack as easily as a false positive.
Finally, monitoring and security operations increasingly overlap: anomalous traffic spikes, unusual error patterns, or unexpected saturation on a service can be the earliest signal of an active incident, whether the cause is a bug or an intrusion. Treating monitoring and security telemetry as separate systems means slower detection either way.
Machine learning and anomaly detection in microservices monitoring
Static thresholds struggle in a microservices fleet because normal traffic patterns vary by time of day, deployment cadence, and seasonal load, making a single fixed alert threshold either too noisy or too slow. Anomaly detection models trained on historical metric behavior can flag deviations that a fixed threshold would miss entirely, catching a slow degradation before it crosses any hard-coded line.
The practical value shows up in three areas. First, baseline learning: instead of a human guessing what “normal” latency looks like for two hundred services, a model can learn each service’s own pattern and flag deviations from it. Second, correlation across signals: machine learning approaches can surface which metrics moved together during an incident, narrowing root cause analysis faster than manually scanning dashboards one service at a time. Third, noise reduction: by learning which combinations of symptoms typically resolve on their own versus which precede real incidents, anomaly detection can reduce the volume of low-value alerts reaching on-call engineers.
This does not replace SLO-driven alerting, it complements it. SLOs still define what “acceptable” means for the business. Anomaly detection helps catch the cases where something is clearly abnormal even though it has not yet crossed an SLO threshold, giving teams earlier warning on emerging issues. The practical limit is that models need enough historical data and stable enough patterns to be useful, so a newly launched service with erratic early traffic is a poor candidate for anomaly-based alerting until it settles into a predictable baseline.
Bringing it together with a unified operational layer
Most teams following this playbook end up with OpenTelemetry-instrumented services, an SLO dashboard, and alerting rules, but the actual incident investigation still means jumping between the tracing backend, the log search tool, and the cloud console to find the root cause. That gap between having the telemetry and actually using it fast during an incident is where a lot of mean-time-to-resolution gets lost.
Opsphere connects to the observability stack, cloud providers, and CI/CD tools you already run, without requiring you to replace any of them, and gives engineers a single interface to investigate incidents using correlated context from across the fleet. Instead of manually stitching together a trace ID, a log query, and a Kubernetes event, an SRE can ask a direct question about a service’s health and get an answer built from real operational data, with the underlying tools remaining the source of truth and every read staying strictly read-only for security.
For platform teams managing dozens of services and a growing pile of point tools, this means less time spent context-switching during an incident and more time spent on the fix. Teams can start on the Developer plan at €19 per month or explore the Community tier at no cost before scaling to Team or Enterprise pricing as the fleet grows.

Explore the Opsphere platform to see how unified operational context fits into your existing OpenTelemetry and SLO workflow, or check current pricing and plan details to find the right starting point for your team.
Sources
- Kubernetes: Observability
- Google SRE Book: Monitoring distributed systems
- OpenTelemetry: Semantic conventions
FAQ
What is the difference between monitoring and observability?
Monitoring is the practice of collecting and alerting on predefined metrics, while observability is the broader capacity to ask new questions about system behavior using metrics, logs, and traces together. The Kubernetes observability documentation frames these three pillars as the foundation for understanding distributed systems beyond fixed dashboards.
What are the four golden signals in SRE monitoring?
The four golden signals are latency, traffic, errors, and saturation, as defined in the Google SRE book. They give teams a consistent baseline for evaluating any service’s health regardless of what it does internally.
How do I set an SLO for a microservice?
Start by picking an SLI that reflects user experience, such as request latency or success rate, then set a target threshold and time window that matches your reliability goals. Google’s SRE guidance recommends choosing a small number of indicators rather than trying to monitor everything, since too many diluted metrics make paging decisions harder, not easier.
Why is OpenTelemetry recommended for microservices tracing?
OpenTelemetry provides a vendor-neutral protocol, OTLP, and a shared set of semantic conventions that let traces, metrics, and logs from different services and languages correlate consistently. The OpenTelemetry semantic conventions standardize naming across HTTP, messaging, and database spans, which is what makes cross-service root cause analysis practical at scale.
How does Opsphere fit into an existing observability stack?
Opsphere connects to your existing observability, cloud, and CI/CD tools without replacing them, giving engineers a unified interface to investigate incidents using correlated operational context. It maintains strict read-only access to the connected tools, which remain the source of truth for their own data.
Recommended
This article is provided for general informational purposes only and does not constitute professional, legal, security, or compliance advice. Please evaluate recommendations against your organization’s specific environment and requirements.
