Microservices Observability: OpenTelemetry, Metrics, Logs & Distributed Tracing
Learn microservices observability with OpenTelemetry, Node.js instrumentation, W3C Trace Context, distributed tracing, metrics, logs, and production monitoring.
System Design: From Zero to Production
A comprehensive engineering series guiding backend developers from single-node instances to highly resilient distributed architectures.
Microservices Observability: The Three Pillars, OpenTelemetry, and Distributed Tracing
Modern microservices architectures distribute application logic across multiple services, databases, APIs, and asynchronous workers. Although this architecture improves modularity and independent scalability, it also makes troubleshooting production incidents more challenging.
When an API becomes slow, a database query times out, or a background worker stops processing messages, engineers need more than basic application logs to identify the underlying problem.
This is where microservices observability becomes essential.
Observability enables engineering teams to understand application behavior using telemetry collected from the system. The three foundational telemetry signals are metrics, logs, and distributed traces. Together, they help developers detect anomalies, investigate failures, analyze latency, and improve system reliability.
OpenTelemetry provides a vendor-neutral framework for generating, collecting, and exporting this telemetry across distributed applications.
In this guide, you will learn how the three pillars of observability work, how W3C Trace Context propagates tracing information across services, and how to instrument a Node.js application with OpenTelemetry.
What Is Observability in Microservices?
Microservices observability is the practice of collecting and analyzing telemetry to understand the internal behavior of distributed applications.
Unlike basic monitoring, which typically checks predefined conditions such as CPU usage or HTTP error rates, observability helps engineers investigate unfamiliar failures by examining the evidence produced by the system.
For example, suppose a request passes through an API gateway, a user service, PostgreSQL, and a Kafka consumer. If the request experiences unexpected latency, observability tools can help identify which operation consumed the most time and whether the issue originated in the application, database, or messaging layer.
A production observability strategy typically combines:
- Metrics: Numerical measurements that reveal system-wide behavior.
- Logs: Structured event records that provide diagnostic details.
- Distributed traces: Records of operations and their relationships across service boundaries.
These signals become more useful when they share consistent service metadata and trace identifiers.
The Three Pillars of Observability
1. Metrics: Detecting Performance Problems
Metrics are numerical measurements collected over time. They provide an aggregated view of application health and infrastructure performance.
Common microservices metrics include:
- Request throughput and requests per second.
- HTTP 4xx and 5xx error rates.
- Request latency at p50, p95, and p99.
- CPU utilization and memory consumption.
- Database connection-pool saturation.
- Kafka consumer lag and queue depth.
- Garbage collection frequency and duration.
Metrics are particularly useful for dashboards, alerting, capacity planning, and service-level objectives (SLOs).
For example, increasing p99 latency combined with a rising HTTP 5xx rate may indicate a downstream dependency failure or resource bottleneck.
Production tip: Avoid high-cardinality metric labels such as user IDs, trace IDs, and unrestricted URL paths. These can significantly increase storage and processing costs.
2. Logs: Understanding Application Events
Logs capture individual events generated by applications and infrastructure. Structured logging makes these events easier to search, filter, and correlate with other telemetry.
Example structured log:
json{ "timestamp": "2026-10-03T10:15:30.123Z", "severity": "ERROR", "service.name": "user-service", "message": "Database query failed", "trace_id": "4bf92f3577b34da6a3ce929d0e0e4736", "span_id": "00f067aa0ba902b7", "error.type": "DatabaseTimeout", "duration_ms": 3000 }
With trace and span identifiers, an engineer can navigate from a failed distributed trace to the corresponding application logs.
Useful logging practices include:
- Prefer structured JSON logs over unstructured messages.
- Include service names, deployment environments, and timestamps.
- Record meaningful exception details and operation outcomes.
- Use appropriate log levels to reduce unnecessary noise.
- Redact credentials, access tokens, and sensitive personal information.
Logs explain what happened at a particular point in time. Distributed traces complement them by revealing how operations relate across services.
3. Distributed Tracing: Finding the Root Cause of Latency
Distributed tracing follows a request as it moves through multiple services and dependencies.
A trace represents the overall execution path, while a span represents an individual operation, such as an HTTP request, a database query, or a message-processing task.
Consider a request with the following execution times:
| Operation | Duration |
|---|---|
| API gateway | 20 ms |
| User service | 45 ms |
| PostgreSQL query | 800 ms |
| Response processing | 15 ms |
In this example, the database operation is the most significant latency contributor.
Distributed tracing helps engineers identify such bottlenecks without relying exclusively on manually inserted log statements.
It is particularly valuable when applications use:
- Multiple REST or gRPC services.
- PostgreSQL, MySQL, or other database systems.
- Kafka, RabbitMQ, or asynchronous job queues.
- Background workers and scheduled tasks.
- Retries, timeouts, and circuit breakers.
Learn more in the official OpenTelemetry tracing documentation.
OpenTelemetry: A Vendor-Neutral Observability Framework
OpenTelemetry (OTel) is an open-source observability framework for instrumenting applications and collecting, processing, and exporting telemetry.
It provides APIs, SDKs, instrumentation libraries, semantic conventions, and the OpenTelemetry Protocol (OTLP).
Its main components include:
| Component | Purpose |
|---|---|
| OpenTelemetry API | Defines interfaces for generating telemetry |
| OpenTelemetry SDK | Implements telemetry creation, processing, and export |
| Auto-instrumentation | Captures supported operations from frameworks and libraries |
| OTLP | Transports telemetry between compatible components |
| OpenTelemetry Collector | Receives, processes, and exports telemetry |
| Observability backend | Stores and visualizes telemetry and supports queries and alerts |
OpenTelemetry reduces dependence on vendor-specific instrumentation by allowing compatible backends to consume standardized telemetry.
However, dashboards, alert rules, query languages, and backend-specific features may still require changes when migrating between providers.
For an overview, see the official OpenTelemetry documentation.
W3C Trace Context Propagation in Distributed Systems
Collecting traces from individual services is not enough. Those traces must also be connected across network requests and asynchronous operations.
The W3C Trace Context specification defines standard HTTP headers for propagating tracing information between compatible systems.
The two primary headers are:
traceparent: Carries the trace ID, parent span ID, and trace flags.tracestate: Carries optional vendor-specific tracing information.
Example:
httptraceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
The header contains four fields:
| Field | Example | Description |
|---|---|---|
| Version | 00 | Trace Context format version |
| Trace ID | 4bf92f3577b34da6a3ce929d0e0e4736 | Identifies the distributed trace |
| Parent ID | 00f067aa0ba902b7 | Identifies the caller's span |
| Trace flags | 01 | Indicates that the sampled flag is set |
When an instrumented service receives a request, it extracts the incoming context and creates the appropriate span. When it calls another service, the tracing instrumentation propagates the context.
This allows compatible tracing systems to connect operations into a distributed execution graph.
How Trace Context Works Across Kafka
Asynchronous messaging introduces another propagation boundary. A producer may publish an event and return an HTTP response before a consumer processes the event.
Without explicit or automatic message-context propagation, the consumer's work may appear as an unrelated trace.
A typical workflow looks like this:
Architecture DiagramMermaid Flow
The producer propagates trace context through Kafka message headers, and the consumer extracts it during processing.
Depending on the messaging instrumentation and processing model, the relationship may be represented through parent-child spans or span links.
For more details, consult the OpenTelemetry JavaScript documentation and the W3C Trace Context specification.
How to Set Up OpenTelemetry in Node.js
OpenTelemetry provides Node.js instrumentation for supported frameworks and libraries. With automatic instrumentation, many HTTP requests, outbound calls, and supported database operations can generate spans without extensive manual code changes.
Step 1: Install the Required Packages
Install the Node.js SDK, tracing exporter, and automatic instrumentation package:
bashnpm install \ @opentelemetry/api \ @opentelemetry/sdk-node \ @opentelemetry/exporter-trace-otlp-http \ @opentelemetry/auto-instrumentations-node
These packages provide the API, SDK, OTLP HTTP trace exporter, and supported automatic instrumentation libraries.
Step 2: Configure the OpenTelemetry SDK
Create an initialization file named instrumentation.ts:
typescriptimport { NodeSDK } from "@opentelemetry/sdk-node"; import { OTLPTraceExporter } from "@opentelemetry/exporter-trace-otlp-http"; import { getNodeAutoInstrumentations } from "@opentelemetry/auto-instrumentations-node"; const traceExporter = new OTLPTraceExporter({ url: process.env.OTEL_EXPORTER_OTLP_TRACES_ENDPOINT ?? "http://localhost:4318/v1/traces", }); const sdk = new NodeSDK({ traceExporter, instrumentations: [getNodeAutoInstrumentations()], }); sdk.start();
The endpoint shown is an example. Configure it to point to an OpenTelemetry Collector or another compatible OTLP HTTP endpoint in your environment.
Step 3: Initialize Instrumentation Before Application Code
OpenTelemetry must initialize before the application imports libraries that need automatic instrumentation.
For a compatible CommonJS build, an example launch command is:
bashnode --require ./dist/instrumentation.js ./dist/server.js
Adapt the paths and initialization method to your TypeScript build process and Node.js module system. ESM applications may require a different preload configuration.
For additional setup instructions, follow the official OpenTelemetry Node.js getting-started guide.
Step 4: Configure the OpenTelemetry Collector
For production deployments, an OpenTelemetry Collector can receive application telemetry, process it, and export it to compatible observability backends.
Architecture DiagramMermaid Flow
The Collector can support batching, filtering, resource enrichment, and export configuration. Its buffering and retry behavior depend on the deployed components and configuration, so monitor resource limits and dropped telemetry.
Official reference: OpenTelemetry Collector documentation.
Important: The Node.js example above configures trace export. Metrics and logs require their respective collection and export configuration; they are not automatically enabled merely by installing the trace exporter.
Measuring OpenTelemetry Performance Overhead
Observability adds some processing, memory, and network overhead. The actual impact depends on request throughput, the number of spans generated, attribute volume, sampling configuration, exporter behavior, and the collector architecture.
Rather than assuming that instrumentation has negligible overhead, measure it under representative workloads.
| Metric | What to measure |
|---|---|
| CPU overhead | Additional CPU consumption with instrumentation enabled |
| Memory overhead | Additional application and collector memory usage |
| Request latency | Changes in p50, p95, and p99 latency |
| Export latency | Time spent processing and exporting telemetry |
| Dropped telemetry | Spans, metric points, and log records lost under load |
| Backend ingestion | Telemetry throughput and ingestion delays |
| Incident detection time | Time from a detectable issue to its identification |
A meaningful benchmark should document hardware, runtime versions, request distribution, sampling configuration, telemetry volume, collector topology, test duration, and baseline measurements.
Avoid publishing fixed overhead percentages or claims of reduced incident time without a reproducible evaluation. Instrumentation overhead and incident detection improvements are workload-dependent.
Production Best Practices for Microservices Observability
Use Consistent Service Metadata
Configure resource attributes such as service name, service version, and deployment environment. Consistent metadata makes it easier to filter telemetry and compare service behavior across deployments.
Correlate Logs and Traces
Include trace and span identifiers in structured logs when available. This makes it easier to move from a failed trace to the detailed application events associated with the operation.
Apply Sampling Strategically
Sampling can reduce telemetry costs, but sampled-out requests may be unavailable during an investigation. Consider the trade-off between telemetry volume and diagnostic coverage.
Tail sampling can retain traces based on characteristics such as errors or latency, but requires an appropriate collector architecture and sufficient trace data to make the sampling decision.
Protect Sensitive Information
Avoid exporting passwords, access tokens, confidential request bodies, or unnecessary personal information. Apply redaction and access controls throughout the telemetry pipeline.
Monitor the Telemetry Pipeline
The Collector and observability backend can also experience failures. Monitor exporter errors, queue utilization, resource consumption, ingestion latency, and dropped telemetry.
Build Alerts Around Service-Level Objectives
Use actionable alerts tied to latency, error rates, availability, and service-level objectives rather than alerting on every small metric fluctuation.
For additional guidance, see the official OpenTelemetry instrumentation documentation.
Common Observability Mistakes to Avoid
- Collecting logs without correlation: Unrelated log entries make distributed incidents harder to investigate.
- Ignoring asynchronous boundaries: Kafka consumers and background workers need appropriate trace-context propagation.
- Using excessive metric labels: High cardinality increases resource consumption and storage costs.
- Enabling instrumentation without testing: Unsupported or incorrectly initialized instrumentation may produce incomplete telemetry.
- Assuming more telemetry is always better: Excessive data can increase costs without improving incident diagnosis.
- Ignoring access control: Telemetry can expose sensitive operational or user information.
- Failing to test under load: Export queues, collectors, and backends can become bottlenecks.
Frequently Asked Questions
What are the three pillars of observability?
The three foundational pillars are metrics, logs, and distributed traces. Metrics reveal aggregate system behavior, logs record individual events, and traces connect operations across distributed services.
What is OpenTelemetry used for?
OpenTelemetry is used to instrument applications and collect, process, and export telemetry. It supports traces and metrics, with logs also supported by the broader framework; the maturity and implementation details vary by language and component.
What is the difference between monitoring and observability?
Monitoring typically evaluates predefined signals and conditions to detect known problems. Observability uses telemetry to investigate system behavior, including failures whose causes were not known in advance. The practices complement each other.
How does distributed tracing work in microservices?
Distributed tracing assigns a trace identity to related operations and records spans across services. Context propagation carries the tracing relationship between compatible components, allowing engineers to inspect the request's execution path and latency.
What is W3C Trace Context?
W3C Trace Context is a standard for propagating tracing context across compatible systems. Its traceparent header identifies the trace and calling span, while tracestate can carry optional vendor-specific tracing information.
Does OpenTelemetry automatically collect metrics, logs, and traces?
Not necessarily. Available signals depend on the language implementation, SDK configuration, exporters, and instrumentation libraries. Each signal must be configured and verified for the intended application and deployment.
Conclusion
Microservices observability provides the visibility needed to understand complex distributed applications. Metrics help detect system-wide problems, logs provide diagnostic context, and distributed traces reveal how operations interact across service boundaries.
OpenTelemetry offers a vendor-neutral foundation for instrumenting applications, while W3C Trace Context supports interoperable trace propagation across compatible services.
For production systems, successful observability depends on more than installing an SDK. Teams must configure instrumentation correctly, propagate context across asynchronous boundaries, protect sensitive telemetry, monitor collection infrastructure, and measure performance under realistic workloads.
The goal is actionable observability: detect problems, locate their causes, and resolve incidents using reliable evidence.
Further Reading
- OpenTelemetry Official Documentation
- OpenTelemetry Node.js Getting Started
- OpenTelemetry Collector
- W3C Trace Context Specification
- OpenTelemetry Instrumentation Concepts
Related Articles on Your Blog
Add contextual internal links to relevant articles on your own website:
- Node.js Performance Optimization
- Microservices Architecture and Design Patterns
- PostgreSQL Performance Tuning
- Docker and Production Deployment Guide
Replace each placeholder with the actual published URL. Remove any link for which you do not have a relevant article.
SEO Publishing Checklist
- Primary keyword: microservices observability
- Secondary keywords: OpenTelemetry Node.js, distributed tracing in microservices, metrics logs and traces, W3C Trace Context
- Suggested URL slug:
microservices-observability-opentelemetry - Meta description: Learn microservices observability with OpenTelemetry, Node.js instrumentation, W3C Trace Context, distributed tracing, metrics, logs, and production monitoring.
- Suggested image alt text: OpenTelemetry observability architecture showing metrics, logs, distributed traces, and microservices.
- Schema markup: Use
ArticleorTechArticlestructured data where supported by your publishing platform. Include accurate author, headline, publication date, and image information. - Canonical URL: Set the canonical URL to the preferred published version of the article.
- Internal linking: Link to related articles using descriptive anchor text, and link back to this article from relevant existing posts.
- External references: Retain authoritative documentation links to OpenTelemetry and W3C.
- Indexing: Ensure the page is indexable and included in your XML sitemap.
SEO note: Internal links help search engines discover and understand relationships between your articles. External links to authoritative references improve usefulness and verifiability, but they do not automatically create backlinks to your website. To earn backlinks, publish original technical examples, reproducible benchmarks, diagrams, or research that other websites have a reason to cite.
Editorial Transparency & Verification Standards
Provenance, research methodology & primary citations
Exhaustive deep dive authored by Nexus staff engineers covering low-level protocol mechanics, source code analysis, and edge failure modes.
All architectural diagrams, code snippets, and distributed protocol assertions are technically reviewed prior to release.
Dev
@krish
Core technical contributor to NexusBlog.
Discussion & Technical Notes0
Peer architectural reviews, benchmark insights, and implementation Q&A
Sign in to ask questions, share benchmark findings, or participate in architecture reviews.
Loading discussions...