ObservabilityOriginal Technical Deep DiveINTERMEDIATEPart 9 of 10 in Series

Microservices Observability: OpenTelemetry, Metrics, Logs & Distributed Tracing

Learn microservices observability with OpenTelemetry, Node.js instrumentation, W3C Trace Context, distributed tracing, metrics, logs, and production monitoring.

D
Dev
@krish
September 22, 2026 11 min read
0
Part 9 of 10Technical Learning Track

System Design: From Zero to Production

A comprehensive engineering series guiding backend developers from single-node instances to highly resilient distributed architectures.

Microservices Observability: OpenTelemetry, Metrics, Logs & Distributed Tracing

Microservices Observability: The Three Pillars, OpenTelemetry, and Distributed Tracing

Modern microservices architectures distribute application logic across multiple services, databases, APIs, and asynchronous workers. Although this architecture improves modularity and independent scalability, it also makes troubleshooting production incidents more challenging.

When an API becomes slow, a database query times out, or a background worker stops processing messages, engineers need more than basic application logs to identify the underlying problem.

This is where microservices observability becomes essential.

Observability enables engineering teams to understand application behavior using telemetry collected from the system. The three foundational telemetry signals are metrics, logs, and distributed traces. Together, they help developers detect anomalies, investigate failures, analyze latency, and improve system reliability.

OpenTelemetry provides a vendor-neutral framework for generating, collecting, and exporting this telemetry across distributed applications.

In this guide, you will learn how the three pillars of observability work, how W3C Trace Context propagates tracing information across services, and how to instrument a Node.js application with OpenTelemetry.

What Is Observability in Microservices?

Microservices observability is the practice of collecting and analyzing telemetry to understand the internal behavior of distributed applications.

Unlike basic monitoring, which typically checks predefined conditions such as CPU usage or HTTP error rates, observability helps engineers investigate unfamiliar failures by examining the evidence produced by the system.

For example, suppose a request passes through an API gateway, a user service, PostgreSQL, and a Kafka consumer. If the request experiences unexpected latency, observability tools can help identify which operation consumed the most time and whether the issue originated in the application, database, or messaging layer.

A production observability strategy typically combines:

  • Metrics: Numerical measurements that reveal system-wide behavior.
  • Logs: Structured event records that provide diagnostic details.
  • Distributed traces: Records of operations and their relationships across service boundaries.

These signals become more useful when they share consistent service metadata and trace identifiers.

The Three Pillars of Observability

1. Metrics: Detecting Performance Problems

Metrics are numerical measurements collected over time. They provide an aggregated view of application health and infrastructure performance.

Common microservices metrics include:

  • Request throughput and requests per second.
  • HTTP 4xx and 5xx error rates.
  • Request latency at p50, p95, and p99.
  • CPU utilization and memory consumption.
  • Database connection-pool saturation.
  • Kafka consumer lag and queue depth.
  • Garbage collection frequency and duration.

Metrics are particularly useful for dashboards, alerting, capacity planning, and service-level objectives (SLOs).

For example, increasing p99 latency combined with a rising HTTP 5xx rate may indicate a downstream dependency failure or resource bottleneck.

Production tip: Avoid high-cardinality metric labels such as user IDs, trace IDs, and unrestricted URL paths. These can significantly increase storage and processing costs.

2. Logs: Understanding Application Events

Logs capture individual events generated by applications and infrastructure. Structured logging makes these events easier to search, filter, and correlate with other telemetry.

Example structured log:

json
{
  "timestamp": "2026-10-03T10:15:30.123Z",
  "severity": "ERROR",
  "service.name": "user-service",
  "message": "Database query failed",
  "trace_id": "4bf92f3577b34da6a3ce929d0e0e4736",
  "span_id": "00f067aa0ba902b7",
  "error.type": "DatabaseTimeout",
  "duration_ms": 3000
}

With trace and span identifiers, an engineer can navigate from a failed distributed trace to the corresponding application logs.

Useful logging practices include:

  • Prefer structured JSON logs over unstructured messages.
  • Include service names, deployment environments, and timestamps.
  • Record meaningful exception details and operation outcomes.
  • Use appropriate log levels to reduce unnecessary noise.
  • Redact credentials, access tokens, and sensitive personal information.

Logs explain what happened at a particular point in time. Distributed traces complement them by revealing how operations relate across services.

3. Distributed Tracing: Finding the Root Cause of Latency

Distributed tracing follows a request as it moves through multiple services and dependencies.

A trace represents the overall execution path, while a span represents an individual operation, such as an HTTP request, a database query, or a message-processing task.

Consider a request with the following execution times:

OperationDuration
API gateway20 ms
User service45 ms
PostgreSQL query800 ms
Response processing15 ms

In this example, the database operation is the most significant latency contributor.

Distributed tracing helps engineers identify such bottlenecks without relying exclusively on manually inserted log statements.

It is particularly valuable when applications use:

  • Multiple REST or gRPC services.
  • PostgreSQL, MySQL, or other database systems.
  • Kafka, RabbitMQ, or asynchronous job queues.
  • Background workers and scheduled tasks.
  • Retries, timeouts, and circuit breakers.

Learn more in the official OpenTelemetry tracing documentation.

OpenTelemetry: A Vendor-Neutral Observability Framework

OpenTelemetry (OTel) is an open-source observability framework for instrumenting applications and collecting, processing, and exporting telemetry.

It provides APIs, SDKs, instrumentation libraries, semantic conventions, and the OpenTelemetry Protocol (OTLP).

Its main components include:

ComponentPurpose
OpenTelemetry APIDefines interfaces for generating telemetry
OpenTelemetry SDKImplements telemetry creation, processing, and export
Auto-instrumentationCaptures supported operations from frameworks and libraries
OTLPTransports telemetry between compatible components
OpenTelemetry CollectorReceives, processes, and exports telemetry
Observability backendStores and visualizes telemetry and supports queries and alerts

OpenTelemetry reduces dependence on vendor-specific instrumentation by allowing compatible backends to consume standardized telemetry.

However, dashboards, alert rules, query languages, and backend-specific features may still require changes when migrating between providers.

For an overview, see the official OpenTelemetry documentation.

W3C Trace Context Propagation in Distributed Systems

Collecting traces from individual services is not enough. Those traces must also be connected across network requests and asynchronous operations.

The W3C Trace Context specification defines standard HTTP headers for propagating tracing information between compatible systems.

The two primary headers are:

  • traceparent: Carries the trace ID, parent span ID, and trace flags.
  • tracestate: Carries optional vendor-specific tracing information.

Example:

http
traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01

The header contains four fields:

FieldExampleDescription
Version00Trace Context format version
Trace ID4bf92f3577b34da6a3ce929d0e0e4736Identifies the distributed trace
Parent ID00f067aa0ba902b7Identifies the caller's span
Trace flags01Indicates that the sampled flag is set

When an instrumented service receives a request, it extracts the incoming context and creates the appropriate span. When it calls another service, the tracing instrumentation propagates the context.

This allows compatible tracing systems to connect operations into a distributed execution graph.

How Trace Context Works Across Kafka

Asynchronous messaging introduces another propagation boundary. A producer may publish an event and return an HTTP response before a consumer processes the event.

Without explicit or automatic message-context propagation, the consumer's work may appear as an unrelated trace.

A typical workflow looks like this:

Architecture Diagram
Mermaid Flow

The producer propagates trace context through Kafka message headers, and the consumer extracts it during processing.

Depending on the messaging instrumentation and processing model, the relationship may be represented through parent-child spans or span links.

For more details, consult the OpenTelemetry JavaScript documentation and the W3C Trace Context specification.

How to Set Up OpenTelemetry in Node.js

OpenTelemetry provides Node.js instrumentation for supported frameworks and libraries. With automatic instrumentation, many HTTP requests, outbound calls, and supported database operations can generate spans without extensive manual code changes.

Step 1: Install the Required Packages

Install the Node.js SDK, tracing exporter, and automatic instrumentation package:

bash
npm install \
  @opentelemetry/api \
  @opentelemetry/sdk-node \
  @opentelemetry/exporter-trace-otlp-http \
  @opentelemetry/auto-instrumentations-node

These packages provide the API, SDK, OTLP HTTP trace exporter, and supported automatic instrumentation libraries.

Step 2: Configure the OpenTelemetry SDK

Create an initialization file named instrumentation.ts:

typescript
import { NodeSDK } from "@opentelemetry/sdk-node";
import { OTLPTraceExporter } from "@opentelemetry/exporter-trace-otlp-http";
import { getNodeAutoInstrumentations } from "@opentelemetry/auto-instrumentations-node";

const traceExporter = new OTLPTraceExporter({
  url:
    process.env.OTEL_EXPORTER_OTLP_TRACES_ENDPOINT ??
    "http://localhost:4318/v1/traces",
});

const sdk = new NodeSDK({
  traceExporter,
  instrumentations: [getNodeAutoInstrumentations()],
});

sdk.start();

The endpoint shown is an example. Configure it to point to an OpenTelemetry Collector or another compatible OTLP HTTP endpoint in your environment.

Step 3: Initialize Instrumentation Before Application Code

OpenTelemetry must initialize before the application imports libraries that need automatic instrumentation.

For a compatible CommonJS build, an example launch command is:

bash
node --require ./dist/instrumentation.js ./dist/server.js

Adapt the paths and initialization method to your TypeScript build process and Node.js module system. ESM applications may require a different preload configuration.

For additional setup instructions, follow the official OpenTelemetry Node.js getting-started guide.

Step 4: Configure the OpenTelemetry Collector

For production deployments, an OpenTelemetry Collector can receive application telemetry, process it, and export it to compatible observability backends.

Architecture Diagram
Mermaid Flow

The Collector can support batching, filtering, resource enrichment, and export configuration. Its buffering and retry behavior depend on the deployed components and configuration, so monitor resource limits and dropped telemetry.

Official reference: OpenTelemetry Collector documentation.

Important: The Node.js example above configures trace export. Metrics and logs require their respective collection and export configuration; they are not automatically enabled merely by installing the trace exporter.

Measuring OpenTelemetry Performance Overhead

Observability adds some processing, memory, and network overhead. The actual impact depends on request throughput, the number of spans generated, attribute volume, sampling configuration, exporter behavior, and the collector architecture.

Rather than assuming that instrumentation has negligible overhead, measure it under representative workloads.

MetricWhat to measure
CPU overheadAdditional CPU consumption with instrumentation enabled
Memory overheadAdditional application and collector memory usage
Request latencyChanges in p50, p95, and p99 latency
Export latencyTime spent processing and exporting telemetry
Dropped telemetrySpans, metric points, and log records lost under load
Backend ingestionTelemetry throughput and ingestion delays
Incident detection timeTime from a detectable issue to its identification

A meaningful benchmark should document hardware, runtime versions, request distribution, sampling configuration, telemetry volume, collector topology, test duration, and baseline measurements.

Avoid publishing fixed overhead percentages or claims of reduced incident time without a reproducible evaluation. Instrumentation overhead and incident detection improvements are workload-dependent.

Production Best Practices for Microservices Observability

Use Consistent Service Metadata

Configure resource attributes such as service name, service version, and deployment environment. Consistent metadata makes it easier to filter telemetry and compare service behavior across deployments.

Correlate Logs and Traces

Include trace and span identifiers in structured logs when available. This makes it easier to move from a failed trace to the detailed application events associated with the operation.

Apply Sampling Strategically

Sampling can reduce telemetry costs, but sampled-out requests may be unavailable during an investigation. Consider the trade-off between telemetry volume and diagnostic coverage.

Tail sampling can retain traces based on characteristics such as errors or latency, but requires an appropriate collector architecture and sufficient trace data to make the sampling decision.

Protect Sensitive Information

Avoid exporting passwords, access tokens, confidential request bodies, or unnecessary personal information. Apply redaction and access controls throughout the telemetry pipeline.

Monitor the Telemetry Pipeline

The Collector and observability backend can also experience failures. Monitor exporter errors, queue utilization, resource consumption, ingestion latency, and dropped telemetry.

Build Alerts Around Service-Level Objectives

Use actionable alerts tied to latency, error rates, availability, and service-level objectives rather than alerting on every small metric fluctuation.

For additional guidance, see the official OpenTelemetry instrumentation documentation.

Common Observability Mistakes to Avoid

  • Collecting logs without correlation: Unrelated log entries make distributed incidents harder to investigate.
  • Ignoring asynchronous boundaries: Kafka consumers and background workers need appropriate trace-context propagation.
  • Using excessive metric labels: High cardinality increases resource consumption and storage costs.
  • Enabling instrumentation without testing: Unsupported or incorrectly initialized instrumentation may produce incomplete telemetry.
  • Assuming more telemetry is always better: Excessive data can increase costs without improving incident diagnosis.
  • Ignoring access control: Telemetry can expose sensitive operational or user information.
  • Failing to test under load: Export queues, collectors, and backends can become bottlenecks.

Frequently Asked Questions

What are the three pillars of observability?

The three foundational pillars are metrics, logs, and distributed traces. Metrics reveal aggregate system behavior, logs record individual events, and traces connect operations across distributed services.

What is OpenTelemetry used for?

OpenTelemetry is used to instrument applications and collect, process, and export telemetry. It supports traces and metrics, with logs also supported by the broader framework; the maturity and implementation details vary by language and component.

What is the difference between monitoring and observability?

Monitoring typically evaluates predefined signals and conditions to detect known problems. Observability uses telemetry to investigate system behavior, including failures whose causes were not known in advance. The practices complement each other.

How does distributed tracing work in microservices?

Distributed tracing assigns a trace identity to related operations and records spans across services. Context propagation carries the tracing relationship between compatible components, allowing engineers to inspect the request's execution path and latency.

What is W3C Trace Context?

W3C Trace Context is a standard for propagating tracing context across compatible systems. Its traceparent header identifies the trace and calling span, while tracestate can carry optional vendor-specific tracing information.

Does OpenTelemetry automatically collect metrics, logs, and traces?

Not necessarily. Available signals depend on the language implementation, SDK configuration, exporters, and instrumentation libraries. Each signal must be configured and verified for the intended application and deployment.

Conclusion

Microservices observability provides the visibility needed to understand complex distributed applications. Metrics help detect system-wide problems, logs provide diagnostic context, and distributed traces reveal how operations interact across service boundaries.

OpenTelemetry offers a vendor-neutral foundation for instrumenting applications, while W3C Trace Context supports interoperable trace propagation across compatible services.

For production systems, successful observability depends on more than installing an SDK. Teams must configure instrumentation correctly, propagate context across asynchronous boundaries, protect sensitive telemetry, monitor collection infrastructure, and measure performance under realistic workloads.

The goal is actionable observability: detect problems, locate their causes, and resolve incidents using reliable evidence.

Further Reading


Add contextual internal links to relevant articles on your own website:

Replace each placeholder with the actual published URL. Remove any link for which you do not have a relevant article.

SEO Publishing Checklist

  • Primary keyword: microservices observability
  • Secondary keywords: OpenTelemetry Node.js, distributed tracing in microservices, metrics logs and traces, W3C Trace Context
  • Suggested URL slug: microservices-observability-opentelemetry
  • Meta description: Learn microservices observability with OpenTelemetry, Node.js instrumentation, W3C Trace Context, distributed tracing, metrics, logs, and production monitoring.
  • Suggested image alt text: OpenTelemetry observability architecture showing metrics, logs, distributed traces, and microservices.
  • Schema markup: Use Article or TechArticle structured data where supported by your publishing platform. Include accurate author, headline, publication date, and image information.
  • Canonical URL: Set the canonical URL to the preferred published version of the article.
  • Internal linking: Link to related articles using descriptive anchor text, and link back to this article from relevant existing posts.
  • External references: Retain authoritative documentation links to OpenTelemetry and W3C.
  • Indexing: Ensure the page is indexable and included in your XML sitemap.

SEO note: Internal links help search engines discover and understand relationships between your articles. External links to authoritative references improve usefulness and verifiability, but they do not automatically create backlinks to your website. To earn backlinks, publish original technical examples, reproducible benchmarks, diagrams, or research that other websites have a reason to cite.

Editorial Transparency & Verification Standards

Provenance, research methodology & primary citations

Original Technical Deep Dive
Research Methodology

Exhaustive deep dive authored by Nexus staff engineers covering low-level protocol mechanics, source code analysis, and edge failure modes.

Technical Peer Review

All architectural diagrams, code snippets, and distributed protocol assertions are technically reviewed prior to release.

Spotted a technical inaccuracy or outdated code sample?
0
D

Dev

@krish

Core technical contributor to NexusBlog.

Discussion & Technical Notes0

Peer architectural reviews, benchmark insights, and implementation Q&A

Join the Technical Discussion

Sign in to ask questions, share benchmark findings, or participate in architecture reviews.

Loading discussions...

Microservices Observability: OpenTelemetry, Metrics, Logs & Distributed Tracing | NexusBlog