Why Is OpenTelemetry Important?
Modern applications rarely run as a single process on a single server. Applications may span microservices, containers, APIs, databases, Kubernetes environments, serverless functions and multiple cloud providers.
Each technology can generate its own operational data. Without standardization, organizations may end up maintaining different instrumentation libraries, agents, formats and integrations across their environments.
OpenTelemetry addresses this problem by providing a common framework for generating and moving telemetry. Instead of instrumenting an application specifically for one monitoring or observability vendor, teams can use OTel-compatible instrumentation and send the resulting telemetry to one or more supported destinations.
This separation between instrumentation and analysis can make observability architectures more portable. It also gives teams a common telemetry model across different languages, frameworks and infrastructure environments.
What Problem Does OpenTelemetry Solve?
Before standardized observability instrumentation, telemetry collection often depended heavily on proprietary agents and vendor-specific libraries.
This could create several problems:
- Different applications generated telemetry in different formats.
- Development teams had to learn multiple instrumentation systems.
- Switching observability backends could require reinstrumenting applications.
- Logs, metrics and traces were difficult to correlate consistently.
- Application and infrastructure teams used different naming conventions.
- Telemetry pipelines became increasingly complicated as environments grew.
OpenTelemetry separates telemetry generation from the backend used to analyze it. Developers instrument applications using a standardized approach. OTel then generates and transports telemetry that compatible platforms can ingest. The result is a more portable observability architecture.
How Does OpenTelemetry Work?
OpenTelemetry sits between the systems generating telemetry and the platforms that ultimately store and analyze it.
A simplified OTel data flow looks like this:
Application or infrastructure → OTel instrumentation → OTel SDK → OTLP → OTel Collector → observability backend
The exact architecture varies by environment, but the process generally follows five stages.
1. Instrument the Application or System
Instrumentation determines what telemetry an application produces.
Teams can use:
- Automatic instrumentation
- Manual instrumentation
- Framework and library instrumentation
- Native OpenTelemetry support built into applications or services
Instrumentation creates telemetry such as spans, metrics and logs while adding useful attributes about the service and operation.
2. Generate Telemetry Through the OTel SDK
OpenTelemetry SDKs implement the OTel APIs for supported programming languages. The SDK controls how telemetry is created, processed and exported. It can also handle functions such as:
- Sampling
- Resource identification
- Span processing
- Metric aggregation
- Export configuration
3. Transmit Telemetry Using OTLP
The OpenTelemetry Protocol, or OTLP, is the native protocol used to transmit telemetry within the OpenTelemetry ecosystem. OTLP can carry traces, metrics and logs between applications, Collectors and observability backends. Using a standardized transport format simplifies interoperability among systems that support OpenTelemetry.
4. Process Telemetry With the OpenTelemetry Collector
The OpenTelemetry Collector is a vendor-neutral service used to receive, process and export telemetry. Instead of configuring every application to communicate directly with each destination, applications can send telemetry to a Collector. The Collector can then transform, filter, enrich, batch or route the data.
This makes the Collector an important component in a broader telemetry pipeline.
5. Export Telemetry to a Backend
OpenTelemetry itself is not responsible for long-term telemetry storage, visualization or analysis. Processed telemetry is typically exported to an observability platform, analytics system or other backend. That backend provides capabilities such as:
- Searching
- Querying
- Dashboards
- Visualization
- Alerting
- Correlation
- Root-cause investigation
- Long-term retention
What Are the Main Components of OpenTelemetry?
OpenTelemetry is more than a collection agent. It is an ecosystem of standards and components designed to work together.
OpenTelemetry APIs
The APIs define how applications create telemetry. Developers use the APIs to create traces, metrics and other telemetry without coupling application code to a specific backend. This abstraction is important because application instrumentation can remain relatively stable even if the downstream observability architecture changes.
OpenTelemetry SDKs
SDKs implement OpenTelemetry APIs for specific programming languages. They manage how telemetry is processed and exported and provide configuration for functions such as sampling, processors and exporters.
OpenTelemetry Collector
The Collector receives telemetry from applications and infrastructure, processes it and exports it to one or more destinations. A Collector pipeline generally contains three types of components:
Receivers → Processors → Exporters
Receivers
Receivers ingest telemetry. They can accept OTLP and other supported telemetry formats.
Processors
Processors modify telemetry as it passes through the Collector. Common processing operations can include:
- Batching
- Filtering
- Sampling
- Attribute modification
- Resource detection
- Memory management
- Data transformation
Exporters
Exporters send processed telemetry to downstream destinations. A single Collector can route telemetry to one or multiple compatible systems.
OpenTelemetry Protocol (OTLP)
OTLP defines how telemetry is encoded and transported between OpenTelemetry components. It supports telemetry transmission over standardized transports and allows applications, Collectors and backends to communicate using a common data model.
Semantic Conventions
Semantic conventions establish standardized names for common telemetry attributes and operations.
For example, two services should ideally describe similar HTTP requests, database operations or cloud resources using consistent terminology. This consistency matters because observability data becomes considerably more useful when teams can search and correlate it reliably across different services.
Instrumentation Libraries
OpenTelemetry provides an ecosystem of instrumentation for commonly used frameworks, databases, libraries and runtimes. These integrations help teams capture telemetry without manually instrumenting every low-level operation.
What Telemetry Does OpenTelemetry Collect?
OpenTelemetry supports multiple telemetry signals that describe system behavior from different perspectives.
For a broader comparison of the core signals, see Logs vs. Metrics vs. Traces.
Traces
A trace represents the path of a request or transaction as it moves through a distributed system. Each trace consists of spans representing individual operations.
For example:
Web request → API gateway → checkout service → inventory service → database
A trace can reveal how much time was spent in each operation, which dependencies participated and where an error occurred. This makes OpenTelemetry particularly useful for distributed tracing.
Metrics
Metrics are numerical measurements collected over time. Examples include:
- Request rate
- Error rate
- CPU utilization
- Memory consumption
- Response latency
- Queue depth
- Database connection count
Metrics are useful for dashboards, alerting, trend analysis and capacity planning.
Logs
Logs are timestamped records describing individual events. They may contain information about:
- Errors
- Authentication
- Transactions
- Configuration changes
- Application events
- Infrastructure activity
OpenTelemetry can add shared context to logs, making it easier to connect a log event with the trace or service associated with it.
Baggage
Baggage allows contextual information to propagate across service boundaries. This context can help connect activity that spans multiple services. Because baggage can propagate across systems, organizations should carefully control what information is placed in it and avoid unnecessarily including sensitive data.
Profiles
Profiling provides code-level information about resource consumption, such as where CPU time is being spent. OpenTelemetry support for profiling continues to evolve, making profiling an emerging addition to the broader telemetry model.
How Do Logs, Metrics and Traces Work Together in OpenTelemetry?
Each telemetry signal answers a different type of question.
Read What Is Observability? for a broader explanation.
Is OpenTelemetry an Observability Platform?
No. OpenTelemetry is not an observability backend or analytics platform. It does not replace the systems responsible for storing, querying, analyzing and visualizing telemetry.
Instead, OTel standardizes how telemetry is produced and transported to those systems. This distinction is important when evaluating an observability architecture.
A typical architecture includes:
- Instrumentation — OpenTelemetry generates telemetry.
- Collection and processing — the OpenTelemetry Collector or another telemetry pipeline processes it.
- Storage and analytics — an observability platform stores and analyzes it.
- Investigation — engineers query and correlate the data to understand system behavior.
OpenTelemetry vs. Monitoring: What Is the Difference?
Monitoring evaluates defined indicators and conditions, such as whether CPU utilization exceeds a threshold. OpenTelemetry does not replace monitoring. Instead, it creates and transports the telemetry that monitoring and observability platforms can use.
The broader distinction between observability and monitoring is that monitoring is often focused on known conditions, while observability enables broader investigation into system behavior.
What Are the Benefits of OpenTelemetry?
Reduces Instrumentation Lock-In
Application code can use vendor-neutral OTel APIs instead of being written specifically around proprietary instrumentation. Changing the downstream backend may therefore require less application-level reinstrumentation.
Standardizes Telemetry
Shared APIs, semantic conventions and protocols help organizations produce more consistent telemetry across services and teams.
Improves Telemetry Correlation
Consistent context across logs, metrics and traces makes it easier to connect telemetry from the same request or service.
Supports Distributed Environments
OpenTelemetry was designed for modern distributed systems where requests may move across multiple services, infrastructure layers and programming languages.
Simplifies Multi-Backend Architectures
The Collector can route telemetry to multiple destinations, reducing the need for each application to manage numerous backend integrations.
Improves Portability
Organizations can evolve their observability platforms while retaining a standardized instrumentation layer.
Supports Telemetry Cost Management
Collector processors and sampling strategies can help organizations decide which telemetry to retain, transform or discard before sending it downstream. This becomes particularly important as telemetry volumes and high cardinality data grow.
Automatic vs. Manual OpenTelemetry Instrumentation
OpenTelemetry supports multiple instrumentation approaches.
Automatic Instrumentation
Automatic instrumentation captures common operations with little or no application-code modification. It can provide quick visibility into:
- HTTP requests
- Database operations
- Framework activity
- Common libraries
- Service dependencies
Automatic instrumentation is often a practical starting point for OTel adoption.
Manual Instrumentation
Manual instrumentation lets developers add telemetry specific to the application or business. For example, a commerce application might instrument events such as:
- Checkout started
- Payment authorized
- Order completed
- Inventory lookup failed
Manual instrumentation provides additional context that generic framework instrumentation may not capture.
Which Should You Use?
Most organizations benefit from combining both approaches. Use automatic instrumentation to establish broad technical visibility, then add manual instrumentation for important workflows and business-specific operations.
OpenTelemetry Collector Deployment Patterns
The Collector can be deployed in several ways depending on scale, architecture and processing requirements.
Agent Pattern
A Collector runs close to the workload, such as on the same host or node. This pattern is useful when telemetry needs local collection or enrichment.
Gateway Pattern
Applications or local Collectors send telemetry to centralized Collector instances. Gateway deployments are useful for:
- Centralized processing
- Data routing
- Credential management
- Sampling
- Backend control
- Scaling telemetry processing independently from applications
Agent-to-Gateway Pattern
Larger environments may combine both approaches. Local agents collect telemetry near workloads and send it to centralized gateways for more advanced processing and export. The right deployment model depends on traffic volume, network architecture, security requirements and the amount of telemetry processing required.
OpenTelemetry and Telemetry Pipelines
OpenTelemetry and telemetry pipelines overlap, but they are not identical concepts. A telemetry pipeline manages how operational data moves from sources to destinations. The OpenTelemetry Collector can serve as an important component within that pipeline by:
- Receiving telemetry
- Filtering unwanted data
- Enriching attributes
- Sampling traces
- Batching records
- Transforming telemetry
- Routing data to different destinations
At enterprise scale, telemetry pipeline design becomes increasingly important because sending all generated telemetry directly to every backend can become expensive and difficult to govern.
OpenTelemetry and High Cardinality
OpenTelemetry attributes can add detailed dimensions to telemetry, such as:
- Service name
- Region
- Container
- Endpoint
- Customer tier
- Deployment version
These dimensions make telemetry more useful during investigations. However, attributes containing extremely large numbers of unique values can create high cardinality. High cardinality is not inherently bad. Detailed attributes can be essential for troubleshooting.
The challenge is determining which dimensions provide useful investigative value without creating unnecessary telemetry volume, query complexity or cost.
OpenTelemetry in Cloud-Native Environments
OpenTelemetry is particularly useful in dynamic, distributed environments. In cloud native observability, workloads may move between hosts, containers may exist for only short periods and applications may depend on dozens of distributed services.
Traditional host-centric monitoring can struggle to preserve these relationships. OpenTelemetry allows teams to instrument services consistently while maintaining contextual information about resources, dependencies and requests. This can help teams investigate behavior across:
- Microservices
- Containers
- Cloud infrastructure
- APIs
- Databases
- Serverless applications
- Distributed applications
How to Implement OpenTelemetry
OpenTelemetry adoption does not need to happen all at once. A practical implementation approach is:
1. Define the Questions You Need to Answer
Start with operational questions.
For example:
- Why are requests slow?
- Which service is generating errors?
- Which deployment caused a regression?
- Where are requests failing?
- Are critical services meeting reliability targets?
2. Identify Critical Services
Prioritize the applications and workflows where improved observability will provide the greatest value.
3. Start With Automatic Instrumentation
Use available automatic instrumentation to establish baseline telemetry with minimal application changes.
4. Add Manual Instrumentation Where Needed
Instrument important business processes and application-specific operations that automatic instrumentation cannot understand.
5. Standardize Resource Attributes
Define consistent naming for:
- Services
- Environments
- Versions
- Regions
- Teams
- Workloads
Consistent metadata makes telemetry easier to correlate.
6. Deploy the Collector
Choose an agent, gateway or combined deployment pattern based on the environment.
7. Configure OTLP and Export Destinations
Determine where telemetry should be sent and how it should be routed.
8. Establish Sampling and Filtering Policies
Avoid simply collecting everything. Prioritize telemetry based on operational value, reliability requirements, retention needs and cost.
9. Connect Telemetry to Reliability Goals
Telemetry becomes more meaningful when connected with measurable service objectives. Teams can use observability telemetry alongside service level objectives and the broader SLA, SLO and SLI framework.
10. Continuously Review the Telemetry Strategy
Identify:
- Unused telemetry
- Missing telemetry
- Expensive high-cardinality attributes
- Inconsistent naming
- Duplicate collection
- Unnecessary retention
- Gaps in important workflows
Instrumentation should evolve as applications and infrastructure change.
OpenTelemetry Best Practices
Use Semantic Conventions Consistently
Standardized attributes make telemetry easier to query and correlate across services.
Instrument Critical User Journeys
Do not focus only on infrastructure.
Capture enough context to understand important transactions and application workflows.
Correlate Logs, Metrics and Traces
Avoid treating each signal as a separate monitoring silo.
Preserve Useful Context
Include service, environment, deployment and resource information that engineers need during investigations.
Control High Cardinality
Use detailed attributes intentionally rather than adding unbounded identifiers everywhere.
Use Sampling Strategically
Sampling can reduce telemetry volume, but aggressive sampling may also remove information needed during incidents.
Process Telemetry Before Exporting It
Use Collectors or a broader telemetry pipeline to filter, enrich and route data efficiently.
Protect Sensitive Information
Review application attributes, logs and propagated context to prevent credentials, personal information or other sensitive data from being collected unnecessarily.
Monitor the Observability Pipeline
Instrumentation and Collectors are production infrastructure. Monitor their throughput, errors, resource consumption and delivery reliability.
Common OpenTelemetry Mistakes
Treating OpenTelemetry as the Observability Backend
OTel collects and transports data; another system is needed to analyze and visualize it.
Collecting Everything
More telemetry does not automatically produce better observability. High-volume, low-value data can increase cost and noise.
Ignoring Semantic Conventions
Inconsistent attribute names make cross-service investigation much harder.
Using Only Automatic Instrumentation
Automatic instrumentation is valuable, but it cannot understand every business-specific operation.
Adding Unbounded Attributes to Metrics
Identifiers such as user IDs or request IDs can create extreme metric cardinality.
Deploying the Collector Without a Scaling Strategy
Centralized Collectors can become bottlenecks if they are not sized and scaled for telemetry volume.
Failing to Correlate Signals
Logs, metrics and traces provide substantially more value when they share service and request context.
OpenTelemetry and Cortex XCOR
OpenTelemetry provides a standardized way to generate and collect telemetry. An observability platform is still needed to turn that telemetry into operational insight.
Cortex XCOR provides AI-driven observability capabilities for analyzing telemetry across applications and infrastructure, helping teams investigate system behavior, troubleshoot issues and improve reliability.
Organizations using OpenTelemetry can maintain standardized instrumentation while sending telemetry into an observability architecture designed for analysis, correlation and operational decision-making.
Explore Cortex XCOR to learn how Palo Alto Networks approaches AI-driven observability.
OpenTelemetry FAQs