# The problem in plain terms Observability tools exist to make incidents cheaper and faster to resolve. But when teams instrument everything at high resolution and retain it indefinitely, the cost of that telemetry can exceed the estimated business impact of downtime. The author reports talking to teams paying six- and seven-figure annual bills for observability—more than the outage costs they hoped to avoid.
# How this happens Each sensible decision compounds into a telemetry firehose. Teams add metrics, dashboards, high-resolution sampling, and distributed tracing across service boundaries because each step seems defensible in isolation. Vendors charge by data volume, so the product incentives push toward collecting and retaining more. The result is large invoices for logs nobody reads, traces nobody queries, and dashboards nobody watches.
# A practical, data-driven alternative Treat observability like any other engineering investment: run a cost-benefit analysis and optimize for the signals you actually use.
- Start with usage, not theory. Identify the specific metrics, logs, and traces engineers open during real incidents. Focus on what your team actually looks at, not what you might want in a hypothetical scenario.
- Preserve high resolution for those signals that directly support detection, diagnosis, and resolution. Keep full fidelity where it produces measurable incident-response gains.
- For the remaining telemetry, apply lower-cost strategies: sampling, aggregation, shorter time-to-live (TTL), or lower cardinality. These choices reduce ingestion and storage costs while keeping a searchable record at reduced fidelity.
# OpenTelemetry helps, but it's not a silver bullet OpenTelemetry removes vendor coupling by standardizing instrumentation and making it easier to switch backends. That can limit future vendor-driven cost surprises. However, decoupling does not address the root issue: over-collection. Engineering decisions about what to collect, how often, and how long to keep it remain necessary.
# How to decide retention and resolution
Use explicit policies rather than ad-hoc cuts. Sampling and filtering are valid tools, but they should be applied deliberately after you understand what the team relies on. Reducing fidelity to meet a byte target risks removing evidence needed to resolve future incidents.
# Vendor economics matter Observability vendors monetize data volume. Features and marketing emphasize "full observability," which can encourage collecting everything. Expect that pricing models reward higher ingestion and retention. That makes internal governance and engineering discipline the primary defenses against runaway spend.
# Bottom line Less data, better curated, is a mature observability strategy. Prioritize the telemetry that materially improves incident outcomes and reduce the resolution or retention of the rest. OpenTelemetry can ease vendor portability, but cost control requires deliberate engineering trade-offs around collection, resolution, and retention so observability reduces incident costs rather than exceeding them.