LinkMesh

Search docs, blog and changelog

ENDE
One standard flange face on a bench, surrounded in order by the several older patterns it replaces
LinkMeshObservability Data Collection Management
OpenTelemetryFleet Management

Standard or Burden?

Both — and the burden is a different shape than the one you dropped.

linkmesh.io
Philippe BraxmeierPhilippe Braxmeier← Back to blog
5 min read

Two things are said about OpenTelemetry, usually by different people in the same meeting. The first is that it is the end of proprietary agents: one standard, instrument once, send anywhere. The second is that it replaced one vendor’s agent with a YAML file, a contrib distribution with three hundred components, and a fleet nobody signed up to operate.

Both are true. The interesting question is not which one wins the argument but what exactly the burden consists of — because roughly half of it is inherent to the standard and the other half is a fleet-operations problem that predates OpenTelemetry and has known answers.

Who this guide is for

Platform leads and architects who have adopted OpenTelemetry or are about to, and want an honest accounting before committing headcount. Prerequisites: you know what a receiver, processor and exporter are.

What you genuinely gain

Worth stating plainly, because the rest of this post is about costs:

  • Instrumentation stops being a vendor asset. The SDK in your application and the wire format on the network are open. Changing backend no longer means re-instrumenting a thousand services, which is the single most expensive thing about leaving an observability vendor.
  • One agent instead of several. Most estates that adopt OTel are consolidating three or four vendor agents into one collector. That is a genuine reduction in surface area, not just a swap.
  • The processing layer becomes yours. Filtering, redaction, sampling and routing happen in a component you configure, before egress — which is what makes cost and compliance controllable at all.

None of that is marketing. It is also not free.

What the burden actually is

Five distinct costs, and they are not equally hard:

  • You now run a fleet. The collector is a process on every host that has to be installed, configured, upgraded and monitored. Fifty of them is an inventory problem; five thousand is an operations discipline. This is the largest and most under-estimated item, and it is covered on its own in keeping a collector fleet up to date.
  • Configuration sprawl. YAML per host drifts. Someone hand-edits a config during an incident and never reverts it; a rollout half-lands. The result is a fleet where you cannot state what is running without going and looking — config drift is the specific failure mode.
  • Component churn. The contrib distribution moves quickly. Components change stability level, semantic conventions evolve, and an upgrade can rename an attribute your dashboards depend on. This is the price of a young, fast-moving standard and it is real work at every upgrade.
  • Cardinality and volume are now your problem. A vendor agent shipped an opinionated, pre-tuned set of metrics. A collector will happily emit whatever you configure, including a per-container label set that quietly multiplies your series count.
  • No batteries included. OpenTelemetry gives you collection and transport. Dashboards, alert rules, and the curated content a commercial agent shipped with are not in the box. That gap is the subject of what OpenTelemetry can and cannot replace.

Which half is inherent, and which is not

This is the useful cut. Of the five:

Inherent to adopting a standard: component churn and the missing curated content. Both come from OpenTelemetry being a specification and an implementation rather than a product. They can be planned for — pin versions, gate upgrades on health, budget for rebuilding dashboards — but they cannot be designed away.

Not inherent — these are fleet-operations problems: configuration sprawl, drift, upgrade risk, and volume control. Every one of them existed with vendor agents too; the vendor’s console just hid them. Puppet-managing three hundred Splunk forwarders had the same drift problem, and the same 3 a.m. rollback.

That distinction matters because the second group has a structural answer rather than a willpower answer. If the running configuration is authoritative, pushed from one place, versioned, previewed before it ships and health-gated on rollout, then drift is not something you detect — it is something that mostly cannot happen. That is the argument in governance and enforcement, and it is what turns a fleet from a liability into an asset.

The failure mode to avoid

There is a specific way OpenTelemetry adoption goes wrong, and it is worth naming: the team adopts the collector as a drop-in agent replacement, manages it with the same config-management tooling they used for the old agent, and never builds the control layer.

Six months later they have all of the burden and none of the leverage. The collector is just a different agent, with a less familiar config format and no vendor support line. The reason to adopt OTel — that the processing layer is now yours to use — is precisely the part that never gets used, because nobody can safely change a pipeline across three hundred hosts.

The tell is simple: if changing one processing rule fleet-wide is a change-management ticket rather than an afternoon, the leverage was never realised.

An honest recommendation

  • Adopt the standard for instrumentation immediately. The lock-in argument is real and the cost of delay compounds with every service you instrument against a vendor SDK.
  • Do not adopt the collector without deciding who operates it. A fleet with no owner drifts into the failure mode above. This is an organisational decision, not a technical one.
  • Treat the fleet as a fleet from day one. Central configuration, versioned, rolled out deliberately. Retrofitting this after two years of hand-managed YAML is a migration in itself.
  • Budget for the content gap. Dashboards and alerts have to come from somewhere. Pretending otherwise is how a technically successful migration gets judged a failure by the people who lost their dashboards.

So: open standard and new operational burden. The burden is smaller than the lock-in it replaces, but only if the fleet is operated rather than merely deployed.

Adopted OpenTelemetry and inherited a fleet?

LinkMesh is a self-hosted control plane for OpenTelemetry collectors: one place to compose pipelines, preview changes against real records, roll them out health-gated over OpAMP, and see which nodes are running what. Priced per managed collector, not per gigabyte. See what it does, or set one up in a few minutes.