Semantic conventions are the part of OpenTelemetry that everyone agrees with
and almost nobody has fully applied. The specification tells you an HTTP
response status belongs under http.response.status_code, a service name
under service.name, a pod under k8s.pod.name. Your fleet, meanwhile,
contains a service written in 2021 that emits http.status_code, a vendor
component that emits status, and a batch job whose author was not thinking
about conventions at all.
TL;DR — Naming drift does not announce itself. It produces queries that return less than they should, with no error anywhere, and the dashboard looks fine. Fixing it in every service is correct and slow, because it needs a code change, a release and a deploy in every repository, and it cannot reach components whose source you do not own. Fixing it at the collector is one rename rule applied to everything already flowing through the pipeline. The hard part is not writing that rule — it is keeping it true on every collector as the fleet grows, which is a config-distribution problem rather than an OpenTelemetry one. Decide the target names, normalize at the collector, version the rule, roll it to the fleet as one action, and verify against real traffic rather than against the config you intended to apply.
The failure mode is a query that silently under-reports
Here is what naming drift actually costs, and why it takes months to notice.
Three services emit what a human would call “the HTTP status code”. One uses the current convention. Two do not. All three arrive at the backend intact — nothing is dropped, nothing errors, every ingest pipeline reports healthy.
Then someone builds the 5xx panel and filters on
http.response.status_code >= 500. That filter matches one service out of
three. The panel renders, the number looks plausible, and it is wrong by two
thirds. Nobody is paged, because under-reporting is not a failure any part of
the stack is designed to detect.
This is what makes naming drift different from most pipeline problems. A dropped batch shows up in a queue metric. A broken exporter shows up in the collector’s own health. A misnamed attribute shows up as a slightly disappointing dashboard, and it gets attributed to low traffic.
The same mechanism quietly degrades everything downstream that keys on attributes: alert rules that never fire for half the fleet, cost attribution that credits the wrong team, and attribute-based routing that sends a subset of the traffic it was meant to send. Routing rules are worth calling out specifically, because a route that matches on a key two services spell differently fails in the direction that is hardest to see: the data goes to the default destination and keeps flowing.
Why the fix does not belong only in your services
The principled answer is that each service should emit the right key in the first place. That is true, and you should still be driving it. It is also, on its own, a plan that does not converge.
Fixing it in the services means finding every repository that emits the wrong key, changing the instrumentation, and getting that through review, release and deploy — per service, per language, per team, and then again for every service created after you finished. Each of those steps is somebody else’s sprint.
And there is a category it cannot reach at all: third-party components, vendor agents and anything whose source you do not control. Those emit what they emit. No amount of internal convention work changes them.
Normalizing at the collector inverts the arithmetic. The collector already sees every signal from every service — that is what it is for. One rename rule there applies to all of them at once, including the services whose teams have not changed a line, and including the components you could never have patched.
The two approaches are not in competition. Normalize at the collector so your queries are right today; keep fixing instrumentation upstream so that the rename rules can eventually be retired. What you should not do is block the first on the second.
The three moves that cover most drift
In collector terms this is mostly the attributes and transform
processors, and three operations account for nearly all real-world drift:
- Rename — the common case.
http.status_codebecomeshttp.response.status_code. Both keys mean the same thing; one of them is the convention. In the LinkMesh processor library this is an Attributes step, which sets, updates, deletes, hashes or extracts attribute keys through the rule builder rather than hand-written configuration. - Promote — a value buried in the log body or in a nested JSON field becomes a first-class attribute under the conventional key. A JSON Log Parser step that extracts top-level keys from a structured body is usually where this starts; see parsing messy logs for the harder cases.
- Enrich — the attribute is not wrong, it is absent. Kubernetes metadata
(
k8s.pod.name,k8s.namespace.name) and cloud resource attributes are added at the collector from the environment, because the service genuinely does not know them.
A caution on the first one, because it is where the damage gets done. Renaming a key changes what every existing query, alert and route matching that key will find. Before you rename, know who is reading the old key. The safe sequence on a live fleet is to add the new key alongside the old one, migrate the queries and alerts that read it, and only then delete the original — rather than performing a rename and discovering which dashboards depended on the previous spelling.
And keep the target list short. Semantic conventions cover an enormous
surface, and trying to conform on all of it at once produces a processor chain
nobody will maintain. The attributes worth normalizing first are the ones
something actually queries, routes or alerts on: service.name,
deployment.environment.name, the HTTP and status keys, and whatever your own
routing rules match against.
Making it a fleet property rather than a per-collector edit
This is the part that decides whether a convention survives contact with a growing fleet, and it is not an OpenTelemetry problem at all.
Write the rule once. Store it as versioned configuration, so the change has an author, a diff, and a previous version you can return to when a rename turns out to have broken a panel. Then push it to every collector in the group as a single action.
Without that, normalization is a per-host edit. You apply the rule on the collectors you reach, and it quietly does not apply on the ones you miss — which is the original drift problem, moved one layer down and made harder to see, because now the same service reports different keys depending on which collector handled it.
Worth being plain about the boundary here: LinkMesh does not decide the target names for you, and it does not detect convention violations on your behalf. You choose what conforms. What the control plane changes is the cost of applying that choice — a normalization rule becomes one versioned change rolled out to the fleet, rather than an edit repeated by hand on every host and re-derived by whoever onboards the next one. That is a deliberate scope boundary, not a gap waiting to be filled: the conventions are a published standard, and which parts of them matter is a decision about your own telemetry.
Verify against real traffic, not against the config
A normalization rule that is present in the configuration is not the same as a normalization rule that is working. The processor may sit after the step that drops the attribute. The key may arrive with different casing. The service may emit it on the resource rather than on the record.
So check the arriving data, not the intended data. Capture live traffic on a collector after the rollout and read the attribute keys that are actually present — the technique in debugging telemetry pipelines with live capture. The question you are answering is narrow and concrete: after this rule, does every stream carry the conventional key?
Then re-run the query that was under-reporting. If the 5xx panel moves when you normalize, that is your confirmation — and also a measure of how wrong it had been.
What this does not fix
Normalizing names makes attributes consistent. It does not make them correct. If a service reports the wrong status code, renaming the key gives you a reliably wrong number under a conventional name.
It also does not retroactively repair data already in the backend. Everything stored under the old key stays under the old key, so queries spanning the rollout date need to match both spellings until the retention window rolls past it. Plan for that rather than being surprised by it — a dashboard that appears to show a sharp change in behaviour on the day of the rollout is usually showing the rollout.
And conventions themselves move. The HTTP conventions have already been through renames on the way to stable, which is precisely how fleets end up with two spellings of the same idea. Normalization rules are not a one-time cleanup; they are a small amount of standing maintenance that follows the specification.
Where to start
Pick the one attribute that already costs you something. Not the full conventions list — one key that a dashboard, a route or an alert depends on.
- Query your backend for the concept, not the key. Find the spellings actually present in the last week.
- Pick the conventional target from the OpenTelemetry semantic conventions registry.
- Add the new key alongside the old one at the collector, on one collector group.
- Capture live traffic and confirm the key is arriving.
- Migrate the queries, alerts and routes that read the old key. Then delete the old key.
- Roll the rule to the rest of the fleet as one versioned change.
The whole loop is an afternoon for the first attribute and minutes for the ones after it, because step 6 is the same action every time. If it is not — if rolling a processor change to the fleet is itself a project — that, rather than the conventions, is the problem worth solving first.
If your collectors are still individually managed, the background on why this layer exists is in what a telemetry pipeline actually is, and LinkMesh is the self-hosted control plane we build for exactly this: one place to define a processor, version it, and roll it to every collector you run — on your own infrastructure.
