When the OpenTelemetry Collector arrived, it was a translation shim. Receive in one format, export in another, maybe batch on the way through. Its job was to stop you having to run a Jaeger agent and a Prometheus scraper and a vendor’s log forwarder. Useful, unglamorous, obviously a piece of plumbing.
That is not what it is any more. Somewhere between OTTL, the routing connector, tail sampling and OpAMP, the Collector quietly became the place where decisions about your telemetry are enforced — which cost, which privacy rule, which destination. It is worth being precise about what that shift does and does not mean, because “the Collector is the control plane” is half right in a way that matters operationally.
Architects and platform leads deciding how much responsibility to place on their Collector tier. Prerequisites: familiarity with receivers, processors and exporters; knowing roughly what OpAMP is helps but is not required.
What changed
Four capabilities, added over several years, changed the Collector’s role from transport to policy:
- A transformation language. OTTL turned “drop this, mask that, rename the other” from a fixed set of processors into an expressible policy. Redaction rules that used to live in application code, or nowhere, can now live in the pipeline.
- Content-based routing. The routing connector made destination a per-record decision rather than a per-agent one. One collected stream, split by attribute, is the mechanic behind routing telemetry by attribute — and it is what makes “regulated data stays here, everything else goes there” enforceable.
- Tail sampling. A decision that needs to see the whole trace has to happen somewhere central. Putting it in the Collector made the Collector the arbiter of what evidence survives.
- A management protocol. OpAMP gave the Collector a way to be configured remotely and report what it is running — the seam that lets configuration come from somewhere other than the host’s filesystem. The protocol mechanics are covered in OpAMP explained.
Individually these look like features. Together they relocate three decisions — how much telemetry costs, what leaves your network, and where it lands — out of the backend and out of application code, into a component you run.
The distinction the phrase obscures
Here is where “the Collector is the control plane” needs qualifying. In every other infrastructure domain, the control plane is the thing that decides and distributes policy; the data plane is the thing that enforces it on traffic. A router forwards packets; something else tells it what the routing table is.
By that definition the Collector is a superb data plane and not a control plane at all. It enforces policy on telemetry beautifully. It has no opinion about where its configuration came from, whether that configuration was reviewed, whether the other four hundred collectors got the same one, or whether someone hand-edited it during an incident three weeks ago.
The OpenTelemetry project left that out deliberately — OpAMP specifies the protocol between a management server and an agent, and stops short of specifying the server. That gap is not an oversight; it is a boundary. But it means adopting the Collector as a policy enforcement point without adopting something above it produces a specific and familiar problem: policy that is enforced inconsistently is not policy.
What that means operationally
If the Collector is where cost, privacy and routing are enforced, three things follow that were not true when it was a protocol shim:
- The configuration becomes the valuable artefact. Not the collector binary — the pipeline definition. It encodes your redaction rules, your sampling rates, your destination map. It should be versioned, reviewed and diffable like any other infrastructure code, because it now carries compliance weight.
- Drift stops being an operations annoyance and becomes a control failure. A node running last quarter’s config is not a tidiness problem if last quarter’s config lacked the masking rule. It is an unredacted-data incident that nobody noticed. This is why detecting config drift reads differently once the pipeline carries policy.
- Changes need previewing, not just validating. A config can be syntactically valid and still drop the records you needed, because a filter matched more than intended. The meaningful check is what a real captured record does when it goes through the new rule — not whether the YAML parses.
Where this is going
The prediction implicit in the trend: the interesting question stops being which agent do we run and becomes who configures the fleet, and how is that authority governed.
Agent choice is converging — OTLP is the wire format, the Collector and Alloy are the two serious runtimes, and the difference between them is narrowing. What is not converging, and what has no default answer, is the layer above: who is allowed to change a pipeline, how a change reaches four hundred nodes, what happens when it is wrong, and how you prove after the fact what each node was running on a given day.
That layer is a control plane in the strict sense, and it is the part every team building on OpenTelemetry ends up either buying or building. The Collector will not grow into it, because the project deliberately drew the line there.
The practical reading
If you are running Collectors today:
- Treat the pipeline config as policy, with the review and versioning that implies.
- Make central config authoritative, so local edits are not the authority — this is the single change that converts drift from a detection problem into a prevention one.
- Preview changes against real records before they go fleet-wide, especially anything that drops or masks.
- Keep the fleet’s version and config state visible. You cannot govern what you cannot enumerate — collector fleet management.
The Collector is not the control plane. It is the thing a control plane controls — and recognising that is what turns a large collector deployment from a liability into the most useful layer in the stack.
LinkMesh is the control plane above the fleet: compose the pipeline once, preview every rule against a real captured sample, roll it out over OpAMP or remotecfg, and keep a versioned, audited record of what each node was told to run. Self-hosted; telemetry never flows through it. See what it does, or set one up in a few minutes.
