Skip to content

Self-telemetry

The per-component numbers you see on the topology canvas — records/sec on each receiver, errors/sec on each exporter, queue depth on the backed-up edge — come from the collector itself. Every managed collector pushes its own internal metrics (Alloy’s prometheus.exporter.self, otelcol-contrib’s otelcol_* family) to LinkMesh over standard OTLP. LinkMesh maps those metrics onto the pipeline shape and renders them.

The chain has no proprietary middleware in the data path. The collector exposes its metrics the same way it always has; LinkMesh just becomes one of the destinations.

For every collector enrolled via OpAMP or remotecfg, the topology canvas shows:

  • Per-edge throughput labels — records/sec on each connection between a receiver, processor, and exporter. Updated every ~30s.
  • Per-component error rates — errors/sec, computed from the collector’s refused / send_failed counters.
  • Exporter queue depth — current queue size vs. capacity, useful for spotting back-pressure before it becomes data loss.
  • Host metrics on the collector detail page — CPU, memory, uptime, restart count.

For what each sparkline, queue bar and error flag on the canvas actually means — element by element — see Reading throughput on the canvas.

If a number is missing where you expect one, see Self-telemetry troubleshooting.

flowchart LR
    subgraph host["Managed host"]
        collector["Collector<br/>(Alloy or otelcol-contrib)"]
        scrape["prometheus.exporter.self<br/>+ service.telemetry"]
        internal["Internal metrics<br/>otelcol_receiver_accepted_logs<br/>otelcol_exporter_sent_metric_points<br/>…"]
        collector --> internal
        internal --> scrape
    end
    scrape -- "OTLP/HTTP + Bearer<br/>POST /v1/metrics" --> server["LinkMesh server<br/>OTLP receiver"]
    server -- "translates otelcol_*<br/>into ComponentThroughput" --> ui["Topology canvas<br/>per-edge numbers + dashboards"]

Three pieces compose:

  1. The collector scrapes its own internal metrics. Alloy via prometheus.exporter.self; otelcol-contrib via its built-in service.telemetry.metrics exporter. These metrics have been part of every OTel collector since v0.85 — nothing LinkMesh-specific.

  2. The collector pushes the scrape result over OTLP to a destination LinkMesh configured at enrollment time. The destination URL is <externalUrl>/v1/metrics — which is why the server’s externalUrl must be set: with no public address to hand out, the offer is never made and the whole fleet reports zeros. The Authorization header is a per-collector bearer token (one token per collector, covering both OTLP push and the native remote-config poll).

  3. The LinkMesh server translates the standard otelcol_* metrics into per-component throughput records. The metric names (otelcol_receiver_accepted_*, otelcol_exporter_sent_*, etc.) map onto receivers/processors/exporters by their OTel SDK instance name; rate computation happens at read time so the UI stays responsive.

The bearer is minted once at enrollment and lives in the collector’s local config — nothing to set up on the collector host. Two paths, depending on runtime:

  • otelcol-contrib via OpAMP: LinkMesh’s OpAMP server mints the token immediately after the collector’s first handshake and pushes it via OpAMP’s ConnectionSettings.OwnMetrics offer. The supervisor receives it and installs it on the collector’s own_metrics pipeline. Operator sees a registered then own_metrics_offered event on the collector’s Events tab.
  • Alloy via remotecfg: the bootstrap config.alloy written at enrollment time contains the bearer + endpoints. Subsequent token rotations happen by regenerating that bootstrap.

The bearer is opaque to the operator — there’s no UI surface to copy it around. To rotate, deregister + re-enroll the collector; this revokes the old token and mints a fresh one.

The collector’s otelcol_* self-metrics, mapped per OTel SDK convention:

What Collector metric LinkMesh field
Receiver accepted otelcol_receiver_accepted_* recordsPerSec
Receiver refused otelcol_receiver_refused_* errorsPerSec
Processor incoming otelcol_processor_incoming_items incomingPerSec
Processor outgoing otelcol_processor_outgoing_items recordsPerSec
Exporter sent otelcol_exporter_sent_* recordsPerSec
Exporter send failed otelcol_exporter_send_failed_* errorsPerSec
Exporter queue size otelcol_exporter_queue_size queueSize
Exporter queue capacity otelcol_exporter_queue_capacity queueCapacity

Plus host metrics on the collector detail page:

What Collector metric LinkMesh field
CPU utilisation system.cpu.utilization hostCpuPercent
Memory utilisation system.memory.utilization hostMemoryPercent
Process CPU process.cpu.utilization cpuPercent
Process memory process.memory.usage memoryMb
Uptime process.uptime uptimeSeconds

The OTLP push is scoped to internal collector metrics only. None of the following ever leaves the collector via the LinkMesh OTLP endpoint:

  • Customer telemetry content (logs, metrics, traces flowing through the pipeline). That goes to the destinations you configured — Grafana Cloud, Loki, Mimir, whatever.
  • Collector application logs (Alloy’s stderr, otelcol-contrib’s stdout). Stays on the host’s journal / log file.
  • Collector configuration. LinkMesh already knows it — it generated it. The collector doesn’t re-emit it.
  • Any host-process data outside the collector’s own service.telemetry scope.

The push is exactly the same payload an operator would get by configuring a prometheus.scrape against the collector’s /metrics endpoint and forwarding to their own backend. LinkMesh just happens to be one such backend.

Self-telemetry works the same way regardless of which runtime a collector runs — both push their internal metrics directly to LinkMesh, so there’s no “limited visibility” runtime.

  • otelcol-contrib via OpAMP: LinkMesh’s OpAMP server offers the own_metrics endpoint + bearer as part of bringing the collector online; the supervisor wires it onto the collector’s own_metrics pipeline.
  • Alloy via remotecfg: the bootstrap config.alloy written at enrollment time wires the same self-telemetry exporter on first start.

Either way the operator sees throughput numbers without touching the collector’s local config, and the topology canvas renders identically. The runtime affects how config arrives at the collector (see Native remote config), not how self-telemetry reports back.

LinkMesh keeps these figures for 24 hours — enough to drive the topology canvas and the last-day charts on the collector detail page. Long-term retention belongs in your own backends, the same ones your pipeline destinations write to. To keep a collector’s CPU and memory history for weeks, route a Host Metrics source on the collector to your metrics backend; see When to resize a collector.

  • Native remote config — the other half of the enrollment story: how config gets to the collector. Same per-collector bearer covers both endpoints.
  • Enrol a collector — walkthrough; the self-telemetry path lights up automatically as part of step 4.
  • Self-telemetry troubleshooting — when the numbers don’t show up.
  • Set up alerting — get notified when these metrics cross a threshold, instead of watching them.
  • Credentials & tokens — the per-collector own-metrics token in the context of LinkMesh’s other credentials.
  • Security & encryption — how the own-metrics push fits the one-TLS-front-door trust model.