Deploy a collector fleet on Kubernetes
This guide deploys a collector fleet on Kubernetes: one collector pod per
node, enrolled with LinkMesh for central config and throughput. The collector
is upstream — either Grafana Alloy (via its official Helm chart) or
otelcol-contrib + opampsupervisor (as a DaemonSet from the upstream OTel
release image). LinkMesh ships no collector distribution of its own; it’s the
control plane that configures whichever runtime you pick.
If your hosts are plain Linux (VMs, bare metal, EC2 instances), the Add a collector flow with the per-host installer is simpler — this page is the Kubernetes substrate for the same two runtimes.
Pick a runtime
Section titled “Pick a runtime”A collector runs one of two runtimes. The runtime decides how its config arrives — pick one per fleet (you can run different runtimes in different clusters or namespaces).
| Runtime | managementMode |
How config arrives | Deploy as |
|---|---|---|---|
| Grafana Alloy + remotecfg | alloy-remotecfg |
Alloy pulls config over Bearer-authenticated HTTPS | Upstream grafana/alloy Helm chart |
| otelcol-contrib + OpAMP | opamp |
Server pushes config over OpAMP (WSS); opampsupervisor applies it |
DaemonSet on the upstream OTel release image |
Both give you central config push, fleet status, and per-component throughput on the topology canvas. Grafana Alloy via its Helm chart is the recommended Kubernetes default — it’s a single upstream image with a first-class chart.
Prerequisites
Section titled “Prerequisites”- A Kubernetes cluster (any flavour — kind, k3s, EKS, GKE, AKS, on-prem).
kubectlconfigured for the cluster;helmfor the Alloy path.- A running LinkMesh server, reachable from your pods over HTTPS (Alloy remotecfg + own_metrics) or WSS (OpAMP). One instance is enough; to make the server itself resilient, run it highly available on an external MongoDB database.
- Your server’s public base URL handy, e.g.
https://linkmesh.example.com. - That same URL configured on the server as
externalUrl— see Quickstart step 2. It must be an address your pods resolve (an Ingress hostname or a Service DNS name, never a pod IP). Skip it and the fleet enrols and ships data correctly while every collector shows throughput 0 / CPU 0 / memory 0.
1. Mint a reusable enrollment token
Section titled “1. Mint a reusable enrollment token”Open Collectors → + Add Collector in the LinkMesh UI, pick your runtime, and choose the Kubernetes snippet — it’s pre-filled with your server URL and a token.
Because a DaemonSet enrols many pods (and reschedules them), use a reusable enrollment token rather than a single-use one: one token enrols every pod, and a rescheduled pod re-attaches without minting anything new. Mint reusable, scoped tokens under Settings → Enrollment Tokens; see Enrollment tokens for TTL, scope, and revocation.
2. Deploy the fleet
Section titled “2. Deploy the fleet”The commands below are the ones the Add Collector wizard shows for the
Kubernetes snippet, with https://linkmesh.example.com in place of your server
URL and <ENROLLMENT_TOKEN> in place of the token. Copy them from the wizard
instead and both are already filled in.
Deploy the upstream grafana/alloy chart as a DaemonSet. Alloy pulls its
pipeline config from LinkMesh via remotecfg and pushes its own metrics back
so the topology canvas shows per-component throughput.
# Deploy a Grafana Alloy collector fleet on Kubernetes (UPSTREAM Helm chart).
# Alloy pulls its config from LinkMesh via remotecfg; one reusable token
# enrols every pod (rescheduled pods re-attach — no churn).
helm repo add grafana https://grafana.github.io/helm-charts && helm repo update
cat > linkmesh-alloy-values.yaml <<'EOF'
controller:
type: daemonset
volumes:
extra:
# One directory per node, reused by whichever Alloy pod runs there next.
# Alloy runs as root in this image, so the directory needs no ownership
# step. Remove it by hand after uninstalling if the disk matters.
- name: linkmesh-state
hostPath: { path: /var/lib/linkmesh/alloy, type: DirectoryOrCreate }
alloy:
# The chart defaults to "generally-available", which refuses the file-tail
# and storage components and so cannot load a File Tail source. Same level
# the Linux installer uses.
stabilityLevel: experimental
# Where Alloy keeps state, including a File Tail or Kubernetes Pod Logs
# source's read position and a destination's durable queue. It is a
# directory on the NODE (the hostPath below), so it survives the pod being
# rescheduled, drained or bumped to a new image. The chart's own default,
# /tmp/alloy, lives in the container and is lost with it.
storagePath: /var/lib/alloy/data
mounts:
# Pod logs are files on the node, under /var/log/pods. The chart does not
# mount /var/log by default, and without it a Kubernetes Pod Logs source
# matches nothing inside the container and collects zero records while
# every component reports healthy. The chart mounts it read-only.
# containerd/CRI-O write the real files there; on a Docker Engine node,
# where they are symlinks into /var/lib/docker/containers, also set
# dockercontainers: true.
varlog: true
extra:
- { name: linkmesh-state, mountPath: /var/lib/alloy/data }
configMap:
content: |
logging { level = "info" }
remotecfg {
url = "https://linkmesh.example.com"
id = "k8s-fleet"
poll_frequency = "60s"
// remotecfg auth is Bearer-only — basic_auth is rejected (401).
bearer_token = "<ENROLLMENT_TOKEN>"
}
// own_metrics -> LinkMesh (per-component throughput on the canvas)
prometheus.exporter.self "default" { }
prometheus.scrape "linkmesh_self" {
targets = prometheus.exporter.self.default.targets
forward_to = [otelcol.receiver.prometheus.linkmesh.receiver]
scrape_interval = "30s"
}
otelcol.receiver.prometheus "linkmesh" {
output { metrics = [otelcol.exporter.otlphttp.linkmesh.input] }
}
otelcol.exporter.otlphttp "linkmesh" {
client {
endpoint = "https://linkmesh.example.com"
headers = { "Authorization" = "Bearer <ENROLLMENT_TOKEN>" }
}
}
EOF
helm install linkmesh-alloy grafana/alloy \
--namespace linkmesh --create-namespace \
--values linkmesh-alloy-values.yamlWhat the values file sets beyond the chart’s defaults:
stabilityLevel: experimental— the chart’s default level refuses the file-tail and storage components, so a File Tail source could not load.- A state directory on the node.
storagePathpoints Alloy at/var/lib/alloy/data, which is mounted from/var/lib/linkmesh/alloyon the node. A File Tail or Kubernetes Pod Logs source’s read position and a destination’s durable queue live there, so they survive the pod being rescheduled, drained or moved to a new image: a log line written in between is read, not skipped. The chart’s own default lives inside the container and is lost with it. varlog: true— mounts the node’s/var/logread-only, which is where pod logs are. On a Docker Engine node, where/var/log/podsholds symlinks into/var/lib/docker/containers, also setdockercontainers: trueunderalloy.mounts.
All pods share the same id and token, so the fleet registers as one logical
collector. Want each node as a distinct collector? Template a per-pod id
(e.g. from the node name) instead of the shared k8s-fleet id.
The line-by-line meaning of this config — and the standalone-host version — is in Onboard Grafana Alloy via remotecfg.
Run the upstream otelcol-contrib image with the upstream opampsupervisor
as its main process. No upstream image contains both binaries, so an init
container copies the supervisor out of its own release image into the pod; the
supervisor is statically linked and runs unchanged in the collector image. It
connects to LinkMesh over OpAMP, starts the image’s own /otelcol-contrib, and
applies the config the server pushes. Both images are pinned to the same
version.
The manifest is self-contained: it grants the ServiceAccount the read-only
RBAC the Kubernetes receivers need (kubelet metrics, pod and namespace lookup
for attribute enrichment), mounts /var/log/pods for pod logs, and passes the
node name to the collector — so activating a Kubernetes Pod Logs or node
metrics source in the UI works with no further edits to the cluster.
# Deploy an otelcol-contrib + OpAMP collector fleet on Kubernetes.
# An initContainer stages the upstream supervisor binary into the upstream
# collector image (neither image contains both); the supervisor then applies
# the config LinkMesh pushes over OpAMP to the collector alongside it.
kubectl create namespace linkmesh --dry-run=client -o yaml | kubectl apply -f -
kubectl apply -n linkmesh -f - <<'EOF'
apiVersion: v1
kind: ServiceAccount
metadata: { name: linkmesh-otelcol, namespace: linkmesh }
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRole
metadata: { name: linkmesh-otelcol }
rules:
# The k8sattributes processor LinkMesh attaches to every Kubernetes source
# reads pods and namespaces, and resolves a pod's workload name through its
# ReplicaSet owner reference. Without these the collector runs but logs
# "pods is forbidden" continuously and enriches nothing.
- apiGroups: [""]
resources: ["pods", "namespaces", "nodes"]
verbs: ["get", "list", "watch"]
- apiGroups: ["apps"]
resources: ["replicasets"]
verbs: ["get", "list", "watch"]
# Kubernetes Node Metrics (kubeletstats) reads its own node's kubelet.
- apiGroups: [""]
resources: ["nodes/stats", "nodes/proxy"]
verbs: ["get"]
---
apiVersion: rbac.authorization.k8s.io/v1
kind: ClusterRoleBinding
metadata: { name: linkmesh-otelcol }
roleRef: { apiGroup: rbac.authorization.k8s.io, kind: ClusterRole, name: linkmesh-otelcol }
subjects:
- { kind: ServiceAccount, name: linkmesh-otelcol, namespace: linkmesh }
---
apiVersion: v1
kind: ConfigMap
metadata:
name: linkmesh-supervisor
data:
supervisor.yaml: |
server:
endpoint: "wss://linkmesh.example.com/v1/opamp"
headers:
Authorization: "Bearer <ENROLLMENT_TOKEN>"
capabilities:
accepts_remote_config: true
reports_effective_config: true
reports_health: true
reports_remote_config: true
reports_own_metrics: true
reports_available_components: true
agent:
# Native path inside the collector image this pod runs (see the
# initContainer below — the supervisor is the staged binary, the
# collector is the image's own).
executable: /otelcol-contrib
storage:
directory: /var/lib/otelcol-supervisor
---
apiVersion: apps/v1
kind: DaemonSet
metadata:
name: linkmesh-otelcol
spec:
selector:
matchLabels: { app: linkmesh-otelcol }
# A new pod must not start on a node until the old one has gone: both would
# open the same state files under /var/lib/linkmesh/otelcol, and the
# second would wait on the first one's lock. This is the DaemonSet default,
# stated so nobody turns on maxSurge without reading this.
updateStrategy:
type: RollingUpdate
rollingUpdate: { maxUnavailable: 1, maxSurge: 0 }
template:
metadata:
labels: { app: linkmesh-otelcol }
spec:
serviceAccountName: linkmesh-otelcol
initContainers:
# Copies the supervisor out of its own release image into a volume the
# collector image can execute it from. Required because neither upstream
# image is a complete runtime on its own: this image has no collector in
# it anywhere, and the collector image has no supervisor. The binary is
# statically linked, so it runs unchanged in the (distroless) collector
# image. Keep this tag and the collector tag below equal.
- name: stage-supervisor
image: ghcr.io/open-telemetry/opentelemetry-collector-releases/opentelemetry-collector-opampsupervisor:0.153.0
command: ["cp", "/usr/local/bin/opampsupervisor", "/opt/linkmesh/opampsupervisor"]
volumeMounts:
- { name: supervisor-bin, mountPath: /opt/linkmesh }
# The node directory below is created by the kubelet as root, and the
# collector runs as uid 10001, so without this it could not write its
# read positions there and would refuse to start. Root for this one
# command only, holding no capability but CHOWN.
- name: own-node-state
image: ghcr.io/open-telemetry/opentelemetry-collector-releases/opentelemetry-collector-opampsupervisor:0.153.0
command: ["sh", "-c", "mkdir -p /state/filelog-storage /state/sending-queue && chown 10001:10001 /state/filelog-storage /state/sending-queue"]
securityContext:
runAsUser: 0
runAsNonRoot: false
capabilities: { drop: ["ALL"], add: ["CHOWN"] }
volumeMounts:
- { name: node-state, mountPath: /state }
containers:
- name: supervisor
# Upstream OTel collector release image, one pinned version (matches
# what this server is tested against). PID 1 is the staged supervisor;
# it starts and manages this image's own /otelcol-contrib.
image: otel/opentelemetry-collector-contrib:0.153.0
command: ["/opt/linkmesh/opampsupervisor"]
args: ["--config", "/etc/otelcol-supervisor/supervisor.yaml"]
env:
# Which node this pod is on. Kubernetes Node Metrics targets this
# node's kubelet with it, and the k8sattributes processor scopes its
# pod cache to this node with it. It is NOT optional: once any
# Kubernetes source is activated, a collector without it refuses to
# start ("'node_from_env_var' is configured but envvar K8S_NODE_NAME
# is not set") and takes every other source on it down too.
- name: K8S_NODE_NAME
valueFrom: { fieldRef: { fieldPath: spec.nodeName } }
volumeMounts:
- { name: supervisor-bin, mountPath: /opt/linkmesh, readOnly: true }
- { name: cfg, mountPath: /etc/otelcol-supervisor }
- { name: data, mountPath: /var/lib/otelcol-supervisor }
# A File Tail or Kubernetes Pod Logs source's read position, and a
# destination's durable queue, live in these two directories. They are
# on the NODE, so they survive the pod being rescheduled, drained or
# bumped to a new image — a pod log line written meanwhile is read,
# not skipped. Only these two: the rest of the data volume holds the
# supervisor's own identity, which stays with the pod.
- { name: node-state, mountPath: /var/lib/otelcol-supervisor/filelog-storage, subPath: filelog-storage }
- { name: node-state, mountPath: /var/lib/otelcol-supervisor/sending-queue, subPath: sending-queue }
# The collector image is DISTROLESS and its whole filesystem is three
# files — /otelcol-contrib, /etc/otelcol-contrib/config.yaml and the CA
# bundle. There is no /tmp in it, so without this mount the collector
# has nowhere to write a temp file at all. A File destination left on
# its catalogue default (/tmp/otel-output.json) then fails to start the
# pipeline, and because the supervisor is PID 1 and absorbs the crash
# the POD stays Running while the agent subprocess restarts forever.
- { name: tmp, mountPath: /tmp }
# Pod logs are files on the node. Without this mount a Kubernetes Pod
# Logs source matches nothing inside the container and collects zero
# records while looking perfectly healthy. Read-only on purpose.
# containerd/CRI-O write the real files here; on a node where
# /var/log/pods holds symlinks, mount their targets as well.
- { name: podlogs, mountPath: /var/log/pods, readOnly: true }
volumes:
# Where the initContainer drops the supervisor binary. Pod-scoped on
# purpose — it is re-staged from the image on every pod start, so it can
# never drift from the image tag above.
- { name: supervisor-bin, emptyDir: {} }
- { name: cfg, configMap: { name: linkmesh-supervisor } }
- { name: podlogs, hostPath: { path: /var/log/pods } }
# Scratch space, pod-scoped and not persisted. A File destination writing
# here is for first-setup inspection, which is what the catalogue entry
# says it is for — point it at a hostPath or a PVC if the output has to
# outlive the pod.
- { name: tmp, emptyDir: {} }
# The supervisor's own state: its effective config and its persistent
# instance id. Pod-scoped, unchanged — only the two directories mounted
# from node-state above outlive the pod.
- { name: data, emptyDir: {} }
# Read positions and durable queues (mounted above). One directory per
# node, reused by whichever collector pod runs there next. Remove it by
# hand after uninstalling if the disk matters.
- { name: node-state, hostPath: { path: /var/lib/linkmesh/otelcol, type: DirectoryOrCreate } }
EOFRead positions survive the pod. A File Tail or Kubernetes Pod Logs source’s
read position, and a destination’s durable queue, are kept in
/var/lib/linkmesh/otelcol on the node, so a pod that is rescheduled,
drained or moved to a new image resumes where the previous one stopped instead
of skipping what was written in between. The second init container gives that
directory to the collector’s user (uid 10001); it runs as root for that one
command and holds no capability but CHOWN. The rest of the supervisor’s state
stays with the pod.
The updateStrategy is part of that: a new pod does not start on a node until
the old one is gone, because both would open the same state files. Leave
maxSurge at 0.
The standalone-host version of this runtime, plus how the two binaries fit together, is in Onboard otelcol-contrib via OpAMP.
3. Verify
Section titled “3. Verify”# Wait for an Alloy pod on every node
kubectl -n linkmesh rollout status ds/linkmesh-alloy
# Watch remotecfg fetch its config
kubectl -n linkmesh logs -l app.kubernetes.io/name=alloy --tail=20 | grep -i remotecfg# Wait for a supervisor pod on every node
kubectl -n linkmesh rollout status ds/linkmesh-otelcol
# Watch the OpAMP handshake + config apply
kubectl -n linkmesh logs -l app=linkmesh-otelcol --tail=20In your LinkMesh UI, open Collectors — the fleet appears with its runtime
(alloy-remotecfg or opamp) and leaves awaiting_connection within ~60s
(Alloy’s first poll) or ~30s (the OpAMP handshake). The topology canvas renders
throughput once own_metrics start landing.
Host metrics: report the node, not the pod
Section titled “Host metrics: report the node, not the pod”The Host Metrics source reads /proc, /sys, and the root filesystem to
report CPU, memory, disk, filesystem, and network. Inside a container that’s
the collector’s own filesystem — a few processes, a thin overlay disk —
not the node it runs on. Activate Host Metrics on either fleet above with no
further changes, and it reports the pod, not the node. Nothing errors; the
numbers are just wrong, and they look plausible.
Fixing it takes two changes, and both are required — the mount without the field, or the field without the mount, each silently do nothing:
-
Mount the node’s root filesystem, read-only, into the collector container — the container that runs
alloy, or on the OpAMP DaemonSet thesupervisorcontainer, which runs the collector image.The chart’s
alloy.mountstoggles (varlog,dockercontainers) don’t cover an arbitrary host path, so add the volume and its mount as extra entries inlinkmesh-alloy-values.yaml. Bothextralists already hold the node state directory: add to them, don’t write a secondextra:key, which replaces the first and drops the read positions’ mount.controller: volumes: extra: - name: linkmesh-state hostPath: { path: /var/lib/linkmesh/alloy, type: DirectoryOrCreate } - name: hostfs hostPath: { path: / } alloy: mounts: extra: - { name: linkmesh-state, mountPath: /var/lib/alloy/data } - { name: hostfs, mountPath: /hostfs, readOnly: true }Then
helm upgrade linkmesh-alloy grafana/alloy --namespace linkmesh --values linkmesh-alloy-values.yaml.Add a
hostfsvolume and mount it into the supervisor container (theotelcol-contribit starts runs in the same container), then apply the manifest again:# containers[0].volumeMounts (supervisor), alongside the existing entries: - { name: hostfs, mountPath: /hostfs, readOnly: true }# volumes, alongside the existing entries: - { name: hostfs, hostPath: { path: / } } -
Set
root_pathon the source — Sources → your Host Metrics source → Root Path →/hostfs(or whatever mount path you chose above). Leave it empty for a collector running directly on a host; only a containerised or DaemonSet collector needs it.
Both runtimes accept the field once the mount exists; without a matching mount, the collector reads an empty or absent directory at that path instead — check the field is set and the volume above is present before concluding the source is broken.
Production hardening
Section titled “Production hardening”The manifests above are deliberately minimal. Before promoting to production:
- Pin the Alloy chart version (
helm install --version). The chart picks thegrafana/alloyimage, so an unpinned chart moves the collector forward on the nexthelm repo updateand pod cycle. The OpAMP manifest already pins its images to one version; when you change it, change all threeimage:lines together. - Add a NetworkPolicy on the
linkmeshnamespace allowing egress only to your LinkMesh server’s HTTPS / WSS port. - Move the token into sealed-secrets, SOPS, or Vault rather than templating it into values/manifests. A reusable enrollment token is a fleet credential — treat it accordingly, and revoke + re-mint to rotate.
- Set resource requests/limits sized to your telemetry volume; the Alloy
chart exposes
alloy.resources, and you can add aresources:block to the OpAMP DaemonSet container. - Review the cluster access. The RBAC above is read-only (
get/list/watch, never write, never cluster-admin) — the least the Kubernetes receivers need. - Decide about tainted nodes. Neither fleet carries
tolerations, so no collector runs on a tainted node — typically the control-plane nodes — and that node’s pod logs and metrics are not collected. Addtolerationsto the DaemonSet (orcontroller.tolerationsin the Alloy values) if you want them.
Trying it on kind
Section titled “Trying it on kind”For a quick local evaluation:
# Create a kind cluster
kind create cluster --name linkmesh-eval
# Run a LinkMesh server reachable from the kind cluster
# (host.docker.internal works from kind pods on macOS/Windows)
# Then follow steps 1-3 with your server URL set to, e.g.:
# https://host.docker.internal:8080 (Alloy remotecfg / own_metrics)
# wss://host.docker.internal:8080/v1/opamp (OpAMP)
kind nodes share a Docker network, so any service reachable from your host on
host.docker.internal:<port> is reachable from the collector pods too.
Uninstall
Section titled “Uninstall”helm uninstall linkmesh-alloy --namespace linkmesh
kubectl delete namespace linkmeshkubectl delete namespace linkmeshRead positions and durable queues are kept on each node, outside the
namespace, so deleting it leaves them behind: /var/lib/linkmesh/alloy or
/var/lib/linkmesh/otelcol. Remove that directory on every node if the disk
matters.
Then revoke the fleet’s enrollment token under Settings → Enrollment Tokens if you’re decommissioning the cluster.
Onboard from the cluster inventory
Section titled “Onboard from the cluster inventory”The steps above enrol a collector fleet you then wire by hand. To instead browse
your namespaces and workloads in the UI and onboard their logs in one click,
install the linkmesh-agent alongside the fleet — it reports the cluster
inventory the onboarding view reads. The agent’s packaging/k8s kustomize base
installs both the agent and an OpAMP collector fleet in one apply, sharing a
single reusable fleet token.
The Enroll Agent wizard (Agents → Enroll Agent → Kubernetes) hands you this manifest ready to apply, with the reusable fleet token already filled in:
Or apply the same kustomize base from the command line:
kubectl -n linkmesh-system create secret generic linkmesh-agent-bootstrap \
--from-literal=enrollment-token="$LINKMESH_TOKEN"
kubectl apply -k <path-to>/linkmesh-agent/packaging/k8s
Within ~60s the fleet enrols and your namespaces, workloads, and services appear on the group’s Kubernetes tab. Each workload has a one-click Onboard logs action, and a Kubernetes Pod Logs source starts tailing it with no further cluster edits — the manifest already carries the pod-log mounts and read-only RBAC:
Cluster-wide telemetry is one click too. The Collect cluster metrics card
enables node metrics (kubeletstats, on every DaemonSet member), plus
cluster-object metrics (k8s_cluster) and Kubernetes events (k8s_events) as
cluster-scoped singletons:
See The agent for what it does and does not do.
Next steps
Section titled “Next steps”- Build your first pipeline — wire your new K8s-enrolled collectors through a processing pipeline to a destination.
- Concepts → Collector — what a collector is + the two collector runtimes.
- Concepts → Collector group — organise a fleet into a group on the canvas.
- Troubleshooting → Enrollment — when a pod doesn’t show up.