Skip to content

Deploy a collector fleet on Kubernetes

This guide deploys a collector fleet on Kubernetes: one collector pod per node, enrolled with LinkMesh for central config and throughput. The collector is upstream — either Grafana Alloy (via its official Helm chart) or otelcol-contrib + opampsupervisor (as a DaemonSet from the upstream OTel release image). LinkMesh ships no collector distribution of its own; it’s the control plane that configures whichever runtime you pick.

If your hosts are plain Linux (VMs, bare metal, EC2 instances), the Add a collector flow with the per-host installer is simpler — this page is the Kubernetes substrate for the same two runtimes.

A collector runs one of two runtimes. The runtime decides how its config arrives — pick one per fleet (you can run different runtimes in different clusters or namespaces).

Runtime managementMode How config arrives Deploy as
Grafana Alloy + remotecfg alloy-remotecfg Alloy pulls config over Bearer-authenticated HTTPS Upstream grafana/alloy Helm chart
otelcol-contrib + OpAMP opamp Server pushes config over OpAMP (WSS); opampsupervisor applies it DaemonSet on the upstream OTel release image

Both give you central config push, fleet status, and per-component throughput on the topology canvas. Grafana Alloy via its Helm chart is the recommended Kubernetes default — it’s a single upstream image with a first-class chart.

  • A Kubernetes cluster (any flavour — kind, k3s, EKS, GKE, AKS, on-prem).
  • kubectl configured for the cluster; helm for the Alloy path.
  • A running LinkMesh server, reachable from your pods over HTTPS (Alloy remotecfg + own_metrics) or WSS (OpAMP). One instance is enough; to make the server itself resilient, run it highly available on an external MongoDB database.
  • Your server’s public base URL handy, e.g. https://linkmesh.example.com.
  • That same URL configured on the server as externalUrl — see Quickstart step 2. It must be an address your pods resolve (an Ingress hostname or a Service DNS name, never a pod IP). Skip it and the fleet enrols and ships data correctly while every collector shows throughput 0 / CPU 0 / memory 0.

Open Collectors → + Add Collector in the LinkMesh UI, pick your runtime, and choose the Kubernetes snippet — it’s pre-filled with your server URL and a token.

Because a DaemonSet enrols many pods (and reschedules them), use a reusable enrollment token rather than a single-use one: one token enrols every pod, and a rescheduled pod re-attaches without minting anything new. Mint reusable, scoped tokens under Settings → Enrollment Tokens; see Enrollment tokens for TTL, scope, and revocation.

The commands below are the ones the Add Collector wizard shows for the Kubernetes snippet, with https://linkmesh.example.com in place of your server URL and <ENROLLMENT_TOKEN> in place of the token. Copy them from the wizard instead and both are already filled in.

Deploy the upstream grafana/alloy chart as a DaemonSet. Alloy pulls its pipeline config from LinkMesh via remotecfg and pushes its own metrics back so the topology canvas shows per-component throughput.

# Deploy a Grafana Alloy collector fleet on Kubernetes (UPSTREAM Helm chart).
# Alloy pulls its config from LinkMesh via remotecfg; one reusable token
# enrols every pod (rescheduled pods re-attach — no churn).
helm repo add grafana https://grafana.github.io/helm-charts && helm repo update

cat > linkmesh-alloy-values.yaml <<'EOF'
controller:
  type: daemonset
  volumes:
    extra:
    # One directory per node, reused by whichever Alloy pod runs there next.
    # Alloy runs as root in this image, so the directory needs no ownership
    # step. Remove it by hand after uninstalling if the disk matters.
    - name: linkmesh-state
      hostPath: { path: /var/lib/linkmesh/alloy, type: DirectoryOrCreate }
alloy:
  # The chart defaults to "generally-available", which refuses the file-tail
  # and storage components and so cannot load a File Tail source. Same level
  # the Linux installer uses.
  stabilityLevel: experimental
  # Where Alloy keeps state, including a File Tail or Kubernetes Pod Logs
  # source's read position and a destination's durable queue. It is a
  # directory on the NODE (the hostPath below), so it survives the pod being
  # rescheduled, drained or bumped to a new image. The chart's own default,
  # /tmp/alloy, lives in the container and is lost with it.
  storagePath: /var/lib/alloy/data
  mounts:
    # Pod logs are files on the node, under /var/log/pods. The chart does not
    # mount /var/log by default, and without it a Kubernetes Pod Logs source
    # matches nothing inside the container and collects zero records while
    # every component reports healthy. The chart mounts it read-only.
    # containerd/CRI-O write the real files there; on a Docker Engine node,
    # where they are symlinks into /var/lib/docker/containers, also set
    # dockercontainers: true.
    varlog: true
    extra:
    - { name: linkmesh-state, mountPath: /var/lib/alloy/data }
  configMap:
    content: |
      logging { level = "info" }

      remotecfg {
        url            = "https://linkmesh.example.com"
        id             = "k8s-fleet"
        poll_frequency = "60s"
        // remotecfg auth is Bearer-only — basic_auth is rejected (401).
        bearer_token   = "<ENROLLMENT_TOKEN>"
      }

      // own_metrics -> LinkMesh (per-component throughput on the canvas)
      prometheus.exporter.self "default" { }
      prometheus.scrape "linkmesh_self" {
        targets         = prometheus.exporter.self.default.targets
        forward_to      = [otelcol.receiver.prometheus.linkmesh.receiver]
        scrape_interval = "30s"
      }
      otelcol.receiver.prometheus "linkmesh" {
        output { metrics = [otelcol.exporter.otlphttp.linkmesh.input] }
      }
      otelcol.exporter.otlphttp "linkmesh" {
        client {
          endpoint = "https://linkmesh.example.com"
          headers  = { "Authorization" = "Bearer <ENROLLMENT_TOKEN>" }
        }
      }
EOF

helm install linkmesh-alloy grafana/alloy \
  --namespace linkmesh --create-namespace \
  --values linkmesh-alloy-values.yaml

What the values file sets beyond the chart’s defaults:

  • stabilityLevel: experimental — the chart’s default level refuses the file-tail and storage components, so a File Tail source could not load.
  • A state directory on the node. storagePath points Alloy at /var/lib/alloy/data, which is mounted from /var/lib/linkmesh/alloy on the node. A File Tail or Kubernetes Pod Logs source’s read position and a destination’s durable queue live there, so they survive the pod being rescheduled, drained or moved to a new image: a log line written in between is read, not skipped. The chart’s own default lives inside the container and is lost with it.
  • varlog: true — mounts the node’s /var/log read-only, which is where pod logs are. On a Docker Engine node, where /var/log/pods holds symlinks into /var/lib/docker/containers, also set dockercontainers: true under alloy.mounts.

All pods share the same id and token, so the fleet registers as one logical collector. Want each node as a distinct collector? Template a per-pod id (e.g. from the node name) instead of the shared k8s-fleet id.

The line-by-line meaning of this config — and the standalone-host version — is in Onboard Grafana Alloy via remotecfg.

# Wait for an Alloy pod on every node
kubectl -n linkmesh rollout status ds/linkmesh-alloy

# Watch remotecfg fetch its config
kubectl -n linkmesh logs -l app.kubernetes.io/name=alloy --tail=20 | grep -i remotecfg

In your LinkMesh UI, open Collectors — the fleet appears with its runtime (alloy-remotecfg or opamp) and leaves awaiting_connection within ~60s (Alloy’s first poll) or ~30s (the OpAMP handshake). The topology canvas renders throughput once own_metrics start landing.

Host metrics: report the node, not the pod

Section titled “Host metrics: report the node, not the pod”

The Host Metrics source reads /proc, /sys, and the root filesystem to report CPU, memory, disk, filesystem, and network. Inside a container that’s the collector’s own filesystem — a few processes, a thin overlay disk — not the node it runs on. Activate Host Metrics on either fleet above with no further changes, and it reports the pod, not the node. Nothing errors; the numbers are just wrong, and they look plausible.

Fixing it takes two changes, and both are required — the mount without the field, or the field without the mount, each silently do nothing:

  1. Mount the node’s root filesystem, read-only, into the collector container — the container that runs alloy, or on the OpAMP DaemonSet the supervisor container, which runs the collector image.

    The chart’s alloy.mounts toggles (varlog, dockercontainers) don’t cover an arbitrary host path, so add the volume and its mount as extra entries in linkmesh-alloy-values.yaml. Both extra lists already hold the node state directory: add to them, don’t write a second extra: key, which replaces the first and drops the read positions’ mount.

    controller:
      volumes:
        extra:
        - name: linkmesh-state
          hostPath: { path: /var/lib/linkmesh/alloy, type: DirectoryOrCreate }
        - name: hostfs
          hostPath: { path: / }
    alloy:
      mounts:
        extra:
        - { name: linkmesh-state, mountPath: /var/lib/alloy/data }
        - { name: hostfs, mountPath: /hostfs, readOnly: true }

    Then helm upgrade linkmesh-alloy grafana/alloy --namespace linkmesh --values linkmesh-alloy-values.yaml.

  2. Set root_path on the source — Sources → your Host Metrics source → Root Path → /hostfs (or whatever mount path you chose above). Leave it empty for a collector running directly on a host; only a containerised or DaemonSet collector needs it.

Both runtimes accept the field once the mount exists; without a matching mount, the collector reads an empty or absent directory at that path instead — check the field is set and the volume above is present before concluding the source is broken.

The manifests above are deliberately minimal. Before promoting to production:

  • Pin the Alloy chart version (helm install --version). The chart picks the grafana/alloy image, so an unpinned chart moves the collector forward on the next helm repo update and pod cycle. The OpAMP manifest already pins its images to one version; when you change it, change all three image: lines together.
  • Add a NetworkPolicy on the linkmesh namespace allowing egress only to your LinkMesh server’s HTTPS / WSS port.
  • Move the token into sealed-secrets, SOPS, or Vault rather than templating it into values/manifests. A reusable enrollment token is a fleet credential — treat it accordingly, and revoke + re-mint to rotate.
  • Set resource requests/limits sized to your telemetry volume; the Alloy chart exposes alloy.resources, and you can add a resources: block to the OpAMP DaemonSet container.
  • Review the cluster access. The RBAC above is read-only (get/list/watch, never write, never cluster-admin) — the least the Kubernetes receivers need.
  • Decide about tainted nodes. Neither fleet carries tolerations, so no collector runs on a tainted node — typically the control-plane nodes — and that node’s pod logs and metrics are not collected. Add tolerations to the DaemonSet (or controller.tolerations in the Alloy values) if you want them.

For a quick local evaluation:

# Create a kind cluster
kind create cluster --name linkmesh-eval

# Run a LinkMesh server reachable from the kind cluster
# (host.docker.internal works from kind pods on macOS/Windows)

# Then follow steps 1-3 with your server URL set to, e.g.:
#   https://host.docker.internal:8080   (Alloy remotecfg / own_metrics)
#   wss://host.docker.internal:8080/v1/opamp   (OpAMP)

kind nodes share a Docker network, so any service reachable from your host on host.docker.internal:<port> is reachable from the collector pods too.

helm uninstall linkmesh-alloy --namespace linkmesh
kubectl delete namespace linkmesh

Read positions and durable queues are kept on each node, outside the namespace, so deleting it leaves them behind: /var/lib/linkmesh/alloy or /var/lib/linkmesh/otelcol. Remove that directory on every node if the disk matters.

Then revoke the fleet’s enrollment token under Settings → Enrollment Tokens if you’re decommissioning the cluster.

The steps above enrol a collector fleet you then wire by hand. To instead browse your namespaces and workloads in the UI and onboard their logs in one click, install the linkmesh-agent alongside the fleet — it reports the cluster inventory the onboarding view reads. The agent’s packaging/k8s kustomize base installs both the agent and an OpAMP collector fleet in one apply, sharing a single reusable fleet token.

The Enroll Agent wizard (Agents → Enroll Agent → Kubernetes) hands you this manifest ready to apply, with the reusable fleet token already filled in:

The Enroll Agent wizard's Kubernetes tab: a zero-edit DaemonSet manifest carrying the reusable fleet token, pod-log mounts, and read-only RBAC. One kubectl apply enrols one agent per node.

Or apply the same kustomize base from the command line:

kubectl -n linkmesh-system create secret generic linkmesh-agent-bootstrap \
  --from-literal=enrollment-token="$LINKMESH_TOKEN"
kubectl apply -k <path-to>/linkmesh-agent/packaging/k8s

Within ~60s the fleet enrols and your namespaces, workloads, and services appear on the group’s Kubernetes tab. Each workload has a one-click Onboard logs action, and a Kubernetes Pod Logs source starts tailing it with no further cluster edits — the manifest already carries the pod-log mounts and read-only RBAC:

The group's Kubernetes tab shows the live cluster the agent discovered — reporting members, namespaces, workloads, and services — with a one-click Onboard logs action on any workload.

Cluster-wide telemetry is one click too. The Collect cluster metrics card enables node metrics (kubeletstats, on every DaemonSet member), plus cluster-object metrics (k8s_cluster) and Kubernetes events (k8s_events) as cluster-scoped singletons:

One click enables cluster telemetry for the whole fleet — per-node kubeletstats, plus cluster-object metrics and events collected once cluster-wide. Multi-member fleets disable the cluster-singleton sources so only one replica collects them.

See The agent for what it does and does not do.