Skip to content

Upgrade LinkMesh

A LinkMesh deployment has parts that version independently: the server (with its embedded UI), the collector runtime on each managed host (otelcol-contrib supervised by opampsupervisor, or Grafana Alloy), the optional agent on hosts where you run one for onboarding, and the server’s stored state (its database and its internal certificate authority). This page covers how to move all of them from one version to the next without interrupting the fleet.

The good news up front: telemetry never travels through the server, so even while the server is being upgraded your collectors keep running and your data keeps flowing. What an upgrade briefly affects is the control plane — the UI, the API, config push, and fleet status.

When the server comes back, every collector reconnects its control channel on its own within a short backoff — the enrollment token each host holds is a durable reconnect credential, so a restart or upgrade needs no manual re-enrollment. You’ll see the fleet flip from offline back to healthy within a minute or two of the server reporting healthy.

“State” is more than the database. On a package install it spans three places, and an upgrade or a restore that carries only the first one loses data:

What Where (package install) Lose it and…
Datastore /data/linkmesh/state.db (embedded), or your MongoDB everything below is orphaned
Secrets key /data/linkmesh/secrets.key destination secrets in the datastore stay encrypted and cannot be recovered
GitOps config repos /data/linkmesh-config/ (working and production) your pipeline configuration and its history

On Kubernetes all three live under the single /data PersistentVolumeClaim, so carrying that one volume carries all of them.

The server’s own /etc/linkmesh/config.yaml holds the JWT signing secret. A package upgrade preserves it; a rebuild onto a fresh host does not, and everyone has to sign in again. Authentication is Bearer-over-TLS and the server holds no private certificate authority, so there is nothing else to keep in sync.

Always upgrade in this order:

  1. Server — first, on its own. A newer server understands older collectors and agents, but an older server does not understand newer ones. Upgrading the server first is the only safe direction.

  2. Collector runtimes — next, and whenever convenient. Their configuration is delivered by the server through each runtime’s own channel (Alloy over remotecfg, otelcol-contrib over OpAMP), so the runtime binary and its config move independently — there is no coordinated step.

  3. Agents (where you run one) — last, at your own pace. The agent is the optional onboarding edge connector, not every host has one; where present it may lag the server by one minor version and keep working normally. The fleet UI flags any agent that has drifted further behind so you know which hosts to attend to.

  1. Take a state backup. Stop the server first (systemctl stop linkmesh-server) so the snapshot is consistent — copying a live embedded datastore can capture a half-written file. Then copy both /data/linkmesh/ (the datastore and secrets.key) and /data/linkmesh-config/ (the git config repos). On Kubernetes, snapshot the /data PVC, and your MongoDB if it is external.

    This backup is your only way back across a release that changes the data format. That is not a figure of speech — see Rollback below, because nothing in the server will stop you from rolling back into a corrupt state.

  2. Roll the new version. On Kubernetes, point the deployment at the new image tag and let it roll. On a package install, upgrade from the same APT or YUM repository you added at install time — the commands that add the repository live on Install from APT or YUM and are not repeated here, so they cannot drift. Then apt upgrade linkmesh-server / dnf upgrade linkmesh-server (or yum update linkmesh-server). The package restarts the service itself — you do not need a separate systemctl restart, though one is harmless. If you originally installed from a one-shot .deb / .rpm rather than the repository, either add the repository first or install the new package the same way you installed the old one.

  3. Watch the first boot. On startup the server applies the forward data migrations the new version needs. Confirm it reports healthy (curl -fsS http://localhost:8080/health, which echoes the running version) and that the UI loads.

  4. Write down the version you came from. The server does not record it anywhere — see the warning below. Your own note is the only record that exists, and a rollback needs it.

On hosts where you run the optional onboarding agent, it’s upgraded by the operator on that host — there is no remote “upgrade” button, by design. Use whichever you installed with:

  • Package install: apt upgrade linkmesh-agent or dnf upgrade linkmesh-agent / yum update linkmesh-agent from the same repository as Install from APT or YUM, then restart the service.
  • Install script: re-run the same install one-liner you used to enrol the host; it installs the current version in place. Existing enrollment (the durable token) is preserved.

Either way the enrollment token in /etc/linkmesh-agent/config.yaml is left alone, so the agent reconnects on its own after the restart — there is no re-enrollment step.

You don’t need to do every host at once. Within the one-minor skew window the fleet runs mixed versions happily; work through the hosts the collector list flags as out-of-date first.

Because the server delivers collector configuration, the runtime upgrades on its own track:

  • otelcol-contrib (OpAMP): re-run the install-opamp.sh one-liner on the host. It installs the runtime version LinkMesh currently pins and re-applies the supervised setup; the server keeps driving config across the change.
  • Grafana Alloy (remotecfg): upgrade Alloy from its own repository (apt upgrade alloy) and restart it. Its remotecfg block keeps pulling configuration from the server unchanged.

No server-side step is required for either.

Server rollback is a version pin: point back at the previous image tag, or downgrade the package, and restart.

Agents and collector runtimes roll back the same way you rolled them forward: install the previous package version, or re-run an installer pinned to the prior version.

  1. You’ve read the changelog entries between the version you run and the one you’re moving to, and carried out any upgrade steps a breaking change asks for.

  2. You have a fresh backup taken with the server stopped, covering /data/linkmesh/ (datastore and secrets.key) and /data/linkmesh-config/ — or the /data PVC on Kubernetes, plus an external MongoDB if you run one.

  3. You’re upgrading the server first, agents and collectors after.

  4. You’ve written down the version you are on. The server does not record it, so this note is the only thing that tells you what to pin back to — and the only way to know whether a rollback needs the backup restored.