LinkMesh

Search docs, blog and changelog

ENDE
Four different pipe runs, each with its own tap fitting, all landing on one common collecting header
LinkMeshObservability Data Collection Management
OpenTelemetryObservability

Proxmox Monitoring

Nodes, guests, Ceph and backups — through one collector.

linkmesh.io
Roman HüslerRoman Hüsler← Back to blog
7 min read

Teams arriving at Proxmox VE from VMware usually arrive for two reasons: the licensing maths changed, and they wanted the hypervisor to be something they actually control. (What each vCenter and Aria capability maps to on the way over is in VMware to Proxmox: what happens to monitoring.) Then they discover that the monitoring story is not a drop-in replacement either. There is no vCenter, no vRealize, and — the detail that catches most people — no native Prometheus endpoint.

What Proxmox does have is enough to build a proper observability picture, provided you know which four sources to combine. This guide covers those sources, the collector configuration that unifies them, and the signals that actually predict a bad night on a Proxmox cluster.

Who this guide is for

Platform and virtualisation engineers running Proxmox VE in production — particularly in regulated or sovereignty-sensitive environments where the monitoring stack has to stay on your own infrastructure. Prerequisites: a Proxmox cluster you can install packages on, and somewhere to send metrics.

Where Proxmox metrics actually come from

There is no single endpoint. A complete picture needs four sources, and each has a different mechanism:

  • The node OS. Proxmox VE is Debian, so CPU, memory, disk, filesystem and network for the host itself come from the OpenTelemetry Collector’s own hostmetrics receiver. No extra exporter required — this is the same collection you would run on any Linux host.
  • The PVE API — cluster, guests, storage. Guest state, per-VM CPU and memory, storage pool usage, HA state and node membership live in the Proxmox API. The usual bridge is the community prometheus-pve-exporter, scraped by the collector’s prometheus receiver. It is not an official Proxmox component; treat it as a dependency you own.
  • The built-in metric server. Proxmox can push its own metrics out natively, but it speaks InfluxDB or Graphite line protocol, not Prometheus. That is usually read as a dead end. It isn’t: the collector’s influxdb receiver accepts InfluxDB line protocol, so you can point Datacenter → Metric Server straight at a collector and get PVE’s own metrics without the exporter in the path at all.
  • Logs. journald on each node covers the system and the Proxmox daemons. Task logs under /var/log/pve/tasks/ are where backup, migration and replication outcomes actually land, and they are a filelog source.

Two viable designs follow. Either scrape the PVE API with the exporter, or have Proxmox push line protocol into the collector. The push path has fewer moving parts; the scrape path exposes more per-guest detail. Most clusters end up running both.

A collector on every node, plus a small gateway

The shape that works is the standard two-tier one: an agent collector on each Proxmox node doing local collection, forwarding to a small gateway that handles egress and routing.

receivers:
  hostmetrics:
    collection_interval: 30s
    scrapers:
      cpu:
      memory:
      load:
      disk:
      filesystem:
      network:
      paging:
  # Proxmox's built-in metric server, pushing line protocol to us.
  influxdb:
    endpoint: 0.0.0.0:8086
  # The community PVE exporter, and Ceph's own mgr module.
  prometheus:
    config:
      scrape_configs:
        - job_name: pve
          scrape_interval: 30s
          static_configs:
            - targets: ['127.0.0.1:9221']
        - job_name: ceph
          scrape_interval: 30s
          static_configs:
            - targets: ['127.0.0.1:9283']
  journald:
    units: [pvedaemon, pveproxy, pvestatd, corosync, ceph-mon, ceph-osd]

processors:
  resourcedetection:
    detectors: [system]
  batch:

exporters:
  otlp/gateway:
    endpoint: otel-gateway.internal:4317

Point Datacenter → Metric Server at http://<node>:8086 with InfluxDB selected, enable the Ceph manager’s Prometheus module (ceph mgr module enable prometheus) if you run hyperconverged storage, and the node now emits one coherent stream instead of four disconnected ones.

The signals that actually predict trouble

Proxmox clusters fail in a small number of characteristic ways, and generic host monitoring misses most of them:

  • Quorum and Corosync link health. A cluster that loses quorum stops allowing changes and HA stops working. Corosync link status and ring errors are the leading indicator; by the time a node shows as offline in the UI you are already in the incident.
  • Storage latency, not storage capacity. A full pool is obvious. The failure that actually bites is rising IO wait on a ZFS pool or a Ceph OSD approaching saturation, which manifests as guests going sluggish long before anything alerts. Latency and utilisation beat a capacity threshold — the same argument as in storage metrics worth tracking.
  • ZFS ARC pressure. Proxmox nodes commonly run ZFS with an ARC that competes with guest memory. ARC shrinking under memory pressure is a slow, silent performance regression.
  • CPU steal inside guests. Host CPU can look fine while guests starve. Steal time is measured in the guest, which means an agent inside the guest — the host view alone cannot see it.
  • Backup job outcomes. A backup that silently stopped running is the failure nobody notices until a restore. Proxmox Backup Server task results are the signal; treat a missing success as an alert, not just a failure as an alert.
  • Replication lag on ZFS replication jobs, which drift quietly when a target fills up.

What this does not give you

Being straight about the gaps, because they change the design:

  • There is no official OpenTelemetry receiver for Proxmox. The exporter is a community project. It is widely used and it works, but it is a dependency with its own release cadence, and it sits in your monitoring path.
  • Guest-internal visibility needs a guest agent. Host-side metrics tell you what the hypervisor sees. Filesystem usage inside a VM, application logs and CPU steal all require collection inside the guest — usually the same collector, deployed as a normal host agent.
  • Per-guest cardinality adds up fast. A few hundred guests with per-VM series across CPU, memory, disk and network is a large metric surface. Decide early which per-guest series you actually alert on, and aggregate the rest, or the storage bill becomes the problem you migrated to avoid.
  • hostmetrics on the hypervisor counts the hypervisor. It does not decompose usage by guest; that comes from the PVE API.

Keeping it on your own infrastructure

If Proxmox was chosen partly for control, it is worth not undoing that in the monitoring layer. Two things keep the position intact:

  • Terminate the pipeline where you want it. The collector fleet exports to whatever you run — Prometheus/Mimir and Loki on-prem, or a cloud backend for the parts that carry no sensitive data. That is a routing decision, made per record, not a platform commitment.
  • Manage the fleet without a third party in the data path. A control plane needs to see configuration and health, not telemetry — see how data is handled for what that separation means in practice.

Where to start

  1. One node, hostmetrics only. Prove the collector, the gateway and the backend before adding Proxmox-specific sources.
  2. Add the built-in metric server pointed at the influxdb receiver — it is the lowest-effort source of genuinely Proxmox-shaped data.
  3. Add the PVE exporter if you need per-guest detail the push path does not carry.
  4. Add journald and the task-log filelog source, then alert on backup success before you alert on anything else.
  5. Roll the same config to every node from one place, so node six is not subtly different from node one — fleet management.

Running OpenStack alongside? The telemetry of the two platforms is identical at the node and divergent above it — OpenStack vs. Proxmox observability covers running one collector fleet across both.

Running Proxmox and want one monitoring config across every node?

LinkMesh manages the OpenTelemetry collectors on your Proxmox cluster from a self-hosted control plane — compose the sources once, preview the rendered config, roll it to every node over OpAMP, and see per-edge throughput as it flows. Telemetry goes straight from your nodes to your backend. See what it does, or set one up in a few minutes.