Teams arriving at Proxmox VE from VMware usually arrive for two reasons: the licensing maths changed, and they wanted the hypervisor to be something they actually control. (What each vCenter and Aria capability maps to on the way over is in VMware to Proxmox: what happens to monitoring.) Then they discover that the monitoring story is not a drop-in replacement either. There is no vCenter, no vRealize, and — the detail that catches most people — no native Prometheus endpoint.
What Proxmox does have is enough to build a proper observability picture, provided you know which four sources to combine. This guide covers those sources, the collector configuration that unifies them, and the signals that actually predict a bad night on a Proxmox cluster.
Platform and virtualisation engineers running Proxmox VE in production — particularly in regulated or sovereignty-sensitive environments where the monitoring stack has to stay on your own infrastructure. Prerequisites: a Proxmox cluster you can install packages on, and somewhere to send metrics.
Where Proxmox metrics actually come from
There is no single endpoint. A complete picture needs four sources, and each has a different mechanism:
- The node OS. Proxmox VE is Debian, so CPU, memory, disk, filesystem and network for
the host itself come from the OpenTelemetry Collector’s own
hostmetricsreceiver. No extra exporter required — this is the same collection you would run on any Linux host. - The PVE API — cluster, guests, storage. Guest state, per-VM CPU and memory, storage
pool usage, HA state and node membership live in the Proxmox API. The usual bridge is
the community
prometheus-pve-exporter, scraped by the collector’sprometheusreceiver. It is not an official Proxmox component; treat it as a dependency you own. - The built-in metric server. Proxmox can push its own metrics out natively, but it
speaks InfluxDB or Graphite line protocol, not Prometheus. That is usually read as
a dead end. It isn’t: the collector’s
influxdbreceiver accepts InfluxDB line protocol, so you can point Datacenter → Metric Server straight at a collector and get PVE’s own metrics without the exporter in the path at all. - Logs.
journaldon each node covers the system and the Proxmox daemons. Task logs under/var/log/pve/tasks/are where backup, migration and replication outcomes actually land, and they are afilelogsource.
Two viable designs follow. Either scrape the PVE API with the exporter, or have Proxmox push line protocol into the collector. The push path has fewer moving parts; the scrape path exposes more per-guest detail. Most clusters end up running both.
A collector on every node, plus a small gateway
The shape that works is the standard two-tier one: an agent collector on each Proxmox node doing local collection, forwarding to a small gateway that handles egress and routing.
receivers:
hostmetrics:
collection_interval: 30s
scrapers:
cpu:
memory:
load:
disk:
filesystem:
network:
paging:
# Proxmox's built-in metric server, pushing line protocol to us.
influxdb:
endpoint: 0.0.0.0:8086
# The community PVE exporter, and Ceph's own mgr module.
prometheus:
config:
scrape_configs:
- job_name: pve
scrape_interval: 30s
static_configs:
- targets: ['127.0.0.1:9221']
- job_name: ceph
scrape_interval: 30s
static_configs:
- targets: ['127.0.0.1:9283']
journald:
units: [pvedaemon, pveproxy, pvestatd, corosync, ceph-mon, ceph-osd]
processors:
resourcedetection:
detectors: [system]
batch:
exporters:
otlp/gateway:
endpoint: otel-gateway.internal:4317
Point Datacenter → Metric Server at http://<node>:8086 with InfluxDB selected, enable
the Ceph manager’s Prometheus module (ceph mgr module enable prometheus) if you run
hyperconverged storage, and the node now emits one coherent stream instead of four
disconnected ones.
The signals that actually predict trouble
Proxmox clusters fail in a small number of characteristic ways, and generic host monitoring misses most of them:
- Quorum and Corosync link health. A cluster that loses quorum stops allowing changes and HA stops working. Corosync link status and ring errors are the leading indicator; by the time a node shows as offline in the UI you are already in the incident.
- Storage latency, not storage capacity. A full pool is obvious. The failure that actually bites is rising IO wait on a ZFS pool or a Ceph OSD approaching saturation, which manifests as guests going sluggish long before anything alerts. Latency and utilisation beat a capacity threshold — the same argument as in storage metrics worth tracking.
- ZFS ARC pressure. Proxmox nodes commonly run ZFS with an ARC that competes with guest memory. ARC shrinking under memory pressure is a slow, silent performance regression.
- CPU steal inside guests. Host CPU can look fine while guests starve. Steal time is measured in the guest, which means an agent inside the guest — the host view alone cannot see it.
- Backup job outcomes. A backup that silently stopped running is the failure nobody notices until a restore. Proxmox Backup Server task results are the signal; treat a missing success as an alert, not just a failure as an alert.
- Replication lag on ZFS replication jobs, which drift quietly when a target fills up.
What this does not give you
Being straight about the gaps, because they change the design:
- There is no official OpenTelemetry receiver for Proxmox. The exporter is a community project. It is widely used and it works, but it is a dependency with its own release cadence, and it sits in your monitoring path.
- Guest-internal visibility needs a guest agent. Host-side metrics tell you what the hypervisor sees. Filesystem usage inside a VM, application logs and CPU steal all require collection inside the guest — usually the same collector, deployed as a normal host agent.
- Per-guest cardinality adds up fast. A few hundred guests with per-VM series across CPU, memory, disk and network is a large metric surface. Decide early which per-guest series you actually alert on, and aggregate the rest, or the storage bill becomes the problem you migrated to avoid.
hostmetricson the hypervisor counts the hypervisor. It does not decompose usage by guest; that comes from the PVE API.
Keeping it on your own infrastructure
If Proxmox was chosen partly for control, it is worth not undoing that in the monitoring layer. Two things keep the position intact:
- Terminate the pipeline where you want it. The collector fleet exports to whatever you run — Prometheus/Mimir and Loki on-prem, or a cloud backend for the parts that carry no sensitive data. That is a routing decision, made per record, not a platform commitment.
- Manage the fleet without a third party in the data path. A control plane needs to see configuration and health, not telemetry — see how data is handled for what that separation means in practice.
Where to start
- One node,
hostmetricsonly. Prove the collector, the gateway and the backend before adding Proxmox-specific sources. - Add the built-in metric server pointed at the
influxdbreceiver — it is the lowest-effort source of genuinely Proxmox-shaped data. - Add the PVE exporter if you need per-guest detail the push path does not carry.
- Add
journaldand the task-logfilelogsource, then alert on backup success before you alert on anything else. - Roll the same config to every node from one place, so node six is not subtly different from node one — fleet management.
Running OpenStack alongside? The telemetry of the two platforms is identical at the node and divergent above it — OpenStack vs. Proxmox observability covers running one collector fleet across both.
LinkMesh manages the OpenTelemetry collectors on your Proxmox cluster from a self-hosted control plane — compose the sources once, preview the rendered config, roll it to every node over OpAMP, and see per-edge throughput as it flows. Telemetry goes straight from your nodes to your backend. See what it does, or set one up in a few minutes.
