LinkMesh

Search docs, blog and changelog

ENDE
A row of identical contactors all correctly closed, linked by one long connecting rod bowed under load
LinkMeshObservability Data Collection Management
OpenTelemetryObservability

OpenStack: Monitoring vs Observability

Every service green, and the launch still takes four minutes.

linkmesh.io
Roman HüslerRoman Hüsler← Back to blog
6 min read

An OpenStack operator’s dashboard says everything is fine. nova-api is up, neutron is up, RabbitMQ queues are shallow, Galera has quorum, every compute node is reporting. And a tenant is telling you that launching an instance takes four minutes when it used to take forty seconds.

That gap is the entire distinction between monitoring and observability on OpenStack. Monitoring tells you the services are running. It cannot tell you where four minutes went, because a single instance launch crosses six services, two message queues and a database, and no per-service health check sees the request as one thing.

Who this guide is for

Operators running OpenStack in production — private cloud teams, telcos, research computing, and regulated organisations running their own IaaS. Prerequisites: you know which services your deployment runs and where their logs land.

What OpenStack monitoring covers well

Conventional monitoring on OpenStack is mature and worth keeping. It answers component questions:

  • Control-plane liveness. API endpoints responding, workers registered, scheduler running.
  • Infrastructure health. Compute node CPU, memory and disk; Ceph OSD state; Galera cluster size and replication lag; RabbitMQ queue depth and consumer counts.
  • Capacity. Placement’s view of allocated versus available resources per cell, which is what actually determines whether a scheduling request can succeed.

This layer is well served by exporters — the community openstack-exporter for service and resource state, Ceph’s manager module, RabbitMQ’s own Prometheus plugin, plus hostmetrics on every node. An OpenTelemetry Collector’s prometheus receiver scrapes all of them into one pipeline, which is a real simplification over four separate scrape configs.

What none of it does is explain a slow request.

Why request-level questions are hard here

An instance launch is a distributed transaction in the shape OpenStack has always had:

nova-api → placement (candidate hosts) → nova-scheduler → nova-conductor → nova-compute → neutron (port binding) → cinder (volume attach) → glance (image fetch)

Each hop is a separate process, often on a separate host, communicating over oslo.messaging on RabbitMQ. The four minutes could be image download, a slow Cinder backend, a Neutron agent that is retrying port binding, or scheduler retries against a host that keeps rejecting the claim. Every one of those services can be individually “healthy” while the transaction crawls.

Monitoring gives you six green lights. Observability gives you the timeline.

The correlation ID OpenStack already has

The useful thing — and the piece most OpenStack operators under-use — is that OpenStack already stamps a correlation identifier on requests. Every API call carries an X-Openstack-Request-Id, and oslo.log writes a request ID into log lines as they pass through each service.

That is not a trace, but it behaves like one for the purpose of reconstructing a request. If you parse the request ID out of the log line into a structured attribute at collection time, you can select every log record from every service belonging to a single launch, in order, without touching application code.

Practically that means a filelog or journald source per service with a parse step that lifts the request ID, severity and service name into attributes — the same onboarding problem covered in parsing messy logs. It is the cheapest step from “we have logs” to “we can follow a request”, and it needs no OpenStack changes.

Where real tracing sits, honestly

Two things are worth being straight about:

  • OpenStack services are not OpenTelemetry-instrumented out of the box. There is no flag that turns nova into an OTLP trace producer. OSProfiler exists and does produce request traces for OpenStack, but it is its own subsystem with its own backends, it is typically enabled selectively because of overhead, and it is not an OTLP pipeline you can point at any backend.
  • So the realistic ladder is metrics → correlated logs → selective tracing. Most operators get the majority of the benefit at rung two. Full distributed tracing across the control plane is a project, not a configuration change, and it is worth scoping honestly rather than promising it in a design document.

Where OpenTelemetry does help immediately is as the collection and routing layer: one agent per node collecting host metrics, scraping the exporters, tailing the oslo logs, parsing request IDs, and shipping everything to wherever you keep it — instead of four agents each owned by a different subsystem.

What to instrument first

If you are moving from monitoring toward observability on OpenStack, the order that pays off fastest:

  1. One collector per node, doing hostmetrics plus the Prometheus scrapes you already have. This changes nothing operationally and consolidates the agent estate.
  2. Add the service logs with request-ID parsing. This is the step that turns “green dashboards, angry tenant” into an answerable question.
  3. Add RabbitMQ and Galera detail. Queue depth per queue and replication lag are the two signals that most often explain control-plane slowness.
  4. Measure the launch path end to end as a synthetic: launch an instance on a schedule and record how long each phase took. A synthetic transaction is a poor substitute for tracing and an excellent substitute for nothing.
  5. Then consider OSProfiler for the specific flows you still cannot explain.

The multi-tenant wrinkle

OpenStack is multi-tenant by construction, and the observability layer inherits that. Project (tenant) identity belongs on the telemetry as a resource attribute, and once it is there, the same routing and quota questions apply that apply to any shared pipeline: which tenant’s data goes where, who is allowed to see it, and how one noisy project is stopped from consuming the collection budget for everyone. That is the subject of multi-tenant collector architecture, and on OpenStack it is not optional.

The honest summary

Monitoring OpenStack is a solved problem with good exporters. Observability on OpenStack is a partially solved problem: you can get correlated, queryable, request-scoped evidence without instrumenting the services, and you can get real traces only with more effort than most deployments will spend.

The layer that makes either affordable is the same one: a collection tier you control, consolidating the exporters and the logs, parsing enough structure to make the data answerable, and routing it to backends you chose — which is the telemetry pipeline argument applied to a private cloud. If Proxmox is also in the estate, OpenStack vs. Proxmox observability sets the two side by side.

Running OpenStack and tired of four agents per node?

LinkMesh manages the OpenTelemetry collectors across your control plane and compute nodes from one self-hosted control plane — compose the exporters, log sources and parsing once, preview the rendered config, and roll it everywhere over OpAMP. Telemetry flows straight from your nodes to your backend. See what it does, or set one up in a few minutes.