LinkMesh

Search docs, blog and changelog

ENDE
A chart recorder drum carrying one continuous unbroken trace from edge to edge
LinkMeshObservability Data Collection Management
ObservabilityCompliance

Monitoring → Observability

In a regulated estate, the hard part isn't the dashboards.

linkmesh.io
Philippe BraxmeierPhilippe Braxmeier← Back to blog
6 min read

Every regulated IT department already has monitoring. Nagios or Checkmk watching hosts, a few hundred threshold alerts, a Grafana wall in the NOC, and an on-call rota that has learned which alerts to ignore. It works, in the sense that outages get noticed.

Then an auditor asks a question that monitoring cannot answer: show me that client data never left Switzerland during the incident on 12 March, and show me who could have seen it. Nobody’s threshold alert has an opinion about that. This is the gap observability is supposed to close, and in a regulated Swiss environment — FINMA circular 2018/3 outsourcing, revDSG, a cantonal bank’s own internal control framework — closing it is a different exercise than it is at a startup.

Who this guide is for

Platform leads, heads of infrastructure, and IT risk officers in Swiss banks, insurers, hospitals and public administration who already run monitoring and are being asked for observability — often by an auditor rather than by an engineer. Prerequisites: none beyond knowing what your current monitoring stack does.

The distinction that actually matters

The textbook definition — monitoring tells you whether a known thing broke, observability lets you ask why something you never anticipated broke — is true but not very useful for a budget conversation. In a regulated estate the practical difference is narrower and sharper:

  • Monitoring answers questions you wrote down in advance. Every check is a hypothesis someone had: disk over 90%, service not responding, certificate expiring. If the failure mode was not imagined, there is no check for it.
  • Observability keeps enough evidence to answer questions you have not thought of yet. That is the entire proposition. It is also, precisely, why it costs more and why compliance cares: you are now retaining a much richer record of what your systems did, and some of that record is regulated data.

That second point is the one most migration plans skip. Observability is not just better operations — it is a data retention decision, and in a regulated environment a data retention decision is a legal one.

What an auditor actually asks for

Engineering teams tend to prepare for the wrong review. The questions that come up in a FINMA-relevant audit or an internal ICS review are rarely about dashboard quality:

  • Where is this data processed, and by whom? Not “which SaaS vendor” — which legal entity, in which jurisdiction, under which contract, with which sub-processors.
  • What is in it? Whether telemetry contains personal data under revDSG, or client-identifying data under banking secrecy, is a question about content, not about the tool. See PII vs CID for why those are two different problems and why a scanner only solves one of them.
  • Who can change what is collected? If any engineer can edit a collector config on a host and start shipping a new field to a third party, your control is documentation, not enforcement.
  • Can you prove what was running on 12 March? Not what the repository said should be running — what the node actually ran. That is config drift, and it is the question that turns a clean audit into a finding.

None of those four are answered by adding traces. They are answered by controlling the collection layer.

Why the usual migration order is backwards

The common plan is: pick an observability backend, ship everything into it, then worry about governance. In a regulated estate this fails in a specific and expensive way.

Once telemetry is in a third-party backend, every subsequent compliance question is retroactive. You are now negotiating deletion, field-level access controls and data residency with a vendor whose product was not designed around your cantonal supervisor. The cheapest moment to decide what leaves your network is before it leaves your network — which means the decision belongs in the collection layer, not the backend.

The order that survives an audit is:

  1. Standardise collection. One agent, one wire format. OpenTelemetry is the obvious choice because it is a standard rather than a vendor’s agent — see what OpenTelemetry can and cannot replace for an honest scope.
  2. Put a pipeline in the middle. A layer you control, between the systems producing telemetry and whatever stores it. This is the piece most stacks are missing; it is where redaction, routing and volume control physically happen.
  3. Then choose backends. Possibly several. Once collection is standardised and the pipeline is yours, the backend becomes a replaceable component rather than a ten-year commitment.

Doing it in that order means the compliance controls are structural — a field that is masked at the host cannot leak from a backend, because it never arrived.

What this looks like in practice

Concretely, three controls do most of the work in a regulated estate:

  • Mask at the source. Redaction runs in the collector on the host, before egress, so the sensitive value never crosses the network boundary. That is a materially different claim from “the vendor encrypts at rest” — see masking PII in logs.
  • Route by classification. Client-identifying records go to an on-premises destination; anonymous operational metrics can go to a cloud backend. One pipeline, different destinations, decided by attribute rather than by hoping teams configure their agents correctly — routing by attribute.
  • Make the configuration authoritative and auditable. Central config, versioned, pushed over OpAMP, with a record of who changed what and when. Local edits stop being the authority, which converts drift from something you detect into something that largely cannot happen — governance and enforcement.

The honest trade-offs

This is not free, and a plan that pretends otherwise will not survive contact with a steering committee:

  • You now operate a pipeline. It is infrastructure that can fail, sits between your telemetry and your backends, and needs its own availability story — see collector high availability.
  • Observability data volume grows. Richer evidence is more bytes. The cost control has to be designed in from the start rather than discovered on an invoice.
  • OpenTelemetry does not replace your backend. It replaces collection and transport. Dashboards, alerting rules and long-term storage still have to come from somewhere, and rebuilding them is real work.
  • Traces are the hardest part. Metrics and logs migrate readily. Distributed tracing requires application instrumentation, which means development capacity, which means a roadmap conversation rather than an infrastructure one.

Where to start

If your monitoring works and your auditor is asking harder questions, you do not need to rip anything out. Start where the compliance risk and the cost both live:

  1. Inventory what leaves the network today — which agents ship what, to whom, under which contract. Most estates discover at least one surprise here.
  2. Put a collector in the path for one high-volume, low-sensitivity signal. Prove the pattern without touching regulated data.
  3. Add redaction and routing for one genuinely sensitive source, and verify the masking against a real captured sample rather than a unit test.
  4. Then extend backend by backend, keeping the collection layer constant.
Being asked to prove where your telemetry goes?

LinkMesh is a self-hosted control plane for OpenTelemetry collectors: mask at the source, route by classification, and keep a versioned, audited record of exactly what each node was configured to run. Telemetry never flows through LinkMesh — it stays on your infrastructure. See how data is handled, or set one up in a few minutes.