The case for sovereign telemetry is made elsewhere on this blog — who controls your observability data is the why: residency is not jurisdiction, jurisdiction is not operational control, and telemetry is the data class nobody classified. This post is the how. It is the reference architecture a Swiss bank, insurer, hospital or federal office can put in front of an architecture board, with each component tied to the control it provides.
It builds on the generic two-tier shape in OpenTelemetry architecture for enterprises — agents, a gateway tier, a control plane — and adds what a sovereignty requirement changes: where each tier physically runs, what happens to a record before it may leave, and what evidence the whole thing produces.
Architects and IT risk officers who need a defensible design, not a product list. The legal anchors referenced — revDSG, FINMA circulars 2018/3 (outsourcing) and 2023/1 (operational risk and resilience), banking secrecy — are named so your compliance team can map controls to them; this is an engineering document, not legal advice. Prerequisites: familiarity with the agent/gateway model.
The design in one paragraph
Every host runs an agent collector that classifies and masks records before they leave it. Agents send only to a gateway tier running on Swiss infrastructure you control, across two sites. The gateway routes each record by its classification: client-identifying and personal data to on-premises destinations, anonymous operational data wherever is economical, including cloud. A self-hosted control plane configures every collector, keeps a versioned record of what each was told to run, and never sees telemetry. Egress from the estate exists only at the gateway, only to the allow-listed destinations. That is the whole thing; the rest is detail and evidence.
Layer 1 — the agent: classify and mask before egress
The sovereignty decision is made at the first hop, because every later hop only sees what the agent let through.
- Classification as a resource attribute. Each source is tagged at the agent with a
data class —
data.class = cid | pii | ops— derived from what the source is (the core banking system’s logs arecidby definition, node metrics areops) rather than from scanning content. Content scanning finds PII; it structurally cannot find client-identifying data, which is the harder Swiss category — PII vs CID. Classify by source, scan as a second line. - Masking on the host. Known-sensitive fields are masked in the agent, so the value never crosses even the internal network to the gateway. This is the control that makes “the data never left the host” a true statement rather than a hope — masking PII in logs.
- No direct egress. The agent’s only exporter is the gateway. It has no credentials for any backend and no route to the internet. A host compromised at the application layer cannot exfiltrate through its collector.
Layer 2 — the gateway tier: terminate in Switzerland, route by class
- Physically Swiss, operationally yours. The gateway tier runs on infrastructure whose operator, jurisdiction and physical location you can each name. Two sites, active-active, load-balanced — the availability design is in collector high availability, and in a regulated estate it is also the resilience evidence FINMA 2023/1 asks for.
- Routing on
data.class. The gateway is where the stream splits:cidandpiirecords to on-premises storage;opsrecords to the cheapest destination, which may be a cloud backend. One pipeline, classified once, destinations chosen per record — routing by attribute. - The only egress point. Firewall rules allow outbound connections from the gateway addresses only, to the allow-listed destination endpoints only. Adding a destination is therefore a firewall change with a change record, not something an engineer does in a config file.
- Heavy processing lives here. Tail sampling, aggregation, fan-out — the work that needs to see the whole trace or the whole fleet — is done once at the gateway, keeping agents light.
Layer 3 — destinations: split, not consolidated
The temptation is one backend. The sovereign design is deliberately two or more:
- On-premises for regulated classes. Whatever you run — a Loki/Mimir/Tempo stack,
Elastic, Splunk — for
cidandpii, on hardware or a Swiss provider whose jurisdiction and operator you have named. Retention set to the regulatory minimum and no longer. - Wherever is economical for
ops. Node metrics, health signals and platform logs carry no regulated content and can go to a cloud backend at cloud prices. This is where the cost saving that funds the rest lives. - Nothing decides at the destination. The backend receives already-classified, already-masked data. Its access controls are a second line, never the first.
Layer 4 — the control plane: hosted by you, blind to data
This is the layer most designs forget, and the one auditors increasingly ask about directly.
- Self-hosted, in Switzerland. The system that holds every collector’s configuration and the fleet inventory is itself in scope. A SaaS management plane is a third party in your trust boundary regardless of where telemetry flows.
- Configuration only. The control plane distributes config and receives health; it is never in the data path. That separation is what lets you state “no third party can see our telemetry” and mean it — see how LinkMesh handles data for one concrete instance of the model.
- Authoritative and audited. Every config change is versioned with who, what and when, pushed to nodes over OpAMP, and every node reports the hash of what it is actually running. A node that drifts from the intended config — the one with the masking rule — is flagged, not discovered by an auditor — detecting drift.
- Preview before rollout. A masking or drop rule is tested against a real captured sample before it goes fleet-wide. A wrong regex that silently un-masks a field is the failure mode this design exists to prevent.
The evidence the architecture produces
An architecture that satisfies a regulator is one that can show things. This one produces, as a by-product of operating:
| Question | Evidence |
|---|---|
| Where is telemetry processed and by whom? | Gateway and control-plane locations, named operators; the egress allow-list |
| What left the network, and in what form? | Per-source classification; masking rules in versioned config; per-edge throughput by class |
| Who could change what is collected? | Control-plane RBAC and the change audit trail |
| What was node N running on date D? | Config version history plus per-node running-config hash |
| Can we leave a vendor? | Instrumentation and collection are OTLP/OpenTelemetry; destinations are exporter changes |
What this does not cover
- The application layer. Instrumentation, and what an SDK puts on a span, is the application team’s control. The architecture masks what reaches the agent; it cannot un-emit what the code emitted. The Java-agent post covers that side.
- Backup, key management and BCM for the destinations themselves. They are the destinations’ problem and standard practice applies.
- Legal determinations. Which data is
cidunder banking secrecy for your institution is a decision for your compliance function. The architecture gives them a place to record it (the source classification) and a mechanism to enforce it.
Where to start
- Classify sources, not content — a spreadsheet of every telemetry source and its
data.class, agreed with compliance. This is the design’s foundation and takes days, not months. - Stand up the gateway tier on infrastructure you can name, and point one
opssource through it end to end. - Add one
cidsource with masking at the agent and an on-premises destination, and verify the masking against a captured sample. - Close egress so only the gateway can leave, then move remaining sources across.
- Host the control plane before the fleet is large enough that drift becomes invisible.
LinkMesh is a self-hosted control plane for OpenTelemetry collectors, built by OpenSight, an independent Swiss company. It distributes configuration and receives health; telemetry flows from your agents to your gateways to your destinations and never through it. Versioned, audited config with per-node drift detection and preview before rollout. See the data-handling model, or set one up in a few minutes.
