A central observability programme needs a repeatable way to connect workloads, apply standards and prove that operational data reaches the right destination. In a hybrid estate, each new application can otherwise become another negotiation about agents, firewall rules, attributes and ownership.
This guide is for platform engineers and solution architects evaluating collector management across on-premises infrastructure and Azure, including Kubernetes environments. It focuses on the operating model and pilot evidence needed before a wider rollout. For collector topology fundamentals, use the separate enterprise architecture guide.
Start with a service and an accountable owner
Choose a business service with an application owner who can explain what failure looks like. Agree which transactions matter, which indicators expose degradation and who responds. Healthy hosts alone cannot prove that customers can complete a payment or submit an order.
Translate those priorities into a small monitoring contract: required signals, resource attributes, service objectives, retention, destination and operational owner. Keep the first pilot small enough to inspect individual events and trace a change through the entire pipeline.
| Responsibility | Accountable team | Pilot deliverable |
|---|---|---|
| Service impact and acceptable disruption | Service owner | Prioritised user journey and reliability objectives |
| Instrumentation and signal quality | Application team | Representative traces, metrics and logs |
| Collector configuration and operation | Platform team | Reusable blueprint and recovery procedure |
| Data classification and access | Security and privacy teams | Approved data paths and access requirements |
| Response and escalation | Operations / ITSM | Tested alert-to-incident handoff |
These responsibilities can overlap. Record who approves a change and who can implement it; a diagram cannot establish that authority.
Separate four architectural responsibilities
Applications and infrastructure generate telemetry. Collectors receive, process and export it. Observability backends store and analyse it. ITSM systems connect operational signals with incidents, ownership and changes.
LinkMesh provides a management layer for collectors and pipelines. It does not replace backend storage, dashboard provisioning, a CMDB or incident workflows. Put those integrations into the architecture explicitly.
Reference design, not a certification of a particular AKS or OpenShift deployment. Open the full-size diagram.
Use local collectors where host access or early filtering is required. Add gateways where central egress, aggregation or processing justifies another operational dependency. Follow the OpenTelemetry agent and gateway deployment guidance when selecting the topology.
For every environment, verify the collector distribution and version, available receivers, service identity, certificates and outbound routes. On AKS or OpenShift, also validate namespace permissions, host mounts and cluster security policies. A generic Kubernetes manifest is not evidence that privileged host collection is allowed.
Define an onboarding blueprint
A blueprint should let an application team answer a short set of questions and receive a reviewable configuration. Avoid assuming that every server needs every receiver.
| Blueprint field | Example decision |
|---|---|
| Workload identity | service.name, service.namespace, deployment.environment.name |
| Organisational ownership | A documented custom attribute such as example.company.team |
| Sources | Application OTLP, host metrics, selected log paths or event channels |
| Processing | Required enrichment, explicit drop rules and sensitive-data handling |
| Destinations | Approved backend endpoints and environment-specific credentials |
| Operations | Collector owner, version, resource limits and recovery instructions |
The ownership attribute is an organisational convention, not a standard OpenTelemetry resource attribute. Agree naming and validation with the teams that query the data.
Use synthetic secrets and identifiers to test filtering. Check the actual exported log bodies and attributes, including nested fields. A processor being present in the configuration does not prove that every sensitive value is removed.
Connect infrastructure automation with configuration management
Use your IaC process to provision compute, networking, service identities and collector installation. Then document how configuration reaches each supported collector. LinkMesh uses the supported runtime’s remote-configuration mechanism; review native remote configuration and the compatibility matrix before choosing versions.
There are important change-management boundaries:
- Saving applies configuration. Do not treat the UI as a draft with a later approval gate. Rehearse changes in a separate test group before applying them to production.
- The external Git integration is a history mirror. LinkMesh writes history to Git; it does not read repository changes back as deployments. See the Git mirror setup guide.
- Configuration rollback does not restore source/destination settings or secret values. Capture their recovery procedure separately.
- Review current edition requirements for rollback, automation and identity integration. The version-history documentation explains the supported operations.

Use collector inventory to check the pilot’s connected runtimes and ownership before changing configuration.
Make security and legacy coverage testable
Self-hosting gives you a deployment choice. Regulatory acceptance still depends on the complete system: data classification, identity, encryption, retention, operating procedures and external dependencies. Assess BYOK/KMS requirements separately rather than assuming them from the hosting model.
Review management traffic as well as workload telemetry. Configuration, inventory, health and collector self-telemetry need an approved destination. Diagnostic sample capture is an additional data-handling workflow. Start with the maintained security and connectivity documentation.
For legacy replacement, build a coverage map before removing an agent. Compare the inputs to existing dashboards and alerts, not only raw event counts. Keep proprietary monitoring capabilities where they remain necessary. Bound dual shipping to a defined population and observation period, with separate routes and a volume budget; see dual shipping without duplicates.
A proof-of-value checklist with acceptance evidence
Agree numerical thresholds with the service owners before starting. The table below proposes measurements, not promised product results.
| Test | Evidence to collect | Acceptance decision |
|---|---|---|
| Onboard one on-premises and one Kubernetes workload | Elapsed time, manual steps, permissions and connected versions | Within the team’s onboarding target; every dependency documented |
| Validate signal coverage | Counts, timestamps, resource attributes and dashboard/alert inputs | Required coverage and agreed tolerances met |
| Test sensitive-data rules | Synthetic test values checked at each destination | No prohibited test values exported through the tested paths |
| Apply and reverse a configuration change | Diff, effective configuration and recovery log | Previous behaviour restored, including settings outside version history |
| Interrupt control-plane connectivity | Collection continuity and reconnection evidence | Behaviour matches the agreed failure model |
| Interrupt a backend | Queue growth, retries, loss, resource use and recovery time | Loss and recovery remain within agreed limits |
| Exercise access boundaries | Allowed and denied operations for representative roles | Permissions match the approved responsibility model |
| Hand an alert to operations | Service identity, incident routing and owner acknowledgement | Responders can identify the service and act |
Also measure collector CPU and memory, ingest volume and operator effort under a representative load. These observations support capacity and cost decisions; they are not substitutes for longer production validation.
Turn the pilot into an evaluation decision
Finish with a short evidence pack: approved architecture, compatibility results, onboarding blueprint, access model, change/recovery procedure, unresolved gaps and named owners. Decide which gaps block adoption and which can be addressed during a limited rollout.
Use the self-hosted OpenTelemetry evaluation page to map LinkMesh to your requirements. Bring your collector inventory, target backends and pilot criteria to a technical evaluation discussion, or connect your first collector.