LinkMesh

Search docs, blog and changelog

ENDE
A machined orifice plate between two flanges at the near end of a pipe run, its bore smaller than the pipe
LinkMeshObservability Data Collection Management
Cost OptimizationTelemetry Pipelines

Cost Starts at the Edge

By the time the backend can help, you have already paid.

linkmesh.io
Philippe BraxmeierPhilippe Braxmeier← Back to blog
6 min read

The observability cost conversation almost always starts in the wrong place: a renewal quote, a Sentinel commitment tier, a Grafana Cloud usage graph. Someone is asked to reduce spend, and the levers they reach for are the ones the platform offers — shorter retention, a cheaper storage tier, a sampling setting, a bigger commitment for a better unit rate.

Every one of those levers acts on data you have already collected, already transmitted and already committed to ingesting. They are real, they help at the margins, and they are all downstream of the decision that actually set the number.

This post is about where that decision is, and why the backend is structurally the wrong place to make it. For the practical playbook — measure, filter, sample, route — see reducing observability costs. This is the argument for why that sequence is in that order.

Who this guide is for

Platform leads and FinOps partners who have been asked to bring observability spend down and are evaluating where to apply effort. Prerequisites: you have a backend bill and a rough idea of daily ingest volume.

The cost chain

A single log line passes through six stages, and money attaches at different points:

  1. It is written. A developer’s log statement, at a level someone chose, in a loop someone did not think about.
  2. It is collected. An agent, configured months ago, decides this file is in scope.
  3. It is transmitted. Network egress — often invisible on the observability invoice and very visible on the cloud bill.
  4. It is ingested. The per-GB charge on almost every commercial platform. This is the line item everyone sees.
  5. It is indexed. Sometimes separately priced, sometimes the thing that makes the difference between a cheap tier and an expensive one.
  6. It is retained. Storage over time, the lever teams reach for first because it is the easiest to change.

The backend’s controls all live at stages 4–6. Stages 1–3 are where the volume is determined, and nothing in the backend can reach back into them.

That is the whole argument. Shortening retention from 90 days to 30 reduces stage 6. It does not reduce what you paid at stage 4 for the same records, and it does nothing at all about stages 2 and 3.

Why the backend’s levers underperform

Each platform lever has a specific structural limitation:

  • Retention reduction is the cheapest thing to do and the least effective, because ingest usually dominates the bill. It also has a hard floor: your audit or regulatory requirement.
  • Cheap tiers — basic, auxiliary, archive — reduce the unit price but still charge for ingest, and they trade away query capability. They also require something to decide which records belong in which tier, and that decision cannot be made by the backend receiving an undifferentiated stream.
  • Backend-side sampling discards data after you paid to move and ingest it. It reduces storage, not ingest. It is the most commonly misunderstood lever on this list.
  • Commitment tiers are the one worth being genuinely wary of. They lower the unit rate in exchange for a volume floor — which means you have converted a variable cost into a fixed one, and removed your own incentive to reduce. A three-year commitment at 800 GB/day is a three-year commitment to producing 800 GB/day.

None of this makes the platforms wrong. They are pricing a service by the resource it consumes, which is reasonable. It just means the platform is not where your leverage is.

The costs that never appear on the invoice

Three more costs are decided at the edge and never show up in the observability line item:

  • Egress. Shipping hundreds of gigabytes a day out of your network to a SaaS backend is a cloud networking charge that lands on a different team’s budget, which is precisely why nobody optimises it.
  • Compliance exposure. Every record that crosses the boundary is a record in scope for whatever regime applies to it. Unredacted client data in a third-party index is a cost with no invoice attached until it becomes a very large one — masking PII in logs.
  • Coupling. Volume shipped through a vendor’s agent is volume you cannot redirect without touching every host. That is the migration tax, paid later, at a moment you did not choose.

Where the decision actually belongs

If the volume is set between the log statement and the network boundary, that is where the control has to sit. Concretely, four decisions made before egress:

  • Drop what will never be queried. Health checks, successful liveness probes, debug chatter from a service in steady state. In most estates this alone is a double-digit percentage, and it is not a subtle judgement call.
  • Sample the high-volume, low-information majority — while keeping errors and slow traces unconditionally, which is the rule that makes sampling safe rather than frightening (sampling governance).
  • Route by value. Security-relevant and regulated records to the expensive, query-capable destination; bulk operational logs to cheap storage or a different backend entirely (routing by attribute).
  • Aggregate at the edge where per-instance detail is not worth its cardinality.

The important property is not that these are clever. It is that they happen before stages 3, 4, 5 and 6 — so the saving compounds across every one of them at once, including the egress and compliance costs that never appear on the observability bill.

Measure before you cut

One caveat, because the failure mode here is real: cutting volume without measuring it first is how teams delete the one log source an incident needed.

Start by knowing which sources produce the bytes. Most estates find a very skewed distribution — a handful of sources generating the majority of volume, several of which are queried approximately never. That is a measurement exercise, not a guess, and it is what makes the subsequent filtering defensible rather than nervous. See seeing real throughput.

Then preview every drop rule against real captured records before it goes fleet-wide. A filter that matches more than intended is indistinguishable from a working one until the day you need the records it silently removed.

The summary

The backend decides the rate. The edge decides the volume. Negotiating the rate is a procurement exercise with a floor; changing the volume is an engineering exercise with far more headroom — and it is the only one that also reduces egress, exposure and coupling.

Do the edge work first. Then negotiate, from a much better position.

Want the volume decision to happen before the invoice?

LinkMesh manages the OpenTelemetry collectors in front of your backends from one self-hosted control plane — measure per-route throughput live, filter and sample before egress, route records to the destination their value justifies, and preview every rule against a real captured sample first. Priced per collector, not per gigabyte. See what it does, or set one up in a few minutes.