Skip to content

When to resize a collector

Collector sizing gives a collector a starting size from an estimate. This page is the other half: once the collector runs, its own figures replace the estimate. It tells you where those figures are, what level means “about to be too small”, and what to do about it — in order.

You don’t set anything up for this. LinkMesh configures every managed collector to report on itself and push those figures to the server; they arrive as soon as the collector is enrolled. See Self-telemetry for how.

LinkMesh keeps a collector’s own figures for 24 hours. The collector detail page and the API both stop there, and there is no rollup of older data.

That is enough to judge today. It is not enough for a capacity decision, which is a question about last week and last month: a Monday-morning peak or a month-end batch is easy to miss in a single day. For the long view, send the figures to your own metrics backend as well — see Keep a longer history.

The sizing model has two drivers, and they are different quantities: CPU follows log throughput, memory follows metric cardinality (active series). Read each against its own figure.

You want to know Read Where
How much CPU the collector uses Process CPU Collector → Health → Process CPU (current value). History: cpuPercent from the API below.
How much memory the collector uses Process memory Collector → Health → Process Memory (current value, MB). History: memoryMb from the API below.
How much it is moving Records per second per component Topology canvas edges (throughput overlay), and the throughput API.
How many bytes it is moving Bytes per second on each destination Turn on Detailed telemetry (collector → Overview → Telemetry), then read the destination exporter’s bytesPerSec from the throughput API.
Whether a destination keeps up Sender queue fill The queue bar on the topology edge, with the throughput overlay on. Hover it for the figures: Sender queue 640/1000 (64%).
Whether anything is being lost Dropped and refused The throughput history API: droppedPerSec on the destination exporter, errorsPerSec on the receiver.

Process CPU is a percentage of one core. 100 means one full core, 250 means two and a half. Compare it with the number of cores the collector has, not with 100.

The Health tab charts describe the host, not the collector. On a collector that reports host metrics, the CPU Usage chart plots the host’s CPU utilisation and the Memory Usage chart plots the host’s memory utilisation as a percentage, even though its axis is labelled MB. On a host that runs nothing but the collector, that is close enough. On a shared host, use the process figures.

Grafana Alloy collectors report per-component throughput and queue figures to LinkMesh, but not process CPU, process memory or host metrics. Read those from Alloy’s own metrics endpoint (http://127.0.0.1:12345/metrics on the collector host): process_cpu_seconds_total and process_resident_memory_bytes.

Bytes per second is measured at the destination, as each batch is encoded for export. It is the closest LinkMesh gets to the “MiB/s of logs” the CPU coefficient is written in. If a route sends the same records to two destinations, each reports the full volume; take one, not the sum. Detailed telemetry increases the collector’s self-telemetry volume roughly tenfold, so turn it on to measure and off again when you have your number.

The detail page shows current values; the percentiles below need the history. With a service-account token that has collectors:read (see Authenticate with a service account):

export LINKMESH="https://linkmesh.example.com/api/v1"
export LMSAT="lmsat_…"
ID="<collector id>"

# One sample per minute for the last 24 h: cpuPercent, memoryMb, host CPU/memory, network.
curl -fsS -H "Authorization: Bearer $LMSAT" "$LINKMESH/collectors/$ID/metrics?range=24h" \
  | jq '{samples: (.data|length),
         p95_cpu_percent: ([.data[].cpuPercent] | sort | .[(length*0.95|floor)]),
         p95_memory_mb:   ([.data[].memoryMb]   | sort | .[(length*0.95|floor)]),
         max_memory_mb:   ([.data[].memoryMb]   | max)}'

Per-component throughput, bucketed, for the last 24 h (rangeSeconds stops at 86400). Component keys are the ones the throughput API lists for the collector, such as otlp/output_dst_… for a destination:

KEY="otlp/output_dst_…"
curl -fsS -H "Authorization: Bearer $LMSAT" \
  "$LINKMESH/collectors/$ID/throughput/series?key=$KEY&rangeSeconds=86400&bucketSeconds=300" \
  | jq --arg k "$KEY" '.series[$k]
      | {p95_records_per_sec: ([.[].recordsPerSec] | sort | .[(length*0.95|floor)]),
         p95_bytes_per_sec:   ([.[].bytesPerSec]   | sort | .[(length*0.95|floor)])}'

GET $LINKMESH/collectors/$ID/throughput/history returns the raw rows, including queueSize, queueCapacity and droppedPerSec.

Measure against the row the collector was sized for — its cores and its memory. Use the 95th percentile over the last 24 hours of one-minute samples, not the highest sample.

A percentile, because a single spike is not a capacity signal. A collector that touches its limit for one minute while a destination reconnects is behaving normally; one that spends more than an hour a day near it is not. The 95th percentile over a day is the level exceeded for about 72 minutes of that day.

Signal Watch Act Why these levels
Process CPU (p95, % of one core) ≥ 70 % of the collector’s cores ≥ 85 % of its cores 70 % is the target utilisation the sizing table already assumes (U_cpu). Past 85 % at the 95th percentile, the peaks are already at the limit. CPU saturation shows up as latency and a growing queue, not as a crash.
Process memory (p95) ≥ 80 % of the memory limit ≥ 90 % of the limit, or any single sample ≥ 95 % 80 % is the sizing table’s U_mem. Memory is the one driver where a single peak does matter: running out of memory is an immediate kill, not a slowdown.
Sender queue (fill, on any destination) Stays above zero while traffic is normal Reaches 80 % A queue should drain as fast as it fills. A standing queue means the destination can’t keep up. The topology edge marks the queue as a warning at 80 % and critical at 95 %.
Dropped or refused — Any non-zero value This is loss. It is not a sizing threshold; act on it now.

For a collector on row C (4 cores, 8 GiB), that works out to:

  • Process CPU: watch at a p95 of 280 %, act at 340 %.
  • Process memory: watch at a p95 of 6,550 MB, act at 7,370 MB or any sample above 7,780 MB.

Take the p95 from LinkMesh’s 24 hours to judge today, and from your own backend over a week or more to judge the trend. A collector that crosses watch every day for a week is due for the decision order below; one that crosses it once is due for a look at what happened that day.

When a collector crosses act, work through these in order. Most operators run the list backwards: they add a replica first and trim the data last.

  1. Reduce the demand. Cut the cardinality: drop labels nobody queries, stop scraping targets nobody reads. Drop or sample logs nobody looks at. This is the only step that makes the collector’s job smaller rather than giving it more to do the same job with, and it usually costs nothing downstream either. See Drop noisy logs and Filter records on a route.

  2. Then scale out. Add a collector and split the sources between them. This relieves CPU and throughput. It does not relieve a collector that is bound by memory unless each series goes to exactly one instance: under round-robin balancing, every instance ends up holding nearly every series. See Scaling out and memory. If memory is the driver that crossed act, skip to step 3 unless you have arranged that split.

  3. Then scale up. Give the collector the next row in the selection table. This is the fix for a memory-bound collector, and the simplest fix for a CPU-bound one you can’t split.

  4. Only then change storage and queues. A larger sending queue buys time while a destination is away; it doesn’t add throughput, and it holds memory while it fills. It is the answer to “my destination has outages”, not to “my collector is too small”.

Re-check on change, not only on a schedule

Section titled “Re-check on change, not only on a schedule”

Capacity problems arrive with a change, not on a calendar. Re-read the figures above — and re-run the sizing selection with the new workload — whenever you:

  • Onboard a materially larger source, or many small ones at once.
  • Add a destination. Each destination adds its own export work and its own queue.
  • Widen a processor’s scope — a parser, masking rule or transform that now runs on more records than before.
  • Turn on a metrics source with many targets, or change a scrape interval: a shorter interval multiplies datapoints and CPU; new targets multiply series and memory.

Compare the day after the change with the day before, while both are still inside LinkMesh’s 24 hours.

Gauges are easy to read in only one direction. Before acting on a number, check what it counts.

  • Records per second is records, not bytes. A thousand one-line records and a thousand stack traces are the same number here and very different CPU. Use bytes per second for the CPU driver.
  • A destination’s sent rate means “the destination acknowledged it”, not “the destination stored it”. Records a backend drops after acknowledging them — rate limits, validation, retention rules applied on its side — don’t appear here. Check the backend’s own ingestion figures for those.
  • An error rate on a destination arrives late. While a destination is unreachable the collector keeps retrying, and the error figure stays at zero. The first sign is the sender queue filling. In a test with the default retry settings, the first destination errors appeared about five minutes after the destination went away, when the collector stopped retrying the oldest batches. By then those batches are gone.
  • Dropped means discarded at the collector. When a destination’s queue is full, new batches are thrown away and counted as dropped. In the same test, from the moment the queue filled, dropped matched the whole input rate.
  • Refused happens at the sender. Once a collector can’t pass data on, its receivers start refusing it, and the refused figure appears on the receiver, not the destination. Whether the refused data is lost or retried depends on the sender: an OTLP client may retry, a plain TCP or syslog sender usually does not.
  • Queue size counts batches, not records. 640/1000 is 640 batches waiting out of a capacity of 1,000 batches.
  • Components named …/self are the collector observing itself — hostmetrics/self, resource/self, otlphttp/self. Their few records per second are the self-telemetry on this page. Leave them out when you add up throughput.

LinkMesh routes telemetry; it doesn’t store it long-term. To judge a week or a month, send the collector’s resource figures to your own metrics backend, the same way you send everything else:

  1. Add a Host Metrics source to the collector. It ships in the source catalog; create a source from it with the cpu and memory scrapers.

  2. Route it to your metrics destination with a route whose filter matches it.

The collector then sends the host’s CPU time (system.cpu.time) and memory use (system.memory.usage) to your backend alongside everything else it forwards, and you can take the 95th percentile over any window your backend retains. On a host that runs only the collector, those are its figures.

For Alloy collectors, the process figures and the active-series count live on Alloy’s own metrics endpoint (127.0.0.1:12345/metrics). Scrape it from your monitoring stack to keep them.