Skip to content

Collector sizing

This page answers one question: how much CPU and memory does a collector need? It covers the collectors themselves — the data plane. Sizing the LinkMesh server is a separate question with a separate answer, on System requirements & sizing.

The most useful thing to know before sizing anything:

Resource Driven by Rough cost
CPU log throughput (MiB/s of log ingest) ~1 core per 1 MiB/s of logs
Memory metric cardinality (number of active series) ~11 GiB per 1 million active series

These are not two views of the same number. A collector moving a large volume of logs needs cores and very little memory; a collector carrying a large set of metric series needs memory and comparatively little CPU.

The coefficients below are Grafana’s, published for Grafana Alloy (see Sources). LinkMesh did not measure them. They are the input to everything else on this page.

Signal Per CPU Memory Network
Logs 1 MiB/s ingested 1 core 120 MiB —
Metrics 1 million active series, default scrape interval 0.4 cores 11 GiB 1.5 MiB/s (send + receive)
Traces 1 MB/s ingested ~0.13 cores ~0.17 GiB —

Two more conversions the table below uses:

  • Datapoints per second is series divided by scrape interval. The metrics coefficients assume Alloy’s default 60 s interval, where 1 million series is ~16,700 datapoints/s. Metrics CPU follows datapoints/s; metrics memory follows series. Scraping every 15 s quadruples the datapoints and the CPU; the memory barely moves.
  • Spans per second is traces MB/s divided by your average encoded span size. The table assumes 1 KB per span, so 1 MB/s is 1,000 spans/s. Real spans range from a few hundred bytes to several kilobytes; scale the spans column by 1 KB ÷ your average.

Where the figures come from, and where they may not fit

Section titled “Where the figures come from, and where they may not fit”
  • They were measured on Grafana Alloy. LinkMesh also manages otelcol-contrib, which is a different implementation. The shape (logs cost CPU, series cost memory) follows from what the work is, but the coefficients themselves are Alloy’s. Use them as a starting point for otelcol-contrib and measure.
  • Grafana notes the metrics figures come from clustered deployments and “broadly apply” to other modes.
  • Grafana also notes logs can cost more per MiB/s on many small nodes, because each collector carries a fixed baseline. A per-host collector moving a trickle of logs is dominated by that baseline, not by the coefficient.
  • Your processor chain adds cost on top: parsing, masking and transforms spend CPU per record, and an export queue holds memory while a destination is slow.

The table does not hide padding inside its numbers. Every reserve is a named factor you can see and change, and each has one reason:

Factor Value Why
P — peak-to-average 1.5 You size for the peak, but you usually know the daily volume. Grafana’s reference plan pairs 1 TB/day of logs (an average of ~11.6 MB/s) with a 17.5 MB/s peak, a ratio of ~1.5.
U_cpu — CPU target utilisation 0.70 CPU saturation shows up as latency and a growing queue, never as a crash, so it’s found late. Grafana’s reference gateway autoscales at 70 % CPU.
U_mem — memory target utilisation 0.80 The Go garbage collector needs room to work before the memory limit is reached. Grafana sets the Go memory limit (GOMEMLIMIT) at ~80 % of the container limit for the same reason.
B_net — network burst allowance 2.0 When a destination comes back after an outage, the collector drains its queue on top of live traffic. The link must carry that without becoming the new bottleneck.

A fifth factor is yours to set, because only you know it: growth. Size for the workload at the end of your planning horizon, not today’s. Put that number in as your input; the table doesn’t add any growth of its own.

If your peaks are sharper than 1.5× the daily average, or you want more reserve than 70 %, change the factor and recompute. That is the point of stating it.

Each row is a collector size. Each column is the most that size carries of one driver on its own, with U_cpu, U_mem and P from the table above already applied.

Row CPU Memory Logs, peak Logs, per day (at P = 1.5) Active series Datapoints/s Spans/s (at 1 KB/span)
A 1 core 2 GiB 0.7 MiB/s ~40 GB 145k 29k 5k
B 2 cores 4 GiB 1.4 MiB/s ~80 GB 290k 58k 10k
C 4 cores 8 GiB 2.8 MiB/s ~165 GB 580k 115k 21k
D 8 cores 16 GiB 5.6 MiB/s ~335 GB 1.16M 230k 43k
E 16 cores 32 GiB 11.2 MiB/s ~675 GB 2.3M 465k 86k

How each column is derived, so you can extend the ladder or change a factor:

  • Logs, peak = cores × U_cpu ÷ 1 core per MiB/s. Memory is never the limit for logs at these shapes.
  • Logs, per day = logs peak ÷ P, converted to GB/day.
  • Active series = memory × U_mem ÷ 11 GiB per million.
  • Datapoints/s = cores × U_cpu ÷ 0.4 cores per 16.7k datapoints/s. At a 60 s interval, the series column binds long before this one. The datapoints column only matters when you scrape at short intervals.
  • Spans/s = cores × U_cpu ÷ 0.13 cores per MB/s, at 1 KB per span.
  1. Find the row each driver needs on its own. Look up your logs, series, datapoints/s and spans/s separately and note the smallest row that carries each one.

  2. The largest of those rows wins. Not the average, and not the driver you care about most. One driver that doesn’t fit is enough to make the whole collector not fit.

  3. Add up what shares a resource. The columns are limits for each driver alone, but logs, datapoints and spans all draw on the same cores, and series, logs and traces all draw on the same memory. For the row from step 2, divide each need by its column and add up the fractions for CPU and for memory separately. If either sum is over 1, move up one row and check again.

Step 3 matters whenever two drivers land in the same row, which is common. The worked example below is exactly that case.

Before adding a replica to fix a memory problem, check that the series will actually divide between instances. If they won’t, size the single instance up instead, or split the sources.

This workload is invented for illustration: an aggregation collector receiving from a group of application hosts.

Input Value
Logs 120 GB/day
Metrics 400,000 active series, scraped every 30 s
Traces 6,000 spans/s, average span ~1 KB
Growth none. Today’s volume is the horizon for this example.

1. Convert to the table’s units.

  • Logs: 120 GB/day ÷ 86,400 s ≈ 1.39 MB/s ≈ 1.32 MiB/s average. × P (1.5) = 1.99 MiB/s peak.
  • Datapoints/s: 400,000 series ÷ 30 s = 13,300/s.
  • Traces: 6,000 spans/s × 1 KB = 6 MB/s.

2. Find the row each driver needs on its own.

Driver Need Smallest row that carries it
Logs, peak 1.99 MiB/s C (2.8), because B carries 1.4
Active series 400k C (580k), because B carries 290k
Datapoints/s 13.3k A (29k)
Spans/s 6k B (10k)

3. The largest row wins: C. Logs and series both land there.

4. Add up what shares a resource, on row C.

  • CPU: logs 1.99 ÷ 2.8 = 0.71, datapoints 13.3k ÷ 115k = 0.12, spans 6k ÷ 21k = 0.29. Sum 1.11. That’s over 1.
  • Memory: series 400k ÷ 580k = 0.69, plus logs (1.99 × 120 MiB ≈ 0.23 GiB) and traces (6 × 0.17 ≈ 1.0 GiB) against C’s 6.4 GiB usable (8 GiB × U_mem) → 0.88. Fits.

Row C fits the memory but not the CPU: logs and traces together use more than its cores allow. Move up one row.

5. Check row D.

  • CPU: 1.99 ÷ 5.6 + 13.3k ÷ 230k + 6k ÷ 43k = 0.36 + 0.06 + 0.14 = 0.55. Fits.
  • Memory: the same ~5.65 GiB against D’s 12.8 GiB usable = 0.44. Fits.

Result: row D — 8 cores, 16 GiB. CPU chose the row. Memory would have been content with C.

6. Check the network. Assuming data leaves roughly as large as it arrives (compression on export reduces this):

  • Logs: 1.99 MiB/s in + out ≈ 4.0 MiB/s
  • Metrics: 1.5 MiB/s per million series at a 60 s interval → 0.6 MiB/s for 400k; doubled for the 30 s interval → 1.2 MiB/s
  • Traces: 6 MB/s in + out ≈ 11.4 MiB/s

≈ 16.6 MiB/s at peak, × B_net (2.0) ≈ 33 MiB/s, about 280 Mbit/s the host’s link must sustain.

The alternative, two row-C instances behind a load balancer, would halve the CPU per instance (0.56, fits). But if the metrics arrive round-robin, each instance still holds nearly all 400k series, so each still needs the full ~5.65 GiB. That fits in C’s 6.4 GiB at 0.88, but with little margin, and a third replica would buy no memory at all. See Scaling out and memory.

The coefficients are an estimate. Once a collector runs, its own figures replace them. LinkMesh collects each managed collector’s per-component throughput and sender queues and, for otelcol-contrib collectors, its process CPU and memory, with nothing to set up.

When to resize a collector shows where to read those figures against the row you chose, the watch and act thresholds, the order to try fixes in, and how to keep a longer history than the 24 hours LinkMesh holds.

Figures on this page that LinkMesh did not measure are Grafana’s:

The headroom factors, the selection rule, the derived traces row, the scale-out reasoning and the worked example are LinkMesh’s. No LinkMesh lab measurements are used on this page.