Collector sizing
This page answers one question: how much CPU and memory does a collector need? It covers the collectors themselves — the data plane. Sizing the LinkMesh server is a separate question with a separate answer, on System requirements & sizing.
The two drivers are different quantities
Section titled “The two drivers are different quantities”The most useful thing to know before sizing anything:
| Resource | Driven by | Rough cost |
|---|---|---|
| CPU | log throughput (MiB/s of log ingest) | ~1 core per 1 MiB/s of logs |
| Memory | metric cardinality (number of active series) | ~11 GiB per 1 million active series |
These are not two views of the same number. A collector moving a large volume of logs needs cores and very little memory; a collector carrying a large set of metric series needs memory and comparatively little CPU.
Per-unit costs
Section titled “Per-unit costs”The coefficients below are Grafana’s, published for Grafana Alloy (see Sources). LinkMesh did not measure them. They are the input to everything else on this page.
| Signal | Per | CPU | Memory | Network |
|---|---|---|---|---|
| Logs | 1 MiB/s ingested | 1 core | 120 MiB | — |
| Metrics | 1 million active series, default scrape interval | 0.4 cores | 11 GiB | 1.5 MiB/s (send + receive) |
| Traces | 1 MB/s ingested | ~0.13 cores | ~0.17 GiB | — |
Two more conversions the table below uses:
- Datapoints per second is series divided by scrape interval. The metrics coefficients assume Alloy’s default 60 s interval, where 1 million series is ~16,700 datapoints/s. Metrics CPU follows datapoints/s; metrics memory follows series. Scraping every 15 s quadruples the datapoints and the CPU; the memory barely moves.
- Spans per second is traces MB/s divided by your average encoded
span size. The table assumes 1 KB per span, so 1 MB/s is
1,000 spans/s. Real spans range from a few hundred bytes to several
kilobytes; scale the spans column by
1 KB ÷ your average.
Where the figures come from, and where they may not fit
Section titled “Where the figures come from, and where they may not fit”- They were measured on Grafana Alloy. LinkMesh also manages otelcol-contrib, which is a different implementation. The shape (logs cost CPU, series cost memory) follows from what the work is, but the coefficients themselves are Alloy’s. Use them as a starting point for otelcol-contrib and measure.
- Grafana notes the metrics figures come from clustered deployments and “broadly apply” to other modes.
- Grafana also notes logs can cost more per MiB/s on many small nodes, because each collector carries a fixed baseline. A per-host collector moving a trickle of logs is dominated by that baseline, not by the coefficient.
- Your processor chain adds cost on top: parsing, masking and transforms spend CPU per record, and an export queue holds memory while a destination is slow.
Headroom: the named factors
Section titled “Headroom: the named factors”The table does not hide padding inside its numbers. Every reserve is a named factor you can see and change, and each has one reason:
| Factor | Value | Why |
|---|---|---|
P — peak-to-average |
1.5 | You size for the peak, but you usually know the daily volume. Grafana’s reference plan pairs 1 TB/day of logs (an average of ~11.6 MB/s) with a 17.5 MB/s peak, a ratio of ~1.5. |
U_cpu — CPU target utilisation |
0.70 | CPU saturation shows up as latency and a growing queue, never as a crash, so it’s found late. Grafana’s reference gateway autoscales at 70 % CPU. |
U_mem — memory target utilisation |
0.80 | The Go garbage collector needs room to work before the memory limit is reached. Grafana sets the Go memory limit (GOMEMLIMIT) at ~80 % of the container limit for the same reason. |
B_net — network burst allowance |
2.0 | When a destination comes back after an outage, the collector drains its queue on top of live traffic. The link must carry that without becoming the new bottleneck. |
A fifth factor is yours to set, because only you know it: growth. Size for the workload at the end of your planning horizon, not today’s. Put that number in as your input; the table doesn’t add any growth of its own.
If your peaks are sharper than 1.5× the daily average, or you want more reserve than 70 %, change the factor and recompute. That is the point of stating it.
Selection table
Section titled “Selection table”Each row is a collector size. Each column is the most that size
carries of one driver on its own, with U_cpu, U_mem and P from
the table above already applied.
| Row | CPU | Memory | Logs, peak | Logs, per day (at P = 1.5) |
Active series | Datapoints/s | Spans/s (at 1 KB/span) |
|---|---|---|---|---|---|---|---|
| A | 1 core | 2 GiB | 0.7 MiB/s | ~40 GB | 145k | 29k | 5k |
| B | 2 cores | 4 GiB | 1.4 MiB/s | ~80 GB | 290k | 58k | 10k |
| C | 4 cores | 8 GiB | 2.8 MiB/s | ~165 GB | 580k | 115k | 21k |
| D | 8 cores | 16 GiB | 5.6 MiB/s | ~335 GB | 1.16M | 230k | 43k |
| E | 16 cores | 32 GiB | 11.2 MiB/s | ~675 GB | 2.3M | 465k | 86k |
How each column is derived, so you can extend the ladder or change a factor:
- Logs, peak =
cores × U_cpu ÷ 1 core per MiB/s. Memory is never the limit for logs at these shapes. - Logs, per day = logs peak ÷
P, converted to GB/day. - Active series =
memory × U_mem ÷ 11 GiB per million. - Datapoints/s =
cores × U_cpu ÷ 0.4 cores per 16.7k datapoints/s. At a 60 s interval, the series column binds long before this one. The datapoints column only matters when you scrape at short intervals. - Spans/s =
cores × U_cpu ÷ 0.13 cores per MB/s, at 1 KB per span.
The selection rule
Section titled “The selection rule”-
Find the row each driver needs on its own. Look up your logs, series, datapoints/s and spans/s separately and note the smallest row that carries each one.
-
The largest of those rows wins. Not the average, and not the driver you care about most. One driver that doesn’t fit is enough to make the whole collector not fit.
-
Add up what shares a resource. The columns are limits for each driver alone, but logs, datapoints and spans all draw on the same cores, and series, logs and traces all draw on the same memory. For the row from step 2, divide each need by its column and add up the fractions for CPU and for memory separately. If either sum is over 1, move up one row and check again.
Step 3 matters whenever two drivers land in the same row, which is common. The worked example below is exactly that case.
Scaling out and memory
Section titled “Scaling out and memory”Before adding a replica to fix a memory problem, check that the series will actually divide between instances. If they won’t, size the single instance up instead, or split the sources.
Worked example
Section titled “Worked example”This workload is invented for illustration: an aggregation collector receiving from a group of application hosts.
| Input | Value |
|---|---|
| Logs | 120 GB/day |
| Metrics | 400,000 active series, scraped every 30 s |
| Traces | 6,000 spans/s, average span ~1 KB |
| Growth | none. Today’s volume is the horizon for this example. |
1. Convert to the table’s units.
- Logs: 120 GB/day ÷ 86,400 s ≈ 1.39 MB/s ≈ 1.32 MiB/s average.
×
P(1.5) = 1.99 MiB/s peak. - Datapoints/s: 400,000 series ÷ 30 s = 13,300/s.
- Traces: 6,000 spans/s × 1 KB = 6 MB/s.
2. Find the row each driver needs on its own.
| Driver | Need | Smallest row that carries it |
|---|---|---|
| Logs, peak | 1.99 MiB/s | C (2.8), because B carries 1.4 |
| Active series | 400k | C (580k), because B carries 290k |
| Datapoints/s | 13.3k | A (29k) |
| Spans/s | 6k | B (10k) |
3. The largest row wins: C. Logs and series both land there.
4. Add up what shares a resource, on row C.
- CPU: logs 1.99 ÷ 2.8 = 0.71, datapoints 13.3k ÷ 115k = 0.12, spans 6k ÷ 21k = 0.29. Sum 1.11. That’s over 1.
- Memory: series 400k ÷ 580k = 0.69, plus logs (1.99 × 120 MiB ≈
0.23 GiB) and traces (6 × 0.17 ≈ 1.0 GiB) against C’s 6.4 GiB usable
(8 GiB ×
U_mem) → 0.88. Fits.
Row C fits the memory but not the CPU: logs and traces together use more than its cores allow. Move up one row.
5. Check row D.
- CPU: 1.99 ÷ 5.6 + 13.3k ÷ 230k + 6k ÷ 43k = 0.36 + 0.06 + 0.14 = 0.55. Fits.
- Memory: the same ~5.65 GiB against D’s 12.8 GiB usable = 0.44. Fits.
Result: row D — 8 cores, 16 GiB. CPU chose the row. Memory would have been content with C.
6. Check the network. Assuming data leaves roughly as large as it arrives (compression on export reduces this):
- Logs: 1.99 MiB/s in + out ≈ 4.0 MiB/s
- Metrics: 1.5 MiB/s per million series at a 60 s interval → 0.6 MiB/s for 400k; doubled for the 30 s interval → 1.2 MiB/s
- Traces: 6 MB/s in + out ≈ 11.4 MiB/s
≈ 16.6 MiB/s at peak, × B_net (2.0) ≈ 33 MiB/s, about 280 Mbit/s
the host’s link must sustain.
The alternative, two row-C instances behind a load balancer, would halve the CPU per instance (0.56, fits). But if the metrics arrive round-robin, each instance still holds nearly all 400k series, so each still needs the full ~5.65 GiB. That fits in C’s 6.4 GiB at 0.88, but with little margin, and a third replica would buy no memory at all. See Scaling out and memory.
Check against your own collectors
Section titled “Check against your own collectors”The coefficients are an estimate. Once a collector runs, its own figures replace them. LinkMesh collects each managed collector’s per-component throughput and sender queues and, for otelcol-contrib collectors, its process CPU and memory, with nothing to set up.
When to resize a collector shows where to read those figures against the row you chose, the watch and act thresholds, the order to try fixes in, and how to keep a longer history than the 24 hours LinkMesh holds.
Sources
Section titled “Sources”Figures on this page that LinkMesh did not measure are Grafana’s:
- Grafana Labs, Estimate Grafana Alloy resource usage (Alloy documentation). The source of the logs and metrics coefficients and of the clustering and small-node caveats.
- Grafana Labs, How to scale Alloy as a central telemetry gateway:
capacity planning, load testing, and production
lessons
(blog, 2026). The source of the reference gateway plan the traces row
is derived from, the 1 TB/day-to-17.5 MB/s peak ratio behind
P, the 70 % CPU autoscaling target behindU_cpu, and the 80 % memory limit behindU_mem.
The headroom factors, the selection rule, the derived traces row, the scale-out reasoning and the worked example are LinkMesh’s. No LinkMesh lab measurements are used on this page.
Related
Section titled “Related”- When to resize a collector — reading a running collector’s figures against this table, and what to do when it outgrows its row
- System requirements & sizing — sizing the LinkMesh server, and why it scales with fleet size rather than telemetry volume
- Self-telemetry — the per-collector CPU, memory and throughput figures LinkMesh collects
- Manage collector groups — running several collectors as one unit