LinkMesh

Search docs, blog and changelog

ENDE
One master lever connected by a single linkage bar to a bank of identical valve actuators
LinkMeshObservability Data Collection Management
OpenTelemetryOperations

Java Agent in Production

One flag instruments everything. That is the problem as much as the point.

linkmesh.io
Roman HüslerRoman Hüsler← Back to blog
6 min read

The OpenTelemetry Java agent is the most effective single artefact in the whole project: one -javaagent: flag, and a Spring Boot service you have never touched emits traces for every HTTP request, JDBC call, Kafka message and outbound client call, plus JVM metrics, with W3C context propagation across all of it. No code change. That is why it is the default answer whenever a Java estate has to leave a vendor APM agent.

It is also why teams get burned in the first week of production. The agent’s defaults are tuned for coverage, not for your traffic — and the moment it meets a real request rate you discover which of the hundred instrumentations you did not want, what a 100 % sampling rate costs, and that your logs are now going two places. This post is the production checklist: what the agent does, what it costs, and what to change before it goes live.

Who this guide is for

Engineers rolling the Java agent onto services with real traffic — typically as part of a vendor-APM migration. Prerequisites: you can change JVM flags and environment variables for your services, and you have an OTLP endpoint to send to. This is the application-side counterpart to the collector-side posts on this blog; it stops at the collector’s front door.

What you get with one flag

java -javaagent:/opt/otel/opentelemetry-javaagent.jar \
     -Dotel.service.name=payments-api \
     -Dotel.exporter.otlp.endpoint=http://localhost:4317 \
     -jar payments-api.jar

Attached at JVM start, the agent rewrites bytecode as classes load and instruments whichever of its supported libraries it finds — servlet containers, Spring, JDBC drivers, HTTP clients, Kafka, gRPC, Redis clients, and a long tail. Out of that you get:

  • Traces for every server request and every outbound call, with parent–child spans and W3C traceparent propagation, so a request crossing four services is one trace.
  • JVM runtime metrics — heap and non-heap usage, GC pause counts and durations, thread counts, class loading — under the jvm.* semantic conventions.
  • Log correlation: trace and span IDs injected into the MDC, so existing Logback or Log4j2 output can carry them; and optionally the logs themselves exported over OTLP.

All of it in the OTLP wire format, to wherever you point it. That last property is the point of the exercise — see what OpenTelemetry can and cannot replace for the honest scope of the migration it enables.

The production checklist

Six things to settle before the first real deployment, in roughly the order they bite:

1. Send to a local collector, never straight to the backend. The agent’s exporter is fine, but it is inside your process: retries, queueing and backpressure happen in your service’s memory and on its threads. A collector on localhost (or the node) absorbs that, batches, and gives you one egress point to redact, sample and route through — the whole argument in what is a telemetry pipeline. It also means a backend outage degrades the collector, not your service.

2. Set sampling explicitly. The default is to sample everything a parent did not decide on, which at 100 % is the right choice on day one and the wrong one at scale. The production-safe pattern is parent-based ratio sampling at the agent, with the keep errors and slow traces rule enforced by tail sampling at the gateway — the agent cannot see the outcome of a trace it is at the root of. Rates belong under central governance, not per service: sampling governance.

-Dotel.traces.sampler=parentbased_traceidratio
-Dotel.traces.sampler.arg=0.1

3. Turn off the instrumentations you do not want. The agent enables everything it recognises. Common candidates to disable on day one: internal-only libraries that produce a span per method, health-check endpoints that produce a trace per probe, and any instrumentation whose spans you have never looked at. Each is one property:

-Dotel.instrumentation.<name>.enabled=false

Do this by allow-listing rather than hunting: disable the common defaults and enable only what you use, then widen. The disabled set is configuration worth versioning with the service.

4. Decide where logs go — once. Two paths exist and running both is the classic mistake: the agent can export logs over OTLP via the Logback/Log4j2 appender bridge, and your existing file or stdout logs are probably already being tailed by a collector. Pick one. If the collector already tails files, keep that and let the agent only inject trace IDs into the MDC; if you want structured logs with resource attributes attached, use the bridge and stop tailing. Both is duplicated ingest with a duplicated bill.

5. Budget the startup cost. Bytecode instrumentation happens at class load, so startup is slower — commonly a few seconds on a large Spring application, occasionally more. On Kubernetes that interacts with readiness probes and rolling deploys; on serverless or scale-to-zero it is a real problem. Measure it on your slowest service before it surprises a deploy. Steady-state overhead is usually low single-digit CPU, but only after sampling and instrumentation scope are set — with defaults it can be noticeably more.

6. Sanitise what leaves the process. JDBC instrumentation captures statements; by default the agent sanitises literals, and you want that on. HTTP instrumentation can capture headers if asked; do not ask for Authorization. Anything the agent puts on a span attribute is on its way to a backend, so the same masking discipline as for logs applies — masking PII — and the collector is the place to enforce it fleet-wide, because you will not audit every service’s flags.

Things that will surprise you

  • Version skew across services. The agent is a jar you ship with each service, so a hundred services means a hundred copies at whatever version each team last bumped. Semantic-convention changes between agent versions can rename attributes your dashboards use. Pin it, and roll it deliberately — the same discipline as for the collector fleet, one layer down.
  • The agent and the SDK are different things. Teams that later add manual spans with the API get the agent’s auto-instrumentation and their own spans in one trace, which is the intended model. Adding the SDK separately on top of the agent is not, and produces duplicate exporters.
  • Frameworks the agent does not know. In-house RPC layers, unusual thread pools and reactive code that hops threads can lose context. The fix is usually a small amount of manual propagation, not a different agent — but it should be found in staging, not by an incident with a truncated trace.

The honest position

The Java agent is the best thing OpenTelemetry ships and the reason most Java estates can leave a vendor APM at all. It is also a component with production-grade blast radius that gets deployed with tutorial-grade defaults. Treat the six items above as the go-live gate, put the agent behind a collector so the fleet-wide decisions happen in one place, and version the agent like the dependency it is.

Agents on every JVM and a collector on every host?

LinkMesh manages the OpenTelemetry collectors your Java agents export to, from one self-hosted control plane — tail sampling, redaction and routing composed once and rolled to every node, so the per-service agent flags stay simple and the fleet-wide decisions happen where you can see them. See what it does, or set one up in a few minutes.