Fractal Techware

Guides /

Span metrics and service graph with OpenTelemetry Collector connectors

Tested with otel/opentelemetry-collector-contrib:0.161.0 and prom/prometheus:v3.14.0: otelcol validate, a smoke test that sent a client/server span pair, and a query of the generated series in Prometheus.

Traces tell you what happened to one request. Dashboards and alerts need rates, errors and durations (RED) per service and endpoint, and a map of who calls whom. You can compute both in the collector from spans you already have, before sampling throws most of them away.

Two connectors do this. A connector is an exporter in one pipeline and a receiver in another.

Older configs call these spanmetrics and servicegraph. In 0.161.0 those names still load but log a deprecation warning.

The config

receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317
      http:
        endpoint: 0.0.0.0:4318

connectors:
  span_metrics:
    histogram:
      unit: s
      explicit:
        buckets: [10ms, 50ms, 100ms, 250ms, 500ms, 1s, 2500ms, 5s]
    dimensions:
      - name: http.request.method
      - name: http.response.status_code
      - name: http.route
    metrics_flush_interval: 15s
    metrics_expiration: 5m
    aggregation_cardinality_limit: 2000

  service_graph:
    latency_histogram_buckets: [10ms, 50ms, 100ms, 250ms, 500ms, 1s, 2500ms, 5s]
    store:
      ttl: 5s
      max_items: 10000
    metrics_flush_interval: 15s

processors:
  memory_limiter:
    check_interval: 1s
    limit_percentage: 80
    spike_limit_percentage: 20
  batch:
    timeout: 5s

exporters:
  otlp_grpc/traces:
    endpoint: tempo.observability.svc:4317
    tls:
      insecure: true
  otlp_http/metrics:
    endpoint: http://mimir.observability.svc:8080/otlp

service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [memory_limiter, batch]
      exporters: [otlp_grpc/traces, span_metrics, service_graph]
    metrics/from_spans:
      receivers: [span_metrics, service_graph]
      processors: [batch]
      exporters: [otlp_http/metrics]

Replace the two endpoints with your trace backend and an OTLP-capable metrics backend. Mimir, Prometheus 3 (with --web.enable-otlp-receiver, endpoint http://prometheus:9090/api/v1/otlp), and most vendors accept OTLP metrics.

The important parts

Wiring. The traces pipeline sends every span to three places: the trace backend and both connectors. The connectors appear again as receivers of metrics/from_spans. A connector must be used on both sides, or validation fails.

histogram.unit: s. The default unit is milliseconds. With seconds, the Prometheus name becomes traces_span_metrics_duration_seconds_bucket, which matches the convention most Grafana dashboards expect. The bucket values accept durations (10ms, 2500ms) either way.

Dimensions are cardinality. Every distinct combination of service.name × span name × kind × status × your dimensions is a separate series. http.route (/orders/{id}) is safe. url.full or user.id would create one series per URL or user. aggregation_cardinality_limit is a hard cap. Extra combinations are folded into an overflow series instead of growing forever.

metrics_expiration forgets series that stopped receiving spans, such as old routes after a deploy. Without it, cumulative series stay in memory as long as the collector runs.

service_graph.store.ttl is how long an unpaired client or server span waits for its partner. Both halves must arrive at the same collector within that time. For most in-cluster calls, 5 s is plenty. Raise it for slow async hops, and remember that memory scales with max_items.

Where to put it. Put the connectors before tail or probabilistic sampling. Sampled traces produce sampled request rates, so a 10% sample makes your traffic look 10× smaller.

Verify it

Use this overlay (debug.yaml). It removes the trace backend from the traces pipeline and prints the generated metrics:

exporters:
  debug:
    verbosity: detailed
service:
  pipelines:
    traces:
      exporters: [span_metrics, service_graph]
    metrics/from_spans:
      exporters: [debug]
docker run --rm -p 4318:4318 -v "$PWD:/cfg" otel/opentelemetry-collector-contrib:0.161.0 \
  --config=/cfg/config.yaml --config=/cfg/debug.yaml

Send one trace with two spans: a CLIENT span (kind: 3) from frontend, and a SERVER span (kind: 2) from checkout whose parentSpanId is the client span’s ID. Put each under its own resource with its own service.name. After the 15 s flush, the debug output in our test contained:

-> Name: traces.span.metrics.calls
-> Name: traces.span.metrics.duration
-> Name: traces_service_graph_request_total
-> Name: traces_service_graph_request_client
-> Name: traces_service_graph_request_server
-> client: Str(frontend)
-> server: Str(checkout)
-> http.route: Str(/checkout)

Sent to Prometheus 3.14 over OTLP, they arrived as traces_span_metrics_calls_total, traces_span_metrics_duration_seconds_bucket, traces_service_graph_request_total and traces_service_graph_request_server_seconds_bucket. Labels included service_name, span_kind, span_name, status_code, http_route and http_response_status_code. A typical RED query:

sum by (service_name, http_route) (rate(traces_span_metrics_calls_total{span_kind="SPAN_KIND_SERVER"}[5m]))

Pitfalls

Next steps