💻 Coding

OpenTelemetry Collector Config Writer for otelcol-contrib: OTLP, Host Metrics, and Filelog Receivers, memory_limiter First, Batch, PII Redaction, Tail Sampling, Agent vs Gateway Pipelines, and otelcol validate

Generate an OpenTelemetry Collector (contrib distribution) YAML config you can validate before deploying: OTLP gRPC and HTTP, hostmetrics, and filelog receivers, processors in a safe order with memory_limiter first, attribute based PII redaction, tail sampling policies on the gateway tier, exporters, health_check and zpages extensions, and separate service.pipelines for traces, metrics, and logs.

0.0
0Reviews
P
October 7, 2026

Prompt

Act as a site reliability engineer who runs the OpenTelemetry Collector contrib distribution (otelcol-contrib) in production as both a per host agent and a central gateway, and who has debugged dropped spans, memory kills, and leaked user emails in log bodies.

Inputs:
- Deployment shape: agent only, gateway only, or agent plus gateway; platform (VM, Kubernetes DaemonSet and Deployment, Docker); Collector version: [Deployment]
- Signals and sources: which apps send OTLP over gRPC or HTTP, which hosts need host metrics, which log files to tail and their format: [Signals]
- Backend exporters: vendor or self hosted endpoints per signal, auth header names (never values), TLS needs: [Backends]
- Sensitive attributes to remove or hash, such as user.email, http.request.header.authorization, client IPs: [PiiFields]
- Sampling goals: keep all errors, keep slow traces above a threshold, a baseline percentage for the rest: [SamplingPolicy]
- Container or host memory limit for each Collector: [MemoryBudget]
- Output format: [Format]

Generate:
1. A topology note: what runs on the agent and what runs on the gateway. Tail sampling must run where all spans of a trace arrive, so if there is more than one gateway replica, route traces by trace ID with the loadbalancing exporter on the agent tier.
2. Receivers: otlp with grpc and http endpoints set explicitly (bind to 0.0.0.0 only where the network requires it), hostmetrics with an interval and only the scrapers needed, filelog with include paths, start_at, and a parser operator matching the log format in Signals.
3. Processors in order: memory_limiter first with check_interval and limits derived from MemoryBudget, then resource or attributes processors that delete or hash every PiiFields key, then tail_sampling on the gateway only, then batch last before export.
4. tail_sampling policies from SamplingPolicy: status_code for errors, latency with threshold_ms, probabilistic for the baseline, plus decision_wait and num_traces with a short reason for each value.
5. Exporters per signal from Backends, with auth read from environment variables using the env syntax, retry and sending queue left on.
6. Extensions: health_check and zpages, listed under service.extensions, with ports noted for probes.
7. service.pipelines: separate traces, metrics, and logs pipelines; every component referenced exists and every defined component is used.
8. Validation steps: otelcol-contrib validate --config with the file path, then a smoke test that sends one span and checks zpages or the debug exporter.
9. A risk list: what fails first if MemoryBudget is too small, and how to spot dropped data in the Collector's own metrics.

Constraints:
- Only use components that exist in otelcol-contrib; if a field name may differ by version, mark it CHECK against the Deployment version.
- Never put secrets in YAML. No em dashes. Keep comments short and inside the YAML.

Instructions

Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.

Generated Output

This image was generated using the prompt above.

OpenTelemetry Collector Config Writer for otelcol-contrib: OTLP, Host Metrics, and Filelog Receivers, memory_limiter First, Batch, PII Redaction, Tail Sampling, Agent vs Gateway Pipelines, and otelcol validate - Result

Examples

Example Input

Deployment: agent plus gateway on Kubernetes; agent as DaemonSet, gateway as Deployment with 3 replicas behind a headless Service otel-gateway-headless; Collector contrib 0.11x
Signals: Java and Node services send OTLP gRPC to the node agent; nginx JSON access logs at /var/log/nginx/access.json; host metrics for cpu, memory, filesystem
Backends: Grafana Cloud style OTLP HTTP endpoint for all signals, header Authorization from env OTLP_AUTH
PiiFields: user.email (delete), enduser.id (hash), http.request.header.authorization (delete)
SamplingPolicy: keep every error trace, keep traces slower than 800 ms, 10 percent of everything else
MemoryBudget: agent 512Mi, gateway 2Gi
Format: topology note, gateway YAML, agent notes, validation

Example Output

Topology: agents receive OTLP, host metrics, and nginx logs, redact PII, and send traces through the loadbalancing exporter keyed on traceID so every span of a trace lands on one of the 3 gateways. Gateways run tail_sampling and export to the backend.

gateway.yaml
extensions:
  health_check:
    endpoint: 0.0.0.0:13133   # readiness and liveness probe
  zpages:
    endpoint: 0.0.0.0:55679
receivers:
  otlp:
    protocols:
      grpc:
        endpoint: 0.0.0.0:4317
processors:
  memory_limiter:            # always first
    check_interval: 1s
    limit_percentage: 80     # of the 2Gi container limit
    spike_limit_percentage: 20
  attributes/pii:            # second line of defense if an agent missed it
    actions:
      - key: user.email
        action: delete
      - key: http.request.header.authorization
        action: delete
      - key: enduser.id
        action: hash
  tail_sampling:
    decision_wait: 10s       # longer than the slowest normal trace
    num_traces: 50000        # traces held in memory; size against 2Gi, CHECK
    policies:
      - name: errors
        type: status_code
        status_code:
          status_codes: [ERROR]
      - name: slow
        type: latency
        latency:
          threshold_ms: 800
      - name: baseline
        type: probabilistic
        probabilistic:
          sampling_percentage: 10
  batch: {}
exporters:
  otlphttp:
    endpoint: https://otlp.example-backend.net/otlp
    headers:
      Authorization: ${env:OTLP_AUTH}
service:
  extensions: [health_check, zpages]
  pipelines:
    traces:
      receivers: [otlp]
      processors: [memory_limiter, attributes/pii, tail_sampling, batch]
      exporters: [otlphttp]
    metrics:
      receivers: [otlp]
      processors: [memory_limiter, attributes/pii, batch]
      exporters: [otlphttp]
    logs:
      receivers: [otlp]
      processors: [memory_limiter, attributes/pii, batch]
      exporters: [otlphttp]

Agent notes (agent.yaml)
- receivers: otlp grpc 0.0.0.0:4317; hostmetrics with collection_interval 30s and scrapers cpu, memory, filesystem; filelog include /var/log/nginx/access.json, start_at end, operator json_parser.
- processors: memory_limiter (limit_percentage 80 of 512Mi), the same attributes/pii block, batch.
- exporters: loadbalancing for traces with routing_key traceID, protocol otlp, resolver dns hostname otel-gateway-headless; otlp to the gateway Service for metrics and logs.
- No tail_sampling on agents: an agent sees only part of each trace.

Validation
1. otelcol-contrib validate --config=gateway.yaml, then the same for agent.yaml.
2. Send one test span with telemetrygen or an app, then open zpages /debug/tracez on a gateway.
3. Confirm user.email is absent on the backend for a test request that set it.

Risks
- Too little memory: memory_limiter starts refusing data first; watch the Collector's refused and dropped span metrics and the exporter queue size.
- num_traces too low for traffic: traces get evicted before decision_wait ends.

Reviews (0)

Please login to leave a review.
Loading reviews...