💻 Coding
OpenTelemetry Collector Config Writer for otelcol-contrib: OTLP, Host Metrics, and Filelog Receivers, memory_limiter First, Batch, PII Redaction, Tail Sampling, Agent vs Gateway Pipelines, and otelcol validate
Generate an OpenTelemetry Collector (contrib distribution) YAML config you can validate before deploying: OTLP gRPC and HTTP, hostmetrics, and filelog receivers, processors in a safe order with memory_limiter first, attribute based PII redaction, tail sampling policies on the gateway tier, exporters, health_check and zpages extensions, and separate service.pipelines for traces, metrics, and logs.
0Reviews
Prompt
Act as a site reliability engineer who runs the OpenTelemetry Collector contrib distribution (otelcol-contrib) in production as both a per host agent and a central gateway, and who has debugged dropped spans, memory kills, and leaked user emails in log bodies. Inputs: - Deployment shape: agent only, gateway only, or agent plus gateway; platform (VM, Kubernetes DaemonSet and Deployment, Docker); Collector version: [Deployment] - Signals and sources: which apps send OTLP over gRPC or HTTP, which hosts need host metrics, which log files to tail and their format: [Signals] - Backend exporters: vendor or self hosted endpoints per signal, auth header names (never values), TLS needs: [Backends] - Sensitive attributes to remove or hash, such as user.email, http.request.header.authorization, client IPs: [PiiFields] - Sampling goals: keep all errors, keep slow traces above a threshold, a baseline percentage for the rest: [SamplingPolicy] - Container or host memory limit for each Collector: [MemoryBudget] - Output format: [Format] Generate: 1. A topology note: what runs on the agent and what runs on the gateway. Tail sampling must run where all spans of a trace arrive, so if there is more than one gateway replica, route traces by trace ID with the loadbalancing exporter on the agent tier. 2. Receivers: otlp with grpc and http endpoints set explicitly (bind to 0.0.0.0 only where the network requires it), hostmetrics with an interval and only the scrapers needed, filelog with include paths, start_at, and a parser operator matching the log format in Signals. 3. Processors in order: memory_limiter first with check_interval and limits derived from MemoryBudget, then resource or attributes processors that delete or hash every PiiFields key, then tail_sampling on the gateway only, then batch last before export. 4. tail_sampling policies from SamplingPolicy: status_code for errors, latency with threshold_ms, probabilistic for the baseline, plus decision_wait and num_traces with a short reason for each value. 5. Exporters per signal from Backends, with auth read from environment variables using the env syntax, retry and sending queue left on. 6. Extensions: health_check and zpages, listed under service.extensions, with ports noted for probes. 7. service.pipelines: separate traces, metrics, and logs pipelines; every component referenced exists and every defined component is used. 8. Validation steps: otelcol-contrib validate --config with the file path, then a smoke test that sends one span and checks zpages or the debug exporter. 9. A risk list: what fails first if MemoryBudget is too small, and how to spot dropped data in the Collector's own metrics. Constraints: - Only use components that exist in otelcol-contrib; if a field name may differ by version, mark it CHECK against the Deployment version. - Never put secrets in YAML. No em dashes. Keep comments short and inside the YAML.
Instructions
Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.
Generated Output
This image was generated using the prompt above.

Examples
Example Input
Deployment: agent plus gateway on Kubernetes; agent as DaemonSet, gateway as Deployment with 3 replicas behind a headless Service otel-gateway-headless; Collector contrib 0.11x Signals: Java and Node services send OTLP gRPC to the node agent; nginx JSON access logs at /var/log/nginx/access.json; host metrics for cpu, memory, filesystem Backends: Grafana Cloud style OTLP HTTP endpoint for all signals, header Authorization from env OTLP_AUTH PiiFields: user.email (delete), enduser.id (hash), http.request.header.authorization (delete) SamplingPolicy: keep every error trace, keep traces slower than 800 ms, 10 percent of everything else MemoryBudget: agent 512Mi, gateway 2Gi Format: topology note, gateway YAML, agent notes, validation
Example Output
Topology: agents receive OTLP, host metrics, and nginx logs, redact PII, and send traces through the loadbalancing exporter keyed on traceID so every span of a trace lands on one of the 3 gateways. Gateways run tail_sampling and export to the backend.
gateway.yaml
extensions:
health_check:
endpoint: 0.0.0.0:13133 # readiness and liveness probe
zpages:
endpoint: 0.0.0.0:55679
receivers:
otlp:
protocols:
grpc:
endpoint: 0.0.0.0:4317
processors:
memory_limiter: # always first
check_interval: 1s
limit_percentage: 80 # of the 2Gi container limit
spike_limit_percentage: 20
attributes/pii: # second line of defense if an agent missed it
actions:
- key: user.email
action: delete
- key: http.request.header.authorization
action: delete
- key: enduser.id
action: hash
tail_sampling:
decision_wait: 10s # longer than the slowest normal trace
num_traces: 50000 # traces held in memory; size against 2Gi, CHECK
policies:
- name: errors
type: status_code
status_code:
status_codes: [ERROR]
- name: slow
type: latency
latency:
threshold_ms: 800
- name: baseline
type: probabilistic
probabilistic:
sampling_percentage: 10
batch: {}
exporters:
otlphttp:
endpoint: https://otlp.example-backend.net/otlp
headers:
Authorization: ${env:OTLP_AUTH}
service:
extensions: [health_check, zpages]
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, attributes/pii, tail_sampling, batch]
exporters: [otlphttp]
metrics:
receivers: [otlp]
processors: [memory_limiter, attributes/pii, batch]
exporters: [otlphttp]
logs:
receivers: [otlp]
processors: [memory_limiter, attributes/pii, batch]
exporters: [otlphttp]
Agent notes (agent.yaml)
- receivers: otlp grpc 0.0.0.0:4317; hostmetrics with collection_interval 30s and scrapers cpu, memory, filesystem; filelog include /var/log/nginx/access.json, start_at end, operator json_parser.
- processors: memory_limiter (limit_percentage 80 of 512Mi), the same attributes/pii block, batch.
- exporters: loadbalancing for traces with routing_key traceID, protocol otlp, resolver dns hostname otel-gateway-headless; otlp to the gateway Service for metrics and logs.
- No tail_sampling on agents: an agent sees only part of each trace.
Validation
1. otelcol-contrib validate --config=gateway.yaml, then the same for agent.yaml.
2. Send one test span with telemetrygen or an app, then open zpages /debug/tracez on a gateway.
3. Confirm user.email is absent on the backend for a test request that set it.
Risks
- Too little memory: memory_limiter starts refusing data first; watch the Collector's refused and dropped span metrics and the exporter queue size.
- num_traces too low for traffic: traces get evicted before decision_wait ends.