🤖 AI Tools

Langfuse Trace Score Checklist from Observability Notes (No Invented Token Totals)

Compile a Langfuse trace-score checklist from pasted observability notes only. No invented token totals, latency ranks, or cost scoreboards. Not a live Langfuse sync.

0.0
0Reviews
P
September 28, 2026

Prompt

Act as a Langfuse trace-score checklist coordinator who only uses pasted observability notes. You compile a trace-score checklist the notes already support. You do not invent token totals, latency ranks, cost scoreboards, or accuracy percent claims. This is not a live Langfuse sync, not Helicone cost merge, and not observability scoring advice.
You work only from Inputs. Do not invent stats, citations, quotes, URLs, names, IDs, or records that are not in Inputs.

Inputs:
- Observability notes I lock (trace stubs, score cues, session fragments): [ObservabilityNotes]
- Langfuse project or SDK version notes I lock: [Version]
- Project or env label I may quote (or UNKNOWN): [ProjectLabel]
- Trace labels already present (or UNKNOWN): [TraceLabels]
- Score cues already present (or UNKNOWN): [ScoreCues]
- Session cues already present (or UNKNOWN): [SessionCues]
- Eval cues already present (or UNKNOWN): [EvalCues]
- Words I must not use: [Banned]
- What I must never invent (token totals, latency ranks, cost scoreboards, accuracy percent claims): [Never]
- Output format: [Format]
- Language: [Lang]

Generate:
1. Honesty ledger: ObservabilityNotes nouns, Version, ProjectLabel, TraceLabels, ScoreCues, SessionCues, EvalCues, Lang. Banner: not observability scoring advice; not a live Langfuse sync. Forbidden: invented token totals, latency ranks, cost scoreboards, accuracy percent claims.
2. Trace-score checklist: one checkbox row per TraceLabels entry. Attach only ScoreCues and SessionCues named beside that trace in ObservabilityNotes. Missing cue write NOT IN INPUTS.
3. Eval sketch: for each EvalCues entry, list traces that name it. Do not invent an 18k token-total claim if absent.
4. Session caution block: quote SessionCues only. Helicone packs not in ObservabilityNotes stay NOT IN INPUTS.
5. Refuse list: inventing 18k token totals, inventing p99 latency ranks, inventing cost scoreboards, inventing accuracy percent claims.
6. Compliance pass: quote Banned and Never hits. Cut them. Print trace and score counts from ObservabilityNotes only. Format as Format.

Constraints:
- Trace-score checklist from ObservabilityNotes only. No invented token totals.
- Honor Version. No emojis. Not a live Langfuse dashboard. Not observability scoring advice.

Instructions

Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.

Generated Output

This image was generated using the prompt above.

Langfuse Trace Score Checklist from Observability Notes (No Invented Token Totals) - Result

Examples

Example Input

ObservabilityNotes: trace label Harbor Chat Turn as pasted score Relevance as pasted session Pilot A as pasted; trace label Quay Tool Call as pasted score ToolSuccess as pasted. Token invent NONE. Latency invent NONE.
Version: Langfuse as pasted (do not invent unreleased AI cost coach)
ProjectLabel: Cedar Pier eval project as pasted
TraceLabels: Harbor Chat Turn as pasted; Quay Tool Call as pasted. Background Rerank UNKNOWN.
ScoreCues: score Relevance as pasted for Harbor Chat Turn; score ToolSuccess as pasted for Quay Tool Call. Toxicity pack UNKNOWN.
SessionCues: session Pilot A as pasted for Harbor Chat Turn; Quay Tool Call session NOT IN INPUTS.
EvalCues: eval HumanRubric as pasted for Harbor Chat Turn; Quay Tool Call eval NOT IN INPUTS.
Banned: 18k tokens, p99 latency rank #1, guaranteed accuracy percent
Never: invent token totals, invent latency ranks, invent cost scoreboards, invent accuracy percent claims
Format: ledger + trace-score checklist + eval sketch + session caution + refuse + compliance
Lang: English

Example Output

1. Ledger. ObservabilityNotes: trace label Harbor Chat Turn + score Relevance + session Pilot A; trace label Quay Tool Call + score ToolSuccess. Token invent NONE. Latency invent NONE. Version Langfuse. ProjectLabel Cedar Pier eval project. TraceLabels Harbor Chat Turn; Quay Tool Call. ScoreCues Relevance for Harbor Chat Turn; ToolSuccess for Quay Tool Call. SessionCues Pilot A for Harbor Chat Turn; Quay Tool Call session NOT IN INPUTS. EvalCues HumanRubric for Harbor Chat Turn; Quay Tool Call eval NOT IN INPUTS. Lang English. Banner: not observability scoring advice; not a live Langfuse sync. Forbidden: invented token totals, latency ranks, cost scoreboards, accuracy percent claims, 18k tokens, p99 latency rank #1, guaranteed accuracy percent.

2. Trace-score checklist.
[ ] Harbor Chat Turn | score Relevance | session Pilot A as pasted.
[ ] Quay Tool Call | score ToolSuccess | session NOT IN INPUTS.
Background Rerank not attached. Toxicity pack not added.

3. Eval sketch.
eval HumanRubric | trace Harbor Chat Turn as pasted.
Quay Tool Call eval | NOT IN INPUTS.
Token totals NOT IN INPUTS so do not invent 18k tokens. Second HumanRubric cue not invented.

4. Session caution. session Pilot A as pasted for Harbor Chat Turn. Quay Tool Call session NOT IN INPUTS. Helicone pack NOT IN INPUTS. Do not invent cost packs.

5. Refuse. 18k token totals: refused. p99 latency ranks: refused. cost scoreboards: refused. accuracy percent claims: refused. Unreleased AI cost coach: refused.

6. Compliance. Banned hits none. Traces 2. Scores named 2. Format ledger+trace-score checklist+eval sketch+session caution+refuse+compliance. Gaps: Quay Tool Call session, Quay Tool Call eval, Background Rerank, Toxicity pack, token totals.

Missing-data policy: if a field was blank, write NOT IN INPUTS rather than guessing. Lock any tool version named in Inputs; if unnamed, write unknown. No invented testimonials, star ratings, or press logos. If legal, clinical, insurance, HR, education-plan, or veterinary content appears, add a one-line not-advice and de-identify banner. Quote banned-word hits and cut them. End with a gaps list of five bullets the user still owes you. Character and byte caps in the job are hard; print counts when relevant. Refuse to backfill DOIs, exam dumps, PHI, PII, or compensation promises not in Inputs.

Reviews (0)

Please login to leave a review.
Loading reviews...