🤖 AI Tools
Langfuse Trace Scorecard from Experiment Notes (No Invented Accuracy Percent)
Compile a Langfuse trace scorecard from pasted experiment notes only. No invented accuracy percentages, F1 scores, or latency percentiles.
0Reviews
Prompt
Act as a Langfuse eval engineer who only uses pasted experiment notes. You compile a trace scorecard the notes already support. You do not invent accuracy percentages, F1 scores, latency percentiles, or leaderboard ranks. This is not a live Langfuse cloud sync and not a model-card certificate. You work only from Inputs. Do not invent stats, citations, quotes, URLs, names, IDs, or records that are not in Inputs. Inputs: - Experiment notes I lock (trace stubs, score cues, experiment labels): [ExperimentNotes] - Langfuse version or project notes I lock: [Version] - Project or experiment label I may quote (or UNKNOWN): [ExperimentLabel] - Trace names already present (or UNKNOWN): [TraceNames] - Score names already present (or UNKNOWN): [ScoreNames] - Session names already present (or UNKNOWN): [SessionNames] - Model or prompt cues already present (or UNKNOWN): [ModelCues] - Words I must not use: [Banned] - What I must never invent (accuracy percentages, F1 scores, latency percentiles, leaderboard ranks): [Never] - Output format: [Format] - Language: [Lang] Generate: 1. Honesty ledger: ExperimentNotes nouns, Version, ExperimentLabel, TraceNames, ScoreNames, SessionNames, ModelCues, Lang. Forbidden: invented accuracy percentages, F1 scores, latency percentiles, leaderboard ranks. 2. Trace scorecard checklist: one checkbox row per TraceNames entry. Attach only ScoreNames named beside that trace in ExperimentNotes. Missing score write NOT IN INPUTS. 3. Session sketch: for each SessionNames entry, list traces that name it. Do not invent a 91.8% accuracy if absent. 4. Model caution block: quote ModelCues only. Benchmark packs not in ExperimentNotes stay NOT IN INPUTS. 5. Refuse list: inventing 91.8% accuracy, inventing F1 0.88, inventing p95 310ms latency, inventing leaderboard rank #1. 6. Compliance pass: quote Banned and Never hits. Cut them. Print trace and score counts from ExperimentNotes only. Format as Format. Constraints: - Trace scorecard from ExperimentNotes only. No invented accuracy percentages. - Honor Version. No emojis. Not a live Langfuse console.
Instructions
Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.
Generated Output
This image was generated using the prompt above.

Examples
Example Input
ExperimentNotes: trace Harbor Quay Chat as pasted score Exact Match as pasted; trace Pier Edge Toolcall as pasted score Rubric Score as pasted. Accuracy invent NONE. F1 invent NONE. Version: Langfuse cloud as pasted (do not invent unreleased AI accuracy coach) ExperimentLabel: Harbor Quay eval suite as pasted TraceNames: Harbor Quay Chat as pasted; Pier Edge Toolcall as pasted. Cedar Smoke UNKNOWN. ScoreNames: score Exact Match as pasted; score Rubric Score as pasted. Midpoint BLEU UNKNOWN. SessionNames: session Quay v4 as pasted for Harbor Quay Chat; Pier Edge Toolcall session NOT IN INPUTS. ModelCues: model cue gpt-4.1-mini as pasted for Harbor Quay Chat. Full benchmark pack UNKNOWN. Banned: 91.8% accuracy, F1 0.88, guaranteed p95 310ms Never: invent accuracy percentages, invent F1 scores, invent latency percentiles, invent leaderboard ranks Format: ledger + trace scorecard + session sketch + model caution + refuse + compliance Lang: English
Example Output
1. Ledger. ExperimentNotes: trace Harbor Quay Chat + score Exact Match; trace Pier Edge Toolcall + score Rubric Score. Accuracy invent NONE. F1 invent NONE. Version Langfuse cloud. ExperimentLabel Harbor Quay eval suite. TraceNames Harbor Quay Chat; Pier Edge Toolcall. ScoreNames Exact Match; Rubric Score. SessionNames Quay v4 for Harbor Quay Chat; Pier Edge Toolcall session NOT IN INPUTS. ModelCues gpt-4.1-mini for Harbor Quay Chat. Full benchmark pack UNKNOWN. Lang English. Forbidden: invented accuracy percentages, F1 scores, latency percentiles, leaderboard ranks, 91.8% accuracy, F1 0.88, guaranteed p95 310ms. 2. Trace scorecard checklist. [ ] Harbor Quay Chat | score Exact Match as pasted. [ ] Pier Edge Toolcall | score Rubric Score as pasted. Midpoint BLEU not attached. Cedar Smoke not added. 3. Session sketch. session Quay v4 | trace Harbor Quay Chat as pasted. Pier Edge Toolcall session | NOT IN INPUTS. Accuracy NOT IN INPUTS so do not invent 91.8% accuracy. Second Harbor Quay Chat cue not invented. 4. Model caution. gpt-4.1-mini as pasted for Harbor Quay Chat. Full benchmark pack UNKNOWN. Benchmark pack NOT IN INPUTS. Do not invent leaderboard packs. 5. Refuse. 91.8% accuracy: refused. F1 0.88: refused. p95 310ms latency: refused. leaderboard rank #1: refused. Unreleased AI accuracy coach: refused. 6. Compliance. Banned hits none. Traces 2. Scores 2. Format ledger+trace scorecard+session sketch+model caution+refuse+compliance. Gaps: Pier Edge Toolcall session, Cedar Smoke, Midpoint BLEU, Full benchmark pack, accuracy percent. Missing-data policy: if a field was blank, write NOT IN INPUTS rather than guessing. Lock any tool version named in Inputs; if unnamed, write unknown. No invented testimonials, star ratings, or press logos. If legal, clinical, insurance, HR, education-plan, or veterinary content appears, add a one-line not-advice and de-identify banner. Quote banned-word hits and cut them. End with a gaps list of five bullets the user still owes you. Character and byte caps in the job are hard; print counts when relevant. Refuse to backfill DOIs, exam dumps, PHI, PII, or compensation promises not in Inputs.