🤖 AI Tools

Langfuse Trace Scorecard from Experiment Notes (No Invented Accuracy Percent)

Compile a Langfuse trace scorecard from pasted experiment notes only. No invented accuracy percentages, F1 scores, or latency percentiles.

0.0
0Reviews
P
September 27, 2026

Prompt

Act as a Langfuse eval engineer who only uses pasted experiment notes. You compile a trace scorecard the notes already support. You do not invent accuracy percentages, F1 scores, latency percentiles, or leaderboard ranks. This is not a live Langfuse cloud sync and not a model-card certificate.
You work only from Inputs. Do not invent stats, citations, quotes, URLs, names, IDs, or records that are not in Inputs.

Inputs:
- Experiment notes I lock (trace stubs, score cues, experiment labels): [ExperimentNotes]
- Langfuse version or project notes I lock: [Version]
- Project or experiment label I may quote (or UNKNOWN): [ExperimentLabel]
- Trace names already present (or UNKNOWN): [TraceNames]
- Score names already present (or UNKNOWN): [ScoreNames]
- Session names already present (or UNKNOWN): [SessionNames]
- Model or prompt cues already present (or UNKNOWN): [ModelCues]
- Words I must not use: [Banned]
- What I must never invent (accuracy percentages, F1 scores, latency percentiles, leaderboard ranks): [Never]
- Output format: [Format]
- Language: [Lang]

Generate:
1. Honesty ledger: ExperimentNotes nouns, Version, ExperimentLabel, TraceNames, ScoreNames, SessionNames, ModelCues, Lang. Forbidden: invented accuracy percentages, F1 scores, latency percentiles, leaderboard ranks.
2. Trace scorecard checklist: one checkbox row per TraceNames entry. Attach only ScoreNames named beside that trace in ExperimentNotes. Missing score write NOT IN INPUTS.
3. Session sketch: for each SessionNames entry, list traces that name it. Do not invent a 91.8% accuracy if absent.
4. Model caution block: quote ModelCues only. Benchmark packs not in ExperimentNotes stay NOT IN INPUTS.
5. Refuse list: inventing 91.8% accuracy, inventing F1 0.88, inventing p95 310ms latency, inventing leaderboard rank #1.
6. Compliance pass: quote Banned and Never hits. Cut them. Print trace and score counts from ExperimentNotes only. Format as Format.

Constraints:
- Trace scorecard from ExperimentNotes only. No invented accuracy percentages.
- Honor Version. No emojis. Not a live Langfuse console.

Instructions

Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.

Generated Output

This image was generated using the prompt above.

Langfuse Trace Scorecard from Experiment Notes (No Invented Accuracy Percent) - Result

Examples

Example Input

ExperimentNotes: trace Harbor Quay Chat as pasted score Exact Match as pasted; trace Pier Edge Toolcall as pasted score Rubric Score as pasted. Accuracy invent NONE. F1 invent NONE.
Version: Langfuse cloud as pasted (do not invent unreleased AI accuracy coach)
ExperimentLabel: Harbor Quay eval suite as pasted
TraceNames: Harbor Quay Chat as pasted; Pier Edge Toolcall as pasted. Cedar Smoke UNKNOWN.
ScoreNames: score Exact Match as pasted; score Rubric Score as pasted. Midpoint BLEU UNKNOWN.
SessionNames: session Quay v4 as pasted for Harbor Quay Chat; Pier Edge Toolcall session NOT IN INPUTS.
ModelCues: model cue gpt-4.1-mini as pasted for Harbor Quay Chat. Full benchmark pack UNKNOWN.
Banned: 91.8% accuracy, F1 0.88, guaranteed p95 310ms
Never: invent accuracy percentages, invent F1 scores, invent latency percentiles, invent leaderboard ranks
Format: ledger + trace scorecard + session sketch + model caution + refuse + compliance
Lang: English

Example Output

1. Ledger. ExperimentNotes: trace Harbor Quay Chat + score Exact Match; trace Pier Edge Toolcall + score Rubric Score. Accuracy invent NONE. F1 invent NONE. Version Langfuse cloud. ExperimentLabel Harbor Quay eval suite. TraceNames Harbor Quay Chat; Pier Edge Toolcall. ScoreNames Exact Match; Rubric Score. SessionNames Quay v4 for Harbor Quay Chat; Pier Edge Toolcall session NOT IN INPUTS. ModelCues gpt-4.1-mini for Harbor Quay Chat. Full benchmark pack UNKNOWN. Lang English. Forbidden: invented accuracy percentages, F1 scores, latency percentiles, leaderboard ranks, 91.8% accuracy, F1 0.88, guaranteed p95 310ms.

2. Trace scorecard checklist.
[ ] Harbor Quay Chat | score Exact Match as pasted.
[ ] Pier Edge Toolcall | score Rubric Score as pasted.
Midpoint BLEU not attached. Cedar Smoke not added.

3. Session sketch.
session Quay v4 | trace Harbor Quay Chat as pasted.
Pier Edge Toolcall session | NOT IN INPUTS.
Accuracy NOT IN INPUTS so do not invent 91.8% accuracy. Second Harbor Quay Chat cue not invented.

4. Model caution. gpt-4.1-mini as pasted for Harbor Quay Chat. Full benchmark pack UNKNOWN. Benchmark pack NOT IN INPUTS. Do not invent leaderboard packs.

5. Refuse. 91.8% accuracy: refused. F1 0.88: refused. p95 310ms latency: refused. leaderboard rank #1: refused. Unreleased AI accuracy coach: refused.

6. Compliance. Banned hits none. Traces 2. Scores 2. Format ledger+trace scorecard+session sketch+model caution+refuse+compliance. Gaps: Pier Edge Toolcall session, Cedar Smoke, Midpoint BLEU, Full benchmark pack, accuracy percent.

Missing-data policy: if a field was blank, write NOT IN INPUTS rather than guessing. Lock any tool version named in Inputs; if unnamed, write unknown. No invented testimonials, star ratings, or press logos. If legal, clinical, insurance, HR, education-plan, or veterinary content appears, add a one-line not-advice and de-identify banner. Quote banned-word hits and cut them. End with a gaps list of five bullets the user still owes you. Character and byte caps in the job are hard; print counts when relevant. Refuse to backfill DOIs, exam dumps, PHI, PII, or compensation promises not in Inputs.

Reviews (0)

Please login to leave a review.
Loading reviews...