🤖 AI Tools

Braintrust Eval Dataset Scorecard from Experiment Notes (No Invented Accuracy Percent)

Compile a Braintrust eval dataset scorecard from pasted experiment notes only. No invented accuracy percentages, F1 scores, or latency percentiles.

0.0
0Reviews
P
September 27, 2026

Prompt

Act as a Braintrust eval engineer who only uses pasted experiment notes. You compile an eval dataset scorecard the notes already support. You do not invent accuracy percentages, F1 scores, latency percentiles, or leaderboard ranks. This is not a live Braintrust experiment sync and not a model-card certificate.
You work only from Inputs. Do not invent stats, citations, quotes, URLs, names, IDs, or records that are not in Inputs.

Inputs:
- Experiment notes I lock (dataset stubs, scorer cues, experiment labels): [ExperimentNotes]
- Braintrust version or project notes I lock: [Version]
- Project or experiment label I may quote (or UNKNOWN): [ExperimentLabel]
- Dataset names already present (or UNKNOWN): [DatasetNames]
- Scorer names already present (or UNKNOWN): [ScorerNames]
- Experiment names already present (or UNKNOWN): [ExperimentNames]
- Prompt or model cues already present (or UNKNOWN): [ModelCues]
- Words I must not use: [Banned]
- What I must never invent (accuracy percentages, F1 scores, latency percentiles, leaderboard ranks): [Never]
- Output format: [Format]
- Language: [Lang]

Generate:
1. Honesty ledger: ExperimentNotes nouns, Version, ExperimentLabel, DatasetNames, ScorerNames, ExperimentNames, ModelCues, Lang. Forbidden: invented accuracy percentages, F1 scores, latency percentiles, leaderboard ranks.
2. Eval dataset scorecard checklist: one checkbox row per DatasetNames entry. Attach only ScorerNames named beside that dataset in ExperimentNotes. Missing scorer write NOT IN INPUTS.
3. Experiment sketch: for each ExperimentNames entry, list datasets that name it. Do not invent a 94.2% accuracy if absent.
4. Model caution block: quote ModelCues only. Benchmark packs not in ExperimentNotes stay NOT IN INPUTS.
5. Refuse list: inventing 94.2% accuracy, inventing F1 0.91, inventing p95 220ms latency, inventing leaderboard rank #1.
6. Compliance pass: quote Banned and Never hits. Cut them. Print dataset and scorer counts from ExperimentNotes only. Format as Format.

Constraints:
- Eval dataset scorecard from ExperimentNotes only. No invented accuracy percentages.
- Honor Version. No emojis. Not a live Braintrust console.

Instructions

Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.

Generated Output

This image was generated using the prompt above.

Braintrust Eval Dataset Scorecard from Experiment Notes (No Invented Accuracy Percent) - Result

Examples

Example Input

ExperimentNotes: dataset Harbor Quay Gold as pasted scorer Exact Match as pasted; dataset Pier Edge as pasted scorer Rubric Score as pasted. Accuracy invent NONE. F1 invent NONE.
Version: Braintrust cloud as pasted (do not invent unreleased AI accuracy coach)
ExperimentLabel: Harbor Quay eval suite as pasted
DatasetNames: Harbor Quay Gold as pasted; Pier Edge as pasted. Cedar Smoke UNKNOWN.
ScorerNames: scorer Exact Match as pasted; scorer Rubric Score as pasted. Midpoint BLEU UNKNOWN.
ExperimentNames: experiment Quay v3 as pasted for Harbor Quay Gold; Pier Edge experiment NOT IN INPUTS.
ModelCues: model cue gpt-4.1-mini as pasted for Harbor Quay Gold. Full benchmark pack UNKNOWN.
Banned: 94.2% accuracy, F1 0.91, guaranteed p95 220ms
Never: invent accuracy percentages, invent F1 scores, invent latency percentiles, invent leaderboard ranks
Format: ledger + eval dataset scorecard + experiment sketch + model caution + refuse + compliance
Lang: English

Example Output

1. Ledger. ExperimentNotes: dataset Harbor Quay Gold + scorer Exact Match; dataset Pier Edge + scorer Rubric Score. Accuracy invent NONE. F1 invent NONE. Version Braintrust cloud. ExperimentLabel Harbor Quay eval suite. DatasetNames Harbor Quay Gold; Pier Edge. ScorerNames Exact Match; Rubric Score. ExperimentNames Quay v3 for Harbor Quay Gold; Pier Edge experiment NOT IN INPUTS. ModelCues gpt-4.1-mini for Harbor Quay Gold. Full benchmark pack UNKNOWN. Lang English. Forbidden: invented accuracy percentages, F1 scores, latency percentiles, leaderboard ranks, 94.2% accuracy, F1 0.91, guaranteed p95 220ms.

2. Eval dataset scorecard checklist.
[ ] Harbor Quay Gold | scorer Exact Match as pasted.
[ ] Pier Edge | scorer Rubric Score as pasted.
Midpoint BLEU not attached. Cedar Smoke not added.

3. Experiment sketch.
experiment Quay v3 | dataset Harbor Quay Gold as pasted.
Pier Edge experiment | NOT IN INPUTS.
Accuracy NOT IN INPUTS so do not invent 94.2% accuracy. Second Harbor Quay Gold cue not invented.

4. Model caution. gpt-4.1-mini as pasted for Harbor Quay Gold. Full benchmark pack UNKNOWN. Benchmark pack NOT IN INPUTS. Do not invent leaderboard packs.

5. Refuse. 94.2% accuracy: refused. F1 0.91: refused. p95 220ms latency: refused. leaderboard rank #1: refused. Unreleased AI accuracy coach: refused.

6. Compliance. Banned hits none. Datasets 2. Scorers 2. Format ledger+eval dataset scorecard+experiment sketch+model caution+refuse+compliance. Gaps: Pier Edge experiment, Cedar Smoke, Midpoint BLEU, Full benchmark pack, accuracy percent.

Missing-data policy: if a field was blank, write NOT IN INPUTS rather than guessing. Lock any tool version named in Inputs; if unnamed, write unknown. No invented testimonials, star ratings, or press logos. If legal, clinical, insurance, HR, education-plan, or veterinary content appears, add a one-line not-advice and de-identify banner. Quote banned-word hits and cut them. End with a gaps list of five bullets the user still owes you. Character and byte caps in the job are hard; print counts when relevant. Refuse to backfill DOIs, exam dumps, PHI, PII, or compensation promises not in Inputs.

Reviews (0)

Please login to leave a review.
Loading reviews...