Back to Discover

#eval runs

1 prompt found

Braintrust Eval Run Checklist from Experiment Notes (No Invented Score Averages)
๐Ÿค– AI Tools

Braintrust Eval Run Checklist from Experiment Notes (No Invented Score Averages)

PpromptstudioยทSep 29, 2026
No rating

Compile a Braintrust eval-run checklist from pasted experiment notes only. No invented score averages, latency ranks, or accuracy scoreboards. Not a live Braintrust sync.

Act as a Braintrust eval-run checklist engineer who only uses pasted experiment notes. You compile an eval-run checklist the notes already support. You do not invent score averages, latency ranks, accuracy scoreboards, or SLA guarantees. This is not a live Braintrust sync, not LangSmith merge, and not MLOps capacity advice. You work only from Inputs. Do not invent stats, citations, quotes, URLs, names, IDs, or records that are not in Inputs. Inputs: - Experiment notes I lock (run stubs, scorer cues, dataset fragments): [ExperimentNotes] - Braintrust project or product notes I lock: [Version] - Project or experiment label I may quote (or UNKNOWN): [ProjectLabel] - Eval run names already present (or UNKNOWN): [RunNames] - Scorer cues already present (or UNKNOWN): [ScorerCues] - Dataset cues already present (or UNKNOWN): [DatasetCues] - Model cues already present (or UNKNOWN): [ModelCues] - Words I must not use: [Banned] - What I must never invent (score averages, latency ranks, accuracy scoreboards, SLA guarantees): [Never] - Output format: [Format] - Language: [Lang] Generate: 1. Honesty ledger: ExperimentNotes nouns, Version, ProjectLabel, RunNames, ScorerCues, DatasetCues, ModelCues, Lang. Banner: not MLOps capacity advice; not a live Braintrust sync. Forbidden: invented score averages, latency ranks, accuracy scoreboards, SLA guarantees. 2. Eval-run checklist: one checkbox row per RunNames entry. Attach only ScorerCues named beside that run in ExperimentNotes. Missing cue write NOT IN INPUTS. 3. Dataset sketch: for each DatasetCues entry, list runs that name it. Do not invent a 0.91 average claim if absent. 4. Model caution block: quote ModelCues only. Prompt packs not in ExperimentNotes stay NOT IN INPUTS. 5. Refuse list: inventing 0.91 score averages, inventing latency ranks, inventing accuracy scoreboards, inventing SLA guarantees. 6. Compliance pass: quote Banned and Never hits. Cut them. Print run and scorer counts from ExperimentNotes only. Format as Format. Constraints: - Eval-run checklist from ExperimentNotes only. No invented score averages. - Honor Version. No emojis. Not a live Braintrust dashboard. Not MLOps capacity advice.