🤖 AI Tools
Braintrust Eval Run Checklist from Experiment Notes (No Invented Score Averages)
Compile a Braintrust eval-run checklist from pasted experiment notes only. No invented score averages, latency ranks, or accuracy scoreboards. Not a live Braintrust sync.
0Reviews
Prompt
Act as a Braintrust eval-run checklist engineer who only uses pasted experiment notes. You compile an eval-run checklist the notes already support. You do not invent score averages, latency ranks, accuracy scoreboards, or SLA guarantees. This is not a live Braintrust sync, not LangSmith merge, and not MLOps capacity advice. You work only from Inputs. Do not invent stats, citations, quotes, URLs, names, IDs, or records that are not in Inputs. Inputs: - Experiment notes I lock (run stubs, scorer cues, dataset fragments): [ExperimentNotes] - Braintrust project or product notes I lock: [Version] - Project or experiment label I may quote (or UNKNOWN): [ProjectLabel] - Eval run names already present (or UNKNOWN): [RunNames] - Scorer cues already present (or UNKNOWN): [ScorerCues] - Dataset cues already present (or UNKNOWN): [DatasetCues] - Model cues already present (or UNKNOWN): [ModelCues] - Words I must not use: [Banned] - What I must never invent (score averages, latency ranks, accuracy scoreboards, SLA guarantees): [Never] - Output format: [Format] - Language: [Lang] Generate: 1. Honesty ledger: ExperimentNotes nouns, Version, ProjectLabel, RunNames, ScorerCues, DatasetCues, ModelCues, Lang. Banner: not MLOps capacity advice; not a live Braintrust sync. Forbidden: invented score averages, latency ranks, accuracy scoreboards, SLA guarantees. 2. Eval-run checklist: one checkbox row per RunNames entry. Attach only ScorerCues named beside that run in ExperimentNotes. Missing cue write NOT IN INPUTS. 3. Dataset sketch: for each DatasetCues entry, list runs that name it. Do not invent a 0.91 average claim if absent. 4. Model caution block: quote ModelCues only. Prompt packs not in ExperimentNotes stay NOT IN INPUTS. 5. Refuse list: inventing 0.91 score averages, inventing latency ranks, inventing accuracy scoreboards, inventing SLA guarantees. 6. Compliance pass: quote Banned and Never hits. Cut them. Print run and scorer counts from ExperimentNotes only. Format as Format. Constraints: - Eval-run checklist from ExperimentNotes only. No invented score averages. - Honor Version. No emojis. Not a live Braintrust dashboard. Not MLOps capacity advice.
Instructions
Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.
Generated Output
This image was generated using the prompt above.

Examples
Example Input
ExperimentNotes: eval run Harbor Pier Judge v2 as pasted scorer cue factuality rubric as pasted dataset cue pier-faq-50 as pasted; eval run Quay Pier Judge smoke as pasted scorer cue toxicity gate as pasted. Score invent NONE. Latency invent NONE. Version: Braintrust as pasted (do not invent unreleased AI eval coach) ProjectLabel: Cedar Pier eval project as pasted RunNames: Harbor Pier Judge v2 as pasted; Quay Pier Judge smoke as pasted. Nightly regression run UNKNOWN. ScorerCues: scorer cue factuality rubric as pasted for Harbor Pier Judge v2; scorer cue toxicity gate as pasted for Quay Pier Judge smoke. LLM-as-judge pack UNKNOWN. DatasetCues: dataset cue pier-faq-50 as pasted for Harbor Pier Judge v2; Quay Pier Judge smoke dataset NOT IN INPUTS. ModelCues: model cue gpt-class as pasted for Harbor Pier Judge v2. Prompt pack UNKNOWN. Banned: 0.91 average, latency rank #1, guaranteed accuracy scoreboard Never: invent score averages, invent latency ranks, invent accuracy scoreboards, invent SLA guarantees Format: ledger + eval-run checklist + dataset sketch + model caution + refuse + compliance Lang: English
Example Output
1. Ledger. ExperimentNotes: eval run Harbor Pier Judge v2 + scorer cue factuality rubric + dataset cue pier-faq-50; eval run Quay Pier Judge smoke + scorer cue toxicity gate. Score invent NONE. Latency invent NONE. Version Braintrust. ProjectLabel Cedar Pier eval project. RunNames Harbor Pier Judge v2; Quay Pier Judge smoke. ScorerCues factuality rubric for Harbor Pier Judge v2; toxicity gate for Quay Pier Judge smoke. DatasetCues pier-faq-50 for Harbor Pier Judge v2; Quay Pier Judge smoke dataset NOT IN INPUTS. ModelCues gpt-class for Harbor Pier Judge v2. Prompt pack UNKNOWN. Lang English. Banner: not MLOps capacity advice; not a live Braintrust sync. Forbidden: invented score averages, latency ranks, accuracy scoreboards, SLA guarantees, 0.91 average, latency rank #1, guaranteed accuracy scoreboard. 2. Eval-run checklist. [ ] Harbor Pier Judge v2 | scorer factuality rubric as pasted. [ ] Quay Pier Judge smoke | scorer toxicity gate as pasted. LLM-as-judge pack not attached. Nightly regression run not added. 3. Dataset sketch. dataset cue pier-faq-50 | run Harbor Pier Judge v2 as pasted. Quay Pier Judge smoke dataset | NOT IN INPUTS. Score averages NOT IN INPUTS so do not invent 0.91 average. Second pier-faq-50 cue not invented. 4. Model caution. model cue gpt-class as pasted for Harbor Pier Judge v2. Prompt pack UNKNOWN. Do not invent SLA guarantee packs. 5. Refuse. 0.91 score averages: refused. latency ranks: refused. accuracy scoreboards: refused. SLA guarantees: refused. Unreleased AI eval coach: refused. 6. Compliance. Banned hits none. Runs 2. Scorer cues named 2. Format ledger+eval-run checklist+dataset sketch+model caution+refuse+compliance. Gaps: Quay Pier Judge smoke dataset, LLM-as-judge pack, Nightly regression run, Prompt pack, score averages. Missing-data policy: if a field was blank, write NOT IN INPUTS rather than guessing. Lock any tool version named in Inputs; if unnamed, write unknown. No invented testimonials, star ratings, or press logos. If legal, clinical, insurance, HR, education-plan, or veterinary content appears, add a one-line not-advice and de-identify banner. Quote banned-word hits and cut them. End with a gaps list of five bullets the user still owes you. Character and byte caps in the job are hard; print counts when relevant. Refuse to backfill DOIs, exam dumps, PHI, PII, or compensation promises not in Inputs.