🤖 AI Tools
Braintrust Eval Experiment Map from Dataset Inventory (No Invented Scores)
Turn a Braintrust dataset inventory into an eval experiment map only. No invented scores, latency p95, or token counts beyond the inventory.
0Reviews
Prompt
Act as a Braintrust eval engineer who only uses a pasted dataset inventory. You write an eval experiment map the inventory already supports. You do not invent scores, latency p95, token counts, or API keys. This is not a live Braintrust run and not a model benchmark claim. You work only from Inputs. Do not invent stats, citations, quotes, URLs, names, IDs, or records that are not in Inputs. Inputs: - Dataset inventory I lock (experiment stubs, scorer cues, status notes): [Inventory] - Braintrust / SDK version notes I lock: [Version] - Project or workspace label I may quote (or UNKNOWN): [Workspace] - Required experiment or dataset names I may quote (or UNKNOWN): [SpecNames] - Words I must not use: [Banned] - What I must never invent (scores, latency p95, token counts, API keys): [Never] - Output format: [Format] - Language: [Lang] Generate: 1. Honesty ledger: Inventory nouns, Version, Workspace, SpecNames, Lang. Forbidden: invented scores, latency p95, token counts, API keys. Banner: not a live Braintrust run; not a model benchmark claim. 2. Eval experiment map table: one row per Inventory experiment stub or scorer cue. Missing status notes write NOT IN INPUTS. Use Braintrust Dataset, Experiment, Scorer, and Project language when Inventory supports it. Never print Eval() result VALUES not in Inventory. 3. Spec name set: only names in SpecNames. Unnamed experiments stay NOT IN INPUTS. Never print score or token VALUES not in Inventory. 4. Version lock: print Version. Refuse Braintrust APIs newer than Version if Version is named (no unreleased Autoevals claims). 5. Refuse list: inventing scores, inventing latency p95, inventing token counts, inventing API keys. 6. Compliance pass: quote Banned and Never hits. Cut them. Format as Format. Constraints: - Map from Inventory only. No invented score VALUES. Teach Braintrust experiment mapping, not a generic eval swap. - Honor Version. No emojis.
Instructions
Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.
Generated Output
This image was generated using the prompt above.

Examples
Example Input
Inventory: experiment Harbor Quay Golden Set as pasted scorer cue Factuality as pasted status note draft as pasted; experiment River Ops Regression as pasted scorer cue ExactMatch as pasted. Score NOT IN INPUTS. Latency p95 NONE. Version: Braintrust SDK 0.0.177 as pasted (do not invent unreleased Autoevals APIs) Workspace: Harbor Quay Eval Project as pasted SpecNames: Golden Set as pasted; Regression as pasted. Canary UNKNOWN. Banned: guaranteed score, latency invent, token invent, key invent Never: invent scores, invent latency p95, invent token counts, invent API keys Format: ledger + table + specs + version + refuse + compliance Lang: English
Example Output
1. Ledger. Inventory: experiment Harbor Quay Golden Set scorer cue Factuality status note draft; experiment River Ops Regression scorer cue ExactMatch. Score NOT IN INPUTS. Latency p95 NONE. Version Braintrust SDK 0.0.177. Workspace Harbor Quay Eval Project. SpecNames Golden Set; Regression; Canary UNKNOWN. Lang English. Banner: not a live Braintrust run; not a model benchmark claim. Forbidden: invented scores, latency p95, token counts, API keys, guaranteed score, latency invent, token invent, key invent. 2. Eval experiment map. 1. Harbor Quay Golden Set | scorer Factuality as pasted | status draft as pasted | score NOT IN INPUTS | p95 NONE | Braintrust Dataset / Experiment map only. Eval() results not printed. 2. River Ops Regression | scorer ExactMatch as pasted | status NOT IN INPUTS | score NOT IN INPUTS | p95 NONE. Token counts not printed. API keys not invented. 3. Spec name set. Golden Set; Regression as SpecNames. Canary UNKNOWN so write Canary NOT IN INPUTS. No score VALUES printed. No third experiment invented. 4. Version lock. Braintrust SDK 0.0.177 as pasted. Unreleased Autoevals APIs not used. prompt-playground schema NOT IN INPUTS. 5. Refuse. Score invent: refused. Latency invent: refused. Token invent: refused. Key invent: refused. Guaranteed score: refused. 6. Compliance. Banned hits none. Format ledger+table+specs+version+refuse+compliance. Gaps: status for River Ops Regression, Canary decision, dataset split, model name, scorer weights if any. Missing-data policy: if a field was blank, write NOT IN INPUTS rather than guessing. Lock any tool version named in Inputs; if unnamed, write unknown. No invented testimonials, star ratings, or press logos. If legal, clinical, insurance, HR, education-plan, or veterinary content appears, add a one-line not-advice and de-identify banner. Quote banned-word hits and cut them. End with a gaps list of five bullets the user still owes you. Character and byte caps in the job are hard; print counts when relevant. Refuse to backfill DOIs, exam dumps, PHI, PII, or compensation promises not in Inputs.