🤖 AI Tools
LangSmith Eval Dataset Checklist from Trace Notes (No Invented Score Totals)
Compile a LangSmith eval-dataset checklist from pasted trace notes only. No invented score totals, latency ranks, or token scoreboards. Not a live LangSmith sync.
0Reviews
Prompt
Act as a LangSmith eval-dataset checklist engineer who only uses pasted trace notes. You compile an eval-dataset checklist the notes already support. You do not invent score totals, latency ranks, token scoreboards, or pass-rate guarantees. This is not a live LangSmith sync, not Langfuse merge, and not ML capacity advice. You work only from Inputs. Do not invent stats, citations, quotes, URLs, names, IDs, or records that are not in Inputs. Inputs: - Trace notes I lock (dataset stubs, example cues, scorer fragments): [TraceNotes] - LangSmith project or version notes I lock: [Version] - Project or desk label I may quote (or UNKNOWN): [ProjectLabel] - Dataset labels already present (or UNKNOWN): [DatasetLabels] - Example cues already present (or UNKNOWN): [ExampleCues] - Scorer cues already present (or UNKNOWN): [ScorerCues] - Split cues already present (or UNKNOWN): [SplitCues] - Words I must not use: [Banned] - What I must never invent (score totals, latency ranks, token scoreboards, pass-rate guarantees): [Never] - Output format: [Format] - Language: [Lang] Generate: 1. Honesty ledger: TraceNotes nouns, Version, ProjectLabel, DatasetLabels, ExampleCues, ScorerCues, SplitCues, Lang. Banner: not ML capacity advice; not a live LangSmith sync. Forbidden: invented score totals, latency ranks, token scoreboards, pass-rate guarantees. 2. Eval-dataset checklist: one checkbox row per DatasetLabels entry. Attach only ExampleCues named beside that dataset in TraceNotes. Missing cue write NOT IN INPUTS. 3. Scorer sketch: for each ScorerCues entry, list datasets that name it. Do not invent a 0.94 score claim if absent. 4. Split caution block: quote SplitCues only. Holdout packs not in TraceNotes stay NOT IN INPUTS. 5. Refuse list: inventing 0.94 score totals, inventing latency ranks, inventing token scoreboards, inventing pass-rate guarantees. 6. Compliance pass: quote Banned and Never hits. Cut them. Print dataset and example counts from TraceNotes only. Format as Format. Constraints: - Eval-dataset checklist from TraceNotes only. No invented score totals or latency ranks. - Honor Version. No emojis. Not a live LangSmith dashboard. Not ML capacity advice.
Instructions
Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.
Generated Output
This image was generated using the prompt above.

Examples
Example Input
TraceNotes: dataset label Harbor Quay FAQ Eval as pasted example cue refund question row as pasted scorer cue exact match note as pasted; dataset label Pier Deck Tool Call Eval as pasted example cue calendar tool row as pasted. Score invent NONE. Latency invent NONE. Version: LangSmith as pasted (do not invent unreleased AI auto-scorer) ProjectLabel: Cedar Pier LLM desk as pasted DatasetLabels: Harbor Quay FAQ Eval as pasted; Pier Deck Tool Call Eval as pasted. Safety dataset UNKNOWN. ExampleCues: example cue refund question row as pasted for Harbor Quay FAQ Eval; example cue calendar tool row as pasted for Pier Deck Tool Call Eval. Multilingual row UNKNOWN. ScorerCues: scorer cue exact match note as pasted for Harbor Quay FAQ Eval. Pier Deck Tool Call Eval scorer NOT IN INPUTS. SplitCues: split cue 80 20 holdout as pasted for Harbor Quay FAQ Eval. Pier Deck Tool Call Eval split NOT IN INPUTS. Banned: 0.94 score, latency rank #1, guaranteed token board Never: invent score totals, invent latency ranks, invent token scoreboards, invent pass-rate guarantees Format: ledger + eval-dataset checklist + scorer sketch + split caution + refuse + compliance Lang: English
Example Output
1. Ledger. TraceNotes: dataset label Harbor Quay FAQ Eval + example cue refund question row + scorer cue exact match note; dataset label Pier Deck Tool Call Eval + example cue calendar tool row. Score invent NONE. Latency invent NONE. Version LangSmith. ProjectLabel Cedar Pier LLM desk. DatasetLabels Harbor Quay FAQ Eval; Pier Deck Tool Call Eval. ExampleCues refund question row for Harbor Quay FAQ Eval; calendar tool row for Pier Deck Tool Call Eval. ScorerCues exact match note for Harbor Quay FAQ Eval. Pier Deck Tool Call Eval scorer NOT IN INPUTS. SplitCues 80 20 holdout for Harbor Quay FAQ Eval; Pier Deck Tool Call Eval split NOT IN INPUTS. Lang English. Banner: not ML capacity advice; not a live LangSmith sync. Forbidden: invented score totals, latency ranks, token scoreboards, pass-rate guarantees, 0.94 score, latency rank #1, guaranteed token board. 2. Eval-dataset checklist. [ ] Harbor Quay FAQ Eval | example refund question row as pasted | scorer exact match note as pasted. [ ] Pier Deck Tool Call Eval | example calendar tool row as pasted | scorer NOT IN INPUTS. Multilingual row not attached. Safety dataset not added. 3. Scorer sketch. scorer cue exact match note | dataset Harbor Quay FAQ Eval as pasted. Pier Deck Tool Call Eval scorer NOT IN INPUTS so do not invent a second scorer. Score totals NOT IN INPUTS so do not invent 0.94 score. 4. Split caution. split cue 80 20 holdout as pasted for Harbor Quay FAQ Eval. Pier Deck Tool Call Eval split NOT IN INPUTS. Do not invent holdout packs. 5. Refuse. 0.94 score totals: refused. latency ranks: refused. token scoreboards: refused. pass-rate guarantees: refused. Unreleased AI auto-scorer: refused. 6. Compliance. Banned hits none. Datasets 2. Example cues named 2. Format ledger+eval-dataset checklist+scorer sketch+split caution+refuse+compliance. Gaps: Pier Deck Tool Call Eval scorer, Pier Deck Tool Call Eval split, Multilingual row, Safety dataset, Score totals. Missing-data policy: if a field was blank, write NOT IN INPUTS rather than guessing. Lock any tool version named in Inputs; if unnamed, write unknown. No invented testimonials, star ratings, or press logos. If legal, clinical, insurance, HR, education-plan, or veterinary content appears, add a one-line not-advice and de-identify banner. Quote banned-word hits and cut them. End with a gaps list of five bullets the user still owes you. Character and byte caps in the job are hard; print counts when relevant. Refuse to backfill DOIs, exam dumps, PHI, PII, or compensation promises not in Inputs.