🤖 AI Tools
LangSmith Run Eval Dataset Row Checklist from Trace Notes (No Invented Latency Scores)
Compile a LangSmith run eval dataset-row checklist from pasted trace notes only. No invented latency scores, token ranks, or eval scoreboards. Not a live LangSmith sync.
0Reviews
Prompt
Act as a LangSmith run eval dataset-row checklist engineer who only uses pasted trace notes. You compile a run eval dataset-row checklist the notes already support. You do not invent latency scores, token ranks, eval scoreboards, or cost claims. This is not a live LangSmith sync, not a Weights and Biases invent, and not ML ops consulting advice. You work only from Inputs. Do not invent stats, citations, quotes, URLs, names, IDs, or records that are not in Inputs. Inputs: - Trace notes I lock (run stubs, eval cues, dataset-row fragments): [TraceNotes] - LangSmith project or version notes I lock: [Version] - Project or workspace label I may quote (or UNKNOWN): [ProjectLabel] - Run titles already present (or UNKNOWN): [RunTitles] - Eval cues already present (or UNKNOWN): [EvalCues] - Dataset-row cues already present (or UNKNOWN): [DatasetRowCues] - Feedback cues already present (or UNKNOWN): [FeedbackCues] - Words I must not use: [Banned] - What I must never invent (latency scores, token ranks, eval scoreboards, cost claims): [Never] - Output format: [Format] - Language: [Lang] Generate: 1. Honesty ledger: TraceNotes nouns, Version, ProjectLabel, RunTitles, EvalCues, DatasetRowCues, FeedbackCues, Lang. Banner: not ML ops consulting advice; not a live LangSmith sync. Forbidden: invented latency scores, token ranks, eval scoreboards, cost claims. 2. Run eval dataset-row checklist: one checkbox row per RunTitles entry. Attach only EvalCues and DatasetRowCues named beside that run in TraceNotes. Missing eval or dataset-row write NOT IN INPUTS. 3. Feedback sketch: for each FeedbackCues entry, list runs that name it. Do not invent a 240ms latency-score claim if absent. 4. Project caution block: quote ProjectLabel and Version only. Prompt-hub packs not in TraceNotes stay NOT IN INPUTS. 5. Refuse list: inventing 240ms latency scores, inventing token ranks, inventing eval scoreboards, inventing cost claims. 6. Compliance pass: quote Banned and Never hits. Cut them. Print run and eval counts from TraceNotes only. Format as Format. Constraints: - Run eval dataset-row checklist from TraceNotes only. No invented latency scores. - Honor Version. No emojis. Not a live LangSmith dashboard. Not ML ops consulting advice.
Instructions
Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.
Generated Output
This image was generated using the prompt above.

Examples
Example Input
TraceNotes: run title Harbor Quay Chain as pasted eval cue faithfulness as pasted dataset-row cue row-hq-01 as pasted; run title Quay Storm Retriever as pasted eval cue groundedness as pasted. Latency invent NONE. Token invent NONE. Version: LangSmith as pasted (do not invent unreleased AI eval coach) ProjectLabel: Cedar Pier rag project as pasted RunTitles: Harbor Quay Chain as pasted; Quay Storm Retriever as pasted. Annotation pack UNKNOWN. EvalCues: eval cue faithfulness as pasted for Harbor Quay Chain; eval cue groundedness as pasted for Quay Storm Retriever. Judge pack UNKNOWN. DatasetRowCues: dataset-row cue row-hq-01 as pasted for Harbor Quay Chain; Quay Storm Retriever dataset-row NOT IN INPUTS. FeedbackCues: feedback cue thumbs-up as pasted for Harbor Quay Chain. Experiment pack UNKNOWN. Banned: 240ms latency, token rank #1, guaranteed eval scoreboard Never: invent latency scores, invent token ranks, invent eval scoreboards, invent cost claims Format: ledger + run eval dataset-row checklist + feedback sketch + project caution + refuse + compliance Lang: English
Example Output
1. Ledger. TraceNotes: run title Harbor Quay Chain + eval cue faithfulness + dataset-row cue row-hq-01; run title Quay Storm Retriever + eval cue groundedness. Latency invent NONE. Token invent NONE. Version LangSmith. ProjectLabel Cedar Pier rag project. RunTitles Harbor Quay Chain; Quay Storm Retriever. EvalCues faithfulness for Harbor Quay Chain; groundedness for Quay Storm Retriever. DatasetRowCues row-hq-01 for Harbor Quay Chain; Quay Storm Retriever dataset-row NOT IN INPUTS. FeedbackCues thumbs-up for Harbor Quay Chain. Experiment pack UNKNOWN. Lang English. Banner: not ML ops consulting advice; not a live LangSmith sync. Forbidden: invented latency scores, token ranks, eval scoreboards, cost claims, 240ms latency, token rank #1, guaranteed eval scoreboard. 2. Run eval dataset-row checklist. [ ] Harbor Quay Chain | eval faithfulness as pasted | dataset-row row-hq-01 as pasted. [ ] Quay Storm Retriever | eval groundedness as pasted | dataset-row NOT IN INPUTS. Judge pack not attached. Annotation pack not added. 3. Feedback sketch. feedback cue thumbs-up | run Harbor Quay Chain as pasted. Quay Storm Retriever feedback | NOT IN INPUTS. Latency scores NOT IN INPUTS so do not invent 240ms latency. Second feedback cue not invented. 4. Project caution. ProjectLabel Cedar Pier rag project. Version LangSmith. Experiment pack UNKNOWN. Do not invent cost claim packs. 5. Refuse. 240ms latency scores: refused. token ranks: refused. eval scoreboards: refused. cost claims: refused. Unreleased AI eval coach: refused. 6. Compliance. Banned hits none. Runs 2. Evals named 2. Format ledger+run eval dataset-row checklist+feedback sketch+project caution+refuse+compliance. Gaps: Quay Storm Retriever dataset-row, Quay Storm Retriever feedback, Judge pack, Annotation pack, Experiment pack. Missing-data policy: if a field was blank, write NOT IN INPUTS rather than guessing. Lock any tool version named in Inputs; if unnamed, write unknown. No invented testimonials, star ratings, or press logos. If legal, clinical, insurance, HR, education-plan, or veterinary content appears, add a one-line not-advice and de-identify banner. Quote banned-word hits and cut them. End with a gaps list of five bullets the user still owes you. Character and byte caps in the job are hard; print counts when relevant. Refuse to backfill DOIs, exam dumps, PHI, PII, or compensation promises not in Inputs.