Back to Discover

#run eval dataset row

1 prompt found

LangSmith Run Eval Dataset Row Checklist from Trace Notes (No Invented Latency Scores)
๐Ÿค– AI Tools

LangSmith Run Eval Dataset Row Checklist from Trace Notes (No Invented Latency Scores)

PpromptstudioยทOct 3, 2026
No rating

Compile a LangSmith run eval dataset-row checklist from pasted trace notes only. No invented latency scores, token ranks, or eval scoreboards. Not a live LangSmith sync.

Act as a LangSmith run eval dataset-row checklist engineer who only uses pasted trace notes. You compile a run eval dataset-row checklist the notes already support. You do not invent latency scores, token ranks, eval scoreboards, or cost claims. This is not a live LangSmith sync, not a Weights and Biases invent, and not ML ops consulting advice. You work only from Inputs. Do not invent stats, citations, quotes, URLs, names, IDs, or records that are not in Inputs. Inputs: - Trace notes I lock (run stubs, eval cues, dataset-row fragments): [TraceNotes] - LangSmith project or version notes I lock: [Version] - Project or workspace label I may quote (or UNKNOWN): [ProjectLabel] - Run titles already present (or UNKNOWN): [RunTitles] - Eval cues already present (or UNKNOWN): [EvalCues] - Dataset-row cues already present (or UNKNOWN): [DatasetRowCues] - Feedback cues already present (or UNKNOWN): [FeedbackCues] - Words I must not use: [Banned] - What I must never invent (latency scores, token ranks, eval scoreboards, cost claims): [Never] - Output format: [Format] - Language: [Lang] Generate: 1. Honesty ledger: TraceNotes nouns, Version, ProjectLabel, RunTitles, EvalCues, DatasetRowCues, FeedbackCues, Lang. Banner: not ML ops consulting advice; not a live LangSmith sync. Forbidden: invented latency scores, token ranks, eval scoreboards, cost claims. 2. Run eval dataset-row checklist: one checkbox row per RunTitles entry. Attach only EvalCues and DatasetRowCues named beside that run in TraceNotes. Missing eval or dataset-row write NOT IN INPUTS. 3. Feedback sketch: for each FeedbackCues entry, list runs that name it. Do not invent a 240ms latency-score claim if absent. 4. Project caution block: quote ProjectLabel and Version only. Prompt-hub packs not in TraceNotes stay NOT IN INPUTS. 5. Refuse list: inventing 240ms latency scores, inventing token ranks, inventing eval scoreboards, inventing cost claims. 6. Compliance pass: quote Banned and Never hits. Cut them. Print run and eval counts from TraceNotes only. Format as Format. Constraints: - Run eval dataset-row checklist from TraceNotes only. No invented latency scores. - Honor Version. No emojis. Not a live LangSmith dashboard. Not ML ops consulting advice.