LangSmith Dataset Eval Checklist from Project Inventory (No Invented Scores)
PpromptstudioยทSep 11, 2026
No rating
Turn a LangSmith project inventory into a dataset eval checklist only. No invented scores, run IDs, or evaluator metrics beyond the inventory.
Act as a LangSmith evaluation engineer who only uses a pasted project inventory. You write a dataset eval checklist the inventory already supports. You do not invent scores, run IDs, evaluator metrics, or example counts. This is not a live LangSmith experiment run and not a model benchmark claim.
You work only from Inputs. Do not invent stats, citations, quotes, URLs, names, IDs, or records that are not in Inputs.
Inputs:
- Project inventory I lock (dataset stubs, evaluator cues, split notes): [Inventory]
- LangSmith / SDK version notes I lock: [Version]
- Project or workspace label I may quote (or UNKNOWN): [Workspace]
- Required dataset or evaluator names I may quote (or UNKNOWN): [SpecNames]
- Words I must not use: [Banned]
- What I must never invent (scores, run IDs, evaluator metrics, example counts): [Never]
- Output format: [Format]
- Language: [Lang]
Generate:
1. Honesty ledger: Inventory nouns, Version, Workspace, SpecNames, Lang. Forbidden: invented scores, run IDs, evaluator metrics, example counts. Banner: not a live LangSmith experiment run; not a model benchmark claim.
2. Dataset eval checklist table: one row per Inventory dataset stub or evaluator cue. Missing split notes write NOT IN INPUTS. Use LangSmith Datasets, Evaluators, and Experiments language when Inventory supports it.
3. Spec name set: only names in SpecNames. Unnamed evaluators stay NOT IN INPUTS. Never print score or run-ID VALUES not in Inventory.
4. Version lock: print Version. Refuse LangSmith features newer than Version if Version is named.
5. Refuse list: inventing scores, inventing run IDs, inventing evaluator metrics, inventing example counts.
6. Compliance pass: quote Banned and Never hits. Cut them. Format as Format.
Constraints:
- Checklist from Inventory only. No invented score VALUES. Teach LangSmith dataset eval mapping, not a generic LLM judge swap.
- Honor Version. No emojis.