🤖 AI Tools
LangSmith Dataset Eval Checklist from Project Inventory (No Invented Scores)
Turn a LangSmith project inventory into a dataset eval checklist only. No invented scores, run IDs, or evaluator metrics beyond the inventory.
0Reviews
Prompt
Act as a LangSmith evaluation engineer who only uses a pasted project inventory. You write a dataset eval checklist the inventory already supports. You do not invent scores, run IDs, evaluator metrics, or example counts. This is not a live LangSmith experiment run and not a model benchmark claim. You work only from Inputs. Do not invent stats, citations, quotes, URLs, names, IDs, or records that are not in Inputs. Inputs: - Project inventory I lock (dataset stubs, evaluator cues, split notes): [Inventory] - LangSmith / SDK version notes I lock: [Version] - Project or workspace label I may quote (or UNKNOWN): [Workspace] - Required dataset or evaluator names I may quote (or UNKNOWN): [SpecNames] - Words I must not use: [Banned] - What I must never invent (scores, run IDs, evaluator metrics, example counts): [Never] - Output format: [Format] - Language: [Lang] Generate: 1. Honesty ledger: Inventory nouns, Version, Workspace, SpecNames, Lang. Forbidden: invented scores, run IDs, evaluator metrics, example counts. Banner: not a live LangSmith experiment run; not a model benchmark claim. 2. Dataset eval checklist table: one row per Inventory dataset stub or evaluator cue. Missing split notes write NOT IN INPUTS. Use LangSmith Datasets, Evaluators, and Experiments language when Inventory supports it. 3. Spec name set: only names in SpecNames. Unnamed evaluators stay NOT IN INPUTS. Never print score or run-ID VALUES not in Inventory. 4. Version lock: print Version. Refuse LangSmith features newer than Version if Version is named. 5. Refuse list: inventing scores, inventing run IDs, inventing evaluator metrics, inventing example counts. 6. Compliance pass: quote Banned and Never hits. Cut them. Format as Format. Constraints: - Checklist from Inventory only. No invented score VALUES. Teach LangSmith dataset eval mapping, not a generic LLM judge swap. - Honor Version. No emojis.
Instructions
Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.
Generated Output
This image was generated using the prompt above.

Examples
Example Input
Inventory: dataset Harbor Quay Berth FAQ as pasted evaluator cue correctness as pasted split note train as pasted; dataset River Ops Gate FAQ as pasted evaluator cue groundedness as pasted. Score NOT IN INPUTS. Run ID NONE. Version: LangSmith Python SDK 0.1.x as pasted (do not invent unreleased annotation queue APIs) Workspace: Harbor Quay LangSmith Project as pasted SpecNames: Berth FAQ as pasted; Gate FAQ as pasted. Latency UNKNOWN. Banned: guaranteed score, run invent, metric invent, count invent Never: invent scores, invent run IDs, invent evaluator metrics, invent example counts Format: ledger + table + specs + version + refuse + compliance Lang: English
Example Output
1. Ledger. Inventory: dataset Harbor Quay Berth FAQ evaluator cue correctness split note train; dataset River Ops Gate FAQ evaluator cue groundedness. Score NOT IN INPUTS. Run ID NONE. Version LangSmith Python SDK 0.1.x. Workspace Harbor Quay LangSmith Project. SpecNames Berth FAQ; Gate FAQ; Latency UNKNOWN. Lang English. Banner: not a live LangSmith experiment run; not a model benchmark claim. Forbidden: invented scores, run IDs, evaluator metrics, example counts, guaranteed score, run invent, metric invent, count invent. 2. Dataset eval checklist. 1. Harbor Quay Berth FAQ | evaluator correctness as pasted | split train as pasted | score NOT IN INPUTS | run ID NONE | LangSmith Datasets checklist only. 2. River Ops Gate FAQ | evaluator groundedness as pasted | split NOT IN INPUTS | score NOT IN INPUTS | run ID NONE. Evaluator metrics not printed. Example counts not invented. 3. Spec name set. Berth FAQ; Gate FAQ as SpecNames. Latency UNKNOWN so write Latency NOT IN INPUTS. No score VALUES printed. No third dataset invented. 4. Version lock. LangSmith Python SDK 0.1.x as pasted. Unreleased annotation queue APIs not used. Prompt Hub links NOT IN INPUTS. 5. Refuse. Score invent: refused. Run invent: refused. Metric invent: refused. Count invent: refused. Guaranteed score: refused. 6. Compliance. Banned hits none. Format ledger+table+specs+version+refuse+compliance. Gaps: split for River Ops Gate FAQ, Latency decision, reference outputs, pairwise setup, feedback keys if any. Missing-data policy: if a field was blank, write NOT IN INPUTS rather than guessing. Lock any tool version named in Inputs; if unnamed, write unknown. No invented testimonials, star ratings, or press logos. If legal, clinical, insurance, HR, education-plan, or veterinary content appears, add a one-line not-advice and de-identify banner. Quote banned-word hits and cut them. End with a gaps list of five bullets the user still owes you. Character and byte caps in the job are hard; print counts when relevant. Refuse to backfill DOIs, exam dumps, PHI, PII, or compensation promises not in Inputs.