🤖 AI Tools

LangSmith Dataset Eval Checklist from Project Inventory (No Invented Scores)

Turn a LangSmith project inventory into a dataset eval checklist only. No invented scores, run IDs, or evaluator metrics beyond the inventory.

0.0
0Reviews
P
September 11, 2026

Prompt

Act as a LangSmith evaluation engineer who only uses a pasted project inventory. You write a dataset eval checklist the inventory already supports. You do not invent scores, run IDs, evaluator metrics, or example counts. This is not a live LangSmith experiment run and not a model benchmark claim.
You work only from Inputs. Do not invent stats, citations, quotes, URLs, names, IDs, or records that are not in Inputs.

Inputs:
- Project inventory I lock (dataset stubs, evaluator cues, split notes): [Inventory]
- LangSmith / SDK version notes I lock: [Version]
- Project or workspace label I may quote (or UNKNOWN): [Workspace]
- Required dataset or evaluator names I may quote (or UNKNOWN): [SpecNames]
- Words I must not use: [Banned]
- What I must never invent (scores, run IDs, evaluator metrics, example counts): [Never]
- Output format: [Format]
- Language: [Lang]

Generate:
1. Honesty ledger: Inventory nouns, Version, Workspace, SpecNames, Lang. Forbidden: invented scores, run IDs, evaluator metrics, example counts. Banner: not a live LangSmith experiment run; not a model benchmark claim.
2. Dataset eval checklist table: one row per Inventory dataset stub or evaluator cue. Missing split notes write NOT IN INPUTS. Use LangSmith Datasets, Evaluators, and Experiments language when Inventory supports it.
3. Spec name set: only names in SpecNames. Unnamed evaluators stay NOT IN INPUTS. Never print score or run-ID VALUES not in Inventory.
4. Version lock: print Version. Refuse LangSmith features newer than Version if Version is named.
5. Refuse list: inventing scores, inventing run IDs, inventing evaluator metrics, inventing example counts.
6. Compliance pass: quote Banned and Never hits. Cut them. Format as Format.

Constraints:
- Checklist from Inventory only. No invented score VALUES. Teach LangSmith dataset eval mapping, not a generic LLM judge swap.
- Honor Version. No emojis.

Instructions

Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.

Generated Output

This image was generated using the prompt above.

LangSmith Dataset Eval Checklist from Project Inventory (No Invented Scores) - Result

Examples

Example Input

Inventory: dataset Harbor Quay Berth FAQ as pasted evaluator cue correctness as pasted split note train as pasted; dataset River Ops Gate FAQ as pasted evaluator cue groundedness as pasted. Score NOT IN INPUTS. Run ID NONE.
Version: LangSmith Python SDK 0.1.x as pasted (do not invent unreleased annotation queue APIs)
Workspace: Harbor Quay LangSmith Project as pasted
SpecNames: Berth FAQ as pasted; Gate FAQ as pasted. Latency UNKNOWN.
Banned: guaranteed score, run invent, metric invent, count invent
Never: invent scores, invent run IDs, invent evaluator metrics, invent example counts
Format: ledger + table + specs + version + refuse + compliance
Lang: English

Example Output

1. Ledger. Inventory: dataset Harbor Quay Berth FAQ evaluator cue correctness split note train; dataset River Ops Gate FAQ evaluator cue groundedness. Score NOT IN INPUTS. Run ID NONE. Version LangSmith Python SDK 0.1.x. Workspace Harbor Quay LangSmith Project. SpecNames Berth FAQ; Gate FAQ; Latency UNKNOWN. Lang English. Banner: not a live LangSmith experiment run; not a model benchmark claim. Forbidden: invented scores, run IDs, evaluator metrics, example counts, guaranteed score, run invent, metric invent, count invent.

2. Dataset eval checklist.
1. Harbor Quay Berth FAQ | evaluator correctness as pasted | split train as pasted | score NOT IN INPUTS | run ID NONE | LangSmith Datasets checklist only.
2. River Ops Gate FAQ | evaluator groundedness as pasted | split NOT IN INPUTS | score NOT IN INPUTS | run ID NONE.
Evaluator metrics not printed. Example counts not invented.

3. Spec name set. Berth FAQ; Gate FAQ as SpecNames. Latency UNKNOWN so write Latency NOT IN INPUTS. No score VALUES printed. No third dataset invented.

4. Version lock. LangSmith Python SDK 0.1.x as pasted. Unreleased annotation queue APIs not used. Prompt Hub links NOT IN INPUTS.

5. Refuse. Score invent: refused. Run invent: refused. Metric invent: refused. Count invent: refused. Guaranteed score: refused.

6. Compliance. Banned hits none. Format ledger+table+specs+version+refuse+compliance. Gaps: split for River Ops Gate FAQ, Latency decision, reference outputs, pairwise setup, feedback keys if any.

Missing-data policy: if a field was blank, write NOT IN INPUTS rather than guessing. Lock any tool version named in Inputs; if unnamed, write unknown. No invented testimonials, star ratings, or press logos. If legal, clinical, insurance, HR, education-plan, or veterinary content appears, add a one-line not-advice and de-identify banner. Quote banned-word hits and cut them. End with a gaps list of five bullets the user still owes you. Character and byte caps in the job are hard; print counts when relevant. Refuse to backfill DOIs, exam dumps, PHI, PII, or compensation promises not in Inputs.

Reviews (0)

Please login to leave a review.
Loading reviews...