🤖 AI Tools

LangSmith Run Eval Checklist (No Invented Token Costs)

Build a LangSmith run eval checklist from a pasted run-trace excerpt and eval notes. No invented token costs, latency averages, or score totals beyond Inputs.

0.0
0Reviews
P
September 19, 2026

Prompt

Act as a LangSmith eval ops engineer who only uses a pasted run-trace excerpt. You write a run eval checklist the excerpt already supports. You do not invent token costs, latency averages, score totals, or throughput claims not in Inputs. This is not a LangSmith billing pitch; not a guaranteed accuracy promise.
You work only from Inputs. Do not invent stats, citations, quotes, URLs, names, IDs, or records that are not in Inputs.

Inputs:
- Run-trace excerpt I lock (run names, datasets, statuses present): [TraceExcerpt]
- Eval notes I lock (metrics or gates named only if present): [EvalNotes]
- LangSmith / workspace version notes I lock: [Version]
- Project or dataset label I may quote (or UNKNOWN): [ProjectLabel]
- Run pair labels already present (or UNKNOWN): [RunPairs]
- Eval field names already present (or UNKNOWN): [EvalFields]
- Flag rules already present (or UNKNOWN): [FlagRules]
- Words I must not use: [Banned]
- What I must never invent (token costs, latency averages, score totals, or throughput claims): [Never]
- Output format: [Format]
- Language: [Lang]

Generate:
1. Honesty ledger: TraceExcerpt, EvalNotes, Version, ProjectLabel, RunPairs, EvalFields, FlagRules, Lang. Forbidden: invented token costs, latency averages, score totals, or throughput claims. Banner: not a LangSmith billing pitch; not a guaranteed accuracy promise.
2. Run eval checklist: one row per RunPairs item. Columns: run, TraceExcerpt facts, EvalFields, FlagRules, missing cells. Missing FlagRules write NOT IN INPUTS. Use LangSmith Traces, Datasets, Evaluators, Feedback, and Experiments language when TraceExcerpt supports it. Never print token-cost or latency VALUES not in TraceExcerpt or EvalNotes.
3. Eval lock: quote EvalNotes only. Unnamed metrics stay NOT IN INPUTS.
4. Version lock: print Version. Refuse features newer than Version if Version is named.
5. Refuse list: inventing token costs, inventing latency averages, inventing score totals, inventing throughput claims.
6. Compliance pass: quote Banned and Never hits. Cut them. Format as Format.
Constraints:
- Checklist from TraceExcerpt and EvalNotes only. No invented token-cost VALUES. Teach LangSmith run eval, not a Weights and Biases or Phoenix swap.
- Honor Version. No emojis.

Instructions

Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.

Generated Output

This image was generated using the prompt above.

LangSmith Run Eval Checklist (No Invented Token Costs) - Result

Examples

Example Input

TraceExcerpt: Harbor Quay Permit Classifier run as pasted dataset Dock FAQs as pasted status errored as pasted. Token invent NONE. Latency invent NONE.
EvalNotes: hallucination gate as pasted; groundedness UNKNOWN. Score invent NONE.
Version: LangSmith as pasted (do not invent unreleased auto-cost pack)
ProjectLabel: Harbor Quay LLM ops desk as pasted
RunPairs: Harbor Quay Permit Classifier as pasted; River Ops summarizer UNKNOWN.
EvalFields: hallucination gate as pasted. Throughput NOT IN INPUTS.
FlagRules: status errored flag as pasted.
Banned: guaranteed accuracy, token invent, latency invent, score invent
Never: invent token costs, invent latency averages, invent score totals, invent throughput claims
Format: ledger + eval checklist + eval lock + version + refuse + compliance
Lang: English

Example Output

1. Ledger. TraceExcerpt: Harbor Quay Permit Classifier run dataset Dock FAQs status errored. Token invent NONE. Latency invent NONE. EvalNotes hallucination gate; groundedness UNKNOWN. Score invent NONE. Version LangSmith as pasted. ProjectLabel Harbor Quay LLM ops desk. RunPairs Harbor Quay Permit Classifier; River Ops summarizer UNKNOWN. EvalFields hallucination gate. Throughput NOT IN INPUTS. FlagRules status errored flag. Lang English. Banner: not a LangSmith billing pitch; not a guaranteed accuracy promise. Forbidden: invented token costs, latency averages, score totals, throughput claims.

2. Run eval checklist.
1. Harbor Quay Permit Classifier | TraceExcerpt Dock FAQs errored | EvalFields hallucination gate | FlagRules status errored triggered | token VALUES NONE | LangSmith eval only.
2. River Ops summarizer | RunPairs UNKNOWN | write NOT IN INPUTS | do not invent a second run.
groundedness UNKNOWN so write groundedness NOT IN INPUTS. Throughput NOT IN INPUTS so write throughput VALUES NOT IN INPUTS.

3. Eval lock. hallucination gate as pasted. groundedness UNKNOWN so NOT IN INPUTS. Score invent NONE.

4. Version lock. LangSmith as pasted. Refuse unreleased auto-cost pack.

5. Refuse list. inventing token costs; inventing latency averages; inventing score totals; inventing throughput claims.

6. Compliance. Quoted Banned: guaranteed accuracy, token invent, latency invent, score invent. Cut. Quoted Never: invent token costs, invent latency averages, invent score totals, invent throughput claims. Cut. Format ledger + eval checklist + eval lock + version + refuse + compliance.

Missing-data policy: if a field was blank, write NOT IN INPUTS rather than guessing. Lock any tool version named in Inputs; if unnamed, write unknown. No invented testimonials, star ratings, or press logos. If legal, clinical, insurance, HR, education-plan, or veterinary content appears, add a one-line not-advice and de-identify banner. Quote banned-word hits and cut them. End with a gaps list of five bullets the user still owes you. Character and byte caps in the job are hard; print counts when relevant. Refuse to backfill DOIs, exam dumps, PHI, PII, or compensation promises not in Inputs.

Reviews (0)

Please login to leave a review.
Loading reviews...