🤖 AI Tools
LangSmith Run Eval Checklist from Trace Inventory (No Invented Token Costs)
Turn a LangSmith trace inventory into a run eval checklist only. No invented token costs, latency ms, or API keys beyond the inventory.
0Reviews
Prompt
Act as a LangSmith evaluation engineer who only uses a pasted trace inventory. You write a run eval checklist the inventory already supports. You do not invent token costs, latency ms, API keys, or dataset row counts. This is not a live LangSmith API sync and not a pricing quote. You work only from Inputs. Do not invent stats, citations, quotes, URLs, names, IDs, or records that are not in Inputs. Inputs: - Trace inventory I lock (trace stubs, run cues, evaluator notes): [Inventory] - LangSmith / SDK version notes I lock: [Version] - Project or workspace label I may quote (or UNKNOWN): [Workspace] - Required run or evaluator names I may quote (or UNKNOWN): [SpecNames] - Run names already present (or UNKNOWN): [RunNames] - Evaluator labels already present (or UNKNOWN): [EvalLabels] - Trace labels already present (or UNKNOWN): [TraceLabels] - Words I must not use: [Banned] - What I must never invent (token costs, latency ms, API keys, dataset row counts): [Never] - Output format: [Format] - Language: [Lang] Generate: 1. Honesty ledger: Inventory nouns, Version, Workspace, SpecNames, RunNames, EvalLabels, TraceLabels, Lang. Forbidden: invented token costs, latency ms, API keys, dataset row counts. Banner: not a live LangSmith API sync; not a pricing quote. 2. Run eval checklist table: one row per Inventory trace stub or run cue. Missing evaluator notes write NOT IN INPUTS. Use LangSmith runs, traces, evaluators, and feedback language when Inventory supports it. Quote RunNames and EvalLabels only when present. 3. Spec name set: only names in SpecNames. Unnamed traces stay NOT IN INPUTS. Never print token-cost or latency VALUES not in Inventory. 4. Version lock: print Version. Refuse LangSmith features newer than Version if Version is named. 5. Refuse list: inventing token costs, inventing latency ms, inventing API keys, inventing dataset row counts. 6. Compliance pass: quote Banned and Never hits. Cut them. Format as Format. Constraints: - Map from Inventory only. No invented token-cost VALUES. Teach LangSmith run eval checklist mapping, not a generic Haystack or Instructor swap. - Honor Version. No emojis.
Instructions
Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.
Generated Output
This image was generated using the prompt above.

Examples
Example Input
Inventory: trace Harbor Quay Cleat Trace as pasted run cue Correctness score as pasted evaluator note dataset_run as pasted; trace River Ops Gate Trace as pasted run cue Faithfulness score as pasted. Token cost NOT IN INPUTS. Latency invent NONE. Version: LangSmith 0.1.x as pasted (do not invent unreleased AI eval APIs) Workspace: Harbor Quay LangSmith as pasted SpecNames: Cleat Trace as pasted; Gate Trace as pasted. LatencyGate UNKNOWN. RunNames: Cleat Trace as pasted; Gate Trace as pasted. Archive UNKNOWN. EvalLabels: Correctness score as pasted; Faithfulness score as pasted. Toxicity UNKNOWN. TraceLabels: dataset_run as pasted. SpanHint UNKNOWN. Banned: guaranteed token cost, latency invent, api key invent, dataset invent Never: invent token costs, invent latency ms, invent API keys, invent dataset row counts Format: ledger + table + specs + version + refuse + compliance Lang: English
Example Output
1. Ledger. Inventory: trace Harbor Quay Cleat Trace run cue Correctness score evaluator note dataset_run; trace River Ops Gate Trace run cue Faithfulness score. Token cost NOT IN INPUTS. Latency invent NONE. Version LangSmith 0.1.x as pasted. Workspace Harbor Quay LangSmith. SpecNames Cleat Trace; Gate Trace; LatencyGate UNKNOWN. RunNames Cleat Trace; Gate Trace; Archive UNKNOWN. EvalLabels Correctness score; Faithfulness score; Toxicity UNKNOWN. TraceLabels dataset_run; SpanHint UNKNOWN. Lang English. Banner: not a live LangSmith API sync; not a pricing quote. Forbidden: invented token costs, latency ms, API keys, dataset row counts, guaranteed token cost, latency invent, api key invent, dataset invent. 2. Run eval checklist. 1. Harbor Quay Cleat Trace | trace Cleat Trace as pasted (RunNames) | Correctness score as pasted (EvalLabels) | dataset_run as pasted (TraceLabels) | token cost NOT IN INPUTS | Latency NONE | LangSmith run eval checklist only. 2. River Ops Gate Trace | trace Gate Trace as pasted (RunNames) | Faithfulness score as pasted (EvalLabels) | evaluator note beyond dataset_run NOT IN INPUTS | token cost NOT IN INPUTS | Latency NONE. Toxicity not printed beyond EvalLabels. SpanHint not invented beyond TraceLabels. 3. Spec name set. Cleat Trace; Gate Trace as SpecNames. LatencyGate UNKNOWN so write LatencyGate NOT IN INPUTS. No token-cost VALUES beyond Inventory. No third trace invented. 4. Version lock. LangSmith 0.1.x as pasted. Unreleased features beyond Version not used. Credential path NOT IN INPUTS. 5. Refuse. Token costs invent: refused. Latency ms invent: refused. API keys invent: refused. Dataset row counts invent: refused. Guaranteed token cost: refused. 6. Compliance. Banned hits none. Format ledger+table+specs+version+refuse+compliance. Gaps: LatencyGate decision, Archive RunNames, Toxicity EvalLabels, SpanHint TraceLabels, credential path if any. Missing-data policy: if a field was blank, write NOT IN INPUTS rather than guessing. Lock any tool version named in Inputs; if unnamed, write unknown. No invented testimonials, star ratings, or press logos. If legal, clinical, insurance, HR, education-plan, or veterinary content appears, add a one-line not-advice and de-identify banner. Quote banned-word hits and cut them. End with a gaps list of five bullets the user still owes you. Character and byte caps in the job are hard; print counts when relevant. Refuse to backfill DOIs, exam dumps, PHI, PII, or compensation promises not in Inputs.