🤖 AI Tools
LangSmith Dataset Eval Scorecard from Experiment Notes (No Invented Accuracy Percent)
Compile a LangSmith dataset and eval scorecard from pasted experiment notes only. No invented accuracy percentages, F1 scores, or latency percentiles.
0Reviews
Prompt
Act as a LangSmith evaluation engineer aide who only uses pasted experiment notes. You compile a dataset and eval scorecard the notes already support. You do not invent accuracy percentages, F1 scores, latency percentiles, or leaderboard ranks. This is not a live LangSmith sync and not a model-certification certificate. You work only from Inputs. Do not invent stats, citations, quotes, URLs, names, IDs, or records that are not in Inputs. Inputs: - Experiment notes I lock (dataset stubs, evaluator cues, run labels): [ExperimentNotes] - LangSmith version or project notes I lock: [Version] - Project or dataset label I may quote (or UNKNOWN): [ProjectLabel] - Dataset names already present (or UNKNOWN): [DatasetNames] - Evaluator names already present (or UNKNOWN): [EvaluatorNames] - Example or split cues already present (or UNKNOWN): [ExampleCues] - Trace or feedback cues already present (or UNKNOWN): [TraceCues] - Words I must not use: [Banned] - What I must never invent (accuracy percentages, F1 scores, latency percentiles, leaderboard ranks): [Never] - Output format: [Format] - Language: [Lang] Generate: 1. Honesty ledger: ExperimentNotes nouns, Version, ProjectLabel, DatasetNames, EvaluatorNames, ExampleCues, TraceCues, Lang. Forbidden: invented accuracy percentages, F1 scores, latency percentiles, leaderboard ranks. 2. Dataset eval scorecard: one checkbox row per DatasetNames entry. Attach only EvaluatorNames named beside that dataset in ExperimentNotes. Missing evaluator write NOT IN INPUTS. 3. Example sketch: for each ExampleCues entry, list datasets that name it. Do not invent a 94.2% accuracy if absent. 4. Trace caution block: quote TraceCues only. Benchmark packs not in ExperimentNotes stay NOT IN INPUTS. 5. Refuse list: inventing 94.2% accuracy percentages, inventing 0.91 F1 scores, inventing p95 latency 120ms, inventing leaderboard rank #1. 6. Compliance pass: quote Banned and Never hits. Cut them. Print dataset and evaluator counts from ExperimentNotes only. Format as Format. Constraints: - Dataset eval scorecard from ExperimentNotes only. No invented accuracy percent. - Honor Version. No emojis. Not a live LangSmith console. Not a model-certification audit.
Instructions
Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.
Generated Output
This image was generated using the prompt above.

Examples
Example Input
ExperimentNotes: dataset Harbor Quay Support FAQ as pasted evaluator Pier Exact Match as pasted; dataset Pier Edge Tool Calls as pasted evaluator NOT IN INPUTS. Accuracy invent NONE. F1 invent NONE. Version: LangSmith as pasted (do not invent unreleased AI accuracy coach) ProjectLabel: Harbor Quay eval lab as pasted DatasetNames: Harbor Quay Support FAQ as pasted; Pier Edge Tool Calls as pasted. Cedar Multilingual UNKNOWN. EvaluatorNames: evaluator Pier Exact Match as pasted for Harbor Quay Support FAQ; Pier Edge Tool Calls evaluator NOT IN INPUTS. ExampleCues: example split Harbor Quay gold as pasted for Harbor Quay Support FAQ. Full benchmark pack UNKNOWN. TraceCues: trace feedback Harbor Quay thumbs as pasted. Leaderboard rank UNKNOWN. Banned: 94.2% accuracy, 0.91 F1, guaranteed p95 120ms Never: invent accuracy percentages, invent F1 scores, invent latency percentiles, invent leaderboard ranks Format: ledger + dataset eval scorecard + example sketch + trace caution + refuse + compliance Lang: English
Example Output
1. Ledger. ExperimentNotes: dataset Harbor Quay Support FAQ + evaluator Pier Exact Match; dataset Pier Edge Tool Calls + evaluator NOT IN INPUTS. Accuracy invent NONE. F1 invent NONE. Version LangSmith. ProjectLabel Harbor Quay eval lab. DatasetNames Harbor Quay Support FAQ; Pier Edge Tool Calls. EvaluatorNames Pier Exact Match for Harbor Quay Support FAQ; Pier Edge Tool Calls evaluator NOT IN INPUTS. ExampleCues Harbor Quay gold for Harbor Quay Support FAQ. TraceCues Harbor Quay thumbs. Full benchmark pack UNKNOWN. Lang English. Forbidden: invented accuracy percentages, F1 scores, latency percentiles, leaderboard ranks, 94.2% accuracy, 0.91 F1, guaranteed p95 120ms. 2. Dataset eval scorecard. [ ] Harbor Quay Support FAQ | evaluator Pier Exact Match as pasted. [ ] Pier Edge Tool Calls | evaluator NOT IN INPUTS. Cedar Multilingual not added. 3. Example sketch. example split Harbor Quay gold | dataset Harbor Quay Support FAQ as pasted. Pier Edge Tool Calls example | NOT IN INPUTS. Accuracy NOT IN INPUTS so do not invent 94.2% accuracy. Second Harbor Quay Support FAQ cue not invented. 4. Trace caution. Harbor Quay thumbs as pasted. Full benchmark pack UNKNOWN. Leaderboard rank NOT IN INPUTS. Do not invent benchmark packs. 5. Refuse. 94.2% accuracy percentages: refused. 0.91 F1 scores: refused. p95 latency 120ms: refused. leaderboard rank #1: refused. Unreleased AI accuracy coach: refused. 6. Compliance. Banned hits none. Datasets 2. Evaluators 1. Format ledger+dataset eval scorecard+example sketch+trace caution+refuse+compliance. Gaps: Pier Edge Tool Calls evaluator, Cedar Multilingual, Pier Edge Tool Calls example, Full benchmark pack, accuracy percentages. Missing-data policy: if a field was blank, write NOT IN INPUTS rather than guessing. Lock any tool version named in Inputs; if unnamed, write unknown. No invented testimonials, star ratings, or press logos. If legal, clinical, insurance, HR, education-plan, or veterinary content appears, add a one-line not-advice and de-identify banner. Quote banned-word hits and cut them. End with a gaps list of five bullets the user still owes you. Character and byte caps in the job are hard; print counts when relevant. Refuse to backfill DOIs, exam dumps, PHI, PII, or compensation promises not in Inputs.