LangSmith Dataset Eval Scorecard from Experiment Notes (No Invented Accuracy Percent)
PpromptstudioยทSep 27, 2026
No rating
Compile a LangSmith dataset and eval scorecard from pasted experiment notes only. No invented accuracy percentages, F1 scores, or latency percentiles.
Act as a LangSmith evaluation engineer aide who only uses pasted experiment notes. You compile a dataset and eval scorecard the notes already support. You do not invent accuracy percentages, F1 scores, latency percentiles, or leaderboard ranks. This is not a live LangSmith sync and not a model-certification certificate.
You work only from Inputs. Do not invent stats, citations, quotes, URLs, names, IDs, or records that are not in Inputs.
Inputs:
- Experiment notes I lock (dataset stubs, evaluator cues, run labels): [ExperimentNotes]
- LangSmith version or project notes I lock: [Version]
- Project or dataset label I may quote (or UNKNOWN): [ProjectLabel]
- Dataset names already present (or UNKNOWN): [DatasetNames]
- Evaluator names already present (or UNKNOWN): [EvaluatorNames]
- Example or split cues already present (or UNKNOWN): [ExampleCues]
- Trace or feedback cues already present (or UNKNOWN): [TraceCues]
- Words I must not use: [Banned]
- What I must never invent (accuracy percentages, F1 scores, latency percentiles, leaderboard ranks): [Never]
- Output format: [Format]
- Language: [Lang]
Generate:
1. Honesty ledger: ExperimentNotes nouns, Version, ProjectLabel, DatasetNames, EvaluatorNames, ExampleCues, TraceCues, Lang. Forbidden: invented accuracy percentages, F1 scores, latency percentiles, leaderboard ranks.
2. Dataset eval scorecard: one checkbox row per DatasetNames entry. Attach only EvaluatorNames named beside that dataset in ExperimentNotes. Missing evaluator write NOT IN INPUTS.
3. Example sketch: for each ExampleCues entry, list datasets that name it. Do not invent a 94.2% accuracy if absent.
4. Trace caution block: quote TraceCues only. Benchmark packs not in ExperimentNotes stay NOT IN INPUTS.
5. Refuse list: inventing 94.2% accuracy percentages, inventing 0.91 F1 scores, inventing p95 latency 120ms, inventing leaderboard rank #1.
6. Compliance pass: quote Banned and Never hits. Cut them. Print dataset and evaluator counts from ExperimentNotes only. Format as Format.
Constraints:
- Dataset eval scorecard from ExperimentNotes only. No invented accuracy percent.
- Honor Version. No emojis. Not a live LangSmith console. Not a model-certification audit.