🤖 AI Tools
PromptLayer Eval Trace Scorecard from Experiment Notes (No Invented Accuracy Percent)
Compile a PromptLayer eval and trace scorecard from pasted experiment notes only. No invented accuracy percentages, F1 scores, or latency percentiles.
0Reviews
Prompt
Act as a PromptLayer eval engineer who only uses pasted experiment notes. You compile an eval and trace scorecard the notes already support. You do not invent accuracy percentages, F1 scores, latency percentiles, or leaderboard ranks. This is not a live PromptLayer cloud sync and not a model-card certificate. You work only from Inputs. Do not invent stats, citations, quotes, URLs, names, IDs, or records that are not in Inputs. Inputs: - Experiment notes I lock (trace stubs, score cues, prompt labels): [ExperimentNotes] - PromptLayer version or workspace notes I lock: [Version] - Project or experiment label I may quote (or UNKNOWN): [ExperimentLabel] - Trace names already present (or UNKNOWN): [TraceNames] - Score names already present (or UNKNOWN): [ScoreNames] - Prompt template names already present (or UNKNOWN): [PromptNames] - Model or release cues already present (or UNKNOWN): [ModelCues] - Words I must not use: [Banned] - What I must never invent (accuracy percentages, F1 scores, latency percentiles, leaderboard ranks): [Never] - Output format: [Format] - Language: [Lang] Generate: 1. Honesty ledger: ExperimentNotes nouns, Version, ExperimentLabel, TraceNames, ScoreNames, PromptNames, ModelCues, Lang. Forbidden: invented accuracy percentages, F1 scores, latency percentiles, leaderboard ranks. 2. Eval trace scorecard checklist: one checkbox row per TraceNames entry. Attach only ScoreNames named beside that trace in ExperimentNotes. Missing score write NOT IN INPUTS. 3. Prompt sketch: for each PromptNames entry, list traces that name it. Do not invent a 93.2% accuracy if absent. 4. Model caution block: quote ModelCues only. Benchmark packs not in ExperimentNotes stay NOT IN INPUTS. 5. Refuse list: inventing 93.2% accuracy, inventing F1 0.91, inventing p95 280ms latency, inventing leaderboard rank #1. 6. Compliance pass: quote Banned and Never hits. Cut them. Print trace and score counts from ExperimentNotes only. Format as Format. Constraints: - Eval trace scorecard from ExperimentNotes only. No invented accuracy percentages. - Honor Version. No emojis. Not a live PromptLayer console.
Instructions
Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.
Generated Output
This image was generated using the prompt above.

Examples
Example Input
ExperimentNotes: trace Harbor Quay Chat as pasted score Exact Match as pasted; trace Pier Edge Toolcall as pasted score Rubric Score as pasted. Accuracy invent NONE. F1 invent NONE. Version: PromptLayer cloud as pasted (do not invent unreleased AI accuracy coach) ExperimentLabel: Harbor Quay eval suite as pasted TraceNames: Harbor Quay Chat as pasted; Pier Edge Toolcall as pasted. Cedar Smoke UNKNOWN. ScoreNames: score Exact Match as pasted; score Rubric Score as pasted. Midpoint BLEU UNKNOWN. PromptNames: prompt Quay System v3 as pasted for Harbor Quay Chat; Pier Edge Toolcall prompt NOT IN INPUTS. ModelCues: model cue gpt-4.1-mini as pasted for Harbor Quay Chat. Full benchmark pack UNKNOWN. Banned: 93.2% accuracy, F1 0.91, guaranteed p95 280ms Never: invent accuracy percentages, invent F1 scores, invent latency percentiles, invent leaderboard ranks Format: ledger + eval trace scorecard + prompt sketch + model caution + refuse + compliance Lang: English
Example Output
1. Ledger. ExperimentNotes: trace Harbor Quay Chat + score Exact Match; trace Pier Edge Toolcall + score Rubric Score. Accuracy invent NONE. F1 invent NONE. Version PromptLayer cloud. ExperimentLabel Harbor Quay eval suite. TraceNames Harbor Quay Chat; Pier Edge Toolcall. ScoreNames Exact Match; Rubric Score. PromptNames Quay System v3 for Harbor Quay Chat; Pier Edge Toolcall prompt NOT IN INPUTS. ModelCues gpt-4.1-mini for Harbor Quay Chat. Full benchmark pack UNKNOWN. Lang English. Forbidden: invented accuracy percentages, F1 scores, latency percentiles, leaderboard ranks, 93.2% accuracy, F1 0.91, guaranteed p95 280ms. 2. Eval trace scorecard checklist. [ ] Harbor Quay Chat | score Exact Match as pasted. [ ] Pier Edge Toolcall | score Rubric Score as pasted. Midpoint BLEU not attached. Cedar Smoke not added. 3. Prompt sketch. prompt Quay System v3 | trace Harbor Quay Chat as pasted. Pier Edge Toolcall prompt | NOT IN INPUTS. Accuracy NOT IN INPUTS so do not invent 93.2% accuracy. Second Harbor Quay Chat cue not invented. 4. Model caution. gpt-4.1-mini as pasted for Harbor Quay Chat. Full benchmark pack UNKNOWN. Benchmark pack NOT IN INPUTS. Do not invent leaderboard packs. 5. Refuse. 93.2% accuracy: refused. F1 0.91: refused. p95 280ms latency: refused. leaderboard rank #1: refused. Unreleased AI accuracy coach: refused. 6. Compliance. Banned hits none. Traces 2. Scores 2. Format ledger+eval trace scorecard+prompt sketch+model caution+refuse+compliance. Gaps: Pier Edge Toolcall prompt, Cedar Smoke, Midpoint BLEU, Full benchmark pack, accuracy percent. Missing-data policy: if a field was blank, write NOT IN INPUTS rather than guessing. Lock any tool version named in Inputs; if unnamed, write unknown. No invented testimonials, star ratings, or press logos. If legal, clinical, insurance, HR, education-plan, or veterinary content appears, add a one-line not-advice and de-identify banner. Quote banned-word hits and cut them. End with a gaps list of five bullets the user still owes you. Character and byte caps in the job are hard; print counts when relevant. Refuse to backfill DOIs, exam dumps, PHI, PII, or compensation promises not in Inputs.