🤖 AI Tools
Phoenix Eval Trace Checklist from Experiment Notes (No Invented Accuracy Percents)
Compile an Arize Phoenix eval-trace checklist from pasted experiment notes only. No invented accuracy percents, latency ranks, or eval scoreboards. Not a live Phoenix sync.
0Reviews
Prompt
Act as an Arize Phoenix eval-trace checklist engineer who only uses pasted experiment notes. You compile an eval-trace checklist the notes already support. You do not invent accuracy percents, latency ranks, eval scoreboards, or savings claims. This is not a live Phoenix sync, not LangSmith merge, and not ML capacity advice. You work only from Inputs. Do not invent stats, citations, quotes, URLs, names, IDs, or records that are not in Inputs. Inputs: - Experiment notes I lock (trace stubs, eval cues, span fragments): [ExperimentNotes] - Phoenix version or project notes I lock: [Version] - Project or dataset label I may quote (or UNKNOWN): [ProjectLabel] - Trace IDs already present (or UNKNOWN): [TraceIds] - Eval cues already present (or UNKNOWN): [EvalCues] - Span cues already present (or UNKNOWN): [SpanCues] - Dataset cues already present (or UNKNOWN): [DatasetCues] - Words I must not use: [Banned] - What I must never invent (accuracy percents, latency ranks, eval scoreboards, savings claims): [Never] - Output format: [Format] - Language: [Lang] Generate: 1. Honesty ledger: ExperimentNotes nouns, Version, ProjectLabel, TraceIds, EvalCues, SpanCues, DatasetCues, Lang. Banner: not ML capacity advice; not a live Phoenix sync. Forbidden: invented accuracy percents, latency ranks, eval scoreboards, savings claims. 2. Eval-trace checklist: one checkbox row per TraceIds entry. Attach only EvalCues named beside that trace in ExperimentNotes. Missing eval write NOT IN INPUTS. 3. Span sketch: for each SpanCues entry, list traces that name it. Do not invent a 91-percent accuracy claim if absent. 4. Dataset caution block: quote DatasetCues only. Golden packs not in ExperimentNotes stay NOT IN INPUTS. 5. Refuse list: inventing 91-percent accuracy percents, inventing latency ranks, inventing eval scoreboards, inventing savings claims. 6. Compliance pass: quote Banned and Never hits. Cut them. Print trace and eval counts from ExperimentNotes only. Format as Format. Constraints: - Eval-trace checklist from ExperimentNotes only. No invented accuracy percents. - Honor Version. No emojis. Not a live Phoenix dashboard. Not ML capacity advice.
Instructions
Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.
Generated Output
This image was generated using the prompt above.

Examples
Example Input
ExperimentNotes: trace id tr-harbor-01 as pasted eval cue faithfulness as pasted span cue retrieval as pasted; trace id tr-quay-02 as pasted eval cue answer relevance as pasted. Accuracy invent NONE. Latency invent NONE. Version: Arize Phoenix as pasted (do not invent unreleased AI accuracy coach) ProjectLabel: Cedar Pier RAG eval as pasted TraceIds: tr-harbor-01 as pasted; tr-quay-02 as pasted. Batch nightly UNKNOWN. EvalCues: eval cue faithfulness as pasted for tr-harbor-01; eval cue answer relevance as pasted for tr-quay-02. Toxicity pack UNKNOWN. SpanCues: span cue retrieval as pasted for tr-harbor-01; tr-quay-02 span NOT IN INPUTS. DatasetCues: dataset cue pier-faq-v2 as pasted for tr-harbor-01. Golden pack UNKNOWN. Banned: 91% accuracy, latency rank #1, guaranteed eval scoreboard Never: invent accuracy percents, invent latency ranks, invent eval scoreboards, invent savings claims Format: ledger + eval-trace checklist + span sketch + dataset caution + refuse + compliance Lang: English
Example Output
1. Ledger. ExperimentNotes: trace id tr-harbor-01 + eval cue faithfulness + span cue retrieval; trace id tr-quay-02 + eval cue answer relevance. Accuracy invent NONE. Latency invent NONE. Version Arize Phoenix. ProjectLabel Cedar Pier RAG eval. TraceIds tr-harbor-01; tr-quay-02. EvalCues faithfulness for tr-harbor-01; answer relevance for tr-quay-02. SpanCues retrieval for tr-harbor-01; tr-quay-02 span NOT IN INPUTS. DatasetCues pier-faq-v2 for tr-harbor-01. Golden pack UNKNOWN. Lang English. Banner: not ML capacity advice; not a live Phoenix sync. Forbidden: invented accuracy percents, latency ranks, eval scoreboards, savings claims, 91% accuracy, latency rank #1, guaranteed eval scoreboard. 2. Eval-trace checklist. [ ] tr-harbor-01 | eval faithfulness as pasted. [ ] tr-quay-02 | eval answer relevance as pasted. Toxicity pack not attached. Batch nightly not added. 3. Span sketch. span cue retrieval | trace tr-harbor-01 as pasted. tr-quay-02 span | NOT IN INPUTS. Accuracy percents NOT IN INPUTS so do not invent 91% accuracy. Second retrieval cue not invented. 4. Dataset caution. dataset cue pier-faq-v2 as pasted for tr-harbor-01. Golden pack UNKNOWN. Do not invent savings claim packs. 5. Refuse. 91-percent accuracy percents: refused. latency ranks: refused. eval scoreboards: refused. savings claims: refused. Unreleased AI accuracy coach: refused. 6. Compliance. Banned hits none. Traces 2. Evals named 2. Format ledger+eval-trace checklist+span sketch+dataset caution+refuse+compliance. Gaps: tr-quay-02 span, Toxicity pack, Batch nightly, Golden pack, accuracy percents. Missing-data policy: if a field was blank, write NOT IN INPUTS rather than guessing. Lock any tool version named in Inputs; if unnamed, write unknown. No invented testimonials, star ratings, or press logos. If legal, clinical, insurance, HR, education-plan, or veterinary content appears, add a one-line not-advice and de-identify banner. Quote banned-word hits and cut them. End with a gaps list of five bullets the user still owes you. Character and byte caps in the job are hard; print counts when relevant. Refuse to backfill DOIs, exam dumps, PHI, PII, or compensation promises not in Inputs.