🤖 AI Tools
Galileo Eval Dataset Span Checklist from Ops Notes (No Invented Latency Scores)
Compile a Galileo eval-dataset-span checklist from pasted ops notes only. No invented latency scores, accuracy ranks, or eval scoreboards. Not a live Galileo sync.
0Reviews
Prompt
Act as a Galileo eval-dataset-span checklist engineer who only uses pasted ops notes. You compile an eval-dataset-span checklist the notes already support. You do not invent latency scores, accuracy ranks, eval scoreboards, or guaranteed quality claims. This is not a live Galileo sync, not a LangSmith invent, and not MLOps consulting advice. You work only from Inputs. Do not invent stats, citations, quotes, URLs, names, IDs, or records that are not in Inputs. Inputs: - Ops notes I lock (eval stubs, dataset cues, span fragments): [OpsNotes] - Galileo project or version notes I lock: [Version] - Project or workspace label I may quote (or UNKNOWN): [ProjectLabel] - Eval titles already present (or UNKNOWN): [EvalTitles] - Dataset cues already present (or UNKNOWN): [DatasetCues] - Span cues already present (or UNKNOWN): [SpanCues] - Trace cues already present (or UNKNOWN): [TraceCues] - Words I must not use: [Banned] - What I must never invent (latency scores, accuracy ranks, eval scoreboards, guaranteed quality claims): [Never] - Output format: [Format] - Language: [Lang] Generate: 1. Honesty ledger: OpsNotes nouns, Version, ProjectLabel, EvalTitles, DatasetCues, SpanCues, TraceCues, Lang. Banner: not MLOps consulting advice; not a live Galileo sync. Forbidden: invented latency scores, accuracy ranks, eval scoreboards, guaranteed quality claims. 2. Eval-dataset-span checklist: one checkbox row per EvalTitles entry. Attach only DatasetCues and SpanCues named beside that eval in OpsNotes. Missing dataset or span write NOT IN INPUTS. 3. Trace sketch: for each TraceCues entry, list evals that name it. Do not invent a 195ms latency-score claim if absent. 4. Project caution block: quote ProjectLabel and Version only. Annotation packs not in OpsNotes stay NOT IN INPUTS. 5. Refuse list: inventing 195ms latency scores, inventing accuracy ranks, inventing eval scoreboards, inventing guaranteed quality claims. 6. Compliance pass: quote Banned and Never hits. Cut them. Print eval and dataset counts from OpsNotes only. Format as Format. Constraints: - Eval-dataset-span checklist from OpsNotes only. No invented latency scores. - Honor Version. No emojis. Not a live Galileo dashboard. Not MLOps consulting advice.
Instructions
Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.
Generated Output
This image was generated using the prompt above.

Examples
Example Input
OpsNotes: eval title Harbor Quay Faithfulness Eval as pasted dataset cue dataset:pier-gold as pasted span cue span:fog-retriever as pasted; eval title Quay Storm Toxicity Eval as pasted dataset cue dataset:storm-cases as pasted. Latency invent NONE. Rank invent NONE. Version: Galileo as pasted (do not invent unreleased AI eval coach) ProjectLabel: Cedar Pier llm observability as pasted EvalTitles: Harbor Quay Faithfulness Eval as pasted; Quay Storm Toxicity Eval as pasted. Annotation pack UNKNOWN. DatasetCues: dataset cue dataset:pier-gold as pasted for Harbor Quay Faithfulness Eval; dataset cue dataset:storm-cases as pasted for Quay Storm Toxicity Eval. Experiment pack UNKNOWN. SpanCues: span cue span:fog-retriever as pasted for Harbor Quay Faithfulness Eval; Quay Storm Toxicity Eval span NOT IN INPUTS. TraceCues: trace cue trace:quay-session as pasted for Harbor Quay Faithfulness Eval. Session pack UNKNOWN. Banned: 195ms latency, accuracy rank #1, guaranteed eval scoreboard Never: invent latency scores, invent accuracy ranks, invent eval scoreboards, invent guaranteed quality claims Format: ledger + eval-dataset-span checklist + trace sketch + project caution + refuse + compliance Lang: English
Example Output
1. Ledger. OpsNotes: eval title Harbor Quay Faithfulness Eval + dataset cue dataset:pier-gold + span cue span:fog-retriever; eval title Quay Storm Toxicity Eval + dataset cue dataset:storm-cases. Latency invent NONE. Rank invent NONE. Version Galileo. ProjectLabel Cedar Pier llm observability. EvalTitles Harbor Quay Faithfulness Eval; Quay Storm Toxicity Eval. DatasetCues dataset:pier-gold for Harbor Quay Faithfulness Eval; dataset:storm-cases for Quay Storm Toxicity Eval. SpanCues span:fog-retriever for Harbor Quay Faithfulness Eval; Quay Storm Toxicity Eval span NOT IN INPUTS. TraceCues trace:quay-session for Harbor Quay Faithfulness Eval. Session pack UNKNOWN. Lang English. Banner: not MLOps consulting advice; not a live Galileo sync. Forbidden: invented latency scores, accuracy ranks, eval scoreboards, guaranteed quality claims, 195ms latency, accuracy rank #1, guaranteed eval scoreboard. 2. Eval-dataset-span checklist. [ ] Harbor Quay Faithfulness Eval | dataset dataset:pier-gold as pasted | span span:fog-retriever as pasted. [ ] Quay Storm Toxicity Eval | dataset dataset:storm-cases as pasted | span NOT IN INPUTS. Experiment pack not attached. Annotation pack not added. 3. Trace sketch. trace cue trace:quay-session | eval Harbor Quay Faithfulness Eval as pasted. Quay Storm Toxicity Eval trace | NOT IN INPUTS. Latency scores NOT IN INPUTS so do not invent 195ms latency. Second trace cue not invented. 4. Project caution. ProjectLabel Cedar Pier llm observability. Version Galileo. Session pack UNKNOWN. Do not invent guaranteed quality claim packs. 5. Refuse. 195ms latency scores: refused. accuracy ranks: refused. eval scoreboards: refused. guaranteed quality claims: refused. Unreleased AI eval coach: refused. 6. Compliance. Banned hits none. Evals 2. Datasets named 2. Format ledger+eval-dataset-span checklist+trace sketch+project caution+refuse+compliance. Gaps: Quay Storm Toxicity Eval span, Quay Storm Toxicity Eval trace, Experiment pack, Annotation pack, Session pack. Missing-data policy: if a field was blank, write NOT IN INPUTS rather than guessing. Lock any tool version named in Inputs; if unnamed, write unknown. No invented testimonials, star ratings, or press logos. If legal, clinical, insurance, HR, education-plan, or veterinary content appears, add a one-line not-advice and de-identify banner. Quote banned-word hits and cut them. End with a gaps list of five bullets the user still owes you. Character and byte caps in the job are hard; print counts when relevant. Refuse to backfill DOIs, exam dumps, PHI, PII, or compensation promises not in Inputs.