🤖 AI Tools

Weights and Biases Run Trace Checklist from Experiment Notes (No Invented Token Totals)

Compile a Weights and Biases run-trace checklist from pasted experiment notes only. No invented token totals, accuracy ranks, or GPU-hour scoreboards. Not a live W&B sync.

0.0
0Reviews
P
September 28, 2026

Prompt

Act as a Weights and Biases run-trace checklist coordinator who only uses pasted experiment notes. You compile a run-trace checklist the notes already support. You do not invent token totals, accuracy ranks, GPU-hour scoreboards, or SOTA claims. This is not a live W&B sync, not MLflow merge, and not model-performance scoring advice.
You work only from Inputs. Do not invent stats, citations, quotes, URLs, names, IDs, or records that are not in Inputs.

Inputs:
- Experiment notes I lock (run stubs, metric cues, config fragments): [ExperimentNotes]
- Weights and Biases project or SDK version notes I lock: [Version]
- Project or team label I may quote (or UNKNOWN): [ProjectLabel]
- Run labels already present (or UNKNOWN): [RunLabels]
- Metric cues already present (or UNKNOWN): [MetricCues]
- Config cues already present (or UNKNOWN): [ConfigCues]
- Artifact cues already present (or UNKNOWN): [ArtifactCues]
- Words I must not use: [Banned]
- What I must never invent (token totals, accuracy ranks, GPU-hour scoreboards, SOTA claims): [Never]
- Output format: [Format]
- Language: [Lang]

Generate:
1. Honesty ledger: ExperimentNotes nouns, Version, ProjectLabel, RunLabels, MetricCues, ConfigCues, ArtifactCues, Lang. Banner: not model-performance scoring advice; not a live W&B sync. Forbidden: invented token totals, accuracy ranks, GPU-hour scoreboards, SOTA claims.
2. Run-trace checklist: one checkbox row per RunLabels entry. Attach only MetricCues and ConfigCues named beside that run in ExperimentNotes. Missing cue write NOT IN INPUTS.
3. Artifact sketch: for each ArtifactCues entry, list runs that name it. Do not invent a 2M token-total claim if absent.
4. Metric caution block: quote MetricCues only. MLflow packs not in ExperimentNotes stay NOT IN INPUTS.
5. Refuse list: inventing 2M token totals, inventing accuracy ranks, inventing GPU-hour scoreboards, inventing SOTA claims.
6. Compliance pass: quote Banned and Never hits. Cut them. Print run and metric counts from ExperimentNotes only. Format as Format.

Constraints:
- Run-trace checklist from ExperimentNotes only. No invented token totals.
- Honor Version. No emojis. Not a live W&B dashboard. Not model-performance scoring advice.

Instructions

Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.

Generated Output

This image was generated using the prompt above.

Weights and Biases Run Trace Checklist from Experiment Notes (No Invented Token Totals) - Result

Examples

Example Input

ExperimentNotes: run label Harbor RAG Baseline as pasted metric val_loss as pasted config lr=2e-5 as pasted; run label Quay Rerank Sweep as pasted metric mrr as pasted. Token invent NONE. Accuracy invent NONE.
Version: Weights and Biases as pasted (do not invent unreleased AI sweep coach)
ProjectLabel: Cedar Pier retrieval project as pasted
RunLabels: Harbor RAG Baseline as pasted; Quay Rerank Sweep as pasted. Background Ablation UNKNOWN.
MetricCues: metric val_loss as pasted for Harbor RAG Baseline; metric mrr as pasted for Quay Rerank Sweep. BLEU pack UNKNOWN.
ConfigCues: config lr=2e-5 as pasted for Harbor RAG Baseline; Quay Rerank Sweep config NOT IN INPUTS.
ArtifactCues: artifact checkpoint.pt as pasted for Harbor RAG Baseline; Quay Rerank Sweep artifact NOT IN INPUTS.
Banned: 2M tokens, accuracy rank #1, guaranteed SOTA
Never: invent token totals, invent accuracy ranks, invent GPU-hour scoreboards, invent SOTA claims
Format: ledger + run-trace checklist + artifact sketch + metric caution + refuse + compliance
Lang: English

Example Output

1. Ledger. ExperimentNotes: run label Harbor RAG Baseline + metric val_loss + config lr=2e-5; run label Quay Rerank Sweep + metric mrr. Token invent NONE. Accuracy invent NONE. Version Weights and Biases. ProjectLabel Cedar Pier retrieval project. RunLabels Harbor RAG Baseline; Quay Rerank Sweep. MetricCues val_loss for Harbor RAG Baseline; mrr for Quay Rerank Sweep. ConfigCues lr=2e-5 for Harbor RAG Baseline; Quay Rerank Sweep config NOT IN INPUTS. ArtifactCues checkpoint.pt for Harbor RAG Baseline; Quay Rerank Sweep artifact NOT IN INPUTS. Lang English. Banner: not model-performance scoring advice; not a live W&B sync. Forbidden: invented token totals, accuracy ranks, GPU-hour scoreboards, SOTA claims, 2M tokens, accuracy rank #1, guaranteed SOTA.

2. Run-trace checklist.
[ ] Harbor RAG Baseline | metric val_loss | config lr=2e-5 as pasted.
[ ] Quay Rerank Sweep | metric mrr | config NOT IN INPUTS.
Background Ablation not attached. BLEU pack not added.

3. Artifact sketch.
artifact checkpoint.pt | run Harbor RAG Baseline as pasted.
Quay Rerank Sweep artifact | NOT IN INPUTS.
Token totals NOT IN INPUTS so do not invent 2M tokens. Second checkpoint.pt cue not invented.

4. Metric caution. metric val_loss as pasted for Harbor RAG Baseline. metric mrr as pasted for Quay Rerank Sweep. MLflow pack NOT IN INPUTS. Do not invent SOTA packs.

5. Refuse. 2M token totals: refused. accuracy ranks: refused. GPU-hour scoreboards: refused. SOTA claims: refused. Unreleased AI sweep coach: refused.

6. Compliance. Banned hits none. Runs 2. Metrics named 2. Format ledger+run-trace checklist+artifact sketch+metric caution+refuse+compliance. Gaps: Quay Rerank Sweep config, Quay Rerank Sweep artifact, Background Ablation, BLEU pack, token totals.

Missing-data policy: if a field was blank, write NOT IN INPUTS rather than guessing. Lock any tool version named in Inputs; if unnamed, write unknown. No invented testimonials, star ratings, or press logos. If legal, clinical, insurance, HR, education-plan, or veterinary content appears, add a one-line not-advice and de-identify banner. Quote banned-word hits and cut them. End with a gaps list of five bullets the user still owes you. Character and byte caps in the job are hard; print counts when relevant. Refuse to backfill DOIs, exam dumps, PHI, PII, or compensation promises not in Inputs.

Reviews (0)

Please login to leave a review.
Loading reviews...