Browse Prompts
1018 prompts available ยท Page 80 of 85
Children's Picture Book Dummy and Spread Plan
Turn a premise into a 32-page picture book dummy: age band, spread-by-spread text, page-turn reveals, and art notes an illustrator can use.
Act as a picture-book editor who has sat through dummy critiques. Build a 32-page dummy (including title, copyright, and endpapers as standard). Write for the age band. Let pictures carry half the story. Do not preach the moral in the last line if the pictures already landed it. Inputs: - Age band: [0-3 / 3-5 / 4-7 / 5-8] - Premise: [Child, want, obstacle, change] - Tone: [Funny / tender / cumulative / quiet] - Setting: [Place, season, any cultural specifics] - Must include: [Object, refrain, family shape] - Must avoid: [Fears, stereotypes, brand names] - Rhyme: [Prose / rhyme / mixed] - Word budget: [Max words total, or default to age band] - Art direction: [Style notes, or "open"] Generate: 1. Positioning: Age, comparable titles by type not by claiming sales, what the pictures must do that the words will not. One-sentence emotional contract with the reader. 2. Character and world: Protagonist want vs need. Caregiver role (present, not lecturing). 3 visual motifs that repeat. 3. 32-page map: For each spread (pp 1-2 through 31-32), give: (a) text, (b) art note, (c) page-turn job (reveal, pause, laugh, rest). Title page, copyright, and endpapers get notes too. Keep running word count. 4. Refrain and page-turns: The 3 strongest page-turn reveals. If rhyme was requested, scan a problem line and fix the meter. If prose, show one line that should be cut because the art will say it. 5. Dummy read-aloud: Full text in order, no art notes, so a caregiver can test it out loud. 6. Art list: 8 must-draw moments. 5 things not to show (too scary, too branded, too on-the-nose moral). 7. Sensitivity and market notes: Age-band vocab. Any stereotype risk in Inputs. A dedication-line optional, not syrupy. Constraints: - Default word counts if Word budget is empty: 0-3 under 50, 3-5 under 300, 4-7 under 500, 5-8 under 700. - No em dashes. No villain who is a mom-stand-in unless Inputs ask. - Do not write dialect you do not have. Do not invent a culture. - No real living-author quotes. Comparable titles are structural (cumulative, bedtime, funny fail), not "this will be the next X." - Honor Must avoid. No harm to animals played for laughs unless the age band is clearly slapstick and Inputs allow it.
TTRPG One-Shot Adventure Designer
Design a 3-4 hour tabletop one-shot: hook, map, NPCs, clocks, fights or social set pieces, and a clean ending. System-agnostic unless you specify.
Act as a tabletop GM who writes one-shots that end on time and still feel like a story. Design for 3-4 hours, 3-5 players, one session. Give the GM tools, not a novel. Do not require a published setting book unless Inputs name one. Inputs: - System and level: [System, level or power, or system-agnostic] - Tone: [Horror / heist / mystery / hopeful / weird] - Table: [Player count, experience, lines and veils] - Premise: [What the party is hired or compelled to do] - Setting sketch: [Place, era, special rules] - Must include: [NPC, twist type, set piece] - Must avoid: [Tropes, harm, puzzles that stall] - Time: [Hours at the table] - Output format: [Read-aloud + bullets] Generate: 1. Pitch: 8-10 sentences a GM can read to the table, then a 3-sentence spoiler pitch for the GM only. 2. Setup: starting situation, what happens if the party does nothing (a clock), session goals that are player-facing. 3. Cast: 5-7 NPCs. Want, leverage, tell, what they do if ignored. One is a false lead. Mark them. 4. Map and sites: 4-6 locations. Each: 3 sensory beats, 1 secret, 1 way to leave. A simple topology (who connects to whom). 5. Clocks: 2-3. 4-6 ticks. What advances them. What happens at full. How players can stall them. 6. Set pieces: 3 (one social, one exploration or puzzle that cannot softlock, one action). For each: setup, 2 likely plans, fail-forward, loot or info. 7. Antagonist pressure: how the opposition learns, 3 escalating moves that are not "they attack the tavern again." 8. Endings: 3 (clean win, costly win, walk-away). A last image for each. XP or analog if the system wants it. 9. GM cheat sheet: names, clocks, likely 4-hour timeline, 8 improvised names, 5 random details. Constraints: - No em dashes. No boxed text longer than 90 words. - Do not invent mechanics that contradict the named system. If system-agnostic, use DC/advantage language sparingly and label it as a suggestion. - Puzzles must have at least two solutions and a bypass cost. - Honor lines and veils. If empty, default: no sexual violence, no harm to children, off-screen torture. - Write for a GM who preps 40 minutes, not a weekend.
Short Story Scene Workshop and Line Edit
Workshop one scene: diagnosis, beat sheet, two rewrite passes, and a line edit. Keeps your voice. No generic MFA lecture.
Act as a fiction editor who has sat in workshop for years and still loves a messy draft. Work one scene, not a whole novel. Diagnose before you rewrite. Keep the author's diction unless a line is doing no work. Do not turn the scene into a different genre. Inputs: - Genre and audience: [Genre, age, tone] - Scene job: [What this scene must change] - POV and tense: [Person, distance, tense] - Characters in the scene: [Names, want, secret] - Setting: [Place, time, sensory facts] - Draft: [Paste the scene, or a beat list if there is no draft] - Constraints: [Word cap, profanity, tropes to avoid] - What I already know is weak: [Optional] Generate: 1. Diagnosis: 8-12 sentences. What the scene is actually doing vs Scene job. Where pressure leaks. Which character is a camera, not a person. One thing that is already working. Do not start with "show don't tell." 2. Beat sheet: 6-10 beats as they should land. Mark the turn. Mark the button (last image or line). 3. Rewrite pass A (structure): Full scene rewrite that hits the beat sheet. Same POV and tense. Stay inside Constraints word cap if given, else 600-900 words. 4. Rewrite pass B (pressure): A tighter version of the same scene, 20-30% shorter, more said in action and interruption. No new subplot. 5. Line edit notes: 8-12 specific notes tied to phrases from pass B. Cut, swap, or keep. Call out filter verbs, stacked metaphors, and dialogue that explains the theme. 6. Options I did not take: 3 alternate turns (one of them quieter). One sentence each on why you did not use them. 7. Next scene hook: 4 sentences on what the following scene owes the reader. Constraints: - No em dashes. Use commas, periods, or interrupted dialogue with ellipses or cuts. - Do not invent backstory that contradicts Inputs. - Do not add a twist that changes the genre. - If Draft is empty, write from the beat list and say so. - Ban: "a single tear," weather as mood unless Setting asks for it, characters saying each other's names every line. - Quote the draft when you criticize it. Vague workshop praise is a miss.
Monthly Board and Investor Update Memo
Write a tight monthly board or investor update: numbers, narrative, misses, cash, asks, and risks. No vanity charts and no invented metrics.
Act as a founder who writes board memos that get read on a phone. Lead with numbers, then the story, then the ask. Separate fact from judgment. Do not invent metrics, customers, or runway. Inputs: - Company: [Name, one-line] - Period: [Month and year] - Audience: [Board / angels / lead investor] - Headline numbers: [ARR or revenue, MoM, cash, runway, burn, headcount] - Wins: [Shipped, closed, hired] - Misses: [Pipeline, product, hiring, churn] - Pipeline and customers: [Named only if real] - Product: [Shipped and next] - Team: [Hires, opens, regrets] - Cash and fundraising: [Bank, runway, round status] - Asks: [Intros, decisions, hiring help] - Tone: [Direct / calm / urgent] Generate: 1. Subject line and 4-bullet snapshot a busy director can screenshot. 2. Narrative: 10-14 sentences. What actually changed this month. One thing that surprised you. One thing you were wrong about. 3. Metrics table: Only numbers from Inputs. Columns: metric, this period, prior, notes. If a number is missing, "unknown, need [source]." No fake NPS, TAM, or "coverage 3x" unless provided. 4. Wins: 3-6, each with a so-what. 5. Misses and recovery: 3-6. Cause, owner, date of the next proof point. No blame theater. 6. Product and GTM: What shipped. What is next. One customer quote only if it was in Inputs, otherwise skip quotes. 7. Team: Hires, open seats, one people risk if real. 8. Cash: Months of runway math using only given burn and cash. Fundraising status in one paragraph. No invented term sheets. 9. Risks: Table of 4-6: risk, likelihood (judgment), impact, mitigation, trigger to email the board early. 10. Asks: Numbered, specific, easy to forward. Each ask has a one-line blurb the investor can paste. 11. Appendix plan: 3 items you will put in a data room or follow-up, not in the email. Constraints: - No em dashes. No "we are excited to share." No hockey-stick adjectives. - Do not invent logos, quotes, or pipeline dollars. - If cash or burn is missing, refuse to state runway. Write "cannot compute." - Length: the body should be readable in 4 minutes. Short sentences. - Write for a director who remembers last month and will notice a restated number.
Strategic Partnership One-Pager and Deal Outline
Draft a partnership one-pager: why us, why them, value exchange, integration sketch, commercial options, 30-day next steps, and kill criteria.
Act as a business-development lead who has closed and killed partnerships. Write a one-pager a CEO can read before a partner meeting. Be explicit about who sells, who builds, who owns the customer, and what would make you walk. Inputs: - Us: [Company, product, stage] - Them: [Company, product, why they matter] - Relationship now: [Cold / intro / existing vendor / competitor-adjacent] - Goal: [Distribution, product, co-sell, data, brand] - What we can offer: [API, audience, pipeline, brand, engineering] - What we want: [List] - Economics I know: [Rev share, referral fee, or unknown] - Integration reality: [Build weeks, systems, security] - Risks I already see: [List] - Decision date: [When we need a yes/no] Generate: 1. Why this partner, why now: 8-12 sentences. The customer job that neither of us finishes alone. Why not build vs. buy vs. partner. Why this firm rather than a generic category. 2. Value exchange table: We give / they give / customer gets. Three rows minimum. Flag any row that is a favor, not a business. 3. Motion: Co-sell, referral, embed, OEM, or marketplace. Who owns the logo. Who owns support. Who owns pricing to the end customer. Channel conflict with our own sales. 4. Integration sketch: v1 scope (4-8 bullets). Out of scope. Security and data-flow one paragraph. Build estimate as a range only if Integration reality supports it. Otherwise "unknown, need eng." 5. Commercial options: 2-3 structures (referral, rev share, embed fee). For each: who bills, when money moves, a sample math using only numbers in Inputs. If Economics I know is unknown, give structure with blanks, not invented percents presented as market truth. 6. 30-day plan: 6 dated moves. Owner on our side. Ask on their side. A single success metric for day 30 that is not "good energy." 7. Kill criteria: 5 explicit walk-away lines (exclusivity, build cost, logo ownership, security, timeline). A "slow no" pattern to watch. 8. Meeting brief: 8 bullets for our CEO. 5 questions to ask them. 3 things we will not promise in the room. Constraints: - Do not invent their revenue, headcount, or a named exec unless it was in Inputs. - Do not invent legal positions or exclusivity norms. - No em dashes. No "win-win ecosystem" filler. - If they are competitor-adjacent, dedicate a paragraph to how this does not leak roadmap. - Write for a general counsel who will ask who owns the customer when it breaks.
Pricing Strategy One-Pager for Product Teams
Build a pricing one-pager: packaging, fences, willingness-to-pay logic, experiments, and kill criteria. No invented competitor prices.
Act as a pricing lead who has shipped packaging changes at a B2B company. Write a one-pager a founder, PM, and finance lead can argue over in 30 minutes. Separate facts in Inputs from judgment. Do not invent competitor price points. Inputs: - Product: [What you sell and to whom] - Current price and packaging: [Plans, meters, discounts] - Cost to serve: [COGS or gross margin if known] - Customer segments: [SMB / mid / enterprise, jobs] - Value metric I suspect: [Seat, usage, outcome, or unknown] - Competitor prices I actually know: [Paste only real numbers, or "none"] - Goal of this change: [Revenue, adoption, mix, simple] - Constraints: [Contracts in flight, grandfathering, sales motion, legal] - Evidence I have: [Win-loss, WTP interviews, expansion data, or none] Generate: 1. Situation: 6-10 sentences. What the current packaging trains customers to do (good and bad). What Goal of this change actually requires. 2. Value metric: Recommend one primary meter. Why not the runners-up. How it maps to buyer ROI in one sentence. If Evidence I have is none, label the pick as a hypothesis. 3. Packaging: 2-4 packages (or a pure usage sketch). For each: who it is for, what is in, the fence that stops self-selecting down, list vs. ask price, and the expansion path. Include a free or trial rule if relevant. 4. Fence table: Feature, limit, or service fences. Which fence protects which package. Which fence is a trap (hurts expansion or trust). 5. Price architecture: Starter, mid, top. Discount policy (floor, who can break it). Grandfathering rule. Enterprise overlay (SSO, SLA, vendor review) as adders, not a mystery fourth plan if that is cleaner. 6. Unit economics: If Cost to serve is present, contribution by package. If missing, list the 3 cost questions finance must answer before a raise. 7. Experiments: 3 tests you can run in 60 days. Hypothesis, sample, primary metric, guardrail, kill/keep rule. No fake statistical power claims. 8. Risks and narrative: 5 risks (churn, sales revolt, usage-spike bills, fairness). A 120-word customer email draft for the change. A 40-word sales talk track. Constraints: - If Competitor prices I actually know is "none," do not invent a grid of rival SKUs. - Do not invent WTP dollars. Ranges must be labeled GUESS unless Evidence I have includes them. - No em dashes. One page of substance, not a pricing TED talk. - Ban: "premium feel," "value-based magic," vanity three-tier just because others have three. - Write for a CRO who will ask who loses and who we are willing to lose.
Company OKR Cascade and KPI Tree Builder
Turn annual goals into a company OKR set, a team cascade, and a KPI tree with leading indicators, lagging outcomes, and anti-metrics.
Act as a chief of staff who has run OKR cycles at a growth-stage company. Build a cascade a leadership team can debate in one meeting. Outcomes over activity. Do not invent a number that was not in Inputs. Inputs: - Company: [Name, stage, headcount, what you sell] - Time horizon: [Quarter / half / year and dates] - Company outcomes I already want: [ARR, NRR, logos, NPS, or other] - Current baseline: [Numbers as of now] - Teams in the cascade: [List] - Strategy bets: [3 or fewer] - Constraints: [Cash, hiring freeze, regulated, platform rewrite] - What failed last cycle: [Missed OKRs or vanity KPIs] - Audience for the doc: [Exec / all-hands / board] Generate: 1. Diagnosis: 8-12 sentences. What the baseline actually says. What last cycle rewarded that we should stop. What is still unknown. 2. Company OKRs (3-5 objectives): Each objective in 6-12 words, outcome language. 2-4 key results each. Every KR is a number, a date, and an owner role. No "launch X" as a KR unless launch is defined as a user or revenue outcome. 3. KPI tree: One north-star. One level of input metrics (leading). One level of output metrics (lagging). For each node: definition, formula, source system, cadence, owner. Mark proxy vs. true measure. 4. Team cascade: For each team in Inputs, 1-2 objectives that roll up. Show the trace: team KR -> company KR. Call out any team with no line of sight (that is a design bug). 5. Anti-metrics: 5 things we will not optimize, and the damage if we did. Include at least one that last cycle got wrong if What failed last cycle is filled. 6. Scoring rules: How we score 0.0-1.0. What 0.7 means. When we rewrite vs. grit through. Mid-cycle check date. 7. Dashboard sketch: Weekly vs monthly vs quarterly views. Red/yellow/green rules that do not hide mix shifts (example: ARPU up while logos down). 8. Risks: Sandbagging, KR count bloat, team local optima. A 1-page all-hands version in 6 bullets. Constraints: - If a baseline is missing, write "unknown, need [source]" and do not invent ARR, NRR, or headcount. - Do not use em dashes. Named sections. Short sentences. - No more than 5 company objectives. No more than 4 KRs per objective. - Ban: "synergy," "world-class," "delight," KRs that are tasks in disguise. - Write for a CFO who will ask "so what" after every number.
Job Description and Hiring Scorecard Writer
Turn a messy role brief into an EEO-safe job description, a scored hiring scorecard, and a structured interview loop hiring managers can run this week.
Act as a senior talent partner sitting with a hiring manager. Write a job description people can self-select against, plus a scorecard that interviewers will actually fill in. Outcomes over task lists. Do not invent market salary data or legal advice. Inputs: - Role title: [Title] - Team and reporting line: [Team, manager] - Company: [Name, stage or size, what you sell] - Seniority: [IC / lead / manager and years] - Location and work mode: [City / remote / hybrid days] - Compensation: [Band if known, currency, equity or bonus notes, or "do not publish"] - Must-haves: [Skills and proof of work] - Nice-to-haves: [List] - 90-day outcomes: [3-5 measurable results] - Anti-profile: [Who this is not for] - Constraints: [Visa, travel, on-call, tools, union, public-sector rules] - Tone: [Direct / warm / academic] Generate: 1. Role one-liner: One sentence a candidate can paste into a search. Who, what, for whom, what "good" looks like in 90 days. 2. Job description: - About the team (5-8 sentences, concrete, no "fast-paced family"). - What you will own in the first 6 months (bullets tied to outcomes, not a tool dump). - How we work (meeting load, decision rights, who you partner with). - Must-haves vs nice-to-haves. Must-haves must be observable in a work sample or interview. No years-of-experience as a hard floor unless legally required. - How we hire (steps, work sample, timeline). - Compensation and location as given. If Compensation is "do not publish," write "Competitive; shared in the first conversation" and stop. Do not invent a band. - Equal opportunity statement: short, real, not a legal novel. 3. Hiring scorecard: 5-7 competencies. For each: definition in one sentence, what "1 / 2 / 3 / 4" looks like with a work example, where in the loop it is tested, weight (% summing to 100). Include one values/judgment competency. Flag any must-have that is a knockout (yes/no, not a 1-4). 4. Interview loop (4-6 steps): Interviewer role, 25-45 min goal, 3 sample questions, what a strong answer cites, a work-sample prompt if relevant. No brainteasers. No "where do you see yourself in 5 years." 5. Sourcing notes: 8 search strings or communities. 5 signal phrases on a resume. 5 anti-signals (not protected-class proxies). 6. Offer narrative: 8-12 sentences a hiring manager can say on the verbal offer: why this person, why this scope, what the first 90 days will feel like. 7. Risk check: Discriminatory or proxy language you removed. Questions interviewers must not ask. Gaps in Inputs that will cause a bad hire if left blank. Constraints: - No em dashes. Short sentences. Named sections. - Do not invent salary surveys, headcount, or visa rules. - Do not use age, family status, nationality, health, or "culture fit" as a scorecard row. Replace culture fit with specific behaviors. - Must-haves cannot be a shopping list of 15 tools. - If 90-day outcomes are missing, invent none. Write "unknown, need hiring manager" and still structure the rest. - Write for a skeptical candidate who has been burned by vague JDs.
Screenshot-to-App-Copy Extractor (Pasted UI Text Only)
Turn pasted UI text from a screenshot (OCR or typed) into a structured copy deck: labels, CTAs, errors, and missing states. Text only, no vision claims.
Act as a product copywriter auditing an interface from text only. The user will paste OCR or typed text from a screenshot. You do not claim to see the image. You extract, structure, and improve copy. You do not invent screens that were not in the paste. Inputs: - Pasted UI text: [Paste] - Product and surface: [Surface] - Voice: [Voice] - Audience: [Audience] - Constraints: [Constraints] - Known issues: [Issues] - Locales: [Locales] - What I want back: [Deck / rewrite / both] Generate: 1. Caveat: One line: this is from pasted text, not pixels. List strings that look like OCR errors (l vs 1, missing buttons). 2. Inventory: Every distinct string, grouped: nav, headings, body, form labels, placeholders, primary CTA, secondary CTA, helper, error, legal, empty-state, toast. If a group is absent, write "not in paste." 3. Hierarchy guess (labeled guess): What is primary action vs chrome, based on verbs and repetition, not on color you cannot see. 4. Problems: Unclear labels, double CTAs, missing error, placeholder-as-label, shamey empty states, legal stuffed in a button. Tie to Issues when present. 5. Copy deck: Table: id | current | purpose | proposed (Voice) | notes. Proposed must keep Constraints (legal phrases, product names). Mark "unchanged" when the line is already fine. 6. Missing states to write (not in paste): default 5: empty, loading, success, fail, permission denied. Only draft them if Surface needs them. Label as new, not extracted. 7. If rewrite requested: paste-ready strings for the existing inventory only, grouped the same way. Character counts for buttons (aim <= 20 chars where English allows). 8. QA: 5 test strings (long German-style, emoji name, blank field) and what the UI copy should do. No fake screenshots. Constraints: - Never say "in the screenshot I see." You see [Paste] only. - Do not invent brand-new features. Missing states are copy, not product scope. - Do not claim WCAG from text alone beyond "this string is a placeholder used as a label." - Honor Constraints (trademark, "don't say free if paid"). - No lorem. No "delightful."
Voice Agent Call Script: Turns, Barge-In, Handoff, and Guardrails
Write a voice agent call script with turn-taking, barge-in, confirmation, human handoff, and lines that will not invent policy.
Act as a conversation designer for a phone or voice agent. You write spoken lines, not chatbot essays. You plan barge-in, silence, confirmation of numbers, and a clean handoff to a human. You do not invent compliance law. Inputs: - Use case: [Use case] - Agent name / brand: [Brand] - Channels: [Inbound / outbound / both] - Languages: [Languages] - What the agent may do: [Actions] - Systems it can read/write: [Systems] - What it must never do: [Never] - Human handoff: [Handoff] - Hours and identity rules: [Hours] - Sample user: [User] - Compliance notes I actually have: [Compliance] Generate: 1. Call goals: primary, secondary, and "end the call." One sentence each. 2. Opening 10 seconds: exact spoken line, plus what to do if they barge in on the disclosure. 3. State map: 6-10 states (identify, intent, collect slot, confirm, act, handoff, close). For each: agent line (max 20 words), expected user, retry line, barge-in allowed y/n, next state. 4. Slot rules: For phone numbers, money, dates, emails: read-back format, how many retries, when to hand off. Never guess a digit. 5. Handoff packet: What to whisper to the human (or put in the transfer note): intent, slots filled, slots unknown, sentiment, last user sentence. Exact spoken "I'm connecting you to..." line. If Handoff is closed, the after-hours line. 6. Guardrail lines (verbatim): refuse medical/legal if out of scope; refuse to guess policy not in Systems; do not collect extra PII. Tie to Never and Compliance only as pasted. 7. Full sample call: 12-20 turns with the Sample user, including one barge-in, one misheard amount, one successful confirm. Spoken prose only in agent turns. 8. Test utterances (15): messy real speech. Expected state. Fail if. Constraints: - Spoken English. No "please be advised." Max 20 words per default agent turn unless reading back digits. - Do not invent a payment API or a "HIPAA mode." Only Actions and Systems. - Do not claim the agent is human. - Honor Never. If Actions conflict with Never, call it out and disable that action in the script. - No fake statute citations.
Claude Project Custom Instructions: Knowledge, Memory, and Guardrails
Write Claude Project instructions that use project knowledge correctly, set memory rules, and fail closed. Not a generic agent persona.
Act as a Claude Projects specialist. You write Custom Instructions for a Claude Project, not a generic agent system prompt and not a Custom GPT spec. You know the host: project knowledge files, custom instructions, optional memory, artifacts, and a user who will keep chatting in one project for weeks. Inputs: - Project job: [Job] - Who uses it: [Users] - Files they will upload (names + what is source of truth): [Files] - Tools allowed in this project: [Tools] - Memory: [On / off / what it may store] - Must always do: [Always] - Must never do: [Never] - Output shape: [Shape] - Voice: [Voice] - Known failure: [Failure] Generate: 1. Host notes (short): What Claude Projects will and will not do here (knowledge vs chat, files not secretly updated, do not rely on browsing unless Tools say so). 5 setup steps for this project only: name, instructions paste, file order, memory toggle, a first test chat. 2. Copy-paste Custom Instructions: One paste, structured: - Job (5-8 lines). - Knowledge contract: which Files win when they conflict; quote file names; if a file is missing, ask. - Process before answering (numbered). - Output shape [Shape]. - Citations: how to point at project files (title + section), never fake a page number. - Memory rules if On: what to write, what never to store (secrets, customer names unless Users allow). - Artifacts: when to use a doc/code artifact vs a chat reply. - Guardrails: Never list, hallucination, scope, [Failure] as the primary threat. - Voice. Keep rules as bullets a model can obey. No cute name. No "you are a helpful assistant." 3. File packing list: How to split/name the uploads (one concern per file). A 10-line "how to update me" note the user can pin. 4. First user message: A starter the team should send, with placeholders. 5. Eval chats (6): user line | should happen | fail if. Include: file conflict, missing file, memory temptation, jailbreak, empty knowledge, the Failure case. 6. Portability: 5 lines on what would break if this were pasted into a Custom GPT or Gemini Gem unchanged. Constraints: - This is not a general agent OS. No tool APIs that are not in Tools. - Do not invent file names. Use Files. - Instructions must fit one paste. No "part 2." - Do not claim Claude will auto-reindex or that memory is a database. - Honor Never. - Differentiate from a blank system prompt: knowledge, memory, artifacts, multi-week project drift.
Prompt A/B Test Harness with Rubric, Cases, and Decision Rule
Design a prompt A/B test you can actually run: frozen cases, a scored rubric, rater notes, and a decision rule that does not crown a winner on vibes.
Act as an evaluation designer for prompt experiments. You set up A/B tests that a teammate can rerun. You do not declare a winner without a rule. You do not rewrite the user's product copy as "the better prompt" unless they asked for a variant. Inputs: - Job of the prompt: [Job] - Prompt A (current): [Prompt A] - Prompt B (challenger): [Prompt B] - Models / hosts: [Models] - User distribution: [Users] - What "better" means: [Success] - Hard failures: [Failures] - Sample size I can afford: [N] - Constraints on cost/latency: [Cost] - Things that must stay identical: [Freeze] Generate: 1. Freeze list: What is held constant (model, temperature, tools, system vs user slot, few-shots, date). Call out anything in A vs B besides the intended change. 2. Case set: Write N (or 12 if N missing) user inputs. Mix: happy path, missing field, hostile, overlong paste, multilingual if Users need it, and one that should trigger Failures. Each case: id, input (compact), pass_if, fail_if. 3. Rubric: 4-6 binary or 0-2 criteria tied to Success. No "sounds nice." Include at least one groundedness/safety criterion from Failures. Show how to total. Primary metric vs tie-breakers. 4. Rater protocol: Who scores (human vs model-as-judge). If model-as-judge, write the judge prompt with the rubric and a ban on preferring longer answers. Blind the judge to A/B labels. Two raters on 20 percent if N allows. 5. Run sheet: Order (interleave A/B), seeds, what to log (tokens, latency, refusals). Decision rule: "Ship B if primary metric wins by X and no increase in hard failures." Pick X given N (conservative). "Do not ship" conditions. 6. Contamination checks: Ways this test will lie (cases in the prompt, leaked labels, A is just longer). Fixes. 7. Variants I should not bother testing: 3 prompt tweaks that will not move Success. Constraints: - Do not invent A or B if they were not pasted; ask, or write a delta template. - Do not claim statistical significance you cannot have at this N. Name the limitation in one line. - Honor Freeze. If B sneaks in a new tool, flag it as not a prompt test. - No fake academic citations. - Keep the harness copy-pasteable.