Back to Discover

#ai safety

1 prompt found

Llama Guard 3 Safety Policy Builder: Kept and Custom Hazard Categories, User and Agent Classification Prompt Templates, Allowed vs Unsafe Examples, and an App Eval Set
🤖 AI Tools

Llama Guard 3 Safety Policy Builder: Kept and Custom Hazard Categories, User and Agent Classification Prompt Templates, Allowed vs Unsafe Examples, and an App Eval Set

Ppromptstudio·Oct 9, 2026
No rating

Set up Llama Guard 3 as the input and output safety classifier for your app: choose which default S1 to S14 categories to keep, write custom categories in the same style, build the exact classification prompts for user messages and model responses, and test them with allowed and unsafe examples and a small labeled eval set.

Act as an ML safety engineer who deploys Llama Guard 3 as an input and output classifier for production chat apps, writes category policies the model can follow, and tests them before launch. Inputs: - App description: who uses it, what it does, ages of users: [AppDescription] - Default Llama Guard 3 categories to keep from S1 to S14, and any to drop with a reason: [DefaultCategories] - Custom categories your app needs, in plain words: [CustomCategories] - Edge case messages you are worried about, with what should happen: [EdgeCases] - Where the classifier runs: user prompts, agent responses, or both: [DeploymentPoint] - Output format: [Format] Generate: 1. A final category list for Llama Guard 3: kept defaults with their S codes and names, and custom categories numbered after S14, each with a one or two sentence definition of what is unsafe and what is allowed. 2. The full classification prompt template for user messages: the task line, the BEGIN and END UNSAFE CONTENT CATEGORIES block, the BEGIN and END CONVERSATION block, and the instruction that the first line is safe or unsafe and the second line lists violated categories. 3. The matching template for agent responses if DeploymentPoint includes them, changing the role being judged. 4. Allowed and unsafe example pairs for each custom category from EdgeCases, with the expected label output. 5. A 12 row labeled eval set covering kept and custom categories, with expected output strings. 6. App handling rules: what the app does on unsafe for each category (block, rewrite, or route to a human), and logging. 7. Tuning notes: testing custom categories against the eval set, and using the probability of the unsafe first token as a threshold if needed. Constraints: - Use the official S1 to S14 names for kept categories; do not rename them. - Custom category definitions must say what is allowed as well as what is unsafe, to cut false positives. - No invented accuracy numbers. No em dashes.