Home/Blog/How to Use the Llama Guard 3 Safety Policy Prompt to Set Up Moderation for Your App
Blog

How to Use the Llama Guard 3 Safety Policy Prompt to Set Up Moderation for Your App

P
promptstudio

Set up Llama Guard 3 as the input and output safety classifier for your app: choose which default S1 to S14 categories to keep, write custom categories in the same style, build the exact classification prompts for user messages and model responses, and test them with allowed and unsafe examples and a small labeled eval set.

How to Use the Llama Guard 3 Safety Policy Prompt to Set Up Moderation for Your App

Llama Guard 3 is a classifier model that reads a conversation and returns safe or unsafe, plus the categories that were violated. It ships with fourteen default hazard categories, S1 through S14, and it accepts a custom category list in its prompt. That flexibility is useful, but it means the quality of your moderation depends on how well you write the category policy and how well you test it. The Llama Guard 3 Safety Policy Builder: Kept and Custom Hazard Categories, User and Agent Classification Prompt Templates, Allowed vs Unsafe Examples, and an App Eval Set prompt helps you choose categories, write custom ones, build the exact classification prompts, and test them before launch.

What the prompt produces

  1. A final category list with kept defaults under their official names and custom categories numbered after S14, each with a definition of what is unsafe and what is allowed.
  2. A user message classification template with the task line, the unsafe content categories block, the conversation block, and the output instructions.
  3. An agent response template for checking model outputs.
  4. Allowed and unsafe example pairs for each custom category.
  5. A twelve row labeled eval set with expected outputs.
  6. App handling rules for each category, such as block, rewrite, or route to a human.
  7. Tuning notes for testing custom categories and using a probability threshold.

How to fill the inputs

AppDescription explains who uses the app, what it does, and the ages of users. A tool for adults and a tool for middle school students need different policies.

DefaultCategories lists which of S1 to S14 to keep and which to drop, with reasons. Dropping categories that rarely apply can reduce false positives.

CustomCategories describes, in plain words, the extra risks your app has.

EdgeCases lists messages you are worried about and what should happen with each one. These become test examples.

DeploymentPoint says whether you are checking user prompts, model responses, or both.

Reading the example output

The example sets up moderation for a homework help chatbot used by students aged eleven to fourteen. It keeps nine default categories and adds two custom ones: graded test completion, and off platform contact.

Each custom category says what is allowed as well as what is unsafe. Graded test completion blocks giving full answers to a quiz the student says is in progress, but allows explaining concepts and checking the student's own work. That allowed line matters, because without it a classifier can start flagging ordinary homework help.

The templates follow the Llama Guard 3 format, with the category block, the conversation block, and the instruction that the first line reads safe or unsafe. The eval set includes history topics that should stay safe, which tests whether the hate category is too aggressive. The handling rules treat self harm differently from other categories: instead of a silent block, the app shows crisis resources and alerts a counselor queue.

Tips for better results

  • Write allowed examples for every custom category, not only unsafe ones.
  • Run the eval set before launch and after every policy change.
  • Log category codes so you can see which rules fire most.
  • Review misses with someone who knows your users.

Mistakes to avoid

  • Do not rename the default categories.
  • Do not assume custom categories work as well as the defaults. Test them.
  • Do not quote accuracy numbers you have not measured.

Who it is for

ML engineers adding moderation to chat apps, trust and safety teams, and developers building tools for schools or other sensitive audiences.

A practical routine

Start with the eval set, launch with logging on, and review flagged and missed messages weekly. Add new cases to the eval set and rerun the prompt when your policy changes.

Related PromptDig links

Build your policy with the Llama Guard 3 Safety Policy Builder: Kept and Custom Hazard Categories, User and Agent Classification Prompt Templates, Allowed vs Unsafe Examples, and an App Eval Set prompt. For more AI tool and model prompts, Browse more prompts. If you have a safety testing workflow that works, Share a prompt.