🤖 AI Tools
llama.cpp GBNF Grammar Writer: Force a Local Model to Return Valid Invoice Extraction JSON
Write a GBNF grammar for llama.cpp that forces a local GGUF model to emit JSON matching your invoice extraction schema: fixed key order, bounded strings and whitespace, enums, dates, nullable fields, and the commands to run it with llama-cli or llama-server, plus the validation that still has to happen in code.
0Reviews
Prompt
Act as a machine learning engineer who runs local GGUF models with llama.cpp in document extraction pipelines and writes GBNF grammars by hand. You know a grammar controls the shape of the output, not whether the values are true, and you design both halves.
Inputs:
- Target JSON schema or a sample of the exact JSON wanted: [JsonSchema]
- Field rules (required, nullable, enums, date format, max lengths, max array items): [FieldRules]
- How llama.cpp is run (llama-cli, llama-server /completion, or the OpenAI compatible endpoint) and the build date if known: [LlamaCppSetup]
- Two short sample documents the model will read: [SampleDocs]
- Downstream code that consumes the JSON (language, validator): [Consumer]
- Output format: [Format]
- Language: [Lang]
Generate:
1. Grammar design notes: fixed key order from JsonSchema, which fields are nullable, how enums become literal alternatives, and why strings, arrays, and whitespace get upper bounds (an unbounded ws rule lets the model pad forever).
2. The .gbnf file. Start with root ::= and define rules for objects, arrays, string, number, integer, date, enums, and ws. Write literal quotes as "\"" and use a character class that excludes quote, backslash, and control characters for string content. Use {m,n} repetition for bounds and note that older builds may not support it.
3. Run commands for LlamaCppSetup: llama-cli with --grammar-file, or the grammar field in a llama-server /completion request body. Mention --json-schema and the json_schema_to_grammar.py script as the alternative when the schema changes often, and say to confirm paths in the current repo.
4. Prompt text that goes with the grammar: a short instruction plus the document, telling the model to use null when a field is not on the page.
5. Validation in Consumer: JSON parse, schema validation, and business checks the grammar cannot do (line items times price against total, date is a real calendar date).
6. Test plan against SampleDocs: what to look for, including a document with a missing field to confirm null is produced instead of a guess.
Constraints:
- The grammar must accept every valid output and nothing else; walk through one sample output against the rules.
- Do not claim the grammar improves accuracy; it only guarantees parseable structure.
- No em dashes.Instructions
Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.
Generated Output
This image was generated using the prompt above.

Examples
Example Input
JsonSchema: {"vendor": string, "invoice_number": string or null, "invoice_date": YYYY-MM-DD, "currency": USD|EUR|GBP, "total": number, "line_items": [{"description": string, "qty": integer, "unit_price": number}]}
FieldRules: strings max 120 chars, at most 30 line items, money with up to 2 decimals, no negative values in this pipeline
LlamaCppSetup: llama-server on a workstation, POST to /completion from Python
SampleDocs: a print shop invoice with three lines; a utility bill with no invoice number
Consumer: Python 3.12 with pydantic
Format: notes, grammar, commands, checks
Lang: EnglishExample Output
1. Design notes
- Keys appear in schema order, so the grammar hardcodes that order; your parser does not care, and fixed order makes the grammar small.
- invoice_number is the only nullable field.
- currency becomes three literal alternatives.
- Strings are capped at 120 characters, line items at 30, and whitespace at one optional newline plus a little indentation.
2. invoice.gbnf
root ::= "{" ws "\"vendor\":" ws str "," ws "\"invoice_number\":" ws (str | "null") "," ws "\"invoice_date\":" ws date "," ws "\"currency\":" ws currency "," ws "\"total\":" ws money "," ws "\"line_items\":" ws items ws "}"
items ::= "[" ws (item ("," ws item){0,29})? ws "]"
item ::= "{" ws "\"description\":" ws str "," ws "\"qty\":" ws int "," ws "\"unit_price\":" ws money ws "}"
str ::= "\"" char{0,120} "\""
char ::= [^"\\\x00-\x1F] | "\\" ["\\/nt]
date ::= "\"" [0-9]{4} "-" [0-1] [0-9] "-" [0-3] [0-9] "\""
currency ::= "\"USD\"" | "\"EUR\"" | "\"GBP\""
int ::= "0" | [1-9] [0-9]{0,5}
money ::= int ("." [0-9]{1,2})?
ws ::= (" " | "\n" " "{0,4})?
Walkthrough: {"vendor": "Ink & Co", "invoice_number": null, ...} matches root because null is an allowed alternative and the space after each colon is one ws.
3. Run
Python request body to llama-server:
{"prompt": PROMPT, "grammar": open("invoice.gbnf").read(), "n_predict": 1024, "temperature": 0}
CLI check: llama-cli -m model.gguf --grammar-file invoice.gbnf -f prompt.txt
If the schema changes often, generate the grammar from JSON Schema with json_schema_to_grammar.py or pass --json-schema; confirm both in your current llama.cpp checkout.
4. Prompt
Extract the invoice below as JSON. Use null for invoice_number if it is not printed. Copy amounts exactly as printed. Document:
<document text>
5. Validation in Python 3.12
- json.loads, then a pydantic model with invoice_date as datetime.date (rejects 2026-02-31, which the grammar allows).
- Check sum(qty * unit_price) against total within 0.01; on mismatch, route to human review instead of fixing it.
- Reject vendor values that are empty strings.
6. Tests
- Print shop invoice: expect three line_items and an invoice_number string.
- Utility bill: expect invoice_number null. If the model writes a made-up number, the grammar is doing its job and the prompt needs the null instruction moved closer to the document.
- Add a 31 line invoice to confirm generation stops at 30 items and your code flags it as truncated.