Home/Blog/How to Use the llama.cpp GBNF Grammar Writer Prompt for Reliable Invoice JSON
Blog

How to Use the llama.cpp GBNF Grammar Writer Prompt for Reliable Invoice JSON

P
promptstudio

Write a GBNF grammar for llama.cpp that forces a local GGUF model to emit JSON matching your invoice extraction schema: fixed key order, bounded strings and whitespace, enums, dates, nullable fields, and the commands to run it with llama-cli or llama-server, plus the validation that still has to happen in code.

How to Use the llama.cpp GBNF Grammar Writer Prompt for Reliable Invoice JSON

Local models are great for document extraction when data cannot leave your machine, but they do not always return clean JSON. A missing quote, an extra sentence before the opening brace, or a field in the wrong format breaks the parser and stalls the pipeline. llama.cpp solves the shape problem with GBNF grammars, which restrict what the model is allowed to generate token by token. The llama.cpp GBNF Grammar Writer: Force a Local Model to Return Valid Invoice Extraction JSON prompt writes the grammar, the run commands, and the validation that still has to happen in your code.

What a grammar does and does not do

A GBNF grammar is a set of rules, starting from root, that describes every allowed output. If the next token would break the rules, llama.cpp does not let the model pick it. That means you always get output your parser can read. What a grammar cannot do is make the values true. The model can still write the wrong total or a date that does not exist. The prompt is explicit about this split and designs both halves: the grammar for structure, and checks in code for meaning.

What the prompt produces

  1. Design notes that explain fixed key order, nullable fields, enums, and why strings, arrays, and whitespace get upper limits.
  2. A complete .gbnf file with rules for the object, line items, strings, numbers, dates, enums, and whitespace.
  3. Run commands for llama-cli with a grammar file or a llama-server request with the grammar field, plus the JSON Schema route as an alternative.
  4. A short extraction prompt that tells the model to use null when a field is missing.
  5. Validation steps for your consumer code, such as parsing, schema checks, and business rules.
  6. A test plan using your own sample documents.

How to fill the inputs

JsonSchema is the exact JSON you want back, either as a schema or a sample. FieldRules adds the details that matter for a grammar: which fields can be null, allowed enum values, date format, string length limits, and the maximum number of array items. LlamaCppSetup says whether you use llama-cli, llama-server, or the OpenAI compatible endpoint, and roughly how recent your build is, since older builds may not support every repetition syntax.

SampleDocs describes two real documents the model will read, ideally one complete and one with a missing field. Consumer names the language and validator that receive the JSON, such as Python with pydantic.

Reading the example output

The example targets invoices with a vendor, an optional invoice number, a date, a currency, a total, and up to 30 line items. Several choices are worth copying:

  • Keys are in a fixed order. Your parser does not care, and the grammar stays small and easy to read.
  • Only invoice_number is nullable, written as an alternative between a string and the literal null.
  • Currency is three literal options, so the model cannot write anything else.
  • Whitespace is bounded. An unbounded whitespace rule can let a model pad output for a long time, so the grammar allows at most one newline and a little indentation.
  • Strings exclude quotes, backslashes, and control characters unless properly escaped.
  • Dates are shape checked in the grammar and calendar checked in code, because a grammar happily accepts February 31.

Validation that still matters

Even with perfect structure, check the meaning. Parse the JSON, validate it against a model with real date types, and compare the sum of quantity times unit price to the total. When something does not add up, route it to a person instead of silently correcting it. Run the utility bill sample without an invoice number and confirm you get null, not an invented value. If you do get a made up number, the grammar is working as designed; the fix belongs in the prompt wording.

Mistakes to avoid

  • Treating grammar output as verified data. It is parseable, not proven.
  • Unbounded rules. Cap strings, arrays, and whitespace.
  • Copying syntax from an old example without checking your build. Confirm flags and script paths in your current checkout.
  • Skipping the missing field test. It shows whether the model guesses.

Who it is for

This prompt fits ML engineers, backend developers, and automation builders running local models for extraction, classification, or form filling. The invoice example adapts easily to receipts, purchase orders, or any document with a stable shape.

Related PromptDig links