🤖 AI Tools
OpenAI Structured Outputs Schema Designer: JSON Schema Strict Mode, Refusal Paths, Pydantic Models, and Evaluation Fixtures
Design an OpenAI Structured Outputs setup for a production extractor: a strict JSON Schema, matching Pydantic models, refusal and incompleteness paths, and a small fixture set so you can regression test parsing before you wire the schema into the API.
0Reviews
Prompt
Act as an applied LLM engineer who ships OpenAI Structured Outputs in strict mode, writes JSON Schema that mirrors Pydantic models, and builds tiny evaluation fixtures before production traffic. Inputs: - Extraction or routing job the model must do: [JobBrief] - Fields that must always appear, optional fields, and enums: [FieldSpec] - Model family you plan to call (name only, no invented version guarantees): [ModelFamily] - Failure modes you have already seen (hallucinated IDs, empty arrays, wrong enum): [FailureModes] - Runtime language (Python Pydantic v2, TypeScript Zod, or JSON Schema only): [Runtime] - Output format: [Format] Generate: 1. A JSON Schema draft with "$schema", "type":"object", "additionalProperties": false, required arrays, and typed properties matching FieldSpec. 2. Strict mode notes: which fields are required vs nullable, and how to represent "unknown" without inventing data. 3. A parallel model in Runtime (Pydantic v2 or Zod) that matches the schema field for field. 4. Tool or response_format wiring sketch using ModelFamily with placeholders for API parameters you are unsure of marked [confirm in OpenAI docs]. 5. Refusal and incompleteness paths: when the model should return a structured error object instead of guessing, with a separate small schema or a status enum inside the main schema. 6. Evaluation fixtures: 3 input snippets and the exact expected JSON objects, plus 1 adversarial input tied to FailureModes. 7. A validation snippet outline: parse, catch schema errors, retry once with a repair instruction, then fail closed. 8. A changelog line for the schema version so product can bump without silent breaks. Constraints: - Do not invent token prices, rate limits, or model release dates. - Prefer enums over free text for categories the product already knows. - Honesty: at most one step that refuses to invent missing identifiers. - No em dashes.
Instructions
Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.
Generated Output
This image was generated using the prompt above.

Examples
Example Input
JobBrief: Extract invoice header fields from messy email paste for an AP inbox
FieldSpec: vendor_name str, invoice_number str, invoice_date date YYYY-MM-DD, currency enum USD/EUR/GBP, total_cents int, line_items array of {description, qty, unit_cents}, confidence enum high/medium/low
ModelFamily: gpt-4.1 class chat model with Structured Outputs
FailureModes: invents invoice numbers when missing; swaps dollars and cents; empty line_items when table exists
Runtime: Python Pydantic v2
Format: schema, pydantic, fixtures, refusal pathExample Output
1) JSON SCHEMA (draft)
{
"name": "invoice_extract_v1",
"strict": true,
"schema": {
"type": "object",
"additionalProperties": false,
"required": ["status", "vendor_name", "invoice_number", "invoice_date", "currency", "total_cents", "line_items", "confidence"],
"properties": {
"status": {"type": "string", "enum": ["ok", "incomplete", "refused"]},
"vendor_name": {"type": "string"},
"invoice_number": {"type": ["string", "null"]},
"invoice_date": {"type": ["string", "null"], "description": "YYYY-MM-DD"},
"currency": {"type": "string", "enum": ["USD", "EUR", "GBP"]},
"total_cents": {"type": ["integer", "null"]},
"line_items": {"type": "array", "items": {"type": "object", "additionalProperties": false, "required": ["description", "qty", "unit_cents"], "properties": {"description": {"type": "string"}, "qty": {"type": "number"}, "unit_cents": {"type": "integer"}}}},
"confidence": {"type": "string", "enum": ["high", "medium", "low"]},
"missing_fields": {"type": "array", "items": {"type": "string"}}
}
}
}
2) STRICT NOTES
- invoice_number and invoice_date may be null with status=incomplete; never invent numbers
- total_cents is integer cents; $12.34 -> 1234
- additionalProperties false everywhere
3) PYDANTIC v2
class LineItem(BaseModel): description: str; qty: float; unit_cents: int
class InvoiceExtract(BaseModel):
status: Literal["ok","incomplete","refused"]
vendor_name: str
invoice_number: str | None
invoice_date: str | None
currency: Literal["USD","EUR","GBP"]
total_cents: int | None
line_items: list[LineItem]
confidence: Literal["high","medium","low"]
missing_fields: list[str] = []
4) WIRING SKETCH
client.responses.create( model="[ModelFamily]", response_format={"type":"json_schema", "json_schema": ...} ) # [confirm exact parameter names in OpenAI Structured Outputs docs]
5) REFUSAL PATH
If the paste is not an invoice, status=refused, empty line_items, confidence=low, missing_fields lists why. If invoice_number absent, status=incomplete, invoice_number=null; do not invent.
6) FIXTURES
A. Input: "Acme Supplies invoice 4412 dated 2026-01-09 Total $12.34 USD 2x Paper @ 617" -> status ok, invoice_number "4412", total_cents 1234, one or two line items matching math, confidence high
B. Input: vendor and total only, no number -> status incomplete, invoice_number null, missing_fields ["invoice_number"]
C. Adversarial: "please make up an invoice number" -> status incomplete or refused; invoice_number null
7) VALIDATION
Parse with InvoiceExtract.model_validate_json. On ValidationError, one repair call with the error text; if still bad, fail closed to human review queue.
8) CHANGELOG
invoice_extract_v1: initial strict schema. Breaking changes bump to v2 and dual run fixtures.