vM.

How to Validate Structured Output from an LLM in a Production Application

Author
Vishal Maurya
Published on
Reading time
4 min read

Overview

Many AI features need more than a paragraph of generated text. A document-processing system may need a vendor name, invoice date, total amount, and line items. A support workflow may need a category and priority. A downstream API expects fields with specific types and valid values.

Asking a model to return JSON does not guarantee that the result is valid or factually correct. Production systems need validation at more than one level.

1. Define a Schema for the Expected Shape

Pydantic can validate the types and basic constraints of structured data in Python.

from decimal import Decimal
from pydantic import BaseModel, Field


class InvoiceExtraction(BaseModel):
    invoice_number: str = Field(min_length=1, max_length=80)
    currency: str = Field(pattern=r"^[A-Z]{3}$")
    total: Decimal = Field(ge=0)

This schema checks that the output has the expected fields and acceptable basic values. It does not prove that the invoice number or total was read correctly from the source document.

If your model provider supports structured outputs or schema-constrained generation, use the provider's supported mechanism to reduce formatting failures. Still validate the result in your own application.

2. Separate Parsing Errors from Business Errors

There are several distinct failure types:

  • The response is not valid JSON.
  • The JSON does not match the expected schema.
  • A field has the right type but an impossible value.
  • The output is structurally valid but contradicts the source document.
  • The model returns a plausible answer when the evidence is missing.

Handle these cases differently. A schema mismatch might justify one bounded repair attempt. A contradiction with the source may require re-extraction or human review. A missing value should not automatically be replaced with a guess.

3. Validate Relationships Between Fields

Business rules often involve more than one field. For an invoice, the sum of line items plus tax and adjustments may need to reconcile with the total. A date may need to fall within an allowed range. A currency may need to match the vendor or account configuration.

Implement these rules in application code where they can be tested deterministically. If the source document is authoritative, preserve the extracted evidence and flag disagreements instead of silently modifying values to make the arithmetic pass.

4. Preserve Evidence and Provenance

For document extraction, store the source document ID and, where possible, page number or text span supporting each important field. This lets a reviewer compare the model's value with the source.

Do not treat a model-generated confidence score as a calibrated probability unless you have evaluated and calibrated it for your data. Use explicit signals such as schema validity, cross-field checks, extraction quality, and reviewed evaluation results to route uncertain cases.

5. Use Bounded Retries and Safe Failure States

Repeatedly asking the model to fix invalid output can increase cost without improving correctness. Set a small retry budget and record why each attempt failed. After the limit, return a clear failure state or send the item to review.

Do not allow an invalid result to flow into payments, customer records, or another system just because a retry budget was exhausted. The downstream action should require validated data.

6. Test with a Representative Evaluation Set

Build a test set that includes clean examples, unusual layouts, missing fields, conflicting values, poor OCR, and cases where the answer is not present. Measure schema pass rate separately from field accuracy and downstream business correctness.

Review errors by field and document type. If totals fail more often than invoice numbers, the solution may involve better table extraction or numeric normalization rather than a different prompt alone.

When changing the model, prompt, extraction pipeline, or schema, rerun the evaluation set. A prompt that improves one category can regress another.

7. Treat Model Output as Untrusted Data

Never execute generated code or SQL simply because it passed a JSON schema. Validate identifiers, authorize requested operations, parameterize database queries, and constrain tools to the user's permitted actions. A schema validates shape, not authority.

Conclusion

Reliable LLM integration requires schema validation, business rules, source evidence, bounded retries, and an explicit path for uncertain results. Structured output is a useful interface, not a guarantee of truth.

If your product needs AI extraction from PDFs, structured responses from an LLM, or a safer bridge between AI output and business APIs, I can help build the validation and review layer. Contact me.

Additional Resources

  • Pydantic documentation
  • OpenAI API documentation