vM.

How to Extract and Validate Data from PDF Documents Using Python

Author
Vishal Maurya
Published on
Reading time
5 min read

Overview

PDF automation often fails for a simple reason: a PDF is a page format, not a guarantee that its contents are stored as structured data. One file may contain selectable text, another may be a scan, and a third may combine paragraphs, tables, and images. Choosing an extraction library before checking the files can produce incomplete data without an obvious error.

A reliable pipeline identifies the document type, extracts content using the appropriate method, validates the fields that matter, and preserves enough source information for a reviewer to investigate mistakes.

1. Inspect the PDFs before selecting a tool

Start with representative files from each source system. Check whether text can be selected, whether tables have consistent layouts, whether pages are rotated, and whether documents mix scanned and digital pages.

  • Native text: extract text directly; OCR may be unnecessary.
  • Scanned pages: render the page and use OCR.
  • Tables: use table-aware extraction and inspect row and column alignment.
  • Mixed documents: choose the extraction method per page when necessary.

Do not assume an empty text result means the page is blank. It may be a scan or a page whose layout the parser could not interpret.

2. Extract text page by page

Install pdfplumber with pip install pdfplumber.

from pathlib import Path
import pdfplumber


def extract_pages(path: Path) -> list[dict[str, object]]:
    result = []
    with pdfplumber.open(path) as pdf:
        for page_number, page in enumerate(pdf.pages, start=1):
            result.append({
                "page_number": page_number,
                "text": page.extract_text() or "",
            })
    return result


for page in extract_pages(Path("statement.pdf")):
    print(f"--- page {page['page_number']} ---")
    print(page["text"])

Keeping page boundaries helps with auditability and later citations. It also makes it easier to compare a suspicious field with the source page.

3. Extract tables as a separate step

from pathlib import Path
import pdfplumber


def extract_tables(path: Path) -> list[dict[str, object]]:
    result = []
    with pdfplumber.open(path) as pdf:
        for page_number, page in enumerate(pdf.pages, start=1):
            for table_number, rows in enumerate(page.extract_tables(), start=1):
                result.append({
                    "page_number": page_number,
                    "table_number": table_number,
                    "rows": rows,
                })
    return result

This is a starting point, not a universal solution. Table detection depends on lines, spacing, alignment, and how the PDF was generated. Compare extracted rows with the rendered page, especially when values determine money, eligibility, or compliance outcomes. Preserve headers, units, footnotes, and effective dates with each row.

4. Use OCR when the page is an image

A typical OCR path is: render a page at a suitable resolution, correct rotation or skew if needed, run an OCR engine, and preserve page numbers and text coordinates. Tesseract is one local option. A managed document-processing service may be a better fit when forms, tables, or high document volumes are central to the use case.

The right choice depends on document quality, languages, privacy requirements, throughput, and review costs. OCR output is a transcription candidate, not ground truth: characters such as 0 and O, decimal points, and dates can be misread.

5. Validate extracted values

Parsing answers “what text did we find?” Validation asks “does this value make sense for this business process?” For example, a financial amount should be numeric and finite, while an invoice identifier may need to follow a defined format.

import re
from decimal import Decimal, InvalidOperation


def parse_amount(raw: str) -> Decimal:
    normalized = raw.replace(",", "").strip()
    try:
        amount = Decimal(normalized)
    except InvalidOperation as exc:
        raise ValueError("Amount is not a valid decimal") from exc

    if not amount.is_finite() or amount < 0:
        raise ValueError("Amount must be finite and non-negative")
    return amount


def validate_invoice_number(raw: str) -> str:
    value = raw.strip()
    if not re.fullmatch(r"[A-Za-z0-9/-]{3,40}", value):
        raise ValueError("Unexpected invoice number format")
    return value

These checks validate shape and range, not whether the OCR read the source correctly. Add business rules such as checking subtotal plus tax against the total, matching the vendor to a known record, detecting duplicate invoice numbers, and checking date ranges. Use Decimal rather than binary floating-point arithmetic for financial amounts.

6. Route uncertain results to review

Do not force every document into an automatic success state. Track which fields were extracted, which validation rules passed, and which source page supports each value. Send missing or conflicting fields to a review queue, and let a reviewer correct the result against the original PDF.

A useful record might contain the document ID, page number, field name, extracted value, validation status, and processing version. Avoid storing more sensitive document content than the workflow needs.

7. Separate pipeline stages

A maintainable workflow separates upload, classification, extraction, normalization, validation, and persistence. Each stage should have a status and an error category so one failed document can be retried without reprocessing the entire batch. For larger workloads, run processing asynchronously and store generated outputs in durable object storage rather than relying on a container's temporary filesystem.

When to add AI

Rules work well when layouts and field formats are predictable. An AI model can help interpret variable labels or semi-structured content, but its output should be schema-validated and checked against business rules. For regulated or financial workflows, preserve source references and require review when evidence is incomplete or inconsistent.

Conclusion

Reliable PDF extraction is a pipeline problem, not just a library choice. Start by measuring performance and error types on a representative set of files. Then select the simplest extraction method that meets the accuracy and operational requirements.

If your team processes invoices, loan documents, statements, contracts, or other PDFs at scale, I can help implement extraction, OCR, table handling, validation, and integration with your existing API or database.

Contact me with the types of PDFs you process and the fields you need to extract.

Additional Resources

  • pdfplumber
  • PyMuPDF documentation
  • Tesseract OCR documentation