Skip to content
StrataHub

Financial Services · February 16, 2026 · 6 min read

Intelligent Document Processing for Loan Origination

Loan files are still built by humans re-keying PDFs. We break down how to get document automation past the demo and into a production origination pipeline — with the eval harness that keeps it honest.

Walk the floor of any mid-size lender and you'll find the same bottleneck: loan files assembled by people reading PDFs and re-keying numbers into an LOS. Bank statements, pay stubs, W-2s, tax returns, insurance declarations, articles of incorporation for commercial deals. A single mortgage file routinely contains 300–500 pages. Processors spend most of their day on stare-and-compare, not on judgment.

Document AI has been "solved" in demos for years. LLM-based extraction genuinely changed the ceiling — modern models read a messy scanned bank statement better than most template-based OCR ever did. But origination is a production problem, not an extraction problem, and that's where most initiatives stall.

Why demos don't survive the mailroom

Three things break between the proof-of-concept and go-live:

Document variety is worse than anyone budgets for. A POC runs on 50 clean samples. Production sees 4,000+ layout variants of bank statements alone, phone photos of pay stubs, 1099s stapled behind W-2s in a single scan, and documents in the wrong slot half the time. Classification and splitting — deciding what each page is — is usually harder than extraction and matters more.

Nobody defined "good enough" per field. A 95% accurate extraction system sounds great until you realize which 5% it gets wrong. Borrower income tolerance is not the same as employer-name tolerance. Without field-level accuracy targets tied to downstream risk, you either trust everything (dangerous) or review everything (pointless).

There's no exception path. The question isn't whether the system will be uncertain; it's what happens when it is. If low-confidence extractions just fail silently or dump raw JSON on a processor, adoption dies in week two.

The architecture that holds up

Our production pattern for origination document processing:

  1. Ingest and split. Every inbound file — portal upload, email, scan — is page-classified and split into logical documents. We hold splitting to a higher eval bar than extraction because a misclassified document poisons everything downstream.
  2. Extract with confidence. LLM extraction against a versioned schema per document type, with calibrated per-field confidence. Calibration matters: raw model logprobs are not trustworthy confidence signals, so we fit them against a labeled holdout so that "90% confident" actually means 90% correct.
  3. Cross-document validation. This is where real value lives. Income on the pay stub versus the VOE versus stated income on the application. Deposits on the bank statement versus claimed assets. Name and address consistency across the file. Discrepancies become structured findings, not silent errors.
  4. Route by confidence. High-confidence, internally consistent fields flow straight through to the LOS. Everything else lands in a review queue that shows the processor the extracted value and the source pixels side by side, one keystroke to confirm or correct.
  5. Learn from corrections. Every human correction is a labeled example. The eval set grows from production reality, and retraining targets the document variants that are actually failing.

Straight-through processing rate is the honest KPI — the percentage of fields (and eventually whole documents) that require zero human touches while meeting accuracy targets. Track it per document type. Vanity accuracy numbers averaged across easy fields tell you nothing.

The eval harness is the product

Before we write a line of pipeline code, we build a golden set: 500–1,500 real documents per major type, labeled at field level, stratified across the ugly variants — skewed scans, phone photos, multi-account statements, handwritten additions. Every model change, prompt change, and schema change runs against it in CI.

Reasonable production targets we've hit and hold in this domain: 97%+ field accuracy on structured documents (W-2s, 1003s), 93–96% on bank statements and pay stubs, with 60–75% of fields flowing straight through at launch and climbing as the correction loop feeds back. Processor time per file typically drops 40–60%, and — the number executives actually care about — cycle time from document receipt to underwriting-ready shrinks by days.

Compliance rides along from day one: full lineage from every extracted field back to the source page and coordinates, immutable processing logs, and PII handling that keeps documents inside your security boundary. If you're using hosted models, that means contractual and architectural controls your CISO signs off on before the pilot, not after.

How we scope it

A StrataHub Pilot here is 4–6 weeks: pick the two or three document types that dominate processor time, build the golden set, stand up the classify-extract-validate pipeline, and run it against live volume in shadow. You get measured field-level accuracy and a straight-through-processing projection grounded in your actual document mix — not a vendor benchmark.

Co-Build (3–6 months) integrates with the LOS, builds the review UI and correction loop, extends to the long tail of document types, and hands your team the eval and retraining machinery. Scale engagements keep the system improving as products, forms, and volumes change.

The lenders getting real returns from document AI didn't buy magic. They built a measured pipeline with honest confidence scores, a humane exception path, and an eval set that grows with production. Everything else is a demo.

Work with us

Shipping something like this?

We co-build production AI systems with enterprise teams — pilots in 4-6 weeks.