Decision Matrix — Document AI Model Selection
Four invoice-extraction approaches benchmarked on 2,000 held-out invoices, with accuracy, latency and cost per 1,000 pages measured and a 95% field-accuracy must-have that rules out the current OCR-and-rules pipeline. The fine-tuned open model ranks first at 4.11 of 5, just 0.09 ahead of the small hosted model — a fragile lead that flips if ops effort's weight rises from 10% to 12.5%. Illustrative values.
Make it your own.
title "Invoice extraction pipeline — model selection"
subtitle "Benchmarked on 2,000 held-out invoices from 340 suppliers"
baseline "OCR + rules (current)"
note "Illustrative benchmark results for a fictional accounts-payable team."
note "Accuracy is field-level exact match on 14 fields; latency is p95 per page."
option "Large hosted model"
option "Small hosted model"
option "Fine-tuned open model" "self-hosted, 8B parameters"
option "OCR + rules (current)"
"Field-level accuracy" weight 35 unit "%" must >= 95 : 98.6 96.1 97.8 91.4
"Latency (p95)" weight 10 unit "s" lower : 6.8 2.1 1.4 0.6
"Cost per 1,000 pages" weight 20 unit "$" lower : 41 9 6 2
"Robustness to new layouts" weight 15 : 5 4 3 1
"EU data residency" weight 10 must >= 4 : 4 4 5 5
"Ops effort (1 = minimal)" weight 10 lower : 1 1 4 3