Skip to content
Decision matrix (weighted / Pugh) templates

Decision Matrix — Document AI Model Selection

Four invoice-extraction approaches benchmarked on 2,000 held-out invoices, with accuracy, latency and cost per 1,000 pages measured and a 95% field-accuracy must-have that rules out the current OCR-and-rules pipeline. The fine-tuned open model ranks first at 4.11 of 5, just 0.09 ahead of the small hosted model — a fragile lead that flips if ops effort's weight rises from 10% to 12.5%. Illustrative values.

Template previewDecision matrix (weighted / Pugh)
Invoice extraction pipeline — model selectionInvoice extraction pipeline — model selectionBenchmarked on 2,000 held-out invoices from 340 suppliersWeighted scoringScale 1–56 criteria4 optionsBaseline: OCR + rules (current)CRITERIONWEIGHTLarge hostedmodelSmall hostedmodelFine-tuned openmodelself-hosted, 8BparametersRECOMMENDEDOCR + rules(current)EXCLUDEDField-level accuracymeasured, % · must ≥ 95%35%98.6%score 5+96.1%score 3.6+97.8%score 4.6+91.4%score 1Latency (p95)measured, s · lower is better10%6.8 sscore 1−2.1 sscore 4−1.4 sscore 4.5−0.6 sscore 5Cost per 1,000 pagesmeasured, $ · lower is better20%$41score 1−$9score 4.3−$6score 4.6−$2score 5Robustness to new layouts15%5+4+3+1EU data residencymust ≥ 410%4−4−55Ops effort (1 = minimal)lower is better10%1+1+4−3Weighted scoreΣ weight × score, of 53.7074% of max4.0280.5% of max4.1182.2% of max2.8056% of maxvs baselinechange in weighted score+0.903 better · 3 worse+1.223 better · 3 worse+1.312 better · 3 worsebaselineRankhighest score first3rd2nd1stexcludedWeighted totalsEach bar is split into its criteria's weighted contributions (colours match the dots in the table); ranked best first.012345Fine-tuned open model4.1182.2%Small hosted model4.0280.5%Large hosted model3.7074%OCR + rules (current)2.8056%excluded — fails a must-havebaselineFINDINGS & SENSITIVITYFRAGILEFine-tuned open model ranks first at 4.11 of 5 (82.2%), 0.09 ahead of Small hosted model.Sensitivity (fragile): Small hosted model draws level with Fine-tuned open model if the weight on “Ops effort (1 = minimal)” rises from10% to 12.5% (+2.5 points), the other weights rescaled in proportion — the smallest single-weight change that changes the winner.Small hosted model would draw level by scoring 4.58 instead of 4 on “Robustness to new layouts” — the smallest single-score changethat would.Against the baseline “OCR + rules (current)”: Fine-tuned open model adds 1.31 points, better on 2 criteria and worse on 3 criteria.OCR + rules (current) is excluded — it fails a must-have (“Field-level accuracy” is 91.4%, must be ≥ 95%), so its total of 2.80 is notranked.Illustrative benchmark results for a fictional accounts-payable team.Accuracy is field-level exact match on 14 fields; latency is p95 per page.

Make it your own.

title "Invoice extraction pipeline — model selection"
subtitle "Benchmarked on 2,000 held-out invoices from 340 suppliers"
baseline "OCR + rules (current)"
note "Illustrative benchmark results for a fictional accounts-payable team."
note "Accuracy is field-level exact match on 14 fields; latency is p95 per page."

option "Large hosted model"
option "Small hosted model"
option "Fine-tuned open model" "self-hosted, 8B parameters"
option "OCR + rules (current)"

"Field-level accuracy"       weight 35 unit "%" must >= 95 : 98.6 96.1 97.8 91.4
"Latency (p95)"              weight 10 unit "s" lower      : 6.8 2.1 1.4 0.6
"Cost per 1,000 pages"       weight 20 unit "$" lower      : 41 9 6 2
"Robustness to new layouts"  weight 15                     : 5 4 3 1
"EU data residency"          weight 10 must >= 4           : 4 4 5 5
"Ops effort (1 = minimal)"   weight 10 lower               : 1 1 4 3