Thinklytics

AI Automation · 9 min read · September 2026

What AI document processing actually costs in 2026, and the accuracy gap that decides the business case

By Sean Majidi, Founder, Thinklytics

Vendors quote field-level accuracy in the mid 90s. Whole-document accuracy in a 2026 benchmark of 10,000 filings ran 63% to 76%. That gap is your exception queue, and it is what decides whether the business case works.

Budget between $0.15 and $0.45 per document for the AI processing itself, and expect the software to be the small line. The cost that decides whether the project pays back is the exception queue, and its size is set by a number most vendors do not quote.

The number vendors quote, and the number that decides your case

Extraction accuracy is normally reported per field. A March 2026 benchmark over 10,000 SEC filings measured field-level F1 between 0.903 and 0.943 depending on the architecture. That is the figure that reaches a pitch deck, and it is not wrong.

Document-level strict accuracy for the same systems, meaning every field on the page correct at once, ran between 63.1% and 75.8%.

The two accuracy numbers, and which one you pay for

  • Field-level F1. 0.903 to 0.943. Per-field accuracy across four architectures. The number that reaches a pitch deck.
  • Document-level strict accuracy. 63.1% to 75.8%. Every field on the document correct at once. The number straight-through processing depends on.

Benchmark of 10,000 SEC filings (4,000 10-K, 4,000 10-Q, 2,000 8-K) across 11 GICS sectors, 25 extracted field types. The gap between the two columns is your exception queue.

Source: Kulkarni and Kulkarni, Benchmarking Multi-Agent LLM Architectures for Financial Document Processing, arXiv:2603.22651, March 2026.

The gap between those two numbers is your exception queue. If a document carries 25 extracted fields and each is right 94% of the time, the chance that all 25 are right together is far below 94%. Straight-through processing depends on the whole document being correct, not on the average field, so the business case has to be built on the lower number.

This is the single most common error we see in a document automation business case. A team sizes the saving against 94% and staffs the review queue for 6% of volume, then discovers a third of documents need a human look.

What the documents cost you today

Accounts payable is the best-benchmarked case, so it is the honest place to start even if your documents are contracts or claims.

What an invoice costs to process today

Average performer against best-in-class. The exception rate is the work automation has to survive, not avoid.

MeasureAverageBest-in-class
Cost per invoice$9.8479% lower
Cycle time8.2 days79% faster
Exception rate18.4%47% lower
Touchless processingBaselineMore than 1.8x as many
Staff time on supplier inquiries21.9%About half
Suppliers able to invoice electronically57%1.4x more

Source: Ardent Partners, State of ePayables 2025, Part Nine: AP Benchmarks and Best-in-Class Performance, published January 2026.

Two things in that table matter more than the headline cost. The exception rate of 18.4% is the work that automation has to survive, not avoid. And the best-in-class gap is a multiple, not a rounding: those organisations process more than 1.8 times as many invoices touchlessly and spend roughly half the time on supplier inquiries.

Outside AP the anchors are coarser but still real. Mortgage Bankers Association data for Q1 2026 puts total loan production expense at $11,898 per loan across 324 companies, which is the denominator any lending document project is measured against. In healthcare administration, the 2025 CAQH Index reports $258 billion in administrative cost avoided through automated transactions, with $21 billion of annual opportunity still on the table.

What the processing itself costs

What the processing itself costs

Published benchmark costs and cloud list prices. The software is rarely the expensive part.

ApproachUnit costAccuracy or note
Sequential architecture$0.187 per document0.903 field F1
Hierarchical, cost-optimised$0.148 per document0.924 field F1
Reflexive architecture$0.430 per document0.943 field F1, degrades at volume
Amazon Textract, expense analysis$0.01 per page$0.008 above the first tier
Amazon Textract, text detection$0.0015 per page$0.0006 above 1M pages
Google Enterprise Document OCR$1.50 per 1,000 pages$0.60 above 5M pages a month
Google Custom Extractor$30 per 1,000 pages$20 above 1M pages a month

Source: arXiv:2603.22651, March 2026; AWS Textract and Google Document AI published pricing, retrieved September 2026.

Note what happens at volume. The most accurate architecture in that benchmark degraded from 0.943 to 0.871 F1 as throughput rose from 1,000 to 100,000 documents a day, because its self-correction loops began timing out. A cheaper architecture overtook it above roughly 50,000 documents a day. Accuracy measured at pilot volume is not accuracy at production volume, and the crossover point is worth finding before you commit.

Why input quality moves the number more than model choice

A July 2026 benchmark ran the same extraction models against clean reference text and against the output of a production OCR engine.

What OCR noise costs you, before the model sees anything

Identical extraction models run against clean reference text and against production OCR output.

  • Forms, clean text
  • Forms, OCR input
  • Receipts (SROIE), clean text
  • Receipts (SROIE), OCR input
  • Receipts (CORD), clean text
  • Receipts (CORD), OCR input

Source: Anvari and Athitsos, From Pixels to Pairs, arXiv:2609.17538, July 2026. Values are value-level F1, shown as percentages.

Receipts with dense layouts lost nearly 15 points of F1 purely from OCR noise. The failure mode was not the model failing to find the field. It was the value being corrupted, or the key and value being misaligned, before the model ever saw them.

The practical consequence: if your documents arrive as phone photographs, faxes, or scans of scans, the highest-return work is at the capture step, not the model step. Changing model providers will not recover 15 points. Getting suppliers to send a digital original might.

How to size the case without guessing

Four numbers, in this order.

  • Your current fully loaded cost per document. Not the software cost, the people cost. Ardent's $9.84 is the AP benchmark; your own figure is what matters.
  • Your document-level accuracy on a real sample, not the vendor's field-level accuracy on theirs. Run 200 of your own documents, count how many came out completely correct, and use that number.
  • The cost of a human touch, which is the loaded hourly rate divided by documents reviewed per hour. This is what the exception queue actually costs.
  • The volume at which you will run, because the architecture that wins at 1,000 a day may not be the one that wins at 50,000.

The saving is then arithmetic: volume times current cost, minus volume times processing cost, minus exceptions times cost of a human touch. If that arithmetic only works at the vendor's accuracy figure and fails at your own, the answer is not a different vendor.

What the exception queue needs before you buy anything

The queue is the part that decides the economics, and it is usually designed last.

  • A measured document-level accuracy on your own documents. At least 200, including the difficult ones. Count documents that came out completely correct, not fields.
  • A named owner for every exception type. An unrouted queue becomes a shared inbox, and a shared inbox becomes nobody's job.
  • Confidence thresholds set per field, not per document. A wrong supplier name and a wrong line-item description do not carry the same cost.
  • A cost per human touch. Loaded hourly rate divided by documents reviewed per hour. Without it you cannot size the saving.
  • Volume testing at production scale. One benchmarked architecture lost 7 points of F1 between 1,000 and 100,000 documents a day.
  • A feedback path from correction back into the system. If reviewer corrections do not change future behaviour, the queue never shrinks.

Nothing here is about model selection. Every item is about what happens to the documents the model could not finish.

Source: Thinklytics AI automation practice, 2026.

The clock that is actually running

Electronic invoicing mandates are doing more to force this work than any AI business case. France required all VAT-registered companies established there to be able to receive electronic invoices from 1 September 2026, with large companies and mid-caps also required to issue from that date and smaller companies following in September 2027. Poland's KSeF regime began with the largest taxpayers in February 2026 and extended to other VAT-registered businesses in April 2026.

If you trade in those markets the question is no longer whether to handle structured invoice data. It is whether you also get the extraction benefit while you are rebuilding the pipe anyway. That is a much easier case to fund than an AI project standing on its own.

What we would do first

Before any platform conversation, take 200 documents that represent your real mix, including the ugly ones, and measure three things: how many come out completely correct, where the failures cluster, and what a human touch costs you today. That is a week of work and it turns the business case from a vendor's number into yours.

Our AI automation practice runs that measurement as the first step of any document engagement, and the companion piece on designing the escalation path covers what to do with the documents that will always need a person. If you want a view across functions before narrowing to documents, the AI Opportunity Finder ranks use cases by impact against effort in about three minutes.

Frequently asked questions

How much does AI document processing cost per document?

The processing itself runs roughly $0.15 to $0.45 per document. A March 2026 benchmark over 10,000 SEC filings measured $0.187 to $0.430 per document across four architectures, with a cost-optimised configuration reaching 0.924 field-level F1 at $0.148. Cloud OCR and extraction services price separately and lower: Amazon Textract lists expense analysis at $0.01 per page and Google's Enterprise Document OCR at $1.50 per 1,000 pages. None of those are the real cost, which is the human review of documents the system could not complete.

What accuracy should we expect from document extraction?

Expect field-level accuracy in the low-to-mid 90s and whole-document accuracy far lower. The same 2026 benchmark that measured field-level F1 between 0.903 and 0.943 measured document-level strict accuracy, meaning every field correct at once, between 63.1% and 75.8%. Straight-through processing depends on the document-level figure, so that is the one to build the business case on.

Why is our accuracy worse than the vendor demo?

Usually because of input quality rather than the model. A July 2026 benchmark ran identical models against clean reference text and against production OCR output. Receipt extraction fell from 0.9700 to 0.8267 F1 purely from OCR noise, and form extraction from 0.6371 to 0.5781. The errors came from corrupted values and misaligned key-value pairs, not from the model failing to locate fields. If documents arrive as photographs or scans of scans, the capture step is where the accuracy is lost.

What does a manual invoice cost to process?

Ardent Partners' State of ePayables research, published January 2026, puts the average at $9.84 per invoice with an 8.2 day cycle time and an 18.4% exception rate. Best-in-class performers run 79% lower cost and 79% faster cycle times, process more than 1.8 times as many invoices touchlessly, and spend about half as much staff time on supplier inquiries. Widely repeated figures of $12 to $30 per invoice do not trace to a primary source.

Does document processing accuracy hold up at volume?

Not automatically. In the 2026 benchmark, the most accurate architecture degraded from 0.943 to 0.871 F1 as throughput rose from 1,000 to 100,000 documents a day, because its self-correction loops timed out under load. A simpler and cheaper architecture overtook it above roughly 50,000 documents a day. Accuracy measured during a pilot is not accuracy in production, so test at the volume you will actually run.

Are e-invoicing mandates forcing this work?

In several markets, yes. France required all VAT-registered companies established there to be able to receive electronic invoices from 1 September 2026, with large companies and mid-caps also required to issue from that date and smaller companies following in September 2027. Poland's KSeF regime started with the largest taxpayers in February 2026 and extended to other VAT-registered businesses in April 2026. Where a mandate applies, the pipeline work is happening regardless, which makes the extraction benefit much easier to fund.

How do we size the business case without guessing?

Measure four things: your current fully loaded cost per document, your document-level accuracy on a sample of at least 200 of your own documents including the difficult ones, the cost of a single human touch, and the volume you will run at. The saving is volume times current cost, minus volume times processing cost, minus exceptions times the cost of a human touch. If the case only works using the vendor's accuracy figure rather than your own, a different vendor will not fix it.

Topics covered

  • ai document processing cost
  • intelligent document processing
  • invoice processing cost
  • document extraction accuracy
  • straight-through processing
  • idp business case
  • e-invoicing mandate

Frequently asked questions

How much does AI document processing cost per document?

The processing itself runs roughly $0.15 to $0.45 per document. A March 2026 benchmark over 10,000 SEC filings measured $0.187 to $0.430 per document across four architectures, with a cost-optimised configuration reaching 0.924 field-level F1 at $0.148. Cloud OCR and extraction services price separately and lower: Amazon Textract lists expense analysis at $0.01 per page and Google's Enterprise Document OCR at $1.50 per 1,000 pages. None of those are the real cost, which is the human review of documents the system could not complete.

What accuracy should we expect from document extraction?

Expect field-level accuracy in the low-to-mid 90s and whole-document accuracy far lower. The same 2026 benchmark that measured field-level F1 between 0.903 and 0.943 measured document-level strict accuracy, meaning every field correct at once, between 63.1% and 75.8%. Straight-through processing depends on the document-level figure, so that is the one to build the business case on.

Why is our accuracy worse than the vendor demo?

Usually because of input quality rather than the model. A July 2026 benchmark ran identical models against clean reference text and against production OCR output. Receipt extraction fell from 0.9700 to 0.8267 F1 purely from OCR noise, and form extraction from 0.6371 to 0.5781. The errors came from corrupted values and misaligned key-value pairs, not from the model failing to locate fields. If documents arrive as photographs or scans of scans, the capture step is where the accuracy is lost.

What does a manual invoice cost to process?

Ardent Partners' State of ePayables research, published January 2026, puts the average at $9.84 per invoice with an 8.2 day cycle time and an 18.4% exception rate. Best-in-class performers run 79% lower cost and 79% faster cycle times, process more than 1.8 times as many invoices touchlessly, and spend about half as much staff time on supplier inquiries. Widely repeated figures of $12 to $30 per invoice do not trace to a primary source.

Does document processing accuracy hold up at volume?

Not automatically. In the 2026 benchmark, the most accurate architecture degraded from 0.943 to 0.871 F1 as throughput rose from 1,000 to 100,000 documents a day, because its self-correction loops timed out under load. A simpler and cheaper architecture overtook it above roughly 50,000 documents a day. Accuracy measured during a pilot is not accuracy in production, so test at the volume you will actually run.

Are e-invoicing mandates forcing this work?

In several markets, yes. France required all VAT-registered companies established there to be able to receive electronic invoices from 1 September 2026, with large companies and mid-caps also required to issue from that date and smaller companies following in September 2027. Poland's KSeF regime started with the largest taxpayers in February 2026 and extended to other VAT-registered businesses in April 2026. Where a mandate applies, the pipeline work is happening regardless, which makes the extraction benefit much easier to fund.

How do we size the business case without guessing?

Measure four things: your current fully loaded cost per document, your document-level accuracy on a sample of at least 200 of your own documents including the difficult ones, the cost of a single human touch, and the volume you will run at. The saving is volume times current cost, minus volume times processing cost, minus exceptions times the cost of a human touch. If the case only works using the vendor's accuracy figure rather than your own, a different vendor will not fix it.

Related reading