AI Automation · 8 min read · September 2026
Computer vision inspection in 2026: why the threshold decision matters more than the model
By Sean Majidi, Founder, Thinklytics
A field trial measured 43 to 59% less herbicide used with targeted spray. At the wrong sensitivity setting, the target weed population rose 280% a year. Same equipment. The threshold decision, not the model, is what determines the outcome.
Any computer vision accuracy claim without a false-call rate attached is not a measurement. A system that passes every part is 99.9% accurate on a line where one part in a thousand is defective, and it catches nothing.
That is the first thing to fix in how these projects get evaluated.
What the research shows, and what deployment shows
Published accuracy, and what sits behind it
- Median classification accuracy. Above 95%. Across 196 open-access papers on deep learning for automated visual inspection. Median F1 around 0.85 to 0.90.
- Median dataset size. 1,952 images. Range 45 to over 100,000, mean 14,614. The mean is dragged by a few large datasets. 97% of studies used supervised learning.
The same survey identified a gap of roughly three years between academic breakthrough and industrial deployment.
Source: Hutten et al., Deep Learning for Automated Visual Inspection in Manufacturing and Maintenance, Applied System Innovation 7(1):11, January 2024.
A 2024 systematic survey of 196 open-access papers found median classification accuracy above 95% in most reported cases and median F1 around 0.85 to 0.90. Those are real results. The same survey also found a lag of roughly three years between academic breakthrough and industrial deployment, and that 97% of the studies relied on supervised learning despite limited labelled data being the normal condition in industry.
One detail in that survey is more useful than the accuracy figures. Dataset sizes ranged from 45 images to over 100,000, with a mean of 14,614 and a median of 1,952. The mean is dragged by a handful of very large datasets. Most published work is done on small ones, which means the binding constraint in practice is labelling quality rather than volume.
The failure mode, measured
The same equipment, two sensitivity settings
- Highest sensitivity. 43 to 59% less herbicide. Weed control generally comparable to standard broadcast application, across a three-year field trial.
- Low sensitivity. 280% annual increase. Palmer amaranth populations rose year on year. The opposite of the intended outcome.
The threshold decision, not the model, determined whether the deployment created value or destroyed it.
Source: Arkansas Agricultural Experiment Station, University of Arkansas, three-year soybean field trial from 2022, published March 2025.
The best documented example of what goes wrong is agricultural rather than industrial, and it is unusually honest because it comes from a university experiment station rather than a vendor.
A three-year field trial of targeted spray technology measured 43 to 59% reductions in herbicide use compared with broadcast application. At the highest sensitivity setting, weed control was generally comparable to standard practice. At low sensitivity settings, populations of one target weed rose 280% per year.
The same equipment, tuned differently, produced either a substantial saving or a worse outcome than doing nothing. That is the precision and recall trade-off made concrete, and it is the decision that actually determines whether a vision deployment works.
What a credible result looks like
GE Aerospace reported that AI-assisted borescope inspection raised detection rates by around 34% for high-pressure compressor defects, with recall up 33.6% and precision up 13.5%.
Precision rising alongside recall is the part that makes it credible. Most systems buy recall by accepting more false calls, which moves the cost from missed defects to wasted operator time rather than removing it. A result where both improve means the system is really discriminating better, not just flagging more.
What to specify before anyone quotes
What to specify before a vendor quotes
None of these is about the model.
- The false-call rate your line can absorb per shift. An accuracy figure without this attached is not a measurement.
- Recall on the specific defects that matter. Not overall accuracy. A system that passes everything is 99.9% accurate where one part in a thousand is defective.
- 200 borderline acceptable parts, not just 200 defects. The borderline set is the one nobody collects and the one that decides the threshold.
- Who adjudicates when the system is unsure. The escalation path, before the deployment rather than after.
- How the threshold gets changed, and by whom. This is the control that determines the outcome. It needs an owner and a change record.
GE Aerospace reported recall up 33.6% and precision up 13.5% together on borescope inspection. Both rising is the hard combination and the reason that result is credible.
Source: GE Aerospace and Waygate Technologies, October 2024; Thinklytics AI automation practice.
Notice that none of these are about the model. A vendor who cannot discuss false-call rate and threshold policy in the first meeting is selling a demo.
Adoption is not the same as benefit
Deloitte's survey of 600 manufacturing executives found 80% planning to invest 20% or more of improvement budgets in smart manufacturing and 22% planning to use physical AI within two years, up from 9% today. The intent is real.
A clinical counterpoint is worth holding alongside it. A survey of academic US neuroradiology departments found 81% using AI tools, and the department leaders reported minimal impact on workload so far, with tool performance remaining inconsistent. The sample was 16 departments, so the percentage is indicative rather than definitive, but the qualitative finding is the one that travels: deployment and benefit are separate milestones, and the gap between them is measured in integration and threshold work rather than in model quality.
What we would do first
Before any vendor conversation, assemble 200 images of defects you actually care about and 200 of acceptable parts that look borderline. The second set is the one that matters and the one nobody collects. Then state the false-call rate your line can absorb per shift. Those two artefacts turn a vendor demo into a test.
Our AI readiness work covers whether the data and labelling to support this exist before anyone commits capital, and AI automation covers the integration into the line and the escalation path for the parts the system cannot judge. For the same trade-off in document form, see what AI document processing costs.
Frequently asked questions
How accurate is AI visual inspection?
In published research, very. A 2024 systematic survey of 196 open-access papers on deep learning for automated visual inspection found median classification accuracy above 95% in most reported cases, with median F1 around 0.85 to 0.90. The gap worth knowing is between that and deployment: the same survey identified roughly a three-year lag between academic breakthrough and industrial use, and 97% of the studies relied on supervised learning despite limited labelled data in real settings.
Why do vendor accuracy claims of 99% mean nothing?
Because an accuracy figure without a false-call rate and a defect distribution is not a measurement. If one part in a thousand is defective, a system that passes everything is 99.9% accurate and catches nothing. The numbers that matter are recall on the defects you care about and the false-call rate, because false calls are what destroy line throughput and operator trust.
What does a real deployment achieve?
Meaningful but specific gains. GE Aerospace reported that AI-assisted borescope inspection increased detection rates by around 34% for high-pressure compressor defects, with recall up 33.6% and precision up 13.5%. Note that precision rose alongside recall, which is the hard combination and the reason the result is credible.
What is the biggest risk in a computer vision deployment?
Getting the sensitivity threshold wrong, and the consequences can be worse than not deploying. A three-year university field trial of targeted spray technology measured 43 to 59% reductions in herbicide use against broadcast application. At the highest sensitivity, weed control was generally comparable to standard practice. At low sensitivity settings, populations of one target weed rose 280% per year. The same system, tuned differently, produced the opposite of the intended outcome.
How much training data do we need?
Less than expected, and the published distribution is revealing. Across 196 studies, dataset sizes ranged from 45 images to over 100,000, with a mean of 14,614 but a median of only 1,952. The mean is pulled by a few very large datasets. Most real work is done on small ones, which is why the practical constraint is usually labelling quality rather than volume.
Where is computer vision actually being adopted?
Manufacturing leads and the intent is strong. Deloitte's survey of 600 manufacturing executives found 80% planning to invest 20% or more of improvement budgets in smart manufacturing, and 22% planning to use physical AI within two years against 9% currently. The machine vision market itself is forecast at $15.83 billion in 2025 growing to $23.63 billion by 2030, a more modest 8.3% compound rate than the headline AI numbers.
Does adoption mean value?
Not automatically, and one clinical study makes the point sharply. A survey of academic US neuroradiology departments found 81% using AI tools, but the leaders reported minimal impact on workload so far and that tool performance remains inconsistent. The sample was only 16 departments, so treat the percentage as indicative, but the qualitative finding matches what we see elsewhere: deployment and benefit are different milestones.
Topics covered
- computer vision inspection
- automated visual inspection
- defect detection
- false call rate
- machine vision manufacturing
- quality inspection ai
- precision recall tradeoff
Frequently asked questions
How accurate is AI visual inspection?
In published research, very. A 2024 systematic survey of 196 open-access papers on deep learning for automated visual inspection found median classification accuracy above 95% in most reported cases, with median F1 around 0.85 to 0.90. The gap worth knowing is between that and deployment: the same survey identified roughly a three-year lag between academic breakthrough and industrial use, and 97% of the studies relied on supervised learning despite limited labelled data in real settings.
Why do vendor accuracy claims of 99% mean nothing?
Because an accuracy figure without a false-call rate and a defect distribution is not a measurement. If one part in a thousand is defective, a system that passes everything is 99.9% accurate and catches nothing. The numbers that matter are recall on the defects you care about and the false-call rate, because false calls are what destroy line throughput and operator trust.
What does a real deployment achieve?
Meaningful but specific gains. GE Aerospace reported that AI-assisted borescope inspection increased detection rates by around 34% for high-pressure compressor defects, with recall up 33.6% and precision up 13.5%. Note that precision rose alongside recall, which is the hard combination and the reason the result is credible.
What is the biggest risk in a computer vision deployment?
Getting the sensitivity threshold wrong, and the consequences can be worse than not deploying. A three-year university field trial of targeted spray technology measured 43 to 59% reductions in herbicide use against broadcast application. At the highest sensitivity, weed control was generally comparable to standard practice. At low sensitivity settings, populations of one target weed rose 280% per year. The same system, tuned differently, produced the opposite of the intended outcome.
How much training data do we need?
Less than expected, and the published distribution is revealing. Across 196 studies, dataset sizes ranged from 45 images to over 100,000, with a mean of 14,614 but a median of only 1,952. The mean is pulled by a few very large datasets. Most real work is done on small ones, which is why the practical constraint is usually labelling quality rather than volume.
Where is computer vision actually being adopted?
Manufacturing leads and the intent is strong. Deloitte's survey of 600 manufacturing executives found 80% planning to invest 20% or more of improvement budgets in smart manufacturing, and 22% planning to use physical AI within two years against 9% currently. The machine vision market itself is forecast at $15.83 billion in 2025 growing to $23.63 billion by 2030, a more modest 8.3% compound rate than the headline AI numbers.
Does adoption mean value?
Not automatically, and one clinical study makes the point sharply. A survey of academic US neuroradiology departments found 81% using AI tools, but the leaders reported minimal impact on workload so far and that tool performance remains inconsistent. The sample was only 16 departments, so treat the percentage as indicative, but the qualitative finding matches what we see elsewhere: deployment and benefit are different milestones.