Case Study: Reducing AI Inspection False Positives from 8% to 1.2% on an SMT Line in 90 Days
Manufacturers often buy AI visual inspection, deploy it, and then disable it. The system catches defects, but false positives create rework bottlenecks that operators refuse to trust. This case study covers 90 days on an SMT line. No hardware changes, no algorithm vendor swap. False-positive rates dropped from 8% to 1.2%, and rework headcount fell with it. The difference came down to data feedback loops, on-floor presence, and building back operational trust.
By
Tenten AI 交付團隊
產業交付
Published
November 14, 2025
Read time
5 分鐘

AOI systems already catch defects. The production problem is false positives, good boards flagged as failures, forcing rework crews to manually inspect every flagged board. This line started at 8% false positives. That translates to roughly 12,000 boards per day, 960 sent to rework, two operators spending each shift clicking OK or NG on false calls.
The customer had deployed AI inspection before contacting us. The vendor's system identified opens, tombstones, placement shifts, and cold solder during commissioning. Three months later, the engineering team treated the AI output as reference only, relying on full manual rework. The algorithm was functional but had never trained on this production line's materials or reflow profile.
What changed in 90 days (numbers first)
We did not swap cameras or replace the algorithm vendor. Changes were made to data collection, process, and operator involvement.
| Metric | Baseline | Day 90 | Change |
|---|---|---|---|
| False-positive rate | 8.0% | 1.2% | −85% |
| Manual rework volume per day | ~960 boards | ~144 boards | −85% |
| Rework headcount per shift | 2 | 0.5 | −1.5 |
| Actual miss rate (customer feedback) | 0.03% | 0.02% | Flat, slight improvement |
| Floor engineer adoption rate | ~30% | 92% | Actually in use |
The adoption rate distinguishes this outcome. The other metrics are results. This row indicates whether the system remained in active use.
Days 1 to 30: Diagnose the failures
The field deployment engineer worked on the production line rather than remotely. The first step was retrieving every board flagged NG over the previous three months that manual inspection had determined was acceptable. This yielded over 4,000 images. Pattern analysis revealed that 60% of the false positives fell into two categories: dark-colored capacitors misidentified as tombstones due to light reflection at specific angles, and lead-free solder paste with a matte finish misidentified as insufficient reflow at certain temperatures.
The model was not broken. It had never learned what normal looked like on this line with these materials and these reflow profiles. Generic models train on generic production environments. This one encountered non-standard materials and non-standard thermal curves. This mirrors challenges in customer service and fraud detection: demo environments use clean data, production floors do not.
Days 31 to 60: Close the data feedback loop
We built a mechanism: when an operator overrides a call during manual inspection, the image and decision automatically feed into a labeling queue. The team lead spends 20 minutes each day validating label quality. After two weeks, we fine-tuned the two worst-performing defect categories using the accumulated data without retraining the full model.
The first fine-tune attempt created a problem. We fed all data at once without staged validation, and miss rates increased. The model became too permissive. We shifted to category-specific tuning with separate thresholds for false positives and false negatives. We prioritized holding miss rates stable over rapid reduction of false positives. In inspection, misses ship and customers absorb the cost. That cannot be undone.
Days 61 to 90: Building operator trust
Model tuning addresses only half of the problem. The original system failed because operators had lost trust in the output. Engineers had experienced false positives and stopped treating the system as reliable. Over the final 30 days, we made three changes. We displayed confidence scores and defect-detection bounding boxes in the re-inspect interface so operators could understand the system's reasoning. We posted daily counts of false positives and false negatives in a morning review so operators saw progress. We gave senior engineers direct access to adjust thresholds, with our engineering team present to advise.
By day 90, adoption rose from 30% to 92%. The rework queue dropped from 960 to 144 boards daily, so operators preferred the new system.
What we learned
AI inspection deployments fail mostly because of data feedback and operator trust, not model quality. False positives damage adoption more than false negatives because they exhaust operators consistently. Results in 90 days depend on having engineers work on-site rather than shipping a model remotely.
Engineers work on-site to build data feedback loops, set dual thresholds, and establish dashboards for daily review. We remain until the system is in production and operators trust it enough to use it. Sustained adoption, in this case 92%, is the actual measure of success.

One stuck workflow
is enough to begin
Tell us what the team does today, where it breaks down, and what a better working day should look like.