The gap between the pilot and the line.
Almost every automated visual inspection project demonstrates well. A few hundred images, a model trained over a weekend, accuracy figures that make the business case obvious. Then it goes onto the line and the numbers fall apart. The usual explanation offered is that the model needs more data. Occasionally true, usually not. The pilot was run under conditions the line does not reproduce: consistent lighting, parts presented squarely, a camera nobody had knocked, and a dataset built from defects someone had already found. The line has none of those, and each difference costs accuracy in a way no amount of retraining recovers.
Lighting decides more than the model does.
On most inspection problems, lighting is the single largest determinant of accuracy, and it is settled before any model work begins. Diffuse, direct, backlit or angled illumination each make different defect classes visible, and a surface scratch that is obvious under raking light can be invisible under a flat panel. Ambient light changing across a shift, or a door opening onto daylight, will shift your input distribution in a way the model has no way to reason about. Fixing the lighting and enclosing the inspection station is unglamorous and routinely cheaper than the modelling work it replaces.
Your dataset is probably not what you think.
Inspection datasets are almost always built from the defects that were caught. That is a biased sample in a specific and awkward direction: it excludes the defects that escaped, which are precisely the ones the system exists to find. It also tends to over-represent obvious examples, because those are the ones photographed and filed. A useful dataset needs borderline cases, near-misses and the ambiguous parts two inspectors would grade differently. Collecting those takes longer than collecting clear examples, and it is the difference between a model that agrees with your best inspector and one that agrees with your easiest decisions.
Label disagreement is data, not noise.
Have two experienced inspectors grade the same hundred parts and you will find they disagree on some of them. That disagreement rate is the practical ceiling on what any model can achieve, and knowing it changes the conversation. If your people agree ninety-two percent of the time on borderline parts, a model reporting ninety-six percent accuracy is not outperforming them — it is being graded against labels that were themselves uncertain. Measuring inter-rater agreement early is one of the cheapest things you can do and one of the most frequently skipped.
Edge deployment is bounded by physics, not software.
Inspection at line rate usually means processing on the line, because sending every frame to a server is either too slow or too expensive. That puts you inside real constraints: power budget, thermal envelope, and how much inference a unit can sustain before it throttles. A model that runs comfortably on a development machine may not fit those limits, and the fix is rarely a better model — it is a smaller one, or a two-stage design where a fast filter passes only candidates to a heavier check. Establishing the hardware envelope before selecting an architecture avoids a rewrite.
Measuring accuracy in a way that survives a review.
A single accuracy percentage is not a useful measure for inspection, because the two error types have very different costs. A false negative ships a defect to a customer. A false positive stops a line and consumes an inspector's time. You need both rates, measured separately, against a held-out set that reflects real production mix rather than the training data. And you need to decide in advance which error you would rather make, because tuning the threshold trades one directly for the other. Teams that skip that decision end up with a system nobody trusts: it either cries wolf until operators ignore it, or it misses enough to be worth double-checking, which defeats the point.
Drift is the year-two problem.
A system validated in month one will degrade. Suppliers change, tooling wears, a batch arrives with a different surface finish, a camera is bumped during maintenance. None of these announce themselves, and the model will keep producing confident outputs throughout. The protection is monitoring the inputs as well as the outputs: track the distribution of what the camera sees, flag when it shifts, and hold back a periodically refreshed validation set. This is ordinary operational engineering rather than machine learning, and it is what separates a system still trusted in year two from one quietly bypassed.
Where to start.
Pick one defect class that is expensive, reasonably common, and visually distinct. Fix the lighting. Measure how often your own inspectors agree on it. Build a dataset that includes the borderline cases. Then train, and evaluate against both error rates separately. It is a narrower start than most proposals suggest, and it produces something running on a line rather than an impressive result in a slide deck.
