You Don't Need Thousands of Photos to Train a Reliable Inspection Model
The biggest fear teams have about building their own AI is the labelling mountain. We tested how few labelled photos it actually takes, and a starting model already adapted to field-inspection imagery reached strong accuracy with as few as 15 examples per defect type.
The myth that custom AI means months of labelling
Almost every operations leader we talk to has the same worry before they start: "We'll have to label tens of thousands of photos before the AI is any good." It is a reasonable fear. The story you usually hear about machine learning is that it is hungry, feed it a mountain of examples or don't bother.
That story comes from building models from a cold start, where the AI knows nothing about your world and has to learn everything from your images alone. But that isn't how Forge works, and the difference matters enormously for how quickly your team can get value.
Why a starting model that already knows field imagery changes the math
Forge doesn't start from zero. Every new model you train begins from a starting model that has already been adapted to our library of real field-inspection imagery, cables, fittings, connectors, panels, the kind of scenes your inspectors photograph every day.
Think of it like hiring. A brand-new model is a smart graduate who has never seen your job. The field-adapted starting model is someone who has already spent years on inspection sites; they recognise the equipment, the lighting, the angles. You still have to teach them your specific faults, but they learn far faster because the context is already familiar.
In our internal benchmarks, that adapted starting model reached around 92% validation accuracy across well over a hundred field-inspection categories before any customer even adds their own data. That general grounding is exactly what gives each new defect classifier a head start.
~92% validation accuracy across well over a hundred field-inspection categories, before you add a single photo of your own.
Low-data results: meaningful gains at 50 labels, usable accuracy at 10
We tested how few labelled photos it really takes by training the same checks from both a generic starting point and the field-adapted one, then comparing.
The head start showed up most clearly when labels were scarce. A simple two-class connector check jumped from 94.7% to 98.2% accuracy, a +3.5% gain, with only 50 labelled images per class. The same check still reached 90.3% accuracy with just 10 labelled images per class. For a more demanding multi-class fault classifier, the field-adapted start added +1.7%, lifting accuracy from 83.7% to 85.4% at 200 labels per class.
In plain terms: because even 10 labelled images per class already reached 90.3%, around 15 good examples per fault is usually enough to get a genuinely useful model, with a little headroom above that floor. Even a first handful of labels can produce something worth testing. That is a labelling job your own team can finish in an afternoon, not a quarter.
- Two-class connector check: 94.7% to 98.2% (+3.5%) at 50 labels per class
- Same check: 90.3% accuracy at just 10 labels per class
- Multi-class fault classifier: 83.7% to 85.4% (+1.7%) at 200 labels per class
15 good examples per fault is often enough, a labelling job your team can finish in an afternoon, not a quarter.
Where the head start matters most, and where it stops mattering
The pattern is consistent: the fewer labels you have, the more the field-adapted starting model helps. That is the opposite of the usual AI worry, and it is good news for teams just getting going.
As you add more data, the two starting points naturally converge. With a full dataset behind it, the field-adapted model and the generic one landed in almost the same place, roughly 98.6% versus 98.7%. Once you have enough examples, the model can learn the context itself, so the head start fades.
Crucially, that early advantage came at no extra cost. The field-adapted start added accuracy without adding training time, your one-click train runs just as fast.
An honest note: it isn't magic on every dataset
We averaged every result over three runs with different random starting points, so what you see above is typical performance, not a single lucky outcome. We are not cherry-picking.
And to be straight with you: it did not win on every dataset. On one set of images, a model trained from the generic starting point actually did slightly better than the field-adapted one. That is the reality of working with real-world data, and it is why Forge makes retraining a one-click job, you can try, measure, and adjust without a data scientist in the room.
What this means for your team's first week with Forge
You do not need a labelling marathon, a consultant, or a machine-learning hire to get started. Gather around 15 clear photos of each fault you care about, draw a box around each one, and train. Most teams have a working, testable model on day one and a refined one within the week.
That is the whole promise of Forge: the people who already know your faults are the people who build the AI, and they deploy it to inspectors' phones in days, not months. The labelling mountain was always smaller than it looked.
How we measured this. Accuracy figures are averaged over three training runs with different random starting points, so they reflect typical results rather than a single lucky run. The low-data wins were measured on focused checks with a small number of defect types; gains vary by dataset, and one dataset in our tests did better starting from a generic model than from the field-adapted one. These are internal benchmark results, see the team before citing them externally.