What changes when an inspection system can reason across images?
See how ThreeV is designing inspection agents to connect a utility's questions, images and specialist models, with evidence a human inspector can review.
A utility needs to know which structures need attention and why. An inspection should help its team decide what to check, what to repair, and what can wait.
At ThreeV, we are designing inspection agents to answer those questions by reasoning across images and using specialist vision models as tools, including models a utility already owns. Our agentic harness is the system that gives an agent access to those tools, equipment knowledge and inspection evidence. The aim is an answer supported by images a human inspector can review. The diagram below shows how these roles fit together.
One structure, several questions
For drone inspections, a shot sheet sets out the views the crew should capture, such as views around a pole and close-ups of its equipment.
The photographs of a structure form a collection, whether taken by a drone moving around a pole or a helicopter flying along a line. Each view offers different evidence.
Our example contains 61 drone images of one wood pole. In plain language, the inspection form asks, "How many crossarm levels are present?" It then asks how many crossarms are present at each level. These inventory questions give later condition questions a clear subject. If a crossarm is damaged, the team needs to know which one.
We ran our own detector across the collection. The summary below covers six of the photographs: where the camera was, how many crossarm boxes the detector returned, and what a human inspector marked.
The detector returned two, four, three, six, one and two crossarm boxes in these six views. The human inspector recorded two crossarm levels with one crossarm at each level.
Across all 61 images, the detector returned between zero and six crossarm boxes per image. The most frequent result was four boxes; fourteen photographs returned two. The chart below shows the full collection.
Look at those six views again. Some boxes belong to another pole in the background. In the view from above, two boxes cover parts of the same crossarm. In the photograph of the pole's base, the detector marks part of a fence as a crossarm.
These are different problems: the wrong structure, a duplicate finding and a mistaken identification. Recognizing a crossarm can help locate it, but answering the inventory question requires connecting the views to the same physical equipment. Adding the boxes together does not do that.
The right evidence depends on the question
To answer "How many crossarm levels are present?", we need a view that separates the levels. A photograph from directly above can hide a lower crossarm behind an upper one. A wider view from the side may make both visible.
Now ask, "Is a crossarm broken, split or deteriorated?" The utility needs to identify equipment that may require closer inspection or repair. A photograph that establishes the arrangement may be too distant to show a split in the wood.
In this collection, the human inspector recorded a split crossarm and marked it in the fourth photograph. The orange mark in the summary shows that annotation. Our detector identifies crossarms, but it has no separate class for a split.
We are designing our agents to bring the question, equipment knowledge and vision tools together. A vision-language model can work with images and language together, helping connect the wording of a question to what a photograph shows. Tool choice is meant to follow the question and its context.
Camera metadata and short image descriptions help establish what each view offers. We are building agents that draw on thousands of real inspections, so those descriptions can carry the context of the utility's equipment and questions.
My advice to a utility is to ask for the supporting images alongside each answer. Can a human inspector see the relevant component? Is it on the correct structure? Does the view show enough to judge the condition being reported?
Knowing when another view is needed
Consider a related question: "Is an insulator cracked, chipped or broken?"
A trained model might find the insulator in several photographs. That does not mean any of them shows its surface clearly enough to judge damage. The views could be too distant, partly obscured, or taken from the wrong side.
In that situation, a useful response explains what could not be assessed and what evidence would help. The example below identifies a closer view of the insulator's surface as the missing shot.
The human inspection team sets the standard for what counts as enough evidence. An agent should work within that standard and refer uncertain cases for review. A request for a better photograph can then improve the next shot sheet instead of leaving someone to guess why the question was unanswered.
Fixed workflows can still serve well-defined tasks with consistent questions and capture patterns. Reasoning becomes useful when the system must account for a changed definition, a missing view or terminology that differs from a detector's labels. Specialist models remain tools for the parts they do well.
Bring us the questions your team needs answered
If your utility already collects inspection imagery, we would like to help you assess what it can support. Connect with us for a free assessment of your questions and a sample of your imagery.
A useful starting point is one inventory question, one condition question and one question your team finds difficult to answer from the available views. Together, we can examine the supporting images, what still needs human judgment and where another photograph would help.
We build Vision around those inspection questions. Velocity makes specialist models available as tools, including models a utility brings itself. Bring us the questions that matter to your team, and we can assess what your imagery can support.