/ Blog

The unit of progress is a question, not a model

Model-centric roadmaps produce steady technical progress and unpredictable business progress. Organising the work around individual inspection questions fixes the mismatch, and changes how a team is staffed.

Prajwal PaudyalPhD, machine learning18 August 20265 min readOperating model

Most machine learning roadmaps I have seen are lists of models. Train a detector for this component, fine tune something for that condition, improve the segmentation. Each item is real work and each one can be delivered. The problem is that nobody outside the team can tell what any of it bought.

The alternative is to make the unit of work a question from the inspection form. Not a model, not a dataset, not a capability. One question, with a target, measured against what a qualified inspector says about the same images.

It sounds like a reporting change. It is closer to an organisational one.

One row per question, each with a target and a disposition. The numbers here are an illustration of the shape, not our results.Scroll the diagram sideways

Why the model-shaped roadmap fails

Three failure modes, all of which I have participated in rather than merely observed.

The metric moves and nothing ships. A detector improves. The answers a customer sees do not change, because the errors that mattered were somewhere else in the chain. This is not a hypothetical: it is the normal case once detection is decent, and I have written about why the chain has more links than people expect.

Progress cannot be sequenced against value. If the roadmap is a list of models, there is no principled way to decide which one to do next, because models do not have business value individually. Questions do. Some questions appear on every structure and drive real spend. Others are rare and cheap to answer manually. That ordering is invisible in a model-shaped plan and obvious in a question-shaped one.

Nobody can say what is done. "The vegetation model is at 0.78" is not a statement about readiness. "Question 22 agrees with our inspectors 91 percent of the time, on 400 structures, and abstains on 6 percent" is.

What a question-shaped plan looks like

Each question gets a row, and each row gets three things.

A target, expressed as agreement with a qualified reviewer. Not accuracy against a labelled set assembled by the team that built the model. Agreement with the people whose judgement the program already relies on, on imagery from the program.

A disposition. I use three. Ship means the answer goes out with review by exception. Assist means the answer is shown to a reviewer as a starting point and they decide. Hold means it is not shown at all. The distinction matters because it lets a question be valuable before it is finished. An assist-grade answer that saves a reviewer forty seconds is worth shipping, and it starts generating corrections, which is how it gets better.

An owner. One person who can say why the number is what it is.

Hold, assist, ship. A question can be useful well before it is finished.Scroll the diagram sideways

That is the whole apparatus. What it produces is a board where the state of the product is legible to somebody who does not work on models, which turns out to be most of the people who need to make decisions about it.

The staffing consequence, which is the real point

Here is what changed for us when the unit became the question rather than the model.

Work stopped being organised by technique. A question needs whatever it needs: sometimes a new detector, often not. Frequently what it needs is better subject selection, or a change to how the question is phrased to the model, or an abstention rule, or labelled attributes on data that already exists. When the work item is a model, only the people who train models can pick it up. When the work item is a question, a much wider set of people can make it move, and the bottleneck stops being a single specialism.

It also changes what a product manager does. The interesting question stops being "what should we build" and becomes "which questions, in what order, to what grade" which is answerable with data the customer already has: how often each question appears, what answering it costs today, what a wrong answer costs.

And it makes the review loop a first-class part of the plan rather than an afterthought, because the corrections coming back are the mechanism by which a row moves from hold to assist to ship. Review capacity becomes a planned input, not a cost centre to be squeezed.

Where it gets uncomfortable

Two honest costs, because a method that only has benefits is being sold rather than described.

The numbers are lower, at first, and they are public. Measuring against qualified human agreement at the level of the answer produces smaller numbers than component metrics do. That is because it is measuring something harder and more real. A team used to reporting mAP will find the first question-level scoreboard uncomfortable, and somebody senior has to be willing to say out loud that the lower number is the honest one.

It requires the questions to be written down properly. Inspection programs accumulate questions over years, in spreadsheets, with overlapping meanings and inconsistent vocabulary. Turning that into a set of questions each of which can be answered, scored, and owned is unglamorous work with no model in it. It is also the precondition for everything above, and skipping it is why so many pilots cannot say whether they succeeded.

What I would tell another CTO

If your AI roadmap is a list of models, you will produce technical progress that is real and business progress that is hard to predict, and the gap between those two things will eventually be a credibility problem.

Rewrite it as a list of questions. Put agreement with qualified humans next to each one. Add a disposition so that partial answers can be useful. Then let the models be an implementation detail of the rows, which is what they always were.

Written by

Prajwal Paudyal

PhD, machine learning

Reach out

Want to argue with any of this?

We would rather hear where you think it is wrong than where you think it is right.

Get in touch