Image Checks From a Few Dozen Labelled Photos
Many checks a product needs on an image are small and typed: is this part defective, which of four defect types is it, how ripe is this banana on a four-level scale. For those you want a probability you can trust and a way for the system to say "I have not seen images like this". We built certo-vision, an open-source layer that does both from a few dozen labelled photos per answer. On gross defects and ripeness it was right 93 to 95% of the time, and in our trial it rejected every out-of-scope image. On small, subtle defects it fails, and we show where that line is.
01 — the problem
Small image decisions need a probability and a way to say no
An inspection line asks "is this bottle defective?" of every photo. A generative vision model can answer, but it returns free text you have to parse. Getting a trustworthy probability and a refusal out of it is neither cheap nor built in. We did not compare accuracy against one; the case here is cost and calibration.
So we fixed a contract. Each question declares its type: yes/no, a choice among named options, or a score on a levelled scale. Each answer comes back as that type with probabilities, plus two fields the caller can act on. confidence says how close the image is to the photos the question was set up on; it is not the answer's probability. abstain is true when that confidence is too low to answer.
Every question is answered from the same embedding of the image: a list of 1,152 numbers that summarises it, in a space shared with text, so a photo and a sentence describing it land close together. Fifty questions cost one pass through the encoder plus fifty cheap comparisons. That makes ten or twenty checks on every upload affordable, where a model call per question is not.
02 — how it works
One frozen encoder, then a few cheap steps
The encoder is SigLIP 2, an open image-text model of about 1.1 billion parameters. We never train it. Everything that adapts to a deployment happens after it, and fits in one small file.
Write each option as a sentence. For "is the banana overripe?", the yes option reads "yes, the banana is overripe with brown spots or mostly brown". It is embedded once. The answer's probabilities come from how similar the photo is to each option's sentence.
Blend in labelled photos. Each option's sentence embedding is mixed with the average embedding of the labelled photos for that option: 70% text, 30% photos. It is one line of arithmetic, and section 04 shows it does almost all the work.
Calibrate. On the same labelled photos we fit one temperature and one bias per option, so that an answer given at 80% is right about 80% of the time. That is at most 256 numbers, fitted with plain gradient descent in seconds.
Check scope. We keep the labelled photos' embeddings. For a new image, we find the most similar labelled photo and compare that similarity with how close labelled photos usually are to their own nearest neighbour. The abstain threshold is set so that 95% of the labelled photos would pass. The gate never looks at the labels.
All of this is saved to one calibration file. If someone edits a question's wording after calibration, the server does not guess: it falls back to the uncalibrated answer and reports calibrated: false.
03 — results
Coarse checks pass; subtle defects on a small part fail
We chose nine questions on two public datasets that resemble client work: MVTec AD, industrial inspection photos of bottles, hazelnuts, carpet and screws, and a banana ripeness set. Each set was split 30% for calibration and 70% for testing, five times over with different random splits. Every row reports the majority rate, the score from always giving the most common answer. A model that does not beat it has learned nothing.
Gross defects on bottles, hazelnuts and carpet landed between 0.93 and 0.95 with 19 to 27 labelled photos per answer. The four-level ripeness score was within 0.23 of a level on average. After calibration, the gap between stated probability and actual accuracy was 3 to 6 points on these rows.
The screw set failed: thread damage, scratches and chipped tips on a small part against a plain background. It reached 0.71 on yes/no and chance on defect type, with the worst calibration in the table. One summary of the whole image does not keep that much detail. Methods that compare small patches of the image are the known fix if such questions matter.
The defect-type rows had only five to seven labelled photos per option, and their results moved by up to 5 points between splits. Treat them as indicative. The calibration tool warns below ten photos per option.
04 — compared with a plain encoder
What a plain encoder gets wrong, and what each step fixes
The simplest way to use an image-text encoder is to embed the photo, embed one sentence per answer and pick the closest. That is the "sentences only" circle in the chart above. It has one great property: each image is read once, and every extra question is a cheap comparison. certo-vision keeps that property. Everything it adds sits after the encoder and costs about a millisecond.
| Plain encoder | certo-vision | |
|---|---|---|
| What defines an answer | One sentence per option | The sentence, moved toward your labelled photos |
| Industrial yes/no accuracy | At or near the majority rate | 0.93–0.95 on bottle, hazelnut, carpet |
| Gap between probability and accuracy | 20–40 points | 3–6 points |
| A photo unlike any it has seen | Still picks an answer | Separate check; abstains |
| Cost of one more question | One comparison | One comparison |
| Training the encoder | None | None |
The differences come from four things we learned about what an encoder does and does not know.
The encoder knows what is in the photo, not where your line is
An encoder like SigLIP 2 learns from huge numbers of photos paired with captions. It places images by what they show: bananas near bananas, bottles near bottles. It never learned where your inspection line draws "defective", or how brown a banana must be before your buyers call it overripe.
Our reading of the results: on one production line, every photo shows the same object from the same angle, so they all land in one small patch of the encoder's space. The sentence "a defective bottle" points roughly the right way, but not precisely enough to split that patch in two. Blending in the labelled photos moves each answer toward where your photos actually are.
The numbers fit this picture. With sentences alone, the industrial yes/no questions scored at or within two points of the majority rate, and the banana questions were 25 to 35 points below the version with labelled photos. A synthetic five-level blur scale went from 0.34 to 0.88 only once example photos were blended in. "Blurry" is a word the encoder knows; "level 3 on our scale" is not.
A sentence and its opposite land close together
Yes/no questions are harder for sentences than they look. The "no" answer has to be written as a negation, such as "no, the banana is not overripe". An encoder places text by what it is about, and that sentence is mostly about an overripe banana. With the other encoder we tested, EmbeddingGemma 2, the sentences-only "is it a cat?" said yes to nearly everything: 0.126 accuracy where always saying no scored 0.900. Labelled photos fix this too, because the "no" answer is pulled toward real photos of things that are not cats instead of resting on a sentence about cats.
A similarity score is not a probability
The encoder returns similarities, and for photos of one object they typically differ only slightly between answers. Turned straight into percentages, they mean little: sentences-only probabilities were off from the real accuracy by 20 to 40 points. Calibration fits a temperature, which stretches or squeezes those gaps, and a bias per answer, which corrects for a sentence that sits a little closer to every photo. After that the gap fell to 3 to 6 points.
Calibration has a hard limit: it reshapes probabilities but cannot add information. On a standard ten-class photo set, a calibrated "is it a cat?" without labelled photos scored exactly the majority rate, 0.900. The labelled photos supply the information; calibration makes the numbers honest.
Always answering is not the same as knowing
A plain encoder has no "none of these". Its percentages must add up to 100% across the answers you offered, so a texture close-up gets split between "cat" and "not cat" like any other photo, often with high confidence. certo-vision asks a second, separate question: is this photo near any labelled photo? That uses the encoder's map directly and needs no labels. Section 05 shows how well it works and where it misfires.
The practical lesson from all four: spend effort on collecting labelled photos from the real camera, not on rewording the questions. The sentences are a starting guess that the photos correct.
05 — knowing when to abstain
The scope gate catches unfamiliar images; the answer's probability does not
In the trial, the gate rejected 100% of out-of-scope images for every question set: the other MVTec objects shot in the same style, texture photos and cat photos. The cost was abstaining on 2 to 10% of in-scope images.
The tempting shortcut is to abstain when the answer's own probability is low. It does not work. With "is it a cat?" calibrated on ordinary photos, 97% of texture close-ups and every satellite tile got an answer above 0.9. The model was most certain about images unlike anything it had been shown.
The chart compares the two signals with AUROC, a measure of how well a score separates in-scope from out-of-scope images: 1.0 is perfect, 0.5 is a coin flip. A third signal we tried, agreement between answers from shortened embeddings, was close to a coin flip and we removed it.
The pet photos show the gate's limit. The question was "is it a cat?" and the model answered most of them correctly, yet the gate rejected 93%: they were high-resolution photos, and the labelled set was 96-pixel thumbnails. The gate detects a change in the kind of image, not an off-topic question. Labelled photos must come from the camera you will deploy with, and the abstain rate belongs next to accuracy in every report.
06 — the ceiling
Fine detail needs a crop, and beyond 92% a dedicated model
To put a number on the limit, we built 600 synthetic form snippets: a printed field name and four boxes of handwritten digits. The printed field name was read correctly every time. Asked about the digits through the whole image, the layer got 29 to 50% per digit.
Cropping each box first raised that to 0.918 per digit, and 71% for the whole four-digit number. Doubling the labelled crops did not move it, so the limit is the image summary, not the amount of data. A simple linear model on the same embeddings reached 0.961. The intuition: a prototype is an average that weighs all 1,152 numbers equally, while a trained model learns which of them matter for this question. When fine detail decides the answer and there are enough examples, that pays off. A small digit network fine-tuned on the 720 labelled crops reached 0.979 per digit and 92% for the whole number. It trains in seconds and can sit behind the same typed question.
Our rule of thumb: if a person could answer from a thumbnail, this layer fits. Small defects, text, counting and spatial relations should go to a different model.
07 — the takeaway
Thirty to fifty photos per answer, from the camera you will use
Give it thirty to fifty labelled photos per option from the deployment camera, and you get calibrated, typed answers with an abstain flag, for questions a person could answer from a thumbnail.
Some things are not shown yet. Every trial calibrated and tested on photos from the same source, so we have not measured how much accuracy survives a camera change; that is the first test on real client data. We did not run a generative vision model on the same questions. And the score type has been tested on one real dataset.
The calibration command checks itself before saving: it holds out part of the labelled photos and prints accuracy, calibration error, the majority rate and the abstain rate. There are no trained weights to ship; the only learned state is the per-deployment calibration file. The code, the evaluation harness and the result files behind every number here are open source under Apache-2.0 on GitHub.