All posts
    Engineering

    Image Checks From a Few Dozen Labelled Photos

    AltSlate LabsOctober 9, 202611 min read

    Many checks a product needs on an image are small and typed: is this part defective, which of four defect types is it, how ripe is this banana on a four-level scale. For those you want a probability you can trust and a way for the system to say "I have not seen images like this". We built certo-vision, an open-source layer that does both from a few dozen labelled photos per answer. On gross defects and ripeness it was right 93 to 95% of the time, and in our trial it rejected every out-of-scope image. On small, subtle defects it fails, and we show where that line is.

    93–95%
    Correct on coarse yes/no checks
    gross defects and ripeness, 19–45 labelled photos per answer
    100%
    Out-of-scope images rejected in the trial
    while abstaining on 2–10% of in-scope images
    72 ms
    To read one image, shared by every question
    one GPU, batch size one; two questions add about 1 ms

    01 — the problem

    Small image decisions need a probability and a way to say no

    An inspection line asks "is this bottle defective?" of every photo. A generative vision model can answer, but it returns free text you have to parse. Getting a trustworthy probability and a refusal out of it is neither cheap nor built in. We did not compare accuracy against one; the case here is cost and calibration.

    So we fixed a contract. Each question declares its type: yes/no, a choice among named options, or a score on a levelled scale. Each answer comes back as that type with probabilities, plus two fields the caller can act on. confidence says how close the image is to the photos the question was set up on; it is not the answer's probability. abstain is true when that confidence is too low to answer.

    Every question is answered from the same embedding of the image: a list of 1,152 numbers that summarises it, in a space shared with text, so a photo and a sentence describing it land close together. Fifty questions cost one pass through the encoder plus fifty cheap comparisons. That makes ten or twenty checks on every upload affordable, where a model call per question is not.

    02 — how it works

    One frozen encoder, then a few cheap steps

    The encoder is SigLIP 2, an open image-text model of about 1.1 billion parameters. We never train it. Everything that adapts to a deployment happens after it, and fits in one small file.

    photoimage encoderfrozen · runs oncecompare with each optiontext + example photoscalibrated answeryes 0.94 · no 0.06nearest labelled photohow familiar is this?confidence 0.82abstain: no
    One read of the image, two separate outputs. The top path answers the question; the bottom path decides whether the image looks like the labelled photos at all. The answer's probability and the scope confidence are kept apart on purpose.

    Write each option as a sentence. For "is the banana overripe?", the yes option reads "yes, the banana is overripe with brown spots or mostly brown". It is embedded once. The answer's probabilities come from how similar the photo is to each option's sentence.

    Blend in labelled photos. Each option's sentence embedding is mixed with the average embedding of the labelled photos for that option: 70% text, 30% photos. It is one line of arithmetic, and section 04 shows it does almost all the work.

    Calibrate. On the same labelled photos we fit one temperature and one bias per option, so that an answer given at 80% is right about 80% of the time. That is at most 256 numbers, fitted with plain gradient descent in seconds.

    Check scope. We keep the labelled photos' embeddings. For a new image, we find the most similar labelled photo and compare that similarity with how close labelled photos usually are to their own nearest neighbour. The abstain threshold is set so that 95% of the labelled photos would pass. The gate never looks at the labels.

    All of this is saved to one calibration file. If someone edits a question's wording after calibration, the server does not guess: it falls back to the uncalibrated answer and reports calibrated: false.

    03 — results

    Coarse checks pass; subtle defects on a small part fail

    We chose nine questions on two public datasets that resemble client work: MVTec AD, industrial inspection photos of bottles, hazelnuts, carpet and screws, and a banana ripeness set. Each set was split 30% for calibration and 70% for testing, five times over with different random splits. Every row reports the majority rate, the score from always giving the most common answer. A model that does not beat it has learned nothing.

    with labelled photossentences onlymajority rate0.00.51.00.9hazelnut: defective?0.95carpet: defective?0.94bottle: defective?0.93banana: overripe?0.93hazelnut: defect type (4)0.87banana: ripeness (4 levels)0.83banana: stage (4)0.79screw: defective?0.71screw: defect type (5)0.36
    Accuracy per question. Bars: calibrated with labelled photos, mean of five splits. Circles: the sentences alone (zero-shot). Ticks: the majority rate. The dashed line is 0.9. Labelled photos per option ranged from 19 to 45, except the defect-type rows (5 to 7).

    Gross defects on bottles, hazelnuts and carpet landed between 0.93 and 0.95 with 19 to 27 labelled photos per answer. The four-level ripeness score was within 0.23 of a level on average. After calibration, the gap between stated probability and actual accuracy was 3 to 6 points on these rows.

    The screw set failed: thread damage, scratches and chipped tips on a small part against a plain background. It reached 0.71 on yes/no and chance on defect type, with the worst calibration in the table. One summary of the whole image does not keep that much detail. Methods that compare small patches of the image are the known fix if such questions matter.

    The defect-type rows had only five to seven labelled photos per option, and their results moved by up to 5 points between splits. Treat them as indicative. The calibration tool warns below ten photos per option.

    04 — compared with a plain encoder

    What a plain encoder gets wrong, and what each step fixes

    The simplest way to use an image-text encoder is to embed the photo, embed one sentence per answer and pick the closest. That is the "sentences only" circle in the chart above. It has one great property: each image is read once, and every extra question is a cheap comparison. certo-vision keeps that property. Everything it adds sits after the encoder and costs about a millisecond.

    Plain encodercerto-vision
    What defines an answerOne sentence per optionThe sentence, moved toward your labelled photos
    Industrial yes/no accuracyAt or near the majority rate0.93–0.95 on bottle, hazelnut, carpet
    Gap between probability and accuracy20–40 points3–6 points
    A photo unlike any it has seenStill picks an answerSeparate check; abstains
    Cost of one more questionOne comparisonOne comparison
    Training the encoderNoneNone

    The differences come from four things we learned about what an encoder does and does not know.

    The encoder knows what is in the photo, not where your line is

    An encoder like SigLIP 2 learns from huge numbers of photos paired with captions. It places images by what they show: bananas near bananas, bottles near bottles. It never learned where your inspection line draws "defective", or how brown a banana must be before your buyers call it overripe.

    Our reading of the results: on one production line, every photo shows the same object from the same angle, so they all land in one small patch of the encoder's space. The sentence "a defective bottle" points roughly the right way, but not precisely enough to split that patch in two. Blending in the labelled photos moves each answer toward where your photos actually are.

    THE ENCODER'S SPACEphotos from your cameraripeoverripe★sentence: "overripe banana"★after blending in labelled photos:moved toward the real onestexture photofar from every label: abstain
    An illustration, not measured data. Think of the encoder's space as a map. The sentence is a guess at where "overripe" sits (open star). Blending in labelled photos moves it 30% of the way toward the real overripe photos (filled star). A photo far from every labelled photo, like a texture close-up, is where the scope check abstains.

    The numbers fit this picture. With sentences alone, the industrial yes/no questions scored at or within two points of the majority rate, and the banana questions were 25 to 35 points below the version with labelled photos. A synthetic five-level blur scale went from 0.34 to 0.88 only once example photos were blended in. "Blurry" is a word the encoder knows; "level 3 on our scale" is not.

    A sentence and its opposite land close together

    Yes/no questions are harder for sentences than they look. The "no" answer has to be written as a negation, such as "no, the banana is not overripe". An encoder places text by what it is about, and that sentence is mostly about an overripe banana. With the other encoder we tested, EmbeddingGemma 2, the sentences-only "is it a cat?" said yes to nearly everything: 0.126 accuracy where always saying no scored 0.900. Labelled photos fix this too, because the "no" answer is pulled toward real photos of things that are not cats instead of resting on a sentence about cats.

    A similarity score is not a probability

    The encoder returns similarities, and for photos of one object they typically differ only slightly between answers. Turned straight into percentages, they mean little: sentences-only probabilities were off from the real accuracy by 20 to 40 points. Calibration fits a temperature, which stretches or squeezes those gaps, and a bias per answer, which corrects for a sentence that sits a little closer to every photo. After that the gap fell to 3 to 6 points.

    Calibration has a hard limit: it reshapes probabilities but cannot add information. On a standard ten-class photo set, a calibrated "is it a cat?" without labelled photos scored exactly the majority rate, 0.900. The labelled photos supply the information; calibration makes the numbers honest.

    Always answering is not the same as knowing

    A plain encoder has no "none of these". Its percentages must add up to 100% across the answers you offered, so a texture close-up gets split between "cat" and "not cat" like any other photo, often with high confidence. certo-vision asks a second, separate question: is this photo near any labelled photo? That uses the encoder's map directly and needs no labels. Section 05 shows how well it works and where it misfires.

    The practical lesson from all four: spend effort on collecting labelled photos from the real camera, not on rewording the questions. The sentences are a starting guess that the photos correct.

    05 — knowing when to abstain

    The scope gate catches unfamiliar images; the answer's probability does not

    In the trial, the gate rejected 100% of out-of-scope images for every question set: the other MVTec objects shot in the same style, texture photos and cat photos. The cost was abstaining on 2 to 10% of in-scope images.

    The tempting shortcut is to abstain when the answer's own probability is low. It does not work. With "is it a cat?" calibrated on ordinary photos, 97% of texture close-ups and every satellite tile got an answer above 0.9. The model was most certain about images unlike anything it had been shown.

    The chart compares the two signals with AUROC, a measure of how well a score separates in-scope from out-of-scope images: 1.0 is perfect, 0.5 is a coin flip. A third signal we tried, agreement between answers from shortened embeddings, was close to a coin flip and we removed it.

    nearest-photo gateanswer's probabilitycoin flip1.000.68textures0.990.48satellite tiles0.980.84pet photos
    Separating familiar from unfamiliar images on the yes/no question, 700 in-scope test photos against 200 from each out-of-scope set. The nearest-photo gate separates all three; the answer's probability is near or below a coin flip on two.

    The pet photos show the gate's limit. The question was "is it a cat?" and the model answered most of them correctly, yet the gate rejected 93%: they were high-resolution photos, and the labelled set was 96-pixel thumbnails. The gate detects a change in the kind of image, not an off-topic question. Labelled photos must come from the camera you will deploy with, and the abstain rate belongs next to accuracy in every report.

    06 — the ceiling

    Fine detail needs a crop, and beyond 92% a dedicated model

    To put a number on the limit, we built 600 synthetic form snippets: a printed field name and four boxes of handwritten digits. The printed field name was read correctly every time. Asked about the digits through the whole image, the layer got 29 to 50% per digit.

    Cropping each box first raised that to 0.918 per digit, and 71% for the whole four-digit number. Doubling the labelled crops did not move it, so the limit is the image summary, not the amount of data. A simple linear model on the same embeddings reached 0.961. The intuition: a prototype is an average that weighs all 1,152 numbers equally, while a trained model learns which of them matter for this question. When fine detail decides the answer and there are enough examples, that pays off. A small digit network fine-tuned on the 720 labelled crops reached 0.979 per digit and 92% for the whole number. It trains in seconds and can sit behind the same typed question.

    Our rule of thumb: if a person could answer from a thumbnail, this layer fits. Small defects, text, counting and spatial relations should go to a different model.

    07 — the takeaway

    Thirty to fifty photos per answer, from the camera you will use

    Give it thirty to fifty labelled photos per option from the deployment camera, and you get calibrated, typed answers with an abstain flag, for questions a person could answer from a thumbnail.

    Some things are not shown yet. Every trial calibrated and tested on photos from the same source, so we have not measured how much accuracy survives a camera change; that is the first test on real client data. We did not run a generative vision model on the same questions. And the score type has been tested on one real dataset.

    The calibration command checks itself before saving: it holds out part of the labelled photos and prints accuracy, calibration error, the majority rate and the abstain rate. There are no trained weights to ship; the only learned state is the per-deployment calibration file. The code, the evaluation harness and the result files behind every number here are open source under Apache-2.0 on GitHub.