← Back

Calibrated dog breed classifier

Measured a six-point generalisation gap that the standard benchmark cannot show, because its test images sit in the training set of every pretrained model. The classifier itself names one of 120 dog breeds, refuses the photos it cannot answer for, and reports a confidence that matches how often it is right.

Try it yourself

Upload a photo. The model checks whether it looks like a dog, then names a breed and says how sure it is.

At a glance

Five models compared on two datasets, then the winner exported, checked image by image against the original, and served from a container. Every number below is in the repository.

Calibration error, 8580 test images
5.71% → 0.84%
Real dog photos the gate accepts
98.0%
Cats it accepts
4.47%
Accuracy, 8580 test images
93.11%
Lost by changing the photo source
6.2 points
p95 in production, on the VPS
123 ms

Making the percentage mean something

The demo shows a percentage, which is a promise: say 80% and you should be right four times in five. A raw softmax output does not keep that promise, and expected calibration error is the size of the gap.

Textbooks say networks overstate their confidence and the fix is to flatten the output. This one does the opposite. The temperature that corrects it came out at 0.76, below 1, which sharpens.

Expected calibration errorUncalibratedT = 0.76
Validation, 1800 images5.10%1.22%
Test, 8580 images5.71%0.84%

The temperature is fitted on validation alone, so the test row is measured on 8580 images the fitting never saw. It is one scalar, folded into the exported graph rather than left for the server to remember. Dividing by a positive constant cannot reorder the logits, so not one prediction changes. Only the promise does.

How the temperature was fitted, and why it landed below one →

Refusing what it can't answer for

Softmax normalises over the 120 outputs, so it returns a distribution conditioned on the input already being a dog. A cat gets a peaked one just as readily, and no threshold on the confidence recovers that.

So the decision happens a layer earlier, on the 768 penultimate features (Lee et al., 2018): per-breed means, one shared covariance, and the smallest Mahalanobis distance to any centre. The covariance is fitted on within-class residuals rather than raw features, because on raw features it would describe the distance between breeds and make the gap between two of them look ordinary. That gap is where a cat lands.

Threshold calibrated onVal dogsOxford dogsOxford cats
Stanford validation95.0%87.8%0.25%
Oxford dogs, 95% TPR97.8%95.0%1.18%
Oxford dogs, 98% TPR (used)99.0%98.0%4.47%

The negatives are Oxford's 2371 cat photos, not blank walls. Fur, four legs, a muzzle, the same pet-photo framing. The threshold comes off Oxford's dogs too, because a visitor's photo will look like theirs and not like Stanford's.

A dog detector would be the other way to do this. It was not built, for two reasons. The budget was fixed before any model was trained, 300 ms p95 on CPU and an image under 500 MB, and a detector is a second model inside both. The gate instead reuses features the classifier has already computed, so it adds no forward pass at all. The escalation criterion was written as a number: build the detector if the gate lets through more than 10% of cats. It lets through 4.47%.

The two also answer different questions. A detector reports whether a dog is present. The gate reports whether an image is one this classifier can answer for, and the next section is where those two come apart.

The distance, the shared covariance, and the threshold sweep →

Step back from the dog and the gate stops working

The first version shipped at a 95% true positive rate. Using the deployed demo turned up something neither dataset had shown: photographs taken from a few metres away were being refused.

Stanford ships a bounding box with every image, so the test split already held the explanation. Binning all 8580 test photos by how much of the frame the dog occupies puts a number on it.

Dog fillsImagesAccepted by the gateMedian distance
under 10%23377.3%39.7
10 to 20%62291.6%31.4
20 to 35%133196.1%27.6
35 to 50%156098.3%25.0
50 to 70%217699.0%23.4
over 70%265899.7%22.7

Acceptance falls 22 points across that range, and the median distance climbs from 22.7 to 39.7. Nearly a quarter of the most distant dogs are refused.

Nothing was malfunctioning. Both calibration sets are pet portraits, so a dog filling a twentieth of the frame really is far from the training distribution, and the gate reported that correctly. It answered the question it was asked. The question was wrong, so the threshold moved to 98%, which takes acceptance of those distant dogs from 77.3% to 87.6% and costs 3.3 points of cats.

The full sweep and what each step of true positive rate costs →

The experiments

Five configurations, same data, same seed, same preprocessing. Scored twice, because Stanford Dogs is cut from ImageNet and every pretrained backbone has already seen its test images. Oxford-IIIT Pet shares 21 breeds, photographed by other people.

EvaluationImagesAccuracy
Stanford test, 120 breeds858093.11%
Stanford test, 21 shared breeds163094.72%
Oxford-IIIT Pet, same 21 breeds417888.54%

The middle row is what makes the comparison honest. Those 21 breeds are easier than the average of 120, so without it the drop would mix two causes. Holding the breeds fixed and changing only where the photographs came from leaves 6.2 points, and every configuration shows it, between 4.7 and 8.1. It belongs to the benchmark rather than to a model.

Since all 120 breeds are ImageNet classes, each backbone also ships a classifier for this task among its 1000 outputs. Keeping those 120 rows gives a fourth thing to compare against, and it is the one that came out in front.

ExperimentBackboneTrainedBackbone alone
convnext_t_probeconvnext_tiny89.99%93.11%
baselineresnet5086.86%95.05%
effnet_b0efficientnet_b080.85%93.04%
convnext_tconvnext_tiny77.65%93.11%
effnet_b0_probeefficientnet_b076.31%93.04%

Freezing the backbone helped convnext by 12.3 points and hurt efficientnet by 4.5. The regime is not what decided the outcome. The quality of the pretrained features did, and the right-hand column is how far that goes.

The full grid, and why two of the runs prove little →

What runs in production

convnext_tiny with the 120 dog rows of its own ImageNet classifier. 93.11% on Stanford, 88.54% on Oxford.

resnet50 scores higher on Stanford at 95.05% and was not chosen. On Oxford, the only set whose photos are not ImageNet photos, the two are separated by eight images out of 4178, and resnet50 falls further between the two datasets. Picking it would mean picking on the contaminated number, which is the one this project spent its time measuring.

Swapping it in changed only the last matrix. The backbone was already frozen, so the 768 features are identical and the gate did not have to move. Rebuilding it from scratch produced a byte-for-byte identical file, threshold included, which was a prediction before it was a result.

Why the highest number was not the right one →

Serving it

The deployed service carries neither PyTorch nor timm, which takes about 400 MB out of the image. Parity tests hold the reimplemented preprocessing and distance to the training originals, checked for exact equality.

Scenariop50p95
Model only, onnxruntime, 2 threads104.7 ms126.6 ms
End to end, scaled JPEG decode (dev machine)141.8 ms160.3 ms
The deployed service, over loopback66 ms123 ms
The same, from a laptop in Italy (HTTPS + network)127 ms157 ms

Decoding a 4K JPEG cost more than the inference until Image.draft() let libjpeg decode at 1/8 scale inside the DCT domain. The 300 ms budget was the last number still untested on real hardware, and the service came in at 123 ms. The row below it is the same request through a domestic connection, where the tail is the connection's: its worst case reached 1156 ms while the container's, over the same 200 requests, stopped at 134 ms.

The container and the deployment →