Skip to main content

How We Test Dog Breed Detection

A transparent account of the reviewed photo benchmark, release gates, measured results, and important limits behind Dog Breed Detector.

100

breeds in benchmark v2

500

accepted licensed photos

254

reviewed rejections retained

15

look-alike breed families

What the benchmark measures

The benchmark measures visual breed recognition from reviewed photographs. It records top-1 and top-3 accuracy, response-contract compliance, repeated-run stability, confidence calibration, latency, and estimated inference cost. Subject classification and mixed-dog behavior use separate gates because they answer different questions.

It does not measure genetic ancestry. A model can correctly identify a dog's strongest visible resemblance without knowing the dog's family tree.

Dataset construction and review

  • 100 canonical breeds with five accepted photographs per breed.
  • Three development photos, one regression photo, and one untouched holdout photo per breed.
  • Photographs sourced from Wikimedia Commons with recorded author, license, URL, and SHA-256.
  • Each accepted photo must show one meaningful dog subject with visible breed morphology.
  • Near duplicates, wrong breeds, multiple dogs, weak visibility, and non-photographic media are rejected.
  • Fifteen confusion packs group visually similar breeds for targeted error analysis.

Breed names and aliases are aligned against the AKC directory and FCI nomenclature. Review decisions remain recorded rather than silently discarded.

Measured production release

The production release evaluated on August 19, 2026 used the same prompt and response schema across repeated clear-breed cases. The selected vision stage produced:

  • 94.4% top-1 accuracy across 90 repeated clear-breed calls.
  • 100% top-3 accuracy and 100% response-contract compliance.
  • 96.7% repeated-run stability.
  • 2.645 second median latency and 7.692 second p95 latency.

A separate 108-call subject-classifier gate reached 100% dog recall and 100% non-dog specificity on its reviewed fixtures. A 60-call mixed-dog robustness gate reached 100% dog detection and response-contract compliance without ancestry claims.

These are benchmark measurements, not a promise that 94.4% of every real-world upload will be correct. Phone photos, puppies, unusual grooming, occlusion, lighting, and breeds outside the evaluation set can be harder.

Confidence is not ancestry

The detector returns confidence-ranked visual matches. A 70% result means the first candidate was the strongest visual match in that response; it does not mean that a dog is genetically 70% that breed.

For a mixed dog, compare multiple photos and treat shared traits across the shortlist as clues. Use a laboratory DNA test when ancestry evidence matters.

Known limitations

  • Closely related breeds can share structure, coat, and expression.
  • Puppies and senior dogs may differ from adult reference morphology.
  • Close crops hide body proportions; wide-angle phone lenses can distort the muzzle.
  • Multiple animals or heavy occlusion can make the subject ambiguous.
  • Visible resemblance cannot establish genetic ancestry or predict individual behavior.
  • Provider behavior, models, and pricing can change after a release is measured.

Release and correction policy

Production candidates are compared on the same reviewed fixtures. Prompt changes use development and regression cases; the holdout is reserved for final decisions. A release must pass accuracy, response, stability, and latency gates before replacement.

Report a reproducible problem through support. Methodology and results are updated when the model, prompt, benchmark, or release decision changes.