Skip to main content
dbb.By Dog Breed Detector

Dog Breed Bench.

How well can models identify a dog’s breed?
Results, version by version.

Version 1

Tested Sep 29, 2026

100 photos · 1 attempt per photoReasoning settings are shown with each run.

Model comparison ranked by first-choice breed accuracy. Accuracy uses a fixed 0–100% scale. Response time and cost bars scale to the largest value in their column.
ModelRanked by accuracyAccuracyCorrect first-choice breedResponse timeMedian · secondsCost / 1,000 callsEstimated · USD
Gemini 3.8 FlashGoogle · mediumAccuracy
93.0%
Median response
5.63 s
Cost / 1,000 calls
$4.28
Gemini 3.1 Pro PreviewGoogle · medium1 of 100 calls failedAccuracy
92.0%
Median response
4.08 s
Cost / 1,000 calls
$7.32
Claude Sonnet 5.5Anthropic · medium3 of 100 calls failedAccuracy
91.0%
Median response
3.27 s
Cost / 1,000 calls
$10.77
GPT-6 SolOpenAI · mediumAccuracy
84.0%
Median response
8.20 s
Cost / 1,000 calls
$9.46
GPT-6.1 SolOpenAI · mediumAccuracy
83.0%
Median response
6.48 s
Cost / 1,000 calls
$8.63
GPT-5.6 SolOpenAI · mediumAccuracy
83.0%
Median response
4.71 s
Cost / 1,000 calls
$16.58
GPT-6 LunaOpenAI · mediumAccuracy
75.0%
Median response
5.89 s
Cost / 1,000 calls
$0.53
GPT-6 LunaOpenAI · max15 of 100 calls failedAccuracy
69.0%
Median response
24.43 s
Cost / 1,000 calls
$1.32
GPT-5.6 LunaOpenAI · mediumAccuracy
68.0%
Median response
4.43 s
Cost / 1,000 calls
$0.95
GPT-5.6 LunaOpenAI · max10 of 100 calls failedAccuracy
63.0%
Median response
12.11 s
Cost / 1,000 calls
$2.52
Claude Opus 5.5Anthropic · mediumRun incomplete — 4/100 calls; 3 failedUnavailableProvider access unavailable (HTTP 429).
Accuracy: 0–100%. Shorter time and cost bars are better; each uses its own scale.

Accuracy for the cost

Higher and further left means more correct first choices at a lower cost.

First-choice accuracy (%)

First-choice breed accuracy versus estimated costGemini 3.8 Flash · medium: 93.0% first-choice accuracy, $4.28 estimated per 1,000 calls; 95% confidence interval 88.0% to 97.0%Gemini 3.1 Pro Preview · medium: 92.0% first-choice accuracy, $7.32 estimated per 1,000 calls; 95% confidence interval 86.0% to 97.0%Claude Sonnet 5.5 · medium: 91.0% first-choice accuracy, $10.77 estimated per 1,000 calls; 95% confidence interval 85.0% to 96.0%GPT-6 Sol · medium: 84.0% first-choice accuracy, $9.46 estimated per 1,000 calls; 95% confidence interval 77.0% to 91.0%GPT-6.1 Sol · medium: 83.0% first-choice accuracy, $8.63 estimated per 1,000 calls; 95% confidence interval 75.0% to 90.0%GPT-5.6 Sol · medium: 83.0% first-choice accuracy, $16.58 estimated per 1,000 calls; 95% confidence interval 75.0% to 90.0%GPT-6 Luna · medium: 75.0% first-choice accuracy, $0.53 estimated per 1,000 calls; 95% confidence interval 66.0% to 83.0%GPT-6 Luna · max: 69.0% first-choice accuracy, $1.32 estimated per 1,000 calls; 95% confidence interval 60.0% to 78.0%GPT-5.6 Luna · medium: 68.0% first-choice accuracy, $0.95 estimated per 1,000 calls; 95% confidence interval 59.0% to 77.0%GPT-5.6 Luna · max: 63.0% first-choice accuracy, $2.52 estimated per 1,000 calls; 95% confidence interval 54.0% to 73.0%First-choice breed accuracy versus estimated costGemini 3.8 Flash · medium: 93.0% first-choice accuracy, $4.28 estimated per 1,000 calls; 95% confidence interval 88.0% to 97.0%Gemini 3.1 Pro Preview · medium: 92.0% first-choice accuracy, $7.32 estimated per 1,000 calls; 95% confidence interval 86.0% to 97.0%Claude Sonnet 5.5 · medium: 91.0% first-choice accuracy, $10.77 estimated per 1,000 calls; 95% confidence interval 85.0% to 96.0%GPT-6 Sol · medium: 84.0% first-choice accuracy, $9.46 estimated per 1,000 calls; 95% confidence interval 77.0% to 91.0%GPT-6.1 Sol · medium: 83.0% first-choice accuracy, $8.63 estimated per 1,000 calls; 95% confidence interval 75.0% to 90.0%GPT-5.6 Sol · medium: 83.0% first-choice accuracy, $16.58 estimated per 1,000 calls; 95% confidence interval 75.0% to 90.0%GPT-6 Luna · medium: 75.0% first-choice accuracy, $0.53 estimated per 1,000 calls; 95% confidence interval 66.0% to 83.0%GPT-6 Luna · max: 69.0% first-choice accuracy, $1.32 estimated per 1,000 calls; 95% confidence interval 60.0% to 78.0%GPT-5.6 Luna · medium: 68.0% first-choice accuracy, $0.95 estimated per 1,000 calls; 95% confidence interval 59.0% to 77.0%GPT-5.6 Luna · max: 63.0% first-choice accuracy, $2.52 estimated per 1,000 calls; 95% confidence interval 54.0% to 73.0%

Estimated USD per 1,000 calls · logarithmic scale

  1. Gemini 3.8 FlashGoogle · medium93.0% · $4.28 / 1,000 calls
  2. Gemini 3.1 Pro PreviewGoogle · medium92.0% · $7.32 / 1,000 calls
  3. Claude Sonnet 5.5Anthropic · medium91.0% · $10.77 / 1,000 calls
  4. GPT-6 SolOpenAI · medium84.0% · $9.46 / 1,000 calls
  5. GPT-6.1 SolOpenAI · medium83.0% · $8.63 / 1,000 calls
  6. GPT-5.6 SolOpenAI · medium83.0% · $16.58 / 1,000 calls
  7. GPT-6 LunaOpenAI · medium75.0% · $0.53 / 1,000 calls
  8. GPT-6 LunaOpenAI · max69.0% · $1.32 / 1,000 calls
  9. GPT-5.6 LunaOpenAI · medium68.0% · $0.95 / 1,000 calls
  10. GPT-5.6 LunaOpenAI · max63.0% · $2.52 / 1,000 calls
Each point is a measured model run; whiskers show its 95% confidence interval. Equal spacing on the cost axis represents equal price ratios. Unavailable models are omitted.

Claude Sonnet 5.5’s saved scores include 10 responses that exceeded the prompt’s note-length limit. Strict response-format validation needs re-evaluation; these archived scores have not been changed.

Gemini 3.1 Pro Preview’s saved scores include two responses that exceeded the prompt’s note-length limit by one and three characters. These scores follow the archived scoring rules; strict response-format compliance is reported separately in the repository audit.

Accuracy by dog group

Correct first-choice breed · correct / attempts

0%100%
ModelSporting11 photosHound12 photosWorking21 photosTerrier13 photosToy16 photosNon-Sporting12 photosHerding13 photos
Gemini 3.8 FlashGoogle · medium
90.9%10 / 11 correct
100.0%12 / 12 correct
100.0%21 / 21 correct
84.6%11 / 13 correct
81.3%13 / 16 correct
100.0%12 / 12 correct
92.3%12 / 13 correct
Gemini 3.1 Pro PreviewGoogle · medium
90.9%10 / 11 correct
91.7%11 / 12 correct
95.2%20 / 21 correct
92.3%12 / 13 correct
81.3%13 / 16 correct
100.0%12 / 12 correct
92.3%12 / 13 correct
Claude Sonnet 5.5Anthropic · medium
90.9%10 / 11 correct
100.0%12 / 12 correct
81.0%17 / 21 correct
84.6%11 / 13 correct
100.0%16 / 16 correct
100.0%12 / 12 correct
84.6%11 / 13 correct
GPT-6 SolOpenAI · medium
81.8%9 / 11 correct
91.7%11 / 12 correct
85.7%18 / 21 correct
61.5%8 / 13 correct
87.5%14 / 16 correct
100.0%12 / 12 correct
76.9%10 / 13 correct
GPT-6.1 SolOpenAI · medium
81.8%9 / 11 correct
91.7%11 / 12 correct
81.0%17 / 21 correct
69.2%9 / 13 correct
87.5%14 / 16 correct
100.0%12 / 12 correct
69.2%9 / 13 correct
GPT-5.6 SolOpenAI · medium
90.9%10 / 11 correct
83.3%10 / 12 correct
76.2%16 / 21 correct
76.9%10 / 13 correct
87.5%14 / 16 correct
100.0%12 / 12 correct
69.2%9 / 13 correct
GPT-6 LunaOpenAI · medium
81.8%9 / 11 correct
66.7%8 / 12 correct
81.0%17 / 21 correct
61.5%8 / 13 correct
68.8%11 / 16 correct
100.0%12 / 12 correct
61.5%8 / 13 correct
GPT-6 LunaOpenAI · max
72.7%8 / 11 correct
75.0%9 / 12 correct
81.0%17 / 21 correct
38.5%5 / 13 correct
50.0%8 / 16 correct
100.0%12 / 12 correct
61.5%8 / 13 correct
GPT-5.6 LunaOpenAI · medium
36.4%4 / 11 correct
66.7%8 / 12 correct
61.9%13 / 21 correct
84.6%11 / 13 correct
75.0%12 / 16 correct
91.7%11 / 12 correct
53.8%7 / 13 correct
GPT-5.6 LunaOpenAI · max
54.5%6 / 11 correct
58.3%7 / 12 correct
61.9%13 / 21 correct
69.2%9 / 13 correct
56.3%9 / 16 correct
91.7%11 / 12 correct
46.2%6 / 13 correct
Claude Opus 5.5Anthropic · mediumRun incomplete — 4/100 calls; 3 failed
Unavailable
Unavailable
Unavailable
Unavailable
Unavailable
Unavailable
Unavailable

Same reference photos, grouped by AKC group. Small samples describe this run, not general group accuracy. Unavailable means the run is incomplete or more than half of calls failed in the full run or group.

2 other photos remain in overall scores: Dutch Shepherd (Miscellaneous Class), 1 photo; Australian Kelpie (Foundation Stock Service), 1 photo. Mixed breeds are not part of this run.

Photos & test conditions

Photos
100 source-labeled photographs, one per breed. Every model receives the same frozen JPEG files; the shortest edge is at least 600 pixels and the longest at most 1,536. Natural lighting, framing and visibility vary; these conditions are not yet consistently labeled.
Reference labels
Public photo labels were checked against source descriptions and images; they have not had independent expert verification. This is a curated reference comparison, not a pedigree or DNA test. Mixed breeds and non-dog detection are outside this run.
Settings
The archived prompt, one analysis pass per photo, a 60-second request limit and no retries. Reasoning settings are provider-specific. Accuracy counts the first breed match; failed calls count as incorrect.

One photo per breed; small differences do not establish a clear winner. Costs are estimates from returned token usage, not invoices. Runs used different request pacing; response times may not be directly comparable.