Dog Breed Bench.
How well can models identify a dog’s breed?
Results, version by version.
Version 1
Tested Sep 29, 2026
100 photos · 1 attempt per photoReasoning settings are shown with each run.
| ModelRanked by accuracy | AccuracyCorrect first-choice breed | Response timeMedian · seconds | Cost / 1,000 callsEstimated · USD |
|---|---|---|---|
| Gemini 3.8 FlashGoogle · medium | Accuracy | Median response 5.63 s | Cost / 1,000 calls $4.28 |
| Gemini 3.1 Pro PreviewGoogle · medium1 of 100 calls failed | Accuracy | Median response 4.08 s | Cost / 1,000 calls $7.32 |
| Claude Sonnet 5.5Anthropic · medium3 of 100 calls failed | Accuracy | Median response 3.27 s | Cost / 1,000 calls $10.77 |
| GPT-6 SolOpenAI · medium | Accuracy | Median response 8.20 s | Cost / 1,000 calls $9.46 |
| GPT-6.1 SolOpenAI · medium | Accuracy | Median response 6.48 s | Cost / 1,000 calls $8.63 |
| GPT-5.6 SolOpenAI · medium | Accuracy | Median response 4.71 s | Cost / 1,000 calls $16.58 |
| GPT-6 LunaOpenAI · medium | Accuracy | Median response 5.89 s | Cost / 1,000 calls $0.53 |
| GPT-6 LunaOpenAI · max15 of 100 calls failed | Accuracy | Median response 24.43 s | Cost / 1,000 calls $1.32 |
| GPT-5.6 LunaOpenAI · medium | Accuracy | Median response 4.43 s | Cost / 1,000 calls $0.95 |
| GPT-5.6 LunaOpenAI · max10 of 100 calls failed | Accuracy | Median response 12.11 s | Cost / 1,000 calls $2.52 |
| Claude Opus 5.5Anthropic · mediumRun incomplete — 4/100 calls; 3 failed | UnavailableProvider access unavailable (HTTP 429). | ||
Accuracy for the cost
Higher and further left means more correct first choices at a lower cost.
First-choice accuracy (%)
Estimated USD per 1,000 calls · logarithmic scale
- Gemini 3.8 FlashGoogle · medium93.0% · $4.28 / 1,000 calls
- Gemini 3.1 Pro PreviewGoogle · medium92.0% · $7.32 / 1,000 calls
- Claude Sonnet 5.5Anthropic · medium91.0% · $10.77 / 1,000 calls
- GPT-6 SolOpenAI · medium84.0% · $9.46 / 1,000 calls
- GPT-6.1 SolOpenAI · medium83.0% · $8.63 / 1,000 calls
- GPT-5.6 SolOpenAI · medium83.0% · $16.58 / 1,000 calls
- GPT-6 LunaOpenAI · medium75.0% · $0.53 / 1,000 calls
- GPT-6 LunaOpenAI · max69.0% · $1.32 / 1,000 calls
- GPT-5.6 LunaOpenAI · medium68.0% · $0.95 / 1,000 calls
- GPT-5.6 LunaOpenAI · max63.0% · $2.52 / 1,000 calls
Claude Sonnet 5.5’s saved scores include 10 responses that exceeded the prompt’s note-length limit. Strict response-format validation needs re-evaluation; these archived scores have not been changed.
Gemini 3.1 Pro Preview’s saved scores include two responses that exceeded the prompt’s note-length limit by one and three characters. These scores follow the archived scoring rules; strict response-format compliance is reported separately in the repository audit.
Accuracy by dog group
Correct first-choice breed · correct / attempts
| Model | Sporting11 photos | Hound12 photos | Working21 photos | Terrier13 photos | Toy16 photos | Non-Sporting12 photos | Herding13 photos |
|---|---|---|---|---|---|---|---|
| Gemini 3.8 FlashGoogle · medium | 90.9%10 / 11 correct | 100.0%12 / 12 correct | 100.0%21 / 21 correct | 84.6%11 / 13 correct | 81.3%13 / 16 correct | 100.0%12 / 12 correct | 92.3%12 / 13 correct |
| Gemini 3.1 Pro PreviewGoogle · medium | 90.9%10 / 11 correct | 91.7%11 / 12 correct | 95.2%20 / 21 correct | 92.3%12 / 13 correct | 81.3%13 / 16 correct | 100.0%12 / 12 correct | 92.3%12 / 13 correct |
| Claude Sonnet 5.5Anthropic · medium | 90.9%10 / 11 correct | 100.0%12 / 12 correct | 81.0%17 / 21 correct | 84.6%11 / 13 correct | 100.0%16 / 16 correct | 100.0%12 / 12 correct | 84.6%11 / 13 correct |
| GPT-6 SolOpenAI · medium | 81.8%9 / 11 correct | 91.7%11 / 12 correct | 85.7%18 / 21 correct | 61.5%8 / 13 correct | 87.5%14 / 16 correct | 100.0%12 / 12 correct | 76.9%10 / 13 correct |
| GPT-6.1 SolOpenAI · medium | 81.8%9 / 11 correct | 91.7%11 / 12 correct | 81.0%17 / 21 correct | 69.2%9 / 13 correct | 87.5%14 / 16 correct | 100.0%12 / 12 correct | 69.2%9 / 13 correct |
| GPT-5.6 SolOpenAI · medium | 90.9%10 / 11 correct | 83.3%10 / 12 correct | 76.2%16 / 21 correct | 76.9%10 / 13 correct | 87.5%14 / 16 correct | 100.0%12 / 12 correct | 69.2%9 / 13 correct |
| GPT-6 LunaOpenAI · medium | 81.8%9 / 11 correct | 66.7%8 / 12 correct | 81.0%17 / 21 correct | 61.5%8 / 13 correct | 68.8%11 / 16 correct | 100.0%12 / 12 correct | 61.5%8 / 13 correct |
| GPT-6 LunaOpenAI · max | 72.7%8 / 11 correct | 75.0%9 / 12 correct | 81.0%17 / 21 correct | 38.5%5 / 13 correct | 50.0%8 / 16 correct | 100.0%12 / 12 correct | 61.5%8 / 13 correct |
| GPT-5.6 LunaOpenAI · medium | 36.4%4 / 11 correct | 66.7%8 / 12 correct | 61.9%13 / 21 correct | 84.6%11 / 13 correct | 75.0%12 / 16 correct | 91.7%11 / 12 correct | 53.8%7 / 13 correct |
| GPT-5.6 LunaOpenAI · max | 54.5%6 / 11 correct | 58.3%7 / 12 correct | 61.9%13 / 21 correct | 69.2%9 / 13 correct | 56.3%9 / 16 correct | 91.7%11 / 12 correct | 46.2%6 / 13 correct |
| Claude Opus 5.5Anthropic · mediumRun incomplete — 4/100 calls; 3 failed | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable | Unavailable |
Same reference photos, grouped by AKC group. Small samples describe this run, not general group accuracy. Unavailable means the run is incomplete or more than half of calls failed in the full run or group.
2 other photos remain in overall scores: Dutch Shepherd (Miscellaneous Class), 1 photo; Australian Kelpie (Foundation Stock Service), 1 photo. Mixed breeds are not part of this run.
Photos & test conditions
- Photos
- 100 source-labeled photographs, one per breed. Every model receives the same frozen JPEG files; the shortest edge is at least 600 pixels and the longest at most 1,536. Natural lighting, framing and visibility vary; these conditions are not yet consistently labeled.
- Reference labels
- Public photo labels were checked against source descriptions and images; they have not had independent expert verification. This is a curated reference comparison, not a pedigree or DNA test. Mixed breeds and non-dog detection are outside this run.
- Settings
- The archived prompt, one analysis pass per photo, a 60-second request limit and no retries. Reasoning settings are provider-specific. Accuracy counts the first breed match; failed calls count as incorrect.
One photo per breed; small differences do not establish a clear winner. Costs are estimates from returned token usage, not invoices. Runs used different request pacing; response times may not be directly comparable.