Platform benchmark

claude-fable-5

anthropic · GPQA-Diamond · seed 42 · 2026-07-21 02:29 UTC

ResultsGPQA-Diamond · n=96 · seed 42
81.3%

81.3% correct · 84.4% answered on GPQA-Diamond

A low answered rate means the run starved on timeouts — not that the model is weak.

Score per $
0.16
Run cost
$5.15
Latency p50
11.5 s
Latency p95
39.0 s
Items
96/96 completed
Dropped
0