Platform benchmark

claude-fable-5-1

anthropic · GPQA-Diamond · seed 42 · 2026-09-01 21:10 UTC

ResultsGPQA-Diamond · n=96 · seed 42
83.3%

83.3% correct · 86.5% answered on GPQA-Diamond

A low answered rate means the run starved on timeouts — not that the model is weak.

Score per $
0.25
Run cost
$3.29
Latency p50
9.2 s
Latency p95
25.2 s
Items
96/96 completed
Dropped
0

Dataset: GPQA-Diamond · question and response text withheld under dataset access terms.