Research notes · Evidence update (original paper March 2026)
From the Jul 18, 2026 daily brief
A Meta FAIR-affiliated team's March paper, Principia, argues that current math benchmarks mostly grade a final value or a multiple-choice pick, while real scientific work demands deriving formulas and structure. They built a 2,558-problem derivation benchmark. At the time we only recorded the qualitative claim that frontier models struggle; this week we read the PDF directly and filled in the numbers (arXiv, Mar 2026): OpenAI's o3 scores 62.90 on the derivation benchmark, while the same model scores 85.63 on competition math AIME-2024 — a gap of about 23 points. The most memorable result is a separate experiment, run on the math-and-engineering subset of the existing SuperGPQA benchmark, not on the 2,558 problems themselves: take the same questions, merely remove the answer options, require derivation instead — and strong models drop roughly 6 to 14 points. o3 falls from 69.10 to 62.90; Qwen3-235B from 69.33 to 55.58. Some fraction of those high scores came from answer-choice cues rather than derivation. For anyone using benchmark scores as a selling point or acceptance criterion, the actionable takeaway: never take a single headline number — demand the control condition, the same questions with the options stripped, or what you are buying may be a score propped up by answer-choice cues. Caveat: the grader is o3 itself — the self-built-eval, self-reported-results bias still applies.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
The current mainstay of post-training (the stage after base training where a model is taug…
In the same AMD assessment, SemiAnalysis reports firsthand engineering observations: 2.5 e…
The Jacobian conjecture, posed in 1939, is a famous problem in algebraic geometry. It says…
Our July 21 Research Notes covered this empirical study (gains from optimizing an agent pi…