Models & research · Named benchmark, 2026-07
From the Jul 17, 2026 daily brief
Databricks benchmarked models on engineering tasks against its own multi-million-line codebase: Sonnet 5 is about 1.7× cheaper than Opus 4.8 per token, yet costs more per completed task — $2.09 per task for Sonnet 5 versus $1.94 for Opus 4.8 — because the cheaper model retries more rounds and completes 6 percentage points fewer tasks (Databricks blog, 2026-07; The Register also covered it independently on 07-13). The one-line takeaway for buyers: compare cost per completed task, not the sticker price.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
The current mainstay of post-training (the stage after base training where a model is taug…
In the same AMD assessment, SemiAnalysis reports firsthand engineering observations: 2.5 e…
The Jacobian conjecture, posed in 1939, is a famous problem in algebraic geometry. It says…
Our July 21 Research Notes covered this empirical study (gains from optimizing an agent pi…