Research notes · Retrospective (November 2024 interview)
From the Jul 20, 2026 daily brief
Gwern, the anonymous independent researcher known for his early systematic case for the scaling hypothesis in 2020, diagnosed it this way: web corpora contain only descriptions of agent behavior, never the step-by-step decision traces — in his words, "All the agency there is is an accidental byproduct of somebody training on data" (Dwarkesh Podcast × Gwern, Nov 2024). That still has explanatory power for the 2026 reality of agents that demo brilliantly and wobble in production. It also echoes today's lead 4: the product form is evolving; the root problem of reliability is a separate thread. Time boundary: industry investment in reinforcement learning for agents has risen sharply across 2025–26; this is a diagnosis from a 2024 vantage point.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
The current mainstay of post-training (the stage after base training where a model is taug…
In the same AMD assessment, SemiAnalysis reports firsthand engineering observations: 2.5 e…
The Jacobian conjecture, posed in 1939, is a famous problem in algebraic geometry. It says…
Our July 21 Research Notes covered this empirical study (gains from optimizing an agent pi…