Research notes · This week (paper Jul 15)
From the Jul 19, 2026 daily brief
For problems with no answer key, sample many solutions from the model, take the majority-consensus answer — then train the model verbatim on the solutions that reach that answer. The team (Gkountouras, Jukić, Titov) self-reports: on the authors' chosen math-reasoning benchmarks (pass@1), gains of up to 12 points; with about one-seventh the compute it beats label-free reinforcement learning (training the model to self-adjust by trial and error) by 6 points; and after training the model solves problems it had failed in 32 straight prior attempts. That last claim is aimed squarely at the old objection that self-training only makes the model more confident about what it already knows (arXiv, Jul 15). A single self-report, no third-party reproduction; the "AI improving itself without human labels" research line stays on our watch list.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
The current mainstay of post-training (the stage after base training where a model is taug…
In the same AMD assessment, SemiAnalysis reports firsthand engineering observations: 2.5 e…
The Jacobian conjecture, posed in 1939, is a famous problem in algebraic geometry. It says…
Our July 21 Research Notes covered this empirical study (gains from optimizing an agent pi…