Research Watch · This week Business (preprint, Jul 2026)
From the Jul 21, 2026 daily brief
Three mainstream "don't touch the model, optimize the agent workflow" methods went through a two-phase continuous evaluation: all three beat baseline on the fixed test set; once new tasks arrived they diverged — some fell below baseline, others stopped improving. Only the method with regression control built into the loop (every optimization step re-verifies that old tasks haven't regressed) delivered both positive transfer and continued gains (76.4% lifetime average pass rate versus 58.7% for baseline) (arXiv, 2026-07). Caveat: the winning method is affiliated with the paper's authors; the numbers await independent replication. The portable rule: when you see "agent optimization method lifts benchmark scores," ask one question — a one-time lift, or compounding gains? Read together with the previous item: the ROI of deploying agents rides on process and regression infrastructure, not on model selection.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
The current mainstay of post-training (the stage after base training where a model is taug…
In the same AMD assessment, SemiAnalysis reports firsthand engineering observations: 2.5 e…
The Jacobian conjecture, posed in 1939, is a famous problem in algebraic geometry. It says…
Our July 21 Research Notes covered this empirical study (gains from optimizing an agent pi…