SecondSourceJudgment rebuilt from primary sources
Research · Jul 21, 2026

First head-on test of whether agent-optimization gains compound: by default they are one-shot, and only methods with built-in regression control hold up.

Research Watch · This week Business (preprint, Jul 2026)

From the Jul 21, 2026 daily brief

Three mainstream "don't touch the model, optimize the agent workflow" methods went through a two-phase continuous evaluation: all three beat baseline on the fixed test set; once new tasks arrived they diverged — some fell below baseline, others stopped improving. Only the method with regression control built into the loop (every optimization step re-verifies that old tasks haven't regressed) delivered both positive transfer and continued gains (76.4% lifetime average pass rate versus 58.7% for baseline) (arXiv, 2026-07). Caveat: the winning method is affiliated with the paper's authors; the numbers await independent replication. The portable rule: when you see "agent optimization method lifts benchmark scores," ask one question — a one-time lift, or compounding gains? Read together with the previous item: the ROI of deploying agents rides on process and regression infrastructure, not on model selection.

Subscribe free — first issue lands tomorrow morning

Just an email address, unsubscribe anytime. This is the only thing we ask of you.

More in this section