Research notes · This week (paper Jul 15)
From the Jul 19, 2026 daily brief
The conclusion first: evaluating AI with AI-generated data has a structural blind spot — there is no mechanical way to verify, item by item, that the data itself is right. Solo researcher Serkan Ballı offers a first-person case: while building a multilingual evaluation corpus, a single decode-length parameter shared by two scripts truncated one group of "wrong answers" to a few words, manufacturing out of thin air a "this AI judge's accuracy collapses by 32 points" effect. The effect held as the sample grew from 50 to 500 items, came with a three-layer mechanistic explanation, and was backed by control experiments — and all of it was false: fix the parameter and the effect goes to zero. Only human reading of the raw generated content would have caught it; no statistical check could (arXiv, Jul 15). A single first-person report, but a warning shot for any evaluation pipeline that uses AI as judge: "reproducible" is not "trustworthy." The portable line: if your eval's wrong answers are AI-generated, first ask whether every item is mechanically verified — and if not, sample them and read by hand.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
The current mainstay of post-training (the stage after base training where a model is taug…
In the same AMD assessment, SemiAnalysis reports firsthand engineering observations: 2.5 e…
The Jacobian conjecture, posed in 1939, is a famous problem in algebraic geometry. It says…
Our July 21 Research Notes covered this empirical study (gains from optimizing an agent pi…