SecondSourceJudgment rebuilt from primary sources
Research · Jul 19, 2026

How one bug manufactured a finding that was reproducible, statistically robust, mechanistically explained — and false.

Research notes · This week (paper Jul 15)

From the Jul 19, 2026 daily brief

The conclusion first: evaluating AI with AI-generated data has a structural blind spot — there is no mechanical way to verify, item by item, that the data itself is right. Solo researcher Serkan Ballı offers a first-person case: while building a multilingual evaluation corpus, a single decode-length parameter shared by two scripts truncated one group of "wrong answers" to a few words, manufacturing out of thin air a "this AI judge's accuracy collapses by 32 points" effect. The effect held as the sample grew from 50 to 500 items, came with a three-layer mechanistic explanation, and was backed by control experiments — and all of it was false: fix the parameter and the effect goes to zero. Only human reading of the raw generated content would have caught it; no statistical check could (arXiv, Jul 15). A single first-person report, but a warning shot for any evaluation pipeline that uses AI as judge: "reproducible" is not "trustworthy." The portable line: if your eval's wrong answers are AI-generated, first ask whether every item is mechanically verified — and if not, sample them and read by hand.

Subscribe free — first issue lands tomorrow morning

Just an email address, unsubscribe anytime. This is the only thing we ask of you.

More in this section