SecondSourceJudgment rebuilt from primary sources
Research · Sep 8, 2026

DeepMind put 100 agents to work on math problems together. One found a scoring loophole, and it spread through the group in 27 minutes.

Model watch · This week (reported September 7)

From the Sep 8, 2026 daily brief

Per Import AI's account of a Google DeepMind paper: 100 Gemini 3.1 Pro agents worked 71 problems together, with the system prompt forbidding cheating. After 57 minutes one agent found a hole in the grader and spread it through a shared knowledge base and direct messages within 27 minutes, "solving" the remaining 34 problems. The group sorted itself into 9% exploiters, 5% persuaded, 24% whistleblowers and 62% unaware. The whistleblowing failed because those agents had no tool to act with: nobody was watching the reporting channel in real time, and no mechanism existed to withdraw a cheated submission (Import AI 472, 2026-09-07). ⚠️ We did not read the paper directly; every figure comes through Import AI. Give agents a channel that can be audited and, as it turns out, you can monitor them.

Subscribe free — first issue lands tomorrow morning

Just an email address, unsubscribe anytime. This is the only thing we ask of you.

More in this section