SecondSourceJudgment rebuilt from primary sources
Research · Sep 10, 2026

Same capability tier, 9B to 397B parameters: whether a model can control its own reasoning trace is decided by whoever did the post-training, not by how large the model is.

Model watch · Today (posted September 10)

From the Sep 10, 2026 daily brief

The instrument here belongs to someone else — an evaluation called CoT-Control, which appends instructions about the reasoning trace to a question ("think only in lowercase," "do not use this word") and grades whether the trace complied. The results cluster by lab across three downloadable-weight model families from three different labs: the GLM-5 series from Z.AI (formerly Zhipu AI, a Chinese frontier lab) lands between 6% and 13%, one DeepSeek model at 11%, and every Qwen 3.5 checkpoint — Qwen is Alibaba's open-weight family — from 9B to 397B sits at zero (Alignment Forum, 2026-09-10). ⚠️ Only open-weight models can be measured this way, because only they expose the real reasoning trace — closed models hand you a summary — so this reading cannot reach the model in today's main line. It has no peer review and no named author. Why it is useful: it reclassifies controllability from an inevitable by-product of growing capability into an engineering choice, and engineering choices are the sort of thing procurement contracts and regulations can specify.

Subscribe free — first issue lands tomorrow morning

Just an email address, unsubscribe anytime. This is the only thing we ask of you.

More in this section