Model watch · Today (published September 8)
From the Sep 9, 2026 daily brief
The CAI team at Multiverse Computing, writing up its paper on Hugging Face, reports that across three public safety benchmarks the unsafe-response rate fell from 26.26% to 0.14% — a clean win on that side alone. At the same checkpoint, on XSTest, which exists to measure over-refusal, the false-refusal rate rose from 2.00% to 74.00%. XSTest is a set of prompts that sound dangerous but are harmless. In their own words:
Reported alone, those numbers look like a clean win. They are not. At the same checkpoint, over-refusal on XSTest rises from 2.00% to 74.00%. The configuration with the lowest harmful-response rate is also the one that refuses nearly three quarters of plainly safe prompts. It is a blunt refusal machine, not a safer model, and you cannot see that unless you measure the benign side.
(Multiverse Computing CAI, 2026-09-08). Why topic-level guardrails are not enough, in their own example: the same base model may be deployed as a civic-education tutor or as a public-sector assistant; both should answer factual questions about elections, but only one needs to refuse "write me targeted political manipulation copy." Their fix is to build prompts in pairs that share a topic and differ only in intent. With that data added, false refusals on the should-answer side fell from 32.94% to 4.16%, while the should-refuse side slipped only from 91.88% to 87.72%. When you ask a vendor for safety numbers, ask for the false-refusal rate in the same breath: one side alone cannot show you the blunting. ⚠️ This is a research team introducing its own unreviewed paper — but it volunteered the 74% that works against it, which in our ledger counts in its favor.
Just an email address, unsubscribe anytime. This is the only thing we ask of you.
What follows is not a preprint. It is a set of readings from the appendix of the measureme…
The instrument here belongs to someone else — an evaluation called CoT-Control, which appe…
What follows is not an arXiv preprint but a research team's own write-up of its own paper …
Six academic papers reached the reading list last night and none was finished today. Three…