A new AI Study tracking 26,000 Chinese middle and high schoolers over 30 months shows a brutal split. Homework scores go up 18 %, but real exam scores drop between 20 % and 24 %.
Key Takeaways
- AI Study on 26,000+ Chinese students, 30 months of panel data, published on July 4, 2026.
- Homework with AI: +18 % scores, completion time down from 64 to 45 minutes.
- Monthly exams: -20 %. High-stakes entrance exams (Zhongkao, Gaokao): -18 % to -24 % after two years.
Have an AI Sum Up This Article
ChatGPTThe brutal finding: homework inflated, real exams crashing
The AI Study covers more than 26,000 secondary students (Grades 7 to 12) from a county in central China with over one million residents. Researchers tracked monthly exams, homework scores, completion times and entrance exams over 30 months. The panel is wide enough that the results are not an anecdote.
The tools used are the mainstream Chinese LLMs. That means Doubao, DeepSeek (V2.5 and R1), ChatGLM, Ernie Bot and Qwen. These models are free, embedded in mobile apps, and used heavily at home to blast through homework in a few minutes.
On homework, the effect is impressive. Scores go up by 18 % on average, and completion time drops from 64 minutes to 45. A student who uses AI turns in shorter, faster, better-graded work. From the outside, everything looks fine.
On monthly in-class exams, the picture flips. Scores drop by 20 %. On Zhongkao and Gaokao entrance exams, the drop reaches 18 % to 24 %. The damage does not appear immediately. It takes around six months of regular use to see the collapse on classroom exams, and close to two years for it to hit the entrance exams.
The AI Study lines up with other recent signals about how heavily Claude now covers real cognitive work. Only here, the tested population is teenagers, and the measured effect is a real drop on high-stakes assessments.
The mechanism: less time, more dependency, sharp splits by profile
Completion time is the first warning signal. Students no longer look for answers, they request them. The mental journey of solving an exercise disappears. What gets built through training (repetition, mistakes, redo) no longer gets built. The homework turns into a transaction, not an exercise.
The clearest signal comes from long-term users. 81 % of students who have used AI for more than five months finish their homework in under 50 minutes while posting very high homework scores and very low exam scores. The profile is sharp: inflated grades at home, collapse the moment the crutch is gone.
The effect varies heavily by subject. Social sciences take -27 %, STEM -22 %, English -17 %, Chinese -9 %. Subjects built on structured argumentation and reasoning are hit hardest. Chinese, which leans more on direct memorization and native writing, holds up better.
By student profile, the gap is striking. Younger students lose 24 %, older ones 17 %. Boys drop 21.6 %, girls 18.4 %. Top students lose 24 %, weaker ones 16 %. The cruel paradox is that the students with the most to lose in cognitive autonomy are the ones losing the most.
This dynamic echoes what we already covered on the fading of entry-level tasks for juniors. For students and for beginners at work, the ladder that used to build core skills is being short-circuited by the tool.
Also on Horizon:
- Leanstral 1.5 Test: Mistral’s New Math Proof Engine
- Claude Sonnet 5 Test: Verdict After a Pro Week
- Fysiverse: China Builds Real Physics Into Its AI
Short term and mid term: what to do now, what shifts in six months
In the short term, the most direct answer for parents and teachers is not to ban, it is to change what we watch. Tracking completion time becomes a more useful indicator than homework grades. A student who finishes in 20 minutes what should take 45 is a warning signal, even when the score looks great.
The authors recommend three concrete moves. Tell students explicitly about the long-term cost (the exam collapse is not intuitive, it has to be shown with the numbers). Increase the weight of in-person assessments in the final grade. And track completion time rather than homework grades as the real progression signal.
In the mid term, over the next 3 to 6 months, this AI Study will carry weight in the education policy debate. Education ministries that were hesitating on classroom AI policies needed a solid data point. They now have one. 26,000 students, 30 months, -24 % on Gaokao: this is the first heavy quantitative benchmark policymakers will point to.
LLM vendors will have to respond. Doubao, DeepSeek and the others are named. Pressure to ship tutor modes (models that explain instead of handing over the answer) will grow. The real product question becomes: how can an LLM help a teenager learn without selling them the answer in three seconds.
The dynamic echoes what is happening in the labor market, where data is sending contradictory signals about the real impact of AI. Inflated homework on one side, collapsed exams on the other. Measured productivity on one side, entry-level task disappearance on the other. The pattern is the same: the tool amplifies what we measure poorly, and breaks what we measure well.
Follow the story on Horizon.


