The AI Homework Trap: When Higher Scores Hide Lower Learning
By SUSAN ESSEX
AI Made the Homework Better. Then the Exam Scores Fell
THE promise of artificial intelligence in education has always rested on a simple proposition: if technology can help students understand difficult concepts, complete assignments more efficiently and receive instant feedback, learning should improve.
A large new study complicates that assumption.
Researchers who followed 26,811 Chinese secondary-school students for 30 months found that generative AI made homework look substantially better while producing a very different result when students had to work without the technology.
After adopting generative AI, students’ homework scores rose by about 18 per cent. Their average homework completion time also fell by roughly 30 per cent, from about 64 minutes to 45 minutes.
Yet within six months, their monthly closed-book examination scores had fallen by about 20 per cent.
The findings, contained in a 2026 CEPR discussion paper by David Strömberg of Stockholm University and Victor Lei and Yanhui Wu of the University of Hong Kong, have intensified a central question in education: when AI makes schoolwork easier, does it necessarily make students more knowledgeable?
What the researchers actually studied
The study goes considerably deeper than the brief claim circulating online.
The researchers analysed 30 months of panel data covering students in Grades 7 to 12 in a county in central China. The dataset combined homework scores and completion times with monthly closed-book examinations, county-wide examinations and high-school and college entrance examinations across nine subjects.
The researchers also tracked when individual students began using generative AI. By June 2025, about 80 per cent of the students had reported using such tools, compared with almost none at the beginning of the study period in September 2022.
The most commonly used systems included Doubao and DeepSeek, alongside other general-purpose AI tools such as ChatGLM, Ernie Bot and Qwen. These were not specialised educational tutors designed around a particular curriculum. They were general-purpose systems capable of answering questions, producing text and helping users complete tasks.
That distinction is important because the study was examining AI as students naturally adopted and used it, rather than testing a carefully designed classroom tutoring programme.
The striking split between homework & exams
The central finding is the divergence between assisted work and unaided performance.
Following AI adoption, homework scores increased by 18 per cent relative to the pre-adoption average. At the same time, students spent about 19 fewer minutes on an assignment, reducing average completion time from 64 to 45 minutes.
On the surface, those numbers appear to describe a highly successful educational technology.
Students completed assignments faster and received higher marks.
The picture changed when researchers examined closed-book examinations.
Within six months of AI adoption, monthly exam scores fell by approximately 20 per cent relative to the pre-adoption average. Unlike homework, those examinations required students to demonstrate their knowledge without access to the AI system.
The contrast is significant because homework and examinations measure different things.
Homework can capture the quality of a finished product. A closed-book examination places greater weight on what a student can recall, understand and apply independently.
The researchers therefore distinguish between productivity and learning. AI can make students more productive at completing a task without necessarily increasing the amount they learn from doing it.
The penalty did not stop at monthly examinations
The researchers found that the learning effect became even more consequential when they examined high-stakes entrance examinations.
The losses accumulated more slowly because those examinations cover knowledge acquired over much longer periods.
After roughly two years, the estimated decline reached about 24 per cent of the baseline mean for high-school entrance examinations and 18 per cent for college entrance examinations.
The authors argue that this delayed effect matters because shorter studies can miss the cumulative consequences of changing how students learn. Monthly examinations generally test recently acquired material. Entrance examinations can draw on several years of accumulated knowledge.
Consequently, an AI-related learning problem may not immediately appear in a student’s grades.
A pupil can continue producing impressive homework while gradually losing opportunities to practise independent problem-solving.
The real issue may be how AI is used
The study does not support the conclusion that every form of AI use damages learning.
In fact, one of its most important findings points in the opposite direction.
The researchers found that learning losses were concentrated among roughly 80 per cent of AI users whose behaviour was consistent with homework outsourcing. These students tended to complete assignments unusually quickly while obtaining unusually high homework scores.
Students who used AI but maintained homework completion times similar to those of non-users experienced much smaller learning losses.
That distinction changes the question.
Instead of asking whether students should use AI at all, educators may need to ask what the technology is actually doing during the learning process.
If a student asks an AI system to explain a difficult mathematical concept, challenge an argument, identify a mistake or provide feedback after attempting a problem, the technology can function as a learning aid.
If the student simply submits the question and copies the answer, the machine has completed the central cognitive task.
The homework may look better.
The learner may not be.
Evidence from other research complicates the picture
Other studies help put the Chinese findings into perspective.
A 2025 study published in Proceedings of the National Academy of Sciences examined GPT-4-assisted mathematics learning among high-school students. It found that unrestricted access to generative AI could improve performance on AI-assisted tasks while harming subsequent unaided learning. The researchers also examined a more guarded AI tutor designed to provide hints rather than simply give away solutions.
More recent experimental work points towards a different possibility.
In a randomized experiment involving undergraduates learning an unfamiliar subject, Zara Contractor and Germán Reyes found that access to generative AI increased immediate knowledge-test performance by 0.27 standard deviations, with gains persisting one week later. Their findings also indicated that students who used AI to explain concepts rather than generate text experienced larger delayed benefits.
The contrast suggests that AI use is not one educational behaviour.
The same technology can act as a substitute for thinking or as an additional layer of support for thinking.
What the OECD says about the emerging evidence
The broader international evidence points in a similar direction.
The OECD’s Digital Education Outlook 2026 concludes that generative AI can support learning when educators use it with clear pedagogical objectives. But it warns that simply outsourcing tasks to general-purpose AI can improve students’ performance on those tasks without producing genuine learning gains.
The OECD consequently distinguishes between AI used as a tutor, partner or assistant and AI used to replace cognitive effort.
That distinction may become increasingly important as AI systems become more capable.
The easier it becomes for a student to obtain a polished answer, the less visible the learning process becomes to teachers and parents.
A measurement problem for schools
The findings also expose a problem beyond student behaviour.
Schools traditionally rely heavily on homework as an indicator of progress. Teachers see completed assignments and grades. Parents see improved marks. Students experience faster completion.
All three can therefore receive signals suggesting improvement.
But if the assignment has been completed primarily by a machine, the grade may no longer provide a reliable indication of the student’s independent understanding.
This creates a measurement gap.
The visible product improves while the underlying capability may deteriorate.
That gap becomes apparent only when the student encounters an assessment in which the technology is unavailable.
What the study does — & does not — prove
The scale of the research is significant, but its limitations matter.
The study covers students in one county in China and examines a particular period of rapid AI adoption. Its findings therefore should not automatically be treated as a precise estimate of what will happen to every student in every education system.
It is also a CEPR discussion paper, not a final peer-reviewed journal publication. The researchers used a staggered difference-in-differences design and report that students who later adopted AI had followed similar trends to comparison students before adoption. That strengthens the basis for the analysis, but it does not eliminate every limitation associated with observational administrative data.
The study’s strongest conclusion is therefore narrower than the viral claim.
It shows that, in this large sample, self-directed adoption of general-purpose generative AI was associated with substantially better homework performance but poorer unaided examination performance, particularly where AI use appeared to replace rather than supplement student effort.
The question schools now face
The educational challenge is unlikely to be solved simply by banning every chatbot.
Nor does the evidence justify allowing unrestricted AI use and assuming that higher homework marks represent better learning.
The more difficult task is designing rules around how students use the technology.
Teachers may need to distinguish between AI-generated answers and AI-supported reasoning. Schools may also need more supervised work, oral assessments, handwritten or closed-book exercises and assignments that require students to explain how they reached their conclusions.
At the same time, AI literacy is becoming part of education itself.
Students will need to know not only how to obtain information from AI, but also when relying on it undermines the skill they are supposed to acquire.
From higher marks to deeper learning
The most important lesson from the study is not that artificial intelligence has failed education.
It is that a better-looking piece of homework is not necessarily evidence of better learning.
The technology can save time. It can improve writing. It can explain concepts and provide immediate feedback. Other research shows that carefully structured AI use can improve learning outcomes.
But when the machine performs the intellectual work that the student was expected to practise, the apparent productivity gain can conceal a learning loss.
That is why the 18 per cent rise in homework scores and 20 per cent fall in examination scores matter together.
They describe two different outcomes.
One measures how well the task was completed.
The other asks whether the student can still do the work alone.
For education systems entering an AI era, the distinction may prove more important than either number by itself.
