It is relatively hard to determine whether artificial intelligence is helping some students learn better or stunting their development of learning skills.
But that might not be unexpected.
In the business world, higher-performing firms seemingly are better at adopting and extracting value from new technologies, including AI.
A small minority of high-performing companies capture the bulk of AI-driven returns (top five to 20 percent account for the large majority of gains), according to one analysis. Other studies tend to agree.
Students who are already good at standardized tests may also be better at using AI as a learning technology.
When educational researchers evaluate AI tools (such as automated writing assistants, generative AI tutors, or adaptive platforms), they frequently report positive short-term gains in student output, task completion speeds, or test scores.
However, critics and recent consensus reports argue that these metrics may conflate performance with learning. That’s akin to “teaching to the test,” where the objective is to improve student test scores by focusing on improving test performance, rather than other learning objectives.
Study | Design / students | What it found | Why it matters |
Contractor & Reyes, 2026 – Experimental Evidence on the Learning Impact of Generative AI | Randomized experiment with undergraduates; AI vs. no AI; unaided tests immediately and one week later | AI increased immediate knowledge-test scores by 0.27 SD, with gains persisting one week; essay effects depended on whether AI was used for augmentation versus automation | Strong evidence that AI can improve actual unaided performance, but how students use AI matters greatly. (IZA) |
Fischer, Rau & Rilke, 2025/26 – AI Tutoring Enhances Student Learning | RCT, 334 university students | AI tutoring raised test performance 0.23 SD; effects were largest among students with lower baseline knowledge and stronger self-regulation | Direct evidence for your hypothesis: baseline ability changes the size of the AI effect. (Scale) |
Harvard AI tutor RCT, 2025 | 194 undergraduates; AI tutor vs. active-learning classroom | AI group learned substantially more in less time; researchers measured pre-test knowledge and post-test learning | Shows why pre-testing matters: the relevant outcome is change from baseline, not simply the final score. (PubMed Central (PMC)) |
Noy et al./PNAS high-school mathematics study | High-school students using GPT assistance in math practice | Unrestricted/basic GPT assistance could reduce performance on an unassisted exam; students nevertheless believed they had learned more | Particularly important: perceived learning and measured learning diverged. (DOI) |
ChatGPT cognitive-crutch RCT, 2025 | 120 undergraduates; ChatGPT-assisted vs. traditional study | AI group scored 57.5% vs. 68.5% on a surprise retention test 45 days later | Shows why immediate test performance can be misleading: easier performance during learning can come at the cost of retention. (ScienceDirect) |
Oreopoulos et al., 2026 – AI tutoring/mastery math | >6,000 middle-school students; randomized AI vs. conventional computer-assisted learning | AI students made fewer attempts but were more accurate conditional on attempting; strongest delayed-learning evidence occurred when AI was embedded in a mastery structure | AI may alter the learning process, not simply raise scores. Structure matters. (National Bureau of Economic Research) |
Oreopoulos & Low, 2026 – Khanmigo | Two-year cluster RCT in 18 Tennessee middle schools | AI-tutor assignment raised math achievement about 0.06–0.08 SD per year, with larger effects for full-year active participation | Real-world AI effects can be modest because actual usage is much lower than theoretical availability. (National Bureau of Economic Research) |
Saloojee et al., 2026 – medical students | RCT, final-year medical students, ChatGPT available during clinical exams | No significant improvement in clinical performance; prior academic performance predicted scores | A useful reminder that strong students don't automatically benefit more from AI; task and implementation matter. (PubMed) |
Wu et al., 2026 – meta-analysis of 35 experimental/quasi-experimental studies | Meta-analysis | Finds substantial variation in ChatGPT effects across subjects, educational levels, instructional approaches and knowledge types | The average AI effect conceals considerable heterogeneity. (Nature) |
Deng et al., 2025 – meta-analysis of experimental studies | Systematic review/meta-analysis | Overall positive effects on academic performance and higher-order thinking, but also reduced mental effort | Illustrates the central ambiguity: better outcomes can coexist with less cognitive effort. (DOI) |
Nickow, Oreopoulos & Quan, 2020 – tutoring meta-analysis | Meta-analysis of tutoring experiments | Tutoring produced a large average learning effect (~0.37 SD), but effects varied substantially by program/context | Important pre-AI benchmark: individualized assistance has long produced heterogeneous effects, so AI shouldn't be expected to have one uniform effect. (National Bureau of Economic Research) |
Kraft, Schueler & Falken, 2024 – tutoring generalizability | Meta-analysis of 265 RCTs | Effects fell to roughly one-third to one-half of the original estimates when studies were restricted to settings resembling large-scale standardized-test interventions | A major warning about extrapolating impressive experimental effects to real-world populations. (ERIC) |
Up to a point, AI might help learning by minimizing friction. On the other hand, learning arguably often requires encountering friction and overcoming it. So if AI minimizes cognitive struggle, it might also negatively affect learning.
The other unavoidable problem is that some students are just better equipped to benefit from AI tool use. Students who are already high performers, highly self-regulated, or technologically fluent tend to extract more value out of AI tools.
Struggling students might rely too much on “getting the answers” without developing thinking, research or other skills. In other words, “output” might not reflect student learning so much as AI answers.
And it might be a reasonable assumption that better-equipped learners will also tend to be those that learn most when using AI as well.
Teachers might agree that an in-class essay exam using blue books is a more-reliable test of what a given student might know, though obviously also favoring better writers, compared to any out-of-class essay, which can be a better test of what a given AI engine knows and expresses.
The point is that better-performing students are also likely to be better-performing users of AI.
Study / Report Focus | Authors / Organization | Key Finding Regarding Performance vs. Learning | Source Link |
Children’s and Adolescents’ Learning with Educational Technology | American Psychological Association (APA) (2026) | Warns that metrics like time-on-task, completion rates, and immediate output are frequently mistaken for learning. Highlights that AI can improve immediate performance while simultaneously reducing long-term independent knowledge retention. | APA Report on Educational Tech |
Measuring Student Trust and Over-Reliance on AI Tutors | Educational Cybernetics and Studies (2025) | Investigates how moderate trust improves confidence, but over-trust triggers automation bias—where students accept AI-generated answers without critical evaluation, inflating performance metrics while depressing deep cognitive skill-building. | ECSE Article |
Assessing the Quality of AI-Generated Exams: A Large-Scale Field Study | ArXiv Educational Technology Research (2025) | Evaluates AI-generated assessment items across ~1,200 students using Item Response Theory (IRT), establishing that while AI can efficiently construct valid psychometric tests, student baseline engagement heavily dictates performance outcomes. | arXiv Study |
Improving Student Learning with Hybrid Human-AI Tutoring | Educational Research Working Paper Series (2023–2026) | Explores the nuance of baseline abilities, finding that while lower-achieving students show measurable proficiency gains with hybrid support, unmonitored independent AI usage risks widening the gap due to varying digital literacy and self-regulation skills. | arXiv Working Paper |
I’m not sure how we compensate for learning prowess in general, anymore than we seem to systematically produce high-performing firms.
No comments:
Post a Comment