Friday, September 11, 2026

Both Companies and Students are Unequal in Ability to Use AI Productively

It is relatively hard to determine whether artificial intelligence is helping some students learn better or stunting their development of learning skills.


But that might not be unexpected. 


In the business world, higher-performing firms seemingly are better at adopting and extracting value from new technologies, including AI. 


A small minority of high-performing companies capture the bulk of AI-driven returns (top five to 20 percent account for the large majority of gains), according to one analysis. Other studies tend to agree.  


Students who are already good at standardized tests may also be better at using AI as a learning technology. 


When educational researchers evaluate AI tools (such as automated writing assistants, generative AI tutors, or adaptive platforms), they frequently report positive short-term gains in student output, task completion speeds, or test scores. 


However, critics and recent consensus reports argue that these metrics may conflate performance with learning. That’s akin to “teaching to the test,” where the objective is to improve student test scores by focusing on improving test performance, rather than other learning objectives. 


Study

Design / students

What it found

Why it matters

Contractor & Reyes, 2026 – Experimental Evidence on the Learning Impact of Generative AI

Randomized experiment with undergraduates; AI vs. no AI; unaided tests immediately and one week later

AI increased immediate knowledge-test scores by 0.27 SD, with gains persisting one week; essay effects depended on whether AI was used for augmentation versus automation

Strong evidence that AI can improve actual unaided performance, but how students use AI matters greatly. (IZA)

Fischer, Rau & Rilke, 2025/26 – AI Tutoring Enhances Student Learning

RCT, 334 university students

AI tutoring raised test performance 0.23 SD; effects were largest among students with lower baseline knowledge and stronger self-regulation

Direct evidence for your hypothesis: baseline ability changes the size of the AI effect. (Scale)

Harvard AI tutor RCT, 2025

194 undergraduates; AI tutor vs. active-learning classroom

AI group learned substantially more in less time; researchers measured pre-test knowledge and post-test learning

Shows why pre-testing matters: the relevant outcome is change from baseline, not simply the final score. (PubMed Central (PMC))

Noy et al./PNAS high-school mathematics study

High-school students using GPT assistance in math practice

Unrestricted/basic GPT assistance could reduce performance on an unassisted exam; students nevertheless believed they had learned more

Particularly important: perceived learning and measured learning diverged. (DOI)

ChatGPT cognitive-crutch RCT, 2025

120 undergraduates; ChatGPT-assisted vs. traditional study

AI group scored 57.5% vs. 68.5% on a surprise retention test 45 days later

Shows why immediate test performance can be misleading: easier performance during learning can come at the cost of retention. (ScienceDirect)

Oreopoulos et al., 2026 – AI tutoring/mastery math

>6,000 middle-school students; randomized AI vs. conventional computer-assisted learning

AI students made fewer attempts but were more accurate conditional on attempting; strongest delayed-learning evidence occurred when AI was embedded in a mastery structure

AI may alter the learning process, not simply raise scores. Structure matters. (National Bureau of Economic Research)

Oreopoulos & Low, 2026 – Khanmigo

Two-year cluster RCT in 18 Tennessee middle schools

AI-tutor assignment raised math achievement about 0.06–0.08 SD per year, with larger effects for full-year active participation

Real-world AI effects can be modest because actual usage is much lower than theoretical availability. (National Bureau of Economic Research)

Saloojee et al., 2026 – medical students

RCT, final-year medical students, ChatGPT available during clinical exams

No significant improvement in clinical performance; prior academic performance predicted scores

A useful reminder that strong students don't automatically benefit more from AI; task and implementation matter. (PubMed)

Wu et al., 2026 – meta-analysis of 35 experimental/quasi-experimental studies

Meta-analysis

Finds substantial variation in ChatGPT effects across subjects, educational levels, instructional approaches and knowledge types

The average AI effect conceals considerable heterogeneity. (Nature)

Deng et al., 2025 – meta-analysis of experimental studies

Systematic review/meta-analysis

Overall positive effects on academic performance and higher-order thinking, but also reduced mental effort

Illustrates the central ambiguity: better outcomes can coexist with less cognitive effort. (DOI)

Nickow, Oreopoulos & Quan, 2020 – tutoring meta-analysis

Meta-analysis of tutoring experiments

Tutoring produced a large average learning effect (~0.37 SD), but effects varied substantially by program/context

Important pre-AI benchmark: individualized assistance has long produced heterogeneous effects, so AI shouldn't be expected to have one uniform effect. (National Bureau of Economic Research)

Kraft, Schueler & Falken, 2024 – tutoring generalizability

Meta-analysis of 265 RCTs

Effects fell to roughly one-third to one-half of the original estimates when studies were restricted to settings resembling large-scale standardized-test interventions

A major warning about extrapolating impressive experimental effects to real-world populations. (ERIC)


Up to a point, AI might help learning by minimizing friction. On the other hand, learning arguably often requires encountering friction and overcoming it. So if AI minimizes cognitive struggle, it might also negatively affect learning. 


The other unavoidable problem is that some students are just better equipped to benefit from AI tool use. Students who are already high performers, highly self-regulated, or technologically fluent tend to extract more value out of AI tools. 


Struggling students might rely too much on “getting the answers” without developing thinking, research or other skills. In other words, “output” might not reflect student learning so much as AI answers. 


And it might be a reasonable assumption that better-equipped learners will also tend to be those that learn most when using AI as well. 


Teachers might agree that an in-class essay exam using blue books is a more-reliable test of what a given student might know, though obviously also favoring better writers, compared to any out-of-class essay, which can be a better test of what a given AI engine knows and expresses. 


The point is that better-performing students are also likely to be better-performing users of AI. 


Study / Report Focus

Authors / Organization

Key Finding Regarding Performance vs. Learning

Source Link

Children’s and Adolescents’ Learning with Educational Technology

American Psychological Association (APA) (2026)

Warns that metrics like time-on-task, completion rates, and immediate output are frequently mistaken for learning. Highlights that AI can improve immediate performance while simultaneously reducing long-term independent knowledge retention.

APA Report on Educational Tech

Measuring Student Trust and Over-Reliance on AI Tutors

Educational Cybernetics and Studies (2025)

Investigates how moderate trust improves confidence, but over-trust triggers automation bias—where students accept AI-generated answers without critical evaluation, inflating performance metrics while depressing deep cognitive skill-building.

ECSE Article

Assessing the Quality of AI-Generated Exams: A Large-Scale Field Study

ArXiv Educational Technology Research (2025)

Evaluates AI-generated assessment items across ~1,200 students using Item Response Theory (IRT), establishing that while AI can efficiently construct valid psychometric tests, student baseline engagement heavily dictates performance outcomes.

arXiv Study

Improving Student Learning with Hybrid Human-AI Tutoring

Educational Research Working Paper Series (2023–2026)

Explores the nuance of baseline abilities, finding that while lower-achieving students show measurable proficiency gains with hybrid support, unmonitored independent AI usage risks widening the gap due to varying digital literacy and self-regulation skills.

arXiv Working Paper


I’m not sure how we compensate for learning prowess in general, anymore than we seem to systematically produce high-performing firms.


No comments:

Both Companies and Students are Unequal in Ability to Use AI Productively

It is relatively hard to determine whether artificial intelligence is helping some students learn better or stunting their development of le...