The Evidence Register

AI in education.

Prevalence, detection reliability, and what AI assistance does to learning. 14 verified findings — quoted verbatim, fully cited, independently checked (register v1.3.0, verified 2026-08-24).

AI in Education: What the Evidence Actually Shows

By late 2025, AI use among students was no longer a fringe behaviour but the norm. UK undergraduates reported using generative AI for assessments at 88%, up from 53% the year before [AIE-01]. US teens roughly doubled their use of ChatGPT for schoolwork between 2023 and 2024, from 13% to 26% [AIE-02], and by late 2025 nearly two-thirds of US teens said they had used an AI chatbot at all, with 30% using one daily [AIE-03]. Australia tracks the same trajectory: an Elevate Education survey of over 3,000 Australian high-schoolers in 2025 found three-quarters using AI at least a few times a week and almost a quarter using it daily, with ChatGPT the clear leader at 34% [AIE-13]. On the supply side, Australian teachers are near the front of the pack too — 66% reported using AI in the past year in the OECD's most recent teaching survey, the fourth-highest rate in the OECD and nearly double the OECD average of 36% [AIE-14]. Whatever families decide about AI, they are deciding it against a backdrop where most students and most teachers are already using it.

The most important finding in this set is not a prevalence number but a mechanism. A pre-registered field trial with nearly 1,000 Turkish high-school maths students randomised access to two versions of GPT-4: a plain ChatGPT-style interface and a pedagogically-safeguarded tutor built with teaching guardrails. Both improved in-session practice performance substantially — 48% for the plain interface, 127% for the safeguarded tutor. But the sting was in what happened after access was withdrawn: students who had used the unguarded interface then performed 17% worse on their own than students who had never had AI access at all. The safeguarded tutor's design largely prevented this harm [AIE-07]. This is the closest thing in the current evidence base to a controlled answer to the question every parent is really asking — not 'does AI help', but 'does it help or hollow out my child's actual ability'. The answer this trial gives is that it depends entirely on how the tool is built, not on whether AI is used at all.

That structured/unguarded distinction is not a one-off result. A separate randomised trial at Harvard, this time in undergraduate physics (N=194, crossover design), found students using a custom-built AI tutor — deliberately engineered with pedagogy best practices and given correct solutions to avoid hallucinating wrong answers — learned roughly twice as much as students in an in-class active-learning session on the same material, while reporting higher engagement and motivation [AIE-08]. The explicit caveat researchers attached matters as much as the headline result: this is evidence for a purpose-built tutoring tool, not for generic, unstructured chatbot use, which is a different thing entirely [AIE-08]. Read together, the Turkish and Harvard trials say something consistent — a well-designed, pedagogically-constrained AI tutor can outperform both unguarded AI and, in the Harvard case, live classroom instruction, while an AI tool with no guardrails can actively damage a student's ability to work independently once it is taken away.

Meanwhile, the promise that schools could simply detect AI use and hold the line has not held up. OpenAI discontinued its own AI-text classifier six months after launch, citing a low rate of accuracy — it correctly flagged only 26% of AI-written text while wrongly labelling 9% of genuinely human-written text as AI-generated [AIE-04]. The bias compounds for the students least equipped to fight a wrongful accusation: a peer-reviewed study found GPT detectors misclassified over half of essays written by non-native English speakers as AI-generated (a 61.22% average false-positive rate), against near-perfect accuracy on native-English writers — a gap that dropped to 11.77% only once students were coached to simply reword their own writing to sound less 'formulaic' [AIE-05]. Turnitin, the dominant commercial detector, states in its own published guidance that it tunes for a false-positive rate below 1% by deliberately trading away detection sensitivity — a design choice that sits uneasily against the tool's separately-marketed 98% accuracy headline, a gap independent commentary has flagged as worth scrutiny whenever a student is actually accused [AIE-06]. Regulators have drawn the obvious conclusion. Australia's TEQSA required every higher-education provider to submit an institutional generative-AI action plan by July 2024, and every one of them did — a 100% response rate [AIE-11]. The thrust of that regulatory programme is toward redesigning how assessment works, not toward better detection, because TEQSA's own posture — evidenced by the shift from a compliance ask to an institution-wide action-plan requirement — treats detection as a dead end rather than a fix [AIE-11].

Australian schools have already lived through one detection-era, ban-first cycle and moved past it. NSW banned ChatGPT in state schools in January 2023; within nine months, national Education Ministers had approved the Australian Framework for Generative AI in Schools (5 October 2023), implemented from Term 1, 2024, replacing the ban with an enabling policy built on six principles and 25 guiding statements [AIE-10]. The policy arc, in other words, ran from prohibition to structure in under a year — a faster pivot than most families noticed happening.

None of this settles the harder question of what AI use actually does to a student's mind over time, and the evidence here should be handled carefully. The most viral claim in this space — an MIT Media Lab EEG study reporting 'cognitive debt' from LLM-assisted essay writing, showing reduced neural connectivity in AI-assisted writers — is an unreviewed preprint with a small sample: 54 participants across three sessions, only 18 of whom completed the critical fourth session testing skill transfer [AIE-09]. It has already drawn a critique of its own — also an unreviewed preprint — flagging its small sample size, its EEG methodology, and reproducibility concerns [AIE-09]. That does not mean its worry is wrong — it means the finding is preliminary, not the settled neuroscience it is often presented as in shared social posts, and it should never be cited without that caveat [AIE-09]. A more grounded signal on the same worry comes from Anthropic's own analysis of over half a million university-level conversations on Claude.ai, which found that nearly 47% were 'Direct' interactions — students handing over instructions and taking outputs with minimal engagement — rather than collaborative back-and-forth use [AIE-12]. It is vendor-generated data about the vendor's own product, so it should be read as directional rather than definitive, but it points the same direction as the Turkish RCT: AI used as a shortcut, rather than as a scaffolded thinking partner, is where the risk concentrates.

Put the pieces together and a consistent picture emerges, even though the evidence base is young and still contains a live, unresolved preprint at its centre. AI use among Australian students and teachers is now the default, not the exception [AIE-13] [AIE-14]. Detection was tried and has not worked reliably enough to be trusted as the safeguard, especially for non-native English speakers who bear a disproportionate false-positive burden [AIE-04] [AIE-05] [AIE-06]. Regulators in Australia have already pivoted from banning and detecting to structural redesign [AIE-10] [AIE-11]. And the two best-designed controlled trials available say the same thing from two different angles: AI built with pedagogical guardrails and accountability for the student's actual learning produces real gains, while AI used without that structure can leave a student measurably worse off than if they had never touched it at all [AIE-07] [AIE-08]. For a family deciding how their child should use AI for schoolwork, the evidence does not support waiting for a ban, and it does not support unlimited, unsupervised access either. The live question was never whether a student uses AI — nearly all of them already do. It is whether that use is structured by someone who is actually accountable for whether the student learns, or left to run unguarded until the gap shows up later, once the tool is taken away and the skill was never really built.

GPT-generated-text detectors misclassified over half (average false positive rate 61.22%) of non-native-English (TOEFL) essays as AI-generated, versus near-perfect accuracy on native-English (US 8th-grade) essays; simple rewriting prompts cut this to 11.77%.

The detectors demonstrated near-perfect accuracy for US 8-th grade essays. However, they misclassified over half of the TOEFL essays as "AI-generated" (average false positive rate: 61.22%)... this intervention led to a substantial reduction in misclassification, with the average false positive rate decreasing by 49.45% (from 61.22% to 11.77%).

Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu, James Zou · GPT detectors are biased against non-native English writers · Patterns (Cell Press) / preprint on arXiv · 2023

Peer-reviewed / causalVerified 2026-08-23#AIE-05

A pre-registered field RCT with nearly 1,000 Turkish high-school maths students found GPT-4 access improved in-session practice performance substantially (48% for a plain ChatGPT-style interface, 127% for a pedagogically-safeguarded tutor), but once access was removed, students who had used the plain interface performed 17% worse on their own than students who never had access at all.

having GPT-4 access while solving problems significantly improves performance (48% improvement in grades for GPT Base and 127% for GPT Tutor). However, we additionally find that when access is subsequently taken away, students actually perform worse than those who never had access (17% reduction in grades for GPT Base)—i.e., unfettered access to GPT-4 can harm educational outcomes. These negative learning effects are largely mitigated by the safeguards in GPT Tutor.

Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı, Rei Mariman · Generative AI without guardrails can harm learning: Evidence from high school mathematics · Proceedings of the National Academy of Sciences (PNAS) · 2025

Peer-reviewed / causalVerified 2026-08-23#AIE-07

A Harvard undergraduate physics RCT (N=194, crossover design) found students using a custom-built AI tutor at home learned roughly twice as much as students in an in-class active-learning session covering the same content, and reported greater engagement and motivation.

students learn significantly more in less time when using the AI tutor, compared with the in-class active learning, and they also feel more engaged and more motivated.

Gregory Kestin, Kelly Miller, Anna Klales, Timothy Milbourne, Guido Ponti · AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting · Scientific Reports (Nature Portfolio) · 2025

Peer-reviewed / causalVerified 2026-08-23#AIE-08

The Australian Framework for Generative AI in Schools was approved by national Education Ministers on 5 October 2023 and implemented from Term 1, 2024, replacing the prior state-level ChatGPT bans (e.g. NSW's ban, announced January 2023) with an enabling policy setting 6 principles and 25 guiding statements.

On 5 October 2023, Education Ministers approved the Australian Framework for Generative Artificial Intelligence (AI) in Schools – providing guidance on understanding, using and responding to generative AI in Australian school-based education. ...The Framework will be implemented from Term 1 2024.

National AI in Schools Taskforce (Commonwealth, states/territories, school sectors, national education agencies) · The Australian Framework for Generative Artificial Intelligence (AI) in Schools · Australian Government Department of Education · 2023

Peer-reviewed / causalVerified 2026-08-23#AIE-10

TEQSA required every Australian higher-education provider to submit an institutional generative-AI action plan by July 2024; all 203 providers responded (100% response rate). TEQSA's guidance explicitly states AI detection tools cannot guarantee academic integrity and that structural assessment redesign is the only sustainable response.

In June 2024, TEQSA asked all registered higher education providers for an institutional action plan addressing the risk gen AI poses to the integrity of their awards. The 100% response rate from providers to this request is testament to the partnership TEQSA has received from providers in addressing the impact of gen AI.

Tertiary Education Quality and Standards Agency (TEQSA) · Gen AI strategies for Australian higher education: Emerging practice (TEQSA toolkit) · TEQSA (Australian Government regulator) · 2024

Peer-reviewed / causalVerified 2026-08-23#AIE-11

UK undergraduate use of any AI tool jumped to 92% in the 2025 survey, up from 66% the prior year; use of generative AI specifically for assessments rose to 88% from 53%.

The proportion of students reporting using any AI tool has jumped from 66% last year to 92% this year... The proportion of students using generative AI tools such as ChatGPT for assessments has jumped from 53% last year to 88% this year.

Higher Education Policy Institute (HEPI) with Kortext, fieldwork by Savanta · Student Generative AI Survey 2025 · HEPI (HEPI Policy Note 61) · 2025

Official / governmentVerified 2026-08-23#AIE-01

OpenAI discontinued its own AI-generated-text classifier six months after launch, citing 'low rate of accuracy'; independent reporting states the tool correctly flagged only 26% of AI-written text as 'likely AI-written' while mislabelling human-written text as AI-written 9% of the time.

the AI classifier is no longer available due to its low rate of accuracy

Search Engine Land staff, reporting on OpenAI's own statement · OpenAI's AI Text Classifier no longer available due to 'low rate of accuracy' · Search Engine Land (reporting OpenAI's discontinuation notice) · 2023

Official / governmentVerified 2026-08-23#AIE-04

OECD's most recent international teaching survey found 66% of Australian lower-secondary teachers reported using AI in the past year — the 4th-highest rate among OECD countries and well above the OECD average of 36%. Most common uses were lesson-plan brainstorming and content summarisation; only 15% used AI for assessing student work (vs 30% OECD average) and 9% for reviewing student performance (vs 28% OECD average).

About two-thirds (66%) of lower secondary teachers reported using AI in the past year, putting Australia as the fourth highest country within the OECD, and far above the OECD average of 36%.

The Conversation (reporting OECD international teaching survey data) · Australian teachers are some of the highest users of AI in classrooms around the world – new survey · The Conversation / OECD · 2025

Official / governmentVerified 2026-08-23#AIE-14

Turnitin's own published guidance states its AI writing detection is built to keep the false-positive rate below 1% (explicitly trading off some detection sensitivity to achieve this), which independent commentary contrasts with Turnitin's separately-advertised 98% accuracy headline figure — the gap between the two claims is itself a documented pattern worth noting when Turnitin flags a student's work.

Turnitin's AI writing detection focuses on accuracy—if we say there's AI writing, we're very sure there is. Our efforts have primarily been on ensuring a high accuracy rate accompanied by a less than 1% false positive rate, to ensure that students are not falsely accused of any misconduct.

Turnitin (official company blog) · Understanding false positives in Turnitin AI detection · Turnitin · 2024

Large-N industryVerified 2026-08-23#AIE-06

A widely-circulated MIT Media Lab EEG study on 'cognitive debt' from LLM-assisted essay writing is an unreviewed preprint with a small sample (54 participants across 3 sessions, only 18 completing a 4th session) — its findings should not be treated as settled peer-reviewed science.

arXiv preprint arXiv:2506.08872 (2025)

Nataliya Kosmyna et al. · Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task · MIT Media Lab (preprint, arXiv) · 2025

Large-N industryVerified 2026-08-23#AIE-09

Anthropic's analysis of over half a million higher-education Claude.ai conversations found nearly 47% were 'Direct' interactions (students giving instructions and receiving outputs with minimal engagement) rather than collaborative/dialogic use — raising concern about automation of, rather than augmentation of, student thinking.

An inverted pyramid, after all, can topple over

Anthropic · Anthropic Education Report: How University Students Use Claude · Anthropic · 2025

Large-N industryVerified 2026-08-23#AIE-12

In a 2025 survey of over 3,000 Australian high-school students by Elevate Education, three-quarters reported using AI at least a few times per week, almost a quarter reported daily use, and ChatGPT was the most-used tool (34%).

three-quarters of students reported using AI at least a few times per week, and almost a quarter say they use it every single day... ChatGPT leads the way, used by 34% of students.

National Education Summit / Elevate Education survey · How Students Are Really Using AI in 2025 · National Education Summit (nationaleducationsummit.com.au) · 2025

Large-N industryVerified 2026-08-23#AIE-13

Cite any finding by its anchor — for example #AIE-05 — and the link resolves to the claim, quote and source above. The verification protocol lives on the register’s main page.