← All domains, the protocol, and the register’s through-lines
AI in Education: What the Evidence Actually Shows
By late 2025, AI use among students was no longer a fringe behaviour but the norm. UK undergraduates reported using generative AI for assessments at 88%, up from 53% the year before [AIE-01]. US teens roughly doubled their use of ChatGPT for schoolwork between 2023 and 2024, from 13% to 26% [AIE-02], and by late 2025 nearly two-thirds of US teens said they had used an AI chatbot at all, with 30% using one daily [AIE-03]. Australia tracks the same trajectory: an Elevate Education survey of over 3,000 Australian high-schoolers in 2025 found three-quarters using AI at least a few times a week and almost a quarter using it daily, with ChatGPT the clear leader at 34% [AIE-13]. On the supply side, Australian teachers are near the front of the pack too — 66% reported using AI in the past year in the OECD's most recent teaching survey, the fourth-highest rate in the OECD and nearly double the OECD average of 36% [AIE-14]. Whatever families decide about AI, they are deciding it against a backdrop where most students and most teachers are already using it.
The most important finding in this set is not a prevalence number but a mechanism. A pre-registered field trial with nearly 1,000 Turkish high-school maths students randomised access to two versions of GPT-4: a plain ChatGPT-style interface and a pedagogically-safeguarded tutor built with teaching guardrails. Both improved in-session practice performance substantially — 48% for the plain interface, 127% for the safeguarded tutor. But the sting was in what happened after access was withdrawn: students who had used the unguarded interface then performed 17% worse on their own than students who had never had AI access at all. The safeguarded tutor's design largely prevented this harm [AIE-07]. This is the closest thing in the current evidence base to a controlled answer to the question every parent is really asking — not 'does AI help', but 'does it help or hollow out my child's actual ability'. The answer this trial gives is that it depends entirely on how the tool is built, not on whether AI is used at all.
That structured/unguarded distinction is not a one-off result. A separate randomised trial at Harvard, this time in undergraduate physics (N=194, crossover design), found students using a custom-built AI tutor — deliberately engineered with pedagogy best practices and given correct solutions to avoid hallucinating wrong answers — learned roughly twice as much as students in an in-class active-learning session on the same material, while reporting higher engagement and motivation [AIE-08]. The explicit caveat researchers attached matters as much as the headline result: this is evidence for a purpose-built tutoring tool, not for generic, unstructured chatbot use, which is a different thing entirely [AIE-08]. Read together, the Turkish and Harvard trials say something consistent — a well-designed, pedagogically-constrained AI tutor can outperform both unguarded AI and, in the Harvard case, live classroom instruction, while an AI tool with no guardrails can actively damage a student's ability to work independently once it is taken away.
Meanwhile, the promise that schools could simply detect AI use and hold the line has not held up. OpenAI discontinued its own AI-text classifier six months after launch, citing a low rate of accuracy — it correctly flagged only 26% of AI-written text while wrongly labelling 9% of genuinely human-written text as AI-generated [AIE-04]. The bias compounds for the students least equipped to fight a wrongful accusation: a peer-reviewed United States study, comparing TOEFL essays against United States eighth-grade essays, found GPT detectors misclassified over half of the essays written by non-native English speakers as AI-generated (a 61.22% average false-positive rate), against near-perfect accuracy on the native-English writers — a gap that dropped to 11.77% only once students were coached to simply reword their own writing to sound less 'formulaic' [AIE-05]. Turnitin, the dominant commercial detector, states in its own published guidance that it tunes for a false-positive rate below 1% by deliberately trading away detection sensitivity — a design choice that sits uneasily against the tool's separately-marketed 98% accuracy headline, a gap independent commentary has flagged as worth scrutiny whenever a student is actually accused [AIE-06]. Regulators have drawn the obvious conclusion. Australia's TEQSA required every higher-education provider to submit an institutional generative-AI action plan by July 2024, and every one of them did — a 100% response rate [AIE-11]. The thrust of that regulatory programme is toward redesigning how assessment works, not toward better detection, because TEQSA's own posture — evidenced by the shift from a compliance ask to an institution-wide action-plan requirement — treats detection as a dead end rather than a fix [AIE-11].
Australian schools have already lived through one detection-era, ban-first cycle and moved past it. NSW banned ChatGPT in state schools in January 2023; within nine months, national Education Ministers had approved the Australian Framework for Generative AI in Schools (5 October 2023), implemented from Term 1, 2024, replacing the ban with an enabling policy built on six principles and 25 guiding statements [AIE-10]. The policy arc, in other words, ran from prohibition to structure in under a year — a faster pivot than most families noticed happening.
None of this settles the harder question of what AI use actually does to a student's mind over time, and the evidence here should be handled carefully. The most viral claim in this space — an MIT Media Lab EEG study reporting 'cognitive debt' from LLM-assisted essay writing, showing reduced neural connectivity in AI-assisted writers — is an unreviewed preprint with a small sample: 54 participants across three sessions, only 18 of whom completed the critical fourth session testing skill transfer [AIE-09]. It has already drawn a critique of its own — also an unreviewed preprint — flagging its small sample size, its EEG methodology, and reproducibility concerns [AIE-09]. That does not mean its worry is wrong — it means the finding is preliminary, not the settled neuroscience it is often presented as in shared social posts, and it should never be cited without that caveat [AIE-09]. A more grounded signal on the same worry comes from Anthropic's own analysis of over half a million university-level conversations on Claude.ai, which found that nearly 47% were 'Direct' interactions — students handing over instructions and taking outputs with minimal engagement — rather than collaborative back-and-forth use [AIE-12]. It is vendor-generated data about the vendor's own product, so it should be read as directional rather than definitive, but it points the same direction as the Turkish RCT: AI used as a shortcut, rather than as a scaffolded thinking partner, is where the risk concentrates.
Put the pieces together and a consistent picture emerges, even though the evidence base is young and still contains a live, unresolved preprint at its centre. AI use among Australian students and teachers is now the default, not the exception [AIE-13] [AIE-14]. Detection was tried and has not worked reliably enough to be trusted as the safeguard, especially for non-native English speakers who bear a disproportionate false-positive burden in the United States research that measured it [AIE-04] [AIE-05] [AIE-06]. Regulators in Australia have already pivoted from banning and detecting to structural redesign [AIE-10] [AIE-11]. And the two best-designed controlled trials available say the same thing from two different angles: AI built with pedagogical guardrails and accountability for the student's actual learning produces real gains, while AI used without that structure can leave a student measurably worse off than if they had never touched it at all [AIE-07] [AIE-08]. For a family deciding how their child should use AI for schoolwork, the evidence does not support waiting for a ban, and it does not support unlimited, unsupervised access either. The live question was never whether a student uses AI — nearly all of them already do. It is whether that use is structured by someone who is actually accountable for whether the student learns, or left to run unguarded until the gap shows up later, once the tool is taken away and the skill was never really built.
AI in schools: what the Australian evidence says — These findings read together for parents — the national framework, what the research shows about learning, and what to make of a school flagging work as AI-generated.
GPT-generated-text detectors misclassified over half (average false positive rate 61.22%) of non-native-English (TOEFL) essays as AI-generated, versus near-perfect accuracy on native-English (US 8th-grade) essays; simple rewriting prompts cut this to 11.77%.
The detectors demonstrated near-perfect accuracy for US 8-th grade essays. However, they misclassified over half of the TOEFL essays as "AI-generated" (average false positive rate: 61.22%)... this intervention led to a substantial reduction in misclassification, with the average false positive rate decreasing by 49.45% (from 61.22% to 11.77%).
Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu, James Zou · GPT detectors are biased against non-native English writers · Patterns (Cell Press) / preprint on arXiv · 2023
Peer-reviewed / causalVerified 2026-08-23#AIE-05
A pre-registered field RCT with nearly 1,000 Turkish high-school maths students found GPT-4 access improved in-session practice performance substantially (48% for a plain ChatGPT-style interface, 127% for a pedagogically-safeguarded tutor), but once access was removed, students who had used the plain interface performed 17% worse on their own than students who never had access at all. Students in the safeguarded GPT Tutor arm were statistically indistinguishable from control once access was removed (point estimate -0.004) — the harm, in the paper's own words, "largely mitigated by the safeguards in GPT Tutor".
having GPT-4 access while solving problems significantly improves performance (48% improvement in grades for GPT Base and 127% for GPT Tutor). However, we additionally find that when access is subsequently taken away, students actually perform worse than those who never had access (17% reduction in grades for GPT Base)—i.e., unfettered access to GPT-4 can harm educational outcomes. These negative learning effects are largely mitigated by the safeguards in GPT Tutor.
Hamsa Bastani, Osbert Bastani, Alp Sungu, Haosen Ge, Özge Kabakcı, Rei Mariman · Generative AI without guardrails can harm learning: Evidence from high school mathematics · Proceedings of the National Academy of Sciences (PNAS) · 2025
Peer-reviewed / causalVerified 2026-08-23#AIE-07
A Harvard undergraduate physics RCT (N=194, crossover design) found students using a custom-built AI tutor at home learned roughly twice as much as students in an in-class active-learning session covering the same content, and reported greater engagement and motivation.
students learn significantly more in less time when using the AI tutor, compared with the in-class active learning, and they also feel more engaged and more motivated.
Gregory Kestin, Kelly Miller, Anna Klales, Timothy Milbourne, Guido Ponti · AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting · Scientific Reports (Nature Portfolio) · 2025
Peer-reviewed / causalVerified 2026-08-23#AIE-08
The Australian Framework for Generative AI in Schools was approved by national Education Ministers on 5 October 2023 and implemented from Term 1, 2024, replacing the prior state-level ChatGPT bans (e.g. NSW's ban, announced January 2023) with an enabling policy setting 6 principles and 25 guiding statements.
On 5 October 2023, Education Ministers approved the Australian Framework for Generative Artificial Intelligence (AI) in Schools – providing guidance on understanding, using and responding to generative AI in Australian school-based education. ...The Framework will be implemented from Term 1 2024.
National AI in Schools Taskforce (Commonwealth, states/territories, school sectors, national education agencies) · The Australian Framework for Generative Artificial Intelligence (AI) in Schools · Australian Government Department of Education · 2023
Peer-reviewed / causalVerified 2026-08-23#AIE-10
TEQSA required every Australian higher-education provider to submit an institutional generative-AI action plan by July 2024; all 203 providers responded (100% response rate). TEQSA's guidance explicitly states AI detection tools cannot guarantee academic integrity and that structural assessment redesign is the only sustainable response.
In June 2024, TEQSA asked all registered higher education providers for an institutional action plan addressing the risk gen AI poses to the integrity of their awards. The 100% response rate from providers to this request is testament to the partnership TEQSA has received from providers in addressing the impact of gen AI.
Tertiary Education Quality and Standards Agency (TEQSA) · Gen AI strategies for Australian higher education: Emerging practice (TEQSA toolkit) · TEQSA (Australian Government regulator) · 2024
Peer-reviewed / causalVerified 2026-08-23#AIE-11
In ACARA's 2015 evaluation of automated essay scoring on NAPLAN persuasive writing, four commercial scoring engines agreed with human markers on total score about as closely as the two human markers agreed with each other: human-human quadratic weighted kappa was 0.79, and the eight human-to-machine pairings ranged from 0.72 to 0.82.
Table 5. Quadratic weighted Kappa ... Marker 1 [vs] Marker 2 0.79 (0.63-0.87), AES 1 0.74 (0.61-0.83), AES 2 0.76 (0.62-0.84), AES 3 0.77 (0.63-0.85), AES 4 0.73 (0.56-0.83); Marker 2 [vs] AES 1 0.72 (0.54-082), AES 2 0.77 (0.64-0.85), AES 3 0.75 (0.59-0.84), AES 4 0.82 (0.67-0.89) ... Taken together, these analyses provide comprehensive evidence that the set of automated essay scoring engines provides satisfactory levels of consistency and reliability in marking NAPLAN persuasive writing at the rubric criteria and total score levels.
ACARA NASOP Research Team · An Evaluation of Automated Scoring of NAPLAN Persuasive Writing · Australian Curriculum, Assessment and Reporting Authority (ACARA) · 2015
Peer-reviewed / causalVerified 2026-09-04#AIE-15
The NSW Education Standards Authority, in an official notice effective Term 4 2026, instructs schools that they should not rely on AI detection tools as the main safeguard against malpractice, that such software if used at all is only one input for teacher professional review and never sole proof, and directs schools to check that their malpractice policy does not rely on AI detection tools. The same notice caps Stage 6 school-based assessment at no more than one take-home task worth a maximum weighting of 15%.
Schools should not rely on AI detection tools as the main safeguard against malpractice. If used at all, detection software should be treated as only one input for teacher professional review, not as sole proof of AI use or malpractice. ... For Preliminary and HSC courses, school-based assessment programs can have no more than one take-home assessment task, worth a maximum weighting of 15%. ... Schools must design formal assessment tasks so that marks reflect the student's own knowledge, skills and understanding. Marks must not reflect a student's use of generative AI. ... Check your malpractice policy does not rely on the use of AI tools to detect malpractice.
NSW Education Standards Authority · New limit on take-home assessment tasks from Term 4 2026 · NSW Education Standards Authority, official notice NESA 29/26 · 2026
Peer-reviewed / causalVerified 2026-09-16#AIE-18
UK undergraduate use of any AI tool jumped to 92% in the 2025 survey, up from 66% the prior year; use of generative AI specifically for assessments rose to 88% from 53%.
The proportion of students reporting using any AI tool has jumped from 66% last year to 92% this year... The proportion of students using generative AI tools such as ChatGPT for assessments has jumped from 53% last year to 88% this year.
Higher Education Policy Institute (HEPI) with Kortext, fieldwork by Savanta · Student Generative AI Survey 2025 · HEPI (HEPI Policy Note 61) · 2025
Official / governmentVerified 2026-08-23#AIE-01
US teen use of ChatGPT for schoolwork doubled year-on-year: from 13% in 2023 to 26% in 2024.
the share of U.S. teens ages 13 to 17 who say they have used ChatGPT for schoolwork has increased significantly over the last year, from 13% in fall 2023 to 26% in fall 2024
Pew Research Center · About a quarter of U.S. teens have used ChatGPT for schoolwork – double the share in 2023 · Pew Research Center · 2025
Official / governmentVerified 2026-08-23#AIE-02
In late 2025, 64% of US teens report ever using an AI chatbot, with 30% using one daily and 16% using one 'several times a day or almost constantly.'
Roughly two-thirds of teens (64%) say they ever use an AI chatbot... About three-in-ten (30%) use chatbots daily... 16% use them 'several times a day or almost constantly.'
Pew Research Center · Teens, Social Media and AI Chatbots 2025 · Pew Research Center · 2025
Official / governmentVerified 2026-08-23#AIE-03
OpenAI discontinued its own AI-generated-text classifier six months after launch, citing 'low rate of accuracy'; independent reporting states the tool correctly flagged only 26% of AI-written text as 'likely AI-written' while mislabelling human-written text as AI-written 9% of the time.
the AI classifier is no longer available due to its low rate of accuracy
Search Engine Land staff, reporting on OpenAI's own statement · OpenAI's AI Text Classifier no longer available due to 'low rate of accuracy' · Search Engine Land (reporting OpenAI's discontinuation notice) · 2023
Official / governmentVerified 2026-08-23#AIE-04
OECD's most recent international teaching survey found 66% of Australian lower-secondary teachers reported using AI in the past year — the 4th-highest rate among OECD countries and well above the OECD average of 36%. Most common uses were lesson-plan brainstorming and content summarisation; only 15% used AI for assessing student work (vs 30% OECD average) and 9% for reviewing student performance (vs 28% OECD average).
About two-thirds (66%) of lower secondary teachers reported using AI in the past year, putting Australia as the fourth highest country within the OECD, and far above the OECD average of 36%.
The Conversation (reporting OECD international teaching survey data) · Australian teachers are some of the highest users of AI in classrooms around the world – new survey · The Conversation / OECD · 2025
Official / governmentVerified 2026-08-23#AIE-14
A two-year cluster-randomised trial in 18 United States middle schools, giving randomly assigned students an AI tutor configured to coach rather than answer, raised mathematics achievement by 1.3 national percentile ranks per term -- about 0.06 to 0.08 standard deviations over a school year -- and the authors state those gains resemble what the same practice platform delivers with no AI at all. The constraint they identify is engagement rather than capability: 96% of students tried the tutor at least once, but the median student messaged it on only a third of the days they practised, and in only 17% of the exercise sessions in which they made a mistake.
Assignment raises math achievement by 1.3 national percentile ranks per term, or about 0.06 to 0.08 standard deviations over a school year; the implied effect of a full year of active participation reaches 0.14 standard deviations. These gains resemble those from Khan Academy practice without AI assistance. One explanation is that students used the tutor infrequently and, when they did, rarely engaged it in substantive mathematical dialogue: 96 percent of students tried Khanmigo at least once, but the median student messaged it on only a third of the days they practiced, and in only 17 percent of the exercise sessions in which they made a mistake. Messages that students did send were mostly bare answers or clicks on suggested prompts. The binding constraint appears to be engagement: realizing the promise of AI tutoring will require getting students to use it, not just giving them access.
Philip Oreopoulos and Nina Low · One Click Away: AI Tutoring with Khanmigo in a Two-Year School Experiment · National Bureau of Economic Research, Working Paper 35620 · 2026
Official / governmentVerified 2026-09-16#AIE-16
Turnitin's own published guidance states its AI writing detection is built to keep the false-positive rate below 1% (explicitly trading off some detection sensitivity to achieve this), which independent commentary contrasts with Turnitin's separately-advertised 98% accuracy headline figure — the gap between the two claims is itself a documented pattern worth noting when Turnitin flags a student's work.
Turnitin's AI writing detection focuses on accuracy—if we say there's AI writing, we're very sure there is. Our efforts have primarily been on ensuring a high accuracy rate accompanied by a less than 1% false positive rate, to ensure that students are not falsely accused of any misconduct.
Turnitin (official company blog) · Understanding false positives in Turnitin AI detection · Turnitin · 2024
Large-N industryVerified 2026-08-23#AIE-06
A widely-circulated MIT Media Lab EEG study on 'cognitive debt' from LLM-assisted essay writing is an unreviewed preprint with a small sample (54 participants across 3 sessions, only 18 completing a 4th session) — its findings should not be treated as settled peer-reviewed science.
arXiv preprint arXiv:2506.08872 (2025)
Nataliya Kosmyna et al. · Your Brain on ChatGPT: Accumulation of Cognitive Debt when Using an AI Assistant for Essay Writing Task · MIT Media Lab (preprint, arXiv) · 2025
Large-N industryVerified 2026-08-23#AIE-09
Anthropic's analysis of over half a million higher-education Claude.ai conversations found nearly 47% were 'Direct' interactions (students giving instructions and receiving outputs with minimal engagement) rather than collaborative/dialogic use — raising concern about automation of, rather than augmentation of, student thinking.
An inverted pyramid, after all, can topple over
Anthropic · Anthropic Education Report: How University Students Use Claude · Anthropic · 2025
Large-N industryVerified 2026-08-23#AIE-12
In a 2025 survey of over 3,000 Australian high-school students by Elevate Education, three-quarters reported using AI at least a few times per week, almost a quarter reported daily use, and ChatGPT was the most-used tool (34%).
three-quarters of students reported using AI at least a few times per week, and almost a quarter say they use it every single day... ChatGPT leads the way, used by 34% of students.
National Education Summit / Elevate Education survey · How Students Are Really Using AI in 2025 · National Education Summit (nationaleducationsummit.com.au) · 2025
Large-N industryVerified 2026-08-23#AIE-13
A 30-month panel of 26,811 Chinese students in grades 7 to 12 found generative AI adoption raised homework scores by 18% and cut homework completion time by 30%, while lowering monthly closed-book exam scores by 20% within six months and high-stakes entrance-exam scores by 18 and 24%, with the full penalty emerging only after about two years. The losses concentrate among roughly 80% of AI users whose behaviour is consistent with homework outsourcing -- exceptionally short completion times together with high homework marks -- and the authors report that AI users who keep their homework taking as long as non-users show only small losses.
AI adoption raises homework scores by 18% and reduces completion time by 30%, but lowers monthly exam scores by 20% within six months. High-stakes entrance-exam scores fall by 18 and 24%, with the full penalty emerging only after about two years. The losses are largest in social science subjects, followed by STEM and languages, and are especially large for junior students, high-achieving students, and boys. The learning losses are concentrated among roughly 80% of AI users whose behavior is consistent with homework outsourcing, as indicated by exceptionally short homework completion time coupled with high homework scores. AI users who maintain similar homework completion time as non-AI users experience small learning losses.
David Stromberg, Victor Lei and Yanhui Wu · DP21577 The Generative AI Learning Penalty: Evidence from Chinese Secondary Education · Centre for Economic Policy Research, Discussion Paper 21577 · 2026
Large-N industryVerified 2026-09-16#AIE-17
Cite any finding by its anchor — for example #AIE-05 — and the link resolves to the claim, quote and source above. The verification protocol lives on the register’s main page.