By late 2025, AI use among students was no longer a fringe behaviour but the norm. UK undergraduates reported using generative AI for assessments at 88%, up from 53% the year before [AIE-01]. US teens roughly doubled their use of ChatGPT for schoolwork between 2023 and 2024, from 13% to 26% [AIE-02], and by late 2025 nearly two-thirds of US teens said they had used an AI chatbot at all, with 30% using one daily [AIE-03]. Australia tracks the same trajectory: an Elevate Education survey of over 3,000 Australian high-schoolers in 2025 found three-quarters using AI at least a few times a week and almost a quarter using it daily, with ChatGPT the clear leader at 34% [AIE-13]. On the supply side, Australian teachers are near the front of the pack too — 66% reported using AI in the past year in the OECD's most recent teaching survey, the fourth-highest rate in the OECD and nearly double the OECD average of 36% [AIE-14]. Whatever families decide about AI, they are deciding it against a backdrop where most students and most teachers are already using it.
The most important finding in this set is not a prevalence number but a mechanism. A pre-registered field trial with nearly 1,000 Turkish high-school maths students randomised access to two versions of GPT-4: a plain ChatGPT-style interface and a pedagogically-safeguarded tutor built with teaching guardrails. Both improved in-session practice performance substantially — 48% for the plain interface, 127% for the safeguarded tutor. But the sting was in what happened after access was withdrawn: students who had used the unguarded interface then performed 17% worse on their own than students who had never had AI access at all. The safeguarded tutor's design largely prevented this harm [AIE-07]. This is the closest thing in the current evidence base to a controlled answer to the question every parent is really asking — not 'does AI help', but 'does it help or hollow out my child's actual ability'. The answer this trial gives is that it depends entirely on how the tool is built, not on whether AI is used at all.
That structured/unguarded distinction is not a one-off result. A separate randomised trial at Harvard, this time in undergraduate physics (N=194, crossover design), found students using a custom-built AI tutor — deliberately engineered with pedagogy best practices and given correct solutions to avoid hallucinating wrong answers — learned roughly twice as much as students in an in-class active-learning session on the same material, while reporting higher engagement and motivation [AIE-08]. The explicit caveat researchers attached matters as much as the headline result: this is evidence for a purpose-built tutoring tool, not for generic, unstructured chatbot use, which is a different thing entirely [AIE-08]. Read together, the Turkish and Harvard trials say something consistent — a well-designed, pedagogically-constrained AI tutor can outperform both unguarded AI and, in the Harvard case, live classroom instruction, while an AI tool with no guardrails can actively damage a student's ability to work independently once it is taken away.
Meanwhile, the promise that schools could simply detect AI use and hold the line has not held up. OpenAI discontinued its own AI-text classifier six months after launch, citing a low rate of accuracy — it correctly flagged only 26% of AI-written text while wrongly labelling 9% of genuinely human-written text as AI-generated [AIE-04]. The bias compounds for the students least equipped to fight a wrongful accusation: a peer-reviewed study found GPT detectors misclassified over half of essays written by non-native English speakers as AI-generated (a 61.22% average false-positive rate), against near-perfect accuracy on native-English writers — a gap that dropped to 11.77% only once students were coached to simply reword their own writing to sound less 'formulaic' [AIE-05]. Turnitin, the dominant commercial detector, states in its own published guidance that it tunes for a false-positive rate below 1% by deliberately trading away detection sensitivity — a design choice that sits uneasily against the tool's separately-marketed 98% accuracy headline, a gap independent commentary has flagged as worth scrutiny whenever a student is actually accused [AIE-06]. Regulators have drawn the obvious conclusion. Australia's TEQSA required every higher-education provider to submit an institutional generative-AI action plan by July 2024, and every one of them did — a 100% response rate [AIE-11]. The thrust of that regulatory programme is toward redesigning how assessment works, not toward better detection, because TEQSA's own posture — evidenced by the shift from a compliance ask to an institution-wide action-plan requirement — treats detection as a dead end rather than a fix [AIE-11].
Australian schools have already lived through one detection-era, ban-first cycle and moved past it. NSW banned ChatGPT in state schools in January 2023; within nine months, national Education Ministers had approved the Australian Framework for Generative AI in Schools (5 October 2023), implemented from Term 1, 2024, replacing the ban with an enabling policy built on six principles and 25 guiding statements [AIE-10]. The policy arc, in other words, ran from prohibition to structure in under a year — a faster pivot than most families noticed happening.
None of this settles the harder question of what AI use actually does to a student's mind over time, and the evidence here should be handled carefully. The most viral claim in this space — an MIT Media Lab EEG study reporting 'cognitive debt' from LLM-assisted essay writing, showing reduced neural connectivity in AI-assisted writers — is an unreviewed preprint with a small sample: 54 participants across three sessions, only 18 of whom completed the critical fourth session testing skill transfer [AIE-09]. It has already drawn a critique of its own — also an unreviewed preprint — flagging its small sample size, its EEG methodology, and reproducibility concerns [AIE-09]. That does not mean its worry is wrong — it means the finding is preliminary, not the settled neuroscience it is often presented as in shared social posts, and it should never be cited without that caveat [AIE-09]. A more grounded signal on the same worry comes from Anthropic's own analysis of over half a million university-level conversations on Claude.ai, which found that nearly 47% were 'Direct' interactions — students handing over instructions and taking outputs with minimal engagement — rather than collaborative back-and-forth use [AIE-12]. It is vendor-generated data about the vendor's own product, so it should be read as directional rather than definitive, but it points the same direction as the Turkish RCT: AI used as a shortcut, rather than as a scaffolded thinking partner, is where the risk concentrates.
Put the pieces together and a consistent picture emerges, even though the evidence base is young and still contains a live, unresolved preprint at its centre. AI use among Australian students and teachers is now the default, not the exception [AIE-13] [AIE-14]. Detection was tried and has not worked reliably enough to be trusted as the safeguard, especially for non-native English speakers who bear a disproportionate false-positive burden [AIE-04] [AIE-05] [AIE-06]. Regulators in Australia have already pivoted from banning and detecting to structural redesign [AIE-10] [AIE-11]. And the two best-designed controlled trials available say the same thing from two different angles: AI built with pedagogical guardrails and accountability for the student's actual learning produces real gains, while AI used without that structure can leave a student measurably worse off than if they had never touched it at all [AIE-07] [AIE-08]. For a family deciding how their child should use AI for schoolwork, the evidence does not support waiting for a ban, and it does not support unlimited, unsupervised access either. The live question was never whether a student uses AI — nearly all of them already do. It is whether that use is structured by someone who is actually accountable for whether the student learns, or left to run unguarded until the gap shows up later, once the tool is taken away and the skill was never really built.