Skip to content
The research behind the practice

The Evidence Register.

300 verified findings across 16 domains — quoted verbatim, fully cited, independently checked, and read together in essays and through-lines. Each domain has its own page; every claim is indexed below.

300verified findings
16domains
138peer-reviewed / causal
2026-09-16last verified
300verified entries, 16 domains, sources 1984–2026
  • Government and statutory162
  • Peer-reviewed78
  • Other source types60
The register, countedEvery figure above is computed from the register file itself at build time, so it cannot drift from what it describes. 40further items sit in the frontier — read and quoted, not yet folded into a domain. Counting sources is a statement about provenance, not about quality: a government dataset is not automatically better evidence than a journal article, and the register grades neither.

Looking for something specific rather than the whole register? The questions parents actually ask — whether the boys’ writing gap is real, whether handwriting still matters, whether a selective school changes anything — are each answered from these entries, with the limits of the evidence stated.

Twenty minutes, free, to talk it through — book a discovery call →

How this register is verified — every quote word-for-word, every statistic digit-for-digit, failures removed rather than softened. Read the protocol.

Every entry is collected and then separately re-checked against its primary source: the quote confirmed word-for-word, the statistics matched digit-for-digit, and the claim tested against the source’s actual scope — population, year, jurisdiction. Anything that fails is removed, not softened. Findings that later failed replication are shown with that record, because the caveats are part of the evidence. The domain essays, the chart and the through-lines are views of the same database: every factual sentence and every plotted point cites an entry, and the site refuses to build if a reference drifts. Cite any finding by its anchor — for example #TUT-08; older links to anchors on this page forward automatically. Register version 1.56.2.

The domains

Tutoring (21)What one-to-one teaching measurably does, and how common tutoring is in Australia.Writing (28)How Australian students are writing, and which ways of teaching writing carry evidence.Reading (24)How reading is learned, and where Australian reading results stand.VCE English (58)What VCE assessors report distinguishes strong responses, from the examination reports themselves.Handwriting & typing (16)What the research actually shows — including the famous findings that failed replication.AI in education (18)Prevalence, detection reliability, and what AI assistance does to learning.Selective entry & testing (22)What selective schooling and test preparation demonstrably change, and what they don't.Vocabulary (14)How word knowledge is built, and what it carries.Practice & memory (16)The learning-science mechanisms weekly lessons are built on.Literature (13)What reading fiction and writing craft are actually evidenced to do.Spelling, grammar & punctuation (12)Where Australian students stand on the conventions of language, and what instruction shifts them.Senior English nationally (10)English as the one universal Year 12 subject — enrolments, achievement and standards across the states.EAL/D learners (13)The scale and diversity of students learning English as an additional language, and what supports them.Early childhood (14)Where language gaps begin — development data and the early-years evidence.The teaching workforce (13)Who teaches English in Australia — time, shortages and out-of-field rates, stated systemically.Adult literacy (8)The long-run stakes: where school literacy ends up in adult life.

Frontier — new education research, read as it lands is the register’s forward edge: research read close to publication, where the standing is not yet settled. The register holds only what has already been verified; the frontier shows what is being watched, dated, with its caveats stated.

These findings, applied

The register states what the evidence shows. These pages use it to answer the questions people actually ask — each claim anchored to an entry above, so the working is visible rather than asserted.

The numbers at a glance

Standardised effect sizes reported by the register’s sources (d or g). A bar shows a source’s reported range. Effects from different designs are not directly comparable — each row links to its entry, where the population and design live. One difference does more work here than it looks: what a study measured with. Vocabulary instruction appears twice below, at 0.50 and 0.10, and the gap is the same intervention scored on researcher-built tests against standardised ones. The Grades 4–12 reading meta-analysis reports the same split independently. Treat a number measured on a researcher’s own test as the ceiling, not the expectation.

Through-lines

Read across its domains, the register keeps arriving at a small number of conclusions. Each step below is a claim the database actually carries — follow the anchors.

Structure beats hours

Across tutoring, study technique and AI use, the evidence keeps converging on the same shape: it is the structure around the work — not the volume of it — that moves learning.

  1. Tutoring works in randomised trials — and works best where it is structured: at least three sessions a week, in school, with a teacher or paraprofessional tutor. [TUT-02]
  2. The same pattern inside study technique: of ten common methods, only structured retrieval and spacing rate high-utility. [PRC-01]
  3. Undergraduate intuitions ran the other way — restudiers felt more confident, and a week later remembered less. [PRC-08]
  4. And the newest version of the finding: unguarded AI help boosted practice scores, then left students worse once removed — while a pedagogically-structured AI tutor largely mitigated the harm. [AIE-07]
  5. The common mechanism is feedback quality: high-information feedback roughly doubles the average feedback effect. [TUT-14]

What examiners keep rewarding

VCE assessors publish, every year, a precise description of the difference between competent and excellent English. Read across years, it is a curriculum for excellence — and it matches the writing-instruction evidence.

  1. Assessors reward evidence that is embedded and analysed, not quotes 'tacked on' as proof. [VCE-02]
  2. High scorers analyse a topic's implications — its absolute terms, connotations and silences — rather than restating it. [VCE-05]
  3. Literature assessors are blunter still: paraphrase 'will score very few marks'. [VCE-11]
  4. The instruction evidence agrees: in Writing Next, explicit summarization strategies carried an effect of 0.82 — tied with strategy instruction for the largest of eleven elements, though on only four studies. [WRI-13]
  5. The caution: when the test itself becomes the curriculum, teaching usually narrows — though a significant minority of tests widened it — which is why craft, not formula, is the durable preparation. [TST-10]

Reading compounds

Reading volume behaves like compound interest: exposure drives vocabulary, vocabulary drives comprehension, and the habit's dividend is still measurable decades later.

  1. The volume gap is enormous — a 90th-percentile reader meets roughly two million words a year outside school; a 10th-percentile reader, about eight thousand. US estimates, extrapolated rather than measured. [RDG-09]
  2. Printed school English runs to ~88,500 word families, on the American corpus estimate — too many to teach directly; wide reading, with context and morphology, is the route through most of them. [RDG-18]
  3. Vocabulary and comprehension then feed each other reciprocally. [VOC-10]
  4. Longitudinally, childhood reading-for-pleasure predicts stronger progress to age 16 — with an effect around four times that of having a graduate parent. [LIT-04]
  5. And the arc holds: frequent childhood readers still show a measurable vocabulary advantage at age 42. [LIT-09]

Famous findings that didn't hold

Some of education's most-quoted findings weakened or failed when someone looked again. The register keeps them — with their replication records — because knowing what didn't hold is part of knowing what does.

  1. Bloom's 'two sigma' tutoring effect is real as a 1984 result — [TUT-05]
  2. — but has never replicated at scale; modern trials cluster far below 2.0. [TUT-06]
  3. 'The pen is mightier than the keyboard' was a genuine 2014 finding — [HND-01]
  4. — that a large direct replication could not reproduce. [HND-03]
  5. Reading literary fiction briefly boosted theory-of-mind scores in 2013 — [LIT-01]
  6. — and a three-lab replication found no advantage. [LIT-02]
  7. The '30-million-word gap' — from 42 American families — shaped policy for two decades, [VOC-03]
  8. — before a 2019 replication across five communities found little to no SES difference by its method. [VOC-04]

Every finding, in one index

One line per verified claim. Search or filter, then follow a claim through to its full quote, citation and domain essay.

Tutoring (21) — essay & full entries
Writing (28) — essay & full entries
Reading (24) — essay & full entries
VCE English (58) — essay & full entries
Handwriting & typing (16) — essay & full entries
AI in education (18) — essay & full entries
Selective entry & testing (22) — essay & full entries
Vocabulary (14) — essay & full entries
Practice & memory (16) — essay & full entries
Literature (13) — essay & full entries
Spelling, grammar & punctuation (12) — essay & full entries
Senior English nationally (10) — essay & full entries
EAL/D learners (13) — essay & full entries
Early childhood (14) — essay & full entries
The teaching workforce (13) — essay & full entries
Adult literacy (8) — essay & full entries

The register is a reference, not a sales document — several entries cut against tutoring-industry claims, and they stay in. For how the practice applies this evidence, start with the philosophy; for what a lesson actually looks like, the discovery call is the honest first step.