Start with the least comfortable finding in this domain. When Dunlosky and colleagues reviewed ten of the most common study techniques against the experimental evidence, only two earned a 'high utility' rating: practice testing and distributed practice [PRC-01][PRC-02]. Five of the techniques students reach for by default — rereading, highlighting, summarisation, the keyword mnemonic, and imagery for text learning — were rated low utility, for reasons ranging from inconsistent effects to a lack of evidence that they help beyond the immediate task [PRC-03]. Interleaved practice, elaborative interrogation, and self-explanation sat in between, rated moderate utility [PRC-03]. The gap between what works and what students habitually do is close to the whole story of this domain.
Retrieval practice is the more counterintuitive of the two winners, because it feels worse while it is working better. In a controlled study of 180 Washington University undergraduates, one group took a single recall test on a passage while another restudied it the same number of times. Immediately afterwards, restudying produced better recall (81% vs 75%) [PRC-08]. That advantage reversed within two days (68% vs 54% in the tested group's favour) and held at a similar-sized gap a week later: at the one-week mark the tested group recalled 56% of the material against 42% for the restudy group [PRC-08]. A second experiment in the same study found recall scaled with how many times students had been tested rather than restudied [PRC-08]. The sting is in the confidence data: the restudy-only group, who went on to remember the least, were significantly more confident they would remember the passage in a week than the tested groups were [PRC-08]. Rereading feels productive. Being tested on the material feels harder and less certain. The evidence says the feeling is backwards — which is exactly why students default to rereading and skip retrieval unless someone builds it into the structure for them.
Spacing works on the same logic and compounds with it. The foundational meta-analysis behind the 'high utility' spacing rating pooled 839 assessments of distributed practice across 317 experiments in 184 articles, and found that the ideal gap between study sessions grows as the desired retention interval grows [PRC-07]. That is a genuinely useful design parameter, not just a slogan about not cramming: if you want a technique retained for months, the sessions practising it need to be spread over a period closer to that scale, not squeezed into a single afternoon. It is worth being honest about where this evidence comes from — the underlying studies are overwhelmingly laboratory tasks involving word lists, paired associates, and prose recall, not authentic classroom essay-writing [PRC-02][PRC-07]. Applying it to how long a persuasive-writing technique needs practising is a reasonable extrapolation, not a direct classroom replication.
Worked examples sit next to retrieval and spacing as a third well-evidenced tool, with a caveat attached from the start. A meta-analysis cited in an Australian government review found a moderate positive effect (d = 0.52) for studying fully worked examples rather than solving problems unaided [PRC-04] — though that figure traces back to an unpublished doctoral thesis cited secondhand by the government review, so it should be held a little more loosely than a directly-verified primary result [PRC-04]. The same review is explicit about the boundary condition: the 'expertise reversal effect' means heavy reliance on worked examples helps novices but becomes progressively less useful, and eventually counterproductive, as a learner's competence grows [PRC-05]. The design implication is to model first and fade the modelling as skill develops, not to keep supplying worked examples indefinitely once a student has the technique.
The same government review makes the case for explicit instruction generally — decades of research showing that for novices, direct guidance with practice and feedback beats minimally-guided discovery learning [PRC-06] — but it does not let that claim travel further than the evidence supports. The review states plainly that cognitive load theory is best evidenced by randomised controlled trials in technical domains like mathematics and science, and that 'far less research has been done on whether cognitive load theory is effective for teaching in less technical, or more creative subject areas — such as literature, history, art and other humanities subjects' [PRC-06]. For an English and writing tutoring practice, that caveat is not a footnote; it is close to the headline. The explicit-instruction case is strong, but its strongest trial base sits in maths and science classrooms, not essay-writing ones — a gap this register is not going to paper over.
Interleaving offers the field study most directly comparable to a real classroom. Rohrer, Dedrick, and Stershic ran a genuine school-based experiment: 126 seventh-grade maths students practised the same problems over three months, arranged either by the usual blocked method (one problem type at a time) or interleaved (mixed types within a session) [PRC-10]. Interleaved practice beat blocked practice on both an immediate test (d = 0.42) and a 30-day delayed test, where the advantage nearly doubled (d = 0.79) [PRC-10]. The domain is mathematics, and the specific effect sizes shouldn't be assumed to transfer to essay technique — but the structural principle (mixing skill types within a practice set rather than drilling one technique to exhaustion before moving on) is a transferable design idea for how a practice set between lessons is built.
Two findings guard against overclaiming on either side of 'more practice = more improvement'. First, a meta-analysis spanning multiple performance domains found that accumulated deliberate practice explains only 4% of the variance in achievement within education specifically — far less than in games (26%), music (21%), or sports (18%) [PRC-09]. That is not a case against practice; the same review argues practice quality and structure matter more than raw hours in education, which supports a feedback-driven, error-corrected model of practice over a simply-do-more-of-it one [PRC-09]. Second, whether repeated practice actually improves performance depends heavily on what happens after each attempt: a large meta-analysis of feedback research (435 studies, N>61,000) found feedback carrying real information about the nature of an error and how to fix it was roughly four times more effective than simple praise or correction alone (d = 0.99 vs d = 0.24) [PRC-11][TUT-14]. Practice without information-rich feedback is closer to repetition than to the deliberate-practice model the evidence supports.
Homework's evidence base is the most modest of the group, and the register says so plainly. A synthesis of US research from 1987 to 2003 found a generally consistent positive correlation between homework and achievement, notably stronger for secondary students (Grades 7-12) than for younger children — but the authors themselves note every study reviewed, of any design, had methodological flaws, and the finding is correlational, not causal [PRC-12]. That is enough to support the idea that between-lesson written work matters more for secondary, VCE-track students than for younger primary students, but not enough to claim homework itself directly causes achievement gains.
Put together, these mechanisms map fairly directly onto how a weekly tutoring lesson should run, as design implications rather than measured claims about tutoring itself. Weekly spacing between lessons sits inside the range the distributed-practice literature examined, rather than compressing everything into occasional intensive blocks [PRC-02][PRC-07]. Opening a lesson with retrieval — asking a student to recall last week's technique before it is retaught — trades the comfortable feeling of reviewing notes for the harder, more durable work of pulling the memory back out [PRC-08]. Modelling a technique with a worked example before a student attempts it independently, then deliberately withdrawing that scaffolding as competence builds, follows both the worked-example effect and its expertise-reversal limit [PRC-04][PRC-05]. Writing between lessons, mixing technique types rather than drilling one in isolation, and returning feedback that names the specific error and its fix rather than just marking it right or wrong, follows the interleaving and feedback-quality findings [PRC-10][PRC-11]. None of this is a tutoring RCT — it is inference from adjacent domains, mostly laboratory and maths/science classroom evidence, applied to a humanities subject the source literature itself flags as thin ground. That caveat is carried forward here on purpose, not smoothed over.
Figure · from the registerThe testing effect — recall over a week
Roediger & Karpicke (2006): restudying wins five minutes after learning, then the lines cross — the tested group retains more at two days and one week. This re-draws the study's own retention figure with its published percentages (#PRC-08).
Practice testing (retrieval practice / self-testing) is rated a high-utility learning technique, with effectiveness demonstrated across ages, abilities, and educational contexts.
Practice testing and distributed practice received high utility assessments because they benefit learners of different ages and abilities and have been shown to boost students' performance across many criterion tasks and even in educational contexts.
Distributed (spaced) practice is rated a high-utility learning technique on the same tier as practice testing, in contrast to massed/cramming approaches.
Practice testing and distributed practice received high utility assessments because they benefit learners of different ages and abilities and have been shown to boost students' performance across many criterion tasks and even in educational contexts.
Five commonly-used student study techniques -- rereading, highlighting/underlining, summarization, the keyword mnemonic, and imagery use for text learning -- were rated low utility due to inconsistent effects, narrow conditions of applicability, or lack of evidence for long-term or transfer benefits.
Five techniques received a low utility assessment: summarization, highlighting, the keyword mnemonic, imagery use for text learning, and rereading. These techniques were rated as low utility for numerous reasons.
The worked-example effect -- studying fully-solved example problems rather than solving problems unaided -- produces a moderate positive effect on novice learning, based on a meta-analysis (Crissman, 2006) cited by an AU government cognitive-load-theory review; Crissman's original analysis is an unpublished doctoral thesis, not independently re-verified here.
In a meta-analysis of studies on the effectiveness of worked examples, Crissman (2006) found an effect size of 0.52.
The worked-example effect is moderated by learner expertise: as students move from novice to more expert, heavy reliance on worked examples becomes progressively less effective and can even become counterproductive (the 'expertise reversal effect').
According to the expertise reversal effect, the heavy use of worked examples becomes less and less effective as learners' expertise increases, eventually becoming redundant or even counter-productive to learning outcomes.
Cognitive load theory, evidenced primarily by randomised controlled trials in technical/STEM domains (mathematics, science, technology), supports explicit instruction with modelling, guided practice, and feedback as more effective and efficient than minimally-guided/discovery approaches for novice learners; the source document itself notes far less research has tested this in humanities/creative subjects such as literature.
Decades of research clearly demonstrate that for novices (comprising virtually all students), direct, explicit instruction is more effective and more efficient than partial guidance. So, when teaching new content and skills to novices, teachers are more effective when they provide explicit guidance accompanied by practice and feedback, not when they require students to discover many aspects of what they must learn.
A large-scale meta-analysis of the spacing/distributed-practice effect across verbal recall tasks found consistent benefits of spacing study sessions apart rather than massing them, based on hundreds of experiments, and showed that the optimal gap between study sessions increases as the desired retention interval increases.
The authors performed a meta-analysis of the distributed practice effect to illuminate the effects of temporal variables that have been neglected in previous reviews. This review found 839 assessments of distributed practice in 317 experiments located in 184 articles. Effects of spacing (consecutive massed presentations vs. spaced learning episodes) and lag (less spaced vs. more spaced learning episodes) were examined, as were expanding interstudy interval (ISI) effects. Analyses suggest that ISI and retention interval operate jointly to affect final-test retention; specifically, the ISI producing maximal retention increased as retention interval increased.
In a controlled experiment with 180 Washington University undergraduates, students who took a single initial recall test on a studied prose passage retained substantially more of the material one week later than students who restudied the passage the same number of times, even though repeated studying produced better immediate recall and higher (miscalibrated) confidence.
Post hoc analyses confirmed that on the 5-min retention tests, restudying produced better recall than testing (81% vs. 75%), t(39) = 3.22, d = 0.52. However, the opposite pattern of results was observed on the delayed retention tests. After 2 days, the initially tested group recalled more than the additional-study group (68% vs. 54%), t(39) = 6.97, d = 0.95. The benefits of initial testing were also observed after 1 week: The tested group recalled 56% of the material, whereas the restudy group recalled only 42%, t(39) = 6.41, d = 0.83.
Across a meta-analysis spanning multiple performance domains, the amount of deliberate (structured, feedback-driven) practice a person has accumulated explains only a small proportion of variance in achievement within education specifically -- far less than in games, music, or sports -- indicating that practice quantity alone is not a sufficient explanation for educational attainment.
We found that deliberate practice explained 26% of the variance in performance for games, 21% for music, 18% for sports, 4% for education, and less than 1% for professions. We conclude that deliberate practice is important, but not as important as has been argued.
Interleaved practice (mixing different problem types within a practice session, rather than practicing one type in a block before moving to the next) produced markedly higher test scores than blocked practice for secondary-school mathematics students, with the advantage growing larger on a delayed test than an immediate one.
In the experiment reported here, 126 seventh-grade students received the same practice problems over a 3-month period, but the problems were arranged so that skills were learned by interleaved practice or by the usual blocked approach. The practice phase concluded with a review session, followed 1 or 30 days later by an unannounced test. Compared with blocked practice, interleaved practice produced higher scores on both the immediate and delayed tests (Cohen's ds = 0.42 and 0.79, respectively).
A large-scale meta-analysis of educational feedback research found an overall medium positive effect of feedback on student learning, but effectiveness varied greatly by the type of feedback given: feedback carrying substantial information content (e.g. explaining the nature of an error and how to correct it) was roughly four times more effective than simple reinforcement or praise/punishment. NOTE: this same study is already the primary feedback-mechanism citation in this database's tutoring domain (entry TUT-14) -- retain this practice-domain entry only for its practice-loop framing (feedback quality as the gate on whether repeated practice attempts actually improve performance), not as an independent confirmatory source.
Overall results based on a random-effects model indicate a medium effect (d = 0.48) of feedback on student learning, but the significant heterogeneity in the data shows that feedback cannot be understood as a single consistent form of treatment... Feedback is more effective the more information it contains.
A large research synthesis found generally consistent evidence for a positive CORRELATION (not a causal/RCT-established effect) between homework and achievement in US studies from 1987-2003, and that this correlation is notably stronger for secondary students (Grades 7-12) than for younger elementary students (K-6). The authors note every study reviewed, of any design, had methodological flaws.
In this article, research conducted in the United States since 1987 on the effects of homework is summarized. Studies are grouped into four research designs. The authors found that all studies, regardless of type, had design flaws. However, both within and across design types, there was generally consistent evidence for a positive influence of homework on achievement. Studies that reported simple homework-achievement correlations revealed evidence that a stronger correlation existed (a) in Grades 7-12 than in K-6 and (b) when students rather than parents reported time on homework.
Cite any finding by its anchor — for example #PRC-01 — and the link resolves to the claim, quote and source above. The verification protocol lives on the register’s main page.