Neurogogy is the fusion of neuroscience and pedagogy: designing training around what the brain actually requires rather than around what feels like it is working.
Those two things come apart more often than the industry admits. The techniques with the strongest evidence behind them are often the ones learners rate worst.
This handbook sets out what the evidence supports, what it means for the way you build, and which numbers we should all stop repeating.
Every claim in it is traceable. Where a figure is contested, we say so. Where a popular statistic turns out to be folklore, we name it and retire it, including the ones our own industry has been repeating for forty years.
Each part synthesises the research and then points you to the full article on The Lab, our research library, where the studies, the caveats and the counter-evidence are set out at length. Read the part for the argument. Follow the links when you need the detail. Every number carries a reference to the claims register at the back, which gives the full citation, the sample it came from and a link to the source.
Almost every conversation about training effectiveness reaches the same number sooner or later: only 10% of training transfers to the job.
It is not a research finding. It comes from a 1982 article in Training and Development Journal by David Georgenson1. It does not even appear there as his own conclusion. It appears in the opening paragraph as a line spoken by an unnamed training director: "I would estimate that only 10 percent of content which is presented in the classroom is reflected in behavioral change on the job."
No study. No method. No data. An illustrative aside in a trade magazine, which the profession then spent four decades citing as evidence. In 2011, Ford, Yelon and Billington2 gave it the name it deserves: the 10% delusion.
Only 10% of training transfers to the job. Cited for four decades as evidence of systemic failure.
A line of dialogue attributed to a hypothetical training director in a 1982 trade magazine. No study, no method, no data.
So what does the evidence show? When Saks and Belcourt3 surveyed 150 training and development professionals, respondents estimated that 62% of employees apply what they learned immediately after training, 44% are still applying it six months later, and 34% after a year.
That distinction matters, because the two diagnoses lead to opposite prescriptions. If training fails at the point of delivery, you fix the content. If training works and then decays, you fix everything that happens after the content: the spacing, the retrieval, the reinforcement, the manager who either creates room to practise or does not. One of those is a courseware problem. The other is a design problem, which is where the research keeps pointing.
This handbook is about closing that gap. Not by adding more content, or better video, or a slicker interface, but by designing against how the brain actually encodes, stores, retrieves and loses information. The science has been sitting in the journals for decades. Bliss and Lømo demonstrated the cellular mechanism of memory strengthening in 1973.4 Ebbinghaus charted forgetting in the 1880s. Sweller formalised cognitive load in 1988.5 Very little of it has reached the average compliance module.
A second gap is less comfortable to name. A large amount of what our industry calls brain science is not science. Learners have styles that should be matched. Retention follows a neat pyramid of percentages. Ebbinghaus proved 90% of what you learn is forgotten within a week. Only 10% of training ever reaches the job. Every one of those claims is either unsupported or actively contradicted by the evidence, and every one of them still shapes real training budgets.
Design for the brain first, and learning will follow. That means accepting constraints you would rather not have, including a working memory that holds about four things, a forgetting process that begins the moment learning ends, and an attention system that no amount of enthusiasm will override.
Pedagogy puts the teacher in charge. Andragogy, as Malcolm Knowles6 framed it, recognises that adults direct their own learning. Heutagogy goes further, handing learners the choice of what to learn and why.
Neurogogy asks a different question. Not who leads, but what the brain requires. It is the fusion of neuroscience and pedagogy: evidence about how the brain encodes, consolidates and retrieves information, applied directly to how training gets designed, delivered and reinforced.
It does not replace the frameworks above. It underpins them. You can run a beautifully learner-led programme that still ignores cognitive load, still never revisits anything, and still produces nothing durable. Autonomy is a motivational condition. It is not a memory mechanism.
The question is never what learners enjoy, or what they say helps them. It is what demonstrably produces retention and transfer. These frequently point in opposite directions, and Part 3 explains why: the conditions that feel most productive during learning are often the ones that produce the weakest durable memory.
Neuroscience is only worth anything here if it changes what gets built. Every part ends with what the evidence means for the structure of a programme: how long, how often, in what order, reinforced how, measured by what.
Four rules, applied to every number in this handbook. They are worth stating because they are what makes the difference between a white paper and a brochure with footnotes.
Every figure here was checked against the paper it came from. Not a blog post, not a conference slide, not another vendor's white paper. Five of the sources are books with no online version we can link you to. Each of those entries is marked print only.
A result from sixteen people and a result from ninety-seven thousand are not the same kind of thing, and hiding the difference in a footnote is a choice. Where a study is small, you will see how small it is at the point the number is used.
An effect size of 0.83 means nothing to most readers, and a decimal presented without a scale is a way of sounding precise without being clear. Every effect size in this handbook is given with its plain-language equivalent, and proportions are distinguished from percentage points.
Effects are routinely largest in the laboratory that discovered them and smaller once a field pools its results. Where a single striking result and a pooled one disagree, we have set expectations by the pooled figure, even when the bigger number was the one we would rather have quoted.
Six terms do most of the work in this handbook. If they are already familiar, skip the box.
Applying those rules meant retiring numbers we had used ourselves. These are the ones that did not survive.
Four constraints you cannot design around, and the folklore that has grown up around each of them.
Memory is a process rather than a place. Working memory is far smaller than you think. The brain physically reorganises in response to practice. And forgetting is the default state, not a malfunction.
In September 1953, a 27-year-old man named Henry Molaison underwent bilateral medial temporal lobe resection to control severe epilepsy. Scoville and Milner's 1957 report14 described what the surgery cost him.
Molaison's case established the architecture. Holding something in mind for a moment, storing it for years, and retrieving it later are separate operations, carried out by different structures, and they fail independently. His capacity to acquire new motor skills, demonstrated in later testing, showed that even long-term memory splits into systems that can survive without one another.
The mechanism beneath all of this is physical. In 1973, Bliss and Lømo4 stimulated the perforant path in anaesthetised rabbits and found that brief bursts of high-frequency stimulation strengthened the response of dentate granule cells for anywhere from 30 minutes to 10 hours. Long-term potentiation, as it became known, is the closest thing we have to memory made visible: repeated activation of a pathway makes that pathway easier to activate again.
A single exposure builds almost nothing worth keeping. Encoding, consolidation and retrieval are three jobs, and a one-off training event only attempts the first.
George Miller's 1956 paper15 is the most cited and most misread work in the field. The title gave us "the magical number seven, plus or minus two", and the phrase entered folklore. Read the paper and you find Miller openly sceptical of his own number, describing the recurrence of seven as "only a pernicious, Pythagorean coincidence". His substantive point was subtler and more useful: the limit applies to chunks, not to information, so recoding material into larger meaningful units increases how much you can hold.
Miller's headline number, which Miller himself called a "pernicious, Pythagorean coincidence".
Cowan's central capacity limit16, once rehearsal, chunking and long-term support are stripped out.
That constraint is what cognitive load theory is built on. John Sweller's insight, formalised from 1988, is that instruction which exceeds working memory capacity does not merely slow learning down. It prevents it.
Two kinds of load are worth separating, because you deal with them in opposite ways.
The difficulty built into the material itself: how many pieces have to be held in mind at once to make sense of it. Pricing a single product is light. Pricing one where discount, region and contract length all interact is heavy. You can sequence it, chunk it and build the prerequisites first. You cannot wish it away.
The difficulty the design adds on top, which does nothing for learning: the split-screen diagram that makes learners hold one element while hunting for the other, the narration reading the on-screen text aloud, the decorative animation. Unlike intrinsic load, all of this is yours to remove, and removing it is usually the cheapest improvement available.
One correction is due here, and it applies to a great deal of published L&D writing including our own.
Intrinsic, extraneous and germane load, with germane presented as a third dial you can turn up to increase learning.
Two things to manage. Germane load "redistributes working memory resources from extraneous activities to activities directly relevant to learning", with "a redistributive function… rather than imposing a load in its own right".17
The most direct way to cut extraneous load is to remove the search. Set a novice a problem and most of their working memory goes on hunting for a solution path while trying to hold the problem itself in mind. Give them a worked solution to study instead and that hunt disappears, leaving capacity for the method. It is the cleanest test of the theory available, because the content is identical and only the load changes.
Barbieri and colleagues18 pooled 55 studies across 43 articles and 181 effect sizes, and found worked examples improved mathematics performance at a medium average effect.
Medium. The average learner given worked examples finished ahead of roughly 68% of those who were not.
Ginns' meta-analysis of 50 studies19 found substantial gains from integrating spatially or temporally separated information, with the effect concentrated in complex material and close to nothing for simple material.
Cutting extraneous load works, and it works most where the content is hardest. Which is precisely where most training is thinnest.
Neuroplasticity is the brain physically changing in response to what you repeatedly do. Connections that fire together strengthen and grow new points of contact, the fibres carrying them get better insulated so signals travel faster, and connections that go unused are pruned away. It is the mechanism that makes any training work at all.
The machinery does not switch off with age. What changes with age is what it takes to engage it, and how much of the result you should expect to be able to see.
The best-powered test of this is not a brain scan. The ACTIVE trial randomised 2,832 adults aged 65 to 94 to training in memory, in reasoning or in processing speed, or to nothing at all. Each programme ran for ten sessions. Everyone was then followed for a decade.2021
Ten sessions, and two of the three trained abilities were still measurably better ten years later. That is a stronger durability result than most corporate programmes could claim at ten weeks.
The same trial sets the limit. The gains stayed inside the ability that was trained. Speed training made people faster and left their reasoning where it was, and no arm improved performance on observed everyday tasks at five years or at ten.22 Adults learn well. They do not generalise for free.
Lövdén and colleagues separate two things the industry runs together.23 Flexibility is the brain performing better within the capacity it already has. Plasticity is the harder change, where the structure itself adapts to do something it previously could not. Plasticity is triggered by a mismatch between what the brain can currently supply and what the situation demands. A task comfortably inside current capacity creates no adaptive pressure. One far outside it creates none either. And the mismatch has to last, because plasticity is, in their word, sluggish.
That is a design brief rather than a caveat. Find the gap, and hold people in it.
Zatorre, Fields and Johansen-Berg reviewed what a scanner is picking up when a trained brain looks different from an untrained one.24 Synapses are added and pruned, support cells multiply, axons gain insulation that makes established pathways faster, and blood supply increases where the work is.
Two timings are worth carrying into a plan. Grey-matter change has been detected after as little as seven days of training. White-matter change develops across roughly six weeks.
The pictures are real. They are also small. Trainee London taxi drivers who passed the Knowledge showed hippocampal growth across three years, where those who failed and a control group showed none. The same drivers ended up measurably worse than controls at recalling a complex figure, so capacity was reallocated rather than created.25
Across 33 imaging studies the typical structural effect runs to 2 to 5%. Regions often expand early in learning and partly shrink back while performance carries on improving, and the change fades once practice stops.2627
And the largest test of the lot found nothing at all. Judd and Kievit used the 1972 rise in the UK school leaving age as a natural experiment, comparing brain scans of roughly 5,100 people who were made to stay on an extra year against those who were not.28 Across 117 structural measures the answer was no detectable difference, and the analysis was registered before anyone looked at the data.
That is not the deflating result it first appears. A year of school plainly teaches people things. What it does not do is leave a mark large enough for a brain scan to find. The finding is about the resolution of the instrument, not about whether adults learn, which is exactly why the behavioural evidence above carries more weight than any scan.
One claim here has changed in our favour. Adult humans do grow new neurons in the hippocampus. Two independent studies reading the genetic activity of individual cells have now found the progenitors, after a decade of real dispute.2930 The rates are low, they vary enormously between people, and none of this is what the training scans are showing, so it rescues none of the marketing built on it.
The outcome of a training programme is not a state you reach. It is a state you maintain.
Adults keep the machinery, so three things follow. Set the difficulty at the gap between what somebody can do now and what the job demands. Keep them in that gap long enough for it to register, because a single hard afternoon changes nothing. And train the exact capability you want back, because the gains do not spread.
Hermann Ebbinghaus spent the 1880s learning lists of thirteen nonsense syllables and testing how much effort he saved when relearning them later. His data, reproduced by Murre and Dros in their 2015 replication11, look nothing like the version our industry repeats.
The subject was Ebbinghaus himself. Murre and Dros replicated the shape with a single subject over 70 hours of testing.
The material was chosen specifically to strip out meaning. Job-relevant material decays far more slowly, because it connects to what learners already know.
Savings measures relearning efficiency, not the proportion of items a learner could produce. The two are routinely confused.
So what survives? The shape survives, and the shape is enough. Retention falls steeply at first, then flattens. Each reinforcement resets the slope.
Design against the shape of forgetting and stop quoting its numbers. Reinforcement schedules are the intervention. The percentages are folklore.
Four design consequences follow, and none of them are about content quality.
Sequence and chunk intrinsic complexity, and strip extraneous load ruthlessly, starting with the hardest material rather than the easiest.
Consolidation and retrieval are separate jobs that require separate events at separate times.
No gap between current ability and real demand, no structural change. And what practice builds, the absence of practice takes back.
Reinforcement schedules are the intervention. The percentages attributed to Ebbinghaus are folklore. They are not his.
Emotion, motivation, sleep, stress and attention. Five conditions that decide what a programme achieves, and most of them sit outside the course.
A learner's capacity to encode, consolidate and retrieve is not fixed. It moves with how they feel, how much they slept, how threatened they are and what else is competing for their attention. None of that is soft. All of it is measurable, and most of it sits outside the course.
Part 1 dealt with the machinery. This part deals with its operating conditions. It is also the part of the field with the most folklore in it. Several of the statistics L&D uses to argue for exactly the right conclusions turn out to have no research behind them. We have named those where they arise. A fabricated number is the first thing an informed reader goes after, and it takes the good argument down with it.
The brain does not store what is important. It stores what is tagged, and emotion does the tagging.
The mechanism is modulation rather than storage. Emotional arousal triggers adrenal stress hormones, which drive noradrenergic activation in the amygdala, which in turn strengthens consolidation in the hippocampus and cortex. Cahill, Prins, Weber and McGaugh demonstrated this in humans in 1994 by blocking it.31
The effect shows up as a difference in decay rather than a difference in learning. Anderson and colleagues tested recollection across delays running from fifteen minutes to two weeks.32
Two further details are worth carrying. The effect appeared for recollection and not for vaguer feelings of familiarity. And it did not appear for fearful faces, so emotional intensity on its own is not the lever.
Proximity matters too. Sharot, Martorella, Delgado and Phelps scanned New Yorkers three years after 9/11 and split them by how vivid their recall was.33 The split turned out to track distance: the vivid group had been an average of 2.1 miles from the towers that morning, the rest 4.5 miles.
The curve the field reaches for to describe this, the inverted U of arousal against performance, is not what its source says.
The Yerkes-Dodson curve: performance rises with arousal to an optimum, then falls. Therefore a bit of pressure improves learning.
Forty dancing mice learning a white-black discrimination under electric shock.34 Strong stimulation sped up easy learning and impaired difficult learning. No arousal, no humans, and no curve.
Practitioners should not seek to increase performance through the manipulation of employee stress levels.Corbett (2015), reviewing the law's use in management35
There is a real inverted U in this territory. It is chemical rather than motivational. Salehi, Cordero and Sandi noted in 2010 that despite universal belief in the curve, nobody had demonstrated it under constant experimental conditions, so they did.36 Rats trained in a water maze at a moderate stress level made fewer errors than rats trained under either milder or harsher conditions, tracking their corticosterone.
Which is a statement about a stress hormone in a water maze, not about deadlines in an office. None of this makes emotion the problem. Emotion is the tagging mechanism this whole section is built on, and worth designing for. What the curve does not support is the separate idea that adding pressure will sharpen performance.
Emotion belongs in learning because it determines what gets kept. It does not belong there as pressure. The correction makes the case sharper, not weaker.
Deci's 1971 experiments produced the finding that launched fifty years of argument.37 Participants paid per puzzle solved spent less of a later free-choice period on the puzzles than they had before. Verbal praise produced no such decline. The individual studies were small and the original results were not decisive, but the pattern held up under synthesis.
Which is where most L&D writing stops, because it supports a comfortable conclusion: intrinsic good, extrinsic bad, build for meaning.
The larger and more recent evidence does not support that conclusion. Cerasoli, Nicklin and Ford pooled 183 samples and more than 212,000 people, and found the two work on different things.39
So the honest position is not that rewards corrode motivation. It is that tangible, expected, controlling rewards for work people already find interesting reduce their willingness to keep doing it unpaid, and that incentives buy volume while meaning buys care. If you are driving completion of mandatory compliance content, incentives are the right instrument. If you are trying to change how someone exercises judgement, they are the wrong one, and no amount of them will substitute.
The size of that gap is worth stating plainly. Van den Broeck and colleagues pooled 124 workplace samples and asked how the five types of motivation divide up the variance they jointly explain in job performance.40 Together the five account for a quarter of it.
Read that ranking twice. For performance specifically, believing in the work beats enjoying it, and being paid for it explains almost nothing. The same pattern holds in education. Pooling 344 samples and 223,209 learners, Howard and colleagues found intrinsic motivation related to student success and wellbeing, and personal value particularly strongly related to persistence, while motivation driven by a desire to obtain rewards or avoid punishment was associated with neither performance nor persistence, and was associated with decreased wellbeing.41
Underneath all of this sits a mechanism that is routinely described backwards. Dopamine is not a pleasure signal. It tracks prediction error, the gap between what was expected and what arrived, and it drives the willingness to expend effort. Wanting and liking are separate systems. That distinction is why a reward that arrives predictably stops motivating while an unexpected one still does, and why the feeling of progress does more work than the prize at the end.
The engagement industry treats boredom as the enemy. Boredom is real and consistent. It is also far smaller than the industry assumes.
Qi and colleagues, pooling 21 studies and 240 effect sizes, found negative emotions associated with worse online learning performance at r = −.303, and positive emotions with better performance at r = .478.43 The positive side of the ledger is the larger one, which argues for building interest rather than merely removing tedium.
Match the instrument to the outcome. Incentives move completion. Meaning moves judgement. And the bigger prize is on the positive side: building interest beats stripping out tedium.
Read on The LabIntrinsic Motivation vs Extrinsic Motivation
Dopamine and Learning
Consolidation is not something that happens after learning. It is part of learning, and most of it happens while the learner is unconscious.
Diekelmann and Born's review sets out the mechanism.44 During slow-wave sleep, slow oscillations, spindles and ripples coordinate the reactivation and redistribution of hippocampus-dependent memories towards the neocortex. REM sleep supports their synaptic stabilisation. Sleep does not protect memory. It processes it.
Three findings make this operational.
Put sleep between the input and the point where it has to be used. A workshop that teaches in the morning and applies in the afternoon gives the material no night at all. Learners will still rate it well. They are rating it at the moment their fluency peaks, which is the moment before the forgetting starts.
The scale of this is not in dispute. The International Labour Organization's April 2026 global report puts the losses from psychosocial risks at work at 1.37% of global GDP each year.48 What is in dispute is what to do about it, and most of the received wisdom points at the wrong lever.
Sustained exposure does more. Newcomer and colleagues gave 51 healthy adults placebo or one of two cortisol doses for four days, the higher dose approximating the exposure seen during major stress.51
It matters because it changes what you can do about it.
Stress raises cortisol, cortisol harms memory, therefore lower cortisol and memory improves.
Across 113 studies and 6,216 participants, stress reliably raised cortisol, but the magnitude of the cortisol response was not related to the effect of stress on memory.52
Small, and consistent: the confidence interval runs from −0.346 to −0.085, and the studies agree closely with each other. Stress at retrieval reliably costs you something, but not much.
"Reduce cortisol" is therefore not the lever the wellbeing literature implies. Something about the stressed state impairs retrieval, and cortisol concentration is not a reliable index of it.
Sonia Lupien's Centre for Studies on Human Stress sets out the four ingredients that reliably provoke a stress response.53 The acronym, N.U.T.S., is the centre's. The components come from Mason's work and from Dickerson and Kemeny's meta-analysis of social-evaluative threat.54
Something the learner has not encountered before. An unfamiliar interface counts.
No way of knowing what is coming. Unclear expectations, unannounced assessment.
Competence called into question. Assessment that can embarrass someone in front of peers.
Little or none over the situation. No choice over pace, order or route.
Every one of the four is something a learning designer controls directly. That is a four-item audit you can run before you add anything, worth more than any resilience module.
Read on The LabTaming Cortisol: The Neuroscience of Stress-Free Learning Design
Start with what holds. Three findings are robust, and enough on their own to justify redesigning most corporate training.
Now the clear-out, because this topic carries more folklore than any other in L&D. Every row below is a number in current circulation, and none of them should be.
| The claim | What it actually is | Status |
|---|---|---|
| Willpower is a fuel tank | Ego depletion failed a 23-laboratory preregistered replication of 2,141 people at d = 0.04, with a confidence interval spanning zero.55 A second, larger multi-site test across 36 labs and 3,531 people returned d = 0.06, and its Bayesian analysis found the data four times more likely under the null.56 | Failed replication |
| Task switching costs 40% of productive time | A remark attributed to David Meyer on an APA explainer page, which reports that he "has said" switching can cost as much as 40%. No study cited. The 2001 paper it is attached to measured switching costs in fractions of a second.57 | Not a finding |
| Media multitaskers have worse cognitive control | Two powered replications found a significant effect in 5 of 14 tests, only two surviving a conservative Bayesian analysis. The accompanying meta-analysis turned non-significant once publication bias was corrected.58 | Largely unreplicated |
| 2.5% of people are "supertaskers" | Five people out of 200 tested in a driving simulator, identified by a threshold on difference scores. Never independently replicated.59 | Single study |
None of that weakens the design case. It strengthens it, because the case never needed the numbers. Attention is limited, protecting it is a design decision, and the honest version is more defensible than the dramatic one.
Read on The LabThe Neuroscience of Focus
The Multitasking Myth
Four consequences, none of which are about content.
Sleep, stress and attention account for a large share of what a programme achieves, and all three sit outside the course. If you control none of them, you are optimising the smaller half of the problem.
Relevance, stakes and narrative earn their place because they determine what gets consolidated. Pressure does not, and the evidence usually cited to justify it is a study of mice.
Incentives buy volume. Meaning buys care. Of everything the five motivation types explain about job performance, rewards and punishments account for under 1%, so if judgement is what you need, the reward budget is not where to find it.
Novelty, unpredictability, threat to ego, low control. Removing those is cheaper than anything you could add, and the one thing here you can act on this week.
Seven techniques with evidence behind them, what each one is actually worth, and the reason almost nobody uses them.
The conditions that make learning feel productive and the conditions that make it durable are frequently opposites. Every technique below costs the learner something in the moment and pays them back later. Every technique that feels smooth costs them later instead. Robert Bjork called that category desirable difficulties. It is the thread running through all seven.
Parts 1 and 2 dealt with the machinery and its operating conditions. This part is the shortlist.
In 2013, John Dunlosky and colleagues60 rated ten widely used study techniques against the evidence behind them. Two reached high utility: practice testing and distributed practice, which this part calls retrieval and spacing. Three more reached moderate: elaborative interrogation, self-explanation and interleaved practice. Rereading and highlighting, the two techniques that dominate classrooms and corporate learning alike, came bottom.
That verdict sets the order of this part. None of this is new. Most of it has been sitting in the literature for thirty years or more. The gap is not knowledge.
Trying to pull something out of memory rather than putting it in again. A quiz, a flashcard, or closing the book and writing down what you can remember.
Assessment asks what a learner knows. Retrieval practice uses the asking to build what they know. Same mechanic, different purpose. Retrieval practice is worth considerably more.
Karpicke and Roediger demonstrated the size of it in 2008.61 Students learned forty Swahili-English word pairs. Those who kept being tested on the pairs recalled 80% a week later. Those who kept restudying them recalled 36%. Same material, same total study time, and more than double the retention.
Almost nobody does this. Surveying 177 undergraduates about how they study, Karpicke, Butler and Roediger62 found 84% listed rereading among their strategies and 11% mentioned practising retrieval. Asked to name the single strategy they used most, 55% said rereading. Practising recall was named by 1%. Two students out of 177.
Reread until it feels familiar. 84% report doing it, and 55% name it as their main strategy.
Put the material away and try to produce it. 11% report doing it, and 1% name it as their main strategy.
The reason is not ignorance, it is fluency. Roediger and Karpicke63 found that learners who restudied felt more confident about what they would remember, and then performed dramatically worse than learners who had been tested. Restudying produces the sensation of knowing. Retrieval produces the knowing, and feels worse doing it.
Every quiz in your programme is doing two jobs: measuring what stuck, and making more of it stick. Most organisations bank the first and throw away the second, because the quiz sits at the end of the module rather than through it.
Separating repeated encounters with the same material in time, instead of massing them into one session. The material does not change. Only the calendar does.
Spacing is the other technique that cleared Dunlosky's bar, and the cheapest intervention in this handbook. It needs no new content, no new platform and no new budget. It needs a calendar.
Rohrer and Taylor ran the cleanest demonstration of why in 2007.64 Three groups of undergraduates practised the same kind of maths problem, then sat a test a week later.
How long the gap should be depends on how long you need the material to last. Cepeda and colleagues65 taught facts to more than 1,350 people, reviewed them after gaps of up to three and a half months, and tested up to a year later.
Lengthening the gap raised performance and then lowered it again, so there is an optimum rather than a "longer is better" rule. The optimum tracks the horizon you are designing for, but it does not scale with it: the further out you need the material to survive, the smaller a fraction of that horizon the best gap becomes.
Turned into dates, which is the form you can actually schedule against:
Review roughly one to three days after the first session.
Review roughly a week later.
Review roughly three to five weeks later.
Those dates are our own arithmetic on the proportions Cepeda reports, which run from about 20 to 40% of a one-week delay down to about 5 to 10% of a one-year delay. Treat them as the right order of magnitude, not as a prescription.
Then the part the market would rather not hear. Karpicke and Bauernschmidt66 gave learners three repeated tests on the same items and varied both the total spacing and its pattern. More total spacing produced a 200% improvement in long-term retention over tests with no gap at all. But expanding intervals, the schedule almost every spaced repetition product is built on, performed no better than equal or even shrinking ones.
The relative schedule of repeated tests had no discernible impact.Karpicke & Bauernschmidt, 2011
Total spacing is the active ingredient. Algorithms earn their keep by forcing spacing to happen at all, not because their curve is optimal.
Mixing different problem types within a practice session, so a learner has to work out which method a problem calls for before they can apply it.
Most training is blocked. One topic, practised until it feels solid, then the next. Interleaving mixes the practice instead, so learners have to work out which approach a problem calls for before they can apply it. That second step is the whole point, and blocked practice removes it.
The strongest classroom evidence in this handbook comes from the same team five years later.
Large by any standard. The average interleaved student finished ahead of roughly 80% of the blocked group.
Set against that, the wider literature is more modest. Brunmair and Richter69 pooled 59 studies and found an overall interleaving effect of 0.42, and for mathematics specifically 0.34. For word learning, blocking actually won. One large trial and a whole literature rarely agree, and where they disagree the literature is the safer number to plan against.
There is also a condition, and the paper puts it in the title. Brunmair and Richter called their meta-analysis "Similarity matters". That is the finding in two words. Interleaving pays when the things being mixed are genuinely confusable, because the benefit comes from learning to tell them apart. Mixing unrelated topics does much less, and shuffling a programme for its own sake does nothing.
Block the first exposure, then interleave everything after it. Interleaving before a learner has any foundation is not a desirable difficulty. It is just difficulty.
Robert Bjork's term for conditions that make learning slower and harder while it is happening, and stronger once it is done. The three techniques above are all instances of it.
Retrieval, spacing and interleaving share a property. Each one makes practice harder and makes performance during practice worse. Each one also produces more of what is left a month later. Bjork's point was that these are not two separate findings: the difficulty is doing the work.
The clearest single demonstration is not about any of the three. It is about being wrong on purpose. Kornell and colleagues70 gave learners word pairs they had no way of guessing. One group saw the cue alone, failed to produce the answer, and was then shown it. The other group simply studied the pair. Both groups had the same thirteen seconds per item.
The reach for an answer prepares the brain for the right one, even when the reach fails.
Difficulty works by forcing connections to existing knowledge. With no existing knowledge, learners generate misconceptions and then cement them.
An uncorrected guess hardens into a false belief. The productive sequence is always attempt, then answer.
A learner who can reach the answer with effort is learning. A learner who cannot is being overloaded, and nothing is encoded at all.
Stop treating a struggling learner as evidence of a design fault. Check the three conditions first: do they have a foundation, will they get the answer afterwards, and is the task within reach? If all three hold, the struggle is the mechanism working, and smoothing it out is what would cost you.
Presenting the same idea in words and in pictures, so it is encoded through two channels rather than one. Not to be confused with the myth it is usually mistaken for.
This one needs its ground cleared before it can be used, because it is routinely confused with the most persistent myth in the field.
Learners have a preferred visual, auditory or kinaesthetic channel, and matching material to it improves learning. No credible supporting evidence. Believed by 89.1% of educators across 37 studies and 18 countries.
Every learner has two channels, verbal and visual. Material processed through both is held better than material processed through either alone.
Allan Paivio proposed the theory in 1971.71 The concreteness effect is the everyday evidence for it: words that evoke an image are recalled better than abstract ones.
Richard Mayer took it into instructional design.72 Across eleven controlled experiments in his own laboratory, learners given words and pictures together beat learners given words alone on every single comparison.
The largest effect in this part. Eleven experiments, all in Mayer's own laboratory, and positive in every single comparison.
The counterintuitive corollary is more useful than the principle. Mayer also found that narration paired with graphics and on-screen text produced worse encoding than narration and graphics alone. Two channels beat one. Three inputs across two channels is worse than two, because the verbal channel is now processing the same content twice.
The instruction is not "add visuals". It is: use each channel once, and never say in text what the narration is already saying.
Making the learner produce the answer, the explanation or the example themselves, rather than giving them one that is already written.
Reading an answer and producing one are not the same cognitive event. Only producing builds a memory worth having.
Teaching is generation at full stretch, and the benefit arrives before the lesson does. Preparing to explain something forces a learner to organise scattered facts, anticipate questions and find the gaps in their own understanding, and the preparation alone does most of the work.
Which puts most current technology on the wrong side of the argument. A field experiment in Turkish schools75 gave students access to GPT-4 during maths practice. While they had the tool in front of them, performance rose sharply: 48% for a standard chatbot, 127% for a version built to tutor rather than answer.
Then the researchers removed it and ran an unassisted exam. The students who had used the standard chatbot scored 17% below students who had never had access at all. The practice gains were real. They belonged to the tool rather than to the learner. The tutoring version largely removed the penalty.
Generative AI Can Harm Learning.Bastani et al., paper title
The question is not whether your learners use AI. It is whether your AI does the thinking or makes them do it.
Feedback is the most trusted intervention in corporate learning, and the one most likely to do harm. In 1996, Kluger and DeNisi76 pooled 607 effect sizes from 131 studies covering 12,652 people. On average feedback helped, at 0.41. In more than 38% of cases it made performance worse. They noted in their opening line that these negative effects had been "largely ignored" for most of a century, and they were largely ignored for three decades afterwards.
What separates the two outcomes is not warmth, timing or delivery. It is information.
The cheapest available fix is also the largest. Butler, Karpicke and Roediger78 tested learners on general knowledge facts and gave feedback on half the questions. Testing without feedback produced 41% recall on the final test. Testing with feedback produced 87%.
The most interesting part is what feedback did to the answers learners had got right but doubted: it doubled their retention. Showing the correct answer does not only fix errors. It tells a learner that the thing they hesitated over was right.
Meanwhile the thing teams usually argue about, whether feedback should be instant, makes almost no difference. Across 51 studies, Kandemir and colleagues79 found the gap between immediate and delayed feedback was 0.03, statistically indistinguishable from nothing. The authors note that few of those studies used delays longer than a day, so this settles the instant-versus-end-of-session argument rather than the question of feedback arriving a week later.
The reason organisations ship the weak version is not ignorance either. High-information feedback takes an expert twenty minutes a question and does not scale. Scores are free. That cost asymmetry, rather than any disagreement about the evidence, is why corporate feedback looks the way it does, and the part technology is now genuinely changing.
Write the explanation once, at authoring time, and attach it to the wrong answer. It costs an author twenty minutes and every learner who picks that option gets it for nothing, for the life of the course. A score attached to nothing is the version that carries the risk of making performance worse.
Four consequences, and none of them require new content.
A quiz at the end of a module measures. The same questions distributed through and after it teach, and showing the correct answer afterwards is worth more than forty points of recall for one line of configuration.
Spacing is the highest-return, lowest-cost change available, and the gap matters far more than the pattern. You do not need an adaptive algorithm to start. You need dates in a calendar.
Every technique in this part looks worse than the alternative on an end-of-course quiz and better a month later. If completion and immediate scores are your only measures, your data will consistently recommend the weaker option.
Timing makes almost no difference, warmth makes almost no difference, and information makes all of it. One well-written explanation of why a wrong answer is wrong outperforms any amount of polish on a score.
What order things come in, who else is in the room, and whether any of it reaches the job. Six borrowed models, and the evidence sitting underneath them.
The models are borrowed frameworks. The evidence sits underneath them, in practices with their own literatures. Where the two come apart, we have followed the evidence.
Parts 1 to 3 dealt with the machinery, the conditions it runs in, and the techniques that work. This part deals with architecture. It is also the part of the field where our industry holds its strongest opinions on the thinnest evidence. Several of the models below are genuinely useful, and almost none of them say what they are commonly quoted as saying. Two of the most-cited numbers in corporate learning appear in this part, and both are wrong by a wide margin, in opposite directions.
A six-way classification of the kinds of thinking a learning objective can ask for. In Anderson and Krathwohl's 2001 revision they are stated as verbs: remember, understand, apply, analyse, evaluate, create. Its job is to make an objective specific enough that somebody could mark it.
Benjamin Bloom and four collaborators published the taxonomy in 1956 to give educators a common language for describing what learners should be able to do.80 Anderson and Krathwohl revised it in 2001, turning the categories into verbs and moving creation to the top.81
As a vocabulary it is excellent. It is why a learning objective can be written precisely enough to assess. "Understand the policy" is unassessable. "Apply the policy to an unfamiliar case" can be marked.
Master each level before the next becomes available. Recall, then understand, then apply, and so on up.
Kinds of cognitive work, not a sequence you climb. Learners analyse before they can recall reliably, and they create while still shaky on the fundamentals.
The useful discipline it imposes is on your own ambition. If a programme's objectives all sit at the bottom two levels, it is a knowledge transfer exercise, and no amount of scenario design will turn it into judgement development. That mismatch between stated objective and actual design is one of the most common faults in corporate learning, and Bloom's taxonomy is the cheapest tool for spotting it.
Write the objectives first, in Bloom's verbs, and read them back. The verbs you have chosen tell you what you are actually building, which is not always what you told the business you were building.
The zone of proximal development is the gap between what a learner can already do on their own and what they can do with help. Below the gap the task is trivial. Above it they cannot get there at all. Inside it, help works.
Scaffolding is that help: temporary support that carries the part of the task the learner cannot yet manage. A worked example, a prompt, a partly completed template, a colleague who has done it before.
Both terms get attributed to Lev Vygotsky. Only the zone is his. He defined the zone briefly, in the last two years of his life.
The distance between the actual developmental level as determined by independent problem solving and the level of potential development as determined through problem solving under adult guidance or in collaboration with more capable peers.Vygotsky, defining the zone of proximal development82
He died in 1934, and the English text almost everyone quotes is Mind in Society, a compilation assembled and translated in 1978. Reviewing the primary texts, Chaiklin identifies three assumptions in the popular version that Vygotsky never made.83
That the zone applies to learning any kind of subject matter. Vygotsky was writing about child development, not skills training.
That the active ingredient is a more competent other. The zone describes a state of readiness, not a delivery method.
That readiness is a fixed property of the learner, and the job is to locate it.
Scaffolding is not his at all. Wood, Bruner and Ross introduced it in 1976, forty years later.84 Gredler describes the zone itself as "a minor discussion" in Vygotsky's work about which "inaccurate information… attracted attention early on and became identified as a major aspect of his theory".85
So the attribution is shaky. The practice is not.
Doo and colleagues went narrower, pooling 64 effect sizes from 18 studies of online higher education covering 4,852 learners.87 The overall effect was 0.87. The breakdown inside it is the useful part.
Fading, the gradual withdrawal of support as competence grows, is treated as definitional. It is also the claim with the least behind it.
Scaffolding must fade. Support that persists creates dependency, so plan its removal from the start.
Fading appeared in only 16.5% of the 333 outcomes, and where it did appear it made no significant difference to the result.86
Read that carefully, because it is not a licence to build permanent crutches. Fading is barely tested, the studies that do test it are mostly intelligent tutoring systems, and absence of evidence at this sample size is not evidence of absence. What it does mean is that anyone claiming fading is the active ingredient is ahead of the data, and that a design budget is better spent on what the scaffold asks the learner to do than on the schedule for taking it away.
Gating progression on demonstrated readiness, which is the zone's core claim, has firmer ground. In the literature this is called mastery learning: the learner stays on a unit, with corrective feedback and re-testing, until they can demonstrate it, and only then moves on. Time varies. The standard does not.
The instruction is not "find the zone". It is: hold learners at a task until they can do it, and make the support ask them to think rather than tell them what to do.
In 1984 Benjamin Bloom reported that students who received one-to-one tutoring combined with mastery learning performed two standard deviations above a conventionally taught control group.89
The average tutored student was above 98% of the students in the control class.Bloom (1984)89
About 90% of the tutored students reached the achievement level of the top 20% of the conventional group. The two figures describe the same result from different ends: shift a group by two standard deviations and roughly nine in ten of them clear what used to be the top fifth. It is the most-quoted number in the case for personalised learning, almost always with the mastery learning half dropped.
Bloom did not drop it. His tutoring condition was "followed periodically by formative tests, feedback-corrective procedures, and parallel formative tests", and the same paper puts mastery learning on its own at a full sigma. That is worth sitting with before the debunking starts: mastery learning, with no tutor attached, moved the average student a full standard deviation. Half the famous headline is a technique you can actually afford. Then the other half met a harder test.
Follow that sequence.7 Two sigma is two interventions, not one, in two experiments run by Bloom's doctoral students. Test the surviving half at scale across ninety-six randomised trials and you get 0.37. Nobody falsified anything. The number shrank each time it was asked a harder question, which is what usually happens.
Which leaves 0.37 as the honest figure, and 0.37 is still very good. It is comfortably better than most things an L&D team can buy, and the strongest evidence in this handbook for individual attention as a design principle.
The problem was never that tutoring does not work. It is that one-to-one attention has never been affordable at organisational scale, and quoting 2.0 to justify a platform purchase invites the first informed reader to dismantle the argument.
Social learning is the observation that people acquire behaviour by watching other people do it. The phrase comes from Albert Bandura, the psychologist whose work in the 1960s and 70s established that a behaviour can be picked up through observation alone, with no direct experience, instruction or reinforcement of the observer's own.90
His more useful contribution for L&D was the distinction he drew afterwards: learning and performance are not the same event. Someone can observe a behaviour, understand it completely, and still not do it. Completion proves exposure and nothing else.
Behaviour modelling training is that theory turned into a method, and it has the strongest workplace evidence in this part.
Watching someone competent reliably teaches the skill. Whether it survives contact with the job depends on conditions the authors specify, and all five are things a programme owner controls.
None of which happens in a climate where admitting ignorance is risky.
Social learning is not a feature you switch on. It is a climate you either have or do not, with a platform component. Audit the climate before you buy the platform.
Metacognition is the highest-leverage structural addition available, and the cheapest.
The mechanism is straightforward. Learners who plan, monitor and evaluate their own learning select better strategies and notice their own gaps. This is why the scaffolds that prompted planning and self-monitoring outscored the ones that gave instructions. Same finding, approached from the other side.
The problem is that the whole approach runs on self-assessment, and self-assessment is systematically unreliable.
It is not confined to the laboratory. Zenger surveyed engineers at two technology firms and found 32% at one and 42% at the other believed their skill placed them in the top 5% of performers at their company.95 Svenson asked people to rank themselves against the others in the room and found 93% of his American sample placing themselves above the median for driving skill.96
The fix is specific, and "more reflection" is not it. Self-assessment needs an external reference point attached to it: a quiz result, a peer review, a scored scenario, a manager's observation. Reflection without data returns the learner's existing self-image with more words around it.
One warning about what does not belong here. Growth mindset is routinely offered as the metacognitive intervention. It does not survive its own evidence base. Part 6 deals with it in full.
Everything in this part exists to serve one outcome, which is whether the capability shows up at work. This handbook opened on our industry's favourite number for this, so the short version will do here. It is fictional, and the real shape is more useful anyway.
Traces to a single 1982 article by David Georgenson, where it appears as an off-the-cuff estimate voiced by an unnamed training director. No study, no method, no data.1
62% still being applied immediately after training, 44% after six months, 34% after a year, across 150 training and development professionals.3
The Saks and Belcourt numbers are practitioner estimates rather than measurements, so hold them loosely. The shape is the point. Transfer does not fail at the door. It decays, on a timescale you can intervene on. Those are different problems with different fixes, and the fictional number points at the wrong one.
As for what drives it, the answer is inconvenient for anyone selling a single solution.
One warning belongs with those numbers, and it concerns a figure the chart deliberately leaves off. Pool every study and the work-environment relationship comes out at .54, which would make it far and away the biggest lever in the set. It more than halves, to .23, once you set aside the studies in which the same trainee rated both their own workplace and their own transfer.
Asking one person two questions on the same form and correlating the answers inflates the result, because somebody who feels warmly towards their employer tends to answer both questions warmly. The smaller figure is the honest one, and it sits in the same range as the transfer-climate estimate above.
The largest single factor is something you select for rather than train. Which makes voluntary participation the most actionable finding in the set, because it sits at .34 and is a decision about how you enrol people rather than how you teach them.
Four consequences follow, and three of them are about what you stop doing.
Bloom's taxonomy writes better objectives. Vygotsky's zone names a real phenomenon. Neither tells you how to build anything. Scaffolding, mastery gating and behaviour modelling do, and each has its own literature with its own numbers.
Holding a learner at a task until they can perform it is worth about half a standard deviation across 108 controlled evaluations. Cohort scheduling optimises for administrative convenience, and the learners who need the programme most are the ones who pay for it.
Scaffolds that prompted learners to plan and monitor their own thinking scored 1.10. Scaffolds that told them what to do next scored 0.39. That is the largest single design difference in this part, and it costs nothing to act on.
Learners' judgement of their own competence is not weakly calibrated, it is confidently wrong in a predictable direction. Reflection prompts without an external reference point make people more articulate about an inaccurate self-image, not more accurate.
Five formats dominate the market. Each has a research base, each is sold as though the format were the active ingredient, and none of them is.
Every format here works by making something from Parts 1 to 3 affordable, and every one disappoints when it is bought as a substitute for that thing rather than a route to it. Name the mechanism, or you are buying a wrapper.
Parts 1 to 4 dealt with the machinery, the conditions it runs in, the techniques that work and the architecture around them. This part deals with the formats all of it actually arrives in. Which gives it a single test to apply to any format decision.
| Format | What it actually delivers | What it becomes without that |
|---|---|---|
| Microlearning | Spacing and retrieval, made schedulable | A library of short videos |
| Multimedia | Cognitive load management, mostly by removal | Production value |
| Gamification | Feedback and progress information | Decoration with a scoreboard |
| Storytelling | Encoding, via a dedicated decoder | Entertainment with a logo on it |
| Habit design | Return visits, which every other row depends on | A reminder nobody opens |
The industry argument about microlearning is about duration, and duration is the least interesting thing about it.
Learning delivered in short, self-contained units rather than in one long session. The industry argues about how short. The more useful question is what the shortness buys you.
Let's start with the honest case for brevity, because it is real and it has nothing to do with the brain. A five-minute unit fits into a working day. A sixty-minute course requires someone to defend an hour in the diary against everything else competing for it, and to defend it again for every repeat. Short units get started, get finished, and get come back to. That is a logistics argument rather than a cognitive one, and quite strong enough on its own.
What it is not is a research finding. The specific numbers the industry trades in have no study behind them.
Ask instead what short units buy you, and the answer is the two techniques that cleared Dunlosky's bar in Part 3. Spacing only works if learners come back, and nobody returns to a sixty-minute course five times. Retrieval only works if the material is small enough to be asked about. Microlearning is the format that makes both of those schedulable.
It works when it is a schedule and it disappoints when it is a library of short videos. The unit of design is the return visit, not the runtime.
The evidence filed under the word "microlearning" is thin: a handful of small studies, mostly with students. The evidence for what microlearning actually delivers is not thin at all, and it comes from working professionals.
The scale is there too. A 2026 review of spaced repetition in medical education pooled 21,415 learners and found an effect of 0.78 against conventional study methods, which puts the average spaced learner ahead of about 78% of those who studied the usual way.100 That is a larger body of evidence than anything published under the microlearning label, by roughly thirty to one.
One substitution to refuse, because a careful reader will catch it. Every result above compares spaced delivery against massed delivery. None of them isolates shortness as the active ingredient. What the evidence supports is distribution over time. Brevity is what makes distribution practical, which makes it a supporting argument rather than the finding itself.
Attention spans have collapsed to under a minute, so content must be tiny.
Time on a single screen before switching: about two and a half minutes in 2003, 47 seconds in her recent work. Others have since found 50 and 44 seconds, median around 40.101 A real, replicated finding about how people move between windows.
It is not evidence that human attention capacity has shrunk, and Part 6 deals with the folklore version of that claim. Treat it as a constraint on the environment your content lands in, not a diagnosis of your learner.
Richard Mayer's cognitive theory of multimedia learning rests on three assumptions. The brain has separate but interacting channels for what it hears as words and what it sees. Each channel has limited capacity. And learning requires the learner to select, organise and integrate, which is effortful. From those, Mayer reports fifteen evidence-based principles drawn from more than two hundred experiments run by his own group.102
Read the fifteen together and a pattern appears. Four of them make it plain:
The other eleven are in Mayer's own summary and most of them run the same way. Multimedia design guidance is largely subtraction, which is exactly why it is so rarely followed. Adding things is visible work. Removing them looks like having done less.
Captioning video for people working in a second language is among the highest-return decisions in the entire multimedia literature, and almost nobody in L&D frames it as a learning intervention rather than an accessibility tick-box.
It matters more than the rest, because it explains why best-practice arguments go in circles.
The asymmetry is your tie-breaker. Helping a novice buys more than withholding help from an expert costs. Where the audience is mixed and segmentation is not available, err towards support.
Applying the mechanics of games (points, levels, leaderboards, streaks, badges) to something that is not a game. Nick Pelling coined the term in 2002 for game-like interfaces on cash machines and vending machines, calling it "the deliberately ugly word 'gamification'".105
It reached learning about a decade later, carrying a set of claims that largely cannot be sourced. We tried to trace the best-known gamification statistics in corporate learning back to a primary study. Most lead to a vendor, a press release, or a number that contradicts itself elsewhere on the same website.
That finding has an obvious reading and a better one. The obvious reading is that gamification works. The better one is that some game elements are doing the work while others are along for the ride, and the ones that reliably do the work are the ones carrying information.
Each of those is feedback with a game mechanic attached, and Part 3 established that the information content of feedback, not its delivery, is what moves performance. A badge with nothing attached to it is the bottom-row feedback that made performance worse in Kluger and DeNisi's data.
Agents began trying to squeeze in a couple more calls to receive more points for the day.Microsoft's own description of the mechanism107
The mechanic changed behaviour reliably. Whether it changed the behaviour you wanted is a design question, and yours to answer rather than the mechanic's. Attach information to every mechanic, and check what the information is actually rewarding.
In 2010 a woman lay in an fMRI scanner at Princeton and told an unrehearsed, fifteen-minute account of something that happened to her as a freshman in high school. Uri Hasson's group recorded her brain activity, played the recording to eleven listeners in the same scanner, and found the listeners reproducing the speaker's patterns a few seconds behind. In some regions the listeners ran ahead, anticipating what was coming. When the same story was played in Russian to non-Russian speakers, the alignment vanished.108
That last detail is what makes the study useful rather than merely striking. Neural coupling is not a response to sound, or to paying attention. It tracks comprehension, and it disappears the moment comprehension fails.
Which raises the obvious question for anyone running a training session: if brains fall into step during a story, does being in step predict who actually learns anything? Davidesco and colleagues built a small classroom in a lab to find out. Groups of four students were taught a short science lesson by a real teacher while everyone wore portable EEG, and were then tested immediately and again a week later.109
Synchrony did predict learning, but not between the people you would expect. Students who were in step with each other scored better on both tests. Synchrony with the teacher predicted only the delayed test, and only when the students' brain activity trailed the teacher's by around 300 milliseconds.
The within-brain counterpart is narrative transportation, which Richard Gerrig named in 1993110 and Green and Brock made measurable in 2000: the state in which someone is absorbed enough in a story that the room recedes.111 Two meta-analyses have asked whether it changes anything.
A scenario is not a story because it has a name in it. It is a story when the learner is inside a situation with a stake in the outcome, and it earns its cognitive cost only when the thing the story encodes is the thing you need them to do.
Narrative that carries the objective is the most efficient format available. Narrative that merely entertains is the subject of Part 6.
Every format above needs the learner to come back, which makes habit design the delivery layer for all of them. It is also where the field's appetite for a clean number is strongest and least satisfied.
Domain matters more than elapsed time. Buyalskaya and colleagues applied machine learning to two enormous panels of objective behaviour: over 12 million observations of gym attendance and over 40 million of hospital handwashing.115
Contrary to the popular belief in a "magic number" of days to develop a habit, we find that it typically takes months to form the habit of going to the gym but weeks to develop the habit of handwashing in the hospital.Buyalskaya et al. (2023)115
Same species, same mechanism, an order of magnitude between the two answers. The number you are looking for is a property of the behaviour, not of habit formation.
What does replicate is the structure. A cue triggers a routine that delivers a reward, and what you are engineering is the transfer of a behaviour from effortful prefrontal control to automatic execution. Two findings tell you where to spend.
Duckworth and colleagues randomly assigned students either to modify their situation or to regulate their own response.117 Across two field experiments, one with high school students and one with undergraduates, the students who removed the temptation met more of their academic goals. In the university study the advantage was partly explained by their reporting less temptation during the week, which is the point: they were not resisting better, they had less to resist.
Translated into a learning programme, the obstacles are rarely motivational. They are the things that stand between a person and the first thirty seconds of the task:
A separate password, a VPN, a system nobody is already in. Put the learning where people already are, or make the link open the content rather than a login screen.
Six clicks to find the module somebody has been asked to complete. The reminder should land on the thing itself.
An expectation to fit it in around the job, which means it competes with the job and loses. A held slot in the calendar removes the daily decision.
A 45-minute module that cannot be paused. If the only available gap is ten minutes, nothing starts.
Willpower lost to rearranging the furniture. Ask for the specifics of when and where, and remove the obstacle rather than asking anyone to out-argue it.
Four consequences, and each one is a decision you make before the build starts.
Every format in this part is a delivery route for something established earlier in this handbook. Microlearning delivers spacing and retrieval. Gamification delivers feedback and progress information. Story delivers encoding. If a format decision cannot be traced to a mechanism, it is a preference dressed as a strategy.
Most of Mayer's principles instruct you to remove something, and the highest-return finding in the largest synthesis available is captioning video for second-language learners. Both are cheap. Neither looks like work, which is why neither gets done.
Points, badges and leaderboards move behaviour whether or not they carry meaning. The ones that improve performance tell the learner something about their performance. The ones that do not are the feedback condition that made people worse.
A motivational message moved 35% of people against a control group's 38%. The same message plus a written statement of when and where moved 91%. Ask for the specifics, and remove the obstacle rather than asking anyone to out-argue it.
Almost none of these ideas is simply false. Each contains a true observation and a false inference, and the inference is the part the industry bought.
The next bad idea will arrive in the same shape as the last one. A finding that holds, an extension that does not, and a number somewhere in the middle that nobody has traced.
The five parts before this one were about what to build. This one is about what to stop paying for. The temptation is to read it as a myths list. It is not.
Preferences are real. People will tell you how they like to learn.
Matching instruction to the stated preference improves learning.
People do remember more from doing something than from being told about it.
So the percentages attached to that ladder, 10% of what we read up to 90% of what we do, describe how much is retained.
Experience develops people. Most of it does happen at work.
The split is 70:20:10, and you should treat it as a target.
Beliefs about capability shape behaviour.
A mindset intervention will move attainment at scale.
Interest matters. People remember more from material that holds their attention than from material that bores them.
So make the material more interesting by adding vivid, memorable extras to it.
Part 3 established that learning styles have no credible evidence behind them. The more useful question is which claim, precisely, the evidence refutes, because the industry's usual defence is to retreat to a version nobody disputes.
Preferences exist. People will tell you, sincerely and consistently, that they prefer diagrams or discussion or doing. No researcher disputes this. What Pashler and colleagues tested is narrower, and they gave it a name.
The most common, but not the only, hypothesis about the instructional relevance of learning styles is the meshing hypothesis, according to which instruction is best provided in a format that matches the preferences of the learner.Pashler, McDaniel, Rohrer & Bjork (2008)118
That is the claim that fails, and the only one anything has ever been built on. Their conclusion after looking for studies designed well enough to test it: "there is no adequate evidence base to justify incorporating learning-styles assessments into general educational practice." Willingham and colleagues put it less diplomatically in 2015. Learning styles theories "have not panned out", and telling students so is a professional obligation.119
One further detail makes the position harder to argue with: the same review counts 71 separate learning styles schemes, and notes their authors did not claim the list was exhaustive. A field with 71 competing classifications and no agreed one is telling you something about its foundations.
Stop defending or attacking "learning styles" in general and name the meshing hypothesis instead. Preferences are worth knowing for engagement and consent. They are not a routing instruction.
Edgar Dale published his Cone of Experience in 1946 as a visual metaphor for how abstract different kinds of learning experience are. It carried no numbers, no percentages and no claim about retention.
People remember 10% of what they read, 20% of what they hear, 90% of what they do. Every version disagrees with the others about the exact split, which is the first clue.
Thalheimer, on the original: Dale "included no numbers in his cone" and "warned his readers not to take the cone too literally".122
In 2014 a group of researchers devoted an entire special issue of Educational Technology to the corruption. Subramony and Molenda's introduction describes "the corrupted cone and its attendant 'data'" as "akin to a living organism, a virtual 21st century plague, that continues to spread and mutate all over the World Wide Web".8 Thalheimer traces the numbers backwards through decades of citation and finds them attached to a chain of secondary sources, with percentages of this shape appearing as early as 1914. At no point does the chain terminate in a study.
The most instructive thing about the Cone is not that it is wrong. It is how the error was made. Dale drew a defensible observation, that experience varies in how concrete it is, and someone downstream converted a shape into arithmetic. Nobody fabricated data in a laboratory. They fabricated precision.
The fix is not to memorise that the percentages are false. It is to notice the shape of the claim. A tidy set of round numbers, arranged in a satisfying gradient and cited to nobody in particular, has almost certainly been made up. Real measurements are untidy, and they come with a sample size attached.
There is a real finding underneath, and it deserves to survive the debunking. Active retrieval beats passive review, and Part 3 gave you the evidence for it. That evidence has effect sizes and named studies. It does not need a pyramid.
The model says development splits 70% experience, 20% social, 10% formal. It is probably the single most quoted number in L&D strategy documents, and it has never been a research finding.
It traces to the Center for Creative Leadership's interview work with executives, published in 1988 by McCall, Lombardo and Morrison as The Lessons of Experience.10 The method was to ask successful managers to recall what had developed them. That is a survey of recollection, not a measurement of development, and recollection is exactly the faculty the rest of this handbook has shown to be unreliable about learning.
The 70% rule of informal learning needs to be set aside.Clardy (2018), after examining five literature traditions123
Alan Clardy's review is the most thorough attempt to establish whether the ratio can be supported. He found the apparent convergence illusory, critiquing the traditions for "sloppy scholarship, inconsistent conceptualizations, and fundamental research protocol problems".
The awkward part for critics is that the underlying observation is sound. Most development does happen at work. Johnson, Blackman and Buick interviewed 145 managers across three phases and found four recurring misconceptions about the model.124
Along with a finding worth keeping: in their words, "the social aspect of the framework is the 'glue' that integrates formal and experiential learning". Experience on its own is exposure. Experience plus someone to make sense of it is development.
So the ratio is not a target. Used as one, it produces the familiar failure: an organisation announces a 70:20:10 strategy, changes nothing about how work is designed, and relabels its existing courses as the 10.
Used as a diagnostic it is genuinely useful. Map where your development budget goes, then map where your people say they actually learned. If both land almost entirely in formal delivery, the model has told you something true about your organisation. That is a question, not a quota.
Part 4 noted in passing that mindset interventions do not carry the weight placed on them. This is the working.
Note what that sentence is actually saying. It is not that mindset does nothing. It is that the literature reporting mindset effects is shaped by which studies got published, and that the better the study, the smaller the effect. That is the signature of an effect that is mostly artefact, and a pattern you can look for anywhere.
Yeager and Dweck have engaged with the criticism rather than dismissing it. Their 2020 review argues that "large-scale studies, including preregistered replications and studies conducted by third parties… justify confidence in growth mindset research", while conceding in the same paper that "mindset effects, however, are meaningfully heterogeneous across individuals and contexts".126
Buy a mindset intervention, deploy it at scale, expect performance to move.
Beliefs about your own capability are real and worth attending to. They are shaped by what an organisation rewards, not by a workshop.
Part 5 distinguished narrative that carries the learning from narrative that merely entertains. This is the research on the second kind.
Garner, Gillingham and White named the effect in 1989.128 Seductive details are additions that are interesting to the learner and irrelevant to the objective: the dramatic photograph, the opening anecdote, the arresting statistic that has nothing to do with the task. Harp and Mayer's four experiments with 357 undergraduates found that learners who read a passage carrying such details "recalled fewer main ideas and generated fewer problem-solving transfer solutions" than those who read it without.129
Sit with the gap between those two numbers, because it is the most useful thing in this part. That gap is common in psychology and it usually means the smaller number is closer to the truth. Applied here, it says seductive details are a steady tax on comprehension rather than a catastrophe. Worth removing, because removal is free. Not worth stripping a module of everything that makes a human being want to open it.
One limitation carries the practical weight. Almost all of this research uses captive participants who cannot leave. Your learners can. A module optimised to the last percentage point for encoding is also a module optimised for abandonment, and the retention rate of a course nobody finishes is zero.
The explanation is not that practitioners are careless. It is that learning is one of the few activities where doing it correctly feels worse than doing it badly.
This is why the evaluation habits of the field entrench the problem. Alliger and colleagues found the correlation between reactions of any type and immediate learning sits at .08, and for affective reactions, whether people enjoyed it, at .02.134 Happy sheets measure enjoyment accurately and learning not at all, and every idea in this part scores well on that instrument.
The same meta-analysis contains the fix. Change the question, not the instrument.
Asking "will this help me do my job?" is worth something. Asking "did you enjoy it?" is worth .02.
One more that keeps coming back, and is not covered anywhere else in this handbook. Nielsen and colleagues analysed resting-state scans from 1,011 people aged 7 to 29 and found that lateralisation is real but local. Their data "are not consistent with a whole-brain phenotype of greater 'left-brained' or greater 'right-brained' network strength across individuals".135 Nobody is left-brained. Particular jobs are lateralised, and people are not.
The common thread is not stupidity. It is that all of them are pleasant, memorable, and confirmed by exactly the evidence a busy professional has to hand.
Four consequences, and each one is a habit rather than a decision.
Almost every idea here is built on something true. Preferences exist, experience develops people, beliefs matter, interest matters. Ask what specific claim the evidence tests, then check whether that is the claim you are buying. It usually is not.
10, 20, 70, 90. Fabricated precision has a look, and the look is a tidy series arranged in a satisfying order with no study at the end of it. Ask who measured it, on whom, and how many.
The pattern repeats across growth mindset, seductive details and most of this literature: large in the lab that found it, small once the field pools its results. When a headline effect size and a meta-analytic one disagree, plan against the meta-analysis.
Reactions of any type correlate with learning at .08, and enjoyment at .02. But whether learners think the training will help them do their job correlates at .26. Learners who learn the most frequently report learning the least, because effective study feels harder.
Six parts, twenty-two thousand words and more than a hundred and thirty sources went into this. Here is what this all amounts to.
Each part makes its own case. Some things are only visible from above: patterns that repeat across all six, and that will repeat in the next claim you are asked to believe.
Four chunks of working memory, and forgetting as the default. No amount of content design changes either one.
Sleep, stress, cognitive ability, whether attendance was voluntary, whether it is safe to admit you do not know.
Retrieval, spacing, interleaving, feedback that carries information, generation.
Bloom writes better objectives. Vygotsky names a real phenomenon. Neither tells you how to build.
Microlearning delivers spacing and retrieval, or it is short video. Name the mechanism or you are buying a wrapper.
Preferences are real, but matching to them does nothing. Experience develops people, but the ratio was never measured.
Four patterns repeat often enough to be worth naming, because the next dubious claim you meet will fit one of them.
The most reliable finding in this handbook is not about learning at all. It is about evidence.
When a headline effect size and a meta-analytic one disagree, plan against the meta-analysis. With only the headline, assume you are at the top of a curve nobody has tested yet.
Nine in ten of Kornell's participants learned more from spacing. Seven in ten believed cramming had worked better. Not a calibration error but an inversion, and it explains why our instruments miss it: reactions correlate with learning at .08, enjoyment at .02, and the single-day intensive rates well precisely because fluency peaks when the feedback form appears.
One exception beats the rest of the evaluation apparatus combined. Ask whether learners think the training will help them do their job and the correlation is .26. Change the question, not the instrument.
This one surprised us, and it runs through every part.
Even the large exception proves the point: the highest-return finding in the largest multimedia synthesis available is captioning video for second-language learners, which means attaching a file that usually already exists.
Removal is cheap, fast and invisible. Nobody gets promoted for it, which is exactly why it stays available.
L&D does more sensible things than its evidence base would suggest. Spaced reinforcement, practice, coaching, peer learning: the instincts are broadly right. The arithmetic is not. Nearly every headline figure in the field has no source, has one that says something else, or has one since contradicted. That is a reason to stop quoting the number, not to distrust the practice. The number is what an informed reader will attack.
Ordered by what it costs you, cheapest first. Every one traces to a mechanism rather than a preference.
Novelty, unpredictability, threat to ego, low control. Take them out of your next programme.
Part 2The highest-return finding in the multimedia literature, and the file usually already exists.
Part 5The stock photography, and the narration that reads the on-screen text aloud.
Parts 1 & 5Between input and application. Kill the two-day intensive that runs straight through.
Part 2Same content, same slot in the calendar. The learners produce it instead of reading it.
Part 3From "did you enjoy it?" at .02 to "will this help you do your job?" at .26.
Part 6Wherever you can. It is the largest actionable predictor of transfer in the literature.
Part 4Hold learners at a task until they can do it. Worth about half a standard deviation.
Part 4Prompts that ask a learner to plan scored 1.10. Prompts that tell them what to do scored 0.39.
Part 4Trace every number in your own materials to a primary source, and delete the ones that do not survive it. This handbook cost us more than forty claims. It was worth it.
The specific myths in Part 6 will be replaced. The shape will not.
Five questions, in the order that disposes of a bad claim fastest. Most fail at the first two.
A vendor page, a conference slide or another article is not the source. Keep going until you reach a method section, or stop using the number.
Almost never the claim being sold to you. Preferences exist, but meshing was the hypothesis tested. Tutoring worked, but mastery learning was bundled with it.
It usually is. Use the pooled figure as your planning number.
Undergraduates who cannot walk out are not your workforce. Almost the entire seductive-details literature rests on captive samples.
10, 20, 70, 90. Fabricated precision has a look: a tidy series in a satisfying order with no study at the end of it.
It would be dishonest to end a handbook about unsourced numbers without saying how many of ours did not survive it. More than forty claims in our own drafts were corrected as we wrote. We had a scaffolding finding exactly backwards, load-bearing enough to have produced a design recommendation. We attributed an imaging result to the wrong paper, described a lab-simulated classroom as a live one, and deleted a tutoring effect size we could not trace.
Every one came from a reputable secondary source, several from our own published articles. That is the point. The distance between "widely repeated by serious people" and "true" is wider in this field than in most, and the only way to close it is to open the paper.
Not that you believe this handbook. That you can check it. Every figure in all six parts carries its citation and the sample behind it, which means every one is available to be argued with.
Neurogogy is not a new science. It is the discipline of designing for the brain people actually have, rather than the one the industry finds convenient: small working memory, default forgetting, a state that shifts with sleep and stress, and instincts about its own learning that run almost exactly backwards.
None of the six parts asks you to take anything on trust. That was the whole design.
It is also how we build. L'Oréal Travel Retail runs its beauty advisor programme on exactly the mechanisms this handbook argues for: short units that make spacing and retrieval schedulable, a social layer that gives people a reason to come back, and game mechanics that carry information rather than decoration.
Every mechanism in this handbook ends up as a decision somebody makes on a Tuesday: how long, how often, in what order, reinforced how. The Impact Suite runs spacing, retrieval and social reinforcement as defaults rather than as things you remember to schedule. If you would rather talk it through than read another page about it, that works too.
135 sources, numbered in the order they first appear. Each entry gives the full citation, the sample behind the figure, and a link to the source. Every one was read against the primary source.