
Feedback is the most trusted intervention in learning and development. It sits in every competency framework, every manager training programme, and every review cycle. The assumption underneath all of this is that more feedback is better than less.
In 1996, Avraham Kluger and Angelo DeNisi tested that theory. They pooled 607 effect sizes from 131 studies, covering 12,652 people. It remains the largest analysis of feedback and performance ever conducted.
On average, feedback worked: d = 0.41. That’s a respectable effect size, and exactly what the industry expected.
Then there’s the finding that should have changed everything and didn’t. More than 38% of those interventions made performance worse. That’s right, worse. Kluger and DeNisi noted in their opening line that these negative effects had been “largely ignored” since the beginning of the century.
Feedback is not a mild intervention that helps a little. It’s a powerful one that cuts both ways, and which way it cuts is a matter of design, not luck.
What Feedback Actually Is
Writing in 1989, Royce Sadler adopted a definition that turns on a single clause:
“Feedback is information about the gap between the actual level and the reference level of a system parameter which is used to alter the gap in some way.”
That may sound unnecessarily technical, but the key part is “which is used to alter the gap“. On that definition, information that changes nothing isn’t weak feedback. It isn’t feedback at all.
Sadler had a name for it. Information that is “simply recorded, passed to a third party who lacks either the knowledge or the power to change the outcome, or is too deeply coded (for example, as a summary grade)” leaves the loop open, so “‘dangling data‘ substitute for effective feedback.”
A learner who scores 68% on an assessment knows they scored 68%. They don’t know which answers were wrong, why, or what to do next. The number is too deeply coded to act on.
So the first test isn’t whether feedback was kind, timely, or well received. It’s whether the person receiving it could do something differently as a result.

Feedback Models Worth Knowing
The literature has produced no shortage of models. Lipnevich and Panadero reviewed fourteen of them, spanning 1983 to 2018, and found the field moving decisively in one direction: away from feedback as something delivered, towards feedback as something the learner does.
As they put it: “All feedback that comes from any external source will have to be internalised and converted into self- or inner feedback.” Here’s the breakdown:
| Model | Year | Core Idea |
|---|---|---|
| Ramaprasad | 1983 | Feedback is information about a gap that is used to close the gap. |
| Sadler | 1989 | Information nobody can act on is dangling data, not feedback. |
| Butler & Winne | 1995 | Feedback acts through the learner’s own self-regulation. Its effect depends on how it changes task engagement. |
| Kluger & DeNisi | 1996 | Feedback works by moving attention. Moving it to the self can lower performance. |
| Nicol & Macfarlane-Dick | 2006 | Learners are already generating their own feedback. Good practice builds on this instead of replacing it. |
| Hattie & Timperley | 2007 | Feedback operates at four levels, and the self level is the weakest. |
| Carless & Boud | 2018 | Receiving feedback well is itself a skill. |
Thirty-five years of work, all converging on the same conclusion: feedback is only feedback once the learner does something with it.
What Kind of Feedback Actually Works?
If a third of feedback makes things worse, the useful question isn’t how much to give. It’s what has to be in it.
Wisniewski, Zierer and Hattie pooled 32 meta-analyses covering 435 studies and more than 61,000 participants. Rather than sorting feedback by who delivered it or when, they sorted it by how much information it carried.
| Type | What The Learner Receives | Effect Size |
|---|---|---|
| High-information | What was wrong, why, and what to do differently | d = 0.99 |
| Corrective | Which answers were wrong | d = 0.46 |
| Reinforcement or punishment | A score, a pass, a “well done” | d = 0.24 |
Their conclusion: “feedback is more effective the more information it contains”. Two further findings sharpen it:
- Across the whole dataset, 17% of feedback effects were negative.
- And 86% of the negative effects on motivation came from reinforcement or punishment.
So the harm isn’t randomly distributed. It concentrates in precisely the kind of feedback that is cheapest to automate and most common in corporate learning. A completion percentage is bottom-row feedback. So, unfortunately, is a badge.
Overall, feedback averages d = 0.48, with 83% heterogeneity. The authors were clear about what this implies. Feedback “cannot be understood as a single consistent form of treatment”. It’s a category that contains both your strongest available intervention and one of your weakest.
What Good Feedback Actually Looks Like
We know that good feedback is information-loaded and actionable, but what does that actually look like? Here’s one question from a data protection module, answered wrong, with three possible responses.

What Most Platforms Give:
“You scored 4 out of 5. 80%. Pass.”
Nothing to act on. Bottom row of the table, d = 0.24.

The Corrective Version:
“Question 3 was incorrect. The correct answer is B: verify the requester’s identity before acting on the request.”
That’s better. The learner now knows the right answer. Middle row, d = 0.46.

The High-Information Version:
“Question 3 was incorrect. You chose “delete the records immediately”, which is exactly the instinct the regulation is written to interrupt. Acting on an unverified request is itself a breach. Verify identity first, then act within the statutory deadline.”
Top row, d = 0.99. It names the error, explains why the reasoning failed rather than just which box was wrong, and gives the learner something transferable instead of one corrected answer.
Now the part nobody says out loud. The first version costs nothing and scales easily to thousands of learners. The third takes a subject matter expert 20 minutes per question and doesn’t scale at all. Multiply that across a compliance suite and you have a number no L&D budget has ever contained.
Organisations aren’t choosing the weak version of feedback out of ignorance. They’re simply choosing the most feasible approach. That’s the real reason feedback in corporate learning looks the way it does, and the real problem worth solving.
Does Feedback Transfer To Work?
Most of the research above involves students. Corporate learners are adults with jobs and mortgages, so the stakes are different.
The largest workplace evidence we have is Zyberaj’s 2026 meta-analysis, pooling 24 studies across 595,950 employees. It measured how strongly different characteristics of a manager’s feedback relate to whether employees actually take it on board, on a scale where 0 means no relationship and 1 would mean a perfect one.
Across everything, the average came in at 0.36, a moderate link. Three characteristics scored higher.
- Quality of feedback: 0.55
- Credibility of the person giving it: 0.56
- Whether it was positive or negative: 0.50
Anything above 0.5 in research like this counts as a strong relationship, so those three aren’t just top of the list. They sit in a different band from the average. For corporate learners to take feedback on board, it needs to be information-rich and it needs to come from a source they find credible.
Worth noting what this measures: whether feedback gets taken seriously, not whether it improves performance. On that second question, Kluger and DeNisi is still the best evidence we have.
The Confidence Trap
One finding reverses how most organisations decide who needs help.
Butterfield and Metcalfe tracked what happened to errors after correction, sorted by how confident people had been when they made them. Errors made with low confidence were corrected 6% of the time. Errors made with high confidence were corrected 89% of the time.
This means the learner who confidently gets it wrong isn’t your hardest case. The most difficult fixes are the hesitant, low-confidence errors that quietly survive correction and come back.
Why Praise Backfires
Praise sits in the weakest category of feedback there is, sharing it with punishment. Nobody defends punishment. But praise is close to universal, and it fails for reasons that have nothing to do with kindness.
Hattie and Timperley help to explain this. They distinguish feedback by where it’s aimed:
- The Task: Whether the work was right
- The Process: The approach that produced it
- Self-Regulation: The learner’s ability to judge their own work
- The Self: The person doing it
Their verdict is that feedback directed at the self is the least effective of the four, because it carries almost no information about the task. Claims like “You’re a natural” give a learner nothing to act on. It’s the same defect as a completion percentage, wrapped up in a compliment.
Kluger and DeNisi supply the mechanism for this. Feedback works by moving attention, and attention is finite. Feedback aimed at the self spends it on self-evaluation instead of on the work. As they note: “attention to meta-task goals may lead to disengagement from the task even when the [feedback] is positive.”
In other words, being told you’re brilliant can move you further from the task than being told nothing at all.
What Happens When You Praise Ability
Mueller and Dweck tested this on children across six experiments and got the clearest results in the field.
In the first experiment, after a run of failures, children who had been praised for intelligence solved 0.92 fewer problems than they had before, a drop of around 18%. Those praised for effort solved 1.21 more, a gain of around 23%. Same failures, different responses.
A later experiment in the same paper produced the finding that should end this practice. Eighty-eight children were asked to report their scores to children they had never met. Among those praised for ability, 38% lied about their scores. Among those praised for effort it was 13%, and in the control group, 14%.
In other words, praising ability doesn’t just dent persistence. It gives children something to defend, and they defend it by lying.
Which leaves the industry’s most automated form of feedback as also its least effective. Of course, that doesn’t make recognition worthless. It makes recognition a different thing from feedback, and worth keeping separate from the part that actually teaches.
Feedback on Quizzes and Assessments
Most assessments in corporate learning do one job: recording a score. By Sadler’s test this isn’t feedback. After all, a score tells a learner where they landed and nothing about what to do next.
Butler, Karpicke and Roediger ran an experiment to see what’s being left on the table. Undergraduates read twelve prose passages, sat a 40-question multiple-choice test on them, and then took a final recall test either a day or a week later.
Three things were varied: whether feedback was given at all, what kind it was, and when it arrived. Here are the results, ranked by how much each one mattered.
- Whether The Learner Was Shown the Right Answer: Large impact. Final recall came in at 70% with feedback against 51% without, tested a day later. In a second experiment testing a week later, 65% against 42%.
- When It Arrived: Small impact. Delayed feedback beat immediate by 70% to 60% at one week, and at one day the difference wasn’t statistically significant.
- What Kind It Was: No impact. Standard feedback and answer-until-correct feedback produced identical results: 71% against 71%, then 64% against 65%.
What The Quiz Data Tells Us
It’s worth being precise about what counted as feedback here. Learners were shown the correct answer, nothing more. This rarely costs more effort and is worth nearly twenty points of recall.
This data is important. Teams rewrite feedback copy again and again and agonise over whether it should appear instantly. Meanwhile, the decision that actually moves retention is whether the quiz says anything beyond a percentage.
There’s also a compounding effect worth noticing. Retrieving information strengthens memory on its own, which is why retrieval practice works so well. Feedback doesn’t replace that mechanism, it multiplies it. The test does the encoding and the feedback corrects what the test got wrong.
A quiz that returns only a score delivers only the first half of that process.
The Timing Myth
Ask an L&D team what makes feedback effective and “instant” will come up in the first three answers. Most platforms treat it as a virtue in itself. However, the evidence doesn’t support it.
Kandemir and colleagues published the first meta-analysis in 38 years to compare immediate and delayed feedback directly, across 51 studies. The difference was g = 0.03, with a confidence interval running from -0.08 to 0.13. Their conclusion: “feedback timing does not significantly influence learning outcomes on average”.
The previous direct comparison was Kulik and Kulik in 1988. The field then spent nearly four decades arguing about timing without testing it, and when someone finally did, neither side won.
Kulik and Kulik also explain why both camps can cite evidence. Across 53 studies they found that “applied studies using actual classroom quizzes and real learning materials found immediate rather than delayed feedback to be more effective; experimental studies of acquisition of test content indicate the opposite.”
Choose your studies and you can prove either case. Where timing does matter, it’s conditional. Indeed, Shute’s review inverts most people’s intuition:
| Use Immediate Feedback When: | Use Delayed Feedback When: |
|---|---|
| The task is hard relative to the learner | The task is relatively simple |
| The learner is struggling or low-achieving | The learner is already high-achieving |
| You want retention of procedural or conceptual knowledge | You want transfer to new situations |
In other words, immediate feedback earns its place when people are out of their depth. Delay is a luxury for learners who can already cope.
This is the reverse of how most programmes are built. Instant feedback gets switched on uniformly, as a platform feature, when the evidence suggests using a targeted intervention for the people finding it hardest.
Feedback, Technology, and AI

Digital delivery doesn’t weaken any of this.
Adıgüzel and colleagues pooled 30 studies of feedback in online learning and found an overall effect of g = 0.93 (comfortably large). Split by outcome, it was 1.24 for knowledge and understanding against 0.28 for how learners felt about the course. Feedback teaches far more than it reassures.
Now AI. There’s less evidence than the category deserves, but still plenty to discuss:
- Firstly, It Works: Huang and colleagues pooled 36 studies of AI-powered feedback from 2023 to 2025, covering 4,538 learners. They found a 0.61 effect size on learning outcomes (moderate to large).
- But Not Everywhere: Split by teaching approach: 0.71 in collaborative learning, 0.68 in self-directed, and -0.27 in direct instruction. Bolted onto a linear course, AI feedback stopped helping. After all, you’re simply adding an explanation on top of an explanation, effectively overloading your learners.
- It’s About As Good As a Person: Kaliisa and colleagues set AI against human feedback across 41 studies and 4,813 students. AI edged ahead by a small margin (0.25), but with enough variation between the studies that the two are, statistically, a tie.
Which brings us back to the cost problem. High-information feedback works well, but is expensive to produce, so organisations default to the cheap version: a score, a badge, a well done.
That’s the gap Zavmo is built for, and it isn’t about replacing the conversation. It’s about making the informative version affordable at the scale a workforce needs.
How to Design Feedback That Works
Most feedback advice is about delivery: how to phrase it, how to soften it, when to send it. Very little of it is ranked by how much difference each decision makes. Well, these five tips are ranked in order of effect.
That means you can stop wherever your budget runs out and still have the biggest wins first.
- Never Return A Score On Its Own: Showing which answers were wrong is the cheapest large win available. It moved recall from 51% to 70% in the quiz research, and it costs one line of configuration.
- Explain The Error, Not Just The Answer: Corrective feedback returns an effect size of 0.46. Feedback that explains why the reasoning failed returns 0.99. That’s roughly double the effect and the only extra input is somebody’s thinking, written down once.
- Aim At The Work, Not The Person: “You’re doing really well” is the weakest type of feedback. “You’ve applied the rule before checking whether it applied” is among the strongest. The difference is in the amount of actionable information contained in the sentence.
- Stop Optimising For Timing: Across 51 studies timing made no measurable difference on average. However, there is one exception. Bring feedback forward for learners who are struggling and for material that is genuinely hard for them. Everyone else can wait.
- Chase The Confident Errors First: They get corrected 89% of the time, against 6% for hesitant ones. If your platform captures confidence alongside answers, you already have a remediation list most organisations never think to build.
Final Words
Kluger and DeNisi found that over a third (38%) of feedback interventions made performance worse back in 1996 and it changed almost nothing. Three decades later, most organisations still measure feedback by how much of it they deliver, and most learners still just receive a number.
The gap between those two facts isn’t a knowledge problem. Every finding in this article has been sitting in the literature for years. The gap is that informative feedback has always been expensive to produce, while scores are easy to calculate and automate.
Technology is changing that calculation by taking one expert’s explanation and making it reach 50,000 learners instead of 50. So the only thing left to decide is whether you spend your hard-earned budget on explanations, or more badges.
Thanks for reading. If you’ve enjoyed this content, please connect with me here or find more articles here.
Feedback is one piece of a much larger picture. For the full story, download ‘Your Guide to the Science of Learner Engagement‘ now.