Research 09: What Effective Human Teaching and Tutoring Do, and What an App Can Take From Them
Prepared 2026-09-27 for Numberkit. Scope: classroom and tutoring research across ages (preschool to undergraduate), not educational games. Every numerical claim below was checked against the source abstract, the full text, or a reputable index of the source during this review unless it is marked "[unverified]". Where a claim comes from my prior knowledge of the literature and could not be re-checked in this session (the session's search budget ran out), it is marked "[unverified]" and should be checked before it is quoted in public material.
Summary
- Tutoring works, but far less than "2 sigma", and effects shrink at scale. Across 96 RCTs the pooled effect is 0.37 SD, larger with trained tutors, during the school day, and at three or more sessions a week; once-weekly tutoring rarely works (Nickow et al.). Programs resembling large deployments get a third to a half of that (Kraft et al., 282 RCTs). About half of Bloom's 2 sigma was test, correct, re-test (von Hippel 2024).
- Most of tutoring's mechanism is mechanisable. Step-based computer tutors nearly match human tutors (VanLehn 2011: 0.76 vs 0.79); the benefit is step-level action, feedback, and targeted help. Learning tracks what the student constructs more than what the tutor says (Chi et al. 2001).
- What software cannot do is the relational and open-ended diagnostic part: reading affect, rebuilding confidence, hearing a child's own reasoning, a relationship over months (Lepper & Woolverton 2002). The parent is Numberkit's human, and the evidence is double-edged: a light, structured, positive routine helps (Berkowitz et al. 2015); anxious parents who help a lot with homework can harm (Maloney et al. 2015).
- Explicit instruction beats minimal guidance for novices; guided discovery with elicited explanation can beat both (Alfieri et al. 2011: -0.38 and +0.30). Direct Instruction curricula average d = 0.60 (Stockard et al. 2018). "I do, we do, you do" is a reasonable shape with little direct evidence of its own.
- Feedback and formative assessment help on average and often backfire. Feedback d = 0.48, highly heterogeneous (Wisniewski et al. 2020); over a third of feedback interventions lowered performance, especially self-directed feedback (Kluger & DeNisi 1996). Formative assessment is about 0.20, 0.17 in maths (Kingston & Nash 2011).
- Errors handled well are learning events; confident errors, once corrected, are remembered best, also in 8- to 12-year-olds (Metcalfe & Finn 2012).
- Explaining to someone produces durable learning; the social frame alone mainly adds effort (Chase et al. 2009; Fiorella & Mayer 2014). A parent is the right audience.
- Lab effects shrink in classrooms, mostly through delivery drift and broad outcome tests (EEF SMART Spaces null, 2023; mean large-RCT effect 0.06 SD, Lortie-Forgues & Inglis 2019), while classroom retrieval practice holds up (Agarwal et al. 2021). Software delivers schedules faithfully, which is where teachers' versions drift.
Detailed findings
1. Tutoring effects and why they work
Nickow, Oreopoulos & Quan (2020 working paper; published 2024 in American Educational Research Journal as "The Promise of Tutoring for PreK-12 Learning", DOI 10.3102/00028312231208687). A systematic review and meta-analysis of 96 randomized trials of tutoring (one-to-one or small group, by teachers, paraprofessionals, volunteers, or parents), preK to grade 12. Pooled effect 0.37 SD. Effects are larger for teacher and paraprofessional tutors than for volunteer and parent tutors; strongest in the earliest grades for reading, while maths tutoring effects if anything rise from kindergarten to grades 2 to 5; during-school programs outperform after-school ones; effects rise with sessions per week, and "there is little evidence of once-weekly tutoring sessions generating large effect sizes." Parent-tutoring studies were few and small; the authors note that 0.20 to 0.25 SD can still be worthwhile at parent-level cost. Studies at low risk of bias gave nearly the same pooled estimate. Ages: roughly 4 to 17.
Kraft, Schueler & Falken (EdWorkingPaper 24-1031, 2024; Review of Educational Research 2026, DOI 10.3102/00346543261446660). 282 RCTs. Studies that resemble large U.S. programs aimed at standardized tests show pooled effects "a third to a half" of the full sample, driven by "stark declines in pooled effect sizes as tutoring program scale increases." They identify a bundle of design features that partly protects against the decline [specific list not re-checked here].
High-dosage tutoring in secondary maths. Guryan, Ludwig, Bhatt, Cook, Davis, Dodge, Farkas, Fryer, Mayer, Pollack & Steinberg (2023), "Not Too Late: Improving Academic Outcomes among Adolescents", American Economic Review 113(3), 738-765, DOI 10.1257/aer.20210434. Two large RCTs of Saga Education's daily, two-students-per-tutor, in-school maths tutoring for Chicago 9th and 10th graders: participation raised maths test scores by 0.18 to 0.40 SD and raised maths and non-maths grades. Fryer (2014), "Injecting Charter School Best Practices into Traditional Public Schools: Evidence from Field Experiments", Quarterly Journal of Economics 129(3), 1355-1407: a bundle including high-dosage tutoring raised maths by about 0.21 SD per year in Houston; the tutored grades (4th, 6th, 9th maths) showed the largest gains (about 0.4 SD in comparisons reported in the paper). These are adolescents; the key features were daily sessions, a tiny ratio, paid full-time tutors, a set curriculum, and scheduling during the school day.
Bloom's 2 sigma revisited. Bloom (1984), "The 2 Sigma Problem", Educational Researcher 13(6), 4-16 [volume and pages unverified]. Von Hippel (2024), "Two-Sigma Tutoring: Separating Science Fiction from Science Fact", Education Next 24(2): the evidence was two doctoral dissertations (Anania; Burke) with 4th, 5th, and 8th graders learning probability or cartography for three weeks; tests were narrow and curriculum-aligned (tutoring effects in Cohen, Kulik & Kulik's meta-analysis average 0.84 on narrow tests vs 0.27 on broad ones); tutored students retook quizzes with corrective feedback until they passed, which von Hippel estimates accounts for about 1.1 of the 2 SD; and tutoring replaced rather than added to class instruction. The transferable half of Bloom's effect is mastery learning, which software does well.
VanLehn (2011), "The Relative Effectiveness of Human Tutoring, Intelligent Tutoring Systems, and Other Tutoring Systems", Educational Psychologist 46(4), 197-221, DOI 10.1080/00461520.2011.611369. Reviewing experiments against no-tutoring controls: human tutoring d = 0.79, step-based intelligent tutors d = 0.76, answer-based systems much lower. The "interaction plateau": once a tutor interacts at the level of each step, finer granularity (substeps, free dialogue) adds little.
What expert tutors actually do.
- Lepper & Woolverton (2002), "The Wisdom of Practice: Lessons Learned from the Study of Highly Effective Tutors", in J. Aronson (Ed.), Improving Academic Achievement, Academic Press, 135-158. Studied experienced primary and secondary maths tutors whose students gained most. Summarised as INSPIRE: Intelligent, Nurturant, Socratic, Progressive, Indirect, Reflective, Encouraging. Their tutors attended to motivation and affect as much as cognition, corrected errors indirectly (a question, a hint) rather than bluntly, let students articulate their reasoning, and moved from easy to hard. Caveat: an observational, small-sample book chapter, not an experiment.
- Chi, Siler, Jeong, Yamauchi & Hausmann (2001), "Learning from Human Tutoring", Cognitive Science 25, 471-533, DOI 10.1207/s15516709cog2504_1. Tested tutor-centred, student-centred, and interactive hypotheses in college physiology tutoring. When tutors were stopped from explaining and giving feedback and had to prompt instead, students learned about as well, because they did more constructive work themselves. Learning is better predicted by what students construct and by genuine interaction than by the tutor's explanations.
- Graesser, Person & Magliano (1995), "Collaborative Dialogue Patterns in Naturalistic One-to-One Tutoring", Applied Cognitive Psychology 9(6), 495-522, DOI 10.1002/acp.2350090604. Graduate students tutoring research methods and high-school students tutoring 7th-grade algebra. Untrained tutors rarely used sophisticated strategies; tutoring followed a recurring frame (tutor asks, student answers, tutor gives feedback, the two improve the answer together, tutor checks understanding) [the exact five-step wording is from my knowledge of the paper and was not re-checked]. The benefit came from the collaborative, step-by-step repair of answers, not from pedagogical sophistication.
- Lehman, D'Mello, Cade & Person (2012), "How Do They Do It? Investigating Dialogue Moves within Dialogue Modes in Expert Human Tutoring", ITS 2012, Springer LNCS: expert sessions are organised into modes (lecture, scaffolding, modelling) with characteristic moves; this informed Guru, a biology tutor built on brief interactive lectures and rounds of scaffolding. A 2015 Instructional Science paper, "Expertise Amiss" (DOI 10.1007/s11251-015-9363-8) [authors unverified], reports that expert tutors are often less interactive than is ideal.
- Wang, Ribeiro, Robinson, Loeb & Demszky (2024), "Tutor CoPilot: A Human-AI Approach for Scaling Real-Time Expertise", EdWorkingPaper 24-1054. RCT with 900 tutors and 1,800 K-12 students: tutors given AI suggestions modelled on expert moves raised topic mastery by 4 percentage points, and by 9 points for lower-rated tutors. Relevant because it shows the expert-move repertoire can be written down and handed to a less expert human (here, it could be a parent), with the AI advising the adult, not talking to the child.
Can software do this faithfully?
- Frequent, step-level practice with immediate feedback and targeted help: yes; this is the part of tutoring VanLehn shows machines match.
- Test, correct, re-test until mastery (half of Bloom's effect): yes.
- Dosage (three or more short sessions a week, during a set routine): partly; software can schedule and remind, but a person makes it happen. Parent role: agree a fixed slot with the child.
- Nurturant, indirect, affect-reading moves: partly at best. Software can phrase corrections indirectly and never shame, but cannot see a child's face or know that today is a bad day. Parent role: the parent reads the mood and decides whether to start, stop, or just talk.
- Relationship over months: no. Needs a person.
2. Chi's ICAP framework
Chi & Wylie (2014), "The ICAP Framework: Linking Cognitive Engagement to Active Learning Outcomes", Educational Psychologist 49(4), 219-243, DOI 10.1080/00461520.2014.965823; building on Chi (2009), Topics in Cognitive Science [volume unverified]. ICAP classifies overt learning behaviour as Passive (receiving), Active (manipulating, e.g. highlighting, choosing), Constructive (generating something beyond the material, e.g. explaining, drawing, predicting), or Interactive (co-constructing with a partner who contributes substantively), and predicts I > C > A > P.
Evidence for the ordering. Chi & Wylie review many lab and classroom comparisons consistent with the ordering. Menekse, Stump, Krause & Chi (2013), Journal of Engineering Education [details unverified], found the predicted ordering on engineering materials-science concepts. Wiggins et al. (2017), "The ICAP Active Learning Framework Predicts the Learning Gains Observed in Intensely Active Classroom Experiences", AERA Open, DOI 10.1177/2332858417708567: in an undergraduate biology course, interactive sessions outperformed constructive ones.
Evidence against or complicating it. "Questioning Central Assumptions of the ICAP Framework", npj Science of Learning (2023), DOI 10.1038/s41539-023-00197-4 [authors not verified in this review], argues that overt behaviour is a poor proxy for cognitive engagement (a student listening to a good explanation can be thinking hard; a student "constructing" can be copying), and cites cases where direct instruction (overtly passive) beat carefully designed self-learning materials (overtly constructive) for sixth-grade algebra. The ordering also ignores prior knowledge: for novices, studying a worked example (passive by ICAP's label) often beats solving (constructive), which is the worked-example effect from cognitive load research.
Takeaway. ICAP is a useful design heuristic ("make the child generate, not just select") but not a law. It applies best once the child has enough knowledge to generate something meaningful.
Can software do this faithfully? Passive and active: yes. Constructive: partly; software can require the child to build a representation (place counters, drag a rectangle, type the next step, choose which line of a worked example is wrong), which is constructive in a checkable way, but it cannot evaluate a free-form verbal explanation without AI, which Numberkit's rules keep away from the child's live session. Interactive: no in the true sense; a scripted character does not contribute new ideas. Parent role: the "explain it to me" moment (see section 6).
3. Explicit and direct instruction
Rosenshine (2012), "Principles of Instruction: Research-Based Strategies That All Teachers Should Know", American Educator 36(1), 12-19, 39. Ten principles drawn from cognitive science, studies of master teachers, and research on cognitive supports: daily review; new material in small steps with practice after each; many questions, checking every student; models and worked examples; guided practice; checking for understanding; a high success rate (Rosenshine cites about 80% in guided practice [the figure is from my knowledge of the article, unverified]); scaffolds for difficult tasks; independent practice; weekly and monthly review. Its base is mostly correlational 1970s-80s studies of effective teachers plus experimental cognitive science: a synthesis, not a tested package.
Project Follow Through (1968-1977) and Stockard et al. (2018). Follow Through compared about 20 sponsored models across roughly 352,000 children in 178 projects; the Abt Associates evaluation (Stebbins et al., 1977) found Engelmann's Direct Instruction model gave the highest gains, including on basic skills and on some cognitive and affective measures. Critics (House et al., 1978) noted non-random assignment, weak instruments, and more variation within models than between them. Stockard, Wood, Coughlin & Khoury (2018), "The Effectiveness of Direct Instruction Curricula: A Meta-Analysis of a Half Century of Research", Review of Educational Research 88(4), 479-507, DOI 10.3102/0034654317751919: 328 studies, 413 designs, nearly 4,000 effects, 1966-2016; overall d = 0.60; all effects positive and significant except affective outcomes. DI is a specific scripted curriculum, not "any teacher talk".
Kirschner, Sweller & Clark (2006), "Why Minimal Guidance During Instruction Does Not Work", Educational Psychologist 41(2), 75-86, DOI 10.1207/s15326985ep4102_1. Argues from cognitive architecture (limited working memory for novices) that unguided discovery is inefficient and that guidance should fade only as expertise grows. Response: Hmelo-Silver, Duncan & Chinn (2007), Educational Psychologist 42(2), 99-107: problem-based and inquiry learning as practised are heavily scaffolded and should not be lumped with pure discovery. The meta-analytic resolution is Alfieri, Brooks, Aldrich & Tenenbaum (2011), "Does Discovery-Based Instruction Enhance Learning?", Journal of Educational Psychology 103(1), 1-18: 164 studies; unassisted discovery vs explicit instruction d = -0.38 (580 comparisons); enhanced discovery (with feedback, worked examples, scaffolding, or elicited explanation) vs other instruction d = +0.30 (360 comparisons). So: pure discovery loses; guided discovery with elicited explanation can win.
Gradual release ("I do, we do, you do"). Pearson & Gallagher (1983) [details unverified] proposed modelling, guided practice, and independent application for reading comprehension. I found no experimental test of the gradual-release sequence as such; its support is indirect (worked examples, guided practice, and fading are each well supported). Treat it as a design shape consistent with the evidence, not as a separately validated method.
Can software do this faithfully?
- Small steps, worked examples, guided practice with feedback, fading, high success rate, cumulative review: yes. Software executes these more consistently than most teachers.
- Checking every student's response: yes, for one learner it is automatic.
- Deciding when to re-explain in a new way because the child's face shows confusion: partly; software can branch on errors and latencies, and should offer a second representation, but has a fixed repertoire.
- Parent role: minimal here. The parent should not re-teach a method differently from the app's (see section 7 on anxious help).
4. Formative assessment and feedback
Black & Wiliam (1998), "Assessment and Classroom Learning", Assessment in Education 5(1), 7-74, DOI 10.1080/0969595980050102. A narrative review of about 250 sources [count from my knowledge, unverified] concluding that strengthening classroom formative assessment produces substantial gains, often summarised as 0.4 to 0.7 SD. Wiliam and colleagues later distilled five strategies (Leahy, Lyon, Thompson & Wiliam, 2005, "Classroom Assessment: Minute by Minute, Day by Day", Educational Leadership 63(3)): clarify and share learning intentions and success criteria; engineer discussions, questions, and tasks that elicit evidence of learning; give feedback that moves learners forward; activate students as owners of their learning; activate students as resources for one another. Hinge questions (a diagnostic multiple-choice question at a pivot point in a lesson, whose wrong options each map to a known misconception, answered by everyone at once) and exit tickets are practitioner techniques from this tradition (Wiliam 2011, Embedded Formative Assessment); they are well reasoned but not separately tested in RCTs that I could find.
Kingston & Nash (2011), "Formative Assessment: A Meta-Analysis and a Call for Research", Educational Measurement: Issues and Practice 30(4), 28-37, DOI 10.1111/j.1745-3992.2011.00220.x. Of 300+ studies, only 13 (42 effects) were usable; weighted mean 0.20 (median 0.25); by subject ELA 0.32, maths 0.17, science 0.09. A sober correction.
Hattie & Timperley (2007), "The Power of Feedback", Review of Educational Research 77(1), 81-112, DOI 10.3102/003465430298487. Feedback answers three questions (Where am I going? How am I going? Where to next?) at four levels: task, process (strategy), self-regulation, and self (praise of the person). Task, process, and self-regulation feedback help; self-level praise usually does not.
Wisniewski, Zierer & Hattie (2020), "The Power of Feedback Revisited", Frontiers in Psychology 10, 3087, DOI 10.3389/fpsyg.2019.03087. 435 studies, 994 effects, n > 61,000: d = 0.48, with significant heterogeneity; more informative feedback has more effect; effects larger on cognitive and motor outcomes than on motivation and behaviour.
Kluger & DeNisi (1996), "The Effects of Feedback Interventions on Performance", Psychological Bulletin 119(2), 254-284, DOI 10.1037/0033-2909.119.2.254. 607 effects, 23,663 observations: mean d = 0.41, but over one third of feedback interventions decreased performance. Feedback loses power as it moves attention away from the task toward the self (praise, criticism, normative comparison).
Can software do this faithfully?
- Eliciting evidence from every response, hinge-style diagnostic items whose distractors map to named misconceptions, and adapting what comes next: yes; this is formative assessment's core, done per child per item.
- Task- and process-level feedback (what was right, which step went wrong, what strategy to use): yes when authored against known bug rules.
- Avoiding self-level feedback, normative comparison, and volume badges: yes, by design choice (and required by Numberkit's rules).
- Sharing learning intentions and success criteria: yes, in child language ("today: tens times ones").
- Activating students as resources for one another: no; single learner. Parent role: the parent can be the "other" (section 6).
5. Questioning and responsive teaching
Wait time. Rowe (1986), "Wait Time: Slowing Down May Be a Way of Speeding Up!", Journal of Teacher Education 37(1), 43-50, DOI 10.1177/002248718603700110. Teachers typically wait under one second after a question; at three seconds or more (after the question, and again after the answer), answers become longer, more reasoned, and more students respond. Evidence is mostly 1970s observational work.
Cold call and no opt out. Dallimore, Hertenstein & Platt (2013), Journal of Management Education 37(3), 305-341: in undergraduate classes, high cold-calling was associated with more voluntary participation and growing comfort. No-opt-out (the student who cannot answer is returned to after hearing a peer's answer) is a practitioner technique (Lemov, Teach Like a Champion) without experimental evidence I could find. Its principle, that every learner retrieves, is well supported.
Eliciting and responding to student thinking. Carpenter, Fennema, Peterson, Chiang & Loef (1989), "Using Knowledge of Children's Mathematics Thinking in Classroom Teaching: An Experimental Study", American Educational Research Journal 26(4), 499-531, DOI 10.3102/00028312026004499: 40 first-grade teachers randomly assigned; those who studied a research-based map of children's addition and subtraction strategies (Cognitively Guided Instruction) listened to children's methods more, taught problem solving more and number facts less, and their students did better at problem solving without losing fact recall [outcome details from my knowledge of the paper; the design is verified]. This is the strongest experimental evidence that knowing the typical paths of children's thinking, and listening for them, improves maths learning.
Error handling. Steuer, Rosentritt-Brunn & Dresel (2013), "Dealing with Errors in Mathematics Classrooms: Structure and Relevance of Perceived Error Climate", Contemporary Educational Psychology 38(3), 196-210: students often experience errors as shameful; a perceived constructive error climate (errors discussed, no ridicule, teacher support) predicts adaptive responses to errors and learning. Metcalfe (2017), "Learning from Errors", Annual Review of Psychology 68, 465-489, DOI 10.1146/annurev-psych-010416-044022: making errors followed by corrective feedback is beneficial, contrary to old errorless-learning advice, provided correction is prompt and ideally includes why. Butterfield & Metcalfe (2001) named the hypercorrection effect: high-confidence errors are corrected more readily than low-confidence ones. Metcalfe & Finn (2012), "Hypercorrection of High Confidence Errors in Children", Learning and Instruction 22(4), 253-261: shown in grade 3 to 6 children across three experiments. Booth, Lange, Koedinger & Newton (2013), Learning and Instruction 25, 24-34: in Algebra I Cognitive Tutor classrooms, adding correct and especially incorrect worked examples with self-explanation prompts improved conceptual understanding over practice alone.
Can software do this faithfully?
- Wait time: yes, and better than most humans; never rush, never count down (already Numberkit rule 3).
- Every learner answers, no opt-out: yes; the child must attempt before the answer is shown, and a missed item returns later.
- Diagnosing the child's strategy from the answer: partly; bug rules detect many known misconceptions from the wrong answer, and step items expose where it went wrong, but a child's novel reasoning is invisible without talk. Parent role: a scripted "show me how you did that" prompt for the parent (already in Numberkit's parent report).
- Constructive error climate and hypercorrection: yes; neutral correction, show why, re-ask soon, re-ask again later; optionally ask confidence on some items so confident errors get richer correction.
- Incorrect worked examples: yes (Numberkit's Phase 4 plan already includes them).
6. Learning by teaching and peers
Protégé effect. Chase, Chin, Oppezzo & Schwartz (2009), "Teachable Agents and the Protégé Effect: Increasing the Effort Towards Learning", Journal of Science Education and Technology 18(4), 334-352, DOI 10.1007/s10956-009-9180-4: 8th graders and 5th graders [the grade of the second study is from my knowledge, unverified] using Betty's Brain spent more time on learning and learned more when they believed they were teaching an agent than when learning for themselves, with the largest benefit for lower achievers. Mechanism: effort for someone else, and protection from ego threat.
Learning by teaching. Fiorella & Mayer (2013), Contemporary Educational Psychology 38(4), 281-288: undergraduates who prepared to teach did better on an immediate test but not a delayed one. Fiorella & Mayer (2014), Contemporary Educational Psychology 39(2), 75-85: those who actually taught (recorded a video lesson for a fictitious student) did best on the delayed test. So the act of explaining, not the expectation, produces durable learning. Roscoe & Chi (2007), "Understanding Tutor Learning", Review of Educational Research 77(4), 534-574: peer tutors learn when they engage in knowledge-building (monitoring their understanding, integrating, generating new inferences) but show a pervasive knowledge-telling bias (reciting), which limits their gains.
Peer instruction. Crouch & Mazur (2001), American Journal of Physics 69(9), 970-977: ten years of concept question, vote, peer discussion, re-vote in Harvard introductory physics raised conceptual and quantitative performance. Smith et al. (2009), Science 323(5910), 122-124: discussion improved answers to a new isomorphic question even when no one in the group first knew the answer. Undergraduates.
Reciprocal teaching. Palincsar & Brown (1984), Cognition and Instruction [details unverified]; Rosenshine & Meister (1994), "Reciprocal Teaching: A Review of the Research", Review of Educational Research 64(4), 479-530: 16 studies; median effect 0.32 on standardised comprehension tests and 0.88 on experimenter-made tests. Reading comprehension, not maths.
Cooperative learning. Slavin's reviews (e.g. Slavin, "Cooperative Learning and Achievement: Theory and Research", in Handbook of Psychology, Wiley): cooperative methods raise achievement when they combine group goals with individual accountability; in his count, 37 of 44 comparisons of at least four weeks were significantly positive and none favoured control; informal group work without these features generally was not effective. Slavin & Lake (2008), "Effective Programs in Elementary Mathematics: A Best-Evidence Synthesis", Review of Educational Research 78(3), 427-515 [pages unverified], DOI 10.3102/0034654308317473: the strongest effects in elementary maths came from instructional-process approaches (cooperative learning, classroom management and motivation, supplemental tutoring), more than from curricula or computer-assisted instruction.
Can software do this faithfully?
- A protégé character the child teaches: partly. Software can let the child choose the correct step for a character, spot the character's error, or build the character's array, which captures effort and ego protection and is checkable. It cannot evaluate a free explanation, so the constructive and knowledge-building part is limited to what can be structured. Note the rule that themes are skins and no mascot reactions during practice: a protégé must be a mathematical task (the character's worked example has an error; find it), not a reactive mascot.
- Peer discussion, reciprocal teaching, cooperative structures: no for a single learner.
- Explaining to a real person: no in software, but it is the strongest use of a parent. Fiorella & Mayer's result suggests the durable benefit comes from actually explaining to someone. A short, scripted "teach it back" moment with a parent (the child shows the parent the array for 7 x 8 and explains why it is 56) is the closest faithful replica.
7. Relationships and belief
Teacher expectations. Rosenthal & Jacobson (1968), Pygmalion in the Classroom. Raudenbush (1984), "Magnitude of Teacher Expectancy Effects on Pupil IQ as a Function of the Credibility of Expectancy Induction: A Synthesis of Findings from 18 Experiments", Journal of Educational Psychology 76(1), 85-97: effects on IQ were small on average and appeared mainly when expectations were induced before teachers knew the pupils; a couple of weeks' contact largely removed them. Jussim & Harber (2005), "Teacher Expectations and Self-Fulfilling Prophecies: Knowns and Unknowns, Resolved and Unresolved Controversies", Personality and Social Psychology Review 9(2), 131-155, DOI 10.1207/s15327957pspr0902_3: self-fulfilling prophecies occur but are typically small, do not accumulate much, may be larger for stigmatised groups, and teacher expectations predict outcomes mainly because they are accurate.
Praise and mindset. Mueller & Dweck (1998), Journal of Personality and Social Psychology 75(1), 33-52: fifth graders praised for intelligence rather than effort chose easier tasks, persisted less, and performed worse after failure. Gunderson et al. (2013), Child Development 84(5), 1526-1541: the proportion of parents' process praise to 1- to 3-year-olds predicted children's incremental beliefs and preference for challenge at ages 7 to 8 (correlational). Growth-mindset interventions are weaker than the popular story: weak average effects (Sisk et al. 2018, Psychological Science 29(4), 549-571); modest grade gains for lower achievers in a national trial of over 11,000 9th graders (Yeager et al. 2019, Nature 573, 364-369); and a live dispute over bias and method (Macnamara & Burgoyne 2023, Psychological Bulletin, vs Burnette et al. 2023).
Wise feedback. Yeager, Purdie-Vaughns, Garcia, Apfel, Brzustoski, Master, Hessert, Williams & Cohen (2014), "Breaking the Cycle of Mistrust: Wise Interventions to Provide Critical Feedback Across the Racial Divide", Journal of Experimental Psychology: General 143(2), 804-824, DOI 10.1037/a0033906: 7th graders' essays returned with critical comments plus a note saying "I'm giving you these comments because I have very high expectations and I know that you can reach them": 71% of Black students revised vs 17% in control, with the largest effects for students with low trust. The active ingredient is high standards plus assurance, from a person who has a relationship with the student.
The parent at home. Hill & Tyson (2009), "Parental Involvement in Middle School: A Meta-Analytic Assessment of the Strategies That Promote Achievement", Developmental Psychology 45(3), 740-763: 50 studies, over 50,000 students; academic socialisation (conveying the value of education, linking schoolwork to goals, discussing learning strategies) had the strongest positive association; direct homework help was mixed. Maloney, Ramirez, Gunderson, Levine & Beilock (2015), Psychological Science 26(9), 1480-1488 [pages unverified], DOI 10.1177/0956797615592630: 1st and 2nd graders of maths-anxious parents learned less maths and became more anxious over the year, but only when those parents helped often with homework. Berkowitz, Schaeffer, Maloney, Peterson, Gregor, Levine & Beilock (2015), "Math at Home Adds Up to Achievement in School", Science 350(6257), 196-198 [pages unverified]: 587 first graders; families randomised to a maths or reading story-problem app used together at bedtime; the maths app raised maths achievement over the year, especially for children of maths-anxious parents, even with weekly use. Schaeffer, Rozek, Berkowitz, Levine & Beilock (2018), Journal of Experimental Psychology: General 147(12), 1782-1790 [pages after 1782 unverified]: the benefit persisted through 3rd grade even after use declined, partly by loosening the link between parents' anxiety and their attitudes about maths for their child. A published comment questioned the statistics of the 2015 study, and the authors replied (Science, 2016).
Can software do this faithfully?
- High expectations expressed through content (never capping the child by age, always a next step): yes (Numberkit rule 8).
- Process-level, not person-level, praise in the child's text: yes, by authoring rules.
- Wise feedback's trust component: no; its power comes from a known adult's belief. Parent role: the parent can say the wise-feedback sentence, and the app can give them the words.
- Parent involvement that helps rather than harms: software can shape it. Give the parent a short, positive, structured activity (talking through a problem the child already solved; a "show me" prompt), not a role as a second teacher checking homework. Tell parents explicitly that they do not need to be good at maths themselves and should not re-teach methods.
8. The science of learning in classrooms at scale
Retrieval practice in classrooms holds up. Roediger, Agarwal, McDaniel & McDermott (2011), "Test-Enhanced Learning in the Classroom: Long-Term Improvements from Quizzing", Journal of Experimental Psychology: Applied 17(4), 382-395: in 6th-grade social studies, quizzed items were retained better on unit and end-of-term tests than items not quizzed or merely re-studied. Agarwal, Nunes & Blunt (2021), "Retrieval Practice Consistently Benefits Student Learning: A Systematic Review of Applied Research in Schools and Classrooms", Educational Psychology Review 33, 1409-1453 [pages unverified], DOI 10.1007/s10648-021-09595-9: 50 classroom experiments, 49 effects, n = 5,374; 57% medium or large; benefits across ages, subjects, delays, and formats; only 6% from non-WEIRD countries.
Teacher-delivered spacing can fail. EEF, "SMART Spaces: Spaced Learning Revision Programme", evaluation report (UCL, July 2023): a spaced-learning GCSE science revision programme gave no additional progress in chemistry or combined science, with a high security rating, after a promising pilot. Perry, Lea, Jørgensen, Cordingley, Shapiro & Youdell (2021), Cognitive Science in the Classroom: Evidence and Practice Review, EEF: real potential, thinner and more variable classroom evidence, and faithful implementation as the recurring problem.
Why effects shrink. Lortie-Forgues & Inglis (2019), "Rigorous Large-Scale Educational RCTs Are Often Uninformative: Should We Be Concerned?", Educational Researcher 48(3), 158-166, DOI 10.3102/0013189X19832850: 141 EEF and NCEE trials, over 1.2 million students, mean effect 0.06 SD with wide intervals. Common reasons: broad outcome tests instead of narrow aligned ones, controls that already practise, delivery drift, lower dosage, and heterogeneous students; the Kraft et al. tutoring result shows the same pattern.
Can software do this faithfully? Yes, and this is the strongest argument for an app. Retrieval, spacing, interleaving, and mastery gating are schedules; software executes them identically for every child every day, which is precisely what teacher-delivered versions fail to do. The honest caution is that Numberkit's own gains should be measured on delayed and transfer probes, not on its own practised items, because narrow aligned measures inflate effects (von Hippel's point about Bloom).
What an app can take from great tutors (ranked)
Ranked by strength of evidence times fidelity of software delivery.
- Step-level interaction with immediate, specific feedback and help targeted to the step that went wrong (VanLehn 2011).
- Mastery loops: test, correct, re-test, with a later retrieval before a KC counts as done (von Hippel 2024 on Bloom).
- Retrieval and spacing on a faithful schedule (Agarwal et al. 2021); software removes the delivery drift behind the EEF SMART Spaces null.
- Explicit small-step instruction with worked examples, fading to independent practice at a high success rate (Rosenshine 2012; Alfieri et al. 2011; Stockard et al. 2018).
- Diagnostic items whose wrong answers name the misconception, then a responsive next step (Wiliam's formative assessment; Carpenter et al. 1989), through bug rules over answers and steps.
- Task- and process-level feedback only; no person-level praise, no comparison (Kluger & DeNisi 1996; Hattie & Timperley 2007; Mueller & Dweck 1998).
- A constructive error climate: neutral correction with the reason, the item back soon; optionally ask confidence to exploit hypercorrection (Steuer et al. 2013; Metcalfe & Finn 2012).
- Make the child generate where it can be checked: build the array, type the next step, find the error in a worked example (ICAP; Booth et al. 2013).
- Generous wait time, no countdowns (Rowe 1986).
- Dosage by routine: three or more short sessions a week at a set time, shown to the parent (Nickow et al.).
- Structured protégé tasks (correct a character's worked example), which raise effort most for lower achievers (Chase et al. 2009), kept as mathematics, not a reactive mascot.
- Indirect, nurturant correction wording: a question or hint before the answer (Lepper & Woolverton 2002).
What needs a person, and how to involve one
The realistic person is the parent. The evidence says parents help most through light, structured, positive routines and harm most when anxious parents become intensive homework helpers. So Numberkit's parent role should be small, scripted, and warm, never "second teacher."
| Needs a person | Why software falls short | How Numberkit can involve the parent |
|---|---|---|
| Reading mood and affect; deciding not to practise today | No sight of the child; latency and error patterns are weak proxies | Parent decides when to start and stop; app tells the parent what a normal session looks like and that stopping early is fine |
| Relationship and belief (wise feedback) | Trust comes from a known adult (Yeager et al. 2014) | Parent screen offers one sentence to say after a hard week: high standards plus "I know you can reach them", tied to a specific thing the child mastered |
| Hearing the child's own reasoning | Free explanations cannot be evaluated without live AI, which Numberkit keeps away from the child | Weekly "show me" script (already in the parent report): the child shows the parent the representation for a newly mastered fact and explains it; the parent's only job is to ask "how do you know?" |
| Learning by teaching a real audience | Durable gain needs actual explaining (Fiorella & Mayer 2014) | Same "teach it back" moment; framed as the child being the expert, and optionally the child teaching a sibling-character problem to the parent |
| Interactive co-construction and peer discussion | A scripted character adds no new ideas | A short parent-child talk about a word problem the child already solved, using a provided question ("what would change if there were 9 bags?") |
| Dosage and routine | Apps can remind; only people make it happen | Parent sets the fixed slot; the report shows sessions per week against the recommended three or more |
| Academic socialisation (why maths matters) | Hill & Tyson's strongest factor is the parent conveying value | Parent report suggests one real-life use of this week's maths (not a lecture) |
| Protecting against transmitted anxiety | Maloney et al. 2015 | Plain statement to parents: you do not need to be good at maths; do not re-teach methods; ask, don't tell; the app handles correction |
| Novel misconceptions no bug rule covers | Fixed repertoire | Parent report flags "stuck" KCs (repeated errors not matched by any rule) so the parent can ask the child to show their method, and optionally report back to the maintainers for new bug rules |
An AI model could help the adult rather than the child, as Tutor CoPilot did for tutors (Wang et al. 2024), for example by drafting parent prompts ahead of time for review. Under Numberkit's rules this would have to be generated ahead of time, validated, and reviewed, never live to the child. Any such feature would need its own justification against PLAN section 10.
Contested or weak evidence
- Bloom's 2 sigma: not a tutoring estimate; realistic RCT effects are 0.3 to 0.4 SD, smaller at scale.
- Formative assessment at 0.4 to 0.7: the strict meta-analysis gives about 0.20 (0.17 in maths). Hinge questions, exit tickets, cold call, and no-opt-out lack direct experimental tests.
- ICAP's strict ordering: a heuristic, not a law; overt behaviour is a poor proxy for thinking, and novices often learn more from worked examples than from solving.
- Gradual release: no direct experimental test of the sequence found.
- Direct Instruction's d = 0.60: largely older, non-randomised studies, many by developers; Follow Through was not randomised.
- Pygmalion: small and short-lived once teachers know pupils.
- Growth mindset interventions: weak on average and disputed; process praise remains a safe authoring rule, a mindset module is not justified.
- Expert-tutor studies: observational and small; they describe, not establish, which moves cause learning.
- Bedtime Math: randomised, but its statistics were publicly disputed; the follow-up points the same way.
- Items marked [unverified] must be checked before external use.
References
Agarwal, P. K., Nunes, L. D., & Blunt, J. R. (2021). Retrieval practice consistently benefits student learning: A systematic review of applied research in schools and classrooms. Educational Psychology Review, 33. https://doi.org/10.1007/s10648-021-09595-9
Alfieri, L., Brooks, P. J., Aldrich, N. J., & Tenenbaum, H. R. (2011). Does discovery-based instruction enhance learning? Journal of Educational Psychology, 103(1), 1-18. https://eric.ed.gov/?id=EJ933606
Berkowitz, T., Schaeffer, M. W., Maloney, E. A., Peterson, L., Gregor, C., Levine, S. C., & Beilock, S. L. (2015). Math at home adds up to achievement in school. Science, 350(6257). https://pubmed.ncbi.nlm.nih.gov/26450209/
Black, P., & Wiliam, D. (1998). Assessment and classroom learning. Assessment in Education: Principles, Policy & Practice, 5(1), 7-74. https://doi.org/10.1080/0969595980050102
Bloom, B. S. (1984). The 2 sigma problem: The search for methods of group instruction as effective as one-to-one tutoring. Educational Researcher, 13(6), 4-16. [volume and pages unverified]
Booth, J. L., Lange, K. E., Koedinger, K. R., & Newton, K. J. (2013). Using example problems to improve student learning in algebra: Differentiating between correct and incorrect examples. Learning and Instruction, 25, 24-34. https://www.sciencedirect.com/science/article/abs/pii/S0959475212000904
Carpenter, T. P., Fennema, E., Peterson, P. L., Chiang, C.-P., & Loef, M. (1989). Using knowledge of children's mathematics thinking in classroom teaching: An experimental study. American Educational Research Journal, 26(4), 499-531. https://doi.org/10.3102/00028312026004499
Chase, C. C., Chin, D. B., Oppezzo, M. A., & Schwartz, D. L. (2009). Teachable agents and the protégé effect: Increasing the effort towards learning. Journal of Science Education and Technology, 18(4), 334-352. https://doi.org/10.1007/s10956-009-9180-4
Chi, M. T. H., Siler, S. A., Jeong, H., Yamauchi, T., & Hausmann, R. G. (2001). Learning from human tutoring. Cognitive Science, 25, 471-533. https://doi.org/10.1207/s15516709cog2504_1
Chi, M. T. H., & Wylie, R. (2014). The ICAP framework: Linking cognitive engagement to active learning outcomes. Educational Psychologist, 49(4), 219-243. https://doi.org/10.1080/00461520.2014.965823
Crouch, C. H., & Mazur, E. (2001). Peer Instruction: Ten years of experience and results. American Journal of Physics, 69(9), 970-977. https://pubs.aip.org/aapt/ajp/article/69/9/970/310529
Dallimore, E. J., Hertenstein, J. H., & Platt, M. B. (2013). Impact of cold-calling on student voluntary participation. Journal of Management Education, 37(3), 305-341. https://doi.org/10.1177/1052562912446067
Education Endowment Foundation / UCL (2023). SMART Spaces: Spaced Learning Revision Programme. Evaluation report. https://educationendowmentfoundation.org.uk/projects-and-evaluation/projects/smart-spaces
Fiorella, L., & Mayer, R. E. (2013). The relative benefits of learning by teaching and teaching expectancy. Contemporary Educational Psychology, 38(4), 281-288. https://www.sciencedirect.com/science/article/abs/pii/S0361476X13000209
Fiorella, L., & Mayer, R. E. (2014). Role of expectations and explanations in learning by teaching. Contemporary Educational Psychology, 39(2), 75-85. https://www.sciencedirect.com/science/article/abs/pii/S0361476X14000022
Fryer, R. G. (2014). Injecting charter school best practices into traditional public schools: Evidence from field experiments. Quarterly Journal of Economics, 129(3), 1355-1407. https://academic.oup.com/qje/article-abstract/129/3/1355/1817328
Graesser, A. C., Person, N. K., & Magliano, J. P. (1995). Collaborative dialogue patterns in naturalistic one-to-one tutoring. Applied Cognitive Psychology, 9(6), 495-522. https://doi.org/10.1002/acp.2350090604
Gunderson, E. A., Gripshover, S. J., Romero, C., Dweck, C. S., Goldin-Meadow, S., & Levine, S. C. (2013). Parent praise to 1- to 3-year-olds predicts children's motivational frameworks 5 years later. Child Development, 84(5), 1526-1541. https://onlinelibrary.wiley.com/doi/abs/10.1111/cdev.12064
Guryan, J., Ludwig, J., Bhatt, M. P., Cook, P. J., Davis, J. M. V., Dodge, K., Farkas, G., Fryer, R. G., Mayer, S., Pollack, H., & Steinberg, L. (2023). Not too late: Improving academic outcomes among adolescents. American Economic Review, 113(3), 738-765. https://doi.org/10.1257/aer.20210434
Hattie, J., & Timperley, H. (2007). The power of feedback. Review of Educational Research, 77(1), 81-112. https://doi.org/10.3102/003465430298487
Hill, N. E., & Tyson, D. F. (2009). Parental involvement in middle school: A meta-analytic assessment of the strategies that promote achievement. Developmental Psychology, 45(3), 740-763. https://pmc.ncbi.nlm.nih.gov/articles/PMC2782391/
Hmelo-Silver, C. E., Duncan, R. G., & Chinn, C. A. (2007). Scaffolding and achievement in problem-based and inquiry learning: A response to Kirschner, Sweller, and Clark (2006). Educational Psychologist, 42(2), 99-107. https://eric.ed.gov/?id=EJ772220
House, E. R., Glass, G. V., McLean, L. D., & Walker, D. F. (1978). No simple answer: Critique of the Follow Through evaluation. Harvard Educational Review, 48(2). [issue unverified]
Jussim, L., & Harber, K. D. (2005). Teacher expectations and self-fulfilling prophecies: Knowns and unknowns, resolved and unresolved controversies. Personality and Social Psychology Review, 9(2), 131-155. https://doi.org/10.1207/s15327957pspr0902_3
Kingston, N., & Nash, B. (2011). Formative assessment: A meta-analysis and a call for research. Educational Measurement: Issues and Practice, 30(4), 28-37. https://doi.org/10.1111/j.1745-3992.2011.00220.x
Kirschner, P. A., Sweller, J., & Clark, R. E. (2006). Why minimal guidance during instruction does not work. Educational Psychologist, 41(2), 75-86. https://doi.org/10.1207/s15326985ep4102_1
Kluger, A. N., & DeNisi, A. (1996). The effects of feedback interventions on performance: A historical review, a meta-analysis, and a preliminary feedback intervention theory. Psychological Bulletin, 119(2), 254-284. https://doi.org/10.1037/0033-2909.119.2.254
Kraft, M. A., Schueler, B. E., & Falken, G. T. (2024/2026). What impacts should we expect from tutoring at scale? Exploring meta-analytic generalizability. EdWorkingPaper 24-1031; Review of Educational Research. https://doi.org/10.3102/00346543261446660
Leahy, S., Lyon, C., Thompson, M., & Wiliam, D. (2005). Classroom assessment: Minute by minute, day by day. Educational Leadership, 63(3), 18-24. https://eric.ed.gov/?id=EJ745452
Lehman, B., D'Mello, S., Cade, W., & Person, N. (2012). How do they do it? Investigating dialogue moves within dialogue modes in expert human tutoring. Intelligent Tutoring Systems (ITS 2012), Springer LNCS.
Lepper, M. R., & Woolverton, M. (2002). The wisdom of practice: Lessons learned from the study of highly effective tutors. In J. Aronson (Ed.), Improving Academic Achievement (pp. 135-158). Academic Press.
Lortie-Forgues, H., & Inglis, M. (2019). Rigorous large-scale educational RCTs are often uninformative: Should we be concerned? Educational Researcher, 48(3), 158-166. https://doi.org/10.3102/0013189X19832850
Macnamara, B. N., & Burgoyne, A. P. (2023). Do growth mindset interventions impact students' academic achievement? A systematic review and meta-analysis with recommendations for best practices. Psychological Bulletin. [volume unverified]
Maloney, E. A., Ramirez, G., Gunderson, E. A., Levine, S. C., & Beilock, S. L. (2015). Intergenerational effects of parents' math anxiety on children's math achievement and anxiety. Psychological Science, 26(9). https://doi.org/10.1177/0956797615592630
Metcalfe, J. (2017). Learning from errors. Annual Review of Psychology, 68, 465-489. https://doi.org/10.1146/annurev-psych-010416-044022
Metcalfe, J., & Finn, B. (2012). Hypercorrection of high confidence errors in children. Learning and Instruction, 22(4), 253-261. https://eric.ed.gov/?id=EJ964332
Mueller, C. M., & Dweck, C. S. (1998). Praise for intelligence can undermine children's motivation and performance. Journal of Personality and Social Psychology, 75(1), 33-52. https://pubmed.ncbi.nlm.nih.gov/9686450/
Nickow, A., Oreopoulos, P., & Quan, V. (2020). The impressive effects of tutoring on preK-12 learning: A systematic review and meta-analysis of the experimental evidence. NBER Working Paper 27476 / EdWorkingPaper 20-267, https://doi.org/10.26300/eh0c-pc52. Published 2024 as "The promise of tutoring for preK-12 learning", American Educational Research Journal, https://doi.org/10.3102/00028312231208687
Pearson, P. D., & Gallagher, M. C. (1983). The instruction of reading comprehension. Contemporary Educational Psychology, 8(3), 317-344. [details unverified]
Perry, T., Lea, R., Jørgensen, C. R., Cordingley, P., Shapiro, K., & Youdell, D. (2021). Cognitive Science in the Classroom: Evidence and Practice Review. Education Endowment Foundation.
Questioning central assumptions of the ICAP framework (2023). npj Science of Learning. https://doi.org/10.1038/s41539-023-00197-4 [authors not verified in this review]
Raudenbush, S. W. (1984). Magnitude of teacher expectancy effects on pupil IQ as a function of the credibility of expectancy induction: A synthesis of findings from 18 experiments. Journal of Educational Psychology, 76(1), 85-97. https://eric.ed.gov/?id=EJ304954
Roediger, H. L., Agarwal, P. K., McDaniel, M. A., & McDermott, K. B. (2011). Test-enhanced learning in the classroom: Long-term improvements from quizzing. Journal of Experimental Psychology: Applied, 17(4), 382-395. https://psycnet.apa.org/record/2011-26204-001
Roscoe, R. D., & Chi, M. T. H. (2007). Understanding tutor learning: Knowledge-building and knowledge-telling in peer tutors' explanations and questions. Review of Educational Research, 77(4), 534-574. https://doi.org/10.3102/0034654307309920
Rosenshine, B. (2012). Principles of instruction: Research-based strategies that all teachers should know. American Educator, 36(1), 12-19, 39. https://www.aft.org/sites/default/files/Rosenshine.pdf
Rosenshine, B., & Meister, C. (1994). Reciprocal teaching: A review of the research. Review of Educational Research, 64(4), 479-530. https://doi.org/10.3102/00346543064004479
Rowe, M. B. (1986). Wait time: Slowing down may be a way of speeding up! Journal of Teacher Education, 37(1), 43-50. https://doi.org/10.1177/002248718603700110
Schaeffer, M. W., Rozek, C. S., Berkowitz, T., Levine, S. C., & Beilock, S. L. (2018). Disassociating the relation between parents' math anxiety and children's math achievement: Long-term effects of a math app intervention. Journal of Experimental Psychology: General, 147(12), 1782-. https://cogdevlab.uchicago.edu/files/2025/07/Schaeffer-Rozek-Berkowitz-Levine-2018.pdf
Sisk, V. F., Burgoyne, A. P., Sun, J., Butler, J. L., & Macnamara, B. N. (2018). To what extent and under which circumstances are growth mind-sets important to academic achievement? Two meta-analyses. Psychological Science, 29(4), 549-571. https://doi.org/10.1177/0956797617739704
Slavin, R. E. (2013). Cooperative learning and achievement: Theory and research. In Handbook of Psychology (2nd ed.), Wiley. https://onlinelibrary.wiley.com/doi/abs/10.1002/9781118133880.hop207008 [edition year unverified]
Slavin, R. E., & Lake, C. (2008). Effective programs in elementary mathematics: A best-evidence synthesis. Review of Educational Research, 78(3). https://doi.org/10.3102/0034654308317473
Smith, M. K., Wood, W. B., Adams, W. K., Wieman, C., Knight, J. K., Guild, N., & Su, T. T. (2009). Why peer discussion improves student performance on in-class concept questions. Science, 323(5910), 122-124.
Stebbins, L. B., et al. (1977). Education as Experimentation: A Planned Variation Model (Follow Through evaluation). Abt Associates. [details unverified]
Steuer, G., Rosentritt-Brunn, G., & Dresel, M. (2013). Dealing with errors in mathematics classrooms: Structure and relevance of perceived error climate. Contemporary Educational Psychology, 38(3), 196-210.
Stockard, J., Wood, T. W., Coughlin, C., & Rasplica Khoury, C. (2018). The effectiveness of Direct Instruction curricula: A meta-analysis of a half century of research. Review of Educational Research, 88(4), 479-507. https://doi.org/10.3102/0034654317751919
VanLehn, K. (2011). The relative effectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems. Educational Psychologist, 46(4), 197-221. https://doi.org/10.1080/00461520.2011.611369
von Hippel, P. T. (2024). Two-sigma tutoring: Separating science fiction from science fact. Education Next, 24(2). https://www.educationnext.org/two-sigma-tutoring-separating-science-fiction-from-science-fact/
Wang, R. E., Ribeiro, A. T., Robinson, C. D., Loeb, S., & Demszky, D. (2024). Tutor CoPilot: A human-AI approach for scaling real-time expertise. EdWorkingPaper 24-1054. https://edworkingpapers.com/ai24-1054
Wiggins, B. L., Eddy, S. L., Grunspan, D. Z., & Crowe, A. J. (2017). The ICAP active learning framework predicts the learning gains observed in intensely active classroom experiences. AERA Open, 3(2). https://doi.org/10.1177/2332858417708567
Wiliam, D. (2011). Embedded Formative Assessment. Solution Tree.
Wisniewski, B., Zierer, K., & Hattie, J. (2020). The power of feedback revisited: A meta-analysis of educational feedback research. Frontiers in Psychology, 10, 3087. https://doi.org/10.3389/fpsyg.2019.03087
Yeager, D. S., Hanselman, P., Walton, G. M., Murray, J. S., Crosnoe, R., Muller, C., et al. (2019). A national experiment reveals where a growth mindset improves achievement. Nature, 573, 364-369. https://www.nature.com/articles/s41586-019-1466-y
Yeager, D. S., Purdie-Vaughns, V., Garcia, J., Apfel, N., Brzustoski, P., Master, A., Hessert, W. T., Williams, M. E., & Cohen, G. L. (2014). Breaking the cycle of mistrust: Wise interventions to provide critical feedback across the racial divide. Journal of Experimental Psychology: General, 143(2), 804-824. https://doi.org/10.1037/a0033906