Making a Math App Genuinely Fun for a 9-Year-Old: Literature Review
Prepared 2026-09-24. Scope: motivation, game design, rewards, challenge calibration, mindset, narrative, parent involvement, and evidence on well-known math learning games. Citations are inline; a References list is at the end. Where a citation or number could not be verified against a primary source during this review, it is marked "unverified".
Summary of key findings
- "Chocolate-covered broccoli" has a research counterpart: intrinsic integration. Habgood and Ainsworth (2011) built two versions of the same game (Zombie Division) that differed only in whether the division was inside the combat mechanic or in an end-of-level quiz. Children learned more from the intrinsic version on a delayed test (partial eta-squared = 0.24 for the group effect) and, when given free choice, played it over seven times longer (75.7 vs 10.3 minutes). The fun should come from manipulating the mathematical objects themselves, not from a theme wrapped around a quiz.
- Games do help learning, modestly: g = 0.33 vs non-game instruction (Clark, Tanner-Smith and Killingsworth, 2016), d = 0.29 for learning and 0.36 for retention (Wouters et al., 2013), but only d = 0.13 in the math-specific meta-analysis (Tokac, Novak and Thompson, 2019). Games are not reliably more motivating than ordinary instruction (Wouters et al.: d = 0.26, not significant).
- Design features that move the needle in meta-analyses: multiple play sessions (g = 0.44 vs 0.08 for a single session), enhanced scaffolding such as explanations and hints (g = 0.56 vs 0.26 for plain right/wrong/points feedback), schematic rather than realistic visuals (g = 0.44 vs 0.28), and single-player non-competitive play (g = 0.45) over competitive configurations (g near zero). Strong narrative did not help learning (g = 0.27 vs 0.44 with no narrative).
- Self-Determination Theory holds up in games: perceived competence and autonomy, supported by intuitive controls, predict enjoyment and voluntary continued play (Ryan, Rigby and Przybylski, 2006). Small, incidental choices boost intrinsic motivation, especially in children (Patall, Cooper and Robinson, 2008; Cordova and Lepper, 1996).
- Expected tangible rewards undermine children's intrinsic motivation (d = -0.46 for free-choice behavior in child samples; Deci, Koestner and Ryan, 1999). Verbal praise helps adults (d = 0.33) but not measurably children (d = 0.11, n.s.). Badges plus a leaderboard reduced motivation and exam scores over a semester (Hanus and Fox, 2015). Leaderboards trigger social comparison and can depress the performance of stereotyped groups (Christy and Fox, 2014).
- Streaks work as a usage lever: a 60,000-student RCT found streak-highlighting increased platform use and improved math achievement (Aulagnon et al., 2025). Evidence on the downside of streak-loss for children is anecdotal only.
- Challenge is the strongest game-level predictor of learning (Hamari et al., 2016), but the "85 percent rule" (Wilson et al., 2019) is a result about gradient-descent learners on binary tasks and has not been validated in humans. Large-scale A/B tests in a children's number-line game found that easier conditions were more engaging but produced the slowest learning (Lomas et al., 2013), so engagement and learning must be balanced deliberately.
- Growth-mindset interventions have small average effects (about 0.05 to 0.10 SD), concentrated among lower-achieving students, and the literature shows publication and design bias. Keep explicit mindset messaging minimal; instead design error handling and feedback framing so that mistakes are informative (process praise, not ability praise: Mueller and Dweck, 1998).
- Narrative, characters and fantasy raise interest but not necessarily learning; irrelevant decorative content actively hurts (seductive details g = -0.16; Sundararajan and Adesope, 2020). Contextualization and personalization of the math content itself does help (Cordova and Lepper, 1996).
- Parents matter, but math-anxious parents who help frequently transmit anxiety and lower gains (Maloney et al., 2015). A shared story-problem app (Bedtime Math) helped especially children of math-anxious parents, with effects persisting into third grade (Berkowitz et al., 2015; Schaeffer et al., 2018), though the intent-to-treat effect was disputed (Frank, 2016).
- The acclaimed products have mixed evidence: Motion Math and Zombie Division show positive controlled results; DragonBox produces enjoyment and volume but weak transfer to paper algebra (Long and Aleven, 2017); ST Math's independent RCT was null (Rutherford et al., 2014); Prodigy's own data imply roughly 888 questions per test-score point and an observation logged 16 membership ads to 4 math problems in 19 minutes (Fairplay, 2021).
1. Intrinsic integration: the mechanic is the math
Amy Bruckman (1999) coined the complaint at GDC: "Fun is often treated like a sugar coating to be added to an educational core. Which makes about as much sense as chocolate-dipped broccoli." Malone and Lepper (1987), building on Malone (1980, 1981), proposed the taxonomy still used for intrinsically motivating instruction: individual motivations of challenge (goals, uncertain outcomes, performance feedback, self-esteem), curiosity (sensory and cognitive), control (contingency, choice, power) and fantasy, plus interpersonal motivations of cooperation, competition and recognition. Their distinction between endogenous fantasy (where the skill and the fantasy depend on each other, so the fantasy provides useful representations of the content) and exogenous fantasy (a wrapper that could be swapped for any other subject) is the theoretical root of the broccoli critique.
Habgood and Ainsworth (2011, Journal of the Learning Sciences, 20(2), 169-206, doi:10.1080/10508406.2010.508029) sharpened this from "fantasy" to "core mechanic". Their definition of an intrinsically integrated game has two parts: it (1) delivers learning material through the parts of the game that are the most fun to play, riding on the flow experience rather than interrupting it, and (2) embodies the learning material within the structure of the game world and the player's interactions with it, providing an external representation of the content that is explored through the core mechanics.
What this looked like in Zombie Division (for 7- to 11-year-olds): each skeleton carries a number (the dividend) on its chest. The player's weapons are divisors (each attack embodies a different divisor with its own animation), and an attack only defeats a skeleton if the divisor divides its number exactly. The defeated skeleton's spirit splits into equal portions that become small ghosts bearing the quotient, so the consequence of the math is visible in the world. Deciding which weapon to use, or whether to avoid a skeleton entirely, is the math. The extrinsic version replaced the numbers with symbol combinations (swords, gauntlets, shields) that encoded the same vulnerabilities logically, and moved identical division questions into an end-of-level multiple-choice quiz. A control version had no math.
Results, Study 1 (n = 58, three conditions, 2.25 hours of play plus teacher-led reflection): all groups improved, but the intrinsic group kept improving from post-test to delayed test (+7.0 points) while the extrinsic group plateaued. The delayed test was the only one with a group difference, F(2,48) = 7.49, p < .001, partial eta-squared = 0.24; intrinsic scored about 17 points higher than extrinsic and 21 higher than control. Logs showed the extrinsic group attempted fewer math tasks but was more accurate per attempt; intrinsic children were more willing to risk an attack when a valid divisor existed. Girls learned as much as boys. Study 2 (n = 16, after-school club, free switching between versions with position preserved): 75.7 minutes (SD 35.5) on the intrinsic version versus 10.3 on the extrinsic, t(15) = 7.38, p < .001, r = .89. Children called the intrinsic version easier, quicker and more fun and the quiz slower and "too easy"; one said it was "like mixing paint... the maths in the game with the fun... you don't really think you're doing that much".
Practical reading: the number must be a property of a game object the player reasons about and acts on, and the game's response must be a mathematical consequence. If the content could be swapped for spelling without breaking the game, it is a wrapper.
2. Meta-analyses of game-based learning and gamification
Clark, Tanner-Smith and Killingsworth (2016, Review of Educational Research, 86(1), 79-122, doi:10.3102/0034654315582065; 69 K-16 studies, 2000-2012). Games vs non-game instruction: g = 0.33 [0.19, 0.48], k = 57. Augmented vs standard game designs: g = 0.34 [0.17, 0.51], k = 20. Moderators (full text at PMC4748544): multiple sessions g = 0.44 vs single session g = 0.08 (n.s.); single player without competition g = 0.45 vs single player with competition g = -0.06, team competition g = 0.22, multiplayer g = -0.05; enhanced scaffolding g = 0.56 vs success/fail/points feedback g = 0.26; schematic visuals g = 0.44 vs cartoon 0.27 and realistic 0.28; no narrative g = 0.44 vs strong narrative g = 0.27. Several intervals overlap, so treat the rank orderings, not the point estimates, as the finding. The headline: design matters more than medium.
Wouters, van Nimwegen, van Oostendorp and van der Spek (2013, Journal of Educational Psychology, 105(2), 249-265, doi:10.1037/a0031311; 77 learning comparisons, N = 5,547; 31 motivation comparisons, N = 2,216). Serious games beat conventional instruction for learning (d = 0.29) and retention (d = 0.36) but were not more motivating (d = 0.26, n.s.). Gains were larger with supplementary instruction, multiple sessions, and group play.
Sailer and Homner (2020, Educational Psychology Review, 32, 77-112, doi:10.1007/s10648-019-09498-w). Gamification showed small effects on cognitive (g = 0.49, k = 19, N = 1,686), motivational (g = 0.36, k = 16, N = 2,246) and behavioral outcomes (g = 0.25, k = 9, N = 951). Only the cognitive effect survived a high-rigor subsplit. Game fiction and social interaction moderated behavioral outcomes, and combining competition with collaboration was particularly effective; the authors call the factors behind successful gamification "still somewhat unresolved". Moderator-level g values were paywalled.
Tokac, Novak and Thompson (2019, Journal of Computer Assisted Learning, 35(3), 407-420, doi:10.1111/jcal.12347; 24 PreK-12 math studies): d = 0.13 (p = .02), heterogeneous in magnitude and direction; games are "a slightly effective instructional strategy" for math. Byun and Joung (2018, School Science and Mathematics, 118(3-4), 113-126, doi:10.1111/ssm.12271; 33 empirical studies, 17 with usable data) found drill-and-practice games dominate the literature; the overall effect size was paywalled and is reported secondhand as about 0.37 (unverified).
Two math-game experiments on play configuration: Plass et al. (2013, Journal of Educational Psychology, 105(4), 1050-1066; N = 58 middle schoolers, arithmetic fluency) found competition raised in-game learning, collaboration lowered in-game performance, out-of-game fluency improved equally in all conditions, both social modes raised situational interest, and collaboration produced stronger mastery goals and intention to play again. Ke and Grabowski (2007, British Journal of Educational Technology, 38(2), 249-259; 125 fifth graders) found game play beat paper drills on a standards-based test and on attitudes.
Takeaways: short repeated sessions; explain, do not just mark; schematic visuals; competition optional, cooperative or self-referenced goals for a child at home.
3. Self-Determination Theory in games
Ryan, Rigby and Przybylski (2006, Motivation and Emotion, 30(4), 347-363, doi:10.1007/s11031-006-9051-8) ran four studies. Study 1 (n = 89, 20 minutes of Super Mario 64): intuitive controls predicted in-game competence (beta = .40) and autonomy (beta = .37); competence predicted free-choice continued play (beta = .41) and both needs predicted enjoyment, presence and short-term well-being. Study 4 (MMO players) found autonomy, competence and relatedness independently predicted enjoyment and intention to keep playing. The design translation is well established:
- Competence: clear goals, immediate informational feedback about the math consequence, difficulty that tracks skill, visible mastery (which facts or skills are now "yours"), and controls simple enough that the challenge is the math rather than the interface.
- Autonomy: meaningful choices of what to work on, in which order, with which strategy or representation; choices about incidental context (avatar, world, names). Patall, Cooper and Robinson (2008, Psychological Bulletin, 134(2), 270-300; 41 studies) found choice increased intrinsic motivation, effort, performance and perceived competence, with larger effects for children than adults, for instructionally irrelevant choices, for 2 to 4 successive choices, and when no reward followed the choice. Cordova and Lepper (1996, Journal of Educational Psychology, 88(4), 715-730) found that contextualizing order-of-operations practice in a story, personalizing it with the child's own details, and offering incidental choices each produced large gains in motivation, depth of engagement, learning in a fixed time and perceived competence among elementary students.
- Relatedness: for a 9-year-old at home this means a parent or sibling as co-player or audience, characters who react to the child's mathematical actions, and cooperative goals rather than rankings.
4. Rewards, badges, points, leaderboards, streaks
Lepper, Greene and Nisbett (1973, Journal of Personality and Social Psychology, 28(1), 129-137): 51 preschoolers who already liked drawing; those promised a "Good Player" certificate later drew less in free play than children who got an unexpected award or none. Deci, Koestner and Ryan (1999, Psychological Bulletin, 125(6), 627-668; 128 experiments): expected tangible rewards undermined free-choice intrinsic motivation whether engagement-contingent (d = -0.40), completion-contingent (d = -0.36) or performance-contingent (d = -0.28); in child samples tangible rewards had d = -0.46 (k = 51), worse than for college students. Unexpected rewards had no effect. Verbal rewards enhanced free-choice motivation overall (d = 0.33) but the child-only composite was d = 0.11, not significant. Cameron and Pierce (1994, Review of Educational Research, 64(3), 363-423; 96 studies) argued rewards do not reduce intrinsic motivation overall and praise increases it, while conceding expected tangible rewards for mere completion reduce time on task. Both camps agree that expected tangible rewards for simply doing the activity are the risky category and informational competence feedback is the safe one.
Hanus and Fox (2015, Computers and Education, 80, 152-161; two sections of a 16-week college course, final n = 80): the section with badges and a leaderboard showed declining intrinsic motivation, satisfaction and empowerment, and lower final exam scores mediated by motivation. Christy and Fox (2014, Computers and Education, 78, 66-77; undergraduate women): a female-dominated leaderboard lowered math test performance relative to a male-dominated one, and leaderboards had no positive performance effect. Clark et al. (2016) found competitive single-player configurations had g near zero versus 0.45 for non-competitive play. Both classroom studies were with adults; child evidence is the reward literature plus Plass et al. (2013), where competition raised interest but collaboration raised mastery goals.
Streaks. Aulagnon, Cristia, Cueto and Malamud (2025, Economics of Education Review, doi:10.1016/j.econedurev.2025.102721) randomized about 60,000 Peruvian 4th-6th graders over eight summer weeks. Streak-highlighting messages increased platform use, especially intensive use, and improved math achievement in the roughly 1,500 students assessed afterward. This supports "highlight consistency", not "punish lapses". Claims that streak loss causes disengagement or shifts attention from learning to streak maintenance come from company blogs and commentary, not peer-reviewed research.
Keep: informational feedback tied to the math consequence; unlocks that expand what the child can do mathematically (a new "weapon" is a new divisor or a new representation); unexpected, non-contingent celebrations; personal-best and consistency framing. Avoid: expected tangible or currency rewards for engagement; global leaderboards; badges for volume; any pay-to-progress economy.
5. Flow and challenge calibration
Csikszentmihalyi's flow model predicts engagement when challenge slightly exceeds skill with clear goals and immediate feedback. Hamari et al. (2016, Computers in Human Behavior, 54, 170-179; N = 173 players of two physics/engineering learning games) found challenge predicted learning both directly and via engagement, skill affected learning only via engagement, and immersion did not predict learning; the authors recommend that "challenge of the game should be able to keep up with the learner's growing abilities".
Wilson, Shenhav, Straccia and Cohen (2019, Nature Communications, 10, 4646, doi:10.1038/s41467-019-12552-4) derived that for binary classification learned by stochastic-gradient-descent style learners with Gaussian noise, learning is fastest at about 85 percent training accuracy (15.87 percent error). By their own analysis the rule is not universal: a Bayesian learner with perfect memory learns equally well at any difficulty; the optimum shifts with the noise distribution (82 percent Laplacian, 75 percent Cauchy); multi-class tasks are open; and empirical testing in humans is listed as future work. They note the model yields flow-, boredom- and anxiety-like regions and that psychophysics staircases target roughly 80-85 percent. So 85 percent is a defensible heuristic for a discrimination-type task, not a validated target for children learning multiplication or fractions.
Two cautions from real children's math games. Lomas, Patel, Forlizzi and Koedinger (2013, CHI, doi:10.1145/2470654.2470668; 10,000 to 70,000 players of a number-line estimation game) found the easiest conditions produced the most engagement and longest play, contradicting the inverted-U hypothesis for engagement, while the most engaging conditions had the slowest learning; they propose feedforward (telling players how hard the next challenge is) to encourage challenge-seeking. Habgood and Ainsworth's children preferred the intrinsic version partly because it felt "easier", while learning more from it. The practical target is two-layered: keep moment-to-moment success high (roughly 80-90 percent on core actions) while the mathematical demand advances, and let the child opt into labeled harder challenges.
6. Growth mindset, mistakes and feedback framing
Mueller and Dweck (1998, Journal of Personality and Social Psychology, 75(1), 33-52): six studies with fifth graders (for example n = 128, ages 10-12); praise for intelligence after success led children to prefer performance goals, and after failure to show less persistence, less enjoyment, more low-ability attributions and worse performance than children praised for effort. Moser et al. (2011, Psychological Science, 22(12), 1484-1489; n = 25 adults) linked growth mindset to a larger error positivity (Pe) and better post-error accuracy. Schroder et al. (2017, Developmental Cognitive Neuroscience, 24, 42-50; 123 school-aged children) replicated this in children: growth mindset was associated with larger Pe and higher post-error accuracy.
Intervention evidence is contested. Yeager et al. (2019, Nature, 573, 364-369, doi:10.1038/s41586-019-1466-y; 12,490 ninth graders, 65 schools, two 25-minute online sessions): among lower-achieving students (n = 6,320) core GPA rose 0.10 grade points (standardized 0.11) and advanced-math enrollment rose overall, with effects concentrated where peer norms supported challenge-seeking. Sisk et al. (2018, Psychological Science, 29(4), 549-571): weak mindset-achievement correlation across 273 studies (N = 365,915) and a small intervention effect across 43 studies (N = 57,155; about d = 0.08), with gains not depending on whether mindsets changed. Macnamara and Burgoyne (2023, Psychological Bulletin, 149(3-4), 133-173; 63 studies, N = 97,672): about 0.05 SD overall, smaller in higher-quality studies, larger from authors with financial incentives. Burnette et al. (2023) estimated 0.09 SD overall and 0.16 SD for targeted at-risk students; Tipton et al. (2023) reanalyzed Macnamara and Burgoyne's data and obtained 0.09 and 0.15.
Verdict for a 9-year-old's app: do not build a mindset curriculum; effects are small, adolescent-focused and context-dependent. What is well supported is the micro-level: praise effort, strategy and specific progress rather than ability; make errors information the game responds to (a wrong divisor leaves the skeleton standing and shows why); avoid "You're so smart" and time-pressure framings that make errors feel like exposure. A sentence or two about brains getting stronger on hard problems is harmless but is not the motivational engine.
7. Narrative, characters and fantasy
Cordova and Lepper (1996) show that context and personalization of the problem content itself improves learning, not just interest. Schroeder, Adesope and Gilbert (2013, Journal of Educational Computing Research, 49(1), 1-39; 43 studies, N = 3,088) found pedagogical agents produce a small but significant learning benefit (reported in the paper as g = 0.19, not re-verified here), larger for K-12 than post-secondary, and larger when the agent communicated by on-screen text rather than narration. Clark et al. (2016) found no learning advantage for narrative (strong narrative g = 0.27 vs no narrative g = 0.44), while Sailer and Homner (2020) found game fiction helped behavioral outcomes (doing more). Sundararajan and Adesope (2020, Educational Psychology Review, 32, 707-734; 50 studies, 177 effect sizes) quantified the seductive-details effect: interesting but irrelevant material reduced learning overall (g = -0.16), comprehension (g = -0.19), recall (g = -0.17) and transfer (g = -0.12), mainly by raising extraneous cognitive load. Habgood and Ainsworth argue that the core mechanic, not the fantasy, is what needs to be integrated; Wouters et al. found games not more motivating than instruction on average, so narrative alone does not guarantee engagement.
Implication: use a light, endogenous story in which characters and settings are made of the mathematics (a kingdom whose walls are arrays, a river that is a number line), give the child a companion character that reacts to mathematical moves, and cut decorative animation, cutscenes and collectibles that are not mathematical.
8. Parents, home math, co-play
Maloney, Ramirez, Gunderson, Levine and Beilock (2015, Psychological Science, 26(9), 1480-1488): among first and second graders, parents' math anxiety predicted lower math learning and higher child math anxiety, but only when those parents helped with math homework frequently; no effect on reading. Berkowitz et al. (2015, Science, 350(6257), 196-198, doi:10.1126/science.aac7427; 587 first graders randomized to a Bedtime Math story-problem app or a reading app): the math group gained more, especially children of math-anxious parents, with benefits at about one use per week. Frank (2016, Science, doi:10.1126/science.aad8008) reanalyzed the data and found no significant intent-to-treat effect; the authors stood by their dosage-based conclusions. Schaeffer, Rozek, Berkowitz, Levine and Beilock (2018, Journal of Experimental Psychology: General, doi:10.1037/xge0000490) followed the sample through third grade: the negative link between parental math anxiety and child achievement was broken in the app group and persisted after regular use stopped, mediated by changes in parents' attitudes and expectations. Takeuchi and Stevens (2011, Joan Ganz Cooney Center) define joint media engagement as co-use that prompts interaction and meaning-making.
Implications: involve the parent in a role that does not make them the math authority: story-problem sessions the app poses and the pair solves together, alternating-turn co-op, and digests that report progress in strategy terms rather than scores. Sibling or friend play should be cooperative or turn-based with shared goals (Plass et al., 2013).
9. What acclaimed learning games show
- Zombie Division: see Section 1; the clearest evidence that mechanic-level integration beats a wrapped quiz.
- DragonBox Algebra. Siew, Geofrey and Lee (2016, Research Journal of Mathematics and Technology, 5(1), 66-79; quasi-experiment, 30 vs 30 eighth graders): higher algebraic thinking and attitudes. Long and Aleven (2017, ACM Transactions on Computer-Human Interaction, 24(3), doi:10.1145/3057889; 190 seventh and eighth graders randomized, five class periods): DragonBox players solved many more problems and enjoyed it more, but the Lynnette tutor produced better post-test equation solving. Chan et al. (2023, British Journal of Educational Technology, doi:10.1111/bjet.13304; 253 seventh graders): more in-game progress went with higher algebra knowledge, especially for lower-baseline students. Lesson: an integrated mechanic generates large voluntary practice volume, but transfer to standard notation needs explicit bridging.
- ST Math / JiJi. Rutherford et al. (2014, Journal of Research on Educational Effectiveness, 7(4), 358-383; 52 low-performing schools randomized): negligible, non-significant effects (g about 0.06-0.10) after one and two years. A vendor-commissioned WestEd quasi-experiment (Wendt, Rice and Nakamoto, 2019), rated by SRI as meeting WWC standards with reservations, reported 0.13 (scale scores) and 0.17 (proficiency). Lesson: a wordless, spatial puzzle design is not sufficient at scale; implementation and dosage dominate.
- Motion Math: Fractions (Riconscente, 2013, Games and Culture, 8(4), 186-214; 122 fourth graders, crossover): 20 minutes a day for five days raised fractions scores by over 15 percent and self-efficacy and liking by about 10 percent, all significant versus control. Tilting the device to land a fraction on a number line is the math being the mechanic.
- Slice Fractions. Cyr, Charland, Riopel and Bruyere (2016, EDULEARN16 proceedings): third graders learned as much from the game as from teaching on TIMSS fraction items, with transfer to more abstract items; the publisher cites a 10.5 percent gain after three hours. Conference paper, small sample, partly vendor-reported.
- Prodigy. Fairplay's 2021 FTC complaint and NEPC's 2021 newsletter document 16 membership prompts and 4 math problems in a 19-minute observation, non-members walking in mud while members ride clouds, and Prodigy's own research implying 888 answered questions per standardized-test point. The math is unrelated to the battle narrative (an exogenous wrapper). No independent efficacy study located.
- Times Tables Rock Stars. A British Psychological Society critique states no formal research studies of its effectiveness exist; it relies on generic retrieval-practice evidence plus timers, coins and leaderboards (BPS page not directly accessible; from secondary summaries).
- Number Munchers (MECC, 1986): no efficacy research found.
Design implications: what to build and what to avoid
Build the math into the verbs of the game.
- Number-as-object mechanics. Every enemy, door, bridge or creature carries a number or quantity and responds to mathematical actions: factor-based attacks (Zombie Division), array or area building to make walls of exact size, fraction slicing to share loot or cross gaps, number-line landing (Motion Math, Battleship Numberline), balance mechanics for equations (DragonBox). Test: if you could swap the content for spelling without changing the game, it is a wrapper.
- Mathematical consequences as feedback. A wrong divisor leaves the skeleton standing; a right one splits its spirit into quotient-sized ghosts. Show why, not just whether. Clark et al.'s enhanced-scaffolding result (g = 0.56 vs 0.26) argues for hints, worked examples and explanations inside the fiction.
- Schematic visuals: arrays, number lines, area models, ten-frames. Schematic visuals outperformed cartoon and realistic ones for learning, and they are the representations a 9-year-old needs for multiplication, division and fractions.
- Session structure: 10-15 minute sessions, several per week, with a natural stopping point; multiple sessions g = 0.44 vs single session 0.08. Interleave and space facts and skills rather than block them.
- Challenge: adapt so the child succeeds on roughly 80-90 percent of core actions, but escalate the mathematical demand (larger dividends, more factors, unfamiliar fractions) rather than the motor demand. Label optional challenges ("this gate needs a 7 or a 9") so the child chooses harder content knowingly (feedforward, Lomas et al., 2013).
- Autonomy: two to four incidental choices per session (avatar, world, quest order) and one meaningful mathematical choice (which representation or strategy to use). Personalize contexts with the child's names and interests (Cordova and Lepper, 1996).
- Progression through capability, not currency: unlocking a new divisor, a new fraction denominator or a new tool expands the play space and is informational about competence. Celebrate unexpectedly and non-contingently rather than promising rewards for engagement.
- Consistency framing: show a calendar of days played and personal bests; if a streak is used, make missing a day cost nothing beyond the count resetting, and let the child see that a longer streak means more skills learned. The evidence supports highlighting streaks; punishing lapses is unsupported.
- Feedback language: process-oriented ("you tried 4, saw the remainder, then chose 8"), never ability praise. Errors get a mathematical response, not a buzzer. No countdown timers on conceptual tasks; timed fluency rounds, if any, should be self-referenced (beat your own time) and opt-in.
- Light endogenous narrative: characters and places made of the mathematics; a companion who reacts to mathematical moves; no cutscenes, collectibles or animation that are not mathematical (seductive details, g = -0.16).
- Co-play: a two-player cooperative mode (parent or sibling) with shared goals, alternating turns and roles where the parent does not need to know the answer; a weekly parent digest in strategy terms; occasional story-problem sessions modeled on Bedtime Math for math-anxious households.
- Measurement: embedded assessment on fresh, untutored items and periodic checks with standard-notation problems, since DragonBox and ST Math show that in-game progress does not automatically transfer.
Avoid: end-of-level quizzes that pause the fun; global or class leaderboards; badges for volume; expected tangible rewards or in-game currency for merely logging in; pay-to-progress economies or membership prompts; realistic 3D visuals that add load; strong narrative that delays play; mindset lectures; intelligence praise; heavy time pressure; decorative rewards unrelated to mathematics.
Contested or weak evidence
- Growth mindset interventions: average effects 0.05-0.10 SD, publication and researcher-allegiance bias documented, effects concentrated in lower-achieving adolescents; the child-level neural findings (Schroder et al., 2017) are correlational.
- The 85 percent rule is a theoretical result for gradient-descent learners on binary tasks, with no human validation, and Lomas et al. (2013) show that engagement and learning peak at different difficulties.
- Bedtime Math: the intent-to-treat effect was not significant on Frank's reanalysis; the positive findings depend on dosage analyses and on the anxious-parent subgroup, though the third-grade follow-up strengthens the case.
- Gamification meta-analyses are heterogeneous, mostly college-level, short, and dominated by points-badges-leaderboards; Sailer and Homner's moderator estimates could not be retrieved and their motivational effects were not robust to rigor.
- Clark et al.'s moderator contrasts have overlapping confidence intervals, and the competitive-play categories rest on few studies.
- Byun and Joung's overall effect size (about 0.37) is unverified here; Tokac et al. (d = 0.13) is the more conservative math-specific estimate.
- Product studies are often vendor-commissioned (ST Math WestEd study, Slice Fractions, Prodigy's internal research) or small and quasi-experimental (Siew et al., 2016; Cyr et al., 2016).
- Leaderboard harms (Hanus and Fox, Christy and Fox) were shown in adults; the child evidence is indirect, via the reward literature and Clark et al.'s configuration moderator.
- Streak evidence rests on a single (large) RCT; harms of streak loss in children are undocumented.
- No efficacy research was found for Times Tables Rock Stars or Number Munchers.
References
- Aulagnon, R., Cristia, J., Cueto, S., and Malamud, O. (2025). Streaks to success: The effects of highlighting streaks on student effort and learning. Economics of Education Review. doi:10.1016/j.econedurev.2025.102721
- Berkowitz, T., Schaeffer, M. W., Maloney, E. A., Peterson, L., Gregor, C., Levine, S. C., and Beilock, S. L. (2015). Math at home adds up to achievement in school. Science, 350(6257), 196-198. doi:10.1126/science.aac7427
- Bruckman, A. (1999). Can educational be fun? Game Developers Conference, San Jose. https://faculty.cc.gatech.edu/~asb/papers/bruckman-gdc99.pdf
- Burnette, J. L., et al. (2023). A systematic review and meta-analysis of growth mindset interventions: For whom, how, and why might such interventions work? Psychological Bulletin, 149(3-4). (Effect sizes as cited in Tipton et al., 2023.)
- Byun, J., and Joung, E. (2018). Digital game-based learning for K-12 mathematics education: A meta-analysis. School Science and Mathematics, 118(3-4), 113-126. doi:10.1111/ssm.12271
- Cameron, J., and Pierce, W. D. (1994). Reinforcement, reward, and intrinsic motivation: A meta-analysis. Review of Educational Research, 64(3), 363-423.
- Chan, J. Y.-C., et al. (2023). Keep DRAGging ON: Is solving more problems in DragonBox 12+ associated with higher mathematical performance during the COVID-19 pandemic? British Journal of Educational Technology. doi:10.1111/bjet.13304
- Christy, K. R., and Fox, J. (2014). Leaderboards in a virtual classroom: A test of stereotype threat and social comparison explanations for women's math performance. Computers and Education, 78, 66-77. doi:10.1016/j.compedu.2014.05.005
- Clark, D. B., Tanner-Smith, E. E., and Killingsworth, S. S. (2016). Digital games, design, and learning: A systematic review and meta-analysis. Review of Educational Research, 86(1), 79-122. doi:10.3102/0034654315582065 (open access: PMC4748544)
- Cordova, D. I., and Lepper, M. R. (1996). Intrinsic motivation and the process of learning: Beneficial effects of contextualization, personalization, and choice. Journal of Educational Psychology, 88(4), 715-730. doi:10.1037/0022-0663.88.4.715
- Cyr, S., Charland, P., Riopel, M., and Bruyere, M.-H. (2016). Game-based learning of fractions: Slice Fractions. EDULEARN16 Proceedings. https://library.iated.org/view/CYR2016GAM
- Deci, E. L., Koestner, R., and Ryan, R. M. (1999). A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation. Psychological Bulletin, 125(6), 627-668.
- Fairplay (2021). 7 reasons to say "no" to Prodigy. https://fairplayforkids.org/pf/prodigy/ ; NEPC (2021). Malicious math games? The problem with Prodigy. https://nepc.colorado.edu/publication/newsletter-prodigy-032521
- Frank, M. C. (2016). Comment on "Math at home adds up to achievement in school". Science, 351(6278). doi:10.1126/science.aad8008
- Habgood, M. P. J., and Ainsworth, S. E. (2011). Motivating children to learn effectively: Exploring the value of intrinsic integration in educational games. Journal of the Learning Sciences, 20(2), 169-206. doi:10.1080/10508406.2010.508029
- Hamari, J., Shernoff, D. J., Rowe, E., Coller, B., Asbell-Clarke, J., and Edwards, T. (2016). Challenging games help students learn: An empirical study on engagement, flow and immersion in game-based learning. Computers in Human Behavior, 54, 170-179.
- Hanus, M. D., and Fox, J. (2015). Assessing the effects of gamification in the classroom: A longitudinal study on intrinsic motivation, social comparison, satisfaction, effort, and academic performance. Computers and Education, 80, 152-161. doi:10.1016/j.compedu.2014.08.019
- Ke, F., and Grabowski, B. (2007). Gameplaying for maths learning: Cooperative or not? British Journal of Educational Technology, 38(2), 249-259. doi:10.1111/j.1467-8535.2006.00593.x
- Lepper, M. R., Greene, D., and Nisbett, R. E. (1973). Undermining children's intrinsic interest with extrinsic reward: A test of the "overjustification" hypothesis. Journal of Personality and Social Psychology, 28(1), 129-137.
- Lomas, D., Patel, K., Forlizzi, J. L., and Koedinger, K. R. (2013). Optimizing challenge in an educational game using large-scale design experiments. Proceedings of CHI 2013. doi:10.1145/2470654.2470668
- Long, Y., and Aleven, V. (2017). Educational game and intelligent tutoring system: A classroom study and comparative design analysis. ACM Transactions on Computer-Human Interaction, 24(3). doi:10.1145/3057889
- Macnamara, B. N., and Burgoyne, A. P. (2023). Do growth mindset interventions impact students' academic achievement? A systematic review and meta-analysis with recommendations for best practices. Psychological Bulletin, 149(3-4), 133-173.
- Malone, T. W., and Lepper, M. R. (1987). Making learning fun: A taxonomy of intrinsic motivations for learning. In R. E. Snow and M. J. Farr (Eds.), Aptitude, learning, and instruction: Vol. 3. Conative and affective process analyses (pp. 223-253). Erlbaum.
- Maloney, E. A., Ramirez, G., Gunderson, E. A., Levine, S. C., and Beilock, S. L. (2015). Intergenerational effects of parents' math anxiety on children's math achievement and anxiety. Psychological Science, 26(9), 1480-1488. doi:10.1177/0956797615592630
- Moser, J. S., Schroder, H. S., Heeter, C., Moran, T. P., and Lee, Y.-H. (2011). Mind your errors: Evidence for a neural mechanism linking growth mind-set to adaptive posterror adjustments. Psychological Science, 22(12), 1484-1489. doi:10.1177/0956797611419520
- Mueller, C. M., and Dweck, C. S. (1998). Praise for intelligence can undermine children's motivation and performance. Journal of Personality and Social Psychology, 75(1), 33-52. doi:10.1037/0022-3514.75.1.33
- Patall, E. A., Cooper, H., and Robinson, J. C. (2008). The effects of choice on intrinsic motivation and related outcomes: A meta-analysis of research findings. Psychological Bulletin, 134(2), 270-300. doi:10.1037/0033-2909.134.2.270
- Plass, J. L., O'Keefe, P. A., Homer, B. D., Case, J., Hayward, E. O., Stein, M., and Perlin, K. (2013). The impact of individual, competitive, and collaborative mathematics game play on learning, performance, and motivation. Journal of Educational Psychology, 105(4), 1050-1066. doi:10.1037/a0032688
- Riconscente, M. M. (2013). Results from a controlled study of the iPad fractions game Motion Math. Games and Culture, 8(4), 186-214. doi:10.1177/1555412013496894
- Rutherford, T., Farkas, G., Duncan, G., Burchinal, M., Kibrick, M., Graham, J., Richland, L., Tran, N., Schneider, S., Duran, L., and Martinez, M. E. (2014). A randomized trial of an elementary school mathematics software intervention: Spatial-Temporal Math. Journal of Research on Educational Effectiveness, 7(4), 358-383. doi:10.1080/19345747.2013.856978
- Ryan, R. M., Rigby, C. S., and Przybylski, A. (2006). The motivational pull of video games: A self-determination theory approach. Motivation and Emotion, 30(4), 347-363. doi:10.1007/s11031-006-9051-8
- Sailer, M., and Homner, L. (2020). The gamification of learning: A meta-analysis. Educational Psychology Review, 32, 77-112. doi:10.1007/s10648-019-09498-w
- Schaeffer, M. W., Rozek, C. S., Berkowitz, T., Levine, S. C., and Beilock, S. L. (2018). Disassociating the relation between parents' math anxiety and children's math achievement: Long-term effects of a math app intervention. Journal of Experimental Psychology: General, 147(12). doi:10.1037/xge0000490
- Schenke, K., Rutherford, T., and Farkas, G. (2014). Alignment of game design features and state mathematics standards: Do results reflect intentions? Computers and Education, 76, 215-224.
- Schroder, H. S., Fisher, M. E., Lin, Y., Lo, S. L., Danovitch, J. H., and Moser, J. S. (2017). Neural evidence for enhanced attention to mistakes among school-aged children with a growth mindset. Developmental Cognitive Neuroscience, 24, 42-50. doi:10.1016/j.dcn.2017.01.004
- Schroeder, N. L., Adesope, O. O., and Gilbert, R. B. (2013). How effective are pedagogical agents for learning? A meta-analytic review. Journal of Educational Computing Research, 49(1), 1-39. doi:10.2190/EC.49.1.a
- Siew, N. M., Geofrey, J., and Lee, B. N. (2016). Students' algebraic thinking and attitudes towards algebra: The effects of game-based learning using DragonBox 12+ app. Research Journal of Mathematics and Technology, 5(1), 66-79.
- Sisk, V. F., Burgoyne, A. P., Sun, J., Butler, J. L., and Macnamara, B. N. (2018). To what extent and under which circumstances are growth mind-sets important to academic achievement? Two meta-analyses. Psychological Science, 29(4), 549-571. doi:10.1177/0956797617739704
- Sundararajan, N., and Adesope, O. (2020). Keep it coherent: A meta-analysis of the seductive details effect. Educational Psychology Review, 32, 707-734. doi:10.1007/s10648-020-09522-4
- Takeuchi, L., and Stevens, R. (2011). The new coviewing: Designing for learning through joint media engagement. Joan Ganz Cooney Center.
- Tipton, E., Bryan, C., Murray, J., McDaniel, M., Schneider, B., and Yeager, D. S. (2023). Why meta-analyses of growth mindset and other interventions should follow best practices for examining heterogeneity. Psychological Bulletin, 149(3-4), 229-241.
- Tokac, U., Novak, E., and Thompson, C. G. (2019). Effects of game-based learning on students' mathematics achievement: A meta-analysis. Journal of Computer Assisted Learning, 35(3), 407-420. doi:10.1111/jcal.12347
- Wendt, S., Rice, J., and Nakamoto, J. (2019). A cross-state evaluation of MIND Research Institute's ST Math program and math performance. WestEd. (Vendor-commissioned; SRI review.)
- Wilson, R. C., Shenhav, A., Straccia, M., and Cohen, J. D. (2019). The Eighty Five Percent Rule for optimal learning. Nature Communications, 10, 4646. doi:10.1038/s41467-019-12552-4
- Wouters, P., van Nimwegen, C., van Oostendorp, H., and van der Spek, E. D. (2013). A meta-analysis of the cognitive and motivational effects of serious games. Journal of Educational Psychology, 105(2), 249-265. doi:10.1037/a0031311
- Yeager, D. S., et al. (2019). A national experiment reveals where a growth mindset improves achievement. Nature, 573, 364-369. doi:10.1038/s41586-019-1466-y
- British Psychological Society (n.d.). Times Tables Rock Stars: An academic critique. https://explore.bps.org.uk/content/bpsdeb/1/189/14 (not directly accessible; summarized from secondary sources; unverified)