Adaptive Learning, Mastery, Knowledge Tracing and Feedback for a Solo-Built Math App (age 9)
Literature review compiled 2026-09-24. Verification status: every citation below was checked in this session against a publisher page, ERIC record, Semantic Scholar record, or the full-text PDF unless explicitly marked "(not re-verified)". Numbers quoted from full text are marked as such.
Summary of key findings
- Mastery learning is a real, moderate effect: Kulik, Kulik and Bangert-Drowns (1990; 108 studies) found d = 0.52, larger for weaker students (0.61 vs 0.40) and larger with stricter mastery standards (91-100% correct: d = 0.64 vs 0.44-0.49 below 90%). Bloom's (1984) "2 sigma" rests on two dissertations, not a meta-analysis.
- Step-based tutoring is about as effective as human tutoring (VanLehn 2011: ITS d = 0.76, human d = 0.79). Pooled meta-analytic estimates are 0.42-0.66 (Ma et al. 2014; Kulik and Fletcher 2016), but large field RCTs on standardized tests are ~0.2 (Pane et al. 2014; Roschelle et al. 2016). ALEKS pooled: no better than classroom teaching (Fang et al. 2019).
- With one learner, data-hungry models (DKT, fitted BKT, PFA) are the wrong tools (Gervet et al. 2020; Khajah et al. 2016; Yeung and Yeung 2018). Elo with an uncertainty-based K (Pelanek 2016; Klinkenberg et al. 2011) needs two hand-set parameters and self-corrects online.
- Mastery decisions: threshold and input data matter more than the model; a simple exponential moving average does as well as BKT (Pelanek and Rihak 2018). Cognitive Tutor uses P(learned) >= 0.95.
- Math Garden (Klinkenberg et al. 2011; Brinkhuis et al. 2018) is the closest published model: per-domain Elo ratings for children and items, the "high speed high stakes" scoring rule S = (2x - 1)(d - t) with a visible 20 s limit, items sampled at ~75% expected success (child-selectable 60/75/90), validated on 3,648 children and 3.5 million responses.
- Jansen et al. (2013; N = 207) found higher success rates did not reduce math anxiety but led children to attempt more problems, and practice volume drove gains.
- Elaborated feedback beats correct-answer feedback beats right/wrong (Van der Kleij et al. 2015: 0.49 / 0.32 / 0.05). Immediate feedback wins in applied settings (Kulik and Kulik 1988) but can hurt children who were just taught a correct strategy (Fyfe and Rittle-Johnson 2016).
- Gaming the system is the off-task behaviour most tied to reduced learning (Baker et al. 2004); bottom-out hints work as worked examples when the child pauses to read them (Shih et al. 2008).
- Knowledge components with prerequisites are the unit of adaptation (Koedinger, Corbett and Perfetti 2012); better KC models yield faster mastery and better learning (Cen et al. 2006; Koedinger et al. 2013).
- Interleaving mastered skills: d = 0.79 at 30 days (Rohrer et al. 2015) and d = 0.83 in a 54-class preregistered RCT (Rohrer et al. 2020).
- Parent-dashboard evidence is thin; show practice volume, newly mastered KCs and the current working edge, not peer comparison or per-session error rates. Math-anxious parents who help often with homework depress children's math learning (Maloney et al. 2015).
1. Mastery learning
Bloom (1984), Educational Researcher 13(6), 4-16, DOI 10.3102/0013189X013006004. Summarising the dissertations of Anania and Burke: one-to-one tutoring with mastery procedures scored about 2 SD above conventional classes, group mastery learning (the feedback-corrective cycle) about 1 SD. From the full text: the mastery standard was 80% on parallel formative tests. Caveats: small dissertation samples, locally developed tests, and the 2.0 figure has never been reproduced meta-analytically (VanLehn 2011).
Kulik, Kulik and Bangert-Drowns (1990), Review of Educational Research 60(2), 265-299, DOI 10.3102/00346543060002265. 108 controlled evaluations (72 Keller PSI, 36 Bloom LFM; only two in primary grades). From the full text: mean d = 0.52 (SE 0.033); less able students 0.61 vs more able 0.40 (n.s.); unit mastery level 70-80% d = 0.49 (22 studies), 81-90% 0.44 (34), 91-100% 0.64 (36), with mastery level a significant positive predictor in regression (p = .016); effects larger on local than standardized tests (LFM on standardized tests: none), larger when controls got less quiz feedback; median instructional time rose only 4%.
Guskey (2010), Educational Leadership 68(2), 52-57. Essential features: diagnostic pre-assessment, formative assessment per unit, correctives for non-masters, enrichment for masters, a second parallel assessment; the feedback-corrective cycle is the active ingredient. I did not retrieve his threshold language; Bloom-tradition practice is 80-90% and the Kulik data favour the high end.
Thresholds in computer tutors. Cognitive Tutor / MATHia declares mastery at BKT P(learned) >= 0.95 (Corbett and Anderson 1995; Corbett 2001, UM 2001, existence verified; Fancsali, Nixon and Ritter 2013, EDM, existence verified; Ritter et al. 2016, L@S, existence verified). Israni, Sales and Pane (2018, arXiv 1802.08616) found in the CTAI trial logs that teachers often moved students past sections before mastery and that this appeared to lower posttest scores.
Takeaway. Use a strict standard (>= 0.90-0.95 with a minimum number of correct responses) rather than 80%; treat premature mastery as the costlier error (Pelanek and Rihak weight it 5:1); expect payoff mainly on tests aligned to what was practised.
2. Intelligent tutoring systems: what the evidence says and which features matter
VanLehn (2011), Educational Psychologist 46(4), 197-221, DOI 10.1080/00461520.2011.611369. Verified from the abstract: human tutoring d = 0.79 (not 2.0), ITS d = 0.76. His granularity taxonomy: answer-based, step-based, substep-based; the paper's tabulation gives roughly 0.31 / 0.76 / 0.40 (breakdown from my reading, not re-verified). Interpretation: feedback and help at the step level is what matters, not finer dialogue.
Ma, Adesope, Nesbit and Liu (2014), JEP 106(4), 901-918, DOI 10.1037/a0037123. 107 effect sizes, 14,321 participants. ITS vs teacher-led instruction g = 0.42; vs non-ITS computer instruction 0.57; vs textbooks 0.35; vs human one-to-one tutoring -0.11 (n.s.); vs small groups 0.05. Positive across grade levels and roles; effects did not depend on whether the ITS modelled misconceptions.
Kulik and Fletcher (2016), RER 86(1), 42-78, DOI 10.3102/0034654315581420. 50 evaluations; median 0.66 SD, much smaller on standardized tests; flawed implementations and non-conventional controls gave smaller effects.
Pane, Griffin, McCaffrey and Karam (2014), EEPA 36(2), 127-144, DOI 10.3102/0162373713507480. School-level RCT, seven states: no effect in year one; year two high schools about +0.20 SD; middle schools similar, n.s.
Roschelle, Feng, Murphy and Mason (2016), AERA Open 2(4), DOI 10.1177/2332858416673968. 43 Maine schools, 2,850 seventh graders; ASSISTments homework (immediate feedback, hints, teacher reports): g = 0.18 on a standardized test.
ALEKS. Fang, Ren, Hu and Graesser (2019), Educational Psychology 39(10), 1278-1292: 15 studies, 24 samples; as good as but not better than traditional teaching.
Which features drive the effects? Consistently: step-level immediate feedback and on-demand hints (VanLehn; ASSISTments), mastery-based progression through a KC model (Cognitive Tutor), and practice aligned to what is assessed (the test-type moderator). Misconception modelling was not a significant moderator in Ma et al. (2014): useful, not core.
3. Knowledge tracing, student modelling, and memory models for a solo developer
3.1 Bayesian Knowledge Tracing (BKT)
Corbett and Anderson (1995), User Modeling and User-Adapted Interaction 4(4), 253-278, DOI 10.1007/BF01099821. Two-state HMM per KC with four parameters: P(L0) prior knowledge, P(T) learning rate, P(G) guess, P(S) slip; mastery at P(L) >= 0.95. Illustrative values: P(L0) = 0.36, P(T) = 0.1, P(G) = 0.3, P(S) = 0.05 (van de Sande 2013).
Pitfalls. Beck and Chang (2007), UM 2007, LNCS 4511, DOI 10.1007/978-3-540-73078-1_17: observed performance is consistent with an infinite family of parameter sets that predict identically but make different claims about knowledge; they proposed Dirichlet priors. Van de Sande (2013), JEDM 5(2) (full text read): the HMM curve has only three effective parameters (explaining the P(G)-P(L0) degeneracy), but the per-student tracing update is identifiable given both correct and incorrect responses, and sensible behaviour needs P(G) + P(S) < 1 (commonly both < 0.5). Practical implication: with one learner you cannot fit BKT; run it with hand-set parameters and it is just a smoothed running estimate, and Pelanek and Rihak (2018) show a simple EMA does about as well even on BKT-generated data. pyBKT (Badrinath, Wang and Pardos 2021) exists for later multi-user fitting.
3.2 Performance Factors Analysis and logistic models
Pavlik, Cen and Koedinger (2009), AIED 2009: logit P(correct) = beta_KC + gamma_KC x prior successes + rho_KC x prior failures, fitted by logistic regression; handles multi-KC items. Gervet, Koedinger, Schneider and Mitchell (2020), JEDM 12(3), 31-54 (full text read): across nine datasets, logistic regression with good features ("Best-LR": student ability, item difficulty, log-scaled per-KC and total success/failure counts) led on moderate-size datasets; DKT led only on very large ones (algebra05: Best-LR AUC 0.831, DKT 0.821, PFA 0.769, BKT 0.621). These still require fitting across many learners; for one child, borrow the form with hand-set coefficients.
3.3 Elo-based adaptive practice
Pelanek (2016), Computers & Education 98, 169-179, DOI 10.1016/j.compedu.2016.03.017 (full text read). An answer is a match between learner theta and item difficulty delta: P = 1 / (1 + exp(-(theta - delta))); theta += K (correct - P); delta -= K (correct - P). Earlier work used constant K = 0.4; better is an uncertainty function U(n) = a / (1 + b n), n = prior answers, with a = 1, b = 0.05 as a starting point that can be grid-searched later. Klinkenberg et al. (2011) instead initialise U = 1 and update U := U - 1/40 + D/30 per answer, D = days since the last attempt, so uncertainty shrinks with practice and grows with time away. Pelanek's conclusion: basic Elo plus a/(1 + bn) is the right start; complex variants add little; the extension worth adding for mathematics is response time. He explicitly recommends Elo where expert-built systems are infeasible.
3.4 Deep Knowledge Tracing and why not here
Piech et al. (2015), NIPS 28: LSTM over response sequences with large reported AUC gains over BKT. Khajah, Lindsey and Mozer (2016), EDM, arXiv 1604.02416: BKT extended with forgetting, skill discovery and latent ability matches DKT, so the gain came from ignored regularities, not depth. Yeung and Yeung (2018), L@S, arXiv 1806.02180: DKT can lower a KC's predicted mastery after a correct answer (reconstruction problem) and oscillates across time steps (waviness), requiring extra regularisers. For a one-child app there is no training data, no stability and no interpretability. Do not use it.
3.5 Memory and spacing models
- SM-2 / Leitner / FSRS. SM-2 and Leitner boxes are hand-set expanding schedules. FSRS is a fitted difficulty-stability-retrievability model descended from Ye, Su and Cao (2022), KDD, DOI 10.1145/3534678.3539081 (220 million MaiMemo logs); its default weights are public, so a solo developer can use them without fitting.
- Half-Life Regression. Settles and Meeder (2016), ACL, 1848-1858 (full text read): p = 2^(-Delta / h), h = 2^(Theta . x), features = counts of exposures, correct and incorrect responses, item tags. On 12.9 million Duolingo instances MAE fell to 0.128 vs 0.235 (Leitner) and 0.445 (Pimsleur), a 45%+ reduction, and daily engagement rose 12% in an A/B test. With fixed weights HLR collapses to a Leitner-style rule, which is what one learner's data supports.
- ACT-R scheduling. Pavlik and Anderson (2005), Cognitive Science 29(4), 559-586 (existence verified); Pavlik and Anderson (2008), JEP: Applied 14(2), 101-117: activation-based decay reproduces spacing; picking the next item by expected gain per second gave large recall and latency benefits. Implementable, but parameters were fitted on vocabulary.
- Classroom evidence. Lindsey, Shroyer, Pashler and Mozer (2014), Psychological Science 25(3), 639-647: in a semester-long middle-school Spanish course, personalised spaced review beat massed study by 16.5% and one-size-fits-all spacing by 10.0% on a post-semester cumulative exam. Cepeda et al. (2006), Psychological Bulletin 132(3), 354-380 (839 assessments): optimal spacing grows with the desired retention interval.
- Interleaving in math. Rohrer, Dedrick and Stershic (2015), JEP 107(3), 900-908: n = 126 seventh graders, d = 0.42 (1 day) and 0.79 (30 days). Rohrer, Dedrick, Hartwig and Cheung (2020), JEP, DOI 10.1037/edu0000367 (full text read): preregistered cluster RCT, 54 classes, four months; unannounced test one month later 61% vs 38%, d = 0.83.
4. Math Garden / Prowise Learn (Rekentuin)
Sources: Klinkenberg, Straatemeier and van der Maas (2011), Computers & Education 57(2), 1813-1824, DOI 10.1016/j.compedu.2011.02.003; Maris and van der Maas (2012), Psychometrika 77, 615-633, DOI 10.1007/s11336-012-9288-y; Brinkhuis, Savi, Hofman, Coomans, van der Maas and Maris (2018), Journal of Learning Analytics 5(2), 29-46, DOI 10.18608/jla.2018.52.3 (full text read); Klinkenberg (2014), Springer CCIS 439, DOI 10.1007/978-3-319-08657-6_11 (existence verified).
How it works (Brinkhuis et al. 2018 full text):
- Each domain (addition, multiplication, ...) has its own Elo scale; a child has one rating per domain, every item has a rating; over a billion responses by 2018.
- Scoring rule (Signed Residual Time, "high speed high stakes"): S = (2x - 1)(d - t), x = 0/1 accuracy, t = response time, d = time limit, "generally fixed at 20 seconds". Fast correct = large positive, fast wrong = large negative, slow answers near zero, which discourages guessing and makes the speed-accuracy trade-off explicit. The UI shows coins vanishing one per second; a "?" skip scores 0 and is rate-limited to prevent point-farming.
- Updates: theta <- theta + K (S - E[S]); delta <- delta - K (S - E[S]), with E[S | theta, delta] = d (e^{2d(theta - delta)} + 1) / (e^{2d(theta - delta)} - 1) - 1 / (theta - delta). K is a bias-variance knob tied to an uncertainty term (section 3.3).
- Items are sampled (not max-information CAT) so expected P(correct) is about 0.75; the child can choose easy (~0.90), medium (~0.75) or hard (~0.60); recent items are avoided.
- Validation: 3,648 children, 3.5 million responses in ten months (33% outside school); high reliability, validity and pupil satisfaction; mirror items (2 x 9 vs 9 x 2) track each other over 3.5 years.
- Documented pitfalls: unidimensionality violations; children developing "local strategies" that only work on the item cluster the selector keeps serving (fixed by mixing item types); fast-guess vs slow-solve mixtures the SRT model misfits.
Jansen, Louwerse, Straatemeier, Van der Ven, Klinkenberg and van der Maas (2013), Learning and Individual Differences 24, 190-197, DOI 10.1016/j.lindif.2012.12.014 (ERIC abstract verified). 207 children, grades 3-6, control plus three success-rate conditions, six weeks. Math anxiety improved equally in all conditions; perceived competence gains were limited; performance improved only in the experimental conditions, and higher success rates led to more problems attempted, which produced larger gains. A 2023 Journal of Intelligence study (11(6):108) on the same system again found total tasks completed predicted year-long gains.
Design implications: 75% is a defensible default target; an easy/medium/hard chooser is cheap and supports engagement; response time is informative only if the child sees the timer and understands the rule; item selection must mix item types.
5. Feedback design
Hattie and Timperley (2007), RER 77(1), 81-112, DOI 10.3102/003465430298487. Feedback answers where am I going, how am I going, where next, at four levels: task, process, self-regulation, self. Self-level praise is least effective and can backfire; task feedback works best when it leads to process feedback. Do not quote Hattie's global feedback effect size; it is a contested meta-meta-analysis.
Shute (2008), RER 78(1), 153-189, DOI 10.3102/0034654307313795. Guidelines: nonevaluative, supportive, timely, specific; elaborated response-specific feedback beats verification or answer-until-correct; deliver it in manageable units; avoid normative comparison; immediate and directive for low-achieving learners, delayed and facilitative for high-achieving ones; do not interrupt an engaged learner.
Timing. Kulik and Kulik (1988), RER 58(1), 79-97, DOI 10.3102/00346543058001079: 53 studies; applied classroom studies favour immediate feedback, some laboratory paradigms favour delayed. Fyfe and Rittle-Johnson (2016), JEP 108(1), 82-97 (N = 108 and 101 elementary children): immediate verification helped children not taught a correct strategy but hurt children who had just been taught one; summative end-of-set feedback was a workable alternative. Implication: right after teaching a new procedure, lighten or defer item-by-item feedback for a short exploratory set.
Type. Van der Kleij, Feskens and Eggen (2015), RER 85(4), 475-511, DOI 10.3102/0034654314564881: 40 studies, 70 effects; elaborated feedback 0.49, knowledge of correct response 0.32, knowledge of result 0.05, with the elaborated advantage largest for higher-order outcomes.
Hints and gaming. Baker, Corbett, Koedinger and Wagner (2004), CHI, 383-390, DOI 10.1145/985692.985741: gaming (systematic guessing and hint-clicking) was more strongly associated with reduced learning than any other off-task behaviour. Aleven, McLaren, Roll and Koedinger (2006), IJAIED 16(2), 101-128, and Aleven, Roll, McLaren and Koedinger (2016), IJAIED: tutoring help-seeking durably improved help-seeking but not domain learning. Shih, Koedinger and Scheines (2008), EDM, 117-126: students who pause to read a bottom-out hint are studying a worked example and learn; fast hint-runs are abuse; response time after the hint discriminates.
Error-specific feedback / buggy rules. Brown and Burton (1978), Cognitive Science 2(2), 155-192, DOI 10.1207/s15516709cog0202_4: BUGGY diagnosed systematic subtraction bugs ("smaller-from-larger", "borrow from zero") from a procedural network. VanLehn (1990), Mind Bugs, MIT Press: bugs arise from impasses and repairs. Multiplication-fact errors are mostly operand errors (answers from a neighbouring table, e.g. 3 x 6 = 24), about 87.5% of adult errors per Campbell (1997, secondary source), with a problem-size effect. Fraction misconceptions stem from whole-number bias (Ni and Zhou 2005, Educational Psychologist 40(1), 27-52): adding across numerators and denominators, "bigger denominator means bigger", and "multiplying enlarges, dividing shrinks" (Lortie-Forgues, Tian and Siegler 2015, Developmental Review 38, 201-221; Siegler, Thompson and Schneider 2011 on fraction magnitude as the core competence).
Practical bug catalogue for grade 3-5 (names are mine; categories from the sources above plus practitioner error-pattern references such as Ashlock, not re-verified):
- Facts: table-neighbour errors (7 x 8 = 54); n x 0 = n; n x 1 = 0; inverse confusion (56 / 7 = 7).
- Multi-digit multiplication: forgetting the carry; adding the carry before multiplying; not shifting the second partial product; treating a 0 in the multiplier as 1 or skipping its row.
- Place value: 305 vs 35 for "3 hundreds 5 ones"; treating 0 as nothing when regrouping; comparing by first digit only.
- Division: remainder dropped or appended to the quotient; missing zero in the quotient when the partial dividend is smaller than the divisor; "division makes smaller".
- Fractions: 1/8 > 1/4 because 8 > 4; 1/2 + 1/3 = 2/5; 2/4 != 1/2; "multiplying makes bigger".
Detection: precompute each bug's answer per item, match the child's answer, and after two consecutive matches on a KC deliver the elaborated feedback written for that bug.
6. Knowledge components and the skill graph
Koedinger, Corbett and Perfetti (2012), Cognitive Science 36(5), 757-798, DOI 10.1111/j.1551-6709.2012.01245.x. A knowledge component is an acquired unit of cognitive function inferred from performance on related tasks. KLI's rule: match instruction to the KC's learning process. Math facts are memory/fluency KCs (retrieval practice, spacing); multi-digit procedures are induction/refinement KCs (worked examples, step feedback); fraction magnitude and place-value meaning are sense-making KCs (self-explanation, number lines, contrasting cases).
Learning Factors Analysis. Cen, Koedinger and Junker (2006), ITS 2006, DOI 10.1007/11774303_17: fit the Additive Factor Model (logit = student + sum over KCs of difficulty + learning rate x opportunities) and search KC splits and merges. Koedinger, Stamper, McLaughlin and Nixon (2013), AIED, DOI 10.1007/978-3-642-39112-5_43 (existence verified): a unit redesigned from a data-discovered KC model produced faster mastery and better learning of the targeted skills. Martin, Mitrovic, Koedinger and Mathan (2011), UMUAI 21, 249-283: a well-specified KC shows a smoothly declining error curve; a flat curve means split the KC or add a prerequisite. The solo-developer lesson: design KCs fine enough that each has one error curve, then look at the curves.
Example prerequisite graph (grade 3-5 arithmetic). Edges point from prerequisite to dependent KC. Each leaf should have its own items and its own rating.
Number sense and place value
PV1 read/write numbers to 1,000 -> PV2 to 10,000 -> PV3 to 1,000,000
PV1 -> PV4 compose/decompose (345 = 300 + 40 + 5)
PV4 -> PV5 multiply/divide by 10, 100 (place shift)
Multiplication facts (memory KCs; one KC per table, items per fact)
MF0 skip counting 2s, 5s, 10s -> MF2 x2, MF5 x5, MF10 x10
MF2 -> MF4 x4 (doubling) MF2, MF10 -> MF9 x9 (10n - n)
MF3 x3 -> MF6 x6 (double x3) MF3, MF4 -> MF7 x7, MF8 x8
MF_squares (n x n) as its own KC
MF_all -> MFC commutativity as a rule KC (a x b = b x a)
Division facts
MFk -> DFk (fact family: 6 x 7 = 42 -> 42 / 7 = 6), one KC per table
DFk -> DFr division with remainder (single digit divisor, quotient < 10)
Multi-digit multiplication (procedural KCs)
PV4, PV5, MF_all -> MD1 1-digit x 2-digit, no regrouping (partial products / area model)
MD1 -> MD2 1-digit x 2-digit with carry
MD2, PV3 -> MD3 1-digit x 3-digit
MD2 -> MD4 2-digit x 2-digit (two partial products, shift) [buggy rules: carry-order, no-shift, zero-row]
MD4 -> MD5 estimation and checking (round then multiply)
Multi-digit division
DFk, DFr, PV5 -> LD1 2-digit / 1-digit, no remainder
LD1 -> LD2 with remainder -> LD3 3-digit / 1-digit (with internal zero quotient digits)
LD3, MD4 -> LD4 by 2-digit divisor (grade 5)
Fractions (sense-making first, then procedures)
FR1 unit fractions as equal parts of a whole -> FR2 non-unit fractions a/b
FR2 -> FR3 fractions on a number line (magnitude) [whole-number-bias checks live here]
FR3 -> FR4 compare fractions with same denominator / same numerator
MF_all, DFk -> FR5 equivalent fractions (scale numerator and denominator)
FR5, FR4 -> FR6 compare unlike fractions (common denominator or benchmarks)
FR2 -> FR7 add/subtract like denominators
FR5, FR7 -> FR8 add/subtract unlike denominators
FR2, MF_all -> FR9 fraction of a set (3/4 of 20)
FR9, FR5 -> FR10 multiply fraction by whole number -> FR11 fraction x fraction
FR3, PV5 -> FR12 decimal fractions (tenths, hundredths) and fraction-decimal equivalence
Rule of thumb from LFA: if two KCs always rise and fall together, merge them; if a KC's error rate will not fall after 15-20 opportunities, split it or add a missing prerequisite (wheel-spinning; Wan and Beck 2015).
7. Learning analytics for a parent dashboard
What predicts learning in log data:
- Per-KC error curves. Declining curves are the signature of learning (Martin et al. 2011); flat curves are the earliest sign of wheel-spinning (Wan and Beck 2015: bottom-20% prerequisite knowledge wheel-spun 50% of the time vs 10% for the top 20%) or a bad KC definition.
- Practice volume and persistence. Total problems attempted and problems completed before quitting predicted yearly gains in Math Garden (Jansen et al. 2013; Journal of Intelligence 2023).
- Response-time trends on facts. Falling time at stable accuracy is the fluency signal; Hasselbring et al. (1988) showed drill builds automaticity only once a declarative network exists, so do not drill facts the child cannot derive.
- Domain rating trajectory (Math Garden reports ratings with reference groups to teachers; Brinkhuis et al. 2018).
- Gaming indicators: rapid hint sequences and very fast wrong answers (Baker et al. 2004; Shih et al. 2008).
Showing it to people: Valle, Antonenko, Dawson and Huggins-Manley (2021), BJET 52(4), DOI 10.1111/bjet.13089 (existence verified) reviews learner-facing dashboards; the recurring finding is that dashboards change behaviour more reliably than outcomes, e.g. Hellings and Haelermans (2022, Higher Education, RCT, n = 556): changed online behaviour, no exam effect. I found no peer-reviewed evaluation of a parent-facing dashboard for a children's math app. Maloney, Ramirez, Gunderson, Levine and Beilock (2015), Psychological Science 26(9), 1480-1488: first and second graders of math-anxious parents learned less and became more anxious, but only when those parents helped often with homework. Shute (2008) and Hattie and Timperley (2007) warn against normative comparison and self-level praise.
Design recommendations (evidence-informed, not tested):
- Show: minutes and days practised this week; KCs newly mastered in plain language ("x7 facts"); the current working edge ("two-digit by one-digit with carrying"); one or two "coming next" items.
- Show cautiously: fluency (seconds per fact) as a trend for mastered fact families only.
- Do not show: per-session percent correct (adaptive targeting pins it near 75% by design, which reads as "failing a quarter of the time"), peer or grade comparisons, raw ratings, live answer feeds, "errors today".
- Give the parent a script instead of data: "Ask them to show you how they did 34 x 6", steering involvement toward process talk rather than correction.
Recommended adaptive engine architecture
Principles: one learner, no fitting; hand-set priors that self-correct online; log everything so models can be refitted later.
Components
- Domain model. KC graph as above; each item tagged with 1-3 KCs, a designer difficulty prior, a time limit, generator parameters and applicable bug rules.
- Learner model. Per-KC Elo rating theta_k (logit scale, init 0), count n_k, EMA mastery score m_k, memory half-life h_k; per-fact records for memory KCs.
- Item model. delta_i = KC base prior + feature offsets (problem size for facts, number of carries for procedures), nudged online by Elo. With one learner item ratings barely move; the designer prior does the work.
- Selector. Pick a KC (frontier or review), then an item with predicted P(correct) near target (0.75; child-selectable 0.6/0.75/0.9), avoiding recent items and mixing item types.
- Feedback controller. Immediate verification; bug-rule elaborated feedback; graduated hints ending in a worked example; gaming detector.
- Scheduler. About 60-70% frontier practice, 30-40% interleaved review chosen by predicted recall.
- Event log. One row per response: timestamp, item, KCs, correct, response ms, hint level, bug matched, theta and m_k before/after, session id.
Parameters (starting values, with sources)
- K = a / (1 + b n_k), a = 1.0, b = 0.05 (Pelanek 2016), floored at 0.1; optionally reinflate with days since last practice (Klinkenberg et al. 2011).
- Target success 0.75 (Klinkenberg et al. 2011; Brinkhuis et al. 2018).
- Mastery: EMA alpha = 0.75, threshold 0.90, at least 8 attempts and 3 consecutive correct; by Pelanek and Rihak's bound N >= log_alpha(1 - T), 8 consecutive correct suffices. Stricter Cognitive-Tutor-like: T = 0.95, alpha = 0.7 (9 consecutive). Weight premature mastery 5x over-practice when tuning later.
- Fact fluency: median response time <= 3 s over the last 8 correct (practitioner convention; Hasselbring et al. 1988 for the rationale; not an experimentally derived cutoff).
- Fact time limit: 20 s visible countdown if you adopt the SRT rule; otherwise record time silently and keep it out of the rating update (Pelanek 2016).
- Review: half-life h_k in days, init 1 at mastery; correct review h_k x 2.0 (x 2.5 if fast), wrong review max(0.5, h_k x 0.5); predicted recall 2^(-days / h_k); review when <= 0.85 (HLR/FSRS retention-target logic; multipliers are Leitner-style hand values).
- Never more than two consecutive review items requiring the same procedure (Rohrer et al. 2020).
Pseudocode
def p_correct(theta, delta):
return 1 / (1 + exp(-(theta - delta)))
def K(n): # Pelanek 2016 uncertainty function
return max(0.1, 1.0 / (1 + 0.05 * n))
def update_after_response(child, item, correct, rt_ms, hint_level):
for kc in item.kcs:
theta = child.theta[kc]; delta = item.delta
p = p_correct(theta, delta)
outcome = 1.0 if (correct and hint_level == 0) else 0.0 # hinted correct does not count as a win
k = K(child.n[kc])
child.theta[kc] += k * (outcome - p)
item.delta -= k * (outcome - p) * 0.5 # damp item drift: one learner only
child.n[kc] += 1
# mastery EMA (Pelanek & Rihak 2018)
child.m[kc] = 0.75 * child.m[kc] + 0.25 * outcome
child.streak[kc] = child.streak[kc] + 1 if outcome else 0
if (not child.mastered[kc]
and child.n[kc] >= 8 and child.streak[kc] >= 3 and child.m[kc] >= 0.90
and (kc.type != "fact" or median_rt(child, kc, last=8) <= 3000)):
child.mastered[kc] = True
child.h[kc] = 1.0; child.last_review[kc] = now()
elif child.mastered[kc] and is_review_item:
child.h[kc] = child.h[kc] * (2.5 if outcome and rt_ms < item.limit/3 else 2.0) if outcome \
else max(0.5, child.h[kc] * 0.5)
child.last_review[kc] = now()
log_event(...)
def frontier(child, graph):
return [kc for kc in graph if not child.mastered[kc]
and all(child.mastered[p] for p in kc.prereqs)]
def due_for_review(child):
return [kc for kc in child.mastered_kcs()
if 2 ** (-(days_since(child.last_review[kc])) / child.h[kc]) <= 0.85]
def pick_next_item(child, target=0.75):
review = due_for_review(child)
use_review = review and (random() < 0.35 or session.frontier_run >= 4)
if use_review:
kc = min(review, key=lambda k: predicted_recall(child, k)) # most at risk first
else:
cands = frontier(child, graph)
kc = choose_weighted(cands, weight=lambda k: 1 + child.n[k] * 0.1) # bias to KC already started
pool = [i for i in kc.items if i.id not in session.recent(20)
and i.kind != session.last_item_kind or use_review] # force type mixing
scored = [(abs(p_correct(child.theta[kc], i.delta) - target), i) for i in pool]
scored.sort()
return random.choice([i for _, i in scored[:5]]) # sample near target, not argmax
def feedback(child, item, answer, rt_ms):
if answer == item.answer:
return Verification(correct=True) # brief, non-evaluative
bug = match_bug_rules(item, answer) # precomputed buggy answers
if bug and child.bug_count[bug] >= 2:
return Elaborated(bug.explanation, show_step=True) # Van der Kleij 2015: EF > KCR
if child.attempts_on(item) == 1:
return Verification(correct=False, retry=True)
return KCR(item.answer, worked_example=item.worked_steps) # second miss: show the answer + steps
def on_hint_request(child, item):
level = child.hint_level[item] + 1
if level == item.max_hint: # bottom-out hint = worked example
require_dwell(ms=4000) # Shih et al. 2008: dwell distinguishes reading from gaming
return item.hints[level]
def gaming_detector(session):
# Baker et al. 2004: rapid hint runs and fast repeated wrong answers
recent = session.last_events(6)
fast_wrong = sum(1 for e in recent if not e.correct and e.rt_ms < 1500)
hint_runs = sum(1 for e in recent if e.hint_level >= 2 and e.rt_since_hint_ms < 2000)
if fast_wrong >= 3 or hint_runs >= 2:
session.mode = "worked_example_first" # show a worked example, then a near-transfer item
session.target = 0.85 # ease difficulty briefly
Later, when you have logs from more than one child, refit in this order: (1) grid-search a, b and the EMA alpha/T against the weighted mastery-lag metric; (2) fit item difficulties with Elo or a Rasch model; (3) fit a PFA/Best-LR model per KC; (4) inspect learning curves per KC and split/merge KCs (LFA).
Contested or weak evidence
- Bloom's 2 sigma. Two dissertations with local tests; VanLehn (2011) finds human tutoring at 0.79. A motivating story, not an estimate.
- Meta-analytic ITS effects vs field RCTs. 0.42-0.76 in meta-analyses dominated by local tests and short studies; 0.18-0.20 in large RCTs on standardized outcomes. Plan for the latter.
- Mastery thresholds. 80% (Bloom), 91-100% (Kulik regression) and 0.95 (Cognitive Tutor) are different constructs (percent correct vs posterior probability); no RCT compares thresholds in a children's app; Pelanek and Rihak rely on simulation plus observational data.
- 75% success target. A Math Garden design choice validated by engagement and reliability, not by an outcome RCT; Jansen et al. (2013) found no anxiety benefit and the performance benefit ran through practice volume.
- Response-time scoring. The SRT rule assumes one response process; Brinkhuis et al. (2018) document misfit from fast-guess vs slow-solve mixtures, and a visible timer may pressure some children.
- BKT identifiability. Beck and Chang (2007) vs van de Sande (2013) disagree on framing; either way one learner cannot support parameter estimation.
- Feedback timing for children. Immediate wins in applied settings (Kulik and Kulik 1988) but can hurt right after strategy instruction (Fyfe and Rittle-Johnson 2016); it depends on prior knowledge.
- Hattie's feedback effect size. Meta-meta-analysis with heterogeneous inputs; cite the model, not the number.
- Misconception modelling. Not a significant moderator in Ma et al. (2014); bug-rule feedback rests on Van der Kleij's elaborated-feedback result.
- Parent dashboards. No peer-reviewed evaluation found; recommendations extrapolate from learner-facing dashboard reviews and Maloney et al. (2015).
- Spacing multipliers and fluency cutoffs. Hand values, not fitted; FSRS defaults are fitted on adult flashcard data.
- Not re-verified this session: VanLehn's granularity breakdown (0.31 / 0.76 / 0.40); Campbell (1997) 87.5% figure (secondary source); Ashlock's error-pattern book; Pavlik and Anderson (2005) and Corbett (2001) contents (existence only).
References
- Aleven, V., McLaren, B. M., Roll, I., & Koedinger, K. R. (2006). Toward meta-cognitive tutoring: A model of help seeking with a Cognitive Tutor. International Journal of Artificial Intelligence in Education, 16(2), 101-128.
- Aleven, V., Roll, I., McLaren, B. M., & Koedinger, K. R. (2016). Help helps, but only so much: Research on help seeking with intelligent tutoring systems. International Journal of Artificial Intelligence in Education, 26, 205-223. https://doi.org/10.1007/s40593-015-0089-1
- Badrinath, A., Wang, F., & Pardos, Z. (2021). pyBKT: An accessible Python library of Bayesian Knowledge Tracing models. Proceedings of EDM 2021.
- Baker, R. S., Corbett, A. T., Koedinger, K. R., & Wagner, A. Z. (2004). Off-task behavior in the Cognitive Tutor classroom: When students "game the system". CHI 2004, 383-390. https://doi.org/10.1145/985692.985741
- Beck, J. E., & Chang, K. (2007). Identifiability: A fundamental problem of student modeling. User Modeling 2007, LNCS 4511. https://doi.org/10.1007/978-3-540-73078-1_17
- Bloom, B. S. (1984). The 2 sigma problem: The search for methods of group instruction as effective as one-to-one tutoring. Educational Researcher, 13(6), 4-16. https://doi.org/10.3102/0013189X013006004
- Brinkhuis, M. J. S., Savi, A. O., Hofman, A. D., Coomans, F., van der Maas, H. L. J., & Maris, G. (2018). Learning as it happens: A decade of analyzing and shaping a large-scale online learning system. Journal of Learning Analytics, 5(2), 29-46. https://doi.org/10.18608/jla.2018.52.3
- Brown, J. S., & Burton, R. R. (1978). Diagnostic models for procedural bugs in basic mathematical skills. Cognitive Science, 2(2), 155-192. https://doi.org/10.1207/s15516709cog0202_4
- Campbell, J. I. D. (1997). On the relation between skilled performance of simple division and multiplication. Journal of Experimental Psychology: Learning, Memory, and Cognition, 23(5), 1140-1159. (Cited via secondary sources.)
- Cen, H., Koedinger, K. R., & Junker, B. (2006). Learning Factors Analysis - A general method for cognitive model evaluation and improvement. ITS 2006, LNCS 4053. https://doi.org/10.1007/11774303_17
- Cepeda, N. J., Pashler, H., Vul, E., Wixted, J. T., & Rohrer, D. (2006). Distributed practice in verbal recall tasks: A review and quantitative synthesis. Psychological Bulletin, 132(3), 354-380.
- Corbett, A. T. (2001). Cognitive computer tutors: Solving the two-sigma problem. User Modeling 2001, LNCS 2109. https://doi.org/10.1007/3-540-44566-8_14
- Corbett, A. T., & Anderson, J. R. (1995). Knowledge tracing: Modeling the acquisition of procedural knowledge. User Modeling and User-Adapted Interaction, 4(4), 253-278. https://doi.org/10.1007/BF01099821
- Fancsali, S. E., Nixon, T., & Ritter, S. (2013). Optimal and worst-case performance of mastery learning assessment with Bayesian Knowledge Tracing. Proceedings of EDM 2013.
- Fang, Y., Ren, Z., Hu, X., & Graesser, A. C. (2019). A meta-analysis of the effectiveness of ALEKS on learning. Educational Psychology, 39(10), 1278-1292.
- Fyfe, E. R., & Rittle-Johnson, B. (2016). Feedback both helps and hinders learning: The causal role of prior knowledge. Journal of Educational Psychology, 108(1), 82-97.
- Gervet, T., Koedinger, K., Schneider, J., & Mitchell, T. (2020). When is deep learning the best approach to knowledge tracing? Journal of Educational Data Mining, 12(3), 31-54.
- Guskey, T. R. (2010). Lessons of mastery learning. Educational Leadership, 68(2), 52-57.
- Hasselbring, T. S., Goin, L. I., & Bransford, J. D. (1988). Developing math automaticity in learning handicapped children: The role of computerized drill and practice. Focus on Exceptional Children, 20(6), 1-7.
- Hattie, J., & Timperley, H. (2007). The power of feedback. Review of Educational Research, 77(1), 81-112. https://doi.org/10.3102/003465430298487
- Hellings, J., & Haelermans, C. (2022). The effect of providing learning analytics on student behaviour and performance in programming: A randomised controlled experiment. Higher Education, 83, 1-18.
- Israni, A., Sales, A. C., & Pane, J. F. (2018). Mastery learning in practice: A (mostly) descriptive analysis of log data from the Cognitive Tutor Algebra I effectiveness trial. arXiv:1802.08616.
- Jansen, B. R. J., Louwerse, J., Straatemeier, M., Van der Ven, S. H. G., Klinkenberg, S., & van der Maas, H. L. J. (2013). The influence of experiencing success in math on math anxiety, perceived math competence, and math performance. Learning and Individual Differences, 24, 190-197. https://doi.org/10.1016/j.lindif.2012.12.014
- Khajah, M., Lindsey, R. V., & Mozer, M. C. (2016). How deep is knowledge tracing? Proceedings of EDM 2016. arXiv:1604.02416
- Klinkenberg, S. (2014). High speed high stakes scoring rule: Assessing the performance of a new scoring rule for digital assessment. In CAA 2014, Springer CCIS 439. https://doi.org/10.1007/978-3-319-08657-6_11
- Klinkenberg, S., Straatemeier, M., & van der Maas, H. L. J. (2011). Computer adaptive practice of maths ability using a new item response model for on the fly ability and difficulty estimation. Computers & Education, 57(2), 1813-1824. https://doi.org/10.1016/j.compedu.2011.02.003
- Koedinger, K. R., Corbett, A. T., & Perfetti, C. (2012). The Knowledge-Learning-Instruction framework: Bridging the science-practice chasm to enhance robust student learning. Cognitive Science, 36(5), 757-798. https://doi.org/10.1111/j.1551-6709.2012.01245.x
- Koedinger, K. R., Stamper, J. C., McLaughlin, E. A., & Nixon, T. (2013). Using data-driven discovery of better student models to improve student learning. AIED 2013, LNCS 7926. https://doi.org/10.1007/978-3-642-39112-5_43
- Kulik, C.-L. C., Kulik, J. A., & Bangert-Drowns, R. L. (1990). Effectiveness of mastery learning programs: A meta-analysis. Review of Educational Research, 60(2), 265-299. https://doi.org/10.3102/00346543060002265
- Kulik, J. A., & Fletcher, J. D. (2016). Effectiveness of intelligent tutoring systems: A meta-analytic review. Review of Educational Research, 86(1), 42-78. https://doi.org/10.3102/0034654315581420
- Kulik, J. A., & Kulik, C.-L. C. (1988). Timing of feedback and verbal learning. Review of Educational Research, 58(1), 79-97. https://doi.org/10.3102/00346543058001079
- Lindsey, R. V., Shroyer, J. D., Pashler, H., & Mozer, M. C. (2014). Improving students' long-term knowledge retention through personalized review. Psychological Science, 25(3), 639-647. https://doi.org/10.1177/0956797613504302
- Lortie-Forgues, H., Tian, J., & Siegler, R. S. (2015). Why is learning fraction and decimal arithmetic so difficult? Developmental Review, 38, 201-221.
- Ma, W., Adesope, O. O., Nesbit, J. C., & Liu, Q. (2014). Intelligent tutoring systems and learning outcomes: A meta-analysis. Journal of Educational Psychology, 106(4), 901-918. https://doi.org/10.1037/a0037123
- Maloney, E. A., Ramirez, G., Gunderson, E. A., Levine, S. C., & Beilock, S. L. (2015). Intergenerational effects of parents' math anxiety on children's math achievement and anxiety. Psychological Science, 26(9), 1480-1488. https://doi.org/10.1177/0956797615592630
- Maris, G., & van der Maas, H. (2012). Speed-accuracy response models: Scoring rules based on response time and accuracy. Psychometrika, 77, 615-633. https://doi.org/10.1007/s11336-012-9288-y
- Martin, B., Mitrovic, A., Koedinger, K. R., & Mathan, S. (2011). Evaluating and improving adaptive educational systems with learning curves. User Modeling and User-Adapted Interaction, 21, 249-283. https://doi.org/10.1007/s11257-010-9084-2
- Ni, Y., & Zhou, Y.-D. (2005). Teaching and learning fraction and rational numbers: The origins and implications of whole number bias. Educational Psychologist, 40(1), 27-52. https://doi.org/10.1207/s15326985ep4001_3
- Pane, J. F., Griffin, B. A., McCaffrey, D. F., & Karam, R. (2014). Effectiveness of Cognitive Tutor Algebra I at scale. Educational Evaluation and Policy Analysis, 36(2), 127-144. https://doi.org/10.3102/0162373713507480
- Pavlik, P. I., & Anderson, J. R. (2005). Practice and forgetting effects on vocabulary memory: An activation-based model of the spacing effect. Cognitive Science, 29(4), 559-586.
- Pavlik, P. I., & Anderson, J. R. (2008). Using a model to compute the optimal schedule of practice. Journal of Experimental Psychology: Applied, 14(2), 101-117.
- Pavlik, P. I., Cen, H., & Koedinger, K. R. (2009). Performance Factors Analysis - A new alternative to knowledge tracing. AIED 2009, Frontiers in AI and Applications 200, 531-538.
- Pelanek, R. (2016). Applications of the Elo rating system in adaptive educational systems. Computers & Education, 98, 169-179. https://doi.org/10.1016/j.compedu.2016.03.017
- Pelanek, R., & Rihak, J. (2018). Analysis and design of mastery learning criteria. New Review of Hypermedia and Multimedia, 24(3), 133-159. (Preprint read; journal details from the preprint header "nrhm".)
- Piech, C., Bassen, J., Huang, J., Ganguli, S., Sahami, M., Guibas, L., & Sohl-Dickstein, J. (2015). Deep knowledge tracing. NIPS 28.
- Ritter, S., Yudelson, M., Fancsali, S. E., & Berman, S. R. (2016). How mastery learning works at scale. L@S 2016, 71-79. https://doi.org/10.1145/2876034.2876039
- Rohrer, D., Dedrick, R. F., & Stershic, S. (2015). Interleaved practice improves mathematics learning. Journal of Educational Psychology, 107(3), 900-908.
- Rohrer, D., Dedrick, R. F., Hartwig, M. K., & Cheung, C.-N. (2020). A randomized controlled trial of interleaved mathematics practice. Journal of Educational Psychology, 112(1), 40-52. https://doi.org/10.1037/edu0000367
- Roschelle, J., Feng, M., Murphy, R. F., & Mason, C. A. (2016). Online mathematics homework increases student achievement. AERA Open, 2(4). https://doi.org/10.1177/2332858416673968
- Settles, B., & Meeder, B. (2016). A trainable spaced repetition model for language learning. Proceedings of ACL 2016, 1848-1858. https://aclanthology.org/P16-1174/
- Shih, B., Koedinger, K. R., & Scheines, R. (2008). A response time model for bottom-out hints as worked examples. Proceedings of EDM 2008, 117-126.
- Shute, V. J. (2008). Focus on formative feedback. Review of Educational Research, 78(1), 153-189. https://doi.org/10.3102/0034654307313795
- Siegler, R. S., Thompson, C. A., & Schneider, M. (2011). An integrated theory of whole number and fractions development. Cognitive Psychology, 62(4), 273-296.
- Valle, N., Antonenko, P., Dawson, K., & Huggins-Manley, A. C. (2021). Staying on target: A systematic literature review on learner-facing learning analytics dashboards. British Journal of Educational Technology, 52(4). https://doi.org/10.1111/bjet.13089
- van de Sande, B. (2013). Properties of the Bayesian Knowledge Tracing model. Journal of Educational Data Mining, 5(2), 1-10.
- Van der Kleij, F. M., Feskens, R. C. W., & Eggen, T. J. H. M. (2015). Effects of feedback in a computer-based learning environment on students' learning outcomes: A meta-analysis. Review of Educational Research, 85(4), 475-511. https://doi.org/10.3102/0034654314564881
- VanLehn, K. (1990). Mind Bugs: The Origins of Procedural Misconceptions. MIT Press.
- VanLehn, K. (2011). The relative effectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems. Educational Psychologist, 46(4), 197-221. https://doi.org/10.1080/00461520.2011.611369
- Wan, H., & Beck, J. B. (2015). Considering the influence of prerequisite performance on wheel spinning. Proceedings of EDM 2015.
- Ye, J., Su, J., & Cao, Y. (2022). A stochastic shortest path algorithm for optimizing spaced repetition scheduling. KDD 2022, 4381-4390. https://doi.org/10.1145/3534678.3539081
- Yeung, C.-K., & Yeung, D.-Y. (2018). Addressing two problems in deep knowledge tracing via prediction-consistent regularization. L@S 2018. arXiv:1806.02180