Research 11: Calibrating and evaluating Numberkit with very few learners

Date: 2026-09-27. Scope: two practical questions for an app with one child user now and perhaps a few dozen families and one school class later. (A) How can the hand-set parameters in docs/PLAN.md section 6 be calibrated or validated with so few learners? (B) How can the app tell, honestly, whether it works at n = 1 to a few dozen?

Verification key: [verified] means checked against the primary text or its official page during this review. [from memory] means a well-known source whose details I am confident of but did not re-open. [unverified] means the detail must be checked before it is quoted. The web search budget ran out partway through, so some secondary details carry that mark on purpose.

Summary

  1. Most of the parameters can be checked against their own predictions with one child. The scheduler claims that a review served at predicted recall 0.85 succeeds 85 percent of the time. Elo claims Play succeeds about 75 percent of the time. The mastery rule claims mastered facts survive the two-week delayed probe. Each is a calibration check needing about 100 to 300 responses, which one child supplies within weeks.
  2. Revisit the scheduler first. Its multipliers are hand values, the plan's own simulation shows it over-predicting recall (about 0.95 predicted against 0.73 to 0.80 simulated), and it moves retention more than any other parameter. Pelánek (2016) found the Elo constants barely matter. Pelánek and Řihák (2017) found the mastery threshold matters more than the learner model.
  3. Reviews at 0.85 are good for learning but poor for measurement. By my own Fisher-information calculation, pinning one child's half-life scale to a factor of about 1.4 takes about 200 reviews at 0.85, but about 70 at 0.6. The delayed probe is therefore the richest calibration data the app collects. With several children, hierarchical pooling (Lindsey et al. 2014) is the standard remedy.
  4. FSRS defaults do not transfer. They were fitted on about 10,000 adult Anki collections. Refit FSRS on Numberkit's own log once there are several hundred reviews per child across a few children.
  5. For evidence that the app works, use single-case logic that the app already half-contains. Tables are introduced in a staggered order, which is a multiple-baseline design across fact sets. Add item-level randomisation of features for within-child comparisons. Judge the results against the WWC single-case standards (Kratochwill et al. 2010).
  6. Transfer must be measured off the app, with paper probes in digits correct per minute (Burns et al. 2006), standard-notation items, and delayed retention.
  7. One comparison class supports no causal claim. Frame a class pilot as a feasibility study with within-class designs. Preregister it, get parental opt-in and child assent, and write a DPIA.

A. Calibrating with few learners

A.0 Validate predictions rather than fit parameters

One child does not supply the variation across people that fitting needs. What one child does supply is a stream of predictions (already logged as "model state before") and outcomes. The n = 1 tool is therefore the calibration check: bin responses by predicted probability and compare the prediction with the observed rate. Brinkhuis et al. (2018) plot exactly this for one Math Garden child over three years (20,392 responses): the observed-minus-expected residuals narrowed as the child "increasingly conforms to the scoring rule" [verified].

How much data a check needs follows from binomial arithmetic (my calculation):

  • Near 0.75, the standard error of an observed success rate is 0.061 at 50 responses, 0.043 at 100, and 0.031 at 200.
  • Near 0.85 it is 0.050, 0.036, and 0.025.
  • So a 10-point miscalibration shows at about 100 responses, and a 5-point one needs about 300.

A.1 Elo, K = 1/(1 + 0.05n)

Provenance. The uncertainty function U(n) = a/(1 + bn) with a = 1 and b = 0.05 comes from Papoušek, Pelánek and Stanislav (2014), found by grid search on a large geography system. Pelánek (2016) writes that "this exact choice of parameter values is not important, many different choices of a, b provide very similar results." He adds that items usually get far more answers than students, so "it may be useful to use different uncertainty function for items and for students." He also reports that principled Bayesian uncertainty "do[es] not bring significant advantages" [all verified]. Math Garden's uncertainty falls with answers and rises with days away (Klinkenberg et al. 2011, as summarised by Pelánek 2016) [verified].

What to check. Build a reliability diagram of the Play predictions and track the rolling success rate against 0.75. About 100 responses suffice.

The few-learner problem.

  • With one child, item difficulties never converge; the plan says they "barely move from their designer priors."
  • Items are generated, so difficulty belongs on the template or fact, and the designer prior is effectively the model.
  • Once several children have answered each template, refit template difficulties offline with a Rasch or logistic model, and use the result as the new priors.
  • Give items a smaller K than learners, so one child's bad day does not rewrite difficulties that other children inherit.

Pitfalls.

  • Items are chosen from current estimates, so the data are adaptively collected and offline refits can be biased (Pelánek, Řihák and Papoušek 2016) [from memory].
  • Math Garden found that adaptive selection rewarded local strategies that worked only on the items presented, and fixed it by making item order less adaptive (Brinkhuis et al. 2018) [verified]. Generator families invite the same thing (for example, "the answer ends in 0").

A.2 Target success 0.75

Math Garden lets children choose easy (about 90 percent correct), medium (75), or hard (60) (Brinkhuis et al. 2018) [verified]. Jansen et al. (2013) found no reduction in maths anxiety after 3 to 6 weeks, with performance gains mediated by the number of items attempted [verified through summaries of the abstract; condition-level results unverified]. The plan already states that 0.75 is "validated by engagement and reliability rather than by an outcome RCT."

The target is a motivational design choice, so calibrating it means two checks:

  • Does the system deliver about 0.75? That is the Elo check.
  • Does practice look healthy at 0.75? Watch completion, return rate, "I don't know" use, and gaming-detector firings.

The child-selectable 0.6, 0.75, and 0.9 setting is a latent experiment. Log every change, and at 20 or more children randomise the default. Do not use the "85 percent rule" (Wilson et al. 2019) [from memory] as support for any target: it is a theoretical result for model learners doing binary classification, not evidence about children.

A.3 Half-life scheduler

The rule. The half-life starts at 1 day at mastery. A correct review multiplies it by 2 (by 2.5 if fast). A wrong review halves it, with a floor of 0.5. A review falls due when 2^(-t/h) reaches 0.85. The model form is Half-Life Regression, which Duolingo fitted on 12.9 million instances, cutting prediction error by 45 percent or more against baselines (Settles and Meeder 2016) [verified]. The multipliers are hand values.

The built-in test. Because reviews fire at predicted recall 0.85, first-attempt review accuracy is a direct test of the model:

  • about 95 percent correct means the half-lives are too short and the child is over-reviewing;
  • about 70 percent means they are too long. Two details matter. Log the predicted recall at the moment the review is served, because the cap and absences delay reviews. And split the check by review number (first after mastery, second, third and later), since each multiplier makes its own claim.

How much data one child's curve needs. For p = 2^(-t/h) and theta = ln h, the information one review carries about theta is p (ln p)^2 / (1 - p). This is my derivation, assuming the model form is right and reviews are independent. It gives 0.15 at p = 0.85, 0.25 at 0.75, 0.39 at 0.6, and 0.48 at 0.5. The 95 percent precision on one child's half-life scale follows:

Reviews Precision at p = 0.85 Precision at p = 0.6
50 within a factor of 2.05 within a factor of 1.56
100 1.66 1.37
200 1.43 1.25
400 1.29 not computed

Real data will need more than this idealised bound. The practical lesson: the two-week delayed probe, which hits facts at lower predicted recall, should be logged with its prediction and analysed as calibration data, not only as an outcome.

Borrowing strength.

  • Sense et al. (2016) found that an individual's rate of forgetting is stable over time but differs across materials [verified]. So estimate a per-child half-life multiplier separately for each KC family, not as one global trait.
  • Lindsey, Shroyer, Pashler and Mozer (2014) combined a memory model with hierarchical Bayesian inference across students and items. In a middle-school language course, personalised review raised retention by 16.5 percent over massed study and by 10.0 percent over one-size-fits-all spacing [verified].
  • The structure to copy: log half-life = population mean + child effect + fact effect + a review-history term, with priors on the spreads. With one child this collapses to informative priors. With 5 to 20 children, partial pooling weights each child's estimate by how much data they have.
  • Fit it offline in Stan or brms on an exported log, and let the engine read refreshed priors from configuration. That keeps the engine pure.
  • Van der Velde, Sense, Borst and van Rijn (2021) [from memory] use pooled item estimates to warm-start new learners, which is the same move.

FSRS. FSRS (Ye, Su and Cao 2022) [from memory] has defaults "obtained by running each algorithm on 10 thousand collections of Anki users," about 350 million reviews after filtering. The benchmark page mentions no children and no maths [verified]. On how much data a per-user fit needs:

  • The Anki manual cites "less than a few hundred" reviews as a reason FSRS underperforms [verified].
  • The optimizer's minimum fell from 1,000 reviews to 400.
  • A 2024 community notebook argued that fitting beats the defaults after about 16 reviews (ankitects/anki issue 3094) [verified as a claim; not peer reviewed].
  • Numberkit's simulation found the defaults reviewed about five times less often and lost 15 to 20 points of retention, but only under a hand-set child model.

So keep the ladder. Once there are several hundred reviews per child, fit FSRS and compare it with the ladder on log loss and calibration, using a time-ordered held-out split.

A.4 Mastery rule (EMA 0.75, threshold 0.9, at least 8 attempts, 3 consecutive correct)

Pelánek and Řihák (2017) found that "the choice of data sources used for mastery decision and setting of thresholds are more important than the choice of a learner modeling technique," and recommend an exponential moving average [verified]. Related work compares knowledge tracing with N-consecutive-correct rules (Kelly et al. 2015) and shows how BKT mastery decisions can fail in worst cases (Fancsali, Nixon and Ritter 2013) [both from memory].

A mastery rule is a classifier whose criterion is later performance:

  • Delayed-probe accuracy. This confounds the mastery rule with the scheduler. To separate them, give a random fifth of newly mastered facts one extra probe at a fixed delay (say 3 days) outside the scheduler.
  • Attempts to mastery. If most KCs pass at exactly 8 attempts with perfect accuracy, the floor, not the EMA, is deciding, and known facts are being over-practised.
  • Offline replay. Replay the log at thresholds 0.8, 0.85, 0.9, and 0.95, and compare the later accuracy of KCs that would have been declared mastered at one threshold and not another. This is Pelánek and Řihák's method, and it works on one child's log.

Practice stops at mastery, so later data are truncated. Learning curves drawn from mastery-gated data are biased (Murray et al. 2013) [unverified].

A.5 Knowledge-tracing parameter recovery

Beck and Chang (2007) showed that Bayesian knowledge tracing (BKT) fits the same data with many parameter sets that make different claims about knowledge. They proposed Dirichlet priors as the remedy [verified]. Doroudi and Brunskill (2017) argue the real problem is degenerate fits rather than strict non-identifiability [from memory], and later showed that unmodelled differences in learning rate make mastery decisions inequitable (2019) [from memory]. Either way, per-child BKT parameters cannot be recovered from a few dozen responses; individualised BKT (Pardos and Heffernan 2010; Yudelson et al. 2013) [from memory] needs large data. Simple models match deep knowledge tracing on prediction (Khajah, Lindsey and Mozer 2016) [from memory].

Keep Elo plus the EMA, and keep priors. Evaluate with a time-ordered split, reporting log loss and calibration, because evaluation choices can change which model wins (Pelánek 2018) [from memory].

A.6 Simulation: what it can and cannot settle

"Simulation" covers two different tools.

Simulation-based calibration (Talts et al. 2018) [verified] validates inference code. Draw parameters from the prior, simulate data shaped like the app's log, fit, and check that the ranks of the true values are uniform. Run it on the offline hierarchical half-life model before trusting any refit. It also shows recoverability directly: if 100 simulated reviews leave a child's forgetting multiplier as uncertain as its prior, real data will not pin it either.

Policy simulation (the existing packages/sim) catches bugs and gross design errors, but its answers are conditional on the simulated learner. Decision 0001 says so: the comparison "says which scheduler suits a child who forgets at that rate, not what the rate is." The remedy is Doroudi, Aleven and Brunskill's (2017) robust evaluation matrix [from memory]:

  • evaluate each policy under several plausible learner models;
  • prefer policies that do well under all of them;
  • vary the simulated forgetting, learning, lapse, and guessing rates;
  • where the ranking flips, that child-model parameter is the one to measure on the first child first.

A.7 Response time: fluency cutoffs and the lapse rule

Fluency cutoffs (3 s, 1.5 s). A fixed cutoff mixes retrieval with typing speed. Children also mix strategies whose response times overlap, and Siegler (1987) showed that averaging over strategies misrepresents children's arithmetic [from memory]. Three calibrations work with one child:

  1. A motor floor. Measure times on trivial items (x1, x10, copying a number) and express the cutoffs relative to that floor.
  2. Strategy spot-checks. On about 1 item in 40 among phase-2 and phase-3 facts, ask "How did you get that?" with picture choices: just knew it, counted, used another fact. After about 100 reports, set the cutoff where "just knew it" dominates. This needs the parent's agreement, because it interrupts practice.
  3. The log response-time distribution by phase. Retrieval and procedural answers often show as two overlapping modes.

Formal response-time models.

  • Lognormal speed and time-intensity models (van der Linden 2007) [from memory] make "fast for this child on this fact" a residual, and they suit the self-referenced timing rule.
  • Ex-Gaussian fits (Heathcote, Popiel and Mewhort 1991) [from memory] isolate the slow tail (tau), which is where children with attention difficulties differ (Kofler et al. 2013) [from memory]. They need on the order of 100 or more RTs per condition [unverified figure], and the parameters do not map cleanly onto processes (Matzke and Wagenmakers 2009) [from memory].
  • Diffusion models (Ratcliff and McKoon 2008; EZ-diffusion, Wagenmakers et al. 2007) [from memory] are built for two-choice tasks and need many trials (Lerche, Voss and Nagler 2017) [from memory], so they fit typed answers poorly. Their legacy is Math Garden's scoring rule (Maris and van der Maas 2012) [from memory], which shows time pressure to the child and so conflicts with the plan's rule 3.

The probable-lapse rule (over 3 s and over 2.5 times the running median). Research 06 records that no validated lapse detector exists, but the rule can be checked on its consequences:

  • Predictive check. A real lapse should be followed by a normal answer the next time the fact appears. Compare that next answer after lapse-flagged responses with the next answer after slow responses that were not flagged.
  • Randomised check (at 20 or more children). Among eligible responses, apply the rule to a random half and compare delayed recall. Both arms are defensible practice. It needs about 50 or more flagged events per arm.
  • Mixture model (later). Fit a lognormal retrieval component plus a broad lapse component, and compare its classifications with the rule's.

A.8 How Math Garden and Duolingo tuned theirs

Math Garden combines Elo with a time-based scoring rule and an uncertainty that decays with answers and grows with time away. Children choose among three difficulty levels. Over more than a billion responses, the team found misfit through residual diagnostics per child and per item (guessing, local strategies, strategic skipping). It also ran A/B tests; one showed that delaying the skip button increased effortful practice (Savi et al. 2018) (all in Brinkhuis et al. 2018) [verified]. The residual diagnostics work per child and transfer. The A/B tests need scale.

Duolingo fitted HLR by regression on millions of review traces and tested it with A/B experiments on about 1 million and 3.3 million students. The outcome was daily retention (whether users came back), and HLR raised daily engagement by 12 percent (Settles and Meeder 2016) [verified]. What transfers is the model form and the predicted-versus-observed evaluation. The tuning method does not transfer, and the success metric was engagement, not learning. Duolingo's later "Birdbrain" work is described only on its blog [unverified].

A.9 Sensitivity: where effort pays

The existing evidence points the same way:

  • the Elo constants barely matter (Pelánek 2016);
  • the mastery threshold matters more than the learner model (Pelánek and Řihák 2017);
  • the choice of scheduler moved simulated retention by 15 to 20 points (decision 0001).

To confirm this for Numberkit, run Morris elementary-effects screening (Morris 1991; Saltelli et al. 2008) [from memory] on the simulator:

  • Parameters: the K constants, target rate, EMA alpha and threshold, minimum attempts, consecutive-correct count, initial half-life, correct, fast, and wrong multipliers, review threshold, fluency cutoffs, and lapse multiplier.
  • Child profiles: the robust evaluation matrix of A.6.
  • Outcomes: delayed retention, review load, sessions to mastery, and the share of practice spent on known facts.

The expected ranking, to be tested rather than assumed:

  • highest: the scheduler (review threshold and multipliers) and the mastery threshold with its minimum attempts;
  • middle: the fluency cutoffs;
  • lowest: Elo K.

A parameter whose best value flips across child profiles needs the first child's data. One that moves nothing can be left alone.

B. Evaluating with few learners

Three claims need separate evidence:

  1. The child learns what they practise (acquisition).
  2. He keeps it (retention).
  3. He uses it off the app (transfer).

The log bears on the first two. Only off-app measures answer the third. A claim of being better than TTRS or school practice needs a comparison, and at n = 1 the only credible comparisons are within the child: across skills, across time, or across randomised items.

B.1 Single-case experimental designs

The WWC pilot standards (Kratochwill, Hitchcock, Horner, Levin, Odom, Rindskopf and Shadish 2010) [verified] require the following:

  • Demonstrations. At least three attempts to demonstrate an effect, at three different points in time. AB, ABA, and BAB designs do not qualify.
  • Phase length. A phase needs at least 3 points to count.
  • ABAB. Four phases with at least 5 points each to meet standards, or 3 each to meet them with reservations.
  • Multiple baseline. At least six phases with at least 5 points each (3 with reservations).
  • Alternating treatments. Five repetitions of the alternation to meet standards, four with reservations.
  • Inter-assessor agreement. Checked on at least 20 percent of sessions, reaching 0.80 to 0.90 by percentage agreement or at least 0.60 by kappa.
  • Evidence-based rating across studies. Five studies, from three teams in three locations, with 20 experiments in total.
  • Convention. The three-demonstration criterion is explicitly "based on professional convention."

The journal version is Kratochwill et al. (2013) [from memory]. The WWC handbook has since been revised (version 5.0, 2022) [unverified details].

Multiple baseline across skills (the best fit).

  • What it answers: whether gains on a fact set follow its teaching, rather than time, maturation, or school.
  • Minimum data: three or more tiers, each probed at least 5 times before and 5 times after its introduction; roughly 8 to 12 weekly probes.
  • Event-log implementation:
    • Define the tiers as fact sets closed under commutativity, so that 6 x 7 and 7 x 6 sit in the same tier. For example: the 6s and 7s cross facts, the 8s and 9s, and the 11s and 12s.
    • Run a short weekly probe on all tiers, tagged with probe_tier.
    • Stagger when Forge teaching starts on each tier, and randomise the start points in advance (Kratochwill and Levin 2010, "randomization to the rescue") [from memory].
    • Randomised start points permit a randomisation test (Bulté and Onghena 2008) [from memory].
  • Pitfalls:
    • Derived-fact strategies spread across tiers, which biases toward null.
    • School teaching may introduce a tier early, so ask the parent what the class covered.
    • Deciding phase changes by looking at the data inflates false positives.

Multiple baseline across learners. Families start on randomised, staggered dates, and each child's pre-access paper probes form the baseline. The risk is attrition during a long baseline.

ABAB. Learning does not reverse, so ABAB suits only features with reversible effects, such as the silent latency-spike rule's effect on errors.

Adapted alternating treatments (Sindelar, Rosenberg and Wilson 1985) [from memory]. Two methods are applied to matched, equivalent fact sets in quick alternation, for example with or without a Learn segment, or with the fast-review multiplier at x2.5 or x2. It needs matched sets and five alternations.

Analysis. Visual analysis has only moderate inter-rater agreement (Ninci et al. 2015) [from memory], so supplement it with:

  • randomisation tests;
  • a between-case standardised mean difference once there are 3 or more cases (Hedges, Pustejovsky and Shadish 2012; Pustejovsky et al. 2014) [from memory];
  • Tau-U as a descriptive overlap index (Parker et al. 2011) [from memory];
  • multilevel models across cases (Shadish, Kyse and Rindskopf 2013) [from memory].

There is no agreed single-case effect size (Kratochwill et al. 2010) [verified]. Report against the SCRIBE checklist (Tate et al. 2016) [from memory].

B.2 N-of-1 trials

Medical n-of-1 trials are randomised crossovers within one person (Kravitz, Duan et al. 2014; CENT, Vohra et al. 2015) [from memory]. They assume the effect washes out, which learning does not. Use them only for effects that are immediate and reversible, such as session length or time of day on accuracy, with a preset block structure (for example, six pairs of weeks with the order randomised within each pair). For learning features, use item-level randomisation (B.3).

B.3 Randomised within-learner comparisons

Precedent. Lindsey et al. (2014) compared review strategies within students by assigning material to conditions [verified at abstract level]. Micro-randomised trials (Klasnja et al. 2015) [from memory] randomise options many times per person.

Implementation.

  • When a fact family enters the relevant state (introduction, mastery, or first review), assign an arm with a seeded draw.
  • Store experiment_id, arm, and seed on the fact's model state, so every later response carries them.
  • Randomise by fact family (a x b, b x a, and the related division facts) to limit spillover.
  • Analyse with a permutation test at n = 1. With several children, use a mixed logistic model: correct ~ arm + (1 | fact) + (1 | child).

Minimum data. My rough power reasoning, not a formal calculation: with about 30 fact families per arm for one child, only differences of about 15 to 20 points show; pooled over 20 children, differences of 5 to 10 points become detectable.

Pitfalls.

  • Spillover between arms biases toward null.
  • Features that act on the whole session cannot be randomised per item.
  • Run one or two experiments at a time, preregistered.
  • Both arms must be defensible practice (equipoise).

B.4 Curriculum-based measurement

Curriculum-based measurement (CBM; Deno 1985) [from memory] scores short timed probes in digits correct per minute (DCPM).

  • Criteria. Burns, VanDerHeyden and Jiban (2006) put the instructional range at 14 to 31 DCPM for grades 2 and 3, and 24 to 49 for grades 4 and 5 [verified through secondary citations].
  • Data needed. A progress slope needs about 6 to 10 weekly points; this is best studied for reading (Christ et al. 2013), less so for maths computation (Christ et al. 2008) [both from memory].
  • Implementation.
    • Alternate the plan's weekly one-minute probe between the app and paper, so mode effects can be estimated. Online and paper scores differ (Backes and Cowan 2019) [from memory].
    • Use parallel forms from a fixed blueprint.
    • Log the form, mode, and DCPM, with the parent entering paper answers item by item.
    • Keep a control strand the app has not yet taught (division before Phase 2, column subtraction before Phase 3) to expose practice effects and maturation.
  • Pitfalls. The norms are US and dated. Tablet DCPM includes typing. The personal-best chart makes the probe partly a motivational tool, so it should not be the only evidence.

B.5 Growth modelling for small samples

  • At n = 1, model the weekly series as a short interrupted time series: level and slope, first-order autocorrelation, and breaks at introduction points. Hofman et al. (2018) argue that adaptive practice data make within-person (idiographic) measurement feasible because they give dense, high-quality individual time series [verified].
  • At 5 to 50 children, use multilevel growth models with REML and Kenward-Roger corrections, or Bayesian estimation with considered, weakly informative priors on the variances. Default diffuse priors can do worse than frequentist methods at small n (McNeish 2016; McNeish and Stapleton 2016) [from memory].
  • Reporting. Report intervals rather than significance tests, and treat dropouts as data.

B.6 Transfer to paper and standard notation

Transfer here forms a ladder (following Barnett and Ceci's 2002 taxonomy) [from memory]:

  1. the same facts on paper;
  2. standard notation, including missing-factor forms (7 x ? = 56) and division;
  3. extended facts and word problems (the plan's session-3 set);
  4. school measures such as England's Multiplication Tables Check.

Print a paper pack from the same generators every 4 to 6 weeks and log it with mode = paper. Have a second adult score 20 percent of the sheets for agreement. In-game progress did not transfer for DragonBox or ST Math (Research 03), so a null result on transfer is a finding and should be reported.

B.7 Delayed retention

  • Keep the two-week delayed probe, and add a 6 to 8 week probe on a sample of mastered facts, since two weeks is short against school holidays.
  • Keep probe items out of review for a few days beforehand.
  • Log the time since last practice for every probe item. Accuracy against that lag is both the child's real forgetting curve (A.3) and a retention outcome.

B.8 Regression to the mean and practice effects

Regression to the mean. Facts chosen for practice because they were missed once will partly "improve" by chance (Barnett, van der Pols and Dobson 2005) [from memory]. To control it:

  • define "unknown" by two baseline passes, which the three-session baseline allows;
  • compare practised facts with missed facts not yet practised, which is the multiple-baseline logic again;
  • in a class pilot, never compare the weakest children with the rest.

Practice effects. Use parallel forms, keep probe items out of recent practice, and keep a control strand. A baseline that rises before teaching begins is the signature of a practice effect.

B.9 Preregistration

Preregistration separates confirmatory from exploratory analysis (Nosek et al. 2018) [from memory], and it works for single-case designs (Johnson and Cook 2019) [unverified pages]. Register on OSF, with an embargo if wanted:

  • the primary outcomes (paper DCPM on taught tables at 8 weeks, and delayed-probe accuracy);
  • the tiers and their randomised start points;
  • the item-level experiments;
  • the analysis methods;
  • exclusion and stopping rules;
  • the rules that will trigger a parameter refit.

The last item prevents tuning on the outcome (Gelman and Loken 2013) [from memory]. Commit to reporting whatever the results show, as the plan's section 15.3 already does.

B.10 Ethics and consent

Using an app with your own child is product use. A pilot whose results will be reported is research, and ethical review is hard to obtain after the fact.

UK.

  • BERA's guidelines (4th edition 2018; a 2024 5th edition is [unverified]) [from memory] require informed parental consent, the child's own assent, the right to withdraw without penalty, minimal burden, and honest reporting.
  • The UK age of digital consent is 13 (Data Protection Act 2018, s. 9) [from memory].
  • The ICO Children's Code (in force September 2021) [from memory] requires high-privacy defaults, data minimisation, no nudges, and a DPIA.
  • No statutory ethics committee covers non-NHS educational research, but an independent review (a university partner or a commercial board) strengthens publication and a school pitch.
  • For a school: the headteacher's gatekeeper consent, parental opt-in, a data processing agreement, and a shared DPIA.

US.

  • COPPA (16 CFR 312) requires verifiable parental consent under 13. It was amended in 2025, with compliance in 2026 [dates unverified]. A school may consent only for the educational context [from memory].
  • FERPA's school-official exception needs a contract, and PPRA covers surveys on sensitive topics [from memory]. State laws such as California's SOPIPA apply too [check per state].
  • The Common Rule, with Subpart D for children, binds federally funded or institutionally engaged research. Exemption 46.104(d)(1), for normal educational practice, may fit a school pilot, but whether it applies is the review board's decision [from memory].

What consent materials must say.

  • Response time is recorded silently, and why.
  • Item-level randomisation happens, and both arms are ordinary practice.
  • The comparison class continues its usual practice and is offered the app afterwards (a waitlist design).
  • Results are reported whatever they show.
  • Families can export and delete their data.

The one comparison class. Two intact classes give n = 2 at the level of assignment. Teacher and class are confounded with the treatment, and intraclass correlations for maths achievement run around 0.2 (Hedges and Hedberg 2007) [from memory]. Report the class comparison descriptively, as feasibility and fidelity. Get causal evidence from within-class designs: staggered starts by randomly ordered table groups, and item-level randomisation. Use the comparison class to estimate probe practice effects. The plan's section 15 design is sound, but it should say that the between-class difference is not causal.

A staged plan for Numberkit

Stage 0: before the first child's next sessions (engineering only)

  1. Make sure the log carries what calibration needs.
    • Every response: the Elo prediction, the predicted recall at the moment of serving, the time since last practice, the EMA before and after, the phase, the response time, and a flag for trivial items.
    • Every probe: the form, mode, tier, and time since the item was last practised.
  2. Write an offline calibration report from an export. It should show:
    • the Elo reliability diagram and Play success rate;
    • review accuracy against 0.85, split by review number;
    • accuracy of mastered KCs at the delayed probe;
    • attempts to mastery;
    • response-time distributions by phase;
    • next-occurrence outcomes after lapse flags.
  3. Extend the simulator into a robust evaluation matrix with Morris screening, and list the parameters whose best value flips.
  4. Run SBC on the planned hierarchical half-life model, to learn how many reviews recover a child's multiplier.
  5. Define the commutativity-closed tiers, randomise their start points in advance, and write a private one-page preregistration.

Stage 1: the first child, weeks 1 to 12 (n = 1)

Measure:

  • At baseline: the planned three-session baseline, a paper probe in week 1, and a second pass on missed facts.
  • Weekly: tier probes on parallel forms, and the DCPM probe, alternating app and paper.
  • Every two weeks: the delayed probe, logged with its predicted recall.
  • With the parent's agreement: a strategy report on about 1 item in 40.

Decide, in this order:

  1. Scheduler. After about 100 scheduled reviews, if first-attempt accuracy is outside roughly 0.75 to 0.93, adjust the global half-life scale, not the individual multipliers. After about 200 reviews plus delayed probes, check the split by review number.
  2. Fluency cutoffs. After about 100 strategy reports, reset them relative to the child's motor floor.
  3. Mastery. Replay the log at thresholds 0.8 to 0.95, and inspect KCs that passed at exactly 8 attempts.
  4. Elo. Confirm that Play sits near 0.75, and leave K alone.
  5. Evidence that it works. The first real evidence is a multiple-baseline graph with 3 tiers and at least 5 points per phase, plus paper DCPM rising into or above the 24 to 49 range.

Stage 2: about 5 learners

  • Staggered, randomised family starts, with 2 to 3 weeks of paper-probe baseline where families accept it.
  • The hierarchical half-life model: a per-child multiplier for each KC family, partially pooled, to warm-start new children.
  • Template difficulties refit from pooled first attempts, with a smaller K for items.
  • The first item-level experiment: fast-review multiplier x2.5 against x2, randomised by fact family at mastery. It is low risk and targets the most important parameter.
  • Consent and assent forms, and an OSF preregistration before the second family starts.

Stage 3: about 20 learners

  • Refit the scheduler on pooled data (HLR-style regression with child and fact effects, and FSRS fitted on this log). Compare with the ladder on held-out later reviews, and switch only if the winner is better and stays under the review cap.
  • The randomised lapse-rule check.
  • A randomised default target level, with the child still free to change it.
  • Multilevel growth models, and a between-case effect size for the multiple-baseline data.

Stage 4: about 50 learners, or the school class

  • Lognormal speed and time-intensity models to replace the fixed fluency cutoffs, if Stage 3 showed they misclassify, plus a mixture-model lapse check.
  • For the class pilot:
    • independent ethics review, a DPIA, a data processing agreement, and parental opt-in;
    • a stepped start by table group within the class;
    • paper probes before and after in both classes, with the MTC mock as a secondary measure;
    • a protocol registered before term;
    • a descriptive between-class report.
  • The first external report: preregistered outcomes, multiple-baseline graphs, experiment results, transfer results, dropouts, and nulls.

Which parameter to revisit first

The scheduler's global half-life scale, then its later-review multipliers. It moves retention most in the simulation, the simulation already shows it over-predicting recall, it has the least external support, and it is the only parameter with a clean self-test in one child's own data within a few weeks. Second: the mastery threshold and minimum attempts. Last: Elo K.

Contested or weak evidence

  • The 0.75 target rests on Math Garden's design and engagement data, not an outcome RCT. The "85 percent rule" is a result about model learners.
  • FSRS for children and maths is untested. The claim that fitting beats defaults after about 16 reviews is a community analysis on adult data.
  • Duolingo's HLR test measured engagement, not learning.
  • Beck and Chang's identifiability claim has been reinterpreted as a problem of degenerate fits. The practical advice (use priors; do not trust per-child BKT parameters) stands either way.
  • Lapse detection has no validated method. The proposed checks test the rule's consequences, not its truth.
  • Fixed fluency cutoffs are conventions, and mixed strategies blur them.
  • Ex-Gaussian and diffusion parameters lack clean process interpretations. Diffusion models suit two-choice tasks.
  • CBM norms are US, mid-2000s, and on paper.
  • Single-case visual analysis is only moderately reliable, no single-case effect size is agreed, and the WWC thresholds are conventions.
  • Policy simulation is only as good as its learner model.
  • The Fisher-information figures (A.3) are my own derivation under an idealised model. Real data will need more reviews.
  • Small-sample Bayesian results depend on the priors, so preregister them.
  • Two-class comparisons cannot support causal claims.

References

Verification key: [verified] = checked against the primary text or official page this session; [from memory] = well-established reference not re-opened this session; [unverified] = details need checking before quoting.

  • Backes, B., and Cowan, J. (2019). Is the pen mightier than the keyboard? The effect of online testing on measured student achievement. Economics of Education Review, 68, 89-103. [from memory]
  • Barnett, A. G., van der Pols, J. C., and Dobson, A. J. (2005). Regression to the mean: what it is and how to deal with it. International Journal of Epidemiology, 34(1), 215-220. [from memory]
  • Barnett, S. M., and Ceci, S. J. (2002). When and where do we apply what we learn? A taxonomy for far transfer. Psychological Bulletin, 128(4), 612-637. [from memory]
  • Beck, J. E., and Chang, K. (2007). Identifiability: A fundamental problem of student modeling. In User Modeling 2007, LNCS 4511, 137-146. Springer. https://link.springer.com/chapter/10.1007/978-3-540-73078-1_17 [verified: existence and claim]
  • British Educational Research Association (2018). Ethical Guidelines for Educational Research (4th ed.). London: BERA. [from memory; a 5th edition (2024) is unverified]
  • Brinkhuis, M. J. S., Savi, A. O., Hofman, A. D., Coomans, F., van der Maas, H. L. J., and Maris, G. (2018). Learning as it happens: A decade of analyzing and shaping a large-scale online learning system. Journal of Learning Analytics, 5(2), 29-46. https://files.eric.ed.gov/fulltext/EJ1187391.pdf [verified]
  • Bulté, I., and Onghena, P. (2008). An R package for single-case randomization tests. Behavior Research Methods, 40(2), 467-478. [from memory]
  • Burns, M. K., VanDerHeyden, A. M., and Jiban, C. L. (2006). Assessing the instructional level for mathematics: A comparison of methods. School Psychology Review, 35(3), 401-418. [verified via secondary citation of the ranges]
  • Christ, T. J., Scullin, S., Tolbize, A., and Jiban, C. L. (2008). Implications of recent research: Curriculum-based measurement of math computation. Assessment for Effective Intervention, 33(4), 198-205. [from memory]
  • Christ, T. J., Zopluoglu, C., Monaghen, B. D., and Van Norman, E. R. (2013). Curriculum-based measurement of oral reading: Multi-study evaluation of schedule, duration, and dataset quality on progress monitoring outcomes. Journal of School Psychology, 51(1), 19-57. [from memory]
  • Codding, R. S., Burns, M. K., and Lukito, G. (2011). Meta-analysis of mathematic basic-fact fluency interventions: A component analysis. Learning Disabilities Research and Practice, 26(1), 36-47. [from memory]
  • Deno, S. L. (1985). Curriculum-based measurement: The emerging alternative. Exceptional Children, 52(3), 219-232. [from memory]
  • Doroudi, S., Aleven, V., and Brunskill, E. (2017). Robust evaluation matrix: Towards a more principled offline exploration of instructional policies. Proceedings of Learning at Scale 2017. [from memory]
  • Doroudi, S., and Brunskill, E. (2017). The misidentified identifiability problem of Bayesian knowledge tracing. Proceedings of EDM 2017. https://files.eric.ed.gov/fulltext/ED577166.pdf [verified: existence]
  • Doroudi, S., and Brunskill, E. (2019). Fairer but not fair enough: On the equitability of knowledge tracing. Proceedings of LAK 2019. [from memory]
  • Expertium (n.d., accessed 2026-09-27). Benchmark of spaced repetition algorithms. https://expertium.github.io/Benchmark.html [verified]
  • Anki manual (accessed 2026-09-27). Deck options: FSRS. https://docs.ankiweb.net/deck-options.html [verified]
  • ankitects/anki issue 3094 (2024). FSRS: Research suggests the minimum limit can be reduced to 16. https://github.com/ankitects/anki/issues/3094 [verified: claim made; community analysis]
  • Fancsali, S. E., Nixon, T., and Ritter, S. (2013). Optimal and worst-case performance of mastery learning assessment with Bayesian knowledge tracing. Proceedings of EDM 2013. [from memory]
  • Gelman, A., and Loken, E. (2013). The garden of forking paths. Unpublished manuscript, Columbia University. [from memory]
  • Gelman, A., Vehtari, A., Simpson, D., et al. (2020). Bayesian workflow. arXiv:2011.01808. [from memory]
  • Glickman, M. E. (1999). Parameter estimation in large dynamic paired comparison experiments. Applied Statistics, 48, 377-394. [verified: cited in Brinkhuis et al. 2018]
  • Heathcote, A., Popiel, S. J., and Mewhort, D. J. K. (1991). Analysis of response time distributions: An example using the Stroop task. Psychological Bulletin, 109(2), 340-347. [from memory]
  • Hedges, L. V., and Hedberg, E. C. (2007). Intraclass correlation values for planning group-randomized trials in education. Educational Evaluation and Policy Analysis, 29(1), 60-87. [from memory]
  • Hedges, L. V., Pustejovsky, J. E., and Shadish, W. R. (2012). A standardized mean difference effect size for single case designs. Research Synthesis Methods, 3(3), 224-239. (And 2013, multiple baseline designs, Research Synthesis Methods, 4(4), 324-341.) [from memory]
  • Hofman, A. D., Jansen, B. R. J., de Mooij, S. M. M., Stevenson, C. E., and van der Maas, H. L. J. (2018). A solution to the measurement problem in the idiographic approach using computer adaptive practicing. Journal of Intelligence, 6(1), 14. https://doi.org/10.3390/jintelligence6010014 [verified]
  • Jansen, B. R. J., Louwerse, J., Straatemeier, M., Van der Ven, S. H. G., Klinkenberg, S., and Van der Maas, H. L. J. (2013). The influence of experiencing success in math on math anxiety, perceived math competence, and math performance. Learning and Individual Differences, 24, 190-197. [verified: citation and headline findings; condition-level details unverified]
  • Johnson, A. H., and Cook, B. G. (2019). Preregistration in single-case design research. Exceptional Children, 86(1). [unverified pages]
  • Kelly, K., Wang, Y., Thompson, T., and Heffernan, N. (2015). Defining mastery: Knowledge tracing versus N-consecutive correct responses. Proceedings of EDM 2015. [from memory]
  • Khajah, M., Lindsey, R. V., and Mozer, M. C. (2016). How deep is knowledge tracing? Proceedings of EDM 2016. [from memory]
  • Klasnja, P., Hekler, E. B., Shiffman, S., et al. (2015). Microrandomized trials: An experimental design for developing just-in-time adaptive interventions. Health Psychology, 34(Suppl), 1220-1228. [from memory]
  • Klinkenberg, S., Straatemeier, M., and van der Maas, H. L. J. (2011). Computer adaptive practice of maths ability using a new item response model for on the fly ability and difficulty estimation. Computers and Education, 57(2), 1813-1824. [verified: citation via Brinkhuis et al. 2018]
  • Kofler, M. J., Rapport, M. D., Sarver, D. E., et al. (2013). Reaction time variability in ADHD: A meta-analytic review of 319 studies. Clinical Psychology Review, 33(6), 795-811. [from memory]
  • Kratochwill, T. R., Hitchcock, J., Horner, R. H., Levin, J. R., Odom, S. L., Rindskopf, D. M., and Shadish, W. R. (2010). Single-case designs technical documentation (Version 1.0, pilot). What Works Clearinghouse. https://ies.ed.gov/ncee/wwc/Docs/referenceresources/wwc_scd.pdf [verified]
  • Kratochwill, T. R., Hitchcock, J. H., Horner, R. H., Levin, J. R., Odom, S. L., Rindskopf, D. M., and Shadish, W. R. (2013). Single-case intervention research design standards. Remedial and Special Education, 34(1), 26-38. [from memory]
  • Kratochwill, T. R., and Levin, J. R. (2010). Enhancing the scientific credibility of single-case intervention research: Randomization to the rescue. Psychological Methods, 15(2), 124-144. [from memory]
  • Kravitz, R. L., Duan, N., and the DEcIDE Methods Center N-of-1 Guidance Panel (2014). Design and implementation of N-of-1 trials: A user's guide. AHRQ Publication No. 13(14)-EHC122-EF. [from memory]
  • Lerche, V., Voss, A., and Nagler, M. (2017). How many trials are required for parameter estimation in diffusion modeling? A comparison of different optimization criteria. Behavior Research Methods, 49(2), 513-537. [from memory]
  • Lindsey, R. V., Shroyer, J. D., Pashler, H., and Mozer, M. C. (2014). Improving students' long-term knowledge retention through personalized review. Psychological Science, 25(3), 639-647. https://pubmed.ncbi.nlm.nih.gov/24444515/ [verified]
  • Maris, G., and van der Maas, H. L. J. (2012). Speed-accuracy response models: Scoring rules based on response time and accuracy. Psychometrika, 77(4), 615-633. [from memory]
  • Matzke, D., and Wagenmakers, E.-J. (2009). Psychological interpretation of the ex-Gaussian and shifted Wald parameters: A diffusion model analysis. Psychonomic Bulletin and Review, 16(5), 798-817. [from memory]
  • McNeish, D. (2016). On using Bayesian methods to address small sample problems. Structural Equation Modeling, 23(5), 750-773. [from memory]
  • McNeish, D., and Stapleton, L. M. (2016). The effect of small sample size on two-level model estimates: A review and illustration. Educational Psychology Review, 28, 295-314. [from memory]
  • Morris, M. D. (1991). Factorial sampling plans for preliminary computational experiments. Technometrics, 33(2), 161-174. [from memory]
  • Murray, R. C., Ritter, S., Nixon, T., et al. (2013). Revealing the learning in learning curves. Proceedings of AIED 2013. [unverified]
  • Ninci, J., Vannest, K. J., Willson, V., and Zhang, N. (2015). Interrater agreement between visual analysts of single-case data: A meta-analysis. Behavior Modification, 39(4), 510-541. [from memory]
  • Nosek, B. A., Ebersole, C. R., DeHaven, A. C., and Mellor, D. T. (2018). The preregistration revolution. PNAS, 115(11), 2600-2606. [from memory]
  • Papoušek, J., Pelánek, R., and Stanislav, V. (2014). Adaptive practice of facts in domains with varied prior knowledge. Proceedings of EDM 2014. [verified: existence; parameters via Pelánek 2016]
  • Pardos, Z. A., and Heffernan, N. T. (2010). Modeling individualization in a Bayesian networks implementation of knowledge tracing. UMAP 2010. [from memory]
  • Parker, R. I., Vannest, K. J., Davis, J. L., and Sauber, S. B. (2011). Combining nonoverlap and trend for single-case research: Tau-U. Behavior Therapy, 42(2), 284-299. [from memory]
  • Pavlik, P. I., and Anderson, J. R. (2008). Using a model to compute the optimal schedule of practice. Journal of Experimental Psychology: Applied, 14(2), 101-117. [from memory]
  • Pelánek, R. (2016). Applications of the Elo rating system in adaptive educational systems. Computers and Education, 98, 169-179. https://www.fi.muni.cz/~xpelanek/publications/CAE-elo.pdf [verified]
  • Pelánek, R. (2017). Bayesian knowledge tracing, logistic models, and beyond: An overview of learner modeling techniques. User Modeling and User-Adapted Interaction, 27, 313-350. [verified: existence]
  • Pelánek, R. (2018). The details matter: Methodological nuances in the evaluation of student models. User Modeling and User-Adapted Interaction, 28, 207-235. [from memory]
  • Pelánek, R., and Řihák, J. (2017). Experimental analysis of mastery learning criteria. Proceedings of UMAP 2017, 156-163. https://www.fi.muni.cz/~xpelanek/publications/mastery-detection.pdf [verified: findings]
  • Pelánek, R., Řihák, J., and Papoušek, J. (2016). Impact of data collection on interpretation and evaluation of student models. Proceedings of LAK 2016. [from memory]
  • Pustejovsky, J. E., Hedges, L. V., and Shadish, W. R. (2014). Design-comparable effect sizes in multiple baseline designs. Journal of Educational and Behavioral Statistics, 39(5), 368-393. [from memory]
  • Ratcliff, R., and McKoon, G. (2008). The diffusion decision model: Theory and data for two-choice decision tasks. Neural Computation, 20(4), 873-922. [from memory]
  • Saltelli, A., Ratto, M., Andres, T., et al. (2008). Global sensitivity analysis: The primer. Wiley. [from memory]
  • Savi, A. O., Ruijs, N. M., Maris, G. K. J., and van der Maas, H. L. J. (2018). Delaying access to a problem-skipping option increases effortful practice: Application of an A/B test in large-scale online learning. Computers and Education, 119, 84-94. [verified: via Brinkhuis et al. 2018]
  • Sense, F., Behrens, F., Meijer, R. R., and van Rijn, H. (2016). An individual's rate of forgetting is stable over time but differs across materials. Topics in Cognitive Science, 8(1), 305-321. https://onlinelibrary.wiley.com/doi/10.1111/tops.12183 [verified]
  • Settles, B., and Meeder, B. (2016). A trainable spaced repetition model for language learning. Proceedings of ACL 2016, 1848-1858. https://aclanthology.org/P16-1174.pdf [verified]
  • Shadish, W. R., Kyse, E. N., and Rindskopf, D. M. (2013). Analyzing data from single-case designs using multilevel models. Psychological Methods, 18(3), 385-405. [from memory]
  • Siegler, R. S. (1987). The perils of averaging data over strategies: An example from children's addition. Journal of Experimental Psychology: General, 116(3), 250-264. [from memory]
  • Sindelar, P. T., Rosenberg, M. S., and Wilson, R. J. (1985). An adapted alternating treatments design for instructional research. Education and Treatment of Children, 8(1), 67-76. [from memory]
  • Tabibian, B., Upadhyay, U., De, A., Zarezade, A., Schölkopf, B., and Gomez-Rodriguez, M. (2019). Enhancing human learning via spaced repetition optimization. PNAS, 116(10), 3988-3993. [from memory]
  • Talts, S., Betancourt, M., Simpson, D., Vehtari, A., and Gelman, A. (2018). Validating Bayesian inference algorithms with simulation-based calibration. arXiv:1804.06788. [verified]
  • Tate, R. L., Perdices, M., Rosenkoetter, U., et al. (2016). The Single-Case Reporting Guideline In BEhavioural Interventions (SCRIBE) 2016 statement. (Published in several journals.) [from memory]
  • UK Information Commissioner's Office (2020). Age appropriate design: a code of practice for online services. [from memory]
  • US Federal Trade Commission. Children's Online Privacy Protection Rule, 16 CFR Part 312 (amended 2025). [from memory; amendment dates unverified]
  • US Department of Health and Human Services. Protection of Human Subjects, 45 CFR 46, including Subpart D and section 46.104(d)(1). [from memory]
  • van der Linden, W. J. (2007). A hierarchical framework for modeling speed and accuracy on test items. Psychometrika, 72(3), 287-308. [from memory]
  • van der Velde, M., Sense, F., Borst, J., and van Rijn, H. (2021). Alleviating the cold start problem in adaptive learning using data-driven difficulty estimates. Computational Brain and Behavior, 4, 231-249. [from memory]
  • Vohra, S., Shamseer, L., Sampson, M., et al. (2015). CONSORT extension for reporting N-of-1 trials (CENT) 2015 statement. BMJ, 350, h1738. [from memory]
  • Wagenmakers, E.-J., van der Maas, H. L. J., and Grasman, R. P. P. P. (2007). An EZ-diffusion model for response time and accuracy. Psychonomic Bulletin and Review, 14(1), 3-22. [from memory]
  • Wilson, R. C., Shenhav, A., Straccia, M., and Cohen, J. D. (2019). The Eighty Five Percent Rule for optimal learning. Nature Communications, 10, 4646. [from memory]
  • Ye, J., Su, J., and Cao, Y. (2022). A stochastic shortest path algorithm for optimizing spaced repetition scheduling. Proceedings of KDD 2022. [from memory]
  • Yudelson, M. V., Koedinger, K. R., and Gordon, G. J. (2013). Individualized Bayesian knowledge tracing models. AIED 2013. [from memory]

Numberkit internal sources: docs/PLAN.md sections 6, 13, 15; docs/decisions/0001-scheduler.md; docs/research/06-attention-adhd-and-math-practice.md.