Research 13: Adaptive placement with a handful of questions
Date: 2026-10-07. Written for Task AC.3 of docs/IMPLEMENTATION-ADAPTIVE-CHECKS.md, before its constants were fixed in code. Scope: what is established about (A) assessment over a prerequisite graph and (B) adaptive testing with very few items, and what that means for the numbers the adaptive check uses.
Verification key, as in Research 11: [verified] means the primary text or the official page was fetched and read during this review, on 2026-10-07. [recalled] means a well-known source whose details were not re-checked. [unverified] means the detail must be checked before it is quoted. The two main sources (Falmagne et al. 2006; Cosyn et al. 2021) were read in full text. Both are written by the people who built ALEKS, three of the four authors of the second being employees, so their accuracy figures are the vendor's own.
Summary
- Nothing found contradicts section 2 of the adaptive-checks plan. Its three structural choices (walk the prerequisite graph, ask where the answer splits what is open in half, count a right answer for what it rests on) are the design of the one large system built on such an assessment, described in its own papers.
- The half-split is the published rule for a chain with nothing known beforehand. In knowledge space theory the next question is the one the child has "about a 50% chance" of answering, judged over the states still possible (Falmagne et al. 2006) [verified]. On a tree, or with some states likelier than others, counting open KCs only approximates it (section C.10).
- Right is strong, wrong is weak, "I don't know" is strongest. ALEKS weights its update at about 35 for a right answer, 5 for a wrong one, and 50 for "I don't know" (Cosyn et al. 2021) [verified]. That is the same ordering as the plan's one right answer, two misses, and one declined answer. The published system never settles an item on one answer, so the plan's "thin" placement and its quick take-back carry real weight.
- The weakest part is the graph, not the walk. ALEKS's precedence relation was built by asking experts a specific question about every pair and then corrected against the answers of thousands of students. Numberkit's graph is one designer's teaching order and has never been checked against an answer. Every inferred KC is only as good as the edge it was inferred along.
- There is no evidence for six questions, for twelve, or for three in a row. They are hand values. The sources bound them loosely and the arithmetic in section C shows what each can and cannot do.
A. Assessment over a prerequisite graph
A.1 Knowledge space theory
Doignon and Falmagne (1985; book 1999) [recalled] describe a domain as a set of problem types and a child's knowledge as a state: the set of types he can solve. A precedence relation between types ("the mastery of Problem (e) implies that of (b), (c) and (a)") cuts the possible states from every subset to the ones consistent with it (Falmagne et al. 2006) [verified]. Their six-problem example has 10 states out of 64 subsets, and 88 problem types of beginning algebra have "around 60,000" [verified].
Two ideas in it are the plan's own, under other names [verified, same source]:
- The outer fringe of a state is the set of problems a child can learn next ("what the student is ready to learn"), and the inner fringe is the most advanced of what he knows. The outer fringe is the engine's
frontier. - The result of an assessment is those two short lists, not a score.
A.2 How the ALEKS assessment asks and updates
From Falmagne et al. (2006) [verified]:
- The questioning rule. The first problem "is chosen so as to be 'maximally informative.' This is interpreted to mean that, on the basis of the current likelihoods of the states, the student has about a 50% chance of knowing how to solve" it: "the sum of the likelihoods of all the states containing p1 is as close to .5 as possible." Ties are broken at random.
- The updating rule is soft. A right answer raises the likelihood of every state containing the problem and a wrong answer lowers it. No state is ever ruled out by one answer, so "the final state may very well contain a problem to which the student gave a false response. Such a response is thus regarded as due to a careless error."
- "I don't know" is an answer. It "results in a substantial increase in the likelihood of the states not containing" the problem, "thereby decreasing the total number of questions required".
- Guessing is designed out. "All the problems have open-ended responses (no multiple choice), with a large number of possible solutions", so "the probability of lucky guesses is negligible."
- The prior may use the school year and "play[s] no role in the final result of the assessment but may be helpful in shortening it."
- Length. The worked example settled one state among 57,147 in 24 questions, "close to the average for Arithmetic" (108 problem types). It stopped when the uncertainty was low and the best remaining question had only a 19 percent chance of being solved.
From Cosyn et al. (2021) [verified]:
- The weights. "The update parameters in ALEKS are about 35 for a correct answer, 5 for an incorrect answer, and 50 for 'I don't know'". The wrong-answer update "cannot be very aggressive given that there is a non-negligible chance for a careless error", and this is given as the reason the "I don't know" button exists. The button is also labelled "I haven't learned it yet", the plan's wording.
- Slips are common and lucky guesses rare. "In ALEKS, careless errors are much more common than lucky guesses; this is due to the fact that the majority of ALEKS items require an open-ended response".
- The cap. A noise-free halving search of the placement course would need 77 questions. "Based on the feedback from students and instructors over the years, the number of questions in an initial assessment has been capped at 30", to balance information against "the risk of overwhelming the student with too many questions", with a note that students show a fatigue effect in assessments. "Most often" the assessment ends at the cap, not because it converged.
- What is left over. Items still unsettled at the cap are classed "uncertain", left out of the state (so the state "may underestimate"), and their learning is fast-tracked: a lower score is needed to learn them. This is the plan's guard (what is still open is "left as it is for practice to decide") together with its clean start.
- Most of the information comes early. Over 2.69 million placement assessments of the full 29 questions, accuracy measured after each question "converge[s] early in the assessment, with the changes being minimal after question 10". The course has 314 items.
- Children know what they know. The rate of "I don't know" on items the assessment classed as known is "almost constant" however far inside the state the item lies, which the authors read as "students appear to recognize when they do in fact know how to solve an item."
A.3 How accurate it is
ALEKS adds one randomly chosen extra problem to every assessment and does not use its answer, so the answer tests the assessment's verdict (Falmagne et al. 2006; Cosyn et al. 2021) [verified]. In the 2021 data (Table 3, read from the text):
| Course | Items classed as known answered right | Items classed as not known answered right | Unsettled items answered right |
|---|---|---|---|
| Sixth-Grade Math | 0.794 | 0.110 | 0.428 |
| College Algebra | 0.740 | 0.110 | 0.427 |
| College Placement | 0.838 | 0.102 | 0.449 |
Two readings matter here.
- A child misses about one item in five of those he is judged to know. For items just inside the edge the rate of right answers is "roughly 0.7", against "roughly 0.27" for items just outside it [verified]. The plan's lapse rate of one answer in ten (PLAN section 13) is optimistic by this measure, at least near the edge.
- The top of what a child knows is the least secure. In over six million assessments, "the content representing the 'high points' of a student's learning sees a more precipitous drop" in memory than the rest (Matayoshi et al. 2018) [verified: abstract only, through a search listing; the paper was not read].
Doble et al. (2019) report a simulation study of the reliability of this assessment [citation verified; not read].
A.4 Where the graph comes from
Falmagne et al. (2006) [verified] build the precedence relation by asking experts one question about pairs of problems, "Suppose that a student is not capable of solving problem p. Could this student nevertheless solve problem p'?", which takes "a few thousand questions" a subject. They then say this "does not guarantee the validity of the knowledge structure" and that "actual student data are also needed": states that never occur are deleted, and the extra problem's agreement with the verdict (a correlation "between .7 and .8" for beginning algebra) monitors the structure continuously.
Numberkit's graph answers a different question: what should be taught before what. The two agree where a prerequisite is truly used inside the later skill (column addition uses the addition facts). They can differ where an edge is a sensible teaching order and nothing more. An inference down such an edge is a guess.
A.5 Other products
- Math Academy (vendor's own description, mathacademy.com, read 2026-10-07) [verified as a description; no evidence offered]. Its diagnostic "repeatedly selects the topic whose assessment provides the most information"; "each correct answer provides positive evidence that the student knows the topic, its prerequisites, and other related topics"; and topics a child was only just placed out of are "conditionally completed": work proceeds as if they were known, "but if the student struggles, then the system will immediately begin 'falling backwards.'" That is the plan's thin placement and take-back, as a product description.
- Math Garden starts a child from a rating and adapts item by item with Elo, without a prerequisite graph or a placement test (Klinkenberg et al. 2011) [recalled; see Research 04 section 4]. A question skipped scores as wrong there (Research 04). It offers nothing on placing over a graph.
A.6 Searching an order with unreliable answers
The half-split on a chain is binary search, which finds one of n + 1 positions in about log2(n + 1) answers when answers are reliable. Searching a partial order by the question that best halves the remaining possibilities is a studied problem (Linial and Saks 1985), and binary search with answers that are sometimes wrong still works at the cost of repeating questions, a constant factor more for a fixed error rate (Feige et al. 1994; Karp and Kleinberg 2007) [all recalled, not checked]. Repeating a question after a surprising answer, as the plan does after a first miss, is the simplest form of that.
B. Adaptive testing with very few items
- A handful of items cannot give a score. In a Rasch-model adaptive test the standard error is 1 / sqrt(sum of P(1 - P)) (Linacre 2006) [verified through a summary of the page], so six perfectly targeted items give 1 / sqrt(6 x 0.25) = 0.82 logits (calculated for this review). Linacre's table puts the fewest items for a standard error of 0.5 logits at 16. Health questionnaires that stop at 4 to 12 items use items with several response levels, and shortening them to 8 or fewer lowered reliability [unverified: from a search summary of a 2025 paper in JMIR Formative Research; the page did not load].
- The check does not estimate a score. It locates an edge on a short order. A strand of 5 to 8 KCs in a chain has 6 to 9 possible edges, which is 2.6 to 3.2 reliable yes-or-no answers. Six questions are enough for that, with one or two repeats. They would not be enough to place a child on a scale, and nothing in the plan asks them to.
- Short mastery decisions are an old problem. Sequential tests that stop as soon as the evidence clears a bar go back to Wald, and were applied to pass or fail decisions with adaptive tests by Reckase (1983) [recalled]. Knowledge tracing puts typical guess and slip probabilities for typed tutor steps near 0.1 to 0.3 (Corbett and Anderson 1995; Baker, Corbett and Aleven 2008) [recalled]. None of this was re-read, and none of it gives a number for a child of nine answering one question.
C. What this means for the plan's constants
Each line says what the sources support, what is only arithmetic, and what has no support.
- One right typed answer shows a KC (decision 3). The direction is supported: an open-ended right answer is strong evidence (the weight of 35 against 5; guesses "negligible"). The amount is not: ALEKS combines about 29 answers and never settles an item on one, and its "not known" items were still answered right about one time in ten. So one answer is a provisional placement and no more. The plan says so, and the engine marks it thin.
- A second question after a miss (decision 3). Supported. Slips are common (one item in five of those judged known, three in ten at the edge), and the published update for a wrong answer is deliberately weak. With a slip rate of 0.2, one miss wrongly fails a child who knows the KC one time in five and two misses one time in twenty-five (arithmetic, assuming the two are independent).
- "I haven't learned this yet" counts at once (decision 10). Supported: it is the strongest single response in ALEKS (50), it exists there to shorten the assessment, and children's use of it tracks what they know. One caution: lower-performing students in ALEKS were less willing to attempt a problem the second time they saw it (Matayoshi et al. 2018) [verified: abstract only], so a child who feels he is failing may decline what he knows. The plan's clean start is the remedy and should be watched.
- Three in a row for a choice item (decision 3). No source gives three. The arithmetic: three right in a row by guessing between two options happens one time in eight, which is weaker than one typed answer. The sources' own answer is to avoid choice items in an assessment. A choice KC shown this way is not thin by the plan's rule (three direct answers), and it is still the least secure placement the walk makes.
- Six questions a sitting (decision 8). No evidence for six. The sources support short over long (a cap chosen from complaints and fatigue; most information by question 10 of 29) and section B shows six is enough for one strand's edge. It is the owner's number and it is safe.
- The guard of twelve questions a strand (section 2.3). No evidence; arithmetic only. A typed chain of up to 8 KCs settles in at most about 2 x 4 = 8 questions if every miss is asked twice, so twelve is slack there. A strand of choice KCs is different: each needs three right answers and there is nothing to infer from one answer, so four such KCs fill the guard. Any walk with more than four choice KCs will end at the guard with some of them open, which the plan allows ("left as it is for practice to decide"). Task AC.4 should count the choice KCs in each strand before relying on the guard to be slack.
- A thin placement taken back by its first miss (decision 6). This is not in section 2, and the evidence cuts against it a little. If a child misses one item in five of what he knows, a KC he really knows will be taken back at its first practice answer about one time in five, and the shown KC is the likeliest to be forgotten (A.3). The cost is small as the plan stands (the check reopens below with a question or two, and a clean start promotes again in three answers), and Math Academy describes the same design. It is the constant most worth logging: count the take-backs that a clean start reverses.
- Counting a right answer for the KCs beneath it (decision 2). Supported as a method and unsupported for this graph (A.4). Log the first practice answer on every inferred KC, so the edges that mislead can be found. The plan's claim ("the app's own graph read backwards, and nothing more is claimed for it") is the right size.
- No start from the year band (decision 12). ALEKS does use the school year as a prior and says it only shortens the assessment. The plan gives that up on purpose, at about one question a strand. No conflict.
- Counting KCs in place of weighing states. The plan's half-split maximises the smaller of two counts of open KCs. On a chain with nothing known beforehand this is the published rule exactly. On a tree it is an approximation, since the published rule weighs states by likelihood and the number of states is not the number of KCs. No source measures how many questions the approximation costs, and none is claimed here. What is known is the engine's own count: on the real place-value strand, a tree of 6 KCs, a child who gets everything right is asked 5 questions, because a right answer on one branch says nothing about another.
Contested or weak evidence
- Every accuracy figure here is ALEKS's, about ALEKS, from its own staff, on secondary and college mathematics and one sixth-grade course. None is about nine-year-olds or arithmetic.
- ALEKS's outcomes are ordinary. Fang et al. (2019) found it as good as, not better than, traditional teaching (Research 04). A sound assessment is not evidence of better learning.
- The slip and guess rates do not behave as the simple model says. Cosyn et al. (2021) show the rates vary with how far an item is from the edge, so "one in five" is an average and is higher at the edge.
- Nothing was found on the length of a sitting for a child, as Research 06 found nothing on session length.
- The sources on noisy search and on sequential mastery testing are recalled, and are used only to say that repeating a question is a known remedy.
- The plan's graph has never met data. Until it has, inference is a designer's judgement applied quickly.
References
- Baker, R. S. J. d., Corbett, A. T., and Aleven, V. (2008). More accurate student modeling through contextual estimation of slip and guess probabilities in Bayesian knowledge tracing. In Intelligent Tutoring Systems 2008, LNCS 5091. Springer. [recalled]
- Corbett, A. T., and Anderson, J. R. (1995). Knowledge tracing: Modeling the acquisition of procedural knowledge. User Modeling and User-Adapted Interaction, 4, 253-278. [recalled]
- Cosyn, E., Uzun, H., Doble, C., and Matayoshi, J. (2021). A practical perspective on knowledge space theory: ALEKS and its data. Journal of Mathematical Psychology, 101, 102512. Full text read at https://www.mheducation.com/content/dam/mhe/research/prek-12/essa-tier-4/aleks-practical-perspective-knowledge-space-theory.pdf [verified; vendor-authored; article number recalled]
- Doble, C., Matayoshi, J., Cosyn, E., Uzun, H., and Karami, A. (2019). A data-based simulation study of reliability for an adaptive assessment based on knowledge space theory. International Journal of Artificial Intelligence in Education, 29(2), 258-282. [citation verified; not read]
- Doignon, J.-P., and Falmagne, J.-C. (1985). Spaces for the assessment of knowledge. International Journal of Man-Machine Studies, 23, 175-196. [recalled]
- Doignon, J.-P., and Falmagne, J.-C. (1999). Knowledge Spaces. Berlin: Springer. [recalled; cited in Falmagne et al. 2006]
- Falmagne, J.-C., Cosyn, E., Doignon, J.-P., and Thiéry, N. (2006). The assessment of knowledge, in theory and in practice. Full text read at https://www.aleks.com/about_aleks/Science_Behind_ALEKS.pdf [verified; vendor-hosted; the year and its publication in Formal Concept Analysis, LNCS 3874, Springer, are recalled]
- Feige, U., Raghavan, P., Peleg, D., and Upfal, E. (1994). Computing with noisy information. SIAM Journal on Computing, 23(5), 1001-1018. [recalled]
- Karp, R. M., and Kleinberg, R. (2007). Noisy binary search and its applications. Proceedings of SODA 2007. [recalled]
- Klinkenberg, S., Straatemeier, M., and van der Maas, H. L. J. (2011). Computer adaptive practice of maths ability using a new item response model for on the fly ability and difficulty estimation. Computers and Education, 57(2), 1813-1824. [recalled; see Research 04]
- Linacre, J. M. (2006). Computer adaptive tests (CAT), item selection, standard errors and stopping rules. Rasch Measurement Transactions, 20(2), 1062. https://www.rasch.org/rmt/rmt202f.htm [verified through a summary of the page]
- Linial, N., and Saks, M. (1985). Searching ordered structures. Journal of Algorithms, 6(1), 86-103. [recalled]
- Matayoshi, J., Granziol, U., Doble, C., Uzun, H., and Cosyn, E. (2018). Forgetting curves and testing effect in an adaptive learning and assessment system. Proceedings of the 11th International Conference on Educational Data Mining. [verified: abstract only]
- Math Academy (accessed 2026-10-07). How our AI works. https://mathacademy.com/how-our-ai-works [verified; vendor description]
- Reckase, M. D. (1983). A procedure for decision making using tailored testing. In D. J. Weiss (Ed.), New Horizons in Testing. New York: Academic Press. [recalled]