Course material
Custom Prompt AI_tutoring_cognitive_offloading_literature_map
Details
Filename
Custom Prompt AI_tutoring_cognitive_offloading_literature_map.pdf
Size
2.9 MB
Type
application/pdf
Published
Preview
Extracted text
Structured AI Tutoring vs. Cognitive Offloading in Undergraduate STEM A citation-lineage and evidence map centered on Kestin et al. (2025) Research question Research question: Does using AI tools for problem-solving improve undergraduate STEM students' learning and long-term retention, or does it create a cognitive-offloading effect that reduces skill-building - and does the outcome depend on how the AI is used, such as answer-giving versus Socratic tutoring? Conceptual illustration: the same AI interface can sit on a learning path that preserves cognitive work or a shortcut path that displaces it; the empirical literature increasingly suggests the difference is pedagogical design and user behavior, not AI access alone. Seed paper Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2025). AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports, 15, 17458. doi:10.1038/s41598-025-97652-6 Scope 20 related papers selected for conceptual lineage and research-question relevance: 10 confirmed antecedents cited by the seed + 10 recent descendants/adjacent studies. The seed itself is shown separately and is not counted in the 20-paper limit. Literature check Recent-source verification performed through 18 August 2026. Direct citation/extension relationships are labeled separately from inferred thematic relationships. AI tutoring, cognitive offloading, and undergraduate STEM learning 2Evidence map centered on Kestin et al. (2025) | literature checked through 18 Aug 2026 Executive synthesis Bottom line The evidence does not support a simple "AI helps" or "AI harms" conclusion. Structured AI that keeps the learner cognitively active - prompting attempts, providing targeted explanations or hints, asking Socratic questions, and withholding complete solutions when appropriate - can improve learning. By contrast, open-ended answer-giving can raise assisted task performance while reducing independent performance, persistence, study time, or retention once the tool is removed. What the seed establishes. Kestin et al. randomized 194 undergraduates in introductory physics to a carefully engineered GPT-4 tutor versus a strong active-learning classroom condition. The AI condition produced substantially higher immediate post-test performance in less median time and higher engagement/motivation, demonstrating that an AI tutor can outperform current best-practice instruction when its interaction is tightly scaffolded. What the seed does not establish. The intervention spans two lessons and measures immediate post-tests; the authors explicitly identify retention, longer course integration, higher-order synthesis, and the decision of when to give an answer versus guide reflection as open research directions. That limitation is decisive for your question. What subsequent evidence adds. A direct extension (Miller, 2026) shows that pedagogical system prompts change engagement and cognitive load. Later/parallel experiments sharpen the mechanism: guarded or Socratic tutoring preserves cognitive work, while unrestricted answer access can create dependence. Contractor and Reyes (2026) add one-week undergraduate evidence in which explanation-seeking "augmentation" users retain gains better than text-generating "automation" users. Where cognitive offloading becomes visible. Bastani et al. (2025) observe a reversal from higher assisted performance to lower unassisted performance under unrestricted GPT-4; Liu et al. (2026) find reduced independent performance and persistence after brief AI exposure; Rismanchian et al. (2026) report less time on AI-susceptible college-math problems and weaker later proctored retention in a large longitudinal dataset. Current answer to the research question. The most defensible conclusion is conditional: AI can improve learning when it behaves like a tutor and is used as an augmentation; it can reduce skill-building when it behaves like an answer engine and is used as a replacement for reasoning. What remains under-tested is the exact combination you care most about: a long-duration, undergraduate STEM RCT that randomizes answer-giving versus Socratic scaffolding and measures delayed, unaided transfer and retention. AI tutoring, cognitive offloading, and undergraduate STEM learning 3Evidence map centered on Kestin et al. (2025) | literature checked through 18 Aug 2026 Relationship map FOUNDATIONS SEED RECENT EVIDENCE Tutoring Bloom Scaffolding Wood et al. Cognitive load Sweller Feedback Shute Active learning Hake/Freeman Early LLM Kumar/Krupp/Forero Kestin et al. (2025) Structured GPT-4 tutor Undergraduate physics RCT Short-term learning: strong + Scaffolded / Socratic / augmentation Answer-giving / offloading risk Cross-cutting evaluation / synthesis Miller 2026 LearnLM 2025 Degen 2025 Sun & Liu 2026 Contractor & Reyes 2026 Bastani 2025 Liu et al. 2026 Rismanchian et al. 2026 SafeTutors 2026 Wang et al. 2026 Solid arrows indicate a verified citation or explicit extension relationship. Dashed arrows indicate thematic adjacency selected for the research question; those papers were not confirmed as direct citations of the seed in the accessible full text searched. AI tutoring, cognitive offloading, and undergraduate STEM learning 4Evidence map centered on Kestin et al. (2025) | literature checked through 18 Aug 2026 1. The seed paper as the pivot Seed: Kestin, G., Miller, K., Klales, A., Milbourne, T., & Ponti, G. (2025). AI tutoring outperforms in-class active learning: an RCT introducing a novel research-based design in an authentic educational setting. Scientific Reports, 15, 17458. doi:10.1038/s41598-025-97652-6. Dimension What the seed paper actually shows Implication for your question Population / context 194 eligible students in Harvard Physical Sciences 2, an introductory physics course for life sciences; crossover design across two lessons. Strong undergraduate STEM relevance, but one institution and a narrow learning window. Comparator Research-based in-class active learning, not passive lecture. The effect is not merely "AI beats lecture"; the tutor outperformed an already effective pedagogy. AI design GPT-4 tutor with instructor-crafted prompts, sequential scaffolding, brief responses, step-by-step reference solutions, targeted feedback, and self-pacing. The intervention is closer to an engineered Socratic/scaffolded tutor than to unrestricted ChatGPT. Outcome Immediate post-test learning; median AI post-score 4.5 versus 3.5 for active learning, with large adjusted effects; AI students also reported greater engagement and motivation. Strong evidence for immediate acquisition and user experience, not durable retention. Critical gap The authors call for studies of spacing/retention, extended course use, higher-order synthesis, and decisions about when the tutor should openly provide answers versus guide reflection. Your question is a direct continuation of the seed paper's stated limitations and future-work agenda. Interpretive caution Because the seed tutor explicitly discourages giving away the full solution and embeds scaffolding, it should not be generalized to "ChatGPT use" in the abstract. The treatment is a particular pedagogy delivered through an LLM. That distinction explains why later studies of unrestricted answer access can produce opposite effects without actually contradicting Kestin et al. AI tutoring, cognitive offloading, and undergraduate STEM learning 5Evidence map centered on Kestin et al. (2025) | literature checked through 18 Aug 2026 2. Foundational papers the seed builds on All ten papers below are confirmed antecedents: they appear in the seed paper's reference list and are used to motivate its tutoring design, active-learning comparator, learning/perception measures, or contrast with unguided LLM use. CONFIRMED ANTECEDENT 1. Bloom, B. S. (1984). The 2 sigma problem: The search for methods of group instruction as effective as one-to-one tutoring. Educational Researcher, 13(6), 4-16. doi:10.3102/0013189X013006004 Relationship to seed: Directly cited by Kestin et al. as the benchmark for the effectiveness - and scalability problem - of expert one-to-one tutoring. One-sentence summary: Bloom framed individualized tutoring as a high-performing instructional benchmark, motivating the seed paper's attempt to reproduce tutor-like personalization at scale with AI. Research-question fit: Foundational to the tutoring side of the question, but it predates generative AI and does not address offloading. CONFIRMED ANTECEDENT 2. Wood, D., Bruner, J. S., & Ross, G. (1976). The role of tutoring in problem-solving. Journal of Child Psychology and Psychiatry, 17(2), 89-100. Relationship to seed: Directly cited by the seed paper as the basis for scaffolding content and sequencing support during problem solving. One-sentence summary: This classic tutoring study formalized scaffolding as contingent support that helps learners perform beyond current capability while preserving learner participation in the solution process. Research-question fit: Provides the clearest pre-AI rationale for Socratic or hint-based assistance rather than simply supplying answers. CONFIRMED ANTECEDENT 3. Sweller, J. (2011). Cognitive load theory. In J. P. Mestre & B. H. Ross (Eds.), The Psychology of Learning and Motivation: Cognition in Education (pp. 37-76). Elsevier Academic Press. Relationship to seed: Directly cited and explicitly implemented in the AI tutor design through brief, targeted exchanges intended to manage working-memory demands. One-sentence summary: Cognitive load theory argues that instruction should reduce extraneous load and manage task complexity so limited working memory can be devoted to schema construction. Research-question fit: Central to distinguishing productive offloading of irrelevant burden from harmful offloading of the reasoning practice itself. CONFIRMED ANTECEDENT 4. Shute, V. J. (2008). Focus on formative feedback. Review of Educational Research, 78(1), 153-189. doi:10.3102/0034654307313795 Relationship to seed: Directly cited by Kestin et al. to justify targeted, timely feedback as a core design principle of the tutor. One-sentence summary: Shute synthesizes evidence that formative feedback is most useful when it is timely, specific, non-evaluative, and oriented toward closing the gap between current and desired performance. Research-question fit: Supports AI as a feedback partner that responds to student thinking rather than a solution generator. AI tutoring, cognitive offloading, and undergraduate STEM learning 6Evidence map centered on Kestin et al. (2025) | literature checked through 18 Aug 2026 CONFIRMED ANTECEDENT 5. Hake, R. R. (1998). Interactive-engagement versus traditional methods: A six-thousand-student survey of mechanics test data for introductory physics courses. American Journal of Physics, 66(1), 64-74. doi:10.1119/1.18809 Relationship to seed: Directly cited as evidence that interactive engagement outperforms traditional lecture in introductory physics, helping define the strong control condition the AI tutor must beat. One-sentence summary: Across thousands of introductory physics students, interactive-engagement methods produced substantially larger conceptual gains than traditional instruction. Research-question fit: Establishes active problem solving as the benchmark against which AI-supported learning should be judged. CONFIRMED ANTECEDENT 6. Freeman, S., Eddy, S. L., McDonough, M., Smith, M. K., Okoroafor, N., Jordt, H., & Wenderoth, M. P. (2014). Active learning increases student performance in science, engineering, and mathematics. Proceedings of the National Academy of Sciences, 111(23), 8410-8415. doi:10.1073/pnas.1319030111 Relationship to seed: Directly cited to establish active learning as an evidence-based STEM standard rather than a weak lecture-only comparator. One-sentence summary: This meta-analysis of 225 STEM studies found that active learning improves exam performance and lowers failure rates relative to traditional lecturing. Research-question fit: Makes the seed RCT more consequential: its comparison is against an established effective pedagogy, not passive instruction. CONFIRMED ANTECEDENT 7. Deslauriers, L., McCarty, L. S., Miller, K., Callaghan, K., & Kestin, G. (2019). Measuring actual learning versus feeling of learning in response to being actively engaged in the classroom. Proceedings of the National Academy of Sciences, 116(39), 19251-19257. doi:10.1073/pnas.1821936116 Relationship to seed: Directly cited and authored partly by the same team; it motivates the seed paper's separation of objective learning gains from students' subjective experience. One-sentence summary: Students in active classrooms learned more even when they felt they learned less, demonstrating that perceived fluency or enjoyment can diverge from actual learning. Research-question fit: Important for AI studies because a frictionless, confidence-inducing tool can feel helpful while reducing independent skill formation. CONFIRMED ANTECEDENT 8. Kumar, H., Rothschild, D. M., Goldstein, D. G., & Hofman, J. M. (2023). Math education with large language models: Peril or promise? SSRN Electronic Journal. doi:10.2139/ssrn.4641653 Relationship to seed: Directly cited by Kestin et al. as early mixed evidence on LLM-supported learning and as a precursor to testing how instructional structure changes outcomes. One-sentence summary: In a preregistered experiment with about 1,200 participants, LLM explanations improved later math performance over answers alone, with the largest gains when learners attempted problems before consulting the explanation. Research-question fit: One of the most direct antecedents for the answer-giving versus explanation/attempt-first dimension in the research question. AI tutoring, cognitive offloading, and undergraduate STEM learning 7Evidence map centered on Kestin et al. (2025) | literature checked through 18 Aug 2026 CONFIRMED ANTECEDENT 9. Krupp, L., Steinert, S., Kiefer-Emmanouilidis, M., Avila, K. E., Lukowicz, P., Kuhn, J., Küchemann, S., & Karolus, J. (2024). Unreflected acceptance - Investigating the negative consequences of ChatGPT-assisted problem solving in physics education. In HHAI 2024: Hybrid Human AI Systems for the Social Good (pp. 199-212). doi:10.3233/FAIA240195 Relationship to seed: Directly contrasted with the seed results as evidence that unrestricted ChatGPT use can elicit shallow interaction and uncritical acceptance. One-sentence summary: Physics-background participants frequently over-trusted incorrect ChatGPT-supported solutions and used copy-and-paste prompting far more than search-engine users, indicating limited reflection during problem solving. Research-question fit: Directly supports the cognitive-offloading concern in higher-education STEM problem solving. CONFIRMED ANTECEDENT 10. Forero, M. G., & Herrera-Suárez, H. J. (2023). ChatGPT in the classroom: Boon or bane for physics students' academic performance? arXiv preprint arXiv:2312.02422. arXiv:2312.02422 Relationship to seed: Directly contrasted by Kestin et al. as a prior physics study in which relatively unstructured ChatGPT integration was associated with poorer academic performance. One-sentence summary: A second-semester engineering-physics cohort encouraged to use ChatGPT earned lower exam grades than a prior control cohort, while many students simultaneously reported convenience and concerns about reduced critical or independent thinking. Research-question fit: A close pre-seed warning that unguided AI access may harm learning even when students perceive it as useful. AI tutoring, cognitive offloading, and undergraduate STEM learning 8Evidence map centered on Kestin et al. (2025) | literature checked through 18 Aug 2026 3. Recent papers that extend, cite, or directly test the seed paper's thesis The first five are verified citation-lineage papers. The final five are deliberately included as adjacent evidence because they address the exact causal mechanisms in your question even though I did not confirm a direct citation to Kestin et al. in the accessible full text. CONFIRMED EXTENSION 1. Miller, K. (2026). Prompt Matters: How Pedagogical Engineering Shapes Behavior and Engagement with AI Tutors. Technology, Knowledge and Learning. doi:10.1007/s10758-026-09983-6 Relationship to seed: The paper explicitly states that it uses the same framework as Kestin et al. and acknowledges the work as an extension of the earlier experiment. One-sentence summary: In a randomized study (N=125), an enhanced tutor prompt incorporating active learning, cognitive-load management, and growth-mindset support produced more student-AI interaction, higher engagement and motivation, and less cognitive overload than a minimal prompt, while both groups learned. Research-question fit: Directly isolates prompt-level pedagogy as a mechanism, although it does not establish long-term retention. CONFIRMED DESCENDANT 2. LearnLM Team, Google & Eedi. (2025). AI tutoring can safely and effectively support students: An exploratory RCT in UK classrooms. arXiv preprint arXiv:2512.23633. arXiv:2512.23633 Relationship to seed: The paper directly cites Kestin et al. in its review of evidence for effective generative-AI tutoring. One-sentence summary: In an exploratory RCT with 165 secondary-school students, human-supervised LearnLM tutoring matched human tutoring and improved success on novel subsequent-topic problems by 5.5 percentage points, with tutors specifically praising its Socratic questions. Research-question fit: Strong evidence for guided dialogue and transfer, but the population is K-12 rather than undergraduates. CONFIRMED DESCENDANT 3. Hazra, R., Ghuku, B., Marchenko, I., Tokarieva, Y., Layek, S., Banerjee, S., Stoyanovich, J., & Pechenizkiy, M. (2026). SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems. arXiv preprint arXiv:2603.17373. arXiv:2603.17373 Relationship to seed: The benchmark paper directly cites Kestin et al. while reframing educational AI safety around preserving learning effort and scaffolding, not merely avoiding harmful content. One-sentence summary: Across mathematics, physics, and chemistry, SafeTutors finds widespread pedagogical harms such as answer over-disclosure and collapsed scaffolding, with multi-turn failure rates rising sharply in its evaluation scenarios. Research-question fit: Does not measure student learning directly, but operationalizes the exact failure mode behind answer-giving and cognitive offloading. AI tutoring, cognitive offloading, and undergraduate STEM learning 9Evidence map centered on Kestin et al. (2025) | literature checked through 18 Aug 2026 CONFIRMED PREPRINT-LINEAGE 4. Degen, P.-B., & Asanov, I. (2025). Beyond Automation: Socratic AI, Epistemic Agency, and the Implications of the Emergence of Orchestrated Multi-Agent Learning Architectures. arXiv preprint arXiv:2508.05116. arXiv:2508.05116 Relationship to seed: The paper cites the 2024 Kestin et al. preprint version and extends the design logic toward explicitly Socratic, dialogue-driven AI support. One-sentence summary: In a controlled experiment with 65 pre-service teacher students, a Socratic tutor was rated as providing more support for critical, independent, and reflective thinking than an uninstructed chatbot. Research-question fit: Directly tests Socratic versus generic chatbot interaction, but outcomes are largely self-reported rather than delayed performance measures. CONFIRMED DESCENDANT 5. Sun, Y., & Liu, F. (2026). The impact of an AI Digital Teacher on human-AI collaborative learning in higher education. Smart Learning Environments, 13, 30. doi:10.1186/s40561-026-00454-0 Relationship to seed: The authors directly cite Kestin et al. as evidence of strong gains in structured physics and frame their study as testing whether such benefits generalize to a complex non-STEM domain. One-sentence summary: A semester-long RCT with 84 undergraduates found that adding a bespoke AI teacher to traditional literature instruction improved objective tests and analytical essays while reducing extraneous and increasing germane cognitive load. Research-question fit: Useful evidence for duration and cognitive-load mechanisms, but not a STEM problem-solving test. INFERRED / PARALLEL 6. Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), e2422633122. doi:10.1073/pnas.2422633122 Relationship to seed: I did not confirm a direct citation to the seed paper in accessible full text; it is included because it is the closest causal test of unrestricted answer-giving versus a learning-protective tutor design. One-sentence summary: In a field experiment with nearly 1,000 high-school math students, unrestricted GPT-4 raised assisted performance but led to 17% lower unassisted grades after access was removed, whereas a guarded GPT Tutor largely mitigated that learning penalty. Research-question fit: Probably the strongest current causal evidence that design and answer access determine whether AI becomes a crutch. INFERRED / ADJACENT 7. Contractor, Z., & Reyes, G. (2026). Experimental Evidence on the Learning Impact of Generative AI. arXiv preprint arXiv:2607.08849. arXiv:2607.08849 Relationship to seed: No direct citation to Kestin et al. was confirmed in the accessible version; it is selected because it directly fills the seed paper's retention and usage-mode gap. One-sentence summary: In a randomized experiment with 211 undergraduates, AI access improved unaided knowledge immediately and about one week later, but delayed gains were concentrated among augmentation users who sought explanations, while automation users' short-run gains vanished once AI was removed. Research-question fit: The clearest undergraduate evidence so far that augmentation versus automation predicts durable benefit, though the task is not a standard STEM course problem set. AI tutoring, cognitive offloading, and undergraduate STEM learning 10Evidence map centered on Kestin et al. (2025) | literature checked through 18 Aug 2026 INFERRED / ADJACENT 8. Liu, G., Christian, B., Dumbalska, T., Bakker, M. A., & Dubey, R. (2026). AI Assistance Reduces Persistence and Hurts Independent Performance. arXiv preprint arXiv:2604.04721. arXiv:2604.04721 Relationship to seed: The accessible full text contains no Kestin citation; it is included as mechanism-level causal evidence that immediate AI assistance can undermine independent skill and persistence. One-sentence summary: Across randomized experiments totaling 1,222 participants, AI assistance improved performance while available but reduced later unassisted performance and increased giving up after only brief exposure, especially when users obtained direct answers. Research-question fit: Strong causal evidence for a cognitive-offloading/persistence mechanism, with less direct correspondence to authentic undergraduate STEM courses. INFERRED / ADJACENT 9. Rismanchian, S., Uzun, H., Matayoshi, J., Cosyn, E., & Kurd-Misto, E. (2026). Faster Completion, Less Learning: Generative AI Reduced Study Time on Math Problems and the Knowledge They Build. arXiv preprint arXiv:2605.21629. arXiv:2605.21629 Relationship to seed: The paper does not cite Kestin et al. in the accessible version; it is included because it tests long-run college-math behavior and delayed proctored retention at scale. One-sentence summary: Using 3.2 million ALEKS learning interactions over a decade, the authors estimate a 26.9% post-ChatGPT decline in time spent on AI-susceptible college math problems and a 25% cumulative decline in the odds of correct responses on later proctored retention items. Research-question fit: Highly relevant to durable STEM learning, but the quasi-experimental design infers AI use from task susceptibility rather than observing prompts directly. INFERRED / SYNTHESIS 10. Wang, G., Wang, W., Yang, D., & Ren, J. (2026). Generative AI, Cognitive Offloading, and Learner Agency in Higher Education: A Scoping Review. Behavioral Sciences, 16(7), 1150. doi:10.3390/bs16071150 Relationship to seed: I did not find a direct Kestin citation in the accessible article; it is included as a current higher-education synthesis of the augmentation-versus-replacement mechanism central to the seed paper's design logic. One-sentence summary: Synthesizing 123 higher-education studies, the review concludes that scaffolded, self-regulated, augmentation-oriented AI use is associated with learner agency, whereas replacement-oriented use clusters with offloading, overreliance, dependence, and weakened judgment. Research-question fit: A useful map of the broader literature, but its conclusions are configurative rather than causal and not STEM-specific. AI tutoring, cognitive offloading, and undergraduate STEM learning 11Evidence map centered on Kestin et al. (2025) | literature checked through 18 Aug 2026 4. What the 20-paper map says about your research question Sub-question Best current reading Most probative papers Confidence / gap Does AI improve immediate problem-solving performance? Often yes, especially when explanations, feedback, or tutoring are available; assisted performance is not the same as learning. Kestin 2025; Kumar 2023; Bastani 2025; Miller 2026 High confidence for immediate performance; interpretation must separate assisted output from unaided knowledge. Does AI improve durable, unaided learning? Conditionally. One-week undergraduate gains can persist under augmentation, but unrestricted solution access can reduce later unassisted performance or retention. Contractor & Reyes 2026; Bastani 2025; Liu et al. 2026; Rismanchian et al. 2026 Moderate confidence; direct long-duration undergraduate STEM RCTs remain scarce. Is there a cognitive-offloading / deskilling effect? Yes under some usage patterns. Evidence appears as copying/over-trust, reduced persistence, reduced time-on-task, and weaker unassisted performance. Krupp 2024; Bastani 2025; Liu et al. 2026; Rismanchian et al. 2026 Moderate-to-high for mechanism; ecological and causal strength varies by study. Does answer-giving versus Socratic/scaffolded tutoring matter? Very likely yes. Attempt-before-help, pedagogical prompts, guardrails, and Socratic questions consistently point toward greater engagement, reflection, transfer, or preserved independent ability. Kumar 2023; Miller 2026; LearnLM 2025; Degen 2025; Bastani 2025; SafeTutors 2026 Strong convergent evidence, but few studies randomize the exact same model into answer-giving versus Socratic modes with delayed STEM tests. Can we claim long-term retention benefits in undergraduate STEM? Not yet. The seed is short-term; the strongest delayed evidence either uses non-STEM tasks, younger students, broad participant samples, or quasi-experimental college math data. Kestin 2025; Contractor & Reyes 2026; Rismanchian et al. 2026 Low-to-moderate; this is the central research gap. The emerging mechanism A useful working model for a new study AI as cognitive amplifier: student attempts first -> AI diagnoses or asks a question -> targeted hint/explanation -> student produces the next reasoning step -> delayed unaided retrieval/transfer. AI as cognitive substitute: student delegates the reasoning step -> AI returns a complete solution -> student verifies superficially or copies -> immediate task success rises while practice opportunities, persistence, and durable schema construction fall. A rigorous next RCT would hold content, model, time, and interface constant while randomizing only the assistance policy: (A) full answer on request, (B) explanation after an independent attempt, and (C) Socratic/hint-based tutoring that withholds full solutions until predefined thresholds. Outcomes should include immediate mastery, unaided transfer, delayed retention at one week and several weeks, time-on-task, persistence, hint/answer-seeking behavior, and process traces that distinguish productive self-explanation from cognitive surrender. The most informative primary endpoint for your exact question is not homework completion or post-test performance while AI is present; it is delayed unaided performance on novel or isomorphic problems after AI removal. A secondary endpoint should capture whether students continue attempting difficult problems before seeking an answer. AI tutoring, cognitive offloading, and undergraduate STEM learning 12Evidence map centered on Kestin et al. (2025) | literature checked through 18 Aug 2026 5. Selection method, relationship audit, and limitations Selection rule. I prioritized papers that are closest to the target construct: undergraduate or STEM problem solving; objective learning or retention; manipulation or observation of AI assistance style; scaffolding/Socratic interaction; and evidence of cognitive offloading, overreliance, or independent performance after AI removal. Canonical papers were retained only when they directly explain the seed tutor's design. Relationship verification. "Confirmed antecedent" means the work is cited by the uploaded seed paper. "Confirmed extension/descendant" means the newer paper directly cites the published seed or its 2024 preprint, or explicitly calls itself an extension. "Inferred / adjacent" means the paper was selected because it tests the same causal mechanism, but a direct citation relationship was not confirmed in the accessible full text searched. Publication-status caveat. Several of the most directly relevant 2026 studies are arXiv preprints. Their results should be treated as provisional until peer review and replication. Conversely, peer-reviewed status does not remove limitations in population, intervention duration, or causal identification. Bibliometric caveat. This report is a targeted evidence map, not an exhaustive citation census. The seed article has accumulated a fast-moving citing literature, and some recent papers may be indexed unevenly across databases. The ten recent selections are therefore optimized for your research question, not for raw citation count. Recent paper Relationship status used here Verification basis Miller, K et al. CONFIRMED EXTENSION Direct statement/citation located in accessible full text. LearnLM Team, Google & Eedi et al. CONFIRMED DESCENDANT Direct statement/citation located in accessible full text. Hazra, R et al. CONFIRMED DESCENDANT Direct statement/citation located in accessible full text. Degen, P et al. CONFIRMED PREPRINT-LINEAGE Direct statement/citation located in accessible full text. Sun, Y et al. CONFIRMED DESCENDANT Direct statement/citation located in accessible full text. Bastani, H et al. INFERRED / PARALLEL No direct seed citation confirmed in accessible full text; thematic relationship is an inference. Contractor, Z et al. INFERRED / ADJACENT No direct seed citation confirmed in accessible full text; thematic relationship is an inference. Liu, G et al. INFERRED / ADJACENT No direct seed citation confirmed in accessible full text; thematic relationship is an inference. Rismanchian, S et al. INFERRED / ADJACENT No direct seed citation confirmed in accessible full text; thematic relationship is an inference. Wang, G et al. INFERRED / SYNTHESIS No direct seed citation confirmed in accessible full text; thematic relationship is an inference. Primary source trail for recent papers • Miller (2026): Springer article and notes explicitly identify the same AI tutoring framework as Kestin et al. and call the study an extension. • LearnLM Team (2025): arXiv full text directly lists Kestin et al. (2025) in its references and discusses mixed evidence on guarded versus unguarded AI tutoring. AI tutoring, cognitive offloading, and undergraduate STEM learning 13Evidence map centered on Kestin et al. (2025) | literature checked through 18 Aug 2026 • Sun & Liu (2026): Springer/Smart Learning Environments article directly cites Kestin et al. when motivating the gap beyond structured physics. • SafeTutors (2026): arXiv full text directly cites Kestin et al. and defines answer over-disclosure and lost scaffolding as pedagogical safety risks. • Degen & Asanov (2025): arXiv PDF directly cites the 2024 Research Square preprint version of the Kestin study. • Bastani et al. (2025), Contractor & Reyes (2026), Liu et al. (2026), Rismanchian et al. (2026), and Wang et al. (2026): included for topical/causal relevance; no direct Kestin citation was confirmed in the accessible full text searched. 6. Compact bibliography: the 20 selected related papers Seed paper is listed on the cover and in Section 1; the 20 entries below are the selected antecedent and recent papers requested. 1. Bloom, B. S. (1984). The 2 sigma problem: The search for methods of group instruction as effective as one-to-one tutoring. Educational Researcher, 13(6), 4-16. doi:10.3102/0013189X013006004 2. Wood, D., Bruner, J. S., & Ross, G. (1976). The role of tutoring in problem-solving. Journal of Child Psychology and Psychiatry, 17(2), 89-100. 3. Sweller, J. (2011). Cognitive load theory. In J. P. Mestre & B. H. Ross (Eds.), The Psychology of Learning and Motivation: Cognition in Education (pp. 37-76). Elsevier Academic Press. 4. Shute, V. J. (2008). Focus on formative feedback. Review of Educational Research, 78(1), 153-189. doi:10.3102/0034654307313795 5. Hake, R. R. (1998). Interactive-engagement versus traditional methods: A six-thousand-student survey of mechanics test data for introductory physics courses. American Journal of Physics, 66(1), 64-74. doi:10.1119/1.18809 6. Freeman, S., Eddy, S. L., McDonough, M., Smith, M. K., Okoroafor, N., Jordt, H., & Wenderoth, M. P. (2014). Active learning increases student performance in science, engineering, and mathematics. Proceedings of the National Academy of Sciences, 111(23), 8410-8415. doi:10.1073/pnas.1319030111 7. Deslauriers, L., McCarty, L. S., Miller, K., Callaghan, K., & Kestin, G. (2019). Measuring actual learning versus feeling of learning in response to being actively engaged in the classroom. Proceedings of the National Academy of Sciences, 116(39), 19251-19257. doi:10.1073/pnas.1821936116 8. Kumar, H., Rothschild, D. M., Goldstein, D. G., & Hofman, J. M. (2023). Math education with large language models: Peril or promise? SSRN Electronic Journal. doi:10.2139/ssrn.4641653 9. Krupp, L., Steinert, S., Kiefer-Emmanouilidis, M., Avila, K. E., Lukowicz, P., Kuhn, J., Küchemann, S., & Karolus, J. (2024). Unreflected acceptance - Investigating the negative consequences of ChatGPT-assisted problem solving in physics education. In HHAI 2024: Hybrid Human AI Systems for the Social Good (pp. 199-212). doi:10.3233/FAIA240195 10. Forero, M. G., & Herrera-Suárez, H. J. (2023). ChatGPT in the classroom: Boon or bane for physics students' academic performance? arXiv preprint arXiv:2312.02422. arXiv:2312.02422 11. Miller, K. (2026). Prompt Matters: How Pedagogical Engineering Shapes Behavior and Engagement with AI Tutors. Technology, Knowledge and Learning. doi:10.1007/s10758-026-09983-6 12. LearnLM Team, Google & Eedi. (2025). AI tutoring can safely and effectively support students: An exploratory RCT in UK classrooms. arXiv preprint arXiv:2512.23633. arXiv:2512.23633 13. Hazra, R., Ghuku, B., Marchenko, I., Tokarieva, Y., Layek, S., Banerjee, S., Stoyanovich, J., & Pechenizkiy, M. (2026). SafeTutors: Benchmarking Pedagogical Safety in AI Tutoring Systems. arXiv preprint arXiv:2603.17373. arXiv:2603.17373 14. Degen, P.-B., & Asanov, I. (2025). Beyond Automation: Socratic AI, Epistemic Agency, and the Implications of the Emergence of Orchestrated Multi-Agent Learning Architectures. arXiv preprint arXiv:2508.05116. arXiv:2508.05116 15. Sun, Y., & Liu, F. (2026). The impact of an AI Digital Teacher on human-AI collaborative learning in higher education. Smart Learning Environments, 13, 30. doi:10.1186/s40561-026-00454-0 16. Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. Proceedings of the National Academy of Sciences, 122(26), e2422633122. doi:10.1073/pnas.2422633122 17. Contractor, Z., & Reyes, G. (2026). Experimental Evidence on the Learning Impact of Generative AI. arXiv preprint arXiv:2607.08849. arXiv:2607.08849 18. Liu, G., Christian, B., Dumbalska, T., Bakker, M. A., & Dubey, R. (2026). AI Assistance Reduces Persistence and Hurts Independent Performance. arXiv preprint arXiv:2604.04721. arXiv:2604.04721 19. Rismanchian, S., Uzun, H., Matayoshi, J., Cosyn, E., & Kurd-Misto, E. (2026). Faster Completion, Less Learning: Generative AI Reduced Study Time on Math Problems and the Knowledge They Build. arXiv preprint arXiv:2605.21629. arXiv:2605.21629 20. Wang, G., Wang, W., Yang, D., & Ren, J. (2026). Generative AI, Cognitive Offloading, and Learner Agency in Higher Education: A Scoping Review. Behavioral Sciences, 16(7), 1150. doi:10.3390/bs16071150