LearnLM Team, Google & Eedi (2025). AI tutoring can safely and effectively support students: An exploratory RCT in UK classrooms. + Full-text exploratory RCT / technical report (not abstract-based). + Exploratory classroom RCT over seven consecutive weeks with two levels of randomization: students were assigned to static-hint control (N=91) or interactive tutoring (N=74); tutoring sessions were then randomized to a human tutor alone or human-supervised LearnLM. Quantitative outcomes were analyzed with Bayesian regression; the study also audited LearnLM messages and used tutor/student surveys plus tutor interviews. + N=165 Year 9-10 students (ages 13-15) from five UK secondary schools using the Eedi mathematics platform; N=17 expert tutors delivered/supervised interactive tutoring. Post-trial student survey N=27; tutor interview subset N=5. + Secondary-school mathematics on Eedi diagnostic study units; curriculum-aligned multiple-choice problems with identified misconceptions. + Student-level condition: static pre-written hint vs interactive tutoring. Within tutoring: session-level assignment to human tutor alone vs supervised LearnLM. Baseline performance was included as a covariate in learning-outcome regressions. + LearnLM, a pedagogically fine-tuned model based on Gemini 2.0 Flash. The system prompt instructed a concise, Socratic response aimed at helping the student self-correct a misconception without revealing the answer. Every draft was reviewed by a human tutor, who could approve, edit, or replace it. + Static, pre-written misconception-specific hints; human tutor alone. The transfer analysis also reports a natural benchmark for students who needed no intervention on the preceding unit. + Primary learning outcomes: mistake remediation; misconception resolution; knowledge transfer (correctness on the first question of the next study unit). Additional outcomes: message approval/edit/error rates, tutor perceptions/comfort, student helpfulness ratings, and qualitative tutor interview themes. + Seven-week trial (May 13-June 30, 2025). Support interventions were triggered after an incorrect first question in a study unit; frequency and timing therefore varied with each student's platform activity/performance. + Mistake remediation was measured on the retry immediately after an intervention; misconception resolution was measured within the same unit. Knowledge transfer was measured on the initial question of the next study unit; the transfer analysis was restricted to cases where the next sequential unit was attempted on the same day. + Knowledge transfer to the next study unit. Model-estimated success: LearnLM 66.2% [61.1%, 71.2%], human tutor 60.7% [55.8%, 65.4%], static hint 56.2% [54.2%, 58.2%]. The paper describes this as its best within-RCT test of durable effects, while stating that substantive longer-term effects require a different longitudinal design. + Tutors consistently highlighted LearnLM's Socratic dialogue. In the subset of drafts they edited or rewrote, the most frequent motivation was moderating pedagogical pacing (44.3% of edits); tutors reported that persistent Socratic questioning could frustrate some students and require intervention. + Interactive tutoring substantially outperformed static hints on immediate learning outcomes. Supervised LearnLM and human tutoring were similar on immediate remediation/resolution. For knowledge transfer, LearnLM was estimated at +5.5 percentage points versus human tutoring (95% credible interval -1.4 to +12.4; posterior probability LearnLM > human 93.6%) and +10.1 points versus static hints (+4.6 to +15.4; >99.9%). + LearnLM was explicitly instructed to use Socratic, answer-withholding guidance. The study compares this design with human tutoring and static hints rather than with an answer-giving AI. It reports strong immediate learning and next-unit transfer, while stating that longer-term cumulative learning was not isolated by the session-level design. + Session-by-session randomization prevented isolation of the cumulative impact of working with LearnLM over time. Tutors reported learning from LearnLM, creating possible crossover into human-only sessions. The authors state that longer-term effects require consistent support over several months and external standardized assessments. They also state that the mathematics setting, with precise answers and validated misconception explanations, provides limited evidence for more interpretive subjects. The trial design also precluded rigorous throughput/efficiency measurement. + Uploaded: AI tutoring can safely and effectively support students.pdf; arXiv:2512.23633. + Akçapınar, G., & Sidan, E. (2024). AI chatbots in programming education: guiding success or encouraging plagiarism. + Full-text peer-reviewed research article (not abstract-based). + One-group pretest-posttest quasi-experimental design. Students completed the same 13-question programming exam in two 30-minute sessions: first without AI, then with optional AI assistance. Overall scores were compared with a paired-samples t-test and Cohen's d. For one question known to elicit an incorrect AI response, the authors used McNemar's test plus qualitative analysis of student answers and AI interaction logs. + 45 first-year Educational Technologies students at a public university in Türkiye (26 male, 19 female) enrolled in Introductory Programming; most had no prior programming knowledge and this was their first university programming course. + Introductory C# programming; topics included operators, conditionals, type conversions, loops, arrays, methods, file operations, and debugging. + Within-student attempt/AI condition: first exam attempt without AI; second attempt with access to the AI assistant at the student's discretion. The misinformation analysis examined student behavior when the assistant returned a known incorrect answer on a type-conversion question. + Custom assistant using OpenAI GPT-3.5 Turbo, restricted to C# topics. The interface supported conceptual explanations, code creation, code completion, and code explanation; all interactions were logged. + Each student's first exam attempt without AI assistance. + Programming exam score; correctness on the selected type-conversion question; qualitative coding of whether students directly used/copied AI-generated misinformation. + The assistant was introduced two weeks before data collection and students practiced using it on small tasks. Main data collection occurred in week 12 of the course as two 30-minute exam sessions. + The score comparison uses the first exam attempt without AI and the second, otherwise identical, exam attempt in which AI assistance was available. + For the selected question, 42 of 45 students used the assistant; it gave an incorrect answer to 36. Of those 36, 22 (61%) copied and pasted the AI response directly, 11 (31%) did not copy it verbatim but still answered incorrectly based on the suggestion, and 3 (8%) ignored the incorrect response and answered correctly. + Mean exam score increased from 48.33 without AI to 74.47 with AI; the paired difference was statistically significant with a large effect size (d=1.56). On the known-error item, 33 of 36 students (92%) who received the incorrect AI response answered incorrectly; 22 of those 36 (61%) copied the erroneous response directly. + The assistant could generate and complete code, and the paper directly documents uncritical use of AI-provided content: most students exposed to a clearly incorrect response accepted it, including direct copy/paste. The authors state that this behavior suggests students transfer the critical-thinking process to the AI tool when they use AI answers without questioning them. + The conclusion characterizes the work as a preliminary investigation conducted with a small group of students over a limited period of time. The authors call for future research on how students use AI-generated content and on learning designs/tools that enable students to learn from or alongside AI. + Uploaded: AI Chatbots in Programming Education.pdf; https://doi.org/10.1007/s44163-024-00203-7 + Lehmann, M., Cornelius, P. B., & Sting, F. J. (2025). AI Meets the Classroom: When Do Large Language Models Harm Learning? + Full-text arXiv preprint (not abstract-based). + Three-study design. Study 1: field panel data from two graduate programming courses, using code similarity to ChatGPT-generated solutions as a proxy for substitutive LLM use, with two-way fixed effects and ChatGPT outages as instruments in FE2SLS. Studies 2 and 3: pre-registered, incentivized randomized laboratory experiments comparing LLM access during a 45-minute learning phase with no LLM; Study 3 replicated Study 2 with copy/paste enabled. Combined exploratory analyses manually coded prompts by usage behavior. + Study 1: 113 graduate students in two programming courses at a public Dutch university (56 information systems; 57 business analytics), with 6,594 analyzed student-question observations after exclusions. Study 2: 107 analyzed enrolled students at a public German university; 79% had no prior Python experience and 17% were beginners. Study 3: 69 analyzed students; 78% had no prior Python experience and 14% were beginners. + Python programming. Field courses ran from introductory Python through simple machine-learning models; laboratory learning covered basic Python through a sequence of explanations, examples, and coding exercises. + Study 1: current and cumulative ChatGPT Similarity, with ChatGPT outages/cumulative outages as instruments and multiple controls. Studies 2/3: randomized LLM access during the learning phase. Study 3/combined exploratory analyses also examine copy/paste availability, counts of solution requests, counts of explanation requests, and prior-knowledge interactions. + Lab treatment used unrestricted ChatGPT (gpt-3.5-turbo-0125) with arbitrary chat messages. Coded usage distinguished Solutions (asking the LLM to solve an exercise; substitutive use) from Explanations (asking the LLM to explain a concept; complementary use). Study 1's observed proxy captures solution-substitution behavior. + Laboratory experiments: no-LLM control condition. Field study: within-student and within-question fixed-effects comparisons using variation in ChatGPT-similar code, supplemented by outage-based instrumental variables. + Study 1: normalized grade on each coding question and effects of cumulative prior LLM use on later question grades. Labs: number of learning-phase practice questions solved (topic volume), post-test score (overall learning), post-test conditional on learning-phase progress and Covered Post-test (understanding), perceived learning, message/usage counts, and heterogeneous effects by pre-test knowledge. + Study 1: weekly homework across four- or five-week graduate programming courses. Labs: 20-minute pre-test, 45-minute learning phase, and 20-minute post-test; LLM access was available only during the learning phase. Study 3 occurred three weeks after Study 2 and enabled copy/paste. + In the laboratory experiments, the LLM was unavailable in the pre-test and post-test and available only during the 45-minute learning phase. In the field study, cumulative prior ChatGPT-similar submissions were used to estimate effects on grades for subsequent questions. + Study 1 estimates how cumulative prior solution-substitution is associated with performance on subsequent questions and describes the field result as a long-term decline in learning outcomes. The laboratory post-tests occur immediately after the learning phase. The authors state that their learning measures abstract away from memory. + Across treated lab subjects, 54.1% of coded messages were Solutions and 24.4% were Explanations. Enabling copy/paste significantly increased Solution requests but not Explanations. In Study 3, 42% of solution-request messages were sent without a single attempt to solve the corresponding question. + Randomized LLM access had no statistically significant overall learning effect in Study 2 or Study 3. In exploratory usage analyses, solution-seeking was associated with lower understanding while explanation-seeking was associated with higher understanding. In Study 1, cumulative solution-substitution predicted lower grades on subsequent questions (FE coefficient -0.02; FE2SLS coefficient -0.06). + The paper explicitly separates solution-seeking as substitution from explanation-seeking as complementarity. Solution requests reduced understanding, while explanation requests increased understanding; easier copy/paste increased solution-seeking. These measured usage categories directly address the distinction between getting solutions and using AI to support one's own learning activity. + The authors state that the exploratory findings should be replicated in controlled experiments and with additional field data. They also state that their outcome framework abstracts away from neurological building blocks of understanding, including conceptual/factual understanding and memory. + Uploaded: AI Meets the Classroom.pdf; arXiv:2409.09047v2. + Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. + Uploaded file is a PNAS HTML reader shell containing the article DOI/metadata rather than article body text. Matrix extraction used the full text of the same paper/DOI from the author-hosted PDF; not abstract-based. + Large-scale randomized controlled field experiment. About 50 9th-, 10th-, and 11th-grade classes completed four 90-minute mathematics sessions. Classrooms were assigned to one of three arms (control, GPT Base, GPT Tutor) and kept the same arm across all four sessions. Each session included teacher review, an assisted practice period, and a subsequent unassisted closed-book/closed-laptop exam. Primary analysis was pre-registered and used intention-to-treat regression with classroom-clustered standard errors. + Nearly 1,000 high-school mathematics students at one large high school in Turkey; main analysis sample N=839 students (277 GPT Tutor, 320 control, 242 GPT Base) after excluding honors-designated classes and students without the baseline survey. + High-school mathematics review and practice. Across the four sessions, the material collectively represented about 15% of each grade's semester mathematics curriculum. + Class-level randomized treatment arm: control (books/notes, no device), GPT Base, or GPT Tutor. Previous-year normalized GPA was included as a covariate; the main model also included session, grader, grade-level, and teacher fixed effects. + Both tools used GPT-4. GPT Base mimicked a standard ChatGPT-like interface and followed student instructions. GPT Tutor added teacher-designed problem solutions/common mistakes plus guardrails: do not provide the full solution, require students to show work before help, point out mistakes, give minimal hints first, and guide step-by-step. + Business-as-usual control with course books and notes and no AI/device during practice; GPT Base and GPT Tutor are also directly compared. + Normalized grade on assisted practice problems; normalized grade on the subsequent unassisted exam (pre-registered primary learning outcome); student perceptions; message/conversation behavior; robustness/heterogeneity and grade-dispersion outcomes. + Four 90-minute in-class sessions during Fall 2023-2024; each session contained the treatment-relevant assisted practice period followed by the unassisted exam. + Practice performance was measured while GPT access was available in the two GPT arms. Learning was measured in the third part of each same session with a closed-book, closed-laptop exam and no resources; each exam problem corresponded to a conceptually similar practice problem. + Short-term unassisted exam performance after AI removal. The paper explicitly states that its analysis focuses on short-term exam performance rather than long-term learning and defers long-term learning to future work. + Mechanism analysis concludes that using GPT Base as a crutch was the main mechanism impeding learning. The most common GPT Base message was "What is the answer?"; interaction analysis found students often asked for/copied solutions with GPT Base, whereas GPT Tutor produced more substantive interactions such as attempted answers and requests for help. + Relative to control, GPT Base increased assisted-practice performance by 48% and GPT Tutor by 127%. On the subsequent unassisted exam, GPT Base reduced performance by 17% relative to control; GPT Tutor was statistically indistinguishable from control in the main regression, so its guardrails largely mitigated the negative learning effect but did not produce a positive unassisted-exam effect. + This study directly contrasts a ChatGPT-like, solution-access condition with a guardrailed hint/step-by-step tutor that withholds full solutions. The unrestricted GPT Base arm improved assisted task performance but reduced subsequent unassisted performance, while GPT Tutor largely removed that penalty. The authors also state that GPT Tutor remains passive and does not proactively ask the probing questions characteristic of effective human tutors. + The authors state that the study covers two tutor designs, one topic area (mathematics), and one high school in Turkey; math has objective evaluation criteria that may not transfer to creative subjects. The deployment occurred in Fall 2023, when generative AI was newer and models/users may differ today. Generalizability to other tutor designs and deployment contexts requires further study. Outcomes are short-term due to partner-school constraints; long-term outcomes remain future work. The authors also call for controlled lab experiments to clarify learning mechanisms. + Uploaded: Generative AI without guardrails can harm learning_ Evidence from high school mathematics.html; DOI: https://doi.org/10.1073/pnas.2422633122; same-paper full text used for extraction: https://hamsabastani.github.io/education_llm.pdf + Structured AI Tutoring vs. Cognitive Offloading in Undergraduate STEM: A citation-lineage and evidence map centered on Kestin et al. (2025). + Full-text secondary citation-lineage/evidence map, not a primary empirical study (not abstract-based). + Targeted evidence map centered on Kestin et al. (2025). It selected 20 related papers for conceptual lineage and research-question relevance: 10 confirmed antecedents cited by the seed paper and 10 recent descendants/adjacent studies. Recent-source verification was performed through 18 Aug 2026, with direct citation/extension relationships labeled separately from inferred thematic relationships. + Evidence-map scope prioritizes undergraduate or STEM problem-solving, objective learning/retention, AI assistance style, scaffolding/Socratic interaction, and cognitive-offloading/overreliance or independent performance after AI removal. + The map organizes the literature around structured/scaffolded/Socratic/augmentation-oriented use versus answer-giving/offloading risk, while also distinguishing AI access from pedagogical design and user behavior. + The map defines an "AI as cognitive amplifier" pathway in which students attempt first and AI diagnoses/asks questions before hints/explanations, and an "AI as cognitive substitute" pathway in which students delegate reasoning, receive a complete solution, verify/copy superficially, and lose practice/persistence opportunities. + Stated bottom line of the evidence map: structured AI that keeps learners cognitively active can improve learning, whereas open-ended answer-giving can raise assisted performance while reducing independent performance, persistence, study time, or retention after the tool is removed. + The map explicitly identifies the exact long-duration undergraduate STEM experiment of interest - randomizing answer-giving versus Socratic scaffolding and measuring delayed, unaided transfer/retention - as under-tested. It recommends distinguishing immediate AI-assisted performance from delayed unaided performance. + The report states that it is a targeted evidence map, not an exhaustive citation census; recent citing literature is fast-moving and may be unevenly indexed. Several highly relevant 2026 studies are arXiv preprints. The recent selections were optimized for the research question rather than raw citation count. + Uploaded: Custom Prompt AI_tutoring_cognitive_offloading_literature_map.pdf. +