Leveraging AI for Simulations Finkelstien

Details

Filename
Leveraging AI for Simulations Finkelstien.pdf
Size
953.5 KB
Type
application/pdf
Published

Extracted text

Leveraging generative artificial intelligence for simulation-based physics experiments:
A new approach to virtual learning about the real world
Yossi Ben-Zion ,1,* Turhan K. Carroll ,2,*,† Colin G. West ,3,*
Jesse Wong ,3 and Noah D. Finkelstein 3
1Department of Physics, Bar-Ilan University, Ramat Gan IL, 52900, Israel
2Department of Workforce Education and Instructional Technology, University of Georgia,
Athens, Georgia 30602, USA
3Department of Physics, University of Colorado Boulder, Boulder, Colorado 80309, USA
(Received 26 September 2025; accepted 22 December 2025; published 26 January 2026)
This study investigates the impact of a novel application of generative artificial intelligence (AI) in
physics instruction: engaging students in prompting, refining, and validating AI-constructed simulations of
physical phenomena. In a second-semester physics course for life science majors, we conducted a
comparative study of three instructional approaches in a laboratory focused on electric potentials:
(i) students using physical equipment, (ii) students using a prebuilt simulator, and (iii) students using AI to
generate a simulation. Among the groups, we found significant differences in performance on conceptual
assessments of the laboratory content (η2 ¼ 0.359). Post hoc analysis showed that students in both the
AI-generated and prebuilt simulation conditions scored significantly higher on the conceptual assessments
than students in the physical equipment condition. Students in these groups also reported more favorable
perceptions of the learning experience. Finally, this preliminary study highlights opportunities
for developing students’ modeling skills through the processes of designing, refining, and validating
AI-generated simulations.
DOI: 10.1103/s8dy-kqy5
I. INTRODUCTION
Generative artificial intelligence (AI) is rapidly trans-
forming our educational practices. Already these tools
appear to be able to solve introductory physics problems
at a passing level [1], to compete with or outperform typical
students at various undergraduate conceptual assessments
[1–3], and to support student learning and success—in
some cases even better than humans working with state-of-
the-art curricula and pedagogical practices [4]. At the same
time, there is increasing attention to where, when, and how
to use these tools within the physics teaching community
[5–8], and we have observed a wide variation in the
approaches that both faculty and students take to using
generative AI. Given the apparent inevitability of the
changes brought by AI tools, the physics education
community ought to consider where these technologies
are headed and how they can effectively be used to support
new ways of learning in our course environments.
To such ends, we present one research-validated
approach to using generative AI productively in our physics
classrooms. At heart, this approach draws from a long-
standing notion of learning-by-teaching [8,9]. In this case,
the students are “teaching” generative AI, or more pre-
cisely, prompting generative AI to produce effective sim-
ulations for modeling physics phenomena. A cornerstone
of the process is having students validate both the technical
and scientific aspects of these simulations, and iteratively
prompt (or “teach”) the AI to produce increasingly accurate
models of physics phenomena. Students can then use these
simulations to further explore the physics content. We hope
that this pedagogical approach is more-or-less evergreen
and applicable no matter the capacities of generative AI.
This present work may be seen as an extension of early
work with the PhET simulations [10], where we documented
that working with simulations supported student learning
about the physical world. In fact, students working first
with simulators and then with real circuit components
outperformed students who only worked with the real
equipment. Students using the simulators did better both on
measures of conceptual survey of related topics (series and
parallel circuits) and on the ability to conduct an experi-
ment—physically manipulate equipment and describe the
experiment and its outcomes. The takeaway here is not that
simulations are more effective than laboratory equipment
per se, but that exposure to simulations can be a valuable
*These authors contributed equally to this work.
†Contact author: tkcarroll@uga.edu
Published by the American Physical Society under the terms of
the Creative Commons Attribution 4.0 International license.
Further distribution of this work must maintain attribution to
the author(s) and the published article’s title, journal citation,
and DOI.
PHYSICAL REVIEW PHYSICS EDUCATION RESEARCH 22, 010109 (2026)
Editors' Suggestion
2469-9896=26=22(1)=010109(14) 010109-1 Published by the American Physical Society

approach to prepare students for working effectively with
laboratory equipment (which is a skill we still value). In this
work, we similarly demonstrate a technique based on
simulation usage—this time updated with the new capabil-
ities of generative AI—which supports both conceptual
understanding and desirable affective outcomes, while also
offering the potential for developing scientific modeling
skills. In particular, the involvement of the AI tool creates a
space in which the student is both exploring conceptual
physics and evaluating the quality and limitations of the
physical simulation being used. In some ways, this context
could be even better than working with purely prebuilt
simulations as preparation for work with real laboratory
equipment, though, of course, still without offering many of
the valuable experiences that only a true hands-on lab can
replicate.
The present study explores the possibility of using
generative AI to support student engagement and under-
standing of basic physics phenomena. In particular, this
paper examines:
(1) How does the approach documented here (having
students prompt, validate, and refine a simulation
produced by generative AI) impact students’ con-
ceptual understanding as compared to using a
prebuilt (AI-designed) simulation or compared to
using physical equipment in a traditional physics
laboratory focused on equipotential lines?
(2) How do students reflect on the value, ease of use,
and enjoyment of using an AI-produced simulation
around equipotential lines (again, as compared to
prebuilt simulations and physical equipment)?
In this work, we further explore preliminary evidence
about the potential impact on modeling skills and other
laboratory learning goals. Ultimately, this approach to
deploy generative AI in our classes is designed to support
some of the core learning objectives in undergraduate
physics classes, but with the added benefits of teaching
students how to use these new emerging technologies
effectively and engaging them in authentic science, tech-
nology, engineering, and mathematics practices related to
model-building and evaluation.
II. METHODS
A. Research design and topic selection
The study was conducted during the spring of 2025 at a
large, R1 public university in the western United States.
Data were collected in the second semester of an intro-
ductory physics for life sciences (“IPLS”) course, taught
without calculus. Initial enrollment for the course was 187
students, of whom 54.5% were nonmale identifying.
Students represented a mixture of different majors, pri-
marily from life science disciplines and pre-med tracks; the
largest of these groups was integrated physiology students
(44%), followed by molecular, cellular, and develop-
mental biology (“MCDB,” 24%). Notably, despite being
an “introductory” physics course, the student population
contained no freshmen, consisting instead of sophomores
(6%), juniors (28%), seniors (38%) and a significant
“postbaccalaureate” cohort (28%), already possessing an
undergraduate degree, but returning to school to complete
physics as a requirement for further postgraduate study
such as med school, vet school, or dental school.
The intervention took place in the laboratory component
of the course, which comprised two of the five credits of the
course. Weekly labs were run by graduate student teaching
assistants graded primarily for active participation and
constituted 10% of the students’ final course grade. The
AI-enabled methods studied here were deployed during the
third such laboratory meeting of the semester, with data
collected immediately at the end of the lab period, on the
subsequent midterm exam one week later, and on the final
exam approximately two and a half months later.
To examine differences in conceptual understanding and
student perceptions resulting from different approaches to
teaching the same physics content with varying forms of
AI-engagement, a comparative study was designed with
three quasi-experimental groups.
The study investigated the effects of:
• The unmodified “physical” laboratory which involved
hands-on use of equipment to generate and measure
real, physical quantities. This approach served as a
control group.
• A lab involving the use of a “prebuilt” digital simu-
lation, designed by author Ben-Zion using AI tools, but
shared with the students only in its final form.
• An “AI-engagement” lab activity in which students
used AI tools to generate and test their own simu-
lation, but otherwise undertook the same tasks as the
students with the “prebuilt” simulation.
The three lab conditions are depicted in Fig. 1. The study
created three independent groups, where each course par-
ticipant was exposed to only one of the instructional
approaches under investigation [11]. This design choice
was made to ensure that comparisons reflect the unique
FIG. 1. Three conditions of the laboratory: Physical equipment
(top left), prebuilt simulation (top right), AI engaged simulation
design (bottom).
YOSSI BEN-ZION et al. PHYS. REV. PHYS. EDUC. RES. 22, 010109 (2026)
010109-2

influence of each approach without confounding from other
course or student factors. There were seven laboratory
sections in the course taught by four teaching assistants.
Three teaching assistants were responsible for two sections
each, with the fourth handling the remaining section. This
fourth TA, with only one section, was assigned to teach the
lab in the traditional physical laboratory format; the remain-
ing three were randomly assigned to one section using the
prebuilt simulation and one section using the AI-engagement
method.
The focal concepts for the labs during this period were
electric potential and equipotential lines. The choice to
deploy the interventions during this week was motivated by
both the fundamental importance of electrostatic potential
in the undergraduate physics curriculum and the specific
pedagogical challenges it presents, along with the oppor-
tunities of simulation and AI-assisted outcomes. Unlike
more tangible physical concepts such as force or motion,
electric potential is an abstract quantity that students may
conceptualize in a variety of ways [12], presenting sub-
stantial pedagogical challenges. The topic is also often
identified as one of the most important pathways to
understanding subsequent topics in physics and engineer-
ing, such as circuit analysis [13,14], and failure to develop
familiarity with electric potential concepts can be a sig-
nificant obstacle to understanding subsequent course-
work [15].
B. Activity objectives and description
The activities in each of the three variations (physical,
prebuilt sim, and AI-engaged sim development) aimed to
introduce students to the concept of electric potential
through practical investigation of its spatial distribution
around various charge configurations. The activity focused
on three core learning objectives common to all groups.
The first objective of the lab was to develop an under-
standing of the quantitative relationship between electric
potential and distance from a point charge. Students were
expected to examine the formula V ¼ kq=r through mea-
surements at various points, compare measured values with
theoretical results, and understand how potential varies
with distance for both positive and negative charges.
The second objective focused on understanding equi-
potential lines and their properties for single charges and
multiple charge configurations. Students were required to
identify and map lines of constant potential, understand the
relationship between line shapes and charge type and
location, and investigate how multiple charges affect the
resulting patterns. Additionally, they examined locations
where potential becomes zero in systems with both positive
and negative charges.
The third objective addressed understanding the relation-
ship between equipotential line density and potential
differences, as well as the physical significance of motion
along these lines.
The physical laboratory group additionally investigated
the effects of different charge geometries (nonpoint charges)
on potential distribution and the explicit relationship between
equipotential lines and electric field lines. Both simulation
groups focused exclusively on point charges but included
additional digital manipulation capabilities.
Across all groups, the activity concluded with an
identical conceptual assessment consisting of five multi-
ple-choice questions addressing the common concepts and
learning objectives, and a feedback questionnaire regarding
the students’ perspectives on the given laboratory medium.
1. Physical equipment
Students in the traditional physical lab group took part in a
“tried and true” activity which has been employed in
essentially the same form as part of this coursework for a
decade or more. The activity had been designed by members
of the CU PER group, building on effective approaches in
the physics community at the time. The equipment for the
lab consisted of a large, shallow plastic storage tub filled
with a mildly conductive solution. By placing metal objects
at various points in the tub and connecting them to a battery
to establish potential differences, students can create a
variety of potential differences throughout the liquid, which
can then be measured by a multimeter. Using graph paper
positioned under the translucent base of the storage tub, they
can then map out equipotential lines from various charge
distributions to study their shape and properties.
For example, using a metal ring at the edges of the tub as
“ground,” and a narrow metal cylinder in the center as an
approximate “point charge,” students can observe the classic
1=r concentric-circular equipotential pattern of a lone point
charge on a 2D plane. In the activity, students first begin with
this exploration, followed by further investigation of what
happens under various changes to the system, such as
reversing the polarity of the battery or increasing the
magnitude of the voltage difference between the central
charge and the grounding ring. Subsequently, students were
prompted to explore concepts relating potential difference to
concepts like electric potential energy. In this instance,
students used an LED light with its leads connected to
various points in the tub, identifying, for example, that if both
leads fall on an equipotential, the bulb will not light up.
Finally, students explored more complex arrangements
of the equipment to create scenarios equivalent to the
presence of multiple point charges, exploring both quali-
tative and quantitative effects on the equipotential lines,
including identifying scenarios and specific points where
the electric potential might vanish relative to ground.
2. Prebuilt digital simulation
The preexisting simulation provided students with an
interactive digital interface for exploring electric potential.
The tool enabled students to add point charges of both
positive and negative signs to the workspace, adjust their
LEVERAGING GENERATIVE ARTIFICIAL … PHYS. REV. PHYS. EDUC. RES. 22, 010109 (2026)
010109-3

magnitudes using a slider control, and reposition them by
dragging. The simulation displayed real-time potential
values and distances from charges as students moved the
cursor across the screen. A key feature was the ability to
generate and visualize equipotential lines by clicking on
specific points, with options to display multiple lines
simultaneously in different colors. The interface included
reset and clear functions to remove charges or equipotential
lines and restart the investigation.
The activity, while not running in identical sequence to the
traditional lab activity, followed the same conceptual flow
and covered the same three activity objectives listed above,
with modest variation in the nature and focus on each
subtopic. The activity began with quantitative verification,
where students placed a single charge and measured poten-
tial values at various distances, comparing these measure-
ments with theoretical calculations using V ¼ kq=r. This
initial phase established familiarity with the simulation
interface while reinforcing the mathematical relationship
between potential and distance. Students then explored
superposition effects by adding multiple charges and inves-
tigating locations where the net potential becomes zero,
discovering that such points exist only when charges of
opposite signs are present. The investigation proceeded to
mapping equipotential lines, where students learned to
visualize regions of constant potential by clicking on points
and observing the resulting contours. Students systemati-
cally mapped multiple equipotential lines at regular voltage
intervals for both positive and negative single charges,
compared the resulting patterns, and subsequently inves-
tigated how the addition of a second positive charge altered
the equipotential line geometry near each charge, between
the charges, and far from both charges.
This approach emphasized guided exploration using a
ready-made tool, allowing students to focus immediately
on the physics concepts without technical barriers. The
preexisting simulation enabled rapid investigation of multi-
ple scenarios and parameter variations, facilitating pattern
recognition and conceptual understanding through iterative
experimentation. Students could concentrate on interpret-
ing results and making connections between mathematical
relationships and visual representations; however, they
remained users rather than creators of the simulation
environment.
3. AI simulation design
Students in the final group constructed their own electric
potential simulation through structured interaction with
Claude AI, running on its Sonnet 4 model, free version.1
The activity required students to generate code through
prompt engineering, systematically validate both interface
functionality and physical accuracy, and iteratively refine
the simulation through natural language feedback. This
approach combined a conceptual physics focus with
computational framing, as students needed to articulate
physical requirements precisely and verify that the resulting
simulation adhered to established electrostatic principles.
Notably, because of the use of the AI tools, the
simulation design process required no prior programming
knowledge—and indeed, given their fields of study, it
would be quite uncommon for students in this course to
have had formal programming instruction. Instead, we
employ an approach validated in a pilot study [5] among
students who similarly lacked any programming back-
ground. Students interacted with the AI model entirely
through natural language prompts, evaluating and refining
the generated simulations iteratively without needing to
read or edit code directly. We note that in courses with more
computational focus, or in contexts where more program-
ming background might be assumed, a hybrid approach
combining AI generation with direct code correction and
editing by the student might be more appropriate. However,
we caution that the code involved, even in a relatively
“simple” simulation, can consist of hundreds of lines,
parsing which can be a distraction from the underlying
physics, even for students with more coding experience.
The activity began with students receiving a compre-
hensive initial prompt designed to generate a complete
electric potential simulation. This prompt served as the
foundation for the entire learning experience, requiring
students to copy and paste detailed specifications into
Claude AI. The prompt read:
Task: Write a single HTML5 file (including JavaScript and
CSS) that displays an interactive simulation of electric
potential generated by multiple charges.
Requirements: Canvas Display: Create an element to
display the simulation. Use a two-pixel grid resolution
for accurate equipotential lines.
User interface: “Add Charge” button to place a new
charge at the center of the canvas (þ1 nC default).
Slider to adjust the selected charge value (−10 to
þ10 nC, 1-nC steps). Display all numerical values
with three decimal places (not in scientific notation).
Display the current charge value dynamically next to
the slider. Allow dragging charges to reposition them.
“Reset” button to clear all charges.
Electric potential: Calculate the potential at any point
using the formula: V ¼ kq=r, where k is Coulomb’s
constant (9 × 109), q is the charge value in nano-
coulombs (convert to SI units), and r is the distance
from the point to the charge. Define minimum allowed
distance from charges for potential calculation (to
handle near-field behavior). Slider controls the value
of the selected charge. Use colors to distinguish
1For those students who exhausted the number of allowed
prompts (varying from 5 to 12) in the free version, Claude
reverted to using Haiku 3.5. Notably, this caused some challenges
for students, who ended up working with their peers or the
slower, less powerful version of the AI engine.
YOSSI BEN-ZION et al. PHYS. REV. PHYS. EDUC. RES. 22, 010109 (2026)
010109-4

positive (red) and negative (blue) charges. Show each
charge’s value (e.g., “þ1 nC”) next to the charge.
When moving the cursor over the canvas, display in
the control panel the potential value (V), and in the
case of single charge only, also display the distance
from the charge (r) (in volts and meters).
Charge interaction: Click a charge to select it. Slider
controls the value of the selected charge. Use colors to
distinguish positive (red) and negative (blue) charges.
Show each charge’s value (e.g., “þ1 nC”) next to the
charge.
Following code generation from this initial prompt, the
activity evolved into a structured validation and enhance-
ment process, while otherwise tracking the form of the
activity with the prebuilt simulation. Students proceeded
through systematic testing phases to verify both technical
functionality and physical accuracy, followed by iterative
refinement through natural language communication with
the AI [5]. Notably, while the initial prompt used to produce
the simulation was given to the students, all subsequent
prompts were left open for the students to generate.
Following this design-and-verification phase, students
were required to independently formulate additional
prompts to extend the simulation’s capabilities and address
any issues that emerged during testing. This progression
transformed the initial code output into a fully functional
educational tool tailored to their specific learning objec-
tives. Apart from the interaction with artificial intelligence
and the creation and validation processes, the activities
within the lab and the physics content explored were
identical to those of the preexisting simulation.
C. Assessment of impacts
Students’ understanding of the focal concepts of the
laboratory was assessed at the end of the laboratory session.
Five conceptual questions (“CQs”) were attached to the
laboratory and completed by the students immediately
following their lab activity. These questions are included
in Appendix A.
Immediately after the lab, in addition to the conceptual
questions described above, students were surveyed about
their experiences in the laboratory and the particular
instructional approach used (physical equipment, prebuilt
simulation, or AI-engaged development). Five Likert-scale
questions probed student views:
1. How did you like this lab compared to last week’s
lab?
Response options ranged from “1. way worse”
to “5. way better.”
2. How easy was it for you to use the [physical
equipment/simulation/AI]?
Response options ranged from “1. extremely
difficult” to “5. extremely easy.”
3. Do you feel like the [physical equipment/AI/Sim]
helped you understand voltage/electric potential?
Response options ranged from “1. Not at all” to
“5. Enormously.”
4. Did you enjoy working with the [physical equip-
ment/AI/Simulator]?
Response options ranged from “1. Really didn’t
like” to “5. A great deal.”
5. Would you suggest we do this again?
Response options ranged from “1. Definitely
no” to “5. Definitely yes.”
Students were also given space to write either explanations
for their Likert-scale responses or to share other
unprompted sentiments.
Student performance on the midterm, 3 weeks later, and on
the final examination, 12 weeks later, was also collected.
Data were collected both on overall performance on the
midterm (MQs) and final (FQs). These data were desig-
ned to document any overall differences between the
samples. Given the common homework, interactive lectures,
and study sessions provided to all students after the labo-
ratory experience, we expected no differences in student
performance on the midterm and final examinations. These
measures were used to document the similarity of samples.
III. DATA ANALYSIS AND RESULTS
A. Analysis of student performance on the lab review,
midterm exam, and final exam
We performed several statistical tests to understand the
relationship between lab type and performance on the three
sets of content questions described above (CQs, MQs, and
FQs). Since this was an exploratory study with the goal of
TABLE I. Variables used for omnibus tests.
Variable Level of measurement Definition Value range
Lab type Nominal The type of lab a student participated in. AI (student-generated AI sim); Sim
(prebuilt AI sim); Lab (traditional lab)
CQTot Interval Total score on the review questions students completed
about electric potential as part of their lab activity.
0–100
MQTot Interval Student’s total score on the midterm exam. 0–100
FQTot Interval Student’s total score on the final exam. 0–100
LEVERAGING GENERATIVE ARTIFICIAL … PHYS. REV. PHYS. EDUC. RES. 22, 010109 (2026)
010109-5

comparing more than two groups along a categorical
variable (the type of lab experience, “Lab Type”), we
performed an omnibus test to assess whether there were any
statistically significant differences between groups in the
outcomes of interest. Table I shows the variables we used
for our omnibus tests.
Given that research question 1 aims to compare assess-
ment performance across lab types, we planned to use a
one-way analysis of variance (ANOVA) to assess mean
performance differences across lab types. We performed a
preliminary analysis to see whether our data satisfied the
six assumptions of ANOVA. The details of this preliminary
analysis are provided in Appendix B. We concluded that,
while most of the assumptions were satisfied, the residuals
of the dependent variables were not normally distributed.
As a result, we utilized nonparametric statistical techniques
in our analysis. We compared median differences across lab
type, as the median is the appropriate measure of central
tendency for nonparametric data [16].
In order to visually represent the medians, CQtot, MQtot,
and FQtot, across lab groups and highlight potential median
differences, we created bar charts (with error bars repre-
senting the standard error of the median). These bar charts
are shown in Figs. 2–4 below. Standard errors of the
medians were calculated using bootstrapping with replace-
ment [17,18]. Our procedure used 1000 bootstrap repli-
cates. This method for calculating confidence intervals was
used because of the non-normality of our data.
The figures above suggest that there are large differences
in medians of CQtot between the physical lab and other
conditions, and small differences in the medians for MTtot
and FEtot. In order to quantify the significance of these
differences, we used the Kruskal-Wallis test, a nonpara-
metric omnibus test that determines whether there are
statistically significant differences between the medians
of three or more groups and does not assume normality of
residuals [19].
The results of our Kruskal-Wallis analyses are in
Table II below:
We found that median performance on CQs differed
significantly across our three lab types, HCQTotð2Þ ¼
58.718, p < 0.001. We found that there was no significant
FIG. 3. Median scores for the midterm exam (MQtot). Error bars
represent the standard error of the median.
FIG. 4. Median scores for the final exam (FQtot). Error bars
represent the standard error of the median.
TABLE II. Kruskal-Wallis test results for each dependent
variable.
d.o.f. H statistic p value η2
CQTot 2 58.718 0.000 0.359
MQTot 2 0.837 0.658
FQTot 2 3.302 0.192
TABLE III. Results of Dunn’s test for CQTot.
Comparison Adjusted p value
AI vs Lab 0.000
AI vs Sim 0.258
Sim vs Lab 0.000
FIG. 2. Median scores for the conceptual questions (CQtot).
Error bars represent the standard error of the median.
YOSSI BEN-ZION et al. PHYS. REV. PHYS. EDUC. RES. 22, 010109 (2026)
010109-6

difference in median MQ performance across lab type,
HMQTotð2Þ ¼ 0.837, p ¼ 0.658, and the differences in FQ
scores across lab types were not found to be significant at
the .05 significance-level (HFQTotð2Þ ¼ 3.302, p ¼ 0.192).
The analysis in Table II establishes that the median CQ
scores–that is, the conceptual questions answered by
students immediately after their labs–showed significant
differences across the three groups. We used η2 to assess the
effect size of these differences [20] and found that the effect
was strong (η2 ¼ 0.359). In order to assess which groups
showed significant differences in performance, we used
Dunn’s post hoc test [21]:
We can see from Table III that there are no significant
group differences in performance on CQs between students
in the two simulation lab sections (AI and prebuilt sim).
However, there are statistically significant differences
between the AI and lab students (pAI-lab Adj < 0.001)
and between Sim and Lab students (pSim-lab Adj < 0.001).
We can conclude that students in the AI and Sim lab groups
both performed significantly better than students in the
traditional lab section.
Since this was an exploratory study, we also performed a
question-by-question analysis of the conceptual questions
to see if the results were consistent with our omnibus
analysis. The results of this analysis, which are included in
Appendix C, are consistent with the results reported above.
B. Analysis of student attitudes toward
laboratory conditions
We then examined responses to ascertain student attitudes
and perspectives on the laboratories (AQs), which were
included at the end of their lab activity. These questions were
analyzed individually as they were not designed to form a
singular construct. There were five questions, and each was
meant to probe students’ attitudes after completing their lab
activity. As appropriate, the wording used for the questions
was varied to address the particular lab setting in which the
student worked. The questions used a five-point Likert scale
(strongly disagree, disagree, neutral, agree, and strongly
agree), and each question was coded so that a higher score
indicated a more positive effect. For our analysis, we
collapsed the strongly disagree and disagree categories into
one category and collapsed the strongly agree and agree
categories into one category because, upon discussion, the
research team believed the varying “agree” and “disagree”
categories to be redundant. Prior work has suggested that
collapsing the five-point ordinal scale to a three-point ordinal
scale is appropriate in cases where respondents may use the
various “agree” and “disagree” categories redundantly [22].
Table IV lists the variables used for this analysis:
This analysis was done in three stages. First, contingency
tables were developed depicting the frequency of each
response option for each AQ within each lab type to inform
us about how many students in each lab type selected
disagree/neutral/agree. This analysis of each question
resulted in a contingency table. For example, the contin-
gency table for AQ1 is shown in Table V below (other
contingency tables are available upon request).
For the second phase of the analysis, we statistically
tested the association between lab type and student
response. This was done using Fisher’s exact test because
our contingency tables had cells containing a frequency that
was less than 5 [23]. Cramer’s V [20] was used as a
measure of effect size for this test. Our results for each AQ
are listed in Table VI below.
As the p value (two-tailed) obtained from Fisher’s exact
test is significant for each AQ (see p values and effect sizes
in Table VI), we see that there is a statistically significant
association between lab type and student response for each
of our student attitude questions.
Though Fisher’s test established a significant association
between lab type and student attitudes, it does not reveal
which lab types have a strong association. Given that each
of our contingency tables in this analysis was 3 × 3 (three
lab types and three possible outcomes), in the third phase of
our analysis, we performed a post hoc pairwise Fisher’s
exact test to compare each lab type. The results of the
pairwise comparisons are tabulated in Table VII:
TABLE IV. Variables used for our question-by-question analysis of the student attitude questions.
Variable Level of measurement Definition Value range
Lab type Nominal The type of lab a student participated in. AI; Sim; Lab
AQn Ordinal Review question n, where n ¼ 1, 2, 3, 4, or 5 1 ¼ Disagree; 2 ¼ Neutral; 3 ¼ Agree
TABLE V. Contingency table for AQ1.
Agree Disagree Neutral
AI 48 4 14
Lab 5 6 16
Sim 45 0 23
TABLE VI. Fisher’s test results for each AQ.
Question p value Cramer’s V
AQ1 3.581 × 10−7 0.323
AQ2 7.257 × 10−14 0.475
AQ3 0.005 0.215
AQ4 0.002 0.25
AQ5 8.969 × 10−7 0.323
LEVERAGING GENERATIVE ARTIFICIAL … PHYS. REV. PHYS. EDUC. RES. 22, 010109 (2026)
010109-7

Fisher’s pairwise exact tests indicate that students
who did the AI lab reported more positive attitudes
than students who did the traditional lab for all
AQs (padjAQ1;AI-Lab < 0.001, padjAQ2;AI-Lab < 0.001,
padjAQ3;AI-Lab ¼ 0.037, padjAQ4;AI-Lab ¼ 0.002,
padjAQ5;AI-Lab < 0.001). Similarly, students who did the
lab using prebuilt simulations also reported more posi-
tive attitudes than students who did the traditional
lab (padjAQ1;Sim-Lab < 0.001, padjAQ2;Sim-Lab < 0.001,
padjAQ3;Sim-Lab ¼ 0.002, padjAQ4;Sim-Lab ¼ 0.002,
padjAQ5;Sim-Lab < 0.001). Comparing between the two sim-
ulation groups, however (AI and prebuilt), there is a sig-
nificant difference only for AQ1 (padjAQ1;AI-Sim ¼ 0.035).
We can conclude that students who did the lab via prebuilt
simulations demonstrated more positive self-reported atti-
tudes on AQ1 than students who did the lab using AI.
IV. DISCUSSION
A. Developing conceptual understanding
The interactive use of a generative AI platform, where
students prompt, validate, and refine a simulation, can
support conceptual understanding of equipotential lines as
effectively as a prebuilt simulation and significantly better
than using physical equipment. From the results, we see a
significant and sizable impact on conceptual understanding
based on the approach taken (p < 0.001, η2 ¼ 0.359) among
the three experimental conditions. In pairwise comparisons,
we observe significant differences between the physical
equipment group and each of the other two approaches,
preexisting sim and AI-engaged laboratory. There was no
significant difference between the preexisting sim group and
the AI-engaged group. Furthermore, in a similar analysis (not
reported here), we found no significant differences in student
midterm performance or final exam performance across lab
sections or lab instructors (indicating a solid basis for
comparison). Hence, there is a strong indication that both
the use of the prebuilt simulation and the engaged use of
generative AI to develop, refine, and validate a simulation
were comparably productive in developing student under-
standing of concepts related to the laboratory.
In one sense, these results replicate the main results of
earlier studies showing that simulations can support greater
conceptual understanding of physics concepts in the
immediate aftermath of a learning activity [24]. In these
earlier studies, the learning differences were also found to
persist through the end of the term, which we notably did
not replicate here. As observed above, the student groups
did not perform differently on the midterm or final overall;
in fact, there was also no variation in performance between
groups, even considering only the midterm and final
questions focused on electric potential as a concept.
That is, the differences in conceptual performance among
the treatment groups that showed up following the lab
vanished on the subsequent midterm and final exam
questions. However, a significant difference in the broader
course context between this and the 2005 study offers a
plausible explanation of this difference: in the prior studies,
the simulation lab, which was studied, took place only after
the relevant topic had been fully discussed in lecture, and
was part of a course that did not employ modern interactive
teaching methods. In this current work, the lab we studied
was followed by continued discussion of the content in
lectures, homework, and exam reviews, including inter-
active and peer-instructional methods associated with
improved learning outcomes [25–27].
And yet, beyond the confirmation of the prior results for
the simulation groups, our comparable results for the AI-
engaged groups are striking. It is far from given that
students’ use of generative AI to build simulations would
support their conceptual development. For one thing,
building the simulation with the aid of the AI is an
additional layer of work on top of the actual use of the
simulation itself. For another, AI is prone to hallucination
or presenting material not-well matched to undergraduate
learners [28]. Nor are these learners particularly well
prepared in simulation development or the use of generative
AI, and we did not provide any instruction on these topics
prior to the lab activities documented here. However, we do
find that student-prompted, refined, and validated simu-
lation development did promote conceptual understanding
just as well as when they used a prebuilt, vetted, and
validated simulation, and to a greater extent than when they
used real equipment. In short, it appears that at least under
appropriate conditions, the opportunity to work with AI
tools can be added without cost to conceptual learning.
B. Student attitudes
Paralleling what was observed in the CQs, AQ res-
ponses indicated that students appreciated both the prebuilt
TABLE VII. Pairwise comparisons for affective questions. *p ¼ 0.05; **p ¼ 0.001; ***p < 0.001. Pairwise p values indicating
whether there was a significant difference between the indicated pair of treatment groups (rows) on a particular attitude question
(columns).
padjAQ1 padjAQ2 padjAQ3 padjAQ4 padjAQ5
AI vs Sim 0.035* 0.245 0.343 0.688 0.658
AI vs Lab 9.23 × 10−6*** 4.8 × 10−10*** 0.037* 0.002** 1.38 × 10−5***
Sim vs Lab 3.57 × 10−6*** 2.81 × 10−13*** 0.002** 0.002** 9.66 × 10−7***
YOSSI BEN-ZION et al. PHYS. REV. PHYS. EDUC. RES. 22, 010109 (2026)
010109-8

simulation and building their own simulation with an AI
more than working with physical equipment. On each of
the AQs, students using the digital technologies reported
more favorable responses than those using the physical lab
equipment. Of course, using the prebuilt simulation was
faster and easier than either of the other two conditions. So,
some students noted this affordance and appreciated being
done with the laboratory sooner. This may also account for
the one instance where students reported a preference for
the use of the prebuilt sim over the use of the AI generation
of a sim, on AQ1, comparing this approach to prior weeks.
In other cases, students reflected that the prebuilt
simulation and the design-based approach supported their
conceptual learning, whereas the equipment was less useful
to such ends. A student noted, “One of the reasons I like the
simulation compared to a hands-on experiment is that it
tends to be more difficult for me to carry out and understand
a hands-on experiment, whereas with the simulation, I feel
like more time is spent thinking through and understanding
the material.” In this sense, we agree with them. Working
with physical equipment may be better used to facilitate
experimental and modeling skills rather than for reinforcing
theory or concepts.
Finally, both the group building an AI-based simulation
and the group using physical equipment sometimes
reflected frustration in “getting it right” or “fixing the
equipment/AI”. While this can be problematic if students
are overly distracted or disengage as a result, it can also be a
positive side, capturing the productive frustration that
learning can entail. While students expressed frustrations
in developing their AI simulations, students’ frustrations
were similar to some of the frustrations of students working
with physical equipment—around design and manipulation
of the materials. None of the student concerns focused on a
limited opportunity to learn or engage with physics con-
cepts. This sentiment is in contrast to the group using
physical equipment, which expressed similar frustrations
with the equipment, with a clear indication that they
believed their learning was negatively impacted. Such
sentiments also highlight the capacity of an AI environment
to develop modeling and debugging skills with less
negative impact on the students’ conceptual understanding
of the material, though, of course, without the opportunity
to develop facility with troubleshooting physical equipment
as offered in a traditional lab activity.
C. An opportunity for modeling and experimental skills
Our study has focused on conceptual learning, which, to
many, is a valuable learning goal in the kind of lab- or
recitation-like settings that might make use of simulations
as an instructional tool. However, another perspective
suggests the focal learning outcomes in such settings could
be the development of scientific modeling practices, an
even more natural objective for laboratory instruction
[29–31]. While beyond the scope of this paper, we suspect
that the approach taken to prompting, refining, and vali-
dating AI-generated simulations can develop relevant
scientific and modeling skills that we seek to promote in
our laboratory environments, including among other things
a level of healthy scientific skepticism and an awareness
that data must be judged as a product of the models or
equipment that produced it. In particular, the iterative
approach invites the students to specifically consider the
nature of “simulation-as-model” as a part of the assign-
ment, which might instead simply be taken for granted
when the simulation is prebuilt.
In other words, rather than considering the traditional
laboratory with students using physical equipment as a
control condition, here we might consider the use of the
prebuilt simulation as a control condition, as both the
equipment group and the simulation creation group have
the opportunity to develop experimental skills. We are not
arguing that the same skills are practiced in each group,
though they may be analogous. For example, students
interacting with the prebuilt simulation and students inter-
acting with the AI-generated simulation both reported (in
free-response feedback) an appreciation for the ability to
visualize the intangible or unseen, such as equipotential
lines. However, the students using the prebuilt simulation
did not have the opportunity to validate (technically or
scientifically) the tool they were using. This validation may
parallel the modeling skills which Lewandowski et al. [30]
and others call for: technical validation as analogous to the
modeling of a physical apparatus, and scientific validation
as analogous to the physical modeling. By interacting with
the AI to build these representations, students building a
simulation with AI had more opportunities to learn skills
similar to the experiment group by developing and applying
models and validating them to aid their understanding.
Though the data presented in this work are not intended
to speak directly to the question, we speculate about
the mechanism by which iterative, validation phase of
AI-based lab work, which was not present for students
working with the prebuilt simulation, could contribute to
improved modeling skills in the AI-engaged groups. This
phase specifically invited students to view the simulations
generated from their prompts with skepticism. In this case,
the fact that AI is based on a probabilistic engine (and even
that it is subject to hallucinations) can serve as an
affordance. By constructing imperfect results, or perhaps
just the possibility of producing imperfect results, this
approach has the natural affordance that requires (or
encourages) students to engage in validation and verifica-
tion in ways that other mediums may not. In the prebuilt
simulation groups, students were essentially asked to treat
the simulations as directly reflecting the true physical
principles governing the underlying phenomena. In doing
so, the nature of the simulation as a “model” is obscured,
since the students do not necessarily consider any dis-
tinction between the simulation and the physical reality it is
meant to represent. A model, by its nature, should have
LEVERAGING GENERATIVE ARTIFICIAL … PHYS. REV. PHYS. EDUC. RES. 22, 010109 (2026)
010109-9

predictive capability, but also has inherent limitations in its
sphere of applicability [32]; when students use preexisting
simulations in coursework, they may overlook these
aspects of the simulation-as-model if the simulation accu-
racy is simply taken for granted. By contrast, students in the
AI-engaged lab activity were asked to treat each iteration of
simulation output by the AI as a possible model of the real
world, but one whose fidelity remained to be established,
even in the simplest contexts, before it could be responsibly
used to explore more complicated scenarios.
D. Limitations
Several methodological and contextual factors constrain
the interpretation and generalizability of these findings.
First, while all three instructional approaches addressed
electric potential concepts through active engagement,
systematic differences existed in the level of abstraction
required. In the physical laboratory approach, for example,
the “basic” objects presented to the students for use and
consideration were the source charges and the conductive
bath around them, in which potential differences could
be measured, and from there, equipotential shapes inferred.
In the simulations, the equipotentials themselves were
presented to students as part of the “basic” objects available
for their contemplation—made instantly at the push of a
button, just like the source charges that create them.
Although each approach incorporated both conceptual
reasoning and quantitative analysis, this fundamental dif-
ference in abstraction may have systematically influenced
performance on the conceptual assessment instrument
used in this study. The utility of visualizations may be
considered a limitation on the one hand, or a finding on
the other.
Second, technical limitations constrained the AI-
engaged group’s implementation fidelity. Students encoun-
tered restrictions imposed by the free-tier service limits of
the AI platform, including constraints on prompt frequency
and session duration. Additionally, some participants
required instructional assistance to resolve technical issues
during simulation development or equipment use, poten-
tially affecting the consistency of the intervention delivery
across participants.
Third, this investigation examined a single physics
concept (electric potential) within a homogeneous student
population at a single research institution. Participants
comprised primarily life sciences majors with similar
academic backgrounds and career trajectories. The extent
to which these findings generalize to other physics topics,
more diverse student populations, or different institutional
contexts remains an empirical question requiring system-
atic investigation across varied educational settings.
V. CONCLUSIONS
The use of generative AI to design, refine, and validate a
simulation can support student conceptual mastery of key
physics topics in a laboratory environment. Students
who developed simulations through AI-guided processes
achieved conceptual understanding equivalent to those
using prebuilt simulations, with both digital approaches
significantly outperforming traditional physical labora-
tory methods. Notably, despite the additional cognitive
demands of prompting, validating, and refining AI-
generated simulations, students’ conceptual learning was
not compromised.
These preliminary findings indicate three complemen-
tary pathways for physics laboratory education, which by
its nature addresses a wide variety of learning outcomes
beyond the basic development of theoretical knowledge.
Prebuilt simulations effectively engage students in learning
key concepts while providing immediate access to complex
visualizations. AI-guided simulation development supports
conceptual learning, awareness of emerging technologies,
and opportunities to model systems. Traditional physical
equipment remains essential for developing experimental
skills with physical equipment, modeling, and connecting
students to the material world that physics seeks to
understand.
Given that generative artificial intelligence will likely
constitute a significant component of students’ future
professional and academic lives, developing competencies
in productive AI use represents an important educational
objective. Our results demonstrate that this can be achieved
without compromising traditional physics learning goals.
Physics fundamentally concerns the study of the material
world, and the use of experimental equipment remains
essential for developing authentic scientific practices.
Generative AI now provides a complementary tool for
use alongside physical equipment and prebuilt simulations.
The chosen mixture of these approaches should align with
context-specific learning objectives, recognizing that differ-
ent methods foster distinct forms of active learning and skill
development.
Future investigations should examine the pedagogical
framing required for optimal implementation of these tools.
Neither excessive optimism nor cynicism is warranted; these
tools require careful implementation with attention to their
limitations but can demonstrably serve student learning and
engagement when thoughtfully deployed.
As physics education evolves to meet the demands of an
increasingly digital world, the integration of AI-assisted
learning alongside traditional experimental work offers a
promising framework. This complementary approach may
help prepare students with the diverse competencies needed
for contemporary scientific practices.
DATA AVAILABILITY
The data that support the findings of this article are not
publicly available because they contain sensitive personal
information. The data are available from the authors upon
reasonable request.
YOSSI BEN-ZION et al. PHYS. REV. PHYS. EDUC. RES. 22, 010109 (2026)
010109-10

APPENDIX A: SURVEY OF CONCEPTUAL
UNDERSTANDING
Circle your answers to the following questions to check
your understanding. In all questions, use the convention
that the potential is zero infinitely far from the charges.
1. A large positive charge sits alone in the center of my
lab (no other charges are present). At a certain point
“P,” which is 2 m away from the positive charge, I
measure the potential to be 40 V. What is the
potential at point “Q” which is one meter away
from the positive charge?
A. 10 V.
B. 20 V.
C. 40 V.
D. 80 V.
E. 160 V.
2. Four different (nonzero) electric charges are placed
in a square. At the exact center of the square, the
potential is exactly zero. What can we say about the
signs of the charges? Select one.
A. All of the charges must be positive.
B. At least one of the charges must be negative.
C. Exactly two of the changes must be negative.
D. All of the charges must be negative.
E. None of the above is necessarily true.
3. In which of the following cases would the equipo-
tential lines around the charges form perfect circles?
Assume there are no other charges present besides
the ones described. Select one
A. A single positive charge is located at the
origin.
B. A single negative charge is located at the origin.
C. A single positive charge is located at (x ¼ 3 m,
y ¼ 3 m).
D. Both (A) and (B) but not (C).
E. (A), (B), and also (C).
4. What happens to the shape of equipotential lines
when switching the sign of every charge
(positive ⇄ negative) in a system?
A. The equipotential lines flip and completely
change their shape.
B. The lines remain the exact same shape, but their
values switch from positive to negative (or
vice versa).
C. The equipotential lines for a negative charge are
more closely spaced than those for a positive
charge.
D. There is no change in equipotential lines.
E. We can’t tell what happens.
F. None of the above.
5. Consider the two positive charges depicted below,
along with the equipotential line shown around them.
What must be true about the magnitude of the
charges?
A. Q1 ¼ Q2.
B. Q1 > Q2.
C. Q1 < Q2.
D. Cannot determine without more information.
APPENDIX B: TESTING ANOVA ASSUMPTIONS
The assumptions for a one-way ANOVA test are:
1. Our dependent variable is continuous: this is true for
our variables.
2. Our independent variable consists of two or more
categorical, independent groups: this is true for
our data.
3. Independence of observations: This is true to the best
of our knowledge.
4. No significant outliers in the data:
5. The residuals of each dependent variable are approx-
imately normally distributed for each category.
6. Homogeneity of variances across categories.
Assumptions 1–3 are matters of design. Assumptions 1
and 2 follow directly from our study design. Assumption 3
TABLE VIII. Descriptive statistics for each lab type.
Lab type CQTot MQTot FQTot
AI (N ¼ 66) Median 80 80 76.6
Mean 83.64 81.82 74.86
Standard deviation 14.43 12.998 14.10
Skewness −0.543 −0.441 −0.420
Kurtosis 0.008 −0.198 −0.573
Sim (N ¼ 68) Median 80 85 78.1
Mean 85.59 81.69 77.11
Standard deviation 13.76 13.43 14.63
Skewness −0.428 −0.883 −1.390
Kurtosis −0.815 0.753 3.492
Lab (N ¼ 27) Median 60 80 78.1
Mean 48.89 78.15 68.98
Standard deviation 20.25 17.05 19.42
Skewness −1.517 −0.430 −0.585
Kurtosis 1.907 −0.786 −1.106
LEVERAGING GENERATIVE ARTIFICIAL … PHYS. REV. PHYS. EDUC. RES. 22, 010109 (2026)
010109-11

is approximately true to the best of our knowledge.
Graphical inspection of our data showed that there were
8 outliers in our dataset across all groups. This represents
∼5% of our data, and it was determined that these outliers
did not cause a significant distortion of our results, thus
assumption 4 held for our data. In order to test assumption
5, we calculated the skewness and kurtosis for our
dependent variables across each category. These results
are documented in Table VIII:
The skewness and kurtosis for each dependent variable
are between −2 and 2, with one exception [the kurtosis for
the FQTot for the Sim (prebuilt simulation) group]. Upon
seeing this inconsistency, we statistically tested for normal-
ity for each dependent variable using the Shapiro-Wilk
(SW) test [33,34]. For the SW test, the null hypothesis is
that a sample is drawn from a normally distributed
population, so if p < 0.05, we conclude that our sample
does not satisfy the assumption of normality. The SW test
indicated that none of these variables follow a normal
distribution for our sample (WCQ ¼ 0.87, pSW-CQ < 0.001;
WMQ ¼ 0.96, pSW-MQ < 0.001; WFQ ¼ 0.95, pSW-FQ <
0.001). In order to test for homogeneity of variances
(assumption 6), we performed Levene’s test of homo-
geneity of variances for our dependent variables. The
results are shown in Table IX:
CQTot, MQTot, and FQTot all show homogeneity of vari-
ances according to Levene’s test (pLev-CQTotð2158Þ ¼ 0.117,
pLev-MQTotð2158Þ ¼ 0.169, and pLev-FQTotð2158Þ ¼ 0.008).
APPENDIX C: QUESTION-BY-QUESTION
ANALYSIS OF CQs
We performed analysis of both the CQs and the affective
questions on a question-by-question basis using a three-
phase method outlined below. In this appendix, we dem-
onstrate the technique in the context of the CQs; equivalent
data for the affective questions is available on request.
Since our analysis of the CQTot indicated significant
differences in performance, among the three lab types,
on the lab review, we decided to analyze each review
question individually to examine whether there were
differences in the number of correct responses across
each lab type. The variables for this analysis are listed in
Table X:
For each question, there were cells in our contingency
table that had a frequency of less than 5, so in order to
perform such an analysis. We used the same three-phase
analysis that was used for the question-by-question analysis
of the student attitude questions detailed in the Data
Analysis and Results section above. This analysis was
done in three stages. First, contingency tables were devel-
oped depicting the frequency of correct and incorrect
responses for each CQ within each lab type to inform us
about how many students in each lab type got the question
correct/incorrect. For example, the contingency table for
CQ1 is shown in Table XI (other contingency tables are
available upon request).
The results of Fisher’s test for each CQ are in Table XII:
As the p value (two-tailed) obtained from Fisher’s exact
test is significant for each CQ (see p values and effect sizes
in Table XII), we reject the null hypothesis (p < 0.05) and
conclude that there is a statistically significant association
between lab type and student performance for each of our
review questions.
The results of the pairwise comparisons are tabulated in
Table XIII:
Results of Fisher’s pairwise exact tests indicate that there
is a significant association between the AI and Sim lab
groups for CQ4 (padjCQ4;AI-Sim ¼ 0.025). We can conclude
that students who did the lab via prebuilt simulations were
more likely to answer CQ4 correctly than students who did
TABLE IX. Levene’s tests results.
F value d:o:f:1 d:o:f:2 p value
CQTot 0.335 2 158 0.717
MQTot 1.686 2 158 0.189
FQTot 2.45 2 158 0.09
TABLE X. Variables used for question-by-question analysis of conceptual questions.
Variable Level of measurement Definition Value range
Lab type Nominal The type of lab a student participated in. AI; Sim; Lab
CQn Nominal Conceptual question n, where n ¼ 1, 2, 3, 4, or 5 0 ¼ Incorrect; 1 ¼ Correct
TABLE XI. Contingency table for CQ1.
Correct Incorrect
AI 63 3
Lab 5 22
Sim 66 2
TABLE XII. Fisher’s test results for each CQ.
Question p value Cramer’s V
CQ1 2.2 × 10−16 0.778
CQ2 0.002 0.268
CQ3 0.037 0.2
CQ4 0.009 0.22
CQ5 3.7 × 10−7 0.481
YOSSI BEN-ZION et al. PHYS. REV. PHYS. EDUC. RES. 22, 010109 (2026)
010109-12

the lab using AI. Students who did the AI lab were
more likely to answer CQ1, CQ2, and CQ5 correctly than
students who did the traditional lab (padjCQ1;AI-Lab < 0.001,
padjCQ2;AI-Lab ¼ 0.003, padjCQ5;AI-Lab ¼ 0.005). Students
who did the lab using prebuilt simulations were more
likely to answer CQ1, CQ2, CQ4, and CQ5 correctly
than students who did the traditional lab (padjCQ1;Sim-Lab <
0.001, padjCQ2;Sim-Lab ¼ 0.003, padjCQ4;Sim-Lab ¼ 0.033,
padjCQ5;Sim-Lab < 0.001).
[1] G. Kortemeyer, Could an artificial-intelligence agent pass
an introductory physics course?, Phys. Rev. Phys. Educ.
Res. 19, 010132 (2023).
[2] C. G. West, AI and the FCI: Can ChatGPT project an
understanding of introductory physics?, arXiv:2303
.01067.
[3] G. Polverini, J. Melin, E. Onerud, and B. Gregorcic,
Performance of ChatGPT on tasks involving physics visual
representations: The case of the brief electricity and
magnetism assessment, Phys. Rev. Phys. Educ. Res. 21,
010154 (2025).
[4] G. Kestin, K. Miller, A. Klales, T. Milbourne, and G.
Ponti, AI tutoring outperforms in-class active learning:
An RCT introducing a novel research-based design in
an authentic educational setting, Sci. Rep. 15, 17458
(2025).
[5] Y. Ben-Zion, R. E. Zarzecki, J. Glazer, and N. D.
Finkelstein, Leveraging AI for rapid generation of physics
simulations in education: Building your own virtual lab,
Phys. Teach. 63, 424 (2025).
[6] K. T. Kotsis, ChatGPT as teacher assistant for physics
teaching, EIKI J. Eff. Teach. Methods 2, 4 (2024).
[7] F. Mahligawati, E. Allanas, M. H. Butarbutar, and N. A. N.
Nordin, Artificial intelligence in physics education: A
comprehensive literature review, J. Phys. Conf. Ser.
2596, 012080 (2023).
[8] E. E. Ukoh and J. Nicholas, AI adoption for teaching and
learning of physics, Int. J. Infonomics 15, 2121 (2022).
[9] J. Dewey, Experience and Education (Macmillan, New
York, NY, 1938).
[10] N. D. Finkelstein et al., When learning about the real world
is better done virtually: A study of substituting computer
simulations for laboratory equipment, Phys. Rev. ST Phys.
Educ. Res. 1, 1 (2005).
[11] N. M. Seel, Experimental and quasi-experimental de-
signs for research on learning, in Encyclopedia of the
Sciences of Learning (Springer, Boston, MA, 2012),
pp. 1223–1229.
[12] M. Chhabra and R. Das, Conceptualization of electrostatic
potential: Resource theory perspective, Phys. Educ. 60,
055019 (2025).
[13] J. P. Burde and T. Wilhelm, Teaching electric circuits with a
focus on potential differences, Phys. Rev. Phys. Educ. Res.
16, 020153 (2020).
[14] D. Psillos, A. Tiberghien, and P. Koumaras, Voltage
presented as a primary concept in an introductory teaching
sequence on dc circuits, Int. J. Sci. Educ. 10, 29 (1988).
[15] R. Millar and K. L. Beh, Students’ understanding of
voltage in simple parallel electric circuits, Int. J. Sci. Educ.
15, 351 (1993).
[16] E. Ostertagová, O. Ostertag, and J. Kováč, Methodology
and application of the Kruskal-Wallis test, Appl. Mech.
Mat. 611, 115 (2014).
[17] B. Efron and R. Tibshirani, Bootstrap methods for standard
errors, confidence intervals, and other measures of stat-
istical accuracy, Stat. Sci. 1, 1 (1986).
[18] J. S. Haukoos and R. J. Lewis, Advanced statistics: boot-
strapping confidence intervals for statistics with “difficult”
distributions, Acad. Emerg. Med. 12, 360 (2005).
[19] W. H. Kruskal and W. A. Wallis, Use of ranks in one-
criterion variance analysis, J. Am. Stat. Assoc. 47, 583
(1952).
[20] J. Cohen, Statistical Power Analysis for the Behavioral
Sciences, 2nd ed. (Lawrence Erlbaum Associates, Hillsdale,
NJ, 1988).
[21] O. J. Dunn, Multiple comparisons using rank sums, Tech-
nometrics 6, 241 (1964).
[22] B. Van Dusen and J. M. Nissen, Criteria for collapsing
rating scale responses: A case study of the CLASS,
presented at PER Conf. 2019, Provo, UT 2019,
10.1119/perc.2019.pr.Van_Dusen.
[23] A. Agresti, An Introduction to Categorical Data Analysis,
2nd ed. (John Wiley & Sons, Hoboken, NJ, 2007).
[24] N. Finkelstein, Learning physics in context: A study of
student learning about electricity and magnetism, Int. J.
Sci. Educ. 27, 1187 (2005).
TABLE XIII. Pairwise comparisons for conceptual questions. *p ¼ 0.05; **p ¼ 0.001; ***p < 0.001.
padjCQ1 padjCQ2 padjCQ3 padjCQ4 padjCQ5
AI vs Sim 0.68 0.86 0.37 0.025* 1.00
AI vs Lab 1.86 × 10−13*** 0.003** 0.07 1.00 4.65 × 10−6***
Sim vs Lab 2.84 × 10−13*** 0.003** 0.33 0.033* 5.02 × 10−6***
LEVERAGING GENERATIVE ARTIFICIAL … PHYS. REV. PHYS. EDUC. RES. 22, 010109 (2026)
010109-13

[25] C. H. Crouch and E. Mazur, Peer instruction: Ten years of
experience and results, Am. J. Phys. 69, 970 (2001).
[26] S. Freeman, S. L. Eddy, McM. Donough, M. K. Smith, N.
Okoroafor, H. Jordt, and M. P. Wenderoth, Active learn-
ing increases student performance in science, engineer-
ing, and mathematics, Proc. Natl. Acad. Sci. U.S.A. 111,
8410 (2014).
[27] C. Turpen and N. D. Finkelstein, Not all interactive
engagement is the same: Variations in physics professors’
implementation of peer instruction, Phys. Rev. ST Phys.
Educ. Res. 5, 020101 (2009).
[28] H. Elsayed, The impact of hallucinated information in
large language models on student learning outcomes: A
critical examination of misinformation risks in AI-assisted
education, North. Rev. Algorithm Res. Theor. Comput.
Complexity 9, 8 (2024), https://northernreviews.com/index
.php/NRATCC/article/view/2024-08-07.
[29] N. G. Holmes and C. E. Wieman, Introductory physics
labs: We can do better, Phys. Today 71, No. 1, 38 (2018).
[30] J. Kozminski et al., AAPT Recommendations for the
Undergraduate Physics Laboratory Curriculum (American
Association of Physics Teachers, College Park, MD, 2014).
[31] B. R. Wilcox and H. J. Lewandowski, Students’ views
about the nature of experimental physics, Phys. Rev. Phys.
Educ. Res. 13, 020110 (2017).
[32] E. Etkina, A. Warren, and M. Gentile, The role of models in
physics instruction, Phys. Teach. 44, 34 (2006).
[33] S. Keskin, Comparison of several univariate normality tests
regarding type I error rate and power of the test in simu-
lation based small samples, J. Appl. Sci. Res. 2, 5
(2006), https://www.academia.edu/83340328/Comparison_
of_Several_Univariate_Normality_Tests_Regarding_Type_I_
Error_Rate_and_Power_of_the_Test_in_Simulation_based_
Small_Samples.
[34] S. S. Shapiro and M. B. Wilk, An analysis of variance test
for normality (complete samples), Biometrika 52, 591
(1965).
YOSSI BEN-ZION et al. PHYS. REV. PHYS. EDUC. RES. 22, 010109 (2026)
010109-14