NeuroChat: AI for Personalized Learning
NeuroChat: AI for Personalized Learning
Learning Experiences
Dünya Baradari Nataliya Kosmyna Oscar Petrov
MIT Media Lab MIT Media Lab Brown University
Cambridge, MA, United States Cambridge, MA, United States Providence, RI, United States
Figure 1: Overview of the NeuroChat neuroadaptive LLM system. A wearable dry-electrode EEG headband collects data from
the brain and sends it to the NeuroChat web app, which computes the user’s level of engagement. The engagement score is sent
with each request to the LLM, allowing it to adapt its response style to the user in real-time.
ABSTRACT pacing using a closed-loop system. We evaluate this approach in a
Generative AI is transforming education by enabling personalized, pilot study (n=24), comparing NeuroChat to a standard LLM-based
on-demand learning experiences. However, AI tutors lack the abil- chatbot. Results indicate that NeuroChat enhances cognitive and
ity to assess a learner’s cognitive state in real time, limiting their subjective engagement but does not show an immediate effect on
adaptability. Meanwhile, electroencephalography (EEG)-based neu- learning outcomes. These findings demonstrate the feasibility of
roadaptive systems have successfully enhanced engagement by real-time cognitive feedback in LLMs, highlighting new directions
dynamically adjusting learning content. This paper presents Neu- for adaptive learning, AI tutoring, and human-AI interaction.
roChat, a proof-of-concept neuroadaptive AI tutor that integrates
real-time EEG-based engagement tracking with generative AI. Neu- KEYWORDS
roChat continuously monitors a learner’s cognitive engagement Brain-computer interface, electroencephalography (EEG), chatbot,
and dynamically adjusts content complexity, response style, and conversational AI, closed-loop
Baradari, Kosmyna, Petrov, Kaplun, and Maes
the perspective of flow theory, learner engagement can be enhanced markers include the Cognitive Load Index (Theta Fz / Alpha Pz)
by designing learning activities that promote autonomy and provide [12, 37] and alpha peak frequency [62], which have both been ex-
appropriate challenges to learners’ skill level [75]. Sinatra et al. [76] plored as indicators of cognitive effort and attentional processing
further distinguish between microlevel engagement, which refers efficiency. ERPs, in contrast, offer time-locked neural responses to
to moment-to-moment cognitive focus on a task, and macro-level stimuli, with key components such as P300 (reflecting attentional al-
engagement, which applies to larger social and educational con- location and task relevance), N200 (linked to conflict detection and
texts, such as classrooms or institutions. Micro-level engagement executive control), and error-related negativity (ERN) (indicating
can be assessed through physiological techniques such as brain engagement in performance monitoring) [8, 63, 85].
imaging, skin conductivity, or eye tracking, whereas macrolevel
engagement is typically measured through sociocultural analysis, 2.1.3 EEG-Based Engagement in Learning. EEG-based engagement
observations, or ratings. metrics have been applied in various educational contexts, from
In cognitive neuroscience, engagement is closely linked to sus- providing feedback to presenters [34] to tracking cognitive effort
tained attention and tonic alertness, reflecting a person’s sustained in classroom and workplace environments [30, 33]. EEG has been
cognitive effort [57]. However, engagement extends beyond atten- shown to capture distinct patterns of student attention that dif-
tion, incorporating factors such as intrinsic motivation and task fer from self-reports and teacher observations, offering a more
involvement [44, 68]. Unlike cognitive load, which reflects the men- objective measure of engagement across instructional activities
tal demands on working memory, engagement captures both effort [30]. Studies have also linked higher engagement to better learning
and motivation in a task-driven context. For this study, we define performance. For example, EEG monitoring during video lectures re-
engagement as the sustained allocation of cognitive and attentional vealed significant fluctuations in attentional focus, suggesting that
resources toward a task, influenced by motivation and mental effort. lecture design should account for these variations [18]. Similarly, in
a reasoning task with medical students, engagement correlated with
2.1.2 Physiological Measures of Engagement. Various physiologi- task performance, though the highest engagement was observed
cal technologies have been explored to assess engagement in digital in students who struggled, likely reflecting heightened cognitive
learning environments, including video analysis, eye tracking, and effort despite insufficient expertise. This aligns with Vygotsky’s
biosensors. Classroom video analysis has been used to monitor Zone of Proximal Development, suggesting that while students
student attention, with Raca and Dillenbourg [65] utilizing video were highly engaged, the task was beyond their current skill level,
recordings and later incorporating a computer vision model to leading to cognitive overload rather than effective learning [47].
approximate eye gaze [66]. However, these models have limited EEG has also been explored for cognitive workload classification,
accuracy in estimating attention levels, as they attempt to infer com- with implications for future learning technologies. Andreessen et al.
plex cognitive states from external behavioral approximations that [4] trained an EEG-based model to distinguish high and low mental
are often ambiguous and context-dependent. Eye-tracking systems workload, suggesting its potential for adaptive learning systems
provide a more granular measure of attention shifts and mind- that adjust reading materials based on cognitive load. Similarly, Api-
wandering, but are often costly, complex, and prone to calibration cella et al. [5] demonstrated a low-channel, wearable EEG system
and accuracy issues [40, 41]. for detecting engagement, proposing its use as an input channel for
More direct physiological measures include heart-rate variability adaptive teaching platforms. While these studies focus on monitor-
(HRV) [12], skin conductance (EDA) [11], and electroencephalogra- ing engagement rather than adapting learning in real time, they lay
phy (EEG) [51, 64, 87]. Among these, EEG stands out as the only the groundwork for neuroadaptive systems that dynamically adjust
method that directly measures neural activity, providing real-time instruction based on cognitive states. The next section explores
insights into alertness, attention, and cognitive workload in both how such systems leverage EEG engagement data to personalize
controlled and real-world settings [10, 28]. Since learning is funda- learning experiences.
mentally a neurological process, EEG offers a unique advantage by
capturing dynamic brain responses during information processing, 2.2 Neuroadaptive Learning Systems
making it particularly well-suited for assessing engagement beyond Neuroadaptive systems leverage real-time neurophysiological data,
behavioral proxies. particularly from electroencephalography (EEG), to dynamically
Engagement can be measured using EEG through oscillatory adjust instructional content or interaction modalities based on a
activity (frequency-based markers) and event-related potentials learner’s cognitive and emotional states. These closed-loop systems
(ERPs). Frequency-based markers provide continuous insights into aim to optimize learning outcomes by continuously monitoring
attention and cognitive workload, with alpha power (8–12 Hz) engagement and adapting pedagogical strategies accordingly.
linked to relaxation and disengagement [30, 31], beta power (13–30 Early approaches to neuroadaptive learning focused on adapt-
Hz) associated with sustained attention and active problem-solving ing presentation styles based on user engagement. For instance,
[64], and theta power (4–8 Hz) indicative of fatigue or reduced Pay Attention! [79] employed an embodied storytelling agent that
vigilance [24]. A widely used composite metric is the Engage- adjusted its voice volume and gestures in real time to recapture
ment Index, defined as Beta / (Alpha + Theta), where higher val- students’ attention when EEG signals indicated a drop in engage-
ues indicate greater attentional focus and cognitive engagement ment. This approach significantly enhanced the recall performance
[5, 22, 34, 47, 51, 64]. Alpha asymmetry reflects differences in alpha of students, demonstrating the potential of adaptive presentation to
power between the two brain hemispheres and is often associated influence learning outcomes. Similarly, EngageMeter [34] provided
with approach motivation and active engagement [24, 83]. Other real-time feedback to keynote presenters about their audience’s
Baradari, Kosmyna, Petrov, Kaplun, and Maes
engagement levels, enabling dynamic adjustments in delivery style. dynamic content modulation and interactive adaptation. Early in-
However, while effective in maintaining attention, these systems vestigations propose that integrating LLMs with BCIs could sig-
were limited to modifying delivery methods without altering the nificantly enhance human-computer interaction, benefiting both
learning content itself. individuals with neurological conditions and healthy users [13].
Thinking Cap [55] extends these ideas into an Intelligent Tu-
2.3.1 Using Generative AI to Analyze EEG. A major focus in AI-BCI
toring System (ITS) featuring an animated tutor agent that dy-
research has been EEG-based brain decoding, where generative AI
namically adjusts the complexity of its instructional dialogue with
and machine learning models are used to encode and decode the
the student based on EEG-derived cognitive load measures. This
neural signals underlying visual or auditory information processing
approach ensures that learners are neither underwhelmed nor over-
[7, 32, 84]. While these methods advance neural signal processing,
whelmed. The authors pre-scripted easy and difficult versions of
they remain limited in real-time user interaction. Readers inter-
the instructional content by altering text complexity dimensions
ested in these approaches can refer to a comprehensive review by
such as narrativity, syntactic ease, and referential cohesion. The
Sabharwal and Rama (2024) [70]. Beyond decoding, LLMs have
more recent Online Continuous Adaptation Mechanism (OCAM)
been increasingly applied to EEG for brain state classification and
[22] builds on these principles by continuously monitoring not just
assistive communication [88]. In clinical applications, language
engagement but also concentration, cognitive load, and emotional
model-enhanced BCI communication systems have significantly
arousal to dynamically adjust content difficulty, pacing, and presen-
improved typing accuracy for ALS patients by up to 84% in online
tation style. This system has been shown to significantly increase
BCI spelling sessions [77]. Subsequent approaches have demon-
learner concentration and engagement, highlighting the value of
strated that LLMs can classify brain states at the word level from
multi-dimensional cognitive measures in adaptive learning.
EEG data during reading tasks [36, 89].
Beyond academic learning environments, Learning Piano with
Recent research has extended these applications to foundation
BACh [87] dynamically adapts the difficulty of piano exercises
models that generalize across EEG tasks. NeuroLM [43] and Neuro-
based on cognitive workload (measured via functional near-infrared
GPT [15] function as foundation models, pre-trained on large EEG
spectroscopy), guiding learners into their zone of proximal develop-
datasets using self-supervised learning and task-based fine-tuning
ment to optimize skill acquisition. Closed-loop systems have also
to develop multi-purpose EEG processing models. NeuroLM, trained
been found effective for enhancing learning in perceptual-cognitive
on over 25,000 hours of EEG recordings, aligns brain signals with
tasks, as demonstrated by Parsons et al. [62], who improved perfor-
text-based representations, enabling multi-task analysis in areas
mance by manipulating a 3D multiple object tracking (3D-MOT)
like cognitive workload detection, emotion recognition, and sleep
task through real-time neurofeedback.
staging. Neuro-GPT, trained on the TUH EEG corpus, applies GPT-
Other systems focused on providing real-time biofeedback to
style tokenization to EEG data, improving feature extraction and
help users self-regulate their engagement and attention. AttentivU
adaptability to small datasets. In contrast, EEG-GPT [48] and Lee
[51], for instance, combines EEG headband with haptic feedback
& Chung (2024) [52] focus on task-specific applications—EEG-GPT
devices that vibrate subtly when engagement levels drop, effec-
applies few-shot learning for EEG-based brain state classification,
tively redirecting attention in both online and in-person learning
while Lee & Chung fine-tune GPT-3.5 Turbo for intracranial EEG
contexts. Unlike content-adaptive systems, these approaches rely
(iEEG) interpretation, mapping neural signals to cognitive states.
on external cues to prompt re-engagement rather than altering the
Other approaches have explored personal health and well-being.
learning material itself. Similarly, Joie [83] introduces a joy-based
For instance, Sano et al. (2024) [71] used LLMs to interpret EEG
brain-computer interface (BCI) that uses prefrontal alpha asymme-
signals for sleep quality assessment, providing tailored recommen-
try—an EEG marker linked to positive emotional states—to control
dations. Similarly, EEG Emotion Copilot [14] integrates EEG with a
an endless runner game. By training users to consciously modulate
lightweight (0.5B parameter) LLM to analyze EEG signals, identify
their brain activity through strategies like imagining joyful scenar-
emotional states, and generate automated clinical insights. [38]
ios, Joie highlights the potential of neuroadaptive systems to foster
propose MultiEEG-GPT, a model that integrates EEG with mul-
affective engagement alongside cognitive performance.
timodal data—such as facial expressions and audio—to enhance
Across these systems, a shared limitation is that all content-
mental health assessments using LLM-based classification.
driven systems rely on pre-scripted content that needs to be pre-
Additionally, generative AI techniques have been leveraged for
pared by the researchers to allow for the adaptation. NeuroChat
data augmentation to enhance EEG-based model training [21, 90].
overcomes this barrier by integrating generative AI, which can
However, while these approaches highlight generative AI’s ability
create new content adapted in complexity and presentation style
to process EEG data for individual adaptation, they focus on recog-
to the reader’s cognitive state and specific questions on the fly.
nizing states and have yet to support real-time user interaction.
2.3.2 Artistic Applications Using EEG to Modulate Generative AI
2.3 Generative AI-BCI Systems Outputs. While most research has focused on analyzing EEG data,
The integration of generative AI with brain-computer interfaces a growing field explores EEG as a control mechanism for real-
(BCIs) is an emerging research area. While machine learning has time generative AI adaptation. Early explorations have emerged
long been used to analyze EEG data, most AI-enhanced BCI systems in artistic and creative applications, where EEG signals influence
have focused on brain state classification rather than interactive, AI-generated media production. For example, Imagination Engine
real-time adaptive applications. The introduction of generative AI [1] translates EEG activity into abstract visual art, while Real-Time
expands the possibilities of BCIs beyond passive decoding, enabling Neuro-Augmented Cinema [9] enables cinematic modifications
NeuroChat: A Neuroadaptive AI Chatbot for Customizing Learning Experiences
Figure 2: Left. The Muse 2 EEG system made by InteraXon Inc. Right. Electrode locations of Muse 2 headband according to
10-20 System. CC Teixeria, Gomes, and Brito-Costa (2023).
via neurofeedback. Similarly, the Bio-Mechanical Poet [82] maps data in real time to adjust content complexity, response
real-time EEG signals to symbolic representations, creating immer- style, and pacing, ensuring an adaptive and personalized
sive poetic audiovisual experiences. These projects demonstrate learning experience.
EEG’s potential to actively modulate generative outputs rather than (3) D3: Seamless EEG Integration in Natural User Envi-
merely classifying brain states. However, these applications remain ronments: Given that most users interact with NeuroChat
limited to artistic expression, with little research on EEG-driven on a laptop or desktop computer, the system accommodates
adaptation of linguistic content. The potential to use real-time neu- a stationary, movement-minimizing environment, which
rofeedback to shape AI-generated textual interactions—particularly is ideal for EEG signal acquisition. This design choice min-
in education—remains largely unexplored. imizes motion artifacts, resulting in more reliable neuro-
feedback processing.
2.3.3 Neuroadaptive Generative AI Systems. The most advanced (4) D4: Web-Based Accessibility and Low-Cost Implemen-
neuroadaptive system integrating generative AI for real-time adap- tation: NeuroChat is designed to be fully browser-based,
tation is AdaptiveCoPilot [86], designed for expert pilots in virtual eliminating the need for complex server-side infrastruc-
reality. AdaptiveCoPilot continuously adjusts visual, auditory, and ture and enabling plug-and-play usability. Users can access
textual cues based on real-time cognitive load assessments, optimiz- the system on any platform with minimal setup, making it
ing performance in high-stakes environments. However, it relies scalable and accessible for a broad audience.
on functional near-infrared spectroscopy (fNIRS) rather than EEG
and is tailored for high-performance cognitive tasks rather than
3.1 Interaction Flow
learning applications.
Despite rapid advancements in AI-enhanced BCI research, no NeuroChat integrates real-time EEG data with an LLM chatbot via
existing system has leveraged EEG data to dynamically modulate a web interface to create adaptive, personalized responses. The user
LLM-driven chatbot interactions. This gap underscores the nov- flow is as follows (Figure 3):
elty of NeuroChat, one of the first systems to integrate EEG-based (1) Connection: The user fits on the Muse 2 EEG headband
cognitive state tracking with generative AI in real-time. and connects it to the NeuroChat web app via Web Blue-
tooth for real-time data streaming.
3 SYSTEM DESIGN (2) Calibration: The user completes a 2-minute relaxation task
We set out four core design goals to ensure that NeuroChat is an to determine the engagement minimum (E_min) and a 2-
accessible, responsive, and effective neuroadaptive learning system: minute word association task for the engagement maximum
(E_max). These values are stored in the browser’s session
(1) D1: Wearable, Non-Invasive Brain Sensing: NeuroChat storage for normalization.
employs a consumer-grade, non-invasive EEG headband (3) Interaction: During interaction, the system continuously
to measure engagement in real time. We opted for the 4- computes the normalized engagement score using a 15-
channel Muse EEG headband by InteraXon [60], balancing second sliding window. The last score before the user begins
signal reliability with ease of use. This design ensures that typing is captured and embedded in the query to the chatbot,
users can engage with the system without complex elec- hidden from the user, ensuring that typing doesn’t interfere
trode setups or invasive procedures. with the engagement metric.
(2) D2: Real-Time Adaptive Personalization: To maximize (4) Interactive Response: The query, along with the embed-
learning effectiveness, NeuroChat provides continuous neu- ded engagement score, is sent to the LLM provider, which
rofeedback, dynamically tailoring chatbot responses based returns a response tailored to the user’s cognitive state.
on real-time engagement levels. The system processes EEG
Baradari, Kosmyna, Petrov, Kaplun, and Maes
Figure 3: Overview of the NeuroChat system and user flow. The user connects the Muse headband, undergoes calibration, and
interacts with the neurofeedback-driven LLM. Engagement scores are computed and inserted into the prompt unnoticed by the
user.
3.2 EEG Signal Processing activity typically correspond to lower cognitive engagement, with
3.2.1 Device. Our system uses the Muse 2 EEG headband, building alpha waves linked to relaxation or passive states of rest [24, 30, 31].
on prior research that has leveraged consumer-grade devices with The engagement index has been widely validated across various
1 to 6 channels to assess cognitive engagement in learning contexts applications, including cognitive load assessments [27], visual pro-
(e.g., [34, 51, 79, 83]). The Muse 2 samples at 256 Hz and includes cessing studies, and sustained attention tasks [10]. It has also been
electrodes at Fpz, AF7, AF8, TP9, and TP10, following the 10-20 applied in complex task environments such as the multi-attribute
System (Figure 2) [42]. The Fpz electrode serves as the reference. task battery (MATB) [64], which involves tasks like tracking, re-
EEG data is streamed to a web browser using the open-source source management, and communication. These studies demon-
MuseJS library [74], which enables real-time streaming via Web strate the engagement index’s effectiveness in detecting attention
Bluetooth. shifts and fluctuations in cognitive state triggered by external stim-
uli [3, 16].
3.2.2 Preprocessing. The EEG data processing pipeline follows Following our preprocessing pipeline, we extract frequency bands
established methods from Hassib et al. [34], Kosmyna and Maes for each epoch and average them over a 15-second sliding window,
[51], Szafir and Mutlu [79] and others. A bandpass filter (1–30 as established by Szafir and Mutlu [79]. Averaging over a time win-
Hz) is applied to retain relevant neural activity while minimizing dow allows us to assess a user’s engagement over a meaningful
noise, and a 60 Hz notch filter removes power line interference. The duration while they read and process the LLM’s output, rather than
data is then segmented into 1-second epochs with 250 ms intervals capturing momentary fluctuations. We selected a 15-second win-
to enable continuous analysis with sufficient temporal resolution. dow to account for variations in reading speed, ensuring sufficient
Power spectral density is computed via fast Fourier transform (FFT), time for users to engage with the response. Unlike previous studies,
and band power is extracted for each frequency range to derive we opted against exponentially weighted moving averages, as our
meaningful neural features. focus is on sustained engagement throughout a task rather than
transient cognitive spikes.
3.2.3 Engagement Score. The engagement index (or engagement Finally, we normalize the engagement score following Kosmyna
score) serves as the core metric of our system, enabling real-time and Maes [51]. Normalization requires determining a minimum
quantification of cognitive engagement during mentally demanding and maximum engagement score for each user, which we obtain
tasks. First introduced by Pope et al. [64], this metric is computed from the calibration task conducted before the main experiment.
as a ratio of key EEG frequency bands using the formula: During calibration, users engaged in two tasks, each lasting two
minutes:
𝐸=
𝛽
(1) (1) Relaxation: Participants remain still, minimizing cognitive
𝛼 +𝜃 effort while we record baseline EEG data.
where 𝛽 (11–20 Hz), 𝛼 (7–11 Hz), and 𝜃 (4–7 Hz) correspond to (2) Mental word association: Participants perform a cognitive
EEG-derived neural oscillations. The index is based on the principle task that requires generating words based on the final let-
that higher beta power reflects heightened brain activity during ter of the previous word (e.g., "elephant" → "tiger"). This
cognitive tasks [11]. The beta frequency band is particularly as- method has been shown to effectively induce cognitive
sociated with cognitive processes such as visual attention, motor activation in non-ALS participants [50].
planning, and active information processing, all of which indicate The lowest and highest engagement scores from the two tasks,
an engaged mental state. Conversely, increased alpha and theta respectively, are taken as the normalization minimum 𝐸 min and
NeuroChat: A Neuroadaptive AI Chatbot for Customizing Learning Experiences
Figure 4: NeuroChat user interface with exposed EEG metrics in the user prompt and experimenter control menu. The
connection to the Muse EEG device happens through the Brain Widget in the top right corner. “Mood mode” activates the
EEG metric injection into the user’s prompts, and turning off “Debug mode” allows the experimenter to hide these from the
user. Chats, raw and filtered EEG data, and computed EEG metrics from the Muse device are stored in the browser’s native
IndexedDB and can be exported from the Settings panel.
maximum 𝐸 max . The calibration task differs from the main task in scientific depth, response format (e.g., bullet points versus long-
that it uses a 10-second sliding window, balancing the 5-second form text), and use of Socratic questioning. Additionally, framing
interval used in prior studies [34, 51, 79] and the 15-second window its role as a “good tutor” improved response quality.
applied in our main experiment. The normalized engagement score
𝐸 norm is then calculated using: 3.4 User Interface
The NeuroChat user interface (UI) consists of four key components
𝐸 − 𝐸 min (Figure 4):
𝐸 norm = (2)
𝐸 max − 𝐸 min
(1) Brain Connect Widget (Top Right Corner): Allows users
where E represents the engagement score averaged over the past to connect or disconnect the EEG device, calibrate or re-
15 seconds. calibrate the system, and start or stop EEG recording to
compute the engagement index.
3.3 LLM Adaptation (2) Calibration Modal (Full Screen) – Appears only after the
The mechanism through which NeuroChat responds adaptively to Muse headset is connected and provides instructions for a
the user’s cognitive state is by embedding their engagement score 2-minute relaxation phase followed by a 2-minute mental
into each query submitted to a Large Language Model. We used word association task. If the EEG connection is lost, users
OpenAI’s GPT-4-turbo model, the latest at the point of study. can restart or resume from the completed relaxation phase.
The system prompt provides a guideline as to how the LLM (3) Chatbot Interface (Main Screen Area) – Functions simi-
should adapt its response style to the user. We developed it through larly to ChatGPT, displaying an alternating conversation
careful evaluation of pilot testing and insights from OpenAI’s Teach- between the user and the AI tutor.
ing with AI guide [59] (see Appendix A.1). One key insight was (4) Menu Sidebar (Left-Hand Side) – Contains chat history
that prompting GPT-4 to increase the user’s “engagement” often re- where users can manage past LLM conversations by cre-
sulted in overly casual, upbeat responses, as the model interpreted ating folders, renaming chat titles, and deleting individual
the term informally. Reframing the engagement score as a “cogni- chats. Also features the Settings panel, which provides op-
tive load metric” helped maintain a neutral tone while allowing tions to toggle “Mood Mode” (enabling LLM adaptation),
responses to adjust dynamically based on neurofeedback. Based on activate “Debug Mode” (hiding EEG metrics from the UI),
the engagement index, the LLM was instructed to modulate detail, import/export chat history, download EEG data (from the
Baradari, Kosmyna, Petrov, Kaplun, and Maes
browser’s IndexedDB), switch between dark and light mode, while still remaining accessible for participants. Condition and
and reset the chat history for a new user session. The terms study topic order were counterbalanced using a Latin square design.
“Mood Mode” and “Debug Mode” were intentionally chosen Before the session, participants signed a consent form and turned
to provide visual cues to the experimenters while being off their electronic devices. They were fitted with a Muse EEG
vague enough to the participants. headband, and signal quality was verified via the Muse EEG app
[112]. Participants were instructed to minimize movement to reduce
4 METHODOLOGY motion artifacts.
The study lasted about 2 hours and proceeded as follows (Fig-
4.1 Hypotheses
ure 5):
Based on prior research, we formulate the following hypotheses:
(1) Pre-Session Measures: Participants completed a background
• (H1) Objective Engagement: NeuroChat will elicit higher questionnaire assessing their alertness and previous ex-
engagement levels than interaction with a standard GPT perience with AI chatbots. A brief EEG calibration phase
model, as measured by EEG-derived engagement scores. followed (2 minutes relaxation, 2 minutes mental exercise).
• (H2) Subjective Engagement: Participants will report greater (2) AI Chatbot Interaction: Participants engaged with the chat-
subjective engagement and satisfaction with NeuroChat, bot for 20 minutes on their first assigned topic, with the
perceiving it as more engaging and effective than a tradi- goal of “learning as much as possible.” To guide exploration,
tional AI tutoring model. they received starting pointers—e.g., characteristics, behav-
• (H3) Learning Outcomes: Participants using NeuroChat will ior, and archaeological research for T. rex and historical
achieve higher scores on post-interaction learning assess- context, significance, and consequences for the Taiping
ments compared to those using the standard GPT model. Rebellion. Participants were free to focus on aspects they
found interesting.
4.2 Participants (3) Knowledge Assessment: Immediately after the chatbot inter-
Thirty participants (15 female, 13 male, 2 other), predominantly action, participants completed a quiz consisting of fill-in-
from academic backgrounds, were recruited for this study (M = 32.4 the-blank and multiple-choice (MCQ) questions, followed
years, median = 30) and compensated with a $50 Amazon gift card. by a 15-minute essay to assess understanding. To prevent
The study received approval from MIT’s institute’s ethical review preparatory bias, participants were not informed about the
board (protocol no. 21070000428). quiz beforehand. The same quiz was used across conditions.
(4) Break & Condition Switch: Participants took a short break
4.3 Study Design and Protocol before repeating the process with the second topic and
We adopted a within-subject study design after pilot studies re- condition. EEG data and chatbot interaction logs were con-
vealed significant individual differences in interactions with the AI tinuously recorded.
chatbot. This design allowed each participant to serve as their own (5) Final Survey & Interview: Participants completed a post-
control, minimizing variability and enabling direct performance study user survey and a semi-structured interview focusing
comparisons between the NeuroChat experimental condition and on their subjective engagement and experience across con-
the control condition. The control condition consisted of a regular ditions. Interviews were thematically analyzed.
GPT chatbot, which was prompted to act within an AI tutoring task
via its system prompt for fair comparison (see Appendix A.2). 4.4 Evaluation
As study topics, we selected the Tyrannosaurus rex (T. rex) and the Assessing learning outcomes requires a multifaceted approach. The
Taiping Rebellion. These topics were chosen to minimize prior topic quizzes incorporated recall-based and synthesis-based questions to
bias while allowing room for facts and explorative interpretation. capture different cognitive processes. Recall was assessed through
Although the T. rex is widely recognized, most people lack in- multiple-choice (MCQ) and fill-in-the-blank questions, while cre-
depth knowledge about the dinosaur. Similarly, despite its historical ative synthesis was evaluated via mini-essays requiring critical
significance, the Taiping Rebellion is rarely emphasized in Western thinking and analysis. MCQ and fill-in-the-blank questions were
education. Both topics provided sufficient complexity and depth designed to assess factual recall, covering information likely en-
for meaningful engagement within the 20-minute learning session countered during topic exploration. Question complexity varied to
NeuroChat: A Neuroadaptive AI Chatbot for Customizing Learning Experiences
(a) Model summary for Engagement (normalized) by Condition, accounting for Order. (b) Model fit statistics.
reflect a range of difficulty levels appropriate for participants with high between-subject variability (intra-class correlation coefficient
minimal prior knowledge engaging for under 20 minutes. For the (ICC) = 36.5%, p < 0.05). Therefore, to account for individual baseline
essay task, participants had 15 minutes to write and could choose differences while preserving within-subject variability, we applied
from a set of prompts or create their own. z-score normalization, adjusting each participant’s engagement
One author manually scored the quiz blind. Fill-in-the-blank scores based on their mean and standard deviation across both con-
questions were graded with 2 points each, with partial scores for ditions. This allowed for direct comparison of relative engagement
semi-correct answers, and multiple choice questions (MCQ) were differences between NeuroChat and the control condition.
given 1 point per correct option. Since we only wanted to grade Given the repeated-measures design, where each participant con-
responses that had come up in the chat interaction, an automated tributed data across multiple conditions, and occasionally missing
keyword detection script scanned a participant’s message history data, we fit a Linear Mixed Model (LMM) to examine the effect of
for the presence or absence of a question, which was checked man- the experimental condition (Condition) on normalized engagement
ually. Answers not covered were excluded from the participant’s (𝐸 norm ) (Table 1). For such conditions, an LMM is more appropriate
total score, leaving us with proportional participant scores for com- than a paired t-test, as it accounts for within-subject correlations
parison. and random variability across participants. We used the statsmodels
The 15-minute mini-essays were graded blind by a professional package in Python [73].
high school English teacher on a 5-point scale across 4 categories: An initial model with Condition as the sole fixed effect did not
Content, Structure and Organization, Language and Style, and Ac- yield a statistically significant relationship with engagement (𝛽 =
curacy (Spelling and Grammar). The Content category was given 0.186, p = 0.063). Since task order could introduce confounding ef-
double weighting when calculating the final score. fects, we extended the model by adding Order as an additional fixed
In addition to objective assessments, participants completed a effect. This refined model revealed a significant effect of Condition
post-study user survey and were interviewed by 1 of 3 of the au- (𝛽 = 0.186 = 0.216, p = 0.029) and a marginal effect of Order (𝛽 =
thors in a semi-structured interview lasting no more than 5 minutes. 0.186 = 0.181, p = 0.068) on normalized engagement (Figure 6). This
The interviewers took note of key quotes and sentiments and sub- finding suggests that task order may have influenced engagement
sequently cross-read each other’s notes and discussed additional levels, warranting its inclusion in the final model. As expected,
takeaways. Survey responses and interview notes were compiled the random effect variance for participants was low (𝜎2 = 0.001),
and thematically analyzed by the first author over multiple rounds reflecting the impact of z-score normalization, which minimized
of descriptive coding. inter-individual differences before running the model, leaving only
within-condition variation.
5 RESULTS Extending the model to test for additional effects of study topic,
age, education level, chatbot experience, chatbot familiarity, chat-
To evaluate the effects of NeuroChat on engagement and learning
bot usage frequency, and prompt engineering skill revealed no
outcomes, we conducted analyses on EEG engagement scores, learn-
significant influence of these factors (p > 0.1). In conclusion, when
ing assessments, and user feedback. Our results address three key
accounting for order effects during the experiment, participants in
areas: (1) cognitive engagement (EEG-derived engagement index),
the NeuroChat condition were, on average, relatively more cogni-
(2) user-reported engagement, and (3) learning performance (quiz
tively engaged.
and essay scores).
Figure 6: Distribution of engagement score means in the control and experimental conditions by order (right: z-score normalized).
5.3 Perceived Engagemenet & Subjective more factual and concise responses. They felt that NeuroChat’s con-
Evaluations versational style detracted from the focus on information, making
it harder to digest the content. Similarly, P26 found the control’s
We analyzed user responses based on the post-questionnaires and
more nuanced, fact-driven responses preferable, describing Neu-
informal interviews regarding reported levels of engagement and
roChat’s conversational prompts as distracting rather than helpful.
learning preferences. Participants were asked in writing and ver-
Despite this, many participants noted that the control chatbot often
bally about their perceived engagement, noticeable differences, and
lacked the personal touch, with P31 describing the control chatbot
learning preferences between the chats. We categorized their feed-
as feeling like “a regular chatbot” that lacked awareness of the
back into five main themes: personalized feedback, and response
user’s emotions or engagement level.
style (factual vs. conversational), density of information, follow-
up questions, and additional feedback, each contributing to the 5.3.2 Factual vs. Casual Response Style. The response style be-
perceived engagement and satisfaction. tween NeuroChat and the control chatbot was another point of
divergence in subjective feedback. Participants like P28 and P30
enjoyed the more conversational and engaging tone of NeuroChat.
5.3.1 Personalized Feedback and Human-Like Responses. A promi- P28 described the experimental chatbot as “very fun, like a tour
nent theme in the subjective evaluations was NeuroChat’s more guide,” with a more interactive and fluid exchange, while the control
human-like responses, which many participants found engaging. felt “like a textbook” in comparison. These participants appreci-
P1 noted that the chatbot “mimics a real person” and provided ated NeuroChat’s ability to dive deeper into topics, making the
feedback that made it seem more interactive and lifelike, such as learning experience more dynamic and enjoyable. In contrast, some
saying “Great question” after user input. P31 also expressed sat- participants preferred the control chatbot’s more formal, factual
isfaction with the experimental chatbot, stating, “Oh, I loved the style. P7 found that the control condition allowed for more focused
second one! I really liked how it was saying how I was feeling.” learning, noting that NeuroChat was more prone to casual conver-
This participant emphasized the importance of NeuroChat’s ability sation that made it harder to focus on the core information. P17
to provide feedback tailored to their emotional state, suggesting reflected that while the experimental chatbot was more enjoyable
that it was responding to affect and engagement levels in real time. due to its fluidity and tendency to present fun facts, the control chat-
P19 added that NeuroChat had “more personality” compared to bot’s responses were better structured and felt more educational,
the standard GPT, making the interaction feel more tutor-like and comparing the control to a “blog post” with dense information.
conversational, rather than merely factual. However, some participants also criticized the control chatbot for
Furthermore, NeuroChat’s personalized prompts made the ex- being too rigid and not encouraging exploration. P32 mentioned
perience feel more responsive and adaptive for some participants. that while the control condition was more “analytical,” it felt more
P30 described the experimental chatbot as “always responding to like attending “a serious lecture,” with little room for the more
my prompt,” and noted how it felt more dynamic than the control enjoyable, exploratory exchanges that NeuroChat provided. P19
chatbot, which often came across as rigid and formal. Similarly, similarly remarked that the control chatbot seemed “less eager to
P28 enjoyed the depth of engagement, stating that NeuroChat al- engage,” contributing to a less immersive and personalized learning
lowed for “more meaningful topics to ask,” creating a richer, more experience.
exploratory interaction.
However, some participants preferred the more straightforward 5.3.3 Density of Information. The verbosity of NeuroChat’s re-
approach of the control chatbot. P7 and P10, who identified them- sponses was a double-edged sword for participants. Some, like P14
selves as “scientific minds,” favored the control condition for its and P12, appreciated the deeper exploration of topics provided by
NeuroChat: A Neuroadaptive AI Chatbot for Customizing Learning Experiences
NeuroChat, which allowed them to learn more than the control (10 participants). This gives us supporting evidence that partic-
chatbot. P14 commented that while they would choose the con- ipants using NeuroChat report higher levels of engagement and
trol chatbot for quick learning, NeuroChat was more suitable for satisfaction than those using a chatbot without neurofeedback (H2).
in-depth exploration due to its comprehensive answers. Similarly,
P28 noted that the experimental chatbot fostered a more nuanced
6 DISCUSSION
understanding, allowing them to focus on the context and reasons
behind a topic rather than just absorbing arbitrary facts. However, This study examined whether NeuroChat, a neuroadaptive AI chat-
other participants found the sheer volume of information over- bot, enhances engagement and learning outcomes by adapting its
whelming. P27 described NeuroChat’s responses as “paragraph responses based on real-time EEG feedback. Our results confirm
after paragraph,” and P32 admitted to skipping parts of its ver- that NeuroChat successfully increased cognitive (EEG-measured)
bose answers, opting instead for the control’s more concise and and self-reported engagement, demonstrating its feasibility as a
digestible responses. P19 and P33 both noted that NeuroChat had a neuroadaptive tutoring system. However, no significant differences
tendency to provide redundant information, which diminished the were found in learning performance, indicating challenges in trans-
clarity and relevance of the responses over time. While the control lating engagement into measurable knowledge gains. Below, we
chatbot was favored for its brevity, some participants like P30 found discuss these findings in the broader context of adaptive learn-
that it occasionally oversimplified complex topics, limiting deeper ing, generative AI, and brain-computer interfaces (BCIs) before
understanding. exploring key challenges and future directions.
5.3.4 Follow-Up Questions. Feedback on the follow-up questions 6.1 Key Findings: Engagement Gains, Learning
varied widely between the NeuroChat and the control condition. Outcomes, and User Perception
Participants appreciated the specific nature of the questions in Neu- We found that NeuroChat significantly increased EEG-measured
roChat. P2 stated, "I liked the questions at the end, they were really engagement (p = 0.029), suggesting that real-time neuroadaptive
specific—much better than Copilot." This sentiment was echoed by feedback can enhance sustained attention. These findings align
P23, who found the prompts helpful, especially since they “didn’t with prior neuroadaptive systems like Pay Attention! [79] and En-
know anything about the topic,” and appreciated the guidance. gageMeter [34], which successfully modulated engagement through
However, not all participants found the follow-up questions use- adaptive presentation techniques. However, NeuroChat goes fur-
ful. P18 felt that while they were prompted with questions, the ther by modifying conversational flow, response complexity, and
responses didn’t lead anywhere meaningful, causing frustration. interaction depth, positioning it as an interactive, closed-loop sys-
P26 provided a particularly nuanced critique, noting that the experi- tem rather than a passive monitoring tool.
mental chatbot’s prompts felt superficial and it didn’t seem to “care” User feedback also reflected higher perceived engagement with
the way a human would. They felt that the questions prompted by NeuroChat, with many participants finding its responses more
NeuroChat often missed their actual interests, leading to a sense human-like, responsive, and personalized. However, individual dif-
of disconnection from the conversation. P33 also noted that Neu- ferences emerged—some users preferred a factual, concise style,
roChat’s tendency to drive the conversation in a specific direction while others favored conversational, exploratory interactions. This
was problematic, as they had to repeatedly bring it back to their suggests that adaptive tutoring must account for personalized learn-
original question, which disrupted the flow of engagement. ing preferences to be effective.
In contrast, participants found the control chatbot’s lack of Despite increased engagement, NeuroChat did not significantly
follow-up questions to be limiting in terms of engagement. P19 improve quiz or essay scores. This result mirrors previous studies
noted that while the control was concise, it didn’t prompt any on learning with LLMs [53], which found that while personalization
follow-up questions, making the conversation feel more transac- enhances motivation, it does not always yield better performance
tional and less interactive. This lack of conversational depth in the on traditional assessments. Possible explanations include:
control condition was also mentioned by P23, who described it as
having “more general, broad questions,” which felt less engaging (1) Engagement ≠ Effective Learning – EEG-based engagement
than the more creative, tailored prompts offered by NeuroChat. captures sustained attention, but not necessarily deep learn-
ing or knowledge retention.
(2) Task Design Limitations – Unlike structured adaptive sys-
5.3.5 Overall Engagement and Satisfaction. In summary, subjective tems like BACh [87] (which progressively increased pi-
feedback on NeuroChat’s engagement and satisfaction levels varied, ano sheet music difficulty), NeuroChat allowed open-ended
with the conversational and human-like elements appealing to par- learning, making structured difficulty progression harder
ticipants seeking a more interactive, engaging experience. However, to implement.
those who preferred straightforward, fact-focused learning found (3) Short Study Duration – Neuroadaptive learning benefits
the control chatbot more aligned with their needs. NeuroChat’s may emerge over multiple sessions, but our study measured
neurofeedback-driven prompts were effective for some, but oth- learning in a single interaction.
ers found them intrusive or misaligned with their interests, which
could detract from overall satisfaction. Overall, more participants Future work should explore long-term retention, conceptual un-
provided positive feedback on their engagement in their experimen- derstanding, and scaffolding techniques that could better translate
tal condition (23 participants) compared to the control condition engagement gains into measurable learning improvements.
Baradari, Kosmyna, Petrov, Kaplun, and Maes
6.2 Implications for Neuroadaptive Learning & mechanisms, personalize content, and balance engagement with
AI-Powered Tutoring cognitive load. By integrating multimodal sensing and long-term
user modeling, AI tutors could one day provide truly personalized,
Traditional neuroadaptive learning systems relied on pre-scripted
lifelong learning experiences.
content, where researchers manually assigned learning materials
to high- or low-engagement conditions. NeuroChat overcomes this
limitation by leveraging generative AI to create content dynam-
ACKNOWLEDGMENTS
ically, enabling real-time adaptation tailored to individual users. We thank Treyden Chiaravalloti for his valuable piloting support
This is a fundamental shift in adaptive learning—moving from rule- and insightful feedback. We also appreciate Protyasha Nishat’s
based, pre-mapped content to generative, personalized tutoring. expertise in signal processing and Nathan Whitmore’s comments
Most LLMs require users to explicitly communicate their needs on study design. Luisa Heiss’s thorough grading of the essays was
(e.g., "Explain this differently", "Make it simpler"). NeuroChat re- instrumental in evaluating participant test performance without
moves this barrier by inferring user engagement levels directly from bias. This research was supported by the MIT J-WEL Education
EEG data, reducing the need for manual prompt engineering. This Innovation Grant.
has implications for personalized AI assistants, where cognitive
state tracking could enhance adaptability without user effort. REFERENCES
NeuroChat has particular relevance for self-directed learners, [1] 2023. Imagination Engine I: Generating Abstract Art through EEG.
[Link]
who often struggle with maintaining engagement. In 2021, over 220 abstract-art-through-eeg.
million students enrolled in MOOCs, yet the average completion [2] 2023. Technology in education. Technical Report. UNESCO.
rate remains at 13% [2, 58]. A neuroadaptive AI tutor could help [3] Yomna Abdelrahman, Mariam Hassib, Maria Guinea Marquez, Markus Funk,
and Albrecht Schmidt. 2015. Implicit engagement detection for interactive
sustain motivation and prevent disengagement, particularly in open- museums using brain-computer interfaces. In Proceedings of the 17th International
ended, autonomous learning environments. Conference on Human-Computer Interaction with Mobile Devices and Services
Adjunct. ACM. doi:10.1145/2786567.2793709
Beyond education, NeuroChat’s EEG-driven AI system could [4] Lena M Andreessen, Peter Gerjets, Detmar Meurers, and Thorsten O Zander.
support knowledge workers, particularly those who struggle with 2021. Toward neuroadaptive support technologies for improving digital reading:
focus and information retention. Additionally, LLM-based BCIs a passive BCI-based assessment of mental workload imposed by text difficulty
and presentation speed during reading. User Model. User-adapt Interact. 31 (March
have been proposed for aiding individuals with learning challenges, 2021), 75–104. doi:10.1007/s11257-020-09273-5
including ADHD [13, 39, 62]. These applications highlight the po- [5] Andrea Apicella, Pasquale Arpaia, Mirco Frosolone, Giovanni Improta, Nicola
tential of neuroadaptive AI beyond the classroom. Moccaldi, and Andrea Pollastro. 2022. EEG-based measurement system for
monitoring student engagement in learning 4.0. Sci. Rep. 12 (7 April 2022), 5857.
doi:10.1038/s41598-022-09578-y
6.3 Challenges & Limitations of NeuroChat [6] Roger Azevedo. 2015. Defining and measuring engagement and learning in
science: Conceptual, theoretical, methodological, and analytical issues. Educ.
Participants showed high variability in engagement and preference Psychol. 50 (2 Jan. 2015), 84–94. doi:10.1080/00461520.2015.1004069
for different interaction styles. Some learners thrived in guided, ex- [7] Yunpeng Bai, Xintao Wang, Yan-Pei Cao, Yixiao Ge, Chun Yuan, and Ying Shan.
2025. DreamDiffusion: High-quality EEG-to-image generation with temporal
ploratory conversations, while others preferred concise, fact-driven masked signal modeling and CLIP alignment. In Lecture Notes in Computer Science.
responses. Future systems should incorporate user preference set- Springer Nature Switzerland, 472–488. doi:10.1007/978-3-031-72751-1_27
[8] N P Bechtereva and V B Gretchin. 1968. Physiological foundations of mental
tings, such as: activity. Int. Rev. Neurobiol. 11 (1968), 329–352. doi:10.1016/s0074-7742(08)60392-
• Preferred interaction style (e.g., structured vs. exploratory). x
[9] Antoine Bellemare-Pepin, Philipp Thölke, Yann Harel, and Karim Jerbi. 2024.
• Response format (e.g., bullet points vs. narratives). Real-Time Neuro-Augmented Cinema via Generative AI. NeurIPS Workshop on
• Memory-based personalization (as seen in OpenAIś user Creativity & Generative AI.
[10] C Berka, D J Levendowski, M N Lumicao, A Yau, G Davis, V T Zivkovic, R E Olm-
memory feature [61]). stead, P D Tremoulet, and P L Craven. 2007. EEG correlates of task engagement
and mental workload in vigilance, learning, and memory tasks. Aviation, space,
Moreover, high engagement is not always beneficial. Medical and environmental medicine 78 (May 2007).
reasoning studies [22] found that struggling learners showed the [11] Wolfram Boucsein, Andrea Haarmann, and Florian Schaefer. 2007. Combining
highest engagement, suggesting that increased engagement can skin conductance and heart rate variability for adaptive automation during sim-
ulated IFR flight. In Engineering Psychology and Cognitive Ergonomics. Springer
sometimes signal cognitive overload rather than productive learn- Berlin Heidelberg, 639–647. doi:10.1007/978-3-540-73331-7_70
ing. Future neuroadaptive tutors must ensure users remain in their [12] E A Byrne and R Parasuraman. 1996. Psychophysiology and adaptive automation.
zone of proximal development rather than pushing them beyond Biol. Psychol. 42 (5 Feb. 1996), 249–268. doi:10.1016/0301-0511(95)05161-9
[13] Andrea Caria. 2024. Towards predictive communication with brain-computer
their capabilities. interfaces integrating large language models. arXiv [[Link]] (10 Dec. 2024).
Consumer EEG devices are fundamentally noisy, and EEG signals [14] Hongyu Chen, Weiming Zeng, Chengcheng Chen, Luhui Cai, Fei Wang, Yuhu
Shi, Lei Wang, Wei Zhang, Yueyang Li, Hongjie Yan, Wai Ting Siok, and Nizhuan
contain biometric markers that can uniquely identify individuals Wang. 2024. EEG Emotion Copilot: Optimizing lightweight LLMs for emo-
[69]. As LLMs process data externally, privacy concerns must be tional EEG interpretation with assisted medical record generation. arXiv [[Link]]
addressed before large-scale adoption of EEG-driven AI tutors. (30 Sept. 2024).
[15] Wenhui Cui, Woojae Jeong, Philipp Thölke, Takfarinas Medani, Karim Jerbi,
Anand A Joshi, and Richard M Leahy. 2024. Neuro-GPT: Towards A foundation
7 CONCLUSION model for EEG. In 2024 IEEE International Symposium on Biomedical Imaging
(ISBI), Vol. 35. IEEE, 1–5. doi:10.1109/isbi56570.2024.10635453
We provide evidence that EEG-driven AI chatbots can enhance [16] Alex Dan and Miriam Reiner. 2017. Real time EEG based measurements of
engagement, bringing users closer into a zone of proximal develop- cognitive load indicates mental states during learning. JEDM 9 (23 Dec. 2017),
31–44. doi:10.5281/ZENODO.3554719
ment, but highlight challenges in translating engagement into learn- [17] David Game College. 2024. GCSE AI Adaptive Learning Programme.
ing gains. Future neuroadaptive AI systems must refine adaptation [Link]
NeuroChat: A Neuroadaptive AI Chatbot for Customizing Learning Experiences
[61] OpenAI. 2024. Memory and new controls for ChatGPT. [Link] [85] Łukasz Warchoł and Ludmiła Zając-Lamparska. 2023. The relationship of N200
index/memory-and-new-controls-for-chatgpt/. and P300 amplitudes with intelligence, working memory, and attentional control
[62] Brendan Parsons and Jocelyn Faubert. 2021. Enhancing learning in a perceptual- behavioral measures in young healthy individuals. Adv. Cogn. Psychol. 19 (2023),
cognitive training paradigm using EEG-neurofeedback. Sci. Rep. 11 (18 Feb. 2021), 63–75. doi:10.5709/acp-0404-2
4061. doi:10.1038/s41598-021-83456-x [86] Shaoyue Wen, Michael Middleton, Songming Ping, Nayan N Chawla, Guande
[63] Salil H Patel and Pierre N Azzam. 2005. Characterization of N200 and P300: Wu, Bradley S Feest, Chihab Nadri, Yunmei Liu, David Kaber, Maryam Zahabi,
selected studies of the Event-Related Potential. Int. J. Med. Sci. 2 (1 Oct. 2005), Ryan P McMahan, Sonia Castelo, Ryan Mckendrick, Jing Qian, and Claudio Silva.
147–154. doi:10.7150/ijms.2.147 2025. AdaptiveCoPilot: Design and testing of a NeuroAdaptive LLM cockpit
[64] A T Pope, E H Bogart, and D S Bartolome. 1995. Biocybernetic system evaluates guidance system in both novice and expert pilots. arXiv [[Link]] (7 Jan. 2025).
indices of operator engagement in automated task. Biol. Psychol. 40 (May 1995), [87] Beste F Yuksel, Kurt B Oleson, Lane Harrison, Evan M Peck, Daniel Afergan,
187–195. doi:10.1016/0301-0511(95)05116-3 Remco Chang, and Robert J K Jacob. 2016. Learn Piano with BACh: An Adaptive
[65] Mirko Raca and Pierre Dillenbourg. 2013. System for assessing classroom atten- Learning Interface that Adjusts Task Difficulty Based on Brain State. In Proceed-
tion. In Proceedings of the Third International Conference on Learning Analytics ings of the 2016 CHI Conference on Human Factors in Computing Systems (CHI ’16).
and Knowledge. ACM. doi:10.1145/2460296.2460351 Association for Computing Machinery, 5372–5384. doi:10.1145/2858036.2858388
[66] Mirko Raca, Lukasz Kidzinski, and Pierre Dillenbourg. 2015. Translating Head [88] Xiayin Zhang, Ziyue Ma, Huaijin Zheng, Tongkeng Li, Kexin Chen, Xun Wang,
Motion into Attention - Towards Processing of Student’s Body-Language. Inter- Chenting Liu, Linxi Xu, Xiaohang Wu, Duoru Lin, and Haotian Lin. 2020. The
national Educational Data Mining Society (June 2015). combination of brain-computer interfaces and artificial intelligence: applications
[67] Johnmarshall Reeve and Ching-Mei Tseng. 2011. Agency as a fourth aspect of and challenges. Ann. Transl. Med. 8 (June 2020), 712. doi:10.21037/atm.2019.11.109
students’ engagement during learning activities. Contemp. Educ. Psychol. 36 [89] Yuhong Zhang, Qin Li, Sujal Nahata, Tasnia Jamal, Shih-Kuen Cheng, Gert
(1 Oct. 2011), 257–267. doi:10.1016/[Link].2011.05.002 Cauwenberghs, and Tzyy-Ping Jung. 2024. Integrating large language model,
[68] Lauren E Reinerman, Gerald Matthews, Joel S Warm, Lisa K Langheim, Kelley EEG, and eye-tracking for word-level neural state classification in reading com-
Parsons, Christina A Proctor, Tazeen Siraj, Lloyd D Tripp, and Robert M Stutz. prehension. IEEE Trans. Neural Syst. Rehabil. Eng. PP (14 Aug. 2024), 1–1.
2006. Cerebral blood flow velocity and task engagement as predictors of vigilance doi:10.1109/TNSRE.2024.3435460
performance. Proc. Hum. Factors Ergon. Soc. Annu. Meet. 50 (Oct. 2006), 1254–1258. [90] Tong Zhou, Xuhang Chen, Yanyan Shen, Martin Nieuwoudt, Chi-Man Pun,
doi:10.1177/154193120605001210 and Shuqiang Wang. 2023. Generative AI enables EEG data augmentation for
[69] Maria V Ruiz-Blondet, Zhanpeng Jin, and Sarah Laszlo. 2016. CEREBRE: A novel Alzheimer’s disease detection via diffusion model. In 2023 IEEE International
method for very high accuracy event-related potential biometric identification. Symposium on Product Compliance Engineering - Asia (ISPCE-ASIA). IEEE, 1–6.
IEEE Trans. Inf. Forensics Secur. 11 (July 2016), 1618–1629. doi:10.1109/tifs.2016. doi:10.1109/ispce-asia60405.2023.10365931
2543524
[70] Yashvir Sabharwal and Balaji Rama. 2024. Comprehensive review of EEG-to-
output research: Decoding neural signals into images, videos, and audio. arXiv
[[Link]] (27 Dec. 2024). A SYSTEM PROMPTS
[71] Akane Sano, Judith Amores, and Mary Czerwinski. 2024. Exploration of LLMs,
EEG, and behavioral data to measure and support attention and sleep. arXiv A.1 NeuroChat system prompt
[[Link]] (1 Aug. 2024).
[72] Brooke Schultz. 2025. This School Will Have Artificial Intelligence Teach Kids NeuroChat System Prompt
(With Some Human Help). Education Week (6 Jan. 2025).
[73] Skipper Seabold and Josef Perktold. 2010. Statsmodels: Econometric and statis-
tical modeling with python. In Proceedings of the Python in Science Conference. You are an encouraging tutor who helps students across
SciPy, 92–96. doi:10.25080/majora-92bf1922-011 various subjects and skill levels understand concepts by
[74] Uri Shaked. 2021. muse-js: Muse 2016 EEG Headset JavaScript Library (using
Web Bluetooth). [Link]
explaining ideas and asking students questions. Start by
[75] David J Shernoff, Mihaly Csikszentmihalyi, Barbara Schneider, and Elisa Steele introducing yourself to the student as their AI-Tutor who
Shernoff. 2014. Student engagement in high school classrooms from the perspec- is happy to help them with any questions.
tive of flow theory. In Applications of Flow in Human Development and Education.
Springer Netherlands, 475–494. doi:10.1007/978-94-017-9094-9_24 Additionally, you will be provided with the student’s
[76] Gale M Sinatra, Benjamin C Heddy, and Doug Lombardi. 2015. The challenges of
defining and measuring student engagement in science. Educ. Psychol. 50 (2 Jan.
cognitive load values while they were reading any
2015), 1–13. doi:10.1080/00461520.2014.1002924 previous responses of yours as measured by EEG. Your
[77] W Speier, C Arnold, and N Pouratian. 2016. Integrating language models into goal is to act like a good tutor, using the insights from
classifiers for BCI communication: a review. J. Neural Eng. 13 (6 June 2016),
031002. doi:10.1088/1741-2560/13/3/031002 these metrics to adapt your responses to the student’s
[78] Ricarda Steinmayr, Anne F Weidinger, Malte Schwinger, and Birgit Spinath. cognitive load dynamically. The value you will be given:
2019. The importance of students’ motivation for their academic achievement
- replicating and extending previous findings. Front. Psychol. 10 (31 July 2019), **Normalized engagement score:** This represents the
1730. doi:10.3389/fpsyg.2019.01730 user’s level of engagement or arousal on a normalized
[79] Daniel Szafir and Bilge Mutlu. 2012. Pay attention!: designing adaptive agents that
monitor and improve user engagement. In Proceedings of the SIGCHI Conference scale from 0 to 1. The engagement index is a ratio of the
on Human Factors in Computing Systems. ACM. doi:10.1145/2207676.2207679 student’s beta/(theta+alpha) bands.
[80] The Abdul Latif Jameel Poverty Action Lab (J-PAL). 2022. Teaching at the
Right Level to improve learning. [Link] Do not ever disclose the EEG metrics to the user since
teaching-right-level-improve-learning.
[81] The Open Innovation Team and Department for Education. 2024. Generative AI they are hidden to them. Also, never make direct
in education - Educator and expert views. Technical Report. UK Department for comments on their metrics and don’t mention the names
Education. of the metrics.
[82] Philipp Tholke, Antoine Bellemare-Pepin, Yann Harel, Francois Lespinasse, and
Karim Jerbi. 2024. Bio-Mechanical Poet: An Immersive Audiovisual Playground Give students explanations, examples, and analogies
for Brain Signals and Generative AI. In Proceedings of the 15th international
conference on computational creativity, In Kazjon Grace, Maria Teresa Llano, about the concept to help them understand.
Martins Pedro, and Maria M Hedblom (Eds.). 65–74.
[83] Angela Vujic, Shreyas Nisal, and Pattie Maes. 2023. Joie: a Joy-based Brain- Adaptations Based on Cognitive Load:
Computer Interface (BCI). In Proceedings of the 36th Annual ACM Symposium on You need to learn how the user reacted to your
User Interface Software and Technology (UIST ’23). Association for Computing
Machinery, 1–14. doi:10.1145/3586183.3606761 adaptations. Based on their cognitive load, modulate the
[84] Yansen Wang and Zilong Wang. 2024. EEG2Video: Towards Decoding Dynamic response length, factual vs. storytelling, ease of text
Visual Perception from EEG Signals. The Thirty-eighth Annual Conference on
Neural Information Processing Systems (NeurIPS 2024).
(explain like I’m 5 vs. explain like I’m a PhD), bullet points
NeuroChat: A Neuroadaptive AI Chatbot for Customizing Learning Experiences
Essay Prompt. Answer one of the following in the form of a mini- Essay Prompt. Answer one of the following in the form of an essay
essay (introduction, main section, conclusion): (introduction, main section, conclusion):
(1) Discuss T. Rex’s physical and behavioral characteristics (1) Discuss the socio-economic factors that contributed to the
which enabled it to dominate its environment. outbreak of the Taiping Rebellion.
(2) Analyze paleobiological discoveries that have changed our (2) Analyze the role of religion in the Taiping Rebellion.
understanding of T. Rex. (3) Discuss the legacy of the Taiping Rebellion on subsequent
(3) Evaluate the theories regarding the function of T. Rex’s Chinese history.
small arms. What are some of the proposed explanations, (4) Create your own essay question related to the Taiping Re-
and which do you find most convincing? bellion based on your conversation.
(4) Create your own essay question related to the T. Rex based
on your conversation.
Taiping Rebellion
Fill-in-the-gap Questions.
(1) Q1: The goal of the Taiping Rebellion was to establish the
. (2 points): Taiping Heavenly Kingdom of Great
Peace, or similar.
(2) Q2: One of the distinctive aspects of the Taiping ideol-
ogy was its spiritual blend of and , which
appealed to the disaffected rural populace. (2 points): Chris-
tianity (1) and Chinese spiritual traditions (Buddhism,
Taoism, Confucianism) or similar (1).
Multiple Choice Questions. Q5: What was the name of the leader
of the Taiping Rebellion?
• Zeng Guofan
• Hong Xiuquan
• Hong Tianguifu
• Feng Yunshan
• Yang Xiuqing
• This did not come up in my conversation
• I don’t know
(1 point): Hong Xiuquan
Q6: The Qing government was administered by leaders of which
ethnic minority?
• Manchu
• Hakka
• Han
• Zhuang
• Hui
• This did not come up in my conversation
• I don’t know
(1 point): Manchu
Ranking Question. Q10: Here is a list of death tolls in wars by lowest
estimate. Where would the Taiping Rebellion sit?
(1) World War II - 80 million
(2) World War I - 17 million
(3) Spanish conquest of Mexico - 10.5 million
(4) Russian Civil War - 7 million
(5) Napoleonic Wars - 3.5 million
(6) Vietnam War - 1.3 million
(7) Taiping Rebellion
(2 points): Between World War II and World War I, with an
estimated toll of 20-30 million.
NeuroChat uses EEG data to adapt its responses in real-time by modifying conversational flow, response complexity, and interaction depth based on the learner's cognitive state. This approach offers a more personalized and adaptive learning experience compared to traditional systems that relied on pre-scripted content. The key benefit is the reduction of manual input from the user, enhancing engagement by dynamically tailoring content to individual needs. However, the limitations include challenges in translating increased engagement into measurable learning outcomes, as evidenced by no significant improvement in quiz or essay scores .
The Chatbot study highlights that user engagement preferences can significantly influence adaptive learning strategies. While some users favored conversational and human-like interactions offered by NeuroChat, others preferred factual and concise responses. This diversity in preferences suggests that effective adaptive learning systems must be flexible enough to tailor engagement strategies that cater to individual user preferences, potentially enhancing both perceived and actual engagement .
Neuroadaptive AI systems like NeuroChat have significant implications for future AI-powered tutoring. They mark a shift from rule-based, pre-mapped content to generative, personalized tutoring that adapts in real-time based on user engagement levels inferred from EEG data. This could enhance adaptability and reduce user effort in communicating needs. Additionally, it suggests a new avenue for creating self-directed learning environments that accommodate varied learning preferences by modifying engagement styles based on real-time feedback .
NeuroChat's distinctive features include its ability to provide human-like, responsive interactions by adapting its responses in real-time based on EEG-measured engagement. Users perceive it as more interactive due to its personalized feedback and conversational, tutor-like approach, which contrasts with standard large language models that offer more factual and structured responses. This human-like adaptability significantly enhances perceived engagement and interaction depth .
EEG monitoring contributes to the personalization of educational experiences by providing real-time feedback on a learner's cognitive states such as engagement, concentration, and cognitive load. This data is used to dynamically adjust instructional content or interaction modalities, ensuring the learner is neither underwhelmed nor overwhelmed. Techniques like EEG-based mental workload classification allow adaptive learning systems to tailor content complexity, improving engagement and potentially optimizing learning outcomes .
EEG for cognitive workload classification enhances adaptive learning systems by enabling them to adjust educational content based on the learner's mental state. By differentiating between high and low cognitive loads, these systems can modify content complexity and presentation style to maintain optimal engagement and prevent cognitive overload, thus creating a more effective learning environment .
Neuroadaptive systems like Thinking Cap differ from earlier adaptive learning technologies primarily in their responsive and dynamic adjustment of instructional content in real-time based on EEG-derived cognitive load measures. Unlike earlier systems that modified delivery methods but not content complexity, Thinking Cap involves creating varied difficulty levels of the content itself. This ensures learners are continuously challenged at an appropriate level, avoiding under- or overwhelming them, thus aligning with real-time cognitive states .
Challenges in converting engagement gains from neuroadaptive systems into measurable learning improvements include: (1) The engagement measured by EEG reflects sustained attention but may not translate to deep learning or knowledge retention. (2) Task design limitations, as NeuroChat's open-ended approach makes it difficult to implement structured difficulty progression that aids learning. (3) Short study duration, as benefits of neuroadaptive learning might require multiple sessions to manifest in measurable performance improvements. These challenges indicate the complexity of aligning engagement with effective learning strategies .
Vygotsky’s Zone of Proximal Development (ZPD) relates to EEG studies on student engagement by illustrating that heightened engagement occurs when tasks are slightly beyond a student's current skill level, demanding higher cognitive effort. EEG studies reveal significant attentional focus fluctuations, indicating tasks are within the ZPD when engagement is high, yet expertise is insufficient, leading to cognitive overload rather than effective learning. This suggests the importance of aligning task complexity with a learner's developmental stage for optimal learning outcomes .
Generative AI capabilities in neuroadaptive systems could revolutionize educational settings by allowing for real-time creation of personalized content that adapts seamlessly to a student's cognitive state. This reduces the need for pre-scripted materials and manual adjustments, facilitating more natural and responsive learning experiences. The AI could infer engagement levels from EEG data and adjust content complexity dynamically, providing a tailored learning path that enhances motivation and engagement without overloading learners .