0% found this document useful (0 votes)
39 views16 pages

NeuroChat: AI for Personalized Learning

NeuroChat is a neuroadaptive AI chatbot that integrates real-time EEG-based engagement tracking with generative AI to enhance personalized learning experiences. It continuously monitors learners' cognitive engagement and dynamically adjusts content complexity and response style, aiming to improve engagement and learning outcomes. A pilot study indicates that while NeuroChat enhances cognitive and subjective engagement, it does not show immediate effects on learning outcomes, highlighting the potential for real-time cognitive feedback in AI tutoring systems.

Uploaded by

vvce22cse0095
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
39 views16 pages

NeuroChat: AI for Personalized Learning

NeuroChat is a neuroadaptive AI chatbot that integrates real-time EEG-based engagement tracking with generative AI to enhance personalized learning experiences. It continuously monitors learners' cognitive engagement and dynamically adjusts content complexity and response style, aiming to improve engagement and learning outcomes. A pilot study indicates that while NeuroChat enhances cognitive and subjective engagement, it does not show immediate effects on learning outcomes, highlighting the potential for real-time cognitive feedback in AI tutoring systems.

Uploaded by

vvce22cse0095
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

NeuroChat: A Neuroadaptive AI Chatbot for Customizing

Learning Experiences
Dünya Baradari Nataliya Kosmyna Oscar Petrov
MIT Media Lab MIT Media Lab Brown University
Cambridge, MA, United States Cambridge, MA, United States Providence, RI, United States

Rebecah Kaplun Pattie Maes


MIT Media Lab MIT Media Lab
Cambridge, MA, United States Cambridge, MA, United States
arXiv:2503.07599v1 [[Link]] 10 Mar 2025

Figure 1: Overview of the NeuroChat neuroadaptive LLM system. A wearable dry-electrode EEG headband collects data from
the brain and sends it to the NeuroChat web app, which computes the user’s level of engagement. The engagement score is sent
with each request to the LLM, allowing it to adapt its response style to the user in real-time.
ABSTRACT pacing using a closed-loop system. We evaluate this approach in a
Generative AI is transforming education by enabling personalized, pilot study (n=24), comparing NeuroChat to a standard LLM-based
on-demand learning experiences. However, AI tutors lack the abil- chatbot. Results indicate that NeuroChat enhances cognitive and
ity to assess a learner’s cognitive state in real time, limiting their subjective engagement but does not show an immediate effect on
adaptability. Meanwhile, electroencephalography (EEG)-based neu- learning outcomes. These findings demonstrate the feasibility of
roadaptive systems have successfully enhanced engagement by real-time cognitive feedback in LLMs, highlighting new directions
dynamically adjusting learning content. This paper presents Neu- for adaptive learning, AI tutoring, and human-AI interaction.
roChat, a proof-of-concept neuroadaptive AI tutor that integrates
real-time EEG-based engagement tracking with generative AI. Neu- KEYWORDS
roChat continuously monitors a learner’s cognitive engagement Brain-computer interface, electroencephalography (EEG), chatbot,
and dynamically adjusts content complexity, response style, and conversational AI, closed-loop
Baradari, Kosmyna, Petrov, Kaplun, and Maes

1 INTRODUCTION However, existing neuroadaptive systems are limited by pre-


The rise of Generative Artificial Intelligence (AI) is reshaping edu- scripted content, where adaptations are constrained to pre-defined
cation, with Large Language Models (LLMs) offering new oppor- difficulty levels rather than generating new instructional mate-
tunities for personalized, on-demand learning. AI-powered tutors rial dynamically. The Online Continuous Adaptation Mechanism
such as ChatGPT Edu [60] and Khanmigo [46] have been inte- (OCAM) [22] attempted to overcome this limitation by continu-
grated into educational settings, promising a future where learners ously adjusting learning content based on EEG-derived measures
can interact dynamically with AI, receive customized explanations, of cognitive load, concentration, and emotional arousal. Yet, even
and engage in self-directed inquiry. AI-generated content has al- in OCAM, content had to be pre-designed and categorized by hu-
ready been shown to improve learning motivation [53] and enhance man experts before it could be adapted. This limitation raises a
teaching efficiency [56], as educators leverage these tools to tailor fundamental question: Can we combine real-time neuroadaptive
lesson plans, streamline administrative tasks, and develop adaptive EEG feedback with the content-generation capabilities of LLMs to
instructional materials [81]. create an AI tutor that is responsive to a learner’s cognitive state?
At the same time, the introduction of AI-only schools in the A truly adaptive learning system should not only be capable of
United Kingdom and United States [17, 72] demonstrates the grow- generating adaptive content but also of recognizing engagement
ing acceptance of AI-driven adaptive learning platforms. These levels and cognitive load—an ability that is fundamental to effective
systems claim to dynamically adjust educational content to indi- teaching [49]. Pedagogical theories such as Teaching at the Right
vidual learners’ strengths and weaknesses, promising a level of Level (TaRL) [80] and Cognitive Load Theory (CLT) [19] empha-
personalization that traditional classroom settings often struggle to size that cognitive adaptation is crucial for learning success. When
achieve [29]. However, despite this enthusiasm, there remain unre- educational tools present material that exceeds a learner’s cogni-
solved challenges regarding the integration of LLMs into education. tive capacity, learning is significantly hindered [25]. Research has
A key issue is that LLMs lack awareness of a learner’s cognitive shown that maintaining optimal cognitive load enhances knowl-
and attentional state unless explicitly communicated. Without this edge retention and critical thinking skills [87], yet current AI tu-
feedback, LLMs may overwhelm learners with excessive cognitive tors lack mechanisms to assess and adjust for cognitive strain in
load, present material at an inappropriate difficulty level, or fail real time. To address this challenge, we introduce NeuroChat, a
to detect engagement fluctuations, ultimately hindering learning real-time neuroadaptive AI tutor that integrates EEG-based engage-
effectiveness. ment tracking with LLM-driven content generation. NeuroChat
Additionally, concerns about AI-generated hallucinations, data continuously monitors EEG-derived engagement levels and dynam-
privacy, and cognitive offloading—where students may become ically adjusts the depth, complexity, and style of content based on a
overly reliant on AI for information rather than developing in- learner’s cognitive state (Figure 1). This work makes the following
dependent research and critical thinking skills—remain pressing contributions:
issues [45]. While some studies have shown that ChatGPT enhances (1) A novel integration of neuroadaptive learning with genera-
engagement [35] and improves critical thinking skills [23], a meta- tive AI, bridging the gap between EEG-based engagement
analysis of its effects on education suggests that its primary impact tracking and dynamic content generation.
is on academic performance, motivation, and cognitive load, while (2) A closed-loop adaptation mechanism, where real-time EEG
self-efficacy remains largely unaffected [20]. Moreover, AI-driven data informs and modifies the interaction with an LLM
education lacks the human ability to detect subtle emotional and tutor to optimize cognitive load and engagement.
cognitive cues, which are crucial in effective teaching and adaptive (3) High accessibility and usability, achieved by implement-
learning. ing a lightweight, browser-based, wearable system with
Meanwhile, research in neuroadaptive learning systems has dry-electrode EEG headbands, ensuring usability beyond
demonstrated that real-time physiological feedback, particularly laboratory settings.
from electroencephalography (EEG), can significantly enhance en- (4) Empirical evaluation of NeuroChat’s effectiveness, examin-
gagement and cognitive adaptation. EEG-based brain-computer ing how real-time neuroadaptive AI tutoring impacts learn-
interfaces (BCIs) can measure engagement, attention, and cognitive ing outcomes, cognitive engagement, and user experience
load, dynamically adjusting learning materials based on a learner’s compared to non-adaptive AI tutoring.
real-time neural state [51, 62]. Several closed-loop EEG systems
have been developed to optimize learning by modulating content 2 RELATED WORK
complexity based on neurophysiological signals. For example, the
2.1 Engagement in Learning
BRAVO system [54] detects fluctuations in attention and adjusts e-
learning materials accordingly, while Thinking Cap [55] integrates 2.1.1 Defining and Conceptualizing Engagement. Engagement is a
EEG-based cognitive load assessment into an Intelligent Tutoring widely used term in education and psychology, though its definition
System (ITS), modifying text complexity based on engagement lev- varies across disciplines [6]. In educational settings, the term can
els. Similarly, FOCUS [39] adapts learning materials for children be conceptualized as a ‘multidimensional construct encompassing
by integrating EEG-driven interventions during reading sessions. behavioral, emotional, and cognitive dimensions,’ [26] which has
These studies show that adaptive educational environments, when been shown to be linked to positive learning outcomes, including in-
informed by physiological feedback, can improve learning retention creasing student motivation [26, 78]. Expanding on this, Reeve and
and engagement. Tseng [67] introduce a fourth dimension, agentic engagement, to
describe students’ contributions to their learning experience. From
NeuroChat: A Neuroadaptive AI Chatbot for Customizing Learning Experiences

the perspective of flow theory, learner engagement can be enhanced markers include the Cognitive Load Index (Theta Fz / Alpha Pz)
by designing learning activities that promote autonomy and provide [12, 37] and alpha peak frequency [62], which have both been ex-
appropriate challenges to learners’ skill level [75]. Sinatra et al. [76] plored as indicators of cognitive effort and attentional processing
further distinguish between microlevel engagement, which refers efficiency. ERPs, in contrast, offer time-locked neural responses to
to moment-to-moment cognitive focus on a task, and macro-level stimuli, with key components such as P300 (reflecting attentional al-
engagement, which applies to larger social and educational con- location and task relevance), N200 (linked to conflict detection and
texts, such as classrooms or institutions. Micro-level engagement executive control), and error-related negativity (ERN) (indicating
can be assessed through physiological techniques such as brain engagement in performance monitoring) [8, 63, 85].
imaging, skin conductivity, or eye tracking, whereas macrolevel
engagement is typically measured through sociocultural analysis, 2.1.3 EEG-Based Engagement in Learning. EEG-based engagement
observations, or ratings. metrics have been applied in various educational contexts, from
In cognitive neuroscience, engagement is closely linked to sus- providing feedback to presenters [34] to tracking cognitive effort
tained attention and tonic alertness, reflecting a person’s sustained in classroom and workplace environments [30, 33]. EEG has been
cognitive effort [57]. However, engagement extends beyond atten- shown to capture distinct patterns of student attention that dif-
tion, incorporating factors such as intrinsic motivation and task fer from self-reports and teacher observations, offering a more
involvement [44, 68]. Unlike cognitive load, which reflects the men- objective measure of engagement across instructional activities
tal demands on working memory, engagement captures both effort [30]. Studies have also linked higher engagement to better learning
and motivation in a task-driven context. For this study, we define performance. For example, EEG monitoring during video lectures re-
engagement as the sustained allocation of cognitive and attentional vealed significant fluctuations in attentional focus, suggesting that
resources toward a task, influenced by motivation and mental effort. lecture design should account for these variations [18]. Similarly, in
a reasoning task with medical students, engagement correlated with
2.1.2 Physiological Measures of Engagement. Various physiologi- task performance, though the highest engagement was observed
cal technologies have been explored to assess engagement in digital in students who struggled, likely reflecting heightened cognitive
learning environments, including video analysis, eye tracking, and effort despite insufficient expertise. This aligns with Vygotsky’s
biosensors. Classroom video analysis has been used to monitor Zone of Proximal Development, suggesting that while students
student attention, with Raca and Dillenbourg [65] utilizing video were highly engaged, the task was beyond their current skill level,
recordings and later incorporating a computer vision model to leading to cognitive overload rather than effective learning [47].
approximate eye gaze [66]. However, these models have limited EEG has also been explored for cognitive workload classification,
accuracy in estimating attention levels, as they attempt to infer com- with implications for future learning technologies. Andreessen et al.
plex cognitive states from external behavioral approximations that [4] trained an EEG-based model to distinguish high and low mental
are often ambiguous and context-dependent. Eye-tracking systems workload, suggesting its potential for adaptive learning systems
provide a more granular measure of attention shifts and mind- that adjust reading materials based on cognitive load. Similarly, Api-
wandering, but are often costly, complex, and prone to calibration cella et al. [5] demonstrated a low-channel, wearable EEG system
and accuracy issues [40, 41]. for detecting engagement, proposing its use as an input channel for
More direct physiological measures include heart-rate variability adaptive teaching platforms. While these studies focus on monitor-
(HRV) [12], skin conductance (EDA) [11], and electroencephalogra- ing engagement rather than adapting learning in real time, they lay
phy (EEG) [51, 64, 87]. Among these, EEG stands out as the only the groundwork for neuroadaptive systems that dynamically adjust
method that directly measures neural activity, providing real-time instruction based on cognitive states. The next section explores
insights into alertness, attention, and cognitive workload in both how such systems leverage EEG engagement data to personalize
controlled and real-world settings [10, 28]. Since learning is funda- learning experiences.
mentally a neurological process, EEG offers a unique advantage by
capturing dynamic brain responses during information processing, 2.2 Neuroadaptive Learning Systems
making it particularly well-suited for assessing engagement beyond Neuroadaptive systems leverage real-time neurophysiological data,
behavioral proxies. particularly from electroencephalography (EEG), to dynamically
Engagement can be measured using EEG through oscillatory adjust instructional content or interaction modalities based on a
activity (frequency-based markers) and event-related potentials learner’s cognitive and emotional states. These closed-loop systems
(ERPs). Frequency-based markers provide continuous insights into aim to optimize learning outcomes by continuously monitoring
attention and cognitive workload, with alpha power (8–12 Hz) engagement and adapting pedagogical strategies accordingly.
linked to relaxation and disengagement [30, 31], beta power (13–30 Early approaches to neuroadaptive learning focused on adapt-
Hz) associated with sustained attention and active problem-solving ing presentation styles based on user engagement. For instance,
[64], and theta power (4–8 Hz) indicative of fatigue or reduced Pay Attention! [79] employed an embodied storytelling agent that
vigilance [24]. A widely used composite metric is the Engage- adjusted its voice volume and gestures in real time to recapture
ment Index, defined as Beta / (Alpha + Theta), where higher val- students’ attention when EEG signals indicated a drop in engage-
ues indicate greater attentional focus and cognitive engagement ment. This approach significantly enhanced the recall performance
[5, 22, 34, 47, 51, 64]. Alpha asymmetry reflects differences in alpha of students, demonstrating the potential of adaptive presentation to
power between the two brain hemispheres and is often associated influence learning outcomes. Similarly, EngageMeter [34] provided
with approach motivation and active engagement [24, 83]. Other real-time feedback to keynote presenters about their audience’s
Baradari, Kosmyna, Petrov, Kaplun, and Maes

engagement levels, enabling dynamic adjustments in delivery style. dynamic content modulation and interactive adaptation. Early in-
However, while effective in maintaining attention, these systems vestigations propose that integrating LLMs with BCIs could sig-
were limited to modifying delivery methods without altering the nificantly enhance human-computer interaction, benefiting both
learning content itself. individuals with neurological conditions and healthy users [13].
Thinking Cap [55] extends these ideas into an Intelligent Tu-
2.3.1 Using Generative AI to Analyze EEG. A major focus in AI-BCI
toring System (ITS) featuring an animated tutor agent that dy-
research has been EEG-based brain decoding, where generative AI
namically adjusts the complexity of its instructional dialogue with
and machine learning models are used to encode and decode the
the student based on EEG-derived cognitive load measures. This
neural signals underlying visual or auditory information processing
approach ensures that learners are neither underwhelmed nor over-
[7, 32, 84]. While these methods advance neural signal processing,
whelmed. The authors pre-scripted easy and difficult versions of
they remain limited in real-time user interaction. Readers inter-
the instructional content by altering text complexity dimensions
ested in these approaches can refer to a comprehensive review by
such as narrativity, syntactic ease, and referential cohesion. The
Sabharwal and Rama (2024) [70]. Beyond decoding, LLMs have
more recent Online Continuous Adaptation Mechanism (OCAM)
been increasingly applied to EEG for brain state classification and
[22] builds on these principles by continuously monitoring not just
assistive communication [88]. In clinical applications, language
engagement but also concentration, cognitive load, and emotional
model-enhanced BCI communication systems have significantly
arousal to dynamically adjust content difficulty, pacing, and presen-
improved typing accuracy for ALS patients by up to 84% in online
tation style. This system has been shown to significantly increase
BCI spelling sessions [77]. Subsequent approaches have demon-
learner concentration and engagement, highlighting the value of
strated that LLMs can classify brain states at the word level from
multi-dimensional cognitive measures in adaptive learning.
EEG data during reading tasks [36, 89].
Beyond academic learning environments, Learning Piano with
Recent research has extended these applications to foundation
BACh [87] dynamically adapts the difficulty of piano exercises
models that generalize across EEG tasks. NeuroLM [43] and Neuro-
based on cognitive workload (measured via functional near-infrared
GPT [15] function as foundation models, pre-trained on large EEG
spectroscopy), guiding learners into their zone of proximal develop-
datasets using self-supervised learning and task-based fine-tuning
ment to optimize skill acquisition. Closed-loop systems have also
to develop multi-purpose EEG processing models. NeuroLM, trained
been found effective for enhancing learning in perceptual-cognitive
on over 25,000 hours of EEG recordings, aligns brain signals with
tasks, as demonstrated by Parsons et al. [62], who improved perfor-
text-based representations, enabling multi-task analysis in areas
mance by manipulating a 3D multiple object tracking (3D-MOT)
like cognitive workload detection, emotion recognition, and sleep
task through real-time neurofeedback.
staging. Neuro-GPT, trained on the TUH EEG corpus, applies GPT-
Other systems focused on providing real-time biofeedback to
style tokenization to EEG data, improving feature extraction and
help users self-regulate their engagement and attention. AttentivU
adaptability to small datasets. In contrast, EEG-GPT [48] and Lee
[51], for instance, combines EEG headband with haptic feedback
& Chung (2024) [52] focus on task-specific applications—EEG-GPT
devices that vibrate subtly when engagement levels drop, effec-
applies few-shot learning for EEG-based brain state classification,
tively redirecting attention in both online and in-person learning
while Lee & Chung fine-tune GPT-3.5 Turbo for intracranial EEG
contexts. Unlike content-adaptive systems, these approaches rely
(iEEG) interpretation, mapping neural signals to cognitive states.
on external cues to prompt re-engagement rather than altering the
Other approaches have explored personal health and well-being.
learning material itself. Similarly, Joie [83] introduces a joy-based
For instance, Sano et al. (2024) [71] used LLMs to interpret EEG
brain-computer interface (BCI) that uses prefrontal alpha asymme-
signals for sleep quality assessment, providing tailored recommen-
try—an EEG marker linked to positive emotional states—to control
dations. Similarly, EEG Emotion Copilot [14] integrates EEG with a
an endless runner game. By training users to consciously modulate
lightweight (0.5B parameter) LLM to analyze EEG signals, identify
their brain activity through strategies like imagining joyful scenar-
emotional states, and generate automated clinical insights. [38]
ios, Joie highlights the potential of neuroadaptive systems to foster
propose MultiEEG-GPT, a model that integrates EEG with mul-
affective engagement alongside cognitive performance.
timodal data—such as facial expressions and audio—to enhance
Across these systems, a shared limitation is that all content-
mental health assessments using LLM-based classification.
driven systems rely on pre-scripted content that needs to be pre-
Additionally, generative AI techniques have been leveraged for
pared by the researchers to allow for the adaptation. NeuroChat
data augmentation to enhance EEG-based model training [21, 90].
overcomes this barrier by integrating generative AI, which can
However, while these approaches highlight generative AI’s ability
create new content adapted in complexity and presentation style
to process EEG data for individual adaptation, they focus on recog-
to the reader’s cognitive state and specific questions on the fly.
nizing states and have yet to support real-time user interaction.
2.3.2 Artistic Applications Using EEG to Modulate Generative AI
2.3 Generative AI-BCI Systems Outputs. While most research has focused on analyzing EEG data,
The integration of generative AI with brain-computer interfaces a growing field explores EEG as a control mechanism for real-
(BCIs) is an emerging research area. While machine learning has time generative AI adaptation. Early explorations have emerged
long been used to analyze EEG data, most AI-enhanced BCI systems in artistic and creative applications, where EEG signals influence
have focused on brain state classification rather than interactive, AI-generated media production. For example, Imagination Engine
real-time adaptive applications. The introduction of generative AI [1] translates EEG activity into abstract visual art, while Real-Time
expands the possibilities of BCIs beyond passive decoding, enabling Neuro-Augmented Cinema [9] enables cinematic modifications
NeuroChat: A Neuroadaptive AI Chatbot for Customizing Learning Experiences

Figure 2: Left. The Muse 2 EEG system made by InteraXon Inc. Right. Electrode locations of Muse 2 headband according to
10-20 System. CC Teixeria, Gomes, and Brito-Costa (2023).

via neurofeedback. Similarly, the Bio-Mechanical Poet [82] maps data in real time to adjust content complexity, response
real-time EEG signals to symbolic representations, creating immer- style, and pacing, ensuring an adaptive and personalized
sive poetic audiovisual experiences. These projects demonstrate learning experience.
EEG’s potential to actively modulate generative outputs rather than (3) D3: Seamless EEG Integration in Natural User Envi-
merely classifying brain states. However, these applications remain ronments: Given that most users interact with NeuroChat
limited to artistic expression, with little research on EEG-driven on a laptop or desktop computer, the system accommodates
adaptation of linguistic content. The potential to use real-time neu- a stationary, movement-minimizing environment, which
rofeedback to shape AI-generated textual interactions—particularly is ideal for EEG signal acquisition. This design choice min-
in education—remains largely unexplored. imizes motion artifacts, resulting in more reliable neuro-
feedback processing.
2.3.3 Neuroadaptive Generative AI Systems. The most advanced (4) D4: Web-Based Accessibility and Low-Cost Implemen-
neuroadaptive system integrating generative AI for real-time adap- tation: NeuroChat is designed to be fully browser-based,
tation is AdaptiveCoPilot [86], designed for expert pilots in virtual eliminating the need for complex server-side infrastruc-
reality. AdaptiveCoPilot continuously adjusts visual, auditory, and ture and enabling plug-and-play usability. Users can access
textual cues based on real-time cognitive load assessments, optimiz- the system on any platform with minimal setup, making it
ing performance in high-stakes environments. However, it relies scalable and accessible for a broad audience.
on functional near-infrared spectroscopy (fNIRS) rather than EEG
and is tailored for high-performance cognitive tasks rather than
3.1 Interaction Flow
learning applications.
Despite rapid advancements in AI-enhanced BCI research, no NeuroChat integrates real-time EEG data with an LLM chatbot via
existing system has leveraged EEG data to dynamically modulate a web interface to create adaptive, personalized responses. The user
LLM-driven chatbot interactions. This gap underscores the nov- flow is as follows (Figure 3):
elty of NeuroChat, one of the first systems to integrate EEG-based (1) Connection: The user fits on the Muse 2 EEG headband
cognitive state tracking with generative AI in real-time. and connects it to the NeuroChat web app via Web Blue-
tooth for real-time data streaming.
3 SYSTEM DESIGN (2) Calibration: The user completes a 2-minute relaxation task
We set out four core design goals to ensure that NeuroChat is an to determine the engagement minimum (E_min) and a 2-
accessible, responsive, and effective neuroadaptive learning system: minute word association task for the engagement maximum
(E_max). These values are stored in the browser’s session
(1) D1: Wearable, Non-Invasive Brain Sensing: NeuroChat storage for normalization.
employs a consumer-grade, non-invasive EEG headband (3) Interaction: During interaction, the system continuously
to measure engagement in real time. We opted for the 4- computes the normalized engagement score using a 15-
channel Muse EEG headband by InteraXon [60], balancing second sliding window. The last score before the user begins
signal reliability with ease of use. This design ensures that typing is captured and embedded in the query to the chatbot,
users can engage with the system without complex elec- hidden from the user, ensuring that typing doesn’t interfere
trode setups or invasive procedures. with the engagement metric.
(2) D2: Real-Time Adaptive Personalization: To maximize (4) Interactive Response: The query, along with the embed-
learning effectiveness, NeuroChat provides continuous neu- ded engagement score, is sent to the LLM provider, which
rofeedback, dynamically tailoring chatbot responses based returns a response tailored to the user’s cognitive state.
on real-time engagement levels. The system processes EEG
Baradari, Kosmyna, Petrov, Kaplun, and Maes

Figure 3: Overview of the NeuroChat system and user flow. The user connects the Muse headband, undergoes calibration, and
interacts with the neurofeedback-driven LLM. Engagement scores are computed and inserted into the prompt unnoticed by the
user.

3.2 EEG Signal Processing activity typically correspond to lower cognitive engagement, with
3.2.1 Device. Our system uses the Muse 2 EEG headband, building alpha waves linked to relaxation or passive states of rest [24, 30, 31].
on prior research that has leveraged consumer-grade devices with The engagement index has been widely validated across various
1 to 6 channels to assess cognitive engagement in learning contexts applications, including cognitive load assessments [27], visual pro-
(e.g., [34, 51, 79, 83]). The Muse 2 samples at 256 Hz and includes cessing studies, and sustained attention tasks [10]. It has also been
electrodes at Fpz, AF7, AF8, TP9, and TP10, following the 10-20 applied in complex task environments such as the multi-attribute
System (Figure 2) [42]. The Fpz electrode serves as the reference. task battery (MATB) [64], which involves tasks like tracking, re-
EEG data is streamed to a web browser using the open-source source management, and communication. These studies demon-
MuseJS library [74], which enables real-time streaming via Web strate the engagement index’s effectiveness in detecting attention
Bluetooth. shifts and fluctuations in cognitive state triggered by external stim-
uli [3, 16].
3.2.2 Preprocessing. The EEG data processing pipeline follows Following our preprocessing pipeline, we extract frequency bands
established methods from Hassib et al. [34], Kosmyna and Maes for each epoch and average them over a 15-second sliding window,
[51], Szafir and Mutlu [79] and others. A bandpass filter (1–30 as established by Szafir and Mutlu [79]. Averaging over a time win-
Hz) is applied to retain relevant neural activity while minimizing dow allows us to assess a user’s engagement over a meaningful
noise, and a 60 Hz notch filter removes power line interference. The duration while they read and process the LLM’s output, rather than
data is then segmented into 1-second epochs with 250 ms intervals capturing momentary fluctuations. We selected a 15-second win-
to enable continuous analysis with sufficient temporal resolution. dow to account for variations in reading speed, ensuring sufficient
Power spectral density is computed via fast Fourier transform (FFT), time for users to engage with the response. Unlike previous studies,
and band power is extracted for each frequency range to derive we opted against exponentially weighted moving averages, as our
meaningful neural features. focus is on sustained engagement throughout a task rather than
transient cognitive spikes.
3.2.3 Engagement Score. The engagement index (or engagement Finally, we normalize the engagement score following Kosmyna
score) serves as the core metric of our system, enabling real-time and Maes [51]. Normalization requires determining a minimum
quantification of cognitive engagement during mentally demanding and maximum engagement score for each user, which we obtain
tasks. First introduced by Pope et al. [64], this metric is computed from the calibration task conducted before the main experiment.
as a ratio of key EEG frequency bands using the formula: During calibration, users engaged in two tasks, each lasting two
minutes:

𝐸=
𝛽
(1) (1) Relaxation: Participants remain still, minimizing cognitive
𝛼 +𝜃 effort while we record baseline EEG data.
where 𝛽 (11–20 Hz), 𝛼 (7–11 Hz), and 𝜃 (4–7 Hz) correspond to (2) Mental word association: Participants perform a cognitive
EEG-derived neural oscillations. The index is based on the principle task that requires generating words based on the final let-
that higher beta power reflects heightened brain activity during ter of the previous word (e.g., "elephant" → "tiger"). This
cognitive tasks [11]. The beta frequency band is particularly as- method has been shown to effectively induce cognitive
sociated with cognitive processes such as visual attention, motor activation in non-ALS participants [50].
planning, and active information processing, all of which indicate The lowest and highest engagement scores from the two tasks,
an engaged mental state. Conversely, increased alpha and theta respectively, are taken as the normalization minimum 𝐸 min and
NeuroChat: A Neuroadaptive AI Chatbot for Customizing Learning Experiences

Figure 4: NeuroChat user interface with exposed EEG metrics in the user prompt and experimenter control menu. The
connection to the Muse EEG device happens through the Brain Widget in the top right corner. “Mood mode” activates the
EEG metric injection into the user’s prompts, and turning off “Debug mode” allows the experimenter to hide these from the
user. Chats, raw and filtered EEG data, and computed EEG metrics from the Muse device are stored in the browser’s native
IndexedDB and can be exported from the Settings panel.

maximum 𝐸 max . The calibration task differs from the main task in scientific depth, response format (e.g., bullet points versus long-
that it uses a 10-second sliding window, balancing the 5-second form text), and use of Socratic questioning. Additionally, framing
interval used in prior studies [34, 51, 79] and the 15-second window its role as a “good tutor” improved response quality.
applied in our main experiment. The normalized engagement score
𝐸 norm is then calculated using: 3.4 User Interface
The NeuroChat user interface (UI) consists of four key components
𝐸 − 𝐸 min (Figure 4):
𝐸 norm = (2)
𝐸 max − 𝐸 min
(1) Brain Connect Widget (Top Right Corner): Allows users
where E represents the engagement score averaged over the past to connect or disconnect the EEG device, calibrate or re-
15 seconds. calibrate the system, and start or stop EEG recording to
compute the engagement index.
3.3 LLM Adaptation (2) Calibration Modal (Full Screen) – Appears only after the
The mechanism through which NeuroChat responds adaptively to Muse headset is connected and provides instructions for a
the user’s cognitive state is by embedding their engagement score 2-minute relaxation phase followed by a 2-minute mental
into each query submitted to a Large Language Model. We used word association task. If the EEG connection is lost, users
OpenAI’s GPT-4-turbo model, the latest at the point of study. can restart or resume from the completed relaxation phase.
The system prompt provides a guideline as to how the LLM (3) Chatbot Interface (Main Screen Area) – Functions simi-
should adapt its response style to the user. We developed it through larly to ChatGPT, displaying an alternating conversation
careful evaluation of pilot testing and insights from OpenAI’s Teach- between the user and the AI tutor.
ing with AI guide [59] (see Appendix A.1). One key insight was (4) Menu Sidebar (Left-Hand Side) – Contains chat history
that prompting GPT-4 to increase the user’s “engagement” often re- where users can manage past LLM conversations by cre-
sulted in overly casual, upbeat responses, as the model interpreted ating folders, renaming chat titles, and deleting individual
the term informally. Reframing the engagement score as a “cogni- chats. Also features the Settings panel, which provides op-
tive load metric” helped maintain a neutral tone while allowing tions to toggle “Mood Mode” (enabling LLM adaptation),
responses to adjust dynamically based on neurofeedback. Based on activate “Debug Mode” (hiding EEG metrics from the UI),
the engagement index, the LLM was instructed to modulate detail, import/export chat history, download EEG data (from the
Baradari, Kosmyna, Petrov, Kaplun, and Maes

Figure 5: Overview of study procedure (not to scale).

browser’s IndexedDB), switch between dark and light mode, while still remaining accessible for participants. Condition and
and reset the chat history for a new user session. The terms study topic order were counterbalanced using a Latin square design.
“Mood Mode” and “Debug Mode” were intentionally chosen Before the session, participants signed a consent form and turned
to provide visual cues to the experimenters while being off their electronic devices. They were fitted with a Muse EEG
vague enough to the participants. headband, and signal quality was verified via the Muse EEG app
[112]. Participants were instructed to minimize movement to reduce
4 METHODOLOGY motion artifacts.
The study lasted about 2 hours and proceeded as follows (Fig-
4.1 Hypotheses
ure 5):
Based on prior research, we formulate the following hypotheses:
(1) Pre-Session Measures: Participants completed a background
• (H1) Objective Engagement: NeuroChat will elicit higher questionnaire assessing their alertness and previous ex-
engagement levels than interaction with a standard GPT perience with AI chatbots. A brief EEG calibration phase
model, as measured by EEG-derived engagement scores. followed (2 minutes relaxation, 2 minutes mental exercise).
• (H2) Subjective Engagement: Participants will report greater (2) AI Chatbot Interaction: Participants engaged with the chat-
subjective engagement and satisfaction with NeuroChat, bot for 20 minutes on their first assigned topic, with the
perceiving it as more engaging and effective than a tradi- goal of “learning as much as possible.” To guide exploration,
tional AI tutoring model. they received starting pointers—e.g., characteristics, behav-
• (H3) Learning Outcomes: Participants using NeuroChat will ior, and archaeological research for T. rex and historical
achieve higher scores on post-interaction learning assess- context, significance, and consequences for the Taiping
ments compared to those using the standard GPT model. Rebellion. Participants were free to focus on aspects they
found interesting.
4.2 Participants (3) Knowledge Assessment: Immediately after the chatbot inter-
Thirty participants (15 female, 13 male, 2 other), predominantly action, participants completed a quiz consisting of fill-in-
from academic backgrounds, were recruited for this study (M = 32.4 the-blank and multiple-choice (MCQ) questions, followed
years, median = 30) and compensated with a $50 Amazon gift card. by a 15-minute essay to assess understanding. To prevent
The study received approval from MIT’s institute’s ethical review preparatory bias, participants were not informed about the
board (protocol no. 21070000428). quiz beforehand. The same quiz was used across conditions.
(4) Break & Condition Switch: Participants took a short break
4.3 Study Design and Protocol before repeating the process with the second topic and
We adopted a within-subject study design after pilot studies re- condition. EEG data and chatbot interaction logs were con-
vealed significant individual differences in interactions with the AI tinuously recorded.
chatbot. This design allowed each participant to serve as their own (5) Final Survey & Interview: Participants completed a post-
control, minimizing variability and enabling direct performance study user survey and a semi-structured interview focusing
comparisons between the NeuroChat experimental condition and on their subjective engagement and experience across con-
the control condition. The control condition consisted of a regular ditions. Interviews were thematically analyzed.
GPT chatbot, which was prompted to act within an AI tutoring task
via its system prompt for fair comparison (see Appendix A.2). 4.4 Evaluation
As study topics, we selected the Tyrannosaurus rex (T. rex) and the Assessing learning outcomes requires a multifaceted approach. The
Taiping Rebellion. These topics were chosen to minimize prior topic quizzes incorporated recall-based and synthesis-based questions to
bias while allowing room for facts and explorative interpretation. capture different cognitive processes. Recall was assessed through
Although the T. rex is widely recognized, most people lack in- multiple-choice (MCQ) and fill-in-the-blank questions, while cre-
depth knowledge about the dinosaur. Similarly, despite its historical ative synthesis was evaluated via mini-essays requiring critical
significance, the Taiping Rebellion is rarely emphasized in Western thinking and analysis. MCQ and fill-in-the-blank questions were
education. Both topics provided sufficient complexity and depth designed to assess factual recall, covering information likely en-
for meaningful engagement within the 20-minute learning session countered during topic exploration. Question complexity varied to
NeuroChat: A Neuroadaptive AI Chatbot for Customizing Learning Experiences

Table 1: Model summary and model fit statistics.

(a) Model summary for Engagement (normalized) by Condition, accounting for Order. (b) Model fit statistics.

Predictor 𝛽 SE z-value p-value 95% CI Statistic Value


Intercept -0.382 0.171 -2.230 0.026 [-0.718, 0.046] Log-Likelihood -19.621
Condition (E) 0.216 0.099 2.185 0.029 [0.022, 0.410] AIC 47.242
Order 0.181 0.099 1.828 0.068 [-0.013, 0.375] BIC 54.727

reflect a range of difficulty levels appropriate for participants with high between-subject variability (intra-class correlation coefficient
minimal prior knowledge engaging for under 20 minutes. For the (ICC) = 36.5%, p < 0.05). Therefore, to account for individual baseline
essay task, participants had 15 minutes to write and could choose differences while preserving within-subject variability, we applied
from a set of prompts or create their own. z-score normalization, adjusting each participant’s engagement
One author manually scored the quiz blind. Fill-in-the-blank scores based on their mean and standard deviation across both con-
questions were graded with 2 points each, with partial scores for ditions. This allowed for direct comparison of relative engagement
semi-correct answers, and multiple choice questions (MCQ) were differences between NeuroChat and the control condition.
given 1 point per correct option. Since we only wanted to grade Given the repeated-measures design, where each participant con-
responses that had come up in the chat interaction, an automated tributed data across multiple conditions, and occasionally missing
keyword detection script scanned a participant’s message history data, we fit a Linear Mixed Model (LMM) to examine the effect of
for the presence or absence of a question, which was checked man- the experimental condition (Condition) on normalized engagement
ually. Answers not covered were excluded from the participant’s (𝐸 norm ) (Table 1). For such conditions, an LMM is more appropriate
total score, leaving us with proportional participant scores for com- than a paired t-test, as it accounts for within-subject correlations
parison. and random variability across participants. We used the statsmodels
The 15-minute mini-essays were graded blind by a professional package in Python [73].
high school English teacher on a 5-point scale across 4 categories: An initial model with Condition as the sole fixed effect did not
Content, Structure and Organization, Language and Style, and Ac- yield a statistically significant relationship with engagement (𝛽 =
curacy (Spelling and Grammar). The Content category was given 0.186, p = 0.063). Since task order could introduce confounding ef-
double weighting when calculating the final score. fects, we extended the model by adding Order as an additional fixed
In addition to objective assessments, participants completed a effect. This refined model revealed a significant effect of Condition
post-study user survey and were interviewed by 1 of 3 of the au- (𝛽 = 0.186 = 0.216, p = 0.029) and a marginal effect of Order (𝛽 =
thors in a semi-structured interview lasting no more than 5 minutes. 0.186 = 0.181, p = 0.068) on normalized engagement (Figure 6). This
The interviewers took note of key quotes and sentiments and sub- finding suggests that task order may have influenced engagement
sequently cross-read each other’s notes and discussed additional levels, warranting its inclusion in the final model. As expected,
takeaways. Survey responses and interview notes were compiled the random effect variance for participants was low (𝜎2 = 0.001),
and thematically analyzed by the first author over multiple rounds reflecting the impact of z-score normalization, which minimized
of descriptive coding. inter-individual differences before running the model, leaving only
within-condition variation.
5 RESULTS Extending the model to test for additional effects of study topic,
age, education level, chatbot experience, chatbot familiarity, chat-
To evaluate the effects of NeuroChat on engagement and learning
bot usage frequency, and prompt engineering skill revealed no
outcomes, we conducted analyses on EEG engagement scores, learn-
significant influence of these factors (p > 0.1). In conclusion, when
ing assessments, and user feedback. Our results address three key
accounting for order effects during the experiment, participants in
areas: (1) cognitive engagement (EEG-derived engagement index),
the NeuroChat condition were, on average, relatively more cogni-
(2) user-reported engagement, and (3) learning performance (quiz
tively engaged.
and essay scores).

5.1 Cognitive Engagement


Since the system processes EEG data in real time, engagement
scores were computed dynamically during each session. Six partici- 5.2 Learning Test Performance
pants were excluded due to missing or poor-quality EEG signals, To assess whether NeuroChat improves learning outcomes (H3),
and one additional participant was removed due to non-compliance we compared participants’ quiz and essay performance. For the
with instructions. This resulted in 24 participants for analysis. For quiz, mean proportional scores showed no significant difference
preprocessing, we removed missing values and extreme outliers between conditions (E = 61.02%, C = 60.66%). Likewise, for the essay,
(beyond 3× standard deviation) and manually inspected the engage- participants in the experimental condition averaged 18.27 points,
ment data, excluding segments with signal disconnections. Despite compared to 17.69 points in the control condition, indicating no
participant-level calibration, engagement scores exhibited relatively notable difference between groups.
Baradari, Kosmyna, Petrov, Kaplun, and Maes

Figure 6: Distribution of engagement score means in the control and experimental conditions by order (right: z-score normalized).

5.3 Perceived Engagemenet & Subjective more factual and concise responses. They felt that NeuroChat’s con-
Evaluations versational style detracted from the focus on information, making
it harder to digest the content. Similarly, P26 found the control’s
We analyzed user responses based on the post-questionnaires and
more nuanced, fact-driven responses preferable, describing Neu-
informal interviews regarding reported levels of engagement and
roChat’s conversational prompts as distracting rather than helpful.
learning preferences. Participants were asked in writing and ver-
Despite this, many participants noted that the control chatbot often
bally about their perceived engagement, noticeable differences, and
lacked the personal touch, with P31 describing the control chatbot
learning preferences between the chats. We categorized their feed-
as feeling like “a regular chatbot” that lacked awareness of the
back into five main themes: personalized feedback, and response
user’s emotions or engagement level.
style (factual vs. conversational), density of information, follow-
up questions, and additional feedback, each contributing to the 5.3.2 Factual vs. Casual Response Style. The response style be-
perceived engagement and satisfaction. tween NeuroChat and the control chatbot was another point of
divergence in subjective feedback. Participants like P28 and P30
enjoyed the more conversational and engaging tone of NeuroChat.
5.3.1 Personalized Feedback and Human-Like Responses. A promi- P28 described the experimental chatbot as “very fun, like a tour
nent theme in the subjective evaluations was NeuroChat’s more guide,” with a more interactive and fluid exchange, while the control
human-like responses, which many participants found engaging. felt “like a textbook” in comparison. These participants appreci-
P1 noted that the chatbot “mimics a real person” and provided ated NeuroChat’s ability to dive deeper into topics, making the
feedback that made it seem more interactive and lifelike, such as learning experience more dynamic and enjoyable. In contrast, some
saying “Great question” after user input. P31 also expressed sat- participants preferred the control chatbot’s more formal, factual
isfaction with the experimental chatbot, stating, “Oh, I loved the style. P7 found that the control condition allowed for more focused
second one! I really liked how it was saying how I was feeling.” learning, noting that NeuroChat was more prone to casual conver-
This participant emphasized the importance of NeuroChat’s ability sation that made it harder to focus on the core information. P17
to provide feedback tailored to their emotional state, suggesting reflected that while the experimental chatbot was more enjoyable
that it was responding to affect and engagement levels in real time. due to its fluidity and tendency to present fun facts, the control chat-
P19 added that NeuroChat had “more personality” compared to bot’s responses were better structured and felt more educational,
the standard GPT, making the interaction feel more tutor-like and comparing the control to a “blog post” with dense information.
conversational, rather than merely factual. However, some participants also criticized the control chatbot for
Furthermore, NeuroChat’s personalized prompts made the ex- being too rigid and not encouraging exploration. P32 mentioned
perience feel more responsive and adaptive for some participants. that while the control condition was more “analytical,” it felt more
P30 described the experimental chatbot as “always responding to like attending “a serious lecture,” with little room for the more
my prompt,” and noted how it felt more dynamic than the control enjoyable, exploratory exchanges that NeuroChat provided. P19
chatbot, which often came across as rigid and formal. Similarly, similarly remarked that the control chatbot seemed “less eager to
P28 enjoyed the depth of engagement, stating that NeuroChat al- engage,” contributing to a less immersive and personalized learning
lowed for “more meaningful topics to ask,” creating a richer, more experience.
exploratory interaction.
However, some participants preferred the more straightforward 5.3.3 Density of Information. The verbosity of NeuroChat’s re-
approach of the control chatbot. P7 and P10, who identified them- sponses was a double-edged sword for participants. Some, like P14
selves as “scientific minds,” favored the control condition for its and P12, appreciated the deeper exploration of topics provided by
NeuroChat: A Neuroadaptive AI Chatbot for Customizing Learning Experiences

NeuroChat, which allowed them to learn more than the control (10 participants). This gives us supporting evidence that partic-
chatbot. P14 commented that while they would choose the con- ipants using NeuroChat report higher levels of engagement and
trol chatbot for quick learning, NeuroChat was more suitable for satisfaction than those using a chatbot without neurofeedback (H2).
in-depth exploration due to its comprehensive answers. Similarly,
P28 noted that the experimental chatbot fostered a more nuanced
6 DISCUSSION
understanding, allowing them to focus on the context and reasons
behind a topic rather than just absorbing arbitrary facts. However, This study examined whether NeuroChat, a neuroadaptive AI chat-
other participants found the sheer volume of information over- bot, enhances engagement and learning outcomes by adapting its
whelming. P27 described NeuroChat’s responses as “paragraph responses based on real-time EEG feedback. Our results confirm
after paragraph,” and P32 admitted to skipping parts of its ver- that NeuroChat successfully increased cognitive (EEG-measured)
bose answers, opting instead for the control’s more concise and and self-reported engagement, demonstrating its feasibility as a
digestible responses. P19 and P33 both noted that NeuroChat had a neuroadaptive tutoring system. However, no significant differences
tendency to provide redundant information, which diminished the were found in learning performance, indicating challenges in trans-
clarity and relevance of the responses over time. While the control lating engagement into measurable knowledge gains. Below, we
chatbot was favored for its brevity, some participants like P30 found discuss these findings in the broader context of adaptive learn-
that it occasionally oversimplified complex topics, limiting deeper ing, generative AI, and brain-computer interfaces (BCIs) before
understanding. exploring key challenges and future directions.

5.3.4 Follow-Up Questions. Feedback on the follow-up questions 6.1 Key Findings: Engagement Gains, Learning
varied widely between the NeuroChat and the control condition. Outcomes, and User Perception
Participants appreciated the specific nature of the questions in Neu- We found that NeuroChat significantly increased EEG-measured
roChat. P2 stated, "I liked the questions at the end, they were really engagement (p = 0.029), suggesting that real-time neuroadaptive
specific—much better than Copilot." This sentiment was echoed by feedback can enhance sustained attention. These findings align
P23, who found the prompts helpful, especially since they “didn’t with prior neuroadaptive systems like Pay Attention! [79] and En-
know anything about the topic,” and appreciated the guidance. gageMeter [34], which successfully modulated engagement through
However, not all participants found the follow-up questions use- adaptive presentation techniques. However, NeuroChat goes fur-
ful. P18 felt that while they were prompted with questions, the ther by modifying conversational flow, response complexity, and
responses didn’t lead anywhere meaningful, causing frustration. interaction depth, positioning it as an interactive, closed-loop sys-
P26 provided a particularly nuanced critique, noting that the experi- tem rather than a passive monitoring tool.
mental chatbot’s prompts felt superficial and it didn’t seem to “care” User feedback also reflected higher perceived engagement with
the way a human would. They felt that the questions prompted by NeuroChat, with many participants finding its responses more
NeuroChat often missed their actual interests, leading to a sense human-like, responsive, and personalized. However, individual dif-
of disconnection from the conversation. P33 also noted that Neu- ferences emerged—some users preferred a factual, concise style,
roChat’s tendency to drive the conversation in a specific direction while others favored conversational, exploratory interactions. This
was problematic, as they had to repeatedly bring it back to their suggests that adaptive tutoring must account for personalized learn-
original question, which disrupted the flow of engagement. ing preferences to be effective.
In contrast, participants found the control chatbot’s lack of Despite increased engagement, NeuroChat did not significantly
follow-up questions to be limiting in terms of engagement. P19 improve quiz or essay scores. This result mirrors previous studies
noted that while the control was concise, it didn’t prompt any on learning with LLMs [53], which found that while personalization
follow-up questions, making the conversation feel more transac- enhances motivation, it does not always yield better performance
tional and less interactive. This lack of conversational depth in the on traditional assessments. Possible explanations include:
control condition was also mentioned by P23, who described it as
having “more general, broad questions,” which felt less engaging (1) Engagement ≠ Effective Learning – EEG-based engagement
than the more creative, tailored prompts offered by NeuroChat. captures sustained attention, but not necessarily deep learn-
ing or knowledge retention.
(2) Task Design Limitations – Unlike structured adaptive sys-
5.3.5 Overall Engagement and Satisfaction. In summary, subjective tems like BACh [87] (which progressively increased pi-
feedback on NeuroChat’s engagement and satisfaction levels varied, ano sheet music difficulty), NeuroChat allowed open-ended
with the conversational and human-like elements appealing to par- learning, making structured difficulty progression harder
ticipants seeking a more interactive, engaging experience. However, to implement.
those who preferred straightforward, fact-focused learning found (3) Short Study Duration – Neuroadaptive learning benefits
the control chatbot more aligned with their needs. NeuroChat’s may emerge over multiple sessions, but our study measured
neurofeedback-driven prompts were effective for some, but oth- learning in a single interaction.
ers found them intrusive or misaligned with their interests, which
could detract from overall satisfaction. Overall, more participants Future work should explore long-term retention, conceptual un-
provided positive feedback on their engagement in their experimen- derstanding, and scaffolding techniques that could better translate
tal condition (23 participants) compared to the control condition engagement gains into measurable learning improvements.
Baradari, Kosmyna, Petrov, Kaplun, and Maes

6.2 Implications for Neuroadaptive Learning & mechanisms, personalize content, and balance engagement with
AI-Powered Tutoring cognitive load. By integrating multimodal sensing and long-term
user modeling, AI tutors could one day provide truly personalized,
Traditional neuroadaptive learning systems relied on pre-scripted
lifelong learning experiences.
content, where researchers manually assigned learning materials
to high- or low-engagement conditions. NeuroChat overcomes this
limitation by leveraging generative AI to create content dynam-
ACKNOWLEDGMENTS
ically, enabling real-time adaptation tailored to individual users. We thank Treyden Chiaravalloti for his valuable piloting support
This is a fundamental shift in adaptive learning—moving from rule- and insightful feedback. We also appreciate Protyasha Nishat’s
based, pre-mapped content to generative, personalized tutoring. expertise in signal processing and Nathan Whitmore’s comments
Most LLMs require users to explicitly communicate their needs on study design. Luisa Heiss’s thorough grading of the essays was
(e.g., "Explain this differently", "Make it simpler"). NeuroChat re- instrumental in evaluating participant test performance without
moves this barrier by inferring user engagement levels directly from bias. This research was supported by the MIT J-WEL Education
EEG data, reducing the need for manual prompt engineering. This Innovation Grant.
has implications for personalized AI assistants, where cognitive
state tracking could enhance adaptability without user effort. REFERENCES
NeuroChat has particular relevance for self-directed learners, [1] 2023. Imagination Engine I: Generating Abstract Art through EEG.
[Link]
who often struggle with maintaining engagement. In 2021, over 220 abstract-art-through-eeg.
million students enrolled in MOOCs, yet the average completion [2] 2023. Technology in education. Technical Report. UNESCO.
rate remains at 13% [2, 58]. A neuroadaptive AI tutor could help [3] Yomna Abdelrahman, Mariam Hassib, Maria Guinea Marquez, Markus Funk,
and Albrecht Schmidt. 2015. Implicit engagement detection for interactive
sustain motivation and prevent disengagement, particularly in open- museums using brain-computer interfaces. In Proceedings of the 17th International
ended, autonomous learning environments. Conference on Human-Computer Interaction with Mobile Devices and Services
Adjunct. ACM. doi:10.1145/2786567.2793709
Beyond education, NeuroChat’s EEG-driven AI system could [4] Lena M Andreessen, Peter Gerjets, Detmar Meurers, and Thorsten O Zander.
support knowledge workers, particularly those who struggle with 2021. Toward neuroadaptive support technologies for improving digital reading:
focus and information retention. Additionally, LLM-based BCIs a passive BCI-based assessment of mental workload imposed by text difficulty
and presentation speed during reading. User Model. User-adapt Interact. 31 (March
have been proposed for aiding individuals with learning challenges, 2021), 75–104. doi:10.1007/s11257-020-09273-5
including ADHD [13, 39, 62]. These applications highlight the po- [5] Andrea Apicella, Pasquale Arpaia, Mirco Frosolone, Giovanni Improta, Nicola
tential of neuroadaptive AI beyond the classroom. Moccaldi, and Andrea Pollastro. 2022. EEG-based measurement system for
monitoring student engagement in learning 4.0. Sci. Rep. 12 (7 April 2022), 5857.
doi:10.1038/s41598-022-09578-y
6.3 Challenges & Limitations of NeuroChat [6] Roger Azevedo. 2015. Defining and measuring engagement and learning in
science: Conceptual, theoretical, methodological, and analytical issues. Educ.
Participants showed high variability in engagement and preference Psychol. 50 (2 Jan. 2015), 84–94. doi:10.1080/00461520.2015.1004069
for different interaction styles. Some learners thrived in guided, ex- [7] Yunpeng Bai, Xintao Wang, Yan-Pei Cao, Yixiao Ge, Chun Yuan, and Ying Shan.
2025. DreamDiffusion: High-quality EEG-to-image generation with temporal
ploratory conversations, while others preferred concise, fact-driven masked signal modeling and CLIP alignment. In Lecture Notes in Computer Science.
responses. Future systems should incorporate user preference set- Springer Nature Switzerland, 472–488. doi:10.1007/978-3-031-72751-1_27
[8] N P Bechtereva and V B Gretchin. 1968. Physiological foundations of mental
tings, such as: activity. Int. Rev. Neurobiol. 11 (1968), 329–352. doi:10.1016/s0074-7742(08)60392-
• Preferred interaction style (e.g., structured vs. exploratory). x
[9] Antoine Bellemare-Pepin, Philipp Thölke, Yann Harel, and Karim Jerbi. 2024.
• Response format (e.g., bullet points vs. narratives). Real-Time Neuro-Augmented Cinema via Generative AI. NeurIPS Workshop on
• Memory-based personalization (as seen in OpenAIś user Creativity & Generative AI.
[10] C Berka, D J Levendowski, M N Lumicao, A Yau, G Davis, V T Zivkovic, R E Olm-
memory feature [61]). stead, P D Tremoulet, and P L Craven. 2007. EEG correlates of task engagement
and mental workload in vigilance, learning, and memory tasks. Aviation, space,
Moreover, high engagement is not always beneficial. Medical and environmental medicine 78 (May 2007).
reasoning studies [22] found that struggling learners showed the [11] Wolfram Boucsein, Andrea Haarmann, and Florian Schaefer. 2007. Combining
highest engagement, suggesting that increased engagement can skin conductance and heart rate variability for adaptive automation during sim-
ulated IFR flight. In Engineering Psychology and Cognitive Ergonomics. Springer
sometimes signal cognitive overload rather than productive learn- Berlin Heidelberg, 639–647. doi:10.1007/978-3-540-73331-7_70
ing. Future neuroadaptive tutors must ensure users remain in their [12] E A Byrne and R Parasuraman. 1996. Psychophysiology and adaptive automation.
zone of proximal development rather than pushing them beyond Biol. Psychol. 42 (5 Feb. 1996), 249–268. doi:10.1016/0301-0511(95)05161-9
[13] Andrea Caria. 2024. Towards predictive communication with brain-computer
their capabilities. interfaces integrating large language models. arXiv [[Link]] (10 Dec. 2024).
Consumer EEG devices are fundamentally noisy, and EEG signals [14] Hongyu Chen, Weiming Zeng, Chengcheng Chen, Luhui Cai, Fei Wang, Yuhu
Shi, Lei Wang, Wei Zhang, Yueyang Li, Hongjie Yan, Wai Ting Siok, and Nizhuan
contain biometric markers that can uniquely identify individuals Wang. 2024. EEG Emotion Copilot: Optimizing lightweight LLMs for emo-
[69]. As LLMs process data externally, privacy concerns must be tional EEG interpretation with assisted medical record generation. arXiv [[Link]]
addressed before large-scale adoption of EEG-driven AI tutors. (30 Sept. 2024).
[15] Wenhui Cui, Woojae Jeong, Philipp Thölke, Takfarinas Medani, Karim Jerbi,
Anand A Joshi, and Richard M Leahy. 2024. Neuro-GPT: Towards A foundation
7 CONCLUSION model for EEG. In 2024 IEEE International Symposium on Biomedical Imaging
(ISBI), Vol. 35. IEEE, 1–5. doi:10.1109/isbi56570.2024.10635453
We provide evidence that EEG-driven AI chatbots can enhance [16] Alex Dan and Miriam Reiner. 2017. Real time EEG based measurements of
engagement, bringing users closer into a zone of proximal develop- cognitive load indicates mental states during learning. JEDM 9 (23 Dec. 2017),
31–44. doi:10.5281/ZENODO.3554719
ment, but highlight challenges in translating engagement into learn- [17] David Game College. 2024. GCSE AI Adaptive Learning Programme.
ing gains. Future neuroadaptive AI systems must refine adaptation [Link]
NeuroChat: A Neuroadaptive AI Chatbot for Customizing Learning Experiences

ai-adaptive-learning-programme. International Joint Conference on Pervasive and Ubiquitous Computing. ACM,


[18] Ido Davidesco, Noah Glaser, Ian H Stevenson, and Or Dagan. 2023. Detecting 412–417. doi:10.1145/3675094.3678494
fluctuations in student engagement and retention during video lectures using [39] Jin Huang, Chun Yu, Yuntao Wang, Yuhang Zhao, Siqi Liu, Chou Mo, Jie Liu, Lie
electroencephalography. Br. J. Educ. Technol. 54 (Nov. 2023), 1895–1916. doi:10. Zhang, and Yuanchun Shi. 2014. FOCUS: enhancing children’s engagement in
1111/bjet.13330 reading by using contextual BCI training sessions. In Proceedings of the SIGCHI
[19] Ton de Jong. 2010. Cognitive load theory, educational research, and instructional Conference on Human Factors in Computing Systems. ACM. doi:10.1145/2556288.
design: some food for thought. Instr. Sci. 38 (March 2010), 105–134. doi:10.1007/ 2557339
s11251-009-9110-0 [40] Stephen Hutt, Kristina Krasich, James R. Brockmole, and Sidney K. D’Mello. 2021.
[20] Ruiqi Deng, Maoli Jiang, Xinlu Yu, Yuyan Lu, and Shasha Liu. 2025. Does Breaking out of the lab: Mitigating mind wandering with gaze-based attention-
ChatGPT enhance student learning? A systematic review and meta-analysis of aware technology in classrooms. In Proceedings of the 2021 CHI Conference on
experimental studies. Comput. Educ. 227 (1 April 2025), 105224. doi:10.1016/j. Human Factors in Computing Systems. ACM. doi:10.1145/3411764.3445269
compedu.2024.105224 [41] Stephen Hutt, Caitlin Mills, Nigel Bosch, Kristina Krasich, James Brockmole, and
[21] Seif Eldawlatly. 2024. On the role of generative artificial intelligence in the Sidney D’Mello. 2017. Out of the fr-eye-ing pan: Towards gaze-based models
development of brain-computer interfaces. BMC Biomed. Eng. 6 (2 May 2024), 4. of attention during learning with technology in the classroom. In Proceedings
doi:10.1186/s42490-024-00080-2 of the 25th Conference on User Modeling, Adaptation and Personalization. ACM.
[22] Atef Eldenfria and Hosam Al-Samarraie. 2019. Towards an online continuous doi:10.1145/3079628.3079669
adaptation mechanism (OCAM) for enhanced engagement: An EEG study. Int. J. [42] InteraXon. 2025. Muse: the brain sensing headband Store with Worldwide
Hum. Comput. Interact. 35 (14 Dec. 2019), 1960–1974. doi:10.1080/10447318.2019. Shipping. [Link]
1595303 [43] Wei-Bang Jiang, Yansen Wang, Bao-Liang Lu, and Dongsheng Li. 2024. NeuroLM:
[23] Harry Barton Essel, Dimitrios Vlachopoulos, Albert Benjamin Essuman, and A universal multi-task foundation model for bridging the gap between language
John Opuni Amankwa. 2024. ChatGPT effects on cognitive skills of undergradu- and EEG signals. arXiv [[Link]] (27 Aug. 2024).
ate students: Receiving instant responses from AI-based conversational large [44] A Kamzanova, G Matthews, A Kustubayeva, and S Jakupov. 2011. EEG indices
language models (LLMs). Computers and Education: Artificial Intelligence 6 (1 June to time-on-task effects and to a workload manipulation (cueing). International
2024), 100198. doi:10.1016/[Link].2023.100198 Scholarly and Scientific Research & Innovation 5 (23 Aug. 2011), 928–931.
[24] Stephen H Fairclough, Liverpool John Moores, Katie C Ewing, and Jenna Roberts. [45] Enkelejda Kasneci, Kathrin Sessler, Stefan Küchemann, Maria Bannert, Daryna
2009. Measuring task engagement as an input to physiological computing. In 2009 Dementieva, Frank Fischer, Urs Gasser, Georg Groh, Stephan Günnemann,
3rd International Conference on Affective Computing and Intelligent Interaction Eyke Hüllermeier, Stephan Krusche, Gitta Kutyniok, Tilman Michaeli, Clau-
and Workshops. IEEE. doi:10.1109/acii.2009.5349483 dia Nerdel, Jürgen Pfeffer, Oleksandra Poquet, Michael Sailer, Albrecht Schmidt,
[25] Shalom M Fisch. 2017. Bridging theory and practice: Applying cognitive and Tina Seidel, Matthias Stadler, Jochen Weller, Jochen Kuhn, and Gjergji Kas-
educational theory to the design of educational media. In Cognitive Development neci. 2023. ChatGPT for good? On opportunities and challenges of large lan-
in Digital Contexts. Elsevier, 217–234. doi:10.1016/b978-0-12-809481-5.00011-0 guage models for education. Learn. Individ. Differ. 103 (1 April 2023), 102274.
[26] Jennifer A Fredricks, Phyllis C Blumenfeld, and Alison H Paris. 2004. School doi:10.1016/[Link].2023.102274
engagement: Potential of the concept, state of the evidence. Rev. Educ. Res. 74 [46] KhanAcademy. 2025. Meet Khanmigo: Khan Academy’s AI-powered teaching
(1 March 2004), 59–109. doi:10.3102/00346543074001059 assistant & tutor. [Link]
[27] F G Freeman, P J Mikulka, L J Prinzel, and M W Scerbo. 1999. Evaluation of an [47] Asma Ben Khedher, Imène Jraidi, and Claude Frasson. 2019. Tracking students’
adaptive automation system using three EEG indices with a visual tracking task. mental engagement using EEG signals during an interaction with a virtual
Biological psychology 50 (May 1999). doi:10.1016/s0301-0511(99)00002-2 learning environment. J. Intell. Learn. Syst. Appl. 11 (2019), 1–14. doi:10.4236/
[28] Jérémy Frey, May Grabli, Ronit Slyper, and Jessica R Cauchard. 2018. Breeze: jilsa.2019.111001
Sharing biofeedback through wearable technologies. In Proceedings of the 2018 [48] Jonathan W Kim, Ahmed Alaa, and Danilo Bernardo. 2024. EEG-GPT: Exploring
CHI Conference on Human Factors in Computing Systems. ACM. doi:10.1145/ capabilities of large language models for EEG classification and interpretation.
3173574.3174219 arXiv [[Link]] (31 Jan. 2024). doi:10.48550/arXiv.2401.18006
[29] Sabine Graf, Tzu-Chien Liu, Kinshuk, Nian-Shing Chen, and Stephen J H Yang. [49] Eckart Klieme, Frank Lipowsky, Katrin Rakoczy, and Nadja Ratzka. 2006. Qual-
2009. Learning styles and cognitive traits – Their relationship and its benefits itätsdimensionen und wirksamkeit von mathematikunterricht. Untersuchungen
in web-based educational systems. Comput. Human Behav. 25 (1 Nov. 2009), zur Bildungsqualität von Schule (2006), 127–146.
1280–1289. doi:10.1016/[Link].2009.06.005 [50] Nataliya Kosmyna, Eugene Hauptmann, and Yasmeen Hmaidan. 2023. A brain-
[30] Jennie K Grammer, Keye Xu, and Agatha Lenartowicz. 2021. Effects of context controlled quadruped robot: A proof-of-concept demonstration. Sensors (Basel)
on the neural correlates of attention in a college classroom. NPJ Sci. Learn. 6 24 (22 Dec. 2023), 80. doi:10.3390/s24010080
(6 July 2021), 15. doi:10.1038/s41539-021-00094-8 [51] Nataliya Kosmyna and Pattie Maes. 2019. AttentivU: An EEG-Based Closed-Loop
[31] Simone Grassini, Giulia Virginia Segurini, and Mika Koivisto. 2022. Watching Biofeedback System for Real-Time Monitoring and Improvement of Engagement
nature videos promotes physiological restoration: Evidence from the modulation for Personalized Learning. Sensors 19 (27 Nov. 2019). doi:10.3390/s19235200
of alpha waves in electroencephalography. Front. Psychol. 13 (7 June 2022), [52] Dong Hyeok Lee and Chun Kee Chung. 2024. Enhancing neural decoding with
871143. doi:10.3389/fpsyg.2022.871143 large language models: A GPT-based approach. In 2024 12th International Winter
[32] Sven Guenther, Nataliya Kosmyna, and Pattie Maes. 2024. Image classification Conference on Brain-Computer Interface (BCI). IEEE, 1–4. doi:10.1109/bci60775.
and reconstruction from low-density EEG. Sci. Rep. 14 (16 July 2024), 16436. 2024.10480499
doi:10.1038/s41598-024-66228-1 [53] Joanne Leong, Pat Pataranutaporn, Valdemar Danry, Florian Perteneder, Yaoli
[33] Mariam Hassib, Mohamed Khamis, Susanne Friedl, Stefan Schneegass, and Flo- Mao, and Pattie Maes. 2024. Putting things into context: Generative AI-enabled
rian Alt. 2017. Brainatwork: logging cognitive engagement and tasks in the work- context personalization for vocabulary learning improves learning motivation.
place using electroencephalography. In Proceedings of the 16th International Con- In Proceedings of the CHI Conference on Human Factors in Computing Systems,
ference on Mobile and Ubiquitous Multimedia. ACM. doi:10.1145/3152832.3152865 Vol. 32. ACM, 1–15. doi:10.1145/3613904.3642393
[34] Mariam Hassib, Stefan Schneegass, Philipp Eiglsperger, Niels Henze, Albrecht [54] Marco Marchesi and Bruno Riccò. 2013. BRAVO: a brain virtual operator for
Schmidt, and Florian Alt. 2017. EngageMeter: A system for implicit audience education exploiting brain-computer interfaces. In CHI ’13 Extended Abstracts on
engagement sensing using electroencephalography. In Proceedings of the 2017 Human Factors in Computing Systems. ACM. doi:10.1145/2468356.2479618
CHI Conference on Human Factors in Computing Systems. ACM. doi:10.1145/ [55] Caitlin Mills, Igor Fridman, Walid Soussou, Disha Waghray, Andrew M Olney,
3025453.3025669 and Sidney K D’Mello. 2017. Put your thinking cap on: detecting cognitive load
[35] Yuk Mui Elly Heung and Thomas K F Chiu. 2025. How ChatGPT impacts student using EEG during learning. In Proceedings of the Seventh International Learning
engagement from a systematic review and meta-analysis study. Computers and Analytics & Knowledge Conference. ACM. doi:10.1145/3027385.3027431
Education: Artificial Intelligence 8 (1 June 2025), 100361. doi:10.1016/[Link].2025. [56] Heni Mulyani, Mohamad Azim Istiaq, Elvia R Shauki, Fitrina Kurniati, and Hanifia
100361 Arlinda. 2025. Transforming education: exploring the influence of generative AI
[36] Nora Hollenstein, Jonathan Rotsztejn, Marius Troendle, Andreas Pedroni, Ce on teaching performance. Cogent Educ. 12 (31 Dec. 2025). doi:10.1080/2331186x.
Zhang, and Nicolas Langer. 2018. ZuCo, a simultaneous EEG and eye-tracking 2024.2448066
resource for natural sentence reading. Sci Data 5 (11 Dec. 2018), 180291. doi:10. [57] B S Oken, M C Salinsky, and S M Elsas. 2006. Vigilance, alertness, or sustained
1038/sdata.2018.291 attention: physiological basis and measurement. Clin. Neurophysiol. 117 (Sept.
[37] Anu Holm, Kristian Lukander, Jussi Korpela, Mikael Sallinen, and Kiti M I Müller. 2006), 1885–1901. doi:10.1016/[Link].2006.01.017
2009. Estimating brain load from the EEG. ScientificWorldJournal 9 (14 July [58] D F O Onah, J E Sinclair, and R Boyatt. 2014. Dropout rates of massive open online
2009), 639–651. doi:10.1100/tsw.2009.83 courses: Behavioural patterns. Unpublished. doi:10.13140/RG.2.1.2402.0009
[38] Yongquan Hu, Shuning Zhang, Ting Dang, Hong Jia, Flora D Salim, Wen Hu, [59] OpenAI. 2023. Teaching with AI. [Link]
and Aaron J Quigley. 2024. Exploring large-scale language models to evaluate [60] OpenAI. 2024. Introducing ChatGPT Edu. [Link]
EEG-based multimodal data for mental health. In Companion of the 2024 on ACM chatgpt-edu/.
Baradari, Kosmyna, Petrov, Kaplun, and Maes

[61] OpenAI. 2024. Memory and new controls for ChatGPT. [Link] [85] Łukasz Warchoł and Ludmiła Zając-Lamparska. 2023. The relationship of N200
index/memory-and-new-controls-for-chatgpt/. and P300 amplitudes with intelligence, working memory, and attentional control
[62] Brendan Parsons and Jocelyn Faubert. 2021. Enhancing learning in a perceptual- behavioral measures in young healthy individuals. Adv. Cogn. Psychol. 19 (2023),
cognitive training paradigm using EEG-neurofeedback. Sci. Rep. 11 (18 Feb. 2021), 63–75. doi:10.5709/acp-0404-2
4061. doi:10.1038/s41598-021-83456-x [86] Shaoyue Wen, Michael Middleton, Songming Ping, Nayan N Chawla, Guande
[63] Salil H Patel and Pierre N Azzam. 2005. Characterization of N200 and P300: Wu, Bradley S Feest, Chihab Nadri, Yunmei Liu, David Kaber, Maryam Zahabi,
selected studies of the Event-Related Potential. Int. J. Med. Sci. 2 (1 Oct. 2005), Ryan P McMahan, Sonia Castelo, Ryan Mckendrick, Jing Qian, and Claudio Silva.
147–154. doi:10.7150/ijms.2.147 2025. AdaptiveCoPilot: Design and testing of a NeuroAdaptive LLM cockpit
[64] A T Pope, E H Bogart, and D S Bartolome. 1995. Biocybernetic system evaluates guidance system in both novice and expert pilots. arXiv [[Link]] (7 Jan. 2025).
indices of operator engagement in automated task. Biol. Psychol. 40 (May 1995), [87] Beste F Yuksel, Kurt B Oleson, Lane Harrison, Evan M Peck, Daniel Afergan,
187–195. doi:10.1016/0301-0511(95)05116-3 Remco Chang, and Robert J K Jacob. 2016. Learn Piano with BACh: An Adaptive
[65] Mirko Raca and Pierre Dillenbourg. 2013. System for assessing classroom atten- Learning Interface that Adjusts Task Difficulty Based on Brain State. In Proceed-
tion. In Proceedings of the Third International Conference on Learning Analytics ings of the 2016 CHI Conference on Human Factors in Computing Systems (CHI ’16).
and Knowledge. ACM. doi:10.1145/2460296.2460351 Association for Computing Machinery, 5372–5384. doi:10.1145/2858036.2858388
[66] Mirko Raca, Lukasz Kidzinski, and Pierre Dillenbourg. 2015. Translating Head [88] Xiayin Zhang, Ziyue Ma, Huaijin Zheng, Tongkeng Li, Kexin Chen, Xun Wang,
Motion into Attention - Towards Processing of Student’s Body-Language. Inter- Chenting Liu, Linxi Xu, Xiaohang Wu, Duoru Lin, and Haotian Lin. 2020. The
national Educational Data Mining Society (June 2015). combination of brain-computer interfaces and artificial intelligence: applications
[67] Johnmarshall Reeve and Ching-Mei Tseng. 2011. Agency as a fourth aspect of and challenges. Ann. Transl. Med. 8 (June 2020), 712. doi:10.21037/atm.2019.11.109
students’ engagement during learning activities. Contemp. Educ. Psychol. 36 [89] Yuhong Zhang, Qin Li, Sujal Nahata, Tasnia Jamal, Shih-Kuen Cheng, Gert
(1 Oct. 2011), 257–267. doi:10.1016/[Link].2011.05.002 Cauwenberghs, and Tzyy-Ping Jung. 2024. Integrating large language model,
[68] Lauren E Reinerman, Gerald Matthews, Joel S Warm, Lisa K Langheim, Kelley EEG, and eye-tracking for word-level neural state classification in reading com-
Parsons, Christina A Proctor, Tazeen Siraj, Lloyd D Tripp, and Robert M Stutz. prehension. IEEE Trans. Neural Syst. Rehabil. Eng. PP (14 Aug. 2024), 1–1.
2006. Cerebral blood flow velocity and task engagement as predictors of vigilance doi:10.1109/TNSRE.2024.3435460
performance. Proc. Hum. Factors Ergon. Soc. Annu. Meet. 50 (Oct. 2006), 1254–1258. [90] Tong Zhou, Xuhang Chen, Yanyan Shen, Martin Nieuwoudt, Chi-Man Pun,
doi:10.1177/154193120605001210 and Shuqiang Wang. 2023. Generative AI enables EEG data augmentation for
[69] Maria V Ruiz-Blondet, Zhanpeng Jin, and Sarah Laszlo. 2016. CEREBRE: A novel Alzheimer’s disease detection via diffusion model. In 2023 IEEE International
method for very high accuracy event-related potential biometric identification. Symposium on Product Compliance Engineering - Asia (ISPCE-ASIA). IEEE, 1–6.
IEEE Trans. Inf. Forensics Secur. 11 (July 2016), 1618–1629. doi:10.1109/tifs.2016. doi:10.1109/ispce-asia60405.2023.10365931
2543524
[70] Yashvir Sabharwal and Balaji Rama. 2024. Comprehensive review of EEG-to-
output research: Decoding neural signals into images, videos, and audio. arXiv
[[Link]] (27 Dec. 2024). A SYSTEM PROMPTS
[71] Akane Sano, Judith Amores, and Mary Czerwinski. 2024. Exploration of LLMs,
EEG, and behavioral data to measure and support attention and sleep. arXiv A.1 NeuroChat system prompt
[[Link]] (1 Aug. 2024).
[72] Brooke Schultz. 2025. This School Will Have Artificial Intelligence Teach Kids NeuroChat System Prompt
(With Some Human Help). Education Week (6 Jan. 2025).
[73] Skipper Seabold and Josef Perktold. 2010. Statsmodels: Econometric and statis-
tical modeling with python. In Proceedings of the Python in Science Conference. You are an encouraging tutor who helps students across
SciPy, 92–96. doi:10.25080/majora-92bf1922-011 various subjects and skill levels understand concepts by
[74] Uri Shaked. 2021. muse-js: Muse 2016 EEG Headset JavaScript Library (using
Web Bluetooth). [Link]
explaining ideas and asking students questions. Start by
[75] David J Shernoff, Mihaly Csikszentmihalyi, Barbara Schneider, and Elisa Steele introducing yourself to the student as their AI-Tutor who
Shernoff. 2014. Student engagement in high school classrooms from the perspec- is happy to help them with any questions.
tive of flow theory. In Applications of Flow in Human Development and Education.
Springer Netherlands, 475–494. doi:10.1007/978-94-017-9094-9_24 Additionally, you will be provided with the student’s
[76] Gale M Sinatra, Benjamin C Heddy, and Doug Lombardi. 2015. The challenges of
defining and measuring student engagement in science. Educ. Psychol. 50 (2 Jan.
cognitive load values while they were reading any
2015), 1–13. doi:10.1080/00461520.2014.1002924 previous responses of yours as measured by EEG. Your
[77] W Speier, C Arnold, and N Pouratian. 2016. Integrating language models into goal is to act like a good tutor, using the insights from
classifiers for BCI communication: a review. J. Neural Eng. 13 (6 June 2016),
031002. doi:10.1088/1741-2560/13/3/031002 these metrics to adapt your responses to the student’s
[78] Ricarda Steinmayr, Anne F Weidinger, Malte Schwinger, and Birgit Spinath. cognitive load dynamically. The value you will be given:
2019. The importance of students’ motivation for their academic achievement
- replicating and extending previous findings. Front. Psychol. 10 (31 July 2019), **Normalized engagement score:** This represents the
1730. doi:10.3389/fpsyg.2019.01730 user’s level of engagement or arousal on a normalized
[79] Daniel Szafir and Bilge Mutlu. 2012. Pay attention!: designing adaptive agents that
monitor and improve user engagement. In Proceedings of the SIGCHI Conference scale from 0 to 1. The engagement index is a ratio of the
on Human Factors in Computing Systems. ACM. doi:10.1145/2207676.2207679 student’s beta/(theta+alpha) bands.
[80] The Abdul Latif Jameel Poverty Action Lab (J-PAL). 2022. Teaching at the
Right Level to improve learning. [Link] Do not ever disclose the EEG metrics to the user since
teaching-right-level-improve-learning.
[81] The Open Innovation Team and Department for Education. 2024. Generative AI they are hidden to them. Also, never make direct
in education - Educator and expert views. Technical Report. UK Department for comments on their metrics and don’t mention the names
Education. of the metrics.
[82] Philipp Tholke, Antoine Bellemare-Pepin, Yann Harel, Francois Lespinasse, and
Karim Jerbi. 2024. Bio-Mechanical Poet: An Immersive Audiovisual Playground Give students explanations, examples, and analogies
for Brain Signals and Generative AI. In Proceedings of the 15th international
conference on computational creativity, In Kazjon Grace, Maria Teresa Llano, about the concept to help them understand.
Martins Pedro, and Maria M Hedblom (Eds.). 65–74.
[83] Angela Vujic, Shreyas Nisal, and Pattie Maes. 2023. Joie: a Joy-based Brain- Adaptations Based on Cognitive Load:
Computer Interface (BCI). In Proceedings of the 36th Annual ACM Symposium on You need to learn how the user reacted to your
User Interface Software and Technology (UIST ’23). Association for Computing
Machinery, 1–14. doi:10.1145/3586183.3606761 adaptations. Based on their cognitive load, modulate the
[84] Yansen Wang and Zilong Wang. 2024. EEG2Video: Towards Decoding Dynamic response length, factual vs. storytelling, ease of text
Visual Perception from EEG Signals. The Thirty-eighth Annual Conference on
Neural Information Processing Systems (NeurIPS 2024).
(explain like I’m 5 vs. explain like I’m a PhD), bullet points
NeuroChat: A Neuroadaptive AI Chatbot for Customizing Learning Experiences

B QUIZ QUESTIONS AND ESSAY PROMPTS


vs. long-form text, level of depth and detail, Socratic
questions, and styling of text (bolding of keywords). T. rex
Every person is different; however, here are some general Fill-in-the-blank Questions.
pointers: (1) Q1: T. rex fossils have primarily been found in . (2
- Students with higher cognitive loads enjoy more points): North America
complex, scientific, and in-depth explorations of a topic. (2) Q2: T. Rex likely obtained its food from hunting as well as
Students with low to medium cognitive load may prefer . (2 points): Scavenging
explanations that prompt more of their curiosity or (3) Q3: One of the most complete T. rex fossils, known as
represent a challenge. , was discovered in South Dakota in 1990. (2 points):
- Students with lower cognitive load need to discover a Sue
question they are curious about; hence provide Multiple Choice Questions. Q4: Which era did the Tyrannosaurus
explanations that prompt more of their curiosity or Rex live in?
represent a challenge. You may also give them interesting
• Jurassic
facts or narrative examples, or ask them questions.
• Cretaceous
- Once a student shows an appropriate level of
• Triassic
understanding given their learning level and cognitive
• Permian
load, ask them to explain the concept in their own words;
• Devonian
this is the best way to show you know something, or ask
• This did not come up in my conversation
them for examples.
• I don’t know
- Encourage learners to explain their thinking.
- If the learner needs more engagement, provide (2 points): Cretaceous
thought-provoking questions or exercises. Also, suggest Q5: How many fingers did the Tyrannosaurus Rex have on each
questions to explore together. Students with higher hand?
cognitive load may also find these interesting. • Two
- Provide positive reinforcement but also critical feedback. • Three
- Offer clarification or examples if the user seems to need • Four
more understanding. • Five
- Analogies or storytelling can help raise cognitive load. • Six
- Sounding more energetic or scientific can help raise • This did not come up in my conversation
cognitive load. • I don’t know
- Bolding important keywords can help raise and maintain (2 points): Two
cognitive load. Q6: Choose all popular myths about T. Rex that are likely wrong.
Remember, your role is to support the user’s learning • It had poor vision.
journey, adapt to their needs, and ensure a positive, • It was a slow, clumsy mover.
effective, and engaging educational experience. Be patient, • Its arms were likely useless.
encouraging, and responsive to the user’s cognitive state • It had one of the most powerful bites ever known.
and feedback. • It was the dominant dinosaur in its environment.
• It may have had feathers.
By following these guidelines, you will help users achieve (6 points): Correct answers are the first three options: It had poor
their learning goals effectively and enjoyably. Respond in vision; It was a slow, clumsy mover; Its arms were likely
Markdown. useless. For partial credit, 1 point was awarded for each correct
choice, and 1 point for not selecting each incorrect choice.
Ranking Question. Q9: Sort the T. Rex into this list of dinosaurs by
length.
A.2 Control condition prompt
(1) Argentinosaurus - 98 ft (30 m)
Control Condition Prompt (2) Brachiosaurus - 82 ft (25 m)
(3) Spinosaurus - 59 ft (18 m)
You are an AI Tutor designed to assist learners across a (4) Iguanodon - 33 ft (10 m)
variety of subjects and skill levels. Your primary goal is to (5) Pachycephalosaurus - 16 ft (5 m)
provide clear, accurate, and engaging explanations tailored (6) Velociraptor - 7 ft (2.5 m)
to each learner’s needs. You should strive to be patient, (7) Tyrannosaurus Rex - ?
encouraging, and adaptive in your teaching style.
(2 points): Tyrannosaurus Rex (40 ft) should be placed between
Respond in Markdown. Iguanodon and Spinosaurus. (1 point): Placement is one off, in
position 3 or 5.
Baradari, Kosmyna, Petrov, Kaplun, and Maes

Essay Prompt. Answer one of the following in the form of a mini- Essay Prompt. Answer one of the following in the form of an essay
essay (introduction, main section, conclusion): (introduction, main section, conclusion):
(1) Discuss T. Rex’s physical and behavioral characteristics (1) Discuss the socio-economic factors that contributed to the
which enabled it to dominate its environment. outbreak of the Taiping Rebellion.
(2) Analyze paleobiological discoveries that have changed our (2) Analyze the role of religion in the Taiping Rebellion.
understanding of T. Rex. (3) Discuss the legacy of the Taiping Rebellion on subsequent
(3) Evaluate the theories regarding the function of T. Rex’s Chinese history.
small arms. What are some of the proposed explanations, (4) Create your own essay question related to the Taiping Re-
and which do you find most convincing? bellion based on your conversation.
(4) Create your own essay question related to the T. Rex based
on your conversation.

Taiping Rebellion
Fill-in-the-gap Questions.
(1) Q1: The goal of the Taiping Rebellion was to establish the
. (2 points): Taiping Heavenly Kingdom of Great
Peace, or similar.
(2) Q2: One of the distinctive aspects of the Taiping ideol-
ogy was its spiritual blend of and , which
appealed to the disaffected rural populace. (2 points): Chris-
tianity (1) and Chinese spiritual traditions (Buddhism,
Taoism, Confucianism) or similar (1).
Multiple Choice Questions. Q5: What was the name of the leader
of the Taiping Rebellion?
• Zeng Guofan
• Hong Xiuquan
• Hong Tianguifu
• Feng Yunshan
• Yang Xiuqing
• This did not come up in my conversation
• I don’t know
(1 point): Hong Xiuquan
Q6: The Qing government was administered by leaders of which
ethnic minority?
• Manchu
• Hakka
• Han
• Zhuang
• Hui
• This did not come up in my conversation
• I don’t know
(1 point): Manchu
Ranking Question. Q10: Here is a list of death tolls in wars by lowest
estimate. Where would the Taiping Rebellion sit?
(1) World War II - 80 million
(2) World War I - 17 million
(3) Spanish conquest of Mexico - 10.5 million
(4) Russian Civil War - 7 million
(5) Napoleonic Wars - 3.5 million
(6) Vietnam War - 1.3 million
(7) Taiping Rebellion
(2 points): Between World War II and World War I, with an
estimated toll of 20-30 million.

Common questions

Powered by AI

NeuroChat uses EEG data to adapt its responses in real-time by modifying conversational flow, response complexity, and interaction depth based on the learner's cognitive state. This approach offers a more personalized and adaptive learning experience compared to traditional systems that relied on pre-scripted content. The key benefit is the reduction of manual input from the user, enhancing engagement by dynamically tailoring content to individual needs. However, the limitations include challenges in translating increased engagement into measurable learning outcomes, as evidenced by no significant improvement in quiz or essay scores .

The Chatbot study highlights that user engagement preferences can significantly influence adaptive learning strategies. While some users favored conversational and human-like interactions offered by NeuroChat, others preferred factual and concise responses. This diversity in preferences suggests that effective adaptive learning systems must be flexible enough to tailor engagement strategies that cater to individual user preferences, potentially enhancing both perceived and actual engagement .

Neuroadaptive AI systems like NeuroChat have significant implications for future AI-powered tutoring. They mark a shift from rule-based, pre-mapped content to generative, personalized tutoring that adapts in real-time based on user engagement levels inferred from EEG data. This could enhance adaptability and reduce user effort in communicating needs. Additionally, it suggests a new avenue for creating self-directed learning environments that accommodate varied learning preferences by modifying engagement styles based on real-time feedback .

NeuroChat's distinctive features include its ability to provide human-like, responsive interactions by adapting its responses in real-time based on EEG-measured engagement. Users perceive it as more interactive due to its personalized feedback and conversational, tutor-like approach, which contrasts with standard large language models that offer more factual and structured responses. This human-like adaptability significantly enhances perceived engagement and interaction depth .

EEG monitoring contributes to the personalization of educational experiences by providing real-time feedback on a learner's cognitive states such as engagement, concentration, and cognitive load. This data is used to dynamically adjust instructional content or interaction modalities, ensuring the learner is neither underwhelmed nor overwhelmed. Techniques like EEG-based mental workload classification allow adaptive learning systems to tailor content complexity, improving engagement and potentially optimizing learning outcomes .

EEG for cognitive workload classification enhances adaptive learning systems by enabling them to adjust educational content based on the learner's mental state. By differentiating between high and low cognitive loads, these systems can modify content complexity and presentation style to maintain optimal engagement and prevent cognitive overload, thus creating a more effective learning environment .

Neuroadaptive systems like Thinking Cap differ from earlier adaptive learning technologies primarily in their responsive and dynamic adjustment of instructional content in real-time based on EEG-derived cognitive load measures. Unlike earlier systems that modified delivery methods but not content complexity, Thinking Cap involves creating varied difficulty levels of the content itself. This ensures learners are continuously challenged at an appropriate level, avoiding under- or overwhelming them, thus aligning with real-time cognitive states .

Challenges in converting engagement gains from neuroadaptive systems into measurable learning improvements include: (1) The engagement measured by EEG reflects sustained attention but may not translate to deep learning or knowledge retention. (2) Task design limitations, as NeuroChat's open-ended approach makes it difficult to implement structured difficulty progression that aids learning. (3) Short study duration, as benefits of neuroadaptive learning might require multiple sessions to manifest in measurable performance improvements. These challenges indicate the complexity of aligning engagement with effective learning strategies .

Vygotsky’s Zone of Proximal Development (ZPD) relates to EEG studies on student engagement by illustrating that heightened engagement occurs when tasks are slightly beyond a student's current skill level, demanding higher cognitive effort. EEG studies reveal significant attentional focus fluctuations, indicating tasks are within the ZPD when engagement is high, yet expertise is insufficient, leading to cognitive overload rather than effective learning. This suggests the importance of aligning task complexity with a learner's developmental stage for optimal learning outcomes .

Generative AI capabilities in neuroadaptive systems could revolutionize educational settings by allowing for real-time creation of personalized content that adapts seamlessly to a student's cognitive state. This reduces the need for pre-scripted materials and manual adjustments, facilitating more natural and responsive learning experiences. The AI could infer engagement levels from EEG data and adjust content complexity dynamically, providing a tailored learning path that enhances motivation and engagement without overloading learners .

You might also like