0% found this document useful (0 votes)
3 views16 pages

Thinking Skills and Creativity: Chokri Kooli, Nadia Yusuf, Mohammed Y. Sarhan

The document presents the Colloquial Engagement Theory with AI Awareness (CET-AIA), a new pedagogical framework aimed at addressing academic integrity challenges posed by generative AI in higher education. CET-AIA incorporates informal, context-sensitive questioning in assessments to discourage AI-assisted cheating and enhance student engagement, demonstrating its effectiveness through a quasi-experimental study. The framework is designed to promote critical thinking and authentic learning experiences while ensuring assessments remain resistant to AI exploitation.

Uploaded by

Lethidieulinh Lx
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views16 pages

Thinking Skills and Creativity: Chokri Kooli, Nadia Yusuf, Mohammed Y. Sarhan

The document presents the Colloquial Engagement Theory with AI Awareness (CET-AIA), a new pedagogical framework aimed at addressing academic integrity challenges posed by generative AI in higher education. CET-AIA incorporates informal, context-sensitive questioning in assessments to discourage AI-assisted cheating and enhance student engagement, demonstrating its effectiveness through a quasi-experimental study. The framework is designed to promote critical thinking and authentic learning experiences while ensuring assessments remain resistant to AI exploitation.

Uploaded by

Lethidieulinh Lx
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Thinking Skills and Creativity 60 (2026) 102051

Contents lists available at ScienceDirect

Thinking Skills and Creativity


journal homepage: [Link]/locate/tsc

Colloquial engagement theory with AI awareness (CET-AIA): A


new creative pedagogical framework for ethical assessment in the
age of artificial intelligence
Chokri Kooli a,d,* , Nadia Yusuf b , Mohammed Y. Sarhan c
a
Graduate school of Public and International affairs, University of Ottawa, Canada
b
Department of Economics, Faculty of Economics and Administration, King Abdulaziz University, Jeddah 21589, Saudi Arabia
c
Department of Management Information System, Faculty of Economics and Administration, King Abdulaziz University, Jeddah 21589, Saudi
Arabia
d
Royal military college of Canada, Kingston, Canada

A R T I C L E I N F O A B S T R A C T

Keywords: This study introduces the Colloquial Engagement Theory with AI Awareness (CET-AIA), a novel
Colloquial engagement theory pedagogical framework developed to address academic integrity challenges in higher education
AI in education posed by generative AI systems. CET-AIA integrates informal, context-sensitive questioning into
Academic integrity
assessment design to discourage AI-assisted cheating and enhance authentic student engagement.
Colloquial questioning
A quasi-experimental design was implemented across three undergraduate courses, comparing
Authentic assessment
CET-AIA student and AI (ChatGPT-3.5) performance on traditional versus colloquial multiple-choice as­
ChatGPT sessments. Theoretical foundations were drawn from Constructivism, Cognitive Load Theory,
Ethical pedagogy Sociocultural Theory, and Authentic Assessment. Performance trends were analyzed using t-tests,
ANOVA, regression models, and mixed-effects modeling. Results indicate a significant initial
decline in student scores under colloquial questioning, followed by gradual improvement, con­
firming both the cognitive challenge and adaptation process. In contrast, AI models showed a
persistent performance drop when confronted with colloquial, context-specific questions. These
outcomes demonstrate CET-AIA’s effectiveness in fostering deeper learning and in shielding as­
sessments from AI exploitation. CET-AIA offers a scalable framework for designing AI-resistant
assessments that promote critical thinking and real-world comprehension. The approach aligns
with current educational priorities in fostering academic honesty and student-centered learning in
digitally enhanced environments. This research is among the first to offer a theoretically
grounded, empirically validated framework specifically designed to neutralize AI-assisted
cheating through linguistic and contextual innovation. CET-AIA bridges the gap between AI
ethics, pedagogy, and assessment, presenting a future-ready model for ethical education.

1. Introduction

In recent years, the increasing integration of artificial intelligence (AI) into various sectors has introduced both opportunities and
challenges. In the educational sphere, intelligent tutoring systems and other AI solutions are being successfully applied to enhance

* Correspondence.
E-mail addresses: ckooli@[Link] (C. Kooli), nyusuf@[Link] (N. Yusuf), mysarhan@[Link] (M.Y. Sarhan).

[Link]
Received 12 July 2025; Received in revised form 28 September 2025; Accepted 3 November 2025
Available online 4 November 2025
1871-1871/© 2025 The Authors. Published by Elsevier Ltd. This is an open access article under the CC BY-NC-ND license
([Link]
C. Kooli et al. Thinking Skills and Creativity 60 (2026) 102051

teaching outcomes, facilitate learning, and streamline assessment (Alrakhawi et al.,., 2023). Although such technologies have the
capacity to offer considerable benefits for both students and educators, the integration of AI also raises pertinent questions concerning
academic integrity (Yusuf et al., 2024) and the reliability of assessment processes (Kooli, 2023). Recent research performed by Kooli,
Kooli and Kooli (2025) described the Generative AI Addiction Syndrome (GAID) as a newly identified behavioral disorder marked by
compulsive, co-creative interaction with AI systems, which can foster dependence among students by impairing their critical thinking,
emotional self-regulation, and creative autonomy. The rapid advancements in AI present both challenges and opportunities within
education. While AI offers tools to enhance teaching, its integration presents a significant challenge to academic integrity and, more
critically, to the development of higher-order thinking skills. The ability of services like ChatGPT to generate plausible answers
instantly is a genuine issue that threatens assessment validity and can detrimentally affect students’ capacity for independent
problem-solving and critical thinking (Cotton et al., 2024). This technological shift necessitates a fundamental rethinking of the core
capabilities learners must develop to effectively collaborate with and critically evaluate AI systems (Markauskaite et al., 2022). This
erosion of essential cognitive skills, coupled with the profound ethical implications of AI misuse, creates an urgent need for peda­
gogical innovation. Educators are thus prompted to explore new assessment strategies that ensure academic honesty while simulta­
neously fostering the creative and flexible thinking required for students to thrive in an AI-saturated world. In fact, findings support the
efficacy of creative courses in preparing students to thrive in any discipline in the age of AI, highlighting the critical role of such
education in a world increasingly influenced by these technological shifts (Habib, Vogel & Thorne, 2025). From this standpoint, the
profound ethical implications of the covered subject prompt educators to explore new options that could ensure fair and honest as­
sessments in spite of readily-accessible advanced AI systems.
Foundational educational theories such as Constructivism, Cognitive Load Theory, Sociocultural Theory, and Authentic Assessment
provide the bedrock for modern pedagogy but were not designed to address the unique challenges of generative AI. For instance, while
Authentic Assessment calls for real-world tasks, these can often be formulated in ways that AI can easily solve. This challenge is
compounded by the rapid evolution of LLMs, with newer models like GPT-4 demonstrating human-level performance on academic
benchmarks and producing text that is increasingly difficult for both educators and technological tools to identify as machine-
generated (Perkins, 2023; Ray, 2023). Consequently, none of the existing theories provide a fully comprehensive framework to
address the specific challenges of AI-enabled cheating and the design of AI-resistant tests (Holmes et al., 2022; Su & Yang, 2023; Xu &
Ouyang, 2022). This gap highlights the need for a new framework designed not to replace these theories, but to adapt their principles
for an AI-rich environment.
The potential for AI to enable cheating and undermine the reliability of assessments is a critical concern that is covered in the extant
literature but lacks an adequate solution (Griesbeck et al., 2024; Oravec, 2022). Due to the current gap in theoretical knowledge on the
subject, a new framework is needed to tackle these specific challenges. Consequently, this research aims to propose the Colloquial
Engagement Theory with AI Awareness (CET-AIA), a novel approach that would integrate colloquial questioning into assessments to
mitigate AI reliance and enhance student engagement. In this framework, colloquial questioning refers to the use of informal language,
regular expressions, and conversational styles when designing multiple-choice questions (MCQs). Unlike conventional MCQs that are
typically reliant on formal and structured language, colloquial questions entail a linguistic and semantic barrier that prevents AI from
generating accurate responses. Intended to create assessments resistant to interpretations by AI models, the proposed approach can
ultimately reduce students’ ability to use AI solutions for cheating. The development and evaluation of this framework are key steps
toward ensuring the integrity and effectiveness of educational assessments in light of the growing acceptance of AI.
For these reasons, the primary objective of this research is to develop and evaluate the CET-AIA framework in educational as­
sessments. By coupling colloquial questioning with AI awareness, we intend to create an engaging and ethical assessment environment
offering a path toward reducing dependency on ChatGPT and other AI solutions. Accordingly, the study is set to investigate the impact
of colloquial questioning on student performance and engagement; evaluate the effectiveness of colloquial questioning in mitigating
AI-assisted cheating; and explore how the CET-AIA framework can enhance the validity and reliability of assessment outcomes. The
significance of this work lies in its potential to transform educational assessment practices by addressing the ethical challenges posed
by AI. As the CET-AIA framework is presented as a proactive solution that can promote authentic learning experiences, this research is
expected to contribute to the broader discourse on AI in education. Moreover, we aim to produce empirical evidence and offer practical
solutions for multiple stakeholders, including educators, policymakers, and AI developers. This study will next present a detailed
literature review, followed by the theoretical framework development. The methodology section will then outline the research design
and data collection procedures before concluding the paper with a discussion of the findings, implications, and future research
directions.

1.1. Literature review

The integration of Artificial Intelligence into education is associated with an ongoing transformative shift and its numerous benefits
need to be embraced along with significant ethical challenges. The current literature review section explores current perspectives on AI
services and products through the lens of disadvantages and advantages of their application in education.

1.2. Current perspectives on AI tools in education

Similar to the situation observed in other sectors, AI technologies are expected to revolutionize traditional pedagogical approaches
in a multitude of ways. Namely, several authors highlighted the long-anticipated adoption of personalizing learning experiences,
accessible tools for enhancing student engagement, and software capable of streamlining administrative tasks (Adıgüzel et al., 2023;

2
C. Kooli et al. Thinking Skills and Creativity 60 (2026) 102051

Igbokwe, 2023; Limo et al., 2023). Described as one of the most significant advantages of AI in education, personalized learning is
enabled by AI-driven systems that adapt instructional content to meet individual students’ learning styles and paces (Rane et al., 2023).
One example of a positive outcome pertains to immediate feedback and tailored educational experiences capable of addressing stu­
dents’ specific needs and enhancing their overall academic outcomes. From the administrative standpoint, AI solutions allow educators
to focus on meaningful interactions with students by automating routine tasks and processes (Parycek et al., 2023). Common use cases
highlighting this aspect of AI in education include grading, scheduling, and monitoring student progress (Jukiewicz, 2024) . Through
the use of generative AI, both students and professionals can engage in an interactive development of skills such as coding,
problem-solving, and critical thinking (Bankins & Formosa, 2023; Mogawi et al., 2023). Encompassing ChatGPT and other chatbots,
this type of AI technology proves to be immensely useful, while also being particularly prone to misuse such as academic cheating and
plagiarism.

1.3. Ethical challenges of AI in education

Despite the array of already recognized as well as theorized benefits, the integration of AI in education is prone to unprecedented
ethical challenges stemming from unintended uses of popular AI-based tools. One of the concerns frequently discussed in academic
sources is data privacy (Hridi et al., 2024; Nguyen et al., 2023). As AI models often require large amounts of data to maintain their
operations, multiple fault lines are being created, leading to concerns regarding how data is collected, stored, and used (Sebastian,
2023). As a result, there is a consistent risk that sensitive information could be misused or inadequately protected from unauthorized
access. Another prominent ethical problem is the recognized role of AI models in perpetuating biases (O’Connor & Liu, 2023). In cases
of inadequately designed AI algorithms that operate without proper monitoring, models are often reported responding with biased or
incorrect answers. This scenario introduces disproportionate risks to students from marginalized or underprivileged backgrounds
(Roshanaei, 2024). Such negative experiences with AI tools not only fail to mitigate educational inequalities but further exacerbate
them and necessitate additional oversight.
In contrast to the aforementioned ethical challenges that often result from unintended problems, malicious use, and design faults of
AI solutions, concerns surrounding academic integrity are inherently related to the intended utilization of such models. Specifically,
the use of AI for cheating and plagiarism is a prominent example of academic dishonesty, as students might exploit AI tools to obtain
correct answers to assignments and exams (Oravec, 2023; Sweeney, 2023). Although many AI models are designed with consideration
for factual accuracy and knowledge sharing as key underlying principles, their unrestricted use undermines the educational process
and the value of academic credentials (Humble et al., 2024; Kooli, 2023). For this reason, scholars emphasize the need for a failsafe
mechanism to ensure that students use AI tools as a complement to learning rather than as a substitute for their own cognitive efforts.

1.4. The use and ethical gaps of AI in educational assessments

Recent research has explored the transformative potential of AI particularly applications based on large language models (LLMs)
like ChatGPT in educational assessments (Kooli & Yusuf, 2025). On one hand, AI can enhance grading efficiency and provide
personalized feedback. On the other hand, this integration introduces profound ethical challenges that current assessment theories fail
to adequately address. While the literature extensively covers the benefits of AI for administrative efficiency or personalized learning
(Owan et al., 2023; Smolansky et al., 2023), there is a critical gap concerning the ethical design of assessments in an age where AI can
mimic student mastery. This study addresses that specific gap.

1.5. Inadequate strategies to mitigate AI-Assisted cheating

Current assessment methods and theories do not provide sufficiently adequate and robust strategies to prevent AI-assisted cheating.
With conventional theories being primarily intended for designing engaging and cognitively appropriate assessments, the data is
scarce regarding the capacity of such approaches to accommodate challenges posed by AI (Ouyang et al., 2022; Tuomi, 2024). In other
words, these theories emphasize the importance of active learning and cognitive engagement but lack specific mechanisms to make
assessments resistant to AI manipulation. This inadequacy is particularly acute because text produced by modern LLMs is uniquely
created for each prompt, allowing it to evade traditional text-matching software designed to detect copy-and-paste plagiarism (Perkins,
2023). Furthermore, research has consistently shown that academic staff struggle to reliably distinguish between human-written and
AI-generated text, a difficulty that will likely increase as AI’s computational potential and training on larger datasets grows (Creely &
Blannin, 2025; Perkins, 2023). This creates an ongoing "arms race" where conventional, detection-based strategies for ensuring aca­
demic integrity are becoming increasingly unenforceable (Perkins, 2023), reinforcing the need for novel assessment designs rather
than purely technological solutions.
Moreover, Kooli (2023) highlights the implications of chatbots on the reliability and validity of assessment outcomes. The study
elucidates how reliance on AI tools can distort the evaluation of students’ knowledge and skills, leading to inaccurate and biased
assessments. By providing instantaneous answers to questions, AI models threaten the development of essential skills such as critical
thinking, creativity, and independent problem-solving. This gradual degradation of conventional learning processes raises concerns
about the integrity and fairness of academic assessments.

3
C. Kooli et al. Thinking Skills and Creativity 60 (2026) 102051

1.6. Need for authentic and engaging assessments

In both scholarly and professional sources, there is a call for more authentic and engaging assessment methods that can challenge
students’ understanding and critical thinking skills (Kooli, 2023; Yusuf et al., 2024; Bicer 2024). One example is Authentic Assessment
theory advocating for real-world applications and tasks that reflect actual challenges students may encounter in real-life situations
(Ifelebuegu, 2023). However, assessments designed based on the above-mentioned approach often still follow traditional formats that
AI models can easily interpret and respond to. Furthermore, Kooli (2023) emphasizes the importance of promoting active learning and
student engagement to counteract the attractive ease of use and convenience of AI services. Through interactive discussions, group
work, and authentic assessments, educators can cultivate a learning environment that values intellectual curiosity and independent
inquiry.
Taking into account the knowledge gaps highlighted above and the current state of research on the subject, the CET-AIA framework
is planned to address critical gaps pertaining to AI in educational assessments by providing a structured approach for mitigating AI-
assisted cheating and promoting student engagement. The incorporation of colloquial questioning, one of the theory’s main concepts,
not only promotes integrity in assessments but also partially aligns with other educational theories to create more authentic, inclusive,
and engaging learning experiences.

1.7. Conceptual framework CET-AIA

The Colloquial Engagement Theory with AI Awareness (CET-AIA) is a novel pedagogical framework designed to enhance student
engagement, comprehension, and academic integrity in the age of artificial intelligence. The theory posits that once students un­
derstand that AI cannot reliably solve assessment questions anchored in unique classroom experiences, they will shift their behavior
from depending on AI to relying on their own cognitive engagement. At its core, CET-AIA posits that utilizing a colloquial teaching
style anchored in real-life examples from students’ immediate local or global environments is the key to fostering this shift. Fig. 1
below provides a visual summary of the conceptual structure of CET-AIA, outlining its key theoretical and practical dimensions.
As illustrated, the framework is built around several interconnected elements. The Core Philosophy is to enhance academic
integrity and shift students from AI dependency toward genuine cognitive engagement. This philosophy is grounded in established
Pedagogical Foundations which are adapted to the AI context and is operationalized through its Core Components: a Colloquial

Fig. 1. Conceptual Framework of CET-AIA.

4
C. Kooli et al. Thinking Skills and Creativity 60 (2026) 102051

Teaching Style and AI Awareness in Assessment. The Rationale for Transition stems from the observation that colloquial questions
disrupt AI interpretation patterns, thereby promoting deeper student reflection. This is achieved via specific Design Principles,
including Contextual Anchoring and Linguistic Obfuscation, which lead to several Key Theoretical Claims, namely that CET-AIA limits
AI utility and fosters independent learning. Finally, the Development Motivation arose from the observed misuse of AI and the need for
more human-centric assessments.
CET-AIA is grounded in four key pedagogical frameworks: Constructivism, Cognitive Load Theory, Sociocultural Theory, and
Authentic Assessment. Each contributes a critical dimension to understanding how informal, context-sensitive assessment can promote
ethical student engagement and reduce AI manipulation. Fig. 2 presents the educational foundations upon which CET-AIA is built,
showing how traditional theories are repurposed to address AI-related assessment challenges.
Constructivism: While constructivism posits that learners build meaning through active interaction with content (MacLeod et al.,
2022), this process is threatened when AI can generate sophisticated outputs without requiring the learner to engage. CET-AIA extends
constructivist practice by creating assessment questions tied to the unique, shared context of the classroom, compelling students to
construct meaning from their own lived academic experiences—a process that AI cannot replicate.
Cognitive Load Theory (CLT): CLT suggests that effective learning occurs when extraneous cognitive demands are minimized.
Colloquial questioning initially increases intrinsic load due to its novelty, but CET-AIA reconfigures the application of CLT by using this
challenge as a pedagogical tool. By anchoring abstract content in familiar, colloquial language, it ultimately helps anchor learning and
supports retention (O’Connor, 2022; Skulmowski & Xu, 2022), discouraging the cognitively passive act of using AI.
Sociocultural Theory: This theory emphasizes that learning is shaped by social and cultural context (Taylor, 2022; Kilag et al.,
2024). CET-AIA operationalizes this principle by embedding assessment within the specific subculture of the classroom, using its
unique linguistic norms and shared references. It thus extends sociocultural theory into a defense of academic integrity, making the
assessment a reflection of the student’s participation within that specific learning community.
Authentic Assessment: The goal of authentic assessment is to mirror real-life tasks and challenges that require complex problem-
solving and knowledge application (McArthur, 2023; Ifelebuegu, 2023). However, many traditional "real-world" scenarios can be
solved by AI. CET-AIA enhances this theory by introducing a layer of authenticity that is AI-resistant: the use of informal, conver­
sational language that simulates the nuanced communication found in real-life human interaction, thereby assessing the human skill of
interpreting context.
These informal elements reflect the kind of linguistic and conceptual ambiguity found in real-world communication. In doing so,
CET-AIA not only tests knowledge but also evaluates students’ capacity to interpret nuance and context skills essential for real-life
problem solving. Furthermore, this approach helps educators distinguish between students who have internalized course material
and those who depend on external tools like AI. Based on these theoretical foundations, we hypothesize that:
H1. Students’ performance on assessments will initially decrease with the introduction of colloquial questioning, indicating that they
were unprepared for this type of questioning and relied more on AI assistance.

Fig. 2. Foundations of CET-AIA.

5
C. Kooli et al. Thinking Skills and Creativity 60 (2026) 102051

1.8. Rationale for transition

The rationale behind transitioning to a colloquial approach is a strategic response to the vulnerabilities of classical teaching
methods in an AI-rich environment. The following Table 1 contrasts the two approaches:
As generative AI systems become increasingly proficient at answering standardized, formalized MCQs, educators face the challenge
of designing assessments that prioritize student authenticity and independent thinking. The conventional format of assessment
questions typically formal, concise, and syntactically consistent aligns well with the input expectations of AI models like ChatGPT.
Colloquial questioning disrupts this alignment by introducing a "contextual barrier" that current generalized AI models are not
equipped to overcome. While modern LLMs are increasingly adept at parsing informal language, their effectiveness depends on access
to public training data. CET-AIA leverages this limitation by grounding questions in the closed, dynamic context of a specific classroom
referencing a spontaneous discussion or a shared experience that exists outside of the AI’s knowledge base. This approach creates
semantic and contextual ambiguity that AI struggles to resolve accurately. The barrier is therefore not merely linguistic but situational,
requiring a level of localized awareness that AI lacks. Moreover, this transition serves to promote deeper learning: when faced with
colloquial questions, students must engage with the content in a more reflective, context-sensitive manner.
Equally important, this shift promotes academic self-reliance. As students realize that AI cannot reliably decode colloquial prompts,
they are encouraged to invest more effort in attending lectures, participating in discussions, and developing personal interpretations of
course material. Thus, the rationale for colloquial questioning is not only defensive preventing cheating but also pedagogical fostering
ownership of learning. In line with this rationale, we propose:
H2. Students’ performance with colloquial questioning will gradually improve as they adjust to this new assessment type and rely
less on AI assistance.

1.9. Core components of CET-AIA

The CET-AIA framework is comprised of two interconnected components designed to promote academic integrity in the age of AI.
Firstly, the Colloquial Teaching Style emphasizes delivering content through informal, culturally relevant language, weaving in
specific classroom discussions, events, and references. This approach personalizes learning by embedding course material within
unique, lived experiences that standard AI models cannot replicate. Secondly, AI Awareness in Assessment leverages this colloquial
teaching style by constructing evaluations based on these specific in-class discussions and situational examples. This creates a form of
"assessment shielding" because AI tools like ChatGPT cannot access or interpret these dynamic, classroom-specific interactions. As a
result, students are required to rely on their own understanding, memory, and analytical reasoning, thereby directly addressing the
challenge of maintaining academic integrity. Fig. 3 compares the behavior of students and AI systems in responding to both traditional
and colloquial assessment formats, reinforcing CET-AIA’s ability to expose AI limitations and promote deeper student reasoning.

1.10. Key theoretical claims

Based on its core components, the CET-AIA framework makes four central claims about its impact on student learning and behavior.
Firstly, by implementing an AI-resistant assessment design that relies on classroom-specific content rather than general knowledge,
CET-AIA limits the utility of AI-generated answers, thereby protecting the integrity of assessments. This strategy directly fosters ac­
ademic self-reliance, discouraging dependence on generative AI tools and promoting independent thought and problem-solving.
Secondly, the framework posits that the informal and relatable nature of its colloquial teaching style enhances student engagement
by encouraging active listening and deeper cognitive processing during lectures. Lastly, this approach is claimed to cultivate authentic
comprehension, as students are better able to connect with instructional content that reflects their own linguistic and cultural realities,
which leads to improved understanding and retention of the material.

1.11. Development of CET-AIA

The CET-AIA framework was developed as a practical response to observed discrepancies in student performance and engagement
in AI-rich academic environments. The design integrates linguistic strategies with pedagogical theory to construct assessments that are
robust against AI exploitation while still aligning with educational goals.
At the core of CET-AIA is the use of colloquial multiple-choice questions; that employ informal, idiomatic, or culturally grounded

Table 1
Comparison of Classical and CET-AIA Methods.
Aspect Classical Method CET-AIA Method

Assessment Basis General content (books, slides, standardized tests) In-class, contextual, spontaneous and colloquial content
AI Vulnerability High – AI can often solve classical MCQs Low – AI cannot answer context-specific, colloquial questions
Student Learning Focus Memorization, grammar drills, formal structures Critical thinking, listening, engagement, situational understanding
Cultural Relevance Often decontextualized from student reality Tied to students’ lived experiences, idioms, and social language use
Student Motivation External (grades, exams) Internal (participation, presence, curiosity, ownership of learning)
Skill Development Formal writing, translation Functional communication, interpretation, real-life language application

6
C. Kooli et al. Thinking Skills and Creativity 60 (2026) 102051

Fig. 3. AI vs. Student Assessment Interaction.

language. These questions often include references to recent course discussions, specific examples, or shared experiences within a
class. For instance, a prompt might begin with, "In the last class, we talked about the prime minister visiting a manufacturing site.
Which factory was it?" This contextual dependency presents a significant barrier for AI systems, which typically lack access to such
dynamic, course-specific inputs.
The framework is organized around four key design principles. First, Contextual Anchoring ensures that assessments are rooted in
specific course content or events that only students attending the class would recognize. By referencing in-class examples, topical
discussions, or institutional settings, questions become context-specific and thus inaccessible to external AI tools. This practice also
strengthens students’ ability to make associations between learning and real-time classroom experiences.
Second, Linguistic Obfuscation involves the deliberate use of conversational phrasing, idioms, and informal or regionally specific
expressions. Such linguistic features are less predictable than formal academic language and therefore more difficult for AI models to
parse. For human students, however, these linguistic elements often enhance relatability and comprehension when tied to familiar
discourse patterns.
Third, Critical Inference is prioritized over rote recall. Questions are framed to require interpretation, judgment, and the drawing of
connections rather than simple factual recall. This promotes deeper learning and the development of higher-order thinking skills,
aligning with Bloom’s taxonomy and fostering cognitive engagement.
Finally, Cultural Relevance is emphasized through the incorporation of local or culturally significant examples, metaphors, and
scenarios. These elements help students from diverse backgrounds see their identities reflected in the learning process and create an
inclusive learning environment. Simultaneously, such references are difficult for AI systems trained on generalized datasets to interpret
accurately.
Together, these principles form a robust framework for designing assessments that foster higher-order thinking and are resistant to
technological shortcuts. The interplay of these design principles is expected to produce distinct and measurable outcomes for both
students and AI systems, leading to our primary research hypotheses. (Hypotheses H1 and H2 are stated previously)

7
C. Kooli et al. Thinking Skills and Creativity 60 (2026) 102051

1.12. Accordingly, we also hypothesize

H3. The AI model’s performance on assessments will significantly decrease with the introduction of colloquial questioning compared
to its performance on traditional questions.
To empirically test these three research hypotheses (H1, H2, and H3), our study employs a comparative quantitative design. For the
purpose of statistical validation, we formulated the following corresponding null hypotheses:
H0a: There is no significant difference in student engagement between conventional assessments and assessments using colloquial
questioning.
H0b: There is no significant difference in AI performance between conventional and colloquial assessment formats.
H0c: There is no significant difference in the reliability of student assessment results between traditional and CET-AIA formats.
By testing these null hypotheses, the study aims to statistically evaluate the effectiveness of CET-AIA in enhancing engagement,
reducing AI manipulation, and improving the validity of educational assessments. Fig. 4 illustrates the flow of CET-AIA imple­
mentation—from designing colloquial assessments to observing student and AI response behavior—highlighting how the framework

Fig. 4. CET-AIA Implementation Flow.

8
C. Kooli et al. Thinking Skills and Creativity 60 (2026) 102051

safeguards academic integrity while fostering engagement.


The following section presents the methods used to evaluate this framework empirically.

2. Methods

2.1. Research design

The current study employed a quasi-experimental design to evaluate the CET-AIA framework in a real-world educational setting.
This design, which included multiple cohorts subjected to alternating testing conditions, was chosen for its ecological validity, though
we acknowledge its limitations. Specifically, the lack of random assignment means that cohort effects cannot be entirely ruled out and
represent a potential confounding variable.
By analyzing both student and AI performance through quantitative methods, the study systematically assessed the effect of the
CET-AIA framework and its alignment with the theoretical constructs introduced earlier.

2.2. Data collection and selection rationale

Participants were divided into three distinct cohorts drawn from different academic disciplines and levels of study. Course 1
included 111 s-year business administration students. Course 2 comprised 85 s-year social sciences students. Course 3 involved 33
fourth-year social sciences students. This stratified sampling approach allowed for the examination of CET-AIA’s effects across varied
academic contexts, promoting generalizability.
ChatGPT 3.5 was the AI model selected for this study, reflecting current real-world use cases in educational settings. All cohorts
were assessed using multiple-choice quizzes administered at scheduled intervals throughout the academic term. Course 1 delivered
three quizzes, with colloquial questioning introduced in the second quiz without prior notice. Courses 2 and 3 administered ten quizzes
each, with colloquial questioning first appearing in the sixth quiz. In both cases, the following quizzes maintained the colloquial
format, allowing students to gradually adapt to the new style. To ensure fair comparisons, quiz content remained equivalent in dif­
ficulty and scope across formats.
Each quiz was also administered to ChatGPT 3.5 using identical questions and rubrics. This allowed for a direct performance
comparison between human students and AI under matched conditions.

2.3. Procedure

Students were informed at the start of the course that quizzes would include a new form of informal, context-sensitive questioning
designed to encourage deeper engagement. No restrictions were placed on the use of AI tools during the quizzes, which were completed
online and unsupervised. This open-access environment was intentionally maintained to observe the real-world impact of CET-AIA
under conditions where students might naturally resort to AI assistance.
In parallel, each quiz and midterm was submitted to ChatGPT-3.5 under controlled prompts. AI performance was measured for both
traditional and colloquial question types. This allowed for direct comparison of human and AI response patterns across assessment
formats.

2.4. Colloquial questioning

Colloquial questioning refers to the construction of multiple-choice items using informal language, real-world expressions, and a
conversational tone. These questions departed from traditional academic formats by referencing specific in-class examples, current
events, or content unique to course discussions. Designed to challenge AI capabilities, colloquial questions required contextual
interpretation and recall of materials unavailable to AI models.
For instance, one quiz asked: "In the last lecture, we discussed the example of the support offered by the federal government to
Canada’s provinces. Which of the following is not among the conditions set by the federal government to support the provinces in
financing their health care systems?" Another example included: "In the last lecture, we mentioned the name of a country that has
fallen into the trap of conditional development loans, creating disguised colonialism. The country in question is:"
These examples demonstrate the CET-AIA approach. Crucially, such questions are designed to assess content mastery through the
vehicle of a classroom anecdote, not merely to test recall of the anecdote itself. For instance, the question about "disguised colonialism"
requires the student to connect a memorable classroom example back to the broader academic concept being taught. This design
reinforces student engagement by valuing their presence and requiring critical reflection on course material, while simultaneously
reducing AI utility.

2.5. Data collection and analysis

Quantitative data were collected on student scores for each quiz item and midterm question. These data were used to assess trends
in performance over time, across both traditional and colloquial questions. Repeated-measures ANOVA was employed to evaluate the
significance of performance changes. Engagement levels were indirectly measured through consistency in quiz participation, response
completeness, and variation in response selection.

9
C. Kooli et al. Thinking Skills and Creativity 60 (2026) 102051

To gauge student engagement, this study utilized indirect behavioral indicators, including consistency in quiz participation,
response completeness, and variation in response selection. It is important to note that these metrics serve as a proxy for engagement,
as they primarily measure participation and persistence rather than the cognitive and affective dimensions of student involvement.
This approach was chosen to provide a preliminary, quantitative snapshot of student behavior in response to the new assessment
format. Future studies could build upon these findings by incorporating more direct measures such as student surveys, qualitative
feedback, or think-aloud protocols to gain a richer understanding of the cognitive engagement fostered by the CET-AIA framework.
AI performance was similarly analyzed, with correct response rates and error types coded for both question formats. The perfor­
mance of ChatGPT-3.5 was benchmarked against student averages for each assessment point.

2.6. Statistical tests were used to assess the null hypotheses

• H0a: Comparing student engagement metrics between traditional and colloquial questions.
• H0b: Comparing AI performance on traditional vs. colloquial MCQs.
• H0c: Assessing the reliability (e.g., discriminatory power and internal consistency) of traditional vs. colloquial assessments.

3. Data analysis

3.1. Descriptive statistics

Descriptive statistics were calculated for each quiz to summarize performance, including mean scores, medians, standard de­
viations, and score distributions. This step provided an overview of both student and AI performance under each question format.

3.2. Inferential statistics

Inferential analyses included paired and independent t-tests to compare student performance before and after the introduction of
colloquial questioning. One-way and repeated-measures ANOVA were used to detect trends across quizzes and cohorts. These tests
helped determine whether differences in performance were statistically significant.

3.3. Correlation and regression analysis

Correlation analysis was used to assess the relationship between early quiz performance and outcomes on later assessments,
including midterms. Regression models evaluated the predictive value of early performance indicators for student adaptation under
colloquial formats.

3.4. Mixed-Effects models

To control for inter-cohort and individual variability, mixed-effects models were applied. These models accounted for both fixed
effects (question type, cohort, timing) and random effects (individual performance patterns), offering a nuanced view of student
adaptation over time.
Cluster Analysis and AI Performance Evaluation. Cluster analysis was employed to identify patterns in how students adapted to
colloquial questioning, categorizing learners by response trends and improvement curves. AI performance was separately analyzed by
comparing accuracy across traditional and colloquial formats. Discrepancies in AI scores provided empirical evidence of the CET-AIA
framework’s effectiveness in reducing AI utility for assessment purposes.

3.5. Ethical considerations

This study was reviewed and approved by the Institutional Review Board of the university of Ottawa under protocol number [S-
11–24–11,042]. All participants were informed of the study’s purpose and provided informed consent for the use of anonymized data.
Students were advised that AI tools could be used during unsupervised assessments, and transparency regarding the experimental
framework was maintained throughout the academic term.

Table 2
Student Performance Summary by Course and Quiz Format.
Course Quiz Format Mean Score (%) Standard Deviation Sample Size

Course 1 Traditional 71.3 8.2 111


Course 1 Colloquial 62.5 7.9 111
Course 2 Traditional 75.2 9.1 85
Course 2 Colloquial 70.8 → 74.6 8.5 → 7.6 85
Course 3 Traditional 77.5 6.8 33
Course 3 Colloquial 71.4 → 76.1 7.2 → 6.5 33

10
C. Kooli et al. Thinking Skills and Creativity 60 (2026) 102051

4. Results

This section presents the statistical outcomes of the study, structured around the hypotheses of the CET-AIA framework. The
findings cover student performance trends, AI response accuracy, student adaptation patterns, and the psychometric properties of
colloquial assessments.

4.1. Student performance trends: an initial drop and gradual adaptation

To test H1, which predicted that student performance would initially decrease with the introduction of colloquial questioning, we
analyzed quiz scores immediately following the transition. As detailed in Table 2, the data support this hypothesis, as a drop in
performance was observed across all three courses. This was especially pronounced in Course 1, where the second quiz administered
without advance notice and containing colloquial questions—resulted in a significant reduction in average scores (mean score drop
from 71.3 % to 62.5 %, p< 0.05). This result supports H1, indicating that students were unprepared for the shift and may have
previously relied on AI assistance or habitual expectations associated with traditional MCQ formats.
In Courses 2 and 3, similar performance declines were observed when colloquial questions were introduced in the sixth quiz. This
initial decline was followed by an adjustment period. To test H2, which predicted that student performance would gradually improve,
we tracked scores over subsequent colloquial quizzes. The data show a clear recovery trend, with means stabilizing by the final quizzes
at levels statistically comparable to pre-colloquial periods. For example, as shown in Table 2, scores in Course 2 rose from 70.8 % to
74.6 %. A repeated measures ANOVA confirmed this adaptation curve was statistically significant (F(4, 224) = 5.87, p< 0.001),
supporting H2 by validating the presence of a student adjustment process.

4.2. AI performance and assessment integrity

In line with H3, which hypothesized that the AI model’s performance would decrease on colloquial questions, our analysis revealed
a stark performance gap. Testing the corresponding null hypothesis (H0b), the study confirmed a significant difference in accuracy
between question formats.
As highlighted in Table 3, ChatGPT-3.5′s performance was notably higher on traditional questions compared to colloquial ones
across all quizzes. In Course 2, for example, the AI’s accuracy dropped from an average of 87 % on conventional MCQs to just 52 % on
colloquial ones. This performance gap was consistent across disciplines and question themes. A t-test comparing AI scores between
question types yielded a significant result (t(18) = 4.11, p< 0.01), providing strong support for H3 and refuting H0b.
Further error analysis showed that AI misinterpretations of colloquial phrasing, idiomatic references, or context-specific prompts
contributed to a higher rate of incorrect responses. These results validate the CET-AIA framework’s assumption that informal language
can disrupt AI pattern recognition.

4.3. Student engagement and adaptation profiles

In testing H0a, behavioral metrics like quiz participation and question completion rates remained stable across both assessment
formats, indicating no significant drop in student engagement. However, a more nuanced picture emerged from a cluster analysis,
which identified three distinct student adaptation profiles:

• Quick Adapters (38 %): Showed performance recovery within two quizzes.
• Gradual Improvers (44 %): Needed at least three colloquial quizzes to stabilize their scores.
• Performance Decliners (18 %): Experienced an ongoing decline, suggesting a continued reliance on ineffective methods or a lack
of adjustment.

The results shown in Table 4 suggest that while overall engagement was consistent, the quality of cognitive adaptation varied across
the student population.

4.4. Assessment reliability and validity

Table 5 results reject the null hypothesis H0c, demonstrating that colloquial quizzes designed under the CET-AIA framework are
psychometrically robust. The colloquial assessments showed stronger internal consistency (Cronbach’s alpha ¼ 0.81) compared to
traditional formats (alpha ¼ 0.73). Furthermore, their item discrimination indices were higher, indicating they were better at
differentiating between high- and low-performing students. Regression analysis also revealed that performance on colloquial quizzes

Table 3
AI vs. Student Performance on Traditional vs. Colloquial MCQs.
Format Students’ Avg. Score AI (ChatGPT 3.5) Score

Traditional 74.3 % 87.0 %


Colloquial 72.1 % (end) 52.0 %

11
C. Kooli et al. Thinking Skills and Creativity 60 (2026) 102051

Table 4
Cluster Analysis of Student Adaptation.
Cluster % of Students Performance Pattern

Quick Adapters 38 % Decline followed by recovery within 2 quizzes


Gradual Improvers 44 % Slower improvement, stabilizing after 3 quizzes
Performance Decliners 18 % Ongoing performance decline

Table 5
Reliability and Predictive Validity Measures.
Format Cronbach’s Alpha Item Discrimination Index Predictive Power (β)

Traditional 0.73 0.38 0.12 (NS)


Colloquial 0.81 0.49 0.43 (p< 0.01)

was a significant predictor of midterm results (β ¼ 0.43, p< 0.01), a correlation not found with traditional quizzes. This suggests that
the CET-AIA framework not only enhances assessment validity but also promotes a more transferable understanding of course
material.
Overall, the results strongly support the CET-AIA framework’s theoretical and practical contributions. The next section discusses
the implications of these findings for assessment design, student learning strategies, and educational policy.

5. Discussion

This study evaluated the Colloquial Engagement Theory with AI Awareness (CET-AIA) framework by assessing the impacts of
colloquial multiple-choice questioning on student performance, academic integrity, and AI model accuracy. The results support the
framework’s theoretical assumptions and validate its practical implementation across diverse academic settings.

5.1. Interpretation of findings in light of theories

The initial drop in student scores provides preliminary support for H1. One interpretation is that this performance disruption
suggests students were unprepared for the linguistic shift, potentially because they were accustomed to leveraging AI tools (MacLeod
et al., 2022). However, other factors must be considered. The performance dip could also be attributed to a novelty effect, where any
unfamiliar format would temporarily increase cognitive load. Acknowledging these alternative explanations provides a more balanced
interpretation.
As predicted in H2, a gradual improvement in scores was observed in the quizzes following the transition, indicating that students
began adjusting to the colloquial style. This recovery, evidenced in the repeated-measures ANOVA and regression outputs, supports
Cognitive Load Theory (Skulmowski & Xu, 2022), particularly the notion that once students overcome the intrinsic complexity
introduced by novel phrasing, they are better able to reallocate cognitive resources to deeper learning. The pattern also aligns with
Authentic Assessment theory (McArthur, 2023), which advocates for real-world, context-anchored evaluations that require inter­
pretive thinking rather than rote recall.
In parallel, the performance of ChatGPT-3.5 dropped markedly on colloquial MCQs, confirming H3. The AI model performed well
on traditionally formatted questions, as found in previous studies (Humble et al., 2024; Kooli & Yusuf, 2025), but struggled with
context-bound or informal phrasing, consistent with findings by Babaian and Xu (2024) and Mei et al. (2024). This result highlights a
fundamental limitation of current AI: its reliance on pattern recognition and formal grammar constrains its ability to process natu­
ralistic or situation-specific prompts.
Importantly, while behavioral indicators such as participation and completion rates remained stable, indicating no substantial loss
of engagement, performance trends revealed three adaptation profiles: quick adapters, gradual improvers, and consistent underper­
formers. This segmentation underscores that engagement, though superficially uniform, varied in qualitative depth, supporting the
sociocultural argument that context, identity, and cognitive framing influence learning behaviors (Taylor, 2022; Kilag et al., 2024).

5.2. Assessment quality and diagnostic power

From a psychometric perspective, the results from this pilot study suggest that colloquial language has the potential to enhance
assessment reliability and validity. The observed higher internal consistency (α = 0.81 vs. 0.73) and stronger item discrimination
indices suggest that the CET-AIA framework can contribute to a more authentic measure of student performance, thereby helping to
safeguard the validity of the assessment. This effectiveness stems from several key factors. Firstly, Authentic Engagement as Authentic
Assessment means that CET-AIA evaluates students’ true internalization of knowledge acquired through active classroom participation
and listening. Because colloquial content and examples are often spontaneous and context-bound, students must demonstrate genuine
understanding rather than mere recall of memorized facts, thereby tapping into deeper cognitive processing. Secondly, AI Resistance as
a Validity Shield directly addresses the vulnerability of traditional assessments to AI tools like ChatGPT. By embedding questions within

12
C. Kooli et al. Thinking Skills and Creativity 60 (2026) 102051

non-replicable in-class moments, CET-AIA ensures that only attentive and engaged students can succeed, thus safeguarding the validity
of the performance measure and restoring academic integrity. Finally, Increased Student Accountability is a crucial outcome, as students
quickly learn that their assessments will be based on in-class discussions and content. This fosters greater responsibility for their
learning, promoting improved listening skills, more thorough note-taking, and real-time comprehension—all vital components of
genuine academic development. These combined factors account for the demonstrably higher internal consistency (α = 0.81 vs. 0.73)
and stronger item discrimination indices observed in colloquial quizzes during our study, suggesting enhanced differentiation between
high- and low-performing students. These results reject H0c and support the idea that CET-AIA fosters more meaningful evaluation by
minimizing overreliance on automated aids and emphasizing applied comprehension. Furthermore, colloquial quiz performance
significantly predicted midterm success (β = 0.43, p< 0.01), unlike traditional formats, which aligns with earlier research emphasizing
the value of context-rich tasks in reinforcing learning transfer (Celik et al., 2022; Ouyang et al., 2022).
This contribution also strengthens emerging work by Griesbeck et al. (2024), who call for pedagogical models that both harness and
resist AI in assessment settings. CET-AIA responds to this dual imperative by embedding cultural relevance, conversational syntax, and
context-specific cues into question design elements that simultaneously engage learners and resist automation. As shown in Fig. 5
CET-AIA represents a comprehensive framework integrating ethical pedagogy, AI resistance, and student engagement into assessment
practice.

5.3. Practical implications and scalability

The findings of this exploratory study offer several practical implications. For instructors, CET-AIA provides a concrete strategy for
designing assessments that encourage class attendance and active listening. It shifts the focus from policing AI use to redesigning
pedagogy to reward human engagement. However, the framework’s successful implementation requires a significant investment of
time and creative effort from instructors.
To address the critical issues of workload and scalability, institutions should consider providing targeted support, such as pro­
fessional development and resources for faculty. The creation of collaborative platforms or shared question banks could also help
educators implement this framework more efficiently. For policymakers, this research underscores the need to move beyond

Fig. 5. CET-AIA Framework (Overview).

13
C. Kooli et al. Thinking Skills and Creativity 60 (2026) 102051

technological solutions and toward supporting pedagogical frameworks that foster a culture of academic integrity from the ground up.

5.4. Limitations

Despite the strength of the findings, several limitations must be acknowledged. First, the quasi-experimental design limited full
control over external variables such as instructor behavior or content pacing. Second, the study’s findings are drawn from courses in
the social sciences and business administration, which limits their immediate generalizability to other academic fields, particularly
STEM disciplines. While the core principles of CET-AIA could be adapted, further research is required to determine its effectiveness in
these less language-dependent contexts. Third, the AI performance benchmark was ChatGPT-3.5, which was representative of widely
accessible AI at the time of testing. However, the field of generative AI is advancing at an exceptional pace. The effectiveness of CET-
AIA hinges on a ’moving target,’ as newer models may be less susceptible to the colloquial and context-dependent phrasing used in this
study. Therefore, the long-term viability of this framework depends on continuous innovation in question design to stay ahead of AI
capabilities.
Additionally, the short adjustment window (i.e., only a few quizzes after the transition) limits conclusions about long-term skill
acquisition or retention. While adaptation occurred, it is unclear whether it reflects deeper conceptual learning or tactical test
familiarity.
Finally, the manuscript does not address the practical implications of instructor workload and scalability. Designing effective
colloquial questions that are context-specific, culturally relevant, unambiguous for students, and yet challenging for AI requires
considerable time, training, and creative effort beyond that of traditional multiple-choice question design. For the CET-AIA framework
to be adopted widely, institutions would need to consider providing faculty with professional development and resources. Future
studies should also investigate methods for creating question banks or collaborative platforms to help educators implement this
framework more efficiently without placing an undue burden on individuals.

5.5. Future research directions

Future research should move beyond quantitative metrics to include qualitative and mixed-methods approaches. Incorporating
student surveys, focus groups, and think-aloud protocols would provide invaluable insight into the cognitive and affective dimensions
of student engagement. Longitudinal applications of CET-AIA are also needed to determine whether it cultivates durable improve­
ments in critical thinking and academic honesty. Furthermore, the adaptability of CET-AIA should be explored by applying its prin­
ciples to other formats, such as oral or multimedia assessments. Finally, continued experimentation with other AI models (e.g., GPT-4,
Claude, Gemini) is essential to test the durability of colloquial resistance strategies as machine understanding improves.

6. Conclusion

This study introduced and evaluated the Colloquial Engagement Theory with AI Awareness (CET-AIA) as a novel approach to
assessment design in the age of generative artificial intelligence. By embedding informal language, contextual cues, and culturally
relevant phrasing into multiple-choice questions, CET-AIA seeks to achieve two complementary goals: fostering deeper student
engagement and reducing the effectiveness of AI-based cheating.
Empirical findings from three distinct cohorts demonstrated that students initially performed worse on colloquial quizzes, con­
firming that this assessment format disrupted habitual answering patterns and potentially limited the utility of AI assistance. However,
consistent with theoretical expectations from Constructivism and Cognitive Load Theory, students adapted over time, showing gradual
performance improvement and, in some cases, outperforming earlier benchmarks. In parallel, ChatGPT-3.5 struggled to interpret
colloquial phrasing, displaying significantly lower accuracy than on traditionally formatted items. This performance gap affirmed the
framework’s core design assumption: that linguistic and contextual complexity can serve as an ethical barrier against AI misuse.
Moreover, assessments developed under CET-AIA exhibited higher reliability, better discrimination power, and stronger predictive
validity for student outcomes. These results suggest that the framework is not only a defense against automation but also a forward-
looking model for more meaningful and inclusive evaluation. Rather than attempting to eliminate AI from the classroom an
increasingly impractical goal this study presents CET-AIA as an initial, adaptable framework for reframing the assessment process to be
more resilient, engaging, and human-centered. It is not a fixed solution, but a pedagogical approach designed to evolve in response to
technological advancements.
In conclusion, this research provides preliminary validation for colloquial questioning as a promising strategy for enhancing
assessment integrity in the digital age. By prioritizing the unique, shared context of the classroom, CET-AIA offers a pathway toward
assessments that not only mitigate AI misuse but also foster the critical and flexible thinking skills essential for genuine learning. As
educational institutions continue to navigate their relationship with AI, adaptable pedagogical frameworks like CET-AIA will be crucial
for building ethically resilient and cognitively rich learning environments.

Declaration of generative ai and ai-assisted technologies in the writing process

While preparing this work, the authors used ChatGPT to improve readability and language. After using this tool, the authors
reviewed and edited the content as needed. The author takes full responsibility for the publication’s content.

14
C. Kooli et al. Thinking Skills and Creativity 60 (2026) 102051

Declaration of generative AI in writing

The authors used ChatGPT to improve readability and language. After using this tool, the authors reviewed and edited the content
as needed and take full responsibility for the publication’s content.

Ethics approval

This study was reviewed and approved by the Institutional Review Board of the University of Ottawa under protocol number [S-
11–24–11,042].

CRediT authorship contribution statement

Chokri Kooli: Writing – review & editing, Formal analysis, Data curation, Conceptualization. Nadia Yusuf: Writing – original
draft, Validation, Methodology, Investigation. Mohammed Y. Sarhan: Visualization, Validation, Software, Resources.

Declaration of competing interest

The authors declare that they have no known competing financial interests or personal relationships that could have appeared to
influence the work reported in this paper.

Data availability

Data supporting the findings of this study are available from the corresponding author on reasonable request.

References

Adıgüzel, T., Kaya, M. H., & Cansu, F. K (2023). Revolutionizing education with AI: Exploring the transformative potential of ChatGPT. Contemporary Educational
Technology.
Alrakhawi, H.A., Jamiat, N.U.R.U.L.L.I.Z.A.M., & Abu-Naser, S.S. (2023). Intelligent Tutoring Systems in education: A systematic review of usage, too.
Babaian, T., & Xu, J. (2024). Entity recognition from colloquial text. Decision Support Systems, 179, Article 114172.
Bankins, S., & Formosa, P. (2023). The ethical implications of artificial intelligence (AI) for meaningful work. Journal of Business Ethics, 185(4), 725–740.
Bicer, A., Aldemir, T., Krall, G., Quiroz, F., Chamberlin, S., Nelson, J. L., & Kwon, H. (2024). Exploring creativity in mathematics assessment: An analysis of
standardized tests. Thinking Skills and Creativity, 54, Article 101652.
Celik, I., Dindar, M., Muukkonen, H., & Järvelä, S. (2022). The promises and challenges of artificial intelligence for teachers: A systematic review of research.
TechTrends, 66(4), 616–630.
Cotton, D. R., Cotton, P. A., & Shipway, J. R (2024). Chatting and cheating: Ensuring academic integrity in the era of ChatGPT. Innovations in education and teaching
international, 61(2), 228–239.
Creely, E., & Blannin, J. (2025). Creative partnerships with generative AI. Possibilities for education and beyond. Thinking Skills and Creativity, 56, Article 101727.
Griesbeck, A., Zrenner, J., Moreira, A., & Au-Yong-Oliveira, M. (2024). AI in higher education: Assessing acceptance, learning enhancement, and ethical considerations
among university students. in world conference on information systems and technologies (pp. 214-227). March. Cham: Springer Nature Switzerland.
Habib, S., Vogel, T., & Thorne, E. (2025). Student perspectives on creative pedagogy: Considerations for the age of AI. Thinking Skills and Creativity, 56, Article 101767.
Holmes, W., Porayska-Pomsta, K., Holstein, K., Sutherland, E., Baker, T., Shum, S. B., & Koedinger, K. R (2022). Ethics of AI in education: Towards a community-wide
framework. International Journal of Artificial Intelligence in Education, 1–23.
Hridi, A. P., Sahay, R., Hosseinalipour, S., & Akram, B. (2024). Revolutionizing AI-assisted education with federated learning: A pathway to distributed, privacy-
preserving, and debiased learning ecosystems. In , 3. Proceedings of the AAAI Symposium Series (pp. 297–303).
Humble, N., Boustedt, J., Holmgren, H., Milutinovic, G., Seipel, S., & Östberg, A. S (2024). Cheaters or ai-enhanced learners: Consequences of chatgpt for
programming education. Electronic Journal of e-Learning, 22(2), 16–29.
Ifelebuegu, A. (2023). Rethinking online assessment strategies: Authenticity versus AI chatbot intervention. Journal of Applied Learning and Teaching, 6(2).
Igbokwe, I. C (2023). Application of artificial intelligence (AI) in educational management. International Journal of Scientific and Research Publications, 13(3), 300–307.
Jukiewicz, M. (2024). The future of grading programming assignments in education: The role of ChatGPT in automating the assessment and feedback process. Thinking
Skills and Creativity, 52, Article 101522.
Kilag, O. K. T., Maghanoy, D. A. F., Calzada-Seraña, K. R. D. D., & Ponte, R. B (2024). Integrating Lev Vygotsky’s sociocultural theory into online instruction: A case
Study. European Journal of Learning on History and Social Sciences, 1(1), 8–15.
Kooli, C. (2023). Chatbots in education and research: A critical examination of ethical implications and solutions. Sustainability, 15(7), 5614.
Kooli, C., & Yusuf, N. (2025). Transforming educational assessment: Insights into the use of ChatGPT and large language models in grading. International Journal of
Human–Computer Interaction, 41(5), 3388–3399.
Kooli, C., Kooli, Y., & Kooli, E. (2025). Generative artificial intelligence addiction syndrome: A new behavioral disorder? Asian Journal of Psychiatry, 107, Article
104476.
Limo, F. A. F., Tiza, D. R. H., Roque, M. M., Herrera, E. E., Murillo, J. P. M., Huallpa, J. J., & Gonzáles, J. L. A (2023). Personalized tutoring: ChatGPT as a virtual tutor
for personalized learning experiences. Przestrzeń Społeczna (Social Space), 23(1), 293–312.
MacLeod, A., Burm, S., & Mann, K. (2022). Constructivism: Learning theories and approaches to research. Researching medical education, 25–40.
Markauskaite, L., Marrone, R., Poquet, O., Knight, S., Martinez-Maldonado, R., Howard, S., & Siemens, G. (2022). Rethinking the entwinement between artificial
intelligence and human learning: What capabilities do learners need for a world with AI? Computers and Education: Artificial Intelligence, 3, Article 100056.
McArthur, J. (2023). Rethinking authentic assessment: Work, well-being, and society. Higher education, 85(1), 85–101.
Mei, L., Liu, S., Wang, Y., Bi, B., & Chen, X. (2024). SLANG: New concept comprehension of large language models. arXiv preprint. arXiv:2401.12585.
Mogavi, R. H., Deng, C., Kim, J. J., Zhou, P., Kwon, Y. D., Metwally, A. H. S., & Hui, P. (2023). Exploring user perspectives on chatgpt: Applications, perceptions, and
implications for ai-integrated education. arXiv preprint. arXiv:2305.13114.
Nguyen, A., Ngo, H. N., Hong, Y., Dang, B., & Nguyen, B. P. T (2023). Ethical principles for artificial intelligence in education. Education and Information Technologies,
28(4), 4221–4241.
O’Connor, K. (2022). Constructivism, curriculum and the knowledge question: Tensions and challenges for higher education. Studies in Higher Education, 47(2),
412–422.

15
C. Kooli et al. Thinking Skills and Creativity 60 (2026) 102051

O’Connor, S., & Liu, H. (2023). Gender bias perpetuation and mitigation in AI technologies: Challenges and opportunities. AI & SOCIETY, 1–13.
Oravec, J. A (2022). AI, biometric analysis, and emerging cheating detection systems: The engineering of academic integrity? Education Policy Analysis Archives, 30
(175), n175.
Oravec, J. A (2023). Artificial intelligence implications for academic cheating: Expanding the dimensions of responsible human-AI collaboration with ChatGPT.
Journal of Interactive Learning Research, 34(2), 213–237.
Ouyang, F., Zheng, L., & Jiao, P. (2022). Artificial intelligence in online higher education: A systematic review of empirical research from 2011 to 2020. Education and
Information Technologies, 27(6), 7893–7925.
Owan, V. J., Abang, K. B., Idika, D. O., Etta, E. O., & Bassey, B. A (2023). Exploring the potential of artificial intelligence tools in educational measurement and
assessment. EURASIA Journal of Mathematics, Science and Technology Education, 19(8), em2307.
Parycek, P., Schmid, V., & Novak, A. S (2023). Artificial intelligence (AI) and automation in administrative procedures: Potentials, limitations, and framework
conditions. Journal of the Knowledge Economy, 1–26.
Perkins, M. (2023). Academic integrity considerations of AI large language models in the post-pandemic era: ChatGPT and beyond. Journal of University Teaching &
Learning Practice, 20(2).
Ray, P. P (2023). ChatGPT: A comprehensive review on background, applications, key challenges, bias, ethics, limitations and future scope. Internet of Things and
Cyber-Physical Systems, 3, 121–154.
Rane, N., Choudhary, S., & Rane, J. (2023). Education 4.0 and 5.0: Integrating artificial Intelligence (AI) for personalized and adaptive learning. Available at SSRN
4638365.
Roshanaei, M. (2024). Towards best practices for mitigating artificial intelligence implicit bias in shaping diversity, inclusion and equity in higher education.
Education and Information Technologies, 1–26.
Sebastian, G. (2023). Privacy and data protection in chatgpt and other ai chatbots: Strategies for securing user information. International Journal of Security and Privacy
in Pervasive Computing (IJSPPC), 15(1), 1–14.
Skulmowski, A., & Xu, K. M (2022). Understanding cognitive load in digital and online learning: A new perspective on extraneous cognitive load. Educational
psychology review, 34(1), 171–196.
Smolansky, A., Cram, A., Raduescu, C., Zeivots, S., Huber, E., & Kizilcec, R. F (2023). Educator and student perspectives on the impact of generative AI on assessments
in higher education. In Proceedings of the tenth ACM conference on Learning@ Scale (pp. 378–382).
Su, J., & Yang, W. (2023). Unlocking the power of ChatGPT: A framework for applying generative AI in education. ECNU Review of Education, 6(3), 355–366.
Sweeney, S. (2023). Who wrote this? Essay mills and assessment–Considerations regarding contract cheating and AI in higher education. The International Journal of
Management Education, 21(2), Article 100818.
Taylor, C. S (2022). Culturally and socially responsible assessment: Theory, research, and practice. Teachers College Press.
Tuomi, I. (2024). Beyond mastery: Toward a broader understanding of ai in education. International Journal of Artificial Intelligence in Education, 34(1), 20–30.
Xu, W., & Ouyang, F. (2022). A systematic review of AI role in the educational system based on a proposed conceptual framework. Education and Information
Technologies, 27(3), 4195–4223.
Yusuf, A., Bello, S., Pervin, N., & Tukur, A. K (2024). Implementing a proposed framework for enhancing critical thinking skills in synthesizing AI-generated texts.
Thinking Skills and Creativity, 53, Article 101619.

16

You might also like