0% found this document useful (0 votes)
10 views3 pages

ESL Study: MLLM Impact on Learning

The document outlines a study focused on the effectiveness of Multilingual Large Language Models (MLLM) in ESL classrooms, utilizing a pre- and post-treatment survey to compare explanations given by an ESL teacher and those generated by ChatGPT-4o. The study was conducted with participants from the Chester County Opportunities Industrialization Center, targeting a specific topic of superlatives and comparatives, and involved multiple choice tests to assess understanding. Limitations include a small sample size and varying participant skill levels, which may affect the generalizability of the findings.

Uploaded by

usatu ding
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
10 views3 pages

ESL Study: MLLM Impact on Learning

The document outlines a study focused on the effectiveness of Multilingual Large Language Models (MLLM) in ESL classrooms, utilizing a pre- and post-treatment survey to compare explanations given by an ESL teacher and those generated by ChatGPT-4o. The study was conducted with participants from the Chester County Opportunities Industrialization Center, targeting a specific topic of superlatives and comparatives, and involved multiple choice tests to assess understanding. Limitations include a small sample size and varying participant skill levels, which may affect the generalizability of the findings.

Uploaded by

usatu ding
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Methods Section Guide & Assignment

Methods:
Explanation and alignment with study
Because my study’s focus was on Multilingual Large Language Models used in ESL classrooms, I
decided that a pre- and post-treatment survey would be the best course of action to effectively
obtain data for which I could compare MLLM explanations to ones given in ESL classrooms. By
implementing a pre- and post-treatment survey similar to (Leinonen et al.), I was able to find an
effective way to consolidate testing into one session and also maximize data from responses. I
decided to use multiple choice questions instead of Likert scale as they are a more concrete
way to gauge participant understanding of a topic, and decided to use explanations from the
curriculum as opposed to student explanations because the students would not be experts in
the topics they are studying.
Multilingual Large Language Models
In order to maximize practicality of the study and attain as much in-depth data as possible, I
chose Open AI’s ChatGPT-4o model as the only Multilingual Large Language Model used to
generate explanations for the survey. We chose ChatGPT-4o because of its strong recognition
as a leading MLLM for teaching in the ESL classroom (cite here
[Link] and its relatively high
accessibility among [internet users]. [Maybe add some more here]

In this study, I reached out to the Chester County Opportunities Industrialization Center (OIC), a
non-profit organization providing for economically disadvantaged individuals throughout the
community. The OIC offers free English Second Language (ESL) classes to adults with locations
throughout Pennsylvania, teaching English to improve their daily lives or abilities to retain a job.
After establishing contact with the Director of the Workforce Education department, he
connected me to an ESL teacher and data administrator to work out the finer details of the
study. I worked with all three to plan, prepare, and execute the necessary components of the
study.

Survey
I initially planned to create a pre-post treatment survey for my evaluation. After discussing with
the ESL teacher and data administrator, we collectively decided that to comply with privacy
restrictions and time limitations, we would choose a single topic in ESL education that students
in the specific class were struggling with, design similar pre and post surveys to evaluate
students’ understanding and provide two different treatments in the form of explanations.
The first explanation was a lesson designed to mimic the ESL class lecture as closely as possible.
We decided to choose the topic of superlatives and comparatives, something that the ESL
teacher noticed students clearly struggled with, and constructed an explanation of the subject
using classroom resources (Burlington Core’s High Beginners textbook) in both English and
Spanish.
The second explanation was a lesson generated by ChatGPT-4o. To minimize any prior data
ChatGPT-4o retained from previous work done on my device, I logged out of all accounts and
accessed the website using the free no-login page on an incognito tab. In order to best
represent the average prompt teachers may use for designing lesson plans or creating
materials, I aimed to design the prompt to be simple yet efficient. The prompt used for the
second explanation was “Generate a multilingual explanation of comparative and superlative
adjective grammar for a high level beginner English speaker in Spanish”.
The survey was structured as follows: the first explanation, followed by a multiple choice test,
then the second explanation, followed by a second multiple choice test, and finally a free
response question asking participants to describe in their preferred language which explanation
was more useful and why.

Multiple Choice
Each multiple choice test consisted of 7 multiple choice questions from the Oxford University
Press (cite [Link]
cc=us&selLanguage=en). To ensure balanced difficulty between the two tests, I divided the total
fifteen questions in Test 7 (Comparatives and Superlatives) systematically. By alternating
questions between the two tests, I split the fifteen questions into seven for the first test and
seven for the second test, omitting question fifteen. Each test had parallel question structure
and was approved by the ESL teacher.

Randomization and Anonymity of Participants


To adhere to the regulations of both Conestoga High School’s Internal Review Board and the
OIC, participants were required to sign forms of consent containing an overview of the study,
possible benefits and risks, confidentiality clauses and a voluntary participation statement
(appendix). To ensure complete protection of participant privacy, we assigned randomized
participant IDs to each participant at the time of the survey, which was the sole identification
label on any of the documents signed by each participant. We also removed any form of written
signature needed and replaced it with a check box to maintain complete anonymity. No photos
or media were taken at the facility. As a final precaution, all documents were immediately
stored in a locked cabinet and shredded once transcribed by me.

Study execution
The study was executed at the Spanish ESL class hosted at a Wester Chester public high school
at 6:30PM. I printed eight copies of each survey and consent form for each participant and
carefully filed them in a manila folder. After arriving into the classroom, the ESL teacher and I
met up to let in the students and began to explain the study immediately after all students
came into the classroom. After the ESL teacher translated my introduction and explanation of
the study’s purpose to the students, I randomly assigned participant IDs for each student and
handed out forms of consent. After collecting forms from students who consented, I distributed
surveys labeled with participant ID. Participants were allowed to ask the ESL teacher for any
clarification on survey questions related to the meaning of the question, however the ESL
teacher was not able to advise or help in any way. All eight students completed and submitted
their survey within the allocated thirty minutes.

Limitations
This study was performed using a convenience sample targeting ESL participants of all
backgrounds. [As the sample size (8) of the first survey was small and less than 30, the Central
Limit theorem states that I am unable to approximate the sample mean as a normal
distribution.] Additionally, I would have preferred to sample from multiple ESL classes, and I am
still working on that with my contacts at the OIC. Results from the survey may vary among
different topics and languages, meaning the findings from the first sample in a Spanish-speaking
class may not be interpreted the same way among other language speakers, and
interpretations of my data will have limited generalizability to other cases not matching these
circumstances.
Another limitation was the differing skill levels of each participant; all participants were from
the same level class (high beginner), however some participants may have had higher fluency in
English or Spanish than others. Not being able to understand parts of the survey would have
resulted in false results that did not match the reality of the participants, and thus affected the
data I collected.

You might also like