According to Ali Khhan et al.
(2023), a well-defined research methodology is crucial for
ensuring the validity, reliability, and credibility of findings in order to answer research
questions or investigate specific problems. In this chapter, the research methodology used in
the study was mentioned. It is divided into … parts which include
2.1 Research question
This study aims to measure the lexical complexity of texts generated by human
writers and ChatGPT 3.5. Specifically, it seeks to address the following research
questions:
How do the lexical complexity levels of human-written texts compare to those
generated by ChatGPT 3.5?
What are the most significant differences in lexical diversity, density, and
sophistication between these two text sources?
How do these differences vary across different writing prompts or academic
genres?
2.3 Research design
Research can be categorized into three main types: quatitative, qulitative, and mixed methods.
This paper adopting a quantitative method and corpus-based comparative study stems from
the objective nature of the research question. The quantitative method can be defined as a
method used to gather data that can be measured using objective and measurement tools.
(Barella et al.,2024), According to Rashid (2022), a quantitative research is also carried out
using both traditional and modern scientific methods, generating numerical data and typically
aiming to establish causal relationships between two or more variables. It aims to prove the
hypothesis produced in a specific study by collecting data, analyzing, and interpreting
quantifiable data.
Quantitative analysis provides the methodological rigor necessary for systematically
evaluating linguistic features. This approach allows for the measurement of specific variables,
enabling clear comparisons between the two types of essays. Additionally, a corpus-based
approach facilitates empirical examination of authentic texts that are representative of both
human and machine-generated outputs. By analyzing a large and diverse set of essays, the
study can yield robust insights into the differences and similarities in lexical complexity,
enhancing the validity of the findings. This combination of quantitative rigor and empirical
grounding is essential for drawing meaningful conclusions from the research.
In short, this study employs a quantitative, corpus-based comparative study to analyze the
lexical complexity in essay writing between human-generated texts and those produced by an
AI language model, specifically ChatGPT 3.5. By focusing on the same set of writing
prompts across both groups, the study ensures high validity, controlling for topical and task-
related variation. This sectione will involve two main phases:
- Data Collection: Collecting a balanced dataset of human-written texts from the
Cambridge IELTS writing tasks and ChatGPT 3.5 generated texts based on the same
prompts.
- Data Analysis: Utilizing lexical complexity measurement tools like Lexical Complexity
Analyzer (LCA) and the Lexical Complexity Calculator (LCC) to assess parameters such as
lexical diversity (e.g., TTR, MTLD), lexical density (e.g., content-to-function word ratios),
and lexical sophistication (e.g., proportion of advanced vocabulary).
2.3.1. Data Collection
1. [Link]
373809840_Research_Methodology_Methods_Approaches_And_Techniques