Using Corpus Linguistics to Inform Vocabulary
Instruction in the EFL Context
Chapter One
Introduction
1.1 Introduction
Vocabulary knowledge is widely recognized as one of the strongest predictors of
success in second-language learning. Regardless of a learner’s proficiency level,
academic discipline, or specific communicative goals, vocabulary forms the
foundation on which the four language skills—reading, listening, speaking, and
writing—are built. Research has consistently demonstrated that vocabulary size
correlates strongly with reading comprehension, writing quality, oral fluency, and
even academic achievement in content subjects delivered in English (Laufer, 2017;
Nation, 2013; Webb & Nation, 2017). Yet despite its central importance, vocabulary
remains an area in which many EFL learners struggle significantly. Learners often
know the basic meanings of words but cannot use them appropriately in context, lack
awareness of collocations and lexical chunks, or fail to understand discipline-
specific vocabulary required for academic or professional purposes. These issues are
not unique to one cultural or educational context; they appear consistently across
EFL settings, including the Middle East, East Asia, Latin America, and Europe.
For decades, EFL vocabulary instruction was shaped largely by intuition, tradition,
and textbook writers’ experience rather than empirical linguistic evidence. Teachers
and curriculum designers often chose vocabulary based on subjective impressions
of importance, thematic convenience, or available materials. Traditional textbooks
sometimes introduced low-frequency or decontextualized words that offered little
value for long-term language development, while essential high-frequency
vocabulary and multi-word units—especially those used in academic discourse—
were underrepresented or taught without adequate contextualization. As a result,
learners frequently memorized lists of words without gaining a deep understanding
of how those words behaved syntactically, semantically, or pragmatically in
authentic contexts.
However, the landscape of applied linguistics has changed dramatically with the rise
of corpus linguistics. Over the past three decades, corpus linguistics has become one
of the most influential methodologies in linguistics, contributing profoundly to
research on vocabulary, phraseology, discourse analysis, and language pedagogy.
Corpora—large, principled collections of naturally occurring language—allow
researchers to observe how language is actually used rather than how it is imagined
to be used. This shift from intuition to evidence has reshaped our understanding of
frequency, collocation, phraseology, and register variation (Biber et al., 1999;
McEnery & Hardie, 2012). Corpus findings have shown that language use is highly
patterned; much of the language we produce consists of formulaic sequences and
recurrent multi-word units rather than isolated words (Wray, 2002; Sinclair, 1991).
Such discoveries have major implications for vocabulary teaching.
A growing body of research demonstrates that high-frequency vocabulary accounts
for a significant proportion of the lexical load required for efficient comprehension.
Nation (2001, 2013) argues that the most frequent 2,000–3,000 word families cover
a substantial percentage of general texts, making them essential for learners at all
levels. Coxhead’s (2000) Academic Word List (AWL) revealed that academic
English relies on recurring lexical items that differ from those found in general
language. More recently, Gardner and Davies (2014) developed the Academic
Vocabulary List (AVL), using the 120-million-word Corpus of Contemporary
American English (COCA) to identify academic vocabulary based on more rigorous
criteria. Other studies have highlighted the importance of subject-specific corpora,
which reveal discipline-based vocabulary and phraseology essential for learners in
fields such as engineering, medicine, law, and applied sciences (Hyland & Tse,
2007; Chen & Baker, 2010).
1.2 Problem Statement
Although corpus linguistics has transformed vocabulary research, its influence on
everyday EFL teaching remains limited. Many EFL learners continue to struggle
with vocabulary comprehension and production because instructional materials
often fail to reflect authentic linguistic usage. In numerous contexts, including Iraq,
vocabulary instruction relies heavily on textbook lists, teacher intuition, and general
teaching traditions rather than frequency data or corpus evidence. As a result,
learners may remain unfamiliar with the high-frequency words, lexical bundles, and
collocations that dominate academic and professional English.
Additionally, teachers often lack the training necessary to integrate corpus-informed
materials into their instruction. They may feel uncertain about using corpus tools or
converting corpus findings into classroom activities. Students may also experience
difficulty engaging with corpus-based tasks, especially if such methods differ greatly
from traditional practices.
Thus, the primary problem is the absence of a systematic, corpus-informed
vocabulary teaching approach, coupled with limited understanding of how teachers
and learners perceive corpus-based instruction in EFL settings.
1.3 Background of the Study
Vocabulary knowledge has long been recognized as a fundamental component of
second and foreign language learning, yet its complexity and the challenges it
presents to learners continue to generate significant attention in applied linguistics
research. Scholars generally agree that vocabulary forms the foundation upon which
communication skills, academic literacy, and broader language competence are built
(Nation, 2013; Schmitt, 2020). Without a sufficient lexical repertoire, learners
struggle to comprehend texts, express ideas accurately, and develop fluency. In
many English as a Foreign Language (EFL) contexts, vocabulary becomes an even
greater obstacle because exposure to natural English input is limited. As a result,
learners often rely heavily on classroom instruction, textbooks, and exam-driven
memorization, which may not reflect how language is used in authentic
communicative settings (Webb & Nation, 2017; Laufer, 2017).
In the last three decades, the field of corpus linguistics has provided powerful tools
to address this gap by offering empirical descriptions of real language use. A
corpus—defined as a large, structured collection of authentic texts—allows
researchers and teachers to examine patterns, frequencies, collocations, and
phraseological units that occur naturally in speech and writing (Biber et al., 1998;
McEnery & Hardie, 2012). Corpus-based approaches have reshaped the
understanding of vocabulary from a list of individual words to a much broader
construct that includes multi-word units, formulaic sequences, chunks, and phrasal
verbs. Research consistently shows that high-frequency vocabulary and recurrent
lexical bundles are essential for fluent processing and successful communication,
especially in academic and professional contexts (Hyland, 2008; Simpson-Vlach &
Ellis, 2010). This shift reflects a growing consensus that vocabulary learning should
be grounded in actual language usage rather than intuition or tradition.
Despite these advances, many EFL classrooms still rely on outdated or intuition-
based materials that do not represent real English. Traditional vocabulary instruction
often centers on decontextualized word lists, dictionary definitions, and translation
exercises. While these methods may provide short-term gains, they rarely enable
learners to understand how words behave in authentic contexts or how lexical
patterns contribute to meaning (Schmitt & Schmitt, 2020). This disconnection
between classroom materials and real-world usage leads to a persistent gap between
learners’ receptive knowledge (what they can recognize) and productive command
(what they can use accurately). Research also suggests that EFL learners frequently
misjudge frequency, assuming some rare or specialized words are common while
overlooking core vocabulary items that occur very frequently in everyday
communication (Coxhead, 2000; Nation, 2013).
Corpus-informed instruction offers a promising alternative by ensuring that teaching
materials reflect actual language patterns. When teachers draw on corpus data, they
can highlight high-frequency vocabulary, common collocations, and recurring
phraseological units that learners are most likely to encounter in academic or general
English contexts (Chen & Baker, 2016). For example, corpora reveal that certain
multi-word sequences—such as on the other hand, as a result of, or according to
the—play a critical role in academic discourse. Textbooks, however, often fail to
emphasize these items systematically. Incorporating corpus findings into
instructional materials helps ensure that learners acquire vocabulary that is useful,
frequent, and contextually appropriate.
At the same time, the pedagogical application of corpora raises important questions
about curriculum design, teacher preparedness, and learner engagement. While
corpora provide valuable insights, teachers may feel unprepared to work with corpus
tools or may lack training in data-driven learning (DDL) approaches (Boulton, 2010;
Breyer, 2009). Students, especially at lower proficiency levels, may find
concordance lines and corpus interfaces overwhelming without proper scaffolding.
These challenges highlight the need for research on how corpus-informed materials
can be adapted for practical classroom use in EFL settings.
Another significant concern is the relevance of existing corpora to specific learner
populations. Many widely used corpora—such as the British National Corpus (BNC)
or the Corpus of Contemporary American English (COCA)—represent general
English usage rather than the linguistic needs of particular EFL groups. In some
educational contexts, learners require specialized vocabulary related to academic
writing, scientific texts, or professional discourse. This has led to the development
of specialized academic corpora such as the Michigan Corpus of Academic Spoken
English (MICASE), the British Academic Written English corpus (BAWE), and
discipline-specific corpora in engineering, medicine, or applied sciences (Hyland &
Tse, 2007). Understanding which high-frequency lexical items appear in such
corpora can help teachers design more targeted vocabulary instruction aligned with
students’ academic or professional goals.
The increasing interest in corpus-based vocabulary teaching reflects a broader
movement in applied linguistics toward evidence-based pedagogy. Rather than
relying on intuition, teachers and curriculum designers are encouraged to examine
empirical data about how English is actually used. This aligns with the principles of
data-driven learning, which promote learner autonomy by encouraging students to
explore authentic examples and derive patterns themselves (Johns, 1991; Boulton &
Cobb, 2017). However, the effectiveness of such approaches depends on learners’
perceptions and attitudes as well as the extent to which teachers feel confident
integrating corpus-based activities into their lessons.
In many EFL contexts—including those in the Middle East and Asia—corpus-based
instruction is still relatively new. Teachers may be unfamiliar with corpus tools, or
institutions may lack the technological resources needed to implement corpus-driven
activities. Students, on the other hand, may be accustomed to rote memorization and
may initially resist inductive or exploratory learning approaches (Tsui, 2004).
Understanding the perceptions of both teachers and learners is therefore essential for
evaluating the practicality and acceptability of corpus-informed vocabulary
instruction.
1.4 Aims of the Study
The present study aims to investigate how corpus linguistics can effectively inform
vocabulary instruction within an English as a Foreign Language (EFL) context.
Although a large body of research highlights the value of corpus-based approaches
to understanding real language use, many EFL classrooms still depend on intuition-
based or traditional vocabulary teaching methods that do not fully reflect authentic
linguistic patterns. This study therefore seeks to bridge the gap between empirical
corpus evidence and practical vocabulary pedagogy by focusing on three
interconnected aims.
First, the study aims to identify the high-frequency lexical items—including single
words, multi-word chunks, collocations, and phrasal verbs—that appear in a
specialized corpus relevant to the target EFL learners. By examining lexical patterns
in academic or general English corpora, the study seeks to determine which
vocabulary items are most essential for learners’ communicative and academic
needs. This analysis is intended to provide an evidence-based foundation for
selecting vocabulary that is genuinely useful and frequently encountered in real
contexts.
Second, the study aims to design corpus-informed vocabulary materials tailored for
EFL classroom use. Drawing on the frequency data and phraseological patterns
revealed through corpus analysis, the study seeks to develop practical instructional
materials that can support learners in understanding how vocabulary functions in
authentic texts. These materials are expected to incorporate examples from corpora,
highlight recurrent lexical bundles, and provide structured activities that help
learners notice patterns, build awareness, and develop stronger receptive and
productive skills.
Third, the study aims to explore the perceptions of EFL teachers and students
regarding the usefulness, feasibility, and challenges of corpus-informed vocabulary
instruction. Understanding how teachers and learners respond to corpus-based
materials is crucial for evaluating their pedagogical value and identifying potential
obstacles to implementation. By gathering perspectives from both groups, the study
aims to provide insights into the practicality of integrating corpus tools into
classroom practice and to suggest ways of enhancing teacher readiness and learner
engagement.
1.5 Research Questions
Based on the aims of the study, the research is guided by the following questions:
1. What are the high-frequency lexical items—such as single words, multi-word
chunks, collocations, and phrasal verbs—found in a specialized academic or
general English corpus that are relevant to the target EFL learners?
2. How can corpus-informed vocabulary materials be effectively designed and
adapted for use in the EFL classroom?
3. What are the perceptions of EFL teachers and students regarding the
usefulness, practicality, and challenges of implementing corpus-informed
vocabulary learning activities?
Chapter Two
2. Literature Review
2.1 The Role of Vocabulary in Second Language Acquisition
Vocabulary knowledge is widely regarded as a core component of second language
acquisition (SLA) and a fundamental prerequisite for effective language use.
Without an adequate lexical repertoire, learners are unable to comprehend input or
express meaning successfully in the target language (Qian, 1999; Schmitt, 2000).
Unlike grammatical knowledge, which operates within a limited set of rules,
vocabulary constitutes an open-ended system that continuously expands as learners
encounter new communicative contexts. As a result, vocabulary development is
often considered one of the greatest challenges for second language learners.
A substantial body of SLA research highlights the close relationship between
vocabulary knowledge and the development of the four language skills: listening,
speaking, reading, and writing. Lexical competence plays a critical role in listening
comprehension, as unfamiliar words can hinder learners’ ability to process spoken
input in real time. Similarly, reading comprehension has been shown to depend
heavily on lexical coverage, with studies suggesting that learners need to know
approximately 95–98% of the words in a text to achieve adequate understanding
(Qian, 1999; Nation, 2001). In productive skills, vocabulary limitations often result
in hesitant speech, reduced fluency, and restricted written expression (Smith, 1998;
Ng & Naim, 2023). Consequently, vocabulary knowledge is widely viewed as a key
component of communicative competence.
Contemporary SLA research distinguishes between two major dimensions of
vocabulary knowledge: breadth and depth. Vocabulary breadth refers to the number
of words a learner knows at least at a basic level, while depth of vocabulary
knowledge encompasses how well those words are known. Depth includes multiple
aspects such as semantic associations, grammatical behavior, collocational patterns,
register, morphological structure, and pragmatic use (Qian, 1999; Schmitt, 2000; Ng
& Naim, 2023). Studies have demonstrated that depth of vocabulary knowledge is
particularly important for advanced language use, as it enables learners to select
appropriate lexical items in context and to interpret nuanced meanings in authentic
texts.
Vocabulary acquisition in SLA occurs through both incidental and intentional
learning processes. Incidental vocabulary learning typically takes place when
learners are exposed to meaningful language input, such as reading or listening,
without an explicit focus on learning new words. This type of learning is often
associated with extensive reading and exposure to rich, contextualized input (Coady
& Huckin, 1997).
2.2 Corpus Linguistics in Language Teaching
Corpus linguistics is a research approach that investigates language use through
large, systematically compiled electronic collections of authentic spoken and written
texts, commonly referred to as corpora. These corpora allow researchers and
educators to examine linguistic patterns based on empirical evidence rather than
intuition or prescriptive rules (McEnery & Hardie, 2012; Biber, Conrad & Reppen,
1998). In the context of language teaching, corpus linguistics offers valuable insights
into how language is actually used by native and non-native speakers across different
registers, genres, and communicative situations.
One of the most significant contributions of corpus linguistics to language teaching
lies in its ability to reveal frequency information and usage patterns that are often
underrepresented or inaccurately portrayed in traditional teaching materials. Corpus-
based studies have shown that high-frequency words, multi-word units, and
collocations play a crucial role in fluent and natural language use (Biber et al., 2004;
Ellis, 2012). By prioritizing frequently occurring lexical items and patterns, teachers
can design instruction that better reflects real-world language use, thereby increasing
the relevance and effectiveness of vocabulary instruction in EFL contexts.
A key pedagogical application of corpus linguistics is Data-Driven Learning (DDL),
an approach originally proposed by Johns (1991), in which learners are encouraged
to explore corpus data directly through concordance lines and other corpus tools.
Rather than being presented with pre-selected rules or examples, learners engage in
inductive learning by identifying patterns, meanings, and usage constraints on their
own. Research suggests that DDL promotes learner autonomy, critical thinking, and
deeper lexical processing, particularly in relation to collocation, phraseology, and
grammatical patterns (Boulton, 2010; Vyatkina, 2016). This approach is especially
beneficial for advanced and university-level EFL learners, who are capable of
analyzing language data and drawing generalizations from authentic input.
2.3 Previous Studies on Corpus Linguistics in EFL
A growing body of research in English as a Foreign Language (EFL) contexts has
examined the role of corpus linguistics in enhancing vocabulary learning and
teaching practices. These studies consistently suggest that corpus-based approaches
contribute positively to learners’ lexical development, awareness of authentic
language use, and overall engagement with vocabulary learning. Unlike traditional
instruction, which often relies on intuition-driven examples, corpus-informed
pedagogy exposes learners to real language patterns, thereby improving the accuracy
and depth of vocabulary knowledge.
Experimental and quasi-experimental studies provide strong evidence for the
effectiveness of corpus-based vocabulary instruction. For instance, Alenizi and
Adawi (2024) conducted an experimental study with Saudi EFL university students
and reported significantly higher vocabulary gains among learners exposed to
corpus-based materials compared to those receiving conventional instruction. Their
findings indicate that repeated exposure to authentic concordance lines enhanced
learners’ understanding of word meanings, collocations, and contextual usage.
Similarly, Harahap et al. (2025) demonstrated that corpus analysis helped learners
identify frequent word forms and usage patterns more efficiently, leading to
improved retention and more accurate language production. These results support
the argument that corpus tools facilitate meaningful and data-rich vocabulary
learning experiences.
Several studies have focused specifically on Data-Driven Learning (DDL) in EFL
settings, highlighting its impact on learners’ autonomy and analytical skills.
Research conducted with university-level EFL learners has shown that DDL
activities promote deeper cognitive processing by encouraging learners to notice
patterns, infer meanings, and test hypotheses about language use (Boulton, 2017;
Vyatkina, 2016). More recent work by Lee and Lin (2022) found that learners using
guided DDL tasks demonstrated improved collocational competence and greater
confidence in vocabulary use compared to those taught through textbook-based
methods. These findings suggest that corpus-based learning is particularly effective
when learners receive appropriate scaffolding and pedagogical support.
In addition to learner outcomes, previous studies have also explored teachers’
perceptions and practices regarding corpus use in EFL classrooms. Çalışkan and
Kuru Gönen (2018) reported generally positive attitudes among EFL teachers toward
corpus linguistics, particularly its potential to provide authentic examples and clarify
problematic language points. However, their study also highlighted challenges such
as limited technical knowledge, time constraints, and a lack of ready-made corpus-
based materials. More recent research confirms that while teachers recognize the
pedagogical value of corpora, effective implementation often depends on
professional training and institutional support (O’Keeffe & Mark, 2017; Zareva,
2023).
Learner corpus studies have further contributed to EFL vocabulary research by
identifying common lexical errors, patterns of overuse and underuse, and
developmental trends in learner language. Such analyses have informed vocabulary
instruction by highlighting areas that require explicit attention, such as collocations,
preposition use, and formulaic language (Granger, 2015; Gablasova et al., 2019).
Recent learner corpus research in Asian and Middle Eastern EFL contexts has shown
that corpus-informed feedback can significantly improve learners’ lexical accuracy
and awareness of register (Yoon, 2021; Al-Khazraji, 2023).
2.4 Vocabulary Teaching Methodologies
Vocabulary teaching methodologies have undergone significant development over
the past decades, shifting from decontextualized memorization toward approaches
that emphasize meaningful use, learner engagement, and strategic competence. Early
vocabulary instruction largely relied on word lists, translation, and rote
memorization, focusing primarily on form–meaning associations. While such
approaches can contribute to short-term retention, research has shown that they often
fail to promote deep lexical knowledge or long-term transfer to communicative use
(Stahl & Nagy, 2005; Nation, 2001).
Contemporary vocabulary pedagogy increasingly supports contextualized learning,
in which learners encounter and practice new words through meaningful input such
as reading, listening, and task-based activities. Contextual exposure allows learners
to develop a richer understanding of lexical items, including their semantic nuances,
grammatical behavior, and pragmatic constraints (Beck, McKeown & Kucan, 2013).
Studies have demonstrated that repeated encounters with words in varied contexts
significantly enhance retention and facilitate the development of depth of vocabulary
knowledge, particularly in EFL settings where exposure is otherwise limited
(Nation, 2013; Schmitt, 2019).
In parallel with contextual approaches, explicit vocabulary instruction remains an
important component of effective methodology. Explicit teaching includes direct
explanation of word meanings, word formation, collocations, and usage restrictions,
often combined with focused practice activities. Research suggests that explicit
instruction is particularly beneficial for low-frequency words and complex lexical
items that learners are unlikely to acquire incidentally (Laufer & Hulstijn, 2001;
Nation, 2001). When combined with contextualized input, explicit instruction
contributes to more robust and durable vocabulary learning.
A key development in modern vocabulary pedagogy is the emphasis on vocabulary
learning strategies, which empower learners to take an active role in expanding their
lexical knowledge. Strategy-based instruction includes techniques such as guessing
meaning from context, using morphological analysis, consulting dictionaries
effectively, and organizing vocabulary through semantic mapping or lexical
notebooks. Studies indicate that learners who are trained in vocabulary strategies
demonstrate greater autonomy, improved retention, and more effective application
of new words in communicative tasks (Stahl & Nagy, 2005; Ng & Naim, 2023).
Strategy instruction is particularly valuable in EFL contexts, where learners must
often compensate for limited exposure to the target language.
Lexical and communicative approaches to language teaching further underscore the
importance of vocabulary as the foundation of language proficiency. The Lexical
Approach, in particular, emphasizes the teaching of chunks, collocations, and
formulaic sequences rather than isolated words, reflecting the way language is stored
and processed by proficient users (Lewis, 2000; Ellis, 2012). This perspective aligns
closely with corpus linguistics, which provides empirical evidence of the frequency
and communicative importance of multi-word units in authentic language use.
Chapter Three
3. Conclusion
This study set out to explore the role of corpus linguistics in informing vocabulary
instruction in the EFL context and to highlight its potential contribution to more
effective and evidence-based language teaching. Vocabulary knowledge has long
been recognized as a core component of language proficiency, yet traditional
vocabulary instruction in many EFL classrooms continues to rely heavily on
intuition, isolated word lists, and textbook-driven input. Such approaches often fail
to expose learners to authentic language use and meaningful lexical patterns. In
response to these challenges, corpus linguistics offers a valuable alternative by
providing access to real language data and systematic insights into how words are
actually used in context.
The findings of this study indicate that corpus-informed vocabulary instruction can
significantly enhance learners’ understanding of word frequency, collocation, and
contextual usage. By engaging with authentic corpus data, learners become more
aware of how vocabulary functions in real communicative situations rather than
memorizing words in isolation. This increased awareness supports both receptive
and productive vocabulary development and contributes to more natural and
accurate language use. Moreover, corpus-based activities encourage learners to take
a more active role in the learning process, fostering learner autonomy and critical
thinking through data-driven learning.
From a pedagogical perspective, the integration of corpus linguistics into EFL
vocabulary instruction offers practical benefits for teachers as well. Corpus tools can
assist teachers in selecting high-frequency and pedagogically relevant vocabulary,
designing more authentic teaching materials, and addressing learners’ common
lexical errors. As a result, vocabulary instruction becomes more systematic,
transparent, and grounded in empirical evidence rather than subjective judgment.
However, the successful implementation of corpus-based instruction requires
adequate teacher training and access to appropriate technological resources, which
remain challenges in some EFL contexts.
Despite its contributions, this study is not without limitations. The scope of
participants and instructional duration may limit the generalizability of the findings.
Future research could examine corpus-informed vocabulary instruction across
different proficiency levels, educational settings, and longer instructional periods.
Additionally, further studies may explore teachers’ attitudes toward corpus use and
the integration of learner corpora in addressing specific learner needs.
REFERENCES
A systematic review of corpus-based instruction in EFL classroom, 2025. Heliyon, 11(2), e42016.
Alenizi, A. & Adawi, R., 2024. Investigating corpus-based developed materials in Saudi EFL
vocabulary learning. Forum for Linguistic Studies, 6(3), pp.1–18.
Aston, G. (2001). Learning with corpora. Athelstan.
Beck, I.L., McKeown, M.G. & Kucan, L., 2013. Bringing Words to Life: Robust Vocabulary
Instruction. New York: Guilford Press.
Biber, D., Conrad, S., & Cortes, V. (2004). If you look at…: Lexical bundles in university teaching
and textbooks. Applied Linguistics, 25(3), 371–405.
Biber, D., Conrad, S., & Reppen, R. (1998). Corpus linguistics: Investigating language structure
and use. Cambridge University Press.
Boulton, A. (2010). Data-driven learning: On paper, in practice. Computer Assisted Language
Learning, 23(2), 83–99.
Boulton, A., & Cobb, T. (2017). Corpus use in language learning: A meta-analysis. Language
Learning, 67(2), 348–393.
Breyer, Y. (2009). Learning and teaching with corpora: Reflections by student teachers. Computer
Assisted Language Learning, 22(2), 153–172.
Çalışkan, G. & Kuru Gönen, S.I., 2018. Teachers’ perceptions of corpus-based vocabulary
instruction. Journal of Language and Linguistic Studies, 14(4), pp.190–210.
Cameron, L., & Larsen-Freeman, D. (2007). Complex systems and applied linguistics.
International Journal of Applied Linguistics, 17(2), 226–240.
Chen, Y. H., & Baker, P. (2016). Investigating lexical bundles in academic writing. Applied
Linguistics, 37(6), 849–880.
Chen, Y., & Baker, P. (2010). Lexical bundles in L1 and L2 academic writing. Language Learning,
60(1), 30–56.
Coady, J. & Huckin, T., 1997. Second Language Vocabulary Acquisition. Cambridge: Cambridge
University Press.
Coxhead, A. (2000). A new academic word list. TESOL Quarterly, 34(2), 213–238.
Gablasova, D., Brezina, V. & McEnery, T., 2019. Exploring corpus linguistics and pedagogy.
Language Teaching, 52(4), pp.529–548.
Gardner, D., & Davies, M. (2014). A new academic vocabulary list. Applied Linguistics, 35(3),
305–327.
Harahap, D.I. et al., 2025. Corpus linguistics in teaching vocabulary for EFL learners. English
Education Journal, 16(1), pp.24–41.
Hyland, K. (2008). Academic clusters: Formulaic language in academic writing. Routledge.
Hyland, K., & Tse, P. (2007). Is there an “academic vocabulary”? TESOL Quarterly, 41(2), 235–
253.
Johns, T. (1991). From printout to handout: Grammar and vocabulary teaching in the context of
data-driven learning. English Language Research Journal, 4, 27–45.
Johns, T. (1991). Should you be persuaded—Two samples of data-driven learning. In Classroom
Concordancing (pp. 1–16).
Johns, T., 1991. ‘From printout to handout: Grammar and vocabulary teaching in the context of
data-driven learning’, in System, 19(4), pp.243–250.
Laufer, B. (2017). From word parts to full texts: Searching for effective vocabulary instruction.
Language Teaching Research, 21(5), 514–527.
McEnery, T., & Hardie, A. (2012). Corpus linguistics: Method, theory and practice. Cambridge
University Press.
Meunier, F., 2011. Learner Corpora and Language Teaching. Amsterdam: John Benjamins.
Nation, I. S. P. (2001). Learning vocabulary in another language. Cambridge University Press.
Nation, I. S. P. (2013). Learning vocabulary in another language (2nd ed.). Cambridge University
Press.
Nation, I.S.P., 2001. Learning Vocabulary in Another Language. Cambridge: Cambridge
University Press.
Ng, W.L. & Naim, R.M., 2023. Vocabulary learning strategies in SLA. Journal of Language,
Literacy and Translation, 6(2), pp.223–241.
Ng, W.L. & Naim, R.M., 2023. Vocabulary learning strategies in SLA. Journal of Language,
Literacy and Translation, 6(2), pp.223–241.
Qian, D.D., 1999. Assessing the roles of depth and breadth of vocabulary knowledge in reading
comprehension. The Canadian Modern Language Review, 56(2), pp.282–308.
Schmitt, N. (2010). Researching vocabulary: A vocabulary research manual. Palgrave Macmillan.
Schmitt, N. (2020). Vocabulary in language teaching (2nd ed.). Cambridge University Press.
Schmitt, N., & Schmitt, D. (2020). Vocabulary and the four skills. Language Teaching, 53(4), 1–
20.
Schmitt, N., 2000. Vocabulary in Language Teaching. Cambridge: Cambridge University Press.
Boulton, A. & Cobb, T., 2017. Corpus Use in Language Learning: A Meta-Analysis. Language
Learning, 67(2), pp.348–393.
Simpson-Vlach, R., & Ellis, N. C. (2010). Formulaic language in academic speech and writing:
Collocational, pragmatic, and rhetorical perspectives. Applied Linguistics, 31(4), 487–512.
Sinclair, J. (1991). Corpus, concordance, collocation. Oxford University Press.
Stahl, S.A. & Nagy, W.E., 2005. Teaching Vocabulary: Perspectives From Research and Practice.
Mahwah: Lawrence Erlbaum Associates.
Timmis, I., 2015. Corpus Linguistics for ELT: Research and Practice. London: Routledge.
Timmis, I., 2015. Corpus Linguistics for ELT: Research and Practice. London: Routledge.
Tsui, A. B. M. (2004). Classroom discourse and the teacher. Routledge.
Webb, S., & Nation, P. (2017). How vocabulary is learned. Oxford University Press.
Wray, A. (2002). Formulaic language and the lexicon. Cambridge University Press.