Method 3 Notes
Method 3 Notes
Language testing and assessment play crucial roles in both language teaching and
learning. They provide valuable information about learners' proficiency, inform
instructional decisions, and contribute to overall program evaluation. This chapter
explores the fundamental differences of the two concepts of language testing and
assessment, examining their purposes, and impacts on language education.
The objectives of language testing can be multifaceted, depending on the context and
purpose of the assessment. However, some common objectives include:
Language testing plays a crucial role in educational settings for several reasons:
Language assessment plays a critical role in the language learning process. It provides
valuable information about learners' progress, informs instruction, and helps to ensure
that teaching and learning are aligned with desired outcomes. Language assessment
includes both formal and informal methods.
• Formal assessments:
o Standardized tests: These are typically high-stakes tests with standardized
procedures and scoring. Examples include TOEFL, IELTS, and Cambridge English
exams.
o Teacher-made tests: These are tests designed by teachers to assess specific learning
objectives and monitor student progress. They may include multiple-choice questions,
short answer questions, writing tasks, and oral presentations.
• Informal assessments:
o Observations: Observing learners' language use in natural settings, such as
classroom interactions or group work.
o Portfolios: Collections of student work that demonstrate their language
development over time.
o Interviews: One-on-one conversations with learners to assess their speaking and
listening skills.
o Self-assessment: Techniques that involve learners in reflecting on their own
language learning and progress.
o Peer assessment: Learners provide feedback on each other's work.
While the terms "testing" and "assessment" are often used interchangeably, there are
important distinctions between them in the context of language learning. Testing
generally refers to the use of formal procedures to measure a specific aspect of a learner's
language ability. It typically involves the administration of a standardized test with a
predetermined set of questions or tasks, followed by a scoring process that yields a
quantitative result. Assessment, on the other hand, is a broader term that encompasses a
wider range of procedures for gathering information about learners' language proficiency.
It involves a more holistic approach that considers a variety of evidence, including
observations, interviews, portfolios, and informal assessments. Table 1.1 below presents
the differences between language testing and assessment:
Table 1.1 Language testing vs. assessment
In summary, while testing provides valuable information about specific language skills,
assessment offers a more comprehensive and holistic view of learners' language
proficiency. Both testing and assessment play important roles in language learning, and a
balanced approach that incorporates both formal and informal assessment methods can
provide valuable insights into learners' progress and inform effective instruction. The
following digagram illustrates the interrelationship of language teaching, assessment and
testing.
Backwash refers to the influence that tests have on teaching and learning practices. When
tests are perceived as high-stakes, they can have a significant impact on what is taught,
how it is taught, and how students learn. Positive washback occurs when tests promote
beneficial teaching practices and learning outcomes, while negative washback can lead to
teaching to the test, where instruction becomes narrowly focused on test content at the
expense of broader language skills (Fulcher & Davidson, 2007).
• Narrowing the curriculum: Teachers may focus on teaching only what is tested,
neglecting other important aspects of language learning, such as creativity and
critical thinking.
• Teaching to the test: Teachers may prioritize test preparation over meaningful
language learning, leading to rote learning and a focus on test-taking strategies.
• Test anxiety: High-stakes tests can cause anxiety and stress, which can negatively
impact learners' performance.
• Inequitable assessment: Tests may not accurately reflect the diverse range of
language abilities, particularly for learners from marginalized backgrounds.
• Test validity and reliability: Ensure that tests are valid and reliable measures of
language proficiency.
• Balanced assessment: Use a variety of assessment methods, including formative
and summative assessments, to avoid over-reliance on high-stakes tests.
• Teacher training: Provide teachers with training on effective test preparation
strategies and how to avoid negative backwash effects.
• Learner awareness: Educate learners about the purpose of tests and how to
approach them effectively.
• Authentic tasks: Incorporate authentic language tasks into teaching and assessment
to promote real-world language use.
By understanding the potential positive and negative impacts of tests, language educators
can make informed decisions about assessment practices and strive to create a positive
learning environment that fosters language development and critical thinking.
1.4 Consolidation activities
“Căn cứ đánh giá là các yêu cầu cần đạt về phẩm chất và năng lực được quy định trong
chương trình tổng thể và các chương trình môn học, hoạt động giáo dục. Phạm vi đánh
giá bao gồm các môn học và hoạt động giáo dục bắt buộc, môn học và chuyên đề học tập
lựa chọn và môn học tự chọn. Đối tượng đánh giá là sản phẩm và quá trình học tập, rèn
luyện của học sinh.
Kết quả giáo dục được đánh giá bằng các hình thức định tính và định lượng thông qua
đánh giá thường xuyên, định kì ở cơ sở giáo dục, các kì đánh giá trên diện rộng ở cấp
quốc gia, cấp địa phương và các kì đánh giá quốc tế. Cùng với kết quả các môn học và
hoạt động giáo dục bắt buộc, các môn học và chuyên đề học tập lựa chọn, kết quả các
môn học tự chọn được sử dụng cho đánh giá kết quả học tập chung của học sinh trong
từng năm học và trong cả quá trình học tập.
Việc đánh giá thường xuyên do giáo viên phụ trách môn học tổ chức, kết hợp đánh giá
của giáo viên, của cha mẹ học sinh, của bản thân học sinh được đánh giá và của các học
sinh khác.
Việc đánh giá định kì do cơ sở giáo dục tổ chức để phục vụ công tác quản lí các hoạt
động dạy học, bảo đảm chất lượng ở cơ sở giáo dục và phục vụ phát triển chương trình.
Việc đánh giá trên diện rộng ở cấp quốc gia, cấp địa phương do tổ chức khảo thí cấp quốc
gia hoặc cấp tỉnh, thành phố trực thuộc trung ương tổ chức để phục vụ công tác quản lí
các hoạt động dạy học, bảo đảm chất lượng đánh giá kết quả giáo dục ở cơ sở giáo dục,
phục vụ phát triển chương trình và nâng cao chất lượng giáo dục.
Phương thức đánh giá bảo đảm độ tin cậy, khách quan, phù hợp với từng lứa tuổi, từng
cấp học, không gây áp lực lên học sinh, hạn chế tốn kém cho ngân sách nhà nước, gia
đình học sinh và xã hội.
Nghiên cứu từng bước áp dụng các thành tựu của khoa học đo lường, đánh giá trong giáo
dục và kinh nghiệm quốc tế vào việc nâng cao chất lượng đánh giá kết quả giáo dục, xếp
loại học sinh”
Procedure: Students work in pairs to sort the following key phrases into a T-table and
then explain why you decide whether the phrases align more closely with testing or
assessment.
o Multiple-choice test
o Portfolio assessment
o Oral interview
o Writing sample
o Observation of classroom interaction
o Standardized proficiency test
o Peer feedback
o Self-assessment
o Teacher-made quizzes
o Project presentation
Procedure: Students in groups examine the prescribed structure of the national English
exam in the link provided which specifies six types of test items. They then discuss the
washback effects of this exam.
Cấu trúc đề thi Tiếng Anh THPT Quốc gia từ năm 2025 gồm có 6 dạng bài chính
([Link]
[Link]). Cấu trúc bao gồm:
Procedure:
Students in groups search on the Internet one standardized proficiency test such as
like VSTEP, TOEIC, IELTS, Cambridge English exams; a high-stake test; a
teacher-made test. They then discuss the following quesitons:
Procedure: Students work in pairs to review various informal assessment methods (e.g.,
observations, interviews, student journals, peer feedback, self-assessment) and discuss
the purposes of informal assessment and answer the following questions:
Fulcher, G., & Davidson, F. (2007). Language testing and assessment: an advanced
resource book. Routledge.
Hughes, A. (2003). Testing for language teachers. 2nd ed. Cambridge University Press.
MOET (2018). Thông tư ban hành giáo dục phổ thông. Available at:
[Link] giao-
[Link]
Shohamy, E. (2020). The power of tests: A critical perspective on the uses of language
tests. Routledge.
Based on Circular 32 from the Ministry of Education and Training (MOET, 2018, pp. 32-33),
language assessment for upper secondary schools in Vietnam is guided by the following
principles:
1. Assessment Criteria:
Evaluation is based on the required learning outcomes in the national curriculum, including
subject knowledge, skills, and competencies.
Assessment covers compulsory subjects, elective subjects, and extracurricular educational
activities.
2. Scope of Assessment:
Includes both the learning process and learning products (i.e., students' assignments, projects,
and test results).
3. Assessment Methods:
Regular assessment: Conducted by subject teachers using a mix of teacher assessments, peer
assessments, self-assessments, and parental feedback.
Periodic assessment: Conducted by schools to manage teaching quality and curriculum
implementation.
Large-scale assessment: Conducted at national and provincial levels to evaluate overall
educational effectiveness and inform policy development.
4. Forms of Assessment:
Qualitative and Quantitative Methods: Assessments include both qualitative (e.g., descriptive
feedback, portfolios) and quantitative methods (e.g., standardized tests).
Combination of Formative and Summative Assessments: Encourages a balance between
continuous assessment (formative) and final evaluations (summative).
5. Reliability and Objectivity:
Assessment methods should ensure fairness, reliability, and validity.
They must be appropriate for each grade level and not create undue pressure on students.
6. Alignment with International Standards:
Vietnam aims to gradually incorporate international assessment standards and best practices to
improve educational quality.
7. Financial Considerations:
Assessment implementation should minimize costs for the state, families, and society.
This assessment framework aims to ensure that English language education in Vietnam aligns
with national curriculum goals while integrating modern evaluation approaches.
Activity 3: Possible Washback Effects of Assessment Forms
The possible washback effects (both positive and negative) for different forms of assessment in
the national English exam:
3. What are the potential advantages and disadvantages of this type of test?
4. What are the ethical considerations associated with this type of test?
- Fairness: Tests should be unbiased and accessible to students from different backgrounds.
- Validity & Reliability: Tests should measure what they are intended to assess consistently.
- Stress & Anxiety: High-stakes testing should avoid unnecessary pressure on students.
- Cost & Accessibility: Standardized tests should not disadvantage students who cannot afford
them.
Summary:
Standardized tests are useful for broad evaluation but can be stressful and costly.
High-stakes exams influence curriculum design but may lead to excessive test
preparation.
Teacher-made tests are flexible but may lack reliability across different contexts.
By analyzing these aspects, students will develop critical thinking about formal language
assessments and their role in education.
Activity 5: Informal Assessment - Beyond the Test
3. How can the data collected through this method be used to inform instruction?
Identifies individual student needs – Observations help teachers adjust lesson plans.
Tracks progress over time – Portfolios reveal strengths and weaknesses in different skills.
Personalizes feedback – Interviews allow teachers to provide direct support.
Encourages student involvement – Self-assessment helps learners set personal goals.
Enhances peer learning – Peer feedback fosters collaborative skill development.
4. What ethical considerations need to be taken into account when using this method?
Fairness & Objectivity – Teachers must ensure unbiased observations and feedback.
Student Privacy – Portfolio content should be confidential and used for learning purposes
only.
Guidance for Peer & Self-Assessment – Students need clear criteria to avoid misleading
evaluations.
Avoiding Over-Reliance on Informal Methods – Combining informal and formal
assessments ensures accuracy.
Conclusion:
- Informal assessment is essential for holistic language evaluation.
- It complements formal tests by promoting deeper learning and engagement.
- Best practice: Use a mix of observations, self/peer assessments, portfolios, and interviews to
support language development.
This key helps students understand how informal assessments contribute to meaningful learning!
Chapter 2: Approaches to Language Testing and Assessment
Language testing and assessment have evolved significantly over the years, reflecting
changes in language teaching methodologies and the broader understanding of language
learning. This chapter will explore the primary approaches to language assessment and
testing: the traditional approach and the communicative approach, along with the
distinction between formative and summative assessment. It also presents the new trend
in language testing and assessment which is technology-enhanced approach.
Key Characteristics:
Limitations:
Key Characteristics:
• Discrete-point tasks:
o Multiple-choice questions
o Fill-in-the-blank exercises
o Cloze tests
o Dictation
o Grammar translation exercises
• Emphasis on accuracy:
o Correct usage of grammar rules, vocabulary, and pronunciation
• Standardized testing:
o Standardized tests with fixed formats and scoring criteria
• Teacher-centered:
o Teachers as the primary authority in the assessment process
Limitations:
o High-stakes nature of traditional assessments can lead to test anxiety and negatively
impact learners' performance
Key characteristics:
Advantages:
Key Characteristics:
Authentic tasks:
o Role-plays
o Simulations
o Interviews
o Presentations
o Writing task
Focus on communicative competence:
o Assessment of learners' ability to use language fluently, accurately, appropriately,
and meaningfully
Holistic scoring:
o Consideration of multiple aspects of language performance, including content,
organization, grammar, vocabulary, and mechanics
Learner-centered:
o Learners actively participate in the assessment process
Subjectivity in scoring:
o Intra-rater reliability: Even the same rater may assess the same performance differently
on different occasions.
Practical constraints:
o Isolation of skills: It can be challenging to isolate and assess specific language skills,
such as grammar or vocabulary, in a purely communicative context.
Cultural bias:
o Cultural differences: Test tasks may not be equally accessible to learners from different
cultural backgrounds, potentially leading to unfair assessment.
Limited standardization:
To mitigate these limitations, language testers often combine communicative tasks with
more traditional, discrete-point tasks. Additionally, the use of rating scales, rubrics, and
clear scoring criteria can help to improve the reliability and validity of communicative
language assessments.
Formative and summative assessment are two types of assessment that serve different
purposes in language education.
Formative assessment:
Purpose: To monitor learner progress and provide feedback to inform instruction
Examples:
o Quizzes
o Self-assessment
o Peer assessment
o Teacher observation
o Portfolio assessment
Benefits:
o Identifies learners' strengths and weaknesses
o Provides timely feedback
o Allows for adjustments to teaching and learning
o Promotes learner autonomy and self-directed learning
Summative assessment:
Examples:
o Final exams
o Standardized tests
o Project-based assessments
Benefits:
o Measures overall learning outcomes
o Provides a benchmark for progress
o Informs decisions about future learning
The differences between formative and summative assessments are presented in Table 2.1
below:
Feature
Purpose Monitor student progress and Evaluate student achievement at the
provide feedback to inform end of a learning period.
instruction.
Timing Ongoing throughout the learning Occurs at specific intervals, such as
process. the end of a unit or semester.
Examples Quizzes, exit slips, homework Final exams, standardized tests,
assignments, class discussions, projects, presentations.
peer reviews.
Feedback Frequent and specific feedback Less frequent and more general
to guide learning. feedback, often focused on overall
performance.
Grading Often ungraded or used for Typically graded and used for
informal assessment. formal evaluation.
Impact on Informs instructional decisions Used to evaluate the effectiveness of
Instruction and adjustments. instruction and curriculum.
assessment.
Computer-based tests (CBTs): These tests are administered on computers and can
include multiple-choice questions, fill-in-the-blank exercises, and short-answer
questions.
Computer-adaptive tests (CATs): CATs tailor questions to the test-taker's ability
level, adjusting the difficulty as the test progresses. CATs adjust the difficulty of
questions based on the test-taker's performance, ensuring a more accurate
assessment of their ability level.
Example: A language proficiency test might start with easier questions. If the test-
taker answers correctly, the next question will be more difficult. If they answer
incorrectly, the next question will be easier. This ensures that the test is both
challenging and fair.
Automated speech recognition (ASR): ASR technology can assess oral language
skillsbyanalyzingspeechpatterns,pronunciation,andfluency. ASRtechnology can
assess oral language skills by analyzing speech patterns, pronunciation, and
fluency.
Example: A language learner can practice speaking into a computer or
smartphone, and the software can provide feedback on pronunciation, intonation,
and grammar.
Example: Duolingo and Babbel use adaptive learning algorithms to tailor exercises
to each learner's needs. They provide immediate feedback on grammar,
vocabulary, and pronunciation.
Example: A student can submit an essay to an OWL, which will check for
grammar, spelling, and punctuation errors. The software can also provide
suggestions for improving sentence structure and style.
These are just a few examples of how technology can be used to enhance language
assessment. As technology continues to evolve, we can expect to see even more
innovative and effective ways to assess language skills.
Additional considerations
Procedure: Students work in pairs to examine the following excersies and asnwer the
following quesitons
Presentation: The teacher presents a grammar rule, such as the past tense of
regular verbs.
Translation: Students translate sentences from their L1 (native language) into the
target language, applying the grammar rule.
Correction: The teacher corrects the exercises and provides feedback. Example
grammar rule:
• Past tense of regular verbs: add -ed to the base form of the verb.
o Example: play -> played, work -> worked
Example Exercise:
Fill in the blanks with the correct past tense form of the verb.
Procedure:
1. How do you feel when you get a good grade? a) happy b) sad c) angry d) tired
2. What do you feel when you fail a test? a) happy b) sad c) excited d) tired
What do you think about the activities? How should they be applied in language classes
for high school students in Vietnam.
Example scenario:
• Restaurant role-play:
o Student A: You are a customer at a restaurant. Order your favorite dish and a
drink.
o Student B: You are a waiter/waitress. Take the customer's order and respond
appropriately.
Procedure: Students examine the following project which aims to assess students'
comprehensive understanding of a specific topic or theme, their ability to apply
knowledge and skills, and their communication skills. Students decide if formative or
summative assessment should be used to evaluate students’ project and give reasons.
Task: Students will work in groups of 3-4 to create a multimedia presentation on the topic
of global warming. The presentation should include the following components:
Introduction: A brief overview of global warming, including its causes and effects.
Effects: A discussion of the impacts of global warming on the environment, society, and
economy.
Procedure: Students examine the following role play for communicative language
assessment which aims to assess students' ability to use language in a real-world context,
focusing on fluency, accuracy, and appropriateness. Then choose activity in an English
textbook currently in use at high schools in Vietnam and design a similar activity for
communicative language assessment
Task: Students will be assigned roles in a specific scenario. They will be required to
interact with each other, using target language to accomplish a task or solve a problem.
Traveler: Explain the situation to the airline representative, provide details about
the lost luggage, and ask for assistance.
Airline Representative: Listen to the traveler's problem, offer solutions, and
provide information about the lost luggage procedures.
Assessment criteria:
Feedback: The teacher will observe the students' performance and provide feedback
on their strengths and weaknesses. Feedback can be given orally or in written form.
References
Why?
o This exercise focuses on discrete-point tasks such as filling in blanks and
translating sentences.
o Emphasizes accuracy in grammar and vocabulary.
o Uses teacher-centered correction instead of communicative or contextualized
feedback.
o Relies on standardized testing formats like cloze tests and translation exercises.
Application in Vietnamese high school classrooms:
o Can be useful for reinforcing grammar rules in controlled practice.
o Should be supplemented with communicative tasks to ensure students develop
practical language use.
Why?
o Uses multiple-choice questions, a discrete-point testing method.
o Focuses on isolated vocabulary knowledge rather than language use in context.
o Uses standardized scoring without assessing communication skills.
Application in Vietnamese high school classrooms:
o Effective for testing vocabulary knowledge quickly.
o Can be useful for formative assessment, allowing teachers to identify areas where
students need support.
o Should be combined with contextualized vocabulary activities (e.g., sentence
creation, role-plays) for better retention.
Why?
o Uses authentic tasks (e.g., role-plays, discussions).
o Focuses on communicative competence rather than just linguistic accuracy.
o Involves holistic scoring, considering fluency, coherence, and interaction.
o Encourages learner-centered assessment, where students actively participate.
Conclusion
This activity effectively applies the communicative approach to language assessment. It helps
students develop real-world English proficiency while making the assessment process more
engaging and interactive.
Activity 3: Analyzing Language Assessment Approach(es) for a Unit in an
English Textbook in Vietnam
The best assessment approach depends on the learning objectives of the unit.
A combination of traditional and communicative approaches is often ideal for
comprehensive assessment.
Conclusion
A hybrid assessment approach is recommended, integrating traditional for accuracy and
communicative for fluency and application.
Teachers should align assessments with real-world tasks to ensure students develop
practical English skills.
Activity 4: Deciding Whether to Use Formative or Summative Assessment
The project on Global Warming involves research, presentation, and collaboration. The decision
to use formative or summative assessment depends on the purpose of evaluation.
Assessment
Purpose Examples When to Use?
Type
Quizzes, self-assessments, During the learning process
Provides feedback
Formative peer reviews, teacher to help students adjust their
for improvement
observations work
At the end of a unit or term
Evaluates final Final exams, standardized
Summative to measure overall learning
achievement tests, project presentations
outcomes
Recommended
Project Component Why?
Assessment Type
Brainstorming ideas and Allows students to refine their ideas based
Formative
research phase on feedback.
Drafting and creating Helps students improve content, structure,
Formative
presentation materials and clarity before final submission.
Final multimedia Assesses students' ability to apply
Summative
presentation knowledge and communicate effectively.
Self-reflection on the Encourages students to evaluate their
Formative
learning process progress and set goals for improvement.
Conclusion
A combination of formative and summative assessment is ideal for this project.
Formative assessment helps improve student performance before the final product.
Summative assessment ensures that students demonstrate their overall understanding and
skills.
Activity 5: Unpacking Communicative Language Assessment
1. Role A (Traveler): Explains the situation about lost luggage and asks for help.
2. Role B (Airline Representative): Responds appropriately, asks clarifying questions, and
provides solutions.
Criteria Description
Fluency Speaks smoothly and naturally without excessive pauses.
Accuracy Uses correct grammar, vocabulary, and pronunciation.
Appropriateness Uses polite and contextually appropriate expressions.
Problem-Solving Identifies the problem and suggests logical solutions.
Interactivity Engages naturally in conversation, responds appropriately.
Conclusion
3.1 Reliability
Reliability in language testing refers to the consistency and dependability of test scores
(Bachman & Palmer, 1996; Liu, et al., 2020). A reliable test produces consistent results
across different administrations, raters, or test items. In other words, a reliable test
minimizes random error or fluctuations in scores that are not due to true differences in
language proficiency. Ensuring reliability is crucial for making valid inferences about
test-takers' language abilities based on their test scores.
Types of reliability
Test-retest reliability: This method involves administering the same test to the
same group of test-takers on two separate occasions and then correlating the two
sets of scores. A high correlation coefficient indicates that the test produces
consistent results over time. However, this method can be affected by factors such
as practice effects and memory.
Procedure: A group of applicants takes the test. After a suitable interval (e.g., two
weeks), the same group of applicants takes the exact same test again under similar
conditions. The scores from both administrations are compared for each
individual.
High reliability: If the scores from the first and second administrations are highly
correlated (e.g., a correlation coefficient close to 1.0), it suggests that the test
produces consistent results over time. This indicates that the test is reliable
because it is measuring the same underlying language proficiency consistently,
regardless of when it is administered.
Low reliability: If the scores show little or no correlation between the two
administrations, it suggests that the test is unreliable. This could be due to factors
like:
Overall, test-retest reliability assesses the consistency of test scores over time. A reliable
test should produce similar results when administered to the same group of test-takers
under comparable conditions.
Parallel forms reliability: This method requires the creation of two equivalent
versions of the same test, which are then administered to the same group of test-
takers. The scores on the two versions are then correlated. A high correlation
coefficient suggests that the two versions of the test are equivalent and measure
the same construct consistently.
Scenario:
Parallel Forms: Two equivalent versions of the test (Form A and Form B) are
created.
o Equivalence: Both forms cover the same range of English language skills
(reading, grammar, vocabulary) with a similar level of difficulty.
Example:
Reading Comprehension: Form A might include a passage
about historical figures, while Form B includes a passage
about scientific discoveries, but both passages assess similar
reading comprehension skills (main idea, supporting details,
inference).
Grammar: Both forms might include questions on verb
tenses and subject-verb agreement, but with different
sentence structures and contexts.
Vocabulary: Both forms might include vocabulary items
related to common academic topics, but with different word
choices and question formats (e.g., multiple choice, fill-in-
the-blank).
Procedure:
High reliability: If the scores from Form A and Form B are highly correlated
(e.g., a correlation coefficient close to 1.0), it suggests that the two forms are
equivalent and measure the same underlying English language proficiency
consistently. This indicates that the test is reliable because it produces similar
results regardless of which version is administered.
Low reliability: If the scores show little or no correlation between the two forms,
it suggests that the two forms are not equivalent and may not be measuring the
same construct. This could be due to differences in difficulty, content coverage, or
question format between the two forms.
In short, parallel forms reliability in this context ensures that the English achievement test
consistently measures students' English language proficiency, regardless of which version
they take. This is crucial for making fair and accurate assessments of student learning and
for identifying areas where students may need additional support.
Scenario:
Test: A standardized English proficiency test for university applicants, designed
to assess reading comprehension.
Example:
Overall, internal consistency reliability ensures that the items within a test section are
measuring the same underlying construct and are not measuring unrelated or inconsistent
skills. In this example, a high Cronbach's alpha for the reading comprehension section
would suggest that the test is reliably measuring general reading comprehension ability.
Scenario:
Test: An English speaking exam for university students
Inter-rater reliability consideration: Since the assessment involves subjective
judgment by human raters (interviewers), it is crucial to ensure consistent scoring
across different raters.
Example:
Inter-rater reliability analysis: The scores given by the two raters for each
applicant are compared. Statistical methods like Cohen's kappa or intraclass
correlation coefficients are used to calculate the level of agreement between the
raters.
Overall, inter-rater reliability ensures that the scores assigned to test-takers are consistent
across different raters. In the case of oral proficiency interviews, high inter-rater
reliability is essential to ensure that the assessment is fair and objective.
Improving reliability
3.2 Validity
In the realm of language assessment, the concept of validity holds paramount importance.
Validity, in its essence, refers to the extent to which a language test accurately measures
what it purports to measure (Chapelle & Lee, 2021). It is a multifaceted concept that
encompasses various aspects, including content, construct, criterion, and consequential
validity. This section delves into the different types of language test validity, exploring
their significance in ensuring the accuracy, fairness, and meaningfulness of language
assessments.
Content validity
Content validity focuses on the degree to which a test comprehensively covers the
relevant content domain (Brown & Abeywickrama, 2019; Hughes, 2020). In language
testing, this entails ensuring that the test tasks and items adequately represent the
knowledge, skills, and abilities that are considered essential for the target language use
domain (Dinh, 2019; Siddiek, 2010). For instance, a test designed to assess academic
writing proficiency should include tasks that reflect the types of writing required in
academic settings, such as essays, research papers, and literature reviews. The following
scenario illustrates content validity:
Scenario:
Test purpose: An English test aims to assess the reading comprehension skills of
university students in an academic setting.
Content Validity Consideration: To ensure content validity, the test developers
would need to carefully select reading materials that are representative of the types
of texts students will encounter in their university studies.
Example:
Instead of using simplified or general interest texts, the test should include:
By including these types of materials, the test directly measures the skills students need
to succeed in their academic reading. It avoids irrelevant content (like casual
conversations or simple narratives) that wouldn't accurately reflect their ability to handle
university-level texts.
Overall, content validity is about ensuring the test's content aligns with the specific skills
or knowledge it intends to measure. In this case, the test content is valid because it
reflects the reading demands of a university environment.
Construct Validity
Construct validity examines whether a test effectively measures the underlying construct
or attribute it is intended to assess (Hill & McNamara, 2015). Constructs in language
testing can include language proficiency, communicative competence, or specific
language skills like reading comprehension or listening ability. To establish construct
validity, researchers often employ various methods, such as correlating test scores with
other established measures of the same construct or examining the internal structure of
the test through factor analysis. The followoing scenario illustrates construct validity:
Scenario:
Test purpose: An English test aims to assess the writing ability of students,
specifically their ability to construct well-formed sentences.
Construct validity consideration: To ensure construct validity, the test needs to
accurately measure the underlying construct of "sentence construction ability."
This means it should effectively differentiate between students who have a good
grasp of sentence structure and those who struggle with it.
Example:
The test includes various tasks designed to elicit different aspects of sentence
construction:
By using a variety of tasks that target different facets of sentence construction, the test
provides a more comprehensive assessment of this underlying construct. It avoids relying
on just one type of task, which might only measure a narrow aspect of sentence
construction ability.
In general, construct validity is about ensuring the test accurately measures the theoretical
construct it intends to measure. In this case, the test is designed to measure the construct
of "sentence construction ability" by using tasks that effectively tap into different aspects
of this skill.
Scenario:
Example:
Test: The university administers its own English language proficiency test to
applicants.
Criterion: The university tracks the academic performance (e.g., GPA) of the
admitted students in their first year of study.
Analysis: The university then correlates the test scores with the students' academic
performance.
If the test scores show a strong positive correlation with the students' academic
performance (i.e., higher test scores predict better academic outcomes), it demonstrates
that the test has good criterion-related validity. It means the test is effective at predicting
how well students will perform in their academic studies, which is the intended purpose
of the test.
In short, criterion-related validity focuses on how well a test's scores correlate with an
external criterion (in this case, academic performance). A strong correlation indicates that
the test is useful for predicting future performance on that criterion.
Consequential validity
Consequential validity considers the social and personal consequences of test use. It
examines the impact of a test on individuals, groups, and society as a whole. This
includes evaluating the potential benefits and drawbacks of test use, such as its effect on
learning, teaching, and educational policy. For instance, a test that is used for high-stakes
decisions, such as university admission or job placement, should be carefully scrutinized
for its potential impact on test takers' lives and opportunities. The following scenario
illustrates consequential validity”
Scenario:
Example:
Positive consequence: The test accurately identifies students with strong English
skills, allowing them to be placed in challenging classes where they can thrive and
learn at a faster pace. This can lead to increased motivation, higher academic
achievement, and greater opportunities for these students.
Negative consequence: The test may unfairly disadvantage certain groups of
students, such as those from low-income backgrounds or those who are English
language learners. If the test is culturally biased or does not adequately account for
their unique learning experiences, it could lead to misplacement, frustration, and
lower self-esteem. This could also perpetuate existing educational inequalities.
Conduct a thorough needs assessment: Identify the specific needs and learning
goals of the target population.
Minimize potential bias: Ensure the test is culturally fair and does not
disadvantage any particular group of students.
Monitor the impact of test use: Track the long-term consequences of test-based
placement decisions on student learning and well-being.
Make adjustments as needed: Based on the monitoring data, make adjustments
to the test or the placement process to mitigate any negative consequences.
The significance of language test validity lies in its ability to ensure that language
assessments are accurate, fair, and meaningful. A valid language test provides reliable
information about test takers' language abilities, enabling stakeholders to make informed
decisions based on the test results. For test takers, a valid test ensures that their language
skills are assessed fairly and accurately, providing them with a true reflection of their
abilities. For educators, a valid test helps to identify students' strengths and weaknesses,
guiding instructional practices and curriculum development. For policymakers, a valid
test can inform decisions about language education programs and policies, ensuring that
they are effective and equitable.
Overall, language test validity is a complex and multifaceted concept that plays a crucial
role in ensuring the quality and meaningfulness of language assessments. By
understanding the different types of validity and their significance, language testers and
educators can strive to develop and utilize assessments that are accurate, fair, and
beneficial for all stakeholders.
3.3 Practicality
Time constraints: Time is a valuable resource, both for test-takers and test
administrators. A practical test should be administered within a reasonable
timeframe. Test-takers should not feel rushed or pressured, but the test should also
not be excessively long, as this can lead to fatigue and reduced performance. Test
administrators should also consider the time required for test preparation,
administration, scoring, and reporting.
Cost-effectiveness: The cost of developing, administering, and scoring a language
test can vary significantly. Test development requires time, expertise, and
resources, such as item writing, pilot testing, and statistical analysis.
Administration costs include venue rental, equipment, and personnel. Scoring
costs can be substantial, especially for subjective assessments that require human
raters. A practical test should be cost-effective, considering the available budget
and the value derived from the test results.
Logistical feasibility: Logistical considerations include the availability of
appropriate testing venues, equipment, and personnel. Test administrators need to
ensure that the testing environment is conducive to optimal test performance, with
adequate space, lighting, and ventilation. They also need to ensure that sufficient
personnel are available to administer the test, monitor test-takers, and provide any
necessary assistance.
Ease of administration and scoring: A practical test should be easy to administer
and score. Clear and concise instructions should be provided to test-takers, and the
testing procedures should be straightforward and easy to follow. Scoring
procedures should be well-defined and easy to apply, reducing the potential for
subjectivity and inconsistency. The use of automated scoring systems can
significantly improve the efficiency and objectivity of the scoring process. Below
is an example:
Scenario:
Example:
Overall, the use of a multiple-choice format enhances the practicality of the test by
making it easy to administer and score. This is especially important for high-stakes tests
that need to be administered to a large number of students efficiently. A practical test
should be accessible to all test-takers, regardless of their physical or cognitive abilities.
Reasonable accommodations should be made for test-takers with disabilities, such as
providing extra time, using assistive technology, or modifying test formats. The test
should also be culturally sensitive and appropriate for test-takers from diverse
backgrounds.
Test results should be easily interpretable and reported in a clear and concise manner.
Test reports should provide meaningful information about test-takers' language
proficiency, such as their strengths and weaknesses. They should also be easy to
understand and use by test-takers, educators, and other stakeholders.
Scenario:
Example:
Standardized Scoring Scale: The test uses a standardized scoring scale, such as a
score range from 1 to 100 or a band score system (e.g., 1-9).
Benefits:
o Easy interpretation: Standardized scores are easy to understand and
compare. Applicants and universities can easily see a student's overall
English proficiency level.
o Clear reporting: Test reports can be generated quickly and easily,
providing clear and concise information about the applicant's performance
in different areas of English language proficiency (reading, writing,
listening, speaking).
o Efficient decision-making: Universities can easily use the standardized
scores to make admissions decisions, compare applicants, and identify
students who may need additional language support.
Using a standardized scoring scale enhances the practicality of the test by making it easy
to interpret and report results. This facilitates efficient decision-making for both
universities and applicants.
Test bias
Test bias occurs when a test systematically favors or disadvantages certain groups of test-
takers based on factors such as gender, ethnicity, socioeconomic status, or cultural
background. One common type of bias is cultural bias, where test items or tasks are
culturally inappropriate or unfamiliar to test-takers from certain cultural backgrounds.
For example, a reading comprehension passage that references cultural concepts or events
that are unfamiliar to test-takers from a different cultural background may disadvantage
them.
Use culturally diverse materials: Include reading passages, listening materials, and
test items that are relevant and accessible to test-takers from diverse cultural
backgrounds.
Avoid culturally loaded language: Use language that is clear, concise, and free
from cultural idioms or slang that may be unfamiliar to certain groups.
Conduct pilot testing: Pilot test the test with diverse groups of test-takers to
identify and address any potential cultural biases.
Test security
Maintaining test security is crucial to ensure the integrity of the testing process. Test
materials should be kept confidential and should not be accessible to unauthorized
individuals. Measures should be taken to prevent cheating, such as proctoring exams
carefully, using secure testing environments, and implementing measures to detect and
prevent cheating.
Test results should be used responsibly and ethically. Test scores should not be used to
make high-stakes decisions without careful consideration of their limitations. It is
important to remember that test scores are just one piece of information about a test-
taker's language ability, and they should not be the sole determinant of important
decisions such as university admissions or job placement. Test results should be
interpreted in conjunction with other relevant information, such as academic records,
recommendations, and interviews.
Test-taker rights
Test-takers have certain rights that must be respected. These rights include:
Right to clear and concise instructions: Test-takers should be provided with clear
and unambiguous instructions for all test tasks.
Right to a fair and equitable testing environment: Test-takers should be provided
with a comfortable and supportive testing environment that is free from
distractions and disruptions.
Right to privacy and confidentiality: Test-takers have the right to expect that their
test scores and personal information will be kept confidential and used only for the
intended purposes.
Right to appeal test scores: Test-takers should have the right to appeal their test
scores if they believe there has been an error in scoring or administration.
In general, ethical considerations are paramount in all aspects of language testing. Test
developers and administrators have a responsibility to ensure that tests are fair, equitable,
and unbiased. By adhering to ethical principles and best practices, we can ensure that
language tests are used responsibly and that test-takers are treated with respect and
dignity.
Instruction: Carefully read each statement in the table below. For each statement, rate
your understanding using the following scale:
Objective: To have students analyze scenarios to identify possible causes for unreliability
of an English test.
Procedure: Divide students into small groups. Provide each group with a handout
containing 3-4 scenarios related to different types of reliability in language testing.
o Example Scenarios:
Scenario 1 (Test-retest): "A student takes an online English
proficiency test. A week later, they retake the same test. Their scores
differ significantly. What factors might have contributed to this
difference?"
Scenario 2 (Parallel Forms): "A school uses two different versions
of a reading comprehension test. Students who perform well on one
version consistently score poorly on the other. What could be the
potential issues with the test?"
Scenario 3 (Internal Consistency): "A vocabulary test includes
items that seem unrelated to each other. How might this affect the
internal consistency of the test?"
Scenario 4 (Inter-rater): "Two teachers score student essays. Their
scores for the same essays differ significantly. What steps can be
taken to improve inter-rater reliability?"
Each group analyzes the scenarios and discusses possible causes for unreliability.
Objective: To have students analyze a sample English test for high school students in
Vietnam and identify potential threats to its reliability.
Material: A sample English test (e.g., a reading comprehension test with multiple-choice
questions)
Procedure:
Test presentation: Present students with a sample English test (or students search
from the Internet). Briefly discuss the purpose of the test and the target language
skills it aims to assess.
Group analysis : Divide students into small groups. Provide each group with a
handout containing the following questions:
o Content: Does the test adequately cover the range of language skills taught
in the English curriculum?
o Construct: Does the test accurately measure the intended language skills
(e.g., reading comprehension, writing, speaking)?
Objective: To have students apply the principles of practicality (time, cost, logistics, ease
of administration/scoring, accessibility, interpretation/reporting) in the design of a short
language test.
Procedure: Present students with a scenario requiring the creation of a language test. For
example, "Design a short English test for incoming exchange students to assess their
basic conversational skills." Each group designs a short language test (e.g., 2-3 tasks)
based on the given scenario. Students discuss in groups to consider the following:
Objective: To have students analyze ethical dilemmas in language testing and develop
strategies for ensuring fair and equitable assessment practices.
Procedure: Divide students into small groups. Each group analyzes the following
questions and discusses potential threats to its reliability. Below is a list containing 3-4
ethical dilemmas related to language testing.
Scenario 1 (Test bias): "A reading comprehension test
includes a passage about a specific cultural event that would
be unfamiliar to students from certain cultural backgrounds.
How does this create bias, and how can it be addressed?"
Scenario 2 (Test security): "A student discovers a copy of
the upcoming English proficiency test online. What are the
ethical implications of this situation, and what steps should be
taken to address it?"
Scenario 3 (Responsible use of test results): "A university
uses an English proficiency test as the sole criterion for
admission to its English program. What are the potential
ethical concerns with this approach?"
Scenario 4 (Test-taker rights): "A student with a learning
disability requests extra time for an English test. How should
the school handle this request to ensure fairness and equity
for all students?"
Each group analyzes the scenarios, discussing the ethical implications and potential
solutions.
References
Bachman, L. F., & Palmer, A. S. (1996). Language testing in practice. Oxford University
Press.
Hill, K., & McNamara, T. (2015). Validity inferences under high-stakes conditions: A
response from language testing. Measurement: Interdisciplinary Research &
Perspectives, 13(1), 39-43.
Liu, Z., Li, T., & Diao, H. (2020). Analysis on the Reliability and Validity of Teachers'
Self-designed English Listening Test. Journal of Language Teaching and
Research, 11(5), 801-808.
Siddiek, A. G. (2010). The Impact of Test Content Validity on Language Teaching and
Learning. Online Submission, 6(12), 133-143.
Tang, W., Cui, Y., & Babenko, O. (2014). Internal consistency: Do we really know what
it is and how to assess it. Journal of Psychology and Behavioral Science, 2(2), 205-
220.
Chapter 4: Types of Language Tests
This chapter delves into the diverse landscape of language tests, examining the various
types employed to measure language skills and knowledge. To provide a structured
framework, the chapter will categorize these tests into two broad categories: classroom-
based language tests and standardized language tests. Classroom-based tests, often
developed and administered within specific educational contexts, reflect the immediate
needs of learners and instructors, emphasizing classroom-based assessments and teacher-
created materials. In contrast, standardized tests are externally developed, rigorously
validated, and designed to provide consistent and comparable measures of language
proficiency across diverse populations. Understanding the distinctions and applications of
these two categories is essential for educators, researchers, and test developers seeking to
make informed decisions about language assessment.
• Focus on isolated skills: These tests typically break down language into smaller
components, such as individual grammatical rules, vocabulary items, or
phonological features.
• Objective scoring: Discrete-point tests often lend themselves to objective scoring
methods, such as multiple-choice questions, true/false items, or matching
exercises. This makes scoring relatively quick and efficient.
• Emphasis on accuracy: The primary focus is on the accuracy of the learner's
response, rather than the overall communicative effectiveness.
• Grammar tests: These tests assess learners' knowledge of grammatical rules, such
as verb tenses, noun phrases, and sentence structure. Examples include fill-in-the-
blank exercises, error correction tasks, and multiple-choice questions that test
grammatical structures.
• Vocabulary tests: These tests measure learners' knowledge of individual words and
phrases. Common formats include multiple-choice vocabulary tests, matching
exercises, and cloze tests where learners fill in missing words in a text.
• Phonology tests: These tests assess learners' pronunciation, intonation, and other
aspects of spoken language. Examples include minimal pair discrimination tasks,
picture-cued elicitation tasks, and tests of stress and intonation.
Discrete-point tests have their place in language assessment, particularly for diagnosing
specific areas of difficulty and providing quick feedback on the mastery of particular
language skills. However, it is crucial to recognize their limitations and to use them in
conjunction with other assessment methods that more comprehensively assess language
proficiency in communicative contexts.
Integrative tests move beyond the isolated assessment of discrete language skills
(Alderson et al. 2014), such as grammar or vocabulary, and instead focus on assessing
how learners use language in a more holistic and communicative manner. These tests aim
to simulate real-life language use, requiring learners to integrate multiple language skills
to achieve a communicative goal (Brown, 2004; Weir, 2005).
Characteristics of integrative tests:
• Reading comprehension tests with complex tasks: These tests go beyond simple
multiple-choice questions and require learners to analyze, synthesize, and evaluate
information from longer texts. Tasks may include summarizing, paraphrasing,
comparing and contrasting different texts, and answering open-ended questions
that require critical thinking and interpretation.
• Writing tasks: These tasks may include essay writing, report writing, letter writing,
or creative writing. They assess learners' ability to organize ideas, develop
arguments, use appropriate language, and produce coherent and effective written
communication.
• Speaking tests: These tests often involve interactive communication tasks, such as
role-plays, discussions, and interviews. They assess learners' fluency, accuracy,
pronunciation, and ability to interact effectively with others.
• Listening comprehension tests: These tests may involve listening to lectures,
conversations, or other authentic audio materials and then answering questions
that require learners to understand and interpret the information.
• Difficulty in scoring: Scoring integrative tests can be more subjective and time-
consuming than scoring discrete-point tests.
• Potential for rater bias: Subjective scoring can be influenced by rater bias, leading
to inconsistencies in scoring.
• Difficulty in standardizing: It can be challenging to standardize the administration
and scoring of integrative tests, which can make it difficult to ensure fairness and
consistency across different administrations.
In general, integrative language tests play a crucial role in providing a more holistic and
authentic assessment of learners' language proficiency. While they may present some
challenges in terms of administration and scoring, their value in assessing communicative
competence makes them an essential component of any comprehensive language
assessment program.
Diagnostic tests play a crucial role in effective language teaching and learning. Unlike
summative assessments that primarily focus on measuring overall achievement,
diagnostic tests aim to pinpoint specific areas of language difficulty for individual
learners (Joughin, 2007). This information allows teachers to tailor instruction, provide
targeted support, and create a more personalized learning experience for each student.
Overall, diagnostic tests play a vital role in effective language teaching and learning. By
providing valuable insights into learners' strengths and weaknesses, diagnostic tests
enable teachers to provide more effective and personalized instruction, ultimately leading
to improved learning outcomes.
Progress tests are a crucial component of effective language teaching and learning. They
provide valuable feedback on students' learning and help teachers monitor the
effectiveness of their instruction. Unlike summative assessments, which typically occur at
the end of a course or unit, progress tests are administered periodically throughout the
learning process to track students' progress and identify areas for improvement.
• Formative in nature: Progress tests are primarily used for formative assessment,
providing ongoing feedback to both teachers and learners.
• Aligned with learning objectives: They are designed to assess specific learning
objectives covered in the course or unit.
• Regular administration: Progress tests are typically administered at regular
intervals, such as at the end of each week, month, or unit.
• Short and focused: They are usually shorter than summative tests and focus on a
specific set of skills or knowledge areas.
• Actionable feedback: The results of progress tests should provide specific and
actionable feedback to both teachers and learners, identifying areas of strength and
weakness.
Types of progress tests:
• Improved learning: Regular feedback from progress tests can help students
identify areas for improvement and adjust their learning strategies accordingly.
• Informed instruction: Progress test results provide valuable information to
teachers, allowing them to identify areas where students may be struggling and to
adjust their teaching methods accordingly.
• Increased motivation: Regular feedback can help to motivate students by
demonstrating their progress and celebrating their achievements.
• Early identification of difficulties: Progress tests can help to identify potential
learning difficulties early on, allowing teachers to provide targeted support and
prevent students from falling behind.
Ethical considerations:
• Use of results: Progress test results should be used constructively and should not
be used for high-stakes decisions, such as grading or placement.
• Feedback: Feedback from progress tests should be provided in a timely and
constructive manner.
• Confidentiality: Student performance on progress tests should be treated with
confidentiality.
Generally, progress tests play a vital role in supporting effective language learning. By
providing regular feedback and identifying areas for improvement, they help both
teachers and learners to monitor progress, adjust instruction, and ensure that students are
on track to achieve their learning goals.
Achievement tests are designed to measure how well learners have mastered specific
language skills or knowledge that have been taught within a particular course or
curriculum. Unlike proficiency tests, which assess overall language ability, achievement
tests focus on evaluating learners' progress in relation to specific learning objectives
(Brown, 2004; Joughin, 2007).
• Fairness and equity: Achievement tests should be fair and equitable for all
learners, regardless of their background or learning style.
• Use of results: Achievement test results should be used responsibly and ethically,
primarily for formative purposes and to provide constructive feedback to learners.
• Test security: Maintaining the security of achievement tests is crucial to ensure the
integrity of the assessment process.
Standardized language tests are formal assessments with standardized procedures and
scoring criteria. They are widely used for various purposes, including university
admissions, immigration, professional certification, and program placement. These tests
typically involve a large-scale administration to a large number of test-takers, ensuring
consistency across different testing locations and administrations.
• Standardized procedures:
o Consistent administration protocols across all testing locations.
o Clear and unambiguous instructions for test-takers.
o Controlled testing environment to minimize distractions.
• Objective scoring:
o Use of scoring rubrics and standardized scoring keys to minimize rater
bias.
o Often involve machine-scoring or computer-assisted scoring.
• Large-scale administration:
o Designed to be administered to a large number of test-takers
simultaneously.
o Efficient and cost-effective for large-scale testing programs.
• High-stakes implications:
o Often used for high-stakes decisions, such as university admissions or
immigration applications.
o Scores have significant implications for test-takers' future opportunities.
• Limited flexibility: May not adequately assess the full range of language skills and
abilities.
• Potential for bias: May not be culturally fair or sensitive to the needs of all test-
takers.
• Limited individualization: May not account for individual learning styles and
needs.
• High-stakes nature: Can create undue pressure and anxiety for test-takers.
Ethical considerations
• Test security: Maintaining test security is crucial to ensure the integrity of the
testing process.
• Fairness and equity: Standardized tests must be fair and equitable for all test-
takers, regardless of their background, culture, or learning style.
• Appropriate use of scores: Test scores should be used responsibly and ethically,
and should not be the sole determinant of important decisions.
Overall, standardized language tests play an important role in various contexts, from
education and immigration to professional certification. While they offer several
advantages, such as objectivity and comparability, it is crucial to recognize their
limitations and to use them responsibly and ethically. By carefully considering the
strengths and weaknesses of standardized tests and using them in conjunction with other
assessment methods, we can ensure that they provide a fair and accurate assessment of
language proficiency.
Instruction: Carefully read each statement in the table below. For each statement, rate
your understanding using the following scale:
Objectives:
Objectives:
Procedure:
• Students in groups briefly review the key characteristics of progress tests. They
then choose a short reading passage suitable for high school students
(approximately 200-300 words) from an English textbook for high schools in
Vietnam and design some reading comprehension test items (e.g., multiple-choice
questions, true/false statements, short answer questions)
Activity 3: Designing an achievement test for Vietnamese high school students
Objectives:
Procedure: Students work in groups briefly review the key characteristics of achievement
tests, emphasizing their purpose in measuring student learning in relation to specific
curriculum objectives. Each group designs an outline for a short achievement test. The
outline should include:
Objectives:
Procedure: Students discuss in groups the guidelines for designing portfolio projects (e.g.,
focus on learning objectives, variety of assessment methods, student reflection) and
relevant sections of the Vietnamese National Curriculum for English (e.g. Chương trình
giáo dục phổ thông, MOET, 2018). Students must include a variety of assessment
methods (e.g., essays, presentations, creative writing, reflective journals), clarity of the
assessment criteria and scoring rubrics.
Activity 5: Designing a role-play for communication skills
Objectives:
References:
Alderson, J. C., Clapham, C., & Wall, D. (2014). Language test construction and
evaluation. Cambridge University Press.
MOET (2018). Thông tư ban hành giáo dục phổ thông. Available at:
[Link]
[Link]
Crafting a language test requires careful planning, meticulous execution, and a deep
understanding of both the material (language) and the purpose of the creation
(assessment). This chapter, delves into the intricate art of language test construction,
guiding you through the process of transforming theoretical knowledge into practical
assessment tools.
Designing a language test requires planning, systematic design, and quality control. This
section will present the key steps involved in developing an effective language test, from
defining its purpose and objectives to administering it in a standardized manner. The
following these steps can ensure that a language test is valid, reliable, and fair to test-
takers.
• Deciding a test’s purpose: Clearly establish the reason for the test (e.g., placement,
proficiency, achievement). Identify the target test-takers, including their age,
background, and language learning experience.
• Deciding the test’s objectives: Determine the specific language skills and
knowledge that the test should assess.
• Designing the test’s specification: Create a detailed blueprint of the test, outlining
the content, format, and scoring criteria. Specify the language skills to be tested,
the types of test items to be used, and the weighting of different sections.
• Writing the test items: Create test items that align with the test specifications and
accurately measure the intended language skills. Ensure that the items are clear,
unambiguous, and appropriate for the target population.
• Reviewing and moderating test items: Have experts review the test items for
clarity, accuracy, and fairness. Revise or discard items that are problematic.
• Administer the test in a standardized and controlled environment: Ensure that all
test-takers have equal opportunities to demonstrate their language
skills/competence.
Objective test items, with their clear-cut right or wrong answers, offer a world of
efficiency and objectivity in language assessment (Hughes, 2020). This section delves
into the realm of these structured assessment tools, exploring their diverse forms and
applications. From the familiar multiple-choice questions to true-false, matching, and
gap-fill items, the following section will unravel the characteristics that make them a
staple in language testing.
Multiple-choice questions are a common and versatile format used in various language
tests, from classroom quizzes to high-stakes standardized examinations. They offer
several advantages in terms of objectivity, efficiency, and ease of scoring, making them a
popular choice for assessing knowledge and skills in a wide range of language areas,
including grammar, vocabulary, reading comprehension, and listening comprehension
(Currie & Chiramanee, 2010).
• Stem: The initial part of the question, which presents the problem or situation to
be addressed.
• Options: A set of possible answers, usually consisting of three or four choices.
• Key: The correct answer among the options.
• Distractors: The incorrect answer options, designed to be plausible but incorrect.
The following diagram illustrates the components of a multiple choice test item
Below are examples of multiple choice items from the English textbook Ilearn Smart
World 9
Faulty MC items
When designing MC test items it is essential to avoid faulty MC items. Below are some
examples of faulty MC items:
Ambiguous question
Question: "What is the meaning of the word 'bank' in this sentence: 'I went to the
bank to deposit my paycheck.'"
(A) A financial institution
(B) The side of a river
(C) A row of seats
Fault: The word "bank" has multiple meanings, and the sentence does not provide
enough context to determine the intended meaning. This makes the question
ambiguous and potentially confusing for students.
More than one correct answer
Question: Which of the following are topics covered in Unit 3 of the English 10 textbook
in Vietnam?
(A) Environmental pollution
(B) Natural disasters
(C) Protecting endangered species
Fault: Environmental pollution, natural disasters, and Protecting endangered species are
often interconnected and could be addressed together.
Unclear instructions
Question: Read the following passage and...
(A) Identify the main idea.
(B) Find the supporting details.
(C) Determine the author's purpose.
Fault: The instructions are too broad. Students need more specific guidance on what they
are supposed to do with the passage.
Misleading distractors
Question: What is the correct spelling of the word?
(A) Accommodate
(B) Acommodate
(C) Accomodate
Fault: The distractors are too similar to the correct answer, making it difficult to identify
the correct spelling. This tests students' ability to spot minor differences rather than their
actual knowledge of spelling rules.
• Clear and concise stem: The stem should be clear, concise, and unambiguous. It
should present the problem or situation to be addressed in a straightforward
manner.
• Plausible distractors: Distractors should be plausible but incorrect. They should be
grammatically correct and relevant to the stem, but not the correct answer.
• Grammatically correct options: All options, including the key and the distractors,
should be grammatically correct.
• Avoid ambiguity: The stem and options should be free from ambiguity and double
negatives.
• Test a single skill: Each item should focus on a single skill or concept.
In general, multiple-choice items are a valuable tool for language assessment, offering
several advantages in terms of objectivity, efficiency, and ease of scoring. However, it is
crucial to recognize their limitations and use them appropriately in conjunction with other
assessment methods that more comprehensively assess language proficiency. When
designing multiple-choice items, careful attention should be paid to clarity, plausibility,
and the overall quality of the distractors to ensure that the items are effective and reliable.
• Comprehend explicitly stated information: Identify key details and understand the
literal meaning of the text.
• Draw inferences: Make logical deductions based on the information provided.
• Identify the scope of the text: Recognize what information is included and,
crucially, what is not.
To maximize the effectiveness and reliability of T/F/NG items, test developers should
consider the following:
• Clear and concise statements: Avoid complex language and ensure each statement
focuses on a single idea.
• Unambiguous wording: Statements should have only one possible interpretation
based on the text.
• Variety of difficulty levels: Include a range of items that assess both explicit and
implicit understanding.
• Avoid verbatim copying: Rephrase information from the text to prevent simple
matching exercises.
• Thorough review and piloting: Ensure items are accurate, clear, and free of bias
before including them in a test.
Fault: This statement is too general and relies on a potentially inaccurate cultural
stereotype. Musical preferences vary greatly among individuals, and there is no
definitive answer without specific data or context.
Fault: This detail might be mentioned in the text, but it's not crucial to the main
storyline or character development. Focusing on such trivial information does not
effectively assess comprehension of the key themes or messages.
Statement: "It is not impossible to learn English fluently with dedicated practice."
(T/F/NG)
Fault: The double negative ("not impossible") makes this statement unnecessarily
convoluted. It's better to rephrase it in a more straightforward way (e.g., "It is
possible to learn English fluently with dedicated practice.")
Overall, T/F/NG items are a valuable tool in language testing, providing an efficient and
objective means of assessing reading comprehension. By understanding their
characteristics, applications, and potential limitations, test developers can effectively
utilize this item type to create reliable and valid language assessments. However, careful
attention must be paid to item construction and piloting to ensure clarity, avoid
ambiguity, and promote accurate measurement of language proficiency.
Cloze tests are a versatile and widely used tool in language education. They involve
presenting learners with a text where certain words have been systematically removed
(typically every fifth, sixth, or seventh word) and replaced with blanks. Learners are then
tasked with filling in these blanks with appropriate words based on their understanding of
the context and their language knowledge. This chapter explores the nature of cloze tests,
their applications in language learning, their advantages and limitations, and best
practices for constructing and implementing them effectively.
By requiring learners to actively engage with the text and make informed choices about
missing words, cloze tests encourage deep processing of language and promote both
receptive and productive skills.
Cloze tests can be used for a variety of purposes in language teaching and assessment:
• Fixed-ratio deletion: Words are removed at regular intervals (e.g., every fifth
word). This is the most common type of cloze test.
• Variable-ratio deletion: Words are removed based on specific criteria, such as
targeting particular grammatical structures or vocabulary items.
• Rational deletion: Words are removed based on their importance for understanding
the text, creating a more challenging and nuanced assessment.
• C-test: The second half of every second word is deleted, requiring learners to
complete the words based on the remaining letters.
• Potential for ambiguity: Some blanks may have multiple possible answers,
making scoring subjective.
• Limited scope: May not fully capture complex aspects of language use, such as
pragmatic understanding or communicative competence.
• Artificiality: The artificial nature of deleting words can sometimes disrupt the
natural flow of the text.
• Choose appropriate texts: Select texts that are relevant to learners' interests and
proficiency levels.
• Determine deletion rate: Consider the difficulty level desired and the specific
skills being assessed.
• Provide clear instructions: Ensure learners understand the task and how to
respond.
• Pilot test the cloze: Administer the test to a small group of learners to identify any
ambiguities or issues.
• Provide feedback: Use the cloze test as a learning opportunity by providing
feedback on learners' responses and discussing the rationale behind correct
answers.
Below are the sample faulty items in a cloze test that should be avoided.
Fault: This sentence lacks sufficient context to determine the missing word. It could
be "museum," "zoo," "park," "historical site," or any other place that students might
visit on a field trip. This makes the item ambiguous and potentially frustrating for
test-takers.
Sentence: "Although it was raining ____, we decided to go for a walk in the park."
Fault: Placing the gap in the middle of the adverbial clause disrupts the natural flow
of the sentence. It makes it harder for students to understand the sentence structure
and choose the correct word (e.g., "heavily," "outside," etc.).
Fault: This item assumes knowledge of Vietnamese culture that might not be
relevant to all students, especially in an EFL context. It's better to choose words or
concepts that are more universally known or covered within the textbook's scope.
Fault: Including grammatically incorrect options can confuse students and hinder
their ability to identify the correct form of the verb (in this case, "does").
In short, cloze tests are a valuable tool in language education, offering a flexible and
efficient means of assessing and developing various language skills. By carefully
considering the principles of cloze test construction and implementation, educators can
harness their potential to enhance language learning and promote learner engagement and
autonomy.
Matching items are a common and effective question type found in English language
tests for high school students in Vietnam. They require students to connect related pieces
of information from two separate columns, testing their ability to recognize relationships,
analyze information, and make accurate connections. This chapter explores the
characteristics, applications, advantages, and limitations of matching items in English
tests, along with guidelines for constructing effective matching activities.
Understanding matching items
Students are tasked with identifying the correct relationships between the items in the two
columns and indicating the matches. This process assesses their ability to:
Matching items are versatile and can be used to assess various aspects of English
language proficiency:
• Efficiency: Can assess a wide range of knowledge and skills in a concise format.
• Objectivity: Scoring is straightforward and less prone to subjective interpretation.
• Clarity: The format is generally easy for students to understand and follow.
• Versatility: Can be adapted to assess different language skills and levels of
difficulty.
Limitations of matching items
• Limited cognitive demand: May primarily test recognition and recall rather than
higher-order thinking skills.
• Potential for guessing: Students may resort to guessing if they are unsure of the
correct answers, especially if the number of items in each column is equal.
• Difficulty in constructing effective items: Creating meaningful and unambiguous
matches can be challenging.
To maximize the effectiveness and validity of matching items, test developers should
consider the following:
Column A Column B
a) happy 1. big
b) large 2. joyful
c) sad 3. unhappy
d) angry 4. furious
e) small 5. tiny
Fault: Some words in Column A could have multiple synonyms in Column B. For
example, "happy" could be matched with both "joyful" and "glad" (if it were an
option), while "large" could be matched with both "big" and "huge" (if it were an
option). This ambiguity makes it difficult for students to determine the single best
match.
Column A Column B
a) photosynthesis 1. the process of making food in plants using sunlight
b) gravity 2. a type of animal that lives in the ocean
c) whale 3. a force that pulls objects towards each other
Fault: The item "whale" and its definition "a type of animal that lives in the ocean"
are not relevant to the other vocabulary words and definitions, which are related to
science and social studies. This creates a mismatch and can distract students.
Instructions: Match the sentences in Column A with the correct question tags in
Column B.
Column A Column B
a) She is a doctor, 1. isn't she?
b) They are playing football, 2. aren't they?
c) He has finished his work, 3. hasn't he?
d) We will go to the party, 4. won't we?
Overall, matching items are a valuable tool in English language tests for high school
students in Vietnam. They provide an efficient and objective way to assess various
language skills and knowledge areas. By adhering to the principles of effective item
construction and considering the potential limitations, educators can utilize matching
activities to create reliable and valid assessments that contribute to meaningful language
learning.
5.2.5 Designing gap-fill items
Gap-fill items, also known as fill-in-the-blank questions, are a staple in English language
tests for high school students in Vietnam. These questions require students to complete a
sentence or passage by filling in missing words or phrases, demonstrating their
understanding of grammar, vocabulary, and overall language structure. This section
delves into the nature of gap-fill items, their applications in English tests, their
advantages and limitations, and guidelines for effective construction and implementation.
Gap-fill items present students with a text where specific words or phrases have been
removed and replaced with blanks. Students must analyze the context and utilize their
language knowledge to determine the missing elements and complete the text coherently
and accurately. This process assesses their ability to:
• Apply grammatical knowledge: Identify the correct tense, form, and agreement
of verbs, nouns, pronouns, adjectives, and adverbs.
• Utilize vocabulary: Select appropriate words that fit the context and convey the
intended meaning.
• Understand sentence structure: Recognize the syntactic roles of words and phrases
within a sentence.
• Comprehend discourse: Maintain coherence and cohesion within the text by
using appropriate linking words and phrases.
Gap-fill items can be used to assess a wide range of language skills and knowledge areas:
• Open-ended: Students are free to choose any word or phrase that fits the context.
• Closed-ended: Students are provided with a word bank or a limited set of options
to choose from.
• Targeted: Gaps are strategically placed to assess specific grammar points or
vocabulary items.
• Sentence completion: Students complete a sentence by filling in a missing word
or phrase.
• Passage completion: Students fill in multiple gaps within a longer text.
• Potential for ambiguity: Some gaps may have multiple possible answers, making
scoring subjective.
• Limited scope: May not fully capture complex aspects of language use, such as
pragmatic understanding or communicative competence.
• Difficulty in constructing effective items: Creating gaps that have a single
correct answer and effectively assess the intended skill can be challenging.
• Choose appropriate texts: Select texts that are relevant to learners' interests and
proficiency levels.
• Determine the type and number of gaps: Consider the difficulty level desired
and the specific skills being assessed.
• Provide clear instructions: Ensure learners understand the task and how to
respond.
• Focus on key language points: Target specific grammar structures or vocabulary
items for assessment.
• Avoid excessive gaps: Too many gaps can disrupt the flow of the text and make
the task overwhelming.
• Thorough review and piloting: Ensure items are accurate, clear, and free of bias
before including them in a test.
Faulty gap-fill items
Fault: This sentence is too open-ended. Students could fill the gap with various
words like "store," "supermarket," "shop," "market," etc., making it difficult to
determine a single correct answer.
Fault: This item requires knowledge of Vietnamese culture that might not be
covered in the English textbook or familiar to all students.
Fault: This sentence provides no context or clues for the missing words. Students
are left to guess randomly, which doesn't effectively assess their language skills.
Fault: Using a highly specific or obscure word like "emblazoned" might not be
appropriate for a high school level gap-fill test, especially if it's not a word
commonly encountered in the textbook.
In general, gap-fill items are a valuable component of English language tests for high
school students in Vietnam. They provide a flexible and focused way to assess various
language skills, particularly grammar, vocabulary, and sentence construction. By
adhering to the principles of effective item construction and considering the potential
limitations, educators can utilize gap-fill activities to create meaningful assessments that
contribute to effective language learning.
Subjective test items, unlike their objective counterparts, require students to construct
their own responses rather than selecting from pre-defined options. These items, such as
essays, short answer questions, and oral interviews, offer a valuable means of assessing
deeper levels of language proficiency and cognitive skills in high school English
language learners in Vietnam. This chapter explores the characteristics, applications,
advantages, and limitations of subjective test items, along with guidelines for their
effective design and implementation.
Subjective test items typically present students with open-ended prompts or questions
that require them to generate unique responses using their own language and ideas. These
items assess a broader range of skills and knowledge, including:
Subjective test items are valuable for assessing various aspects of English language
proficiency:
• Clear and concise prompts: Provide specific and unambiguous instructions that
clearly define the task and expectations.
• Relevant to curriculum: Align prompts with the learning objectives and content
covered in the curriculum.
• Appropriate difficulty level: Ensure tasks are challenging yet attainable for the
students' proficiency level.
• Defined evaluation criteria: Develop clear scoring rubrics that outline the criteria
for evaluating responses and awarding marks.
• Rater training and standardization: Provide training to ensure consistent and
reliable scoring across different raters.
In the realm of English language assessment for high school students in Vietnam, the
choice of scoring methods plays a crucial role in determining the accuracy, fairness, and
effectiveness of evaluations. This section explores the key distinctions between objective
and subjective scoring methods, highlighting their respective advantages, limitations, and
applications in various assessment contexts.
Objective scoring methods are characterized by their clear-cut criteria and standardized
procedures, leaving little room for individual interpretation or bias. These methods are
typically employed for assessing responses to objective test items, such as multiple-
choice, true-false, and matching questions.
• Efficiency: Objective scoring is generally quick and efficient, allowing for rapid
evaluation of large numbers of test papers.
• Reliability: With standardized procedures and answer keys, objective scoring
produces consistent results, minimizing variations between different raters.
• Objectivity: The scoring process is free from personal biases or interpretations,
ensuring fairness and equal treatment for all test-takers.
• Ease of analysis: Objective scores are easily quantifiable, facilitating statistical
analysis and reporting of results.
• Subjectivity and bias: Raters' personal biases and interpretations can influence
scoring, potentially leading to inconsistencies and unfairness.
• Time-consuming: Evaluating open-ended responses requires careful reading and
analysis, making subjective scoring more time-consuming than objective scoring.
• Rater reliability: Ensuring consistency and agreement between different raters
requires thorough training and standardization procedures.
• Difficulty in providing specific feedback: Providing detailed and specific
feedback on subjective responses can be challenging.
• Clarity of Criteria: Establish clear and specific scoring criteria or rubrics that
outline the expectations for each assessment task and the characteristics of
different performance levels.
• Transparency: Communicate the scoring criteria to students beforehand to ensure
they understand the expectations and can focus their efforts accordingly.
• Consistency: Apply the scoring criteria consistently across all student work to
ensure fairness and equity in the evaluation process.
• Objectivity: Minimize personal biases and subjective interpretations when
evaluating student work, particularly for subjective assessments.
• Feedback: Provide constructive feedback to students that highlights their strengths
and areas for improvement, guiding their future learning.
• Accurate answer keys: Develop accurate and unambiguous answer keys for
objective test items, such as multiple-choice, true-false, and matching questions.
• Automated scoring: Utilize technology, where feasible, to automate the scoring
of objective tests, ensuring efficiency and accuracy.
• Item analysis: Analyze student responses to identify problematic items that may
need revision or removal from future assessments.
• Detailed rubrics: Develop detailed scoring rubrics that outline the specific criteria
for evaluating different aspects of performance, such as content, organization,
language use, and mechanics.
• Rater training: Provide thorough training to raters on the scoring rubrics to
ensure inter-rater reliability.
• Multiple raters: When possible, use multiple raters to evaluate subjective
responses, such as essays or oral presentations, and average their scores to
minimize bias.
• Anonymity: Mask student identities during the scoring process to prevent
unconscious biases from influencing evaluations.
• Holistic scoring: Consider the overall quality of the response, rather than focusing
solely on individual errors, to provide a more comprehensive assessment.
• Specificity: Provide specific and detailed feedback that pinpoints areas of strength
and weakness in the student's work.
• Actionable advice: Offer actionable advice on how the student can improve their
performance in the future.
• Timeliness: Provide feedback promptly while the assessment is still fresh in the
student's mind.
• Encouragement: Balance constructive criticism with positive reinforcement to
motivate students and foster a growth mindset.
Implementing best scoring practices is essential for ensuring accurate, fair, and
meaningful assessments of English language proficiency in Vietnam. By adhering to the
principles of clarity, consistency, objectivity, and constructive feedback, educators can
provide valuable evaluations that support student learning and contribute to the overall
effectiveness of the educational system.
In summary, the choice between objective and subjective scoring methods depends on
the specific learning objectives, the type of assessment tasks, and the desired level of
detail in the evaluation. By understanding the strengths and limitations of each approach,
educators can make informed decisions about scoring procedures, ensuring fair, reliable,
and meaningful assessments of English language proficiency for high school students in
Vietnam. Table 5.1 compares objective vs. subjective scoring:
Instruction: Carefully read each statement in the table below. For each statement, rate
your understanding using the following scale:
1: I do not understand this at all.
2: I understand this a little, but I need more help.
3: I understand this fairly well, but I have some questions.
4: I understand this very well and can explain it to others.
Objectives:
• To develop students' understanding of multiple-choice question design principles.
• To enhance students' critical thinking and analytical skills by analyzing a given
text and formulating appropriate test items.
Procedure: Students in groups/pairs read the following passage from Unit 2, the English
textbook Global Success for Grade 11. The students first read the passage carefully and
identify the main ideas, supporting details, any challenging vocabulary or grammatical
structures, and the overall tone and purpose of the passage. The items should include a
variety of question types:
o Factual questions: Test for recall of specific information (e.g., "What is the
main cause of...?").
o Inferential questions: Test for understanding of implied meanings and
making deductions (e.g., "What can be inferred about the author's attitude
towards...?").
o Vocabulary questions: Test for understanding of vocabulary in context.
o Main idea questions: Test for comprehension of the overall theme or central
message.
“Over the past two centuries, different generations were born and given different
names. Each generation comes with its characteristics, which are largely influenced
by the historical, economic, and social conditions of the country they live in.
However, in many countries the following three generations have common
characteristics.
Generation X refers to the generation born between 1965 and 1980. When Gen Xers
grew up, they experienced many social changes and developments in history. As a
result, they are always ready for changes and prepared to work through changes.
Gen Xers are also known as critical thinkers because they achieved higher levels of
education than previous generations.
Generation Y, also known as Millennials, refers to those born between the early
1980s and late 1990s. They are curious and ready to accept changes. If there is a
faster, better way of doing something, Millennials want to try it out. They also value
teamwork. When working in a team, Millennials welcome different points of view
and ideas from others.
Generation Z includes people born between the late 1990s and early 2010s, a time
of great technological developments and changes. That is why Gen Zers are also
called digital natives. They grew up online and never knew the world before digital
and social media. They are very creative and able to experiment with platforms to
suit their needs. Many Gen Zers are also interested in starting their own businesses
and companies. They saw so many people lose their jobs, so they think it is safer to
be your own boss than relying on someone else to hire you.
Soon a new generation, labelled Gen Alpha, will be on the scene. Let's wait and see
if we will notice the generation gap.” (From Unit 2, Global Success, Grade 11,
Hoang et al., 2018)
Cities around the world are becoming smarter, and you can do many things that
seemed impossible in the past.
In Singapore, the mobile app [Link] allows you to locate a nearby car park
easily, book a parking space, and make a payment. You can also extend your
booking or receive a refund if you leave early.
New York City (US) has one of the largest bike-sharing systems called Citi Bike.
Using a mobile app, you can unlock bikes from one station and return them to any
other station in the system, making them ideal for one-way trips.
In Copenhagen (Denmark), you can use a mobile app to guide you through the city
streets and tell how fast you need to pedal to make the next green light. The app can
also give you route recommendations and work out the calories you burn.
In London(UK), you don’t have to buy public transport tickets. You can just touch
your bank card on the card reader when you get on and off the bus or the
underground to pay for your trip.
In Toronto (Canada), you can book an appointment and see a doctor online a from
your own home. You can also receive prescriptions and any other documents you
need, all online.
When reading the passage, students need to highlight key information and identifying
main ideas. They review the characteristics of each type of statement:
Each pair/group should create at least 2 statements for each category (True, False, Not
Given), resulting in a minimum of 6 statements. Students should vary the difficulty level
of their statements and to avoid simply copying sentences directly from the text.
Procedure: Students into small groups search for a set of subjective test items in the
English tests for high school students in Vietnam (e.g. essay prompts, short answer
questions, oral interview questions from previous English tests or textbooks in use). They
discuss the following questions:
Procedure: Students in groups begin by reviewing the key learning objectives and
content areas covered in a unit or several units an English textbook for high school
students in Vietnam, then identify the knowledge and skills that should be assessed in an
achievement test. They then discuss the different types of subjective items commonly
used in achievement tests (essays, short answer questions, etc.). During group discussion,
students
Objective: To enable students to analyze and compare two test matrices (one for Grade 6
and one for Grade 8 in Vietnam as in the images below) and identify the key differences
in terms of skills, knowledge, and complexity.
Procedure: Students in groups spend time reviewing the content and structure of each
matrix. They use highlighters or colored pens to mark any noticeable differences between
the two matrices and then report the findings to the whole class. They should focus on
aspects like:
§ Skills: Are there any skills assessed in one matrix that are not
present in the other? Are the same skills assessed at different
levels of complexity?
§ Knowledge: Are there differences in the types or depth of
knowledge expected at each grade level?
§ Weighting: Are certain skills or knowledge areas given more
emphasis in one matrix compared to the other?
Chapter 6: Innovations in Language Assessment
The landscape of language testing in the world and in Vietnam is rapidly evolving, with
online and computer-based testing (CBT) gaining significant traction. This shift is driven
by technological advancements, increased internet accessibility, and the need for more
efficient and flexible assessment methods.
The COVID-19 pandemic accelerated the adoption of online and CBT in Vietnam, as
educational institutions sought alternative assessment solutions during school closures
(Bui, 2023). This period highlighted the potential of technology to deliver assessments
remotely, cater to diverse learning needs, and provide immediate feedback. Furthermore,
the Vietnamese government has been actively promoting digital transformation in
education, encouraging the integration of technology into teaching and assessment
practices.
• Remote delivery: Online and CBT can be administered remotely, allowing for
greater flexibility and accessibility for test-takers across geographical locations.
• Automated scoring: Many CBT platforms offer automated scoring for objective
test items, increasing efficiency and reducing human error.
• Multimedia integration: Online and CBT can incorporate multimedia elements,
such as audio, video, and interactive simulations, to create more engaging and
authentic assessments.
• Adaptive testing: Some CBT platforms utilize adaptive testing algorithms that
adjust the difficulty level of questions based on the test-taker's performance,
providing a more personalized assessment experience.
• Data-driven insights: Online and CBT generate valuable data on student
performance, enabling educators to track progress, identify learning gaps, and
tailor instruction accordingly.
Advantages of online and CBT
• Increased accessibility: Online and CBT can reach a wider range of test-takers,
including those in remote areas or with disabilities, promoting inclusivity and
equity in assessment.
• Enhanced efficiency: Automated scoring and online test administration streamline
the assessment process, saving time and resources for educators.
• Improved engagement: Multimedia elements and interactive features can enhance
student engagement and motivation during the testing process.
• Personalized learning: Adaptive testing and data-driven insights enable
personalized feedback and targeted interventions to support individual learning
needs.
• Reduced environmental impact: Online and CBT reduce reliance on paper-based
materials, contributing to environmental sustainability.
The adoption of online and CBT has significant implications for language education in
Vietnam:
In summary, online and computer-based language testing are transforming the assessment
landscape in Vietnam, offering numerous advantages in terms of accessibility, efficiency,
and personalization. However, addressing the challenges related to technology, security,
and equity is crucial to ensure the successful and ethical implementation of these
innovative assessment methods. By embracing best practices and investing in teacher
training and infrastructure, Vietnam can harness the power of technology to enhance
language learning and assessment for all students.
• Automated essay scoring (AES): AI algorithms can analyze and score written
responses, providing feedback on grammar, vocabulary, organization, and
coherence (Wilson & Roscoe, 2022).
• Computerized adaptive testing (CAT): AI-powered CAT platforms adjust the
difficulty level of questions based on the test-taker's performance, leading to more
efficient and precise assessments.
• Automated speech recognition (ASR): ASR technology can evaluate spoken
language proficiency, assessing pronunciation, fluency, and accuracy in real-
time.
• Chatbots for language assessment: AI-powered chatbots can engage in
conversations with test-takers, simulating real-life interactions and assessing
communicative competence.
• AI-driven feedback and remediation: AI systems can analyze student performance
and provide personalized feedback and recommendations for improvement.
Benefits of AI in language testing
• Data privacy and security: Protecting student data and ensuring ethical use of AI
are critical considerations.
• Bias and fairness: AI algorithms can reflect biases present in the data they are
trained on, requiring careful development and validation to ensure fairness.
• Validity and reliability: Ensuring the accuracy and consistency of AI-powered
assessments requires rigorous validation and ongoing monitoring.
• Technological infrastructure: Reliable internet access and appropriate devices are
needed for widespread adoption of AI-powered testing.
• Teacher training and acceptance: Educators need training and support to
effectively integrate AI tools into their teaching and assessment practices.
Language testing is a dynamic field, constantly evolving to meet the changing needs of
learners, educators, and society. This chapter explores the emerging trends and future
directions that are shaping the landscape of language testing, with a particular focus on
their potential impact on high school students in Vietnam.
Technology-enhanced assessments
Language tests are evolving to assess not only traditional language skills but also 21st-
century skills, such as:
• Critical thinking and problem-solving: Tasks that require analysis, evaluation, and
creative solutions are becoming more common in language assessments.
• Collaboration and communication: Assessing students' ability to work effectively
in groups, communicate ideas clearly, and engage in collaborative problem-
solving.
• Digital literacy: Evaluating students' ability to navigate digital environments, use
technology effectively for learning, and communicate in online contexts.
• Intercultural competence: Assessing students' ability to understand and interact
effectively with people from different cultural backgrounds.
• Data privacy and security: Protecting student data and ensuring responsible use of
AI and other technologies in assessment.
• Bias and fairness: Addressing potential biases in assessment design and algorithms
to ensure fairness and equity for all test-takers.
• Accessibility and inclusivity: Designing assessments that are accessible to students
with diverse learning needs and disabilities.
• Transparency and accountability: Providing clear information about assessment
practices, scoring methods, and the use of technology.
Instruction: Carefully read each statement in the table below. For each statement, rate
your understanding using the following scale:
Objective: To help students understand and reflect on online English tests by analyzing a
real online test designed for high school students ([Link]
or [Link]
Procedure: Students work in groups/pairs to access the online English test(s) in the links
provided and explore the test interface, navigation, and different question types. They
then answer the following questions:
Procedure: Students in groups obtain a copy of the paper-based test and access to an
online test (see the links in Activity 1). They then analyze and compare the two tests
across various dimensions, such as:
§ Test content: Are the topics and skills assessed similar in both tests?
§ Question types: What types of questions are used in each format
(e.g., multiple-choice, essay, listening comprehension)?
§ Delivery and format: How are the tests administered? How does the
format affect the test-taking experience?
§ Accessibility: Are there differences in how accessible each test is to
different learners?
§ Interaction: How interactive is each test? Are there opportunities for
multimedia or adaptive testing in the online version?
§ Feedback and Scoring: How is feedback provided? How is the test
scored?
Objective: To enable students to understand the principles of online language test design
by creating an outline for an online English test for high school students in Vietnam.
Procedure: Students begin by reviewing the key learning objectives and content areas
covered in an English textbook or an English program. They then discuss the specific
skills and knowledge that should be assessed in an online test. Each group should define
the specific objectives of their online test. What skills and knowledge will it assess? What
level of proficiency will it target? Students consider the following test specifications in
their outline:
Objective: To foster critical thinking and discussion among students about the potential
applications, benefits, and challenges of using AI in English language tests for high
school students in Vietnam.
References
Bui, H. P. (2023). Vietnamese university EFL teachers’ and students’ beliefs and
teachers’ practices regarding classroom assessment. Language Testing in
Asia, 13(1), 10.
Koraishi, O. (2023). Teaching English in the age of AI: Embracing ChatGPT to optimize
EFL materials and assessment. Language Education and Technology, 3(1), 55-72.
Wilson, J., & Roscoe, R. D. (2020). Automated writing evaluation and feedback:
Multiple metrics of efficacy. Journal of Educational Computing Research, 58(1),
87-125.
Chapter 7: Language Assessment Practices at High Schools in Vietnam
Language assessment plays a critical role in the Vietnamese education system, serving
various purposes such as measuring student progress, evaluating teaching effectiveness,
and informing educational policy. This chapter provides an in-depth analysis of language
assessment practices at the high school level in Vietnam, examining the types of
assessments, their purposes, the challenges faced, and the ongoing efforts to improve
assessment quality and fairness.
• Measuring student achievement: Tests are used to assess students' progress and
achievement in English, providing feedback on their strengths and weaknesses.
• Placement and selection: Language tests are used for placement purposes,
assigning students to appropriate English classes based on their proficiency levels.
They are also used for selection, determining eligibility for scholarships, study
abroad programs, and university admission.
• Evaluating teaching effectiveness: Tests can provide insights into the effectiveness
of teaching methodologies and identify areas where improvements are needed.
• Informing educational policy: Large-scale language assessments provide valuable
data for policymakers to evaluate the overall effectiveness of English language
education programs and make informed decisions about curriculum development
and resource allocation.
Classroom-based language tests are an integral part of the English language learning
experience for high school students in Vietnam. These assessments, designed and
administered by teachers, play a crucial role in monitoring student progress, providing
feedback, and guiding instructional decisions. This chapter explores the key features,
types, benefits, challenges, and best practices associated with classroom-based language
tests in the Vietnamese high school context.
In Vietnam, where high-stakes national exams often dominate the educational landscape,
classroom-based assessments provide a valuable counterbalance. They offer a more
nuanced and personalized view of student learning, allowing teachers to:
• Monitor progress: Track individual student progress and identify areas of strength
and weakness.
• Provide feedback: Offer timely and specific feedback to students on their language
development.
• Inform instruction: Adjust teaching strategies and activities based on student
performance and needs.
• Promote learning: Encourage student engagement, motivation, and self-reflection.
• Create a positive learning environment: Foster a supportive and growth-oriented
classroom culture where assessment is seen as a tool for learning.
• Teacher training and expertise: Some teachers may lack adequate training in
language assessment principles and practices, which can affect the quality and
fairness of their assessments.
• Time constraints: Designing, administering, and grading assessments can be time-
consuming for teachers, especially in large classes.
• Subjectivity in scoring: Subjective assessments, such as essays and oral
presentations, can be challenging to score consistently and fairly.
• Balancing formative and summative assessment: Finding the right balance
between formative and summative assessments can be challenging, as over-
reliance on summative assessments can create pressure and neglect ongoing
learning.
Classroom-based language tests are essential for effective English language instruction in
Vietnamese high schools. By implementing best practices, providing ongoing teacher
training, and embracing innovative assessment approaches, educators can create a more
engaging, supportive, and learner-centered assessment environment that promotes
language development and academic success for all students.
High-stakes English exams play a crucial role in Vietnam's education system and society.
They are called “high-stakes” because the results from these tests result in significant
consequences for the test taker related to academic decisions, professional
opportunities, and immigration or citizenship. In other words, high-stakes English
exams are significant for the following reasons:
• Pressure and anxiety: The intense competition for university places and good jobs
creates immense pressure on students to perform well in these exams. This can
lead to test anxiety and mental health concerns.
• Access to resources: While English language learning resources are becoming
more widely available, there's still a disparity between urban and rural areas in
terms of access to quality instruction and preparation materials.
• Emphasis on standardized testing: There's ongoing debate about the over-reliance
on standardized tests as the sole measure of English proficiency, with some
advocating for more holistic assessment methods that consider real-world
communication skills.
Recent developments:
Despite ongoing efforts to improve language testing practices, several challenges persist:
Language testing practices in Vietnamese high schools are dynamic and evolving,
reflecting the growing importance of English language proficiency in the country. By
addressing the challenges, embracing innovations, and focusing on assessment for
learning, Vietnam can create a more effective and equitable language assessment system
that supports student learning and prepares them for success in the 21st century.
Backwash refers to the influence that tests have on teaching and learning practices. When
tests are perceived as high-stakes, they can have a significant impact on what is taught,
how it is taught, and how students learn. Positive washback occurs when tests promote
beneficial teaching practices and learning outcomes, while negative washback can lead to
teaching to the test, where instruction becomes narrowly focused on test content at the
expense of broader language skills (Fulcher & Davidson, 2007).
• Narrowing the curriculum: Teachers may focus on teaching only what is tested,
neglecting other important aspects of language learning, such as creativity and
critical thinking.
• Teaching to the test: Teachers may prioritize test preparation over meaningful
language learning, leading to rote learning and a focus on test-taking strategies.
• Test anxiety: High-stakes tests can cause anxiety and stress, which can negatively
impact learners' performance.
• Inequitable assessment: Tests may not accurately reflect the diverse range of
language abilities, particularly for learners from marginalized backgrounds.
Mitigating negative backwash effects
• Test validity and reliability: Ensure that tests are valid and reliable measures of
language proficiency.
• Balanced assessment: Use a variety of assessment methods, including formative
and summative assessments, to avoid over-reliance on high-stakes tests.
• Teacher training: Provide teachers with training on effective test preparation
strategies and how to avoid negative backwash effects.
• Learner awareness: Educate learners about the purpose of tests and how to
approach them effectively.
• Authentic tasks: Incorporate authentic language tasks into teaching and assessment
to promote real-world language use.
By understanding the potential positive and negative impacts of tests, language educators
can make informed decisions about assessment practices and strive to create a positive
learning environment that fosters language development and critical thinking.
Instruction: Carefully read each statement in the table below. For each statement, rate
your understanding using the following scale:
Procedure: Students begin by reviewing the concept of formative assessment with the
students:
Students then brainstorm a variety of formative assessment activities that could be used
in an English language classrooms, considering activities like:
Objectives:
Procedure: Students work in groups/pairs and examine the following an English test for
high school students in Vietnam in 2024. They need to answer the following guiding
questions:
What test items are included?
What do they aim to measure?
Which parts reflects the English teaching curriculum the most? Why?
Which parts do you suggest to be added? Why?
Activity 4: Comparing English tests for high school students in two different years
Procedure: Students in groups examine the high-stake English test for 2025 below and
compare with the one in 2024 as in Activity 1. They need to point out the similarities and
differences and explain for the changes.
Activity 5: VSTEP debate showdown
Objectives:
• To encourage students to critically analyze the VSTEP exam format and its
components.
• To enhance students' understanding of the skills and strategies needed to succeed
in the VSTEP exam.
References
Bui, H. P., & Nguyen, T. T. T. (2024). Classroom assessment and learning motivation:
insights from secondary school EFL classrooms. International Review of Applied
Linguistics in Language Teaching, 62(2), 275-300.
Ngo, X. M. (2024). English assessment in Vietnam: status quo, major tensions, and
underlying ideological conflicts. Asian Englishes, 26(1), 280-292.
Nguyen, T. L., & Nguyen, T. N. (2020). The role of learners’test perception in changing
English learning practices: A case of a high-stakes English test at Vietnam national
university, Hanoi. VNU Journal of Foreign Studies, 35(6), 2525-2445.
Chapter 8: Interpreting and Reporting Test Results
In the realm of language assessment, a test score is more than just a numerical value; it is
a window into a learner's linguistic journey, a snapshot of their capabilities at a specific
point in time. This chapter, "Interpreting and Reporting Test Results," serves as a guide
to deciphering the intricate language of English test scores. This chapter delves into the
nuances of various score types, from raw scores and percentiles to proficiency levels. It is
also important to contextualize these scores (Levi & Inbar-Lourie, 2020), considering
factors like test purpose, learner characteristics, and test reliability to paint a
comprehensive picture of individual achievement.
An English test score is a coded message containing valuable insights into a learner's
language proficiency. This section presents the tools to crack that code, enabling you to
interpret English test results accurately and report them effectively.
English tests employ various scoring methods, each with its own interpretation:
• Raw scores: These represent the total number of correct answers. While seemingly
straightforward, raw scores lack context. They don't indicate how a student
compares to others or against a specific standard.
• Scaled scores: These convert raw scores onto a standardized scale, allowing for
comparison across different test versions or administrations. They often have a
predefined mean and standard deviation.
• Percentile ranks: These indicate the percentage of test-takers who scored lower
than a particular individual. For example, a percentile rank of 70 means the
student performed better than 70% of those who took the test.
• Proficiency levels: Many tests, like the CEFR (Common European Framework of
Reference for Languages), categorize scores into proficiency levels (e.g., A1, B2,
C1). These describe the learner's abilities in real-world contexts.
Interpreting test results requires going beyond the numbers and considering the context:
• Purpose of the test: What was the test designed to measure? Was it for university
admission, job applications, or general proficiency assessment? The purpose
influences how scores should be interpreted.
• Test-taker characteristics: Factors like age, educational background, first language,
and learning experiences can all affect test performance.
• Test reliability and validity: How reliable and valid is the test itself? A reliable test
produces consistent results, while a valid test measures what it intends to
measure.
• Individual strengths and weaknesses: Analyze performance across different test
sections (e.g., reading, writing, listening, speaking) to identify areas of strength
and areas needing improvement.
Clear and informative reporting is crucial for conveying test results to stakeholders:
• Target audience: Tailor the report to the audience (e.g., students, parents, teachers,
administrators). Use language and visuals that are appropriate for their
understanding.
• Key information: Include essential details like the test name, date of
administration, and the student's score(s).
• Interpretation: Explain what the scores mean in plain language, avoiding technical
jargon. Relate the scores to the test's purpose and the student's individual context.
• Recommendations: Provide specific and actionable recommendations for
improvement. This might include targeted language learning activities, further
testing, or support services.
• Ethical considerations: Maintain confidentiality and ensure that test results are
used fairly and responsibly.
While test scores provide valuable information, they do not capture the whole picture of a
learner's language proficiency (Shohamy, 2020). Consider these additional factors:
• Classroom performance: Observe the student's active language use in class, their
participation in discussions, and their completion of assignments.
• Portfolio assessment: Collect samples of the student's work over time to track their
progress and showcase their achievements.
• Self-assessment: Encourage students to reflect on their own language learning
journey and set goals for improvement.
English test results are not merely an endpoint, a final grade to be recorded and filed
away. Instead, they should be viewed as a starting point, a springboard for informed
educational decision-making that benefits both students and educators (Coombe, et al.,
2020). This chapter explores how to effectively use test outcomes to drive improvements
in English language teaching and learning.
Test results can illuminate individual strengths and weaknesses in language skills. This
knowledge allows educators to:
• Tailor instruction: Adapt teaching methods and materials to address specific needs
identified in the test results. For example, a student struggling with reading
comprehension might benefit from targeted exercises on vocabulary building and
inferencing.
• Provide differentiated support: Offer individualized support and resources, such as
extra practice materials, tutoring sessions, or online learning tools, to help students
overcome their challenges.
• Set personalized goals: Work with students to set realistic and achievable goals
based on their test performance and learning aspirations.
Classroom-level interventions
Analyzing test results at the classroom level can reveal patterns and trends that inform
instructional adjustments (Ismail, et al., 2022). This might involve:
Aggregating and analyzing test data across different classes and grade levels can provide
valuable insights for school-wide improvement efforts. This can lead to:
• Identifying areas of need: Pinpointing areas where students across the school
require additional support, such as vocabulary development, grammar accuracy, or
oral fluency.
• Allocating resources effectively: Directing resources and professional
development opportunities towards addressing the identified needs. This might
involve investing in new learning materials, providing teacher training on specific
teaching strategies, or establishing peer tutoring programs.
• Monitoring progress over time: Tracking student performance on English tests
over time can help assess the effectiveness of school-wide interventions and
inform continuous improvement efforts.
On a broader scale, large-scale English test results can inform educational policy
decisions at the district, regional, or national level (Todd, et al., 2021). This data can be
used to:
• Set benchmarks and standards: Establish clear benchmarks and standards for
English language proficiency at different educational levels.
• Evaluate program effectiveness: Assess the effectiveness of language teaching
programs and initiatives.
• Allocate funding and resources: Guide the allocation of funding and resources to
support English language education.
• Promote educational equity: Identify and address disparities in English language
proficiency among different student populations, ensuring that all students have
access to quality language education.
Ethical considerations
• Avoid high-stakes decisions based solely on test scores. Consider other factors,
such as classroom performance, student effort, and individual learning styles.
• Protect student privacy and confidentiality. Handle test data responsibly and
ensure that it is used ethically and in accordance with relevant regulations.
• Communicate results clearly and transparently. Provide students, parents, and
other stakeholders with clear and understandable explanations of test results and
their implications.
• Use test data to support, not label, students. Focus on using test outcomes to
identify areas for improvement and provide targeted support, rather than labeling
students based on their scores.
Instruction: Carefully read each statement in the table below. For each statement, rate
your understanding using the following scale:
Objectives:
• To help students understand the difference between raw scores and percentile
scores in the context of English language tests.
• To develop their ability to interpret and compare these two types of scores.
Procedure: Students begin by reviewing the concept of raw scores or the number of
correct answers on a test. They then discuss the percentile scores indicating the
percentage of test-takers who scored lower than a particular individual. Students then
discuss in pairs:
o Scenario 1: Student A got a raw score of 30. Student B got a raw score of
40. Who performed better?
o Scenario 2: Student C got a percentile score of 75. Student D got a
percentile score of 90. Who performed better?
o Scenario 3: Student E got a raw score of 35, which is equivalent to the 50th
percentile. What does this mean?
o Scenario 4: Student F got a raw score of 50, which is equivalent to the 99th
percentile. What does this mean?
Activity 2: Decoding the data: Analyzing high school English test scores
Objectives:
(Source: [Link]
[Link])
Objectives:
o Identify the key skills and abilities associated with that level in each
domain (listening, speaking, reading, writing).
o Paraphrase the descriptors in simpler terms, making them easier to
understand.
Objectives:
• To familiarize students with the content of the National English Test score report
in Vietnam.
• To develop their ability to interpret the different sections of the report and
understand their implications.
“Môn Tiếng Anh có 906.549 thí sinh dự thi, ít hơn 139,094 bài thi so với môn
Toán, có thể do một số thí sinh dùng môn ngoại ngữ khác (Pháp, Trung, Nga,
Hàn,…) hoặc dùng kết quả chứng chỉ IELTS, TOEIC, TOEFL,… thay thế bài thi.
Có tới 145 bài thi bị điểm liệt, chiếm tỉ lệ 0,016%, là giảm so với năm 2023
(0,022%)14 bài thi 0 điểm. Điểm của môn Tiếng Anh nhìn chung gần giống năm
ngoái, điểm trung bình là 5,51 và điểm trung vị là 5,2. Số điểm dưới 5 là 386.861
thí sinh, chiếm 42,674% và điểm học sinh đạt nhiều nhất là 4,6. Năm nay ghi nhận
565 thí sinh đạt điểm 10. Tiếng Anh vẫn là môn thi có kết quả thấp nhất trong các
môn từ trước tới nay.
Tổ hợp thí sinh thi các môn Khoa học tự nhiên so với tổ hợp thi Khoa học xã hội
năm nay là 52,77%, có cao hơn tỉ lệ chọn năm 2023 là 50,75%. Kết quả có thể
nhận định về xu hướng chọn ngành nghề học đại học tuy vẫn nghiêng mạnh về các
ngành xét tuyển Khoa học xã hội, nhưng đã có sự dịch chuyển sang khối Khoa học
tự nhiên.” ([Link]
can-cu-quan-trong-lua-chon-dang-ky-nguyen-vong-chuan-
[Link])
• What is the overall English score of high school students in the report?
• What are some implications of the report for the teaching and learning English at
high schools in Vietnam?
Activity 5: Test outcomes in action: A simulation game
Objectives:
• To help students understand how English test outcomes can be used for
educational decision-making in a realistic scenario.
• To develop their ability to analyze test data, identify student needs, and propose
appropriate interventions.
Procedure:
Set the scene: Students begin by describing the educational context for the activity.
For example, they might present a scenario of a high school English teacher who has
just received the results of a recent English test for their class.
Introduce the students: Provide students with a set of fictional student profiles, each
with their own unique background and English test scores. These profiles should
include information about their strengths, weaknesses, learning preferences, and
goals.
Analyze the data: Divide the class into small groups and assign each group a subset
of the student profiles. Have them analyze the test data and answer guiding questions
such as:
References
Coombe, C., Vafadar, H., & Mohebbi, H. (2020). Language assessment literacy: What do
we need to learn, unlearn, and relearn?. Language Testing in Asia, 10(1), 3.
Ismail, S. M., Rahul, D. R., Patra, I., & Rezvani, E. (2022). Formative vs. summative
assessment: impacts on academic motivation, attitude toward learning, test anxiety,
and self-regulation skill. Language Testing in Asia, 12(1), 40.
Levi, T., & Inbar-Lourie, O. (2020). Assessment literacy or language assessment literacy:
Learning from the teachers. Language Assessment Quarterly, 17(2), 168-182.
Shohamy, E. (2020). The power of tests: A critical perspective on the uses of language
tests. Routledge.
Todd, R. W., Pansa, D., Jaturapitakkul, N., Chanchula, N., Pojanapunya, P.,
Tepsuriwong, S., ... & Trakulkasemsuk, W. (2021). Assessment in Thai ELT: What
Do Teachers Do, Why, and How Can Practices Be Improved?. LEARN Journal:
Language Education and Acquisition Research Network, 14(2), 627-649.