0% found this document useful (0 votes)
3 views139 pages

Method 3 Notes

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views139 pages

Method 3 Notes

Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Chapter 1: Introduction to Language Testing and Assessment

Language testing and assessment play crucial roles in both language teaching and
learning. They provide valuable information about learners' proficiency, inform
instructional decisions, and contribute to overall program evaluation. This chapter
explores the fundamental differences of the two concepts of language testing and
assessment, examining their purposes, and impacts on language education.

1.1 Language testing

Language testing refers to the systematic evaluation of an individual’s language abilities,


encompassing skills such as reading, writing, listening, and speaking. It is a process that
involves the development and administration of tasks designed to measure language
proficiency and competence (Shohamy, 2020). The primary focus of language testing is
to assess how well individuals can use a language for communication and to determine
their understanding of linguistic structures (Brown, 2005; Fulcher & Davidson, 2007).

Objectives of language testing

The objectives of language testing can be multifaceted, depending on the context and
purpose of the assessment. However, some common objectives include:

• Measurement of proficiency: One of the primary goals of language testing is to


measure a learner’s language proficiency accurately. This involves assessing their
ability to utilize language skills effectively in real-world contexts (Council of
Europe, 2001).
• Identifying learner needs: Language tests can help identify particular learner
needs, such as specific areas of language difficulty that require additional focus
and instruction. Understanding these needs allows educators to create targeted
learning interventions (Brown, 2005).
• Evaluating instructional effectiveness: By linking assessment outcomes with
educational practices, language testing can help evaluate the effectiveness of
instructional programs, curricula, and teaching strategies. It serves as a feedback
mechanism for educators to refine their approaches to teaching (Fulcher &
Davidson, 2007).
• Facilitating decision-making: Language tests inform decision-makers at various
levels, from classroom teachers to institutional leaders. The data derived from
these assessments support decisions regarding curriculum development,
instructional resource allocation, and program enhancements (Weir, 2005).
• Promoting accountability: Standardized language testing holds schools and
educators accountable for student learning outcomes. By evaluating student
performance against established standards, stakeholders can ensure that
educational institutions are effectively teaching the language and fostering
language proficiency (Hughes, 2003).
• Enhancing learner motivation: Tests, when used appropriately, can actually
enhance learner motivation. Tests can help learners set clear learning goals and
track their progress. Successfully completing a test can provide a sense of
accomplishment and boost motivation. Preparing for a test can encourage focused
study and active learning. Test results can help learners identify areas where they
excel and areas that need improvement. By understanding their strengths and
weaknesses, learners can set realistic goals for future learning.

Importance of language testing

Language testing plays a crucial role in educational settings for several reasons:

• Guiding instruction: Language tests provide valuable information about student


performance, strengths, and weaknesses. Educators use this information to tailor
instruction, identify areas for improvement, and adapt curricular materials to better
meet the needs of learners (Brown, 2005).
• Determining proficiency: Assessments offer a standardized means of measuring
language proficiency, enabling educators, administrators, and policymakers to
make informed decisions about student progress, curriculum effectiveness, and
instructional strategies (Council of Europe, 2001).
• Facilitating placement: Many educational institutions use language tests to place
students in appropriate courses. For instance, a student’s score on a language
proficiency exam can determine whether they are placed in beginner, intermediate,
or advanced classes (Hughes, 2003).
• Monitoring progress: Regular language assessments track individual and group
progress over time. By comparing results from different testing periods, educators
can evaluate the effectiveness of their teaching methods and adjust their
approaches as necessary (Weir, 2005).
• Supporting certification and accreditation: Language testing is often a requirement
for graduation, certification, or admission to advanced education programs.
Proficiency in a language can be crucial for students hoping to pursue further
studies or employment opportunities in multilingual environments (Bachman &
Palmer, 2010).

1.2 Language assessment

Language assessment plays a critical role in the language learning process. It provides
valuable information about learners' progress, informs instruction, and helps to ensure
that teaching and learning are aligned with desired outcomes. Language assessment
includes both formal and informal methods.

Types of language assessment

Language assessment encompasses a wide range of methods, including:

• Formal assessments:
o Standardized tests: These are typically high-stakes tests with standardized
procedures and scoring. Examples include TOEFL, IELTS, and Cambridge English
exams.
o Teacher-made tests: These are tests designed by teachers to assess specific learning
objectives and monitor student progress. They may include multiple-choice questions,
short answer questions, writing tasks, and oral presentations.
• Informal assessments:
o Observations: Observing learners' language use in natural settings, such as
classroom interactions or group work.
o Portfolios: Collections of student work that demonstrate their language
development over time.
o Interviews: One-on-one conversations with learners to assess their speaking and
listening skills.
o Self-assessment: Techniques that involve learners in reflecting on their own
language learning and progress.
o Peer assessment: Learners provide feedback on each other's work.

1.3 Language testing vs. assessment

While the terms "testing" and "assessment" are often used interchangeably, there are
important distinctions between them in the context of language learning. Testing
generally refers to the use of formal procedures to measure a specific aspect of a learner's
language ability. It typically involves the administration of a standardized test with a
predetermined set of questions or tasks, followed by a scoring process that yields a
quantitative result. Assessment, on the other hand, is a broader term that encompasses a
wider range of procedures for gathering information about learners' language proficiency.
It involves a more holistic approach that considers a variety of evidence, including
observations, interviews, portfolios, and informal assessments. Table 1.1 below presents
the differences between language testing and assessment:
Table 1.1 Language testing vs. assessment

Feature Testing Assessment


On a broader range of learning
Primarily on measurement and
outcomes, including communication
quantification of specific skills
Focus skills, language use in real-world
(e.g., grammar, vocabulary, reading
contexts, and overall language
comprehension).
development.
More flexible and varied, including
Standardized procedures, often
observations, interviews, portfolios,
Procedures with limited flexibility in
classroom activities, and informal
administration.
checks.
Qualitative and quantitative data, such
Quantitative data, such as scores,
Outcomes as descriptive feedback, observations,
grades, or percentiles.
and scores.
To provide information for teaching and
Often used for high-stakes
learning, to monitor progress, and to
Purposes decisions, such as admissions,
make informed decisions about
certification, or placement.
instruction.

In summary, while testing provides valuable information about specific language skills,
assessment offers a more comprehensive and holistic view of learners' language
proficiency. Both testing and assessment play important roles in language learning, and a
balanced approach that incorporates both formal and informal assessment methods can
provide valuable insights into learners' progress and inform effective instruction. The
following digagram illustrates the interrelationship of language teaching, assessment and
testing.

Figure 1.1 Language teaching, assessment and testing (Sagagih, 2017)

1.4 Impact of testing on langauge teaching and learning: Backwash effects

Backwash refers to the influence that tests have on teaching and learning practices. When
tests are perceived as high-stakes, they can have a significant impact on what is taught,
how it is taught, and how students learn. Positive washback occurs when tests promote
beneficial teaching practices and learning outcomes, while negative washback can lead to
teaching to the test, where instruction becomes narrowly focused on test content at the
expense of broader language skills (Fulcher & Davidson, 2007).

Positive backwash effects:

• Focus on specific skills: Tests can encourage teachers to focus on specific


language skills, such as reading, writing, listening, and speaking.
• Motivation: Well-designed tests can motivate learners to study harder and improve
their language skills.
• Standardization: Standardized tests can promote consistency in language teaching
and learning across different institutions.

Negative backwash effects:

• Narrowing the curriculum: Teachers may focus on teaching only what is tested,
neglecting other important aspects of language learning, such as creativity and
critical thinking.
• Teaching to the test: Teachers may prioritize test preparation over meaningful
language learning, leading to rote learning and a focus on test-taking strategies.
• Test anxiety: High-stakes tests can cause anxiety and stress, which can negatively
impact learners' performance.
• Inequitable assessment: Tests may not accurately reflect the diverse range of
language abilities, particularly for learners from marginalized backgrounds.

Mitigating negative backwash effects:

• Test validity and reliability: Ensure that tests are valid and reliable measures of
language proficiency.
• Balanced assessment: Use a variety of assessment methods, including formative
and summative assessments, to avoid over-reliance on high-stakes tests.
• Teacher training: Provide teachers with training on effective test preparation
strategies and how to avoid negative backwash effects.
• Learner awareness: Educate learners about the purpose of tests and how to
approach them effectively.
• Authentic tasks: Incorporate authentic language tasks into teaching and assessment
to promote real-world language use.

By understanding the potential positive and negative impacts of tests, language educators
can make informed decisions about assessment practices and strive to create a positive
learning environment that fosters language development and critical thinking.
1.4 Consolidation activities

Activity 1: Understanding the requirements of language testing and assessment for


high schools in Vietnam

Objective: To have students understand what is prescribed in educational documents of


Vietnam about language assessment.

Procedure: In pairs, students read the following description in Circular 32 about


implementing high school education in Vietnam by the Ministry of Education and
Training (MOET, 2018, pp. 32-33) and discuss how English language education in
Vietnam should be assessed based on the objectives of the curriculum.

“Căn cứ đánh giá là các yêu cầu cần đạt về phẩm chất và năng lực được quy định trong
chương trình tổng thể và các chương trình môn học, hoạt động giáo dục. Phạm vi đánh
giá bao gồm các môn học và hoạt động giáo dục bắt buộc, môn học và chuyên đề học tập
lựa chọn và môn học tự chọn. Đối tượng đánh giá là sản phẩm và quá trình học tập, rèn
luyện của học sinh.

Kết quả giáo dục được đánh giá bằng các hình thức định tính và định lượng thông qua
đánh giá thường xuyên, định kì ở cơ sở giáo dục, các kì đánh giá trên diện rộng ở cấp
quốc gia, cấp địa phương và các kì đánh giá quốc tế. Cùng với kết quả các môn học và
hoạt động giáo dục bắt buộc, các môn học và chuyên đề học tập lựa chọn, kết quả các
môn học tự chọn được sử dụng cho đánh giá kết quả học tập chung của học sinh trong
từng năm học và trong cả quá trình học tập.

Việc đánh giá thường xuyên do giáo viên phụ trách môn học tổ chức, kết hợp đánh giá
của giáo viên, của cha mẹ học sinh, của bản thân học sinh được đánh giá và của các học
sinh khác.

Việc đánh giá định kì do cơ sở giáo dục tổ chức để phục vụ công tác quản lí các hoạt
động dạy học, bảo đảm chất lượng ở cơ sở giáo dục và phục vụ phát triển chương trình.

Việc đánh giá trên diện rộng ở cấp quốc gia, cấp địa phương do tổ chức khảo thí cấp quốc
gia hoặc cấp tỉnh, thành phố trực thuộc trung ương tổ chức để phục vụ công tác quản lí
các hoạt động dạy học, bảo đảm chất lượng đánh giá kết quả giáo dục ở cơ sở giáo dục,
phục vụ phát triển chương trình và nâng cao chất lượng giáo dục.

Phương thức đánh giá bảo đảm độ tin cậy, khách quan, phù hợp với từng lứa tuổi, từng
cấp học, không gây áp lực lên học sinh, hạn chế tốn kém cho ngân sách nhà nước, gia
đình học sinh và xã hội.
Nghiên cứu từng bước áp dụng các thành tựu của khoa học đo lường, đánh giá trong giáo
dục và kinh nghiệm quốc tế vào việc nâng cao chất lượng đánh giá kết quả giáo dục, xếp
loại học sinh”

Activity 2: Testing vs. assessment: Sorting and comparing Objectives:

• To enhance students' understanding of the key differences between language


testing and assessment.
• To enable students to identify and categorize various assessment methods
according to whether they align more closely with testing or assessment.

Procedure: Students work in pairs to sort the following key phrases into a T-table and
then explain why you decide whether the phrases align more closely with testing or
assessment.

o Multiple-choice test
o Portfolio assessment
o Oral interview
o Writing sample
o Observation of classroom interaction
o Standardized proficiency test
o Peer feedback
o Self-assessment
o Teacher-made quizzes
o Project presentation

Activity 3: Identifying possible washback effects Objectives:

• To enhance students' understanding of the overall structrue of the national English


exam
• To enable students to identify washback effects of the national exam on English
teaching and learning

Procedure: Students in groups examine the prescribed structure of the national English
exam in the link provided which specifies six types of test items. They then discuss the
washback effects of this exam.

Cấu trúc đề thi Tiếng Anh THPT Quốc gia từ năm 2025 gồm có 6 dạng bài chính
([Link]
[Link]). Cấu trúc bao gồm:

Ngữ âm (Phát âm - Trọng âm)


Điền từ để hoàn thành câu
Điền từ để hoàn thành phiếu thông tin Điền từ để hoàn thành đoạn văn
Sắp xếp câu đúng thứ tự

Activity 4: Unpacking formal language assessment: A closer look Objectives:

• To enhance students' understanding of the characteristics, purposes, and


implications of formal language assessments.
• To enable students to identify and analyze examples of different types of formal
language tests.

Procedure:

Students in groups search on the Internet one standardized proficiency test such as
like VSTEP, TOEIC, IELTS, Cambridge English exams; a high-stake test; a
teacher-made test. They then discuss the following quesitons:

• What are the key features of this type of test?


• What language skills does it assess?
• What are the potential advantages and disadvantages of this type of test?
• What are the ethical considerations associated with this type of test?

Activity 5: Informal assessment: Beyond the test Objectives:

• To enhance students' understanding of the characteristics, purposes, and


applications of informal language assessment methods.
• To enable students to identify and analyze examples of informal language
assessment in real-world contexts.

Procedure: Students work in pairs to review various informal assessment methods (e.g.,
observations, interviews, student journals, peer feedback, self-assessment) and discuss
the purposes of informal assessment and answer the following questions:

• What are the strengths and weaknesses of this method?


• In what situations would this method be most appropriate?
• How can the data collected through this method be used to inform instruction?
• What ethical considerations need to be taken into account when using this
method?
References

Bachman, L. F., & Palmer, A. S. (2010). Language assessment in practice: Developing


language assessments and justifying their use in the classroom. Oxford University Press.

Brown, J. D. (2005). Language assessment: Principles and classroom practices. Pearson


Education.

Council of Europe (2001). Common European framework of reference for languages:


learning, teaching, assessment. Cambridge University Press.

Fulcher, G., & Davidson, F. (2007). Language testing and assessment: an advanced
resource book. Routledge.

Hughes, A. (2003). Testing for language teachers. 2nd ed. Cambridge University Press.

MOET (2018). Thông tư ban hành giáo dục phổ thông. Available at:

[Link] giao-
[Link]

Saragih, F. H. (2016). Testing and assessment in English language instruction. Jurnal


Bahas Unimed, 27(1), 74656.

Shohamy, E. (2020). The power of tests: A critical perspective on the uses of language
tests. Routledge.

Weir, C. J. (2005). Language testing and validation: An evidence-based approach.


Palgrave Macmillan.
Activity 1: Assessment Requirements for Upper Secondary Schools in Vietnam

Based on Circular 32 from the Ministry of Education and Training (MOET, 2018, pp. 32-33),
language assessment for upper secondary schools in Vietnam is guided by the following
principles:
1. Assessment Criteria:
 Evaluation is based on the required learning outcomes in the national curriculum, including
subject knowledge, skills, and competencies.
 Assessment covers compulsory subjects, elective subjects, and extracurricular educational
activities.
2. Scope of Assessment:
Includes both the learning process and learning products (i.e., students' assignments, projects,
and test results).
3. Assessment Methods:
 Regular assessment: Conducted by subject teachers using a mix of teacher assessments, peer
assessments, self-assessments, and parental feedback.
 Periodic assessment: Conducted by schools to manage teaching quality and curriculum
implementation.
 Large-scale assessment: Conducted at national and provincial levels to evaluate overall
educational effectiveness and inform policy development.
4. Forms of Assessment:
 Qualitative and Quantitative Methods: Assessments include both qualitative (e.g., descriptive
feedback, portfolios) and quantitative methods (e.g., standardized tests).
 Combination of Formative and Summative Assessments: Encourages a balance between
continuous assessment (formative) and final evaluations (summative).
5. Reliability and Objectivity:
 Assessment methods should ensure fairness, reliability, and validity.
 They must be appropriate for each grade level and not create undue pressure on students.
6. Alignment with International Standards:
Vietnam aims to gradually incorporate international assessment standards and best practices to
improve educational quality.
7. Financial Considerations:
 Assessment implementation should minimize costs for the state, families, and society.
 This assessment framework aims to ensure that English language education in Vietnam aligns
with national curriculum goals while integrating modern evaluation approaches.
Activity 3: Possible Washback Effects of Assessment Forms

The possible washback effects (both positive and negative) for different forms of assessment in
the national English exam:

Assessment Form Positive Washback Effects Negative Washback Effects


Phonetics (Pronunciation & Encourages better May lead to excessive focus on
Stress) pronunciation practice and memorization rather than real
listening skills. communication.
Fill-in-the-Blank (Sentence Improves grammar and May encourage rote
Completion) vocabulary knowledge. memorization of rules rather than
active language use.
Fill-in-the-Blank Develops comprehension May limit critical thinking if
(Information Completion) and attention to detail in students focus only on textbook
reading. examples.
Fill-in-the-Blank (Passage Encourages understanding Students may focus on guessing
Completion) of text coherence and words instead of overall
context clues. comprehension.
Sentence Ordering Improves syntax and Can lead to overemphasis on test-
(Rearranging Words to sentence structure taking strategies rather than real
Form Correct Sentences) awareness. writing practice.

General Washback Effects


Positive Effects
 Promotes structured learning of essential language skills (grammar, pronunciation, and
reading).
 Helps standardize English proficiency evaluation across schools.
 Encourages teachers to cover essential grammar and vocabulary systematically.
Negative Effects
 May lead to teaching to the test, where teachers focus only on test formats instead of
comprehensive English skills.
 Creates test anxiety among students, reducing confidence in communication skills.
 Overemphasis on accuracy may limit creativity and real-life application of English.
By understanding these effects, educators can balance assessment methods to promote
meaningful language learning while minimizing negative impacts.
Activity 4: Unpacking Formal Language Assessment

Key Answers to Discussion Questions:


1. What are the key features of this type of test?
- Standardized Proficiency Tests (e.g., IELTS, TOEFL, VSTEP, Cambridge English Exams)
 Fixed format and scoring criteria.
 Assesses multiple language skills (Listening, Reading, Writing, Speaking).
 Used for academic, professional, or immigration purposes.
- High-Stakes Tests (e.g., National English Exam for High School Graduation)
 Determines major academic or career decisions.
 Usually has strict grading and validity standards.
 Can create pressure on students and teachers.
- Teacher-Made Tests (e.g., quizzes, midterms, final exams in schools)
 Designed for a specific class curriculum.
 Flexible in format and difficulty level.
 Assesses knowledge retention and application.
2. What language skills does it assess?
- Standardized Proficiency Tests:
 Listening (comprehension of spoken English).
 Reading (understanding of texts and vocabulary).
 Writing (structured essays, reports, or letters).
 Speaking (oral fluency and coherence in conversation).
- High-Stakes Tests:
Usually reading and writing-focused, with limited speaking and listening components.
- Teacher-Made Tests:
Often focus on grammar, vocabulary, and reading but can include speaking and writing
components.

3. What are the potential advantages and disadvantages of this type of test?

Type of Test Advantages Disadvantages


Standardized Reliable, internationally Expensive, stressful, and may not
Proficiency Tests recognized, ensures fairness. reflect real-world language use.
High-Stakes Tests Helps standardize educational Can create test anxiety, may
outcomes, motivates learning. encourage "teaching to the test."
Teacher-Made Tests Customizable, aligned with class May lack standardization and
objectives, low cost. reliability across different classes.

4. What are the ethical considerations associated with this type of test?
- Fairness: Tests should be unbiased and accessible to students from different backgrounds.
- Validity & Reliability: Tests should measure what they are intended to assess consistently.
- Stress & Anxiety: High-stakes testing should avoid unnecessary pressure on students.
- Cost & Accessibility: Standardized tests should not disadvantage students who cannot afford
them.
Summary:
 Standardized tests are useful for broad evaluation but can be stressful and costly.
 High-stakes exams influence curriculum design but may lead to excessive test
preparation.
 Teacher-made tests are flexible but may lack reliability across different contexts.
By analyzing these aspects, students will develop critical thinking about formal language
assessments and their role in education.
Activity 5: Informal Assessment - Beyond the Test

Key Answers to Discussion Questions:

1. What are the strengths and weaknesses of this method?

Informal Strengths Weaknesses


Assessment
Method
Observations Provides real-time insights into student Can be subjective without
interactions and communication skills. clear criteria.
Portfolios Tracks long-term student progress, Requires time and effort to
showcases individual strengths. collect and assess
effectively.
Interviews Engages students in authentic Can be time-consuming and
communication; allows for difficult to standardize.
personalized feedback.
Self-assessment Encourages metacognition and Students may lack accuracy
responsibility for learning. in evaluating themselves.
Peer assessment Promotes collaborative learning and Risk of bias or unhelpful
deeper engagement. feedback if not guided
properly.

2. In what situations would this method be most appropriate?

Method Appropriate Situations


Observations Evaluating classroom participation, student behavior, and group work.
Portfolios Measuring progress in writing, speaking, or project-based learning.
Interviews Assessing speaking and listening skills in an interactive setting.
Self-assessment Encouraging student reflection before exams or after a project.
Peer Used in group activities or when giving feedback on writing/speaking
assessment tasks.

3. How can the data collected through this method be used to inform instruction?
 Identifies individual student needs – Observations help teachers adjust lesson plans.
 Tracks progress over time – Portfolios reveal strengths and weaknesses in different skills.
 Personalizes feedback – Interviews allow teachers to provide direct support.
 Encourages student involvement – Self-assessment helps learners set personal goals.
 Enhances peer learning – Peer feedback fosters collaborative skill development.

4. What ethical considerations need to be taken into account when using this method?
 Fairness & Objectivity – Teachers must ensure unbiased observations and feedback.
 Student Privacy – Portfolio content should be confidential and used for learning purposes
only.
 Guidance for Peer & Self-Assessment – Students need clear criteria to avoid misleading
 evaluations.
 Avoiding Over-Reliance on Informal Methods – Combining informal and formal
assessments ensures accuracy.
Conclusion:
- Informal assessment is essential for holistic language evaluation.
- It complements formal tests by promoting deeper learning and engagement.
- Best practice: Use a mix of observations, self/peer assessments, portfolios, and interviews to
support language development.
This key helps students understand how informal assessments contribute to meaningful learning!
Chapter 2: Approaches to Language Testing and Assessment
Language testing and assessment have evolved significantly over the years, reflecting
changes in language teaching methodologies and the broader understanding of language
learning. This chapter will explore the primary approaches to language assessment and
testing: the traditional approach and the communicative approach, along with the
distinction between formative and summative assessment. It also presents the new trend
in language testing and assessment which is technology-enhanced approach.

2.1 Traditional approach to language assessment

The traditional approach to language assessment, often referred to as the discrete-point


approach, focuses on testing individual language components such as grammar,
vocabulary, and pronunciation in isolation. This approach emphasizes accuracy and the
ability to manipulate language forms correctly (McCallum, 2020; Settles, et al., 2020)

Key Characteristics:

 Discrete-point tasks: Tests often consist of multiple-choice questions, fill-in-the-


blank exercises, and cloze tests that target specific language points.
 Emphasis on accuracy: Correctness of form and mechanics is prioritized over
communicative competence.
 Standardized testing: Traditional assessments tend to be standardized, with fixed
formats and scoring criteria.

Limitations:

 Limited reflection of real-world language use: Traditional assessments may not


adequately assess learners' ability to use language in authentic communicative
contexts.
 Overemphasis on grammar and vocabulary: This approach may neglect other
important aspects of language proficiency, such as fluency and discourse skills.
 Potential for test anxiety: The high-stakes nature of traditional assessments can
lead to test anxiety and negatively impact learners' performance.

Key Characteristics:

• Discrete-point tasks:
o Multiple-choice questions
o Fill-in-the-blank exercises
o Cloze tests
o Dictation
o Grammar translation exercises
• Emphasis on accuracy:
o Correct usage of grammar rules, vocabulary, and pronunciation
• Standardized testing:
o Standardized tests with fixed formats and scoring criteria
• Teacher-centered:
o Teachers as the primary authority in the assessment process

Limitations:

• Limited reflection of real-world language use:


o Does not assess learners' ability to use language in authentic contexts

• Overemphasis on grammar and vocabulary:


o Neglects other important aspects of language proficiency, such as fluency, discourse
skills, and cultural competence

• Potential for test anxiety:

o High-stakes nature of traditional assessments can lead to test anxiety and negatively
impact learners' performance

2.2. Communicative approach to language assessment

The communicative approach to language assessment emphasizes the assessment of


learners' ability to use language effectively for communication. It focuses on how
language is used in real-world contexts and aims to measure communicative competence
(Bachman & Adrian, 2022; Fulcher, 2014).

Key characteristics:

 Authentic tasks: Communicative assessments often employ authentic tasks, such


as role-plays, discussions, and writing tasks, that simulate real-life language use.
 Focus on communicative competence: This approach assesses learners' ability to
use language fluently, accurately, appropriately, and meaningfully.
 Holistic scoring: Communicative assessments often use holistic scoring, which
considers multiple aspects of language performance, including content,
organization, grammar, vocabulary, and mechanics.

Advantages:

 Better reflection of real-world language use: Communicative assessments provide


a more accurate picture of learners' language abilities in authentic contexts.
 Reduced test anxiety: The authentic nature of communicative tasks can help
reduce test anxiety and create a more positive testing experience.
 Motivation for language learning: Communicative assessments can motivate
learners by focusing on their ability to use language for meaningful purposes.

Key Characteristics:

 Authentic tasks:
o Role-plays
o Simulations
o Interviews
o Presentations
o Writing task
 Focus on communicative competence:
o Assessment of learners' ability to use language fluently, accurately, appropriately,
and meaningfully
 Holistic scoring:
o Consideration of multiple aspects of language performance, including content,
organization, grammar, vocabulary, and mechanics
 Learner-centered:
o Learners actively participate in the assessment process

Limitations of communicative language testing

While communicative language testing is a significant improvement over traditional,


discrete-point testing, it does have certain limitations:

Subjectivity in scoring:

o Inter-rater reliability: Different raters may have varying interpretations of performance,


leading to inconsistency in scoring.

o Intra-rater reliability: Even the same rater may assess the same performance differently
on different occasions.

Practical constraints:

o Time-consuming: Designing, administering, and scoring communicative tests can be


time-consuming, especially for large-scale assessments.

o Resource-intensive: It requires significant resources, including trained raters,


appropriate facilities, and authentic materials.
Difficulty in assessing specific skills:

o Isolation of skills: It can be challenging to isolate and assess specific language skills,
such as grammar or vocabulary, in a purely communicative context.

o Focus on overall performance: Communicative tests often prioritize overall


performance, making it difficult to identify specific areas of strength and weakness.

Cultural bias:

o Cultural differences: Test tasks may not be equally accessible to learners from different
cultural backgrounds, potentially leading to unfair assessment.

o Cultural sensitivity: It is important to consider cultural factors when designing and


administering communicative tests.

Limited standardization:

o Lack of standardized procedures: Communicative tests often lack standardized


procedures, making it difficult to compare results across different contexts.

o Variability in test administration: Differences in test administration can affect test-taker


performance.

To mitigate these limitations, language testers often combine communicative tasks with
more traditional, discrete-point tasks. Additionally, the use of rating scales, rubrics, and
clear scoring criteria can help to improve the reliability and validity of communicative
language assessments.

2.3 Formative and summative assessment

Formative and summative assessment are two types of assessment that serve different
purposes in language education.

 Formative assessment: Formative assessment is ongoing assessment that provides


feedback to learners and teachers to inform instruction and guide learning. It helps
identify learners' strengths and weaknesses and allows for timely adjustments to
teaching and learning activities (Chandio & Jafferi, 2015).
 Summative assessment: Summative assessment is used to evaluate learners'
overall achievement at the end of a learning period. It provides a snapshot of
learners' progress and can be used for accountability purposes, such as grading or
certification.

Formative assessment:
 Purpose: To monitor learner progress and provide feedback to inform instruction
Examples:

o Quizzes
o Self-assessment
o Peer assessment
o Teacher observation
o Portfolio assessment

 Benefits:
o Identifies learners' strengths and weaknesses
o Provides timely feedback
o Allows for adjustments to teaching and learning
o Promotes learner autonomy and self-directed learning

Summative assessment:

 Purpose: To evaluate learner achievement at the end of a learning period

Examples:

o Final exams
o Standardized tests
o Project-based assessments

 Benefits:
o Measures overall learning outcomes
o Provides a benchmark for progress
o Informs decisions about future learning

The differences between formative and summative assessments are presented in Table 2.1
below:

Table 2.1 Formative vs. summative assessment

Feature
Purpose Monitor student progress and Evaluate student achievement at the
provide feedback to inform end of a learning period.
instruction.
Timing Ongoing throughout the learning Occurs at specific intervals, such as
process. the end of a unit or semester.
Examples Quizzes, exit slips, homework Final exams, standardized tests,
assignments, class discussions, projects, presentations.
peer reviews.
Feedback Frequent and specific feedback Less frequent and more general
to guide learning. feedback, often focused on overall
performance.
Grading Often ungraded or used for Typically graded and used for
informal assessment. formal evaluation.
Impact on Informs instructional decisions Used to evaluate the effectiveness of
Instruction and adjustments. instruction and curriculum.

Integrating formative and summative assessment:

• Using formative assessment to inform summative assessment:


o Using formative assessment data to identify areas where learners need additional
support
o Adjusting summative assessment to align with the specific needs of learners

• Using summative assessment to inform future learning:


o Analyzing summative assessment results to identify areas for improvement in future
teaching and learning
o Using summative assessment data to inform curriculum development and program
evaluation

By effectively combining traditional and communicative approaches, as well as formative


and summative assessment, language educators can create a comprehensive assessment
framework that supports learner development and promotes language proficiency.

2.4 Technology-based language assessment

Technology has revolutionized language assessment, offering innovative ways to


evaluate language skills. The advent of technology has significantly transformed the
landscape of language assessment. Traditional pen-and-paper tests have been
complemented, and in some cases, replaced by innovative digital tools that offer greater
efficiency, objectivity, and adaptability. This chapter explores the various facets of
technology-based language assessment, examining its advantages, limitations, and future
trends.
Benefits of technology-based language assessment

 Efficiency and scalability: Technology-based assessments can be administered to a


large number of test-takers simultaneously, reducing the time and resources
required for traditional testing.
 Objectivity: Automated scoring systems can minimize human error and bias in the
evaluation process, ensuring a more objective assessment of language skills.
 Adaptivity: Computer-adaptive tests (CATs) can tailor the difficulty level of

questions to the test-taker's ability, providing a more accurate and efficient

assessment.

 Authentic assessment: Technology can facilitate authentic language assessment

tasks, such as simulating real-world communication scenarios through video

conferencing or online discussions.

 Immediate feedback: Automated scoring and feedback systems can provide

learners with timely information on their performance, enabling them to identify


areas for improvement and adjust their learning strategies accordingly.

Types of technology-based language assessment

 Computer-based tests (CBTs): These tests are administered on computers and can
include multiple-choice questions, fill-in-the-blank exercises, and short-answer
questions.
 Computer-adaptive tests (CATs): CATs tailor questions to the test-taker's ability
level, adjusting the difficulty as the test progresses. CATs adjust the difficulty of
questions based on the test-taker's performance, ensuring a more accurate
assessment of their ability level.

Example: A language proficiency test might start with easier questions. If the test-
taker answers correctly, the next question will be more difficult. If they answer
incorrectly, the next question will be easier. This ensures that the test is both
challenging and fair.

 Automated speech recognition (ASR): ASR technology can assess oral language
skillsbyanalyzingspeechpatterns,pronunciation,andfluency. ASRtechnology can
assess oral language skills by analyzing speech patterns, pronunciation, and
fluency.
Example: A language learner can practice speaking into a computer or
smartphone, and the software can provide feedback on pronunciation, intonation,
and grammar.

 Interactive language learning platforms: These platforms offer a variety of


exercises and activities that can be used for formative assessment, providing
immediate feedback on grammar, vocabulary, and pronunciation.

Example: Duolingo and Babbel use adaptive learning algorithms to tailor exercises
to each learner's needs. They provide immediate feedback on grammar,
vocabulary, and pronunciation.

 E-portfolios: E-portfolios allow learners to collect and reflect on their language


learning journey, providing a comprehensive overview of their progress and
achievements. 4. E-Portfolios

Example: Students can upload written assignments, recorded speeches, and


multimedia projects to their e-portfolios. Teachers can provide feedback and track
student progress over time.

 Online writing labs (OWLs): OWLs provide automated feedback on writing


assignments.

Example: A student can submit an essay to an OWL, which will check for
grammar, spelling, and punctuation errors. The software can also provide
suggestions for improving sentence structure and style.

These are just a few examples of how technology can be used to enhance language
assessment. As technology continues to evolve, we can expect to see even more
innovative and effective ways to assess language skills.

Challenges and limitations

 Technical issues: Technical difficulties, such as internet connectivity problems or


software glitches, can disrupt the testing process.
 Digital divide: Not all learners have equal access to technology, which can create
disparities in assessment opportunities.

 Security concerns: Ensuring the security of online assessments and protecting


sensitive information is crucial.
 Ethical considerations: The use of AI and machine learning in language
assessment raises ethical questions about data privacy and algorithmic bias.
Future trends

 Artificial Intelligence (AI): AI-powered language assessment tools can provide


more accurate and personalized feedback, adapting to individual learner needs.
 Virtual and augmented reality: Immersive VR and AR experiences can create
authentic language learning environments, enabling learners to practice language
skills in simulated real-world situations.
 Neurotechnology: Neurotechnology can measure cognitive processes related to
language learning, providing insights into how learners acquire language and
identify areas for improvement.

In conclusion, technology-based language assessment offers numerous


advantages, but it is essential to address the challenges and ethical considerations
associated with its implementation. By harnessing the power of technology, we
can create more effective, efficient, and equitable language assessment practices
that support language learners worldwide.

The choice of assessment approach depends on the specific goals of the


assessment, the context of language learning, and the needs of the learners. It is
important to consider the strengths and limitations of each approach and to use a
combination of formative and summative assessment to provide a comprehensive
picture of learners' language development. As language teaching and learning
continue to evolve, it is likely that assessment practices will also continue to adapt
to meet the changing needs of learners and educators.

Additional considerations

 Technology in language assessment: Explore how technology can be used to


enhance language assessment, including automated scoring, online testing, and
computer-adaptive testing.
 Assessment for learning: Discuss how assessment can be used as a tool for
learning, not just for evaluation.
 Cultural considerations in language assessment: Address the importance of
considering cultural factors in language assessment and testing.

2.5. Consoliation activities

Activity 1: Analyzing language assessment approaches


Objectives: To make students compare and constrast different language assessment
approaches.

Procedure: Students work in pairs to examine the following excersies and asnwer the
following quesitons

 What language assessment approach presented in this chapter do they reflect?


Why?
 How should these activities be used in the English language classrooms for high
school in Vietnam?

Excercise 1: Translation and grammar Procedure:

Presentation: The teacher presents a grammar rule, such as the past tense of
regular verbs.

Translation: Students translate sentences from their L1 (native language) into the
target language, applying the grammar rule.

Exercise 1: Students complete a series of exercises, such as filling in blanks or


translating short passages.

Correction: The teacher corrects the exercises and provides feedback. Example
grammar rule:

• Past tense of regular verbs: add -ed to the base form of the verb.
o Example: play -> played, work -> worked

Example Exercise:
Fill in the blanks with the correct past tense form of the verb.

1. Yesterday, I _____ (play) tennis.


2. She _____ (work) late last night.
3. They _____ (study) for the test.

Exercise 2: Multiple-choice vocabulary quiz

Procedure:

Word list: The teacher provides a list of vocabulary words to be learned.

Multiple-choice questions: The teacher creates a quiz with multiple-choice


questions, each with four answer choices.
Quiz administration: Students take the quiz individually.
Scoring: The teacher scores the quizzes and provides feedback.

Example vocabulary list: happy, sad, angry, excited, tired

Example multiple-choice questions:

1. How do you feel when you get a good grade? a) happy b) sad c) angry d) tired
2. What do you feel when you fail a test? a) happy b) sad c) excited d) tired

What do you think about the activities? How should they be applied in language classes
for high school students in Vietnam.

Activity 2: Assessing students' ability to use language in a real-world context


Objectives:

 To make students aware of communicative language assessment and assessing


language use in a real-world context
 To make students analyze some benefits of using real-world contexts to assess
students’ language performance

Procedure: Some students present a scenario, such as ordering food at a restaurant or


asking for directions. Other students prepare for their roles by practicing vocabulary and
phrases. And perform the role-play in pairs or small groups. The student assessors
observe and assess other students' performance based on their fluency, accuracy, and
appropriateness.

Example scenario:

• Restaurant role-play:
o Student A: You are a customer at a restaurant. Order your favorite dish and a
drink.

o Student B: You are a waiter/waitress. Take the customer's order and respond
appropriately.

Activity 3: Analyzing language assessment approach(es) to be employed for a unit in


an English textbook in Vietnam

Objective: To make students be aware of using an appropriate assessment approach for a


unit in an English textbook.
Procedure: Students read the introduction of an English textbook currently used at high
schools in Vietnam and examine a unit in an English textbook. What language
assessment approach(es) should be used to design language tests? Why?

(Source: Hoang et al., 2018, Global success 10)

Activity 4: Deciding whether to use formative or summative assessment


Objective: To make students distinguish and apply formative or summative assessment

Procedure: Students examine the following project which aims to assess students'
comprehensive understanding of a specific topic or theme, their ability to apply
knowledge and skills, and their communication skills. Students decide if formative or
summative assessment should be used to evaluate students’ project and give reasons.

Topic: Global Warming

Task: Students will work in groups of 3-4 to create a multimedia presentation on the topic
of global warming. The presentation should include the following components:

Introduction: A brief overview of global warming, including its causes and effects.

Causes: A detailed explanation of the factors contributing to global warming, such as


greenhouse gas emissions, deforestation, and industrialization.

Effects: A discussion of the impacts of global warming on the environment, society, and
economy.

Solutions: A proposal of potential solutions to mitigate climate change, such as


renewable energy, energy conservation, and sustainable practices.

Conclusion: A summary of the key points and a call to action.


Activity 5: Unpacking communicative language assessment
Objective: To make students know how to apply communicative language assessment

Procedure: Students examine the following role play for communicative language
assessment which aims to assess students' ability to use language in a real-world context,
focusing on fluency, accuracy, and appropriateness. Then choose activity in an English
textbook currently in use at high schools in Vietnam and design a similar activity for
communicative language assessment

Task: Students will be assigned roles in a specific scenario. They will be required to
interact with each other, using target language to accomplish a task or solve a problem.

Scenario: Airport role-play


o Role A: A traveler who has lost their luggage.

o Role B: An airline representative. Task:

 Traveler: Explain the situation to the airline representative, provide details about
the lost luggage, and ask for assistance.
 Airline Representative: Listen to the traveler's problem, offer solutions, and
provide information about the lost luggage procedures.

Assessment criteria:

 Fluency: Ability to speak smoothly and naturally.


 Accuracy: Correct use of grammar, vocabulary, and pronunciation.
 Appropriateness: Use of appropriate language and register for the situation.
 Problem-solving: Ability to understand the problem and propose solutions.
 Interactivity: Ability to engage in a natural conversation with the partner.

Feedback: The teacher will observe the students' performance and provide feedback
on their strengths and weaknesses. Feedback can be given orally or in written form.

References

Bachman, L., & Adrian, P. (2022). Language assessment in practice: Developing


language assessments and justifying their use in the real world. Oxford University Press.
Chandio, M. T., & Jafferi, S. (2015). Teaching English as a language not subject by
employing formative assessment. Journal of Education and Educational Development,
2(2), 151-171.
Fulcher, G. (2014). Testing second language speaking. Routledge.
Hoang et al., (2018). Tiếng Anh 10: Global success. The educational publishing house.
McCallum, L. (2020). Traditional assessment and encouraging alternative assessment that
promotes learning: Illustrations from EAP. In Perspectives on Language Assessment
Literacy (pp. 33-51). Routledge.
Settles, B., T. LaFlair, G., & Hagiwara, M. (2020). Machine learning–driven language
assessment. Transactions of the Association for computational Linguistics, 8, 247- 263
Exercise 1: Translation and Grammar

Assessment Approach: Traditional Approach

 Why?
o This exercise focuses on discrete-point tasks such as filling in blanks and
translating sentences.
o Emphasizes accuracy in grammar and vocabulary.
o Uses teacher-centered correction instead of communicative or contextualized
feedback.
o Relies on standardized testing formats like cloze tests and translation exercises.
 Application in Vietnamese high school classrooms:
o Can be useful for reinforcing grammar rules in controlled practice.
o Should be supplemented with communicative tasks to ensure students develop
practical language use.

Exercise 2: Multiple-Choice Vocabulary Quiz

Assessment Approach: Traditional Approach

 Why?
o Uses multiple-choice questions, a discrete-point testing method.
o Focuses on isolated vocabulary knowledge rather than language use in context.
o Uses standardized scoring without assessing communication skills.
 Application in Vietnamese high school classrooms:
o Effective for testing vocabulary knowledge quickly.
o Can be useful for formative assessment, allowing teachers to identify areas where
students need support.
o Should be combined with contextualized vocabulary activities (e.g., sentence
creation, role-plays) for better retention.

Overall Reflection on These Activities

 Both exercises belong to the traditional approach to language assessment.


 While these methods help in testing accuracy, they do not fully reflect real-world
language use.
 To make these exercises more effective, they should be integrated with communicative
approaches, such as:
o Using vocabulary quizzes in contextualized dialogues.
o Turning grammar exercises into peer discussion tasks where students apply rules
in conversation.
o Encouraging self-assessment and peer correction to develop learner autonomy.
Activity 2: Assessing Students' Ability to Use Language in a Real-World Context, based on
Chapter 2.

Assessment Approach: Communicative Approach

 Why?
o Uses authentic tasks (e.g., role-plays, discussions).
o Focuses on communicative competence rather than just linguistic accuracy.
o Involves holistic scoring, considering fluency, coherence, and interaction.
o Encourages learner-centered assessment, where students actively participate.

Breakdown of the Activity

Scenario: Restaurant Role-Play

1. Student A (Customer): Orders a dish and a drink.


2. Student B (Waiter/Waitress): Takes the order and responds appropriately.

 Key Language Skills Assessed:


o Fluency: Ability to speak naturally.
o Accuracy: Proper use of grammar and vocabulary.
o Appropriateness: Correct register (e.g., politeness).
o Interaction Skills: Ability to maintain conversation.

Application in Vietnamese High Schools

 How should this be used in the classroom?


o Encourage real-life application: Students practice using English in everyday
situations.
o Use peer and self-assessment: Students can evaluate each other’s fluency and
accuracy.
o Provide immediate feedback: Teachers can guide students on pronunciation and
coherence.
o Incorporate technology: Use AI chatbots or recorded role-plays for reflection.

Comparison with Traditional Approach


Communicative Approach (This
Aspect Traditional Approach
Activity)
Focus Accuracy in grammar and vocabulary Real-world communication and fluency
Role-plays, discussions, interactive
Tasks Multiple-choice, fill-in-the-blanks
scenarios
Assessment
Standardized, teacher-centered Holistic, learner-centered
Type
Limited real-world relevance, high Subjectivity in scoring, requires more
Limitations
test anxiety resources

Conclusion

This activity effectively applies the communicative approach to language assessment. It helps
students develop real-world English proficiency while making the assessment process more
engaging and interactive.
Activity 3: Analyzing Language Assessment Approach(es) for a Unit in an
English Textbook in Vietnam

Assessment Approach Selection

 The best assessment approach depends on the learning objectives of the unit.
 A combination of traditional and communicative approaches is often ideal for
comprehensive assessment.

Step 1: Identifying the Learning Goals

Before selecting an assessment approach, teachers should analyze the unit’s:

 Language focus (Grammar, Vocabulary, Skills)


 Communicative functions (Speaking, Writing)
 Real-life applications (Authentic language use

Step Choosing the Assessment Approach

Unit Content Recommended


Why?
Type Assessment Approach
Grammar-based Uses structured, rule-based exercises to
Traditional Approach
lessons reinforce accuracy.
Vocabulary Traditional + Multiple-choice quizzes (traditional) + role-
development Communicative plays/dialogues (communicative).
Listening and Emphasizes fluency, pronunciation, and
Communicative Approach
speaking interaction through role-plays or discussions.
Reading Uses multiple-choice, matching, and short-
Traditional Approach
comprehension answer questions for comprehension.
Encourages process writing, peer reviews, and
Writing tasks Communicative Approach
real-world writing tasks.

Step 3: Application in Vietnamese High Schools

 Balance between traditional and communicative approaches:


o Use multiple-choice and fill-in-the-blanks for assessing language forms.
o Include real-world writing and speaking tasks for assessing communicative
competence.
 Consider Formative vs. Summative Assessment:
o Formative: Continuous feedback through quizzes, peer assessments, portfolios.
o Summative: Final exams, project-based assessments, oral presentations.

Conclusion
 A hybrid assessment approach is recommended, integrating traditional for accuracy and
communicative for fluency and application.
 Teachers should align assessments with real-world tasks to ensure students develop
practical English skills.
Activity 4: Deciding Whether to Use Formative or Summative Assessment

Assessment Type Selection

The project on Global Warming involves research, presentation, and collaboration. The decision
to use formative or summative assessment depends on the purpose of evaluation.

Step 1: Understanding the Two Assessment Types

Assessment
Purpose Examples When to Use?
Type
Quizzes, self-assessments, During the learning process
Provides feedback
Formative peer reviews, teacher to help students adjust their
for improvement
observations work
At the end of a unit or term
Evaluates final Final exams, standardized
Summative to measure overall learning
achievement tests, project presentations
outcomes

Step 2: Choosing the Assessment for the Global Warming Project

Recommended
Project Component Why?
Assessment Type
Brainstorming ideas and Allows students to refine their ideas based
Formative
research phase on feedback.
Drafting and creating Helps students improve content, structure,
Formative
presentation materials and clarity before final submission.
Final multimedia Assesses students' ability to apply
Summative
presentation knowledge and communicate effectively.
Self-reflection on the Encourages students to evaluate their
Formative
learning process progress and set goals for improvement.

Step 3: Application in Vietnamese High Schools

 Use formative assessment throughout the project:


o Peer reviews on content and clarity.
o Teacher feedback on drafts.
o Self-assessment checklists for organization and language use.
 Summative assessment at the end:
o Evaluate the final presentation based on clarity, accuracy, creativity, and delivery.
o Assign a grade or score for project completion.

Conclusion
 A combination of formative and summative assessment is ideal for this project.
 Formative assessment helps improve student performance before the final product.
 Summative assessment ensures that students demonstrate their overall understanding and
skills.
Activity 5: Unpacking Communicative Language Assessment

Assessment Approach: Communicative Approach

This activity aligns with communicative language assessment as it:

 Simulates real-world interactions (e.g., airport role-play).


 Focuses on fluency, accuracy, appropriateness, problem-solving, and interaction.
 Uses holistic scoring instead of isolated grammar or vocabulary tests.

Breakdown of the Activity

Scenario: Airport Role-Play

1. Role A (Traveler): Explains the situation about lost luggage and asks for help.
2. Role B (Airline Representative): Responds appropriately, asks clarifying questions, and
provides solutions.

Assessment Criteria (Based on Communicative Approach)

Criteria Description
Fluency Speaks smoothly and naturally without excessive pauses.
Accuracy Uses correct grammar, vocabulary, and pronunciation.
Appropriateness Uses polite and contextually appropriate expressions.
Problem-Solving Identifies the problem and suggests logical solutions.
Interactivity Engages naturally in conversation, responds appropriately.

Application in Vietnamese High Schools

 How should this be used in the classroom?


o Pre-activity preparation: Teach key vocabulary and expressions.
o Pair or group role-play: Students take turns in different roles.
o Peer and self-assessment: Students reflect on their communication skills.
o Teacher feedback: Provide constructive feedback on areas for improvement.

holistics, learner-centereComparison with Traditional Approach


Communicative Approach (This
Aspect Traditional Approach
Activity)
Focus Grammar, vocabulary accuracy Real-world communication, fluency
Role-plays, discussions, interactive
Tasks Multiple-choice, fill-in-the-blanks
scenarios
Assessment
Standardized, teacher-centered Holistic, learner-centered
Type
Limited real-world relevance, high Subjectivity in scoring, requires more
Limitations
test anxiety preparation

Conclusion

 This activity effectively applies communicative language assessment principles.


 It enhances real-world communication skills while making assessment engaging.
 A rubric-based assessment ensures fairness and consistency.

Would you like a detailed grading rubric for this role-play?


Chapter 3: Qualities of a Good Language Test

Language assessment plays a pivotal role in various educational and professional


contexts, from classroom settings to high-stakes examinations like university admissions
or professional certifications. To ensure that language tests are fair, accurate, and
meaningful, it is crucial to adhere to a set of fundamental principles. This chapter will
explore four key principles of language assessment: reliability, validity, practicality, and
ethical considerations. By understanding and applying these principles, language testers
can develop and implement assessments that are fair, accurate, meaningful, and ethically
sound, ultimately leading to more effective and equitable language learning and teaching
practices.

3.1 Reliability

Reliability in language testing refers to the consistency and dependability of test scores
(Bachman & Palmer, 1996; Liu, et al., 2020). A reliable test produces consistent results
across different administrations, raters, or test items. In other words, a reliable test
minimizes random error or fluctuations in scores that are not due to true differences in
language proficiency. Ensuring reliability is crucial for making valid inferences about
test-takers' language abilities based on their test scores.

Types of reliability

 Test-retest reliability: This method involves administering the same test to the
same group of test-takers on two separate occasions and then correlating the two
sets of scores. A high correlation coefficient indicates that the test produces
consistent results over time. However, this method can be affected by factors such
as practice effects and memory.

For example: A standardized English proficiency test for university admissions.

Procedure: A group of applicants takes the test. After a suitable interval (e.g., two
weeks), the same group of applicants takes the exact same test again under similar
conditions. The scores from both administrations are compared for each
individual.

Demonstrating test-retest reliability:

High reliability: If the scores from the first and second administrations are highly
correlated (e.g., a correlation coefficient close to 1.0), it suggests that the test
produces consistent results over time. This indicates that the test is reliable
because it is measuring the same underlying language proficiency consistently,
regardless of when it is administered.

Low reliability: If the scores show little or no correlation between the two
administrations, it suggests that the test is unreliable. This could be due to factors
like:

o Test-takers' memory: They might remember answers from the first


administration, artificially inflating their scores on the second attempt.
o Practice effects: They might have improved their language skills between
the two administrations, leading to higher scores on the second attempt.
o Uncontrolled factors: Differences in testing conditions, time of day, or the
emotional state of the test-takers could also influence scores.

Overall, test-retest reliability assesses the consistency of test scores over time. A reliable
test should produce similar results when administered to the same group of test-takers
under comparable conditions.

 Parallel forms reliability: This method requires the creation of two equivalent
versions of the same test, which are then administered to the same group of test-
takers. The scores on the two versions are then correlated. A high correlation
coefficient suggests that the two versions of the test are equivalent and measure
the same construct consistently.

Scenario:

Test: A standardized English achievement test for 8th-grade students, designed to


assess reading comprehension, grammar, and vocabulary.

Purpose: To measure students' overall English language proficiency at the end of


the school year.

Parallel Forms: Two equivalent versions of the test (Form A and Form B) are
created.

o Equivalence: Both forms cover the same range of English language skills
(reading, grammar, vocabulary) with a similar level of difficulty.
 Example:
 Reading Comprehension: Form A might include a passage
about historical figures, while Form B includes a passage
about scientific discoveries, but both passages assess similar
reading comprehension skills (main idea, supporting details,
inference).
 Grammar: Both forms might include questions on verb
tenses and subject-verb agreement, but with different
sentence structures and contexts.
 Vocabulary: Both forms might include vocabulary items
related to common academic topics, but with different word
choices and question formats (e.g., multiple choice, fill-in-
the-blank).

Procedure:

o A group of 8th-grade students takes Form A of the English achievement


test.
o After a suitable interval (e.g., a few weeks), the same group of students
takes Form B of the test.
o The scores from both forms are compared for each student.

Demonstrating parallel forms reliability:

High reliability: If the scores from Form A and Form B are highly correlated
(e.g., a correlation coefficient close to 1.0), it suggests that the two forms are
equivalent and measure the same underlying English language proficiency
consistently. This indicates that the test is reliable because it produces similar
results regardless of which version is administered.

Low reliability: If the scores show little or no correlation between the two forms,
it suggests that the two forms are not equivalent and may not be measuring the
same construct. This could be due to differences in difficulty, content coverage, or
question format between the two forms.

In short, parallel forms reliability in this context ensures that the English achievement test
consistently measures students' English language proficiency, regardless of which version
they take. This is crucial for making fair and accurate assessments of student learning and
for identifying areas where students may need additional support.

 Internal consistency reliability: This method assesses the consistency of items


within a single test. One common measure of internal consistency is Cronbach's
alpha, which estimates the average correlation between all possible pairs of items
on the test. A high alpha coefficient indicates that the items on the test are
measuring the same underlying construct.

Scenario:
Test: A standardized English proficiency test for university applicants, designed
to assess reading comprehension.

Internal consistency consideration: The test aims to measure a single underlying


construct, such as general reading comprehension ability. Therefore, the items
within the reading comprehension section should be measuring the same thing.

Example:

Test content: The reading comprehension section consists of several passages


with multiple-choice questions.

Internal consistency check: To assess internal consistency, we can calculate


Cronbach's alpha, a common statistical measure. Cronbach's alpha estimates the
average correlation between all possible pairs of items on the test.

o High internal consistency: If Cronbach's alpha is high (typically above


0.70), it suggests that the items within the reading comprehension section
are measuring the same construct (i.e., general reading comprehension
ability) and are internally consistent (Tang, et al., 2014). This means that
students who perform well on one item are likely to perform well on other
items as well.
o Low internal consistency: If Cronbach's alpha is low, it suggests that the
items are not measuring the same construct. This could be due to factors
such as:
 Item ambiguity or difficulty: Some items may be poorly written or
too difficult, leading to inconsistent responses.
 Heterogeneity of content: The items may be measuring different
aspects of reading comprehension (e.g., vocabulary, main idea,
inference) rather than a single, unified construct.

Overall, internal consistency reliability ensures that the items within a test section are
measuring the same underlying construct and are not measuring unrelated or inconsistent
skills. In this example, a high Cronbach's alpha for the reading comprehension section
would suggest that the test is reliably measuring general reading comprehension ability.

 Inter-rater reliability: This method is particularly important for subjective


assessments, such as writing or speaking tests. It measures the degree of
agreement between two or more raters on the same set of responses. Inter-rater
reliability can be estimated using statistical methods such as Cohen's kappa or
intraclass correlation coefficients.

Scenario:
 Test: An English speaking exam for university students
 Inter-rater reliability consideration: Since the assessment involves subjective
judgment by human raters (interviewers), it is crucial to ensure consistent scoring
across different raters.

Example:

Interview procedure: A group of applicants is interviewed by two different,


trained raters.

Scoring: Each rater independently scores the applicants' performance on criteria


such as fluency, grammar, vocabulary, pronunciation, and overall communication
effectiveness.

Inter-rater reliability analysis: The scores given by the two raters for each
applicant are compared. Statistical methods like Cohen's kappa or intraclass
correlation coefficients are used to calculate the level of agreement between the
raters.

o High inter-rater reliability: If the correlation between the scores of the


two raters is high, it indicates that they are consistently evaluating the
applicants' performance in a similar manner. This suggests that the scoring
process is reliable and not significantly influenced by the subjective
judgment of individual raters.
o Low inter-rater reliability: If the correlation between the scores is low, it
suggests that the raters are not consistently evaluating the applicants. This
could be due to factors such as:
 Vague scoring rubrics: The criteria for scoring are not clearly
defined or consistently applied.
 Lack of rater training: Raters may not have received adequate
training on how to apply the scoring rubrics consistently.
 Rater bias: Personal biases or preferences of individual raters may
be influencing their scores.

Overall, inter-rater reliability ensures that the scores assigned to test-takers are consistent
across different raters. In the case of oral proficiency interviews, high inter-rater
reliability is essential to ensure that the assessment is fair and objective.

Factors affecting reliability

Several factors can affect the reliability of language tests:


 Test length: Longer tests tend to be more reliable than shorter tests, as they
provide more opportunities to sample the test-taker's language ability.
 Test difficulty: Tests that are too easy or too difficult may not discriminate
effectively between test-takers, leading to lower reliability.
 Test administration: Factors such as testing conditions, time constraints, and
instructions can affect test-taker performance and, consequently, test reliability.
 Rater factors: Rater subjectivity, fatigue, and inconsistency can all impact the
reliability of subjective assessments.

Improving reliability

Several strategies can be used to improve the reliability of language tests:

 Clear instructions: Providing clear and unambiguous instructions to test-takers can


help to minimize confusion and ensure that everyone understands the task
requirements.
 Standardized procedures: Establishing standardized testing procedures can help to
control for extraneous variables and ensure that all test-takers are assessed under
similar conditions.
 Rater training: Providing raters with clear scoring rubrics and training on how to
apply them consistently can improve inter-rater reliability.
 Item analysis: Conducting item analysis can help to identify and eliminate
problematic items that do not contribute to the overall reliability of the test.

In general, reliability is a fundamental aspect of language test quality. By ensuring that


tests are reliable, we can be more confident that the scores they produce accurately reflect
test-takers' language abilities. This is essential for making valid inferences about
language proficiency and for using test scores to make important decisions, such as
admissions to educational programs or job placement.

3.2 Validity

In the realm of language assessment, the concept of validity holds paramount importance.
Validity, in its essence, refers to the extent to which a language test accurately measures
what it purports to measure (Chapelle & Lee, 2021). It is a multifaceted concept that
encompasses various aspects, including content, construct, criterion, and consequential
validity. This section delves into the different types of language test validity, exploring
their significance in ensuring the accuracy, fairness, and meaningfulness of language
assessments.
Content validity

Content validity focuses on the degree to which a test comprehensively covers the
relevant content domain (Brown & Abeywickrama, 2019; Hughes, 2020). In language
testing, this entails ensuring that the test tasks and items adequately represent the
knowledge, skills, and abilities that are considered essential for the target language use
domain (Dinh, 2019; Siddiek, 2010). For instance, a test designed to assess academic
writing proficiency should include tasks that reflect the types of writing required in
academic settings, such as essays, research papers, and literature reviews. The following
scenario illustrates content validity:

Scenario:

 Test purpose: An English test aims to assess the reading comprehension skills of
university students in an academic setting.
 Content Validity Consideration: To ensure content validity, the test developers
would need to carefully select reading materials that are representative of the types
of texts students will encounter in their university studies.

Example:

Instead of using simplified or general interest texts, the test should include:

 Excerpts from academic textbooks: Covering various disciplines like science,


history, or literature.
 Scholarly articles: With complex sentence structures and discipline-specific
vocabulary.
 Graphs and charts: Requiring interpretation of visual data commonly found in
academic materials.

By including these types of materials, the test directly measures the skills students need
to succeed in their academic reading. It avoids irrelevant content (like casual
conversations or simple narratives) that wouldn't accurately reflect their ability to handle
university-level texts.

Overall, content validity is about ensuring the test's content aligns with the specific skills
or knowledge it intends to measure. In this case, the test content is valid because it
reflects the reading demands of a university environment.

Construct Validity

Construct validity examines whether a test effectively measures the underlying construct
or attribute it is intended to assess (Hill & McNamara, 2015). Constructs in language
testing can include language proficiency, communicative competence, or specific
language skills like reading comprehension or listening ability. To establish construct
validity, researchers often employ various methods, such as correlating test scores with
other established measures of the same construct or examining the internal structure of
the test through factor analysis. The followoing scenario illustrates construct validity:

Scenario:

 Test purpose: An English test aims to assess the writing ability of students,
specifically their ability to construct well-formed sentences.
 Construct validity consideration: To ensure construct validity, the test needs to
accurately measure the underlying construct of "sentence construction ability."
This means it should effectively differentiate between students who have a good
grasp of sentence structure and those who struggle with it.

Example:

The test includes various tasks designed to elicit different aspects of sentence
construction:

 Sentence completion: Students fill in missing words or phrases in sentences,


requiring them to understand grammatical relationships and sentence structure.
 Sentence combining: Students combine two or more simple sentences into a more
complex one, testing their ability to use conjunctions, relative clauses, etc.
 Error identification: Students identify grammatical errors in sentences, assessing
their knowledge of correct sentence structure.

Why this demonstrates construct validity:

By using a variety of tasks that target different facets of sentence construction, the test
provides a more comprehensive assessment of this underlying construct. It avoids relying
on just one type of task, which might only measure a narrow aspect of sentence
construction ability.

In general, construct validity is about ensuring the test accurately measures the theoretical
construct it intends to measure. In this case, the test is designed to measure the construct
of "sentence construction ability" by using tasks that effectively tap into different aspects
of this skill.

Criterion-related validity: Criterion-related validity explores the relationship between


test scores and an external criterion or measure of the same construct. This type of
validity is typically divided into two categories: concurrent validity and predictive
validity. Concurrent validity assesses the relationship between test scores and a criterion
measure collected at the same time, while predictive validity examines the relationship
between test scores and a criterion measure collected at a later point in time. For
example, a test designed to predict academic success in a language learning program
would exhibit high predictive validity if its scores strongly correlate with students' actual
academic performance in the program. Another aspect of criterion-related validity is to
Examining how well the test scores correlate with external criteria, such as grades in
English classes or performance on other standardized English assessments. This helps
validate the test's effectiveness in predicting students' actual English language abilities.
Below is an example of criterion-related validity:

Scenario:

 Test purpose: A university wants to assess the English language proficiency of


international students applying for admission.
 Criterion-related validity consideration: The university wants to ensure that the
English test scores accurately predict how well the students will perform in their
academic coursework.

Example:

 Test: The university administers its own English language proficiency test to
applicants.
 Criterion: The university tracks the academic performance (e.g., GPA) of the
admitted students in their first year of study.
 Analysis: The university then correlates the test scores with the students' academic
performance.

Demonstrating criterion-related validity:

If the test scores show a strong positive correlation with the students' academic
performance (i.e., higher test scores predict better academic outcomes), it demonstrates
that the test has good criterion-related validity. It means the test is effective at predicting
how well students will perform in their academic studies, which is the intended purpose
of the test.

In short, criterion-related validity focuses on how well a test's scores correlate with an
external criterion (in this case, academic performance). A strong correlation indicates that
the test is useful for predicting future performance on that criterion.

Consequential validity

Consequential validity considers the social and personal consequences of test use. It
examines the impact of a test on individuals, groups, and society as a whole. This
includes evaluating the potential benefits and drawbacks of test use, such as its effect on
learning, teaching, and educational policy. For instance, a test that is used for high-stakes
decisions, such as university admission or job placement, should be carefully scrutinized
for its potential impact on test takers' lives and opportunities. The following scenario
illustrates consequential validity”

Scenario:

 Test purpose: A standardized English proficiency test is used to determine which


students are eligible for placement in advanced English classes.
 Consequential validity consideration: The test developers and administrators
need to consider the potential consequences of using this test for placement
decisions.

Example:

 Positive consequence: The test accurately identifies students with strong English
skills, allowing them to be placed in challenging classes where they can thrive and
learn at a faster pace. This can lead to increased motivation, higher academic
achievement, and greater opportunities for these students.
 Negative consequence: The test may unfairly disadvantage certain groups of
students, such as those from low-income backgrounds or those who are English
language learners. If the test is culturally biased or does not adequately account for
their unique learning experiences, it could lead to misplacement, frustration, and
lower self-esteem. This could also perpetuate existing educational inequalities.

Demonstrating consequential validity:

To ensure consequential validity, the test developers and administrators should:

 Conduct a thorough needs assessment: Identify the specific needs and learning
goals of the target population.
 Minimize potential bias: Ensure the test is culturally fair and does not
disadvantage any particular group of students.
 Monitor the impact of test use: Track the long-term consequences of test-based
placement decisions on student learning and well-being.
 Make adjustments as needed: Based on the monitoring data, make adjustments
to the test or the placement process to mitigate any negative consequences.

In summary, consequential validity considers the broader social and personal


implications of test use. It's about ensuring that the test has a positive impact on students
and does not create unintended negative consequences.
Significance of language test validity

The significance of language test validity lies in its ability to ensure that language
assessments are accurate, fair, and meaningful. A valid language test provides reliable
information about test takers' language abilities, enabling stakeholders to make informed
decisions based on the test results. For test takers, a valid test ensures that their language
skills are assessed fairly and accurately, providing them with a true reflection of their
abilities. For educators, a valid test helps to identify students' strengths and weaknesses,
guiding instructional practices and curriculum development. For policymakers, a valid
test can inform decisions about language education programs and policies, ensuring that
they are effective and equitable.

Overall, language test validity is a complex and multifaceted concept that plays a crucial
role in ensuring the quality and meaningfulness of language assessments. By
understanding the different types of validity and their significance, language testers and
educators can strive to develop and utilize assessments that are accurate, fair, and
beneficial for all stakeholders.

3.3 Practicality

In addition to validity and reliability, practicality is a crucial consideration in language


test development and administration. A practical test is one that is feasible and efficient
in terms of time, cost, and resources. It should be easy to administer, score, and interpret,
and it should be logistically manageable (Bachman & Palmer, 1996; Brown &
Abeywickrama, 2019).

Factors related practicality

 Time constraints: Time is a valuable resource, both for test-takers and test
administrators. A practical test should be administered within a reasonable
timeframe. Test-takers should not feel rushed or pressured, but the test should also
not be excessively long, as this can lead to fatigue and reduced performance. Test
administrators should also consider the time required for test preparation,
administration, scoring, and reporting.
 Cost-effectiveness: The cost of developing, administering, and scoring a language
test can vary significantly. Test development requires time, expertise, and
resources, such as item writing, pilot testing, and statistical analysis.
Administration costs include venue rental, equipment, and personnel. Scoring
costs can be substantial, especially for subjective assessments that require human
raters. A practical test should be cost-effective, considering the available budget
and the value derived from the test results.
 Logistical feasibility: Logistical considerations include the availability of
appropriate testing venues, equipment, and personnel. Test administrators need to
ensure that the testing environment is conducive to optimal test performance, with
adequate space, lighting, and ventilation. They also need to ensure that sufficient
personnel are available to administer the test, monitor test-takers, and provide any
necessary assistance.
 Ease of administration and scoring: A practical test should be easy to administer
and score. Clear and concise instructions should be provided to test-takers, and the
testing procedures should be straightforward and easy to follow. Scoring
procedures should be well-defined and easy to apply, reducing the potential for
subjectivity and inconsistency. The use of automated scoring systems can
significantly improve the efficiency and objectivity of the scoring process. Below
is an example:

Scenario:

 Test: A standardized English proficiency test for high school students.


 Practicality Consideration: The test needs to be easy to administer and score,
especially for teachers with limited time and resources.

Example:

 Multiple-choice format: The test primarily uses a multiple-choice format for


questions on grammar, vocabulary, and reading comprehension.
 Benefits:
o Easy administration: Multiple-choice questions are straightforward for
teachers to administer. They can be easily printed or displayed on a screen,
and students can mark their answers on a separate answer sheet.
o Objective scoring: Scoring is objective and efficient. Answer keys can be
used to quickly and accurately score the tests, minimizing the potential for
human error.
o Automated scoring: In many cases, answer sheets can be scanned and
scored electronically, further increasing efficiency and reducing the time
and effort required for scoring.

Overall, the use of a multiple-choice format enhances the practicality of the test by
making it easy to administer and score. This is especially important for high-stakes tests
that need to be administered to a large number of students efficiently. A practical test
should be accessible to all test-takers, regardless of their physical or cognitive abilities.
Reasonable accommodations should be made for test-takers with disabilities, such as
providing extra time, using assistive technology, or modifying test formats. The test
should also be culturally sensitive and appropriate for test-takers from diverse
backgrounds.

 Interpretation and reporting

Test results should be easily interpretable and reported in a clear and concise manner.
Test reports should provide meaningful information about test-takers' language
proficiency, such as their strengths and weaknesses. They should also be easy to
understand and use by test-takers, educators, and other stakeholders.

Scenario:

 Test: A standardized English proficiency test for university admissions.


 Practicality consideration: The test results need to be easily interpreted and
reported in a clear and concise manner for both the university and the applicants.

Example:

 Standardized Scoring Scale: The test uses a standardized scoring scale, such as a
score range from 1 to 100 or a band score system (e.g., 1-9).
 Benefits:
o Easy interpretation: Standardized scores are easy to understand and
compare. Applicants and universities can easily see a student's overall
English proficiency level.
o Clear reporting: Test reports can be generated quickly and easily,
providing clear and concise information about the applicant's performance
in different areas of English language proficiency (reading, writing,
listening, speaking).
o Efficient decision-making: Universities can easily use the standardized
scores to make admissions decisions, compare applicants, and identify
students who may need additional language support.

Using a standardized scoring scale enhances the practicality of the test by making it easy
to interpret and report results. This facilitates efficient decision-making for both
universities and applicants.

In summary, practicality is a crucial consideration in language test development and


administration. A practical test is one that is feasible, efficient, and cost-effective, and
that can be administered and scored efficiently. By carefully considering the practical
aspects of test development and administration, we can ensure that language tests are fair,
reliable, and meaningful for all stakeholders.

3.4 Ethical considerations in English language testing


Ethical considerations are paramount in language testing. Fairness, equity, and respect for
test-takers are fundamental principles that must guide all aspects of test development and
administration. This chapter will explore key ethical considerations in English language
testing, including test bias, cultural sensitivity, test security, and the responsible use of
test results.

Test bias

Test bias occurs when a test systematically favors or disadvantages certain groups of test-
takers based on factors such as gender, ethnicity, socioeconomic status, or cultural
background. One common type of bias is cultural bias, where test items or tasks are
culturally inappropriate or unfamiliar to test-takers from certain cultural backgrounds.
For example, a reading comprehension passage that references cultural concepts or events
that are unfamiliar to test-takers from a different cultural background may disadvantage
them.

To minimize cultural bias, test developers should:

 Use culturally diverse materials: Include reading passages, listening materials, and
test items that are relevant and accessible to test-takers from diverse cultural
backgrounds.
 Avoid culturally loaded language: Use language that is clear, concise, and free
from cultural idioms or slang that may be unfamiliar to certain groups.
 Conduct pilot testing: Pilot test the test with diverse groups of test-takers to
identify and address any potential cultural biases.

Test security

Maintaining test security is crucial to ensure the integrity of the testing process. Test
materials should be kept confidential and should not be accessible to unauthorized
individuals. Measures should be taken to prevent cheating, such as proctoring exams
carefully, using secure testing environments, and implementing measures to detect and
prevent cheating.

Responsible use of test results

Test results should be used responsibly and ethically. Test scores should not be used to
make high-stakes decisions without careful consideration of their limitations. It is
important to remember that test scores are just one piece of information about a test-
taker's language ability, and they should not be the sole determinant of important
decisions such as university admissions or job placement. Test results should be
interpreted in conjunction with other relevant information, such as academic records,
recommendations, and interviews.
Test-taker rights

Test-takers have certain rights that must be respected. These rights include:

 Right to clear and concise instructions: Test-takers should be provided with clear
and unambiguous instructions for all test tasks.
 Right to a fair and equitable testing environment: Test-takers should be provided
with a comfortable and supportive testing environment that is free from
distractions and disruptions.
 Right to privacy and confidentiality: Test-takers have the right to expect that their
test scores and personal information will be kept confidential and used only for the
intended purposes.
 Right to appeal test scores: Test-takers should have the right to appeal their test
scores if they believe there has been an error in scoring or administration.

In general, ethical considerations are paramount in all aspects of language testing. Test
developers and administrators have a responsibility to ensure that tests are fair, equitable,
and unbiased. By adhering to ethical principles and best practices, we can ensure that
language tests are used responsibly and that test-takers are treated with respect and
dignity.

3.5 Student self-assessment

Instruction: Carefully read each statement in the table below. For each statement, rate
your understanding using the following scale:

1: I do not understand this at all.


2: I understand this a little, but I need more help.
3: I understand this fairly well, but I have some questions.
4: I understand this very well and can explain it to others.

Learning Objective/Concept Self- Evidence/Notes Action Plan


Assessment (Explain your (What will
Rating (1-4) rating) you do to
improve?)
Key Qualities: I can identify and
explain the four key qualities of a
good language test (reliability,
validity, practicality, ethical
considerations).
Factors Affecting Reliability: I can
identify factors that can affect the
reliability of a language test (e.g.,
test conditions, test-taker
characteristics, rater bias).
Validity: I can define validity and
explain its different types (e.g.,
content validity, construct validity,
criterion-related validity).
Practicality: I can discuss the
concept of practicality in language
testing and explain its components
(e.g., cost, time constraints, ease of
administration).
Ethical Considerations: I can
identify and discuss ethical
considerations related to language
testing (e.g., fairness, bias,
transparency, confidentiality).
Overall Understanding: I feel
confident in my understanding of the
qualities of a good language test and
their importance in language
assessment.

3.6 Consolidation activities

Activity 1: Scenario analysis to identify causes for unreliability

Objective: To have students analyze scenarios to identify possible causes for unreliability
of an English test.

Procedure: Divide students into small groups. Provide each group with a handout
containing 3-4 scenarios related to different types of reliability in language testing.

o Example Scenarios:
 Scenario 1 (Test-retest): "A student takes an online English
proficiency test. A week later, they retake the same test. Their scores
differ significantly. What factors might have contributed to this
difference?"
 Scenario 2 (Parallel Forms): "A school uses two different versions
of a reading comprehension test. Students who perform well on one
version consistently score poorly on the other. What could be the
potential issues with the test?"
 Scenario 3 (Internal Consistency): "A vocabulary test includes
items that seem unrelated to each other. How might this affect the
internal consistency of the test?"
 Scenario 4 (Inter-rater): "Two teachers score student essays. Their
scores for the same essays differ significantly. What steps can be
taken to improve inter-rater reliability?"

Each group analyzes the scenarios and discusses possible causes for unreliability.

Activity 2: Unpacking the reliability of a sample test

Objective: To have students analyze a sample English test for high school students in
Vietnam and identify potential threats to its reliability.

Material: A sample English test (e.g., a reading comprehension test with multiple-choice
questions)

Procedure:

Test presentation: Present students with a sample English test (or students search
from the Internet). Briefly discuss the purpose of the test and the target language
skills it aims to assess.

Group analysis : Divide students into small groups. Provide each group with a
handout containing the following questions:

 Test-retest: "What factors might affect the consistency of scores if


the same students took this test again in a few weeks?"
 Parallel forms: "If a parallel version of this test were created, what
would be important considerations to ensure the two versions are
truly equivalent?"
 Internal consistency: "How could you assess the internal
consistency of this test? What specific measures or analyses would
you use?"
 Inter-rater reliability: "If this test included a writing section, how
could you ensure consistent scoring across different raters?"
 Overall reliability: "Based on your analysis, what are the potential
threats to the reliability of this test?"

Activity 3: Analysis of language test validity


Objective: To have students examine an English test given by the teacher or one on
the Internet to analyze the validity of an English test for high school students in
Vietnam.

Procedure: Students work in pairs to discuss the following questions:

o Content: Does the test adequately cover the range of language skills taught
in the English curriculum?
o Construct: Does the test accurately measure the intended language skills
(e.g., reading comprehension, writing, speaking)?

Activity 4: The practicality puzzle: Designing a feasible test

Objective: To have students apply the principles of practicality (time, cost, logistics, ease
of administration/scoring, accessibility, interpretation/reporting) in the design of a short
language test.

Procedure: Present students with a scenario requiring the creation of a language test. For
example, "Design a short English test for incoming exchange students to assess their
basic conversational skills." Each group designs a short language test (e.g., 2-3 tasks)
based on the given scenario. Students discuss in groups to consider the following:

 Test format: Multiple-choice, short answer, role-play, etc.


 Test length: How long will the test take to administer?
 Scoring procedures: How will the test be scored? Are there clear
scoring rubrics?
 Logistics: Where will the test be administered? What resources are
needed?
 Accessibility: How will the test be made accessible to students with
disabilities?
 Reporting: How will test results be communicated to students and
other stakeholders?

Activity 5: The ethics of assessment: Navigating dilemmas in language testing

Objective: To have students analyze ethical dilemmas in language testing and develop
strategies for ensuring fair and equitable assessment practices.

Procedure: Divide students into small groups. Each group analyzes the following
questions and discusses potential threats to its reliability. Below is a list containing 3-4
ethical dilemmas related to language testing.
 Scenario 1 (Test bias): "A reading comprehension test
includes a passage about a specific cultural event that would
be unfamiliar to students from certain cultural backgrounds.
How does this create bias, and how can it be addressed?"
 Scenario 2 (Test security): "A student discovers a copy of
the upcoming English proficiency test online. What are the
ethical implications of this situation, and what steps should be
taken to address it?"
 Scenario 3 (Responsible use of test results): "A university
uses an English proficiency test as the sole criterion for
admission to its English program. What are the potential
ethical concerns with this approach?"
 Scenario 4 (Test-taker rights): "A student with a learning
disability requests extra time for an English test. How should
the school handle this request to ensure fairness and equity
for all students?"

Each group analyzes the scenarios, discussing the ethical implications and potential
solutions.

References

Bachman, L. F., & Palmer, A. S. (1996). Language testing in practice. Oxford University
Press.

Brown, H. D., & Abeywickrama, P. (2019). Language assessment: Principles and


classroom practices. Pearson.

Chapelle, C. A., & Lee, H. W. (2021). Conceptions of validity. In The Routledge


handbook of language testing (pp. 17-31). Routledge.

Dinh, M. T. (2019). A review on validating language tests. VNU Journal of Foreign


Studies, 35(1).

Hill, K., & McNamara, T. (2015). Validity inferences under high-stakes conditions: A
response from language testing. Measurement: Interdisciplinary Research &
Perspectives, 13(1), 39-43.

Hughes, A. (2020). Testing for language teachers. Cambridge university press.

Liu, Z., Li, T., & Diao, H. (2020). Analysis on the Reliability and Validity of Teachers'
Self-designed English Listening Test. Journal of Language Teaching and
Research, 11(5), 801-808.
Siddiek, A. G. (2010). The Impact of Test Content Validity on Language Teaching and
Learning. Online Submission, 6(12), 133-143.

Tang, W., Cui, Y., & Babenko, O. (2014). Internal consistency: Do we really know what
it is and how to assess it. Journal of Psychology and Behavioral Science, 2(2), 205-
220.
Chapter 4: Types of Language Tests

This chapter delves into the diverse landscape of language tests, examining the various
types employed to measure language skills and knowledge. To provide a structured
framework, the chapter will categorize these tests into two broad categories: classroom-
based language tests and standardized language tests. Classroom-based tests, often
developed and administered within specific educational contexts, reflect the immediate
needs of learners and instructors, emphasizing classroom-based assessments and teacher-
created materials. In contrast, standardized tests are externally developed, rigorously
validated, and designed to provide consistent and comparable measures of language
proficiency across diverse populations. Understanding the distinctions and applications of
these two categories is essential for educators, researchers, and test developers seeking to
make informed decisions about language assessment.

4.1 Classroom-based language tests

4.1.1 Discrete-point tests

Discrete-point language tests focus on isolating and measuring specific, individual


language skills or knowledge points. Unlike integrative tests that assess the use of
language in a more holistic and communicative manner, discrete-point tests aim to
pinpoint learners' strengths and weaknesses in particular areas of grammar, vocabulary,
or phonology (Brown, 2004).

Characteristics of discrete-point tests:

• Focus on isolated skills: These tests typically break down language into smaller
components, such as individual grammatical rules, vocabulary items, or
phonological features.
• Objective scoring: Discrete-point tests often lend themselves to objective scoring
methods, such as multiple-choice questions, true/false items, or matching
exercises. This makes scoring relatively quick and efficient.
• Emphasis on accuracy: The primary focus is on the accuracy of the learner's
response, rather than the overall communicative effectiveness.

Types of discrete-point tests:

• Grammar tests: These tests assess learners' knowledge of grammatical rules, such
as verb tenses, noun phrases, and sentence structure. Examples include fill-in-the-
blank exercises, error correction tasks, and multiple-choice questions that test
grammatical structures.
• Vocabulary tests: These tests measure learners' knowledge of individual words and
phrases. Common formats include multiple-choice vocabulary tests, matching
exercises, and cloze tests where learners fill in missing words in a text.
• Phonology tests: These tests assess learners' pronunciation, intonation, and other
aspects of spoken language. Examples include minimal pair discrimination tasks,
picture-cued elicitation tasks, and tests of stress and intonation.

Strengths of discrete-point tests:

• Ease of administration and scoring: Discrete-point tests are generally easy to


administer and score, particularly with the use of multiple-choice formats and
automated scoring systems.
• Diagnostic value: By isolating specific language skills, these tests can help
identify areas of strength and weakness for individual learners, which can inform
instruction and remediation.
• Efficiency: Discrete-point tests can be used to efficiently assess a wide range of
language skills in a relatively short amount of time.

Weaknesses of discrete-point tests:

• Limited reflection of real-world language use: By focusing on isolated language


skills, discrete-point tests may not accurately reflect how language is used in
authentic communication.
• Potential for overemphasis on isolated skills: An overemphasis on discrete-point
tests may lead to a narrow focus on isolated skills in the classroom, neglecting the
development of communicative competence.
• Limited assessment of higher-order skills: Discrete-point tests may not effectively
assess higher-order skills such as critical thinking, problem-solving, and strategic
language use.

Discrete-point tests have their place in language assessment, particularly for diagnosing
specific areas of difficulty and providing quick feedback on the mastery of particular
language skills. However, it is crucial to recognize their limitations and to use them in
conjunction with other assessment methods that more comprehensively assess language
proficiency in communicative contexts.

4.1.2 Integrative tests

Integrative tests move beyond the isolated assessment of discrete language skills
(Alderson et al. 2014), such as grammar or vocabulary, and instead focus on assessing
how learners use language in a more holistic and communicative manner. These tests aim
to simulate real-life language use, requiring learners to integrate multiple language skills
to achieve a communicative goal (Brown, 2004; Weir, 2005).
Characteristics of integrative tests:

• Focus on communicative competence: Integrative tests emphasize the ability to


use language effectively and appropriately in real-world contexts. They assess not
only grammatical accuracy but also fluency, coherence, and the ability to convey
meaning effectively.
• Emphasis on authentic tasks: Tasks in integrative tests often simulate real-life
communication situations, such as writing emails, giving presentations,
participating in role-plays, or engaging in discussions.
• Integration of multiple skills: Many integrative tests require learners to use
multiple language skills simultaneously, such as reading a text and then
summarizing it in writing or listening to a conversation and then answering
questions about it.

Types of integrative tests:

• Reading comprehension tests with complex tasks: These tests go beyond simple
multiple-choice questions and require learners to analyze, synthesize, and evaluate
information from longer texts. Tasks may include summarizing, paraphrasing,
comparing and contrasting different texts, and answering open-ended questions
that require critical thinking and interpretation.
• Writing tasks: These tasks may include essay writing, report writing, letter writing,
or creative writing. They assess learners' ability to organize ideas, develop
arguments, use appropriate language, and produce coherent and effective written
communication.
• Speaking tests: These tests often involve interactive communication tasks, such as
role-plays, discussions, and interviews. They assess learners' fluency, accuracy,
pronunciation, and ability to interact effectively with others.
• Listening comprehension tests: These tests may involve listening to lectures,
conversations, or other authentic audio materials and then answering questions
that require learners to understand and interpret the information.

Strengths of integrative tests:

• Higher ecological validity: Integrative tests more closely reflect real-world


language use, providing a more authentic assessment of learners' communicative
competence.
• More engaging for learners: Authentic tasks can be more engaging and motivating
for learners than isolated skill-based exercises.
• Assessment of communicative competence: Integrative tests provide a more
comprehensive assessment of learners' ability to use language effectively and
appropriately in real-world contexts.
Weaknesses of integrative tests:

• Difficulty in scoring: Scoring integrative tests can be more subjective and time-
consuming than scoring discrete-point tests.
• Potential for rater bias: Subjective scoring can be influenced by rater bias, leading
to inconsistencies in scoring.
• Difficulty in standardizing: It can be challenging to standardize the administration
and scoring of integrative tests, which can make it difficult to ensure fairness and
consistency across different administrations.

In general, integrative language tests play a crucial role in providing a more holistic and
authentic assessment of learners' language proficiency. While they may present some
challenges in terms of administration and scoring, their value in assessing communicative
competence makes them an essential component of any comprehensive language
assessment program.

4.1.3 Diagnostic tests

Diagnostic tests play a crucial role in effective language teaching and learning. Unlike
summative assessments that primarily focus on measuring overall achievement,
diagnostic tests aim to pinpoint specific areas of language difficulty for individual
learners (Joughin, 2007). This information allows teachers to tailor instruction, provide
targeted support, and create a more personalized learning experience for each student.

Characteristics of diagnostic tests:

• Focus on identifying areas of weakness: Diagnostic tests are designed to identify


specific language skills or knowledge areas where learners are experiencing
difficulties. These areas can include grammar rules, vocabulary, pronunciation,
reading comprehension, or any other aspect of language learning.
• In-depth analysis: Diagnostic tests often involve a more in-depth analysis of
learner performance than broader proficiency tests. They may include detailed
error analysis, individual interviews, or other methods to pinpoint specific areas of
difficulty.
• Formative in nature: Diagnostic tests are primarily used for formative assessment,
providing information that can be used to inform instruction and guide subsequent
learning.

Types of diagnostic tests:

• Pre-tests: These tests are administered at the beginning of a course to assess


learners' existing knowledge and skills. They help teachers identify areas where
students may need additional support or where the curriculum can be adjusted to
meet their needs.
• Interviews: One-on-one interviews can provide valuable insights into learners'
strengths and weaknesses in areas such as fluency, accuracy, and pronunciation.

Uses of diagnostic tests:

• Personalized instruction: Diagnostic tests help teachers to tailor instruction to the


specific needs of individual learners. By identifying areas of difficulty, teachers
can provide targeted support and remediation.
• Grouping students: Diagnostic tests can be used to group students with similar
needs, allowing for more effective and efficient instruction.
• Identifying learning styles: Some diagnostic tests can help to identify learners'
preferred learning styles, which can inform teaching methods and materials.
• Monitoring progress: Diagnostic tests can be administered periodically throughout
a course to track learners' progress and identify any emerging difficulties.

Overall, diagnostic tests play a vital role in effective language teaching and learning. By
providing valuable insights into learners' strengths and weaknesses, diagnostic tests
enable teachers to provide more effective and personalized instruction, ultimately leading
to improved learning outcomes.

4.1.4 Progress tests

Progress tests are a crucial component of effective language teaching and learning. They
provide valuable feedback on students' learning and help teachers monitor the
effectiveness of their instruction. Unlike summative assessments, which typically occur at
the end of a course or unit, progress tests are administered periodically throughout the
learning process to track students' progress and identify areas for improvement.

Characteristics of progress tests:

• Formative in nature: Progress tests are primarily used for formative assessment,
providing ongoing feedback to both teachers and learners.
• Aligned with learning objectives: They are designed to assess specific learning
objectives covered in the course or unit.
• Regular administration: Progress tests are typically administered at regular
intervals, such as at the end of each week, month, or unit.
• Short and focused: They are usually shorter than summative tests and focus on a
specific set of skills or knowledge areas.
• Actionable feedback: The results of progress tests should provide specific and
actionable feedback to both teachers and learners, identifying areas of strength and
weakness.
Types of progress tests:

• Short quizzes: Short quizzes can be used to assess students' understanding of


specific grammar rules, vocabulary items, or other discrete points.
• Classroom activities: Certain classroom activities, such as short presentations,
role-plays, or group discussions, can be adapted to serve as informal progress
tests.
• Short writing assignments: Short writing tasks, such as paragraph writing or essay
outlines, can be used to assess students' writing skills.
• Reading comprehension exercises: Short reading passages with comprehension
questions can be used to assess reading skills and vocabulary.

Benefits of using progress tests:

• Improved learning: Regular feedback from progress tests can help students
identify areas for improvement and adjust their learning strategies accordingly.
• Informed instruction: Progress test results provide valuable information to
teachers, allowing them to identify areas where students may be struggling and to
adjust their teaching methods accordingly.
• Increased motivation: Regular feedback can help to motivate students by
demonstrating their progress and celebrating their achievements.
• Early identification of difficulties: Progress tests can help to identify potential
learning difficulties early on, allowing teachers to provide targeted support and
prevent students from falling behind.

Ethical considerations:

• Use of results: Progress test results should be used constructively and should not
be used for high-stakes decisions, such as grading or placement.
• Feedback: Feedback from progress tests should be provided in a timely and
constructive manner.
• Confidentiality: Student performance on progress tests should be treated with
confidentiality.

Generally, progress tests play a vital role in supporting effective language learning. By
providing regular feedback and identifying areas for improvement, they help both
teachers and learners to monitor progress, adjust instruction, and ensure that students are
on track to achieve their learning goals.

4.1.5 Achievement tests

Achievement tests are designed to measure how well learners have mastered specific
language skills or knowledge that have been taught within a particular course or
curriculum. Unlike proficiency tests, which assess overall language ability, achievement
tests focus on evaluating learners' progress in relation to specific learning objectives
(Brown, 2004; Joughin, 2007).

Characteristics of achievement tests:

• Curriculum-based: Achievement tests are typically closely aligned with the


specific content and objectives of a particular language course or curriculum. They
cover the specific language skills and knowledge that have been taught to the
learners.
• Formative and summative purposes: While primarily used for summative
assessment at the end of a course or unit, achievement tests can also provide
valuable formative feedback to both teachers and learners.
• Varied formats: Achievement tests can take various formats, including multiple-
choice questions, short answer questions, essay writing, oral presentations, and
listening comprehension tasks.
• Focus on specific learning objectives: Achievement tests are designed to assess
specific learning outcomes, such as:
o Knowledge of grammar rules: Accurately using verb tenses, applying
grammatical rules in context.
o Vocabulary acquisition: Understanding and using vocabulary related to
specific themes or topics.
o Reading comprehension: Understanding main ideas, supporting details, and
inferences in texts.
o Writing skills: Producing clear, coherent, and grammatically correct written
texts.
o Speaking skills: Communicating effectively in spoken language, including
fluency, accuracy, and interaction.
o Listening comprehension: Understanding spoken language in various
contexts, such as lectures, conversations, and presentations.

Importance of achievement tests:

• Monitoring student progress: Achievement tests provide valuable information


about students' learning progress and identify areas where they may need
additional support.
• Informing instruction: The results of achievement tests can help teachers identify
areas where instruction needs to be adjusted or supplemented.
• Providing feedback to learners: Achievement tests provide students with feedback
on their learning, helping them to identify their strengths and weaknesses and set
goals for improvement.
• Evaluating the effectiveness of teaching: Achievement tests can help teachers
evaluate the effectiveness of their teaching methods and curriculum.
Ethical considerations:

• Fairness and equity: Achievement tests should be fair and equitable for all
learners, regardless of their background or learning style.
• Use of results: Achievement test results should be used responsibly and ethically,
primarily for formative purposes and to provide constructive feedback to learners.
• Test security: Maintaining the security of achievement tests is crucial to ensure the
integrity of the assessment process.

In summary, achievement tests play a vital role in language education by providing


valuable information about student learning and informing instructional decisions. By
carefully designing and administering achievement tests, teachers can effectively monitor
student progress, identify areas for improvement, and ultimately enhance the learning
experience for all students.

4.2 Standardized language tests

Standardized language tests are formal assessments with standardized procedures and
scoring criteria. They are widely used for various purposes, including university
admissions, immigration, professional certification, and program placement. These tests
typically involve a large-scale administration to a large number of test-takers, ensuring
consistency across different testing locations and administrations.

Characteristics of standardized language tests:

• Standardized procedures:
o Consistent administration protocols across all testing locations.
o Clear and unambiguous instructions for test-takers.
o Controlled testing environment to minimize distractions.
• Objective scoring:
o Use of scoring rubrics and standardized scoring keys to minimize rater
bias.
o Often involve machine-scoring or computer-assisted scoring.
• Large-scale administration:
o Designed to be administered to a large number of test-takers
simultaneously.
o Efficient and cost-effective for large-scale testing programs.
• High-stakes implications:
o Often used for high-stakes decisions, such as university admissions or
immigration applications.
o Scores have significant implications for test-takers' future opportunities.

Types of standardized language tests:


• Proficiency tests:
o Assess overall language proficiency across different skills (reading, writing,
listening, speaking).
o Examples: TOEFL (Test of English as a Foreign Language), IELTS
(International English Language Testing System), Cambridge English
exams (e.g., B2 First, C1 Advanced).
• Achievement tests:
o Measure language skills acquired in a specific educational program or
curriculum.
o Often used for program evaluation and to assess student learning outcomes.
• Diagnostic tests:
o Identify specific areas of language difficulty for individual learners.
o Used for placement, remediation, and individualized instruction.

Strengths of standardized language tests:

• Objectivity: Standardized scoring minimizes rater bias and ensures consistency


across different test-takers.
• Comparability: Allows for meaningful comparisons of test scores across different
individuals and institutions.
• Efficiency: Can be administered and scored efficiently for large groups of test-
takers.
• Reliability: Standardized procedures and scoring contribute to higher test
reliability.

Weaknesses of standardized language tests:

• Limited flexibility: May not adequately assess the full range of language skills and
abilities.
• Potential for bias: May not be culturally fair or sensitive to the needs of all test-
takers.
• Limited individualization: May not account for individual learning styles and
needs.
• High-stakes nature: Can create undue pressure and anxiety for test-takers.

Ethical considerations

• Test security: Maintaining test security is crucial to ensure the integrity of the
testing process.
• Fairness and equity: Standardized tests must be fair and equitable for all test-
takers, regardless of their background, culture, or learning style.
• Appropriate use of scores: Test scores should be used responsibly and ethically,
and should not be the sole determinant of important decisions.
Overall, standardized language tests play an important role in various contexts, from
education and immigration to professional certification. While they offer several
advantages, such as objectivity and comparability, it is crucial to recognize their
limitations and to use them responsibly and ethically. By carefully considering the
strengths and weaknesses of standardized tests and using them in conjunction with other
assessment methods, we can ensure that they provide a fair and accurate assessment of
language proficiency.

4.3 Student self-assessment

Instruction: Carefully read each statement in the table below. For each statement, rate
your understanding using the following scale:

1: I do not understand this at all.


2: I understand this a little, but I need more help.
3: I understand this fairly well, but I have some questions.
4: I understand this very well and can explain it to others.

Learning Objective/Concept Self- Evidence/Notes Action Plan


Assessment (Explain your (What will
Rating (1-4) rating) you do to
improve?)
Test categories: I can differentiate
between classroom-based and
standardized language tests,
explaining their purposes and uses.
Discrete-point Tests: I can define
and describe the characteristics of
discrete-point tests, providing
examples of test items.
Integrative tests: I can define and
describe the characteristics of
integrative tests, providing examples
of test items (e.g., essays, oral
interviews).
Other classroom-based tests: I can
explain other types of classroom-
based tests, such as formative,
summative, diagnostic, and
placement tests.
Standardized tests: I can discuss
the purpose, characteristics, and
examples of standardized language
tests (e.g., TOEFL, IELTS).
Overall understanding: I feel
confident in my understanding of the
various types of language tests and
their applications in language
assessment.

4.4 Consolidation activities

Activity 1: Designing a discrete-point grammar test

Objectives:

• To enhance students' understanding of discrete-point testing principles by having


them design their own test.
• To develop students' skills in creating clear, concise, and objective test items.

Procedure: Students briefly review the key characteristics of discrete-point tests,


emphasizing their focus on isolating and assessing specific language skills (e.g.,
grammar, vocabulary). Students work in pairs/groups to design a short discrete-point test
(5-10 items) to assess learners’ understanding of an assigned grammatical structure (e.g.,
present perfect tense, articles, passive voice). Students are required to use a variety of
item types (e.g., multiple-choice, fill-in-the-blank, true/false).

Activity 2: Designing a reading comprehension progress test

Objectives:

• To enhance students' understanding of progress tests by having them design a


short test for a specific language skill (reading comprehension).
• To develop students' skills in creating clear, concise, and effective reading
comprehension test items.

Procedure:

• Students in groups briefly review the key characteristics of progress tests. They
then choose a short reading passage suitable for high school students
(approximately 200-300 words) from an English textbook for high schools in
Vietnam and design some reading comprehension test items (e.g., multiple-choice
questions, true/false statements, short answer questions)
Activity 3: Designing an achievement test for Vietnamese high school students

Objectives:

• To enhance students' understanding of achievement tests by having them design a


sample test for high school English learners in Vietnam.
• To develop students' skills in creating test items that align with the Vietnamese
national curriculum and address the specific learning objectives of high school
English education.

Procedure: Students work in groups briefly review the key characteristics of achievement
tests, emphasizing their purpose in measuring student learning in relation to specific
curriculum objectives. Each group designs an outline for a short achievement test. The
outline should include:

§ Test sections: Determine the number and type of sections (e.g.,


reading, writing, listening, grammar and vocabulary).
§ Specific learning objectives: Identify 2-3 specific learning objectives
for each section (e.g., "Students will be able to read and understand a
short passage and answer questions about main ideas and supporting
details.").
§ Sample test items: Create 2-3 sample test items for each section,
aligning them with the chosen learning objectives.

Activity 4: Designing a portfolio project for Vietnamese high school students

Objectives:

• To enhance students' understanding of alternative assessment methods by having


them design a portfolio project for high school English students in Vietnam.
• To develop students' skills in creating authentic and meaningful assessment tasks
that align with the Vietnamese National Curriculum (Chương trình giáo dục phổ
thông, MOET, 2018).

Procedure: Students discuss in groups the guidelines for designing portfolio projects (e.g.,
focus on learning objectives, variety of assessment methods, student reflection) and
relevant sections of the Vietnamese National Curriculum for English (e.g. Chương trình
giáo dục phổ thông, MOET, 2018). Students must include a variety of assessment
methods (e.g., essays, presentations, creative writing, reflective journals), clarity of the
assessment criteria and scoring rubrics.
Activity 5: Designing a role-play for communication skills

Objectives:

• To enhance students' understanding of performance-based assessment by having


them design a role-play scenario.
• To develop students' skills in creating authentic and engaging communication
tasks that assess real-world language use.

Procedure: Students in groups briefly review the key characteristics of performance-


based assessment, emphasizing authenticity, real-world application, and the assessment
of communicative competence. Each group will choose a specific theme or context
relevant to the lives of Vietnamese high school students (e.g., ordering food at a
restaurant, making travel arrangements, asking for directions, discussing household
chores) to design a role-play scenario that assesses the following communication skills.
They then discuss the assessment criteria for the role-play (e.g Fluency: Ability to speak
naturally and smoothly and accuracy: Grammatical accuracy and appropriate use of
vocabulary, etc.)

References:

Alderson, J. C., Clapham, C., & Wall, D. (2014). Language test construction and
evaluation. Cambridge University Press.

Brown, H. D. (2004). Language assessment: Principles and classroom practices (2nd


ed.). Longman.

Joughin, G. (2007). Language testing and assessment: An advanced resource book.


Routledge.

MOET (2018). Thông tư ban hành giáo dục phổ thông. Available at:
[Link]
[Link]

Weir, C. J. (2005). Language testing and assessment. Blackwell.


Chapter 5: Language Test Design and Development

Crafting a language test requires careful planning, meticulous execution, and a deep
understanding of both the material (language) and the purpose of the creation
(assessment). This chapter, delves into the intricate art of language test construction,
guiding you through the process of transforming theoretical knowledge into practical
assessment tools.

5.1 Designing a language test

Designing a language test requires planning, systematic design, and quality control. This
section will present the key steps involved in developing an effective language test, from
defining its purpose and objectives to administering it in a standardized manner. The
following these steps can ensure that a language test is valid, reliable, and fair to test-
takers.

Figure 5.1 Steps for designing a language test

Deciding the Developing


Deciding a Writing the Reviewing the Administering
test's the test's
test's purpose test items test items the test
objectives specification

• Deciding a test’s purpose: Clearly establish the reason for the test (e.g., placement,
proficiency, achievement). Identify the target test-takers, including their age,
background, and language learning experience.
• Deciding the test’s objectives: Determine the specific language skills and
knowledge that the test should assess.
• Designing the test’s specification: Create a detailed blueprint of the test, outlining
the content, format, and scoring criteria. Specify the language skills to be tested,
the types of test items to be used, and the weighting of different sections.
• Writing the test items: Create test items that align with the test specifications and
accurately measure the intended language skills. Ensure that the items are clear,
unambiguous, and appropriate for the target population.
• Reviewing and moderating test items: Have experts review the test items for
clarity, accuracy, and fairness. Revise or discard items that are problematic.
• Administer the test in a standardized and controlled environment: Ensure that all
test-takers have equal opportunities to demonstrate their language
skills/competence.

5.2 Designing objective test items

Objective test items, with their clear-cut right or wrong answers, offer a world of
efficiency and objectivity in language assessment (Hughes, 2020). This section delves
into the realm of these structured assessment tools, exploring their diverse forms and
applications. From the familiar multiple-choice questions to true-false, matching, and
gap-fill items, the following section will unravel the characteristics that make them a
staple in language testing.

5.2.1 Designing multiple-choice test items

Multiple-choice questions are a common and versatile format used in various language
tests, from classroom quizzes to high-stakes standardized examinations. They offer
several advantages in terms of objectivity, efficiency, and ease of scoring, making them a
popular choice for assessing knowledge and skills in a wide range of language areas,
including grammar, vocabulary, reading comprehension, and listening comprehension
(Currie & Chiramanee, 2010).

Characteristics of multiple-choice items:

• Stem: The initial part of the question, which presents the problem or situation to
be addressed.
• Options: A set of possible answers, usually consisting of three or four choices.
• Key: The correct answer among the options.
• Distractors: The incorrect answer options, designed to be plausible but incorrect.

The following diagram illustrates the components of a multiple choice test item

Figure 5.2 Key components of a multiple-choice question

1. Stem: presents the problem


2. Keyed response: correct/best answer
3. Distracters: appear to be reasonable answers to the examinee who does
not know the content
4. Options: include the distracters and the keyed response.

Below are examples of multiple choice items from the English textbook Ilearn Smart
World 9

Source: Vo et al., (2024) Tieng Anh 9-I-learn Smart World

Faulty MC items

When designing MC test items it is essential to avoid faulty MC items. Below are some
examples of faulty MC items:

Ambiguous question
Question: "What is the meaning of the word 'bank' in this sentence: 'I went to the
bank to deposit my paycheck.'"
(A) A financial institution
(B) The side of a river
(C) A row of seats
Fault: The word "bank" has multiple meanings, and the sentence does not provide
enough context to determine the intended meaning. This makes the question
ambiguous and potentially confusing for students.
More than one correct answer
Question: Which of the following are topics covered in Unit 3 of the English 10 textbook
in Vietnam?
(A) Environmental pollution
(B) Natural disasters
(C) Protecting endangered species
Fault: Environmental pollution, natural disasters, and Protecting endangered species are
often interconnected and could be addressed together.

Unclear instructions
Question: Read the following passage and...
(A) Identify the main idea.
(B) Find the supporting details.
(C) Determine the author's purpose.
Fault: The instructions are too broad. Students need more specific guidance on what they
are supposed to do with the passage.

Misleading distractors
Question: What is the correct spelling of the word?
(A) Accommodate
(B) Acommodate
(C) Accomodate
Fault: The distractors are too similar to the correct answer, making it difficult to identify
the correct spelling. This tests students' ability to spot minor differences rather than their
actual knowledge of spelling rules.

Culturally biased question


Question: Which of these is a popular American holiday?
(A) Thanksgiving
(B) Diwali
(C) Ramadan
Fault: This question assumes knowledge of American culture, which may not be relevant
or fair to all students, especially in an EFL context.

Grammatically incorrect stem


Question: Which sentence is using the correct tense?
(A) I am go to the store.
(B) He have went to the library.
(C) She will going to the park.
Fault: The question stem itself contains grammatical errors, which can confuse students
and make it difficult to focus on the actual content of the question.
Overall, when creating multiple-choice questions, it is important to be clear, concise, and
avoid ambiguity. The goal is to assess students' language skills accurately, not to trick
them.

Advantages of multiple-choice items:

• Objectivity: Scoring is typically objective and straightforward, minimizing rater


bias.
• Efficiency: Can be administered and scored quickly and efficiently, especially
with the use of automated scoring systems.
• Versatility: Can be used to assess a wide range of language skills and knowledge
areas.
• Reliability: Multiple-choice tests generally exhibit high levels of reliability,
meaning they tend to produce consistent results across different administrations
(Thanyapa & Currie, 2014).

Designing effective multiple-choice items:

• Clear and concise stem: The stem should be clear, concise, and unambiguous. It
should present the problem or situation to be addressed in a straightforward
manner.
• Plausible distractors: Distractors should be plausible but incorrect. They should be
grammatically correct and relevant to the stem, but not the correct answer.
• Grammatically correct options: All options, including the key and the distractors,
should be grammatically correct.
• Avoid ambiguity: The stem and options should be free from ambiguity and double
negatives.
• Test a single skill: Each item should focus on a single skill or concept.

Limitations of multiple-choice items:

• Limited assessment of higher-order thinking skills: Multiple-choice questions may


not effectively assess higher-order thinking skills such as critical thinking,
problem-solving, and creative expression.
• Guessing: Students may be able to guess the correct answer, which can impact test
scores.
• Limited ability to assess writing and speaking skills: Multiple-choice formats are
less suitable for assessing complex skills such as writing and speaking.

In general, multiple-choice items are a valuable tool for language assessment, offering
several advantages in terms of objectivity, efficiency, and ease of scoring. However, it is
crucial to recognize their limitations and use them appropriately in conjunction with other
assessment methods that more comprehensively assess language proficiency. When
designing multiple-choice items, careful attention should be paid to clarity, plausibility,
and the overall quality of the distractors to ensure that the items are effective and reliable.

5.2.2 Designing true-false-not given test items

True-False-Not Given (T/F/NG) questions, also known as True-False-Cannot Say, are a


common item type in language proficiency tests, particularly in reading comprehension
assessments. They require test-takers to evaluate the veracity of a statement based on a
provided text, determining whether the information is explicitly stated (True),
contradicted (False), or not mentioned at all (Not Given). This chapter delves into the
characteristics, applications, advantages, and limitations of T/F/NG items in language
testing.

Understanding T/F/NG items

T/F/NG items assess a test-taker's ability to:

• Comprehend explicitly stated information: Identify key details and understand the
literal meaning of the text.
• Draw inferences: Make logical deductions based on the information provided.
• Identify the scope of the text: Recognize what information is included and,
crucially, what is not.

These items typically consist of a reading passage followed by a series of statements.


Test-takers must carefully analyze each statement and compare it to the information
presented in the text to determine the correct response.

Applications in language testing

T/F/NG items are widely used in various language tests, including:

• Academic proficiency tests: Such as IELTS, TOEFL, and Cambridge exams,


where they assess reading comprehension skills crucial for academic success.
• General proficiency tests: Used in tests like TOEIC to evaluate comprehension
abilities relevant to everyday situations and workplace communication.
• Placement tests: Help determine a learner's language level and assign them to
appropriate classes.

Advantages of T/F/NG items

• Ease of construction and scoring: Relatively straightforward to develop and can be


scored objectively.
• Efficient assessment: Allow for the assessment of a wide range of information
within a single passage.
• Focus on detailed comprehension: Encourage close reading and attention to
specific information.

Limitations of T/F/NG items

• Ambiguity: Poorly written statements can be open to interpretation, leading to


confusion and unreliable results.
• Limited scope: Primarily assess comprehension of factual information and may
not adequately evaluate inferential or critical thinking skills.
• Potential for test-wiseness: Test-takers can sometimes use strategies to guess the
correct answer even without fully understanding the text.

Constructing effective T/F/NG items

To maximize the effectiveness and reliability of T/F/NG items, test developers should
consider the following:

• Clear and concise statements: Avoid complex language and ensure each statement
focuses on a single idea.
• Unambiguous wording: Statements should have only one possible interpretation
based on the text.
• Variety of difficulty levels: Include a range of items that assess both explicit and
implicit understanding.
• Avoid verbatim copying: Rephrase information from the text to prevent simple
matching exercises.
• Thorough review and piloting: Ensure items are accurate, clear, and free of bias
before including them in a test.

Faulty true-false-not given items

Below are some examples of faulty true-false-not given items

Vague statement with cultural bias

Statement: "Most Vietnamese people prefer traditional music to modern music."


(T/F/NG)

Fault: This statement is too general and relies on a potentially inaccurate cultural
stereotype. Musical preferences vary greatly among individuals, and there is no
definitive answer without specific data or context.

Statement with trivial detail


Text: "The story describes a young girl named Lan who lives in a small village in
the Mekong Delta. She loves to help her parents with their rice farm and dreams of
becoming a doctor."

Statement: "Lan's favorite animal is a water buffalo." (T/F/NG)

Fault: This detail might be mentioned in the text, but it's not crucial to the main
storyline or character development. Focusing on such trivial information does not
effectively assess comprehension of the key themes or messages.

Ambiguous statement with negatives

Statement: "It is not impossible to learn English fluently with dedicated practice."
(T/F/NG)

Fault: The double negative ("not impossible") makes this statement unnecessarily
convoluted. It's better to rephrase it in a more straightforward way (e.g., "It is
possible to learn English fluently with dedicated practice.")

Overall, T/F/NG items are a valuable tool in language testing, providing an efficient and
objective means of assessing reading comprehension. By understanding their
characteristics, applications, and potential limitations, test developers can effectively
utilize this item type to create reliable and valid language assessments. However, careful
attention must be paid to item construction and piloting to ensure clarity, avoid
ambiguity, and promote accurate measurement of language proficiency.

5.2.3 Designing cloze test items

Cloze tests are a versatile and widely used tool in language education. They involve
presenting learners with a text where certain words have been systematically removed
(typically every fifth, sixth, or seventh word) and replaced with blanks. Learners are then
tasked with filling in these blanks with appropriate words based on their understanding of
the context and their language knowledge. This chapter explores the nature of cloze tests,
their applications in language learning, their advantages and limitations, and best
practices for constructing and implementing them effectively.

Understanding cloze tests

Cloze tests assess various aspects of language proficiency, including:

• Reading comprehension: Understanding the overall meaning and flow of a text.


• Vocabulary knowledge: Selecting appropriate words that fit the context.
• Grammatical awareness: Choosing words that are grammatically correct within the
sentence structure.
• Discourse competence: Recognizing cohesive devices and understanding how
sentences connect to form a coherent text.

By requiring learners to actively engage with the text and make informed choices about
missing words, cloze tests encourage deep processing of language and promote both
receptive and productive skills.

Applications in language education

Cloze tests can be used for a variety of purposes in language teaching and assessment:

• Assessing language proficiency: Measuring overall language ability or specific


skills like vocabulary or grammar.
• Diagnosing learner needs: Identifying areas of strength and weakness in language
comprehension and production.
• Developing language skills: Enhancing vocabulary acquisition, grammatical
awareness, and reading comprehension through practice and feedback.
• Promoting learner autonomy: Encouraging learners to take responsibility for
their learning by actively engaging with texts and making choices.

Types of cloze tests

• Fixed-ratio deletion: Words are removed at regular intervals (e.g., every fifth
word). This is the most common type of cloze test.
• Variable-ratio deletion: Words are removed based on specific criteria, such as
targeting particular grammatical structures or vocabulary items.
• Rational deletion: Words are removed based on their importance for understanding
the text, creating a more challenging and nuanced assessment.
• C-test: The second half of every second word is deleted, requiring learners to
complete the words based on the remaining letters.

Advantages of cloze tests

• Versatility: Can be adapted to assess different language skills and proficiency


levels.
• Objectivity: Scoring can be relatively straightforward, especially with fixed-ratio
deletion.
• Efficiency: Can assess a wide range of language knowledge in a relatively short
time.
• Authenticity: Can be based on authentic texts, increasing engagement and
relevance for learners.

Limitations of cloze tests

• Potential for ambiguity: Some blanks may have multiple possible answers,
making scoring subjective.
• Limited scope: May not fully capture complex aspects of language use, such as
pragmatic understanding or communicative competence.
• Artificiality: The artificial nature of deleting words can sometimes disrupt the
natural flow of the text.

Constructing effective cloze tests

• Choose appropriate texts: Select texts that are relevant to learners' interests and
proficiency levels.
• Determine deletion rate: Consider the difficulty level desired and the specific
skills being assessed.
• Provide clear instructions: Ensure learners understand the task and how to
respond.
• Pilot test the cloze: Administer the test to a small group of learners to identify any
ambiguities or issues.
• Provide feedback: Use the cloze test as a learning opportunity by providing
feedback on learners' responses and discussing the rationale behind correct
answers.

Faulty items in a cloze test

Below are the sample faulty items in a cloze test that should be avoided.

Faulty items in a cloze test

Missing word with multiple possibilities

Sentence: "The students were excited to go on a field trip to the ______."

Fault: This sentence lacks sufficient context to determine the missing word. It could
be "museum," "zoo," "park," "historical site," or any other place that students might
visit on a field trip. This makes the item ambiguous and potentially frustrating for
test-takers.

Gaps too close together

Sentence: "My friend and I went to the ____ to ____ a movie."


Fault: Having two gaps so close together makes it difficult for students to focus on
each individual word and its grammatical function. It increases the cognitive load
and might lead to guessing rather than demonstrating actual language proficiency.

Gaps that disrupt sentence flow

Sentence: "Although it was raining ____, we decided to go for a walk in the park."

Fault: Placing the gap in the middle of the adverbial clause disrupts the natural flow
of the sentence. It makes it harder for students to understand the sentence structure
and choose the correct word (e.g., "heavily," "outside," etc.).

Gaps requiring specialized knowledge

Sentence: "The ______ is a traditional Vietnamese musical instrument."

Fault: This item assumes knowledge of Vietnamese culture that might not be
relevant to all students, especially in an EFL context. It's better to choose words or
concepts that are more universally known or covered within the textbook's scope.

Gaps with grammatically incorrect options

Sentence: "She ______ her homework every day after school."

Options: (A) do, (B) does, (C) doing, (D) done

Fault: Including grammatically incorrect options can confuse students and hinder
their ability to identify the correct form of the verb (in this case, "does").

In short, cloze tests are a valuable tool in language education, offering a flexible and
efficient means of assessing and developing various language skills. By carefully
considering the principles of cloze test construction and implementation, educators can
harness their potential to enhance language learning and promote learner engagement and
autonomy.

5.2.4 Designing matching items

Matching items are a common and effective question type found in English language
tests for high school students in Vietnam. They require students to connect related pieces
of information from two separate columns, testing their ability to recognize relationships,
analyze information, and make accurate connections. This chapter explores the
characteristics, applications, advantages, and limitations of matching items in English
tests, along with guidelines for constructing effective matching activities.
Understanding matching items

Matching items typically consist of two columns:

• Premise Column: Contains a list of items, such as definitions, descriptions,


questions, or sentence halves.
• Response Column: Contains a corresponding list of items that need to be matched
to the items in the premise column, such as vocabulary words, phrases, answers, or
sentence completions.

Students are tasked with identifying the correct relationships between the items in the two
columns and indicating the matches. This process assesses their ability to:

• Comprehend and analyze information: Understand the meaning and context of


items in both columns.
• Identify relationships: Recognize connections, patterns, and associations
between items.
• Apply knowledge: Utilize their vocabulary, grammar, and reading comprehension
skills to make accurate matches.

Applications in English tests

Matching items are versatile and can be used to assess various aspects of English
language proficiency:

• Vocabulary: Matching words to definitions, synonyms, antonyms, or pictures.


• Grammar: Matching sentence halves, verb forms to tenses, or pronouns to
antecedents.
• Reading comprehension: Matching headings to paragraphs, characters to
descriptions, or events to timelines.
• Functional language: Matching phrases to situations, requests to responses, or
questions to answers.

Advantages of matching items

• Efficiency: Can assess a wide range of knowledge and skills in a concise format.
• Objectivity: Scoring is straightforward and less prone to subjective interpretation.
• Clarity: The format is generally easy for students to understand and follow.
• Versatility: Can be adapted to assess different language skills and levels of
difficulty.
Limitations of matching items

• Limited cognitive demand: May primarily test recognition and recall rather than
higher-order thinking skills.
• Potential for guessing: Students may resort to guessing if they are unsure of the
correct answers, especially if the number of items in each column is equal.
• Difficulty in constructing effective items: Creating meaningful and unambiguous
matches can be challenging.

Constructing effective matching items

To maximize the effectiveness and validity of matching items, test developers should
consider the following:

• Clear and concise instructions: Provide specific directions on how to complete


the matching task.
• Homogeneous content: Ensure all items within a matching set are related to a
common theme or topic.
• Unequal number of items: Include more options in the response column to
reduce the chance of guessing.
• Plausible distractors: Include response options that are similar to the correct
answers to increase the difficulty level.
• Logical arrangement: Arrange items in a clear and organized manner, such as
alphabetically or chronologically.
• Thorough review and piloting: Ensure items are accurate, clear, and free of bias
before including them in a test.

Faulty matching items

Ambiguous matching with multiple possibilities

Instructions: Match the words in Column A with their synonyms in Column B

Column A Column B
a) happy 1. big
b) large 2. joyful
c) sad 3. unhappy
d) angry 4. furious
e) small 5. tiny

Fault: Some words in Column A could have multiple synonyms in Column B. For
example, "happy" could be matched with both "joyful" and "glad" (if it were an
option), while "large" could be matched with both "big" and "huge" (if it were an
option). This ambiguity makes it difficult for students to determine the single best
match.

Irrelevant or mismatched items

Instructions: Match the vocabulary words with their definitions.

Column A Column B
a) photosynthesis 1. the process of making food in plants using sunlight
b) gravity 2. a type of animal that lives in the ocean
c) whale 3. a force that pulls objects towards each other

Fault: The item "whale" and its definition "a type of animal that lives in the ocean"
are not relevant to the other vocabulary words and definitions, which are related to
science and social studies. This creates a mismatch and can distract students.

Grammatically incorrect items

Instructions: Match the sentences in Column A with the correct question tags in
Column B.

Column A Column B
a) She is a doctor, 1. isn't she?
b) They are playing football, 2. aren't they?
c) He has finished his work, 3. hasn't he?
d) We will go to the party, 4. won't we?

Fault: One or more of the question tags in Column B might be grammatically


incorrect. For example, the question tag for "She is a doctor" should be "isn't she?"
not "doesn't she?" This tests students' ability to identify grammatical errors rather
than their ability to match question tags.

Overall, matching items are a valuable tool in English language tests for high school
students in Vietnam. They provide an efficient and objective way to assess various
language skills and knowledge areas. By adhering to the principles of effective item
construction and considering the potential limitations, educators can utilize matching
activities to create reliable and valid assessments that contribute to meaningful language
learning.
5.2.5 Designing gap-fill items

Gap-fill items, also known as fill-in-the-blank questions, are a staple in English language
tests for high school students in Vietnam. These questions require students to complete a
sentence or passage by filling in missing words or phrases, demonstrating their
understanding of grammar, vocabulary, and overall language structure. This section
delves into the nature of gap-fill items, their applications in English tests, their
advantages and limitations, and guidelines for effective construction and implementation.

Understanding gap-fill items

Gap-fill items present students with a text where specific words or phrases have been
removed and replaced with blanks. Students must analyze the context and utilize their
language knowledge to determine the missing elements and complete the text coherently
and accurately. This process assesses their ability to:

• Apply grammatical knowledge: Identify the correct tense, form, and agreement
of verbs, nouns, pronouns, adjectives, and adverbs.
• Utilize vocabulary: Select appropriate words that fit the context and convey the
intended meaning.
• Understand sentence structure: Recognize the syntactic roles of words and phrases
within a sentence.
• Comprehend discourse: Maintain coherence and cohesion within the text by
using appropriate linking words and phrases.

Applications in English tests

Gap-fill items can be used to assess a wide range of language skills and knowledge areas:

• Grammar and syntax: Testing knowledge of verb tenses, prepositions, articles,


conjunctions, and sentence structure.
• Vocabulary: Assessing understanding of word meanings, collocations, and
idiomatic expressions.
• Reading comprehension: Evaluating ability to understand the overall meaning
and flow of a text.
• Writing skills: Measuring ability to produce grammatically correct and
meaningful sentences.

Types of gap-fill items

• Open-ended: Students are free to choose any word or phrase that fits the context.
• Closed-ended: Students are provided with a word bank or a limited set of options
to choose from.
• Targeted: Gaps are strategically placed to assess specific grammar points or
vocabulary items.
• Sentence completion: Students complete a sentence by filling in a missing word
or phrase.
• Passage completion: Students fill in multiple gaps within a longer text.

Advantages of gap-fill items

• Versatility: Can be adapted to assess various language skills and levels of


difficulty.
• Objectivity: Scoring can be relatively straightforward, especially with closed-
ended or targeted items.
• Focus on production: Encourage active language use and demonstrate ability to
construct grammatically correct sentences.
• Authenticity: Can be based on authentic texts, increasing engagement and
relevance for learners.

Limitations of gap-fill items

• Potential for ambiguity: Some gaps may have multiple possible answers, making
scoring subjective.
• Limited scope: May not fully capture complex aspects of language use, such as
pragmatic understanding or communicative competence.
• Difficulty in constructing effective items: Creating gaps that have a single
correct answer and effectively assess the intended skill can be challenging.

Constructing effective gap-fill items

• Choose appropriate texts: Select texts that are relevant to learners' interests and
proficiency levels.
• Determine the type and number of gaps: Consider the difficulty level desired
and the specific skills being assessed.
• Provide clear instructions: Ensure learners understand the task and how to
respond.
• Focus on key language points: Target specific grammar structures or vocabulary
items for assessment.
• Avoid excessive gaps: Too many gaps can disrupt the flow of the text and make
the task overwhelming.
• Thorough review and piloting: Ensure items are accurate, clear, and free of bias
before including them in a test.
Faulty gap-fill items

Gaps with too many possible answers

Sentence: "I went to the ______ to buy some milk."

Fault: This sentence is too open-ended. Students could fill the gap with various
words like "store," "supermarket," "shop," "market," etc., making it difficult to
determine a single correct answer.

Gaps requiring specialized knowledge

Sentence: "The ______ is a traditional Vietnamese musical instrument made of


bamboo."

Fault: This item requires knowledge of Vietnamese culture that might not be
covered in the English textbook or familiar to all students.

Gaps with no clear context

Sentence: "The ______ was very ______."

Fault: This sentence provides no context or clues for the missing words. Students
are left to guess randomly, which doesn't effectively assess their language skills.

Gaps testing obscure vocabulary

Sentence: "The ancient artifact was ______ with intricate carvings."

Fault: Using a highly specific or obscure word like "emblazoned" might not be
appropriate for a high school level gap-fill test, especially if it's not a word
commonly encountered in the textbook.

In general, gap-fill items are a valuable component of English language tests for high
school students in Vietnam. They provide a flexible and focused way to assess various
language skills, particularly grammar, vocabulary, and sentence construction. By
adhering to the principles of effective item construction and considering the potential
limitations, educators can utilize gap-fill activities to create meaningful assessments that
contribute to effective language learning.

5.3 Designing subjective test items

Subjective test items, unlike their objective counterparts, require students to construct
their own responses rather than selecting from pre-defined options. These items, such as
essays, short answer questions, and oral interviews, offer a valuable means of assessing
deeper levels of language proficiency and cognitive skills in high school English
language learners in Vietnam. This chapter explores the characteristics, applications,
advantages, and limitations of subjective test items, along with guidelines for their
effective design and implementation.

Understanding subjective test items

Subjective test items typically present students with open-ended prompts or questions
that require them to generate unique responses using their own language and ideas. These
items assess a broader range of skills and knowledge, including:

• Writing Proficiency: Ability to express ideas clearly and coherently in written


form, demonstrating grammatical accuracy, vocabulary range, and organizational
skills.
• Speaking Proficiency: Ability to communicate effectively in spoken English,
showcasing fluency, pronunciation, grammar, and vocabulary.
• Critical Thinking: Ability to analyze information, form arguments, express
opinions, and provide justifications.
• Creativity: Ability to generate original ideas and express them in a unique and
engaging manner.
• Problem-Solving: Ability to apply knowledge and skills to resolve problems or
respond to complex scenarios (Ross & Okabe, 2006).

Applications in English tests

Subjective test items are valuable for assessing various aspects of English language
proficiency:

• Essay Writing: Evaluating ability to compose well-structured essays on a given


topic, expressing arguments, opinions, or analysis.
• Short Answer Questions: Assessing understanding of specific concepts or details
by requiring concise written responses.
• Oral Interviews: Evaluating spoken communication skills through interactive
conversations and responses to prompts.
• Presentations: Assessing ability to organize and deliver oral presentations on a
chosen topic.
• Creative Writing: Encouraging expression of imagination and storytelling skills
through original compositions.
Advantages of subjective test items

• Depth of assessment: Allow for a more in-depth evaluation of complex language


skills and cognitive abilities.
• Authenticity: Can simulate real-life communication tasks, increasing relevance
and engagement for learners.
• Reduced guessing: Minimize the chance of students obtaining correct answers
through guessing.
• Promotion of higher-order thinking: Encourage critical thinking, analysis, and
problem-solving skills.
• Individualized expression: Allow students to showcase their unique
understanding and perspectives.

Limitations of subjective test items

• Subjectivity in scoring: Evaluating open-ended responses can be subjective and


require clear scoring rubrics and rater training.
• Time-consuming: Marking subjective items can be more time-consuming than
objective items.
• Limited sampling: May not cover as much content as objective tests due to the
time required for each item.
• Potential for bias: Raters' personal biases can influence scoring, requiring careful
training and standardization.

Constructing effective subjective test items

• Clear and concise prompts: Provide specific and unambiguous instructions that
clearly define the task and expectations.
• Relevant to curriculum: Align prompts with the learning objectives and content
covered in the curriculum.
• Appropriate difficulty level: Ensure tasks are challenging yet attainable for the
students' proficiency level.
• Defined evaluation criteria: Develop clear scoring rubrics that outline the criteria
for evaluating responses and awarding marks.
• Rater training and standardization: Provide training to ensure consistent and
reliable scoring across different raters.

In general, subjective test items are an essential component of comprehensive English


language assessment for high school students in Vietnam. They provide valuable insights
into students' higher-order thinking skills, communicative abilities, and creative
expression. By carefully designing prompts, establishing clear scoring criteria, and
ensuring rater reliability, educators can effectively utilize subjective items to gain a more
holistic understanding of students' English language proficiency and promote meaningful
language development.

5.3 Scoring methods: objective vs. subjective

In the realm of English language assessment for high school students in Vietnam, the
choice of scoring methods plays a crucial role in determining the accuracy, fairness, and
effectiveness of evaluations. This section explores the key distinctions between objective
and subjective scoring methods, highlighting their respective advantages, limitations, and
applications in various assessment contexts.

5.3.1 Objective scoring: precision and efficiency

Objective scoring methods are characterized by their clear-cut criteria and standardized
procedures, leaving little room for individual interpretation or bias. These methods are
typically employed for assessing responses to objective test items, such as multiple-
choice, true-false, and matching questions.

Advantages of objective scoring

• Efficiency: Objective scoring is generally quick and efficient, allowing for rapid
evaluation of large numbers of test papers.
• Reliability: With standardized procedures and answer keys, objective scoring
produces consistent results, minimizing variations between different raters.
• Objectivity: The scoring process is free from personal biases or interpretations,
ensuring fairness and equal treatment for all test-takers.
• Ease of analysis: Objective scores are easily quantifiable, facilitating statistical
analysis and reporting of results.

Limitations of objective scoring

• Limited scope: Objective scoring is primarily suited for assessing lower-order


thinking skills, such as knowledge recall and recognition.
• Potential for guessing: Test-takers may obtain correct answers through guessing,
potentially inflating scores and not accurately reflecting their true abilities.
• Inability to assess complex skills: Objective scoring cannot effectively evaluate
complex language skills like writing, speaking, or critical thinking, which require
more nuanced and holistic assessment.

5.3.2 Subjective scoring: depth and nuance

Subjective scoring methods involve human judgment and interpretation to evaluate


responses to open-ended or performance-based tasks, such as essays, oral presentations,
and creative writing. These methods allow for a more in-depth assessment of complex
language skills and higher-order thinking abilities (Bachman & Adrian, 2022).

Advantages of subjective scoring

• Comprehensive assessment: Subjective scoring can evaluate a wider range of


language skills, including writing proficiency, speaking fluency, critical thinking,
and creativity.
• Authenticity: By assessing real-life communication tasks, subjective scoring
provides a more authentic measure of language proficiency.
• Individualized evaluation: Subjective scoring allows for personalized feedback
and recognition of individual strengths and weaknesses.
• Encouragement of higher-order thinking: Open-ended tasks promote critical
thinking, problem-solving, and expression of original ideas.

Limitations of subjective scoring

• Subjectivity and bias: Raters' personal biases and interpretations can influence
scoring, potentially leading to inconsistencies and unfairness.
• Time-consuming: Evaluating open-ended responses requires careful reading and
analysis, making subjective scoring more time-consuming than objective scoring.
• Rater reliability: Ensuring consistency and agreement between different raters
requires thorough training and standardization procedures.
• Difficulty in providing specific feedback: Providing detailed and specific
feedback on subjective responses can be challenging.

5.5 Best scoring practices in English language assessment

Scoring student work is a critical component of English language assessment in Vietnam.


Accurate and fair scoring provides valuable feedback to students, informs instructional
decisions, and contributes to the overall effectiveness of the educational process. This
section outlines best practices for scoring student work in English, encompassing both
objective and subjective assessment methods, to ensure reliable, valid, and meaningful
evaluations.

General principles for effective scoring

• Clarity of Criteria: Establish clear and specific scoring criteria or rubrics that
outline the expectations for each assessment task and the characteristics of
different performance levels.
• Transparency: Communicate the scoring criteria to students beforehand to ensure
they understand the expectations and can focus their efforts accordingly.
• Consistency: Apply the scoring criteria consistently across all student work to
ensure fairness and equity in the evaluation process.
• Objectivity: Minimize personal biases and subjective interpretations when
evaluating student work, particularly for subjective assessments.
• Feedback: Provide constructive feedback to students that highlights their strengths
and areas for improvement, guiding their future learning.

Best practices for objective scoring

• Accurate answer keys: Develop accurate and unambiguous answer keys for
objective test items, such as multiple-choice, true-false, and matching questions.
• Automated scoring: Utilize technology, where feasible, to automate the scoring
of objective tests, ensuring efficiency and accuracy.
• Item analysis: Analyze student responses to identify problematic items that may
need revision or removal from future assessments.

Best practices for subjective scoring

• Detailed rubrics: Develop detailed scoring rubrics that outline the specific criteria
for evaluating different aspects of performance, such as content, organization,
language use, and mechanics.
• Rater training: Provide thorough training to raters on the scoring rubrics to
ensure inter-rater reliability.
• Multiple raters: When possible, use multiple raters to evaluate subjective
responses, such as essays or oral presentations, and average their scores to
minimize bias.
• Anonymity: Mask student identities during the scoring process to prevent
unconscious biases from influencing evaluations.
• Holistic scoring: Consider the overall quality of the response, rather than focusing
solely on individual errors, to provide a more comprehensive assessment.

Providing effective feedback

• Specificity: Provide specific and detailed feedback that pinpoints areas of strength
and weakness in the student's work.
• Actionable advice: Offer actionable advice on how the student can improve their
performance in the future.
• Timeliness: Provide feedback promptly while the assessment is still fresh in the
student's mind.
• Encouragement: Balance constructive criticism with positive reinforcement to
motivate students and foster a growth mindset.

Ethical considerations in scoring


• Fairness: Ensure all students are evaluated fairly and equitably, regardless of their
background or personal characteristics.
• Confidentiality: Protect the confidentiality of student work and scores.
• Professionalism: Maintain a professional and unbiased approach to scoring,
adhering to ethical guidelines and standards.

Implementing best scoring practices is essential for ensuring accurate, fair, and
meaningful assessments of English language proficiency in Vietnam. By adhering to the
principles of clarity, consistency, objectivity, and constructive feedback, educators can
provide valuable evaluations that support student learning and contribute to the overall
effectiveness of the educational system.

In summary, the choice between objective and subjective scoring methods depends on
the specific learning objectives, the type of assessment tasks, and the desired level of
detail in the evaluation. By understanding the strengths and limitations of each approach,
educators can make informed decisions about scoring procedures, ensuring fair, reliable,
and meaningful assessments of English language proficiency for high school students in
Vietnam. Table 5.1 compares objective vs. subjective scoring:

Table 5.1 Objective vs. subjective scoring

Feature Objective scoring Subjective scoring


Closed-ended, selected- Open-ended, constructed-
Nature of Items
response response
Standardized, based on
Scoring Process Human judgment, interpretation
answer keys
Efficiency High Lower
Can be lower, requires rater
Reliability High
training
Objectivity High Can be influenced by rater bias
Higher-order thinking (analysis,
Types of Skills Primarily lower-order
synthesis, evaluation), complex
Assessed thinking (recall, recognition)
skills (writing, speaking)
Multiple-choice, true-false, Essays, oral presentations,
Examples
matching projects
Can be more detailed and
Feedback Limited to correct/incorrect
individualized

5.6 Student self-assessment

Instruction: Carefully read each statement in the table below. For each statement, rate
your understanding using the following scale:
1: I do not understand this at all.
2: I understand this a little, but I need more help.
3: I understand this fairly well, but I have some questions.
4: I understand this very well and can explain it to others.

Learning Objective/Concept Self- Evidence/Notes Action Plan


Assessment (Explain your (What will
Rating (1-4) rating) you do to
improve?)
Steps in test design: I can identify
and explain the key steps involved in
designing a language test (defining
purpose, objectives, specifications,
writing items, reviewing,
administering).
Developing specifications: I can
explain how to create a detailed test
specification, including content,
format, scoring criteria, and
weighting of different sections.
Writing test items: I can discuss the
principles of writing effective test
items, including clarity,
unambiguousness, and
appropriateness for the target
population.
Practical application: I can apply
the steps of test design to develop a
sample language test for a specific
purpose and context.
Overall understanding: I feel
confident in my understanding of the
process of designing and developing
effective language tests.

5.7 Consolidation activities

Activity 1: Crafting multiple-choice questions: A test design exercise

Objectives:
• To develop students' understanding of multiple-choice question design principles.
• To enhance students' critical thinking and analytical skills by analyzing a given
text and formulating appropriate test items.

Procedure: Students in groups/pairs read the following passage from Unit 2, the English
textbook Global Success for Grade 11. The students first read the passage carefully and
identify the main ideas, supporting details, any challenging vocabulary or grammatical
structures, and the overall tone and purpose of the passage. The items should include a
variety of question types:

o Factual questions: Test for recall of specific information (e.g., "What is the
main cause of...?").
o Inferential questions: Test for understanding of implied meanings and
making deductions (e.g., "What can be inferred about the author's attitude
towards...?").
o Vocabulary questions: Test for understanding of vocabulary in context.
o Main idea questions: Test for comprehension of the overall theme or central
message.

“Over the past two centuries, different generations were born and given different
names. Each generation comes with its characteristics, which are largely influenced
by the historical, economic, and social conditions of the country they live in.
However, in many countries the following three generations have common
characteristics.

Generation X refers to the generation born between 1965 and 1980. When Gen Xers
grew up, they experienced many social changes and developments in history. As a
result, they are always ready for changes and prepared to work through changes.
Gen Xers are also known as critical thinkers because they achieved higher levels of
education than previous generations.

Generation Y, also known as Millennials, refers to those born between the early
1980s and late 1990s. They are curious and ready to accept changes. If there is a
faster, better way of doing something, Millennials want to try it out. They also value
teamwork. When working in a team, Millennials welcome different points of view
and ideas from others.

Generation Z includes people born between the late 1990s and early 2010s, a time
of great technological developments and changes. That is why Gen Zers are also
called digital natives. They grew up online and never knew the world before digital
and social media. They are very creative and able to experiment with platforms to
suit their needs. Many Gen Zers are also interested in starting their own businesses
and companies. They saw so many people lose their jobs, so they think it is safer to
be your own boss than relying on someone else to hire you.

Soon a new generation, labelled Gen Alpha, will be on the scene. Let's wait and see
if we will notice the generation gap.” (From Unit 2, Global Success, Grade 11,
Hoang et al., 2018)

Activity 2: Crafting T/F/NG questions: A test design exercise

Objective: To encourage critical thinking and enhance comprehension skills by having


students design their own True-False-Not Given statements for a given text.

Procedure: Students in pairs/groups to read a passage from the section on


communication and culture the Tiếng Anh 11 Unit 3:

Smart Cities Around The World

Cities around the world are becoming smarter, and you can do many things that
seemed impossible in the past.

In Singapore, the mobile app [Link] allows you to locate a nearby car park
easily, book a parking space, and make a payment. You can also extend your
booking or receive a refund if you leave early.

New York City (US) has one of the largest bike-sharing systems called Citi Bike.
Using a mobile app, you can unlock bikes from one station and return them to any
other station in the system, making them ideal for one-way trips.

In Copenhagen (Denmark), you can use a mobile app to guide you through the city
streets and tell how fast you need to pedal to make the next green light. The app can
also give you route recommendations and work out the calories you burn.

In London(UK), you don’t have to buy public transport tickets. You can just touch
your bank card on the card reader when you get on and off the bus or the
underground to pay for your trip.

In Toronto (Canada), you can book an appointment and see a doctor online a from
your own home. You can also receive prescriptions and any other documents you
need, all online.

When reading the passage, students need to highlight key information and identifying
main ideas. They review the characteristics of each type of statement:

o True: Accurately reflects information explicitly stated in the text.


o False: Contradicts information presented in the text.
o Not Given: Refers to information not mentioned or addressed in the text.

Each pair/group should create at least 2 statements for each category (True, False, Not
Given), resulting in a minimum of 6 statements. Students should vary the difficulty level
of their statements and to avoid simply copying sentences directly from the text.

Activity 3: Subjective item explorers

Objective: To deepen students' understanding of subjective test items by actively


engaging in the process of analyzing, evaluating, and constructing these items.

Procedure: Students into small groups search for a set of subjective test items in the
English tests for high school students in Vietnam (e.g. essay prompts, short answer
questions, oral interview questions from previous English tests or textbooks in use). They
discuss the following questions:

§ What skills and knowledge are being assessed? (e.g., writing,


speaking, critical thinking, creativity)
§ What are the key features of a good response? (refer to the
scoring rubrics)
§ What are some common mistakes or challenges students might
face?
§ How can students best prepare for these types of items?

Activity 4: Achievement test architects

Objective: To enable students to understand the construction and purpose of subjective


test items in English achievement tests by having them design their own assessment
questions.

Procedure: Students in groups begin by reviewing the key learning objectives and
content areas covered in a unit or several units an English textbook for high school
students in Vietnam, then identify the knowledge and skills that should be assessed in an
achievement test. They then discuss the different types of subjective items commonly
used in achievement tests (essays, short answer questions, etc.). During group discussion,
students

§ Specify the learning objective being assessed.


§ Formulate a clear and concise prompt or question.
§ Define the expected response format and length.
§ Outline the criteria for evaluating the response.
§ Consider potential challenges or difficulties for test-takers.
Activity 5: Spot the differences: Comparing test matrices

Objective: To enable students to analyze and compare two test matrices (one for Grade 6
and one for Grade 8 in Vietnam as in the images below) and identify the key differences
in terms of skills, knowledge, and complexity.

Procedure: Students in groups spend time reviewing the content and structure of each
matrix. They use highlighters or colored pens to mark any noticeable differences between
the two matrices and then report the findings to the whole class. They should focus on
aspects like:

§ Skills: Are there any skills assessed in one matrix that are not
present in the other? Are the same skills assessed at different
levels of complexity?
§ Knowledge: Are there differences in the types or depth of
knowledge expected at each grade level?
§ Weighting: Are certain skills or knowledge areas given more
emphasis in one matrix compared to the other?
Chapter 6: Innovations in Language Assessment

Language assessment is undergoing a profound transformation, driven by rapid


advancements in technology and a growing understanding of how people learn and use
language. This chapter explores the key forces shaping this evolving landscape, with a
particular focus on the rise of online and computer-based testing, the burgeoning role of
artificial intelligence (AI), and the implications these trends hold for the future of
language assessment.

6.1 Online and computer-based language testing

The landscape of language testing in the world and in Vietnam is rapidly evolving, with
online and computer-based testing (CBT) gaining significant traction. This shift is driven
by technological advancements, increased internet accessibility, and the need for more
efficient and flexible assessment methods.

The rise of online and CBT in Vietnam

The COVID-19 pandemic accelerated the adoption of online and CBT in Vietnam, as
educational institutions sought alternative assessment solutions during school closures
(Bui, 2023). This period highlighted the potential of technology to deliver assessments
remotely, cater to diverse learning needs, and provide immediate feedback. Furthermore,
the Vietnamese government has been actively promoting digital transformation in
education, encouraging the integration of technology into teaching and assessment
practices.

Key features of online and CBT

• Remote delivery: Online and CBT can be administered remotely, allowing for
greater flexibility and accessibility for test-takers across geographical locations.
• Automated scoring: Many CBT platforms offer automated scoring for objective
test items, increasing efficiency and reducing human error.
• Multimedia integration: Online and CBT can incorporate multimedia elements,
such as audio, video, and interactive simulations, to create more engaging and
authentic assessments.
• Adaptive testing: Some CBT platforms utilize adaptive testing algorithms that
adjust the difficulty level of questions based on the test-taker's performance,
providing a more personalized assessment experience.
• Data-driven insights: Online and CBT generate valuable data on student
performance, enabling educators to track progress, identify learning gaps, and
tailor instruction accordingly.
Advantages of online and CBT

• Increased accessibility: Online and CBT can reach a wider range of test-takers,
including those in remote areas or with disabilities, promoting inclusivity and
equity in assessment.
• Enhanced efficiency: Automated scoring and online test administration streamline
the assessment process, saving time and resources for educators.
• Improved engagement: Multimedia elements and interactive features can enhance
student engagement and motivation during the testing process.
• Personalized learning: Adaptive testing and data-driven insights enable
personalized feedback and targeted interventions to support individual learning
needs.
• Reduced environmental impact: Online and CBT reduce reliance on paper-based
materials, contributing to environmental sustainability.

Challenges of online and CBT

• Technological infrastructure: Reliable internet access and appropriate devices are


crucial for successful implementation, which can be a challenge in some areas of
Vietnam.
• Test security: Maintaining test security and preventing cheating can be more
challenging in online environments, requiring robust authentication and proctoring
measures.
• Digital literacy: Test-takers need basic digital literacy skills to navigate online
platforms and utilize assessment tools effectively.
• Validity and reliability: Ensuring the validity and reliability of online and CBT
requires careful test design, item development, and standardization procedures.
• Equity and access: Addressing potential disparities in access to technology and
digital skills is crucial to ensure fairness and equity for all test-takers.

Implications for language education in Vietnam

The adoption of online and CBT has significant implications for language education in
Vietnam:

• Teacher training: Educators need training on designing, administering, and scoring


online and CBT, as well as integrating technology effectively into their teaching
practices.
• Curriculum development: Language curricula need to be adapted to incorporate
digital literacy skills and prepare students for online assessments.
• Assessment policies: Clear policies and guidelines are needed to ensure the
quality, security, and ethical use of online and CBT in language assessment.

In summary, online and computer-based language testing are transforming the assessment
landscape in Vietnam, offering numerous advantages in terms of accessibility, efficiency,
and personalization. However, addressing the challenges related to technology, security,
and equity is crucial to ensure the successful and ethical implementation of these
innovative assessment methods. By embracing best practices and investing in teacher
training and infrastructure, Vietnam can harness the power of technology to enhance
language learning and assessment for all students.

6.2 The use of artificial intelligence in language testing and assessment

Artificial intelligence (AI) is rapidly transforming various sectors including language


education. In the realm of language testing and assessment, AI is being increasingly
utilized to automate tasks, personalize learning experiences, and provide more efficient
and effective evaluations (Koraishi, 2023). This chapter explores the applications,
benefits, challenges, and ethical considerations surrounding the use of AI in language
testing and assessment, with a focus on its potential impact on high school students in
Vietnam.

AI-powered applications in language testing

AI is being integrated into various aspects of language testing, including:

• Automated essay scoring (AES): AI algorithms can analyze and score written
responses, providing feedback on grammar, vocabulary, organization, and
coherence (Wilson & Roscoe, 2022).
• Computerized adaptive testing (CAT): AI-powered CAT platforms adjust the
difficulty level of questions based on the test-taker's performance, leading to more
efficient and precise assessments.
• Automated speech recognition (ASR): ASR technology can evaluate spoken
language proficiency, assessing pronunciation, fluency, and accuracy in real-
time.
• Chatbots for language assessment: AI-powered chatbots can engage in
conversations with test-takers, simulating real-life interactions and assessing
communicative competence.
• AI-driven feedback and remediation: AI systems can analyze student performance
and provide personalized feedback and recommendations for improvement.
Benefits of AI in language testing

• Increased efficiency and scalability: AI can automate time-consuming tasks like


scoring and feedback, allowing educators to assess more students efficiently.
• Enhanced objectivity and fairness: AI algorithms can reduce human bias in
scoring, leading to more objective and fair evaluations.
• Personalized learning experiences: AI can adapt assessments to individual needs
and provide tailored feedback, promoting personalized learning and motivation.
• Improved diagnostic accuracy: AI can analyze student responses in detail,
identifying specific areas of strength and weakness for targeted interventions.
• Real-time feedback and remediation: AI can provide immediate feedback and
suggest personalized learning resources, facilitating timely improvement and
progress.

Challenges and considerations

• Data privacy and security: Protecting student data and ensuring ethical use of AI
are critical considerations.
• Bias and fairness: AI algorithms can reflect biases present in the data they are
trained on, requiring careful development and validation to ensure fairness.
• Validity and reliability: Ensuring the accuracy and consistency of AI-powered
assessments requires rigorous validation and ongoing monitoring.
• Technological infrastructure: Reliable internet access and appropriate devices are
needed for widespread adoption of AI-powered testing.
• Teacher training and acceptance: Educators need training and support to
effectively integrate AI tools into their teaching and assessment practices.

Ethical implications of AI in assessment

• Transparency and explainability: AI algorithms should be transparent and


explainable, allowing educators to understand how they work and make informed
decisions.
• Accountability and responsibility: Clear lines of accountability should be
established for the use of AI in assessment, ensuring responsible and ethical
practices.
• Human oversight and intervention: Human oversight is crucial to ensure that AI is
used ethically and effectively in assessment, with the ability to intervene when
necessary.

Overall, AI has the potential to revolutionize language testing and assessment in


Vietnam, offering numerous benefits in terms of efficiency, personalization, and
objectivity. However, it is essential to address the challenges related to ethics, bias, and
technology to ensure that AI is used responsibly and effectively to support student
learning and development. By embracing best practices and fostering collaboration
between educators, technologists, and policymakers, Vietnam can harness the power of
AI to create a more equitable and effective language assessment system for all students.

6.3 Future trends and directions in language testing

Language testing is a dynamic field, constantly evolving to meet the changing needs of
learners, educators, and society. This chapter explores the emerging trends and future
directions that are shaping the landscape of language testing, with a particular focus on
their potential impact on high school students in Vietnam.

Technology-enhanced assessments

Technology continues to revolutionize language testing, enabling innovative assessment


formats and personalized learning experiences (Kim et al., 2022). Key trends include:

• Artificial Intelligence (AI): As discussed in this chapter, AI is being used for


automated scoring, adaptive testing, and personalized feedback, increasing
efficiency and objectivity in assessment.
• Gamification: Incorporating game-like elements into language tests can enhance
student engagement and motivation, making learning more enjoyable and
effective.
• Virtual reality (VR) and augmented reality (AR): VR and AR technologies can
create immersive simulations for language practice and assessment, providing
authentic and engaging learning environments.
• Mobile-first assessment: With increasing access to smartphones, mobile-first
assessments offer flexibility and convenience for test-takers, allowing them to
complete tests anytime, anywhere.

Focus on real-world communication

There is a growing emphasis on assessing language proficiency in authentic, real-world


contexts. This shift is reflected in trends such as:

• Performance-based assessment: Tasks that simulate real-life communication


scenarios, such as role-playing, presentations, and collaborative projects, are
becoming more prevalent.
• Multimodal assessment: Assessing language skills across different modes of
communication, including speaking, listening, reading, and writing, provides a
more holistic view of language proficiency.
• Assessment for learning: Integrating assessment with instruction, providing
ongoing feedback, and using assessments to guide learning and personalized
instruction.
Emphasis on 21st century skills

Language tests are evolving to assess not only traditional language skills but also 21st-
century skills, such as:

• Critical thinking and problem-solving: Tasks that require analysis, evaluation, and
creative solutions are becoming more common in language assessments.
• Collaboration and communication: Assessing students' ability to work effectively
in groups, communicate ideas clearly, and engage in collaborative problem-
solving.
• Digital literacy: Evaluating students' ability to navigate digital environments, use
technology effectively for learning, and communicate in online contexts.
• Intercultural competence: Assessing students' ability to understand and interact
effectively with people from different cultural backgrounds.

Ethical considerations and fairness

As language testing becomes more technology-driven and personalized, ethical


considerations are paramount. Key areas of focus include:

• Data privacy and security: Protecting student data and ensuring responsible use of
AI and other technologies in assessment.
• Bias and fairness: Addressing potential biases in assessment design and algorithms
to ensure fairness and equity for all test-takers.
• Accessibility and inclusivity: Designing assessments that are accessible to students
with diverse learning needs and disabilities.
• Transparency and accountability: Providing clear information about assessment
practices, scoring methods, and the use of technology.

In short, the future of language testing is marked by innovation, technology integration,


and a focus on real-world communication and 21st-century skills. By embracing these
trends and addressing the ethical considerations, Vietnam can create a robust and
equitable language assessment system that prepares students for success in an
increasingly interconnected and digital world.

6.4 Student self-assessment

Instruction: Carefully read each statement in the table below. For each statement, rate
your understanding using the following scale:

1: I do not understand this at all.


2: I understand this a little, but I need more help.
3: I understand this fairly well, but I have some questions.
4: I understand this very well and can explain it to others.

Learning Objective/Concept Self- Evidence/Notes Action Plan


Assessment (Explain your (What will
Rating (1-4) rating) you do to
improve?)
Online and CBT: I can define and
describe the characteristics of online
and computer-based language testing
(CBT), including their advantages
and disadvantages.
Key features: I can explain the key
features of online and CBT, such as
remote delivery, automated scoring,
multimedia integration, adaptive
testing, and data-driven insights.
Artificial Intelligence in
assessment: I can discuss the use of
artificial intelligence (AI) in
language assessment, including
automated essay scoring, chatbots
for language practice, and AI-
powered feedback systems.
Critical evaluation: I can critically
evaluate the benefits and challenges
of different innovations in language
assessment, considering their impact
on test-takers, educators, and the
field of language testing and
assessment as a whole.
Overall understanding: I feel
confident in my understanding of the
innovations in language assessment
and their implications for the future
of the field.
6.5 Consolidation activities

Activity 1: Online test detectives

Objective: To help students understand and reflect on online English tests by analyzing a
real online test designed for high school students ([Link]
or [Link]

Procedure: Students work in groups/pairs to access the online English test(s) in the links
provided and explore the test interface, navigation, and different question types. They
then answer the following questions:

§ What are the different types of questions included in the test?


§ How does the online format affect the way questions are presented
and answered?
§ What are the advantages and disadvantages of taking this test online
compared to a paper-based test?
§ How user-friendly is the test interface? Are there any aspects that
could be improved?
§ How does this test align with the learning objectives and curriculum
for high school English in Vietnam?
§ What are the technical requirements for taking this test? Are there
any potential barriers to access?

Activity 2: Paper vs. online language test showdown

Objective: To encourage students to critically compare and contrast the features,


advantages, and disadvantages of paper-based and online language tests.

Procedure: Students in groups obtain a copy of the paper-based test and access to an
online test (see the links in Activity 1). They then analyze and compare the two tests
across various dimensions, such as:

§ Test content: Are the topics and skills assessed similar in both tests?
§ Question types: What types of questions are used in each format
(e.g., multiple-choice, essay, listening comprehension)?
§ Delivery and format: How are the tests administered? How does the
format affect the test-taking experience?
§ Accessibility: Are there differences in how accessible each test is to
different learners?
§ Interaction: How interactive is each test? Are there opportunities for
multimedia or adaptive testing in the online version?
§ Feedback and Scoring: How is feedback provided? How is the test
scored?

Activity 3: Online test architects

Objective: To enable students to understand the principles of online language test design
by creating an outline for an online English test for high school students in Vietnam.

Procedure: Students begin by reviewing the key learning objectives and content areas
covered in an English textbook or an English program. They then discuss the specific
skills and knowledge that should be assessed in an online test. Each group should define
the specific objectives of their online test. What skills and knowledge will it assess? What
level of proficiency will it target? Students consider the following test specifications in
their outline:

§ Target audience: (e.g., high school students in Vietnam, specific


grade level)
§ Test purpose: (e.g., achievement test, proficiency test, diagnostic
test)
§ Time allocation: (e.g., overall test duration, time per section)
§ Skills to be assessed: (e.g., reading, writing, listening, speaking,
grammar, vocabulary)
§ Number of items: (total number of questions, distribution across
sections)

Activity 4: Gamify your English test!

Objective: To raise students' awareness of the potential of gamification in English


language tests by having them analyze and brainstorm gamified assessment ideas.

Procedure: Students in groups discuss examples of gamified language learning


activities or apps (e.g., Duolingo, Kahoot!, Quizlet Live) and examples of gamified
language learning activities or apps, highlighting the game mechanics used (e.g., points,
badges, leaderboards, challenges, rewards). They then discuss how these game mechanics
can motivate learners, increase engagement, and promote language acquisition. Below
are the questions to guide the disucssion:

§ Game mechanics used: What game elements are incorporated?


How do they function within the test?
§ Impact on motivation and engagement: How do the game
mechanics affect the test-taking experience?
§ Effectiveness in assessing language skills: How well does the
gamified format assess language proficiency?
§ Potential drawbacks or limitations: Are there any downsides to
using gamification in this test?

Activity 5: AI-powered language testing: Opportunities and challenges

Objective: To foster critical thinking and discussion among students about the potential
applications, benefits, and challenges of using AI in English language tests for high
school students in Vietnam.

Procedure: Students in groups discuss with a set of prompts or questions related to AI


in language testing. Example prompts:

§ How can AI be used to improve the efficiency and objectivity of


English language tests in Vietnam?
§ What are the potential benefits of using AI to provide personalized
feedback and learning recommendations to high school students?
§ What are some challenges or concerns associated with using AI in
language assessment? (e.g., bias, fairness, data privacy, access)
§ How can we ensure that AI is used ethically and responsibly in
language testing?
§ What are your thoughts on using AI-powered chatbots to assess
speaking skills?
§ How might AI change the way English language tests are designed
and administered in the future?

References

Bui, H. P. (2023). Vietnamese university EFL teachers’ and students’ beliefs and
teachers’ practices regarding classroom assessment. Language Testing in
Asia, 13(1), 10.

Kim, A. A., Tywoniw, R. L., & Chapman, M. (2022). Technology-enhanced items in


grades 1–12 English language proficiency assessments. Language Assessment
Quarterly, 19(4), 343-367.

Koraishi, O. (2023). Teaching English in the age of AI: Embracing ChatGPT to optimize
EFL materials and assessment. Language Education and Technology, 3(1), 55-72.

Wilson, J., & Roscoe, R. D. (2020). Automated writing evaluation and feedback:
Multiple metrics of efficacy. Journal of Educational Computing Research, 58(1),
87-125.
Chapter 7: Language Assessment Practices at High Schools in Vietnam

Language assessment plays a critical role in the Vietnamese education system, serving
various purposes such as measuring student progress, evaluating teaching effectiveness,
and informing educational policy. This chapter provides an in-depth analysis of language
assessment practices at the high school level in Vietnam, examining the types of
assessments, their purposes, the challenges faced, and the ongoing efforts to improve
assessment quality and fairness.

7.1 The context of language assessment in Vietnam

English language education has become increasingly important in Vietnam, driven by


globalization, economic development, and the demand for skilled workers with strong
English proficiency (Bui & Nguyen, 2024). The Vietnamese government has
implemented various reforms to enhance English language teaching and learning,
including the adoption of the Common European Framework of Reference for Languages
(CEFR) as a national standard. The national project of Vietnam has created a reference
framework of foreign language competence from which Vietnamese standardized tests of
English proficiency (VSTEP) has been made.

In this context, English language assessment serves multiple purposes:

• Measuring student achievement: Tests are used to assess students' progress and
achievement in English, providing feedback on their strengths and weaknesses.
• Placement and selection: Language tests are used for placement purposes,
assigning students to appropriate English classes based on their proficiency levels.
They are also used for selection, determining eligibility for scholarships, study
abroad programs, and university admission.
• Evaluating teaching effectiveness: Tests can provide insights into the effectiveness
of teaching methodologies and identify areas where improvements are needed.
• Informing educational policy: Large-scale language assessments provide valuable
data for policymakers to evaluate the overall effectiveness of English language
education programs and make informed decisions about curriculum development
and resource allocation.

7.2 Types of language assessments used in high schools in Vietnam

7.2.1 Classroom-based language assessments

Classroom-based language tests are an integral part of the English language learning
experience for high school students in Vietnam. These assessments, designed and
administered by teachers, play a crucial role in monitoring student progress, providing
feedback, and guiding instructional decisions. This chapter explores the key features,
types, benefits, challenges, and best practices associated with classroom-based language
tests in the Vietnamese high school context.

The role of classroom-based assessments

In Vietnam, where high-stakes national exams often dominate the educational landscape,
classroom-based assessments provide a valuable counterbalance. They offer a more
nuanced and personalized view of student learning, allowing teachers to:

• Monitor progress: Track individual student progress and identify areas of strength
and weakness.
• Provide feedback: Offer timely and specific feedback to students on their language
development.
• Inform instruction: Adjust teaching strategies and activities based on student
performance and needs.
• Promote learning: Encourage student engagement, motivation, and self-reflection.
• Create a positive learning environment: Foster a supportive and growth-oriented
classroom culture where assessment is seen as a tool for learning.

Types of classroom-based language assessments

Vietnamese high school teachers employ a variety of classroom-based assessments,


including:

• Formative assessments: These ongoing assessments, such as quizzes, short writing


tasks, and oral presentations, provide continuous feedback and guide learning
throughout the course.
• Summative assessments: These assessments, such as mid-term and end-of-term
tests, evaluate student learning at the end of a unit or semester.
• Performance-based assessments: These assessments require students to
demonstrate their language skills in authentic contexts, such as role-playing,
debates, and project presentations.
• Portfolio assessments: Students compile a collection of their work over time,
showcasing their progress and reflecting on their learning journey.
• Self- and peer-assessments: Students engage in self-reflection and provide
feedback to their peers, promoting metacognitive awareness and learner
autonomy.

Benefits of classroom-based assessments

• Alignment with curriculum: Teachers can tailor assessments to their specific


curriculum and learning objectives, ensuring relevance and alignment.
• Flexibility and adaptability: Classroom-based assessments can be adapted to suit
the needs of diverse learners and cater to different learning styles.
• Timely feedback: Teachers can provide immediate feedback to students,
facilitating timely improvement and addressing misconceptions.
• Reduced test anxiety: Formative assessments, in particular, can reduce test anxiety
by providing frequent and low-stakes opportunities for students to demonstrate
their learning.
• Enhanced teacher-student relationship: Regular assessments and feedback foster a
closer teacher-student relationship, building trust and promoting a positive
learning environment.

Challenges in classroom-based assessments

• Teacher training and expertise: Some teachers may lack adequate training in
language assessment principles and practices, which can affect the quality and
fairness of their assessments.
• Time constraints: Designing, administering, and grading assessments can be time-
consuming for teachers, especially in large classes.
• Subjectivity in scoring: Subjective assessments, such as essays and oral
presentations, can be challenging to score consistently and fairly.
• Balancing formative and summative assessment: Finding the right balance
between formative and summative assessments can be challenging, as over-
reliance on summative assessments can create pressure and neglect ongoing
learning.

Best practices for classroom-based assessments

• Clear learning objectives: Clearly define learning objectives and communicate


them to students so they understand the expectations and can focus their efforts.
• Variety of assessment methods: Use a variety of assessment methods to cater to
different learning styles and provide a more comprehensive view of student
learning.
• Constructive feedback: Provide specific, timely, and actionable feedback that
focuses on learning goals and encourages improvement.
• Student involvement: Involve students in the assessment process through self- and
peer-assessment, promoting learner autonomy and metacognitive awareness.
• Fairness and transparency: Ensure assessments are fair, unbiased, and transparent,
with clear criteria and consistent scoring practices.

Classroom-based language tests are essential for effective English language instruction in
Vietnamese high schools. By implementing best practices, providing ongoing teacher
training, and embracing innovative assessment approaches, educators can create a more
engaging, supportive, and learner-centered assessment environment that promotes
language development and academic success for all students.

7.2.2 High-stakes English exams

High-stakes English exams play a crucial role in Vietnam's education system and society.
They are called “high-stakes” because the results from these tests result in significant
consequences for the test taker related to academic decisions, professional
opportunities, and immigration or citizenship. In other words, high-stakes English
exams are significant for the following reasons:

• University admissions: Vietnamese universities, especially prestigious ones, often


require a minimum English proficiency score for admission. This is increasingly
important as more institutions offer international programs and partnerships.
• Job opportunities: With Vietnam's growing economy and international integration,
English proficiency is highly valued by employers. Many companies, especially in
sectors like tourism, technology, and foreign trade, require job applicants to
demonstrate their English skills through standardized tests.
• Studying abroad: For Vietnamese students aspiring to study abroad, exams like
IELTS or TOEFL are essential for demonstrating their language readiness to
attend universities in English-speaking countries.

Key high-stakes English exams in Vietnam:

• IELTS (International English Language Testing System): Widely recognized by


universities and employers in the UK, Australia, and other Commonwealth
countries.
• TOEFL (Test of English as a Foreign Language): Predominantly accepted by
universities in the US and Canada.
• Cambridge English Exams: A suite of exams offered by Cambridge Assessment
English, including B1 Preliminary, B2 First, and C1 Advanced, recognized by
many institutions worldwide.
• VSTEP (Vietnamese Standardized Test of English Proficiency): Developed by the
Ministry of Education and Training (MOET) as a national English proficiency test,
increasingly used by Vietnamese universities and government agencies.
• National high school exit exam: The national high school exit exam itself includes
a compulsory English component, adding another layer of pressure for students
(Nguyen & Nguyen, 2020).
Challenges and trends:

• Pressure and anxiety: The intense competition for university places and good jobs
creates immense pressure on students to perform well in these exams. This can
lead to test anxiety and mental health concerns.
• Access to resources: While English language learning resources are becoming
more widely available, there's still a disparity between urban and rural areas in
terms of access to quality instruction and preparation materials.
• Emphasis on standardized testing: There's ongoing debate about the over-reliance
on standardized tests as the sole measure of English proficiency, with some
advocating for more holistic assessment methods that consider real-world
communication skills.

Recent developments:

• MOET recognition of Cambridge exams: The Ministry of Education and Training


has recently recognized certain Cambridge English exams, allowing high school
students holding these qualifications to be exempt from the English test in the
national high school exam. This highlights a growing acceptance of international
English proficiency certifications in Vietnam.
• Increasing popularity of VSTEP: The VSTEP exam is gaining traction as a
national standard, potentially reducing reliance on international tests and providing
a more affordable and accessible option for Vietnamese students.

High-stakes English exams are an integral part of Vietnam's education landscape,


reflecting the country's commitment to internationalization and developing a skilled
workforce. As Vietnam continues to integrate with the global economy, these exams will
likely play an even more significant role in shaping the future of its students and its
workforce (Ngo, 2024).

7.3 Challenges and issues in language assessment in Vietnam

Despite ongoing efforts to improve language testing practices, several challenges persist:

• High-stakes testing pressure: The emphasis on high-stakes exams, such as the


National High School Graduation Examination, can create immense pressure on
students and teachers, potentially leading to teaching to the test and neglecting
other important aspects of language learning.
• Assessment literacy: Many teachers lack adequate training in language assessment
principles and practices, which can affect the quality and fairness of classroom-
based assessments.
• Access to resources and technology: unequal access to resources and technology
can create disparities in testing opportunities and impact the validity of
assessments for students from disadvantaged backgrounds.
• Washback effect: The washback effect of high-stakes tests can influence teaching
practices and curriculum design, sometimes leading to a narrow focus on test
preparation and neglecting broader language skills development.

Language testing practices in Vietnamese high schools are dynamic and evolving,
reflecting the growing importance of English language proficiency in the country. By
addressing the challenges, embracing innovations, and focusing on assessment for
learning, Vietnam can create a more effective and equitable language assessment system
that supports student learning and prepares them for success in the 21st century.

7.4 Backwash effects

Backwash refers to the influence that tests have on teaching and learning practices. When
tests are perceived as high-stakes, they can have a significant impact on what is taught,
how it is taught, and how students learn. Positive washback occurs when tests promote
beneficial teaching practices and learning outcomes, while negative washback can lead to
teaching to the test, where instruction becomes narrowly focused on test content at the
expense of broader language skills (Fulcher & Davidson, 2007).

Positive backwash effects

• Focus on specific skills: Tests can encourage teachers to focus on specific


language skills, such as reading, writing, listening, and speaking.
• Motivation: Well-designed tests can motivate learners to study harder and improve
their language skills.
• Standardization: Standardized tests can promote consistency in language teaching
and learning across different institutions.

Negative backwash effects

• Narrowing the curriculum: Teachers may focus on teaching only what is tested,
neglecting other important aspects of language learning, such as creativity and
critical thinking.
• Teaching to the test: Teachers may prioritize test preparation over meaningful
language learning, leading to rote learning and a focus on test-taking strategies.
• Test anxiety: High-stakes tests can cause anxiety and stress, which can negatively
impact learners' performance.
• Inequitable assessment: Tests may not accurately reflect the diverse range of
language abilities, particularly for learners from marginalized backgrounds.
Mitigating negative backwash effects

• Test validity and reliability: Ensure that tests are valid and reliable measures of
language proficiency.
• Balanced assessment: Use a variety of assessment methods, including formative
and summative assessments, to avoid over-reliance on high-stakes tests.
• Teacher training: Provide teachers with training on effective test preparation
strategies and how to avoid negative backwash effects.
• Learner awareness: Educate learners about the purpose of tests and how to
approach them effectively.
• Authentic tasks: Incorporate authentic language tasks into teaching and assessment
to promote real-world language use.

By understanding the potential positive and negative impacts of tests, language educators
can make informed decisions about assessment practices and strive to create a positive
learning environment that fosters language development and critical thinking.

7.5 Student self-assessment

Instruction: Carefully read each statement in the table below. For each statement, rate
your understanding using the following scale:

1: I do not understand this at all.


2: I understand this a little, but I need more help.
3: I understand this fairly well, but I have some questions.
4: I understand this very well and can explain it to others.

Learning Objective/Concept Self- Evidence/Notes Action Plan


Assessment (Explain your (What will
Rating (1-4) rating) you do to
improve?)
Purposes of testing: I can identify
and explain the various purposes of
English language testing and
assessment in Vietnamese high
schools, such as measuring
achievement, placement, evaluating
teaching, and informing policy.
Classroom-based assessments: I
can describe the characteristics and
types of classroom-based language
assessments used in Vietnamese high
schools.
Standardized tests: I can discuss
the use of standardized tests for high
school students in Vietnam,
including their purpose, format, and
examples (e.g., national high school
graduation exams).
Overall understanding: I feel
confident in my understanding of
language assessment practices at
high schools in Vietnam and the
factors that influence them.

7.6 Consolidation activities

Activity 1: Formative assessment designers

Objective: To deepen students' understanding of formative assessment and its purpose in


English language learning for high school students in Vietnam

Procedure: Students begin by reviewing the concept of formative assessment with the
students:

o Focus on learning: Formative assessment is primarily aimed at supporting


and improving student learning, not just assigning grades.
o Ongoing and interactive: It is an ongoing process integrated into
instruction, involving feedback and adjustments along the way.
o Student-centered: It actively involves students in self-assessment and
reflection.

Students then brainstorm a variety of formative assessment activities that could be used
in an English language classrooms, considering activities like:

o Role-plays and simulations


o Presentations and debates
o Games and interactive activities
o Peer assessment and feedback
o Self-reflection journals or portfolios
o Creative projects (e.g., writing stories, creating videos, designing
posters)
o Use of technology (e.g., online quizzes, interactive whiteboards, digital
storytelling)
Activity 2: Textbook assessment in focus

Objectives:

• To encourage critical analysis of summative assessment practices in the context of


a specific English textbook used in Vietnamese high schools.
• To stimulate discussion and debate about the alignment of assessment tasks with
learning objectives and curriculum standards.

Procedure: Students examine an English textbook in use in Vietnamese high schools


(e.g., Tiếng Anh 10, 11, or 12 Global Success), the textbook's summative assessment
sections (e.g., end-of-unit tests/reviews), curriculum guidelines for English language in
Vietnamese high schools. They then choose a specific unit/units from the textbook to
focus on, the review exercises, or sample exam questions. They next discuss specific
aspects of the summative assessment. Some guiding questions are as follows:

• What is the overall purpose of the summative assessment in this textbook? Is it


primarily to measure achievement, assign grades, provide feedback, or a
combination of these?
• How well does the summative assessment align with the learning objectives and
content covered in the textbook? Are all the important topics and skills adequately
assessed?
• Does the summative assessment reflect the emphasis and approach of the
textbook? For example, if the textbook emphasizes communicative skills, does the
assessment provide opportunities for students to demonstrate these skills?
• How does the summative assessment cater to different learning styles and
abilities? Are there a variety of task types to accommodate diverse learners?
• What are the practical considerations for administering and scoring the summative
assessment? Is it feasible to implement within the given time and resource
constraints?

Activity 3: Understanding a high-stakes English test for high school students

Objective: To enable students to understand the components of high-stakes English tests


for high school students and foster critical thinking about the design of such tests.

Procedure: Students work in groups/pairs and examine the following an English test for
high school students in Vietnam in 2024. They need to answer the following guiding
questions:
What test items are included?
What do they aim to measure?
Which parts reflects the English teaching curriculum the most? Why?
Which parts do you suggest to be added? Why?
Activity 4: Comparing English tests for high school students in two different years

Objective: To enhance students’ understanding of the changes in high-stakes English


tests for high school students and the reasons for such changes.

Procedure: Students in groups examine the high-stake English test for 2025 below and
compare with the one in 2024 as in Activity 1. They need to point out the similarities and
differences and explain for the changes.
Activity 5: VSTEP debate showdown

Objectives:

• To encourage students to critically analyze the VSTEP exam format and its
components.
• To enhance students' understanding of the skills and strategies needed to succeed
in the VSTEP exam.

Procedure: Students in groups examine a VSTEP from [Link] They


then examine a test format and component and discuss how it assesses English
language proficiency for Vietnamese students. The following prompts are guiding
questions for the discussion:

§ VSTEP’s alignment with Vietnamese educational context


§ Accessibility and affordability compared to international tests
§ Focus on practical communication skills
§ Role in promoting national standards
§ Potential limitations in assessing real-world English proficiency
§ Concerns about standardization and test bias

References

Bui, H. P., & Nguyen, T. T. T. (2024). Classroom assessment and learning motivation:
insights from secondary school EFL classrooms. International Review of Applied
Linguistics in Language Teaching, 62(2), 275-300.

Ngo, X. M. (2024). English assessment in Vietnam: status quo, major tensions, and
underlying ideological conflicts. Asian Englishes, 26(1), 280-292.

Nguyen, T. L., & Nguyen, T. N. (2020). The role of learners’test perception in changing
English learning practices: A case of a high-stakes English test at Vietnam national
university, Hanoi. VNU Journal of Foreign Studies, 35(6), 2525-2445.
Chapter 8: Interpreting and Reporting Test Results

In the realm of language assessment, a test score is more than just a numerical value; it is
a window into a learner's linguistic journey, a snapshot of their capabilities at a specific
point in time. This chapter, "Interpreting and Reporting Test Results," serves as a guide
to deciphering the intricate language of English test scores. This chapter delves into the
nuances of various score types, from raw scores and percentiles to proficiency levels. It is
also important to contextualize these scores (Levi & Inbar-Lourie, 2020), considering
factors like test purpose, learner characteristics, and test reliability to paint a
comprehensive picture of individual achievement.

8.1 Interpreting and reporting English test results

An English test score is a coded message containing valuable insights into a learner's
language proficiency. This section presents the tools to crack that code, enabling you to
interpret English test results accurately and report them effectively.

Understanding different score types

English tests employ various scoring methods, each with its own interpretation:

• Raw scores: These represent the total number of correct answers. While seemingly
straightforward, raw scores lack context. They don't indicate how a student
compares to others or against a specific standard.
• Scaled scores: These convert raw scores onto a standardized scale, allowing for
comparison across different test versions or administrations. They often have a
predefined mean and standard deviation.
• Percentile ranks: These indicate the percentage of test-takers who scored lower
than a particular individual. For example, a percentile rank of 70 means the
student performed better than 70% of those who took the test.
• Proficiency levels: Many tests, like the CEFR (Common European Framework of
Reference for Languages), categorize scores into proficiency levels (e.g., A1, B2,
C1). These describe the learner's abilities in real-world contexts.

Interpreting scores in context

Interpreting test results requires going beyond the numbers and considering the context:

• Purpose of the test: What was the test designed to measure? Was it for university
admission, job applications, or general proficiency assessment? The purpose
influences how scores should be interpreted.
• Test-taker characteristics: Factors like age, educational background, first language,
and learning experiences can all affect test performance.
• Test reliability and validity: How reliable and valid is the test itself? A reliable test
produces consistent results, while a valid test measures what it intends to
measure.
• Individual strengths and weaknesses: Analyze performance across different test
sections (e.g., reading, writing, listening, speaking) to identify areas of strength
and areas needing improvement.

Reporting test results effectively

Clear and informative reporting is crucial for conveying test results to stakeholders:

• Target audience: Tailor the report to the audience (e.g., students, parents, teachers,
administrators). Use language and visuals that are appropriate for their
understanding.
• Key information: Include essential details like the test name, date of
administration, and the student's score(s).
• Interpretation: Explain what the scores mean in plain language, avoiding technical
jargon. Relate the scores to the test's purpose and the student's individual context.
• Recommendations: Provide specific and actionable recommendations for
improvement. This might include targeted language learning activities, further
testing, or support services.
• Ethical considerations: Maintain confidentiality and ensure that test results are
used fairly and responsibly.

Beyond the numbers

While test scores provide valuable information, they do not capture the whole picture of a
learner's language proficiency (Shohamy, 2020). Consider these additional factors:

• Classroom performance: Observe the student's active language use in class, their
participation in discussions, and their completion of assignments.
• Portfolio assessment: Collect samples of the student's work over time to track their
progress and showcase their achievements.
• Self-assessment: Encourage students to reflect on their own language learning
journey and set goals for improvement.

By combining quantitative test data with qualitative observations and individual


reflection, we can gain a more holistic understanding of a student's English language
abilities and provide them with the support they need to succeed.
8.2 Utilizing test outcomes for educational decision-making

English test results are not merely an endpoint, a final grade to be recorded and filed
away. Instead, they should be viewed as a starting point, a springboard for informed
educational decision-making that benefits both students and educators (Coombe, et al.,
2020). This chapter explores how to effectively use test outcomes to drive improvements
in English language teaching and learning.

Individualized learning paths

Test results can illuminate individual strengths and weaknesses in language skills. This
knowledge allows educators to:

• Tailor instruction: Adapt teaching methods and materials to address specific needs
identified in the test results. For example, a student struggling with reading
comprehension might benefit from targeted exercises on vocabulary building and
inferencing.
• Provide differentiated support: Offer individualized support and resources, such as
extra practice materials, tutoring sessions, or online learning tools, to help students
overcome their challenges.
• Set personalized goals: Work with students to set realistic and achievable goals
based on their test performance and learning aspirations.

Classroom-level interventions

Analyzing test results at the classroom level can reveal patterns and trends that inform
instructional adjustments (Ismail, et al., 2022). This might involve:

• Revisiting teaching strategies: If a significant number of students struggle with a


particular skill (e.g., writing), it may signal a need to revisit teaching strategies or
explore alternative approaches.
• Adjusting curriculum content: Test outcomes can highlight areas where the
curriculum needs to be strengthened or expanded. For example, if students
consistently underperform in listening comprehension, the curriculum might need
to incorporate more authentic listening materials and activities.
• Modifying assessment practices: If test results consistently fail to reflect students'
actual abilities, it might be necessary to revise assessment methods to ensure they
are valid, reliable, and aligned with learning objectives.
School-wide improvement initiatives

Aggregating and analyzing test data across different classes and grade levels can provide
valuable insights for school-wide improvement efforts. This can lead to:

• Identifying areas of need: Pinpointing areas where students across the school
require additional support, such as vocabulary development, grammar accuracy, or
oral fluency.
• Allocating resources effectively: Directing resources and professional
development opportunities towards addressing the identified needs. This might
involve investing in new learning materials, providing teacher training on specific
teaching strategies, or establishing peer tutoring programs.
• Monitoring progress over time: Tracking student performance on English tests
over time can help assess the effectiveness of school-wide interventions and
inform continuous improvement efforts.

Informing educational policies

On a broader scale, large-scale English test results can inform educational policy
decisions at the district, regional, or national level (Todd, et al., 2021). This data can be
used to:

• Set benchmarks and standards: Establish clear benchmarks and standards for
English language proficiency at different educational levels.
• Evaluate program effectiveness: Assess the effectiveness of language teaching
programs and initiatives.
• Allocate funding and resources: Guide the allocation of funding and resources to
support English language education.
• Promote educational equity: Identify and address disparities in English language
proficiency among different student populations, ensuring that all students have
access to quality language education.

Ethical considerations

When utilizing test outcomes for decision-making, it is crucial to uphold ethical


considerations:

• Avoid high-stakes decisions based solely on test scores. Consider other factors,
such as classroom performance, student effort, and individual learning styles.
• Protect student privacy and confidentiality. Handle test data responsibly and
ensure that it is used ethically and in accordance with relevant regulations.
• Communicate results clearly and transparently. Provide students, parents, and
other stakeholders with clear and understandable explanations of test results and
their implications.
• Use test data to support, not label, students. Focus on using test outcomes to
identify areas for improvement and provide targeted support, rather than labeling
students based on their scores.

By embracing a data-driven approach while maintaining ethical practices, educators can


leverage English test outcomes to create a more responsive, effective, and equitable
learning environment for all students.

8.3 Student self-assessment

Instruction: Carefully read each statement in the table below. For each statement, rate
your understanding using the following scale:

1: I do not understand this at all.


2: I understand this a little, but I need more help.
3: I understand this fairly well, but I have some questions.
4: I understand this very well and can explain it to others.

Learning Objective/Concept Self- Evidence/Notes Action Plan


Assessment (Explain your (What will
Rating (1-4) rating) you do to
improve?)
Score types: I can define and
differentiate between various types
of test scores, including raw scores,
scaled scores, percentile ranks, and
proficiency levels.
Reporting test results: I can discuss
the principles of effective reporting
of test results, including clarity,
accuracy, and providing meaningful
feedback to test-takers.
Ethical considerations: I can
identify and discuss the ethical
considerations related to interpreting
and reporting test results, such as
confidentiality, fairness, and
avoiding misinterpretation.
Overall understanding: I feel
confident in my ability to interpret
and report language test results
accurately and ethically, considering
the various score types and
contextual factors involved.

8.4 Consolidation activities

Activity 1: Score showdown: raw vs. percentile

Objectives:

• To help students understand the difference between raw scores and percentile
scores in the context of English language tests.
• To develop their ability to interpret and compare these two types of scores.

Procedure: Students begin by reviewing the concept of raw scores or the number of
correct answers on a test. They then discuss the percentile scores indicating the
percentage of test-takers who scored lower than a particular individual. Students then
discuss in pairs:

o Scenario 1: Student A got a raw score of 30. Student B got a raw score of
40. Who performed better?
o Scenario 2: Student C got a percentile score of 75. Student D got a
percentile score of 90. Who performed better?
o Scenario 3: Student E got a raw score of 35, which is equivalent to the 50th
percentile. What does this mean?
o Scenario 4: Student F got a raw score of 50, which is equivalent to the 99th
percentile. What does this mean?

Activity 2: Decoding the data: Analyzing high school English test scores

Objectives:

• To provide students with hands-on experience in interpreting real English test


scores of high school students in Vietnam.
• To develop their ability to analyze score reports, identify trends, and draw
meaningful conclusions.
• To encourage critical thinking about factors that may influence test performance
and educational decision-making.
Procedure: Students in groups examine the following report of the national English test in
Vietnam in 2024:

(Source: [Link]
[Link])

They then discuss the following questions:

• What type of test is this (e.g., national exam, VSTEP, IELTS)?


• What are the different sections of the test and what skills do they assess?
• What are the overall scores and how are they presented (e.g., raw scores,
percentile ranks, proficiency levels)?
• Are there any noticeable patterns or trends in the scores across different sections
or students?
• What are some possible explanations for the observed score patterns?
• What are some potential implications of these scores for the students' future
educational or career paths?

Activity 3: VSTEP level up: Decoding proficiency descriptors

Objectives:

• To help students understand the proficiency levels of the VSTEP exam


(Vietnamese Standardized Test of English Proficiency) and their associated
descriptors.
Procedure: Students visit the official VSTEP proficiency level descriptors (available on
the VSTEP website [Link] Students then

o Identify the key skills and abilities associated with that level in each
domain (listening, speaking, reading, writing).
o Paraphrase the descriptors in simpler terms, making them easier to
understand.

Activity 4: National exam deconstruction: Analyzing English test reports

Objectives:

• To familiarize students with the content of the National English Test score report
in Vietnam.
• To develop their ability to interpret the different sections of the report and
understand their implications.

Procedure: Students read the following report:

“Môn Tiếng Anh có 906.549 thí sinh dự thi, ít hơn 139,094 bài thi so với môn
Toán, có thể do một số thí sinh dùng môn ngoại ngữ khác (Pháp, Trung, Nga,
Hàn,…) hoặc dùng kết quả chứng chỉ IELTS, TOEIC, TOEFL,… thay thế bài thi.
Có tới 145 bài thi bị điểm liệt, chiếm tỉ lệ 0,016%, là giảm so với năm 2023
(0,022%)14 bài thi 0 điểm. Điểm của môn Tiếng Anh nhìn chung gần giống năm
ngoái, điểm trung bình là 5,51 và điểm trung vị là 5,2. Số điểm dưới 5 là 386.861
thí sinh, chiếm 42,674% và điểm học sinh đạt nhiều nhất là 4,6. Năm nay ghi nhận
565 thí sinh đạt điểm 10. Tiếng Anh vẫn là môn thi có kết quả thấp nhất trong các
môn từ trước tới nay.
Tổ hợp thí sinh thi các môn Khoa học tự nhiên so với tổ hợp thi Khoa học xã hội
năm nay là 52,77%, có cao hơn tỉ lệ chọn năm 2023 là 50,75%. Kết quả có thể
nhận định về xu hướng chọn ngành nghề học đại học tuy vẫn nghiêng mạnh về các
ngành xét tuyển Khoa học xã hội, nhưng đã có sự dịch chuyển sang khối Khoa học
tự nhiên.” ([Link]
can-cu-quan-trong-lua-chon-dang-ky-nguyen-vong-chuan-
[Link])

The students then discuss the following questions:

• What is the overall English score of high school students in the report?
• What are some implications of the report for the teaching and learning English at
high schools in Vietnam?
Activity 5: Test outcomes in action: A simulation game

Objectives:

• To help students understand how English test outcomes can be used for
educational decision-making in a realistic scenario.
• To develop their ability to analyze test data, identify student needs, and propose
appropriate interventions.

Procedure:

Set the scene: Students begin by describing the educational context for the activity.
For example, they might present a scenario of a high school English teacher who has
just received the results of a recent English test for their class.

Introduce the students: Provide students with a set of fictional student profiles, each
with their own unique background and English test scores. These profiles should
include information about their strengths, weaknesses, learning preferences, and
goals.

Analyze the data: Divide the class into small groups and assign each group a subset
of the student profiles. Have them analyze the test data and answer guiding questions
such as:

o What are the overall trends in the class's performance?


o Which students are performing above, at, or below expectations?
o What are the specific strengths and weaknesses of individual students?
o What factors might be contributing to the observed patterns in the data?

References

Coombe, C., Vafadar, H., & Mohebbi, H. (2020). Language assessment literacy: What do
we need to learn, unlearn, and relearn?. Language Testing in Asia, 10(1), 3.

Ismail, S. M., Rahul, D. R., Patra, I., & Rezvani, E. (2022). Formative vs. summative
assessment: impacts on academic motivation, attitude toward learning, test anxiety,
and self-regulation skill. Language Testing in Asia, 12(1), 40.

Levi, T., & Inbar-Lourie, O. (2020). Assessment literacy or language assessment literacy:
Learning from the teachers. Language Assessment Quarterly, 17(2), 168-182.
Shohamy, E. (2020). The power of tests: A critical perspective on the uses of language
tests. Routledge.

Todd, R. W., Pansa, D., Jaturapitakkul, N., Chanchula, N., Pojanapunya, P.,
Tepsuriwong, S., ... & Trakulkasemsuk, W. (2021). Assessment in Thai ELT: What
Do Teachers Do, Why, and How Can Practices Be Improved?. LEARN Journal:
Language Education and Acquisition Research Network, 14(2), 627-649.

You might also like