0% found this document useful (0 votes)
4 views26 pages

Understanding Criterion Validity in Assessment

B.ed assignment

Uploaded by

aasmamunir73
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views26 pages

Understanding Criterion Validity in Assessment

B.ed assignment

Uploaded by

aasmamunir73
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

ASSIGNMENT 2

EDUCATIONAL ASSESSMENT & EVALUATION (8602)

ALLAMA IQBAL OPEN UNIVERSITY ISLAMABAD


ASSIGNMENT 2

ASSIGNMENT 2: EDUCATIONAL ASSESSMENT AND


EVALUATION

SUBMITTED BY: ASIMA MUNIR


STUDENT ID: 0000620244
COURSE CODE: 8602
LEVEL: [Link] 1.5 YEARS
SEMESTER: 1ST (AUTUMN 2023)

ALLAMA IQBAL OPEN UNIVERSITY ISLAMABAD


QUESTION NO. 1

Write a note on criterion validity, concurrents validity and


Predictive validity.

Introduction
Criterion validity, concurrent validity, and predictive validity are important concepts in the field
of psychology and measurement, particularly in the context of assessing the accuracy and validity
of psychological tests and measures. These concepts provide researchers and practitioners with
valuable tools for evaluating the effectiveness and usefulness of assessment instruments in
predicting behavior, performance, or outcomes.

Criterion Validity
Criterion validity, also known as instrumental validity, measures the quality of the measurement
methods. This quality is demonstrated by comparing a measurement with a measure that is already
known to be valid in the real world.
Criterion validity is essential for establishing the credibility and trustworthiness of a test or
measure. It ensures that the assessment instrument accurately captures the construct it aims to
assess, allowing for valid and reliable predictions or estimates of relevant criteria or outcomes. A
strong demonstration of criterion validity is indicative of the test’s ability to effectively measure
the intended construct and its relevance to real-world criteria or outcomes.
Criterion validity refers to the extent to which a measure is related to a particular outcome or
criterion. It evaluates how well an assessment tool can accurately predict or estimate an
individual’s behavior, performance, or other relevant criteria. Criterion validity is essential for
determining whether a test or measure is valid in relation to a specific criterion that it aims to
predict or assess.
Criterion validity shows you how well a test correlates with an established standard of comparison
called a criterion. A measurement instrument, like a questionnaire, has criterion validity if its
results converge with those of some other, accepted instrument, commonly called a “gold
standard.”
A gold standard (or criterion variable) measures:
• The same construct
• Conceptually relevant constructs
• Conceptually relevant behavior or performance

Example: Criterion validity


A researcher wants to know whether a college entrance exam is able to predict future academic
performance. First-semester GPA can serve as the criterion variable, as it is an accepted measure
of academic performance.
The researcher can then compare the college entry exam scores of 100 students to their GPA after
one semester in college. If the scores of the two tests are close, then the college entry exam has
criterion validity. When your test agrees with the criterion variable, it has high criterion validity.
However, criterion variables can be difficult to ffind.

Types of Criterion validity


There are two types of criterion validity. They include concurrent and predictive validity, and both
are used to show how a test compares against the gold standard or the criterion. However, there
are some differences between the two:
• To establish concurrent validity, the test scores and criterion variables have to be measured
at the same time
• To establish predictive validity, test scores need to be obtained at one point in time, and the
criterion scores have to be measured at a later time.

Concurrent validity
Concurrent validity shows you the extent of the agreement between two measures or assessments
taken at the same time. It compares a new assessment with one that has already been tested and
proven to be valid. Concurrent validity is a subtype of criterion validity. It is called “concurrent”
because the scores of the new test and the criterion variables are obtained at the same time.
Concurrent validity is particularly important when comparing a new measure to an established
one. By examining the relationship between scores on the new measure and those on a well-
established measure of the same construct, administered at the same time, researchers can assess
the degree to which the new measure accurately captures the intended construct. This form of
validity provides valuable evidence of the accuracy and usefulness of the new measure in
predicting or assessing the same construct as the established measure.

Example: Concurrent validity


In some cases, researchers assess concurrent validity by administering two tests simultaneously.
They want to determine whether a new test correlates with a previously validated test. Frequently,
they compare two tests to see if they can replace an old test with a new one. The new one might
be easier, cheaper, or faster to implement, but they must ensure it is valid.

For example, a school pays external specialists to evaluate the teachers. However, the
administrators ask the students to assess the teachers at the same time as the specialists because
they are considering a less expensive process. If the student evaluations correlate highly with the
professional assessments, the new process exhibits concurrent validity.

Predictive validity
Predictive validity, another form of criterion validity, examines the extent to which a measure
accurately predicts future performance or behavior. This type of validity is crucial in determining
whether a test or measure can effectively forecast an individual’s future outcomes, behaviors, or
performance based on their current scores on the assessment.
Predictive validity, another form of criterion validity, assesses the extent to which a measure
accurately predicts future performance or behavior. It examines the ability of a test or measure to
forecast an individual’s outcomes, behaviors, or performance based on their current scores on the
assessment, thus demonstrating its predictive validity.
Predictive validity is crucial for determining the ability of a measure to forecast future outcomes,
behaviors, or performance. This type of validity provides researchers and practitioners with
valuable insights into the effectiveness of a test or measure in predicting an individual’s future
behavior or performance based on their current scores. Demonstrating predictive validity is
essential for establishing the practical utility of an assessment instrument in real-world settings.
For predictive criterion validity, researchers often examine how the results of a test predict a
relevant future outcome. For example, the results of an IQ test can be used to predict future
educational achievement. The outcome is, by design, assessed at some point in the future.

Example: Predictive validity


Suppose you want to find out whether a college entrance math test can predict a student’s future
performance in an engineering study program.
A student’s GPA is a widely accepted marker of academic performance and can be used as a
criterion variable. To assess the predictive validity of the math test, you compare how students
scored in that test to their GPA after the first semester in the engineering program. If high test
scores were associated with individuals who later performed well in their studies and achieved a
high GPA, then the math test would have strong predictive validity.

How To Assess Criterion Validity


Criterion validity can be tested in various situations and for various reasons. Some common
methods for testing criterion validity include (Fink, 2010):
• Comparing the results of the test to another similar test that is known to be valid. One
potential pitfall of this method is that both tests may contain measurement errors, which
would make it difficult to determine the validity of either test.
• Ask experts in the field to rate the items on the test according to how well they think each
item measures the construct being tested. This method can be time-consuming and
expensive, and it may be difficult to find experts.
• Using the test results to predict some other outcome that is known to be related to the
construct being measured (e.g., job performance)
• Conducting a factor analysis of the items on the test to see if they cluster together in a way
that makes sense theoretically.
It is important to note that no single method is definitive and that multiple methods should be used
whenever possible. Generally, testing of criterion validity requires a ‘gold standard’ — a definite
example of the thing that a researcher is setting out to measure. However, in psychology and
psychiatry, such “gold standards” are not physical or biological
These three types of validity are crucial in the validation process of psychological tests and
measures. They provide researchers, practitioners, and clinicians with valuable tools for evaluating
the accuracy and effectiveness of assessment instruments in predicting behavior, performance, or
outcomes.
In summary, criterion validity, concurrent validity, and predictive validity play integral roles in the
validation of psychological tests and measures. These concepts provide researchers, practitioners,
and clinicians with essential tools for evaluating the accuracy and effectiveness of assessment
instruments in predicting behavior, performance, or outcomes. Understanding and applying these
types of validity is crucial for ensuring the credibility and trustworthiness of psychological
measures and their ability to accurately capture the constructs they aim to assess.

QUESTION NO. 2

Write a detailed note on scoring objective type test items.

Scoring objective type test items is an essential part of the assessment process in educational,
professional, and research settings. Objective type test items, such as multiple-choice questions,
true/false statements, and matching items, provide a standardized and efficient way to evaluate the
knowledge, understanding, and skills of test takers. Proper scoring of objective type test items is
crucial for ensuring the validity, reliability, and fairness of the assessment results. In this note, we
will explore the various aspects of scoring objective type test items, including understanding the
different types of objective items, best practices for scoring, and considerations for scoring
accuracy.
Types of Objective Type Test Items:
Objective type test items are designed to elicit specific responses from test takers, making them
suitable for measuring factual knowledge, understanding of concepts, and application of
principles. The most common types of objective test items include multiple-choice questions,
true/false statements, and matching items. Each of these types requires specific scoring methods
and considerations.

1) Multiple-Choice Questions:
Multiple-choice questions present test takers with a stem or question, followed by a set of options
from which to choose the correct response. Scoring multiple-choice questions involves assigning
points to test takers based on their selection of the correct answer. In some cases, partial credit may
be awarded for selecting partially correct options.

2) True/False Statements:
True/false statements require test takers to identify whether a given statement is true or false.
Scoring true/false items typically involves awarding full credit for correct responses and no credit
for incorrect responses. In some cases, partial credit may be awarded for partially correct responses
if the scoring rubric allows for it.

3) Matching Items:
Matching items require test takers to correctly pair items from two columns or sets. Scoring
matching items involves determining the number of correct pairs identified by the test taker and
assigning points accordingly. It is essential to ensure that the matching pairs are scored accurately
to reflect the test taker’s understanding of the relationships between the items.
4) Completion type items
This test consists of a series of items which requires the test to fill a word or phrase on the blanks.
This is also called filling the blank type of test.

5) Identification Type Items


An identification type of test in a form of completion test which is defined, described, explained
or indicated by a picture, diagram or a concrete object and the term referred to is supplied by the
pupil or student.

Merits of Objective Type Test:


o Objective type test gives scope for wider sampling of the content.
o It can be scored objectively and easily. The scoring will not vary from
time to time or from examiner to examiner.
o This test reduces (a) the role of luck and (b) cramming of expected
questions. As a result, there is greater reliability and better content
validity.
o This type of question has greater motivational value.
o It possesses economy of time, for it takes less time to answer than an
essay test. Comparatively, many test items can be presented to students.
It also saves a let of time of the scorer.
o It eliminates extraneous (irrelevant) factors such as speed of writing,
fluency of expression, literary style, good handwriting, neatness, etc.
o It measures the higher mental processes of understanding, application,
analysis, prediction and interpretation.
o It permits stencil, machine or clerical scoring. Thus scoring is very easy.
o Linguistic ability is not required.

Limitations of Objective Type Test:


o Objectives like ability to organise matter, ability to present matter logically
and in a coherent fashion, etc., cannot be evaluated.
o Guessing is possible. No doubt the chances of success may be reduced by the
inclusion of a large number of items.
o If a respondent marks all responses as correct, the result may be misleading.
o Construction of the objective test items is difficult while answering them is
quite easy.
o They demand more of analysis than synthesis.
o Linguistic ability of the testee is not at all tested.
o Printing cost considerably greater than that of an essay test.

Best Practices for Scoring Objective Type Test Items:


Scoring objective type test items effectively requires adherence to best practices to ensure
consistency, fairness, and accuracy. The following best practices can help educators, test
administrators, and researchers score objective test items with precision and reliability:

1. Clear Scoring Guidelines:


Before administering the test, clear and detailed scoring guidelines should be established to ensure
that scores are assigned consistently across all test takers. These guidelines should outline the
criteria for awarding points, including the correct responses, partial credit considerations, and the
allocation of points for each item type.
2. Training for Scorers:
Scorers responsible for evaluating objective test items should receive thorough training to
familiarize themselves with the scoring guidelines and procedures. Training should focus on
promoting consistency in scoring and addressing any ambiguities or uncertainties related to
specific test items.

3. Utilize Scoring Rubrics:


Scoring rubrics provide a structured framework for assigning points to objective test items. A well-
designed scoring rubric outlines the criteria for awarding points, including specific requirements
for full credit, partial credit, and zero credit. Using scoring rubrics enhances objectivity and
consistency in scoring across different scorers.

4. Double-Scoring or Cross-Scoring:
In high-stakes assessments or situations where scoring accuracy is critical, employing double-
scoring or cross-scoring methods can help validate the accuracy of scores. This involves having
multiple independent scorers evaluate the same set of test items to ensure inter-rater reliability and
minimize scoring errors.

5. Address Ambiguities and Misconceptions:


During the scoring process, scorers should be vigilant for any ambiguities or misconceptions
present in the test items. It is essential to address any inconsistencies or potential sources of
confusion to ensure that test takers are not penalized unfairly due to unclear wording or misleading
options.

Considerations for Scoring Accuracy:


Scoring objective type test items accurately involves considering several factors that can impact
the reliability and validity of the assessment results. Test administrators and scorers should be
mindful of the following considerations to ensure the accuracy of scoring objective test items:

1. Item Difficulty:
The difficulty level of objective test items can influence scoring accuracy. If an item is significantly
more challenging or ambiguous than intended, it may lead to unintended discrepancies in scores
among test takers. Adjusting scoring guidelines to account for item difficulty can help mitigate
potential scoring inaccuracies.
2. Random Guessing:
In multiple-choice questions, the phenomenon of random guessing can affect the accuracy of
scores. Test administrators should consider strategies to address random guessing, such as
implementing penalties for incorrect responses or utilizing advanced item analysis techniques to
identify and exclude items susceptible to guessing effects.

3. Construct Validity:
Scoring objective test items requires careful consideration of construct validity, ensuring that the
assessed content aligns with the intended learning outcomes. Scoring practices should reflect the
construct being measured, and adjustments may be necessary to accommodate variations in test
item interpretations while maintaining the validity of the assessment.

4. Differential Item Functioning:


Scoring accuracy can be influenced by potential biases in objective test items, leading to
differences in performance based on factors such as gender, ethnicity, or socioeconomic status.
Assessing and addressing any differential item functioning can help ensure that scores accurately
reflect test takers’ knowledge and abilities without being confounded by extraneous variables.

Conclusion:
Effectively scoring objective type test items is an essential component of the assessment process,
contributing to the validity, reliability, and fairness of assessment results. By understanding the
different types of objective test items, adhering to best practices for scoring, and considering
factors that impact scoring accuracy, educators, test administrators, and researchers can ensure that
test takers are evaluated accurately and equitably. Ultimately, meticulous attention to detail, clear
scoring guidelines, and ongoing quality assurance measures are crucial for maintaining the
integrity of objective type test item scoring.

QUESTION NO. 3

What are the measurement scales used for test scores?

How things are measured in research is very important as it provides information about how to
interpret those measurements. ‘Scales of measurement’ is a classification that describes the nature
of the information within the values assigned to variables. That is, measurement is generally
described as the assignment of numbers or labels to a variable. To refresh our memory,
measurement is simply the process of assigning numbers or labels to variables that represent
attributes or properties of subjects or treatments. It is important to note that researchers do not
assign numbers or labels randomly. It is rather done based on some set rule that provides
consistency and helps to compare variables that are measured by other researchers.
In Statistics, the variables or numbers are defined and categorised using different scales of
measurements. Each level of measurement scale has specific properties that determine the various
use of statistical analysis. In this article, we will learn four types of scales such as nominal, ordinal,
interval and ratio scale.

What is the Scale?


A scale is a device or an object used to measure or quantify any event or another object.

Scales of Measurements
There are four different scales of measurement. The data can be defined as being one of the four
scales. The four types of scales are:

✓ Nominal Scale
✓ Ordinal Scale
✓ Interval Scale
✓ Ratio Scale

I. Nominal Scale
A nominal scale is the 1st level of measurement scale in which the numbers serve as “tags” or
“labels” to classify or identify the objects. A nominal scale usually deals with the non-numeric
variables or the numbers that do not have any value. It is the simplest type of measurement that
identifies types rather than the amount of something. Labels can be symbols, words, or even
numbers to classify observations.

Characteristics of Nominal Scale


• A nominal scale variable is classified into two or more categories. In this measurement
mechanism, the answer should fall into either of the classes.
• It is qualitative. The numbers are used here to identify the objects.
• The numbers don’t define the object characteristics. The only permissible aspect of
numbers in the nominal scale is “counting.”

Example:
Here are some examples, below. Notice that all of these scales are mutually exclusive (no overlap)
and none of them have any numerical significance. A good way to remember all of this is that
“nominal” sounds a lot like “name” and nominal scales are kind of like “names” or labels.

II. Ordinal Scale


The ordinal scale is the 2nd level of measurement that reports the ordering and ranking of data
without establishing the degree of variation between them. Ordinal data is known as qualitative
data or categorical data. It can be grouped, named and also ranked. Ordinal scales are typically
measures of non-numeric concepts like satisfaction, happiness, discomfort, etc.
Ordinal scales are typically measures of non-numeric concepts like satisfaction, happiness,
discomfort, etc. “Ordinal” is easy to remember because is sounds like “order” and that’s the key
to remember with “ordinal scales”–it is the order that matters, but that’s all you really get from
these.

Characteristics of the Ordinal Scale


• The ordinal scale shows the relative ranking of the variables
• It identifies and describes the magnitude of a variable
• Along with the information provided by the nominal scale, ordinal scales give the rankings
of those variables
• The interval properties are not known
• The surveyors can quickly analyse the degree of agreement concerning the identified order
of variables

Example:
Take a look at the example below. In each case, we know that a #4 is better than a #3 or #2, but
we don’t know–and cannot quantify–how much better it is. For example, is the difference between
“OK” and “Unhappy” the same as the difference between “Very Happy” and “Happy?” We can’t

say.

III. Interval Scale


The interval scale is the 3rd level of measurement scale. It is defined as a quantitative measurement
scale in which the difference between the two variables is meaningful. In other words, the variables
are measured in an exact manner, not as in a relative way in which the presence of zero is arbitrary.
Interval scales are nice because the realm of statistical analysis on these data sets opens up. For
example, central tendency can be measured by mode, median, or mean; standard deviation can
also be calculated.
Like the others, you can remember the key points of an “interval scale” pretty easily. “Interval”
itself means “space in between,” which is the important thing to remember–interval scales not only
tell us about order, but also about the value between each item.

Characteristics of Interval Scale:


• The interval scale is quantitative as it can quantify the difference between the values
• It allows calculating the mean and median of the variables
• To understand the difference between the variables, you can subtract the values between
the variables
• The interval scale is the preferred scale in Statistics as it helps to assign any numerical
values to arbitrary assessment such as feelings, calendar types, etc.

Example:
The classic example of an interval scale is Celsius temperature because the
difference between each value is the same. For example, the difference between
60 and 50 degrees is a measurable 10 degrees, as is the difference between 80
and 70 degrees.
IV. Ratio Scale
The ratio scale is the 4th level of measurement scale, which is quantitative. It is a type of variable
measurement scale. It allows researchers to compare the differences or intervals. The ratio scale
has a unique feature. It possesses the character of the origin or zero points.
Ratio scales provide a wealth of possibilities when it comes to statistical analysis. These variables
can be meaningfully added, subtracted, multiplied, divided (ratios). Central tendency can be
measured by mode, median, or mean; measures of dispersion, such as standard deviation and
coefficient of variation can also be calculated from ratio scales.

Characteristics of Ratio Scale:


• Ratio scale has a feature of absolute zero
• It doesn’t have negative numbers, because of its zero-point feature
• It affords unique opportunities for statistical analysis. The variables can be orderly added,
subtracted, multiplied, divided. Mean, median, and mode can be calculated using the ratio
scale.
• Ratio scale has unique and useful properties. One such feature is that it allows unit
conversions like kilogram – calories, gram – calories, etc.
Example:
Good examples of ratio variables include height, weight,
and duration.
What is your weight in Kgs?
• Less than 55 kgs
• 55 – 75 kgs
• 76 – 85 kgs
• 86 – 95 kgs
• More than 95 kkg
This Device Provides Two Examples of Ratio Scales
(height and weight)

QUESTION NO. 4

Elaborate the purpose of reporting test scores.

Introduction
In education, reporting test scores serves several critical functions. Firstly, it provides educators
with valuable information about students’ academic strengths and weaknesses, enabling them to
tailor instruction to meet individual learning needs effectively. Moreover, test scores contribute to
the evaluation of educational programs and curriculum effectiveness, informing decisions
regarding instructional strategies and resource allocation. Additionally, these scores play a pivotal
role in educational accountability systems, allowing stakeholders to assess school performance and
identify areas for improvement. By reporting test scores, educational institutions can foster
transparency and accountability while promoting data-driven decision-making processes aimed at
enhancing student learning outcomes.

Purpose of reporting test scores.


The purpose of reporting test scores encompasses everything from gauging student performance
to assessing educational effectiveness, informing instructional decisions, and providing crucial
data for policy-making. By reporting test scores, educational organizations and policymakers aim
to ensure that students receive an equitable education, identify areas for improvement, and track
student progress over time. Additionally, reporting test scores allows for meaningful comparisons
between students, schools, and educational systems, ultimately contributing to the improvement
of education on a broader scale.
➢ One of the primary purposes of reporting test scores is to provide an objective measure of
student learning and achievement. Through the reporting of test scores, educational
stakeholders can evaluate how well students have mastered specific skills, knowledge, and
content. This information is crucial for measuring academic growth, identifying areas of
strength and weakness, and assessing student readiness for further education or the
workforce. Furthermore, it enables educators, parents, and students themselves to gain
insights into individual academic performance and make informed decisions about
academic pathways and support systems.

➢ In addition to evaluating student performance, reporting test scores is essential for


assessing the effectiveness of educational programs, curricula, and instructional methods.
By analyzing test scores at the institutional level, administrators and policymakers can
identify patterns and trends in student achievement, determine the impact of instructional
strategies, and pinpoint areas in need of improvement. This data-driven approach to
educational assessment helps to ensure that resources are allocated effectively and that
educational initiatives are targeted towards areas that will have the most significant impact
on student learning.

➢ Furthermore, reporting test scores plays a crucial role in informing instructional decisions.
Teachers and education professionals rely on test score reports to identify areas where
students may require additional support or enrichment. By understanding students’
strengths and weaknesses as revealed by test scores, educators can tailor their instructional
approaches, differentiate their teaching strategies, and create personalized learning
experiences that better meet the diverse needs of their students. Ultimately, the insights
provided by test score reports can lead to more effective teaching practices and improved
student outcomes.

➢ Another critical purpose of reporting test scores is to provide data for accountability and
transparency within the education system. Stakeholders, including parents, policymakers,
and the public, rely on test score reports to gauge the quality of education provided by
schools and school districts. Transparent reporting of test scores enables stakeholders to
hold educational institutions accountable for student outcomes, identify disparities in
educational opportunities, and advocate for improvements in the educational system. This
accountability and transparency are essential for fostering trust and confidence in the
education system and for driving continuous improvement.

➢ Beyond the local level, the reporting of test scores also serves a broader purpose in enabling
comparisons across schools, districts, and educational systems. By providing standardized
measures of student achievement, test score reports allow for meaningful comparisons that
can highlight disparities, showcase best practices, and inform policy decisions.
Comparative data can help identify effective educational interventions, benchmark student
performance against state or national standards, and provide insights into the factors
contributing to variations in student achievement across different settings. As such,
reporting test scores contributes to efforts aimed at reducing achievement gaps and
improving equity in education.
➢ Moreover, reporting test scores is essential for providing data to inform policy-making and
resource allocation. Policymakers and educational leaders use test score reports to identify
areas in need of improvement, develop evidence-based policies, and allocate resources to
address key challenges in the education system. By analyzing test score data, policymakers
can better understand the impact of educational initiatives, design interventions to support
struggling students, and make informed decisions regarding educational funding and
resource allocation.

➢ In addition to informing policy-making, test score reports serve as a means of evaluating


the overall health and effectiveness of the education system. By tracking student
performance over time and across various demographics, test score reports provide
valuable insights into long-term trends, disparities in achievement, and the impact of
educational policies and reforms. This longitudinal data helps to assess the progress and
effectiveness of educational interventions, identify areas for targeted improvement, and
provide evidence for ongoing efforts to enhance educational outcomes for all students.

➢ Furthermore, reporting test scores supports efforts to ensure the alignment of educational
goals and standards. By providing a means to measure student achievement against
established standards and learning objectives, test score reports enable educators and
policymakers to monitor progress towards educational goals, identify areas where
standards may need to be revised or updated, and ensure that students are being held to
rigorous and relevant expectations. This alignment between assessment, standards, and
curriculum helps to ensure that students are receiving a high-quality education that prepares
them for success in further education, careers, and citizenship.

➢ Additionally, reporting test scores plays a key role in providing feedback to students,
parents, and guardians. Test score reports offer valuable information about students’
strengths and areas in need of improvement, helping to guide conversations about academic
progress and setting goals for future learning. This feedback can empower students to take
ownership of their learning, seek out additional support or enrichment opportunities, and
make informed decisions about their educational pathways. Similarly, parents and
guardians can use test score reports to gain insights into their child’s academic
performance, understand how they can support their child’s learning, and engage with
educators to advocate for their child’s educational needs.

Summary
In summary, reporting test scores serves a multifaceted purpose that extends from assessing student
performance to informing policy-making and ensuring educational accountability. By providing
standardized measures of student achievement, test score reports offer valuable insights that enable
educators, policymakers, and stakeholders to evaluate student learning, identify areas for
improvement, and make informed decisions about educational policies, programs, and resource
allocation. Ultimately, the purpose of reporting test scores is to support the continuous
improvement of educational outcomes for all students and to contribute to the overall effectiveness
and equity of the education system.

QUESTION NO.5

Discuss frequently used measures of variability.


A measure of variability is a summary statistic that represents the amount of dispersion in a dataset.
How spread out are the values? While a measure of central tendency describes the typical value,
measures of variability define how far away the data points tend to fall from the center. We talk
about variability in the context of a distribution of values. A low dispersion indicates that the data
points tend to be clustered tightly around the center. High dispersion signifies that they tend to fall
further away.
In statistics, variability, dispersion, and spread are synonyms that denote the width of the
distribution. Just as there are multiple measures of central tendency, there are several measures of
variability.

What is variability
Variability refers to how spread scores are in a distribution out; that is, it refers to the amount of
spread of the scores around the mean. For example, distributions with the same mean can have
different amounts of variability or dispersion.
In the following two histograms, the distribution of scores for Quiz 1 and Quiz 2 are presented.
Despite the equal means (the mean score for both quizzes is 7), the scores on Quiz 1 are more
packed or clustered around the mean, whilst the scores on Quiz 2 are more spread out. Thus, the
differences within the student group were greater on Quiz 2 than on Quiz 1.
What are Measure of variability
There are four frequently used measures of the variability of a distribution:

1. Range
2. Interquartile range
3. Variance
4. Standard deviation.

Range
Range is one of the simplest measures of variation. It’s the lowest point of data subtracted from
the highest point of data. For example, if your highest point is 10 and your lowest point is three,
your range would be seven. The range tells you a general idea of how widely spread your data is.
Because range is so simple and only uses two pieces of data, consider using it with other measures
of variation so you have a variety of ways to measure and analyze the variability of your [Link]
find the range, simply subtract the lowest value from the highest value in the data set.
Range example
You have 8 data points from Sample A.
Data (minutes)
72
110
134
190
238
287
305
324
The highest value (H) is 324 and the lowest (L) is 72.
R=H–L
R = 324 – 72 = 252
The range of your data is 252 minutes.

Interquartile range
The interquartile range gives you the spread of the middle of your distribution.

For any distribution that’s ordered from low to high, the interquartile range contains half of the
values. While the first quartile (Q1) contains the first 25% of values, the fourth quartile (Q4)
contains the last 25% of values.
The interquartile range is the third quartile (Q3) minus the first quartile (Q1). This gives us the
range of the middle half of a data set.

Interquartile range example


To find the interquartile range of your 8 data points, you first find the values at Q1 and Q3.
Multiply the number of values in the data set (8) by 0.25 for the 25th percentile (Q1) and by 0.75
for the 75th percentile (Q3).
Q1 position: 0.25 x 8 = 2
Q3 position: 0.75 x 8 = 6
Q1 is the value in the 2nd position, which is 110. Q3 is the value in the 6th position, which is 287.
IQR = Q3 – Q1
IQR = 287 – 110 = 177
The interquartile range of your data is 177 minutes.

Variance
Variance is the average squared variations of values from the mean. It compares every piece of
value to the mean, which is why variance differs from the other measures of variation. Variance
also displays the spread of the data set. Typically, the more spread out your data is, the larger the
variance. Statisticians use variance to compare pieces of data to one another to see how they relate.
Variance is standard deviation squared, which denotes that values of variance are larger than the
other values. To calculate the variance, simply square your standard deviation.
Variance reflects the degree of spread in the data set. The more spread the data, the larger the
variance is in relation to the mean.
Variance example
To get variance, square the standard deviation.
S = 95.5
S2 = 95.5 x 95.5 = 9129.14
The variance of your data is 9129.14.
Standard deviation

The standard deviation is the average amount of variability in your dataset. It tells you, on average,
how far each score lies from the mean. The larger the standard deviation, the more variable the
data set is.
There are six steps for finding the standard deviation by hand:
• List each score and find their mean.
• Subtract the mean from each score to get the deviation from the mean.
• Square each of these deviations.
• Add up all of the squared deviations.
• Divide the sum of the squared deviations by n – 1 (for a sample) or N (for a population).
• Find the square root of the number you found.

Standard deviation graph


The concept of standard deviation is pretty useful because it helps us predict how many of the
values in a data set will be at a certain distance from the mean. When carrying out a standard
deviation, we assume that the values in our data set follow a normal distribution. This means that
they are distributed around the mean in a bell-shaped curve, as below.
The x-axis represents the standard deviations around the mean, which in this case is 0. The y-axis
shows the probability density, which means how many of the values in the data set fall between
the standard deviations of the mean. This graph, therefore, tells us that 68.2% of the points in a
normally-distributed data set fall between -1 standard deviation and +1 standard deviation of the
mean, σ.

Standard deviation fformula

Below is the standard deviation formula.

Where:

S: sample standard deviation


∑: sum of
X: each value
X̅: sample mean
N: number of values in the sampl.

Standard deviation example


John and his friend Paul argue about the heights of their dogs to properly categorize them as per
the rules of a dog show where various dogs will compete with different heights based on categories.
John and Paul decided to analyze their dogs’ heights’ variability using the concept of standard
deviation.
They have five dogs with all types of heights, so they noted their heights as given below:
The heights of the dogs are 300mm, 430mm, 170mm, 470mm, and 600mm.

Calculation:

Step 1: Calculate the mean:


Mean ( x ) = 300 + 430 + 170 + 470 + 600 / 5 = 394
The red line in the graph shows the average height of the dogs.

Step 2: Calculate the variance:


Variance ( σ^2 ) = 8836 + 1296 + 50176 + 5776 + 42436 / 5 = 21704
Step 3: Calculate the standard deviation:
Standard Deviation (σ) = √ 21704 = 147
Now, using the empirical method, we can analyze which heights are within one standard deviation
of the mean:
The empirical rule says that 68% of heights fall within + 1 time the SD of mean or ( x + 1 σ ) =
(394 + 1 * 147) = (247, 541). i.e. 68% of heights fluctuate between 247 and 541.

You might also like