Instructional Materials for Statistics
Instructional Materials for Statistics
MATERIALS
FOR
ELEMENTARY STATISTICS
AND PROBABILITY
Course Code:
SEMA 30083
compilers who collected materials from different authors. This is not for sale and the compilers
have no intention to profit from this.
2
INTRODUCTION
COURSE OVERVIEW
This Instructional Material deals with the study of elementary statistics and probability. It
focuses on the descriptive statistics, measures of central tendency, normal distribution, different
counting techniques, probability of an event, conditional probability, and the Bayes theorem. The
course also includes hypothesis testing and correlation and regression analysis.
A reflective journal is a place to write down your daily or weekly reflection entries. You can write
about a positive or negative event that you experienced, what it means or meant to you, and what you
may have learned from that experience.
Directions: Following these steps will guide you in our reflective journal requirement.
1. Take a picture of yourself together with what you work on the day you read and work on the
exercises. Choose a clear photo or make the photo clear and visible.
2. Create a simple narration on what you feel, what you learn, the difficulties you encounter while
reading and answering the exercises, and what you did to overcome those difficulties.
4. The number of photos should not be more that 20, assuming you will be working the lessons per
week.
5. Prepare a soft copy or hard copy of it depending on your situation, if you have internet
connection or none.
6. Submit it together with your answers in modules on the agreed date of submission.
7. Make sure your photo-essay meets the criteria on the rubric. The rubric is provided in the next
page.
3
REFLECTIVE JOURNAL RUBRIC
Criteria Above Average Average Below Average
Achievement of Learner has achieved Learner has achieved Learner has attempted
defined learning all learning outcomes, all of learning all of learning
outcomes demonstrated by valid outcomes. This outcomes. This
solutions. These are achievement is achievement is partially
explored to a deep demonstrated by valid demonstrated by
level, revealing answers with solutions. solutions.
advanced
understanding of
outcomes.
Evidence of Clear progression in Overall progression in Understanding of
progression from week understanding of understanding of learning outcomes is at
to week learning outcomes learning outcomes is the same level from
evident from week to evident but no clear week to week.
week, with advanced progression from week
level of understanding to week.
of the topics
encountered.
Presentation The writings are clear. The writings are clear. The writings are not
Solutions are well Solutions are provided clear. Solutions are
structured. for some answers and provided for some
well structured. answers and not well
structured.
4
TABLE OF CONTENTS
TITLE Pages
5
LESSON 1
Nature of Statistics
• Define statistics.
• Distinguish between descriptive statistics and inferential statistics.
• Identify and explain the types of data
• List and describe the four levels of measurement.
• Identify and explain the sampling techniques.
• Discuss the methods of collecting and presenting data
Statistics is the branch of mathematics that examines and investigates ways to process
and analyze the gathered data. It is a scientific body of knowledge that deals with the collection,
organization, presentation, analysis and interpretation of data. It has evolved rapidly and is now
applied in many fields such as education, governance, researches and other studies.
Statistics can be traced back from ancient times. People compiled statistical data with
regard to all sort of things such as age, taxes, commerce, events and others. As time went by,
statistical work has continued to influence the activities of people in a wider scope from describing
important features of the data and analyzing them to meet specific purpose.
Foundation of Statistics
6
The starting point of statistics can be map out from two fields of interest, namely: the game
of chance otherwise known as gambling and the second is political science. In the early periods,
there were incomplete estimates of the population in the Philippines. Population estimates were
based on church records, births, deaths, marriage and other source of information about the
population such as number of residence certificates issued every year. In American time, a more
systematic way of collecting data was established such as;
• Bureau of Customs – collects, tabulates and disseminates statistics on imports and
exports
• Bureau of Agriculture – keeps records on the number of farms, the cultivated land as
well as irrigated areas
• Bureau of Labor – provides the government with the number of employed and
unemployed citizens as well as the different problems inherent in the work
• National Statistics Office – undertakes the census of population and housing
The game of chance is the second origin of statistics which deals with throwing of dice,
playing cards, and tossing of coins. The early gamblers suspected that the occurrence of events
in various game of chance follow certain laws, but being unschooled, they cannot deduct the laws
from it. It is the famous gambler in the name of Chevalier de Mere who proposed the well-known
problem challenge and so he worked it out with Fermat, another mathematician. They used
different methods and solutions to many problems.
Uses of Statistics
Statistics can be used for many purposes which may be described as follows:
1. Statistics can give a precise description. For instance, in the field of education,
statistical tools are used to get information on enrolment, physical facilities, teachers as
well as the financial aspect which are important for a productive administration and
management.
Another example is the government itself. Statistics can provide data for an effective
management of the affairs of the state. A good record of taxes, cost of living, wages,
population, number of employed and unemployed and other related records which are
useful for intelligent decisions and policy making.
2. Statistics can predict the outcome of experiment or the behaviour of an individual.
In forecasting, statistics is needed so that concerned individuals can plan ahead correctly
and can formulate necessary policies needed with the existing conditions. In business and
economics, statistics play an important role in business forecasting, opening business,
market research as well as quality control.
3. Statistics can be used to test hypothesis. Statistics is used to analyze and interpret
numerical data which will be used for decision making.
Division of Statistics
7
Statistics is divided into two fields.
The table below shows some differences between descriptive statistics and inferential statistics.
1. Last year’s total enrollees in College 1. The chance that the enrolees in College
Algebra was 1,237 Algebra will increase by 100 is to 20%
2. The average kilograms of carrots sold 2. A recent study showed that eating carrots
daily by a certain distributor is 1,200kgs can ease stomach spasms.
3. 45% of the bar examinees this year 3. The chance that every batch of bar
passed the board exam. examinees will pass the bar exam is 42%
4. According to a recent survey, every elite 4. It is predicted that the average number of
individual owns 2-3 mobiles. mobiles each individual upgrades will
increase every year.
Scales of Measurement
In statistics and quantitative research methodology, various attempts have been made to
classify variables (or types of data) and thereby develop taxonomy of levels or scale of
measurement. Perhaps the best known are those developed by the psychologist Stanley Smith
Stevens. He proposed four types: nominal, ordinal, interval, and ratio.
• Nominal Scale - The nominal scale is the lowest form of measurement because it does
not capture information about the focal object other than whether the object belongs or
doesn’t belong to a category; either you are a smoker or not a smoker, you attended
8
college or you didn’t, a subject has some experience with computers, an average amount
of experience with computers, or extensive experience with computers. No data is
captured that can place the measured object on any kind of scale say, for example, on a
continuum from one to ten. Coding of nominal scale data can be accomplished using
numbers, letters, labels, or any symbol that represents a category into which an object
can either belong or not belong.
• Ordinal Scale - The ordinal scale has at least one major advantage over the nominal scale.
The ordinal scale contains all of the information captured in the nominal scale but it also
ranks data from lowest to highest. Rather than simply categorize data by placing an object
either into or not into a category, ordinal data give you some idea of where data lie in
relation to each other.
• Interval Scale - Unlike the nominal scale that simply places objects into or out of a category
or the ordinal scale that rank orders objects, the interval scale indicates the distance one
object is from another. In the social sciences, there is a famous example often taught to
students on this distinction.
• Ratio Scale - The scale that contains the richest information about an object is ratio
scaling. The ratio scale contains all of the information of the previous three levels plus it
contains an absolute zero point.
The distinction between interval and ratio scales is an important one in the social sciences.
Although both can capture continuous data, you have to be careful not to assume that the lowest
possible score in your data collection automatically represents an absolute zero point.
Classification of Variables
Qualitative vs Quantitative
9
Quantitative variable is a variable which involves numbers and can be obtained by
counting
Example: age, height, salary
Qualitative variable is an attribute or characteristic
Example: marital status, sex, educational achievement
Discrete and Continous
Discrete is a variable which can assume finite or at most countably infinite number of
values, it is obtained by counting
Example: number of books, total enrolees, drop-outs
Continous is a variable which can assume infinite values within a specified interval and
can be obtained by measurement
Example: weight, money, time
Independent vs Dependent
Independent variable is a variable that the researcher controls or manipulates in
accordance with the purpose of the study
Dependent variable is a measure of the effect of the independent variable/s.
Independent variable is the predictor while the dependent is the variable whose value is being
predicted.
_____1. Statistics is a branch of mathematics that examines and investigates ways to process
and analyze the data gathered.
_____2. There are two origins of statistics, one is the game of chance and the other one is
lottery.
_____3. The two divisions of statistics are ordinal statistics and nominal statistics.
_____4. Descriptive statistics includes the summarizing and analyzing of all the data collected.
_____5. Data in ordinal scale can be ordered or arranged, but no differences between the data
can be taken that are meaningful
II. Tell whether the following statements require the use of descriptive or inferential
10
6. A forecaster predicts the results of national election using the number of votes cast in 10
out 15 municipalities.
7. Brand A pain medicine brings noticeable relief significantly faster than Brand B pain
medicine.
8. Last year’s total attendance at freshmen orientation was 980 students.
9. A politician would like to predict based on an opinion poll, her chance for winning in the
upcoming election.
10. A businessman wants to determine the average weekly income he had in the past 3
weeks.
11
LESSON 2
Example:
There are 3400 law students who took the bar exam.
12
Example: Out of 3,400 law students who took the bar exam, only 1,257 passed.
Sampling techniques can be grouped into how selection of items are made
such as probability sampling and non- probability sampling.
1. Probability sampling
Each member of the population has an equal chance to be included in the sample
This is also known as lottery or fish bowl technique. There are two ways in
using this technique. First is sampling without replacement in which the drawn
papers are no longer returned in the container. The other procedure called
sampling with replacement involves returning to the container the piece of paper
drawn.
When to use
This is preferable to use if the population is not widely spread
geographically. Also, this is more appropriate to use if the population is
more or less homogeneous with respect to the characteristics of the
population.
b. Systematic sampling
Samples are randomly chosen following certain rules set by the researcher.
This involves using the kth member of the population with,
𝑁
k = 𝑛,
13
but there should be a random start.
Example:
Choose a sample size of 100 from N = 500, using systematic random
sampling
500
Step 1. Determine k (period); k = 100 = 5, so this means that you have to include
every 5th member of N after choosing a random start.
Step 2. Randomly choose a number from 1-10. The randomly chosen number will
serve as the random start. You may also use table of random numbers
to choose the random start.
Step 3. If the chosen random start is 10, then the following will comprise the
sample elements.
When to use
This is advisable to us if the ordering of the population is essentially random
and when stratification with numerous data is used.
c. Stratified random sampling
The population is divided into strata or groupings. The samples from each stratum
is drawn independently from the samples from other strata. Samples from each stratum
may be randomly drawn using simple random sampling techniques.
When to use
If the population is such that the distribution of the characteristics of the
respondents under consideration is concentrated in small and spread segment of
the population. Thus, this is preferred to use if precise estimates are desired for
stratified parts of the population and if sampling problems differ in the various strata
of the population.
d. Cluster Sampling
When to use
e. Multi-Stage Sampling
14
Multi-stage sampling represents a more complicated form of cluster sampling in which
larger clusters are further subdivided into smaller, more targeted groupings for the purposes of
surveying. Despite its name, multi-stage sampling can in fact be easier to implement and can
create a more representative sample of the population than a single sampling technique.
Particularly in cases where a general sampling frame requires preliminary construction, multi-
stage sampling can help reduce costs of large-scale survey research and limit the aspects of a
population which needs to be included within the frame for sampling.
Example:
An interviewer may be told to sample 200 females and 300 males between
the age of 45 and 60. This means that the interviewer can simply select who they
want to sample (targeting)
Example: This method can be done by telephone interview to get the immediate
reactions of a certain group of sample for a certain issue.
Example:
15
secondary data. Primary data are data collected directly by the researcher himself. These are
the first-hand or original sources. Secondary data are published data made by other researchers
or entity. Primary data can be gathered through the following:
Disadvantages
➢ time – consuming
➢ different interviewers may understand and transcribe interviews in different ways.
2. Indirect or Questionnaire Method
This is a very commonly used method of collecting primary data. Here information are
collected through a set of questionnaire. A questionnaire is a document prepared by the
investigator containing a set of questions. These questions relate to the problem of inquiry directly
or indirectly. Here, the questionnaires are mailed to the informants with a formal request to answer
the question and send them back. For better response the investigator should bear the postal
charges. The questionnaire should carry a polite note explaining the aims and objective of the
enquiry, definition of various terms and concepts used there. Besides this the investigator should
ensure the secrecy of the information as well as the name of the informants, if required. Success
of this method greatly depends upon the way in which the questionnaire is drafted. So the
investigator must be very careful while framing the questions.
The Questionnaire
In some instances, the authenticity of the data gathered through the indirect or
questionnaire method depends on the questionnaire. Therefore, it is a must that the
questions be carefully worded, free from ambiguity, and designed to achieve the purpose.
16
The following are some characteristics of a good questionnaire:
Types of Questionnaire:
3. Registration Method
4. Observation Method
17
It is a primary method of collecting data by means of direct or indirect contact. As per
Langley P, “ Observations involve looking and listening very carefully.”
Advantages
➢ Collect data where and when an event or activity is happening
➢ Does not rely on people’s willingness to provide information
➢ Directly see what people do rather than relying on what they say or do
Example:
A group of researchers was tasked to survey whether people
from Metro Manila will vote for Mayor Duterte if he is to run for presidency. If there
are 1,000,000 people and .10 margin of error is expected, compute the sample
size.
n = 1,000,000 e = .10
n = N / 1+Ne2
= 1,000,000 / (1+(1,000,000)(.10)2 ) = 1,000,000/(1+1,000,000 x .0001)
= 1,000,000 /101 = 99.99 = 1,000,000/101
n = 9,901 (sample size)
18
A. Categorical Frequency Distribution
The categorical frequency distribution is used to organize nominal-level or ordinal-
level type of data. Some examples where we can apply this distribution are gender,
business type, political affiliation, and others.
Example:
Twenty five applicants were given a performance evaluation appraisal. The
data is
High High High Low Average
Average Low Average Average Average
Low Average Average High High
Low Average Average Average High
Average Low Low High Low
Solution:
Step 1: Construct a table as shown below.
Class Tally Frequency Percent
High
Average
Low
Step 4: Determine the percentage. The percentage is computed using the formula:
%= f/n x 100%, where f is the frequency of the class and n is the total number of value.
Example
19
Let us consider the results of a long quiz in Statistics.
30 28 21 28 29 25 24 27
24 23 25 23 27 28 30 26
22 25 26 27 28 26 25 27
24 22 25 25 27 30 29 30
29 28 29 27 24 22 25 28
Since the highest score is 30 and the lowest is 21, the range is 9. Thus, an ungrouped
frequency distribution table is still possible
The Ungrouped Frequency Distribution Table for the Results of a Long Quiz in Elementary
Statistics
Example 1: The following is the result of a 50-item test in The Teaching Profession. Construct
a frequency distribution and find the following:
a. Range
b. Interval
c. Class limits
d. Class boundaries
e. Relative frequencies
f. Percentages
g. Cumulative frequencies
h. Midpoints
20 40 35 25 25 20 40 40 36 15
20 25 25 30 30 40 25 15 20 40
35 22 26 50 10 10 20 25 25 35
30 30 25 20 20 10 40 45 45 50
29 25 25 20 25 20 25 15 40 35
20
Step 1: Arrange the raw data in ascending or descending order to make it easier to tally the data.
10 10 10 15 15 15 20 20 20 20
20 20 20 20 20 22 25 25 25 25
25 25 25 25 25 25 25 25 26 29
30 30 30 30 35 35 35 35 36 40
40 40 40 40 40 40 45 45 50 50
Generally, the class interval should be equal for all classes. The class interval is generated using
the formula:
𝑅𝑎𝑛𝑔𝑒 𝐻𝑣−𝐿𝑣 40
Suggested Class Interval (i) = 𝑁𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝐶𝑙𝑎𝑠𝑠𝑒𝑠 = 𝑘 = 6 = 6.67 or 7
Where:
Range = Highest value – Lowest value = HV - LV
Class limits - smallest and largest observations (data, events etc.) in each class. Therefore, each
class has two limits: a lower and upper
Class boundaries - midpoints between the upper class limit of a class and the lower class limit of
the next class in the sequence. Therefore, each class has an upper and lower class boundary
Relative frequency - the ratio of the number of times an event occurs to the number of occasions
on which it might occur in the same period.
21
Select a starting point which can be the smallest data value or any number less than the smallest
data value.
We need to add interval to the starting point to obtain the lower limit of next class. Keep adding
until we reach the 6 classes. 10, 17, 24, 31, 38, and 45.
Class Limits
10 – 16
17 – 23 The upper limit for each class may be divided by adding the
24 – 30 class interval to the lower limit minus 1. The nos. 10,17,24,31,38
31 – 37 and 45 are called lower limits while the nos. 16,23, 30, 37, 44
38 – 44 and 50 are referred to as upper limits.
45 – 51
Set the class boundaries in each class. To do this we just need to subtract 0.5 from each lower
class limit and add 0.5 to each upper class limit.
Class Limits Class Boundaries
10 – 16 9.5 – 16.5
17 – 23 16.5 – 23.5
24 – 30 23.5 – 30.5
31 – 37 30.5 – 37.5
38 – 44 37.5 – 44.5
45 – 51 44.5 – 51.5
Step 7: Determine the relative frequency. It is computed by dividing each frequency by the total
frequency.
Class Limits Class Frequency Relative Found by
Boundaries Frequency
22
10 – 16 9.5 – 16.5 6 0.12 6 / 50
17 – 23 16.5 – 23.5 10 0.20 10 / 50
24 – 30 23.5 – 30.5 18 0.36 18 / 50
31 – 37 30.5 – 37.5 5 0.10 5 / 50
38 – 44 37.5 – 44.5 7 0.14 7 / 50
45 – 51 44.5 – 51.5 4 0.08 4 / 50
Step 8: Determine the percentage. It can be found by multiplying 100% each relative frequency.
Class Limits Class Frequency Percentage Found by
Boundaries
10 – 16 9.5 – 16.5 6 12 (6 / 50) x 100
17 – 23 16.5 – 23.5 10 20 (10 / 50) x 100
24 – 30 23.5 – 30.5 18 36 (18 / 50) x 100
31 – 37 30.5 – 37.5 5 10 (5 / 50) x 100
38 – 44 37.5 – 44.5 7 14 (7 / 50) x 100
45 – 51 44.5 – 50.5 4 8 (4 / 50) x 100
Total 50 100
Step 9: Determine the cumulative frequencies. The cumulative frequency can be found by adding
the frequency in each class to the total of frequencies of the classes preceding that class.
Class Limits Frequency Cumulative Found by
Frequency
10 – 16 6 6 6
17 – 23 10 16 6 + 10
24 – 30 18 34 16 + 18
31 – 37 5 39 34 + 5
38 – 44 7 46 39 + 7
45 – 51 4 50 46 + 4
Step 10: Determine the midpoints. The midpoint can be found by getting the average of the upper
limit and the lower limit in each class.
Class Limits Frequency Midpoints Found by
10 – 16 6 13 (10 + 16) / 2
17 – 23 10 20 (17 + 23) / 2
24 – 30 18 27 (24 + 30) / 2
31 – 37 5 34 (31 + 37) / 2
38 – 44 7 41 (38 + 44) / 2
45 – 51 4 43 (45 + 50) / 2
The midpoint may also be computed by first setting the average of the upper limit and
lower limit of the first class (10 + 16)/2 = 13 and continuously add the class interval for succeeding
classes.
Data presented in a grouped frequency distribution are easier to analyze and describe.
However, the identity of individual score is lost due to grouping. For instance, in a class of 21-
24, no one can identify the test score which falls in the said class unless somebody refers to the
original set of data.
In the final presentation of the table the tally is omitted.
23
Another way of presenting data using table is the stem-and-leaf display. It is called
a stem-and-leaf because the grouping forms a “stem” and the values are listed as
“leaves”. It is a way of listing relatively small sets of numerical data. It has two-column
table in which the stems are written on the left column and the leaves on the second
column.
20 25 25 30 30 40 25 15 20 40
35 22 26 50 10 10 20 25 25 35
30 30 25 20 20 10 40 45 45 50
29 25 25 20 25 20 25 15 40 35
n = 50
Solution
Stem Leaf
1 0, 0, 0, 5, 5
2 0, 0, 0, 0, 0, 0, 0, 0, 0, 2, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 6, 9
3 0, 0, 0, 0, 5, 5, 5, 5, 6
4 0, 0, 0, 0, 0, 0, 0, 5, 5
5 0, 0
The stem is the tens digit (or the leading digits) while the leaf is the unit digit (trailing digits)
Example
Below are scores of freshmen students in English exam
12 40 15 19 40 31
23 34 37 23 33 36
24
25 26 29 30 21 28
32 29 20 34 26 38
The stem-and-leaf display is shown below
Stem L e a v e s
1 2 5 9
2 0 1 3 3 5 6 6 9 9
3 0 2 3 4 4 7
4 0 0
Data Presentation
Presenting data through pictures or graphics is sometimes more appealing than using
texts. Graphs give more comprehensive but plain view of the numerical relationships even without
going into the longer time of textual, tabular and graphical.
In order for the researcher to interpret the data gathered easily, the researcher must organize the
data. There are two ways of organizing data. One is grouped data. This are data that are
organized and arrange into different classes or categories. While the other one is ungrouped
data which are data that are not organized, or if arrange can only be ascending to descending or
descending to ascending.
➢ Textual Presentation
Data presented in phrase, sentences or paragraph form are said to be in the
textual. However, textual presentation would not be of much use since it makes dull
reading and may not give a good interpretation of the meaning.
Among the 17 regions, CALABARZON (Region- IVA) had the largest population with
12.61M, followed by the National Capital Region (NCR) with 11.86M and Central Luzon
(Region III) with 10.41M. The population of these three regions together comprised more
than one-third (37.47 percent) of the Philippine population.
Among provinces, Cavite had the largest population with 3.09M. Bulacan had the
second largest with 2.92M and Pangasinan had the third largest with 2.78M. On the other
hand, the provinces with a population of less than 100,000 persons were Batanes (16,604),
Camiguin (83,807), and Siquijor (91,066).
As projected by NSO, this 2012 the Philippine population is about 97.6M and by
2014, the population may increase to 101.2M.
*Source: National Statistics Office, Manila
April 04, 2012
25
➢ Tabular Presentation
Table 3.0 Table 3.1
Graphical Presentation
Graphical form is the most effective way to present results in a study since it shows the
statistical values and relationship in a pictorial or diagrammatic form. Graphs give more
comprehensive but plain view of the numerical relationships even without going into the longer
time of reading discussions.
BAR GRAPHS
It is a graphical display of data using bars of different heights. It is displayed either horizontally or
vertically and the bars do not touch each other.
26
Note that in the bar graph, the bars do
not touch each other. This indicates
the discrete nature of the variable
being grouped. Bar graphs are often
used to show the frequencies of
various nominal variables. They are
used to compare magnitude.
HISTOGRAM
A histogram is a graphical representation of the distribution of data. It is an estimate of
the probability distribution of a continuous variable and was first introduced by Karl Pearson. The
rectangles of a histogram are drawn so that they touch each other to indicate that the original
variable is continuous.
27
LINEAR GRAPHS
➢ Frequency Polygon
Frequency polygons are a graphical device for understanding the shapes of distributions.
They serve the same purpose as histograms, but are especially helpful for comparing sets of data.
Frequency polygons are also a good choice for displaying cumulative frequency distributions.
➢ Frequency Ogive
28
HUNDRED PERCENT CHARTS
➢ Pie Chart
STATISTICAL MAPS
A special type of map in which the variation in quantity of a factor such as rainfall,
population, or crops in a geographic area is indicated; a dot map is one type.
PICTOGRAMS
an ideogram that conveys its meaning through its pictorial resemblance to a physical
29
object. Pictographs are often used in writing and graphic systems in which the characters
are to a considerable extent pictorial in appearance.
Often mathematical formula require the addition of many variables Summation or sigma
notation is a convenient and simple form of shorthand used to give a concise expression for a
sum of the values of a variable.
Let x1, x2, x3, …xn denote a set of n numbers. x1 is the first number in the set. xi represents the
ith number in the set.
25 10 5 15 15 10 5 5
25 20 25 25 10 15 5 10
20 15 15 10 10 5 5 20
II. Identification
1. A ___________ series graph represents data that occur over specific period of time under
observation. In addition, it shows for a trend orb pattern in increase or decrease over
period of time
2. A __________ is similar to bar histogram. He bases of the rectangles are arbitrary
intervals whose centre are the codes. The height of each rectangle represents the
frequency of that category. It is also applicable
3. A __________ is used to examine possible relationships between two numerical variables.
The two variables are plot in x-axis and y-axis
4. A __________ is a graph that displays the data using points which are connected by lines.
The frequencies are represented by the heights of the points at the midpoint of the classes.
The vertical axis represents the frequency of the distribution while the horizontal axis
represents the midpoint of the frequency distribution
30
5. The ___________ is used to organized nominal-level or ordinal-level type of data. Some
examples where we can apply this distribution are gender, business type, political
affiliation, and others
III. True or False
Write True on the blank if the statement is correct and false if not.
Solution:
Stem L e a v e s
2. One of the leading state universities gathered information about the actual number of enrollees
per colleges. The information in the table below shows the existing colleges and the number
of enrollees respectively. Construct a pie chart and bar graph.
31
College of Education 3, 205
College of Computer and Information Sciences 2, 726
College of Engineering 2, 980
College of Architecture and Fine Arts 1, 517
College of Arts and Letters 2, 892
College of Science 2, 890
a. Pie chart
b. Bar graph
32
LESSON 3
• Calculate the mean, median, mode and midrange as measures of the center of a given
data distribution.
• Discuss the properties of mean, median, mode and midrange.
• Introduce quantiles and the types of distribution
• Compute the weighted mean, geometric mean and combined mean.
• Calculate the mean, median and mode as measures of the center of a given data
distribution.
Measure of central tendency is usually called average. It is a single value that represents
a data set. It is the locator of the center of the data set. We will illustrate how to compute or
calculate it. In this section you can understand what is the purpose of mean, median and mode
for the grouped data, including how to analyze and compute quartile, decile and percentile for the
grouped data also.
Measures of Central Tendency or Average
Average plays a very important role in our daily life and it is an important tool in statistics.
Mean, median, and mode are three kinds of "averages" or sometimes called measures of central
tendency. There are many "averages" in statistics, but these are, I think, the three most common,
and are certainly the three you are most likely to encounter in your pre-statistics courses, if the
topic comes up at all.
MEAN
Most commonly used measures of central tendency. When we speak of getting the
average, we always refer to the mean.
Sample Mean - The sample mean is obtained by adding all the values in your sample and dividing
by the sample size (which is usually denoted by small n). In mathematical notation, we have
Σ𝑥
𝑥̅ =
𝑛
Notice that the symbol for the sample mean is 𝑥̅ , an x with a bar above, and it is
33
read as x bar.
The symbol Σ is the summation sign in mathematics, which means that you should add up all
the values in your sample.
Example 1
Paul Cedric collects the data on the ages of respondents of doctoral degree in
educational management, and his study yields the following:
25 25 24 32 35 35 35 45 43 42 44
Determine the average age of the respondents.
Solution
The mean is the sum of the ages and then dividing by the total number of
respondents.
25+25+ 24+32+35+35+35+45+43+42+44
mean = 𝑥̅ =
11
385
𝑥̅ = = 35
11
Therefore, the average age is 35 years old.
Example 2
If the ages you collected are 22, 21, 48, and 21, then the sample values are written
as
𝑥1 = 22, 𝑥2 = 21, 𝑥3 = 48, 𝑥4 = 21.
Solution
for this particular example, the mean is
𝑥 +𝑥 +𝑥 +𝑥 22+21+48+21 112
𝑥̅ = 1 2 3 4 = = = 28
4 4 4
Example 3
In the case of the population of students in the classroom with the following ages:
20 21 21 24 25 25 25 25 25 27 27 27 27 27 26 26 26 23 23 28 28 29 29 29
29.
Solution
𝜇=
20+21+21+24+25+25+25+25+25+27+27+27+27+27+26+26+26+23+23+28+28+29+29+29+29
25
642
𝜇 = 25
𝜇 = 25.68
Therefore the average age of the class is 25.68 years old.
Example 4
34
Eight students have scores of 15, 12, 12, 11, 10, 14, 13, and 14. The mean is
15+12+12+11+10+14+13+14 101
𝑥̅ = = 8 = 12.63
8
For group data, it involves organizing n observed values into smaller number of disjoint groups
of values and counting the frequency of each group; it is often presented as a frequency table.
The following are the steps in solving for the mean of grouped data.
1. Find the midpoint for each class. Place them in a column.
2. Multiply the frequency by the midpoint for each class. Place them in another column.
3. Find the sum of the resulting column in step 2.
4. Divide the sum obtained in step 3 by the total number of frequencies. That is,
∑𝑓 ∙ 𝑥𝑚
mean =
𝑛
Example 5
Seventy randomly selected televisions were tested to determine their lifetimes (in months).
The following frequency distribution was obtained.
Classes Frequency
21 – 29 4
30 – 38 10
39 – 47 23
48 – 56 16
57 – 65 12
66 – 74 5
Solution
Step 1 Find the midpoints of each class and place the values on the third
column.
21+29
𝑥𝑛 = 2 = 25, etc.
Step 2 Next, multiply the midpoint by the frequency for each class and
place the results on the fourth column.
(25)(4) = 100, etc.
Step 3 Find the sum of the fourth column.
∑𝑓 ∙ 𝑥𝑚 = 3 520
Step 4 Divide the sum by n.
∑𝑓 ∙ 𝑥 3520
𝑥̅ = 70 𝑚 = 70 = 50.29 𝑚𝑜𝑛𝑡ℎ𝑠
Hence, the mean lifetime of the television sets is 50.29 months. These steps are summarized in
the following table.
35
48 - 56 16 55 880
57 - 65 12 65 780
66 - 74 5 75 375
N = 70 Total = 3520
Example 6
Forty randomly chosen patients with dengue were considered for the study. Determine the
average age of patients affected by dengue.
775
𝑥̅ =
40
𝑥̅ = 19.375 ≈ 19
36
Therefore, the average age of the patients in the hospital affected with dengue is
19 years old.
Example 7
To get the profile of freshmen, 62 students were asked for their daily allowance. Determine
the average allowance given the following distribution:
∑𝑓𝑥
Mean = 𝑥̅ = 𝑛
775
𝑥̅ = 40
𝑥̅ = 19.375 ≈ 19
Therefore, the average age of the patients in the hospital affected with dengue is
19 years old.
MEDIAN
The median of a set of data values is the middle value of the data set when it has been
arranged in ascending order. That is, from the smallest value to the highest value.
Example 6 The grades of 8 students in Mathematics test that had a maximum possible mark
of 50 are given below:
45 22 35 30 38 48 29
Find the median of this set of data values.
Solution: Arrange the data values in order from the lowest value to the highest value:
22 29 30 35 38 45 48
Select the middle value.
22 29 30 35 38 45 48
The fourth value, 35, is the middle value in this arrangement.
Median = 35
Example 7 Ten books were randomly selected and the numbers of pages were recorded as
follows:
37
550, 500, 465, 601, 610, 480, 510, 580, 600, 475
Solution Arrange the data values in order from the lowest value to the highest value:
465, 475, 480, 500, 510, 550, 580, 600, 601, 610
The number of values in the data set is 10, which is even. So, the median is the
average of the two middle values.
5𝑡ℎ 𝑑𝑎𝑡𝑎 𝑣𝑎𝑙𝑢𝑒+6𝑡ℎ 𝑑𝑎𝑡𝑎 𝑣𝑎𝑙𝑢𝑒
Median = 2
510+550 1060
= 2
= 2
= 530
MODE
A statistical term that refers to the most frequently occurring number found in a set of
numbers. The mode is found by collecting and organizing the data in order to count the frequency
of each result. The result with the highest occurrences is the mode of the set. The mode for
grouped data is the modal class. If no number is repeated, then there is no mode for the list.
Example 9 Find the mode for the following list values:
7 7 7 5 8 9 9 9 10
The mode is the number repeated most often. This list has two values that are
repeated three times.
mode = 7 and 9
Example 10 From the data in Example 4, the class with the largest frequency is the third class.
Therefore, this is the modal class.
Class Boundaries Frequency
38
20 – 30 4
30 – 40 10
𝑑1 = 23 − 10 = 13
40 – 50 23
modal class 𝑑2 = 23 − 16 = 7
50 – 60 16
60 – 70 12
70 – 80 5
The mode is the only measure of central tendency that can be used in finding the
most typical case when the data are nominal or categorical.
Grouped Mode
𝑑
𝑚𝑜𝑑𝑒 (𝑀𝑜) = 𝐿𝑀𝑜 + (𝑑 +1𝑑 ) 𝑤
1 2
where 𝐿𝑀𝑜 - lower boundary of the modal class
w - class width
𝑑1 - difference of the frequency of the modal class and the class
preceding it
𝑑2 - difference of the frequency of the modal class and the class
succeeding it
MIDRANGE
This is a rough estimate of the middle. It is found by getting the average of the lowest and
highest value of the data.
𝑙𝑜𝑤𝑒𝑠𝑡 𝑣𝑎𝑙𝑢𝑒+ℎ𝑖𝑔ℎ𝑒𝑠𝑡 𝑣𝑎𝑙𝑢𝑒
𝑀𝑖𝑑𝑟𝑎𝑛𝑔𝑒 (𝑀𝑅) =
2
Example 12 The calories per serving of 15 fruit juices are as follows:
80, 76, 120, 110, 105, 35, 150, 135, 85, 65,
100, 120, 145, 138, 130
39
Determine the midrange.
Solution The midrange is the sum of the lowest value, 35 and the highest value,
150. Then these are divided by 2.
𝐿𝑉+𝐻𝑉 35+150
𝑀𝑅 = 2
= 2
= 92.5
WEIGHTED MEAN
This is used to find the mean of values of the data set that are not equally represented.
The weighted average can be found by multiplying the value by its corresponding weight and
dividing the sum of the products by the sum of their weights.
Example 13 A recent survey of a new ice cream reported the following percentages of
people who liked the flavor. Find the weighted mean of the percentages.
Area % favored Number surveyed
1 55 1 100
2 25 700
3 70 1 000
𝑘𝑁
−𝑐𝑓
Deciles: 𝐷𝑘 = 𝐿𝐵 + ( 10 𝑓 ) (𝑖)
𝑘𝑁
−𝑐𝑓
Percentile: 𝑃𝑘 = 𝐿𝐵 + (100𝑓 ) (𝑖)
Whereas:
𝑄𝑘 = Quartile
𝐷𝑘 = Decile
𝑃𝑘 = Percentile
N = Population
40
k = Quartile location
LB = Lower Boundary
f = frequency of the quartile class
cf = cumulative frequency before the quartile class
i = class interval
18-26 3 3
27-35 5 8
36-44 9 17
45-53 14 31
54-62 11 42
63-71 6 48
72-80 2 50
Solution:
𝑁 50
𝑄1 = = = 12.5
4 4
36-44 9 17
𝑄1 class falls under this part of the table (highlighted as green)
𝐿𝐵 = 36 – 0.5 = 35.5
cf = 8 just above 𝑄1 ’s cf. (highlighted as light blue)
f=9
i = still 9
𝑘𝑁 50
−𝑐𝑓 −8
𝐿𝐵 + ( 4 ) (𝑖) Substitute 35.5 + ( 4 ) (9)=40
𝑓 9
Hence, the 𝑄1 is 40, the 𝑄1 will fall within the class boundary of 𝑄1 class.
2𝑁 2(50)
Now for 𝑄2 = = = 25
4 4
41
45-53 14 31
𝑄2 class falls under this part of the table (highlighted as gray)
2𝑁 2(50)
−𝑐𝑓 −17
𝐿𝐵 + ( 4 𝑓 ) (𝑖) Substitute 44.5 + ( 4
) (9) = 49.64
14
3𝑁 3(50)
Lastly for 𝑄3 = = = 37.5
4 4
54-62 11 42
𝑄3 class falls under this part of the table (highlighted as red)
3𝑁 3(50)
−𝑐𝑓 −31
4 4
𝐿𝐵 + ( ) (𝑖) Substitute 53.5 + ( ) (9) = 58.82
𝑓 11
Decile Example:
Class limits Frequency (f) Cumulative frequency (cf)
18-26 3 3
27-35 5 8
36-44 9 17
45-53 14 31
54-62 11 42
63-71 6 48
72-80 2 50
We are going to give you an example
𝐷7 (How did I get 𝐷7? This is usually given)
7𝑁 7(50)
𝐷7 = = = 35
10 10
Percentile Example:
Class limits Frequency (f) Cumulative frequency (cf)
18-26 3 3
27-35 5 8
36-44 9 17
42
45-53 14 31
54-62 11 42
63-71 6 48
72-80 2 50
36-44 9 17
𝑃22 falls under this part of the table (Highlighted as green)
𝑘𝑁 22(50)
−𝑐𝑓 −8
𝐿𝐵 + (100𝑓 ) (𝑖) Substitute 35.5 + ( 100
9
) (9) =38.5
4. Fifteen motors were tested and the following data were obtained for the number of
revolutions per minute the flywheel attached to the motors turned.
216 230 245 225 250
255 245 230 235 245
260 245 245 225 230
5. In a Statistics class, 8 test scores were randomly selected, and the following results
were obtained:
85 75 80 81 89 77 77 82
6. A special aptitude test is given to job applicants. The data shown below represent the
43
scores of 25 applicants.
250 260 285 277 251
265 290 295 290 259
280 275 291 258 277
287 292 278 280 300
298 279 258 261 270
Class limits f X fx cf
1-3 5
4-6 3
7-9 6
10 - 12 5
13 - 15 2
16 - 18 4
𝑄1 =
𝑄2 =
𝑄3 =
44
LESSON 4
Measures of Dispersion
• Compute the different measures of variability for both grouped and ungrouped data.
• Discuss the uses, characteristics, advantages and disadvantages of measures of
spread.
• Compute and interpret the coefficient of variation, skewness and kurtosis.
• Differentiate skewness from kurtosis.
• Determine if a data contains outliers.
In summarizing, a given set of data, sometimes, the measures of central tendency alone
are not sufficient to give useful information. They have to be supplemented by other measures
of description, such as measures of variability which indicate the extent to which values in a
distribution are spread around the central tendency.
In this topic, we shall study the measures of variability for both ungrouped and grouped
data. Three measures of variation, namely, the range, variance, and the standard deviation.
These measures describe how item values cluster or scatter in a distribution.
Example I. Scores of some of the ADPR 3-1D Students in their Accounting Subject
Student A B C D E F G H I J
Scores 19 12 15 10 30 25 17 16 23 17
a.) RANGE
The simplest measure of dispersion is the range. The Range is the difference between
the highest and the lowest values in the set of data. Using the table I, let’s have an example
For example, the range of the set of scores are
45
19, 12, 15, 10, 30, 25, 17, 16, 23, and 17
Formula:
Range = Highest Score – Lowest Score
= 30 - 10
R = 20
It is computed from the lowest and highest score, thus it is a very rough measure of
spread. The range provides useful but limited information, since the range depends only on the
extreme scores.
Example II. Scores of the each contestants who joined the quiz bee
Contestant A B C D E F G H I J
Scores 25 32 30 35 50 44 20 46 27 38
Boxes 21 20 15 30 17 18 15 19 26 25
VARIANCE
The measure of dispersion that removes negative signs by squaring all deviations of each
number from its mean and getting the average of the squared deviations is the variance.
The following are the steps to compute the variance:
1. Determine the mean for the given set of data.
2. Determine the deviation from the mean for each value in the data.
3. Square each deviation.
4. Compute the mean of the squared deviations.
46
∑𝑛
𝑖=𝑙(𝑥1 − 𝑥̅ )
2
Variance =
𝑛
Example 1
The variance of the scores of the 6 female Math students is computed and
shown below.
77 83 90 78 80 87
Solution
Using the formula
(77−82.5)2 +(83−82.5)2 +(90−82.5)2 +(78−82.5)2 +(80−82.5)2 +(87−82.5)2
=
6
133.5
= 6 = 22.25
Example 2
Solve the variance using the following data.
10 11 15 13 13
Solution
Using the formula
(10−12.4)2 + (11−12.4)2 +(15−12.4)2 +(13−12.4)2 +(13−12.4)2
=
5
15.2
= 5 = 3.04
In summation notation, the variance in a population is given by
∑𝑛 (𝑥 − 𝑚)2
𝝈𝟐 = 𝑖=𝑙 𝑛1
Where m is the population mean and n is the number of cases.
The formula for the variance in a sample is given by the statistic
∑𝑛
𝑖=𝑙(𝑥1 − 𝑥̅ )
2
𝑠2 =
𝑛−1
where 𝑥̅ is the mean and n-1 is the number of cases which gives an unbiased estimate
of variance.
Calculating the variance is an important part of many statistical applications and
analyses. It is the first step in calculating the standard deviation.
STANDARD DEVIATION
The standard deviation is the most commonly used measure of variation. The standard
deviation indicates how closely the values of a given data set are clustered around the mean. A
lower value of the standard deviation means that the values of the given data set are spread over
a smaller range around the mean. On the other hand, a large value of the standard deviation
means that values of that data set are spread over a large range around the mean.
In summation notation, the formula for the standard deviation in a sample is given by the
parameter
∑(𝑥− 𝜇)2
𝜎= √
𝑁
where:
47
N = number of population 𝜇 = 𝑝𝑜𝑝𝑢𝑙𝑎𝑡𝑖𝑜𝑛 𝑚𝑒𝑎𝑛
Example 3
Using the given in example 1, the standard deviation of the scores of the 6 female
Math students is computed and shown below.
77 83 90 78 80 87
Solution
Using the formula
(77−82.5)2 +(83−82.5)2 +(90−82.5)2 +(78−82.5)2 +(80−82.5)2 +(87−82.5)2
=√
6
133.5
=√ 6
= √22.25 = 4.72
Example 4
The ages of the family recorded living in a house are as follows (in years).
66 31 29 27 57
Find the variance and standard deviation
Solution
1. Find the mean.
66 + 31 + 29 + 27 + 57 210
= = 42
5 5
2. Find the variance.
(66−42)2 + (31−42)2 + (29−42)2 + (27−42)2 + (57−42)2
5
1 316
=263.2
5
√263.2 = 16.22
Example 5
For 120 randomly selected nursing students, the following IQ frequency distribution
were obtained.
Class Limits Frequency
48
90 – 95 4
95 – 100 8
100 – 105 7
105 – 110 14
110 - 115 16
115 – 120 24
120 – 125 20
125 - 130 17
130 – 135 10
Find the variance and standard deviation
Solution
Step 1 Make a table. Find the midpoints of each class. Multiply the
midpoints by the frequency for each class.
Step 2 Multiply the frequency by the square of the midpoint for each
class.
1 2 3 4 5
Class Limits Frequency 𝑥𝑚 𝑓 ∙ 𝑥𝑚 𝑓 ∙ 𝑥𝑚 2
90 – 95 4 92.5 370 34 225
95 – 100 8 97.5 780 76 050
49
100 – 105 7 102.5 717.5 73 543.75
105 – 110 14 107.5 1505 161 787.5
110 – 115 16 112.5 1800 202 500
115 – 120 24 117.5 2820 331 350
120 – 125 20 122.5 2450 253 125
125 – 130 17 127.5 2167.5 276 356.25
130 - 135 10 132.5 1325 175 562.5
𝑠 = √𝑠 2 = 115.39
___________1. It indicates the extent to which values in a distribution are spread around the
central tendency.
50
___________2. The study of the collection, analysis, interpretation, presentation and
organization of data. It deals with all aspects of data including the planning of data collection in
terms of the design of surveys and experiments.
___________4. The positive square root of the variance measures the spread or dispersion
of each value from the mean of the distribution.
___________6. This distribution contains high scores that have low frequency, but does not
contain extreme low scores with corresponding low frequency.
___________7. It allows the variability of scores in two sets of data that do not necessarily
measure the same thing.
___________8. It is the difference between the first quartile (Q₁) and the third quartile (Q₃)
___________9. The mean, median and the mode are equal.
___________10. This distribution contains extreme low scores that have low frequency, but
does not contain extreme high scores with corresponding low frequency.
51
LESSON 5
Normal Distribution
Normal distribution
The normal distribution is immensely useful because of the central limit theorem, which
states that, under mild conditions, the mean of many random variables independently drawn from
the same distribution is distributed approximately normally, irrespective of the form of the original
distribution: physical quantities that are expected to be the sum of many independent processes
(such as measurement errors) often have a distribution very close to the normal. Moreover, many
results and methods (such as propagation of uncertainty and least squares parameter fitting) can
be derived analytically in explicit form when the relevant variables are normally distributed.
The Gaussian distribution is sometimes informally called the bell curve. However, many
other distributions are bell-shaped (such as Cauchy’s, Student's, and logistic). The
52
terms Gaussian function and Gaussian bell curve are also ambiguous because they
sometimes refer to multiples of the normal distribution that cannot be directly interpreted in terms
of probabilities.
The parameter μ in this definition is the mean or expectation of the distribution (and also
its median and mode). The parameter σ is its standard deviation; its variance is therefore σ 2. A
random variable with a Gaussian distribution is said to be normally distributed and is called
a normal deviate.
If μ = 0 and σ = 1, the distribution is called the standard normal distribution or the unit normal
distribution, and a random variable with that distribution is a standard normal deviate.
The normal distribution is a subclass of the elliptical distributions. The normal distribution
is symmetric about its mean, and is non-zero over the entire real line. As such it may not be a
suitable model for variables that are inherently positive or strongly skewed, such as the weight of
a person or the price of a share. Such variables may be better described by other distributions,
such as the log-normal distribution or the Pareto distribution.
The value of the normal distribution is practically zero when the value x lies more than a
few standard deviations away from the mean. Therefore, it may not be an appropriate model when
one expects a significant fraction of outliers—values that lie many standard deviations away from
the mean — and least squares and other statistical inference methods that are optimal for
normally distributed variables often become highly unreliable when applied to such data. In those
cases, a more heavy-tailed distribution should be assumed and the appropriate robust statistical
inference methods applied.
The Gaussian distribution belongs to the family of stable distributions which are the
attractors of sums of independent, identically distributed distributions whether or not the mean or
variance is finite. Except for the Gaussian which is a limiting case, all stable distributions have
heavy tails and infinite variance
One of the most important theoretical distribution in statistics is the normal distribution.
The graph of the normal distribution is a bell-shaped curve called the normal curve as shown in
the figure below. The normal distribution is a tool that can be used to predict the probability of
natural occurrences of things in many fields such as in natural and social sciences as many real
occurrences behaves like the characteristics of the normal curve.
53
PROPERTIES OF THE NORMAL CURVE:
1.) The normal curve is bell-shaped and extends infinitely in both directions.
2.) The tails in the left and the right becomes closer and closer but never interact with the
horizontal axis. In other words, the tails are asymptotic with respect to the horizontal axis.
3.) The mean, the mode and the median are located in the line that divides the normal curve
into two equal parts. This line may be called the line of symmetry because the curve in the
left side of this line is symmetrical to the right side.
4.) The total area under the normal curve is 1.0 or 100%. Since the curve is symmetrical, the
area to the left and to the right of the line of symmetry under the curve are the same and
are equal to 0.5.
5.) From the center of the distribution, the horizontal axis may be divided into at least three
standard scores(z scores) to the left and to the right.
6.) Let a and b be two points along the horizontal axis. The probability between these two
points a and b in the normal distribution is equal to the area between the two points a and
b under the normal curve.
7.) The distribution of the normal curve is unimodal, meaning, there is only one value that
dominates the distribution of values.
THE STANDARD NORMAL DISTRIBUTION
A special feature of the normal distribution is that it is completely determined by its mean,
and standard deviation is given. A normal distribution can be converted into standard normal
distribution by transforming x-scale into z-scale using the formula,
𝑥−𝜇 𝑥−𝑥
z= or 𝑧=
𝜎 𝑆
Where:
z = z value or z score or standard score
𝑥 = given value of a particular variable
µ = population mean
σ = population standard deviation
𝑥 = sample mean
54
S = sample standard deviation
Z-scale normally has a mean µ = 0 (located at the center of the distribution) and variance
equal to 1. It should be noted that the z value can also be converted into x value by transforming
the formula given above into
x = µ + zσ or 𝑥 = 𝑥 + 𝑧𝑆
Examples:
A. Given the following data:
1.) z = 1.26, 𝑥 = 54 and S = 3, find the value of x
Solution:
x = 54+1.26(3)
x = 57.78
Z = -1.0
3.) x = 90, S = 4 and z = 1.75, find the value of 𝑥
Solution:
𝑥 = x – Sz
𝑥 = 90-4(1.75)
𝑥 = 83
B. The mean and standard deviation of a Quiz in Statistics are 85 and 4 respectively. (1)
Determine the z score of the student having a grade of 93, (2)find the grade of a student
corresponding to a standard score of 1.5.
Solutions:
93−85
(1) 𝑧= (2) 𝑥 = 𝜇 + 𝑧𝜎
4
93−85
𝑧= 𝑥 = 85 + 1.5(4)
4
𝑧 = 2.0 𝑥 = 91
55
AREAS UNDER THE NORMAL CURVE
An Important property of the normal distribution is that the area under the normal curve is
1.0 or 100%. The area under the normal curve bounded by two points a and b represents the
probability that x lies between a and b. We simply find areas under the normal curve by using the
z distribution table/standard normal distribution (Appendix A)
Example 1.
Find the area under the normal curve between z = 0 and z = 1.95.
Solution:
Using the z distribution table, the value at the row containing z = 1.9 and the column
containing z =0.05 is 0.4744. Therefore, the area under the normal curve between z = 0 and z =
1.95 is 0.4744.
Example 2.
Find the area under the normal curve between z = 0 and z = -2.37.
Solution:
Since the normal curve is symmetrical, the area under the negative values of z is equal to
the area under the positive values. Again, by using the z distribution table, the value at z = 2.37
is 0.4991. Therefore, the area under the normal curve between z = 0 and z = -2.37 is 0.4911.
Example 3.
56
Find the area under the normal curve between z = - 0.86 and z = -1.53.
Solution:
Using the z distribution table and by property of symmetry, the area from z = 0 to z = -1.53
is equal to 0.4370, and the area from z = 0 to z = - 0.86 is equal to 0.0351. By subtracting the
areas obtained, the area between z = - 0.86 and z = -1.53 is equal to 0.1319.
EXAMPLES:
1. Suppose that the passing grade for a statistics class containing 50 students is 75. If the
mean grade is 85, the standard deviation is 5 and the grades are normally distributed,
(a)what proportion of class passed the subject?, (b)what proportion of the students
obtained the grades between 82 and 88?
Solution:
(a) Convert the grade of 75 into z score.
75−85
𝑧= = −2
5
The proportion of the students who passed the subject is the same as the shaded area
under the normal curve.
Shaded area = (Area between z = 0 and z = - 2) + (Area between z = 0 and
z = + ∞)
57
Shaded area = 0.4772 + 0.5 = 0.9772
Thus, the students who passed the subject is 97.72%
Or
0.9772 x 50 students = 48.86 ≈ 49 students
Shaded Area = (Area between z = 0 and z = - 0.6) + (Area between z = 0 and z = 0.8)
= 0.2257 + 0.2881
= 0.5138
Thus, the proportion of the students obtaining grades between 82 and 89 is 51.38%
Or
0.5138 x 50 students = 25.69 ≈ 26 students
58
9,000−10,000 10,500−10,000
𝑧1 = 500
= −2, 𝑧2 = 500
=1
45−52
𝑧= = −1.7
4
Then,
P(z< -1.75) = P(- ∞< z < 0) – P(-1.75< z < 0)
= 0.5 – 0.4599
= 0.0401
59
c. between z = 3.0 and z = - 3.0
d. between z = 1.29 and z = 2.88
e. to the right of z = 1.56
f. to the left of z = 2.15
3) A normal distribution has a mean of 45 and a standard deviation of 5. Find the z scores of
the following:
a. 58
b. 35
c. 52
d. 44
4) Find the limits that include the middle 30% of a normal distribution which has a mean of
93 and a standard deviation of 4.
5) Suppose that the weights per bag of cement produced by a company follow a normal
distribution. The mean weight is 40 kilograms and the standard deviation is 0.8 kilogram.
What would be the new standard deviation if the company wants to have only 1% of the
cement produced weighing less than 39 kilograms per bag keeping the mean of 40
kilograms?
SKEWNESS
Skewness is a measure that indicates the general shape of the distribution. It is a tool used to
analyze where most of the values are concentrated and to what extent the distribution departs
from the mean.
TYPES OF SKEWNESS:
1. Symmetrical Distribution or Zero Distribution
In this type of distribution, the mean, the mode and the median are equal and are
located at the center of the distribution. The graph of the distribution is symmetrical best
described as the normal curve.
60
2. Right-skewed Distribution or Positively Skewed Distribution
For this type of distribution, the mode is always less than the mean and the median.
Most of the time, the median is less than the mean and the longer tail of the curve is to
the right of the mode.
For this type of distribution, the mean and the median are always less than the
mode. Most of the time, the mean is less than the median and the longer tail of the curve
is to the left of the mode.
61
COEFFICIENT OF SKEWNESS
There are several formulas in solving the measure of skewness. Karl Pearson developed
the following two formulas:
𝑥̅ −𝑀𝑜
1) 𝑠𝑘1 =
𝑠
and
3(𝑥̅ − 𝑀𝑑 )
2) 𝑠𝑘2 =
𝑠
Where:
sk = coefficient of skewness
𝑥̅ = mean
Mo = mode
Md = median
S = sample standard deviation
Examples:
1. Below are scores of ten students in a short quiz in Biology
8 8 8 8 8 8 8 8 8 8
Since the mean and median are the same which is 8, the distribution of scores is
symmetric because sk = 0.
2. The mean, median and standard deviation of a certain set of scores are 26, 23 and 3
respectively. Using the formula, the coefficient of skewness is
3(26−23)
𝑠𝑘 =
3
𝑠𝑘 = 3
62
Therefore, the set of scores is skewed to the right.
3. Given the data set: 2, 5, 6, 8, 9, 9, 10. Solve for the skewness using Pearson’s coefficient of
skewness.
Solution:
2+5+6+8+9+9+10
𝑥̅ = =7
7
Mo = 9
Md = 8
∑(𝑥−𝑥̅ )2
𝑠=√
𝑛−1
𝑠 = √8
𝑠 = 2.8284
Using formula 1,
7−9
𝑠𝑘1 = 2.8284 = − 0.7071
Or using formula 2,
3(7− 8 )
𝑠𝑘2 = = −1.0607
2.8284
The two formulas obtained the negative value of skewness. Notice that the result of formula 2 is
more skewed to the left and is more advantageous to use considering the given data set wherein
the value for the mode occurs only twice.
63
(𝑄3 − 𝑄2 ) − (𝑄2 − 𝑄1 )
𝑠𝑘3 =
𝑄3 − 𝑄1
or
𝑄3 − 2𝑄2 + 𝑄1
𝑠𝑘3 =
𝑄3 − 𝑄1
where:
𝑄1 = Quartile 1
𝑄3 = Quartile 3
KURTOSIS
Another statistical measure that describes the distribution of data is kurtosis. Kurtosis
measures the peakedness or flatness of the distribution relative to the normal distribution.
TYPES OF KURTOSIS:
1. Mesokurtic
A distribution that is neither too flat nor too peaked is known as mesokurtic
distribution. The graph of this type of distribution is the same as in any normal distribution
and is considered to be the baseline for the two other types.
2. Platykurtic
64
A distribution that has a lower peak compared to the mesokurtic distribution is
known as platykurtic distribution. It is characterized by evenly spread data about the center
with a certain degree of flatness to the peak.
3. Leptokurtic
A simple formula for measuring kurtosis can be based on both quartiles and percentiles
and is given by the formula
1
(𝑄3 − 𝑄1 )
𝐾 =2
𝑃90 − 𝑃10
where:
𝑄1 = Quartile 1
65
𝑄3 = Quartile 3
A standard value of K for a normal distribution is 0.263 or 26.3%. This value serves as the
basis to determine the type of kurtosis for a given distribution. For mesokurtic distribution, K is
equal to 0.263. Platykurtic distribution obtain K is less than 0.263 and leptokurtic distribution
obtain K greater than 0.263.
Where:
k = kurtosis
Example:
Thirty BSA students took an examination in English 1. The variance of the exam is 5 and
the sum of the fourth power of the deviation from the mean is 4 500. Is the distribution of scores
of the students mesokurtic, platykurtic or leptokurtic?
4 500
𝑘=
30(5)2
𝑘=6
66
B. Determine the coefficient of skewness and kurtosis of the following data.
1. 7 7 7 7 7 7
2. 6 8 9 9 10 12
3.
(Class Interval)
40-44 3
35-39 5
30-34 11
25-29 13
20-24 6
15-19 2
67
LESSON 6: INTRODUCTION TO PROBABILITY
LEARNING OBJECTIVES
DISCUSSIONS
68
69
70
71
72
LET’S PRACTICE!
73
REFERENCE:
Introduction to Probability and Probability Distributions(2015),by J.B. Ofosu and C.A. Hesse,
pages 1 − 8.
74
LESSON 7: COUNTING SAMPLE POINTS
LEARNING OBJECTIVES
DISCUSSIONS
75
76
77
LET’S PRACTICE!
78
REFERENCE:
Introduction to Probability and Probability Distributions(2015),by J.B. Ofosu and C.A. Hesse,
pages 8 − 13.
79
LESSON 8: PROBABILITY OF AN EVENT
LEARNING OBJECTIVES
DISCUSSIONS
80
81
82
83
84
85
86
LET’S PRACTICE!
87
REFERENCE:
Introduction to Probability and Probability Distributions(2015),by J.B. Ofosu and C.A. Hesse,
pages 13 − 22.
88
LESSON 9: CONDITIONAL PROBABILITY
LEARNING OBJECTIVES
DISCUSSIONS
89
90
91
92
93
94
LET’S PRACTICE!
95
REFERENCE:
Introduction to Probability and Probability Distributions(2015),by J.B. Ofosu and C.A. Hesse,
pages 22 − 31.
96
LESSON 10: BAYE’S THEOREM
LEARNING OBJECTIVES
DISCUSSIONS
97
98
LET’S PRACTICE!
99
REFERENCE:
Introduction to Probability and Probability Distributions(2015),by J.B. Ofosu and C.A. Hesse,
pages 31 − 35.
100
LESSON 11
Hypothesis Testing
The best way to determine whether a statistical hypothesis is true would be to examine the
entire population. Since that is often impractical, researchers typically examine a random sample
from the population. If sample data are not consistent with the statistical hypothesis, the
hypothesis is rejected.
There are two types of statistical hypotheses.
▪ Null hypothesis. The null hypothesis, denoted by H0, is usually the hypothesis that
sample observations result purely from chance.
101
▪ Alternative hypothesis. The alternative hypothesis, denoted by H1 or Ha, is the
hypothesis that sample observations are influenced by some non-random cause.
For example, suppose we wanted to determine whether a coin was fair and balanced. A null
hypothesis might be that half the flips would result in Heads and half, in Tails. The alternative
hypothesis might be that the number of Heads and Tails would be very different. Symbolically,
these hypotheses would be expressed as
H0: P = 0.5
Ha: P ≠ 0.5
Suppose we flipped the coin 50 times, resulting in 40 Heads and 10 Tails. Given this result,
we would be inclined to reject the null hypothesis. We would conclude, based on the evidence,
that the coin was probably not fair and balanced.
▪ Formulate an analysis plan. The analysis plan describes how to use sample data to
evaluate the null hypothesis. The evaluation often focuses around a single test statistic.
▪ Analyze sample data. Find the value of the test statistic (mean score, proportion, t-
score, z-score, etc.) described in the analysis plan.
▪ Interpret results. Apply the decision rule described in the analysis plan. If the value of
the test statistic is unlikely, based on the null hypothesis, reject the null hypothesis.
Decision Errors
102
▪ Type I error. A Type I error occurs when the researcher rejects a null hypothesis when it
is true. The probability of committing a Type I error is called the significance level. This
probability is also called alpha, and is often denoted by α.
▪ Type II error. A Type II error occurs when the researcher fails to reject a null hypothesis
that is false. The probability of committing a Type II error is called Beta, and is often
denoted by β. The probability of not committing a Type II error is called the Power of the
test.
Decision Rules
The analysis plan includes decision rules for rejecting the null hypothesis. In practice,
statisticians describe these decision rules in two ways - with reference to a P-value or with
reference to a region of acceptance.
The set of values outside the region of acceptance is called the region of rejection. If the
test statistic falls within the region of rejection, the null hypothesis is rejected. In such
cases, we say that the hypothesis has been rejected at the α level of significance.
These approaches are equivalent. Some statistics texts use the P-value approach; others use the
region of acceptance approach. In subsequent lessons, this tutorial will present examples that
illustrate each approach.
A test of a statistical hypothesis, where the region of rejection is on both sides of the sampling
distribution, is called a two-tailed test. For example, suppose the null hypothesis states that the
mean is equal to 10. The alternative hypothesis would be that the mean is less than 10 or greater
than 10. The region of rejection would consist of a range of numbers located on both sides of
sampling distribution; that is, the region of rejection would consist partly of numbers that were less
than 10 and partly of numbers that were greater than 10.
One-Tailed
103
Two-Tailed
104
The probability of not committing a Type II error is called the power of a hypothesis test.
Effect Size
To compute the power of the test, one offers an alternative view about the "true" value of
the population parameter, assuming that the null hypothesis is false. The effect size is the
difference between the true value and the value specified in the null hypothesis.
For example, suppose the null hypothesis states that a population mean is equal to 100.
A researcher might ask: What is the probability of rejecting the null hypothesis if the true
population mean is equal to 90? In this example, the effect size would be 90 - 100, which equals
-10.
▪ Sample size (n). Other things being equal, the greater the sample size, the greater the
power of the test.
▪ Significance level (α). The higher the significance level, the higher the power of the test.
If you increase the significance level, you reduce the region of acceptance. As a result,
you are more likely to reject the null hypothesis. This means you are less likely to accept
the null hypothesis when it is false; i.e., less likely to make a Type II error. Hence, the
power of the test is increased.
▪ The "true" value of the parameter being tested. The greater the difference between the
"true" value of a parameter and the value specified in the null hypothesis, the greater the
power of the test. That is, the greater the effect size, the greater the power of the test.
All hypothesis tests are conducted the same way. The researcher states a hypothesis to be tested,
formulates an analysis plan, analyzes sample data according to the plan, and accepts or rejects
the null hypothesis, based on results of the analysis.
▪ State the hypotheses. Every hypothesis test requires the analyst to state a null
hypothesis and an alternative hypothesis. The hypotheses are stated in such a way that
they are mutually exclusive. That is, if one is true, the other must be false; and vice versa.
▪ Formulate an analysis plan. The analysis plan describes how to use sample data to
accept or reject the null hypothesis. It should specify the following elements.
• Significance level. Often, researchers choose significance levels equal to 0.01,
0.05, or 0.10; but any value between 0 and 1 can be used.
105
• Test method. Typically, the test method involves a test statistic and a sampling
distribution. Computed from sample data, the test statistic might be a mean score,
proportion, difference between means, difference between proportions, z-score, t-
score, chi-square, etc. Given a test statistic and its sampling distribution, a
researcher can assess probabilities associated with the test statistic. If the test
statistic probability is less than the significance level, the null hypothesis is
rejected.
▪ Analyze sample data. Using sample data, perform computations called for in the analysis
plan.
• Test statistic. When the null hypothesis involves a mean or proportion, use either
of the following equations to compute the test statistic.
where Parameter is the value appearing in the null hypothesis, and Statistic is the point
estimate of Parameter. As part of the analysis, you may need to compute the standard deviation
or standard error of the statistic. Previously, we presented common formulas for the standard
deviation and standard error.
When the parameter in the null hypothesis involves categorical data, you may use a chi-square
statistic as the test statistic. Instructions for computing a chi-square test statistic are presented
in the lesson on the chi-square goodness of fit test.
1. What would be the best way to determine whether a statistical hypothesis is true?
2. What would be the 2 possible equations needed in determining the test statistics if the null
hypothesis involves a mean or a proportion?
3. What are the three factors that affect the hypothesis test?
4. What is the effect size?
5. How does a researcher conduct a hypothesis test?
6. Define: One Tailed test; two tailed test.
7. What are the two types of errors that can result from a hypothesis test?
8. What is the use of hypothesis testing?
9. What are the four steps in hypothesis testing?
10. What is denoted by null hypothesis?
106
LESSON 12
• Calculate and interpret the coefficient of correlation and the linear regression equation.
• Distinguish and explain the difference between independent and dependent variables.
• Draw a scatter diagram for a set of ordered pairs.
• Conduct test for spearman rank correlation and explain its applications.
• Calculate and interpret the coefficient of determination and the standard error of
estimate.
The word correlation is used in everyday living to denote some form of association. We
might say that we have noticed a correlation between sunny days and shortage of water supply
in Metro Manila. However, in statistical terms we use correlation to denote association between
two quantitative variables. Let us also assume that association is linear, that one variable
increases or decreases a fixed amount for a unit increase or decrease in the other.
The first way to get some idea about a possible relation between two variables is to do a
scatter plot of the variables.
Let us consider the example of possible correlation between height and ratings of physical
attractiveness (1-10) of 10 persons as shown on the table below:
107
63 9
58 7
58 6
57 8
57 7
59 8
10
9
8
7
6
5
4
3
2
1
0
0 10 20 30 40 50 60 70
Physical Att.
Correlation range from -1 (perfect negative relation) through 0 (no relation) to +1 (perfect
positive relation) Scatterplots would look like:
10
0
0 1 2 3 4 5 6 7 8 9 10
108
2.5
1.5
0.5
0
0 1 2 3 4 5 6 7 8 9
No Correlation
12
10
0
0 2 4 Perfect Positive
6 Correlation 8 10 12
The most widely used measure of correlation is the Pearson Product Moment Correlation
Coefficient, commonly called Pearson r (denoted by r). It is a measure of the linear association
between two variables that have been measured on interval or ratio scales, such as the
relationship between height in inches and weight in pounds. The formula for computing the
correlation given two variables (x and y) is given below:
𝑛𝛴𝑥𝑦 − 𝛴𝑥𝛴𝑦
𝑟=
√[𝑛(𝛴𝑥 2 ) − (∑ 𝑥)2 ][𝑛(𝛴𝑦 2 ) − (∑ 𝑦)2 ]
where:
r = degree of relationship between variables x and y
x = observed data for the independent variable
y = observed data for the dependent variable
n = sample size
109
To interpret the degree of linear relationship, the following range of values of coefficient of
correlation may be used:
Coefficient of Interpretation
Correlation
± 1.00 Perfect correlation, perfect relationship
± 0.76 - ± 0.99 Very high correlation, very dependable relationship
± 0.51 - ±0.75 High correlation, marked relationship
± 0.26 - ± 0.50 Low correlation, definite but small relationship
± 0.01 - ± 0.25 Very low correlation, negligible relationship
0.00 No correlation, no relationship
Example 1:
Determine the relationship between height (in inches) and ratings of physical
attractiveness
(1-10) of 10 persons.
Solution:
The independent variable (x) would be the height and ratings of physical attractiveness
will be the dependent variable (y).
X y 𝑥2 𝑦2 xy
56 7 3136 49 392
52 5 2704 25 260
60 7 3600 49 420
61 8 3721 64 488
63 9 3969 81 567
58 7 3364 49 406
58 6 3364 36 348
57 8 3249 64 456
110
57 7 3249 49 399
59 8 3481 64 472
581 72 33837 530 4208
Σy = 72 Σy2 = 530 n = 10
10 (4208) − (581)(72)
𝑟=
√[10(33837) − (581)2 ][10(530) − (72)2 ]
42080−41832
=
√(338370−337561)(5300−5184)
248
=
√(809)(116)
248
=
√93844
= 0.8096
Example 2:
A biologist assumes that there is a linear relationship between the amount of fertilizer
supplied to tomato plants and the subsequent yield of tomatoes obtained. Seven tomato plants,
of the same variety, were selected at random and treated weekly with a solution in which x grams
of fertilizer was dissolved in a fixed quantity of water. The yield, y kilograms of tomatoes was
recorded.
Plant A B C D E F G
x 1.0 1.5 2.0 2.5 3.0 3.5 4.0
y 3.9 4.4 5.8 6.6 7.0 7.1 7.3
Determine the degree of relationship between the amount of fertilizer supplied to tomato plants
and the yield of tomatoes obtained.
Solution :
x y x2 y2 xy
111
1.0 3.9 1.00 15.21 3.90
1.5 4.4 2.25 19.36 6.60
2.0 5.8 4.00 33.64 11.60
2.5 6.6 6.25 43.56 16.50
3.0 7.0 9.00 81.00 21.00
3.5 7.1 12.25 150.06 24.85
4.0 7.3 16.00 256.00 29.20
17.50 42.10 50.75 589.83 113.65
= 0.1708
where :
rho = Spearman rho
Σ = sum of
d = difference between ranks of x and y or d = Rx – Ry
n = number of pairs of observations
Example :
The table shows a Verbal Reasoning test score, x, and an English test score y, for each
of a random sample of 8 students who took both tests. Compute for Spearman rho and interpret
the result.
112
Student A B C D E F G H
X 112 113 110 112 113 115 108 111
Y 68 65 74 70 70 75 68 76
x y Rx Ry d or Rx -Ry d2
112 68 4.5 2.5 2.0 4.00
113 65 6.5 1 5.5 30.25
110 74 2 6 -4.0 16.00
112 70 4.5 4.5 0.0 0.00
113 70 6.5 4.5 2.0 4.00
115 75 8 7 1.0 1.00
108 68 1 2.5 -1.5 2.25
111 76 3 8 -5.0 25.00
Σ d = 82.50
2
6(82.50)
rho = 1− 8(82 −1)
495
= 1− 8(63)
495
= 1−
504
= 1−0.98
= 0.02
For example, if the GNP levels corresponding to levels of total investment are given, one
can estimate the level of GNP corresponding to some previously unspecified level of investment.
Correlation is used when the end goal is simply to find a number that expresses the
relation between the variables.
Regression is used when the end goal is use the measure of relation to predict values of
the random variable based on values of the fixed variable.
About the example of possible correlation between height and ratings of physical
attractiveness, we could now as, “can we predict a persons’ rating of their attractiveness, based
on their height”.
Let us illustrate this problem assuming that the values of y (ratings of physical
attractiveness) depend on the values assumed by x (height). Let us estimate the value of y when
x is 70.
113
X y
56 7
52 5
60 7
61 8
63 9
58 7
58 6
57 8
57 7
59 8
There are two ways of solving this problem. The first way uses graphical approach which
provides a rough estimate. The second method which will give the exact of y when x is 70 uses
the regression formula.
The first method employs the scatter plot, wherein the points corresponding to the values
of x and y will be plotted on the rectangular coordinate system.
10
8
6
4
2
0
0 20 40 60 80
Physical Att.
After plotting the points, draw the trend line. This line approximates the general direction
of the points. It should be drawn such that the sum of the vertical distances of the points below
the line is approximately equal to the sum of the vertical distances of the points above the line.
Also, the trend line need not pass through the first nor last point. It’s just a matter of
coincidence if the line happen to pass through either or both points.
When x is 70, let us estimate now the value of y using the trend line.
114
10
0
0 10 20 30 Physical Att.
40 50 60 70
This estimate of y will slightly vary depending on how accurately the trend line was drawn.
In this example, the rough estimate of y when x is 70. Using the graphical method is 13.
The second method which is the simplest and probably the most widely used equation for
expressing relationships is the equation of the Least Square Regression Line (LSRL) which
follows the form:
y = a + bx
In this equation, a and b are numerical constant, and calculating a predicted value of y for
any value of x by direct substitution would become easier once they are known. The variable of
which is assumed known and that is being used to predict the value of the other variable is x and
is referred to as the independent, explanatory regressor or predictor variable on the other hand,
the variable which is assumed unknown and that is being predicted on the basis of its relationship
with other variable is y and is referred to as the dependent, explained, regressand, or predicted
variable.
The trend line is the regression line. The line is derived in such a way that the sum of the
squares of the vertical deviations between the line and the individual data plots is minimized. All
other lines will yield a higher result.
In the LSRL equation, values of a and b are derived from “normal” equations which are in
turn derived through graphical techniques and the use of calculus. These normal equations are:
1) Σy = aN + bΣx
2) Σxy = aΣx + bΣ𝑥 2
where :
Σy = sum of the values of the dependent variable
N = number of pairs x and y
Σx = sum of the values of the independent variable x
Σxy = sum of the products of paired values x and y
Σ𝑥 2 = sum of squares of the independent variable x
Solving for the values of a and b using the different methods of solving systems of linear
equations, the formulas for a and b are,
(𝛴𝑦)(𝛴𝑥 2 )−(𝛴𝑥)(𝛴𝑥𝑦)
a=
𝑁(𝛴𝑥 2 )−(𝛴𝑥)2
115
𝑁(𝛴𝑥𝑦)−(𝛴𝑥)(𝛴𝑦)
b=
𝑁(𝛴𝑥 2 )−(𝛴𝑥)2
X y Xy 𝑥2
56 7 392 3136
52 5 260 2704
60 7 420 3600
61 8 488 3721
63 9 567 3969
58 7 406 3364
58 6 348 3364
57 8 456 3249
57 7 399 3249
59 8 472 3481
581 72 4208 33837
(72)(33837)−(581)(4208)
a=
10(33837)−(581)
2,436,264−2,444,848
=
338,370−337,561
−8,584
=
809
= −10.61
(10)(4208)−(581)(72)
b=
10(33837)−(581)2
42080−41832
=
338,370−337,561
248
=
809
= 0.31
This method is very useful in providing a fairly accurate estimate when the values of any
two variables are given.
116
Let’s try this!
_____ 2. The scatter diagram is a chart that portrays the correlation between two variables.
_____ 5. A simple regression is a regression model that contains one dependent and one
independent variable.
B. 1. Ms. Besmonte administered two forms of Mathematics test to ten(10) first year students.
The first form of the test was given in the morning and the second form was given in the
afternoon. Their scores in the first and second forms are presented below. Determine the
degree of correlation of the tests.
2. The following shows the scores of 7 students in their first two quizzes in Algebra:
Quiz 1 (x) 7 8 6 5 8 6 9
Quiz 2 (y) 8 8 7 4 9 8 9
3. The table below shows the price (in peso) and demand (in quantity) of a certain commodity:
117
Demand 10 13 15 9 8 7
Price 8 7 9 6 7 5
References:
Agcaoili, Zenaida, et al. Statistics for Filipino Students 2004 Edition, National Book Store
Aguaviva, Erlinda et al. Statistics and Probability with Business Applications 1998 Edition, Jollence Publication
Cruz, Myrna, et al. Statistics and Probability Theory 2011 Edition.
Pangan, Milagros R. et al. Statistics for College Students. Grandwater Publishing and Research, Makati City, 1996
Sirug, W. Basic Probability and Statistics, 2011
Walpole, E. R., Introduction to Statistics, 3rd ed., EDCA, 1982.
Online References:
[Link]
[Link]
regression-in-statistics
[Link]
justification/sampling/
[Link]
EJg:1582628185856&source=lnms&tbm=isch&sa=X&ved=2ahUKEwjShtG2xeznAhWSHHAKH
bypC3MQ_AUoAXoECAwQAw&biw=865&bih=414#imgrc=2qZ1J6K660XnfM
[Link]
regression/
[Link]
[Link]
118
[Link]
[Link]
[Link]
[Link]
119
120
121
A one-tailed test predicts the direction of the effect (either greater or less than a certain value) and rejects the null hypothesis if the test statistic falls into one specified tail of the distribution. In contrast, a two-tailed test considers deviations in both directions, rejecting the null hypothesis if the test statistic falls into either tail. One-tailed tests are used when the research hypothesis specifies a particular direction of effect, while two-tailed tests are applied when any significant differences from the null hypothesis could be critical .
The midrange is calculated by averaging the lowest and highest values in the dataset, providing a quick and rough estimate of the dataset’s central tendency. While simpler than other central tendency measures, it can be easily skewed by outliers; however, it is useful for understanding the dataset’s range and central position quickly .
Quartiles, deciles, and percentiles are measurements that divide data into intervals; quartiles split data into four parts, deciles into ten, and percentiles into 100. These metrics are valuable for understanding data distribution and variability, particularly in assessing the relative position of individual data points within the entire data set. Quartiles are frequently used in box plots for visualizing data spread, deciles in financial risk assessment, and percentiles in standardized testing or research to show performance relative to peers .
The weighted mean considers the frequency of different data points, giving more importance to certain values based on their corresponding weights. It is calculated by multiplying each value by its weight, summing these products, and dividing by the sum of the weights. Researchers use the weighted mean to account for unequal representation in datasets, ensuring that data points with higher significance have a proportionate impact on the mean .
Hypothesis testing involves four main steps: stating the hypotheses, formulating an analysis plan, analyzing sample data, and interpreting results. The formulation of hypotheses is critical because it guides the entire testing process. Stating null and alternative hypotheses establishes a basis for testing, where the null hypothesis is typically tested against evidence, and the alternative hypothesis represents the researcher’s conjecture. If outcomes refute the null hypothesis, the alternative hypothesis becomes the preferred assertion .
In hypothesis testing, two types of errors can occur: Type I error (rejecting a true null hypothesis) and Type II error (failing to reject a false null hypothesis). The probability of a Type I error is denoted by α (significance level), while the probability of a Type II error is denoted by β. The choice of significance level affects the probability of making these errors. A lower α reduces the likelihood of a Type I error but increases the probability of a Type II error if the amount of data is not adjusted. The complement to β, called the power of the test, indicates the probability of correctly rejecting a false null hypothesis .
P-values express the probability of observing a test statistic as extreme as, or more intense than, the actual statistic, assuming the null hypothesis is true. A small p-value indicates that the observed data are unusual under the null hypothesis, leading to its rejection. The region of acceptance defines the range of values under which the null hypothesis is not rejected. Both concepts help determine whether a test statistic falls into the region of rejection or acceptance, hence guiding the decision to accept or reject the null hypothesis .
To calculate the mean for grouped data, first, the midpoint for each class is determined. These midpoints are then multiplied by the frequencies of their respective classes. The sum of these products is divided by the total number of data points (n). This process organizes the data efficiently when dealing with larger datasets and frequency distributions .
The mode is useful for categorical data because it identifies the most frequently occurring category, providing insight into the most common outcome. In grouped data, the mode is determined by identifying the modal class, which is the class interval with the highest frequency. For more precise calculations, the formula incorporating class boundaries and differences in frequencies with adjacent classes is used .
For an even-numbered dataset, the median is calculated by averaging the two middlemost values after arranging the data in ascending order. This is different from an odd-numbered dataset, where the median is simply the middle value itself. For example, in a dataset of ten numbers, the median is the average of the 5th and 6th values .









