Module 4
Module 4
Data Management
Every lesson that will be included in this module get in the way in statistics to
avoid misused or abused of data or information. Thus, there is a need for you to have
knowledge of statistics and its concepts and principles in order to function well in our
society.
The reliability of a conclusion depends on whether the sample is appropriately
chosen to represent the population. Let us be careful for sampling is one of the crucial
elements of statistical inference.
1.1. Sampling
Let us considered making fruit drinks using powdered flavored drinks in a sachet.
If we pour the powdered drinks in a pitcher with water and get a spoonful from the top
before stirring, then we will get idea that it is bland. On the other hand, if you get a
spoonful of it from the bottom, we will get another misleading idea that it is very sweet.
But if you stir it first before taking a spoonful, what you taste is more representative of
the whole drink.
The process of obtaining samples is called sampling.
Random or probability sampling is a method by which every element of a
population has an equal chance of being included in the sample.
Guided Questions 1
1. Give 4 examples of selecting sampling.
2. Give 3 example of non-random or non-probability sampling.
3. Give 2 examples of census or complete enumeration.
4. Give 3 examples of observations.
5. Give 3 examples of experiments.
6. Give 4 examples of surveys.
3. Mrs. Buzon was requested by the school administrator to conduct a survey about
students’ agreement or disagreement about the proposed reduced days of classes. She
obtained the total number of students through the School Registrar as shown below.
Year Level 1st Year 2nd Year 3rd Year 4th Year
Number of Students 2,400 1,500 1,500 600
In order to determine the number of students per year level that she needed to
survey, she used the table below.
Based from the table, Mrs. Bauzon decided to conduct the survey to the 960 1 st
Year students, 375 students for 2nd and 3rd Year students and 60 for 4th Year students
who first entered the school campus. The way of deciding the number of students in the
sample is what we call the stratified random sampling. In the said sampling technique,
the number of elements of the population is separated into different categories or strata.
The members of the sample are selected proportionally from each category or stratum
by either lottery method or systematic sampling procedure.
4. Mrs. Bauzon may select the 960 students out of 2,400 1 st Year students by lottery
method. She may also select the students by asking every 5 th 1st Year student who
entered the school campus until she arrives with 960 students. Note that Mrs. Bauzon
used stratified random sampling then lottery method or systematic sampling in choosing
the 1st Year students to be included in the sample. Here, she was using a combination
of several sampling techniques called multi-stage sampling.
Question 2:
Give 3 example of non-random or non-probability sampling.
Note: There are times that random sampling is impractical to use for a particular study.
In such case, we are obliged to use a non-random sampling.
Non-random or non-probability sampling is a sampling technique where elements
of a population are drawn based on the judgement of the researcher.
1. Mr. Cruz and his company were hired by a prospective senatorial candidate to study
whether the latter would win the election. They conducted the surveys or interviews in
places where people voted for the winners in the series of previous elections. They did
this because they believed that the people would again vote for the winners in the next
election. They employed purposive sampling in determining the places to administer the
surveys or interviews. They improved their sampling by applying systematic sampling in
determining the respondents in the chosen places.
2. You are asked by your teacher to conduct a study on how six-year-old brother or
sister or any six-year-old child in your neighborhood study their lessons to have ample
time in observing and interviewing your respondents. Here, you are utilizing the
convenience sampling.
3. You are to investigate the relationship of students’ performance in Math and their
attitude towards the subject. However, you are only given limited time to do the study.
You may only consider 25 out 500 students in your school. This method quota sampling
using small samples. However, you may improve your sampling by using any random
sampling technique in determining the respondents.
Question 3:
1.2. Observation/Experiment
There are times when it is appropriate to obtain data through direct observation
or by just merely reading records. We apply observation when data can be collected
without any response from people. The experimental method is used to find out cause
and effects relationship.
Question 4:
Give 3 examples of observation
Question 5:
Give 3 examples of experiments.
1. A drug company wants to test the effectiveness of its new product in treating
influenza. An experiment or clinical test is done by treating 50 persons with the
new product and another 50 persons using existing drug. The result are analyzed
statistically to determine if the new product is significantly effective in treating
influenza.
3. In an oil company, the use of unleaded gasoline was tested to a good condition
car if it will consume gasoline in the same way to the same car by using regular
gasoline. Both cars will be driven by the same driver and will run in a 30 km road
one at a time with constant speed of 80km/hr. After the test, it was found out
through the reading of the fuel gauge that the same car used more unleaded
gasoline than the regular gasoline. This means that regular gasoline burned
faster that unleaded gasoline.
1.3. Survey
We sometimes use survey when data can only be obtained through responses
from people in a sample. We can obtain through questionnaires which may be
distributed by hand, face-to-face or phone interviews.
Question 6:
To avoid bias, questions in the survey should not lead the respondents to answer
towards a desired answer. It is also important to consider the respondents’ emotion so
as to avoid causing embarrassment or unpleasant feeling.
4. The survey question “Do you support the attempt of Congress to reduce the
number of persons who will acquire HIV through unsafe sex by passing the
Reproductive Health Bill?” is biased towards the passing of the bill. The
question may be improved as “Are you in favor of the Reproductive Health
Bill?” Why?
I. DATA GATHERING
Guided Question 2
8 6 4 5 8 8 9 10 10 6
Construct a table with three columns. The first column shows what is being arranged in
ascending order (i.e. the scores).
3. Construct a leaf and stem diagram that will represent the data below for a science
test scores for the third grading period (out of 100%):
97 92 77 82 96 75 68 80 79 96 21 34 55
84 87 68 87 88 97 81
4. The following data represents Peters’ Grades in Science subject for 1st – 4th
quarter. Construct a bar chart based on the table below.
Quarter Grades
First 84
Second 90
Third 89
Fourth 93
5. The following data represent the monthly household expenses of Rich family.
Construct a pie chart based on the table below.
Household
Amount
Expenses
Internet 1,000
Electricity 2,000
Grocery 4,000
Other 3,000
6. Construct a line chart on the following data that show daily temperature in Luna, La
Union, recorded for 5 days in Degrees Celsius.
DAYS °C
MONDAY 29
TUESDAY 33
WEDNESDAY 31
THURSDAY 36
FRIDAY 34
7. The following data represents the number of respondents aged 8-55 who are
disabled.
Age (years) Frequency
8 – 15 10
16 - 23 14
24 - 31 19
32 - 39 12
40 - 47 14
48 - 55 25
1. Construct a table with three columns. Then in the first column, write down all of
the data values in ascending order.
2. To complete the second column, go through the list of data values and place one
tally mark at the appropriate place in the second column for every data value.
When the fifth tally is reached for a mark, draw a diagonal line through the first
four tally marks. We continue this process until all data values in the list are
tallied.
3. Count the number of tally marks for each data value and write it in the third
column.
A. CATEGORICAL/ UNGROUP - Determine the order to list the categories then total
the number of occurrences of each category.
Question 1:
8 6 4 5 8 8 9 10 10 6
Construct a table with three columns. The first column shows what is being arranged in
ascending order (i.e. the scores).
The lowest mark is 4. So, start from 4 in the first column as shown below. The second
column is Tally, third is frequency.
4 I 1
5 I 1
6 II 2
7 0 0
8 III 3
9 I 1
10 II 2
GUIDELINES:
21 26 18 45 32 41 42 22 28 26 33 20 26
44 46 21 24 36 39 30
1. Determine the highest and lowest value and then compute the Range:
Range = Highest value- Lowest value, Range = 46 - 18 = 28.
2. Decide how many numbers of classes (class size) you want to have. Example: 5
classes
or in calculator, you may use the equation:
Log # of observation/log 2 or √ ¿ of observation
3. Compute the Class width or class interval.
4. Lower class limit (Smallest number of each class) and upper class limit (largest
number of each class)
Example: LCL = 18, 24, 30, 36, 42 UCL = 23, 29, 35, 41, 47
5. Class Boundaries – The number that separates the classes from one another by
Subtracting .5 to Lower limit and add 0.5 to upper limit of each class.
Question 3:
Construct a leaf and stem diagram that will represent the data below for a science test
scores for the third grading period (out of 100%):
97 92 77 82 96 75 68 80 79 96
21 34 55 84 87 68 87 88 97 81
STEM LEAVES
2 1
3 4
5 5
6 8 8
7 5 7 9
8 0 1 2 4 7 7 8
9 2 6 6 7 7
2.3.4. GRAPH OR CHART
Graphs or charts condense large amounts of information into easy-to-understand
formats that clearly and effectively communicate important points.
a. Bar Chart
b. Pie Chart
c. Line Chart
d. Histogram
Question 4:
a. Bar chart is composed of discrete bars that represent different categories of data.
The length or height of the bar is equal to the quantity within that category of data.
Bar graphs are best used to compare values across categories.
The following data represents Peters’ Grades in Science subject for 1st – 4th quarter.
Quarter Grades
First 84
Second 90
Third 89
Fourth 93
Household
Amount
Expenses
Internet 1,000
Electricity 2,000
Grocery 4,000
Other 3,000
Question 6:
c. Line chart displays the relationship between two types of information, such as
number of school personnel trained by year. They are useful in illustrating trends
over time.
The following data shows daily temperature in Luna, La Union, recorded for 5 days
in Degrees Celsius.
DAYS °C
MONDAY 29
TUESDAY 33
WEDNESDAY 31
THURSDAY 36
FRIDAY 34
Question 7:
d. Histogram has connected bars that display the frequency or proportion of cases
that fall within defined intervals or columns. The bars on the histogram can be of
varying width and typically display continuous data.
The following data represents the number of respondents aged 8-55 who are
disabled.
Age (years) Frequency
8 - 15 10
16 - 23 14
24 - 31 19
32 - 39 12
40 - 47 14
48 - 55 25
HOW TO CREATE HISTOGRAM?
Key Points:
1. Research data is data that is collected, observed, or created, for purposes of
analysis to produce original research results.
2. Ways of organizing data in research by Frequency Distribution Table, Stem and
Leaf Diagram and Chart.
3. The frequency distribution table mainly composes of three columns namely data
values arranged in ascending or descending order, tally and frequency.
4. There are two types of frequency distribution, the ungroup and group data.
Ungroup data determine the order to list the categories then total the number of
occurrences of each category while group data refers to data being organized
into groups known as classes.
5. Stem-and-leaf diagram is a method used to organize statistical data that helps us
to see values according to their size, so we can order them accordingly. In a
stem-and-leaf diagram, each data value is split into a stem and a leaf. The leaf is
the last digit to the right. The stem is the remaining digits to the left. For the
number 243, the stem is 24 and the leaf is 3. The stem and the leaf are
separated by two-column table which is open on edges or border.
6. Graphs or charts condense large amounts of information into easy-to-understand
formats that clearly and effectively communicate important points. The common
use charts are Bar Chart, Pie Chart, Line Chart and Histogram.
7. Keep chart simple and avoid flashy special effects. Present only essential
information. Avoid using gratuitous options in graphical software programs, such
as three-dimensional bars, that confuse the reader. If the graph or chart is too
complex, it will not clearly communicate the important points.
8. Title your graph or chart clearly to convey the purpose. The title provides the
reader with the overall message you are conveying.
9. Specify the units of measurement on the x and y-axis. Years, number of
participants trained, and type of school personnel are examples of labels for
units of measurement.
Name: _________________________ Date: _______________
Year & Section: __________________ Score: ______________
Exercise B
A. Construct the following. Use another sheet of paper for your answer.
1. Construct a 3-column table showing the scores, tally and frequency of the following
data arrange in descending order.
25, 23, 30, 27, 28, 25, 21, 28, 30, 27, 22, 25, 26, 27, 21, 29, 20, 23, 21, 25, 21
2. Construct a frequency distribution table on the following data that represents the
ages of 20 respondents to group the data:
22 25 17 44 33 42 43 23 27 25 34 21 25
45 48 22 23 35 38 32
3. Construct a leaf and stem diagram that will represent the data below for a science
test scores for the third grading period (out of 100%):
96 93 78 81 97 74 69 82 76 95
20 33 57 83 85 67 84 85 97 83
B. Construct the following chart. Use another sheet of paper for your answer.
1. Construct a bar chart, given the production of medical facemasks of a pharmaceutical
company for 5 consecutive days to be 2000, 1500, 1750, 2150, and 1350.
2. Given the expenses of a family with 6 members in a month: Food – P4,500; Water –
P850; Electricity – P3,500; ICT bills – P2,500; and Others - P1,500. Construct a pie
chart to represent their part in monthly expenses.
3. The following data shows daily temperature in PSAU 32, 38, 34, 37, and 28 recorded
for 5 days in Degrees Celsius. Developed a line chart to determine the differences
that occurs each day.
Lesson 2: MEASURES OF CENTRAL TENDENCY
The goal of central tendency is to identify the single value that is the best
representative for the entire set of data.
In addition, it is possible to compare two (or more) sets of data by simply comparing
the average score (central tendency) for one set versus the average score for
another set.
Σx = x1+x2+x3+ … +xn
One uses this notation because it is more convenient to write the sum in this
fashion.
THE MEAN
The mean is the arithmetic average obtained by adding up all the scores and dividing by
the total number of scores. It is in fact, the numerical average of the set of data.
Formulas for the Mean:
x=
∑x for Ungrouped Data
N
“X bar” equals the sum of all the scores, X, divided by the number of scores, N.
x=
∑ f xm for Grouped Data
N
where:
x m=midpoint of each c lass
f x m = a midpoint multiplied by its frequency
Guide Questions 3
1. Find the mean if the 5 test scores for Calculus I are 95, 83, 92, 81, and 75.
2. Find the mean of the given range of data below;
3. Find the mean of the given grouped data below;
The Hours Spent in Watching TV
Hours Spent f xm fxm
6-7 2
4-5 10
2-3 13
0-1 5
30
What we need to do is find the midpoints of the ranges and then multiply then by
the frequency. So that we can compute the mean.
The midpoints are 16, 19, 21, 23.5, 27.5, and 32.5.
x = 16(94,000) + 19(1,551,000) + 21(1,420,000) + 23.5 (1,091,000) +
27.5(865,000) + 32.5(521,000)] /5,542,000 = 22.94
Question 3:
Find the mean of the given grouped data below;
The Hours Spent in Watching TV
Hours Spent F xm fxm
6-7 2 6.5 13
4-5 10 4.5 45
2-3 13 2.5 32.5
0-1 5 0.5 2.5
30 ∑ f x m=¿ ¿93
x=
∑ fx
N
93
x= = 3.10 hours is the mean
30
THE MEDIAN
The median is the middle value of a distribution of data. It is the score that divides
the distribution into two equal parts, so that half the cases are above it and half
below it. Also, it is the middle score, or average of middle scores in a distribution.
First, if possible or feasible, arrange the data from smallest value to largest
value.
The location of the median can be calculated using this formula: (n+1)/2.
If (n+1)/2 is a whole number then that value gives the location. Just report the
value of that location as the median.
If (n+1)/2 is not a whole number then the first whole number less than the
location value and the first whole number greater than the location value will
be used to calculate the median. Take the data located at those 2 values and
calculate the average, this is the median.
Formulas for the Median:
~ n+1
x= ; for Ungrouped Data
2
n
−cf
~ 2 ; for Grouped Data
x=Lm + i
f
where;
~x - Median
Lm – Class boundary (lower limit) of the median class
n – Number of observation
cf – Cumulative frequency above the median class
f – Frequency of the median class
i – Class interval
Guide Question
1. Find the median of the quiz scores 5, 10, 8, 6, 4, 8, 2, 5, 7, and 7.
2. Find the median of a bunch of 10 points quizzes from MMW; 9, 6, 7, 10, 9, 4, 9, 2, 9,
10, 7, 7, 5, 6, and 7.
3. Find the median of the given grouped table below;
Ages of Adults Participating in Covid-19 Vaccination Trial
Ages f cf
18-23 5
24-29 6
30-35 3
36-41 2
42-47 4
Question 3:
Find the median of the given grouped table below;
n
−cf
~ 2
x=Lm + i
f
n 20
= =10; cf = 5; f = 6; and i = 6
2 2
~ 10−5
x=23.5+ (6)
6
~
x=28.50
Thus, 10 adults participated with ages less than 28.5 years of age and 10 adults
participated with ages more than 28.5 years of age in the vaccination trial.
THE MODE
The mode is the most frequent number in a collection of data. The mode is fairly
useless with data like weights or heights where there are a large number of possible
values. The mode is most commonly used for categorical data, for which median and
mean cannot be computed.
Formulas for the mode:
The MODE is the piece of data that occurs most frequently in the data set that can be
taken by simple judgement of counting.
^x =Lmo +
( d1
)
d1 + d2
(i)
Guide Questions 5
Ages f
18-23 5
24-29 6
30-35 3
36-41 2
42-47 4
Question 2:
Find the mode of the data 2, 5, 1, 5, 1, and 2.
It has no mode because 1, 2, and 5 have a frequency of 2.
Question 3:
Find the mode of the data 5, 7, 9, 1, 7, 5, 0, and 4.
It has two modes 5 and 7. This is said to be bimodal.
Question 4:
Find the mode of the grouped data below;
Ages of Adults Participating in Covid-19 Vaccination Trial
Ages f
18 - 23 5
24 - 29 6
30 - 35 3
36 - 41 2
42- 47 4
20
^x =Lmo +
( d1
)
d1 + d2
(i)
and
Lmo=23.5 , d 1=6−5=1; d 2=6−3=3 ; i=6
For instance, suppose your professor tells you that your grade will be based on a
midterm and a final exam, each of which is based on 100 possible points.
However, the final exam will be worth 60% of the grade and the midterm only 40%. How
could you determine an average score that would reflect these different weights?
Guide Questions 6
1. Suppose your midterm test score is 83 and your final exam score is 95. Using
weights of 40% for the midterm and 60% for the final exam, compute the weighted
average of your scores. If the minimum average for an A is 90, will you earn an A?
2. The table below shows Dillon’s fall semester course grades. Use the weighted mean
formula to find Dillon’s GPA for the fall semester if A = 4, B = 3, C = 2, D = 1, and F = 0.
Course Course Course
Grade Units
Biology A 4
Statistics B 3
Business C 3
Psychology F 2
CAD B 2
2. Find the weighted mean of the table below.
Distribution of Laptop Computers per Household
No. of Laptop Computers No. of Households with Laptop Computers
(x) (f)
0 5
1 12
2 14
3 3
4 2
5 3
6 0
7 1
40
Question 2:
The table below shows Dillon’s fall semester course grades. Use the weighted mean
formula to find Dillon’s GPA for the fall semester if A = 4, B = 3, C = 2, D = 1, and F = 0.
( 4 x 4 ) + ( 3 x 3 ) + ( 2 x 3 )+(3 x 2) 37
Weighted Mean = = =2.64
14 14
Dillon’s GPA for the fall semester is 2.64.
Question 3:
Find the weighted mean of the table below.
Distribution of Laptop Computers per Household
No. of Laptop Computers No. of Households with Laptop Computers
(x) (f)
0 5
1 12
2 14
3 3
4 2
5 3
6 0
7 1
40
( 0 x 5 ) + ( 1 x 12 ) + ( 2 x 14 ) + ( 3 x 3 ) + ( 4 x 2 ) + ( 5 x 3 )+ ( 6 x 0 )+(7 x 1) 79
Weighted Mean = = =1.975
40 40
Weighted Mean=
∑ (x∙w)
∑w
where: ∑ ( x ∙ w ) is the sum of the product formed by multiplying each number by
the assigned weight, and ∑ w is the sum of all the weights.
Name: ________________________ Date: __________
Year & Course:_________________ Score:__________
Exercise C
A. Answer the following. Use other sheet of paper for your answer.
1. Find the mean if the 5 test scores for Calculus I are 93, 80, 95, 84, and 78.
2. Find the median of the 8 test scores in Statistics 12, 16, 15, 18, 20, 15, 17, and 19.
3. Find the mode of the test scores in MMW 10, 8, 10, 8, 6, 5, 8, 7, 9 and 10.
4. Find the weighted mean or GWA of the grades of Alvin for the First Semester given
the table below.
B. Answer the following. Use other sheet of paper for your answer.
2. Find the median of the grouped data below and give your interpretation
Ages of Adults Participating in Math Contest
Ages F
16 – 20 5
21 – 25 4
26 – 30 9
31 – 35 5
36 – 40 2
3. Find the mode of the grouped data below and give your interpretation.
Ages of Adults Participating in Math Contest
Ages F
16 – 20 5
21 – 25 4
26 – 30 9
31 – 35 5
36 – 40 2
Lesson 3: MEASURES OF DISPERSION
Measures of dispersion are descriptive statistics that describe how similar a set of
scores are to each other.
The more similar the scores are to each other, the lower the measure of
dispersion will be.
The less similar the scores are to each other, the higher the measure of
dispersion will be.
In general, the more spread out a distribution is, the larger the measure of
dispersion will be.
THE RANGE
The range is defined as the difference between the largest score in the set of data and
the smallest score in the set of data. The formula is;
R = X L - XS
THE STANDARD DEVIATION
σ=
√ ∑ (x−μ)2
n
; For population
ς=
√ ∑ ( x−x )2 ; For sample
n−1
where: x1, x2, x3, … xn is a sample of n numbers with a mean of x .
THE VARIANCE
2
σ =
∑ (x−μ)2 ; For population
n
ς =
∑
2 (x−x )2
; For sample
n−1
Guide Questions 7
1. What is the range, standard deviation and variance of the sample below?
2 4 7 12 15
2. A consumer group has tested a sample of 8 size-D butteries from each of 3
companies. The results of the test are shown in the table below. According to these
tests, which company produces batteries for which the values representing hours of
constant use have the smallest standard deviation?
Company Hours of constant use per battery
EverSoBright 6.2, 6.4, 7.1, 5.9, 8.3, 5.3, 7.5, 9.3
Dependable 6.8, 6.2, 7.2, 5.9, 7.0, 7.4, 7.3, 8.2
Beacon 6.1, 6.6, 7.3, 5.7, 7.1, 7.6, 7.1, 8.5
ς=
√ ∑ ( x−x )2
n−1
ς=
√ (2−8 )2 + ( 4−8 )2+ ( 7−8 )2+ ( 12−8 )2 + ( 15−8 )2
5−1
ς=√ 29.5=5.43
c) Variance for sample
ς 2=
∑ (x−x )2
n−1
2
ς 2=( √ 29.5 )
2
ς =29.5
Question 2:
A consumer group has tested a sample of 8 size-D butteries from each of 3 companies.
The results of the test are shown in the table below. According to these tests, which
company produces batteries for which the values representing hours of constant use
have the smallest standard deviation?
ς 1=
√( 6.2−7 )2 + ( 6.4−7 )2 +…+ ( 9.3−7 )2
7
ς 1=
√ 12.34
7
=1.328 hours
ς 2=
√( 6.8−7 )2 + ( 6.2−7 )2 +…+ ( 8.2−7 )2
7
ς 2=
√ 3.62
7
=0.719 hours
ς 3=
√ 5.38
7
=0.877 hours
Thus, the batteries from Dependable have the smallest standard deviation. According to
these results, the Dependable company produces the most consistent batteries with
regard to life expectancy under constant use.
Key Points
1. The range of a set of data values is the difference between the greatest data
value and the least data value.
σ=
√ ∑ (x−μ)2
n
and σ 2=
∑ (x−μ)2 .
n
3. The standard deviation and variance. If x1, x2, x3, … xn is a sample of n numbers
with a mean of ς .
ς=
√ ∑ ( x−x )2
n−1
and 2
ς =
∑ (x−x )2 .
n−1
Name: ________________________ Date: __________
Year & Course:_________________ Score:__________
Exercise D
MEASURES OF DISPERSION
1. Find the range, standard deviation and variance of the sample below?
1 5 9 13 17
2. A consumer testing agency has tested the strengths of 3 brands of 1/8-inch rope. The
results of the test are shown in the table below. According to the sample test results,
which company produces 1/8-inch rope for which the breaking point has the smallest
standard deviation?
3. The fuel efficiency in mi/gal of 12 small utility trucks was measured. The results are
recorded in the table below.
Fuel Efficiency (mpg)
22 25 23 27 15 24 24 32 23 22 25 22
Find the mean and sample standard deviation of these data. Round to the nearest
hundredth.
Lesson 4: MEASURES OF RELATIVE POSITIONS
THE z-SCORES
When a set of data values are normally distributed, we can standardize each
score by converting it into a z-Score. z-Scores make it easier to compare data values
measured on different scales. It reflects how many standard deviations above or below
the mean a raw score is. It is positive if the data value lies above the mean and negative
if the data value lies below the mean.
Formula of z-Score:
x−μ
z= ; for population
σ
Where: x = an element of the data set, the mean = μ, and standard deviation = σ .
x−x
z= ; for sample
ς
Where: x = an element of the data set, the mean = x , and standard deviation = ς .
THE PERCENTILE
pth Percentile
A value x is called the pth percentile of a data set provided p% of the data values are
less than x.
Percentile for a Given Data Value:
Given a set of data and a data value x,
number of data values less than x
Percentile of score x = ⋅100
total number of data values
THE QUARTILE
The three numbers Q1, Q2, and Q3 that partition a ranked data set into four
(approximately) equal groups are called quartiles of the data. For instance, for the data
set below, the values Q1 = 11, Q2 = 29, and Q3= 104 are quartiles of the data.
2, 5, 5, 8, 11, 12, 19, 22, 23, 29, 31, 45, 83, 91, 104, 159, 181, 312, 354
↕ ↕ ↕
Q1 Q2 Q3
The quartile Q1 is called the first quartile. The quartile Q2 is called the second quartile. It
is also the median of the data. The quartile Q3 is called the third quartile.
A box-and-whisker plot (sometimes called box plot) is often used to provide a visual
summary of a set of data. It shows the median, the first and third quartiles, and the
minimum and maximum values of a data set.
Q1 Q2 Q3 Whisker
Minimum Maximum
Guide Questions 8
1. Suppose SAT scores among college students are normally distributed with a mean of
500 and a standard deviation of 100. If a student scores a 700, what would be her z-
score?
2. A set of math test scores has a mean of 70 and a standard deviation of 8. A set of
English test scores has a mean of 74 and a standard deviation of 16. For which test
would a score of 78 have a higher standing?
3. In a recent year, the median annual salary for a physical therapist was P360,000. If
the 90th percentile for the annual salary of a physical therapist was P648,000, find the
percent of physical therapist whose annual salary was
a. more than P360,000.
b. less than P648,000.
c. between P360,000 and P648,000.
4. On a reading examination given to 900 students. Elaine’s score of 602 was higher
than the scores of 576 of the students who took the examination. What is the percentile
for Elaine’s score?
5. Find the quartiles Q1, Q2, and Q3 of the following data 20, 30, 25, 23, 22, 32, 36.
6. The following list the weights, in ounces, of 15 avocados in a random sample. Find
the quartiles of the data.
Weights in Ounces of Avocados
12.4 10.8 14.2 7.5 10.2 11.4 12.6 12.8 13.1 15.6
9.8 11.4 12.2 16.4 14.5
7. Construct the box-and-whisker plot of the data set: 85,92,78,88,90,88,89.
A set of math test scores has a mean of 70 and a standard deviation of 8. A set of
English test scores has a mean of 74 and a standard deviation of 16. For which test
would a score of 78 have a higher standing?
x−μ
z=
σ
78−70 8
For Math: z= = =1
8 8
78−74 4 1
For English: z= = = =0.25
16 16 4
The math score would have the highest standing since it is 1 standard deviation above
the mean while the English score is only 0.25 standard deviation above the mean.
Question 3:
In a recent year, the median annual salary for a physical therapist was P360,000. If the
90th percentile for the annual salary of a physical therapist was P648,000, find the
percent of physical therapist whose annual salary was
a. more than P360,000.
b. less than P648,000.
c. between P360,000 and P648,000.
a. By definition, the median is the 50th percentile. Therefore, 50% of the physical
therapists earned more than P360,000 per year.
b. Because P648,000 is the 90th percentile, 90% of all physical therapists made
less than P648,000.
c. From parts a and b, 90% - 50% = 40% of the physical therapists earned
between P360,000 and P648,000.
Question 4:
On a reading examination given to 900 students. Elaine’s score of 602 was higher than
the scores of 576 of the students who took the examination. What is the percentile for
Elaine’s score?
number of data values less than x
Percentile = ⋅100
total number of data values
576
Percentile = ⋅100=64
900
Thus, Elaine’s score of 602 places her at the 64th percentile.
Question 5:
Find the quartiles Q1, Q2, and Q3 of the following data 20, 30, 25, 23, 22, 32, 36.
By ascending arrangement: 20 22 23 25 30 32 36
Question 6:
The following list the weights, in ounces, of 15 avocados in a random sample. Find the
quartiles of the data
Weights in Ounces of Avocados
12.4 10.8 14.2 7.5 10.2 11.4 12.6 12.8 13.1 15.6
9.8 11.4 12.2 16.4 14.5
Arrange data in ascending form, and n = 15 odd number.
7.5 9.8 10.2 10.8 11.4 11.4 12.2 12.4 12.6
12.8 13.1 14.2 14.5 15.6 16.4
The median of these 15 data values has a rank of 8. Thus the median is 12.4 which is
also Q2.
There are 7 data values less than the median and 7 data values greater than the
median. The first quartile is the median of the data values less than the median. Thus,
Q1 is 10.8. While, the third quartile is the median of the data values greater than the
median which is 14.2 which is also the Q3.
Question 7:
Construct the box-and-whisker plot of the data set: 85,92,78,88,90,88,89.
Key Points
1. The z-scores for a given data value x is the number of standard deviations that x
is above or below the mean
x−μ x−x
z= ; for population and z= ; for sample
σ ς
3. Quartile - the three numbers Q1, Q2, and Q3 that partition a ranked data set into
four (approximately) equal groups are called quartiles of the data. The quartile Q 1
is called the first quartile. The quartile Q2 is called the second quartile. It is also
the median of the data. The quartile Q 3 is called the third quartile is the median of
the data values greater than Q2..
1. What will be the miles per gallon for a Toyota Innova when the average km/li is 11, it
has a z value of 1.5 and a standard deviation of 2?
2. Raul has taken two tests in his Chemistry class. He scored 72 on the first test, for
which the mean of all scores was 65 and the standard deviation was 8. He received a
60 on a second test, for which the mean of all scores was 45 and the standard deviation
was 12. In comparison to other students, did Raul do better on the first test or the
second test.
3. Find the quartiles Q1, Q2, and Q3 of the following data 20, 30, 25, 23, 22, 32, 36, 18.
4. Construct the box-and-whisker of the given value data 15, 83, 75, 12, 19, 74, 21.
5. Find the percentiles P8, P50, and P85 of the following data 20, 30, 25, 23, 22, 32, 36 .
Lesson 5: NORMAL DISTRIBUTIONS
Normal Distribution – is a bell-shaped continuous distribution widely used in statistical
inference. The bell-shaped curve is symmetric about a vertical line through the mean of
the data.
Properties of a Normal Distribution:
The graph is symmetric about the vertical line through the mean distribution.
The mean, median, and the mode are equal.
The y-value of each point on the curve is the percent (expressed as a decimal)
of the data at the corresponding x-value.
Areas under the curve that are symmetric about the mean are equal.
The total area under the curve is 1.
It is often helpful to convert data value x to z-scores, as we did in the previous section
by using the z-score formulas;
x−μ x−x
zx= ∨z x =
σ σ
The standard normal distribution is the normal distribution that has a mean of 0 and a
standard deviation of 1.
Guide Questions 9
1. A survey of 1000 U.S. gas stations found that the price charged for a gallon of regular
gas could be closely approximated by a normal distribution with a mean of USS 3.10
and a standard deviation of USS 0.18. How many of the station charge
a. between USS 2.74 and USS 3.46 for a gallon of regular gas?
b. less than USS 3.28 for a gallon of regular gas?
c. more than USS 3.46 for a gallon of regular gas
2. Find the standard normal distribution between z = -1.44 and z = 0.
3. A soda machine dispenses soda into 12-ounce cup. Test show that the actual
amount of soda dispensed is normally distributed with a mean of 11.5 oz and a standard
deviation of 0.2 oz.
a. What percent of cups will receive less than 11.25 oz of soda?
b. What percent of cups will receive between 11.2 oz and 11.55 oz of soda?
c. If a cup is filled at random, what probability that the machine will overflow the
cup?
4. The OnTheGo company manufactures laptop computers. A study indicates that the
life span of its computers are normally distributed with a mean of 4.0 years and a
standard deviation of 1.2 years. How long a warranty period should the company offer if
the company wishes less than 4% of its computer to fail during the warranty period
Answers to Guide Questions 9
Question 1:
A survey of 1000 U.S. gas stations found that the price charged for a gallon of regular
gas could be closely approximated by a normal distribution with a mean of USS 3.10
and a standard deviation of USS 0.18. How many of the station charge
a. between USS 2.74 and USS 3.46 for a gallon of regular gas?
The USS 2.74 per gallon price is 2 standard deviations below the mean. The USS 3.46
price is 2 standard deviations above the mean. In a normal distribution, 95% of all data
lie within 2 standard deviations of the mean. Therefore, approximately
(95%)(1000) = (0.95)(1000) = 950 of the stations charge between USS 2.74 and
USS 3.46 for a gallon of regular gas.
μ−2 σ μ μ+2 σ
95 %
b. less than USS 3.28 for a gallon of regular gas?
The USS 3.28 price is 1 standard deviation above the mean. In a normal distribution,
34% of all data lie between the mean and 1 standard deviation above the mean.
Thus, approximately
(34%)(1000) = (0.34)(1000) = 340 of the stations charge between USS 3.10 and
USS 3.28 for a gallon of regular gasoline. Half of the 1000 stations or 500 stations,
charge less than the mean. Therefore, about 340 + 500 = 840 of the stations charge
less than USS 3.28 for a gallon of regular gas. (Draw the normal distribution graph.)
c. more than USS 3.46 for a gallon of regular gas?
The USS 3.46 price is 2 standard deviations above the mean. In a normal
distribution, 95% of all data are within 2 standard deviation above the mean. This
means that the other 5% of the data will lie either more than 2 standard deviations
above the mean or more than 2 standard deviations below the mean. We are
interested only in the data that are more than 2 standard deviations above the mean
which is ½ of 5% or 2.5% of the data. Thus, approximately
(2.5%)(1000) = (0.025)(1000) = 25 of the stations charge more than USS 3.46 for a
gallon of regular gas. (Draw the normal distribution graph.)
Question 2:
Because the standard normal distribution is symmetrical about the center line z = 0, the
area of the standard normal distribution between z = -1.44 and z = 0 is equal to the area
between z = 0 and z = -1.44. The entry associated with z = 1.44 is 0.425. Thus, the area
of the standard normal distribution between z = -1.44 and z = 0 is 0.425 square unit.
Question 3:
A soda machine dispenses soda into 12-ounce cup. Test show that the actual amount
of soda dispensed is normally distributed with a mean of 11.5 oz and a standard
deviation of 0.2 oz.
a. What percent of cups will receive less than 11.25 oz of soda?
x−x
Recall that the formula for the z-score for a data value x is z x = . Thus, the
σ
z-score for 11.25 oz is;
11.25−11.5
z 11.25= =−1.25
0.2
The table for areas and z-scores indicate that 0.394(39.4%) of the data in a normal
distribution are between z = 0 and z = 1.25. Because the data are normally
distributed, 39.4% of the data is also between z = 0 and z = -1.25. The percent data
to the left of z = -1.25 is 50% - 39.4 = 10.6%. Thus. 10.6% of the cups filled by soda
machine will receive less than 11.25% oz of soda.
b. The z-score of 11.55 oz is
11.55−11.5
z 11.55= =0.25
0.2
It indicates that 0.099(9.9%) of the data in a normal distribution is between z = 0
and z = 0.25.
11.2−11.5
z 11.2 = =−1.5
0.2
It indicates that 0.433 (43.3%) of the data in a normal distribution are between
z = 0 and z = 1.5. Because the data are normally distributed, 43.3% of the data
are also between z = 0 and z = -1.5. Thus, the percent of the cups that the
vending machine will fill with between 11.2 oz and 11.55 oz of soda is 43.3% +
9.9% = 53.2%.
c. A cup will overflow if it receives more than 12 oz of soda. The z-score for 12 oz is
12−11.5
z 12= =2.5
0.2
It indicates that 0.494(49.4%) of the data in the standard normal distribution are
between z = 0 and z = 2.5. The percent of data to the right of z = 2.5 is
determined by subtracting 49.4% from 50%. Thus, 0.6% of the time the machine
produces an overflow, and the probability that a cup filled at random will overflow
is 0.006.
Question 4:
The OnTheGo company manufactures laptop computers. A study indicates that the life
span of its computers are normally distributed with a mean of 4.0 years and a standard
deviation of 1.2 years. How long a warranty period should the company offer if the
company wishes less than 4% of its computer to fail during the warranty period?
The standard normal distribution with 4% of the data to the left of some unknown z-
score and 46% of the data to the right of the z-score but to the left of the mean of 0.
Using the Table Area Under the Standard Normal Curve, we find that the z-score
associated with an area of A = 0.46 is 1.75. Our unknown z-score is to the left of 0, so it
must be negative. Thus zx = -1.75. If we let x the time in years that a computer is use,
then x is related to the z-scores by the formula
x−x
zx=
s
Solving for x with x=4.0 , s=1.2 ,∧z=−1.75 gives us
x−4.0
−1.75=
1.2
(−1.75 ) ( 1.2 )=x−4.0
x=4.0−2.1
x=1.9
Hence, the company can provide a 1.9-year warranty and expect less than 4% of its
computers to fail during the warranty period.
Key Points
1. Frequency Distribution displays a data set by dividing the data into intervals or
classes and the listing the number of data values that fall into each interval. A
relative frequency distribution lists the percent of data in each interval.
3. The empirical rule for a normal distribution states that approximately 68% of the
data lie within 1 standard deviation of the mean, 95% of the data lie within 2
standard deviation of the mean, and 99.7% of the data lie within 3 standard
deviation of the mean.
4. Using the standard normal distribution is the normal distribution that has a mean
of 0 and a standard deviation of 1. Any normal distribution can be converted into
the standard normal distribution by converting data values to their z-scores.
Then, the percent of data values that lie in a given interval can be found as the
area under the standard normal curve between the z-scores of the endpoints of
the given interval. The areas under normal distribution will be given in a table
during the discussion.
Name: ________________________ Date: ___________
Course & Year: _________________ Score: __________
NORMAL DISTRIBUTION
Answer the following. Use other paper sheets for your answers.
1. A vegetable distributor knows that during the month of August, the weights of its
tomatoes are normally distributed with a mean of 0.61 lb and a standard deviation of
0.15 lb. What percent of tomatoes weighs less than 0.76 lb? In a shipment of 6,000
tomatoes, how many tomatoes can be expected to weigh more than 0.31 lb?
2. Human are on average, taller today they were 200 years ago. Today, the mean
height of a 14-year-old is about 65 in. Use the table below answer the following
questions.
3. A highway study of 8,00 vehicles that passed by a checkpoint found that their speeds
were normally distributed with a mean of 61 mph and a standard deviation of 7 mph.
How many vehicles had a speed of more than 68 mph? How many of the vehicles had a
speed of less than 40 mph?
4. Find the area of the standard normal distribution between z = - 1.44 and z = 0.
5. Find the area of the standard normal distribution to the right of z = 0.82. Also draw the
graph showing the area covered.
Lesson 6: LINEAR REGRESSION AND CORRELATION
When performing research studies, scientist often wish to know whether two variables
are related. If the variables are determined to be related, a scientist may then wish to
find an equation that can be used to model the relationship.
Regression model:
Guide Question 10
1. A geologist might want to know whether there is a relationship between the duration
of an eruption of a geyser and the time between eruptions. He collected some and were
given in the table below which gives the bivariate data showing the time between two
eruptions and the duration of the second eruption for 10 eruptions of the geyser. He
wants to know the approximate regression equation and if the time between to eruptions
is 200 seconds, then he wants to find out the estimated duration of the second eruption.
2. Find the equation of the least-squares line for the ordered pairs below and predict the
average speed of an adult man for each of the following stride length 2.8 m and 4.8 m.
The table of the bivariate:
Stride length in meter (x) 2.5 3.0 3.3 3.5 3.8 4.0 4.2 4.5
Speed in m/s (y) 3.4 4.9 5.5 6.6 7.0 7.7 8.3 8.7
3. Find the linear correlation coefficient for stride length versus speed of an adult man in
the table given below. Give your interpretation on the result. Round your result to the
nearest hundredth.
The table of the bivariate:
Stride length in meter (x) 2.5 3.0 3.3 3.5 3.8 4.0 4.2 4.5
Speed in m/s (y) 3.4 4.9 5.5 6.6 7.0 7.7 8.3 8.7
Answer to Guide Question 10
Question 1:
For instance, a geologist might want to know whether there is a relationship between
the duration of an eruption of a geyser and the time between eruptions. A first step in
this determination is to collect some data. Data involving two variables are called
bivariate data. Table below gives the bivariate data showing the time between two
eruptions and the duration of the second eruption for 10 eruptions of the geyser.
The table of the bivariate:
Time between
272 227 237 238 203 270 218 226 250 245
eruptions(in seconds), x
Duration of eruption(in
89 79 83 82 81 85 78 81 85 79
seconds), y
One way to create a model for the relationship between the times between two
eruptions and the duration of the second eruption is to find a line that approximates the
data points plotted in the scatter plot. There are many such lines that can be drawn as
shown in the second figure above. In the figure, all the possible lines that can be drawn-
the one that is usually of most interest (the bold line) is called the line of best fit or the
least-squares regression line. The least-square regression line is the line that fits the
data better than any other line that might be drawn. The least-square regression is
defined as is the line that minimizes the sum of the squares of the vertical deviations
from each data point to the line for a set of bivariate.
The equation for the least-squares line for the n ordered pairs (x 1,y1), (x2,y2), (x3,y3),…,
(xn,yn) is ^y =ax+ b, where
n ∑ xy−∑ x ∑ y
a= and b= y−a x
n ∑ x 2−¿ ( ∑ x ) ¿
2
Apply the formula to the given data above, we first find the value of each summation:
No. of
Eruption X Y xy x2
s
1 272 89 24208 73984
2 227 79 17933 51529
3 237 83 19671 56169
4 238 82 19516 56644
5 203 81 16443 41209
6 270 85 22950 72900
7 218 78 17004 47524
8 226 81 18306 51076
9 250 85 21250 62500
10 245 79 19355 60025
19663 57356
2386 822
Σ 6 0
n = 10
5,692,99
2
(Σx) = 6
10(196,636)−(2,386)(822)
a= =0.1189559666
( 10 ) (573,560 )−(5,692,996)
We find the x and y ,
x=
∑ x = 2,386 =238.6∧ y = ∑ y = 822 =82.2
n 10 n 10
and use them to find the y-intercept, b.
b= y−a x
b=82.2−0.1189559666(238.6)
b=53.81710637
a. The regression equation is ^y =0.1189559666 x +53.81710637. The graph of the
regression equation and a scatter plot of the data are shown below.
^y =0.1189559666 x +53.81710637
Question 2:
Find the equation of the least-squares line for the ordered pairs below and predict the
average speed of an adult man for each of the following stride length 2.8 m and 4.8 m.
Stride length in meter (x) 2.5 3.0 3.3 3.5 3.8 4.0 4.2 4.5
Speed in m/s (y) 3.4 4.9 5.5 6.6 7.0 7.7 8.3 8.7
Apply the formula to the given data above, we first find the value of each summation:
No X y xy x2
1 2.5 3.4 8.5 6.25
2 3 4.9 14.7 9
3 3.3 5.5 18.15 10.89
4 3.5 6.6 23.1 12.25
5 3.8 7 26.6 14.44
6 4 7.7 30.8 16
7 4.2 8.3 34.86 17.64
8 4.5 8.7 39.15 20.25
Σ 28.8 52.1 195.86 106.72
n=8
(Σx)2 = 829.44
8(195.86)−(28.8)(52.1)
a= =2.730263158
( 8 ) (106.72 ) −(829.44 )
We find the x and y ,
x=
∑ x = 28.8 =3.6∧ y = ∑ y = 52.1 =6.5125
n 8 n 8
^y =2.730263158 x−3.316447368
b. The predicted average speed of an adult man with a stride of 2.8 m is,
m
^y =2.730263158 ( 2.8 )−3.316447368 ≈ 4.328 .
sec
The predicted average speed of an adult man with a stride of 4.8 m is,
m
^y =2.730263158 ( 4.8 )−3.316447368 ≈ 9.789 .
sec
Question 3:
Find the linear correlation coefficient for stride length versus speed of an adult man in
the table given below. Give your interpretation on the result. Round your result to the
nearest hundredth.
Stride length in meter (x) 2.5 3.0 3.3 3.5 3.8 4.0 4.2 4.5
Speed in m/s (y) 3.4 4.9 5.5 6.6 7.0 7.7 8.3 8.7
Apply the formula to the given data above, we first find the value of each summation:
No x y Xy x2 y2
1 2.5 3.4 8.5 6.25 11.56
2 3 4.9 14.7 9 24.01
3 3.3 5.5 18.15 10.89 30.25
4 3.5 6.6 23.1 12.25 43.56
5 3.8 7 26.6 14.44 49
6 4 7.7 30.8 16 59.29
7 4.2 8.3 34.86 17.64 68.89
8 4.5 8.7 39.15 20.25 75.69
Σ 28.8 52.1 195.86 106.72 362.25
n=8
(Σx)2 = 829.44 (Σy)2 = 2,714.41
n ( ∑ xy ) −( ∑ x )( ∑ y )
r=
√ n (∑ x )−(∑ x ) ∙ √n (∑ y )−(∑ y )
2 2 2 2
n ∑ xy−∑ x ∑ y
a= and b= y−a x
n ∑ x 2−¿ ( ∑ x ) ¿
2
4. The equation of the least-squares line can be used to predict the value of one
variable when the value of the other variable is known.
n ( ∑ xy ) −( ∑ x )( ∑ y )
r= .
√ 2
√ 2
n ( ∑ x 2 ) −( ∑ x ) ∙ n ( ∑ y 2 )−( ∑ y )
1.
Name: ________________________ Date: ___________
Course & Year: _________________ Score: __________
LINEAR REGRESSION AND CORRELATION
Answer the following. Use other sheets of paper for your answers.
1. Find the equation of the least-squares line, the linear correlation coefficient for the
given data. Round the constants a, b and r to the nearest hundredth.
a. (2,6), (3,6), 4,8), (6,11), (8,18)
b. (2,-3), (3,-4), (4,-9), (5,-10), (7,-12)
c. (2,5), (3,7), (4,8), (6,1), (8,18), (9,21)
2. Given the bivariate data:
x 1 2 3 5 6
y 7 5 3 2 1
a. Draw the scatter plot.
b. Find the equation of the least-squares line.
c. Find to the nearest hundredth, the linear correlation coefficient and make
prediction.
3. The average remaining life-times for women of various ages in the United States are
given in the table below.
Average Remaining Lifetimes for Women
Age (x) 0 15 35 65 75
Years (y) 79.9 65.6 46.2 19.5 12.1
a. Find the equation of the least-squares line.
b. Use the equation of the least-squares to estimate the remaining lifetime of a
woman of age 25.
c. Find to the nearest hundredth, the linear correlation coefficient and make
prediction.