0% found this document useful (0 votes)
4 views70 pages

Module 4

Module 4 focuses on data management, emphasizing the importance of sampling methods, data collection, and organization in statistics. It covers various sampling techniques, including random and non-random sampling, as well as methods for organizing and presenting data through tables and graphs. The module also highlights the significance of proper data handling to ensure accurate representation and analysis in research.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
4 views70 pages

Module 4

Module 4 focuses on data management, emphasizing the importance of sampling methods, data collection, and organization in statistics. It covers various sampling techniques, including random and non-random sampling, as well as methods for organizing and presenting data through tables and graphs. The module also highlights the significance of proper data handling to ensure accurate representation and analysis in research.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

Module 4

Data Management

At the end of this module you are expected to:

1. differentiate a representative sample from biased sample;


2. Identify the appropriate way of collecting data;
3. collect statistical data;
4. organize data in table;
5. construct graphs for sets of data;
6. identify the most appropriate graph for sets of data;
7. determine a single value that accurately describes the center of the distribution
and represents the entire distribution of scores;
8. describe how similar a set of scores are to each other;
9. illustrate the normal distribution of data; and
10. determine the significant relation of data.

Lesson 1. DATA GATHERING AND ORGANIZING DATA

Statistics is often defined as a branch of mathematics that deals with a number


obtained from a set of data that maybe used to represent data. However, the study of
statistics is more merely about numbers or quantity. It is a way of reasoning in dealing
with information, whether numerical or non-numerical to understand the world.

Every lesson that will be included in this module get in the way in statistics to
avoid misused or abused of data or information. Thus, there is a need for you to have
knowledge of statistics and its concepts and principles in order to function well in our
society.
The reliability of a conclusion depends on whether the sample is appropriately
chosen to represent the population. Let us be careful for sampling is one of the crucial
elements of statistical inference.

1.1. Sampling
Let us considered making fruit drinks using powdered flavored drinks in a sachet.
If we pour the powdered drinks in a pitcher with water and get a spoonful from the top
before stirring, then we will get idea that it is bland. On the other hand, if you get a
spoonful of it from the bottom, we will get another misleading idea that it is very sweet.
But if you stir it first before taking a spoonful, what you taste is more representative of
the whole drink.
The process of obtaining samples is called sampling.
Random or probability sampling is a method by which every element of a
population has an equal chance of being included in the sample.

Guided Questions 1
1. Give 4 examples of selecting sampling.
2. Give 3 example of non-random or non-probability sampling.
3. Give 2 examples of census or complete enumeration.
4. Give 3 examples of observations.
5. Give 3 examples of experiments.
6. Give 4 examples of surveys.

Answer to Guide Question 1


Question 1:
1. To avoid bias, Prof. Calaguas used lottery method to select ten students who would
represent the class in a seminar. In the said method, he wrote each name of his
students on a piece of paper and put them in a box. He asked one student to randomly
draw ten pieces of papers to determine the participants.
2. To be fair, Prof. Sigua conducted a systematic sampling to select ten students who
would represent the class in a contest. By counting, she assigned a number to each of
his students. He opened a book and randomly pointed to a certain page and obtained
page 3. He announced that the assigned number to 10 multiple of 3 would be the
representative.

3. Mrs. Buzon was requested by the school administrator to conduct a survey about
students’ agreement or disagreement about the proposed reduced days of classes. She
obtained the total number of students through the School Registrar as shown below.

Year Level 1st Year 2nd Year 3rd Year 4th Year
Number of Students 2,400 1,500 1,500 600

In order to determine the number of students per year level that she needed to
survey, she used the table below.

Year Level Number of Percentage Number of Students


(Strata) Students In the Sample
1st Year 2,400 2,400 ( 2,400 ) (0.40)=960
=0.40∨40 %
6000
2nd Year 1,500 1,500 ( 1,500 ) (0.25)=375
=0.25∨25 %
6000
3rd Year 1,500 1,500 ( 1,500 ) (0.25)=375
=0.25∨25 %
6000
4th Year 600 600 ( 600 ) (0.10)=60
=0.10∨10 %
6000
Total 6,000 1,770

Based from the table, Mrs. Bauzon decided to conduct the survey to the 960 1 st
Year students, 375 students for 2nd and 3rd Year students and 60 for 4th Year students
who first entered the school campus. The way of deciding the number of students in the
sample is what we call the stratified random sampling. In the said sampling technique,
the number of elements of the population is separated into different categories or strata.
The members of the sample are selected proportionally from each category or stratum
by either lottery method or systematic sampling procedure.

4. Mrs. Bauzon may select the 960 students out of 2,400 1 st Year students by lottery
method. She may also select the students by asking every 5 th 1st Year student who
entered the school campus until she arrives with 960 students. Note that Mrs. Bauzon
used stratified random sampling then lottery method or systematic sampling in choosing
the 1st Year students to be included in the sample. Here, she was using a combination
of several sampling techniques called multi-stage sampling.

Question 2:
Give 3 example of non-random or non-probability sampling.

Note: There are times that random sampling is impractical to use for a particular study.
In such case, we are obliged to use a non-random sampling.
Non-random or non-probability sampling is a sampling technique where elements
of a population are drawn based on the judgement of the researcher.

1. Mr. Cruz and his company were hired by a prospective senatorial candidate to study
whether the latter would win the election. They conducted the surveys or interviews in
places where people voted for the winners in the series of previous elections. They did
this because they believed that the people would again vote for the winners in the next
election. They employed purposive sampling in determining the places to administer the
surveys or interviews. They improved their sampling by applying systematic sampling in
determining the respondents in the chosen places.

2. You are asked by your teacher to conduct a study on how six-year-old brother or
sister or any six-year-old child in your neighborhood study their lessons to have ample
time in observing and interviewing your respondents. Here, you are utilizing the
convenience sampling.
3. You are to investigate the relationship of students’ performance in Math and their
attitude towards the subject. However, you are only given limited time to do the study.
You may only consider 25 out 500 students in your school. This method quota sampling
using small samples. However, you may improve your sampling by using any random
sampling technique in determining the respondents.

Question 3:

Give 2 examples of census or complete enumeration.


Note: There are times that the sample in a study requires the whole population.
Census or complete enumeration is a method of data collection from the entire
population.
1. To know the number of persons in different places in our country, the government
conducts a census by taking into consideration the entire population. The data are
stored in the Philippine Statistics Authority.
2. To know if the breaks for cars produced by a car manufacturing company are of good
quality, the company implements a strict rule for testing all the breaks.

1.2. Observation/Experiment

There are times when it is appropriate to obtain data through direct observation
or by just merely reading records. We apply observation when data can be collected
without any response from people. The experimental method is used to find out cause
and effects relationship.
Question 4:
Give 3 examples of observation

1. An observation to investigate the number of students studying in the library from


1 PM to 5 PM.

2. An observation to determine the number of students who go to a mall after


school hours.

3. An observation to find the number of expired food products in a store.

Question 5:
Give 3 examples of experiments.

1. A drug company wants to test the effectiveness of its new product in treating
influenza. An experiment or clinical test is done by treating 50 persons with the
new product and another 50 persons using existing drug. The result are analyzed
statistically to determine if the new product is significantly effective in treating
influenza.

2. At the start of the semester, Mrs. Bismonte administered an examination to two


of her classes in order to determine the entry knowledge of the students. During
the semester, she applied cooperative learning approach to one class and the
usual approach to another class. At the end of the semester, she again
administered an examination and compared the results of the two classes so that
she could know if the cooperative learning approach was more effective than the
usual approach.

3. In an oil company, the use of unleaded gasoline was tested to a good condition
car if it will consume gasoline in the same way to the same car by using regular
gasoline. Both cars will be driven by the same driver and will run in a 30 km road
one at a time with constant speed of 80km/hr. After the test, it was found out
through the reading of the fuel gauge that the same car used more unleaded
gasoline than the regular gasoline. This means that regular gasoline burned
faster that unleaded gasoline.

1.3. Survey

We sometimes use survey when data can only be obtained through responses
from people in a sample. We can obtain through questionnaires which may be
distributed by hand, face-to-face or phone interviews.

Question 6:

Give 4 examples of surveys.

1. A survey to investigate the attitude of the students in studying Mathematics

2. A survey to investigate the opinions of students about changing the grading


system
3. A short face-to-face interview to know the opinions of consumers about a new
bath soap

To avoid bias, questions in the survey should not lead the respondents to answer
towards a desired answer. It is also important to consider the respondents’ emotion so
as to avoid causing embarrassment or unpleasant feeling.

4. The survey question “Do you support the attempt of Congress to reduce the
number of persons who will acquire HIV through unsafe sex by passing the
Reproductive Health Bill?” is biased towards the passing of the bill. The
question may be improved as “Are you in favor of the Reproductive Health
Bill?” Why?

Note: Have you experienced answering a survey questionnaire? Is it not


exhausting to answer it, right? Hence, as much as possible, make sure that
your survey questionnaires are short and precise. Moreover, the questions
must be easy to understand and free from unnecessary technical terms to
avoid misunderstanding or misinterpretation.
Key Points

1. Sampling is the process of obtaining samples.


2. Random or probability sampling is a method by which every element of a
population has an equal chance of being included in the sample. The most
popular random samplings are lottery, systematic, stratified methods and the
combination of two or three sampling methods called multi-stage sampling.
3. There are times that random sampling is impractical to use for a particular study.
In such case, we are obliged to use a non-random sampling. Non-random or
non-probability sampling is a sampling technique where elements of a population
are drawn based on the judgement of the researcher.
4. There are times that the sample in a study requires the whole population. Census
or complete enumeration is a method of data collection from the entire
population.
5. Observation when data can be collected without any response from people.
6. The experimental method is used to find out cause and effects relationship.
7. Survey is use when data can only be obtained through responses from people in
a sample.
Assessment
Name: _________________________ Date: _______________
Year & Section: __________________ Score: ______________

I. DATA GATHERING

1. Give 5 examples of selecting sampling.

2. Give 3 examples of non-random sampling.

3. Give 3 examples of census.

4. Give 2 examples of observation and 2 examples of experiment.

5. Give 3 examples of surveys.


DATA ORGANIZATION
2.1. WHAT IS DATA ORGANIZATION?

 A process organizing collected factual material commonly accepted in the


scientific community as necessary to validate research findings.
 “Research data is data that is collected, observed, or created, for purposes of
analysis to produce original research results” (Boston University Libraries, n.d.a).

2.2. WHY IS DATA IMPORTANT IN RESEARCH?


 Data are intended to represent facts and without proper preservation of the
context of collection and interpretation, may become meaningless (Boston
University Libraries, n.d.a).
 The collection of data and its analysis assists researchers with discovering
answers to their research questions and hypotheses. In some cases, it even
predicts future outcomes (Office of Research Integrity, n.d.a).

2.3. WAYS OF ORGANIZING DATA IN RESEARCH


 Frequency Distribution Table
 Stem and Leaf Diagram
 Chart

Guided Question 2

1. The following data represents the scores of 10 students

8 6 4 5 8 8 9 10 10 6

Construct a table with three columns. The first column shows what is being arranged in
ascending order (i.e. the scores).

2. The following data represents the ages of 20 respondents:


21 26 18 45 32 41 42 22 28 26 33 20 26
44 46 21 24 36 39 30
Construct a table with three columns using the frequency distribution table for grouped
data.

3. Construct a leaf and stem diagram that will represent the data below for a science
test scores for the third grading period (out of 100%):

97 92 77 82 96 75 68 80 79 96 21 34 55
84 87 68 87 88 97 81

4. The following data represents Peters’ Grades in Science subject for 1st – 4th
quarter. Construct a bar chart based on the table below.

Quarter Grades
First 84

Second 90

Third 89

Fourth 93

5. The following data represent the monthly household expenses of Rich family.
Construct a pie chart based on the table below.

Household
Amount
Expenses

Internet 1,000

Electricity 2,000

Grocery 4,000
Other 3,000

6. Construct a line chart on the following data that show daily temperature in Luna, La
Union, recorded for 5 days in Degrees Celsius.

DAYS °C
MONDAY 29
TUESDAY 33
WEDNESDAY 31

THURSDAY 36
FRIDAY 34

7. The following data represents the number of respondents aged 8-55 who are
disabled.
Age (years) Frequency
8 – 15 10
16 - 23 14
24 - 31 19
32 - 39 12
40 - 47 14
48 - 55 25

Answers to Guide Question 2

2.3.1. FREQUENCY DISTRIBUTION TABLE


To construct a frequency table, we use the following steps:

1. Construct a table with three columns. Then in the first column, write down all of
the data values in ascending order.
2. To complete the second column, go through the list of data values and place one
tally mark at the appropriate place in the second column for every data value.
When the fifth tally is reached for a mark, draw a diagonal line through the first
four tally marks. We continue this process until all data values in the list are
tallied.
3. Count the number of tally marks for each data value and write it in the third
column.

2.3.2. TYPES OF FREQUENCY DISTRIBUTION

A. CATEGORICAL/ UNGROUP - Determine the order to list the categories then total
the number of occurrences of each category.

Question 1:

The following data represents the scores of 10 students

8 6 4 5 8 8 9 10 10 6

Construct a table with three columns. The first column shows what is being arranged in
ascending order (i.e. the scores).
The lowest mark is 4. So, start from 4 in the first column as shown below. The second
column is Tally, third is frequency.

Scores Tally Frequency

4 I 1

5 I 1

6 II 2

7 0 0

8 III 3

9 I 1

10 II 2

B. GROUP - It refers to data being organized into groups known as classes.

GUIDELINES:

1. Use between 5 – 20 classes

2. Classes are mutually exclusive

3. Include all classes even if the frequency is zero

4. Use the same width for all classes

5. Use convenient numbers for the class limit

6. The sum of the frequency must total the data set

7. Have enough classes for all the data

8. Remember to use 0 if the class has no data, don’t leave it blank.


Question 2:

The following data represents the ages of 20 respondents:

21 26 18 45 32 41 42 22 28 26 33 20 26
44 46 21 24 36 39 30

1. Determine the highest and lowest value and then compute the Range:
Range = Highest value- Lowest value, Range = 46 - 18 = 28.
2. Decide how many numbers of classes (class size) you want to have. Example: 5
classes
or in calculator, you may use the equation:
Log # of observation/log 2 or √ ¿ of observation
3. Compute the Class width or class interval.

i = Class Interval= Range/# of Classes = 28/5 = 5.6 or 6

4. Lower class limit (Smallest number of each class) and upper class limit (largest
number of each class)

Example: LCL = 18, 24, 30, 36, 42 UCL = 23, 29, 35, 41, 47

5. Class Boundaries – The number that separates the classes from one another by
Subtracting .5 to Lower limit and add 0.5 to upper limit of each class.

Example: (LL) 18 - 0.5 = 17.5 (Class Boundary) and (UP) 23 +


0.5 = 23.5 (Class Boundary)
We proceed as follows:
Age Tally Frequency
18 - 23 IIII 5
24 - 29 IIII - I 6
30 - 35 III 3
36 - 41 II 2
42 - 47 IIII 4

2.3.3. STEM AND LEAF DIAGRAM


A method used to organize statistical data that helps us to see values according
to their size, so we can order them accordingly. In a stem-and-leaf diagram, each data
value is split into a stem and a leaf. The leaf is the last digit to the right. The stem is the
remaining digits to the left. For the number 243, the stem is 24 and the leaf is 3.

Question 3:
Construct a leaf and stem diagram that will represent the data below for a science test
scores for the third grading period (out of 100%):
97 92 77 82 96 75 68 80 79 96
21 34 55 84 87 68 87 88 97 81

STEM LEAVES
2 1
3 4
5 5
6 8 8
7 5 7 9
8 0 1 2 4 7 7 8
9 2 6 6 7 7
2.3.4. GRAPH OR CHART
Graphs or charts condense large amounts of information into easy-to-understand
formats that clearly and effectively communicate important points.

MOST COMMON TYPES OF CHART

a. Bar Chart

b. Pie Chart

c. Line Chart

d. Histogram
Question 4:

a. Bar chart is composed of discrete bars that represent different categories of data.
The length or height of the bar is equal to the quantity within that category of data.
Bar graphs are best used to compare values across categories.
The following data represents Peters’ Grades in Science subject for 1st – 4th quarter.
Quarter Grades
First 84

Second 90

Third 89

Fourth 93

HOW TO CREATE BAR CHART?


Question 5:
b. Pie chart is a circular chart used to compare parts of the whole. It is divided into
sectors that are equal in size to the quantity represented.
The following data represent the monthly household expenses of Rich family.

Household
Amount
Expenses

Internet 1,000

Electricity 2,000

Grocery 4,000

Other 3,000

HOW TO CREATE PIE CHART?

Question 6:
c. Line chart displays the relationship between two types of information, such as
number of school personnel trained by year. They are useful in illustrating trends
over time.
The following data shows daily temperature in Luna, La Union, recorded for 5 days
in Degrees Celsius.
DAYS °C
MONDAY 29
TUESDAY 33
WEDNESDAY 31

THURSDAY 36
FRIDAY 34

HOW TO CREATE LINE CHART?

Question 7:
d. Histogram has connected bars that display the frequency or proportion of cases
that fall within defined intervals or columns. The bars on the histogram can be of
varying width and typically display continuous data.
The following data represents the number of respondents aged 8-55 who are
disabled.
Age (years) Frequency
8 - 15 10
16 - 23 14
24 - 31 19
32 - 39 12
40 - 47 14
48 - 55 25
HOW TO CREATE HISTOGRAM?

Key Points:
1. Research data is data that is collected, observed, or created, for purposes of
analysis to produce original research results.
2. Ways of organizing data in research by Frequency Distribution Table, Stem and
Leaf Diagram and Chart.
3. The frequency distribution table mainly composes of three columns namely data
values arranged in ascending or descending order, tally and frequency.
4. There are two types of frequency distribution, the ungroup and group data.
Ungroup data determine the order to list the categories then total the number of
occurrences of each category while group data refers to data being organized
into groups known as classes.
5. Stem-and-leaf diagram is a method used to organize statistical data that helps us
to see values according to their size, so we can order them accordingly. In a
stem-and-leaf diagram, each data value is split into a stem and a leaf. The leaf is
the last digit to the right. The stem is the remaining digits to the left. For the
number 243, the stem is 24 and the leaf is 3. The stem and the leaf are
separated by two-column table which is open on edges or border.
6. Graphs or charts condense large amounts of information into easy-to-understand
formats that clearly and effectively communicate important points. The common
use charts are Bar Chart, Pie Chart, Line Chart and Histogram.
7. Keep chart simple and avoid flashy special effects. Present only essential
information. Avoid using gratuitous options in graphical software programs, such
as three-dimensional bars, that confuse the reader. If the graph or chart is too
complex, it will not clearly communicate the important points.
8. Title your graph or chart clearly to convey the purpose. The title provides the
reader with the overall message you are conveying.
9. Specify the units of measurement on the x and y-axis. Years, number of
participants trained, and type of school personnel are examples of labels for
units of measurement.
Name: _________________________ Date: _______________
Year & Section: __________________ Score: ______________

Exercise B

A. Construct the following. Use another sheet of paper for your answer.

1. Construct a 3-column table showing the scores, tally and frequency of the following
data arrange in descending order.
25, 23, 30, 27, 28, 25, 21, 28, 30, 27, 22, 25, 26, 27, 21, 29, 20, 23, 21, 25, 21

2. Construct a frequency distribution table on the following data that represents the
ages of 20 respondents to group the data:
22 25 17 44 33 42 43 23 27 25 34 21 25
45 48 22 23 35 38 32

3. Construct a leaf and stem diagram that will represent the data below for a science
test scores for the third grading period (out of 100%):
96 93 78 81 97 74 69 82 76 95
20 33 57 83 85 67 84 85 97 83

B. Construct the following chart. Use another sheet of paper for your answer.
1. Construct a bar chart, given the production of medical facemasks of a pharmaceutical
company for 5 consecutive days to be 2000, 1500, 1750, 2150, and 1350.
2. Given the expenses of a family with 6 members in a month: Food – P4,500; Water –
P850; Electricity – P3,500; ICT bills – P2,500; and Others - P1,500. Construct a pie
chart to represent their part in monthly expenses.
3. The following data shows daily temperature in PSAU 32, 38, 34, 37, and 28 recorded
for 5 days in Degrees Celsius. Developed a line chart to determine the differences
that occurs each day.
Lesson 2: MEASURES OF CENTRAL TENDENCY

 In general terms, central tendency is a statistical measure that determines a single


value that accurately describes the center of the distribution and represents the
entire distribution of scores.

 The goal of central tendency is to identify the single value that is the best
representative for the entire set of data.

 By identifying the "average score," central tendency allows researchers to


summarize or condense a large set of data into a single value.

 Thus, central tendency serves as a descriptive statistic because it allows


researchers to describe or present a set of data in a very simplified, concise form.

 In addition, it is possible to compare two (or more) sets of data by simply comparing
the average score (central tendency) for one set versus the average score for
another set.

THE MEAN, THE MEDIAN AND THE MODE

 It is essential that central tendency be determined by an objective and well-defined


procedure so that others will understand exactly how the "average" value was
obtained and can duplicate the process.

 No single procedure always produces a good, representative value. Therefore,


researchers have developed three commonly used techniques for measuring
central tendency: the mean, the median, and the mode.
Sigma notation Σ
 The sigma notation is a shorthand notation used to sum up a large number of
terms.

 Σx = x1+x2+x3+ … +xn

 One uses this notation because it is more convenient to write the sum in this
fashion.

THE MEAN

The mean is the arithmetic average obtained by adding up all the scores and dividing by
the total number of scores. It is in fact, the numerical average of the set of data.
Formulas for the Mean:

x=
∑x for Ungrouped Data
N
“X bar” equals the sum of all the scores, X, divided by the number of scores, N.

x=
∑ f xm for Grouped Data
N
where:
x m=midpoint of each c lass
f x m = a midpoint multiplied by its frequency
Guide Questions 3

1. Find the mean if the 5 test scores for Calculus I are 95, 83, 92, 81, and 75.
2. Find the mean of the given range of data below;
3. Find the mean of the given grouped data below;
The Hours Spent in Watching TV
Hours Spent f xm fxm
6-7 2
4-5 10
2-3 13
0-1 5
30

Answer to Guide Questions 3


Question 1:
Find the mean if the 5 test scores for Calculus I are 95, 83, 92, 81, and 75.
x = (95+83+92+81+75)/5 = 85.2
(sum up all the tests and divide by the total number of tests.)
Question 2:
Find the mean of the given range of data below;

 What we need to do is find the midpoints of the ranges and then multiply then by
the frequency. So that we can compute the mean.
 The midpoints are 16, 19, 21, 23.5, 27.5, and 32.5.
 x = 16(94,000) + 19(1,551,000) + 21(1,420,000) + 23.5 (1,091,000) +
27.5(865,000) + 32.5(521,000)] /5,542,000 = 22.94
Question 3:
Find the mean of the given grouped data below;
The Hours Spent in Watching TV
Hours Spent F xm fxm
6-7 2 6.5 13
4-5 10 4.5 45
2-3 13 2.5 32.5
0-1 5 0.5 2.5
30 ∑ f x m=¿ ¿93
x=
∑ fx
N
93
x= = 3.10 hours is the mean
30
THE MEDIAN

The median is the middle value of a distribution of data. It is the score that divides
the distribution into two equal parts, so that half the cases are above it and half
below it. Also, it is the middle score, or average of middle scores in a distribution.

How do you find the median?

 First, if possible or feasible, arrange the data from smallest value to largest
value.
 The location of the median can be calculated using this formula: (n+1)/2.
 If (n+1)/2 is a whole number then that value gives the location. Just report the
value of that location as the median.
 If (n+1)/2 is not a whole number then the first whole number less than the
location value and the first whole number greater than the location value will
be used to calculate the median. Take the data located at those 2 values and
calculate the average, this is the median.
Formulas for the Median:
~ n+1
x= ; for Ungrouped Data
2
n
−cf
~ 2 ; for Grouped Data
x=Lm + i
f
where;
~x - Median
Lm – Class boundary (lower limit) of the median class
n – Number of observation
cf – Cumulative frequency above the median class
f – Frequency of the median class
i – Class interval
Guide Question
1. Find the median of the quiz scores 5, 10, 8, 6, 4, 8, 2, 5, 7, and 7.
2. Find the median of a bunch of 10 points quizzes from MMW; 9, 6, 7, 10, 9, 4, 9, 2, 9,
10, 7, 7, 5, 6, and 7.
3. Find the median of the given grouped table below;
Ages of Adults Participating in Covid-19 Vaccination Trial
Ages f cf
18-23 5
24-29 6
30-35 3
36-41 2
42-47 4

Answers to Guide Questions 4


Question 1:
Find the median of the quiz scores 5, 10, 8, 6, 4, 8, 2, 5, 7, and 7.
We start by listing the data in order: 2, 4, 5, 5, 6, 7, 7, 8, 8, and 10
Since there are 10 data values, an even number, there is no one middle number. So we
6+7 13
find the mean of the two middle numbers, 6 and 7, and get ~
x= = =6.5
2 2
Question 2:
Find the median of a bunch of 10 points quizzes from MMW; 9, 6, 7, 10, 9, 4, 9, 2, 9, 10,
7, 7, 5, 6, and 7.
 As you can see there are 15 data points.
 Now arrange the data points in order from smallest to largest 2, 4, 5, 6, 6, 7, 7, 7, 7,
9, 9, 9, 9, 10, and10.
15+1 16
 Calculate the location of the median; = =8. The 8th piece of the data is the
2 2
median. Thus, the median is 7.

Question 3:
Find the median of the given grouped table below;

Ages of Adults Participating in Covid-19 Vaccination Trial


Ages f cf
18-23 5 5
24-29 6 11
30-35 3 14
36-41 2 16
42-47 4 20
20

n
−cf
~ 2
x=Lm + i
f
n 20
= =10; cf = 5; f = 6; and i = 6
2 2
~ 10−5
x=23.5+ (6)
6
~
x=28.50
Thus, 10 adults participated with ages less than 28.5 years of age and 10 adults
participated with ages more than 28.5 years of age in the vaccination trial.

THE MODE
The mode is the most frequent number in a collection of data. The mode is fairly
useless with data like weights or heights where there are a large number of possible
values. The mode is most commonly used for categorical data, for which median and
mean cannot be computed.
Formulas for the mode:
The MODE is the piece of data that occurs most frequently in the data set that can be
taken by simple judgement of counting.

A set of Ungrouped Data can have:


 One mode
 More than one mode
 No mode
For Grouped Data:

^x =Lmo +
( d1
)
d1 + d2
(i)

Guide Questions 5

1. Find the mode of the data 3, 10, 8, 8, 7, 8, 10, 3, 3, and 3.


2. Find the mode of the data 2, 5, 1, 5, 1, and 2.
3. Find the mode of the data 5, 7, 9, 1, 7, 5, 0, and 4.
4. Find the mode of the grouped data below;

Ages of Adults Participating in Covid-19 Vaccination Trial

Ages f
18-23 5
24-29 6
30-35 3
36-41 2
42-47 4

Answers to Guide Questions 5


Question 1:

Find the mode of the data 3, 10, 8, 8, 7, 8, 10, 3, 3, and 3.


The mode of the above example is 3, because 3 has a frequency of 4.

Question 2:
Find the mode of the data 2, 5, 1, 5, 1, and 2.
It has no mode because 1, 2, and 5 have a frequency of 2.

Question 3:
Find the mode of the data 5, 7, 9, 1, 7, 5, 0, and 4.
It has two modes 5 and 7. This is said to be bimodal.
Question 4:
Find the mode of the grouped data below;
Ages of Adults Participating in Covid-19 Vaccination Trial

Ages f
18 - 23 5
24 - 29 6
30 - 35 3
36 - 41 2
42- 47 4
20

^x =Lmo +
( d1
)
d1 + d2
(i)

and
Lmo=23.5 , d 1=6−5=1; d 2=6−3=3 ; i=6

^x =23.5+ ( 1+1 3 )(6)


^x =23.5+ ( ) ( 6 )=25
1
4
Thus, 25 years of age is the most common age to those who participated in the
vaccination trial.
THE WEIGHTED MEAN

Sometimes we wish to average numbers, but we want to assign more importance, or


weight, to some of the numbers.

For instance, suppose your professor tells you that your grade will be based on a
midterm and a final exam, each of which is based on 100 possible points.

However, the final exam will be worth 60% of the grade and the midterm only 40%. How
could you determine an average score that would reflect these different weights?

The average you need is the weighted average.

Guide Questions 6
1. Suppose your midterm test score is 83 and your final exam score is 95. Using
weights of 40% for the midterm and 60% for the final exam, compute the weighted
average of your scores. If the minimum average for an A is 90, will you earn an A?

2. The table below shows Dillon’s fall semester course grades. Use the weighted mean
formula to find Dillon’s GPA for the fall semester if A = 4, B = 3, C = 2, D = 1, and F = 0.
Course Course Course
Grade Units
Biology A 4
Statistics B 3
Business C 3
Psychology F 2
CAD B 2
2. Find the weighted mean of the table below.
Distribution of Laptop Computers per Household
No. of Laptop Computers No. of Households with Laptop Computers
(x) (f)
0 5
1 12
2 14
3 3
4 2
5 3
6 0
7 1
40

Answers to Guide Questions 6


Question 1:
Suppose your midterm test score is 83 and your final exam score is 95. Using weights
of 40% for the midterm and 60% for the final exam, compute the weighted average of
your scores. If the minimum average for an A is 90, will you earn an A?

Your average is high enough to earn an A.

Question 2:
The table below shows Dillon’s fall semester course grades. Use the weighted mean
formula to find Dillon’s GPA for the fall semester if A = 4, B = 3, C = 2, D = 1, and F = 0.

Course Course Course


Grade Units
Biology A 4
Statistics B 3
Business C 3
Psychology F 2
CAD B 2

( 4 x 4 ) + ( 3 x 3 ) + ( 2 x 3 )+(3 x 2) 37
Weighted Mean = = =2.64
14 14
Dillon’s GPA for the fall semester is 2.64.

Question 3:
Find the weighted mean of the table below.
Distribution of Laptop Computers per Household
No. of Laptop Computers No. of Households with Laptop Computers
(x) (f)
0 5
1 12
2 14
3 3
4 2
5 3
6 0
7 1
40

( 0 x 5 ) + ( 1 x 12 ) + ( 2 x 14 ) + ( 3 x 3 ) + ( 4 x 2 ) + ( 5 x 3 )+ ( 6 x 0 )+(7 x 1) 79
Weighted Mean = = =1.975
40 40

The mean number of laptop computers per household is 1.975.


Key Points
1. Mean, Median and Mode. The mean of n numbers is the sum of the numbers
divided by n. The median of a ranked list of n numbers is the middle number if n
is odd, or mean of the two middle numbers if n is even. The mode of a list of
numbers is the number that occurs most frequently.
2. Weighted Mean. The formula for the weighted mean of the n numbers x1, x2, x3,
…, xn is

Weighted Mean=
∑ (x∙w)
∑w
where: ∑ ( x ∙ w ) is the sum of the product formed by multiplying each number by
the assigned weight, and ∑ w is the sum of all the weights.
Name: ________________________ Date: __________
Year & Course:_________________ Score:__________

Exercise C

Measures of Central Tendency

A. Answer the following. Use other sheet of paper for your answer.

1. Find the mean if the 5 test scores for Calculus I are 93, 80, 95, 84, and 78.

2. Find the median of the 8 test scores in Statistics 12, 16, 15, 18, 20, 15, 17, and 19.

3. Find the mode of the test scores in MMW 10, 8, 10, 8, 6, 5, 8, 7, 9 and 10.

4. Find the weighted mean or GWA of the grades of Alvin for the First Semester given
the table below.

Subject Grades of Alvin for the First Semester


Subjects Grades Weighs
Mathematics 88 3
Chemistry 90 2
Social Science 95 1.5
P.E 97 1
Pre-Calculus 86 2
Biology 91 1.5
Values 95 1

B. Answer the following. Use other sheet of paper for your answer.

1. Find the mean of the grouped data below;


Ages of Adults Participating in Covid-19 Vaccination Trial
Ages F
18 – 23 5
24 – 29 6
30 – 35 3
36 – 41 2
42 – 47 4

2. Find the median of the grouped data below and give your interpretation
Ages of Adults Participating in Math Contest
Ages F
16 – 20 5
21 – 25 4
26 – 30 9
31 – 35 5
36 – 40 2

3. Find the mode of the grouped data below and give your interpretation.
Ages of Adults Participating in Math Contest
Ages F
16 – 20 5
21 – 25 4
26 – 30 9
31 – 35 5
36 – 40 2
Lesson 3: MEASURES OF DISPERSION

Measures of dispersion are descriptive statistics that describe how similar a set of
scores are to each other.

 The more similar the scores are to each other, the lower the measure of
dispersion will be.
 The less similar the scores are to each other, the higher the measure of
dispersion will be.
 In general, the more spread out a distribution is, the larger the measure of
dispersion will be.

THE RANGE

The range is defined as the difference between the largest score in the set of data and
the smallest score in the set of data. The formula is;
R = X L - XS
THE STANDARD DEVIATION

The standard deviation is the square root of variance.


Formulas for Standard Deviation:

σ=
√ ∑ (x−μ)2
n
; For population

where: x1, x2, x3, … xn is a population of n numbers with a mean of μ.

ς=
√ ∑ ( x−x )2 ; For sample
n−1
where: x1, x2, x3, … xn is a sample of n numbers with a mean of x .
THE VARIANCE

Variance is defined as the average of the square deviations.

2
σ =
∑ (x−μ)2 ; For population
n

ς =

2 (x−x )2
; For sample
n−1
Guide Questions 7

1. What is the range, standard deviation and variance of the sample below?
2 4 7 12 15
2. A consumer group has tested a sample of 8 size-D butteries from each of 3
companies. The results of the test are shown in the table below. According to these
tests, which company produces batteries for which the values representing hours of
constant use have the smallest standard deviation?
Company Hours of constant use per battery
EverSoBright 6.2, 6.4, 7.1, 5.9, 8.3, 5.3, 7.5, 9.3
Dependable 6.8, 6.2, 7.2, 5.9, 7.0, 7.4, 7.3, 8.2
Beacon 6.1, 6.6, 7.3, 5.7, 7.1, 7.6, 7.1, 8.5

Answers to Guide Questions 7


Question 1:
What is the range, standard deviation and variance of the sample below?
2 4 7 12 15
a) Range
R = X L - XS
R = 15 – 2 = 13
b) Standard deviation for sample
2+4 +7+12+15 40
x= = =8
5 5

ς=
√ ∑ ( x−x )2
n−1
ς=
√ (2−8 )2 + ( 4−8 )2+ ( 7−8 )2+ ( 12−8 )2 + ( 15−8 )2
5−1
ς=√ 29.5=5.43
c) Variance for sample

ς 2=
∑ (x−x )2
n−1
2
ς 2=( √ 29.5 )
2
ς =29.5
Question 2:
A consumer group has tested a sample of 8 size-D butteries from each of 3 companies.
The results of the test are shown in the table below. According to these tests, which
company produces batteries for which the values representing hours of constant use
have the smallest standard deviation?

Company Hours of constant use per battery


EverSoBright 6.2, 6.4, 7.1, 5.9, 8.3, 5.3, 7.5, 9.3
Dependable 6.8, 6.2, 7.2, 5.9, 7.0, 7.4, 7.3, 8.2
Beacon 6.1, 6.6, 7.3, 5.7, 7.1, 7.6, 7.1, 8.5

The mean for each sample of batteries is 7 hours.


a) The batteries from EverSoBright have a standard deviation of

ς 1=
√( 6.2−7 )2 + ( 6.4−7 )2 +…+ ( 9.3−7 )2
7

ς 1=
√ 12.34
7
=1.328 hours

b) The batteries from Dependable have a standard deviation of

ς 2=
√( 6.8−7 )2 + ( 6.2−7 )2 +…+ ( 8.2−7 )2
7

ς 2=
√ 3.62
7
=0.719 hours

c) The batteries from Beacon have a standard deviation of


ς 3=
√ ( 6.1−7 )2 + ( 6.6−7 )2 +…+ ( 8.5−7 )2
7

ς 3=
√ 5.38
7
=0.877 hours

Thus, the batteries from Dependable have the smallest standard deviation. According to
these results, the Dependable company produces the most consistent batteries with
regard to life expectancy under constant use.
Key Points
1. The range of a set of data values is the difference between the greatest data
value and the least data value.

2. The standard deviation and variance. If x1, x2, x3, … xn is a population of n


numbers with a mean of μ.

σ=
√ ∑ (x−μ)2
n
and σ 2=
∑ (x−μ)2 .
n

3. The standard deviation and variance. If x1, x2, x3, … xn is a sample of n numbers
with a mean of ς .

ς=
√ ∑ ( x−x )2
n−1
and 2
ς =
∑ (x−x )2 .
n−1
Name: ________________________ Date: __________
Year & Course:_________________ Score:__________

Exercise D

MEASURES OF DISPERSION

Answer the following.

1. Find the range, standard deviation and variance of the sample below?

1 5 9 13 17

2. A consumer testing agency has tested the strengths of 3 brands of 1/8-inch rope. The
results of the test are shown in the table below. According to the sample test results,
which company produces 1/8-inch rope for which the breaking point has the smallest
standard deviation?

Company Breaking point of 1/8-inch rope (lb)


Trustworthy 122, 141, 151, 114, 108, 149, 125
Brand X 128, 127, 148, 164, 97, 109, 137
NeverSnap 112, 121, 138, 131, 134, 139, 135

3. The fuel efficiency in mi/gal of 12 small utility trucks was measured. The results are
recorded in the table below.
Fuel Efficiency (mpg)
22 25 23 27 15 24 24 32 23 22 25 22

Find the mean and sample standard deviation of these data. Round to the nearest
hundredth.
Lesson 4: MEASURES OF RELATIVE POSITIONS

THE z-SCORES
When a set of data values are normally distributed, we can standardize each
score by converting it into a z-Score. z-Scores make it easier to compare data values
measured on different scales. It reflects how many standard deviations above or below
the mean a raw score is. It is positive if the data value lies above the mean and negative
if the data value lies below the mean.

Formula of z-Score:
x−μ
z= ; for population
σ
Where: x = an element of the data set, the mean = μ, and standard deviation = σ .
x−x
z= ; for sample
ς
Where: x = an element of the data set, the mean = x , and standard deviation = ς .
THE PERCENTILE

Most standardized examinations provide scores in terms of percentile, which are


defined as follows:

pth Percentile
A value x is called the pth percentile of a data set provided p% of the data values are
less than x.
Percentile for a Given Data Value:
Given a set of data and a data value x,
number of data values less than x
Percentile of score x = ⋅100
total number of data values

THE QUARTILE
The three numbers Q1, Q2, and Q3 that partition a ranked data set into four
(approximately) equal groups are called quartiles of the data. For instance, for the data
set below, the values Q1 = 11, Q2 = 29, and Q3= 104 are quartiles of the data.

2, 5, 5, 8, 11, 12, 19, 22, 23, 29, 31, 45, 83, 91, 104, 159, 181, 312, 354
↕ ↕ ↕
Q1 Q2 Q3

The quartile Q1 is called the first quartile. The quartile Q2 is called the second quartile. It
is also the median of the data. The quartile Q3 is called the third quartile.

The Median Procedure for Finding Quartiles:

1. Rank the data


2. Find the median of the data. This is the second quartile, Q2.
3. The first quartile, Q1, is the median of the data values less than Q2. The third
quartile, Q3, is the median of the data values greater than Q2.
BOX-AND-WHISKER PLOTS

A box-and-whisker plot (sometimes called box plot) is often used to provide a visual
summary of a set of data. It shows the median, the first and third quartiles, and the
minimum and maximum values of a data set.

Construction of a Box-and-Whisker Plot:


1. Draw a horizontal scale that extends from minimum data value to the maximum
data value.
2. Above the scale, draw a rectangle (box) with its left side at Q1 and its right side at
Q3.
3. Draw a vertical line segment across the rectangle at the median Q2.
4. Draw horizontal line segment, called a whisker that extends from Q3 to the
maximum.

Q1 Q2 Q3 Whisker

Minimum Maximum

Guide Questions 8
1. Suppose SAT scores among college students are normally distributed with a mean of
500 and a standard deviation of 100. If a student scores a 700, what would be her z-
score?
2. A set of math test scores has a mean of 70 and a standard deviation of 8. A set of
English test scores has a mean of 74 and a standard deviation of 16. For which test
would a score of 78 have a higher standing?
3. In a recent year, the median annual salary for a physical therapist was P360,000. If
the 90th percentile for the annual salary of a physical therapist was P648,000, find the
percent of physical therapist whose annual salary was
a. more than P360,000.
b. less than P648,000.
c. between P360,000 and P648,000.
4. On a reading examination given to 900 students. Elaine’s score of 602 was higher
than the scores of 576 of the students who took the examination. What is the percentile
for Elaine’s score?
5. Find the quartiles Q1, Q2, and Q3 of the following data 20, 30, 25, 23, 22, 32, 36.
6. The following list the weights, in ounces, of 15 avocados in a random sample. Find
the quartiles of the data.
Weights in Ounces of Avocados
12.4 10.8 14.2 7.5 10.2 11.4 12.6 12.8 13.1 15.6
9.8 11.4 12.2 16.4 14.5
7. Construct the box-and-whisker plot of the data set: 85,92,78,88,90,88,89.

Answer to Guide Questions 8


Question 1:
Suppose SAT scores among college students are normally distributed with a mean of
500 and a standard deviation of 100. If a student scores a 700, what would be her z-
score?
x−μ
z=
σ
700−500
z=
100
200
z= =2
100
Her z-score would be 2 which means her score is two standard deviations above the
mean.
Question 2:

A set of math test scores has a mean of 70 and a standard deviation of 8. A set of
English test scores has a mean of 74 and a standard deviation of 16. For which test
would a score of 78 have a higher standing?
x−μ
z=
σ
78−70 8
For Math: z= = =1
8 8
78−74 4 1
For English: z= = = =0.25
16 16 4
The math score would have the highest standing since it is 1 standard deviation above
the mean while the English score is only 0.25 standard deviation above the mean.

Question 3:
In a recent year, the median annual salary for a physical therapist was P360,000. If the
90th percentile for the annual salary of a physical therapist was P648,000, find the
percent of physical therapist whose annual salary was
a. more than P360,000.
b. less than P648,000.
c. between P360,000 and P648,000.
a. By definition, the median is the 50th percentile. Therefore, 50% of the physical
therapists earned more than P360,000 per year.
b. Because P648,000 is the 90th percentile, 90% of all physical therapists made
less than P648,000.
c. From parts a and b, 90% - 50% = 40% of the physical therapists earned
between P360,000 and P648,000.
Question 4:
On a reading examination given to 900 students. Elaine’s score of 602 was higher than
the scores of 576 of the students who took the examination. What is the percentile for
Elaine’s score?
number of data values less than x
Percentile = ⋅100
total number of data values
576
Percentile = ⋅100=64
900
Thus, Elaine’s score of 602 places her at the 64th percentile.

Question 5:
Find the quartiles Q1, Q2, and Q3 of the following data 20, 30, 25, 23, 22, 32, 36.

Arrange data in ascending form, and n = 7 odd number.

By ascending arrangement: 20 22 23 25 30 32 36

q1 = (1/4) x n = (1/4) x 7 = 1.75; q1 = 2 and Q1 = 22

q2 = (2/4) x n = (2/4) x 7 = 3.5 q2 = 4 and Q2 = 25

q3 = (3/4) x n = (3/4) x 7 = 5.25 q3 = 6 and Q3 = 32

Question 6:

The following list the weights, in ounces, of 15 avocados in a random sample. Find the
quartiles of the data
Weights in Ounces of Avocados
12.4 10.8 14.2 7.5 10.2 11.4 12.6 12.8 13.1 15.6
9.8 11.4 12.2 16.4 14.5
Arrange data in ascending form, and n = 15 odd number.
7.5 9.8 10.2 10.8 11.4 11.4 12.2 12.4 12.6
12.8 13.1 14.2 14.5 15.6 16.4

The median of these 15 data values has a rank of 8. Thus the median is 12.4 which is
also Q2.
There are 7 data values less than the median and 7 data values greater than the
median. The first quartile is the median of the data values less than the median. Thus,
Q1 is 10.8. While, the third quartile is the median of the data values greater than the
median which is 14.2 which is also the Q3.

Question 7:
Construct the box-and-whisker plot of the data set: 85,92,78,88,90,88,89.
Key Points

1. The z-scores for a given data value x is the number of standard deviations that x
is above or below the mean

x−μ x−x
z= ; for population and z= ; for sample
σ ς

2. Percentile – a value of x is called pth percentile of a data set provided p% of the


data values are less than x. Given set of data and a data value of x.

number of data values less than x


Percentile of score x = ⋅100
total number of data values

3. Quartile - the three numbers Q1, Q2, and Q3 that partition a ranked data set into
four (approximately) equal groups are called quartiles of the data. The quartile Q 1
is called the first quartile. The quartile Q2 is called the second quartile. It is also
the median of the data. The quartile Q 3 is called the third quartile is the median of
the data values greater than Q2..

4. A box-and-whisker plot is often used to provide a visual summary of a set of


data. It shows the median, the first and third quartiles, and the minimum and
maximum values of a data set.
Name: ________________________ Date: __________
Year & Course:_________________ Score:__________
MEASURES OF RELATIVE POSITION

Answer the following.

1. What will be the miles per gallon for a Toyota Innova when the average km/li is 11, it
has a z value of 1.5 and a standard deviation of 2?

2. Raul has taken two tests in his Chemistry class. He scored 72 on the first test, for
which the mean of all scores was 65 and the standard deviation was 8. He received a
60 on a second test, for which the mean of all scores was 45 and the standard deviation
was 12. In comparison to other students, did Raul do better on the first test or the
second test.

3. Find the quartiles Q1, Q2, and Q3 of the following data 20, 30, 25, 23, 22, 32, 36, 18.

4. Construct the box-and-whisker of the given value data 15, 83, 75, 12, 19, 74, 21.

5. Find the percentiles P8, P50, and P85 of the following data 20, 30, 25, 23, 22, 32, 36 .
Lesson 5: NORMAL DISTRIBUTIONS
Normal Distribution – is a bell-shaped continuous distribution widely used in statistical
inference. The bell-shaped curve is symmetric about a vertical line through the mean of
the data.
Properties of a Normal Distribution:
 The graph is symmetric about the vertical line through the mean distribution.
 The mean, median, and the mode are equal.
 The y-value of each point on the curve is the percent (expressed as a decimal)
of the data at the corresponding x-value.
 Areas under the curve that are symmetric about the mean are equal.
 The total area under the curve is 1.

Empirical Rule for a Normal Distribution:


In a normal distribution, approximately
 68% of the data lie within 1 standard deviation of the mean.
 95% of the data lie within 2 standard deviations of the mean.
 99.7% of the data lie within 3 standard deviations of the mean.

The Normal Distribution


The Standard Normal Distribution:

It is often helpful to convert data value x to z-scores, as we did in the previous section
by using the z-score formulas;
x−μ x−x
zx= ∨z x =
σ σ
The standard normal distribution is the normal distribution that has a mean of 0 and a
standard deviation of 1.

Guide Questions 9
1. A survey of 1000 U.S. gas stations found that the price charged for a gallon of regular
gas could be closely approximated by a normal distribution with a mean of USS 3.10
and a standard deviation of USS 0.18. How many of the station charge
a. between USS 2.74 and USS 3.46 for a gallon of regular gas?
b. less than USS 3.28 for a gallon of regular gas?
c. more than USS 3.46 for a gallon of regular gas
2. Find the standard normal distribution between z = -1.44 and z = 0.
3. A soda machine dispenses soda into 12-ounce cup. Test show that the actual
amount of soda dispensed is normally distributed with a mean of 11.5 oz and a standard
deviation of 0.2 oz.
a. What percent of cups will receive less than 11.25 oz of soda?
b. What percent of cups will receive between 11.2 oz and 11.55 oz of soda?
c. If a cup is filled at random, what probability that the machine will overflow the
cup?
4. The OnTheGo company manufactures laptop computers. A study indicates that the
life span of its computers are normally distributed with a mean of 4.0 years and a
standard deviation of 1.2 years. How long a warranty period should the company offer if
the company wishes less than 4% of its computer to fail during the warranty period
Answers to Guide Questions 9
Question 1:
A survey of 1000 U.S. gas stations found that the price charged for a gallon of regular
gas could be closely approximated by a normal distribution with a mean of USS 3.10
and a standard deviation of USS 0.18. How many of the station charge
a. between USS 2.74 and USS 3.46 for a gallon of regular gas?
The USS 2.74 per gallon price is 2 standard deviations below the mean. The USS 3.46
price is 2 standard deviations above the mean. In a normal distribution, 95% of all data
lie within 2 standard deviations of the mean. Therefore, approximately

(95%)(1000) = (0.95)(1000) = 950 of the stations charge between USS 2.74 and
USS 3.46 for a gallon of regular gas.

μ−2 σ μ μ+2 σ
95 %
b. less than USS 3.28 for a gallon of regular gas?
The USS 3.28 price is 1 standard deviation above the mean. In a normal distribution,
34% of all data lie between the mean and 1 standard deviation above the mean.
Thus, approximately
(34%)(1000) = (0.34)(1000) = 340 of the stations charge between USS 3.10 and
USS 3.28 for a gallon of regular gasoline. Half of the 1000 stations or 500 stations,
charge less than the mean. Therefore, about 340 + 500 = 840 of the stations charge
less than USS 3.28 for a gallon of regular gas. (Draw the normal distribution graph.)
c. more than USS 3.46 for a gallon of regular gas?
The USS 3.46 price is 2 standard deviations above the mean. In a normal
distribution, 95% of all data are within 2 standard deviation above the mean. This
means that the other 5% of the data will lie either more than 2 standard deviations
above the mean or more than 2 standard deviations below the mean. We are
interested only in the data that are more than 2 standard deviations above the mean
which is ½ of 5% or 2.5% of the data. Thus, approximately
(2.5%)(1000) = (0.025)(1000) = 25 of the stations charge more than USS 3.46 for a
gallon of regular gas. (Draw the normal distribution graph.)

Question 2:

Find the standard normal distribution between z = -1.44 and z = 0.

Because the standard normal distribution is symmetrical about the center line z = 0, the
area of the standard normal distribution between z = -1.44 and z = 0 is equal to the area
between z = 0 and z = -1.44. The entry associated with z = 1.44 is 0.425. Thus, the area
of the standard normal distribution between z = -1.44 and z = 0 is 0.425 square unit.

Question 3:

A soda machine dispenses soda into 12-ounce cup. Test show that the actual amount
of soda dispensed is normally distributed with a mean of 11.5 oz and a standard
deviation of 0.2 oz.
a. What percent of cups will receive less than 11.25 oz of soda?
x−x
Recall that the formula for the z-score for a data value x is z x = . Thus, the
σ
z-score for 11.25 oz is;
11.25−11.5
z 11.25= =−1.25
0.2
The table for areas and z-scores indicate that 0.394(39.4%) of the data in a normal
distribution are between z = 0 and z = 1.25. Because the data are normally
distributed, 39.4% of the data is also between z = 0 and z = -1.25. The percent data
to the left of z = -1.25 is 50% - 39.4 = 10.6%. Thus. 10.6% of the cups filled by soda
machine will receive less than 11.25% oz of soda.
b. The z-score of 11.55 oz is
11.55−11.5
z 11.55= =0.25
0.2
It indicates that 0.099(9.9%) of the data in a normal distribution is between z = 0
and z = 0.25.
11.2−11.5
z 11.2 = =−1.5
0.2
It indicates that 0.433 (43.3%) of the data in a normal distribution are between
z = 0 and z = 1.5. Because the data are normally distributed, 43.3% of the data
are also between z = 0 and z = -1.5. Thus, the percent of the cups that the
vending machine will fill with between 11.2 oz and 11.55 oz of soda is 43.3% +
9.9% = 53.2%.
c. A cup will overflow if it receives more than 12 oz of soda. The z-score for 12 oz is
12−11.5
z 12= =2.5
0.2
It indicates that 0.494(49.4%) of the data in the standard normal distribution are
between z = 0 and z = 2.5. The percent of data to the right of z = 2.5 is
determined by subtracting 49.4% from 50%. Thus, 0.6% of the time the machine
produces an overflow, and the probability that a cup filled at random will overflow
is 0.006.

Question 4:
The OnTheGo company manufactures laptop computers. A study indicates that the life
span of its computers are normally distributed with a mean of 4.0 years and a standard
deviation of 1.2 years. How long a warranty period should the company offer if the
company wishes less than 4% of its computer to fail during the warranty period?

The standard normal distribution with 4% of the data to the left of some unknown z-
score and 46% of the data to the right of the z-score but to the left of the mean of 0.
Using the Table Area Under the Standard Normal Curve, we find that the z-score
associated with an area of A = 0.46 is 1.75. Our unknown z-score is to the left of 0, so it
must be negative. Thus zx = -1.75. If we let x the time in years that a computer is use,
then x is related to the z-scores by the formula
x−x
zx=
s
Solving for x with x=4.0 , s=1.2 ,∧z=−1.75 gives us
x−4.0
−1.75=
1.2
(−1.75 ) ( 1.2 )=x−4.0
x=4.0−2.1
x=1.9
Hence, the company can provide a 1.9-year warranty and expect less than 4% of its
computers to fail during the warranty period.
Key Points

1. Frequency Distribution displays a data set by dividing the data into intervals or
classes and the listing the number of data values that fall into each interval. A
relative frequency distribution lists the percent of data in each interval.

2. A normal distribution of data is a bell-shaped curve that is symmetric about a


vertical line through the mean. The y-value of each point on the curve is the
percent of the data at the corresponding x-value. The total area under the curve
is 1.

3. The empirical rule for a normal distribution states that approximately 68% of the
data lie within 1 standard deviation of the mean, 95% of the data lie within 2
standard deviation of the mean, and 99.7% of the data lie within 3 standard
deviation of the mean.

4. Using the standard normal distribution is the normal distribution that has a mean
of 0 and a standard deviation of 1. Any normal distribution can be converted into
the standard normal distribution by converting data values to their z-scores.
Then, the percent of data values that lie in a given interval can be found as the
area under the standard normal curve between the z-scores of the endpoints of
the given interval. The areas under normal distribution will be given in a table
during the discussion.
Name: ________________________ Date: ___________
Course & Year: _________________ Score: __________
NORMAL DISTRIBUTION
Answer the following. Use other paper sheets for your answers.
1. A vegetable distributor knows that during the month of August, the weights of its
tomatoes are normally distributed with a mean of 0.61 lb and a standard deviation of
0.15 lb. What percent of tomatoes weighs less than 0.76 lb? In a shipment of 6,000
tomatoes, how many tomatoes can be expected to weigh more than 0.31 lb?

2. Human are on average, taller today they were 200 years ago. Today, the mean
height of a 14-year-old is about 65 in. Use the table below answer the following
questions.

Heights of a Group of 19th Century Boys, Age 14


Height (in) Percent of Boys
Under 50 0.2
50 – 54 7.0
55 – 59 46.0
60 – 64 41.0
65 – 69 5.8

3. A highway study of 8,00 vehicles that passed by a checkpoint found that their speeds
were normally distributed with a mean of 61 mph and a standard deviation of 7 mph.
How many vehicles had a speed of more than 68 mph? How many of the vehicles had a
speed of less than 40 mph?
4. Find the area of the standard normal distribution between z = - 1.44 and z = 0.
5. Find the area of the standard normal distribution to the right of z = 0.82. Also draw the
graph showing the area covered.
Lesson 6: LINEAR REGRESSION AND CORRELATION

When performing research studies, scientist often wish to know whether two variables
are related. If the variables are determined to be related, a scientist may then wish to
find an equation that can be used to model the relationship.

Regression model:

 Relation between variables where changes in some variables may “explain” or


possibly “cause” changes in other variables.
 Explanatory variables are termed the independent variables and the variables to
be explained are termed the dependent variables.
 Regression model estimates the nature of the relationship between the
independent and dependent variables.
- Change in dependent variables that results from changes in independent
variables, ie. size of the relationship.
- Strength of the relationship.
- Statistical significance of the relationship.
The Formula for the Least-Squares Line:
The equation for the least-squares line for the n ordered pairs (x 1,y1), (x2,y2), (x3,y3),…,
(xn,yn) is ^y =ax+ b, where
n ∑ xy−∑ x ∑ y
a= and b= y−a x
n ∑ x −¿ ( ∑ x ) ¿
2 2

Linear Correlation Coefficient:


For the n ordered pairs (x1,y1), (x2,y2), (x3,y3), …, (xn,yn), the linear correlation coefficient
r is given by,
n ( ∑ xy ) −( ∑ x )( ∑ y )
r=
√ n (∑ x )−(∑ x ) ∙ √n (∑ y )−(∑ y )
2 2 2 2
If the linear correlation coefficient r is positive, the relationship between the variables
has a positive correlation. In this case, if one variable increases, the other variable also
tends to increase. If r is negative, the linear relationship between the variables has a
negative correlation. In this case, if one variable increases, the other variable tends to
decrease.
The table below demonstrates how to interpret the size (strength) of a correlation
coefficient.

Guide Question 10

1. A geologist might want to know whether there is a relationship between the duration
of an eruption of a geyser and the time between eruptions. He collected some and were
given in the table below which gives the bivariate data showing the time between two
eruptions and the duration of the second eruption for 10 eruptions of the geyser. He
wants to know the approximate regression equation and if the time between to eruptions
is 200 seconds, then he wants to find out the estimated duration of the second eruption.

The table of the bivariate:


Time between
272 227 237 238 203 270 218 226 250 245
eruptions(in seconds), x
Duration of eruption(in
89 79 83 82 81 85 78 81 85 79
seconds), y

2. Find the equation of the least-squares line for the ordered pairs below and predict the
average speed of an adult man for each of the following stride length 2.8 m and 4.8 m.
The table of the bivariate:

Stride length in meter (x) 2.5 3.0 3.3 3.5 3.8 4.0 4.2 4.5

Speed in m/s (y) 3.4 4.9 5.5 6.6 7.0 7.7 8.3 8.7
3. Find the linear correlation coefficient for stride length versus speed of an adult man in
the table given below. Give your interpretation on the result. Round your result to the
nearest hundredth.
The table of the bivariate:

Stride length in meter (x) 2.5 3.0 3.3 3.5 3.8 4.0 4.2 4.5

Speed in m/s (y) 3.4 4.9 5.5 6.6 7.0 7.7 8.3 8.7
Answer to Guide Question 10
Question 1:
For instance, a geologist might want to know whether there is a relationship between
the duration of an eruption of a geyser and the time between eruptions. A first step in
this determination is to collect some data. Data involving two variables are called
bivariate data. Table below gives the bivariate data showing the time between two
eruptions and the duration of the second eruption for 10 eruptions of the geyser.
The table of the bivariate:
Time between
272 227 237 238 203 270 218 226 250 245
eruptions(in seconds), x
Duration of eruption(in
89 79 83 82 81 85 78 81 85 79
seconds), y
One way to create a model for the relationship between the times between two
eruptions and the duration of the second eruption is to find a line that approximates the
data points plotted in the scatter plot. There are many such lines that can be drawn as
shown in the second figure above. In the figure, all the possible lines that can be drawn-
the one that is usually of most interest (the bold line) is called the line of best fit or the
least-squares regression line. The least-square regression line is the line that fits the
data better than any other line that might be drawn. The least-square regression is
defined as is the line that minimizes the sum of the squares of the vertical deviations
from each data point to the line for a set of bivariate.

The Formula for the Least-Squares Line:

The equation for the least-squares line for the n ordered pairs (x 1,y1), (x2,y2), (x3,y3),…,
(xn,yn) is ^y =ax+ b, where
n ∑ xy−∑ x ∑ y
a= and b= y−a x
n ∑ x 2−¿ ( ∑ x ) ¿
2

Apply the formula to the given data above, we first find the value of each summation:
No. of
Eruption X Y xy x2
s
1 272 89 24208 73984
2 227 79 17933 51529
3 237 83 19671 56169
4 238 82 19516 56644
5 203 81 16443 41209
6 270 85 22950 72900
7 218 78 17004 47524
8 226 81 18306 51076
9 250 85 21250 62500
10 245 79 19355 60025
19663 57356
2386 822
Σ 6 0

n = 10
5,692,99
2
(Σx) = 6

10(196,636)−(2,386)(822)
a= =0.1189559666
( 10 ) (573,560 )−(5,692,996)
We find the x and y ,

x=
∑ x = 2,386 =238.6∧ y = ∑ y = 822 =82.2
n 10 n 10
and use them to find the y-intercept, b.
b= y−a x
b=82.2−0.1189559666(238.6)
b=53.81710637
a. The regression equation is ^y =0.1189559666 x +53.81710637. The graph of the
regression equation and a scatter plot of the data are shown below.

^y =0.1189559666 x +53.81710637

b. The estimated duration of the eruption after 200 seconds,


^y =0.1189559666 x +53.81710637
^y =0.1189559666(200 sec)+ 53.81710637
^y =78 secondsis the approximate duration of the eruption .

Question 2:
Find the equation of the least-squares line for the ordered pairs below and predict the
average speed of an adult man for each of the following stride length 2.8 m and 4.8 m.

The table of the bivariate:

Stride length in meter (x) 2.5 3.0 3.3 3.5 3.8 4.0 4.2 4.5

Speed in m/s (y) 3.4 4.9 5.5 6.6 7.0 7.7 8.3 8.7

Apply the formula to the given data above, we first find the value of each summation:
No X y xy x2
1 2.5 3.4 8.5 6.25
2 3 4.9 14.7 9
3 3.3 5.5 18.15 10.89
4 3.5 6.6 23.1 12.25
5 3.8 7 26.6 14.44
6 4 7.7 30.8 16
7 4.2 8.3 34.86 17.64
8 4.5 8.7 39.15 20.25
Σ 28.8 52.1 195.86 106.72
n=8
(Σx)2 = 829.44
8(195.86)−(28.8)(52.1)
a= =2.730263158
( 8 ) (106.72 ) −(829.44 )
We find the x and y ,

x=
∑ x = 28.8 =3.6∧ y = ∑ y = 52.1 =6.5125
n 8 n 8

and use them to find the y-intercept, b.


b= y−a x
b=6.5125−2.730263158(3.6)
b=−3.316447368
a. The regression equation is ^y =2.730263158 x−3.316447368 . The graph of the
regression equation and a scatter plot of the data are shown below.

^y =2.730263158 x−3.316447368

b. The predicted average speed of an adult man with a stride of 2.8 m is,
m
^y =2.730263158 ( 2.8 )−3.316447368 ≈ 4.328 .
sec
The predicted average speed of an adult man with a stride of 4.8 m is,
m
^y =2.730263158 ( 4.8 )−3.316447368 ≈ 9.789 .
sec
Question 3:
Find the linear correlation coefficient for stride length versus speed of an adult man in
the table given below. Give your interpretation on the result. Round your result to the
nearest hundredth.

The table of the bivariate:

Stride length in meter (x) 2.5 3.0 3.3 3.5 3.8 4.0 4.2 4.5

Speed in m/s (y) 3.4 4.9 5.5 6.6 7.0 7.7 8.3 8.7

Apply the formula to the given data above, we first find the value of each summation:
No x y Xy x2 y2
1 2.5 3.4 8.5 6.25 11.56
2 3 4.9 14.7 9 24.01
3 3.3 5.5 18.15 10.89 30.25
4 3.5 6.6 23.1 12.25 43.56
5 3.8 7 26.6 14.44 49
6 4 7.7 30.8 16 59.29
7 4.2 8.3 34.86 17.64 68.89
8 4.5 8.7 39.15 20.25 75.69
Σ 28.8 52.1 195.86 106.72 362.25

n=8
(Σx)2 = 829.44 (Σy)2 = 2,714.41
n ( ∑ xy ) −( ∑ x )( ∑ y )
r=
√ n (∑ x )−(∑ x ) ∙ √n (∑ y )−(∑ y )
2 2 2 2

8 ( 195.86 )− (28.8 )( 52.1 )


r=
√ 8 ( 106.72¿ )−829.44 ∙ √8 ( 362.25 )−2,714.41
r =0.9937148574 (Which indicates a very high positive correlation between man’s stride
length and his speed. That is as a man’s stride length increases his speed also
increases.)
Key Points

1. Least-Squares Line. Bivariate data are data given in ordered pairs.


2. The-squares regression line or least squares line for a set of bivariate data is the
line that minimizes the sum of the squares of the vertical deviations from each
data point to the line.
3. The equation of the least-squares line for the n ordered pairs (x 1,y1), (x2,y2),
(x3,y3),…,(xn,yn) is ^y =ax+ b, where

n ∑ xy−∑ x ∑ y
a= and b= y−a x
n ∑ x 2−¿ ( ∑ x ) ¿
2

4. The equation of the least-squares line can be used to predict the value of one
variable when the value of the other variable is known.

5. Linear Correlation Coefficient r measures the strength of a linear relationship


between two variables.
6. The closer the IrI to 1, the stronger the linear relationship is between the variable.
For the n ordered pairs (x1,y1), (x2,y2), (x3,y3), …, (xn,yn), the linear correlation
coefficient r is given by,

n ( ∑ xy ) −( ∑ x )( ∑ y )
r= .
√ 2
√ 2
n ( ∑ x 2 ) −( ∑ x ) ∙ n ( ∑ y 2 )−( ∑ y )
1.
Name: ________________________ Date: ___________
Course & Year: _________________ Score: __________
LINEAR REGRESSION AND CORRELATION
Answer the following. Use other sheets of paper for your answers.
1. Find the equation of the least-squares line, the linear correlation coefficient for the
given data. Round the constants a, b and r to the nearest hundredth.
a. (2,6), (3,6), 4,8), (6,11), (8,18)
b. (2,-3), (3,-4), (4,-9), (5,-10), (7,-12)
c. (2,5), (3,7), (4,8), (6,1), (8,18), (9,21)
2. Given the bivariate data:
x 1 2 3 5 6
y 7 5 3 2 1
a. Draw the scatter plot.
b. Find the equation of the least-squares line.
c. Find to the nearest hundredth, the linear correlation coefficient and make
prediction.
3. The average remaining life-times for women of various ages in the United States are
given in the table below.
Average Remaining Lifetimes for Women
Age (x) 0 15 35 65 75
Years (y) 79.9 65.6 46.2 19.5 12.1
a. Find the equation of the least-squares line.
b. Use the equation of the least-squares to estimate the remaining lifetime of a
woman of age 25.
c. Find to the nearest hundredth, the linear correlation coefficient and make
prediction.

You might also like