95% found this document useful (19 votes)
20K views121 pages

Instructional Materials for Statistics

This document provides an overview of an instructional material on elementary statistics and probability. It begins with defining statistics and distinguishing between descriptive statistics and inferential statistics. It then discusses the different types of data, levels of measurement, sampling techniques, and methods of collecting and presenting data. The document provides a table of contents that outlines the 12 lessons in the material, which cover topics such as measures of central tendency, normal distribution, probability, hypothesis testing, and correlation and regression analysis. Evaluation will be based on practice exercises and a reflective journal.

Uploaded by

Ivan Untalan
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
95% found this document useful (19 votes)
20K views121 pages

Instructional Materials for Statistics

This document provides an overview of an instructional material on elementary statistics and probability. It begins with defining statistics and distinguishing between descriptive statistics and inferential statistics. It then discusses the different types of data, levels of measurement, sampling techniques, and methods of collecting and presenting data. The document provides a table of contents that outlines the 12 lessons in the material, which cover topics such as measures of central tendency, normal distribution, probability, hypothesis testing, and correlation and regression analysis. Evaluation will be based on practice exercises and a reflective journal.

Uploaded by

Ivan Untalan
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
  • Nature of Statistics
  • Presentation and Summarizing Data
  • Measures of Central Tendencies
  • Measures of Dispersion
  • Normal Distribution
  • Introduction to Probability
  • Counting Sample Points
  • Probability of an Event
  • Conditional Probability
  • Baye’s Theorem
  • Hypothesis Testing
  • Correlation and Regression Analysis

INSTRUCTIONAL

MATERIALS

FOR

ELEMENTARY STATISTICS

AND PROBABILITY

Course Code:

SEMA 30083

Lessons compiled by:


JOAN D. RURAL
and
JAY-R A. MANAMTAM
COED DESED Faculty
July 2020
No part of this publication may be reproduced or copied by recording or other electronic/
mechanical methods, without the prior written permission of the publisher/compiler via
jamanamtam@[Link]. Faculty members whose names are printed on the cover are only

compilers who collected materials from different authors. This is not for sale and the compilers
have no intention to profit from this.

2
INTRODUCTION
COURSE OVERVIEW
This Instructional Material deals with the study of elementary statistics and probability. It
focuses on the descriptive statistics, measures of central tendency, normal distribution, different
counting techniques, probability of an event, conditional probability, and the Bayes theorem. The
course also includes hypothesis testing and correlation and regression analysis.

The grading system will be as follows.

Practice Exercises (Class Standing) 70%


Reflective Journal 30%
FINAL GRADE 100%

REFLECTIVE JOURNAL DIRECTIONS

A reflective journal is a place to write down your daily or weekly reflection entries. You can write
about a positive or negative event that you experienced, what it means or meant to you, and what you
may have learned from that experience.
Directions: Following these steps will guide you in our reflective journal requirement.

1. Take a picture of yourself together with what you work on the day you read and work on the
exercises. Choose a clear photo or make the photo clear and visible.

2. Create a simple narration on what you feel, what you learn, the difficulties you encounter while
reading and answering the exercises, and what you did to overcome those difficulties.

3. Organize and Compile it (in a single file).

4. The number of photos should not be more that 20, assuming you will be working the lessons per
week.

5. Prepare a soft copy or hard copy of it depending on your situation, if you have internet
connection or none.

6. Submit it together with your answers in modules on the agreed date of submission.

7. Make sure your photo-essay meets the criteria on the rubric. The rubric is provided in the next
page.

3
REFLECTIVE JOURNAL RUBRIC
Criteria Above Average Average Below Average
Achievement of Learner has achieved Learner has achieved Learner has attempted
defined learning all learning outcomes, all of learning all of learning
outcomes demonstrated by valid outcomes. This outcomes. This
solutions. These are achievement is achievement is partially
explored to a deep demonstrated by valid demonstrated by
level, revealing answers with solutions. solutions.
advanced
understanding of
outcomes.
Evidence of Clear progression in Overall progression in Understanding of
progression from week understanding of understanding of learning outcomes is at
to week learning outcomes learning outcomes is the same level from
evident from week to evident but no clear week to week.
week, with advanced progression from week
level of understanding to week.
of the topics
encountered.
Presentation The writings are clear. The writings are clear. The writings are not
Solutions are well Solutions are provided clear. Solutions are
structured. for some answers and provided for some
well structured. answers and not well
structured.

4
TABLE OF CONTENTS

TITLE Pages

Lesson 1: Nature of Statistics 6 − 11

Lesson 2: Presentation and Summarizing Data 12 − 32

Lesson 3: Measures of Central Tendencies 33 − 44

Lesson 4: Measure of Dispersion 45 − 51

Lesson 5: Normal Distribution 52 − 67

Lesson 6: Introduction to Probability 68 − 74

Lesson 7: Counting Sample Points 75 − 79

Lesson 8: Probability of an Event 80 − 88

Lesson 9: Conditional Probability 89 − 96

Lesson 10: Baye’s Theorem 97 − 100

Lesson 11: hypothesis Testing 101 − 106

Lesson 12: Correlation and Regression Analysis 107 − 118

5
LESSON 1

Nature of Statistics

• Define statistics.
• Distinguish between descriptive statistics and inferential statistics.
• Identify and explain the types of data
• List and describe the four levels of measurement.
• Identify and explain the sampling techniques.
• Discuss the methods of collecting and presenting data

Meaning of Statistics and its Background

Statistics is the branch of mathematics that examines and investigates ways to process
and analyze the gathered data. It is a scientific body of knowledge that deals with the collection,
organization, presentation, analysis and interpretation of data. It has evolved rapidly and is now
applied in many fields such as education, governance, researches and other studies.
Statistics can be traced back from ancient times. People compiled statistical data with
regard to all sort of things such as age, taxes, commerce, events and others. As time went by,
statistical work has continued to influence the activities of people in a wider scope from describing
important features of the data and analyzing them to meet specific purpose.

Foundation of Statistics

6
The starting point of statistics can be map out from two fields of interest, namely: the game
of chance otherwise known as gambling and the second is political science. In the early periods,
there were incomplete estimates of the population in the Philippines. Population estimates were
based on church records, births, deaths, marriage and other source of information about the
population such as number of residence certificates issued every year. In American time, a more
systematic way of collecting data was established such as;
• Bureau of Customs – collects, tabulates and disseminates statistics on imports and
exports
• Bureau of Agriculture – keeps records on the number of farms, the cultivated land as
well as irrigated areas
• Bureau of Labor – provides the government with the number of employed and
unemployed citizens as well as the different problems inherent in the work
• National Statistics Office – undertakes the census of population and housing

The game of chance is the second origin of statistics which deals with throwing of dice,
playing cards, and tossing of coins. The early gamblers suspected that the occurrence of events
in various game of chance follow certain laws, but being unschooled, they cannot deduct the laws
from it. It is the famous gambler in the name of Chevalier de Mere who proposed the well-known
problem challenge and so he worked it out with Fermat, another mathematician. They used
different methods and solutions to many problems.

Uses of Statistics

Statistics can be used for many purposes which may be described as follows:

1. Statistics can give a precise description. For instance, in the field of education,
statistical tools are used to get information on enrolment, physical facilities, teachers as
well as the financial aspect which are important for a productive administration and
management.
Another example is the government itself. Statistics can provide data for an effective
management of the affairs of the state. A good record of taxes, cost of living, wages,
population, number of employed and unemployed and other related records which are
useful for intelligent decisions and policy making.
2. Statistics can predict the outcome of experiment or the behaviour of an individual.
In forecasting, statistics is needed so that concerned individuals can plan ahead correctly
and can formulate necessary policies needed with the existing conditions. In business and
economics, statistics play an important role in business forecasting, opening business,
market research as well as quality control.
3. Statistics can be used to test hypothesis. Statistics is used to analyze and interpret
numerical data which will be used for decision making.

Division of Statistics

7
Statistics is divided into two fields.

• Descriptive Statistics – includes the techniques which are concerned with


summarizing and describing numerical data. This method can either be graphical
or computational. It is used to present and analyze information in convenient,
usable and understandable form. It deals with the collection, organization,
presentation, and analysis of data gathered. Mean, median, mode, standard
deviation, variance, coefficient of variation, skewness, and kurtosis are some of
the measurements under descriptive statistics.
• Inferential Statistics – technique by which decisions about a statistical population
are made based only on a sample which have been observed or a judgment which
have been obtained. This kind of statistics is concerned more with generalizing
information or making inference about the population. It draws inferences about
the population based on the data obtained from the sample using the techniques
applied in the descriptive statistics.

The table below shows some differences between descriptive statistics and inferential statistics.

Descriptive Statistics Inferential Statistics

1. Last year’s total enrollees in College 1. The chance that the enrolees in College
Algebra was 1,237 Algebra will increase by 100 is to 20%

2. The average kilograms of carrots sold 2. A recent study showed that eating carrots
daily by a certain distributor is 1,200kgs can ease stomach spasms.

3. 45% of the bar examinees this year 3. The chance that every batch of bar
passed the board exam. examinees will pass the bar exam is 42%

4. According to a recent survey, every elite 4. It is predicted that the average number of
individual owns 2-3 mobiles. mobiles each individual upgrades will
increase every year.

Scales of Measurement

In statistics and quantitative research methodology, various attempts have been made to
classify variables (or types of data) and thereby develop taxonomy of levels or scale of
measurement. Perhaps the best known are those developed by the psychologist Stanley Smith
Stevens. He proposed four types: nominal, ordinal, interval, and ratio.

• Nominal Scale - The nominal scale is the lowest form of measurement because it does
not capture information about the focal object other than whether the object belongs or
doesn’t belong to a category; either you are a smoker or not a smoker, you attended

8
college or you didn’t, a subject has some experience with computers, an average amount
of experience with computers, or extensive experience with computers. No data is
captured that can place the measured object on any kind of scale say, for example, on a
continuum from one to ten. Coding of nominal scale data can be accomplished using
numbers, letters, labels, or any symbol that represents a category into which an object
can either belong or not belong.
• Ordinal Scale - The ordinal scale has at least one major advantage over the nominal scale.
The ordinal scale contains all of the information captured in the nominal scale but it also
ranks data from lowest to highest. Rather than simply categorize data by placing an object
either into or not into a category, ordinal data give you some idea of where data lie in
relation to each other.
• Interval Scale - Unlike the nominal scale that simply places objects into or out of a category
or the ordinal scale that rank orders objects, the interval scale indicates the distance one
object is from another. In the social sciences, there is a famous example often taught to
students on this distinction.

• Ratio Scale - The scale that contains the richest information about an object is ratio
scaling. The ratio scale contains all of the information of the previous three levels plus it
contains an absolute zero point.

The distinction between interval and ratio scales is an important one in the social sciences.
Although both can capture continuous data, you have to be careful not to assume that the lowest
possible score in your data collection automatically represents an absolute zero point.

Classification of Variables

A variable is any characteristic, number, or quantity that can be measured or counted. A


variable may also be called a data item. Age, sex, business income and expenses, country of
birth, capital expenditure, class grades, eye colour and vehicle type are examples of variables. It
is called a variable because the value may vary between data units in a population, and may
change in value over time.

Qualitative vs Quantitative

9
Quantitative variable is a variable which involves numbers and can be obtained by
counting
Example: age, height, salary
Qualitative variable is an attribute or characteristic
Example: marital status, sex, educational achievement
Discrete and Continous
Discrete is a variable which can assume finite or at most countably infinite number of
values, it is obtained by counting
Example: number of books, total enrolees, drop-outs
Continous is a variable which can assume infinite values within a specified interval and
can be obtained by measurement
Example: weight, money, time
Independent vs Dependent
Independent variable is a variable that the researcher controls or manipulates in
accordance with the purpose of the study
Dependent variable is a measure of the effect of the independent variable/s.
Independent variable is the predictor while the dependent is the variable whose value is being
predicted.

Let’s try this!


I. True or False: Write TRUE if the statement is true and FALSE if not.

_____1. Statistics is a branch of mathematics that examines and investigates ways to process
and analyze the data gathered.
_____2. There are two origins of statistics, one is the game of chance and the other one is
lottery.
_____3. The two divisions of statistics are ordinal statistics and nominal statistics.
_____4. Descriptive statistics includes the summarizing and analyzing of all the data collected.
_____5. Data in ordinal scale can be ordered or arranged, but no differences between the data
can be taken that are meaningful

II. Tell whether the following statements require the use of descriptive or inferential

1. Men perform better than women in Math


2. The chance that a student dropped out from school in a certain private school is 17%.
3. The sample mean is 150.
4. The 95% confidence interval for the population mean is 97 to 103.
5. Forty-five percent of the employees of an organization were recorded late for at least 10
working days.

10
6. A forecaster predicts the results of national election using the number of votes cast in 10
out 15 municipalities.
7. Brand A pain medicine brings noticeable relief significantly faster than Brand B pain
medicine.
8. Last year’s total attendance at freshmen orientation was 980 students.
9. A politician would like to predict based on an opinion poll, her chance for winning in the
upcoming election.
10. A businessman wants to determine the average weekly income he had in the past 3
weeks.

III. Categorize the following variables as qualitative or quantitative. If quantitative, identify


the variable as discrete or continuous

1. Sex 9. political affiliation


2. Time [Link] of school
3. status of employment [Link]
4. weight [Link] status
5. hair color [Link]
6. volume 14. gender
7. total enrollees [Link] allowance
8. occupation
IV. Classify each of the following as nominal, ordinal, interval, or ratio

1. brands of softdrinks 7. economic status


2. number of vehicles registered 8. house number
3. zip code numbers 9. educational achievement
4. monthly income 10. height
5. amount of time spent for computer rentals
6. temperature in Celsius

11
LESSON 2

Presentation and Summarizing Data

• Summarize and present data.


• Make inference about populations.
• Distinguish between quantitative and categorical data.
• Construct and interpret various graphical representations of data.
• Define some basic terms in formulation of frequency distribution.
• Organize data into a frequency distribution.
• Construct a stem-and-leaf plot.
• Represent frequency distribution using graphs and charts.
• Give the importance of graphs in statistics?

Population refers to a large collection of objects, places, or things.

Parameter is any numerical value which describes a population

Example:

There are 3400 law students who took the bar exam.

3400 is the parameter

Sample is a subset of the target population.

Statistics is any numerical value which describes a sample.

12
Example: Out of 3,400 law students who took the bar exam, only 1,257 passed.

Sampling is the process of selecting units (e.g., people, organizations)


from a population of interest so that by studying the sample we may fairly
generalize our results back to the population from which they were chosen.

Sampling techniques can be grouped into how selection of items are made
such as probability sampling and non- probability sampling.

1. Probability sampling

Each member of the population has an equal chance to be included in the sample

Types of Probability Sampling:

a. Simple random sampling

This is also known as lottery or fish bowl technique. There are two ways in
using this technique. First is sampling without replacement in which the drawn
papers are no longer returned in the container. The other procedure called
sampling with replacement involves returning to the container the piece of paper
drawn.

When to use
This is preferable to use if the population is not widely spread
geographically. Also, this is more appropriate to use if the population is
more or less homogeneous with respect to the characteristics of the
population.
b. Systematic sampling
Samples are randomly chosen following certain rules set by the researcher.
This involves using the kth member of the population with,
𝑁
k = 𝑛,

where n is the sample size, and N is the population size

13
but there should be a random start.
Example:
Choose a sample size of 100 from N = 500, using systematic random
sampling
500
Step 1. Determine k (period); k = 100 = 5, so this means that you have to include
every 5th member of N after choosing a random start.
Step 2. Randomly choose a number from 1-10. The randomly chosen number will
serve as the random start. You may also use table of random numbers
to choose the random start.
Step 3. If the chosen random start is 10, then the following will comprise the
sample elements.
When to use
This is advisable to us if the ordering of the population is essentially random
and when stratification with numerous data is used.
c. Stratified random sampling
The population is divided into strata or groupings. The samples from each stratum
is drawn independently from the samples from other strata. Samples from each stratum
may be randomly drawn using simple random sampling techniques.
When to use
If the population is such that the distribution of the characteristics of the
respondents under consideration is concentrated in small and spread segment of
the population. Thus, this is preferred to use if precise estimates are desired for
stratified parts of the population and if sampling problems differ in the various strata
of the population.

d. Cluster Sampling

Cluster sampling is sometimes called area sampling because it is usually applied


when the population is large. In this technique, groups or clusters instead of individuals
are randomly chosen.

When to use

Cluster sampling is typically used when the researcher cannot get a


complete list of the members of a population they wish to study but can get a
complete list of groups or 'clusters' of the population. It is also used when a random
sample would produce a list of subjects so widely scattered that surveying them
would prove to be far too expensive, for example, people who live in different postal
districts in the Visayas.

This sampling technique may well be more practical and/or economical


than simple random sampling or stratified sampling.

e. Multi-Stage Sampling

14
Multi-stage sampling represents a more complicated form of cluster sampling in which
larger clusters are further subdivided into smaller, more targeted groupings for the purposes of
surveying. Despite its name, multi-stage sampling can in fact be easier to implement and can
create a more representative sample of the population than a single sampling technique.
Particularly in cases where a general sampling frame requires preliminary construction, multi-
stage sampling can help reduce costs of large-scale survey research and limit the aspects of a
population which needs to be included within the frame for sampling.

2. Non Probability Sampling

Non-probability sampling is a sampling technique where the samples are gathered


in a process that does not give all the individuals in the population equal chances of being
selected.

Types of non-probability sampling:

a. Quota sampling is a method for selecting survey participants. In quota sampling,


a population is first segmented into mutually exclusive sub-groups, just as
in stratified sampling. Then judgment is used to select the subjects or units from
each segment based on a specified proportion.

Example:

An interviewer may be told to sample 200 females and 300 males between
the age of 45 and 60. This means that the interviewer can simply select who they
want to sample (targeting)

b. Convenience sampling - is a process of picking out people in the most convenient


and fastest way to get reactions immediately.

Example: This method can be done by telephone interview to get the immediate
reactions of a certain group of sample for a certain issue.

c. Purposive Sampling A purposive sample, also commonly called a judgmental


sample, is one that is selected based on the knowledge of a population and the
purpose of the study. The subjects are selected because of specific characteristics.

Example:

If the research will be on the methods and techniques in teaching English


language, then teachers in English must be chosen.

Data Gathering Techniques


The next step after the problem has been defined in the study is data collection. Data are
the values that the variables can assume. There are two types of data; namely, the primary and

15
secondary data. Primary data are data collected directly by the researcher himself. These are
the first-hand or original sources. Secondary data are published data made by other researchers
or entity. Primary data can be gathered through the following:

1. Direct or Interview Method


A conversation between two or more people where questions are asked by the interviewer
to elicit facts or statements from the interviewee
Advantages
➢ Useful to obtain detailed information about personal feelings, perceptions and
opinions.
➢ Allow more detailed questions to be asked.
➢ They usually achieve a high response rate.
➢ Respondents’ own words are recorded.
➢ Ambiguities can be clarified and incomplete answers followed up.
➢ Precise wording can be tailored to respondent and precise meaning of questions
clarified.
➢ Interviewers are not influenced by others in the group.

Disadvantages

➢ time – consuming
➢ different interviewers may understand and transcribe interviews in different ways.
2. Indirect or Questionnaire Method

This is a very commonly used method of collecting primary data. Here information are
collected through a set of questionnaire. A questionnaire is a document prepared by the
investigator containing a set of questions. These questions relate to the problem of inquiry directly
or indirectly. Here, the questionnaires are mailed to the informants with a formal request to answer
the question and send them back. For better response the investigator should bear the postal
charges. The questionnaire should carry a polite note explaining the aims and objective of the
enquiry, definition of various terms and concepts used there. Besides this the investigator should
ensure the secrecy of the information as well as the name of the informants, if required. Success
of this method greatly depends upon the way in which the questionnaire is drafted. So the
investigator must be very careful while framing the questions.

The Questionnaire

In some instances, the authenticity of the data gathered through the indirect or
questionnaire method depends on the questionnaire. Therefore, it is a must that the
questions be carefully worded, free from ambiguity, and designed to achieve the purpose.

16
The following are some characteristics of a good questionnaire:

1. It should contain a short letter to the respondents which includes:


a. the purpose of the survey
b. the assurance of confidentiality
c. the name of the researcher
2. There is a descriptive title for the questionnaire.
3. It is designed to achieve its objectives.
4. The directions are clear.
5. It is designed for easy tabulation.
6. It avoids the used of double negatives.
7. It phrases questions well for all respondents.

Types of Questionnaire:

1. Open - it has an unlimited responses


2. Closed – limits the scope of responses
3. Combination – combination of open and closed types of questionnaire.
Types of Questions:

1. Multiple Choice – allows respondent to select answer/s from the list


2. Ranking – respondent ranks the given items
3. Scales – respondent gives his/her degree of agreement to a statement using Likert
scale
4. Open- ended – essay type

3. Registration Method

Registration method refers to continuous, permanent, compulsory recording of the


occurrence of vital events together with certain identifying or descriptive characteristics
concerning them, as provided through the civil code, laws or regulations of each country. The vital
events may be live births, foetal deaths, deaths, marriages, divorces, judicial separations,
annulments of marriage, adoptions, recognitions (acknowledgements of natural children),
legitimations.

4. Observation Method

17
It is a primary method of collecting data by means of direct or indirect contact. As per
Langley P, “ Observations involve looking and listening very carefully.”
Advantages
➢ Collect data where and when an event or activity is happening
➢ Does not rely on people’s willingness to provide information
➢ Directly see what people do rather than relying on what they say or do

Determining the Sample Size


▪ Most surveys conducted are done on a sample basis because of time and cost involved if
the population is used. We use the Slovin’s formula to determine the statistically
acceptable sample size to be extracted from the given population.

Sample Size n= N / (1 + N e2)


Where: n = sample size
N= population size
e= margin of error

Example:
A group of researchers was tasked to survey whether people
from Metro Manila will vote for Mayor Duterte if he is to run for presidency. If there
are 1,000,000 people and .10 margin of error is expected, compute the sample
size.
n = 1,000,000 e = .10
n = N / 1+Ne2
= 1,000,000 / (1+(1,000,000)(.10)2 ) = 1,000,000/(1+1,000,000 x .0001)
= 1,000,000 /101 = 99.99 = 1,000,000/101
n = 9,901 (sample size)

Frequency Distribution Table


A grouped frequency distribution is used when the range of data set is large; the data
must be grouped into classes whether it is categorical data or interval data. For interval data the
classes is more than one unit in width.

18
A. Categorical Frequency Distribution
The categorical frequency distribution is used to organize nominal-level or ordinal-
level type of data. Some examples where we can apply this distribution are gender,
business type, political affiliation, and others.

Example:
Twenty five applicants were given a performance evaluation appraisal. The
data is
High High High Low Average
Average Low Average Average Average
Low Average Average High High
Low Average Average Average High
Average Low Low High Low

Solution:
Step 1: Construct a table as shown below.
Class Tally Frequency Percent
High
Average
Low

Step 2: Tally the raw data.


Class Tally Frequency Percent
High IIIIIII
Average IIIIIIIIIII
Low IIIIIII

Step 3: Convert the tallied data into numerical frequencies.


Class Tally Frequency Percent
High IIIIIII 7
Average IIIIIIIIIII 11
Low IIIIIII 7

Step 4: Determine the percentage. The percentage is computed using the formula:
%= f/n x 100%, where f is the frequency of the class and n is the total number of value.

Class Tally Frequency Percent Found by


High IIIIIII 7 28% 7 / 25 x 100
Average IIIIIIIIIII 11 44% 11 / 25 x 100
Low IIIIIII 7 28% 7 / 25 x 100

B. Ungrouped Frequency Distribution

Example

19
Let us consider the results of a long quiz in Statistics.
30 28 21 28 29 25 24 27
24 23 25 23 27 28 30 26
22 25 26 27 28 26 25 27
24 22 25 25 27 30 29 30
29 28 29 27 24 22 25 28

Since the highest score is 30 and the lowest is 21, the range is 9. Thus, an ungrouped
frequency distribution table is still possible

The Ungrouped Frequency Distribution Table for the Results of a Long Quiz in Elementary
Statistics

Score Tally Frequency


30 IIII 4
29 IIII 4
28 IIIII – I 6
27 IIIII-I 6
26 III 3
25 IIIII-II 7
24 IIII 4
23 II 2
22 III 3
21 I 1

C. Grouped Frequency Distribution


A grouped frequency distribution is used when the range of data set is large; the data must
be grouped into classes whether it is categorical data or interval data. For interval data the classes
are more than one unit in width.

Example 1: The following is the result of a 50-item test in The Teaching Profession. Construct
a frequency distribution and find the following:
a. Range
b. Interval
c. Class limits
d. Class boundaries
e. Relative frequencies
f. Percentages
g. Cumulative frequencies
h. Midpoints

20 40 35 25 25 20 40 40 36 15
20 25 25 30 30 40 25 15 20 40
35 22 26 50 10 10 20 25 25 35
30 30 25 20 20 10 40 45 45 50
29 25 25 20 25 20 25 15 40 35

20
Step 1: Arrange the raw data in ascending or descending order to make it easier to tally the data.
10 10 10 15 15 15 20 20 20 20
20 20 20 20 20 22 25 25 25 25
25 25 25 25 25 25 25 25 26 29
30 30 30 30 35 35 35 35 36 40
40 40 40 40 40 40 45 45 50 50

Step 2: Determine the range.


Find the highest and lowest value.
Highest value (HV) = 50 Lowest Value (LV) = 10
Range = Highest value (HV) - Lowest Value (LV) = 50 – 10 = 40

Step 3: Determine the number of classes.


We can determine the number of classes (k) using the “2 to the k rule”. This will enable us
to select the smallest number (k) for the number classes such that 2k (2 raised to the power of k)
is greater than the number of observations (n). Using our example, there are 50 students (n= 50).
If we apply k = 6, which means we would use 6 classes the 2k = 26 = 64, which is greater than 50.
Therefore, the recommended number of classes is 6.

General Rule in Determining the Number of Classes


Generally, the number of classes for a frequency distribution table varies from 5 to 20,
depending primarily on the number of observations in the data set. It is preferable to have more
classes as the size of a data set increases. The decision about the number of classes depends
on the method used by the teacher.

Step 4: Determine the class interval (i) or width

Generally, the class interval should be equal for all classes. The class interval is generated using
the formula:
𝑅𝑎𝑛𝑔𝑒 𝐻𝑣−𝐿𝑣 40
Suggested Class Interval (i) = 𝑁𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝐶𝑙𝑎𝑠𝑠𝑒𝑠 = 𝑘 = 6 = 6.67 or 7

Where:
Range = Highest value – Lowest value = HV - LV

Class limits - smallest and largest observations (data, events etc.) in each class. Therefore, each
class has two limits: a lower and upper

Class boundaries - midpoints between the upper class limit of a class and the lower class limit of
the next class in the sequence. Therefore, each class has an upper and lower class boundary
Relative frequency - the ratio of the number of times an event occurs to the number of occasions
on which it might occur in the same period.

Percentage – means one out of a hundred. It is used to describe parts of a whole


Cumulative frequency - the total of a frequency and all frequencies so far in a frequency
distribution

Midpoint – the point that is exactly between two other point

21
Select a starting point which can be the smallest data value or any number less than the smallest
data value.

Set individual class limit

We need to add interval to the starting point to obtain the lower limit of next class. Keep adding
until we reach the 6 classes. 10, 17, 24, 31, 38, and 45.

Class Limits
10 – 16
17 – 23 The upper limit for each class may be divided by adding the
24 – 30 class interval to the lower limit minus 1. The nos. 10,17,24,31,38
31 – 37 and 45 are called lower limits while the nos. 16,23, 30, 37, 44
38 – 44 and 50 are referred to as upper limits.
45 – 51

Set the class boundaries in each class. To do this we just need to subtract 0.5 from each lower
class limit and add 0.5 to each upper class limit.
Class Limits Class Boundaries
10 – 16 9.5 – 16.5
17 – 23 16.5 – 23.5
24 – 30 23.5 – 30.5
31 – 37 30.5 – 37.5
38 – 44 37.5 – 44.5
45 – 51 44.5 – 51.5

Step 5: Tally the raw data


Class Limits Class Boundaries Tally
10 – 16 9.5 – 16.5 IIIII-I
17 – 23 16.5 – 23.5 IIIII-IIIII
24 – 30 23.5 – 30.5 IIIII-IIIII-IIIII-III
31 – 37 30.5 – 37.5 IIIII
38 – 44 37.5 – 44.5 IIIII-II
45 – 51 44.5 – 51.5 IIII

Step 6: Convert the tallied data into numerical frequency


Class Limits Class Tally Frequency
Boundaries
10 – 16 9.5 – 16.5 IIIII-I 6
17 – 23 16.5 – 23.5 IIIII-IIIII 10
24 – 30 23.5 – 30.5 IIIII-IIIII-IIIII-III 18
31 – 37 30.5 – 37.5 IIIII 5
38 – 44 37.5 – 44.5 IIIII-II 7
45 – 51 44.5 – 51.5 IIII 4

Step 7: Determine the relative frequency. It is computed by dividing each frequency by the total
frequency.
Class Limits Class Frequency Relative Found by
Boundaries Frequency

22
10 – 16 9.5 – 16.5 6 0.12 6 / 50
17 – 23 16.5 – 23.5 10 0.20 10 / 50
24 – 30 23.5 – 30.5 18 0.36 18 / 50
31 – 37 30.5 – 37.5 5 0.10 5 / 50
38 – 44 37.5 – 44.5 7 0.14 7 / 50
45 – 51 44.5 – 51.5 4 0.08 4 / 50

Step 8: Determine the percentage. It can be found by multiplying 100% each relative frequency.
Class Limits Class Frequency Percentage Found by
Boundaries
10 – 16 9.5 – 16.5 6 12 (6 / 50) x 100
17 – 23 16.5 – 23.5 10 20 (10 / 50) x 100
24 – 30 23.5 – 30.5 18 36 (18 / 50) x 100
31 – 37 30.5 – 37.5 5 10 (5 / 50) x 100
38 – 44 37.5 – 44.5 7 14 (7 / 50) x 100
45 – 51 44.5 – 50.5 4 8 (4 / 50) x 100
Total 50 100

Step 9: Determine the cumulative frequencies. The cumulative frequency can be found by adding
the frequency in each class to the total of frequencies of the classes preceding that class.
Class Limits Frequency Cumulative Found by
Frequency
10 – 16 6 6 6
17 – 23 10 16 6 + 10
24 – 30 18 34 16 + 18
31 – 37 5 39 34 + 5
38 – 44 7 46 39 + 7
45 – 51 4 50 46 + 4

Step 10: Determine the midpoints. The midpoint can be found by getting the average of the upper
limit and the lower limit in each class.
Class Limits Frequency Midpoints Found by
10 – 16 6 13 (10 + 16) / 2
17 – 23 10 20 (17 + 23) / 2
24 – 30 18 27 (24 + 30) / 2
31 – 37 5 34 (31 + 37) / 2
38 – 44 7 41 (38 + 44) / 2
45 – 51 4 43 (45 + 50) / 2

The midpoint may also be computed by first setting the average of the upper limit and
lower limit of the first class (10 + 16)/2 = 13 and continuously add the class interval for succeeding
classes.
Data presented in a grouped frequency distribution are easier to analyze and describe.
However, the identity of individual score is lost due to grouping. For instance, in a class of 21-
24, no one can identify the test score which falls in the said class unless somebody refers to the
original set of data.
In the final presentation of the table the tally is omitted.

23
Another way of presenting data using table is the stem-and-leaf display. It is called
a stem-and-leaf because the grouping forms a “stem” and the values are listed as
“leaves”. It is a way of listing relatively small sets of numerical data. It has two-column
table in which the stems are written on the left column and the leaves on the second
column.

Steam and Leaf Plot


It is a device for presenting quantitative data into graphical format. Its advantage over the
histogram is that we can see the actual observations.
The stem is the leading digit and the leaf is the trailing digit.
Example: Construct a stem and leaf plot
20 40 35 25 25 20 40 40 36 15

20 25 25 30 30 40 25 15 20 40
35 22 26 50 10 10 20 25 25 35

30 30 25 20 20 10 40 45 45 50
29 25 25 20 25 20 25 15 40 35
n = 50
Solution
Stem Leaf
1 0, 0, 0, 5, 5
2 0, 0, 0, 0, 0, 0, 0, 0, 0, 2, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 5, 6, 9
3 0, 0, 0, 0, 5, 5, 5, 5, 6
4 0, 0, 0, 0, 0, 0, 0, 5, 5
5 0, 0

The stem is the tens digit (or the leading digits) while the leaf is the unit digit (trailing digits)

Example
Below are scores of freshmen students in English exam

12 40 15 19 40 31
23 34 37 23 33 36

24
25 26 29 30 21 28
32 29 20 34 26 38
The stem-and-leaf display is shown below

Stem L e a v e s
1 2 5 9
2 0 1 3 3 5 6 6 9 9
3 0 2 3 4 4 7
4 0 0

Data Presentation
Presenting data through pictures or graphics is sometimes more appealing than using
texts. Graphs give more comprehensive but plain view of the numerical relationships even without
going into the longer time of textual, tabular and graphical.
In order for the researcher to interpret the data gathered easily, the researcher must organize the
data. There are two ways of organizing data. One is grouped data. This are data that are
organized and arrange into different classes or categories. While the other one is ungrouped
data which are data that are not organized, or if arrange can only be ascending to descending or
descending to ascending.

➢ Textual Presentation
Data presented in phrase, sentences or paragraph form are said to be in the
textual. However, textual presentation would not be of much use since it makes dull
reading and may not give a good interpretation of the meaning.

Official Population Count of NSO in the year 2010


The final result of the latest Census of Population (POPCEN 2010) conducted by the
National Statistics Office (NSO) placed the Philippine population at 92,337,852 / 92.34M
persons as of May 1, 2010. The 2010 population is higher by 3,763,238 / 3.76M compared to
2007 population of 88,574,614 / 88.57M persons.

Among the 17 regions, CALABARZON (Region- IVA) had the largest population with
12.61M, followed by the National Capital Region (NCR) with 11.86M and Central Luzon
(Region III) with 10.41M. The population of these three regions together comprised more
than one-third (37.47 percent) of the Philippine population.

Among provinces, Cavite had the largest population with 3.09M. Bulacan had the
second largest with 2.92M and Pangasinan had the third largest with 2.78M. On the other
hand, the provinces with a population of less than 100,000 persons were Batanes (16,604),
Camiguin (83,807), and Siquijor (91,066).

As projected by NSO, this 2012 the Philippine population is about 97.6M and by
2014, the population may increase to 101.2M.
*Source: National Statistics Office, Manila
April 04, 2012

25
➢ Tabular Presentation
Table 3.0 Table 3.1

Paula’s Grades Last Semester Price List of School


Uniform Per College
Subjects Grade Number of Units
Plane Geometry 1.75 3 College Price
P.E. 1.50 2
Accountancy P 390.00
Plane Trigonometry 1.75 3
Ecology 2.00 3 Business P 350.00
Advance Algebra 1.50 3
Education P 350.00
Logic 2.00 3
Finance P 400.00

Graphical Presentation
Graphical form is the most effective way to present results in a study since it shows the
statistical values and relationship in a pictorial or diagrammatic form. Graphs give more
comprehensive but plain view of the numerical relationships even without going into the longer
time of reading discussions.

BAR GRAPHS
It is a graphical display of data using bars of different heights. It is displayed either horizontally or
vertically and the bars do not touch each other.

26
Note that in the bar graph, the bars do
not touch each other. This indicates
the discrete nature of the variable
being grouped. Bar graphs are often
used to show the frequencies of
various nominal variables. They are
used to compare magnitude.
HISTOGRAM
A histogram is a graphical representation of the distribution of data. It is an estimate of
the probability distribution of a continuous variable and was first introduced by Karl Pearson. The
rectangles of a histogram are drawn so that they touch each other to indicate that the original
variable is continuous.

27
LINEAR GRAPHS

➢ Frequency Polygon

Frequency polygons are a graphical device for understanding the shapes of distributions.
They serve the same purpose as histograms, but are especially helpful for comparing sets of data.
Frequency polygons are also a good choice for displaying cumulative frequency distributions.

➢ Frequency Ogive

Ogives do look similar to


frequency polygons. The most important
difference between them is that an ogive
is a plot of cumulative values, whereas a
frequency polygon is a plot of the values
themselves. So, to get from a frequency
polygon to an ogive, we would add up the
counts as we move from left to right in the
graph.

Ogives are useful for determining


the median, percentiles and five number
summary of data. Remember that the
median is simply the value in the middle
when we order the data. A quartile is
simply a quarter of the way from the
beginning or the end of an ordered data
set. With an ogive we already know how many data values are above or below a certain point, so
it is easy to find the middle or a quarter of the data set.

28
HUNDRED PERCENT CHARTS

➢ Pie Chart

A pie chart is a divided into sectors,


illustrating numerical proportion. In a pie chart,
the arc length of each sector (and consequently
it’s central angle and area), is proportional to the
quantity it represents.

STATISTICAL MAPS

A special type of map in which the variation in quantity of a factor such as rainfall,
population, or crops in a geographic area is indicated; a dot map is one type.

PICTOGRAMS

A pictogram, also called a pictogramme, pictograph, or simply picto, and also


an 'icon', is

an ideogram that conveys its meaning through its pictorial resemblance to a physical

29
object. Pictographs are often used in writing and graphic systems in which the characters
are to a considerable extent pictorial in appearance.

Often mathematical formula require the addition of many variables Summation or sigma
notation is a convenient and simple form of shorthand used to give a concise expression for a
sum of the values of a variable.

Let x1, x2, x3, …xn denote a set of n numbers. x1 is the first number in the set. xi represents the
ith number in the set.

Let’s try this!


I. Using the table below construct a frequency distribution table.
20 15 10 10 10 5 20 20

25 10 5 15 15 10 5 5

25 20 25 25 10 15 5 10

20 15 15 10 10 5 5 20

Determine the following:


1. Range
2. Interval
3. Class limits
4. Class boundaries
5. Relative frequencies
6. Percentages
7. Cumulative frequencies
8. Midpoints

II. Identification
1. A ___________ series graph represents data that occur over specific period of time under
observation. In addition, it shows for a trend orb pattern in increase or decrease over
period of time
2. A __________ is similar to bar histogram. He bases of the rectangles are arbitrary
intervals whose centre are the codes. The height of each rectangle represents the
frequency of that category. It is also applicable
3. A __________ is used to examine possible relationships between two numerical variables.
The two variables are plot in x-axis and y-axis
4. A __________ is a graph that displays the data using points which are connected by lines.
The frequencies are represented by the heights of the points at the midpoint of the classes.
The vertical axis represents the frequency of the distribution while the horizontal axis
represents the midpoint of the frequency distribution

30
5. The ___________ is used to organized nominal-level or ordinal-level type of data. Some
examples where we can apply this distribution are gender, business type, political
affiliation, and others
III. True or False
Write True on the blank if the statement is correct and false if not.

1. An open type of questionnaire has an unlimited responses.


2. Respondents has equal chances to be included I non-probability
sampling.
3. Indirect method of collecting data is time-consuming.
4. A question which allows respondent to select answers from the list is
an open-ended type.
5. Simple random is also known as lottery.

IV. Construct a stem-and-leaf plot for the following:


1. ABC e-library is studying the number of times their facilities are used daily.
Following is a list of the number of times the library was used during the last 30 days.
45 46 54 43 67 76 65 66 43 56
43 76 45 48 50 71 73 65 58 49
75 67 66 75 74 70 65 60 56 57

Solution:
Stem L e a v e s

2. One of the leading state universities gathered information about the actual number of enrollees
per colleges. The information in the table below shows the existing colleges and the number
of enrollees respectively. Construct a pie chart and bar graph.

College Number of Enrollees

31
College of Education 3, 205
College of Computer and Information Sciences 2, 726
College of Engineering 2, 980
College of Architecture and Fine Arts 1, 517
College of Arts and Letters 2, 892
College of Science 2, 890

a. Pie chart

b. Bar graph

32
LESSON 3

Measures of Central Tendencies

• Calculate the mean, median, mode and midrange as measures of the center of a given
data distribution.
• Discuss the properties of mean, median, mode and midrange.
• Introduce quantiles and the types of distribution
• Compute the weighted mean, geometric mean and combined mean.
• Calculate the mean, median and mode as measures of the center of a given data
distribution.

Measure of central tendency is usually called average. It is a single value that represents
a data set. It is the locator of the center of the data set. We will illustrate how to compute or
calculate it. In this section you can understand what is the purpose of mean, median and mode
for the grouped data, including how to analyze and compute quartile, decile and percentile for the
grouped data also.
Measures of Central Tendency or Average
Average plays a very important role in our daily life and it is an important tool in statistics.
Mean, median, and mode are three kinds of "averages" or sometimes called measures of central
tendency. There are many "averages" in statistics, but these are, I think, the three most common,
and are certainly the three you are most likely to encounter in your pre-statistics courses, if the
topic comes up at all.

MEAN
Most commonly used measures of central tendency. When we speak of getting the
average, we always refer to the mean.
Sample Mean - The sample mean is obtained by adding all the values in your sample and dividing
by the sample size (which is usually denoted by small n). In mathematical notation, we have
Σ𝑥
𝑥̅ =
𝑛
Notice that the symbol for the sample mean is 𝑥̅ , an x with a bar above, and it is

33
read as x bar.
The symbol Σ is the summation sign in mathematics, which means that you should add up all
the values in your sample.
Example 1
Paul Cedric collects the data on the ages of respondents of doctoral degree in
educational management, and his study yields the following:
25 25 24 32 35 35 35 45 43 42 44
Determine the average age of the respondents.

Solution
The mean is the sum of the ages and then dividing by the total number of
respondents.
25+25+ 24+32+35+35+35+45+43+42+44
mean = 𝑥̅ =
11
385
𝑥̅ = = 35
11
Therefore, the average age is 35 years old.

Example 2
If the ages you collected are 22, 21, 48, and 21, then the sample values are written
as
𝑥1 = 22, 𝑥2 = 21, 𝑥3 = 48, 𝑥4 = 21.
Solution
for this particular example, the mean is
𝑥 +𝑥 +𝑥 +𝑥 22+21+48+21 112
𝑥̅ = 1 2 3 4 = = = 28
4 4 4

Population Mean - If we were to calculate the mean from a population instead of a


sample, then we would still proceed in the same way, we would add up all the values in
the population (more values to add) and we would divide by the population size (denoted
by N).
The symbol for the population mean is the Greek letter μ:mu, so we obtain
Σ𝑥
𝜇=
𝑁

Example 3
In the case of the population of students in the classroom with the following ages:
20 21 21 24 25 25 25 25 25 27 27 27 27 27 26 26 26 23 23 28 28 29 29 29
29.
Solution
𝜇=
20+21+21+24+25+25+25+25+25+27+27+27+27+27+26+26+26+23+23+28+28+29+29+29+29
25
642
𝜇 = 25
𝜇 = 25.68
Therefore the average age of the class is 25.68 years old.

Example 4

34
Eight students have scores of 15, 12, 12, 11, 10, 14, 13, and 14. The mean is
15+12+12+11+10+14+13+14 101
𝑥̅ = = 8 = 12.63
8
For group data, it involves organizing n observed values into smaller number of disjoint groups
of values and counting the frequency of each group; it is often presented as a frequency table.
The following are the steps in solving for the mean of grouped data.
1. Find the midpoint for each class. Place them in a column.
2. Multiply the frequency by the midpoint for each class. Place them in another column.
3. Find the sum of the resulting column in step 2.
4. Divide the sum obtained in step 3 by the total number of frequencies. That is,
∑𝑓 ∙ 𝑥𝑚
mean =
𝑛
Example 5
Seventy randomly selected televisions were tested to determine their lifetimes (in months).
The following frequency distribution was obtained.

Classes Frequency
21 – 29 4
30 – 38 10
39 – 47 23
48 – 56 16
57 – 65 12
66 – 74 5

Determine the mean lifetimes (in months) of these television sets.

Solution

Step 1 Find the midpoints of each class and place the values on the third
column.
21+29
𝑥𝑛 = 2 = 25, etc.
Step 2 Next, multiply the midpoint by the frequency for each class and
place the results on the fourth column.
(25)(4) = 100, etc.
Step 3 Find the sum of the fourth column.
∑𝑓 ∙ 𝑥𝑚 = 3 520
Step 4 Divide the sum by n.
∑𝑓 ∙ 𝑥 3520
𝑥̅ = 70 𝑚 = 70 = 50.29 𝑚𝑜𝑛𝑡ℎ𝑠

Hence, the mean lifetime of the television sets is 50.29 months. These steps are summarized in
the following table.

Class Boundaries Frequency Midpoints 𝑓 ∙ 𝑥𝑚


21 - 29 4 25 100
30 - 38 10 35 350
39 - 47 23 45 1035

35
48 - 56 16 55 880
57 - 65 12 65 780
66 - 74 5 75 375
N = 70 Total = 3520

Example 6
Forty randomly chosen patients with dengue were considered for the study. Determine the
average age of patients affected by dengue.

Age No. of Patients


5–9 5
10 – 14 10
15 – 19 6
20 – 24 7
25 – 29 6
30 – 34 4
35 – 39 2
n = 40
Solution

Step 1 Find the midpoint of each class


Step 2 Multiply the midpoint (x) by the frequency (f) and set the sum for all
classes ( ∑𝑓𝑥 )
Step 3 Divide the sum ( ∑𝑓𝑥 ) by n to get the mean (𝑥̅ )

Age Frequency (f) Midpoints (x) 𝑓 ∙ 𝑥𝑚


5–9 5 7 35
10 – 14 10 12 120
15 – 19 6 17 102
20 – 24 7 22 154
25 – 29 6 27 162
30 – 34 4 32 128
35 – 39 2 37 74
n = 40 ∑𝑓𝑥 = 775
∑𝑓𝑥
Mean = 𝑥̅ =
𝑛

775
𝑥̅ =
40

𝑥̅ = 19.375 ≈ 19

36
Therefore, the average age of the patients in the hospital affected with dengue is
19 years old.

Example 7
To get the profile of freshmen, 62 students were asked for their daily allowance. Determine
the average allowance given the following distribution:

Daily Allowance No of Students (f) Midpoint (x) 𝑓𝑥


50 – 59 2 54.5
60 – 69 4 64.5
70 – 79 6 74.5
80 – 89 10 84.5
90 – 99 18 94.5
100 – 109 12 104.5
110 – 119 6 114.5
120 – 129 4 124.5
n = 62 ∑𝑓𝑥 = 775

∑𝑓𝑥
Mean = 𝑥̅ = 𝑛

775
𝑥̅ = 40

𝑥̅ = 19.375 ≈ 19

Therefore, the average age of the patients in the hospital affected with dengue is
19 years old.

MEDIAN
The median of a set of data values is the middle value of the data set when it has been
arranged in ascending order. That is, from the smallest value to the highest value.

Example 6 The grades of 8 students in Mathematics test that had a maximum possible mark
of 50 are given below:
45 22 35 30 38 48 29
Find the median of this set of data values.
Solution: Arrange the data values in order from the lowest value to the highest value:
22 29 30 35 38 45 48
Select the middle value.
22 29 30 35 38 45 48
The fourth value, 35, is the middle value in this arrangement.
Median = 35

Example 7 Ten books were randomly selected and the numbers of pages were recorded as
follows:

37
550, 500, 465, 601, 610, 480, 510, 580, 600, 475
Solution Arrange the data values in order from the lowest value to the highest value:
465, 475, 480, 500, 510, 550, 580, 600, 601, 610
The number of values in the data set is 10, which is even. So, the median is the
average of the two middle values.
5𝑡ℎ 𝑑𝑎𝑡𝑎 𝑣𝑎𝑙𝑢𝑒+6𝑡ℎ 𝑑𝑎𝑡𝑎 𝑣𝑎𝑙𝑢𝑒
Median = 2
510+550 1060
= 2
= 2
= 530

Example 8 Find the median of the data in Example 5.


Solution Step 1 Make a column for the cumulative frequency.
Class Frequency Midpoints 𝑓 ∙ 𝑥𝑚 cumulative frequency
boundaries
20-30 4 25 100 4
30-40 10 35 350 14 cf
𝑳𝒎𝒅 40-50 23 46 1035 37
median
50-60 16 55 880 53
60-70 12 65 780 65
70-80 5 75 375 70
Step 2 Divide n = 70 by 2 to get the halfway point, which is 35.
Step 3 Find the class that contains the 35th value by using the cumulative
frequency distribution. Since 35 is less than 37, then the median
class is the third class.
Step 4 Using formula,
𝑛 70
−𝑐𝑓 −14
𝑥𝑚𝑑 = ( 2 𝑓 ) (𝑤) + 𝐿𝑚𝑑 = ( 2 23 )(5)+40= 44.57months
The median is 44.57 months.

MODE
A statistical term that refers to the most frequently occurring number found in a set of
numbers. The mode is found by collecting and organizing the data in order to count the frequency
of each result. The result with the highest occurrences is the mode of the set. The mode for
grouped data is the modal class. If no number is repeated, then there is no mode for the list.
Example 9 Find the mode for the following list values:
7 7 7 5 8 9 9 9 10

The mode is the number repeated most often. This list has two values that are
repeated three times.
mode = 7 and 9

Example 10 From the data in Example 4, the class with the largest frequency is the third class.
Therefore, this is the modal class.
Class Boundaries Frequency

38
20 – 30 4
30 – 40 10
𝑑1 = 23 − 10 = 13
40 – 50 23
modal class 𝑑2 = 23 − 16 = 7
50 – 60 16
60 – 70 12
70 – 80 5

The mode is the only measure of central tendency that can be used in finding the
most typical case when the data are nominal or categorical.
Grouped Mode
𝑑
𝑚𝑜𝑑𝑒 (𝑀𝑜) = 𝐿𝑀𝑜 + (𝑑 +1𝑑 ) 𝑤
1 2
where 𝐿𝑀𝑜 - lower boundary of the modal class
w - class width
𝑑1 - difference of the frequency of the modal class and the class
preceding it
𝑑2 - difference of the frequency of the modal class and the class
succeeding it

Example 11 Find the mode of the data in Example 9.


Solution Using the formula.
𝑑 13
𝑀𝑜 = 𝐿𝑀𝑜 + ( 1 ) 𝑤 = 40 + (
𝑑1 + 𝑑2
) (5) = 43.25
13+7
Drill 1 Given the following frequency table, find the mean, median and mode.
Class Boundaries frequency Class Mark 𝑓 ∙ 𝑥𝑚 cf
12.5 – 16.5 8
16.5 – 20.5 35
20.5 – 24.5 55
24.5 – 28.5 86
28.5 – 32.5 60
32.5 – 36.5 2
36.5 – 40.5 1

Mean = ___________ Median = ___________ Mode = ___________

MIDRANGE
This is a rough estimate of the middle. It is found by getting the average of the lowest and
highest value of the data.
𝑙𝑜𝑤𝑒𝑠𝑡 𝑣𝑎𝑙𝑢𝑒+ℎ𝑖𝑔ℎ𝑒𝑠𝑡 𝑣𝑎𝑙𝑢𝑒
𝑀𝑖𝑑𝑟𝑎𝑛𝑔𝑒 (𝑀𝑅) =
2
Example 12 The calories per serving of 15 fruit juices are as follows:
80, 76, 120, 110, 105, 35, 150, 135, 85, 65,
100, 120, 145, 138, 130

39
Determine the midrange.
Solution The midrange is the sum of the lowest value, 35 and the highest value,
150. Then these are divided by 2.
𝐿𝑉+𝐻𝑉 35+150
𝑀𝑅 = 2
= 2
= 92.5

WEIGHTED MEAN
This is used to find the mean of values of the data set that are not equally represented.
The weighted average can be found by multiplying the value by its corresponding weight and
dividing the sum of the products by the sum of their weights.
Example 13 A recent survey of a new ice cream reported the following percentages of
people who liked the flavor. Find the weighted mean of the percentages.
Area % favored Number surveyed
1 55 1 100
2 25 700
3 70 1 000

∑𝑤𝑥 (0.55)(1100)+(0.25)(700)+(0.70)(1 000)


Solution 𝑥̅ = =
∑𝑤 0.55+0.25+0.70
605+175+700 1 480
= = 1.5 = 986.67 ≈ 987
1.5

QUARTILE, DECILE AND PERCENTILE


It is quite helpful when we need to group subjects into several equal groups when
analyzing dataset. Like when we need to create four groups and it needs to be equal. Quartiles
also have three parts and the middle is also called median. The term that is usually used for the
cut off points is Quantiles. Deciles are the one who split the data into 10 while Percentiles split
the data into 100 parts.
Main Formula:
𝑘𝑁
−𝑐𝑓
Quartiles: 𝑄𝑘 = 𝐿𝐵 + ( 4 ) (𝑖)
𝑓

𝑘𝑁
−𝑐𝑓
Deciles: 𝐷𝑘 = 𝐿𝐵 + ( 10 𝑓 ) (𝑖)

𝑘𝑁
−𝑐𝑓
Percentile: 𝑃𝑘 = 𝐿𝐵 + (100𝑓 ) (𝑖)

Whereas:
𝑄𝑘 = Quartile
𝐷𝑘 = Decile
𝑃𝑘 = Percentile
N = Population

40
k = Quartile location
LB = Lower Boundary
f = frequency of the quartile class
cf = cumulative frequency before the quartile class
i = class interval

Quartile Example (Grouped Data)


We still are going to use the example earlier
Class limits Frequency (f) Cumulative frequency (cf)

18-26 3 3
27-35 5 8
36-44 9 17
45-53 14 31
54-62 11 42
63-71 6 48
72-80 2 50

Solution:
𝑁 50
𝑄1 = = = 12.5
4 4

36-44 9 17
𝑄1 class falls under this part of the table (highlighted as green)
𝐿𝐵 = 36 – 0.5 = 35.5
cf = 8 just above 𝑄1 ’s cf. (highlighted as light blue)
f=9
i = still 9
𝑘𝑁 50
−𝑐𝑓 −8
𝐿𝐵 + ( 4 ) (𝑖) Substitute 35.5 + ( 4 ) (9)=40
𝑓 9

Hence, the 𝑄1 is 40, the 𝑄1 will fall within the class boundary of 𝑄1 class.

2𝑁 2(50)
Now for 𝑄2 = = = 25
4 4

41
45-53 14 31
𝑄2 class falls under this part of the table (highlighted as gray)

2𝑁 2(50)
−𝑐𝑓 −17
𝐿𝐵 + ( 4 𝑓 ) (𝑖) Substitute 44.5 + ( 4
) (9) = 49.64
14

3𝑁 3(50)
Lastly for 𝑄3 = = = 37.5
4 4

54-62 11 42
𝑄3 class falls under this part of the table (highlighted as red)
3𝑁 3(50)
−𝑐𝑓 −31
4 4
𝐿𝐵 + ( ) (𝑖) Substitute 53.5 + ( ) (9) = 58.82
𝑓 11

Decile Example:
Class limits Frequency (f) Cumulative frequency (cf)

18-26 3 3
27-35 5 8
36-44 9 17
45-53 14 31
54-62 11 42
63-71 6 48
72-80 2 50
We are going to give you an example
𝐷7 (How did I get 𝐷7? This is usually given)
7𝑁 7(50)
𝐷7 = = = 35
10 10

𝐷7 falls under this part of the table (Highlighted as green)


𝑘𝑁 7(50)
−𝑐𝑓 −31
𝐿𝐵 + ( 10 𝑓 ) (𝑖) Substitute 53.5 + ( 10
11
) (9) = 56.77

Percentile Example:
Class limits Frequency (f) Cumulative frequency (cf)

18-26 3 3
27-35 5 8
36-44 9 17

42
45-53 14 31
54-62 11 42
63-71 6 48
72-80 2 50

We are going to give you again an example


𝑃22
22𝑁 22(50)
𝑃22 = = =1
100 100

36-44 9 17
𝑃22 falls under this part of the table (Highlighted as green)
𝑘𝑁 22(50)
−𝑐𝑓 −8
𝐿𝐵 + (100𝑓 ) (𝑖) Substitute 35.5 + ( 100
9
) (9) =38.5

Let’s try this!


In questions 1 to 6, find the (a) mean, (b) median, (c) mode and (d) midrange of the following.
1. Eight students have the following have scores of 15, 10, 9, 7, 17, 14, 16, and 5.

2. A biologist is studying the gestation period (duration of pregnancy) of domestic dogs.


20 dogs are observed during pregnancy and are found to have the following gestation
periods in days:
59.7 60.5 60.8 61 61.1
59.5 61.3 59.8 60.1 60.6
60.9 61.8 59.6 61.3 61.3

3. A Shampoo manufacturer produces a bottle with an advertised content of 300 ml. A


sample of 15 bottles yielded the following contents:
290 310 300 305 295
300 301 280 285 285
287 298 300 298 298

4. Fifteen motors were tested and the following data were obtained for the number of
revolutions per minute the flywheel attached to the motors turned.
216 230 245 225 250
255 245 230 235 245
260 245 245 225 230

5. In a Statistics class, 8 test scores were randomly selected, and the following results
were obtained:
85 75 80 81 89 77 77 82

6. A special aptitude test is given to job applicants. The data shown below represent the

43
scores of 25 applicants.
250 260 285 277 251
265 290 295 290 259
280 275 291 258 277
287 292 278 280 300
298 279 258 261 270

For questions 7 to 12, use the frequency table.


(a) Identify the class mark for each class interval.
(b) Find the mean, median and mode.
7. Ages of randomly selected residents of Manggahan Pasig.

Ages (years) Frequency 𝑥𝑚 𝑓 ∙ 𝑥𝑚 cf


0–7 65
7 – 12 45
12 – 17 39
17 – 22 33
22 – 27 32
27 – 32 29
32 - 37 27
Mean Median Mode
________________ ________________ ________________

Class limits f X fx cf
1-3 5
4-6 3
7-9 6
10 - 12 5
13 - 15 2
16 - 18 4

𝑄1 =
𝑄2 =
𝑄3 =

44
LESSON 4

Measures of Dispersion

• Compute the different measures of variability for both grouped and ungrouped data.
• Discuss the uses, characteristics, advantages and disadvantages of measures of
spread.
• Compute and interpret the coefficient of variation, skewness and kurtosis.
• Differentiate skewness from kurtosis.
• Determine if a data contains outliers.

In summarizing, a given set of data, sometimes, the measures of central tendency alone
are not sufficient to give useful information. They have to be supplemented by other measures
of description, such as measures of variability which indicate the extent to which values in a
distribution are spread around the central tendency.
In this topic, we shall study the measures of variability for both ungrouped and grouped
data. Three measures of variation, namely, the range, variance, and the standard deviation.
These measures describe how item values cluster or scatter in a distribution.
Example I. Scores of some of the ADPR 3-1D Students in their Accounting Subject
Student A B C D E F G H I J

Scores 19 12 15 10 30 25 17 16 23 17

a.) RANGE
The simplest measure of dispersion is the range. The Range is the difference between
the highest and the lowest values in the set of data. Using the table I, let’s have an example
For example, the range of the set of scores are

45
19, 12, 15, 10, 30, 25, 17, 16, 23, and 17
Formula:
Range = Highest Score – Lowest Score
= 30 - 10
R = 20
It is computed from the lowest and highest score, thus it is a very rough measure of
spread. The range provides useful but limited information, since the range depends only on the
extreme scores.
Example II. Scores of the each contestants who joined the quiz bee
Contestant A B C D E F G H I J

Scores 25 32 30 35 50 44 20 46 27 38

Range = Highest Score – Lowest Score


= 50 – 20
R = 30

Example III. Number of boxes filled by each worker


Worker A B C D E F G H I J

Boxes 21 20 15 30 17 18 15 19 26 25

Range = Highest Score – Lowest Score


= 30 - 15
R = 15

VARIANCE
The measure of dispersion that removes negative signs by squaring all deviations of each
number from its mean and getting the average of the squared deviations is the variance.
The following are the steps to compute the variance:
1. Determine the mean for the given set of data.
2. Determine the deviation from the mean for each value in the data.
3. Square each deviation.
4. Compute the mean of the squared deviations.

46
∑𝑛
𝑖=𝑙(𝑥1 − 𝑥̅ )
2
Variance =
𝑛

Example 1
The variance of the scores of the 6 female Math students is computed and
shown below.
77 83 90 78 80 87
Solution
Using the formula
(77−82.5)2 +(83−82.5)2 +(90−82.5)2 +(78−82.5)2 +(80−82.5)2 +(87−82.5)2
=
6

133.5
= 6 = 22.25
Example 2
Solve the variance using the following data.
10 11 15 13 13
Solution
Using the formula
(10−12.4)2 + (11−12.4)2 +(15−12.4)2 +(13−12.4)2 +(13−12.4)2
=
5
15.2
= 5 = 3.04
In summation notation, the variance in a population is given by

∑𝑛 (𝑥 − 𝑚)2
𝝈𝟐 = 𝑖=𝑙 𝑛1
Where m is the population mean and n is the number of cases.
The formula for the variance in a sample is given by the statistic
∑𝑛
𝑖=𝑙(𝑥1 − 𝑥̅ )
2
𝑠2 =
𝑛−1

where 𝑥̅ is the mean and n-1 is the number of cases which gives an unbiased estimate
of variance.
Calculating the variance is an important part of many statistical applications and
analyses. It is the first step in calculating the standard deviation.

STANDARD DEVIATION
The standard deviation is the most commonly used measure of variation. The standard
deviation indicates how closely the values of a given data set are clustered around the mean. A
lower value of the standard deviation means that the values of the given data set are spread over
a smaller range around the mean. On the other hand, a large value of the standard deviation
means that values of that data set are spread over a large range around the mean.
In summation notation, the formula for the standard deviation in a sample is given by the
parameter

∑(𝑥− 𝜇)2
𝜎= √
𝑁

where:

47
N = number of population 𝜇 = 𝑝𝑜𝑝𝑢𝑙𝑎𝑡𝑖𝑜𝑛 𝑚𝑒𝑎𝑛
Example 3
Using the given in example 1, the standard deviation of the scores of the 6 female
Math students is computed and shown below.
77 83 90 78 80 87
Solution
Using the formula
(77−82.5)2 +(83−82.5)2 +(90−82.5)2 +(78−82.5)2 +(80−82.5)2 +(87−82.5)2
=√
6

133.5
=√ 6
= √22.25 = 4.72
Example 4
The ages of the family recorded living in a house are as follows (in years).
66 31 29 27 57
Find the variance and standard deviation
Solution
1. Find the mean.
66 + 31 + 29 + 27 + 57 210
= = 42
5 5
2. Find the variance.
(66−42)2 + (31−42)2 + (29−42)2 + (27−42)2 + (57−42)2
5
1 316
=263.2
5

3. The standard deviation is the square root of the variance.

√263.2 = 16.22

VARIANCE AND STANDARD DEVIATION FOR GROUPED DATA


The procedure is similar to that of finding the mean for grouped data, and it uses the
midpoints of each class.
𝑛∑𝑓 ∙ 𝑥 2 −(∑𝑓∙𝑥)2
𝑠2 =
𝑛(𝑛−1)

Example 5
For 120 randomly selected nursing students, the following IQ frequency distribution
were obtained.
Class Limits Frequency

48
90 – 95 4
95 – 100 8
100 – 105 7
105 – 110 14
110 - 115 16
115 – 120 24
120 – 125 20
125 - 130 17
130 – 135 10
Find the variance and standard deviation
Solution
Step 1 Make a table. Find the midpoints of each class. Multiply the
midpoints by the frequency for each class.

Class Limits Frequency 𝑥𝑚 𝑓 ∙ 𝑥𝑚


90 – 95 4 92.5 370
95 – 100 8 97.5 780
100 – 105 7 102.5 717.5
105 – 110 14 107.5 1505
110 – 115 16 112.5 1800
115 – 120 24 117.5 2820
120 – 125 20 122.5 2450
125 – 130 17 127.5 2167.5
130 - 135 10 132.5 1325

Step 2 Multiply the frequency by the square of the midpoint for each
class.
1 2 3 4 5
Class Limits Frequency 𝑥𝑚 𝑓 ∙ 𝑥𝑚 𝑓 ∙ 𝑥𝑚 2
90 – 95 4 92.5 370 34 225
95 – 100 8 97.5 780 76 050

49
100 – 105 7 102.5 717.5 73 543.75
105 – 110 14 107.5 1505 161 787.5
110 – 115 16 112.5 1800 202 500
115 – 120 24 117.5 2820 331 350
120 – 125 20 122.5 2450 253 125
125 – 130 17 127.5 2167.5 276 356.25
130 - 135 10 132.5 1325 175 562.5

Step 3 Find the columns 2, 4 and 5. Substitute in the formula for 𝑠 2


1 2 3 4 5
Class Limits Frequency 𝑥𝑚 𝑓 ∙ 𝑥𝑚 𝑓 ∙ 𝑥𝑚 2
90 – 95 4 92.5 370 34 225
95 – 100 8 97.5 780 76 050
100 – 105 7 102.5 717.5 73 543.75
105 – 110 14 107.5 1505 161 787.5
110 – 115 16 112.5 1800 202 500
115 – 120 24 117.5 2820 331 350
120 – 125 20 122.5 2450 253 125
125 – 130 17 127.5 2167.5 276 356.25
130 - 135 10 132.5 1325 175 562.5
Total n = 120 ∑ 𝑓 ∙ 𝑥𝑚 = 13 935 ∑𝑓 ∙ 𝑥𝑚 2 = 1584500
𝑛∑𝑓 ∙ 𝑥 2 −(∑𝑓∙𝑥)2 (120)(1584500)−(13935)2
𝑠2 = 𝑛(𝑛−1)
= (120)(120−1)
= 13 315.13

𝑠 = √𝑠 2 = 115.39

Let’s try this!


I. Identification

___________1. It indicates the extent to which values in a distribution are spread around the
central tendency.

50
___________2. The study of the collection, analysis, interpretation, presentation and
organization of data. It deals with all aspects of data including the planning of data collection in
terms of the design of surveys and experiments.

___________3. The simplest measure of dispersion.

___________4. The positive square root of the variance measures the spread or dispersion
of each value from the mean of the distribution.

___________5. The degree of symmetry, or departures from symmetry of a set of data.

___________6. This distribution contains high scores that have low frequency, but does not
contain extreme low scores with corresponding low frequency.

___________7. It allows the variability of scores in two sets of data that do not necessarily
measure the same thing.

___________8. It is the difference between the first quartile (Q₁) and the third quartile (Q₃)
___________9. The mean, median and the mode are equal.

___________10. This distribution contains extreme low scores that have low frequency, but
does not contain extreme high scores with corresponding low frequency.

51
LESSON 5

Normal Distribution

• Know what Normal Distribution is


• Use tables of the normal distribution to solve problems.
• Use the normal distribution as an approximation to other distributions in appropriate
circumstances.
• Appreciate wide variety of circumstances in which normal distribution can be used.
• Know more about Normal Distribution and the topics related thereto
• Find the area under the standard normal distribution, given various z values.
• Find the z value under the standard normal distribution, given the area.
• Find the probabilities for a normally distributed variable by transforming it into a standard
normal variable.
• Use central limit theorem to solve problems involving sample means for large samples.
• Use the normal approximation to compute probabilities for a binomial variable.

Normal distribution

In probability theory, the normal (or Gaussian) distribution is a very commonly


occurring continuous probability distribution—a function that tells the probability that any real
observation will fall between any two real limits or real numbers, as the curve approaches zero
on either side. Normal distributions are extremely important in statistics and are often used in
the natural and social sciences for real-valued random variables whose distributions are not
known.

The normal distribution is immensely useful because of the central limit theorem, which
states that, under mild conditions, the mean of many random variables independently drawn from
the same distribution is distributed approximately normally, irrespective of the form of the original
distribution: physical quantities that are expected to be the sum of many independent processes
(such as measurement errors) often have a distribution very close to the normal. Moreover, many
results and methods (such as propagation of uncertainty and least squares parameter fitting) can
be derived analytically in explicit form when the relevant variables are normally distributed.

The Gaussian distribution is sometimes informally called the bell curve. However, many
other distributions are bell-shaped (such as Cauchy’s, Student's, and logistic). The

52
terms Gaussian function and Gaussian bell curve are also ambiguous because they
sometimes refer to multiples of the normal distribution that cannot be directly interpreted in terms
of probabilities.

A normal distribution is:

The parameter μ in this definition is the mean or expectation of the distribution (and also
its median and mode). The parameter σ is its standard deviation; its variance is therefore σ 2. A
random variable with a Gaussian distribution is said to be normally distributed and is called
a normal deviate.
If μ = 0 and σ = 1, the distribution is called the standard normal distribution or the unit normal
distribution, and a random variable with that distribution is a standard normal deviate.

The normal distribution is the only absolutely continuous distribution all of


whose cumulants beyond the first two (i.e., other than the mean and variance) are zero. It is also
the continuous distribution with the maximum entropy for a given mean and variance.

The normal distribution is a subclass of the elliptical distributions. The normal distribution
is symmetric about its mean, and is non-zero over the entire real line. As such it may not be a
suitable model for variables that are inherently positive or strongly skewed, such as the weight of
a person or the price of a share. Such variables may be better described by other distributions,
such as the log-normal distribution or the Pareto distribution.

The value of the normal distribution is practically zero when the value x lies more than a
few standard deviations away from the mean. Therefore, it may not be an appropriate model when
one expects a significant fraction of outliers—values that lie many standard deviations away from
the mean — and least squares and other statistical inference methods that are optimal for
normally distributed variables often become highly unreliable when applied to such data. In those
cases, a more heavy-tailed distribution should be assumed and the appropriate robust statistical
inference methods applied.

The Gaussian distribution belongs to the family of stable distributions which are the
attractors of sums of independent, identically distributed distributions whether or not the mean or
variance is finite. Except for the Gaussian which is a limiting case, all stable distributions have
heavy tails and infinite variance

One of the most important theoretical distribution in statistics is the normal distribution.
The graph of the normal distribution is a bell-shaped curve called the normal curve as shown in
the figure below. The normal distribution is a tool that can be used to predict the probability of
natural occurrences of things in many fields such as in natural and social sciences as many real
occurrences behaves like the characteristics of the normal curve.

53
PROPERTIES OF THE NORMAL CURVE:
1.) The normal curve is bell-shaped and extends infinitely in both directions.
2.) The tails in the left and the right becomes closer and closer but never interact with the
horizontal axis. In other words, the tails are asymptotic with respect to the horizontal axis.
3.) The mean, the mode and the median are located in the line that divides the normal curve
into two equal parts. This line may be called the line of symmetry because the curve in the
left side of this line is symmetrical to the right side.
4.) The total area under the normal curve is 1.0 or 100%. Since the curve is symmetrical, the
area to the left and to the right of the line of symmetry under the curve are the same and
are equal to 0.5.
5.) From the center of the distribution, the horizontal axis may be divided into at least three
standard scores(z scores) to the left and to the right.
6.) Let a and b be two points along the horizontal axis. The probability between these two
points a and b in the normal distribution is equal to the area between the two points a and
b under the normal curve.
7.) The distribution of the normal curve is unimodal, meaning, there is only one value that
dominates the distribution of values.
THE STANDARD NORMAL DISTRIBUTION
A special feature of the normal distribution is that it is completely determined by its mean,
and standard deviation is given. A normal distribution can be converted into standard normal
distribution by transforming x-scale into z-scale using the formula,
𝑥−𝜇 𝑥−𝑥
z= or 𝑧=
𝜎 𝑆

Where:
z = z value or z score or standard score
𝑥 = given value of a particular variable
µ = population mean
σ = population standard deviation
𝑥 = sample mean

54
S = sample standard deviation
Z-scale normally has a mean µ = 0 (located at the center of the distribution) and variance
equal to 1. It should be noted that the z value can also be converted into x value by transforming
the formula given above into
x = µ + zσ or 𝑥 = 𝑥 + 𝑧𝑆

Examples:
A. Given the following data:
1.) z = 1.26, 𝑥 = 54 and S = 3, find the value of x
Solution:
x = 54+1.26(3)
x = 57.78

2.) x = 80, 𝑥 = 85 and S = 5, find the value of z


Solution:
80−85
𝑧=
5

Z = -1.0
3.) x = 90, S = 4 and z = 1.75, find the value of 𝑥
Solution:
𝑥 = x – Sz
𝑥 = 90-4(1.75)
𝑥 = 83
B. The mean and standard deviation of a Quiz in Statistics are 85 and 4 respectively. (1)
Determine the z score of the student having a grade of 93, (2)find the grade of a student
corresponding to a standard score of 1.5.

Solutions:
93−85
(1) 𝑧= (2) 𝑥 = 𝜇 + 𝑧𝜎
4
93−85
𝑧= 𝑥 = 85 + 1.5(4)
4

𝑧 = 2.0 𝑥 = 91

55
AREAS UNDER THE NORMAL CURVE

An Important property of the normal distribution is that the area under the normal curve is
1.0 or 100%. The area under the normal curve bounded by two points a and b represents the
probability that x lies between a and b. We simply find areas under the normal curve by using the
z distribution table/standard normal distribution (Appendix A)
Example 1.
Find the area under the normal curve between z = 0 and z = 1.95.

Solution:
Using the z distribution table, the value at the row containing z = 1.9 and the column
containing z =0.05 is 0.4744. Therefore, the area under the normal curve between z = 0 and z =
1.95 is 0.4744.

Example 2.

Find the area under the normal curve between z = 0 and z = -2.37.

Solution:
Since the normal curve is symmetrical, the area under the negative values of z is equal to
the area under the positive values. Again, by using the z distribution table, the value at z = 2.37
is 0.4991. Therefore, the area under the normal curve between z = 0 and z = -2.37 is 0.4911.
Example 3.

56
Find the area under the normal curve between z = - 0.86 and z = -1.53.

Solution:
Using the z distribution table and by property of symmetry, the area from z = 0 to z = -1.53
is equal to 0.4370, and the area from z = 0 to z = - 0.86 is equal to 0.0351. By subtracting the
areas obtained, the area between z = - 0.86 and z = -1.53 is equal to 0.1319.
EXAMPLES:
1. Suppose that the passing grade for a statistics class containing 50 students is 75. If the
mean grade is 85, the standard deviation is 5 and the grades are normally distributed,
(a)what proportion of class passed the subject?, (b)what proportion of the students
obtained the grades between 82 and 88?
Solution:
(a) Convert the grade of 75 into z score.
75−85
𝑧= = −2
5

The proportion of the students who passed the subject is the same as the shaded area
under the normal curve.
Shaded area = (Area between z = 0 and z = - 2) + (Area between z = 0 and
z = + ∞)

Using the z distribution table,

57
Shaded area = 0.4772 + 0.5 = 0.9772
Thus, the students who passed the subject is 97.72%
Or
0.9772 x 50 students = 48.86 ≈ 49 students

(b) First, convert the grades 82 89 into z value.


82 − 85
𝑧1 = = −0.6
5
89 − 85
𝑧2 = = 0.8
5

Shaded Area = (Area between z = 0 and z = - 0.6) + (Area between z = 0 and z = 0.8)
= 0.2257 + 0.2881
= 0.5138
Thus, the proportion of the students obtaining grades between 82 and 89 is 51.38%
Or
0.5138 x 50 students = 25.69 ≈ 26 students

2. The average service life in a batch of manufactured compact fluorescent lamps(CFL)


containing 5,000 CFLs is 10,000 hours. If the standard deviation is 500 hours and the
service life is normally distributed, find the probability that the service life of a bulb taken
from this batch will be in the range between 9,000 and 10, 500 hours.
Solution:
Convert 9,000 and 10,500 into z scores.

58
9,000−10,000 10,500−10,000
𝑧1 = 500
= −2, 𝑧2 = 500
=1

From the graph,


P(- 2.0 < z < 1.0) = P(- 2.0 < z < 0) + P(0 < z < 1.0)
= 0.4772 + 0.3413
= 0.8185
3. The average weight of freshmen students in a university is 52 kilograms. If the standard
deviation is 4 kilograms, find the probability that a student weighs less than 45 kilograms.
Solution:
Convert 45 into z score.

45−52
𝑧= = −1.7
4

Then,
P(z< -1.75) = P(- ∞< z < 0) – P(-1.75< z < 0)
= 0.5 – 0.4599
= 0.0401

Let’s try this!

1) Find the area under the normal curve:


a. between z = 2.57 and z = 0
b. between z = - 1.05 and z = 2.43

59
c. between z = 3.0 and z = - 3.0
d. between z = 1.29 and z = 2.88
e. to the right of z = 1.56
f. to the left of z = 2.15

2) Suppose that the distribution of the result of an examination in mathematics class


containing 45 students is normal. The mean score is 80 and the standard deviation is 3. If
the passing score is 75, what proportion of the students obtained failing score?

3) A normal distribution has a mean of 45 and a standard deviation of 5. Find the z scores of
the following:
a. 58
b. 35
c. 52
d. 44

4) Find the limits that include the middle 30% of a normal distribution which has a mean of
93 and a standard deviation of 4.

5) Suppose that the weights per bag of cement produced by a company follow a normal
distribution. The mean weight is 40 kilograms and the standard deviation is 0.8 kilogram.
What would be the new standard deviation if the company wants to have only 1% of the
cement produced weighing less than 39 kilograms per bag keeping the mean of 40
kilograms?

6) The monthly salaries of employees in a manufacturing company are normally distributed


with mean of P20,000 and standard deviation of P3,000.
a. What percent of the employees earn less than P15,000?
b. What percent of the employees earn more than P26,000?
c. What percent of the employees earn between P18,000 and P23,000?

SKEWNESS
Skewness is a measure that indicates the general shape of the distribution. It is a tool used to
analyze where most of the values are concentrated and to what extent the distribution departs
from the mean.
TYPES OF SKEWNESS:
1. Symmetrical Distribution or Zero Distribution
In this type of distribution, the mean, the mode and the median are equal and are
located at the center of the distribution. The graph of the distribution is symmetrical best
described as the normal curve.

60
2. Right-skewed Distribution or Positively Skewed Distribution
For this type of distribution, the mode is always less than the mean and the median.
Most of the time, the median is less than the mean and the longer tail of the curve is to
the right of the mode.

3. Left-skewed or Negatively Skewed Distribution

For this type of distribution, the mean and the median are always less than the
mode. Most of the time, the mean is less than the median and the longer tail of the curve
is to the left of the mode.

61
COEFFICIENT OF SKEWNESS
There are several formulas in solving the measure of skewness. Karl Pearson developed
the following two formulas:
𝑥̅ −𝑀𝑜
1) 𝑠𝑘1 =
𝑠

and
3(𝑥̅ − 𝑀𝑑 )
2) 𝑠𝑘2 =
𝑠

Where:
sk = coefficient of skewness
𝑥̅ = mean
Mo = mode
Md = median
S = sample standard deviation

To determine the degree of Pearson’s Coefficient of Skewness, use the following


interpretations:
sk = 0 Symmetrical distribution
sk > 0 Positively skewed distribution
sk < 0 Negatively skewed distribution

Examples:
1. Below are scores of ten students in a short quiz in Biology
8 8 8 8 8 8 8 8 8 8
Since the mean and median are the same which is 8, the distribution of scores is
symmetric because sk = 0.
2. The mean, median and standard deviation of a certain set of scores are 26, 23 and 3
respectively. Using the formula, the coefficient of skewness is
3(26−23)
𝑠𝑘 =
3

𝑠𝑘 = 3

62
Therefore, the set of scores is skewed to the right.

3. Given the data set: 2, 5, 6, 8, 9, 9, 10. Solve for the skewness using Pearson’s coefficient of
skewness.
Solution:
2+5+6+8+9+9+10
𝑥̅ = =7
7

Mo = 9
Md = 8

∑(𝑥−𝑥̅ )2
𝑠=√
𝑛−1

(2−7)2 +(5−7)2 +(6−7)2 +(8−7)2 +(9−7)2 +(9−7)2 +(10−7)2


𝑠=√
7−1

(−5)2 +(−2)2 +(−1)2 +(1)2 +(2)2 +(2)2 +(3)2


𝑠= √
6

𝑠 = √8
𝑠 = 2.8284
Using formula 1,
7−9
𝑠𝑘1 = 2.8284 = − 0.7071

Or using formula 2,
3(7− 8 )
𝑠𝑘2 = = −1.0607
2.8284

The two formulas obtained the negative value of skewness. Notice that the result of formula 2 is
more skewed to the left and is more advantageous to use considering the given data set wherein
the value for the mode occurs only twice.

COEFFICIENT OF SKEWNESS BASED ON QUARTILES (BOWLEY’S FORMULA)


If the distribution is normal or symmetrical, the distance between Q1 and Q2(or the median)
must be the same as the distance between Q2 and Q3. For a positively skewed distribution, the
median is closer to Q1 than Q3. On the other hand, the median is closer to Q3 than Q1 for a
negatively skewed distribution. Quartiles can also be used to compute the coefficient of skewness
as shown in the formula below. This formula is known as the Bowley’s Coefficient of skewness
formula.

63
(𝑄3 − 𝑄2 ) − (𝑄2 − 𝑄1 )
𝑠𝑘3 =
𝑄3 − 𝑄1
or

𝑄3 − 2𝑄2 + 𝑄1
𝑠𝑘3 =
𝑄3 − 𝑄1

where:

𝑄1 = Quartile 1

𝑄2 = Quartile 2 or the median

𝑄3 = Quartile 3

KURTOSIS

Another statistical measure that describes the distribution of data is kurtosis. Kurtosis
measures the peakedness or flatness of the distribution relative to the normal distribution.

TYPES OF KURTOSIS:
1. Mesokurtic
A distribution that is neither too flat nor too peaked is known as mesokurtic
distribution. The graph of this type of distribution is the same as in any normal distribution
and is considered to be the baseline for the two other types.

2. Platykurtic

64
A distribution that has a lower peak compared to the mesokurtic distribution is
known as platykurtic distribution. It is characterized by evenly spread data about the center
with a certain degree of flatness to the peak.

3. Leptokurtic

A leptokurtic distribution has a higher peak compared to mesokurtic distribution.


The distribution of data is heavily concentrated or pile up at the center resulting to a peak
that is tall and thin.

PERCENTILE COEFFICIENT OF KURTOSIS

A simple formula for measuring kurtosis can be based on both quartiles and percentiles
and is given by the formula

1
(𝑄3 − 𝑄1 )
𝐾 =2
𝑃90 − 𝑃10

where:

𝑃10 = 10th Percentile

𝑃90 = 90th Percentile

𝑄1 = Quartile 1

65
𝑄3 = Quartile 3

A standard value of K for a normal distribution is 0.263 or 26.3%. This value serves as the
basis to determine the type of kurtosis for a given distribution. For mesokurtic distribution, K is
equal to 0.263. Platykurtic distribution obtain K is less than 0.263 and leptokurtic distribution
obtain K greater than 0.263.

Kurtosis can also be obtained by using the formula

𝑠𝑢𝑚 𝑜𝑓(𝐷𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛 𝑓𝑟𝑜𝑚 𝑡ℎ𝑒 𝑚𝑒𝑎𝑛)4


𝑘=
(𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑐𝑎𝑠𝑒𝑠)(𝑣𝑎𝑟𝑖𝑎𝑛𝑐𝑒)2

Where:

k = kurtosis

• The distribution is mesokurtic if k = 3


• The distribution is platykurtic if k < 3
• The distribution is leptokurtic if k > 3

Example:

Thirty BSA students took an examination in English 1. The variance of the exam is 5 and
the sum of the fourth power of the deviation from the mean is 4 500. Is the distribution of scores
of the students mesokurtic, platykurtic or leptokurtic?

4 500
𝑘=
30(5)2

𝑘=6

Since the kurtosis is greater than 3, then the distribution is leptokurtic.

Let’s try this!


A. Determine if each of the following sets of data is symmetric, negatively or positively skewed or
leptokurtic or platykurtic

1. The scores of students in an easy examination.

2. Scores of bright students in an average quiz.

3. Time to finish a difficult examination.

66
B. Determine the coefficient of skewness and kurtosis of the following data.

1. 7 7 7 7 7 7

2. 6 8 9 9 10 12

3.

Statistics Quiz Scores Frequency

(Class Interval)
40-44 3

35-39 5

30-34 11

25-29 13

20-24 6

15-19 2

67
LESSON 6: INTRODUCTION TO PROBABILITY
LEARNING OBJECTIVES

• Discuss random experiments, sample space and events.


• Know operations on events.
• Define mutually exclusive events.
• State De Morgan’s Laws.

DISCUSSIONS

68
69
70
71
72
LET’S PRACTICE!

73
REFERENCE:
Introduction to Probability and Probability Distributions(2015),by J.B. Ofosu and C.A. Hesse,
pages 1 − 8.

74
LESSON 7: COUNTING SAMPLE POINTS
LEARNING OBJECTIVES

• Manage different counting principles.


• Apply permutation and combination rules.
• Know permutations with repetitions.
• Discuss circular permutations.

DISCUSSIONS

75
76
77
LET’S PRACTICE!

78
REFERENCE:
Introduction to Probability and Probability Distributions(2015),by J.B. Ofosu and C.A. Hesse,
pages 8 − 13.

79
LESSON 8: PROBABILITY OF AN EVENT
LEARNING OBJECTIVES

• Explain the classical probability.


• Discuss relative frequency probability.
• Discuss axioms of probability.
• Apply addition rule of probability.

DISCUSSIONS

80
81
82
83
84
85
86
LET’S PRACTICE!

87
REFERENCE:
Introduction to Probability and Probability Distributions(2015),by J.B. Ofosu and C.A. Hesse,
pages 13 − 22.

88
LESSON 9: CONDITIONAL PROBABILITY
LEARNING OBJECTIVES

• Discuss conditional probability.

• Apply multiplication rule.

• Explain independent events.

DISCUSSIONS

89
90
91
92
93
94
LET’S PRACTICE!

95
REFERENCE:
Introduction to Probability and Probability Distributions(2015),by J.B. Ofosu and C.A. Hesse,
pages 22 − 31.

96
LESSON 10: BAYE’S THEOREM
LEARNING OBJECTIVES

• Explain the total probability rule.

• Discuss Baye’s Theorem.

• Apply Baye’s Theorem.

DISCUSSIONS

97
98
LET’S PRACTICE!

99
REFERENCE:
Introduction to Probability and Probability Distributions(2015),by J.B. Ofosu and C.A. Hesse,
pages 31 − 35.

100
LESSON 11

Hypothesis Testing

• Understand the terms used in hypothesis testing


• Differentiate the methods of the hypothesis testing, traditional method, p-value and
confidence interval.
• State the null and alternative hypotheses.
• Compare and contrast one-tailed and two-tailed test.
• Find the critical values for the z and t test.
• State the steps in hypothesis testing.
• Test means for large samples by means of z test and for small samples by means of t
test.
• Test proportions using the z test.
• Test the hypothesis using p-value and confidence interval.
• Test the difference between two large samples, two means and two proportions.

A statistical hypothesis is an assumption about a


population parameter. This assumption may or may not be true.

The best way to determine whether a statistical hypothesis is true would be to examine the
entire population. Since that is often impractical, researchers typically examine a random sample
from the population. If sample data are not consistent with the statistical hypothesis, the
hypothesis is rejected.
There are two types of statistical hypotheses.
▪ Null hypothesis. The null hypothesis, denoted by H0, is usually the hypothesis that
sample observations result purely from chance.

101
▪ Alternative hypothesis. The alternative hypothesis, denoted by H1 or Ha, is the
hypothesis that sample observations are influenced by some non-random cause.

For example, suppose we wanted to determine whether a coin was fair and balanced. A null
hypothesis might be that half the flips would result in Heads and half, in Tails. The alternative
hypothesis might be that the number of Heads and Tails would be very different. Symbolically,
these hypotheses would be expressed as
H0: P = 0.5
Ha: P ≠ 0.5
Suppose we flipped the coin 50 times, resulting in 40 Heads and 10 Tails. Given this result,
we would be inclined to reject the null hypothesis. We would conclude, based on the evidence,
that the coin was probably not fair and balanced.

Hypothesis testing refers to the formal procedures used by statisticians to


accept or reject statistical hypotheses

Goal of Hypothesis Testing


The goal of hypothesis testing is not to question the computed value of the sample
statistic but to make a judgement about the difference between the sample statistics and
hypothesized population parameter.
Use of hypothesis Testing
Hypothesis testing enables a researcher to generalize population from relatively
small samples. In many instances, a researcher can only rely on the information provided
by a part of population.
Statisticians follow a formal process to determine whether to reject a null hypothesis, based on
sample data. This process, called hypothesis testing, consists of four steps.
▪ State the hypotheses. This involves stating the null and alternative hypotheses. The
hypotheses are stated in such a way that they are mutually exclusive. That is, if one is
true, the other must be false.

▪ Formulate an analysis plan. The analysis plan describes how to use sample data to
evaluate the null hypothesis. The evaluation often focuses around a single test statistic.

▪ Analyze sample data. Find the value of the test statistic (mean score, proportion, t-
score, z-score, etc.) described in the analysis plan.

▪ Interpret results. Apply the decision rule described in the analysis plan. If the value of
the test statistic is unlikely, based on the null hypothesis, reject the null hypothesis.

Decision Errors

Two types of errors can result from a hypothesis test.

102
▪ Type I error. A Type I error occurs when the researcher rejects a null hypothesis when it
is true. The probability of committing a Type I error is called the significance level. This
probability is also called alpha, and is often denoted by α.

▪ Type II error. A Type II error occurs when the researcher fails to reject a null hypothesis
that is false. The probability of committing a Type II error is called Beta, and is often
denoted by β. The probability of not committing a Type II error is called the Power of the
test.

Decision Rules

The analysis plan includes decision rules for rejecting the null hypothesis. In practice,
statisticians describe these decision rules in two ways - with reference to a P-value or with
reference to a region of acceptance.

▪ P-value. The strength of evidence in support of a null hypothesis is measured by the P-


value. Suppose the test statistic is equal to S. The P-value is the probability of observing
a test statistic as extreme as S, assuming the null hypothesis is true. If the P-value is less
than the significance level, we reject the null hypothesis.
▪ Region of acceptance. The region of acceptance is a range of values. If the test statistic
falls within the region of acceptance, the null hypothesis is not rejected. The region of
acceptance is defined so that the chance of making a Type I error is equal to the
significance level.

The set of values outside the region of acceptance is called the region of rejection. If the
test statistic falls within the region of rejection, the null hypothesis is rejected. In such
cases, we say that the hypothesis has been rejected at the α level of significance.

These approaches are equivalent. Some statistics texts use the P-value approach; others use the
region of acceptance approach. In subsequent lessons, this tutorial will present examples that
illustrate each approach.

One-Tailed and Two-Tailed Tests


A test of a statistical hypothesis, where the region of rejection is on only one side of the sampling
distribution, is called a one-tailed test. For example, suppose the null hypothesis states that the
mean is less than or equal to 10. The alternative hypothesis would be that the mean is greater
than 10. The region of rejection would consist of a range of numbers located on the right side of
sampling distribution; that is, a set of numbers greater than 10.

A test of a statistical hypothesis, where the region of rejection is on both sides of the sampling
distribution, is called a two-tailed test. For example, suppose the null hypothesis states that the
mean is equal to 10. The alternative hypothesis would be that the mean is less than 10 or greater
than 10. The region of rejection would consist of a range of numbers located on both sides of
sampling distribution; that is, the region of rejection would consist partly of numbers that were less
than 10 and partly of numbers that were greater than 10.

One-Tailed

103
Two-Tailed

Power of a Hypothesis Test

104
The probability of not committing a Type II error is called the power of a hypothesis test.

Effect Size

To compute the power of the test, one offers an alternative view about the "true" value of
the population parameter, assuming that the null hypothesis is false. The effect size is the
difference between the true value and the value specified in the null hypothesis.

Effect size = True value - Hypothesized value

For example, suppose the null hypothesis states that a population mean is equal to 100.
A researcher might ask: What is the probability of rejecting the null hypothesis if the true
population mean is equal to 90? In this example, the effect size would be 90 - 100, which equals
-10.

Factors That Affect Power

The power of a hypothesis test is affected by three factors.

▪ Sample size (n). Other things being equal, the greater the sample size, the greater the
power of the test.

▪ Significance level (α). The higher the significance level, the higher the power of the test.
If you increase the significance level, you reduce the region of acceptance. As a result,
you are more likely to reject the null hypothesis. This means you are less likely to accept
the null hypothesis when it is false; i.e., less likely to make a Type II error. Hence, the
power of the test is increased.

▪ The "true" value of the parameter being tested. The greater the difference between the
"true" value of a parameter and the value specified in the null hypothesis, the greater the
power of the test. That is, the greater the effect size, the greater the power of the test.

General procedure that can be used to test statistical hypotheses.

How to Conduct Hypothesis Tests

All hypothesis tests are conducted the same way. The researcher states a hypothesis to be tested,
formulates an analysis plan, analyzes sample data according to the plan, and accepts or rejects
the null hypothesis, based on results of the analysis.

▪ State the hypotheses. Every hypothesis test requires the analyst to state a null
hypothesis and an alternative hypothesis. The hypotheses are stated in such a way that
they are mutually exclusive. That is, if one is true, the other must be false; and vice versa.
▪ Formulate an analysis plan. The analysis plan describes how to use sample data to
accept or reject the null hypothesis. It should specify the following elements.
• Significance level. Often, researchers choose significance levels equal to 0.01,
0.05, or 0.10; but any value between 0 and 1 can be used.

105
• Test method. Typically, the test method involves a test statistic and a sampling
distribution. Computed from sample data, the test statistic might be a mean score,
proportion, difference between means, difference between proportions, z-score, t-
score, chi-square, etc. Given a test statistic and its sampling distribution, a
researcher can assess probabilities associated with the test statistic. If the test
statistic probability is less than the significance level, the null hypothesis is
rejected.
▪ Analyze sample data. Using sample data, perform computations called for in the analysis
plan.
• Test statistic. When the null hypothesis involves a mean or proportion, use either
of the following equations to compute the test statistic.

Test statistic = (Statistic - Parameter) / (Standard deviation of statistic)


Test statistic = (Statistic - Parameter) / (Standard error of statistic)

where Parameter is the value appearing in the null hypothesis, and Statistic is the point
estimate of Parameter. As part of the analysis, you may need to compute the standard deviation
or standard error of the statistic. Previously, we presented common formulas for the standard
deviation and standard error.

When the parameter in the null hypothesis involves categorical data, you may use a chi-square
statistic as the test statistic. Instructions for computing a chi-square test statistic are presented
in the lesson on the chi-square goodness of fit test.

• P-value. The P-value is the probability of observing a sample statistic as extreme


as the test statistic, assuming the null hypothesis is true.
▪ Interpret the results. If the sample findings are unlikely, given the null hypothesis, the
researcher rejects the null hypothesis. Typically, this involves comparing the P-value to
the significance level, and rejecting the null hypothesis when the P-value is less than the
significance level.

Let’s try this!


i. Identification

1. What would be the best way to determine whether a statistical hypothesis is true?
2. What would be the 2 possible equations needed in determining the test statistics if the null
hypothesis involves a mean or a proportion?
3. What are the three factors that affect the hypothesis test?
4. What is the effect size?
5. How does a researcher conduct a hypothesis test?
6. Define: One Tailed test; two tailed test.
7. What are the two types of errors that can result from a hypothesis test?
8. What is the use of hypothesis testing?
9. What are the four steps in hypothesis testing?
10. What is denoted by null hypothesis?

106
LESSON 12

Correlation and Regression Analysis

• Calculate and interpret the coefficient of correlation and the linear regression equation.
• Distinguish and explain the difference between independent and dependent variables.
• Draw a scatter diagram for a set of ordered pairs.
• Conduct test for spearman rank correlation and explain its applications.
• Calculate and interpret the coefficient of determination and the standard error of
estimate.

SIMPLE CORRELATION and REGRESSION ANALYSIS

The word correlation is used in everyday living to denote some form of association. We
might say that we have noticed a correlation between sunny days and shortage of water supply
in Metro Manila. However, in statistical terms we use correlation to denote association between
two quantitative variables. Let us also assume that association is linear, that one variable
increases or decreases a fixed amount for a unit increase or decrease in the other.

The first way to get some idea about a possible relation between two variables is to do a
scatter plot of the variables.

Let us consider the example of possible correlation between height and ratings of physical
attractiveness (1-10) of 10 persons as shown on the table below:

Height Physical Att.


(inch)
56 7
52 5
60 7
61 8

107
63 9
58 7
58 6
57 8
57 7
59 8

A scatter plot of these data is shown below:

10
9
8
7
6
5
4
3
2
1
0
0 10 20 30 40 50 60 70
Physical Att.

Correlation range from -1 (perfect negative relation) through 0 (no relation) to +1 (perfect
positive relation) Scatterplots would look like:
10

0
0 1 2 3 4 5 6 7 8 9 10

Perfect Negative Correlation

108
2.5

1.5

0.5

0
0 1 2 3 4 5 6 7 8 9

No Correlation

12

10

0
0 2 4 Perfect Positive
6 Correlation 8 10 12

THE PEARSON PRODUCT MOMENT CORRELATION COEFFICIENT

The most widely used measure of correlation is the Pearson Product Moment Correlation
Coefficient, commonly called Pearson r (denoted by r). It is a measure of the linear association
between two variables that have been measured on interval or ratio scales, such as the
relationship between height in inches and weight in pounds. The formula for computing the
correlation given two variables (x and y) is given below:

𝑛𝛴𝑥𝑦 − 𝛴𝑥𝛴𝑦
𝑟=
√[𝑛(𝛴𝑥 2 ) − (∑ 𝑥)2 ][𝑛(𝛴𝑦 2 ) − (∑ 𝑦)2 ]

where:
r = degree of relationship between variables x and y
x = observed data for the independent variable
y = observed data for the dependent variable
n = sample size

109
To interpret the degree of linear relationship, the following range of values of coefficient of
correlation may be used:

Coefficient of Interpretation
Correlation
± 1.00 Perfect correlation, perfect relationship
± 0.76 - ± 0.99 Very high correlation, very dependable relationship
± 0.51 - ±0.75 High correlation, marked relationship
± 0.26 - ± 0.50 Low correlation, definite but small relationship
± 0.01 - ± 0.25 Very low correlation, negligible relationship
0.00 No correlation, no relationship

Example 1:
Determine the relationship between height (in inches) and ratings of physical
attractiveness
(1-10) of 10 persons.

Height (in inches) Physical Attractiveness (1-10)


56 7
52 5
60 7
61 8
63 9
58 7
58 6
57 8
57 7
59 8

Solution:
The independent variable (x) would be the height and ratings of physical attractiveness
will be the dependent variable (y).

X y 𝑥2 𝑦2 xy
56 7 3136 49 392
52 5 2704 25 260
60 7 3600 49 420
61 8 3721 64 488
63 9 3969 81 567
58 7 3364 49 406
58 6 3364 36 348
57 8 3249 64 456

110
57 7 3249 49 399
59 8 3481 64 472
581 72 33837 530 4208

Σx = 581 Σ𝑥 2 = 33837 Σxy = 4208

Σy = 72 Σy2 = 530 n = 10

Substitute the values to the formula :

10 (4208) − (581)(72)
𝑟=
√[10(33837) − (581)2 ][10(530) − (72)2 ]

42080−41832
=
√(338370−337561)(5300−5184)

248
=
√(809)(116)

248
=
√93844

= 0.8096

Very high correlation

Example 2:

A biologist assumes that there is a linear relationship between the amount of fertilizer
supplied to tomato plants and the subsequent yield of tomatoes obtained. Seven tomato plants,
of the same variety, were selected at random and treated weekly with a solution in which x grams
of fertilizer was dissolved in a fixed quantity of water. The yield, y kilograms of tomatoes was
recorded.

Plant A B C D E F G
x 1.0 1.5 2.0 2.5 3.0 3.5 4.0
y 3.9 4.4 5.8 6.6 7.0 7.1 7.3

Determine the degree of relationship between the amount of fertilizer supplied to tomato plants
and the yield of tomatoes obtained.

Solution :

x y x2 y2 xy

111
1.0 3.9 1.00 15.21 3.90
1.5 4.4 2.25 19.36 6.60
2.0 5.8 4.00 33.64 11.60
2.5 6.6 6.25 43.56 16.50
3.0 7.0 9.00 81.00 21.00
3.5 7.1 12.25 150.06 24.85
4.0 7.3 16.00 256.00 29.20
17.50 42.10 50.75 589.83 113.65

Σx = 17.50 Σ𝑥 2 = 50.75 Σxy = 113.65

Σy = 24.10 Σy2 = 598.83

Substitute the values to the formula :


7 (113.65) − (17.50)(42.10)
𝑟=
√[7(50.75) − (17.50)2 ][7(598.83) − (42.10)2 ]
795.55−736.75
=
√(355.25−306.25)(4191.81−1772.41)
58.8
=
√(49)(2419.4)
58.8
=
√118550.6

= 0.1708

Very low correlation

SPEARMAN RANK CORRELATION COEFFICIENT(Spearman rho)

The Spearman rank correlation coefficient is another example of correlation coefficient. It


is usually calculated on occasions when it is not convenient, economic or even possible to give
actual values to variables, but only to assign a rank order to instances of each variable. This is a
non-parametric measure which may also be a better indicator that a relationship exists between
two variables when the relationship is non-linear.
The formula is :
6𝛴𝑑 2
rho = 1− 𝑛(𝑛2 −1)

where :
rho = Spearman rho
Σ = sum of
d = difference between ranks of x and y or d = Rx – Ry
n = number of pairs of observations

Example :
The table shows a Verbal Reasoning test score, x, and an English test score y, for each
of a random sample of 8 students who took both tests. Compute for Spearman rho and interpret
the result.

112
Student A B C D E F G H
X 112 113 110 112 113 115 108 111
Y 68 65 74 70 70 75 68 76

x y Rx Ry d or Rx -Ry d2
112 68 4.5 2.5 2.0 4.00
113 65 6.5 1 5.5 30.25
110 74 2 6 -4.0 16.00
112 70 4.5 4.5 0.0 0.00
113 70 6.5 4.5 2.0 4.00
115 75 8 7 1.0 1.00
108 68 1 2.5 -1.5 2.25
111 76 3 8 -5.0 25.00
Σ d = 82.50
2

6(82.50)
rho = 1− 8(82 −1)
495
= 1− 8(63)
495
= 1−
504
= 1−0.98
= 0.02

Very low correlation

SIMPLE REGRESSION ANALYSIS

Regression analysis is concerned with the problem of estimation and forecasting. It


involves identifying the relationship between a dependent variable and one or more independent
variables. Given a series of values of two variables, where values of one variable depend on the
other, we can estimate/predict the value of the dependent variable corresponding to a given value
of the independent variable.

For example, if the GNP levels corresponding to levels of total investment are given, one
can estimate the level of GNP corresponding to some previously unspecified level of investment.

Correlation is used when the end goal is simply to find a number that expresses the
relation between the variables.
Regression is used when the end goal is use the measure of relation to predict values of
the random variable based on values of the fixed variable.

About the example of possible correlation between height and ratings of physical
attractiveness, we could now as, “can we predict a persons’ rating of their attractiveness, based
on their height”.

Let us illustrate this problem assuming that the values of y (ratings of physical
attractiveness) depend on the values assumed by x (height). Let us estimate the value of y when
x is 70.

113
X y
56 7
52 5
60 7
61 8
63 9
58 7
58 6
57 8
57 7
59 8

There are two ways of solving this problem. The first way uses graphical approach which
provides a rough estimate. The second method which will give the exact of y when x is 70 uses
the regression formula.

The first method employs the scatter plot, wherein the points corresponding to the values
of x and y will be plotted on the rectangular coordinate system.

10
8
6
4
2
0
0 20 40 60 80

Physical Att.

After plotting the points, draw the trend line. This line approximates the general direction
of the points. It should be drawn such that the sum of the vertical distances of the points below
the line is approximately equal to the sum of the vertical distances of the points above the line.

Also, the trend line need not pass through the first nor last point. It’s just a matter of
coincidence if the line happen to pass through either or both points.

When x is 70, let us estimate now the value of y using the trend line.

114
10

0
0 10 20 30 Physical Att.
40 50 60 70

This estimate of y will slightly vary depending on how accurately the trend line was drawn.
In this example, the rough estimate of y when x is 70. Using the graphical method is 13.
The second method which is the simplest and probably the most widely used equation for
expressing relationships is the equation of the Least Square Regression Line (LSRL) which
follows the form:
y = a + bx

In this equation, a and b are numerical constant, and calculating a predicted value of y for
any value of x by direct substitution would become easier once they are known. The variable of
which is assumed known and that is being used to predict the value of the other variable is x and
is referred to as the independent, explanatory regressor or predictor variable on the other hand,
the variable which is assumed unknown and that is being predicted on the basis of its relationship
with other variable is y and is referred to as the dependent, explained, regressand, or predicted
variable.

The trend line is the regression line. The line is derived in such a way that the sum of the
squares of the vertical deviations between the line and the individual data plots is minimized. All
other lines will yield a higher result.

In the LSRL equation, values of a and b are derived from “normal” equations which are in
turn derived through graphical techniques and the use of calculus. These normal equations are:
1) Σy = aN + bΣx
2) Σxy = aΣx + bΣ𝑥 2
where :
Σy = sum of the values of the dependent variable
N = number of pairs x and y
Σx = sum of the values of the independent variable x
Σxy = sum of the products of paired values x and y
Σ𝑥 2 = sum of squares of the independent variable x

Solving for the values of a and b using the different methods of solving systems of linear
equations, the formulas for a and b are,

(𝛴𝑦)(𝛴𝑥 2 )−(𝛴𝑥)(𝛴𝑥𝑦)
a=
𝑁(𝛴𝑥 2 )−(𝛴𝑥)2

115
𝑁(𝛴𝑥𝑦)−(𝛴𝑥)(𝛴𝑦)
b=
𝑁(𝛴𝑥 2 )−(𝛴𝑥)2

X y Xy 𝑥2
56 7 392 3136
52 5 260 2704
60 7 420 3600
61 8 488 3721
63 9 567 3969
58 7 406 3364
58 6 348 3364
57 8 456 3249
57 7 399 3249
59 8 472 3481
581 72 4208 33837

(72)(33837)−(581)(4208)
a=
10(33837)−(581)

2,436,264−2,444,848
=
338,370−337,561

−8,584
=
809

= −10.61

(10)(4208)−(581)(72)
b=
10(33837)−(581)2

42080−41832
=
338,370−337,561

248
=
809

= 0.31

Substituting these values into the equation of LSRL :


y = -10.61 + 0.31x

This equation will enable us to solve for y when x is 70.


y = -10.61 + 0.31(70)
= -10.61 + 21.70
= 11.09

This method is very useful in providing a fairly accurate estimate when the values of any
two variables are given.

116
Let’s try this!

A. Indicate whether the statement is true or false.

_____ 1. The coefficient of correlation is always positive.

_____ 2. The scatter diagram is a chart that portrays the correlation between two variables.

_____ 3. The trend line is the regression line.

_____ 4. A 0.40 correlation coefficient indicates a stronger correlation than -0.54.

_____ 5. A simple regression is a regression model that contains one dependent and one
independent variable.

B. 1. Ms. Besmonte administered two forms of Mathematics test to ten(10) first year students.
The first form of the test was given in the morning and the second form was given in the
afternoon. Their scores in the first and second forms are presented below. Determine the
degree of correlation of the tests.

1st form (x) 2nd form (y)


60 45
75 58
80 62
82 60
55 43
77 64
63 49
67 50
81 59
64 51

2. The following shows the scores of 7 students in their first two quizzes in Algebra:

Quiz 1 (x) 7 8 6 5 8 6 9
Quiz 2 (y) 8 8 7 4 9 8 9

a. Draw a scatter diagram.


b. Determine the equation of the LSRL.
c. Find the correlation coefficient and interpret the result.
d. What is the expected student score in the second quiz if the score in the first quiz is 10?

3. The table below shows the price (in peso) and demand (in quantity) of a certain commodity:

117
Demand 10 13 15 9 8 7
Price 8 7 9 6 7 5

a. Draw a scatter diagram.


b. Find the equation of the LSRL.
c. Determine the correlation coefficient and interpret the result.
d. What will be the expected price of a commodity when the demand is 5?

References:

Agcaoili, Zenaida, et al. Statistics for Filipino Students 2004 Edition, National Book Store
Aguaviva, Erlinda et al. Statistics and Probability with Business Applications 1998 Edition, Jollence Publication
Cruz, Myrna, et al. Statistics and Probability Theory 2011 Edition.
Pangan, Milagros R. et al. Statistics for College Students. Grandwater Publishing and Research, Makati City, 1996
Sirug, W. Basic Probability and Statistics, 2011
Walpole, E. R., Introduction to Statistics, 3rd ed., EDCA, 1982.

Online References:

[Link]

[Link]
regression-in-statistics

[Link]
justification/sampling/

[Link]
EJg:1582628185856&source=lnms&tbm=isch&sa=X&ved=2ahUKEwjShtG2xeznAhWSHHAKH
bypC3MQ_AUoAXoECAwQAw&biw=865&bih=414#imgrc=2qZ1J6K660XnfM

[Link]
regression/

[Link]

[Link]

118
[Link]

[Link]

[Link]

[Link]

119
120
121

Common questions

Powered by AI

A one-tailed test predicts the direction of the effect (either greater or less than a certain value) and rejects the null hypothesis if the test statistic falls into one specified tail of the distribution. In contrast, a two-tailed test considers deviations in both directions, rejecting the null hypothesis if the test statistic falls into either tail. One-tailed tests are used when the research hypothesis specifies a particular direction of effect, while two-tailed tests are applied when any significant differences from the null hypothesis could be critical .

The midrange is calculated by averaging the lowest and highest values in the dataset, providing a quick and rough estimate of the dataset’s central tendency. While simpler than other central tendency measures, it can be easily skewed by outliers; however, it is useful for understanding the dataset’s range and central position quickly .

Quartiles, deciles, and percentiles are measurements that divide data into intervals; quartiles split data into four parts, deciles into ten, and percentiles into 100. These metrics are valuable for understanding data distribution and variability, particularly in assessing the relative position of individual data points within the entire data set. Quartiles are frequently used in box plots for visualizing data spread, deciles in financial risk assessment, and percentiles in standardized testing or research to show performance relative to peers .

The weighted mean considers the frequency of different data points, giving more importance to certain values based on their corresponding weights. It is calculated by multiplying each value by its weight, summing these products, and dividing by the sum of the weights. Researchers use the weighted mean to account for unequal representation in datasets, ensuring that data points with higher significance have a proportionate impact on the mean .

Hypothesis testing involves four main steps: stating the hypotheses, formulating an analysis plan, analyzing sample data, and interpreting results. The formulation of hypotheses is critical because it guides the entire testing process. Stating null and alternative hypotheses establishes a basis for testing, where the null hypothesis is typically tested against evidence, and the alternative hypothesis represents the researcher’s conjecture. If outcomes refute the null hypothesis, the alternative hypothesis becomes the preferred assertion .

In hypothesis testing, two types of errors can occur: Type I error (rejecting a true null hypothesis) and Type II error (failing to reject a false null hypothesis). The probability of a Type I error is denoted by α (significance level), while the probability of a Type II error is denoted by β. The choice of significance level affects the probability of making these errors. A lower α reduces the likelihood of a Type I error but increases the probability of a Type II error if the amount of data is not adjusted. The complement to β, called the power of the test, indicates the probability of correctly rejecting a false null hypothesis .

P-values express the probability of observing a test statistic as extreme as, or more intense than, the actual statistic, assuming the null hypothesis is true. A small p-value indicates that the observed data are unusual under the null hypothesis, leading to its rejection. The region of acceptance defines the range of values under which the null hypothesis is not rejected. Both concepts help determine whether a test statistic falls into the region of rejection or acceptance, hence guiding the decision to accept or reject the null hypothesis .

To calculate the mean for grouped data, first, the midpoint for each class is determined. These midpoints are then multiplied by the frequencies of their respective classes. The sum of these products is divided by the total number of data points (n). This process organizes the data efficiently when dealing with larger datasets and frequency distributions .

The mode is useful for categorical data because it identifies the most frequently occurring category, providing insight into the most common outcome. In grouped data, the mode is determined by identifying the modal class, which is the class interval with the highest frequency. For more precise calculations, the formula incorporating class boundaries and differences in frequencies with adjacent classes is used .

For an even-numbered dataset, the median is calculated by averaging the two middlemost values after arranging the data in ascending order. This is different from an odd-numbered dataset, where the median is simply the middle value itself. For example, in a dataset of ten numbers, the median is the average of the 5th and 6th values .

INSTRUCTIONAL
MATERIALS
FOR
ELEMENTARY STATISTICS
AND PROBABILITY
Course Code:
SEMA 30083
Lessons compiled by:
JOAN D. RURAL
No part of this publication may be reproduced or copied by recording or other electronic/
mechanical methods, without the pri
INTRODUCTION
COURSE OVERVIEW
This Instructional Material deals with the study of elementary statistics and probability. It
fo
REFLECTIVE JOURNAL RUBRIC 
Criteria 
Above Average 
Average 
Below Average 
Achievement of 
defined learning 
outcomes 
Learn
TABLE OF CONTENTS
TITLE
Pages
Lesson 1:
Nature of Statistics
6 −11
Lesson 2:
Presentation and Summarizing Data
12 −32
Lesson
LESSON 1 
 
Nature of Statistics 
 
 
 
• 
Define statistics. 
• 
Distinguish between descriptive statistics and inferent
The starting point of statistics can be map out from two fields of interest, namely: the game 
of chance otherwise know
Statistics is divided into two fields. 
 
• 
Descriptive Statistics – includes the techniques which are concerned with 
s
college or you didn’t, a subject has some experience with computers, an average amount 
of experience with computers, or
Quantitative variable is a variable which involves numbers and can be obtained by 
counting 
Example: age, height, salary

You might also like