0% found this document useful (0 votes)
780 views34 pages

Data Management in GEC 05 Statistics

This document provides an overview of managing and understanding data through descriptive and inferential statistics. It discusses the importance of statistical thinking in efficiently handling large data sets and making informed decisions. The document also defines key statistical concepts like population, sample, variable, and observation. It explains how descriptive statistics can be used to summarize data through measures like the mean and median. The overall goal is to introduce students to preliminary concepts in data management and statistical analysis.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
780 views34 pages

Data Management in GEC 05 Statistics

This document provides an overview of managing and understanding data through descriptive and inferential statistics. It discusses the importance of statistical thinking in efficiently handling large data sets and making informed decisions. The document also defines key statistical concepts like population, sample, variable, and observation. It explains how descriptive statistics can be used to summarize data through measures like the mean and median. The overall goal is to introduce students to preliminary concepts in data management and statistical analysis.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
  • MODULE 5: Managing and Understanding Data

DISCLAIMER

These unpolished learning modules were compiled and prepared for personal use of students in GEC
05: Mathematics in the Modern World of Southern Luzon State University (SLSU) ONLY, and not
as a reference material. Unauthorized distribution of the modules is not allowed. The topics included
are given in summary form and does not claim to be complete. The instructors do not claim
ownership of all the contents since it was taken from several resources including books, journals, and
the internet.
MODULE 5: MANAGING AND UNDERSTANDING DATA
TOPIC 1: Preliminaries to Data Management

TOPIC 2: Managing Data using Descriptive Statistics

TOPIC 3: Managing Data Using Inferential Statistics

CORE IDEA: Statistical tools derived from Mathematics are useful in processing and
managing numerical data in order to describe a phenomenon and predict values
(Course Syllabus Mathematics in the Modern World by CHED,2016)

LEARNING OUTCOMES: Upon completion of this module, students are expected to:

Knowledge Discussed management and interpretation of data using descriptive and


inferential statistics;
Skills Gathered and interpret ideas used in statistics;
Apply appropriate measure of statistics for a given set of data;
Apply appropriate technology to accurately determine and explain
different measures of tendency, dispersion, correlation in a given context
Values 6. demonstrate honesty and integrity in the application of Statistics
to research

Performance Task Output

The major output in this module is a statistical research paper. You will gather,
process, present, analyze, and interpret data using descriptive and inferential statistics.
TOPIC 1 Preliminaries to Data management

Statistical thinking will one day be as necessary for efficient citizenship


as the ability to read and write
- H. G. Wells

INTRODUCTION

Suppose that we have collected a data set from a group of thousand students. We have
obtained their marks in various subjects. Now we want to calculate how many numbers of students
have scored below and above the average, we also would like to know how far away from the average
are the scores of the students, or we may look on a sample where scores would represent the whole
population of students. To do this calculation efficiently, we have to take the help of statistics.
Our knowledge of statistics enhances our ability to make good decisions. By learning the data
that we have collected as well as the basic properties of probability, we can improve interpretation
of data in which our decisions are based.
Imagine a world without statistics, how are we going to handle huge data set where
vast information can be generated? If we wish to take full advantage of this available
information then there is a necessity for us to understand the very nature of statistics.

DISCUSSION
The term statistics originated from the Latin word “status” which means state. The use of the
term become popular that it is incorporated to the affairs of state. During those times, different states
had to undergo task of collecting taxes, count of births and deaths, count of agricultural products,
livestock, and even resources from every citizen of the land. They realized that these numerical
figures were essential to the governance of their people.

Nowadays, statistics is used predominantly in almost all fields of study. It provides answers
to research problem, helps in decision making, and aids in the process of choosing appropriate
actions to be taken through the analysis of available information.

Uses of Statistics
1. Prediction
2. Testing
3. Forecasting
4. Preparedness
5. Prediction
6. Political
7. Insurance
8. Consumer
9. Financial
10. Sports

Figure 1.1
What possible research questions you might think of with the above applications?

Basic Concepts
Since you will undoubtedly be given statistical information at some point in your life, you
need to know some techniques to analyze the information thoughtfully. Suppose the following data
on the number of hours students sleep in a day were collected. A sample collection of 14 data sets
were given as: 5, 5.5, 6, 6, 6, 6.5, 6.5, 6.5, 6.5, 7, 7, 8, 8, 9 and organized in a form of graph using dots.

If you did the same example in a Statistics class with the


same number of students, do you think the results would
be the same? Why or why not?

Where do your data appear to cluster? How could you


interpret the clustering?

Figure 1.2

If you are asked to analyze and interpret your data, with this example, you have begun your study of
statistics.

Statistics - is the branch of science that deals with the collection, presentation, organization,
analysis, and interpretation of data. It can also be described as a study of variation.

Why is it important to understand how to collect, present, organize, analyze, and interpret data? The
answer is very simple: Information empowers us to make intelligent choices. A statistical inquiry
allows us to answer problems by giving a clear picture of a particular collection of elements which
we call population. What is a population?
Figure 1.3 Population vs. Sample

Population is the collection of all elements under


consideration in a statistical inquiry.

But what if the population is too big to handle or time consuming to collect?

Sample is a subset of a population from which the raw data are being obtained

How do we define the population of a statistical investigation? The specification of the


population of interest depends upon the scope of the study. Suppose we want to determine the
average expenditure of all household living in Quezon province, then the population of interest is the
collection of all household living in Quezon Province which number assignment can be expressed as
1 to N household. But suppose we want to delimit our scope of study maybe because of time
constraint or budget then we would have to redefine the population of interest. That is, we might
consider all household in the First District of Quezon.

We are interested to look on the varied characteristics of element in the population which we
will call variables. And just like in Algebra we denote these characteristics with letters X, Y, or Z
because their realized values may vary for the different elements in the sample or population. These
collection of all the realized values of the variable under study is the information that we are
interested in: the DATA.

Variable- characteristic or attribute of the elements in a collection that can assume different
values for the different elements.

Observation- realized value of a variable.

Data – collection of observations.

Below are illustrations of variables together with their possible values.

Variable Possible Observations


Sex of the students Male, Female
GWA of transferees 2.0, 2.25, 1.75, …
Religious Affiliation Roman Catholic, Seventh Day Adventist, Born
Again, Iglesia ni Cristo …
Educational Attainment Elementary Graduate, Senior High Graduate
College Graduate, Post Graduate
How do we use these observations to our advantage? Regardless of whether we are using
data of population or sample, we have to give meaning to this information by summarizing the bulk
of information to a single numeric value to describe a particular feature of the whole population or
sample.
Data Structures
A data set consist of some basic measurement/s of individual items or elementary units
which may be people, firms, cities, or just anything of interest.

There are three basic ways of classifying data set


1. By the number of variables ( univariate, bivariate, and multivariate)
2. By the kind of information ( quantitative or numbers and qualitative or categories)
3. By whether the data set is a time sequence or cross-sectional data

Example: The following table shows a Human Resource database representing the status of five
employees

Table 1.1 Employment Status of Five Employees

Gender Salary Education Year of


Experience

M P42,300.00 HS 9
F P31,800.00 BA 4
M P29,500.00 MBA 2
F P58,100.00 MBA 15
F P36,000.00 BA 7
HS= High School Diploma BA= College Degree MBA= Master’s Degree in Business

What is an elementary unit for the above data set? EMPLOYEE


What kind of data set is this, univariate, bivariate, or multivariate? MULTIVARIATE
Which of these variables are quantitative? SALARY and YEAR OF EXPERIENCE
qualitative? GENDER and EDUCATION
Which variables, if any, are ordinal qualitative? EDUCATION
Is this a time series or cross sectional? CROSS SECTIONAL

Parameter- summary measure describing a specific characteristic of the population.

Statistic- is a summary measure describing a specific characteristic of the sample.

Example: A marketing company is interested in the proportion of people that will buy a particular
beauty product. Define the following in terms of the study. Give examples where appropriate.

Population: All people (you may consider all people living in Metro Manila)
Sample: A particular group of people (since it is a product endorsement then a particular group
of people that we might be interested in are women with ages 15-40)
Parameter: proportion of all people who will buy the product
Statistic: proportion of sample who will buy the product
Variable: X = number of women who will buy the product
Data: buy or not to buy

Example: The pandemic brought change in the delivery of the lesson for the students of Southern
Luzon State University. The Student Affairs Office wants to know the average (mean) amount of
money a first -year college students spend on internet data a day since lectures and quizzes will be
in online platform. A random survey on 100 first year students at the college was done. Three of those
students spent P150, P50, and P100, respectively. Determine the population, sample, parameter,
statistic, variable, and data of the study.

The population is all first -year students attending SLSU this term.

The sample could be all students enrolled in one section of a beginning statistics course (although
this sample may not represent the entire population).

The parameter is the average (mean) amount of money spent by first year college students at SLSU

The statistic is the average (mean) amount of money spent by first year college students in the
sample.

The variable could be the amount of money spent by one first year student. Let X = the amount of
money spent by one first year student attending at SLSU.

The data are the amounts in peso spent by the first- year students. Examples of the data are P150,
P50, and P100.

On your own

A politician is interested in the proportion of voters in his district that think he is doing a good job.
Define the following in terms of the study. Give examples where appropriate.

• Population
• Sample
• Parameter
• Statistic
• Variable
• Data

Fields of Statistics
In this module, you will also learn how to organize and summarize data. Organizing and
summarizing data is called descriptive statistics. On the other hand, we use methods in inferential
statistics to come up with generalizations or inferences about the population using the information
in the selected sample.

Descriptive Statistics deals with the techniques used in the collection, presentation, organization, and
analysis of the data on hand.

Inferential Statistics- deals with the technique used in analyzing the sample data that will lead to
generalizations about a population from which the sample came from.

Effective interpretation of data (inference) is based on good procedures for producing data
and thoughtful examination of the data. You will encounter what will seem to be too many
mathematical formulas for interpreting data. The goal of statistics is not to perform numerous
calculations using the formulas, but to gain an understanding of your data. The calculations can be
done using a calculator or a computer. The understanding must come from you. If you can thoroughly
grasp the basics of statistics, you can be more confident in the decisions you make in life.

From the raw data which is not so


informative we use the techniques
of descriptive statistics to give
meaning to this observations:
through tables, graphs, or through
summary measures such as
averages and deviations.

When it comes to inferential statistics, there are generally two forms: estimation statistics
and hypothesis testing. Estimation statistics” is a fancy way of saying that you are estimating
population values based on your sample data – confidence interval. While hypothesis testing is
simply another way of drawing conclusion about a population parameter (“parameter” is simply a
number, such as a mean, that includes the full population and not just a sample)

Statistical Inquiry – is a designed research that provides information needed to


solve a research problem.
In a statistical inquiry, the researcher identifies the problem, plans the study, collects the
data, explore the data, analyzes the data and interpret the results. What are the different statistical
inquiries that we can explore with the wealth of data that we have? Some inquiries simply aim to
describe the characteristic of the sample data in the population through computation of estimates of
parameter such as total, averages, and proportion. Other studies would like to focus on relationships
among the different variables of interest. Still other inquiries forecast future values of a variable using
a sequence of observations on the same variable taken over time. There are many types of statistical
inquiries that we can unravel with research objectives ranging from most simple to the most complex.
Four Basic Activities of Statistics

1. Designing a Plan for Data Collection – this involves planning the details of data gathering using a
random sample from a larger population. One advantage of taking randomness in this design phase
is to ensure validity of statistical inferences that will be drawn later.
2. Exploring the Data – this involves looking at your data sets in many ways possible by describing or
summarizing your data. This will entail us to study our observed data in details by verifying data
errors, selecting appropriate analysis, and validating the statistical techniques that are to be used in
further analysis.
3. Estimating an Unknown Quantity – this phase produces the best educated guess based on the
available data. If you knew how accurate your estimates are, then you will have some indication of
the size of the error involved by using the estimated value in place of the actual unknown value.
4. Hypothesis Testing – this phase uses data to decide between two or more different possibilities in
order to resolve an issue. It produces a definite decision about which of the possibilities is correct
based on the data.
Example: Which of the four basic activities of statistics is represented by each of the following
situations.
1. You are trying to determine the quality of the learning modules for distribution based on the
careful observations of the content of the material. ESTIMATING AN UNKNOWN QUANTITY
2. A factory’s quality control division is examining the productivity in order to identify possible
trouble spots. EXPLORING THE DATA
3. A security services agency is being charged with gender discrimination. Data that show salaries of
men and women employees are presented to the court to convince that there is a consistent pattern
of discrimination and that difference is not due to randomness alone. HYPOTHESIS TESTING
4. You are wondering who to interview, how many to interview, and how to process the results so
that your questions can be answered. DESIGNING THE STUDY

Partial List of General Research Objectives

1. Describe the characteristic of the elements in the population through the computation or
estimation of parameter such as the proportion, average, and total
2. Compare the characteristics of the elements in the different subgroups in the population
through contrasts of their respective summary measures.
3. Determine the nature and strength of relationships among the different variables of interest.
4. Determine the effects of one or more variables on a response variable.
5. Classify patterns and trends in the values of the variable over time or space.
6. Predict the value of a variable based upon its relationship with another variable.
7. Forecast future values of a variable using a sequence of observations on the same variable
taken overtime

Determining the Sample size

In doing a research, if the population is too big to handle, a sample is acceptable. Determining
sample size may be the most important step in any statistical study. If the sample fairly represents
the population as a whole, then it is reasonable to make inferences from the sample to the population.
It may also be emphasized that samples that are too large may waste resources and samples that are
too small may lead to inaccurate results.

Approaches to determine sample size


a. Using a census for a small population
b. Using the sample size of similar studies
c. Using published tables by well-established authors (Cochran’s formula)
d. Using sample size calculator
e. Other formulas
The different approaches on determining sample size varies on the types of research and settings.
However, a large sample size leads to more precision of the various assumptions of the population.
We can say that the bigger the sample size the better results you have.
Two Terms that must not be taken for granted
1. confidence interval or the margin of error (e)
2. confidence level (in %)

The margin of error tells us how much a percentage points deviate from the real
population value while confidence interval tells the researcher how sure the responses of
the sample represent the population.
Example: In a survey conducted by [Link] on “Happiness Index” among Filipinos.

Jon Carlos Rodriguez, ABS-CBN News; Posted Au 31, 2016 12:12 PM| Updated as of
Sept 01 2016 09:18

MANILA (UPDATE) – “Filipino employees are the happiest in Southeast Asia and their
positive attitude is likely to boost the economy, results of a [Link] survey
released Aug. 31, 2016, showed.

The Philippines topped the seven nation “Happiness Index” with 73% of the
respondents saying they were happy with their jobs. Indonesia came in second at 71%,
while Malaysia scored the lowest among the seven countries in Southeast Asia, at 41%

Suppose that the researcher used a 5% margin of error, how do we interpret the results given the
data.
Table 1: Selected Results of Job Happiness Index and its Interpretation
Country Job Happiness Index Using Assuming 𝑒 = ∓5%
Samples This implies that in the population …
Philippines 73% 68% to 78% of Filipinos are happy with their jobs
Indonesia 71% 66% to 76% of Indonesians are happy with their jobs.
Thailand 61% ?
Vietnam 60% ?
Hongkong 57% ?
Read the full report on [Link]

What about Slovin’s Formula in determining sample size?


The use of Slovin’s formula is quite popular in determining sample size of a survey research design.
Punzalan and Tejada (2012) suggest that the formula is applicable only when estimating a population
proportion and when the confidence level is 95%. They added that the formula is optimal when the
population proportion is suspected to be close to 0.5. So, if the assumptions given were not met, it is
not advisable to use Slovin’s formula.
𝑁
Slovin’s formula is given by 𝑛 = 1+𝑁𝑒 2
where n is the sample size and N is the population size. It is
widely used because of its simplicity where the computation is based on population size and margin
of error.
Sampling Techniques
• Sampling is the process of selecting a representative group from the population under
study.
• A sampling frame is a list of all the items in your population. It’s a complete list of everyone
or everything you want to study. The difference between a population and a sampling frame
is that the population is general and the frame is specific.
• The target population is the total group of individuals from which the sample might be
drawn.
• A sample is the group of people who take part in the investigation. The people who take
part are referred to as “participants”.
• Generalizability refers to the extent to which we can apply the findings of our research to
the target population we are interested in.

The methods of selecting samples from the given population is shown in the diagram above.
Probability Sampling used random selection wherein each element in the sampled population has
equal chances of being selected. All the elements that belong to the population must be included in
the selection process. On the other hand, non-probability sampling is based on personal choice. It
does not follow the randomization mechanism in identifying the sampling units. It allows the
researcher to choose the elements in the sample subjectively.
Example Categorize the type of sampling used in each of the following situations:
a. To conduct a pre-election opinion poll on a proposed amendment to the Philippine Constitution,
a random sample of 10 cellphone prefixes (first 4 digits of the phone number) was selected, and all
households from the phone prefixes were called.
b. With the COVID- 19 situation worldwide KaPsych conduct a survey on depression among the
elderly, a sample of 30 patients in one nursing home was used.
c. To maintain quality control in a brewery, every 20th bottle of beer coming off the production line
was opened and tested.
d. Subscribers to the magazine Perfect Home were assigned numbers. Then a sample of 30
subscribers was selected by using a random-number table. Then subscribers in the sample were
invited to rate new compact disc players for a ”What Subscribers Think” column
[Link] judge the appeal of a proposed tv sitcom, a random sample of 10 people from each of three
different age categories was selected and those chosen were asked to rate a pilot show.
Answers: a. cluster b. convenience c systematic d. SRS e. stratified
Example The population consists of all CAS students. You plan to obtain a simple random sample of
100 CAS students by using the sampling frame of statistics students.
1. Why does this result in undercoverage? Explain
2. If you did this, might this result in sampling error? Explain. Yes since it has a lot of undercoverage
then the sample will not be a representation of the population
DATA GATHERING TECHNIQUES

Data collection methods fall into four general categories:


1. A census is a survey of a whole population.. Censuses can be very expensive and time-
consuming, if the population is large.
2. A sample survey takes a fraction of the population. Sample surveys are cheaper than censuses,
but are not as accurate. Bias can also be an issue.
3. An experiment is a controlled study of a group. Experiments are very common in the medical
fields. The researcher controls how members are placed study groups and which treatment
each group receives. Bias can be a major issue with experiments.
4. An observational study is about the same as an experiment. However, the researcher does not
use control groups or assign treatments.

What’s best?

There is no one “best” data collection method. Each method has its pros and cons. Which
one you choose depends on what kind of data you have (i.e. qualitative data or quantitative data)
and which pros/cons are important for your study.

Data Gathering Techniques

Quantitative Data Collection Qualitative Data Collection


numerical attributes
Direct or interview In-depth interview
Indirect or Questionnaire (paper/pencil or Indirect or Questionnaire (paper/pencil or
web-based: closed ended) web-based: closed ended)
Registration Document review
Experimental/Clinical Trials Focus group Discussion
Observation Observation
Regardless of the types of data you are going to use, gathering data in a qualitative study
will take a great deal of time because of the wide-ranging themes or record that will be considered
by the researcher.
Discuss the advantages and disadvantages of the data gathering techniques and decide on
what technique will you use in your future research project.

Name ______________________________________
Course & Year ____________________________
EXERCISES1.1

A. Recall that a variable is a characteristic or attribute of the elements in a collection that can
assume the different values for the different elements while an observation is a realized
value of a variable. Below are some variables under study, determine possible observation
values in each of the following.
Variable Possible Observations
[Link]
[Link] Status of an
employee
[Link] Income
[Link] of study
[Link] modality
6. Civil Status
7. IQ scores
8. Blood Type
9. weights of newborn infants
10. Faculty rank

B. Identify the population under study, define your sample and variable/s of interest.

1. The Department of Health is interested in determining the percentage of vulnerable members


of society who are over 59 years old infected by the CORONA Virus in NCR.
Population
Sample
Variable of interest
2. The Office of the Student Admissions is studying the relationship between the score in the
entrance exam during application and the General Weighted Average (GWA) upon graduation
among graduates of the university from 2017-2020
Population
Sample
Variable of interest
………………………………………………………………………………………………………………………………………………….

Determining Sample and Sample Size


3. Suppose we want to determine the opinion of Business Administration students regarding
pre- marital sex. The population consist of N=500 and the sample size is n=50. Out of the 500
there are 300 female and 200 male students. Determine the number of male and female
considered as sample

4. An airline offers a certain flight once per day that usually contains about 250 passengers. The
flight offers seats in first class (most expensive), business class, and economy class (least
expensive). The airline wants to survey 500 passengers of this flight about their overall
satisfaction. The passengers will be selected using a cluster random sample where each flight
is a cluster. Why might the airline choose a cluster random sample instead of a simple random
sample in this setting?

5. Use the Table of random numbers, select 10 distinct numbers from each of the total possible
numbers.

a. 1 to 437
b. 1 to 3495

6. A government official asked the head of a research team on the veracity of the survey results.
He said that the public should not merely rely on presidential race survey despite the latest
result showing one candidate leading the presidential race at 33%. He explains that based on
the estimates, the weights did not seem to tally with the actual percentage of voters.
Complete the table below and see if there is a basis for the contention.

Cluster Number of Percent Number of Correct number of Discrepancy/


registered n/N samples samples using Difference
Voters used in the proportional Correct
survey allocation number of
n = 1800 samples
NCR 6253249 0.1150 300 1800(0.1150)=207 300-207=193
Luzon 24164451 ? 600 ? ?
Visayas 11316789 ? 300 ? ?
Mindanao 12629265 ? 600 ? ?
TOTAL 54363844 ? 1800 1800

7. For the following studies, describe the population, sample, population parameter, and sample
statistic.
a. In order to gauge public opinion on how to handle Philippines growing cases of
COVID-19, the Department of Health surveyed 1001 Filipinos by phone.
b. The Higher Education research team conducts an annual study of attitudes of college
freshmen by surveying approximately 8560 first year students at 45 universities and
colleges in the country. There are approximately 65,000 first-year college students in
the country.

C. Select three different news stories from the past week that involve statistics in some way. For
each case, write a paragraph describing the role of statistics in the story.

References

Bennett, Jeffrey O. et al, Using and Understanding Mathematics: A Quantitative Reasoning Approach,
Pearson Education Inc.,2011

Blay, Basilia E., et al. Mathematical Trips in the Modern World, Outcomes- Based Approach, Anvil
Publishing, Inc, 2020

Earnhart, Richard T., & Adina, Edgar M., Mathematics in the Modern World, Outcome-Based
Module, C&E Publishing, Inc., 2018.
Almeda, JosefinaV., et al., Elementary Statistics, The University of the Philippines Press,2010

Stephanie Glen. "Data Collection Methods" From [Link]: Elementary Statistics for the rest of
us! [Link]

[Link]

[Link]

[Link]

[Link]
statistics%2F&psig=AOvVaw3kO2KPXrpgXygnEuOw66oa&ust=1599051563538000&source=imag
es&cd=vfe&ved=2ahUKEwie-dK0gcjrAhURzYsBHV6fDkUQr4kDegUIARDUAQ

TOPIC 2 Managing and Understanding Data using Descriptive Statistics


One of the ways to see the “big picture” of your observed data sets is through summarization.
This is to use one or more selected or computed values to represent the data set. In statistics our goal
is to identify features that the observed cases are in common or treating the information as a whole.
Descriptive statistics are used to describe the attributes of a particular group of people, places or
things without trying to infer the population.

Data are summarized according to the


estimates of:

Measures of Variability
Measures of Central Tendency
Range Measures of Shapes
Mode
Variance Skewness
Mean
Standard Deviation Kurtosis
median
Coefficient of Variation

The field of statistics is the science of learning from data. When statistical principles are correctly
applied, statistical analyses tend to produce accurate results. When describing or comparing data, a
single value which describes the center is very important. Here are the common measures

Measures of Central Tendency: Mean, Median, and Mode


A measure of central tendency is a summary statistic that represents the center point or typical value
of a dataset. These measures indicate where most values in a distribution fall and are also referred
to as the central location of a distribution.
You can think of it as the tendency of data to cluster around a middle value. In statistics, the three
most common measures of central tendency are the mean, median, and mode. Each of these
measures calculates the location of the central point using a different method.
Choosing the best measure of central tendency depends on the type of data you have.
Locating the Center of Your Data
The three distributions below represent different data conditions. In each distribution, look for the
region where the most common values fall. Even though the shapes and type of data are different,
you can find that central location. That’s the area in the distribution where the most common values
are located.

As the graphs highlight, you can see where most values tend to occur. That’s the concept. Measures
of central tendency represent this idea with a value. You need to know the type of data you have, and
graph it, before choosing a measure of central tendency!
The central tendency of a distribution represents one characteristic of a distribution. Another aspect
is the variability around that central value. While measures of variability describe how far away the
data points tend to fall from the center.

Measures of Variability: Range, Interquartile Range, Variance, and Standard Deviation

A measure of variability is a summary statistic that represents the amount of dispersion in a dataset.
How spread out are the values? While a measure of central tendency describes the typical value,
measures of variability define how far away the data points tend to fall from the center.

A low dispersion indicates that the data points tend to be clustered tightly around the center. High
dispersion signifies that they tend to fall further away.
The two plots above show the difference graphically for distributions with the same mean but more
and less dispersion. The panel on the left shows a distribution that is tightly clustered around the
average, while the distribution in the right panel is more spread out.

Why Understanding Variability is Important

Analysts frequently use the mean to summarize the center of population or a process. While the
mean is relevant, people often react to variability even more. When a distribution has lower
variability, the values in a dataset are more consistent. However, when the variability is higher, the
data points are more dissimilar and extreme values become more likely. Consequently,
understanding variability helps you grasp the likelihood of unusual events.

Recall how to compute for the summary statistics.


Exercises 5.1

Name
Course & Year

1-4 Find the mean, median, and mode for the given set of data and interpret the result.

1. The reported active cases of COVID-19 in a certain barangay of a municipality is shown in the
following: 3,5,6,4,7,8,6,9,10,4,6,7,5,8,9,8,3,4,5, and 5.

2. An investigator found out that the number of dismissal of cases for 12 months were 8, 5, 6, 3,
4, 7, 10, 12, 16, 8, 9, 12.

3. Out of 70 numbers, 15 were 7’s, 10 were 9’s, 12 were 14’s, 8 were 25’s, and the remaining
were 20’s.

4. Consider the quality of cars, as measured by the number of cars requiring extra work after
assembly, in each day’s production for 15 days.
30, 34, 9, 14, 28, 9, 23, 0, 5, 23, 25, 7, 0, 3, 24.

5. The following data are the grades in GEC 05 and GPA (grade point average) of 60 students in
the College of Arts and Sciences.
Male GEC05 Grade GPA Female GEC 05 Grade GPA
1 99 95 1 96 99
2 98 91 2 91 98
3 98 90 3 91 96
4 98 89 4 90 89
5 97 89 5 90 89
6 96 88 6 89 89
7 95 87 7 89 89
8 93 87 8 88 88
9 91 86 9 88 88
10 90 85 10 88 88
11 89 85 11 86 88
12 89 84 12 85 87
13 89 84 13 85 87
14 89 82 14 84 86
15 89 81 15 84 86
16 89 81 16 84 86
17 88 80 17 84 86
18 88 80 18 81 85
19 87 78 19 80 85
20 84 77 20 80 84
21 83 77 21 80 84
22 80 77 22 80 84
23 80 76 23 80 83
24 79 76 24 80 84
25 77 76 25 79 82
26 76 75 26 78 80
27 76 74 27 76 80
28 76 72 28 76 79
29 75 72 29 73 78
30 70 70 30 70 75

Encode your data in Microsoft Excel.


Manage your data Using Microsoft Excel Add Ins under Descriptive Statistics.
Attached your print out in tour answer sheet.
Please refer to your print out to do/answer the following questions. Justify your answers.
Who performed better in GEC 05?
Who showed more uniform set of data in GEC 05?
Who showed better overall academic performance?
Who showed more variable set of performance?
TOPIC 2 Managing and Understanding Data Using Inferential Statistics

While descriptive statistics describes what is going on in our sample data set, inferential
statistics allows us to predict trends about a larger population based on a study of the samples taken
from it. There are also inferences where we examine the relationships among variables within a
sample and then make generalizations or predictions about how those variables will relate to the
larger population.
The prerequisite in following the discussion on test of hypothesis is a clear grasp on the basic
concepts of inferential statistics. Recall the following concepts used in hypothesis testing. Discuss
them with your groupmates.
1. Why do we need to test the hypothesis?
2. Two kinds of statistical hypothesis
3. Level of Significance
4. One-Tailed or Two-Tailed Test, when do we use it?
5. Type 1 / Type II Errors
6. p-value
7. Which hypothesis do we reject/ do not reject?
8. When do we reject H0?

A hypothesis test (or test of significance) is a standard procedure for testing a claim about a property
of a population.

This topic presents individual components of a hypothesis test. We should know and understand the
following:
How to identify the null hypothesis and alternative hypothesis from a given claim, and how to express
both in symbolic form
How to calculate the value of the test statistic, given a claim and sample data
How to identify the critical value(s), given a significance level
How to identify the P-value, given a value of the test statistic
How to state the conclusion about a claim in simple and nontechnical terms

The null hypothesis (denoted by H0) is a statement that the value of a population parameter (such
as proportion, mean, or standard deviation) is equal to some claimed value. This is the one which the
researcher always hopes to reject. In writing, it uses the (=) symbol

The alternative hypothesis (denoted by H1 or Ha or HA) is the statement that the parameter has a
value that somehow differs from the null hypothesis.
The symbolic form of the alternative hypothesis must use one of these symbols:  , <, >.

The critical region (or rejection region) is the set of all values of the test statistic that cause us to
reject the null hypothesis

The significance level (denoted by  ) is the probability that the test statistic will fall in the critical
region when the null hypothesis is actually true. Common choices for  are 0.05, 0.01, and 0.10.
The P-value (or p-value or probability value) is the probability of getting a value of the test statistic
that is at least as extreme as the one representing the sample data, assuming that the null hypothesis
is true.
Critical region in the left tail: P-value = area to the left of the test statistic
Critical region in the right tail: P-value = area to the right of the test statistic
Critical region in two tails:P-value = twice the area in the tail beyond the test statistic

The tails in a distribution are the extreme regions bounded by critical values.
Determinations of P-values and critical values are affected by whether a critical region is in two tails,
the left tail, or the right tail. It therefore becomes important to correctly characterize a hypothesis
test as two-tailed, left-tailed, or right-tailed.

Two-tailed Test Left-tailed test Right-tailed test

We always test the null hypothesis. The initial conclusion will always be one of the following:
1. Reject the null hypothesis.
2. Fail to reject ( Do not Reject) the null hypothesis.

P-value method:
Using the significance level :

If P-value , reject H0.

If P-value > , fail to reject H0.

Never conclude a hypothesis test with a statement of “reject the null hypothesis” or “fail to reject the
null hypothesis.” Always make sense of the conclusion with a statement that uses simple nontechnical
wording that addresses the original claim.
STEPS IN HYPOTHESIS TESTING
1. Formulate H0 and Ha ;
2. Set the level of significance (α); get p-value;
3. Formulate the decision rule (when to reject H0;
4. Give the decision (Reject H0 or Do not reject H0; ; and
5. Draw conclusions

Testing the Significance on Means


There are several types of research problems which utilizes this test statistic depending on the
number of groups we want to compare. But in this module, we only focus on comparing two samples:

1. t-test for two independent samples


2. t-test for dependent or correlated samples

Testing the Significance of Difference Between Proportions

When we are comparing two population means based on our sample, we want to validate
whether the observed difference between the sample means from observations taken from the two
populations is large enough to indicate an actual difference. Let us use the data in Exercises 5.1 #5
for testing hypothesis using t-test for two independent samples.

Problem: Is there a significant difference on the GEC grades of male and female students?

Let us analyze the problem and enumerate the items which will help us test our hypothesis.
Data to investigate Descriptive Statistics Test Statistic
GEC 05 grades of male and Mean t-test on means
female students
Difference on their Mean difference of male and Independent samples
performance in GEC 05 female students
Type of Test 2-tailed test
Go to Data Analysis in Microsoft Excel and perform t-test: Two sample assuming equal variance
In this module, we will be using Microsoft Excel.

5-step solution Statements


1. Formulate H0 and Ha H0: H a :  M   F There is no significant
difference on the Math Grades of male and
female students.
H a : M   F
There is a significant
difference on the Math grades of male and
female students

2. Set the level of significance (α); get p-value;  = 0.05 two-tailed test p-value = 0.067013954

3. Formulate the decision rule (when to reject


if p- value  
H0
H0; Decision Rule: Reject
0.067  0.05
[Link] the decision H p − value  
Decision Do not reject 0
0.067>0.05
[Link] I therefore conclude that there is no significant
difference on Math grades of male and female
students

What is the implication of the result?

Dependent or Correlated Sample

Example: As an aid for improving employee‟s working habits, eight employees were
randomly selected to attend a seminar workshop on the importance of work. The table
shows the number of workload done per week before and after attending the seminar
workshop. At 5% level of significance, did attending the seminar-workshop increase the
performance level of employees?

Before 14 13 9 9 10 10 12 7
After 11 15 10 14 11 13 11 12

1. Open a new MS Excel file, encode the Before and After data set in cells A1 to A8 and B1 to
B8 respectively
2. Click „Data‟ tab then Data Analysis at the right of the toolbar. A dialogue box will appear.
3. From the Data Analysis dialogue box, choose „t-test: paired two sample for means‟ then
click OK.
4. Type the locations for your variable 1 data (before) into the input range of the first text
box, type: “A1:A18”. On the other hand, for variable 2 data (after) into the second text box
which is in cells B1 to B8, type: “B1:B8” into the box.
5. Type “0” into the hypothesized mean difference box and set the alpha level. 6. Choose an
output area example in cell D1 then click OK.

Perform the 5-step on hypothesis testing.

CORRELATION

Correlation is a statistical method that determines the degree of relationship between two
different variables.
The relationship between any two variables can vary from strong, weak, to none.
Correlation coefficient ranges from -1 to 1.
When a relationship is strong, this means that knowing an object’s score on one variable helps
to predict their score on the second variable.
If the correlation or relationship between variables A and B is a weak one, then knowing an
object’s score on variable A does not help to predict their score on variable B.

Positive Correlation: The correlation is said to be positive correlation if the values of two
variables changing with same direction.
Negative Correlation: The correlation is said to be negative correlation when the values of
variables change with opposite direction.
One problem in correlation is that, just because two variables are correlated, it does not mean
that one variable caused the other.

Pearson Correlation coefficient


-the single most common type of correlation
-a measure of the strength of a relationship between two continuous variables
The term 𝑟 is called the Pearson product-moment correlation coefficient, named after Karl
Pearson, an English statistician who developed several coefficients of correlation along with
other significant statistical concepts.

Pearson’s coefficient r
𝑁 ∑ 𝑥𝑦−(∑ 𝑥)(∑ 𝑦)
r= 2
√(𝑁 ∑ 𝑥 2 −(∑ 𝑥)2 )(𝑁 ∑ 𝑦 2 −(∑ 𝑦)

𝑁 is equal to the number of pairs of scores

Pearson’s Correlation Coefficient


Values of 𝜌 Interpretation
0 No linear association
0 < 𝜌 < 0.2 Very weak linear association
0.2 ≤ 𝜌 < 0.4 Weak linear association
0.4 ≤ 𝜌 < 0.6 Moderate linear association
0.6 ≤ 𝜌 < 0.8 Strong linear association
0.8 ≤ 𝜌 < 1 Very strong linear association
1 Perfect linear association
Example #1.
A tobacco company statistician wishes to know whether heavy smoking is related to
longevity. From a sample of recently deceased smokers, the number of cigarettes (estimated
on a per day for their last five years after visits with their surviving relatives) is paired with
the number of years they lived.

Cigarettes Years lived


25 63
35 68
10 72
40 62
85 65
75 46
60 51
45 60
50 55
∑ 𝑥=25+35+…+50=425
∑ 𝑥 2 =252 + 352 + ⋯ + 502 = 24525
∑ 𝑦=63+68+…+55=542
∑ 𝑦 2 = 632 + 682 + ⋯ + 552 =33 188
2
(∑ 𝑥) = 4252 = 180, 625

2
(∑ 𝑦) =5422 =293 764

∑ 𝑥𝑦 = 25(63) + ⋯ + 50(55) = 24640

Note that there 9 pairs of scores, thus we have


𝑁 ∑ 𝑥𝑦−(∑ 𝑥)(∑ 𝑦)
r= 2
√(𝑁 ∑ 𝑥 2 −(∑ 𝑥)2 )(𝑁 ∑ 𝑦 2 −(∑ 𝑦)

9(24 640)−(425)(542)
= = −0.61
√40 100(4928)

Thus, 𝑟 = −0.61 means that there is a strong negative correlation between smoking and
longevity. This indicates that the higher the number of cigarettes smoked in the past five
years, the lower the number of years lived.
Testing of Difference of a Correlation Coefficient

• A correlation coefficient may be tested to determine whether the coefficient significantly


differs from zero. The values r is obtained on a sample. The value ( 𝜌) rho is the
population’s correlation coefficient.

𝐻0 : 𝜌 = 0
𝐻𝑎 : 𝜌 ≠ 0
t-Test Formula

r 𝑁−2
𝑡= = 𝑟 (√ )
1 − r 2 1 − 𝑟2

N−2
• From the previous example, r=-0.61 was obtained which means that there is a strong

negative relationship between smoking and longevity.

𝑟 = −0.61
𝑁=9

9−2
𝑡 = −0.61 (√ ) = −2.042
1 − (−0.61)2

For a 2-tailed test of significance at 𝛼 = 0.05 with df=N-2=7, the critical values of t
are 𝑡 = 2.365 and 𝑡 = −2.365. Now since −2.365 < 𝑡𝑐𝑜𝑚𝑝𝑢𝑡𝑒𝑑 < 2.365, we fail to
reject 𝐻0 . Thus, 𝑟 = −0.61 indicates a non-significant relationship at 𝛼 = 0.05.

Interpretation of 𝒓𝟐

-𝒓𝟐 is called the coefficient of determination, and it has two important interpretations.

-First, it explains the proportion of variance in one variable accounted for by the other
variable.

-for example, if the correlation between two variables 𝐴 and 𝐵 is 𝑟 = 0.25 then 𝑟 2 =
(0.25)2 = 0.0625. This means that the variable 𝐴 explains approximately 6% of the
variation in variable 𝐵. In other words, 6% of the variation in variable 𝐵 can be
explained by its relationship to variable 𝐴, and 94% of the total variance between
variances 𝐴 and 𝐵 remains unexplained.

-Second, it is use as a measure of strength between two or more 𝑟 values. For example,
if given an 𝑟1 = 0.25 and 𝑟2 = 0.50, the second 𝑟 −value is twice as great as the first
𝑟 −value. However, 𝑟 2 values are 6.25% and 25%, respectively. Thus, 𝑟 = 0.50
explains four times as much variance as does an𝑟 = 0.25.
Exercises 5.2

Explain your reasoning.


1. Since the pandemic the price of gasoline varies constantly from the last two [Link],
you observed that local gas stations charges higher price. You want to test your hypothesis that
local gas stations are charging much more than the national average price for gasoline. What is your
claim about this problem? What is the null hypothesis of this study?

2. You carry out a test of the hypothesis described in #1. If the results show that you cannot reject
the null hypothesis, what conclusion can you generate based on your claim?

3. Explain why accepting the null hypothesis is not a possible outcome.

4. What type of correlation would you expect between wages and the unemployment rate?
Exercises 5.3

Test of Hypothesis

The following data are the grades in GEC 05 and GPA (grade point average) of 60 students in the
College of Arts and Sciences.
Male GEC05 Grade GPA Female GEC 05 Grade GPA
1 99 95 1 96 99
2 98 91 2 91 98
3 98 90 3 91 96
4 98 89 4 90 89
5 97 89 5 90 89
6 96 88 6 89 89
7 95 87 7 89 89
8 93 87 8 88 88
9 91 86 9 88 88
10 90 85 10 88 88
11 89 85 11 86 88
12 89 84 12 85 87
13 89 84 13 85 87
14 89 82 14 84 86
15 89 81 15 84 86
16 89 81 16 84 86
17 88 80 17 84 86
18 88 80 18 81 85
19 87 78 19 80 85
20 84 77 20 80 84
21 83 77 21 80 84
22 80 77 22 80 84
23 80 76 23 80 83
24 79 76 24 80 84
25 77 76 25 79 82
26 76 75 26 78 80
27 76 74 27 76 80
28 76 72 28 76 79
29 75 72 29 73 78
30 70 70 30 70 75

Perform the 5 steps in hypothesis testing. Is there a significant correlation between the GEC 05
grades and the GPA of male students?
2. A dietitian wishes to see if a person’s cholesterol level will change if the diet is supplemented by a
certain mineral. Six respondents were pretested and then took the mineral supplement for a six
week period. Can it be concluded that the cholesterol level has been changed at  = 0.10. Assume
that the data is approximately normally [Link] level is measured in milligrams per
deciliter.

Subject 1 2 3 4 5 6
Before 210 235 208 190 172 244
After 190 170 210 188 173 228

Perform the 5-steps in hypothesis testing.

3. A researcher wishes to determine whether the salaries of professional nurses as front liners
employed by private hospitals are higher than those of nurses employed by government-owned
hospitals. She selects a sample of nurses from each type of hospital and calculate the mean and
standard deviations of their salaries. At  = 0.01 Can you conclude that private hospitals pay more
than the government hospitals?

Private Government-owned
x = P26,800 x = P25,400
s=P600 s=P450
n = 10 n= 8

Perform the 5 - steps in hypothesis testing.

DISCLAIMER 
These unpolished learning modules were compiled and prepared for personal use of students in GEC 
05: Mathemati
MODULE 5: MANAGING AND UNDERSTANDING DATA 
 
TOPIC 1:  Preliminaries to Data Management 
TOPIC 2: Managing Data using Descrip
TOPIC 1 Preliminaries to Data management 
INTRODUCTION 
DISCUSSION 
 
 
 
 
 
 
 
Statistical thinking will one day be as nec
1. Prediction 
2. Testing 
3. Forecasting 
4. Preparedness 
5. Prediction 
6. Political 
7. Insurance 
8. Consumer 
9. Fina
Figure 1.3 Population vs. Sample 
 
 
 
 
 
 
Population is the collection of all elements under 
consideration in a statis
How do we use these observations to our advantage? Regardless of whether we are using 
data of population or sample, we hav
Parameter: proportion of all people who will buy the product 
Statistic: proportion of sample who will buy the product 
Varia
In this module, you will also learn how to organize and summarize data. Organizing and 
summarizing data is called descriptiv
In a statistical inquiry, the researcher identifies the problem, plans the study, collects the 
data, explore the data, ana

You might also like