0% found this document useful (0 votes)
6 views48 pages

Introduction to Statistics Concepts

For MAED

Uploaded by

Claire Estimada
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
6 views48 pages

Introduction to Statistics Concepts

For MAED

Uploaded by

Claire Estimada
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd

STATISTICS

DEMER G. PAGLOMUTAN
Central Philippines State University-
Cauayan Campus
Slides Prepared and Compiled by:
Engr. Sylvino v. Tupas
Phd. in mathematics education
University of St. La Salle-Bacolod
Learning Objectives

At the end of the course, students should be able


to:
1. define statistics in its simplest sense
including the terms and concepts;
2. familiarize with the four levels of data
measurements;
3. distinguish one data level from another;
4. Sampling Technique;
5. discuss normally distributed data; and
6. differentiate between parametric and
non-parametric statistics.
COVERAGE
 Statistics defined
 Types of Statistics
 Random Variables; Discrete and Continuous Data
 Level of Measurement
 Measures of Central Tendency:
Mean, Median, Mode, Mid-range
 Frequency Distribution
 Bar Graph and Histograms
 Measures of Dispersion: Variances; Standard Deviation
 Measures of Symmetry
 Parametric and Non-Parametric
Statistics defined
A science dealing with the
collection, presentation, analysis,
and interpretation of numerical
data.
Statistics is the science of learning from
data, and of measuring, controlling, and
communicating uncertainty; and it
thereby provides the navigation essential
for controlling the course of scientific and
societal advances (Davidian, M. and Louis, T.
A).
TERMS and CONCEPTS
ITEM COMPARE CONTRAST
Both describe
A parameter refers to a measure from a
a. Parameter characteristics of a
population, while a statistic refers to a
and Statistic dataset and are used in
inferential statistics. measure from a sample.
Mathematics is a broad field dealing
Both involve numerical
b. Statistics and with abstract concepts, while statistics
analysis and logical
Mathematics focuses on data collection, analysis,
reasoning.
interpretation, and inference.
A quantitative variable is numerical
c. Quantitative Both describe
(e.g., age, height), while a qualitative
and Qualitative characteristics of a
Variable dataset.
variable is categorical (e.g., gender,
color).
A population includes all individuals in
d. Population Both refer to groups in
a study, while a sample is a subset of
and Sample statistical studies.
the population used for analysis.
The independent variable is
e. Independent Both are variables used
manipulated to observe its effect, while
and Dependent in research to analyze
Variable relationships.
the dependent variable is the
outcome being measured.
A Type I error occurs when a true null
hypothesis is rejected (false positive),
f. Type I Error Both are errors in
TERMS and CONCEPTS
ITEM COMPARE CONTRAST
A survey collects data from a sample,
g. Survey and Both are data
while a census collects data from an
Census collection methods.
entire population.
Descriptive statistics summarize and
h. Descriptive Both are branches of
present data, while inferential
and Inferential statistics used to
statistics draw conclusions and make
Statistics analyze data.
predictions based on data.
A discrete variable takes specific,
i. Discrete and countable values (e.g., number of
Both are types of
Continuous students), while a continuous
quantitative variables.
Variable variable can take any value within a
range (e.g., height, weight).
j. Random Random error is unpredictable and
Error and Both are types of errors varies with each measurement, while
Systematic in measurement. systematic error is consistent and
Error caused by a flaw in measurement.
The significance level (α) is the
k. Significance
threshold for rejecting a null hypothesis,
Level and Both relate to
while the p-value measures the
Probability hypothesis testing.
strength of evidence against the null
TERMS and CONCEPTS
ITEM COMPARE CONTRAST
Probability sampling gives every
l. Probability
Both are methods of element a known chance of selection,
and Non-
selecting samples in while non-probability sampling
Probability
research. does not follow a random selection
Sampling
process.
m. Primary Primary data is collected firsthand
Data and Both are sources of for a specific study, while secondary
Secondary data in research. data is obtained from existing
Data sources.
The null hypothesis (H₀) assumes
n. Null and Both are statements no effect or relationship, while the
Alternative used in hypothesis alternative hypothesis (H₁)
Hypotheses testing. suggests the presence of an effect or
relationship.
A discrete variable consists of
o. Discrete
Both are types of distinct, countable values (e.g.,
and
numerical data used number of cars), while a continuous
Continuous
in statistical analysis. variable can take any value within a
Kinds of Statistics

(Enumerative
)
Kinds of Statistics
(Analytical)
Random Variables
Continuous Random Variable
Discrete Random Variables
Activity 1

 Identify at least ten random variables.


 State whether continuous or discrete.
 State its level of measurement
Levels of Measurement
Examples
 Number of white, black, and red cars in the parking
lot
 Result of the recently crown Miss Mass Kara 2015
 weight of under nourish children in Purok 8
 Entrance test scores
 amount of cash in my pocket
 level of understanding in math
 result of Cebu Triathlon 2015
 body temperature in degree Celsius
 temperature of methane gas in Kelvin
 gender
 willing to teach in the senior high
 number of students enrolled in the maritime strand
 extent of understanding about global warming
Sampling Technique
Sampling Classificati
Strength Weaknesses Sample Situation
Procedure on
Simple to implement Risk of periodic
a. Systematic Selecting every 10th customer
Probability and ensures even patterns affecting
Sampling entering a store for a survey
coverage randomness
Requires prior
Ensures Dividing students by grade level
b. Stratified knowledge of the
Probability representation of and randomly selecting from
Sampling population’s
different subgroups each
characteristics
Useful for large Conducting a nationwide survey
c. Multi-stage More complex and
Probability populations and by selecting regions, then cities,
Sampling time-consuming
complex studies then individuals
d. Simple Eliminates bias and
Can be impractical Drawing names from a hat to
Random Probability ensures equal chance
for large populations select participants
Sampling for all
Less precision Selecting entire barangays
e. Cluster Cost-effective for
Probability compared to randomly instead of individuals
Sampling large populations
stratified sampling for a health survey
f.
Non- Quick and easy to High risk of bias and Surveying people at a mall
Convenience
Probability implement not generalizable entrance for a marketing study
Sampling
Ensures data from May not represent
g. Expert Non- Interviewing economists about
knowledgeable the general
Sampling Probability inflation trends
sources population
h. Snowball Non- Useful for hard-to- High risk of bias and Studying drug users by asking
Sampling Probability reach populations lack of randomness participants to refer others
Focuses on specific
i. Purposive Non- Subjective selection Selecting experienced teachers
groups relevant to the
Sampling Probability may lead to bias to assess an education policy
study
Ensures May introduce bias Selecting an equal number of
TECHNIQU
SITUATION JUSTIFICATION
E/S USED
Entire municipalities are
For a survey, 5 samples of municipalities
selected as clusters, and
were selected from every province in the
all individuals in those
country and included 100 child laborers in
clusters are included in
the selected municipalities.
the sample.
The sample is chosen
based on the researcher’s
Familiar people to the researcher are only
accessibility to the
included in the sampling frame.
participants rather than
random selection.
Dr. Psych Ology, who is not part of the The researcher relies on
study, was asked to determine the an expert to select
possible participants of the study to the participants based on
researcher after observing the behavior of their judgment and
the people in the room. expertise.
To select a sample of households in a
province, a sample of provinces were
The sampling is done in
selected, then a sample of municipalities
multiple stages, selecting
were chosen from each of the selected
units at different levels
provinces, then a sample of barangays
(province → municipality
were chosen from each of the selected
→ barangay → household).
municipalities, and all households in the
selected barangays were included.
The sample is deliberately
Clinically proven people with hypertension
chosen based on specific
are being considered in the study that
characteristics
TECHNIQU
SITUATION JUSTIFICATION
E/S USED
The selection is based
Top highest vegetable-producing on specific criteria, in
household per Barangay in Canlaon this case, being the top
City are only included in the study. vegetable-producing
household.
Each ball has an equal
In the game of lotto, 6 balls are chance of being
selected from a container with 42 selected, making it a
balls. purely random
selection.
One participant refers
Mr. X is included in the study because
another, which is
the participant Y gave details to the
commonly used for
researcher that he can be part of the
hard-to-reach
study.
populations.
The population is
A survey obtained a sample of
divided into two strata
laborers by first classifying the
(rural and urban), and
different areas as either rural or
then a random sample
urban area. After which, a sample of
is taken from each
laborers is taken from each area.
stratum.
Every 20th unit is
A car manufacturer conducts quality selected in a
Measures of Central
Tendency
A measure of central tendency (also
referred to as measures of centre or central
location) is a summary measure that
attempts to describe a whole set of data
with a single value that represents the
middle or centre of its distribution.
Mean
Properties of the Mean
Median

The median is the value of the


middle term in a data set that
has been arranged in
increasing order.
Properties of Median

1. The median is less sensitive than


the mean to the presence of a
few extreme scores.

2. In a distribution that are strongly


asymmetrical, the median may
be the better choice for the
measure of the central tendency
of the set of data.
Mode

The mode is the item that is most


popular or common. A distribution
may also be bimodal and multi
modal.

The mode is easy to obtain but is not


very stable from sample to sample.
Mid-Range

Another measure of center that is


not as popular as the mode, mean
and median is called the midrange.

The midrange is the mean of the


maximum and minimum values of
the data set.
Example
Frequency Distribution
Arrangement of gathered data by
categories plus their corresponding
frequencies and class marks or midpoints
(Punsalan, 1989).

Tabular arrangement of data whereby the


data is grouped into different interval and
the number of observations that belongs
to each interval is determined.
Table 1. Distribution of the Respondents

Gender Frequency
Male 23
Female 107
Total 130

Class
Relative Frequency
Frequency = Total
Frequency
Table 2. Preferred Color of Respondents

Color Frequency Relative Percentage


Frequency Frequency
Yellow 23 0.3286 32.86%

Green 7 0.1000 10.00%

Blue 14 0.2000 20.00%

Red 26 0.3714 37.14%

Total 70 1.0000 100.00%


Bar Graph showing the Frequency
Distribution Table
30

25

20

15

10

0
Yellow Green Blue Red
Steps in Constructing a Frequency
Table (Group Data)

 Step 1: Arrange the values (scores) from


H-L/L-H

 Step 2: Identify the Range (Highest – Lowest


value)

 Step 3: Decide on the number of class interval


[equal] (generally between 5 to 15 intervals
depending on the range)

 Step 4: Decide on the values of the first class


limits (make sure that the lowest or highest value
Steps in Constructing a Frequency
Table (Group Data)

 Step 5: Make a tally of the values. Ensure that all the


entries are tallied at various interval.

 Step 6: Under column f, identify the frequency counts in


each class interval.

 Step 7: Compute the cumulative frequencies (CF< and


CF >)

 Step 8: Compute the Relative Frequencies (RF)


Activity
 Draw a Frequency Distribution Table and a
Histogram.
27 95 63 69 43 75 31 50 74
82 79 65 52 60 49 80 50 65
83 58 47 60 70 59 61 39 71
66 59 50 40 58 63 52 60 59
59 61 58 70 63 78 59 53 49
40 31 70 63 59 49 61 72 68
66 55 61 49 53 65 50 89 75
Frequency Distribution Table
Class Interval Tallies f CF < CF > RF %
21 - 30 \ 1 1 63 1.59
31 - 40 \\\\\ 5 6 62 7.94
41 - 50 \\\\\ \\\\\ 10 16 57 15.87
51 - 60 \\\\\ \\\\\ \\\\\ \\ 17 33 47 26.98
61 - 70 \\\\\ \\\\\ \\\\\ \\\ 18 51 30 28.57
71 - 80 \\\\\ \\\ 8 59 12 12.7
81 - 90 \\\ 3 62 4 4.76
91 - 100 \ 1 63 1 1.59
63 100
Histogram (continuous data)
20
18
16
14
12
10
8
6
4
2
0 1

21 - 30 31 - 40 41 - 50 51 - 60 61 - 70 71 - 80 81 - 90 91 - 100
Activity 2:
 Draw a Frequency Distribution Table and a
Histogram.
22 97 63 61 43 75 31 55 74
82 77 65 52 60 49 80 50 65
83 58 47 60 83 59 61 39 71
66 59 50 40 58 63 52 60 59
59 61 58 75 63 78 59 53 49
40 35 70 63 59 49 63 72 68
66 55 61 49 53 65 50 89 75
Variations
Measure of variation is a measure that
describes how spread out or scattered a
set of data. It is also known as measures
of dispersion or measures of spread.

There are 2 measures of variation:


Range and Standard Deviation
Range

- simplest of all the measures of


variation
- difference between the highest
and lowest value in a set of
data
- symbol R
Standard Deviation
- the most popular measures of variation
- square root of the variance
- denoted by SD
The standard deviation measures how
concentrated the data are around the
mean; the more concentrated, the smaller
the standard deviation.

•The standard deviation can never be a negative


number
•The smallest possible value for the standard
deviation is 0, (every single number in the data
set is exactly the same).
•The standard deviation is affected by outliers
(extremely low or extremely high numbers in the
data set).
•The standard deviation has the same units as the
Chebyshev's Theorem
 developed by Russian mathematician
 at least ¾ of the data falls within 2 SD
of the
mean
 at least 8/9 or 88.89% of the data falls
within
3 SD of the mean
 This theorem can be applied to any
distribution regardless of its shape
The Empirical Rule
Chebyshev's theorem as applied to a normal
(bell-shaped) distribution

the empirical rule states that...
- approximately 68% of the data values falls
within 1 SD of the mean
- approximately 95% of the data values falls
within 2 SD of the mean
- approximately 99.7% falls within 3 SD
Normal Distribution Curve
Parametric and Non-
Parametric
 Parametric tests are those that
make assumptions about the
parameters of the population
distribution from which the sample is
drawn. This is often the assumption
that the population data are normally
distributed. Non-parametric tests
are “distribution-free” and, as such,
can be used for non-Normal variables.
Parametric Non- REMARKS
test Parametric
test
Mean Median, Mode Descriptive

Paired t-test Wilcoxon Rank Inferential Statistics


Sum test Dependent Samples

Unpaired t-test Mann-Whitney U Inferential Statistics


test Independent Sample

Pearson Spearman Relationship/


Correlation Correlation Correlation
One-way ANOVA Kruskal Wallis Inferential Statistics
Test 3 or more groups
Thank
You!

You might also like