0% found this document useful (0 votes)
11 views97 pages

Data Management and Statistics Overview

The document provides an overview of statistics, including its types (descriptive and inferential), variables (qualitative and quantitative), and scales of measurement (nominal, ordinal, interval, and ratio). It also discusses the organization of data, measures of central tendency (mean, median, mode), measures of variability (range, variance, standard deviation), and measures of relative position (percentiles and deciles). Overall, it serves as a foundational guide to understanding statistical concepts and methods.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
11 views97 pages

Data Management and Statistics Overview

The document provides an overview of statistics, including its types (descriptive and inferential), variables (qualitative and quantitative), and scales of measurement (nominal, ordinal, interval, and ratio). It also discusses the organization of data, measures of central tendency (mean, median, mode), measures of variability (range, variance, standard deviation), and measures of relative position (percentiles and deciles). Overall, it serves as a foundational guide to understanding statistical concepts and methods.
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

DATA

MANAGEMENT
STATISTICS
STATISTICS is the science of collecting, organizing
and summarizing recorded information or data in
such a way that a valid conclusion and meaningful
predictions can be drawn from them.
TYPES OF STATISTICS
1. Descriptive statistics is consists of methods concerned
with the collection, description and analysis of data
without drawing conclusions or inferences about a larger
set. Its main concern is simply to describe the set of data
such that otherwise obscure information is brought out
clearly.
2. Inferential statistics utilizes sample data to make
estimates, decisions, predictions, or other generalizations
about a larger set of data.
VARIABLES
In statistics, a VARIABLE refers to a specific
characteristic (or attribute) of a subject. Such an
attribute may assume two or more different
values. For example, the “sex” of a person is
variable; its value is either “male” or “female”.
Other examples of variables are your course,
citizenship, age, height and weight.
TYPES OF VARIABLES

1. Qualitative variables are those whose values are


measured not in terms of numbers, but categorically by
means of depression. Examples are “course”,
“citizenship”, “favorite color” and “place of birth”.

2. Quantitative variables are those that are always


associated with numbers or a scale measure. Examples
are “age”, “height”, “weight” and “population”.
SCALES OF MEASUREMENTS

1. Nominal – characterized by data that consists of names,


labels, codes or categories only. These data cannot be
arranged in an ordering scheme and cannot be used for
calculations. Examples are gender, citizenship, religion,
house number, plate number, ID number, and zip code.
SCALES OF MEASUREMENTS

2. Ordinal – it involves data that may be arranged in some


order. Examples are sizes (small, medium, large), socio-
economic class (working, middle, upper), educational
attainment, and the Likert scale (strongly disagree, disagree,
neutral, agree, strongly agree).
SCALES OF MEASUREMENTS

3. Interval – measurements where the difference between


values is meaningful, but ratio is not. In interval scale, zero
does not mean “nothing”. Examples are temperature in
Celsius, pH level, and IQ.
SCALES OF MEASUREMENTS

4. Ratio – measurements are ordered according to the


amount of attribute they possess. Equal differences in the
attribute are represented by equal differences in the
numbers assigned. In ratio, zero means absence of
something. Temperature in Celsius or Fahrenheit is not a
ratio scale because 0⁰C or 0⁰F does not mean the absence of
temperature; while temperature in Kelvin is an example of a
ratio scale since 0⁰K means an absence of heat. Examples are
height, weight, age, and cellphone load.
POPULATION VERSUS SAMPLE
In statistics, a population refers to the entire set of all
objects under study; while a sample refers to any subset of
the population.

EXAMPLE:
A student researcher wants to do a survey among CLSU
students. Instead of doing a survey of all the students in CLSU,
he just chose and surveyed a group of 45 students (five students
per college). In this scenario, the population is all the students
of CLSU, while the sample is the group of 45 students.
ORGANIZING DATA
ORGANIZING DATA

Phase I of organizing data is data collection, where each


element of the data is called a data point. Generally, in this
phase, the raw data may not show any apparent pattern or
trend. Thus, the raw data, suggests nothing but just
numbers. But if we organize the data (Phase II), they
become more meaningful.
ORGANIZING DATA

Frequency Distribution Table - The most common way of


organizing data is using a frequency distribution table or FDT.
It utilizes a table that lists all data points, along with how
many times the data point occurs (frequency, f), and its
percentage of the total number of data (Relative frequency,
𝒇
RF, %𝑹𝑭 = × 𝟏𝟎𝟎)
𝒏
ORGANIZING DATA

Ungrouped data refers to raw data that have not been


organized into groups, classes, or intervals. In other words, it
is data presented as individual values; just as they were
collected.
ORGANIZING DATA

Example: (ungrouped)
The following data are the respective number of kids of 50
families.
ORGANIZING DATA

Example: (ungrouped)
ORGANIZING DATA
Grouped data refers to data that have been organized into
classes or intervals to make them easier to understand and
analyze; especially when there are many observations.

Phase II. We organize the raw data into a frequency distribution.


First, we must decide on how many groups to use. Customarily,
the number of groups is any number from 4 to 8.

ℎ𝑖𝑔ℎ𝑒𝑠𝑡 𝑝𝑜𝑖𝑛𝑡 − 𝑙𝑜𝑤𝑒𝑠𝑡 𝑝𝑜𝑖𝑛𝑡


𝐶𝑙𝑎𝑠𝑠 𝐼𝑛𝑡𝑒𝑟𝑣𝑎𝑙 = 𝑐 =
𝑛𝑢𝑛𝑏𝑒𝑟 𝑜𝑓 𝑔𝑟𝑜𝑢𝑝𝑠
ORGANIZING DATA
Example: (grouped)
The following are examination scores of 42 mathematics
students.

ℎ𝑖𝑔ℎ𝑒𝑠𝑡 𝑝𝑜𝑖𝑛𝑡 − 𝑙𝑜𝑤𝑒𝑠𝑡 𝑝𝑜𝑖𝑛𝑡 62 − 16


𝑐= = ≈ 7.66
𝑛𝑢𝑛𝑏𝑒𝑟 𝑜𝑓 𝑔𝑟𝑜𝑢𝑝𝑠 6
ORGANIZING DATA
Example: (grouped)
HISTOGRAM
&
PIE CHART
HISTOGRAM
Example:
PIE CHART
Example:
MEASURES
OF
CENTRAL
TENDENCY
Measure of central tendency is a value that indicates where
the center of distribution tends to be located, or simply the
average of the data. It is said to form the basis of statistics.
The most common measures of central tendency are the:
mean, median, and mode. On a perfect normal distribution,
all three measures of central tendency are located at the
same score, which is at the center of the normal distribution.
MEAN
The mean is the most commonly used measure of central
tendency. The mean of a data set is the sum of the data points
divided by the number of data points, or simply the average of
the data points. Thus, it is strongly influenced by outliers (data
points that are extremely low or extremely high compared to
other data points). The population mean, denoted by 𝜇, is
estimated by the sample mean denoted by 𝑥ҧ .
SOME CHARACTERISTICS OF THE MEAN ARE THE FOLLOWING:
1. The sum of deviations of the data points from the mean is zero.
2. The sum of the squared deviations of the data points is minimum when
the deviations are taken from the mean.
3. If a constant is added (or subtracted) to every data point, the new mean
is the original mean increase (or decrease) by the said constant.
4. If every data point is multiplied (or divided) by a constant , the new
mean is the original mean multiplied (or divided) by .
5. Since the mean is a calculated number, it may not be an actual value in
the data points.
MEAN
EXAMPLE:
The data below are the current diesel prices (in pesos/liter) in
nearby gas stations, find the mean price.
MEAN
EXAMPLE:
Gabriel has a total of 4 quizzes. One quiz is missing while the scores of his
remaining quizzes are 43, 35 and 39. Calculate the score of the missing quiz
if his mean score is 41.
MEAN (Grouped Data)
MEAN (Grouped Data)
MEAN(Grouped Data)
EXAMPLE:
Find the mean score of 42 students from the following frequency distribution:
MEAN(Grouped Data)
MEDIAN
The median is a value that separates an array of data points
into two equal parts. To find it, the data need first to be
arranged in numerical order. If there is an odd number of data
points, then the median is the middle value. If there is an
even number of values in the data set, then the median is the
average of the two middle values. The median can be denoted
by 𝑀𝑑 or 𝒙෥.
Unlike the mean, median is not affected by extreme values in
data points because it only considers the middle values in the
data set
MEDIAN
EXAMPLE:
Calculate the median age of the seven employees.
MEDIAN
EXAMPLE:
: The current crude oil prices (in pesos/liter) in nearby gas stations are listed
below. Find the median price.
MODE
The mode of a data set is the data point that occurs most
often. If no data point is repeated or every data point is
repeated the same number of times, there is no mode. If the
mode of a data set exists, it may not be unique. A unimodal
data set has one mode, bimodal has two modes, trimodal has
three modes and multimodal has many modes. The mode can
be used for qualitative as well as quantitative data. Mode is
not affected by the extreme values in the data set, since it
only considers the most frequent data. Mode can be denoted
by 𝑀𝑜 or 𝒙ෝ.
MODE
EXAMPLE:
Find the mode of the following data set;
MODE
EXAMPLE:
Thirty students are asked about their favorite color. The data is summarized
by the frequency distribution table below. Find the mode.
MEASURES
OF
VARIABILITY
A measure of variability (or dispersion) is a quantity that
measures the spread of scores in a given population. It
indicates the extent to which observations in a data set are
scattered about the mean. Scores that are relatively close
together have a lower variation as compared to scores that
are spread farther apart. To measure the spread or
dispersion of data, we use statistical values known as the
range, variance and standard deviation, these three
statistical values are the most common measures of
variability.
Suppose that we are choosing between Jerico and Jerwin on who should represent
CLSU to an upcoming Inter-University Math Quiz Bee. To choose, their coach
conducted 6 sessions of quiz-alikes between them, and came up with the following
scores:

Observe that:

Another measure that could help is to look at their consistency. This is about the
measure of variability that is to look at how spread apart or dispersed their scores are.
RANGE

The range, denoted by R, is the difference between the lowest


and the highest values in a data set. A weakness of the range
is that an extreme value (outlier) can greatly alter its value.
RANGE

EXAMPLE:
VARIANCE AND STANDARD DEVIATION

First, we define deviation to be 𝑥 − 𝑥ഥ where 𝑥 is a data point


and 𝒙ഥ is the mean. It is the difference of a data point from the
mean.
VARIANCE AND STANDARD DEVIATION

Variance is the mean of the squared deviation of the data


points. The sample variance (denoted by 𝑠 2 ) is an estimator of
the population variance (denoted by 𝜎 2 ). In symbols, sample
variance of data points 𝑥1 , 𝑥2 , … , 𝑥𝑛 where 𝑛 is the number of
data points is defined as
VARIANCE AND STANDARD DEVIATION

Alternatively, the variance may be computed relatively quicker


and easier by the equivalent formula below. We don’t need
the mean in using this formula.
VARIANCE AND STANDARD DEVIATION

Standard deviation is defined as the square root of the


variance and is denoted by 𝑠 (for sample) or 𝜎 (for
population). Thus,
VARIANCE AND STANDARD DEVIATION

EXAMPLE:
VARIANCE AND STANDARD DEVIATION

EXAMPLE:
MEASURES
OF
RELATIVE POSITION
As earlier discussed, the measures of central tendency
especially the mean and the median describe the “center” of
a distribution. Indeed, such a center is what is usually used
and needed to summarize a distribution. Occasionally
however, a different part of the distribution is of more
interest. The percentile, decile, and quartile are used in
such occasions, as they indicate the location of a data point
relative to the other data points.
PERCENTILE
Percentiles split the whole distribution into 100 subgroups. It is similar to
cutting a long pipe into 100 short pipes of equal lengths. In order to do this,
it is necessary to make 99 cuts. The points where the cuts are done
correspond to percentile ranks or scores. Thus, percentile ranks are from 1
to 99, which we hereby denote by 𝑃1 , 𝑃2 , 𝑃3 , . . , 𝑃99 . There is no sense to
have a 𝑃0 , nor a 𝑃100 .
A percentile is a value that describes the percentage of data that falls
below it. For example, suppose you got a 99-percentile score in an exam. It
means that 99% of the examinees scored lower than you; it doesn’t mean
that you had a score of 99%. In fact, your actual score is not at all indicated.
PERCENTILE
EXAMPLE: Suppose that Sonny is among the 15,000 high school graduates who took
the CLSU Admission Test, and he got a 48 percentile score. His 48 percentile score
means that 48% of the 15,000 examinees (7,200) scored lower than Sonny. It doesn’t
mean that his actual score in the exam is 48. On the other hand, his actual score is lower
than 52% of the 15,000 examinees(7,800).

Suppose another student Nick got a percentile score of 68. This means that 68% of the
15,000 examinees (10,200) scored lower than Nick while 32%or 4,800 examinees scored
higher than him.

The actual scores of Sonny and Nick both remain unknown, until we do some
calculations that also involve the whole distribution of data points, theirpercentile
scores, and the number of data points.
PERCENTILE
PERCENTILE
EXAMPLE:
PERCENTILE
EXAMPLE:
PERCENTILE
EXAMPLE:
PERCENTILE
EXAMPLE:
PERCENTILE
EXAMPLE:
PERCENTILE
EXAMPLE:
PERCENTILE
EXAMPLE:
PERCENTILE
EXAMPLE:
PERCENTILE
EXAMPLE:
DECILE
Deciles split the whole distribution into 10 subgroups. It is similar to cutting
along pipe into 10 shorter pipes of equal lengths. In order to do this, it is
necessary to make 9 cuts. The points where the cuts are done correspond
to decile ranks or scores. Thus, decile ranks are from 1 to 9 denoted by
𝐷1 , 𝐷2 , 𝐷3 , 𝑢𝑝 𝑡𝑜 𝐷9 . It makes no sense to talk about 𝐷0 𝑛𝑜𝑟 𝐷10 .
DECILE
EXAMPLE:

Find the 9th decile or D9 .


DECILE
EXAMPLE:
Solution
• 𝐷9 = 𝑃90
DECILE
EXAMPLE:
Solution
• 𝐷9 = 𝑃90
QUARTILE
Quartiles split the whole distribution into 4 subgroups. It is similar to
cutting along pipe into 4 shorter pipes of equal lengths. In order to do this,
it is necessary to make 3 cuts. The points where the cuts are done
correspond to quartile ranks or scores. Thus, percentile ranks are from 1 to
3 denoted by 𝑄1 , 𝑄2 , 𝑎𝑛𝑑 𝑄3 . It makes no sense to talk about 𝑄0 𝑛𝑜𝑟 𝑄4 .
DECILE
EXAMPLE:
DECILE
EXAMPLE:
NORMAL
DISTRIBUTION
The normal distribution or the Gaussian distribution (in honor
of Gauss, 1777-1835) is the most important distribution in
statistics. Statisticians created an ideal bell-shaped curve (also
called normal curve) to describe such a normally distributed
data. The normal curve is symmetric about a vertical axis
through the mean, with a total are under the curve equal to 1
and the curve is asymptomatic to the x-axis.
STANDARD
NORMAL
DISTRIBUTION
A Problems
B Problems
END OF MATH1100

Thank you for being such


awesome students!

You might also like