Data Management (Module 5)
Measure of Central Tendency of Ungrouped Data
- is the summary measure that describe a whole set of data with a single quantity that
is represented by the middle or center of its distribution the way of which group of
data that cluster around a central value in short this is the measure that tells where
the center of a data set is located.
MEAN
- The mean is the most commonly used measure of central tendency. When we speak
of average, we always refer to the mean.
MEDIAN
- The median is the midpoint of the data array. Before finding this value, the data must
be arranged in order, from least to greatest or vice versa. The median will either be a
specific value or will fall between two values
MODE
- It is the value that occurs most often in the data set.
- The number/value/observation in a data set which appears the MOST number of
times.
THE WEIGHTED MEAN
FREQUENCY DISTRIBUTION
- A frequency distribution, which is a table that list observed events and the
frequency occurrence of each observed event, is often used to organized raw data.
Measures of Central Tendency (Grouped Data)
MEAN
where:
fx = the product of frequency and
class mark.
N = total frequencies
MEDIAN
where:
∑f = total frequencies.
lb_mc = lower boundaries of median class
f_mc = frequency of median class
cw = class width
cf = cumulative frequency before/preceding the
median class
MODE where:
lb mo = lower boundaries of modal class
D_{1} = difference of the modal class and the class
preceding it.
D_{2} = difference of the modal class and the class
succeeding it.
Measures of Dispersion (Ungrouped Data)
MEASURES OF DISPERSION
- A measure of variability of a set of data is a number that conveys the idea of spread
for the data set.
-Range
-Standard Deviation
-Variation
Range
- The range measures the distance between the largest and the smallest values and,
as such, gives an idea of the spread of the data set. However, the range does not use
the concept of deviation. It is affected by outliers but does not consider all values in
the data set. Thus, it is a not a very useful measure of variability.
Variance
- The variance for a given data set is the square of the standard deviation of the data.
THE STANDARD DEVIATION
- The Standard Deviation is a measure of how spread out numbers are.
- Its symbol is (the greek letter sigma).
Measures of Dispersion (Grouped Data)
-RANGE
Scores of 40 Students in a 60-point Quiz
VARIANCE AND STANDARD DEVIATION
Measures of Relative Position (Z-scores)
Z – score
- A z score measures the distance between an observation and the mean, measured
in units of standard deviation.
- If the z score is positive, the positive score is above the mean. If the z score is 0, the
score is the same as the mean. If the z score is negative, the score is below the
mean.
Normal Distribution (Area Under the Curve)
Normal Distribution
- the shape and position of the normal distribution curve depends on two
parameters. The mean and the standard deviation. Each normally distributed
variable has its own normal distribution curve. which depends on the values of the
variables mean and standard deviation. So, the larger the standard deviation, the
more dispersed or spread out the distribution becomes.
Properties of a Normal Distribution
• The distribution curve is bell-shaped
• The curve is symmetrical about its center
• The mean, median, and mode coincide at the center
• The tails of the curve flatten out indefinitely along the horizontal axis but never touch
it (the curve is asymptotic to the base line)
• The area under the curve is 1. thus, it represents the probability or proportion or the
percentage associated with specific sets of measurement values
Empirical Rule for a Normal distribution
In a normal distribution, approximately
• 68% of the data lie within 1 standard deviation of the
mean.
• 95% of the data lie within 2 standard deviations of the
mean.
• 99.7% of the data lie within 3 standard deviations of the
mean.
Four-Step Process
in Finding the Areas Under the Normal Curve given a z-value
1. Express the given z-value into a 3-digit form
2. Using the z-table, find the first two digits on the left column
3. Match the third digit with the appropriate column on the right
4. Read the area (or probability) at the intersection of the row and column.