MMW Statistics: Frequency Distribution
MMW Statistics: Frequency Distribution
MODULE 4
II. OBJECTIVE(S):
III. INTRODUCTION:
Statistics (in the singular sense) is a scientific discipline that deals with the methods and theories
in the manipulation of numerical data. It leads to the analysis and interpretation of the data set so
one can make a sound decision and thorough inferences.
Statistics (in the plural sense) are numerical data. Some examples are revenues, allowed kilograms
for check in luggage, stipend, tuition fee, ID number, military ranks, etc.
IV. DISCUSSION:
DATA MANAGEMENT
Data Management deals with the collection, organization and presentation of the numerical data
or (statistics) in a presentable and usable manner.
Example:
The following are the responses of fifteen students when interviewed on the number of times they
open their chatroom in a day. Create a frequency distribution table.
Students A B C D E F G H I
# of Times 22 23 13 11 25 11 23 17 22
46 57 59 64 56 50 70 62 68 79 63 54
60 51 58 37 68 35 50 74 39 75 67 69
40 52 65 45 59 70 73 84 54 42 44 63
40 63 64 45 70 41 56 49 64 76 80 78
58 54 65 62 55 55 52 81 83 84 85 53
2. Desired number of class interval (dci): 𝑟𝑟𝑟 = 8 (Note that dci is arbitrarily chosen.)
0 45
3. Class size (or class width): 𝑟 = 123 = = 6.25, 𝑟𝑟𝑟𝑟𝑟 𝑟𝑟 𝑟𝑟 7. Since the data set are
6
whole numbers, the class size should be a whole number as well.
SUMMARY MEASURES:
V. SUMMARY
Statistics (in the singular sense) is a scientific discipline that deals with the methods and theories
in the manipulation of numerical data. It leads to the analysis and interpretation of the data set so
one can make a sound decision and thorough inferences.
Statistics (in the plural sense) are numerical data. Some examples are revenues, allowed kilograms
for check in luggage, stipend, tuition fee, ID number, military ranks, etc.
Data Management deals with the collection, organization and presentation of the numerical data or
(statistics) in a presentable and usable manner.
[Link]
[Link]
[Link]
VII. REFERENCE
[Link]
[Link]
Mathematics In The Modern World – Adamson University Textbook
MATHEMATICS IN THE MODERN WORLD
MODULE 4.1
II. OBJECTIVE(S):
III. INTRODUCTION:
Statistics (in the singular sense) is a scientific discipline that deals with the methods and theories
in the manipulation of numerical data. It leads to the analysis and interpretation of the data set so
one can make a sound decision and thorough inferences.
Statistics (in the plural sense) are numerical data. Some examples are revenues, allowed kilograms
for check in luggage, stipend, tuition fee, ID number, military ranks, etc.
IV. DISCUSSION:
A. MEAN
Most common measure of the center. It is also known as arithmetic average. It is the summation
of the data (x) divided by the total number of population.
Formula:
Now, why are we going to get the mean? What are the properties of this:
1. It may not be an actual observation in the data set
2. Can be applied in at least interval level
3. Easy to compute
4. Every observation contributes to the value of the mean
B. MEDIAN
Divides the observations in two equal parts.
a. If the number of observations is odd, the median is the meddle number.
b. If the number of observation is even, the median is the average of the two middle numbers.
Properties of median:
1. May not be an actual observation in the data set
2. Can be applied in at least ordinal level
3. A positional measure; not affect by the extreme values
C. MODE
Mode occurs most frequently in the data set. It is a nominal average, it may be exist or not exist.
V. SUMMARY
A measure of central tendency (also referred to as measures of center or central location) is a summary
measure that attempts to describe a whole set of data with a single value that represents the middle or center
of its distribution. The mode is the most commonly occurring value in a distribution. The mode has an
advantage over the median and the mean as it can be found for both numerical and categorical (non-
numerical) data. The median is the middle value in distribution when the values are arranged in ascending or
descending order. The median is less affected by outliers and skewed data than the mean, and is usually the
preferred measure of central tendency when the distribution is not symmetrical. The mean is the sum of the
value of each observation in a dataset divided by the number of observations. This is also known as the
arithmetic average. The mean can be used for both continuous and discrete numeric data.
VI. REFERENCE
[Link]
[Link] pdf
[Link]
central%[Link]
Mathematics in the Modern World – Adamson University Textbook
MATHEMATICS IN THE MODERN WORLD
MODULE 4.3
II. OBJECTIVE(S):
III. INTRODUCTION:
Statistics (in the singular sense) is a scientific discipline that deals with the methods and theories
in the manipulation of numerical data. It leads to the analysis and interpretation of the data set so
one can make a sound decision and thorough inferences.
Statistics (in the plural sense) are numerical data. Some examples are revenues, allowed kilograms
for check in luggage, stipend, tuition fee, ID number, military ranks, etc.
IV. DISCUSSION:
MEASURE OF VARIATION
A measure of variation is a single value that is used to describe the spread of the distribution. A
measure of central tendency alone does not uniquely describe a distribution.
There are two types of measure of variation; (1) Absolute measures of dispersion and (2) Relative
measure of dispersion.
Absolute measures of dispersion consist of Range, Inter-quartile Range, Variance and Standard
Deviation.
Relative Measure of Variation consist only of coefficient of variation.
A. RANGE
Range is the difference between the maximum and the minimum value in a data set.
R = MAX – MIN
Example:
Pulse rates of 15 male residents of a village
54 58 58 60 62 65 66 71 74 75 78 80 85
Range = 85 – 54 = 31
So, the range is 31.
Properties of range:
1. The karger the value of the range, the more dispersed the observations are.
2. It is quick and easy to understand
3. A rough measure of dispersion
B. INTERQUARTILE RANGE
The difference between the third quartile and the first quartile.
IQR = Q3 – Q1
Properties of the interquartile range:
1. Reduces the influence of extreme values
2. Not as easy to calculate as the range
Example:
Seventy percent of the expemses are higher that 43,000php but only 25% are below it.
Twenty five percent of the expenses are higher that 59,000php but 75% are below it.
Therefore, IQR = 59 – 43 = 14
This means that the middle 50% of the housewives’ expemses has a deviation of 14,000php.
C. VARIANCE
Variance is important measure of variance. It shows variation about the mean
Formula:
Population Variance:
2
( X X )2
N
Sample Variance:
( X X )2
s
2
N 1
D. STANDARD DEVIATION
Most important measure of variation. It is the squareroot of variance. It has the same units as the
original data.
Formula:
( X X )2
2 N
s
s2 ( X X )2
N1
Example:
Consider the following data:
10 12 14 15 17 18 18 24
N=8
Mean = 16
S= 4.309
E. COEFFICIENT OF VARIATION
Measure of relative variaktion. Usually expressed in percent. It shows variation relative to the
mean and used to compare 2 or more groups.
Formula:
𝑆𝐷
𝐶𝑉 = ( ) 𝑋 100%
𝑀𝐸𝐴𝑁
Example:
The data below are the number of latecomers in a week from the three sections in the college if
Liberal Arts. Which section has the highest variability?
Section 1: 5, 4, 2, 1, 3, 1, 2
Section 2: 1, 0, 2, 1, 3, 1, 2
Section 3: 2, 1, 2, 1, 3, 1, 2
The most dispersed section is section 2, since it has the highest variability with a CV of 68.30%.
Section 3 has the least variability with a CV of 44.34%.
NORMAL DISTRIBUTION
Normal distribution is also known as Gaussian distribution, after the mathematician and
astronomer Karl Gauss. It is a continuous distribution which is regarded by many as the most
significant probability distribution in the entire theory of statistics, particularly in the field of
statistical inference.
It is a graphically represented by a symmetrical, bell shaped curve known as the normal curve.
Solution:
a. above 120?
b.
𝑥−𝜇 120 − 100 20
𝑧= = = = 1.33
𝜎 15 15
𝑃 (𝑧 > 1.33) = 0.5 − 0.4082 = 𝟎. 𝟎𝟗𝟏𝟖 = 9.18%
b. below 128?
𝑃(𝑥 < 128)
𝑥−𝜇 128 − 100 28
𝑧= = = = 1.87
𝜎 15 15
c. below 93?
𝑃(𝑥 < 93)
𝑥−𝜇 93 − 100 −7
𝑧= = = = −0.47
𝜎 15 15
𝑥1 − 𝜇 98 − 100
𝑧1= = = −0.13
𝜎 15
𝑥2 − 𝜇 105 − 100
𝑧 2= = = 0.33
𝜎 15
Example: The manager of an art gallery wants to determine the relationship between the auction
of price of paintings, y, and the number of bidders, x. From the data,
a. Determine the regression model
b. Find the estimated price of a painting if there are 20 bidders
c. Find the estimated number of bidders if the price is P13k.
𝑦̂ = 𝑎 + 𝑏𝑥 = 14.3874 − 0.3802𝑥
b. Find the estimated price of a painting if there are 20 bidders
𝑦̂ = 𝑎 + 𝑏𝑥 = 14.3874 − 0.3802(20) = 𝑃𝐻𝑃6.7835
V. SUMMARY
A measure of variability is a summary statistic that represents the amount of dispersion in a
dataset. In statistics, variability, dispersion, and spread are synonyms that denote the width of the
distribution. A range is one of the most basic measures of variation. It is the difference between
the smallest data item in the set and the largest. Quartiles divide your data into quarters: the
lowest 25%, the next lowest 25%, the second highest 25% and the highest 25%. The interquartile
range is one of the most popular measures of variation used in statistics. It is a measure of how
data is spread around the mean. The basic formula is: IQR = Q3 – Q1. Variance tells you how far
a data set is spread out, but it is an abstract number that really is only useful for calculating
the Standard Deviation.
Normal Distribution is a continuous distribution which is regarded by many as the most
significant probability distribution in the entire theory of statistics, particularly in the field of
statistical inference.
Regression determines if the independent variable 𝑥 and the dependent variable 𝑦 show a
positive or negative relationship.
VI. REFERENCES
[Link]
[Link]
[Link] [Link]
ioana/statistics/7.%20Measures%20of%[Link]
[Link]
[Link]
[Link]
MATHEMATICS IN THE MODERN WORLD
TOPIC: STATISTICS
(FREQUENCY DISTRIBUTION,
RELATIVE FREQUENCY)
Learning Outcome:
1. Make a frequency table for a set of data
2. Create a frequency distribution for a data set
3. Understand the relative frequency distribution table
Statistics is a scientific discipline that deals with the methods and theories in
the manipulation of numerical data. It leads to the analysis and interpretation
of the data set so one can make a sound decision and thorough inferences.
Simple linear regression interprets the relationship between two variables by representing it with a straight line, typically described by the equation y = a + bx, where 'y' is the dependent variable, 'x' is the independent variable, 'a' is the y-intercept, and 'b' is the slope. The slope 'b' indicates the direction and strength of the relationship, quantifying how much 'y' changes with a one-unit change in 'x'. Regression analysis thus helps in predicting 'y' based on 'x' and evaluating the significance of the relationship .
The range provides a quick and simple measure of data dispersion by indicating the difference between the maximum and minimum values. However, its limitations include sensitivity to outliers, as extreme values can significantly skew the perception of variability. It also does not reflect the distribution of data within the interval, offering no insight into the spread of most data points, which can lead to misleading interpretations in irregular datasets .
Normal distribution is significant because it is a fundamental probability distribution used extensively in statistical inference. Its properties, such as the symmetrical bell-shaped curve around the mean, median, and mode, allow for the application of a range of statistical tools and methods. The characteristics of the normal curve facilitate predictions and hypothesis testing, contributing to its significance in various fields of research and data analysis .
The interquartile range (IQR) is calculated by subtracting the first quartile (Q1) from the third quartile (Q3). The IQR is significant because it measures the range within which the middle 50% of the data lie, effectively reducing the influence of outliers and extreme values. This makes it more robust compared to the range, providing a clearer understanding of the core spread of the dataset .
The coefficient of variation (CV) is a relative measure of dispersion, expressed as a percentage of the mean. It allows for the comparison of variation between datasets with different units or means, offering a dimensionless measure of variability. In contrast, absolute measures of dispersion like range, interquartile range, and standard deviation use the same units as the data. While absolute measures give a direct sense of spread, they do not allow for direct comparison across different data sets without normalization, unlike the CV .
To create a frequency distribution table, follow these steps: 1) Arrange the data in ascending or descending order to create an array. 2) Count the frequency of each score or variable. 3) Use a two-column table, labeling the first column with the variable's name and the second with 'frequency' for the number of times each score appears. This process organizes data in a coherent form that allows for easy analysis and interpretation .
Frequency and relative frequency distributions support decision-making by organizing data into a clear, summarized format that highlights patterns and trends. Frequency distributions provide actual counts of occurrences, while relative frequency distributions offer percentages. This facilitates comparison, highlights proportions, and assists in making informed decisions by revealing insights about the prevalence of different categories within the dataset .
Relative frequency represents the percentage of the total number of data points that fall into each category, offering a normalized view of the data set. It enhances the understanding of a frequency distribution by providing insights into the proportion each category represents relative to the whole. This allows for comparisons across different data sets or categories to be made more effectively than using absolute frequency counts alone .
In a normal distribution, the mean is a central key metric because it indicates the typical value around which the data is symmetrically distributed. The mean coincides with both the median and mode in a perfectly normal distribution, ensuring that it represents the balance point of the distribution. This centrality underpins its role in further statistical analyses like hypothesis testing and prediction .
Selecting an appropriate number of class intervals is crucial for retaining the data's integrity while ensuring clarity and readability. If too few intervals are chosen, valuable details may be lost, while too many intervals can lead to a sparse table that lacks clarity. The number of class intervals should balance these factors to accurately reflect the distribution's structure and facilitate meaningful analysis .









