0% found this document useful (0 votes)
3 views13 pages

Overview of Statistics and Its Applications

This document provides an overview of data management, focusing on the importance of statistics in various fields such as business, education, psychology, and medicine. It outlines the history of statistics, key figures in its development, and essential concepts such as descriptive and inferential statistics, data types, and methods of data collection. The document emphasizes the application of statistical methods in everyday life and various industries, highlighting their significance in decision-making and analysis.

Uploaded by

Jhon Stackton
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
3 views13 pages

Overview of Statistics and Its Applications

This document provides an overview of data management, focusing on the importance of statistics in various fields such as business, education, psychology, and medicine. It outlines the history of statistics, key figures in its development, and essential concepts such as descriptive and inferential statistics, data types, and methods of data collection. The document emphasizes the application of statistical methods in everyday life and various industries, highlighting their significance in decision-making and analysis.

Uploaded by

Jhon Stackton
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

SESSION 6.

DATA MANAGEMENT

Good day, learner! I hope that you had understood the


essence and rules of elementary logic. The insights that
you had learned from that lesson is very vital to
understanding the languages, especially in English. This
time, you are going to investigate the science of data
management. This is the first lesson for your pre-finals.
The knowledge of statistics is essential for your future
lesson in Research and other research-related topics.

Statistics
Statistics – is a scientific body of knowledge that deals with the collection, organization or presentation, analysis and
interpretation of data

Collection refers to the gathering of information of data

Organization or presentation involves summarizing data or information in textual, graphical or tabular forms.

Analysis involves describing the data by using statistical methods and procedures

Interpretation refers to the process of making conclusions based on the analyzed data.

History of Statistics

Statistics is said to have developed from government records. All cultures with a recorded history also have recorded statistics
done mostly by agents of the government for governmental purpose.

In Egypt, the government prepared registration itself of all the heads of families. In ancient Judea a census of the population was
taken on several occasions, such as in 2010 B.C. when the population was estimated at 3,800,000. In 435 B.C., the first Roman
census was taken in the presence of the censors. Such census was repeated sixty-nine times in the following 470 years of
Roman History. The most famous Roman census was recorded in the Bible. This happened during the time of Joseph and Mary,
the parents of Jesus Christ. On Luke 2:1-4. We can read:

Now in those days was a decree forth from Caesar Augustus for all the inhabited earth to be registered. This
registration took place when Quirinius was governor of Syria and all people went travelling to be registered, each one to his own
City. Of course, Joseph also went from Galilee out of the city of Nazareth into Judea to David’s City, which is called Bethlehem,
because of his being member of the house and the family of David.

In the Middle Ages:

Taxes, military service and custom duties were also recorded. When William the Conqueror took possession of England, he
ordered that a survey be made of the lands of England for purposes of taxation and military service.
In the beginning of 16th century

A large number of statistical handbooks were published. This type of descriptive statistics was referred to as Die Tabellen
Statistik. The first scientific analysis of publicly recorded data may be ascribed to Captain John Graunt (1620-1674). The
registration of deaths was started by Henry VIII in 1532, and weekly bills of mortality were instituted during the period of the
plague.

People who were involved in the development of modern statistics:

Abraham De Moivre (1667–1754) – was a French – born mathematician who pioneered the development of analytic geometry
and the theory of probability. He discovered the equation of the normal curve

Karl Pearson de Laplace (1857–1936) – British statistician, leading founder of the modern field of statistics, prominent
proponent of eugenics, and influential interpreter of the philosophy and social role of science.

(Eugenics is the scientifically erroneous and immoral theory of “racial improvement” and “planned
breeding,” which gained popularity during the early 20th century. Eugenicists worldwide believed that they
could perfect human beings and eliminate so – called social ills through genetics and heredity)

Johann Carl Friedrich Gauss (1777 – 1855) was a German mathematician, geodesist, and physicist who made significant
contributions to many fields in mathematics and science. Gauss ranks among history's most influential mathematicians.

(Geodesists measure and monitor the Earth to determine the exact coordinates of any point. Geodesists
measure and monitor the Earth's size and shape, geodynamic phenomena (e.g., tides and polar motion), and
gravity field to determine the exact coordinates of any point on Earth and how that point will move over time.)
Page 1
Prepared by:
Prof. Ninfa Sua – Sotomil
Cell No: 09338577007/09171022328
Tel No: 5234099
sotomil_ninfas@[Link]
Sir Ronald Aylmer Fisher (1890 – 1962), renowned as "his time's greatest scientist," was a British statistician and biologist who
made significant contributions to experimental design and population genetics. He is widely regarded as the "Father of Modern
Statistics and Experimental Design."

Application of Statistics

In Business
A business firm collects and gathers data or information from its everyday operation. Statistics is used to summarize
and describe those data such as the amount of sales expenditures, and production to enable the management to understand
and determine the status of the firm. Data that have been organized and analyzed provide the management baseline data to
make wise decisions pertaining to the operation of the business.

In Education
Through statistical tools, a teacher can determine the effectiveness of a particular teaching method by analyzing test
scores obtained by their students. Results may be used to improve teaching-learning activities.

In Psychology
Psychologists are able to interpret meaningful aptitude tests. IQ tests, and other psychological tests using statistical
procedures or tools.

In Politics and Government


Public opinion and election polls are commonly used to assess the opinions or preference of the public for issues or
candidates of Internet. Statistics play an important role in conducting surveys or interviews for that purpose.

In Medicine
Statistics is also used in determining the effectiveness of its new drug product in treating tuberculosis. An experiment
or a clinical trial is conducted, then tuberculosis patients are treated using new drug product and another ten are treated using an
existing drug. The results are analyzed statistically to find out if the new product is more effective in treating tuberculosis.

In Agriculture
Through statistical tools, an agriculturist can determine the effectiveness of a new fertilizer in the growth of plants and crops.
Moreover, crop production and yield can be better analyzed through the use of statistical methods.

In Entertainment Industry
The most favorite actresses and actors can be determined by using surveys. Ratings of the members of the board of
judges in a beauty contest are statistically analyzed. Interviews are used to determine the most widely viewed television show.
The top grosser movies for this year are reported based on statistical records of movie houses. All these activities involve the
use of statistics.

In Everyday Life
The number of cars passing through streets or highway is recorded to enable traffic enforcers to manage effectively.
Even the number of pedestrians crossing the street, the number of people entering a warehouse or a department store, and the
number of people engaged in video games involve the use of statistics. In short, statistics is found and used in everyday life.

Descriptive and Inferential Statistics

Descriptive Statistics is a statistical procedure concerned with the describing the characteristics and properties of a
group of persons, places, or things.

Inferential Statistics is a statistical procedure that is used to draw inference or information about the properties or
characteristics by large group of people, places, or things or the basis of the information obtained from a small portion of large
group.

Terminologies in Statistics
 Population (N) – refers to a large collection of objects, persons, places or things. To illustrate this, suppose a
researcher wants to determine the average income of the residents of a certain barangay and there are 1,500
residents in the barangay. Then all of these residents comprise the population. A population is usually denoted by N.
Hence, in this case, N ¿ 1,500.

 Sample – is a small or portion or part of a population. It could also be defined as a subgroup, subset, or representative
of a population. For instance, suppose the above-mentioned researcher does not have enough time and money to
conduct the study using the whole population and he wants to use only 200 residents. These 200 residents comprise
the sample. A sample is usually denoted by n, thus, n ¿ 200.

 Parameter – is any numerical or nominal characteristics of a population. It is a value or measurement obtained from a
population. It is usually referred to as the rule or actual value. If in the preceding illustration, the researcher uses the
whole population (N ¿1500), then the average income obtained is called a parameter.

 Statistic – an estimate of a parameter. It is any value or measurement obtained from a sample. If the researcher in the
preceding illustration makes use of the sample (n ¿ 200), then the average income obtained is called a statistic.

Page 2
Prepared by:
Prof. Ninfa Sua – Sotomil
Cell No: 09171022328
Tel No: 5035955
sotomil_ninfas@[Link]
 Data (singular form is datum) – facts, or a set of information or observations under study. More specifically, data are
gathered by the researcher from a population or a sample. Data may be classified into two categories, qualitative and
quantitative.

Two Categories:
 Qualitative data – data, which can make assume values that manifest the concept of attributes. These are sometimes
called categorical data. Data falling in this category cannot be subjected to meaningful arithmetic. They cannot be
added, subtracted or divided.
 Quantitative data – data which are numerical in nature. These are data obtained from measuring or counting. In
addition, meaningful arithmetic operations can be done with this type of data. Examples: test scores and height

Variable – characteristic or property of a population or sample which makes the members different from each other. If a class
consists boys and girls, then gender is a variable in this class. Height is also a variable because different people have different
heights. Variables may be classified on the basis of whether they are discrete or continuous and whether they are dependent or
independent.

Classification of variables: Discrete and Continuous

 Discrete variable – is one that assumes a finite number of values. In other words, it can assume specific values only.
The values of a discrete variable are obtained through the process of counting.
o Example: The number of students in a class. If there are 40 students in a class, it cannot be reported that
there are 40.2 students or 40.5 students, because it is impossible for a fractional part of a student to be in the
class
 Continuous variable – is one that can assume infinite values within a specified interval. The values of a continuous
variable are obtained though measuring
o Example: Height. If one reports that the height of a building is 15m, it is also possible that another person
reports that the height of the same building is 15.1m or 15.2m, depending on the precision of the measuring
device used. In other words, the height of the building can assume several values.
Dependent and Independent Variable
 Dependent Variable – variable, which is affected or influenced by another variable.
 Independent Variable – one which affects or influences the dependent variable.

Constant – property or characteristics of a population or sample, which makes the members of the group similar to each other.

Scales of Measurement

Measurement – assignment of symbols or numerals to objects or events according to some rules. Since different rules are used
for the assignment of symbols, then this would yield different scales of measurement.

 Nominal Scale – the most primitive level of measurement. It is used when we want to distinguish one object from
another for identification purposes. In this level, we can only say that one object is different from another, but the
amount of difference between them cannot be determined. We cannot tell that one is better or worse than the other.
Example: Gender, Nationality, Civil Status

 Ordinal Scale – data arranged in some specified order or rank. When objects are measured in this level, we can say
that one is better or greater than the other. But we cannot tell how much more or how much less of the characteristics
one object has than the other.
Example: Ranking of contestants in a Beauty Pageant
Siblings in the Family
Honor students in the class

 Interval Scale – if data are measured in this scale, we can say not only one object is greater or lesser than another but
we can also specify the amount of difference
Example: Temperature in Celsius and Fahrenheit

 Ratio Scale – the ratio level of measurement is like the interval level, but with an only difference that ratio level starts
from an absolute or true zero point. In addition, there is always the presence of units of measure.
Example: Suppose Mrs. Sy weighs 50 kg., while her daughter weighs 25 kg. We can say that Mrs. Sy is
twice as heavy as her daughter.

Summation Notation

A symbol denoted by Greek capital letter sigma, which is used to indicate that subscripted variables are to be added.
n

∑ X i read as the summation of X sub i, from i = 1 to i =


i=1

Here, i is the index of summation and its value ranges from 1, the lower limit, to n, the upper limit. Observe that when
we write the sum of values in summation notation, we replace the subscript of the variable by an arbitrary subscript 1 and
indicate in the index, the range of the summation.
Page 3
Prepared by:
Prof. Ninfa Sua – Sotomil
Cell No: 09171022328
Tel No: 5035955
sotomil_ninfas@[Link]
Example:
100
1. X 1 + X 2 + X 3 + …+ X 100 =∑ X i
i=1
11
2. A 21 + A 22 + A23 +…+ A 211=∑ A 2i
i=1

{∑ }
11 2
2
3. { A1 + A2 + A3 + …+ A 11 } = Ai
i=1
20
4. ( Y 4 +5 ) + ( Y 5 +5 )+ ( Y 6 +5 ) + …+ ( Y 20 +5 )=∑ (Y i +5)
i=4
70
5. X 7 Y 7 + X 8 Y 8 + X 9 Y 9 +…+ X 70 Y 70=∑ X i Y i
i=7
Collecting Data
Two Types of Data

 Primary Data – data gathered from primary source. Primary sources of statistical data are the government institutions,
non-governmental organizations (NGOs), business agencies, and other organizations, as well as information derived
from personal interviews.

 Secondary Data – data gathered from secondary source. Secondary sources are books, encyclopedia, journals,
magazines, and research or studies conducted by other individuals.

Ways of collecting or gathering data


1. The Direct or Interview Method
In this method, the researcher has a direct contact with the interviewee. The researcher obtains the information needed
by asking questions and inquiries from the interviewee. This method is usually used in business research.
Example: A business firm would interview residents of a certain barangay regarding their favorite brand of toothpaste,
soap or shoes. TV personnel would ask televiewers about their favorite noontime show. Political analysis uses this method to
determine public opinion or preferences for candidates in upcoming elections.
Using this method, the researcher can get more accurate answers or responses since clarifications can be made if the
interviewee or respondent does not understand the question. However, this method is costly and time-consuming.

2. The Indirect or Questionnaire Method


The researcher makes use of a written questionnaire. The researcher gives or distributes the questionnaire to the respondents
either by personal delivery or by mail. Questionnaires are popular research methods because they offer a fast, efficient and
inexpensive means of gathering large amounts of information from sizeable sample volumes. These tools are particularly
effective for measuring subject behavior, preferences, intentions, attitudes and opinions.

3. The Registration Method


This method of collecting data is governed by laws. For example, birth and death rates are registered in the National
Statistics Office for records and future use. The number of registered cars can be found at the Land Transportations Office
(LTO). The list of registered voters in the Philippines is found at the Commission on Elections (COMELEC). This method of
gathering data is perhaps the most reliable because this is enforced by law.

4. The Experimental Method


The method is usually used to find out cause and effect relationships. Scientific researchers often use this method. For
example, agriculturists would like to know fertilizer on the growth of plants. The new kind of fertilizer will be applied to ten sets of
plants, while another set of ten plants will be given the ordinary. The growth of the plants will then be compared to determine
which fertilizer is better.

5. The Observation Method


The observation method is defined as a technique for observing and describing a subject’s behaviour. It is a method of
gathering pertinent facts and information via observation, as the name would imply. So, the researcher must build a relationship
with the participant and do so by immersing himself in their environment, it is also known as a participatory study.

Determining the Sample Size

In research, we seldom use the entire population because of the cost and time involved. In fact, most researchers do
not use the population in their study. Instead, the sample, which is a small representative of a population, is used. The
characteristic of the whole or entire population is described using the characteristics observed from the sample.

To determine the sample size from a given population size, the Slovin’s formula is used. Below is the Slovin’s formula.

N
n= 2
1+ N e
Where: n=¿ sample size
N=¿ population size
Page 4
Prepared by:
Prof. Ninfa Sua – Sotomil
Cell No: 09171022328
Tel No: 5035955
sotomil_ninfas@[Link]
e=¿ margin of error
Observe that there is a margin of error. When we use a sample, we do not get the actual value but just an estimate of
the parameter. Hence, there is an error associated when using the sample.

Example:
A group of researchers will conduct a survey to find out the opinions of residents of a particular community regarding
the oil price hike. If there are 10,000 residents in the community and the researchers plan to a sample using a 10% margin of
error, what should the sample size be?

Solution:
Given: N=10,000
e=10 %=0.10
Find: n=?
Solution:
N
n= 2
1+ N e
10,000
¿
1+ 10,000¿ ¿
n=99
Sampling Techniques – procedures used to determine the individuals or members of a sample

Two Sampling Techniques


 Probability Sampling – each member or element of the population has an equal chance of being drawn in to the
sample

o Random Sampling – each in the population has an equal chance of being drawn into the sample.
 Lottery Method
 Table of Random Method

o Systematic Sampling – picking every nth element of the population as a member of the sample

o Stratified Random Sampling – the word stratified comes from the word strata, which means groups or
categories (singular form is stratum). When we use this method we are actually dividing the elements of a
population into different categories or subpopulations and then the members of the sample are drawn or
selected proportionally from each subpopulation.

o Cluster Sampling – groups or cluster instead of individuals are randomly chosen. Recall that in the sampling
random sampling we select members of the sample individually. In cluster sampling, we will select or draw
the members of the sample individually. In cluster sampling, we will select or draw the members of the
sample by group and then we select a sample of elements from each cluster or group randomly. Cluster
sampling is sometimes called area sampling because this is usually applied when the population is large.

o Multi-Stage Sampling – combination of several sampling techniques that we have discussed. Usually this
method is used by researchers who are interested in studying a very large population, say the whole island
of Luzon, or even the Philippines. This is done by starting the selection of the members of the sample using
cluster sampling and then dividing each cluster or group into strata. Then, from each stratum individuals are
drawn randomly using simple random sampling.

 Non-Probability Sampling – the members of the sample are drawn from the population based on the judgment of the
researchers. The results of a study using this sampling technique are relatively biased. This technique lacks the
objectivity of selection; hence, it is sometimes called subjective sampling. Inferences made based on the sample
obtained using this technique is not so reliable. Non-probability techniques are used because they are convenient and
economical. Researchers use this method because they are inexpensive and easy to conduct.

o Convenience Sampling – as the name implies, this is used because of the convenience it offers to the
researcher. For example, a researcher who wishes to investigate the most popular noontime show may just
interview the respondents through the telephone. The result of this interview will be biased because the
opinions of those without telephone will not be included. Although convenience sampling may be used
occasionally, we cannot depend on it making inferences about a population.

o Quota Sampling – in this type of sampling, proportions of the various subgroups in the population are
determined and the sample is drawn to have the percentage in it. This is very similar to the stratified random
sampling discussed above. The only difference is that the selection of the members of the sample using
quota sampling is not done randomly. To illustrate this, let us suppose that we want to determine the
teenagers in the favorite brand of T-shirt. If there are 1,000 female and 1,000 male teenagers in the
population and we want to draw 150 members for our sample, we can select 75 female and 75 male
teenagers from the population without using randomization. This is quota sampling.

Page 5
Prepared by:
Prof. Ninfa Sua – Sotomil
Cell No: 09171022328
Tel No: 5035955
sotomil_ninfas@[Link]
o Purposive Sampling – let us suppose that we want to determine or predict the candidate who will win in the
upcoming election. We can conduct the survey or interview in places or precincts where people voted for the
winner in a series of post elections because we feel objectively that they will again vote for the next winner in
the upcoming election.

Presentation of Data
Ungrouped data – data that are neither organized, or if arranged, could only be from highest to lowest or vice versa.
Grouped data – data that are organized and arranged into different classes or categories.

1. Textual method
Ungrouped data can be presented in textual form, as in paragraph form. This involves enumerating the important
characteristics, giving emphasis on significant figures identifying important features of the data.

Example. Below are the test scores of 30 students in statistics.

25 30 18 17 50 12 43 35 40 9
33 37 41 21 20 31 35 46 10 36
28 19 18 13 28 16 42 27 28 31

Arranging the scores from lowest to highest will facilitate the enumeration of important characteristics of the data.

The most immediate statistics to compute are the sum, the minimum value, the maximum value, and the percentage.

For deeper statistical analysis, we compute for the measures of central tendency. The measures of central tendency
are the:
a. mean
b. median
c. mode
These measures give a good representation of the data. For instance, if someone says that the average
income of bank managers is Php 70,000.00, then all income from as low as Php 50,000.00 to Php 100,000.00 are
represented by this average value. The formulas used in computing these, are as follows:
n

a. Mean (
∑ xi where n is the number of scores, x i, the ith score
x ¿= i=1
n

b. Median (~
x ) ¿ is the middlemost score when the scores are arranged in descending or ascending order.
c. Mode ( ^x ¿=¿ is the most frequent score

Example. Suppose there are 5 children with ages 3, 5, 6, 8, and 10. Then the mean age is:
n

∑ xi 3+5+ 6+8+10 32
x= i=1 = = =6.4
n 5 5
The median age is 6 because two scores are below 6 and two scores are above 6. There is no mode.

The mean is the “center of gravity” of the data, whereas, the median is the middle-most score when data is arranged
from highest to lowest or from lowest to highest. The mode is the most frequent score.

The next statistical measures that the statistician computes are the measures of dispersion. The measures of
dispersion tell the statistician how separated the data are from each other. It answers the question on whether the scores in the
data are close to each other or widely separated from the mean. The three measures of dispersion are:

1. standard deviation
2. variance
3. range.
The standard deviation is an over-all distance from the mean to the scores. If this statistical distance is small, then we
say that the scores in the data are close to each other, while if it is large, the data are widely dispersed from the mean.

The variance on the other hand, is the square of the standard deviation. For instance, if the standard deviation is 2
then the variance is 4. Since the variance is from the “square root sign” it is more convenient to use the variance rather than the
standard deviation to describe a mathematical model out of the data.

A mathematical model is a theoretical description of population where the data was extracted from. Examples of
mathematical models are normal density functions, gamma density functions, etc.

The formula for the sample standard deviation is as follows:

Page 6
Prepared by:
Prof. Ninfa Sua – Sotomil
Cell No: 09171022328
Tel No: 5035955
sotomil_ninfas@[Link]

n

∑ (x ¿¿ i−x )2
i=1
Sample standard deviation(s)= ¿
n−1
2
sample variance=s
Example. From the 5 children with ages 3, 5, 6, 8, 10, to compute for the standard deviation, first we compute for the
mean:
3+5+6+ 8+10
x=
5
¿ 6.4
Then we put this value into the formula for the standard deviation:


n

∑ ( x i−x )2
i=1
s=
n−1

¿
√( 3−6.4 )2+ (5−6.4 )2+ ( 6−6.4 )2 + ( 8−6.4 )2 + ( 10−6.4 )2
4

√(−3.4 ) + (−1.4 ) + (−0.4 ) + ( 1.6 )2 + ( 3.6 )2


2 2 2
¿
4
¿
√29.2

¿ √ 7.3
4

¿ 2.70
So, the standard deviation of the ages of children is 2.7.
2
sample variance=s
2
sample variance=( 2.7 )
sample variance=7.3
The range is simply the distance between the highest score from the lowest score. If the range is big, then the difference
between the highest score and the lowest score is wide and hence, the scores in the data are widely spread from each other,
whereas if the range is small, it means that the scores are close to each other. The formula for the range is,
Range ¿ (Highest Score – Lowest Score)
Example. Find the range of the data given above on the ages of children 3, 5, 6, 8, 10.
Range ¿ highest score – lowest score
¿ 10 – 3
¿7
It is important to know measures of dispersion because several sets of data may have the same mean but it differs in
the spread of data. For instance, the mean of 4, 5, 6 is 5 but the mean of 1, 5, 9 is also 5. However, the scores in the first set are
close to each other, the second are not.

Example:
Given the following data set. Fill up the table.
53, 39, 82, 43, 72, 63, 23

1. Arrange the data set from lowest to highest. Add everything for the summation.
2. Get the mean at two decimal places.

x=
∑x
n
375
x=
7
x=53.57
3. Find x−x and its summation is accurately close to 0. However, just write 0.
4. Find (x−x )2 and its summation.

x x−x x−x ( x−x )


2
( x−x )
2

2
1 23 23−53.57=−30.57 −30.57 (−30.57) 934.52
2
2 39 39−53.57=−14.57 −14.57 (−14.57) 212.28
2
3 43 43−53.57=−10.57 −10.57 (−10.57) 111.72

Page 7
Prepared by:
Prof. Ninfa Sua – Sotomil
Cell No: 09171022328
Tel No: 5035955
sotomil_ninfas@[Link]
2
4 53 53−53.57=−0.57 −0.57 (−0.57) 0.32
2
5 63 63−53.57=9.43 9.43 (9.43) 88.92
2
6 72 72−53.57=18.43 18.43 (18.43) 339.66
2
7 82 82−53.57=28.43 28.43 (28.43) 808.26
∑ 375 ¿ ∑ (ABOVE)2495.68

5. Use the ∑ (x−x )2 to solve for the variance and the standard deviation at 2 decimal places.


n

∑ ( xi −x )2
i=1
sd=
n−1
¿

2,495.68
7−1
¿ √ 415.95
¿ 20.39
Final answer must be: Variance=415.95 and Standard deviation=20.39

The third consideration in analyzing a data set is the measures of location. These include percentiles, deciles, and
quartiles. Scores in the board and bar exams are summarized in terms of percentiles.

The figure 50th percentile means that 50 percent of the scores are below this value when the scores are arranged in
descending or ascending order. Thus, when a bar topnotcher obtains 87th percentile, it means that 87% of those who took the
exam are below his/her score. Hence, no one ever get 100th percentile.

Deciles on the other hand divide the array of scores into tens, when arranged into ascending order. The 8th decile
means that 80% of the scores are below this value.

And quartiles divide the score into groups of four when these scores are arranged into ascending order.
Q1, the first quartile, is above 25% of the data; Q2, is above 50% of the data, and Q3 is above 75% of the data.
Because of this, 50% of the data is between Q1 and Q3. This is the basis of the Box-and-Whiskers Plot.

Box – and Whisker Plots

A box – and – whisker plot (sometimes called a box plot) is often used to provide a visual summary of a set of data.
A box – and – whisker plot shows the median, the first and third quartiles, and the minimum and maximum values of a data set.
See the figure below
BOX

WHISKER WHISKER

LOWEST HIGHEST
OBSERVATION OBSERVATION
LOWER MEDIAN, Q2 UPPER
QUARTILE, Q1 QUARTILE, Q3
Box-and-Whiskers Plot
Construction of a Box – and – Whiskers Plot
1. Draw a horizontal scale that extends from the minimum data (lowest observation) value to the maximum data (highest
observation) value.
2. Above the scale, draw a rectangle (box) with its left side as Q 1and its right side as Q 3 .
3. Draw a vertical line segment across the rectangle as the median, Q 2
4. Draw a horizontal line segment, called a whisker, that extends from Q 1in the minimum and another whisker that
extends from Q 3 to the maximum.
Example. Suppose the data is composed of the following scores:
14, 22, 35, 40, 30, 24, 24, 20, 24, 24, 27, 21, 28, 43, 39, 29, 35, 13, 18, 26

To compute for:
1. percentile
2. decile, and
3. quartile,
First, arrange the data in ascending order, as follows:
13, 14, 18, 20, 21, 22, 24, 24, 24, 24, 26, 27, 28, 29, 30, 35, 35, 39, 40, 49
One-fourth of the way from the lowest is the first quartile, half of the way is the median and ¾ of the way is the third quartile.

Please see the figure below.

Page 8
Prepared by:
Prof. Ninfa Sua – Sotomil
Cell No: 09171022328
Tel No: 5035955
sotomil_ninfas@[Link]
21.1 24 30.1 49

First Quartile, Q1 Median Third Quartile, Q3 Highest


Observation

Formula: If N is the number of elements in the data set, the quartile is computed as follows:

N
Q1: Divide N by 4, then count from the beginning of the data set until the quotient .
4
N
From the data above, N=20 , =5.
4
Therefore, Q 1 is a value above the fifth score, or 21. So Q 1=21.1 .

Q3: Divide N by 4, then multiply the quotient by 3. Then count from the beginning of the data set to the product.
3N
From the data above, N=20 and =15 .
4
Therefore, Q 3 is a value above the 15th score, or 30. So Q 3 = 30.1.

Percentiles

P90: Divide N by 100, then multiply by 90. Then count from the beginning of the data set to the product
90 N 90 (20 )
From the data above N = 20 and = =18
100 100
Therefore P90 :is a value above the 18th score, or 39. So P90 =39.1

Normal Distributions and the Empirical Rule

One of the most important statistical distributions of data is known as a normal distribution. This distribution occurs in a
variety of applications. Types of data that may demonstrate a normal distribution include the lengths of leaves on a tree, the
weight of newborns in a hospital, the lengths of time of a student’s trip m from home to school over a period of months, the SAT
scores of a large group of students, and the life spans of light bulbs

A normal distribution forms a bell – shaped curved that is symmetric about a vertical line through the mean of the
data.

A normally distributed population is a distribution where the mean, the median, and the mode are equal. Moreover,
when the mathematical equation of the normal density function is plotted as a graph, its form is bell-shaped, like the one below.

Properties of a Normal Distribution


Every normal distribution has the following properties:
 The graph is symmetric about a vertical line through the mean of the distribution.
 The mean, median, and mode are equal.
 The y – value of each point on the curve is the percent (expressed as a decimal) of the data at the corresponding x –
value.
 Areas under the curve that are symmetric about the mean are equal.
 The total area under the curve is 1.

Empirical Rule for a Normal distribution


In a normal distribution, approximately
 68% of the data lie within 1 standard deviation of the mean.
 95% of the data lie within 2 standard deviations of the mean.
 99.7% of the data lie within 3 standard deviations of the mean.

The Standard Normal Distribution

The exact computation for areas under the normal curve is through the use of the areas of the standard normal
distribution. The standard normal distribution is the normal distribution that has a mean of 0 and a standard deviation of 1. With
Page 9
Prepared by:
Prof. Ninfa Sua – Sotomil
Cell No: 09171022328
Tel No: 5035955
sotomil_ninfas@[Link]
this distribution comes already computed areas in tabular form from negative infinity (−∞) to a z−¿ value. The letter “ z ” is
used to denote standard normal random variable. A few of these values are as follows:
If x is a non-standard normal distribution random variable, then this can be transformed into a standard normal
variable using the following formula:
x−µ
z=
σ
For example, if x=30 , µ=36 , and σ =3, then
30−36
z=
3
−6
¿
3
¿−2
The area from (−∞) to x=30 is the same as the area from (−∞) to z=−2 .

Example: Draw a normal curve and locate the z – distance along the baseline. Shade the area ask in each problem.

1. On the assumption that IQ scores are normally distributed in the population with a mean of 100 and a standard
deviation of 20. Find the proportion of the people with IQ scores:
a. Above 135 d. Between 75 and 125
b. Above 120 e. What is the probability of scoring between 90 and 120?
c. Below 90 f. What percentage of the cases cannot score more than 110?
Solution:
a. Above 135
Given: x=100
x=135
s=20
Find: z=?
x−x
z=
s
135−100
¿
20
z=1.75
A=0.4599 value from the z table
0.5000−0.4599=0.0401
Answer: 4.01% →proportion of people with IQ above 135
b. Above 120
Given: x=100
x=120
s=20
Find: z=?
x−x
z=
s
120−100
¿
20
z=1.0
A=0.3413 value from the z table
0.5000−0.3413=0.1587
Answer: 15.87% →proportion of people with IQ above 120

c. below 90
Given: x=100
x=90
s=20
Find: z=?
x−x
z=
s
90−100
¿
20
z=−0.5
A=0.1915 value from the z table
Page 10
Prepared by:
Prof. Ninfa Sua – Sotomil
Cell No: 09171022328
Tel No: 5035955
sotomil_ninfas@[Link]
0.5000−0.1915=0.3085
Answer: 30.85% →proportion of people with IQ below 90

d. Between 75 and 125


Given: x=100
x=75 ; 125
s=20
Find: z=?
x1−x
z 1=
s
75−100
¿
20
z 1=−1.25; → A 1=0.3944 value from the z table
x2 −x
z 2=
s
125−100
¿
20
z 2=1.25; → A 2=0.3944 value from the z table
AT = A1 + A2 =0.3944+ 0.3944=0.7888
Answer: 78.88% →proportion of people with IQ score between 75 and 125

e. What is the probability of scoring


between 90 and 120?
Given: x=100
x=90 ; 120
s=20
Find: z=?
x1−x
z 1=
s
90−100
¿
20
z 1=−0.5 ; → A 1=0.1915 value from the z table
x2 −x 120−100
z 2= =
s 20
z 2=1.0; → A 2=0.3413 value from the z table
AT = A1 + A2 =0.1915+0.3413=0.5328

Answer: 53.28% →probability of scoring between 90 and 120.

f. What percentage of the cases cannot score more than 110?

Given: x=100
x=110
s=20
Find: z=?
x−x
z=
s
110−100
¿
20
z=0.5 ; → A 1=0.1915 value from the z table
AT =0.5000+ A 1=0.5000+0.1915=0.6915
Answer: 69.15% →percentage of the cases that cannot score more than 110.
Exercises 6
I. Given:
Midterm sample test scores of 10 students in Math 1.
78 70 90 67 85 45 74 68 70 73
Determine the following:
1. Mean 2. Median
Page 11
Prepared by:
Prof. Ninfa Sua – Sotomil
Cell No: 09171022328
Tel No: 5035955
sotomil_ninfas@[Link]
3. Mode 6. Q3
4. Standard deviation
7. D 4
5. Variance
P60
x i−x
P60 x x ( x−x ¿ ( x−x )
2
s z=
s
1 45 72 45−72=−27 729 12.07 −¿2.23695
2 67 72 67−72=−5 25 12.07 −¿0.4142
3 68 72 68−72=−4 16 12.07 −¿0.3314
4 70 72 70−72=−2 4 12.07 −¿0.1657
5 70 72 70−72=−2 4 12.07 −¿ 0.1657
6 73 72 73−72=1 1 12.07 0.0008
7 74 72 74−72=2 4 12.07 0.1657
8 78 72 78−72=6 36 12.07 0.4971
9 85 72 85 −72=13 3 169 12.07 1.0770
10 90 72 90−72=18 324 12.07 1.4913

∑ 327 ¿ ∑ (ABOVE)1312
n
s=12.07384685
1.
∑ xi
720 5. 2
s =12.07384685
2

x= i=1 = =72 2
s =145.78
n 10
70+73 3 N 3 (10)
2. ~
x= =71.5 6. Q 3= = =7.5≈ 8 ; Q3 =78.1
2 4 4
3. ^x ¿ 70 4 N 4 (10)
7. D 4 = = =4 ; D 4=70.1


n 10 10
4. ∑ ( x i−x )2
s= i=1 60 N 60 (10)
8. P60= = =6 ; P 60=73.1
n−1 100 100
¿
√1312
10−1
¿
√1312
9

Page 12
Prepared by:
Prof. Ninfa Sua – Sotomil
Cell No: 09171022328
Tel No: 5035955
sotomil_ninfas@[Link]
II. Find the area under the normal curve from: z 2=2.54 ; A2=0.4945
1. z ¿ −¿ 1.54 to z ¿ −¿ 1.87
2. z ¿ 1.23 to z ¿ 2.54 AT = A2 −A 1
3. z ¿ 1.58 to z ¿ −¿ 2.53 AT =0.4945−0.3907
4. z ¿−¿ 1.32 to z ¿ 1.37
AT =0.1038 ≈ 10.38 %
Solution: 3. z 1=1.58 ; A1=0.4429
1. z 1=−1.54 ; A1=0.4382 z 2=−2.53 ; A2=0.4943
z 2=−1.87 ; A2=0.4693 AT = A2 + A1
AT = A2 −A 1 AT =0.4943+ 0.4429
AT =0.4693−0.4382 AT =0.9372 ≈ 93.72%
AT =0.0311 ≈ 3.11% 4. z 1=−1.32 ; A1=0.4066
z 2=1.37 ; A2=0. 4147
AT = A2 + A1
2. z 1=1.23 ; A1=0.3907 AT =0.4147+ 0.4066
AT =0.8213 ≈ 82.13%

Page 13
Prepared by:
Prof. Ninfa Sua – Sotomil
Cell No: 09171022328
Tel No: 5035955
sotomil_ninfas@[Link]

You might also like