0% found this document useful (0 votes)
10 views48 pages

Understanding Frequency Distributions

Chapter 3 discusses descriptive statistics, focusing on data reduction through frequency distributions, which organize raw data into meaningful tables or graphs. It outlines the types of frequency distributions (ungrouped, categorical, and grouped) and the steps for constructing them, emphasizing the importance of visual representations like histograms and frequency polygons for data analysis. The chapter also highlights the advantages and disadvantages of frequency distributions in statistical analysis.

Uploaded by

mayakalkidan603
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
10 views48 pages

Understanding Frequency Distributions

Chapter 3 discusses descriptive statistics, focusing on data reduction through frequency distributions, which organize raw data into meaningful tables or graphs. It outlines the types of frequency distributions (ungrouped, categorical, and grouped) and the steps for constructing them, emphasizing the importance of visual representations like histograms and frequency polygons for data analysis. The chapter also highlights the advantages and disadvantages of frequency distributions in statistical analysis.

Uploaded by

mayakalkidan603
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

CHAPTER 3

DESCRIPTIVE STATISTICS
Content
3.1. Data reduction
3.1.1. Frequency distribution
3.1.2. Graphing display of data
3.3.3. Shape of frequency distribution
3.2. Describing data
3. 2. 1. Measures of central tendency
3. 2. 2. Measures of variation
3 .2. 3. Measures of relationship
3.1. Data reduction
3.1.1 Frequency distribution
To describe situations, draw conclusions, or make inferences about events, the researcher must
organize the data in some meaningful way. The most convenient method of organizing data is to
construct a frequency distribution. A frequency distribution takes a disorganized set of scores and
places them in order from highest to lowest, grouping together individuals who all have the same
score.

A frequency distribution is the organization of raw data in table form, using classes
and frequencies

Thus, a frequency distribution allows the researcher to see “at a glance” the entire set of scores. It
shows whether the scores are generally high or low, whether they are concentrated in one area or
spread out across the entire scale, and generally provides an organized picture of the data. In addition
to providing a picture of the entire set of scores, a frequency distribution allows you to see the location
of any individual score relative to all of the other scores in the set. A frequency distribution can be
structured either as a table or as a graph, but in either case, the distribution presents the same two
elements:

The reasons for constructing a frequency distribution are as follows:


1. To organize the data in a meaningful, intelligible way.
2. To see “at a glance” the entire set of scores.
3. To enable the reader to determine the nature or shape of the distribution.
4. To facilitate computational procedures for measures of average and spread
5. To enable the researcher to draw charts and
1 graphs for the presentation of data
6. To enable the reader to make comparisons among different data sets.
Organizing Data
Suppose a researcher wished to do a study on the number of miles that the employees of a large
department store traveled to work each day. The researcher first would have to collect the data by
asking each employee the approximate distance the store is from his or her home. When data are
collected in original form, they are called raw data. In this case, the data are

Since little information can be obtained from looking at raw data, the researcher organizes the data
into what is called a frequency distribution.

Frequency table ordered listing of number of individuals having each of the different
values for a particular variable.

3 types of frequency distributions are;


A. Ungrouped frequency distribution when the range of data is small, each class
is only one unit
B. Categorical (nominal) frequency distribution, such as blood type or political
affiliation.
C. Grouped frequency distribution when the range is large and classes several
2
units in width are needed.
A. Ungrouped frequency distribution

B. Categorical Frequency Distributions


The categorical frequency distribution is used for data that can be placed in specific categories, such as
nominal- or ordinal-level data. For example, data such as political affiliation, religious affiliation, or major
field of study would use categorical frequency distributions.

Twenty-five army inductees were given a blood test to determine their blood type. The data set is

Construct a frequency distribution for the data.

Solution
Since the data are categorical, discrete classes can be used. There are four blood types: A, B, O, and
AB. These types will be used as the classes for the distribution. The procedure for constructing a
frequency distribution for categorical data is given next

3
Step 1 Make a table as shown.

Step 2 Tally the data and place the results in column B.


Step 3 Count the tallies and place the results in column C.
Step 4 Find the percentage of values in each class by using the formula in D

where f frequency of the class and n total number of values. For example, in the class of type A blood,
the percentage is

Percentages are not normally part of a frequency distribution, but they can be added since they are
used in certain types of graphs such as pie graphs. Also, the decimal equivalent of a percent is called
a relative frequency.

Step 5 Find the totals for columns C (frequency) and D (percent). The completed table is shown.

4
C. Grouped frequency table
When the range of the data is large, the data must be grouped into classes that are more than one unit in width,
in what is called a grouped frequency distribution. Under these conditions, the individual scores are
usually grouped into class intervals and presented as a frequency distribution of grouped scores.
When you are grouping data, one of the important issues is how wide each interval should be.
Whenever data are grouped, some information is lost. The wider the interval, the more information
lost. we can see that in grouping scores there is a trade-off between losing information and
presenting a meaningful visual display. To have the best of both worlds, we must choose an
interval width neither too wide nor too narrow.

The steps for constructing a frequency distribution of grouped scores are as follows:
1. Find the range of the scores.
2. Determine the width of each class interval (i).
3. List the limits of each class interval, placing the interval containing the lowest score value at the
bottom.
Consideration while determining classes
✓ There should be between 5 and 20 classes. Although there is no hard-and-fast rule
for the number of classes contained in a frequency distribution, it is of the utmost
importance to have enough classes to present a clear description of the collected
data.
✓ Mutually exclusive. Mutually exclusive classes have nonoverlapping class limits so that data
cannot be placed into two classes.
✓ The classes must be continuous. Even if there are no values in a class, the class must be
included in the frequency distribution. There should be no gaps in a frequency distribution.
The only exception occurs when the class with a zero frequency is the first or last class. A
class with a zero frequency at either end can be omitted without affecting the distribution.
✓ The classes must be exhaustive. There should be enough classes to accommodate all
the data.
✓ The classes must be equal in width
One exception occurs when a distribution has a class that is open-ended. That is, the class has no
specific beginning value or no specific ending value. A frequency distribution with an open-ended
class is called an open-ended distribution. Here are two examples of distributions with open-ended
classes.

5
Class limit; lowest and highest score that can be included in the class (e.g. 10-20)
Class boundary; class limit with no gap (9.5-20.5, 20.5-31.5)
Class width; the difference between upper class boundary and lower-class boundary (21-10=11)
5. Tally the raw scores into the appropriate class intervals.
6. Add the tallies for each interval to obtain the interval frequency.
Let us apply these steps to the following data

1. Finding the range.


Range Highest score minus lowest score 99-46 = 53
2. Determining interval width (i). Let’s assume we wish to group the data into approximately 10 class
intervals.

When i has a decimal remainder, we will follow the rule of rounding i to the same number of decimal
places as in the raw scores. Thus, i rounds to 5. (some advice to Rounding up any decimal remainder
than rounding off i.e. 5.3 to 6). When we Rounding up, we reduce one class of interval whereas when
we round off we add one class of interval than the initial class of interval

6
3. Listing the intervals. We begin with the lowest interval. The first step is to determine the lower
limit of this interval. There are two requirements:
A. The lower limit of this interval must be such that the interval contains the lowest score.
B. It is customary to make the lower limit of this interval evenly divisible by i.
Given these two requirements, the lower limit is assigned the value of the lowest score in the
distribution if it is evenly divisible by i. If not, then the lower limit is assigned the next lower value
that is evenly divisible by i. In the present example, the lower limit of the lowest interval begins with
45 because the lowest score (46) is not evenly divisible by 5.
Tallying the scores; in listing the other intervals, we must be sure that the intervals are continuous
and mutually exclusive. By mutually exclusive, we mean that the intervals must be such that no score
can be legitimately included in more than one interval.
Relative Frequency, Cumulative Frequency, and Cumulative Percentage Distributions
■ A relative frequency distribution indicates the proportion of the total number of scores that occur
in each interval.

■ A cumulative frequency distribution indicates the number of scores that fall below the upper real
limit of each interval. It also indicates the percentage of scores that fall below the upper real limit of
each interval.

7
When one is constructing a frequency distribution, the guidelines presented in this section should be
followed. However, one can construct several different but correct frequency distributions for
the same data by using a different class width, a different number of classes, or a different starting
point.
Advantage and Disadvantages of frequency distribution
A. Advantages:
– It condenses a large mass of data into a comparatively small table.
– It attracts the attention of even a layman and gives him an insight into the nature of the distribution.
– It helps for further statistical analysis, like central tendency, scatter, symmetry, of the data.
B. Disadvantages:
– In the grouped frequency distributions, the identity of the observations is lost. We know only the
number of observations in a class and do not know what the values are.
– Because the selection of the class width and the lower-class limit of the first class are to a certain
extent arbitrary, different frequency distributions may be constructed for the same data and hence
may give contradictory impressions.

Check point
[Link] five reasons for organizing data into a frequency distribution.
2. Name the three types of frequency distributions, and explain when each should be used.
3. Find the class boundaries, midpoints, and widths for each class.
a. 12–18
b. 56–74
c. 695–705
4 What are open-ended frequency distributions? Why are they necessary?

8
5. Shown here are four frequency distributions. Each is incorrectly constructed. State the
reason why

6. Organize the data into a frequency distribution with six classes

7. Construct a frequency distribution for the data. (Use your own judgment as to the
number of classes and class size.)

9
3.1.2 Graphing Frequency distributions
After the data have been organized into a frequency distribution, they can be presented in graphical form. The
purpose of graphs in statistics is to convey the data to the viewers in pictorial form. It is easier for most people
to comprehend the meaning of data presented graphically than data presented numerically in tables or
frequency distributions. This is especially true if the users have little or no statistical knowledge.
Statistical graphs can be used to describe the data set or to analyze it. Graphs are also useful in getting the
audience’s attention in a publication or a speaking presentation. They can be used to discuss an issue,
reinforce a critical point, or summarize a data set. They can also be used to discover a trend or pattern in a
situation over a period of time. They have greater attraction, facilitate comparison and are easily
understandable
1. A graph has two axes: vertical and horizontal. The vertical axis is called the ordinate, or Y axis,
and the horizontal axis is the abscissa, or X axis.
2. Very often the independent variable is plotted on the X axis and the dependent variable on the Y
axis. In graphing a frequency distribution, the score values are usually plotted on the X axis and the
frequency of the score values is plotted on the Y axis.
3. Suitable units for plotting scores should be chosen along the axes.
I. Histogram
Bar like graph of a frequency distribution in which the values are plotted along the horizontal axis
and the height of each bar is the frequency of that value; the bars are usually placed next to each other
without spaces, giving the appearance of a city skyline.
To construct a histogram, you first list the numerical scores (the categories of measurement) along
the X-axis. Then you draw a bar above each X value so that
a. The height of the bar corresponds to the frequency for that category.
b. For continuous variables, the width of the bar extends to the real limits of the category. For
discrete variables, each bar extends exactly half the distance to the adjacent category on each
side.

The histogram is a graph that displays the data by using contiguous vertical bars (unless
the frequency of a class is 0) of various heights to represent the frequencies of the classes.

10
Construct a histogram to represent the data shown for the record high temperatures

Solution
Step 1 Draw and label the x and y axes. The x axis is always the horizontal axis, and the y axis is
always the vertical axis.
Step 2 Represent the frequency on the y axis and the class boundaries on the x axis.
Step 3 Using the frequencies as the heights, draw vertical bars for each class.

NB; The histogram does not involve any gaps between the two successive bars.
II. Polygon
The frequency polygon is a graph that displays the data by using lines that connect points plotted for
the frequencies at the midpoints of the classes. The frequencies are represented by the heights of the
points.
A. A dot is centered above each score so that the vertical position of the dot corresponds to the
frequency for the category.
B. A continuous line is drawn from dot to dot to connect the series of dots.

11
C. The graph is completed by drawing a line down to the X-axis (zero frequency) at each end of the
range of scores.
Solution
Step 1 Find the midpoints of each class. Recall that midpoints are found by adding the upper and
lower boundaries and dividing by 2:

Step 2 Draw the x and y axes. Label the x axis with the midpoint of each class, and then use a suitable
scale on the y axis for the frequencies.
Step 3 Using the midpoints for the x values and the frequencies as the y values, plot the points.
Step 4 Connect adjacent points with line segments. Draw a line back to the x axis at
the beginning and end of the graph, at the same distance that the previous and next midpoints would
be located,

12
The frequency polygon and the histogram are two different ways to represent the same data set. The choice
of which one to use is left to the discretion of the researcher.
III The Ogive
This type of graph is called the cumulative frequency graph or ogive. The cumulative frequency is
the sum of the frequencies accumulated up to the upper boundary of a class in the distribution.
Solution
Step 1 Find the cumulative frequency for each class.

Step 2 Draw the x and y axes. Label the x axis with the class boundaries. Use an appropriate scale for
the y axis to represent the cumulative frequencies.
Step 3 Plot the cumulative frequency at each upper-class boundaries are used since the cumulative
frequencies represent the number of data values accumulated up to the upper boundary of each class.
Step 4 Starting with the first upper class boundary, 104.5, connect adjacent points with line segments

Cumulative frequency graphs are used to visually represent how many values are below a certain
upper-class boundary. For example, to find out how many records high temperatures are less than

13
114.5F, locate 114.5F on the x axis, draw a vertical line up until it intersects the graph, and then draw
a horizontal line at that point to the y axis. The y axis value is 28

IV. Bar Chart


It is the most common presentation for nominal, categorical or discrete data. It uses a serious of
separated and equally spaced bars. The heights of the bars represent the frequency or relative
frequency of the classes. But the width of the bars has no meaning; however, all the bars should be
the same width to avoid distortion. The bars are separated by constant distance. A bar graph is
essentially the same as a histogram, except that spaces are left between adjacent bars.
A. Simple Bar Chart: is a diagram in which categories of a variable are marked on the X
axis and the frequencies of the categories are marked on the Y axis. It is applicable for
discrete variables, that is, for data given according to some period, places and timings.

14
B. Multiple Bar Chart: used to display data on more than one variable. In the multiple
bars diagram two or more sets of inter-related data are interpreted.

Check point
1. What questions could be answered more easily by looking at the histogram,
frequency polygon and Ogive
2. Construct a histogram, frequency polygon, and ogive for the data

3. Using the histogram shown here, do Construct a frequency distribution;


include class limits, class frequencies, midpoints, and cumulative frequencies.

15
V. line graph

16
VI. Pie chart
The purpose of the pie graph is to show the relationship of the parts to the whole by visually
comparing the sizes of the sections. Percentages or proportions can be used. The variable is
nominal or categorical.
A pie chart shows how a total amount is divided between levels of a categorical variable as a
circle divided into radial slices. Each categorical value corresponds with a single slice of the
circle, and the size of each slice (both in area and arc length) indicates what proportion of the
whole each category level takes. In order to use a pie chart, you must have some kind of
whole amount that is divided into a number of distinct parts. Your primary objective in a pie
chart should be to compare each group’s contribution to the whole, as opposed to comparing
groups to each other. If the above points are not satisfied, the pie chart is not appropriate, and
a different plot type should be used instead.
Step in construction pie chart
Step 1 Since there are 360 in a circle, the frequency for each class must be converted into a
proportional part of the circle. This conversion is done by using the formula

where f frequency for each class and n sum of the frequencies. Hence, the following conversions are
obtained. The degrees should sum to 360*
*Note: The degrees column does not always sum to 360 due to rounding.
Step 2 Each frequency must also be converted to a percentage. Recall from Example 2–1 that this conversion
is done by using the formula

Step 3 Next, using a protractor and a compass, draw the graph using the appropriate degree measures
found in step 1, and label each section with the name and percentages

17
Pie charts with a large number of slices can be difficult to read. It can be difficult to see the smallest
slices, and it can be difficult to choose enough colors to make all of the slices distinct.
Recommendations vary, but if you have more than about five categories, you might want to think
about using a different chart type.
It is actually very difficult to discern exact proportions from pie charts, outside of small fractions like
1/2 (50%), 1/3 (33%), and 1/4 (25%).

Content of Graphs/charts
A. X and Y Axis/ circle; X and Y Axis for all graph type and circle for pie chart
B. Data; Each item of data on graph/chart ties the dependent variable to an independent
variable represented by bar or line for graphs and segment/portions for pie chart
C. Title; All Graphs/charts have a title that explain what the graph is depicting
D. Legend; If we have two or more variables in the graph/chart distinguished by different
colors, the legend explains what each dependent variable is

18
VII. Stem and Leaf Diagrams
Stem and leaf diagrams were first developed in 1977 by John Tukey, working at Princeton University.
They are a simple alternative to the histogram and are most useful for summarizing and describing
data when the data set includes less than 100 scores. Unlike what happens with a histogram; however,
a stem and leaf diagram does not lose any of the original data.
The stem is placed to the left of the vertical line and the leaf to the right. For example, the stems and
leaf for the first and last original scores are:
Construct a stem and leaf plot for the data.

Solution
Step 1 Arrange the data in order:
02, 13, 14, 20, 23, 25, 31, 32, 32, 32, 32, 33, 36, 43, 44, 44, 45, 51, 52, 57
Step 2 Separate the data according to the first digit, as shown.

Step 3 A display can be made by using the leading digit as the stem and the trailing digit as the leaf.
For example, for the value 32, the leading digit, 3, is the stem and the trailing digit, 2, is the leaf. For
the value 14, the 1 is the stem and the 4 is the leaf.

stem leaf

19
3.1.3 Shapes of Frequency Curves
Frequency distributions can take many different shapes. Some of the more commonly encountered
shapes are shown. Curves are generally classified as symmetrical or skewed.

3.2. Describing data


So far, we discussed how raw data can be organized in terms of tables, charts and frequency
distributions in order to be easily understood and analyzed. Frequency distributions and their
corresponding graphical displays roughly tell us some of the features of a data set. However,
they don’t condense the mass of data in a way that we can easily understand and interpret.
Hence, organizing a data into a frequency is not sufficient, there is a need for further
condensation, particularly when we want to compare two or more distributions we may reduce
the entire distribution into one number that represents the distribution we need.
3. 2. 1. Central tendency
The most common method for summarizing and describing a distribution is to find a single value that
defines the average score and can serve as a representative for the entire distribution. In statistics, the
concept of an average, or representative, score is called central tendency. The goal in measuring
central tendency is to describe a distribution of scores by determining a single value that identifies

20
the center of the distribution. Ideally, this central value is the score that is the best representative value
for all of the individuals in the distribution.
The main objectives of measuring central tendency are:
➢ To get a single value that represent(describe) characteristics of the entire data
➢ To summarizing/reducing the volume of the data
➢ To facilitating comparison within one group or between groups of data
➢ To enable further statistical analysis

Central tendency is a statistical measure that attempts to determine the single


value, usually located in the center of a distribution, that is most typical or most representative
of the entire set of scores.

Types of measures of central tendency


A. The Mean (Arithmetic, Weighted and compound mean)
A. The Mode
B. Median
C. Quantiles
[Link] Mean

A. Arithmetic mean
I. Simple Arithmetic Mean

21
The Summation Notation
Let X1, X2, X3,…, XN be a number of measurements where N is the total number of
observation and Xi is ith observation, then it is very often in statistics an algebraic
expression of the form X1+X2+X3+...+XN is used in a formula to compute a statistic. It is
tedious to write an expression like this very often, so mathematicians have developed a
shorthand notation to represent a sum of scores, called the summation notation.

Example; Find the arithmetic mean for the following data, 2, 4, 5,1,7,18

2+4+5+1+7+18 37 6.16
6 6
II. Arithmetic Mean for Grouped Data
If data are given in the shape of a continuous frequency distribution, then the arithmetic mean
is obtained as follows:

22
Example: Find the arithmetic mean for the following frequency distribution

If a wrong figure has been used when calculating the mean the correct mean can be obtained
without repeating the whole process using:

Example: An average weight of 10 students was calculated to be 65. Latter it was discovered
that one weight was misread as 40 instead of 80 kg. Calculate the correct average weight.

D. Weighted and combined mean


Weighted arithmetic mean is used to calculate the average when the relative importance of the
observations differs. This relative importance is technically known as weight. Weight could be
a frequency or numerical coefficient associated with [Link] also used when the
situation arises in which we know the mean of several groups of scores and we want to calculate
the mean of all the scores combined.

23
I. Weighted mean

Example; Suppose that student X obtained the following grade in the first semester. calculate
Weighted mean of GPA

Solution

II. Combined mean

24
Check point
1. The table shows the 3 candidates result for vacancy of an organization. The
weights of these criteria and scores obtained by 3 candidates (out of 100 in each
criterion) are given in the following table. Who is the appropriate candidate for
this position based on the criteria?
Candidates
Criterion Weight Criterion Weight 1 2 3
Work experience 4 70 89 85
Entrance exam 3 78 83 89
Interview result 2 90 92 90

2. A student obtained the following mark in an examination: English 60,


management 75, Mathematics 63, accounting 59, and economics 55. Find the
students weighted arithmetic mean if weights of course were 4,3,5,3,2 respectively
to the subjects.

Special Properties of Mean


✓ The mean is sensitive to the exact value of all the scores in the distribution.
✓ The mean is very sensitive to extreme scores
✓ The sum of the deviations about the mean equals zero
✓ Under most circumstances, the mean is least subject to sampling variation.
Merits and Demerits of Arithmetic Mean
Merits:
• It is based on all observation.
• It is suitable for further mathematical treatment.
• It is stable average, i.e. it is not affected by fluctuations of sampling to some extent.
• It is easy to calculate and simple to understand.
Demerits:
• It is affected by extreme observations.
• It can not be used in the case of open end classes.
• It can not be determined by the method of inspection.
• It can not be used when dealing with qualitative characteristic

25
[Link] Mode
The mode is a value which occurs most frequently in a set of values, and which occurs more
than once
Example
a) Find the mode of 5, 3, 5, 8, 9 Solution: Mode =5
b) Find the mode of 8, 9, 9, 7, 8, 2, and 5. It is a bimodal Data: 8 and 9
c) Find the mode of 4, 12, 3, 6, and 7. No mode for this data
Mode for Grouped data
In a frequency distribution, the mode is located in the class with highest frequency and that
class is the modal class. Then the formula for mode is

Example: Use the frequency distribution of heights in the following table to find the mode
of height of the 100 male students at X university and interpret the result

Solution:
A class having the highest frequency is considered as a modal class. Thus the 3rd class
(65.5-68.5) is the modal class.

26
Merits and Demerits of Mode
Merits:
• It is not affected by extreme observations.
• Easy to calculate and simple to understand.
• It can be calculated for distribution with open end class.
• Can be used for qualitative data as well.
Demerits:
• It is not rigidly defined.
• It is not based on all observations
• It is not suitable for further mathematical treatment.
• It is not stable average, i.e. it is affected by fluctuations of sampling to some extent.
• Often its value is not unique.
[Link] Median
In a distribution, median is the value of the variable which divides the data in to two equal
halves. In an ordered series of data, the median is an observation lying exactly in the middle
of the series. It is the middle most value in the sense that the number of values less than the
median is equal to the number of values greater than it.

It is an average of location, not the average of the values in the data set and more affected by the
number of observations than the extreme value

27
Example

Median for grouped data: If data are given in the shape of continuous frequency
distribution, the median is defined as:
Median = Lm+ [(n/2) – cf-1)] w
f
Lm=lower limit of the median class
n=the number of observations
cf-1 =the cumulative frequency of class preceding the median class.
f =the frequency of median class
W=the class size/width
Note: The median class is the class with the smallest cumulative frequency greater than
or equal to n/2.

28
Example: Find the median wage of the following distribution

Step 1: First, we find out the total number of observations by summing up all the
frequencies.
Step 2: Then, we need to find the median class, i.e. the class having cumulative frequency
just greater than half of total number of observations.
Step 3: Now, we note the values of lower limit of median class (l), frequency of the median
class (f), cumulative frequency of the class preceding median class (cf-1), and class size
(W).
Step 4: Next, we can substitute these values in the formula

28 is the first cumulative frequency to be greater than or equal to 21.5 therefore 4000-5000 is
the median class

Median = Lm+ [(n/2) – cf-1)] w


f
4000+((43/2)-8) * 1000
20
4000+(21.5-8) *1000
20
4000+.675*1000

4000+675=4675

29
Merits and Demerits of Median
Merits:
• Median is a positional average and hence not influenced by extreme observations.
• Can be calculated in the case of open-end intervals.
• Median can be located even if the data are incomplete.
Demerits:
• It is not a good representative of data if the number of items is small.
• It is not amenable to further algebraic treatment.
• It is susceptible to sampling fluctuations.

Check point
1. Write the correct statistical representation of the following symbols

2. Find the median of the following distribution

Central Tendency and the Shape of the Distribution


A. Symmetrical Distributions
For a symmetrical distribution, the right-hand side of the graph is a mirror image of the left
hand side. If a distribution is perfectly symmetrical, the median is exactly at the center because
exactly half of the area in the graph will be on either side of the center.

30
B. Skewed Distributions
In skewed distributions, especially distributions for continuous variables, there is a strong tendency
for the mean, median, and mode to be located in predictably different positions. Figure 3.12(a), for
example, shows a positively skewed distribution with the peak(highest frequency) on the left-hand
side. Negatively skewed distributions are lopsided in the opposite direction, with the scores piling up
on the right-hand side and the tail tapering off to the left. The grades on an easy exam, for example,
tend to form a negatively skewed distribution (see Figure 3.12(b)).

[Link] Quantiles
The median gives us a value which divides the data set in to two equal parts. There are
also other positional measures that divide a given data set into more than two equal parts.
These measures are collectively known as quantiles. Quantiles include quartiles, deciles and
percentiles.

31
A. Quartiles
are some three points that divide the array in to four parts in away each portion contains equal
number of observations. The first, second and third points are called the first (Q1), second (Q2)
and third (Q3) quartiles respectively. 25% of the data fall below Q1, 50% below Q2 and 75%
below Q3 and Q1 ≤ Q2 ≤ Q3

B. Deciles
Are nine points that divide the array in to ten equal parts. The first, second, . . . , ninth deciles
are denoted by D1, D2, ..., D9 respectively. 10% of the data fall below D1, 20% below D2, . . . ,
90% below D9 and D1 ≤ D2 ≤ . . . ≤ D9

C. Percentiles
Are ninety nine points that divide the array in to 100 equal parts. They are denoted by P1,
P2, ..., P99. Always P1 ≤ P2 ≤ . . . ≤ P99

For ungrouped data

Example: Given the data 420, 430, 435, 438, 441, 449, 490, 500, 510 and 515. Find
(a) all quartiles.

32
(b) the 1st and 7th deciles.

c) the 40th and 75th percentiles.

For data in grouped frequency distribution.

where
Lqi, Ldi, Lpi are the lower-class boundaries of the classes containing the concerned quantile points,
Fqi-1, Fdi-1, Fpi-1 are the LCF of the class which precedes the class containing the concerned
quantile points,

33
fq i, fdi, fpi are frequencies of classes containing the concerned quantile points and
w is the class width of a class containing the concerned quantile point.
Example: Calculate all quartiles, the 5th and 8th deciles, and the 30th and 80th percentiles
for the students score data and interpret the results.

Check point
1. Calculate quartile 3, decile 5 and 65 percentile for the following data
A. 320, 330, 335, 438, 441, 449, 490, 550, 560 and 600.
B.

34
35
3.2.2 Measures of Dispersion (Variation)
he degree to which numerical data tend to spread about an average value is called dispersion
or variation of the data. Measures of dispersions are statistical measures which provide ways
of measuring the extent in which data are dispersed or spread out.

Variability provides a quantitative measure of the differences between scores in a


distribution and describes the degree to which the scores are spread out or clustered together.

Variability describes the distribution. Specifically, it tells whether the scores are clustered close
together or are spread out over a large distance. Usually, variability is defined in terms of distance. It
tells how much distance to expect between one score and another, or how much distance to expect
between an individual score and the mean.
Variability measures how well an individual score (or group of scores) represents the entire
distribution. This aspect of variability is very important for inferential statistics, in which relatively
small samples are used to answer questions about populations.

36
Objectives of measuring Variation:
• To judge the reliability of measures of central tendency
• To control variability itself.
• To compare two or more groups of numbers in terms of their variability
Types of Measures of Dispersion
A. Range (R)and Relative Range (RR)
B. Variance and Standard deviation
A. Range and Relative Range (RR)
Range is the simplest and crudest/rough measure of dispersion. The range is defined as the
difference between the highest and lowest scores in the distribution. In equation form,
Range for ungrouped data R= Highest score - Lowest score
The following two distributions have the same range, 13(45-32) yet appear to
differ greatly in the amount of variability.
Distribution 1: 32 35 36 36 37 38 40 42 42 43 43 45
Distribution 2: 32 32 33 33 33 34 34 34 34 34 35 45
For this reason, among others, the range is not the most important measure of variability.
Relative Range for ungrouped data

Range for grouped data


R =UCLlast − LCLfirst
where UCLlast is the last upper-class limit and LCLfirst is the first lower class limit.
Relative Range/coefficient of range for grouped data

Example 1
Find the R and RR for the following data
monthly income of 10 workers Xi: 347, 420, 500,600,696,710, 835, 850, and 900
R=900-420=480
RR=900-480 480
900+480 1380 .347

37
Example 2

R=35-6= 29
RR= 35-6 29
35+6 41 .707
Merits and Demerits of range
Merits:
• It is easy to calculate and simple to understand.
Demerits:
• It is not based on all observation.
• It is highly affected by extreme observations.
• It is affected by fluctuation in sampling.
• It is not liable to further algebraic treatment.
• It cannot be computed in the case of open-end distribution.
B. Variance and Standard deviation
Variance and standard deviation are the most superior and widely used measures of dispersions and
both measure the average dispersion of the observations around the mean. The variance of a data set is
the sum of the squares of the deviation of each observation taken from the mean divided by total number
of observations in the data set. The positive square root of variance is called standard deviation.

38
Examples 1
1. Consider a sample with data values of 10, 20, 12, 17, and 16. Compute the variance and standard
deviation.
Solution: We are expected to compute the sample mean ¯ x first since the sample variance is a
function the sample mean

Example 2
Calculate the variance and standard deviation for the following frequency distribution.

Solution:
The necessary calculation for calculating variance are as follows

39
Properties of Variance and Standard Deviation
1. If a constant is added (subtracted) to (from) each and every observation, the standard
deviation as well as the variance remains the same.
2. If each and every value is multiplied by a nonzero constant k, the standard deviation
is multiplied by k and the variance is multiplied by k2

1.2.3 Measure of relationship


Correlation is a statistical technique that can show whether and how strongly pairs of variables
are related. Correlation is a mathematical tool desired towards measuring the degree of linear
relationship (degree of association) between the variables.
The correlation coefficient is a value that indicates the strength of the relationship between
variables. The coefficient can take any values from -1 to 1. The interpretations of the values
are:

Correlation Description

Positive • If the Value in Variable (X) is high, the Corresponding Value of Variable (Y) is also
Correlation high. Similarly, If the Value in Variable (X) is Low, the Corresponding Value
of Variable (Y) is also Low. Then it is Positively Correlated.
• The Value of Correlation Coefficient (r) will be Positive.

Negative • If the Value in Variable (X) is high, the Corresponding Value of Variable (Y) is low.
Correlation Similarly, If the Value in Variable (X) is Low, the Corresponding Value of Variable
(Y) is also high. Then it is Negatively Correlated.
• The Value of Correlation Coefficient (r) will be Negative.

No Correlation • There will be no relationship between the two variables (X, Y).
• The Value of the Correlation Coefficient (r) will be Zero

40
The magnitude of the Pearson correlation coefficient determines the strength of the correlation.
Although there are no hard-and-fast rules for assigning strength of association to particular
values, some general guidelines are provided by Cohen (1988):

Coefficient Value Strength of Association

0.1 - .3 small correlation

0.3 - .5 medium/moderate correlation

| > .5 strong correlation

Correlation and Causation


Correlation must not be confused with causality. correlation does not mean causation. If two
variables are correlated, it does not imply that one variable causes the changes in another
variable. Correlation only assesses relationships between variables, and there may be different
factors that lead to the relationships. Causation may be a reason for the correlation, but it is not
the only possible explanation.

41
[Link] Types of Correlation

A. Simple correlation: Under simple correlation problem there are only two variables are studied.
B. Multiple Correlation: Under Multiple Correlation three or more than three variables are studied.
Ex. Qd = f ( P,PC, PS, t, y )
C. Partial correlation: analysis recognizes more than two variables but considers only two variables
keeping the other constant.
D. Linear correlation vs Non-Linear correlation
Linear correlation: Correlation is said to be linear when the amount of change in one variable tends
to bear a constant ratio to the amount of change in the other. The graph of the variables having a linear
relationship will form a straight line.

Example X = 1, 2, 3, 4, 5, 6, 7, 8,

Y = 5, 7, 9, 11, 13, 15, 17, 19,

Non-Linear correlation: The correlation would be nonlinear if the amount of change in one variable
does not bear a constant ratio to the amount of change in the other variable.
E. Total correlation: is based on all the relevant variables, which is normally not feasible.
[Link] Calculating Correlation Coefficient
A. Person’s product moment correlation coefficient
The Person’s correlation coefficient was developed by Karl Pearson in 1886.

42
Assumptions of Pearson’s Correlation Coefficient
✓ the two variables should be measured at the interval or ratio level
✓ There needs to be a linear relationship between the two variables.
✓ There should be no significant outliers.
✓ Data is normally distributed.
✓ the data has homoscedasticity (equal variances between groups)

43
B. Point biserial correlation (rPB)
The point biserial correlation, rpb, is the value of Pearson’s product moment correlation when one of the
variables is dichotomous, taking on only two possible values coded 0 and 1 (see Binary data), and the
other variable is metric (interval or ratio).
The dichotomous variable is the one that can be divided into two sharply distinguished or
mutually exclusive categories. Some examples are, male-female, rural-urban, Indian-American,
diagnosed with illness and not diagnosed with illness, Experimental group and Control Group,
college graduate vs. not a college graduate, first-born child vs. later-born child, success vs.
failure etc.

To compute the point-biserial correlation, the dichotomous variable is first converted to


numerical values by assigning a value of zero (0) to one category and a value of one (1) to the
other category.

• M1 = mean (for the entire test) of the group that received the positive binary
variable (i.e. the “1”).
• M0 = mean (for the entire test) of the group that received the negative binary
variable (i.e. the “0”).
• Sn = standard deviation for the entire test.
• p = Proportion of cases in the “0” group.
• q = Proportion of cases in the “1” group.

It is a data of 20 subjects, out of which 9 are male and 11 are females. Their marks in the final
examination are also provided. We want to correlate marks in the final examination with sex
of the subject. The marks obtained in the final examination are a continuous variable whereas

44
sex is truly dichotomous variable, taking two values male or female. We are using value of 0
for male subject and value of 1 for female subjects. The correlation appropriate for this
purpose is Point-Biserial correlation (rpb).

C. Biserial correlation
It is like the pointbiserial correlation but point-biserial correlation is computed while one of the
variables is dichotomous and do not have any underlying continuity. If a variable has
underlying continuity but measured dichotomously, then the biserial correlation can be
calculated.

An example might be mood (happy-sad) and (low vs. normal mood). Actually, it is fair to
assume that mood is a normally distributed variable.

But this variable is measured discretely and takes only two values, low mood (0) and
normal mood (1).
So biserial correlation is a correlation coefficient between two continuous variables
(X and Y), out of which one is measured dichotomously (X). The formula is very
similar to the point-biserial but yet different:

Where Y0 and Y1 are the Y score means for data pairs with an X score of 0 and 1, respectively,
P0 and P1 are the proportions of data pairs with X scores of 0 and 1,
respectively, and SY is the standard deviation for the Y data, and h is ordinate or the
height of the standard normal distribution at the point which divides the proportions
of P0 and P1. The relationship between the point-biserial and the biserial correlation is as
follows.
D. Spearman’s rank-order correlation or spearman’s rho (rs)
A well-known psychologist and intelligence theorist, Charles Spearman (1904),
developed a correlation procedure called in his honor as Spearman’s rank-order correlation or
Spearman’s rho (rs). It was developed to compute correlation when the data is presented on
two variables for n subjects. It can also be calculated for data of n subjects evaluated by two
judges for inter-judge agreement. It is suitable for the rank-order data. If the data on X or Y or

45
on both the variables are in rank order then Spearman’s rho is applicable. It is used to assess a
monotonic relationship.

r = Spearman’s rank-order correlation


s

D = difference between the pair of ranks of X and Y


n = the number of pairs of ranks
E. PHI coefficient (φ)
When both the variables are dichotomous, then the Pearson’s correlation calculated is called as

Phi Coefficient (φ). For example, let us say that you have to compute correlation between

gender and ownership of the property. The gender takes two levels, male and female. The
ownership of property can be measured as either the person owns a property and the person do
not own property.

[Link] Coefficient of Determination


The coefficient of determination is a statistical measurement that examines how differences in one
variable can be explained by the difference in a second variable, when predicting the outcome of a
given event. In other words, this coefficient, which is more commonly known as R-squared (or R2),
assesses how strong the linear relationship is between two variables, and is heavily relied on by
researchers when conducting trend analysis.

✓ The convenient way of interpreting the value of correlation coefficient is to use of square of
coefficient of correlation which is called Coefficient of Determination.
✓ The Coefficient of Determination = r2.
✓ Suppose r = 0.9, r2 = 0.81 this would mean that 81% of the variation in the dependent variable
has been explained by the independent variable.
✓ The maximum value of r2 is 1 because it is possible to explain all of the variation in y but it
is not possible to explain more than all of it.
✓ Coefficient of Determination = Explained variation / Total variation

46
Example
Suppose: r = 0.60 r = 0.30 It does not mean that the first correlation is twice as strong as the
second the ‘r’ can be understood by computing the value of r2 .

When r = 0.60 r2 = 0.36 -----(1)

r = 0.30 r2 = 0.09 -----(2)

This implies that in the first case 36% of the total variation is explained whereas in second
case 9% of the total variation is explained.

[Link] Zero-order, Partial, and Part Correlations


You also may have come across the terms “zero-order,” “partial,” and “part” in reference to
correlations. These terms refer to correlations that involve more than two variables. More
specifically, these types of correlations are relevant when you have a dependent (outcome)
variable, an independent (explanatory) variable, and one or more confounding (control)
variables. Here we will explain the differences between zero-order, partial, and part
correlations.
A. Zero-order correlation
A zero-order correlation simply refers to the correlation between two variables (i.e., the
independent and dependent variable) without controlling for the influence of any other
variables. Essentially, this means that a zero-order correlation is the same thing as a Pearson
correlation. So why are we discussing the zero-order correlation here? When conducting an
analysis with more than two variables (i.e., multiple independent variables or control variables),
it may be of interest to know the simple bivariable relationships between the variables to get a
better sense of what happens when you begin to control for other variables. This is why SPSS
gives you the option to report zero-order correlations when running a multiple linear regression
analysis.
B. Partial correlation (rP)
The Partial Correlations procedure computes partial correlation coefficients that describe the
linear relationship between two variables while controlling for the effects of one or more
additional variables. Partial correlation is the correlation between two variables after removing the
effect of one or more additional variables. In a partial correlation, the influence of the control
variables on both the independent and dependent variables are taken into account. Suppose we
want to find the correlation between y and x controlling by W. This is called the partial correlation
and its symbol is rYX.W . For example, the researcher is interested in computing the correlation between

47
anxiety and academic achievement controlled from intelligence. Then correlation between academic
achievement (A) and anxiety (B) will be controlled for Intelligence (C).

C. Part correlation
The part correlation, which is sometimes referred to as the “semipartial” correlation. Like the
partial correlation, the part correlation is the correlation between two variables (independent
and dependent) after controlling for one or more other variables. However, for the part
correlation, only the influence of the control variables on the independent variable is taken
into account. In other words, the part correlation does not control for the influence of the
confounding variables on the dependent variable. You might wonder why you would only
want to control for effects on the independent variable and not the dependent variable? The
primary reason for conducting the part correlation would be to see how much unique variance
the independent variable explains in relation to the total variance in the dependent variable,
rather than just the variance unaccounted for by the control variables.

Check point
Calculate Pearson product moment correlation for the following data
X Y
4 2
6 5
7 6
10 9
14 15
9 4
2 6

48

You might also like