0% found this document useful (0 votes)
14 views42 pages

Unit II Descriptive Analytics

Unit II covers descriptive analytics, focusing on frequency distributions, types of data (qualitative, ranked, quantitative), and levels of measurement (nominal, ordinal, interval/ratio). It details the construction of frequency distributions and guidelines for their creation, including the importance of class intervals and handling outliers. Additionally, it discusses variables, including independent and dependent variables, and introduces observational studies and relative frequency distributions.

Uploaded by

Bastin Rogers
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
14 views42 pages

Unit II Descriptive Analytics

Unit II covers descriptive analytics, focusing on frequency distributions, types of data (qualitative, ranked, quantitative), and levels of measurement (nominal, ordinal, interval/ratio). It details the construction of frequency distributions and guidelines for their creation, including the importance of class intervals and handling outliers. Additionally, it discusses variables, including independent and dependent variables, and introduces observational studies and relative frequency distributions.

Uploaded by

Bastin Rogers
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

UNIT II - DESCRIPTIVE ANALYTICS

Frequency distributions – outliers –interpreting distributions – graphs – averages - describing variability –


interquartile range – variability for qualitative and ranked data - Normal distributions – z scores

TYPES OF DATA (NOT INCLUDED IN SYLLABUS)

Data

A collection of actual observations or scores in a survey or an experiment is termed as data. Data can be
categorised into three types. They are qualitative data, ranked data, or quantitative data.

Qualitative Data

A set of observations where any single observation is a word, letter, or numerical code that represents a
class or category is termed as qualitative data. Qualitative data consist of words (Yes or No), letters (Y or
N), or numerical codes (0 or 1) that represent a class or category.

Ranked data

A set of observations where any single observation is a number that indicates relative standing is known
as ranked data. Ranked data consist of numbers (1st, 2nd, . . . 40th place) that represent relative standing
within a group.

Quantitative Data

A set of observations where any single observation is a number that represents an amount or a count is
known as quantitative data. Quantitative data consist of numbers (weights of 238, 170, …., 185 lbs) that
represent an amount or a count.

For example, the weights reported by 53 male students in Table 1.1 are quantitative data, since any
single observation, such as 160 lbs, represents an amount of weight. If the weights in Table 1.1 had been
replaced with ranks, beginning with a rank of 1 for the lightest weight of 133 lbs and ending with a rank
of 53 for the heaviest weight of 245 lbs, these numbers would have been ranked data, since any single
observation represents not an amount, but only relative standing within the group of 53 students.

Finally, the Y and N replies of students in Table 1.2 are qualitative data, since any single observation is a
letter that represents a class of replies. The following figure shows the overview of data and levels of
measurement.

Levels of measurement:

The level of measurement specifies the extent to which a number (or word or letter) actually
represents some attribute and, therefore, has implications for the appropriateness of various arithmetic
operations and statistical procedures. There are three levels of measurement—nominal, ordinal, and
interval/ratio—and these levels are paired with qualitative, ranked, and quantitative data, respectively.
Qualitative Data and Nominal Measurement

The first level of measurement is nominal level of measurement. In this level of measurement, the
numbers in the variable are used only to classify the data. In this level of measurement, words, letters,
and alpha-numeric symbols can be used. Suppose there are data about people belonging to different
gender categories. In this case, the person belonging to the female gender could be classified as F and the
person belonging to the male gender could be classified as M. This type of assigning classification is
nominal level of measurement.

Ranked Data and Ordinal Measurement

The second level of measurement is the ordinal level of measurement. This level of measurement
depicts some ordered relationship among the observations. Suppose a student scores the highest mark of
100 in the class. In this case, he would be assigned the first rank. Then, another classmate scores the
second highest mark of 92; she would be assigned the second rank. A third student scores 81 and he
would be assigned the third rank, and so on. The ordinal level of measurement indicates an ordering of
the measurements.

Quantitative Data and interval/ratio Measurement

The third level of measurement is the interval level of measurement. The interval level of
measurement not only classifies and orders the measurements, but it also specifies that the distances
between each interval on the scale are equivalent along the scale from low interval to high interval. A
popular example of this level of measurement is temperature in centigrade, where, for example, the
distance between 940C and 960C is the same as the distance between 1000C and 1020C.

In this level of measurement, the observations, in addition to having equal intervals, can have a
value of zero as well. The zero in the scale makes this type of measurement unlike the other types of
measurement, although the properties are similar to that of the interval level of measurement. Since true
zero is possible, numerical readings reflect the total amount of a person’s weight, and it’s appropriate to
describe one person’s weight as a certain ratio of another’s. It can be said that the weight of a 140 kg
person is twice that of a 70 kg person.

_____________________________________________________________________________________

TYPES OF VARIABLES

A variable is a characteristic or property that can take on different values. A constant is a


characteristic or property that can take on only one value.

Discrete and Continuous Variables


Quantitative variables can be further distinguished in terms of whether they are discrete or
continuous.

Discrete Variables

A discrete variable consists of isolated numbers separated by gaps. Examples include most counts,
such as the number of children in a family, the number of foreign countries you have visited etc.…

Continuous Variables

A continuous variable consists of numbers whose values have no restrictions. Examples include
amounts, such as weights of students, speed of a car etc.…

Approximate Numbers

Whenever values are rounded off, as is always the case with actual values for continuous variables, the
resulting numbers are approximate, never exact. For example, the weights of the male students are
approximate because they have been rounded to the nearest pound. Someone’s weight, in pounds, might
be 140.01438, and so on, to infinity! For practical consideration it should be rounded off.

Independent variable

An independent variable is the variable which is manipulated or varied in an experimental study to


explore its effects.

It’s called “independent” because it’s not influenced by any other variables in the study.

Independent variables are also called:

• Explanatory variables (they explain an event or outcome)


• Predictor variables (they can be used to predict the value of a dependent variable)
• Right-hand-side variables (they appear on the right-hand side of a regression equation).

Dependent variable

A dependent variable is the variable that changes as a result of the independent variable manipulation. It
“depends” on independent variable.

In statistics, dependent variables are also called:

• Response variables (they respond to a change in another variable)


• Outcome variables (they represent the outcome you want to measure)
• Left-hand-side variables (they appear on the left-hand side of a regression equation)
The dependent variable is a variable recorded after the manipulation of the independent variable.

This measurement data is used to check whether and to what extent independent variable influences the
dependent variable by conducting statistical analyses.

Confounding variables are a type of peripheral variable that are related to a study’s independent and
dependent variables. A variable must meet two conditions to be a confounder:

1. It must be correlated with the independent variable. This may be a causal relationship, but it does
not have to be.
2. It must be causally related to the dependent variable.

Example of a confounding variable

You collect data on sunburns and ice cream consumption. You find that higher ice cream consumption is
associated with a higher probability of sunburn. Does that mean ice cream consumption causes sunburn?

Here, the confounding variable is temperature: hot temperatures cause people to both eat more ice cream
and spend more time outdoors under the sun, resulting in more sunburn.

Observational Studies

An observational study focuses on detecting relationships between variables not manipulated by the
investigator, and it yields less clear-cut conclusions about cause effect relationships than does an
experiment.

_____________________________________________________________________________________

SYLLABUS TOPICS

DESCRIBING DATA WITH TABLES AND GRAPHS

Tables (Frequency Distributions)

Frequency Distributions For Quantitative Data

The table (right) shows one way to organize the performance of students rated in the range of 1 to 10
listed as in table (left). First, arrange a column of consecutive numbers, beginning with the poorest
performance (1) at the bottom and ending with the best performance (10) at the top. Then place a short
vertical stroke or tally next to a number each time its value appears in the original set of data; once this
process has been completed, substitute for each tally count a number indicating the frequency (f) of
occurrence of each value.
A frequency distribution is a collection of observations produced by sorting observations into
classes and showing their frequency (f) of occurrence in each class.

When observations are sorted into classes of single values, the result is referred to as a frequency
distribution for ungrouped data.

Grouped Data:

Table 2.2 shows another way to organize the weights in Table


1.1 according to their frequency of occurrence. When
observations are sorted into classes of more than one value, as
in Table 2.2, the result is referred to as a frequency distribution
for grouped data. Let’s look at the general structure of this
frequency distribution. Data are grouped into class intervals
with 10 possible values each. The bottom class includes the
smallest observation (133), and the top class includes the
largest observation (245). The distance between bottom and top
is occupied by an orderly series of classes. The frequency (f)
column shows the frequency of observations in each class and,
at the bottom, the total number of observations in all classes.

Guidelines

The “Guidelines for Frequency Distributions” box lists seven rules for producing a well-constructed
frequency distribution. The first three rules are essential and should not be violated. The last four rules are
optional and can be modified or ignored.
Guidelines for Frequency Distributions

Essential

1. Each observation should be included in one, and only one, class.

Example: 130–139, 140–149, 150–159, etc. It would be incorrect to use 130–140, 140–150, 150–
160, etc., in which, because the boundaries of classes overlap, an observation of 140 (or 150) could be
assigned to either of two classes.

2. List all classes, even those with zero frequencies.

Example: Listed in Table 2.2 is the class 210–219 and its frequency of zero. It would be incorrect
to skip this class because of its zero frequency.

3. All classes should have equal intervals.

Example: 130–139, 140–149, 150–159, etc. It would be incorrect to use 130–139, 140–159, etc.,
in which the second class interval (140–159) is twice as wide as the first class interval (130–139).

Optional

4. All classes should have both an upper boundary and a lower boundary.

Example: 240–249. Less preferred would be 240–above, in which no maximum value can be assigned to
observations in this class.

5. Select the class interval from convenient numbers, such as 1, 2, 3 . . . 10, particularly 5 and 10 or
multiples of 5 and 10.

Example: 130–139, 140–149, in which the class interval of 10 is a convenient number. Less
preferred would be 130–142, 143–155, etc., in which the class interval of 13 is not a convenient number.

6. The lower boundary of each class interval should be a multiple of the class interval.

Example: 130–139, 140–149, in which the lower boundaries of 130, 140, are multiples of 10, the
class interval. Less preferred would be 135–144, 145–154, etc., in which the lower boundaries of 135 and
145 are not multiples of 10, the class interval.

7. Aim for a total of approximately 10 classes.

Example: The distribution in Table 2.2 uses 12 classes. Less preferred would be the distributions
in Tables 2.3 and 2.4. The distribution in Table 2.3 has too many classes (24), whereas the distribution in
Table 2.4 has too few classes (3).
Gaps between Classes

In well-constructed frequency tables, the gaps between classes, such as between 149 and 150 in Table 2.2,
show clearly that each observation or score has been assigned to one, and only one, class.

The size of the gap should always equal one unit of measurement; that is, it should always equal the
smallest possible difference between scores within a particular set of data. Since the gap is never bigger
than one unit of measurement, no score can fall into the gap

Real Limits of Class Intervals

Gaps cannot be ignored when you are determining the actual width of any class interval. The real limits
are located at the midpoint of the gap between adjacent tabled boundaries; that is, one-half of one unit of
measurement below the lower tabled boundary and one-half of one unit of measurement above the upper
tabled boundary.

For example, the real limits for 140–149 in Table 2.2 are 139.5 (140 minus one-half of the unit of
measurement of 1) and 149.5 (149 plus one-half of the unit of measurement of 1), and the actual width of
the class interval would be 10 (from 149.5 - 139.5 = 10).
Constructing Frequency Distributions

Consider the table 1.1 and construct the frequency distributions using the following steps.

1. Find the range, that is, the difference between the largest and smallest observations. The range of
weights in Table 1.1 is 245 133 = 112.
2. Find the class interval required to span the range by dividing the range by the desired number of
classes (ordinarily 10). In the present example

𝑟𝑎𝑛𝑔𝑒 112
𝐶𝑙𝑎𝑠𝑠 𝐼𝑛𝑡𝑒𝑟𝑣𝑎𝑙 = = = 11.2
𝑑𝑒𝑠𝑖𝑟𝑒𝑑 𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑐𝑙𝑎𝑠𝑠𝑒𝑠 10

3. Round off to the nearest convenient interval (such as 1, 2, 3 . . . 10, particularly 5 or 10 or


multiples of 5 or 10). In the present example, the nearest convenient interval is 10.
4. Determine where the lowest class should begin. (Ordinarily, this number should be a multiple of
the class interval.) In the present example, the smallest score is 133, and therefore the lowest class
should begin at 130, since 130 is a multiple of 10 (the class interval).
5. Determine where the lowest class should end by adding the class interval to the lower boundary
and then subtracting one unit of measurement. In the present example, add 10 to 130 and then
subtract 1, the unit of measurement, to obtain 139—the number at which the lowest class should
end.
6. Working upward, list as many equivalent classes as are required to include the largest observation.
In the present example, list 130–139, 140–149. . . 240–249, so that the last class includes 245, the
largest score.
7. Indicate with a tally the class in which each observation falls. For example, the first score in Table
1.1, 160, produces a tally next to 160–169; the next score, 193, produces a tally next to 190–199;
and so on.
8. Replace the tally count for each class with a number—the frequency (f )—and show the total of all
frequencies. (Tally marks are not usually shown in the final frequency distribution.)
9. Supply headings for both columns and a title for the table.
Outliers

An outlier is a data point that differs significantly from other observations. An outlier may be due
to variability in the measurement or it may indicate experimental error; the latter are sometimes excluded
from the data set. An outlier can cause serious problems in statistical analyses.

Relative Frequency Distributions

An important variation of the frequency distribution is the relative frequency distribution.

Relative frequency distributions show the frequency of each class as a part or fraction of the total
frequency for the entire distribution.

This type of distribution allows us to focus on the relative concentration of observations among different
classes within the same distribution.

Constructing Relative Frequency Distributions

To convert a frequency distribution into a relative frequency distribution, divide the frequency for
each class by the total frequency for the entire distribution. Table 2.5 illustrates a relative frequency
distribution based on the weight distribution of Table 2.2.

The conversion to proportions is straightforward. For instance, to obtain the proportion of .06 for
the class 130–139, divide the frequency of 3 for that class by the total frequency of 53. Repeat this
process until a proportion has been calculated for each class.
To convert the relative frequencies in Table 2.5 from proportions to percentages, multiply each
proportion by 100; that is, move the decimal point two places to the right. For example, multiply .06 (the
proportion for the class 130–139) by 100 to obtain 6 percent.

Cumulative Frequency Distributions

Cumulative frequency distributions show the total number of observations in each class and in
all lower-ranked classes.

This type of distribution can be used effectively with sets of scores, such as test scores for
intellectual or academic aptitude, when relative standing within the distribution assumes primary
importance. Under these circumstances, cumulative frequencies are usually converted, in turn, to
cumulative percentages. Cumulative percentages are often referred to as percentile ranks.

Constructing Cumulative Frequency Distributions

To convert a frequency distribution into a cumulative frequency distribution, add to the frequency
of each class the sum of the frequencies of all classes ranked below it. This gives the cumulative
frequency for that class.

Begin with the lowest-ranked class in the frequency distribution and work upward, finding the
cumulative frequencies in ascending order.

In Table 2.6, the cumulative frequency for the class 130–139 is 3, since there are no classes ranked
lower. The cumulative frequency for the class 140–149 is 4, since 1 is the frequency for that class and 3 is
the frequency of all lower-ranked classes. The cumulative frequency for the class 150–159 is 21, since 17
is the frequency for that class and 4 is the sum of the frequencies of all lower-ranked classes.

Cumulative Percentages
If relative standing within a distribution is particularly important, then cumulative frequencies are
converted to cumulative percentages.

A glance at Table 2.6 reveals that 75 percent of all weights are the same as or lighter than the
weights between 170 and 179 lbs. To obtain this cumulative percentage (75%), the cumulative frequency
of 40 for the class 170–179 should be divided by the total frequency of 53 for the entire distribution.

Percentile Ranks

When used to describe the relative position of any score within its parent distribution, cumulative
percentages are referred to as percentile ranks.

The percentile rank of a score indicates the percentage of scores in the entire distribution with
similar or smaller values than that score. Thus a weight has a percentile rank of 80 if equal or lighter
weights constitute 80 percent of the entire distribution.

The percentile formula determines the performance of a person over others. The percentile
formula is used in finding where a student stands in the test compared to other candidates. A percentile is
a number where a certain percentage of scores fall below the given number.

The percentile of x is the ratio of the number of values below x to the total number of values
multiplied by 100. i.e., the percentile formula is

Percentile = (Number of Values below “x” / Total Number of Values) × 100

Approximate Percentile Ranks (from Grouped Data)

The assignment of exact percentile ranks requires that cumulative percentages be obtained from
frequency distributions for ungrouped data. If we have access only to a frequency distribution for grouped
data, as in Table 2.6, cumulative percentages can be used to assign approximate percentile ranks. In Table
2.6, for example, any weight in the class 170–179 could be assigned an approximate percentile rank of 75,
since 75 is the cumulative percent for this class.

Frequency Distributions for Qualitative (Nominal) Data

When, among a set of observations, any single observation is a word, letter, or numerical code, the data
are qualitative. Frequency distributions for qualitative data are easy to construct. Simply determine the
frequency with which observations occupy each class, and report these frequencies as shown in Table 2.7
for the Facebook profile survey (Whether person is possessing Facebook account or not).
Ordered Qualitative Data

Qualitative data have an ordinal level of measurement because observations can be ordered from
least to most, that order should be preserved in the frequency table, as illustrated in Table 2.8, in which
military ranks are listed in descending order from general to lieutenant.

Relative and Cumulative Distributions for Qualitative Data

Frequency distributions for qualitative variables can always be converted into relative frequency
distributions, as illustrated in Table 2.8. Furthermore, if measurement is ordinal because observations can
be ordered from least to most, cumulative frequencies (and cumulative percentages) can be used.

For example, that a captain has an approximate percentile rank of 63 among officers since 62.5 (or
63) is the cumulative percent for this class. If measurement is only nominal because observations cannot
be ordered, as in Table 2.7, a cumulative frequency distribution is meaningless.

If measurement is only nominal because observations cannot be ordered, as in Table 2.7, a


cumulative frequency distribution is meaningless.

Graphs

Data can be described clearly and concisely with the aid of a well-constructed frequency
distribution. And data can often be described even more vividly, particularly to communicate with a
general audience, by converting frequency distributions into graphs.
Graphs for Quantitative Data

Histograms

Histogram is a bar-type graph for quantitative data. The common boundaries between adjacent
bars emphasize the continuity of the data, as with continuous variables.

The following histogram shows the weight distribution of male students. A casual glance at this
histogram confirms previous conclusions: a dense concentration of weights among the 150s, 160s, and
170s, with a spread in the direction of the heavier weights.

• Equal units along the horizontal axis (the X axis, or abscissa) reflect the various class intervals of
the frequency distribution.
• Equal units along the vertical axis (the Y axis, or ordinate) reflect increases in frequency. (The
units along the vertical axis do not have to be the same width as those along the horizontal axis.)
• The intersection of the two axes defines the origin at which both numerical scales equal 0.
• Numerical scales always increase from left to right along the horizontal axis and from bottom to
top along the vertical axis.
• The body of the histogram consists of a series of bars whose heights reflect the frequencies for
the various classes.

Frequency Polygon

An important variation on a histogram is the frequency polygon, or line graph. Frequency


polygons may be constructed directly from frequency distributions. The step-by-step transformation of a
histogram into a frequency polygon is described in panels A, B, C, and D of Figure 2.2.

A. This panel shows the histogram for the weight distribution.

B. Place dots at the midpoints of each bar top or, in the absence of bar tops, at midpoints for
classes on the horizontal axis, and connect them with straight lines.[To find the midpoint of any
class, such as 160–169, simply add the two tabled boundaries (160 + 169 = 329) and divide this
sum by 2 (329/2 = 164.5).]

C. Anchor the frequency polygon to the horizontal axis. First, extend the upper tail to the midpoint
of the first unoccupied class (250–259) on the upper flank of the histogram. Then extend the lower
tail to the midpoint of the first unoccupied class (120–129) on the lower flank of the histogram.
Now all of the area under the frequency polygon is enclosed completely.

D. Finally, erase all of the histogram bars, leaving only the frequency polygon.

Frequency polygons are particularly useful when two or more frequency distributions or relative
frequency distributions are to be included in the same graph.

Stem and Leaf Displays

Another technique for summarizing quantitative data is a stem and leaf display. It is a device for
sorting quantitative data on the basis of leading and trailing digits. Stem and leaf displays are ideal for
summarizing distributions without destroying the identities of individual observations.

Constructing a Display
The leftmost panel of Table 2.9 re-creates the weights of the 53 male statistics students listed in
Table 1.1.

• To construct the stem and leaf display for these data, first note the weights range. i.e. from
the 130s to the 240s.
• Arrange a column of numbers, the stems, beginning with 13 (representing the 130s) and
ending with 24 (representing the 240s).
• Draw a vertical line to separate the stems, which represent multiples of 10, from the space
to be occupied by the leaves, which represent multiples of 1.
• Next, enter each raw score into the stem and leaf display. As suggested by the shaded
coding in Table 2.9, the first raw score of 160 reappears as a leaf of 0 on a stem of 16.
• The next raw score of 193 reappears as a leaf of 3 on a stem of 19, and the third raw score
of 226 reappears as a leaf of 6 on a stem of 22, and so on, until each raw score reappears as
a leaf on its appropriate stem.

The weight data have been sorted by the stems. All weights in the 130s are listed together; all of
those in the 140s are listed together, and so on. This simple method only works if, as in the present
display, stem values are listed from smallest at the top to largest at the bottom.

Selection of Stems

Stem values are not limited to units of 10. Depending on the data stem value can be varied as 1,
100, 1000, or even .1, .01, .001, and so on. For instance, an annual income of $23,784 could be displayed
as a stem of 23 (thousands) and a leaf of 784. (Leaves consisting of two or more digits, such as 784, are
separated by commas.)

Typical Shapes
An important characteristic of a frequency distribution is its shape. Figure 2.3 shows some of the
more typical shapes for smoothed frequency polygons.

Normal Distribution

A normal distribution, sometimes called the bell curve is a distribution that occurs naturally in
many situations. For example, the bell curve is seen in tests like the SAT and GRE. The bulk of students
will score the average (C), while smaller numbers of students will score a B or D. An even smaller
percentage of students score an F or an A. This creates a distribution that resembles a bell.

Bimodal

Data distributions in statistics can have one peak, or they can have several peaks. The type of
distribution you might be familiar with seeing is the normal distribution, or bell curve, which has one
peak. The bimodal distribution has two peaks.

Example: Scores fall into bimodal distribution with one group of students getting scores between 70 to
75 marks out of 100 and another group of students getting scores between 25 to 30 marks.

Positively Skewed & Negatively skewed Distribution (lopsided distribution)

• A lopsided distribution that includes a few extreme observations in the positive direction (to the
right of the majority of observations) is positively skewed.
• A lopsided distribution caused by a few extreme observations in the negative direction (to the left
of the majority of observations) is negatively skewed.

A Graph for Qualitative (Nominal) Data


The distribution in Table 2.7, based on replies to the question “Do you have a Facebook profile?”
appears as a bar graph in Figure 2.4.

A glance at this graph confirms that Yes replies occur approximately twice as often as No replies.

As with histograms, equal segments along the horizontal axis are allocated to the different words
or classes that appear in the frequency distribution for qualitative data. Likewise, equal segments along
the vertical axis reflect increases in frequency. The body of the bar graph consists of a series of bars
whose heights reflect the frequencies for the various words or classes.

Gaps are placed between adjacent bars of bar graphs to emphasize the discontinuous nature of
qualitative data.

Constructing Graphs

1. Decide on the appropriate type of graph.


2. Draw the horizontal axis, then the vertical axis, remembering that the vertical axis should be about
as tall as the horizontal axis is wide.
3. Identify the string of class intervals that eventually will be superimposed on the horizontal axis.
4. Superimpose the string of class intervals (with gaps for bar graphs) along the entire length of the
horizontal axis.
5. Along the entire length of the vertical axis, superimpose a progression of convenient numbers,
beginning at the bottom with 0 and ending at the top with a number as large as or slightly larger
than the maximum observed frequency. If there is a considerable gap between the origin of 0 and
the smallest observed frequency, use wiggly lines to signal a break in scale.
6. Using the scaled axes, construct bars to reflect the frequency of observations within each class
interval.
7. Supply labels for both axes and a title for the graph.

___________________________________________________________________________________
Describing Data with Averages Averages consist of numbers (or words) about which the data are, in
some sense, centred. They are often referred to as measures of central tendency, the several types of
average yield numbers or words that attempt to describe, most generally, the middle or typical value for a
distribution.

Three different measures of central tendency

• Mode
• Median
• Mean

Mode

The mode reflects the value of the most frequently occurring score.

A mode is defined as the value that has a higher frequency in a given set of values. It is the value that
appears the most number of times.
Example: In the given set of data: 2, 4, 5, 5, 6, 7, the mode of the data set is 5 since it has appeared in the
set twice.

More Than One Mode

Distributions can have more than one mode (or no mode at all). Distributions with two obvious peaks,
even though they are not exactly the same height, are referred to as bimodal.

Distributions with more than two peaks are referred to as multimodal.

Median

The median reflects the middle value when observations are ordered from least to most.

The median splits a set of ordered observations into two equal parts, the upper and lower halves.
In other words, the median has a percentile rank of 50, since observations with equal or smaller values
constitute 50 percent of the entire distribution.

Finding the Median

Table 3.2 shows how to find the median for two different sets of scores. The following steps are followed
to find the median.

1. Order scores from least to most.


2. Find the middle position by adding one to the total number of scores and dividing by 2.
3. If the middle position is a whole number, as in the left-hand panel below, use this number to count
into the set of ordered scores.
4. The value of the median equals the value of the score located at the middle position.
5. If the middle position is not a whole number, as in the right-hand panel below, use the two nearest
whole numbers to count into the set of ordered scores.
6. The value of the median equals the value midway between those of the two middlemost scores; to
find the midway value, add the two given values and divide by 2.

The median term describes the middle-ranked term; the modal term describes the most
frequent term in the distribution.

Mean

The mean is the most common average, one you have doubtless calculated many times. The mean is
found by adding all scores and then dividing by the number of scores.

That is,

sum of all scores


Mean =
number of scores

Example:

In a class there are 20 students and they have secured a percentage of 88, 82, 88, 85, 84, 80, 81, 82, 83,
85, 84, 74, 75, 76, 89, 90, 89, 80, 82, and 83. Find the mean percentage obtained by the class.

Solution: Mean = Total of percentage obtained by 20 students in class/Total number of students

= [88 + 82 + 88 + 85 + 84 + 80 + 81 + 82 + 83 + 85 + 84 + 74 + 75 + 76 + 89 + 90
+ 89 + 80 + 82 + 83]/20
= 1660/20
= 83
Hence, the mean percentage of each student in the class is 83%.

Statisticians distinguish between two types of means—the population mean and the sample
mean—depending on whether the data are viewed as a population (a complete set of scores) or as a
sample (a subset of scores).

Formula for Sample Mean

It’s usually more efficient to substitute symbols for words in statistical formulas, including the
word formula given above for the mean. When symbols are used, 𝑋̅ designates the sample mean, and the
formula becomes

∑𝑋
𝑋̅ =
𝑛

Formula for Population Mean

The formula for the population mean differs from that for the sample mean only because of a
change in some symbols. In statistics, Greek symbols usually describe population characteristics, such as
the population mean, while English letters usually describe sample characteristics, such as the sample
mean. The population mean is represented by μ (pronounced “mu”), the lowercase Greek letter m for
mean,

∑𝑋
𝜇=
𝑁

where the uppercase letter N refers to the population size. Otherwise, the calculations are the same as
those for the sample mean.

The mean reflects the values of all scores, not just those that are middle ranked (as with the median),
or those that occur most frequently (as with the mode).

Averages for Qualitative and Ranked Data

Mode Always Appropriate for Qualitative Data For quantitative data all the three averages can be
used. But when the data are qualitative, the choice among averages is restricted. The mode always can be
used with qualitative data. For instance, Yes qualifies as the modal or most typical response for the
Facebook profile question (Table 2.7).
Median Sometimes Appropriate

The median can be used whenever it is possible to order qualitative data from least to most
because the level of measurement is ordinal.

It’s easiest to determine the median class for ordered qualitative data by using relative frequencies, as in
Table 3.5.

Cumulate the relative frequencies, working up from the bottom of the distribution, until the
cumulative percentage first equals or exceeds 50 percent. Since the corresponding class includes the
median and, roughly speaking, splits the distribution into an upper and a lower half, it is designated as the
median or middle-ranked class.

For example, the qualitative data in Table 3.5 can be ordered from lieutenant to general. Starting
at the bottom of Table 3.5 and cumulating upward, 25. 5 percent is obtained for the class of lieutenant and
a cumulative percent 62.5 for the class of captain. Accordingly, since it includes a cumulative percent of
50, captain is the median rank of officers in the U.S. Army.

Inappropriate Averages

It would not be appropriate to report a median for unordered qualitative data with nominal
measurement. It is inappropriate to report a mean for any qualitative data, such as the ranks of officers in
the U.S. Army. After all, words cannot be added and then divided, as required by the formula for the
mean.
Describing Variability

Measures of variability, that is, measures of the amount by which scores are dispersed or scattered
in a distribution.

1. Standard deviation
2. Variance
3. Range
4. Interquartile range

In Figure 4.1, each of the three frequency distributions consists of seven scores with the same
mean (10) but with different variabilities like distribution A has the least variability, distribution B has
intermediate variability, and distribution C has the most variability.

Range

One of the measures of variability and essential tools in statistics is the range. The range is the
difference between the largest and smallest scores.

In Figure 4.1, distribution A, the least variable, has the smallest range of 0 (from 10 to 10);
distribution B, the moderately variable, has an intermediate range of 2 (from 11 to 9); and distribution C,
the most variable, has the largest range of 6 (from 13 to 7), in agreement with our intuitive judgments
about differences in variability.

The range is a handy measure of variability that can readily be calculated and understood.

Shortcomings

1. First, since its value depends on only two Scores —the largest and the smallest—it fails to use the
information provided by the remaining scores.
2. The value of the range tends to increase with increases in the total number of scores.
Variance

The variance and particularly its square root, the standard deviation, serve as key components for
other important statistical measures.

The term variance refers to a statistical measurement of the spread between numbers in a data set.
More specifically, variance measures how far each number in the set is from the mean (average), and thus
from every other number in the set.

Variance is often depicted by this symbol: σ2

Reconstructing the Variance

In the case of the variance, each original score is re-expressed as a distance or deviation from the
mean by subtracting the mean.

For each of the three distributions in Figure 4.1, the face values of the seven original scores
(shown as numbers along the X axis) have been re-expressed as deviation scores from their mean of 10
(shown as numbers in the boxes). For example, in distribution C, one score coincides with the mean of
10, four scores (two 9s and two 11s) deviate 1 unit from the mean, and two scores (one 7 and one 13)
deviate 3 units from the mean, yielding a set of seven deviation scores: one 0, two –1s, two 1s, one –3,
and one 3. (Deviation scores above the mean are assigned positive signs; those below the mean are
assigned negative signs.)

Standard Deviation

The square root of the mean of all squared deviations from the mean, that is,

Standard deviation variance= √𝒗𝒂𝒓𝒊𝒂𝒏𝒄𝒆


Standard deviation is a statistic that measures the dispersion of a dataset relative to its mean and is
calculated as the square root of the variance. The standard deviation is calculated as the square root of
variance by determining each data point's deviation relative to the mean.

Sum of Squares (SS)

The sum of squares, symbolized by SS can be defined by two formulas the definition formula,
which is easier to understand and remember, and the computation formula, which usually is more
efficient.

Sum of Squares Formulas for Population

The definition formula provides the most accessible version of the population sum of squares:

where SS represents the sum of squares, Σ directs us to sum over the expression to its right, and
(X − μ)2 denotes each of the squared deviation scores.

1. Subtract the population mean, μ, from each original score, X, to obtain a deviation score, X − μ.

2. Square each deviation score, (X − μ)2, to eliminate negative signs.

3. Sum all squared deviation scores, Σ (X − μ)2.

When the mean equals some complex number, such as 169.51, or the number of scores is large the
more efficient computation formula can be used:

where Σ X2 , the sum of the squared X scores, is obtained by first squaring each X score and then
summing all squared X scores; (Σ X)2 , the square of sum of all X scores, is obtained by first adding all X
scores and then squaring the sum of all X scores; and N is the population size.

Calculation of Population Standard Deviation Σ (Definition Formula)


A. Computation Sequence
1. Assign a value to N representing the number of X scores.
2. Sum all X scores.
∑𝑋
3. Obtain the mean of these scores. 𝜇 = 𝑁
4. Subtract the mean from each X score to obtain a deviation score (X-µ).
5. Square each deviation score (X-µ)2.
6. Sum all squared deviation scores to obtain the sum of square, SS= (∑ 𝑋 − 𝜇 2 ).
𝑆𝑆
7. Substitute numbers into the formula to obtain population variance, 𝜎 2 = .
𝑁

𝑠𝑠
8. Take the square root of σ 2 to obtain the population standard deviation, 𝜎 = √ 𝑁 .

Calculation of Population Standard Deviation (Σ) (Computation Formula)


Sum of Squares Formulas for Sample

Sample notation can be substituted for population notation in the above two formulas without causing any
essential changes:

Standard Deviation for Population σ


Recall that, most generally, a mean is defined as the sum of all scores divided by the number of
scores. Since the variance is the mean of all squared deviation scores, it can be defined as the sum of all
squared deviation scores divided by the number of scores:

𝑠𝑢𝑚 𝑜𝑓 𝑎𝑙𝑙 𝑠𝑞𝑢𝑎𝑟𝑒𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛 𝑠𝑐𝑜𝑟𝑒𝑠


𝑣𝑎𝑟𝑖𝑎𝑛𝑐𝑒 =
𝑛𝑢𝑚𝑏𝑒𝑟 𝑜𝑓 𝑠𝑐𝑜𝑟𝑒𝑠

or, in symbols:

where the squared lowercase Greek letter, σ2 (pronounced “sigma squared”), represents the population
variance, SS is the sum of squared deviations for the population, and N is the population size.

The standard deviation is obtained by taking the square root of the variance, that is,

where σ represents the population standard deviation, instructs us to take the square root of the covered
expression.

Standard Deviation for Sample (s)

Although the sum of squares term remains essentially the same for both populations and samples,
there is a small but important change in the formulas for the variance and standard deviation for samples.
This change appears in the denominator of each formula where N, the population size, is replaced by the
sample size, n − 1, as shown:

where s2 and s represent the sample variance and sample standard deviation, SS is the sample sum
of squares.
Calculation of Sample Standard Deviation (S) (Definition Formula)

Calculation of Sample Standard Deviation (S) (Computation Formula)


Degrees of Freedom (df)

Degrees of freedom is a very important notion in inferential statistics. Degrees of freedom (df)
refers to the number of values that are free to vary, given one or more mathematical restrictions, in a
sample being used to estimate a population characteristic.

The concept of degrees of freedom is introduced only because scores in a sample are used to
estimate some unknown characteristic of the population. Typically, when used as an estimate, not all
observed values in the sample are free to vary because of one or more mathematical restrictions.

As has been noted, when n deviations about the sample mean are used to estimate variability in
the population, only n − 1 are free to vary. As a result, there are only n − 1 degrees of freedom, that is, df
= n − 1.

We can use degrees of freedom to rewrite the formulas for the sample variance and standard
deviation:
Interquartile Range (IQR)

The most important spinoff of the range, the interquartile range (IQR), is simply the range for the
middle 50 percent of the scores. More specifically, the IQR equals the distance between the third quartile
(or 75th percentile) and the first quartile (or 25th percentile), that is, after the highest quarter (or top 25
percent) and the lowest quarter (or bottom 25 percent) have been trimmed from the original set of scores.

Calculation of the IQR

1. Order scores from least to most.


2. Add 1 to the total number of scores and divide by 4. If necessary, round the result to the nearest
whole number.
3. Beginning with the largest score, count the requisite number of steps (calculated in step 2) into the
ordered scores to find the location of the third quartile
4. The third quartile equals the value of the score at this location.
5. Beginning with the smallest score, again count the requisite number of steps into the ordered
scores to find the location of the first quartile.
6. The first quartile equals the value of the score at this location.
7. The IQR equals the third quartile minus the first quartile.
Measures of Variability for Qualitative and Ranked Data

Qualitative Data

Measures of variability are virtually non-existent for qualitative or nominal data. It is probably
adequate to note merely whether scores are evenly divided among the various classes (maximum
variability), unevenly divided among the various classes (intermediate variability), or concentrated mostly
in one class (minimum variability).

Ordered Qualitative and Ranked Data

If qualitative data can be ordered because measurement is ordinal (or if the data are ranked), then
it’s appropriate to describe variability by identifying extreme scores (or ranks).

-------------------------------------------------------------------------------------------------------------------------------

Normal Distributions and Standard (z) Scores

The Normal Curve

More accurate generalizations usually can be obtained from distributions based on larger numbers
of data points.

For example: A distribution based on 30,910 men usually is more accurate than one based on 3091, and a
distribution based on 3,091,000 usually is even more accurate.

Normal curve has been superimposed on the original distribution. Any generalizations based on
the smooth normal curve will tend to be more accurate than those based on the original distribution
In Figure 5.2, the idealized normal curve has been superimposed on the original distribution for
3091 men. Irregularities in the original distribution, most likely due to chance, are ignored by the smooth
normal curve.

Accordingly, any generalizations based on the smooth normal curve will tend to be more accurate
than those based on the original distribution.

Interpreting the Shaded Area

The total area under the normal curve in Figure 5.2 can be identified with all data points (Eg.
heights of all men applicants for FBI jobs). Viewed relative to the total area, the shaded area represents
the proportion of applicants who will be eligible because they are shorter than exactly 66 inches.

Properties of the Normal Curve

1. Obtained from a mathematical equation, the normal curve is a theoretical curve defined for a
continuous variable, and is symmetrical bell-shaped curve.
2. Because the normal curve is symmetrical, its lower half is the mirror image of its upper half.
3. Being bell shaped, the normal curve peaks above a point midway along the horizontal spread and
then tapers off gradually in either direction from the peak
4. The values of the mean, median (or 50th percentile), and mode, located at a point midway along
the horizontal spread, are the same for the normal curve.

Types of Normal curves

The various types of normal curves can be produced by an arbitrary change in the value of either
the mean (μ) or the standard deviation (σ).

For example, changing the mean height from 69 to 79 inches produces a new normal curve that, as shown
in panel A of Figure 5.3, is displaced 10 inches to the right of the original curve. Dramatically new
normal curves are produced by changing the value of the standard deviation. As shown in panel B of
Figure 5.3, changing the standard deviation from 3 to 1.5 inches produces a more peaked normal curve
with smaller variability, whereas changing the standard deviation from 3 to 6 inches produces a shallower
normal curve with greater variability.

The differences in appearance among normal curves are less important. Because of their common
mathematical origin, every normal curve can be interpreted in exactly the same way once any distance
from the mean is expressed in standard deviation units.

Once any distance from the mean has been expressed in standard deviation units, the standard normal
table can be used to determine the corresponding proportion of the area under the normal curve.

Z Scores

A z score is a unit-free, standardized score that, regardless of the original units of measurement,
indicates how many standard deviations a score is above or below the mean of its distribution.

To obtain a z score, express any original score, whether measured in inches, milliseconds, dollars,
IQ points, etc., as a deviation from its mean (by subtracting its mean) and then split this deviation into
standard deviation units (by dividing by its standard deviation), that is,

𝑋−𝜇
𝑧=
𝜎

where X is the original score and μ and σ are the mean and the standard deviation, respectively,
for the normal distribution of the original scores. Since identical units of measurement appear in both the
numerator and denominator of the ratio for z, the original units of measurement cancel each other and the
z score emerges as a unit-free or standardized number, often referred to as a standard score.

A z score consists of two parts:


1. A positive or negative sign indicating whether it’s above or below the mean; and

2. A number indicating the size of its deviation from the mean in standard deviation units.

A z score of 2.00 always signifies that the original score is exactly two standard deviations above
its mean. Similarly, a z score of –1.27 signifies that the original score is exactly 1.27 standard deviations
below its mean. A z score of 0 signifies that the original score coincides with the mean.

When X with 66 (the maximum permissible height), μ with 69 (the mean height), and σ with 3
(the standard deviation of heights) and solve for z:

𝑋−𝜇 66−69 −3
The formula for Z score 𝑧 = = = = -1
𝜎 3 3

This shows the score is exactly one standard deviation below the mean.

Standard Normal Curve

The standard normal curve is a special normal distribution where the mean is 0 and the standard
deviation is 1.

To verify (rather than prove) that the mean of a standard normal distribution equals 0, replace X in
the z score formula with μ, the mean of any (nonstandard) normal distribution, and then solve for z:

𝑋−𝜇 𝜇−𝜇 0
𝑀𝑒𝑎𝑛 𝑜𝑓 𝑧 = = = =0
𝜎 𝜎 𝜎

Likewise, to verify that the standard deviation of the standard normal distribution equals 1, replace
X in the z score formula with μ + 1σ, the value corresponding to one standard deviation above the mean
for any (nonstandard) normal distribution, and then solve for z:

𝑋 − 𝜇 𝜇 + 1𝜎 − 𝜇 1𝜎
𝑆𝑡𝑎𝑛𝑑𝑎𝑟𝑑 𝑑𝑒𝑣𝑖𝑎𝑡𝑖𝑜𝑛 𝑜𝑓 𝑧 = = = =1
𝜎 𝜎 𝜎

Although there are an infinite number of different normal curves, each with its own mean and standard
deviation, there is only one standard normal curve, with a mean of 0 and a standard deviation of 1.

Figure 5.4 illustrates the emergence of the standard normal curve from three different normal
curves: that for the men’s heights, with a mean of 69 inches and a standard deviation of 3 inches; that for
the useful lives of 100-watt electric light bulbs, with a mean of 1200 hours and a standard deviation of
120 hours; and that for the IQ scores of fourth graders, with a mean of 105 points and a standard deviation
of 15 points.

To find the proportion for the shaded areas in Figure 5.4, the same z score of –1.00 can be used
when referring to the table for the standard normal curve, the one table for all normal curves.
Standard Normal Table

Essentially, the standard normal table consists of columns of z scores coordinated with columns of
proportions.

Using the Top Legend of the Table

▪ Columns are arranged in sets of three, designated as A, B, and C in the legend at the top of the
table.
▪ When using the top legend, all entries refer to the upper half of the standard normal curve.
▪ The entries in column A are z scores.
▪ Columns B and C indicate how the z score splits the area in the upper half of the normal curve.
▪ Column B indicates the proportion of area between the mean and the z score.
▪ Column C indicates the proportion of area beyond the z score, in the upper tail of the standard
normal curve.

Using the Bottom Legend of the Table

 Columns are designated as A′, B′, and C′ in the legend at the bottom of the table.

 All entries refer to the lower half of the standard normal curve.

 Column B′ indicates the proportion of area between the mean and the negative z score.

 Column C′ indicates the proportion of area beyond the negative z score, in the lower tail of the
standard normal curve.

Solving Normal Curve Problems

There are two main types of normal curve problems. In the first type of problem, a known score
(or scores) is used to find an unknown proportion. In the second type of problem, the procedure is
reversed i.e. a known proportion is used to find an unknown score (or scores).

When using the standard normal table, it is important to remember that for any z score, the
corresponding proportions in columns B and C (or columns B′ and C′) always sum to .5000. Similarly,
the total area under the normal curve always equals 1.0000, the sum of the proportions in the lower and
upper halves, that is, .5000 + .5000. Figure 5.5 summarizes how to interpret the normal curve table.
Finding Proportions

Example: Finding Proportions for One Score

Find the proportion of all FBI applicants who are shorter than exactly 66 inches, given that the
distribution of heights approximates a normal curve with a mean of 69 inches and a standard deviation of
3 inches.

1. Sketch a normal curve and shade in the target area.

2. Plan your solution according to the normal table.

The answer will be obtained from column C′ of the standard normal table, since the target
area coincides with the type of area identified with column C′, that is, the area in the lower tail beyond
a negative z.

3. Convert X to z. Express 66 as a z score.


𝑋 − 𝜇 63 − 66 −3
𝑧= = = = −1
𝜎 3 3
4. Find the target area. Refer to the standard normal table, using the bottom legend, as the z score is
negative.

It can be concluded that only .1587 (or .16) of all of the FBI applicants will be shorter than 66
inches.

Example: Finding Proportions between Two Scores

Finding the proportion between 245 and 255

1. Sketch a normal curve and shade in the target area.


2. Plan your solution according to the normal table.
As shown in above figure, the basic idea is to identify the target area with the difference
between two overlapping areas whose values can be read from column C′ of the standard normal
table. The larger area (less than 255) contains two sectors: the target area (between 245 and 255)
and a remainder (less than 245). The smaller area contains only the remainder (less than 245).
Subtracting the smaller area (less than 245) from the larger area (less than 255), therefore,
eliminates the common remainder (less than 245), leaving only the target area (between 245 and
255).
3. Convert X to z by expressing 255 as
255 − 270 −15
𝑧= = = −1.00
15 15
And by expressing 245 as
245 − 270 −25
𝑧= = = −1.67
15 15
4. Find the target area.

Example: Finding Proportions beyond Two Scores

Assume that high school students’ IQ scores approximate a normal distribution with a mean of 105 and
a standard deviation of 15. What proportion of IQs are more than 30 points either above or below the
mean?

1. Sketch a normal curve and shade in the two target areas, as in the top panel of Figure 5.8.
2. Plan your solution according to the normal table.
The solution to this type of problem is straightforward because each of the target areas can be read
directly from standard normal table. The target area in the tail to the right can be obtained from
column C, and that in the tail to the left can be obtained from column C′.
3. Convert X to z by expressing IQ scores of 135 and 75 as
75 − 105 −30
𝑧= = = −2.00
15 15
135 − 105 30
𝑧= = = 2.00
15 15
4. Find the target area.
In standard normal table, locate a z score of 2.00 in column A, and note the corresponding
proportion of .0228 in column C. Because of the symmetry of the normal curve, there is no need to
enter the table again to find the proportion below a z score of –2.00. Instead, merely double the
above proportion of .0228 to obtain .0456, which represents the proportion of students with IQs
more than 30 points either above or below the mean.

Finding Scores

Example: Finding One Score

Exam scores for a large psychology class approximate a normal curve with a mean of 230 and a standard
deviation of 50. Furthermore, students are graded “on a curve,” with only the upper 20 percent being
awarded grades of A. What is the lowest score on the exam that receives an A?

1. Sketch a normal curve and, on the correct side of the mean, draw a line representing the target score,
as in Figure 5.9.
2. Plan your solution according to the normal table.
Since the target score is on the right side of the mean, concentrate on the area in the upper half of
the normal curve, as described in columns B and C.

3. Find z. Refer to standard normal table.


Scan column C to find .2000. If this value does not appear in column C, as typically will be the
case, approximate the desired value (and the correct score) by locating the entry in column C nearest
to .2000. The entry in column C closest to .2000 is .2005, and the corresponding z score in column A
equals 0.84.
4. Convert z to the target score.
Finally, convert the z score of 0.84 into an exam score, given a distribution with a mean of 230
and a standard deviation of 50.

The following formula is used in the conversion of z score to the target score.

Answer: X=µ + (z) (σ)


= 230+ (0.84) (50)
= 230+42
=272
Therefore, 272 is the lowest score on the exam that receives an A.

Example: Finding Two Scores

Assume that the annual rainfall in the San Francisco area approximates a normal curve with a mean of
22 inches and a standard deviation of 4 inches. What are the rainfalls for the more atypical years,
defined as the driest 2.5 percent of all years and the wettest 2.5 percent of all years?
1. Sketch a normal curve.
On either side of the mean, draw two lines representing the two target scores, as in Figure 5.10.
The smaller (driest) target score splits the total area into .0250 to the left and .9750 to the right, and
the larger (wettest) target score does the exact opposite.

2. Plan your solution according to the normal table.


Since the smaller target score is located on the lower or left side of the mean, concentrate on the
area in the lower half of the normal curve, as described in columns B′ and C′. The target z score can
be found by scanning either column B′ for .4750 or column C′ for .0250. After finding the smaller
target score, find the value of the larger target score.
3. Find z.
Referring to standard table, scan column B′ for .4750, or the entry nearest to .4750. In this case,
.4750 appears in column B′, and the corresponding z score in column A′ equals –1.96. The same z
score of –1.96 would have been obtained if column C′ had been searched for a value of .0250.
4. Convert z to the target score.
When the appropriate numbers are substituted in Formula 5.2, the smaller target score equals
14.16 inches, the amount of annual rainfall that separates the driest 2.5 percent of all years from all
of the other years.

Important Note: For additional problems refer page no. 103 - 106 in II Text Book.

You might also like