MBA Business Statistics E-Content Guide
MBA Business Statistics E-Content Guide
E-CONTENT
BUSINESS STATISTICS
MBA I SEM I
Dr. Anugamini Priya Srivastava
Established under Section 3 of the UGC Act. 1956 I Awarded Category - I by UGC
Gram: Lavale, Tal: Mulshi, Dist: Pune, Maharashtra, India Pin: 412115
E-CONTENT
BUSINESS STATISTICS
MBA I SEM I
Internal Advisory Board (Self-Learning Material)
ISBN: 978-93-95877-05-3
All rights reserved. No part of this content may be reproduced or used in any form or by any means without
prior permission from the publisher. This content is produced only for academic purpose and is for internal
circulation only.
Acknowledgement : Every attempt has been made to trace the copyright holders of material reproduced in
this content. Should any infringement have occurred. We apologize for the same and would be pleased to make
necessary corrections in future editions of this content.
Published by : Symbiosis School for Online and Digital Learning, SIU, Lavale, Pune
STRUCTURE
¡ Statistics and research
¡ Statistics definition
¡ Characteristics of statistics as statistical data
¡ Statistical methods
¡ Stages of statistical investigation
¡ Functions of statistics
¡ Scope of the statistics
¡ Limitations of statistics
¡ Summary
01
From probability sampling to non-probability sampling, statistical knowledge helps researchers
choose the best possible sampling method.
Similarly, statistical understanding helps identify and explore different modes of collection of
data and choose the best possible one from primary and secondary data sources. Along the same
line, statistics provides a wide range of methods to explore the data duly collected. It enables the
recording and cleaning of the data based on missing values and outliers. Further, it has in
identifying the most effective methodology to understand the characteristics of the data per
variable, explore the average or dispersion in the data, the fitness of the model and the
predictability of variables on others. Statistics also support making valuable inferences based on
the statistical thresholds. Using different references and guidelines, researchers can make useful
inferences based on their research ideas and hypothesis. And lastly, in research, statistics can
provide proper formats of data presentation - tabular or diagrammatic to enable researchers to
represent their creativity in report writing.
Thus, statistical analysis in research can prove to be the lifeline.
02
1.4.2. Affected to a marked extent by a multiplicity of causes
generally, the facts and figures are influenced by the number of forces operating together for a
considerable time. Assuming that a particular point or figure works in isolation will be utterly
wrong in statistics. For example, suppose you talk about the rice production quality in a year. In
that case, we need to consider the data on the quantity of rainfall, the quality of soil seeds, and the
cultivation methods. So, in statistics, every piece of information collected is affected by multiple
causes. Different ways and means have been identified for segregating the effects of other forces
on a process. Similarly, statistics also considered the difficulty involved in measuring the
complex impact variety of factors which are not measurable.
03
that enables chronological or region-wise comparison for geographical comparisons, for
example, statistics collected on the per capita income of employees working in developed
nations. All of them can be compared when converted to dollars as the currency all these
characteristics make statistics very crucial in the field of research. Without following these
characteristics, data cannot be called statistics. Like it is well said all statistics are numerical
statements of facts, but all numerical statements of facts and not statistics.
04
where a statistician contains the data based on the research objectives. This data is collected
based explicitly on the purpose of the research projects. Another way to gather information is the
secondary sources. The secondary sources are the published or unpublished sources from where
the investigator collects the data and use it in their projects. However, they need to be very
cautious regarding the use of secondary sources of data. The reliability and suitability of such
data should be tested and only utilised in this study.
05
1.7 FUNCTIONS OF STATISTICS
So now, let's understand the functions of statistics
1.7.1. Definiteness: the statistics specifically require a clear and definite form of general
statements. The numbers and details must be precisely defined and mentioned for accurate
results. For example, saying statements like the GDP of the country has increased. This statement
is neither specific nor definite in any sense. While rewriting the same statement as the country's
GDP has risen from 11% in 2020 to 15% in 2022, this statement provides more clarity to the
readers with a definite meaning.
1.7.2. Simplified mass of figure: along with being definite, statistics helps condense the
gathering of data into a few significant figures with proper and more precise meanings. In a raw
data set, multiple columns and rows of data exist. Using statistical methods, one can condense the
data based on common characteristics and give proper meaning for better interpretations. For
example, by reading the census reports on individual salary structures in a country, a researcher
might not understand the income of the entire population. However, a simplified value of per
capita income can be easily remembered and understood by everyone.
1.7.4. Formulating and testing hypothesis: statistical methods can help investigators to use
different tools and measurements to formulate and test the hypothesis. For example, if the
investigator aims to test the effect of increased dopamine on students' performance in the
classroom, then using statistical understanding, null and alternative hypotheses can be
formulated. If the alternative hypothesis is that dopamine affects students' performance, then a
simple linear regression analysis can be used to test the theory.
1.7.5. Helps in prediction - statistics also enables effective prediction for the future. Using
statistics, researchers can evaluate the overall trend and utilise the information to forecast future
events. For example, data on employee turnover from the last 5 years can be used to predict the
prospective number of employees that might leave the organisation in the next 6 months. Based
on this prediction, good employee retention policies can be developed well in advance.
06
proper data collection, summarization and presentation, but it also helps marketers to develop
new strategies and deal with problems.
1.8.1 Operations
In the field of operations management and supply chain management, statistics support quality
management. On the one hand, it helps manufacture the products with ideal size and
composition; on the other hand, it helps quality check through random selection of items from the
lots. Through a random sample of articles from conveyor belts, managers can check whether the
quality of the product is acceptable or rejected. Similarly, on delivery of raw materials, managers
and executives can verify if the materials delivered are of prescribed quality or not.
1.8.2 Accounting
In accounting, statistics are utilized explicitly in the auditing process. In the auditing process,
auditors cross-verify vouchers or bills with transactions shown in ledgers and journals, also
known as bookkeeping. For cross-verification, the auditor selects vouchers and statements
randomly from the registers and checks them in correspondence to transaction entries. Through
this method of random sampling, there are higher chances that frauds and mistakes can be
detected, if any. We will discuss random samples in upcoming chapters.
1.8.4 Banking
Similarly, in banking, statistics are used to collect and analyze information to understand future
economic conditions and evaluate other external needs to understand every line of business in
which they might be directly or indirectly interested.
1.8.5 Others
Similarly, statistics can be utilized in investment, purchase, credit and control, personnel
management, and research and development. Thus, we can conclude that the scope of the
statistics is vast and continuously increasing. Due to this, it is tough to define statistics. At the
same time, it is unwise to explain the exact usage of statistics as it is permeated almost every
aspect of our lives.
1.9.2. Deals with numbers- statistics involves numeric facts and figures. It does not enable the
usage of qualitative or text data. For example, data on males and females, usually considered
nominal data, can be used for making a diagrammatic presentation. However, it cannot be used
for analysis. For analysis, these labels must be transformed into codes, like male-0 and female-1.
1.9.3. Are true only on an average- the conclusions in statistics are based on the purpose of
research, hypothesis developed, research design adopted and disciplinary background of
researcher. So, it is not true that the conclusion made in one analysis is universally true. A
generalisation of results is required to verify the application of the decision in another context.
1.9.4. Only a means- statistical methods furnish only one way of studying a problem. They may
not be the best under all circumstances. So, it should be carefully noted that statistics is only a
means, not an end. It analyses the facts and throws light on the actual situation. Complete
dependence on statistics male lead to fallacious conclusions in many cases. Proper consideration
of factors affecting and cautious organization and analysis is required for accurate conclusions.
1.9.5. Can be misused- the most significant limitation of statistics is that it is liable to be
misused. It is essential to understand that statistics are likely, and they can be moulded in any
manner to establish right or wrong conclusions if the statistical findings are based on incomplete
information, one may arrive at false conclusions. That's why it requires experience and skills to
draw sensible conclusions from the data. Otherwise, there is every likelihood of a wrong
interpretation.
1.10 SUMMARY
The chapter involved a basic understanding of the term business statistics. Below given is the
summary of this chapter.
¡ Statistics and research: research and statistics go hand-in-hand. However, research can
be conducted without statistics, but statistics without research in the business world will
not help in attaining practical conclusions and decisions.
¡ Statistics definition: statistics can be defined as both statistical methods and data.
Different scholars provided different meanings to explain the crux of statistics and its
characteristics.
¡ Characteristics of statistics as statistical data: significant features of statistics as
statistical data are that it comprises aggregates of facts affected to a marked extent by a
multiplicity of causes, numerically expressed, enumerated or estimated according to
08
reasonable standards of accuracy, collected in a systematic manner for a predetermined
purpose and placed with each other.
¡ Statistical methods: in the singular sense, statistics is also defined as statistical methods.
¡ Stages of statistical investigation: statistics involves 5 stages- collection, organisation,
presentation, analysis and interpretation.
¡ Scope of statistics: statistics are helpful in every discipline, from operations management
to production management to human resource management and marketing. It is also
beneficial in fundamental grounds of accounting and banking.
¡ Limitations of statistics: statistics have certain limitations which restrict their usage in
research work. It only uses aggregates and numerical data for analysis. It is determined to
be correct on average and can lead to misleading results if not conducted with diligence
and understanding of the research purpose.
Short questions
1. How statistics can be helpful in the field of banking
2. How statistics can be misused
3. Define statistics as statistical methods
4. Define statistics as statistical data
09
True/ False
1. Statistics specifically require a clear and definite form of general statements. (True)
2. Statistics help us understand how things have changed over time. (True)
3. To have reliable and accurate data collection for the nuanced analysis, research objectives
need not be kept in mind. (False)
4. In statistics, unless the figures can be compared with another figure of the same kind, they
are devoid of any meaning. (True)
10
1.12 REFERENCES
¡ Black, K. (2019). Business statistics: for contemporary decision making. John Wiley &
Sons.
¡ Anderson, D. R., Sweeney, D. J., Williams, T. A., Camm, J. D., & Cochran, J. J. (2020).
Modern business statistics with Microsoft Excel. Cengage Learning.
11
MODULE - 2 DATA PRESENTATION
STRUCTURE
¡ Introduction
¡ Frequency distributions
¡ Steps to develop frequency distributions
¡ Class midpoint
¡ Relative frequency
¡ Cumulative frequency
¡ Quantitative data graphs
¡ Qualitative data presentations
Learning Objectives
¡ To understand the meaning of diagrammatic presentations
¡ To understand the frequency distributions
¡ To understand quantitative modes of data presentation
¡ To understand the qualitative modes of data presentation.
2.1 INTRODUCTION
In the era of data analytics and data science, it is essential to understand how these data can be
summarized and presented to ensure effective communication of results. Thus, in this chapter,
we will discuss different modes of diagrammatic presentations using qualitative and quantitative
data with the help of examples.
The first step toward analyzing the data is to explore the data using tabular or diagrammatic
presentations.
12
Table 2.1: Data on the number of employees fallen sick in 12 months in the year 2021
34 12
23 34
20 23
56 16
67 11
20 10
Table 2.2: Data on the number of employees fallen sick in 12 months in the year 2021
(grouped)
Classes Frequency
10-20 4
20-30 4
30-40 2
40-50 0
50-60 1
60-70 1
Determine the width of the class interval the next step is to identify the class width. The formula
to calculate the class width = range of data by a number of classes. For the data given in table 1.1,
the approximation would be =
57 = 9.5
6
So the class width can be considered as 10, rounding off to the following whole number, which is
10. The frequency distribution must start at a value equal to or lower than the lowest number of
the ungrouped data and end at a value equal to or higher than the highest number. So, based on the
minimum days of absenteeism as 10 and the highest at 67, the frequency distribution can start at
10 and end at 70. (see table 2.2)
13
For data summarization and presentation, class midpoints are essential. In other words, a class
midpoint can be defined as a value halfway across the class interval and calculated as the average
of the two class endpoints. We can calculate midpoints using the given formula =
Where,
LL represents the lower limit of the class interval
UL represents the upper limit of the class interval
Thus, using this formula, for the data in table 1.3, midpoints can be calculated as
(10+20) / 2 = 15
i.e. for example, for the class 10-20, the relative frequency would be = 4/12
= 0.333333
Table 2.3: Class Midpoints, Relative Frequencies, and Cumulative Frequencies for
absenteeism
14
Similarly, for other classes, the cumulative frequency can be calculated. This process continues
through the last interval, at which point the cumulative total equals the sum of the frequencies. In
other words, to verify the estimates in cumulative frequency, we can cross-check that the final
number is the cumulative frequency column with the total frequency column. In example 2.3, the
final cumulative value is 12, equal to the total frequency.
2.6.1 HISTOGRAM
The histogram is a series of contiguous bars or rectangles that denote the frequency of data in
given class intervals. If the class interval is equal, then the frequency of the values in each class
interval is represented through the height of the bars. While on the other hand, if the class
intervals are unequal, then the relative comparisons of class frequencies are depicted in the areas
of the bars (rectangles). As shown in figure 2.1, a histogram involves an x-axis labelled with class
endpoints and a y-axis marked with their respective frequencies. Here, the x-axis is also known
as abscissa, while the y-axis is named ordinates. Thus, in the histogram are drawn drawing a
horizontal line from the frequency value of one class endpoint to another class endpoint and
interlinking each one vertically from the frequency value to the x-axis to form a series of bars, as
shown in figure 2.1.
15
2.6.2 FREQUENCY POLYGONS
Frequency polygons are another way of presenting a quantitative data set. Theoretically, it is
constructed by scaling class midpoints along the horizontal axis and the frequency scale along
the vertical axis. In other words, it is similar to a histogram. But instead of bars or rectangles, it is
based on plotting a dot at the class midpoint and then connecting each dot by a series of line
segments. As seen in figure 2.2, the dots are plotted and then a line segment connecting each dot
is drawn.
Figure 2.2: Excel produced a frequency polygon for the days of absenteeism
2.6.3 OGIVES
An ogive is another way of presenting quantitative data. However, unlike histograms and
polygons, it is made based on cumulative frequencies. So, a graphical presentation of cumulative
frequency values. In the construction of ogives, the following steps are taken
¡ Labelling x-axis with class endpoints and y-axis with frequencies
¡ Scaling y-axis enough to include the total of frequencies
¡ Starting with plotting a 0 at the beginning of the first class
¡ Preceding with marking each dot at the end of each class interval for the cumulative values.
¡ Lastly, connect all dots to complete the presentation.
The diagrammatic presentation of ogive is most useful when the researcher aims to see running
totals. For example, if researchers are interested in controlling the overall production costs, an
ogive can depict the cumulative costs of a financial year. As shown in figure 2.3, a particularly
steep slope occurs in the 20-30 class interval, signifying a significant jump in class frequency
totals.
16
Figure 2.3: Ogive using excel
17
¡ First, reorganize the data into a "long" format (as shown in table 2.4)
¡ Step 2: Create a dot plot using the "scatterplot" option in excel as shown in figure 2.4.
18
¡ Right-most digits are considered leaves and consist of lower values.
¡ If a set of data has only two digits, the stem is the value on the left, and the leaf is the value
on the right
Let's develop a stem and leaf plot using the example in table 2.1. The stem and leaf plot are
depicted in Figure 2.6
1 0
1 2 6
2 0 0 3 3
3 4 4
5 6
6 7
The primary benefit of using stem and leaf plots is that the investigator can readily see whether
the scores are in the upper or lower end of each bracket and also determine the spread of the
scores. Further, this plot can help determine the spread of the scores. Another advantage of using
this plot is that the values of the original raw data are retained. In contrast, other frequency
distributions and graphic presentations use the class midpoints to depict the values in a class
interval.
FEMALE 13
MALE 7
19
Figure 2.7: Pie chart using excel
Here we can see the total area under the pie is 100%, and the angle is 360 degrees. So, to identify
the proportion, we can use the formula as female = 13/20*100 = 65 % & male = 7/20* 100 = 35%.
Thus, based on the example shown in figure 1.7, we can conclude that the respondents in the
survey were majorly females (65%) while males were only 35%.
Now converting this into degrees, we can calculate the angle as 13/20*360 = 234 degrees for
females. Similarly, for males = 7/20 * 360 = 126 degrees. Thus, the pie chart shows the relative
magnitude of the part to the whole. Pie charts are widely used to depict variables like market
share, resource allocations, investment patterns etc.
20
2.8 COUNTRY MAPS
Apart from the above-given diagrams for qualitative data, we can also use country or state-based
data and use country maps in excel to depict. The steps to develop a country map is stated below
1. Summarize the data in excel
2. Go to insert tab and click on maps option under the charts tab as shown in figure 2.9
Click on map options and the diagram will automatically shown as shown in figure 2.10.
21
However, there are a few points that are needed to be kept in mind.
a. Maps charts can only plot high-level geographic details only. No latitude or longitude, or
street address can be mapped in it.
b. Maps charts also support one-dimensional display only.
c. Online connections are required to create new maps in excel or append data to existing
maps.
d. Existing maps can be viewed offline.
2.10 SUMMARY
In the era of data analytics and data science, it is essential to understand how these data can be
summarized and presented to ensure effective communication of results
Frequency distributions: The frequency distributions refer to a summary of data presented in
class intervals and frequencies.
Class midpoint: a class midpoint can be defined as a value halfway across the class interval and
calculated as the average of the two class endpoints
Relative frequency: It can be defined as the proportion of the total frequency in any given class
interval in a frequency distribution
Cumulative frequency: refers to running the total of frequencies through the classes of a
frequency distribution.
Quantitative data graphs: Quantitative data are the data which are numerically expressed and are
measured by interval and ratio scale of measurement. To effectively present any data, it is
suggested to transform it into graphs and plots. There are five major types of quantitative data
graphs - 1) histogram, 2) frequency polygons, 3) Ogives, 4) dot plot, 5) Stem and leaf plots
Qualitative data presentations can be done using two types of graphs 1. Pie charts, 2. Bar charts,
2. According to the National Retail Federation and Center for Retailing Education at
the University of Florida, the four main sources of inventory shrinkage are employee theft,
shoplifting, administrative error, and vendor fraud. The estimated annual dollar amount in
shrinkage ($ millions) associated with each of these sources follows:
Employee theft $17,918.6
Shoplifting 15,191.9
Administrative error 7,617.6
22
Vendor fraud 2,553.6
Total $43,281.7
Construct a pie chart and a bar chart to depict these data.
3. The following data represent the number of passengers per flight in a sample of 50 flights
from Wichita, Kansas, to Kansas City, Missouri.
23 46 66 67 13 58 19 17 65 17
25 20 47 28 16 38 44 29 48 29
69 34 35 60 37 52 80 59 51 33
48 46 23 38 52 50 17 57 41 77
45 47 49 19 32 64 27 61 70 19
a. Construct a dot plot for these data.
b. Construct a stem-and-leaf plot for these data. What does the stem-and-leaf plot tell you
about the number of passengers per flight?
Short questions
1. Define pie charts.
2. Define Bar charts
3. Define the advantages of stem and leaf plots
4. Explain the uses of country maps.
23
3. Ogive are developed based on cumulative frequency (True)
4. Country maps can be used to depict street addresses (False)
2.12 REFERENCES
¡ Black, K. (2019). Business statistics: for contemporary decision making. John Wiley &
Sons.
¡ Anderson, D. R., Sweeney, D. J., Williams, T. A., Camm, J. D., & Cochran, J. J. (2020).
Modern business statistics with Microsoft Excel. Cengage Learning.
24
MODULE - 3 MEASURES OF CENTRAL TENDENCY
STRUCTURE
¡ Introduction to measures of central tendency
¡ Characteristics of a good average
¡ S Symbol
¡ Different types of measures of central tendency
¡ A.M. of Grouped frequency distribution
¡ Composite A.M.
¡ Advantages and disadvantages of A.M.
¡ Geometric Mean
¡ Advantages and disadvantages of G.M.
¡ Uses of G.M.
¡ Harmonic Mean
¡ Relationship among A.M., G.M. and H.M.
¡ Median
¡ Advantages and Limitations of Median
¡ Mode
¡ Relationship between Mean, Median and Mode
¡ Quartiles, Deciles and Percentiles
3.2 INTRODUCTION
While working with data, we may sometimes need such a numerical expression which can tell us
certain characteristics of the whole data set, such as the central most value, the lowest value, the
highest value or the value which appears the most frequently. One such type of measure which
can tell us the average value of a distribution are known as the 'measures of central tendency' or
25
also known as 'Averages', more popularly. Such a measure which represents the middle most
value should obviously be greater than the smallest value and less than the highest value. A
measure of central tendency or an average of a certain distribution is nothing but a representative
value of that distribution which enables us to comprehend in a single effort the significance of the
whole. It should be a value which lies between the two limits, i.e, the highest and lowest points in
the data set, possibly at the centre, where most of the values of the series cluster.
Measures of central tendency or averages are additionally, also called as measures of central
location. These are arithmetical measures intended to represent the central value of a data set. We
can say that an average of a distribution (of the values) of any variable (say weight of some
students in a class in cms) is a representative value of that variable.
In any observation set, the representative value of a distribution usually lies at or near the centre
of the distribution. This happens due to the inherent tendency of a distribution of data of any kind
that the major part of the values gets concentrated at the centre. Since this average is reflective of
this tendency of the data, hence, the average is called a measure of central tendency.
Depending upon the nature of a distribution, different methods of obtaining the representative
value have been evolved and as a result, we have several averages or measures of central
tendency.
3.4 S SYMBOL
In order to denote sum (i.e, a total of certain quantities), the Greek letter S (capital sigma) is used.
For example, if a variable x takes the values x1, x2, x3…xn, then the sum of these values of the
n n
variable x i.e., (x1 + x2 + x3+… + xn) is denoted by S t =1 xi or S x. The symbol S t = 1xi means
that the lower limit of i is 1 and the upper limit of i is n, i.e., i takes the value 1, 2, 3, …, n and the
symbol S means that all the values of xi for i = 1, 2, 3,…, n are to be added. Again the symbol S x
implies 'sum of the values of x'.
Illustration: Express with the help of S symbol:
a. x4 + x5 + x6 + x7 + x8 + x9 + x10
Solution: x4 + x5 + x6 + x7 + x8 + x9 + x10 = S10i=4 xi
26
3.5 DIFFERENT TYPES OF MEASURES OF CENTRAL TENDENCY
The following three types of averages or measures of central tendency are used:
a) Mean b) Median c) Mode
a) Mean:
Arithmetic mean is the average of a group of observations and is calculated by adding all the
numbers and then dividing the sum so obtained by the number of observations. Because it is the
arithmetic mean out of the three types of means that is most commonly used by statisticians, the
arithmetic mean is more commonly known as only 'mean'.
Here, the mean for the population is represented by the Greek letter mu ( ). And the mean for the
sample is denoted by x ?. While talking of mean, there are three types of means, namely:
i) Arithmetic mean: (AM)
ii) Geometric mean: (GM)
ii) Harmonic mean: (HM)
i) The A.M. of a variable x is denoted by the symbol x ? and is defined to be the sum of the
values of x divided by the number of values of x.
Formula of AM in case of individual series:
x: x1, x2, x3, …, xn, `x will be:
t
`x = (x1+x2+x3+?+xn) = S i=1 xi = Sx
n n n
x x1 x2 x3 … xn
f f1 f2 f3 … fn
27
the A.M. formula for discrete or ungrouped frequency distributions we can find the arithmetic
mean of a grouped frequency distribution.
åfx
\ A.M = = 555/121 = 4.59 (approx.)
N
Note: All measures of central tendency or averages of a distribution will possess the same
unit of the distribution.
Example: The average marks obtained by two groups of students in an examination are 75 and
85. If the average marks of all the students is 80, find the ratio of students in the two groups.
Solution: Let x denote marks of all the students, x1 denote marks of the first group, x2 denote
marks of the second group, n1 denote no. of students of the first group and n2 denote number of
students of the second group.
28
3.8 ADVANTAGES AND DISADVANTAGES OF A.M.
Advantages:
¡ A.M is easy to determine and understand.
¡ The A.M. is based on all values of the distribution.
¡ It can be used for further algebraic treatment.
¡ The formula for A.M is rigidly defined implying that for a given series, the value of A.M.
remains unique.
¡ It provides a good basis for comparison.
¡ The values of a series need not be arranged in any order for calculating the A.M.
¡ If the A.M. and the number of observations in the series are known, then we can also find
out the sum of the distribution.
Disadvantages:
¡ The A.M. is unduly affected by extreme (i.e, very large or small) values.
¡ The A.M. cannot be computed even if one of the values in the series is missing.
¡ The determination of the A.M. in case of a grouped frequency distribution can be
misleading as it is based on an unrealistic assumption that the observations of each class is
concentrated around the centre of that class.
29
1.7.1 Advantages and disadvantages of G.M. :
Advantages:
¡ G.M. is rigidly defined.
¡ G.M. is based on all values of the distribution.
¡ It is possible to do further mathematical treatment in case of G.M.
¡ In comparison to A.M., G.M. is less affected by extreme values.
Disadvantages:
¡ G.M. is not that easy to determine and understand.
¡ G.M. of a distribution cannot be determined if there is even one negative value in the series.
Also, if there is at least one zero value, then the G.M. will be zero.
30
The H.M. is given by:
= 4/ 0.0429 = 93.24.
Disadvantages:
¡ It is difficult to calculate and understand.
¡ If even a single value in the distribution is zero, then the H.M. cannot be computed.
¡ It gives more importance to smaller values.
Note:
¡ Although A.M. is for all values, i.e., positive, negative and zero, G.M. is defined for positive
values only and H.M. is defined for non-zero values.
¡ If x1 = x2, then A.M. = G.M. = H.M. If x1 x2, then A.M > G.M. > H.M.
¡ The above property holds true for any finite number of positive values.
3.12 MEDIAN
Median is the second type of average that is used. Median is the middle most value in a
distribution when the said distribution is arranged either in ascending or descending order. It
31
divides the distribution into two equal parts. Thus there are equal number of observations on the
right and left of the median value, i.e, the number of observations greater than and less than the
median are equal.
For determining the median of an individual series, we have to make sure that all the values in a
distribution are arranged in a definite order, i.e, whether those values are in ascending or
descending order or not. If the values are not in a definite order, then these values have to be
arranged in either an ascending or descending order. For a distribution with odd number of terms
(value), the median is the middlemost value. If there are even number of terms in the array, then
the median is the average of the two middle numbers.
Symbolically, for odd number of terms in the distribution,
Median = ((n+1)/2)th value from the beginning or the end
For even number of terms in the distribution,
Median = average of the n/2th value and the ( n/2 +1 )th value
Example: Determine median for the following series:
i. 77, 73, 72, 70, 75, 79, 78
ii 94, 33, 86, 68, 32, 80, 48, 70
Solution:
i. Arranging the values of the series in ascending order, we get
70, 72, 73, 75, 77, 78, 79
No. of terms in the series = 7
The required median = (7+1)/2 = 4th term = 75.
Note: By arranging the terms in descending order, the same value of median will be obtained.
32
3.12.2 Median of a grouped frequency distribution:
While computing the median of a grouped frequency distribution, the cumulative frequencies of
the various class intervals has to be found out first. Then we have find out the median class. The
median class is the class which contains the median value. We find out the median value by
applying the same principles we did in case of the ungrouped frequency distribution (for odd
number terms, N/2th term is median and for even number of terms, average of the N/2 and [N/2
+1] the term shall be the median). After detecting the median class, the particular median value is
determined by using the following formula:
Disadvantages:
¡ Unlike the other measures of central tendency, determination of median requires the
distribution to be arranged in a definite order if it is not in any order.
¡ Median is not based on all observations of the distribution.
¡ In comparison to mean, it is more affected by fluctuations in sampling.
Use of median:
In order to determine the average in case of distributions having open-end class intervals, median
is the best measure of central tendency. In case of income distribution, median would yield better
results.
33
Example: Determine median for the following distribution:
Solution:
Table for determining median
Weekly wages No. of workers (f) Cumulative frequency ( fc)
50-55 6 6
55-60 10 16
60-65 22 38
65-70 30 68
70-75 16 84
75-80 12 96
80-85 15 111
N = 111
Since no. of classes is 7 (odd), thus median will be ( (N+1)/2)th term. Thus Median will be
(111+1)/2 = 56th term. From the cumulative frequency table, we find that the 56th term lies in the
class 60-70. Therefore 60-70 is the median class.
3.13 MODE
The mode of a distribution is that value which occurs the most frequently in the distribution. It is
that distribution whose frequency is the maximum. It should be noted that mode is not unique
which means that a distribution may have more than one mode. Distributions that have more than
one mode are called bimodal distributions and those that have more than two modes are termed
as multimodal.
Thus, an individual distribution does not have a mode. Even in case of a discrete frequency
distribution, each observation has the same frequency and thus has no mode. In case of an
ungrouped frequency distribution, mode can be determined by observation, in most cases. In
case of a grouped frequency distribution, mode is obtained by using the following formula:
34
f1 = frequency of the modal class
f0 = frequency of the class preceding the modal class
f2 = frequency of the class succeeding the modal class
I = length of the modal class. (Instead of I, the symbol h may also be used.)
Note:
¡ The class (specified by a class interval) whose frequency is the maximum is called the
modal class.
¡ The above formula is applicable when all the classes are of equal length.
Example:
Determine the mode/modes of the following series, if any
i. 3, 4, 5, 2, 3, 4, 1, 6, 4;
ii. 7, 9, 11, 7, 6, 5, 9, 13;
iii. 3, 5, 6, 7, 9, 12, 3, 6, 5, 9, 12, 7
Solution:
i. The number 4 is repeated the maximum number of times (3 times). Hence the mode of the
distribution is 4.
ii. Here we see that both the numbers 7 and 9, appear twice in the distribution. Since the
frequency of these numbers is the highest (2), thus 7 and 9 are the two modes of these
series.
iii. In this series, the frequency of each observation is the same and hence this series has no
mode.
Disadvantages:
¡ Mode is not based on all observations.
¡ Unlike the other averages, it is not capable of further mathematical treatment.
35
Weather forecasts are also based on mode. The mode is an appropriate measure of central
tendency for nominal-level data.
Quartiles
Quartiles are those measures of central tendency which divide the distribution into four equal
parts when the distribution is arranged in an ascending order. There are three quartiles in a
distribution namely Q1, Q2 and Q3.
Deciles
The nine quantities that divide a distribution into ten equal parts are called the deciles of the
distribution. The distribution needs to be arranged in an ascending order. These are denoted by
D1, D2, D3, …, D9.
Percentiles
Percentiles divide a distribution into hundred equal parts. There are 99 percentiles because it
takes 99 dividers to separate a group of data into 100 parts. The nth percentile is the value such
that at least n percent of the data are below that value and at most (100 - n) percent are above that
value.
36
4. Find the missing frequency if the arithmetic mean is Rs 33 thousand.
Loss of sales (Rs in thousand) 0-10 10-20 20-30 30-40 40-50 50-60
No. of families 10 15 30 - 25 20
5. Calculate A.M. and median of the distribution. Hence calculate mode using empirical
relation between the three.
Class intervals 59-61 61-63 63-65 65-67 67-69
Frequency 4 30 45 15 6
6. The following list shows the 15 largest banks in the world by assets according to Standard
and Poor's. Compute the median and the mean assets from this group. Which of these two
measures do you think is most appropriate for summarizing these data, and why? What is the
value of Q2? Determine the 63rd percentile for the data. How could such information on
percentiles potentially help banking decision-makers?
Bank Assets (US$ millions)
Industrial & Commercial Bank of China 4,009
China Construction Bank Corp. 3,400
Agricultural Bank of China 3,236
Bank of China 2,992
Mitsubishi UFJ Financial group 2,785
JP Morgan Chase & Co. 2,534
HSBC Holdings 2,522
BNP Paribas 2,357
Bank of America 2,281
Credit Agricole 2,117
Wells Fargo & Co. 1,952
Japan Post Bank 1,874
Citigroup Inc. 1,842
Sumitomo Mitsui Financial Group 1,175
Deutsche Bank 1,166
No. of men 1 4 2 2 1 10
2. Which of the following relations among the location parameters does not hold?
a. Q2 = median
b. P50 = median
c. D5 = median
d. D6 = median
38
3. Extreme values have no effect on:
a. A.M.
b. Median
c. G.M.
d. H.M.
5. The mean of 8 numbers is 15. After a new number 24 is added, the new mean shall be:
a. 8
b. 16
c. 12
d. 10
40- under 60 13
60 – under 80 10
80 – under 100 19
39
Suggested reading:
¡ Chandan, J.S. Statistics for Business and Economics. New Delhi: Vikas Publishing House
Pvt Ltd., 1998
¡ Gupta, S.C. Fundamentals of Statistics. New Delhi: Himalaya Publoshing House, 2006.
¡ Kothari, C.R. Quantitative Technique. New Delhi: Vikas Publishing House Pvt. Ltd., 1984
¡ Black, K. Business Statistics: Contemporary Decisions Making. Wiley, 2009.
References
¡ Black, K. (2009, December 1). Business Statistics: Contemporary Decision Making (6th
ed.). Wiley.
¡ Padmalochan, H., & Hazarika, P. (ca. 2007, April 4). A Textbook of Business Statistics (1st
ed.) [Print]. S. Chand Limited.
¡ Peck, R., Olsen, C., & Devore, J. (2022, September 30). Introduction to Statistics and Data
Analysis (AP(R) Edition) (4th ed.). Brooks/Cole, Cengage Learning.
¡ Quantitative Techniques (New Format). (2013, January 1). Vikas Publishing House.
40
MODULE - 4 MEASURES OF DISPERSION
STRUCTURE
¡ Measures of dispersion
¡ Different measures of dispersion
¡ Absolute and Relative measures of dispersion
¡ Different measures of dispersion
¡ Range
¡ Interquartile range and Quartile Deviation (Q.D.)
¡ Co-efficient of Quartile Deviation (Q.D.)
¡ Mean Deviation (M.D.)
¡ Standard Deviation
¡ Empirical relationship between Q.D, M.D. and S.D.
¡ Coefficient of Variation
We observe that in the first series, all the values are 40. Thus, the A.M. which is 40, fully
represents the series as well as the individual items. In the second series, although the mean is 40,
the values are not scattered much as the minimum value in the series is 35 and the maximum is 43.
Thus, for the second series too, we can say that the mean is a good representative of the series.
41
Here, the discrepancy between the mean and other values isn't very high.
In case of the third series, we observe that all the values are very different. Though the mean is
same like the other two series, i.e., 40, the values are widely scattered, with 10 being the
minimum value and 80 being the maximum. Clearly, in this case, the mean neither satisfactorily
represents the series generally, nor the individual values of the series in particular.
Thus, measures of central tendency lose their effectiveness and cannot be the representative
when the extent of variation (dispersion or scatteredness) of the individual values of a
distribution in relation to their average or in relation to the other values becomes large. Hence it is
important for a statistician to not only know the average of any type, but also the scatteredness of
a distribution.
Scatteredness of data about an average is termed as dispersion or variation. Quoting Spiegel,
"The degree to which numerical data tend to spread about an average value is called variation or
dispersion."
The study of average alone without knowledge of dispersion may lead to an erroneous
conclusion. A person with the knowledge of average and without the knowledge of variability
was once travelling with his family where he had to cross a river without a boat to reach the
destination. He knew that the average depth of the river was 100cm and the average height of his
family was 130cm. So he and his family decided to swim across the river on foot. But it so
happened that the maximum depth of the river was 150cm and the height of his youngest son was
60cm. We can only fathom what must have happened to the family.
Depending upon the nature of a distribution, different methods of obtaining the representative
value have been evolved and as a result, we have several averages or measures of central
tendency.
42
original distribution is in kilograms, an absolute measure will also be in kilograms. For this
reason an absolute measure of dispersion cannot be used to compare the variability
between two or more distributions.
b. Relative measures: A relative measure of dispersion is one which is calculated as a
percentage or coefficient of an absolute measure. A relative measure of dispersion is also
known as the coefficient of dispersion.
A relative measure of dispersion is free from any unit.
4.5 RANGE
The range in a distribution is the difference between the smallest and the largest value of that
distribution. Thus if L denotes the largest observation and S denotes the smallest observation,
then Range R = L - S.
43
4.6 INTERQUARTILE RANGE AND QUARTILE DEVIATION (Q.D.)
Interquartile range is the difference between the third quartile (Q3) and the first quartile, Q1 of
the distribution.
Quartile deviation, Q.D. is half of the interquartile range of a distribution.
Thus, Interquartile range = Q3- Q1
Quartile deviation = Q3 - Q1 ) / 2
4.6.1 Advantages:
¡ Q.D. is a better measure of dispersion than range as unlike range that takes into account
only two of the values, Q.D. involves 50% of the values.
¡ Q.D. is not affected by extreme values as the lowest 25% and highest 25% of the
observations are not considered while calculating the Q.D.
¡ It is the only measure which can be used for open ended class intervals.
4.6.2 Disadvantages:
¡ Since Q.D. is based on only 50% of the observations, it disregards the other half of the
observations.
¡ It is not amenable for further mathematical treatment.
Interquartile range is specifically useful for those data users who are more interested in values
towards the middle and less interested in the extremes.
44
Here A = mean or median of x and d = x-A
Note:
i. In case of mean deviation from mean, 'A' may be AM, GM or HM but is usually taken as
AM.
ii. |x-A| is called as the absolute value of the deviation x - A. By absolute value we mean the
magnitude of the value without considering the sign.
iii. Since the sum of the deviations measured from the bus is zero, hence in case of mean
deviation we always take absolute deviations.
Disadvantages:
¡ In case of M.D. absolute values are taken and the actual signs of deviations are discarded.
¡ Mean deviation from mode is not considered to be a good measure of dispersion.
¡ For a grouped frequency distribution containing open- end class intervals, one cannot
determine mean deviation.
45
Example: For the following distribution determine the mean deviation (M.D.) from mean and its
coefficient.
46
¡ It is standard deviation which can be considered as the basis of sampling theory and
correlation analysis.
Limitations:
¡ The biggest disadvantage is that S.D. is difficult to calculate in comparison to the other
measures of variability.
¡ Also, as compared to M.D., it is more affected by extreme values.
Example: Find standard deviation of the following observation:
8, 10, 12, 14, 16, 18, 20, 22, 24, 26
Solution:
4.8 VARIANCE
Variance is the square of standard deviation. Thus, we can say that standard deviation is the
positive square root of variance. Thus, variance = 2 where implies standard deviation and
47
4.10 COEFFICIENT OF VARIATION
Coefficient of variation is a relative measure of dispersion developed by Professor Karl Pearson.
Abbreviated as C.V., it is useful in comparing the variability of two or more sets of data especially
if they are expressed in different units of measurement. C.V. is expressed as a percentage. The
formula for C.V. is given as follows:
The coefficient of variation is the most popular relative measure of dispersion. While comparing
the variability between two distributions, the distribution having the minimum C.V., is
considered to less variable, more stable, more consistent, uniform or more homogenous. On the
other hand, the distribution which has higher C.V. is said to be more variable, less stable, less
consistent or homogenous.
Coefficient of variation can be sued in another situation, say, for comparing the relative
consistency of the prices of shares of two companies. It will help an investor to decide which
company's share prices are relatively stable and he can choose to invest in the same. The shares
which are more stable or consistent in the fluctuation of prices will be preferred by the investor.
As an example, suppose the scores of a cricket player A in three matches are 0, 10 and 80. The
scores of another cricket player B in these matches are 28, 30 and 32. Although the mean scores
of both the players are same, both being 30, we can say that B is more consistent than A. even
people without having the knowledge of variability or dispersion will say this. If we calculate the
C.V. for both the series of scores, we shall find that C.V. score of B is less than the C.V. score of A.
In fact, the magnitude of any measure of dispersion will be less in case of scores attained by B
than the scores attained by A.
6. The following frequency distribution gives the height (in inches) of 100students selected at
random from a college having 3000 students. Calculate standard deviation.
Weight (kg) 30-34 35-39 40-44 45-49 50-54
No. of boys 5 11 26 10 8
48
Short Answer Questions
1. Discuss any two objectives of measuring variability.
2. What is variance?
3. What do you meant by coefficient of variation?
4. Explain in what situations range serves as a useful measure of variability.
5. What is the empirical relationship between QD, MD and SD?
6. For a distribution, the coefficient of variation is 22.5% and the value of arithmetic average
is 7.5. Find out the value of standard deviation.
49
2. Which of the following are characteristics of a good measure of dispersion?
a. It should be easy to calculate
b. It should be based on all the observations within a series
c. It should not be affected by the fluctuations within the sampling
d. All of the above
4. While calculating the standard deviation, the deviations are only taken from
a. The mode value of a series
b. The median value of a series
c. The quartile value of a series
d. The mean value of a series
50
Match the following
Column A Column B
1 Measures of dispersion A Open ended class interval
2 Absolute measure B Scatteredness
3 Coefficient of variation C Less consistency
4 Quartile deviation D Karl Pearson
5 Easiest measure of dispersion E Mean deviation
6 Variability F Range
(Answer Key: 1 = B, 2 = E, 3 = D, 4 = A, 5 = F, 6 = C)
51
Suggested reading:
¡ Chandan, J.S. Statistics for Business and Economics. New Delhi: Vikas Publishing House
Pvt Ltd., 1998
¡ Gupta, S.C. Fundamentals of Statistics. New Delhi: Himalaya Publishing House, 2006.
¡ Kothari, C.R. Quantitative Technique. New Delhi: Vikas Publishing House Pvt. Ltd., 1984
¡ Black, K. Business Statistics: Contemporary Decisions Making. Wiley, 2009.
References
¡ Black, K. (2009, December 1). Business Statistics: Contemporary Decision Making (6th
ed.). Wiley.
¡ Padmalochan, H., & Hazarika, P. (ca. 2007, April 4). A Textbook of Business Statistics (1st
ed.) [Print]. S. Chand Limited.
¡ Peck, R., Olsen, C., & Devore, J. (2022, September 30). Introduction to Statistics and Data
Analysis (AP(R) Edition) (4th ed.). Brooks/Cole, Cengage Learning.
¡ Quantitative Techniques (New Format). (2013, January 1). Vikas Publishing House.
52
MODULE - 5 MEASURES OF DISPERSION
STRUCTURE
¡ Z score
¡ Skewness
¡ Kurtosis
¡ Empirical Rule
¡ Chebyshev's Theorem
LEARNING OBJECTIVES
After going through this unit, you will be able to:
¡ Understand a few other characteristics of a distribution
¡ Explain what z score is
¡ Learn the measures of shape of a distribution
¡ Understand the ways of applying standard deviation
5.1 Z-SCORE
A z score represents how far a data point is from the mean. How many standard deviations above
or below the population mean a raw value (x) is when the data is normally distributed is shown
by the z score. The use of z scores enables translating a value's raw distance from the mean into
units of standard deviations.
A negative z score implies that the raw value of x lies below the mean whereas a positive z score
implies the contrary, i.e., the raw value (x) is above the mean. For example, for a data set that is
normally distributed with a mean of 50 and a standard deviation of 10, suppose a statistician
wants to determine the z score for a value of 70. This value (x = 70) is 20 units above the mean, so
the z value is
Here, the z score signifies that the raw value of 70 is two standard deviations above the mean. The
z score is interpreted using the empirical rule which states that 95% of all values in a distribution
53
are within two standard deviations of the mean if the data set follows a normal distribution.
As a z score is the numerical score of standard deviations an individual data point is from the
mean, the empirical rule can also be stated in terms of z score.
Between z = -1.00 and z = +1.00 are approximately 68% of the values.
Between z = -2.00 and z = +2.00 are approximately 95% of the values.
Between z = -3.00 and z = +3.00 are approximately 99.7% of the values.
5.3 SKEWNESS
The concept of skewness can be acquired from the concept of symmetricity. If a distribution of
data is such that the right half is a mirror image of the left half, such a distribution is said to be
symmetrical. Statistically, a distribution is said to be symmetric about it's mean (A.M.) if the
observations of the distribution equidistant from the A.M. have equal frequencies.
Let us consider the following distribution:
It can be observed that the values of x, which are equidistant from 6 are 4 and 8, 2 and 10. 4 and 8
have the same frequency 3, whereas 2 and 10 have the same frequency of 1.
A distribution may also be said to be symmetric about it's A.M. if the deviations of the values
from their A.M. are such that corresponding to each positive deviation there is a negative
deviation of equal magnitude.
A symmetric distribution
54
Coming to skewness, a distribution which is asymmetrical or lacks symmetry is called a skewed
distribution. Thus, in case of a skewed distribution, the magnitudes of the positive and negative
deviations of the values from the mean are unequal or do not balance.
The figure on the left represents a negatively skewed distribution whereas the figure on the right
represents a positively skewed distribution.
Now for a positively skewed distribution which has a longer tail towards the right, like the figure
on the right, A.M. Median Mode. Whereas for a negatively skewed distribution where the tail
lengthens towards the left, A.M. Median Mode.
For a symmetric distribution, A.M. = Median = Mode.
55
Since the absolute value of skewness of (ii) is greater than the absolute value of (I), hence
distribution (ii) is more skewed.
5.4 KURTOSIS
Kurtosis is concerned with the flatness or peakedness of frequency curve - the graphical
representation of a frequency distribution.
From the standpoint of Kurtosis, the normal curve is termed as mesokurtic, i.e, of intermediate
peakedness. A curve which is more peaked than the normal curve is called leptokurtic and a curve
flatter than the normal curve is called platykurtic.
56
Example: A company produces a lightweight valve that is specified to weigh 1365 grams.
Unfortunately, because of imperfections in the manufacturing process not all of the valves
produced weigh exactly 1365 grams. In fact, the weights of the valves produced are normally
distributed with a mean weight of 1365 grams and a standard deviation of 294 grams. Within
what range of weights would approximately 95% of the valve weights fall? Approximately 16%
of the weights would be more than what value? Approximately 0.15% of the weights would be
less than what value?
Solution: Because the valve weights are normally distributed, the empirical rule applies.
According to the empirical rule, approximately 95% of the weights should fall within m ± 2s
=1365 ± 2(294) = 1365 ± 588. Thus, approximately 95% should fall between 777 and 1953.
Approximately 68% of the weights should fall within m ± 1s, and 32% should fall outside this
interval. Because the normal distribution is symmetrical, approximately 16% should lie above
m± 1s = 1365 + 294 = 1659. Approximately 99.7% of the weights should fall within ± 3s, and
.3% should fall outside this interval. Half of these, or .15%, should lie below m - 3s = 1365 -
3(294) = 1365 - 882 = 483.
57
2
.20 = 1/ k
2
=>k = 5.000
k = 2.24
Chebyshev's theorem says that at least .80 of the values are within ±2.24 of the mean. For m = 28
and s = 6, at least .80, or 80%, of the values are within 28 ±2.24(6) = 28 ± 13.4 years of age or
between 14.6 and 41.4 years old.
58
5. For a symmetric distribution, _________ (A.M. = Median = Mode)
6. A curve flatter than the normal curve is called __________ (Platykurtic)
2. If mean, median, and mode are all equal then distribution will be
a. Positive Skewed
b. Negative Skewed
c. Symmetrical
d. None of these
59
5. A curve whose tail is longer to the right is called
a. Negatively skewed
b. Positively skewed
c. Symmetrical
d. None of these option
Suggested reading:
¡ Chandan, J.S. Statistics for Business and Economics. New Delhi: Vikas Publishing House
Pvt Ltd., 1998
¡ Gupta, S.C. Fundamentals of Statistics. New Delhi: Himalaya Publishing House, 2006.
¡ Kothari, C.R. Quantitative Technique. New Delhi: Vikas Publishing House Pvt. Ltd., 1984
¡ Black, K. Business Statistics: Contemporary Decisions Making. Wiley, 2009.
References
¡ Black, K. (2009, December 1). Business Statistics: Contemporary Decision Making (6th
ed.). Wiley.
¡ Padmalochan, H., & Hazarika, P. (ca. 2007, April 4). A Textbook of Business Statistics (1st
ed.) [Print]. S. Chand Limited.
¡ Peck, R., Olsen, C., & Devore, J. (2022, September 30). Introduction to Statistics and Data
Analysis (AP(R) Edition) (4th ed.). Brooks/Cole, Cengage Learning.
¡ Quantitative Techniques (New Format). (2013, January 1). Vikas Publishing House.
60
MODULE - 6 INTRODUCTION TO PROBABILITY
STRUCTURE
¡ Introduction
¡ Usefulness of the theory
¡ Definition of useful terms
¡ Formula
¡ Dependent and Independent events
¡ Approaches
¡ Classical or Mathematical Approach to Probability
¡ Relative frequency of Occurrence
¡ Subjective Probability
¡ Types of probabilities
¡ Conditional Probability
¡ Symbols associated with probability
¡ Addition and Multiplication Theorem
¡ Bayes' Theorem
¡ Questions and Exercises
LEARNING OBJECTIVES
After going through this unit, you will be able to:
¡ Explain what the meaning of probability is
¡ Understand the usefulness of probability
¡ Describe events and their role in probability
¡ Explain the different approaches to probability
¡ Explain marginal, join, union probabilities
¡ Explain conditional probability
¡ Solve probability problems using addition and multiplication rules
¡ Describe Bayes' Theorem Unit Contents
61
certainty.
Consider the following experiments:
a. Zinc is mixed with dilute sulphuric acid.
b. A uniform coin is tossed upwards.
c. A dice is rolled.
The result of experiment (a) is zinc sulphate and hydrogen produced from the chemical reaction
between zinc and sulphuric acid. However, for the experiment (b), the result, i.e., whether head
or tail will turn up, is uncertain. Experiment (c), too, might give any result between 1-6.
Situations such as (b) and (c) are termed as random experiments and the result of such an
experiment is called an event. The likelihood of a particular event occurring is defined as
Probability. It is the measure of certainty of an event. The numerical probability value lies
between 0 and 1, where 0 denotes the impossibility of the event occurring and 1 denotes
assurance of the same event happening.
62
random experiment where each event has got an equal chance of occurring.
Mutually exclusive events: Two or more events are said to be mutually exclusive if one event
prevents the occurrence of the other event. While tossing a coin, the occurrence of either head or
tail is an example of mutually exclusive event.
Complementary event: The complementary to an event A, is denoted by A . All the events of
an experiment not in A are considered its complement. By formula, A = 1 - A.
Exhaustive events: Exhaustive events are those which include all the possible outcomes of a
random experiment. If a coin is tossed, the two events, namely head (H) and tail (T) will be the set
of exhaustive events in this situation.
Sample space: The set of all possible outcomes of a random experiment is called the sample
space. When we roll a die, any of the six faces between 1-6 can turn up and hence the sample
space will be S= {1, 2, 3, 4, 5, 6}.
Subjective Probability
This approach to probability is based on a measure of the feelings and insights of the person
determining the probability. It measures the degree of a person's belief or subjective assessment
that the event will occur. The subjective approach is used when the other probability approaches
cannot be used-examples of situations like electoral outcomes, results of cricket matches, etc. In
a business situation, it can be the likelihood of a successful product launch, investment outcomes
etc. Though it is not a mathematical or scientific approach, this method uses the information,
experience, knowledge, stored and processed in the human mind. Sometimes, it can be merely an
estimate. In other cases, human experience and judgement can yield accurate probabilities.
64
person owns an Audi car. This probability will be obtained by dividing the number of Audi
owners by the total number of car owners. Another example could be the probability of a person
who wears spectacles which can be obtained by dividing the number of spectacle wearers to the
total number of people.
Symbolically, Marginal probability is given by P(E)= n(E)/ n(S)
where n(E) = Number of event favourable to event E
n(S) = Total number of outcomes
Another type of probability is Union probability which denotes the union of two events. The
union probability of two events, A and B, is given by P(A È B), which is the probability that either
A will occur or B will occur or both A and B will occur. An example of union probability is the
probability that a person owns an Audi or a Ford. The condition for union probability is that the
person has to own at least an Audi or a Ford. Another situation could be the probability of
someone having red hair or a tattoo. In this union, all people who have red hair are included,
along with people who have a tattoo and all redheads who have a tattoo. In an organization, the
probability that an employee is a clerical worker or a male is union probability.
Joint probability is the third type of probability which is the intersection of events denoted as
P(AÇ B), for two events A and B. To qualify for a joint probability, both events must occur. An
example of joint probability will be the probability of a person owning both an Audi and a Ford.
Merely owning either an Audi or a Ford will not be sufficient for joint probability. The
probability of a person having red hair and a tattoo is an example of joint probability.
Note: If B does not depend on A, then the symbol used is P(B) and not P(B|A). In this case, the
probability of event B occurring is not dependent on occurrence of event A. P ( B|A) is read as
'probability of B given A' or 'probability B post A'.
As an example, let us suppose that we roll a die and it is known that the number that came up is
greater than 3. We want to find out the probability that the outcome is an even number greater
than 3.
Let Event A = even
And Event B = larger than 3
65
Then P( even| greater than 4) = [P (even and greater than 3)/ P(greater than 3)]
Or
P (A/B) = P (AB)/ P(B)
= (1/6) (2/6) = ½.
66
P (A or B) = P(A) + P (B)
However, if the two events are not mutually exclusive, then the probability of either event A or B
occurring is given by the probability that event A occurs plus the probability that event B occurs,
subtracted by the common probability of both A and B occurring.
Symbolically, it can be written as:
P( AðB) = P(A) + P(B) - P(AðB)
67
Based on this information, the probability that a student picked up at random will be female is
40/60 or 0.67 since there are 40 females out of 60 students. Now let us assume that we are given
additional information that the person picked up at random is Indian, then what is the probability
that this person is a female? This additional information will result in the revised probability or
posterior probability in that it is assigned to the outcome after the additional information is made
available.
Since we want to determine the revised probability of picking a female student at random,
provided we know that the student is an Indian, so let A1 be the event female, A2 be the event
male, and B be the event Indian. Thus, based on our knowledge of conditional probability, Bayes'
Theorem shall be as follows:
P(A1 |B) = {P(A1) P (B| A1)}/{ P(A1) P(B|A1) + P(A2) P (B|A2)}
In the above example, there are 2 basic events which are A1 (female) and A2 (male). However, if
there are n basic events, A, A2, …, An, then Bayes' Theorem can be generalized as,
P(A1|B) = P(A1) P(B|A1)}/{P(A1) P (B|A1) + P(A2) P(B|A2) + … + P(An) P(B|An)}
Solving the previous two events, P(A1|B) =
= {(40/60) (20/40)}/{(40/60) (20/40) + (20/60) (15/20)}
= 20/35 = 4/7 = 0.57
This example demonstrated clearly how the probability of picking up a female student was 0.67
however, after the additional information received that the student is a foreigner, the posterior
probability becomes 0.57.
68
4. Explain Bayes' Theorem.
5. Explain subjective probability.
6. What do you mean by dependent events?
MCQs
1. If A is an uncertain event, then
a. P(A)³0
b. 0£P(A)£1
c. 0<P(A) < 1
d. None of the above
69
c. Equally likely
d. Independent
3. If P(A?B) [or P(AB)] is zero then the two events A and B are:
a. Mutually exclusive
b. Exhaustive
c. Equally likely
d. Independent
70
c. What is the probability that the business is in the construction industry if it is known that
the business is located in the South Atlantic states?
d. What is the probability that the business is located in the South Atlantic states if it is known
to be a construction business?
e. What is the probability that the business is not located in the South Atlantic states if it is
known that it is not a construction business?
Suggested reading:
¡ Chandan, J.S. Statistics for Business and Economics. New Delhi: Vikas Publishing House
Pvt Ltd., 1998
¡ Gupta, S.C. Fundamentals of Statistics. New Delhi: Himalaya Publishing House, 2006.
Kothari, C.R. Quantitative Technique. New Delhi: Vikas Publishing House Pvt. Ltd.,
1984
¡ Black, K. Business Statistics: Contemporary Decisions Making. Wiley, 2009.
References
¡ Black, K. (2009, December 1). Business Statistics: Contemporary Decision Making
(6th ed.). Wiley.
¡ Padmalochan, H., & Hazarika, P. (ca. 2007, April 4). A Textbook of Business Statistics (1st
ed.) [Print]. S. Chand Limited.
¡ Peck, R., Olsen, C., & Devore, J. (2022, September 30). Introduction to Statistics and Data
Analysis (AP(R) Edition) (4th ed.). Brooks/Cole, Cengage Learning.
¡ Quantitative Techniques (New Format). (2013, January 1). Vikas Publishing House.
71
DISCRETE PROBABILITY
MODULE - 7
DISTRIBUTION
STRUCTURE
¡ Introduction to a random variable
¡ Discrete Probability distribution
¡ Expected Value
¡ Variance
¡ Binomial Distribution
¡ Bernoulli Trials
¡ Additional properties of binomial distribution
¡ Hypergeometric Distribution
¡ Poisson Probability Distribution
¡ Questions and Exercises
Sample Point HH HT TH TT
X 2 1 1 0
Corresponding to the sample point HH, we get X=2, to each of HT and TH we get X=2 and for the
sample point TT, X=0.
72
(T, T)
0
(H, T)
(T, H)
(H, H) 2
Through this example, we learn that each sample point may be assigned a numerical value of X,
and more than one sample point can be given the same value. Yet another observation is that
variable X assumes each value with a definite probability, and the sum of all the probabilities
results in unity (1).
We can thereby say that, a random variable is a numerical valued function defined on a sample
space. A random variable can be divided into discrete and continuous random variables.
The result of a random experiment is associated with the random variable, and depending upon
different trials, the value changes and hence is termed a random variable. Say the number of girls
in a three-child family, the number of light bulbs that will turn out defective in a bundle of 150
bulbs, etc. Such a value cannot be predicted with certainty and thus is a random variable.
When a random variable takes on a finite number of values or a countably infinite number of
variables, such a variable is called a discrete random variable. In the majority of statistical
situations, discrete variables take non-negative whole numbers.
When a random variable takes on an uncountably infinite number of values, it is called a
continuous random variable. Such variables take on matters at every point in a given range; the
main distinction between discrete and continuous random is that the former involves counting,
while the latter involves measuring.
A discrete variable may take the form of events like
¡ number of broken eggs in a carton,
¡ the total number of items purchased,
¡ age of individuals
¡ number of defects in a batch of 60 items Continuous variables will include measurements
like
73
¡ Height and weight of individuals
¡ amount of time spent in the store
¡ measuring the time between customer arrivals at a retail outlet
Note: Random variables are also called variates and are denoted by capital letters X, Y…,
whereas their specific values are denoted respectively by small letters x, y…
Example
i) Discrete distribution of the number of power cuts in a day (Table 4.1)
An executive is considering taking leave on a Friday and holding a Zoom meeting online for her
clients' presentation at home. She recognizes that the area of her residence is prone to power cuts
and is concerned about the possibility of the same happening on the said day. The above table
shows a discrete distribution that contains the number of power cuts that could occur on the day
of her presentation and the probability that each number of power cuts will occur. It is apparent
from the table that the most likely number of power cuts is 0 and, 1 with a probability of .37 and
.31, respectively.
74
ii) If two unbiased coins are thrown and if X denotes the number of heads turning up, then we
have
P (HH) = ¼, P(HT) = ¼, P(TH) = ¼, P(TT) = ¼ Then P(X=0) = P(TT) = ¼,
P(X=1) = P(HT+TH) = P(HT) + P(TH) = ¼ + ¼ ½
P(X=2) = P(HH) = ¼
Consequently, we have the following probability distribution of the discrete random variable X:
X=x 0 1 2
P(x) 1/4 1/2 1/4
iii) If three unbiased coins are thrown, then the probability distribution of the discrete random
variable X denoting the number of heads will be as follows
X=x 0 1 2 3
P(x) 1/8 2/8 3/8 1/8
iv) If X denotes the sum of points shown by a pair of dice when they are thrown, then P(X=x) =
P(x) can be displayed as:
X=x 2 3 4 5 6 7 8 9 10 11 12
P(x) 1/36 2/36 3/36 4/36 5/36 6/36 5/36 4/36 3/36 2/36 1/36
The probability values and the different values of X can be represented graphically to obtain the
probability histogram, probability polygon and probability curve.
Expected Value
In the above example, we may observe that the probability of 6/36, where X takes the value of
7, is the highest. But 7 is also the mean or average of the values of X. Thus, the mean or the
average value is the most probable or the most likely. This mean or average value of the random
variable X is the expected value of the random variable X.
This expected value of a discrete distribution is the long-run occurrence of averages. It is also the
mean of the discrete random variable. If a trial is held only once, then a discrete random variable
yield result only once. However, if the trial is repeated long enough, the average of the outcomes
is more likely to yield a mean or an expected value. The mean or expected value is computed by
first multiplying each possible x value by the probability of observing that value and then adding
up the resulting quantities.
Symbolically,
µ = E(x) = S [x . P(x)]
Where
E = long run average
x = an outcome
P(x) = probability of that outcome
75
Computing the expected value from Table 4.1
x P(x) x P(x)
0 .37 0
1 .31 .31
2 .18 .36
3 .09 .27
4 .04 .16
= 1 .1
7.4 VARIANCE
The variance of a discrete random variable is calculated by first subtracting the mean from each
possible value of x to obtain the deviations, then squaring each deviation and multiplying the
result by the probability of the corresponding x value, and then adding all those quantities
together.
Symbolically,
2 2
s x=S(x-m) p(x)
Where x = an outcome
p(x) = probability of a given outcome
m = mean
The standard deviation of x, denoted by ðx, is the square root of the variance. The formula for
standard deviation is given by:
76
The variance s 2x and standard deviation s x (of x) occurs when the probability distribution
describes how x values are distributed among members of a population (so that the probabilities
are population relative frequencies).
Example 1
A person plays a game of throwing a die under the condition that he could get as many rupees as
the number of points on the upper most face. Find the expectation and variance of his winnings.
Solution: Let X denote the amount received by the person. Then X will be a random variable
taking the values Rs 1, Rs 2, Rs 3, Rs 4, Rs 5, Rs 6 with probabilities 1/6 each.
Now the expectation of the person's winnings, i.e, E(X) is given by
77
3. Each trial is independent of other trials.
4. The probability of success 'p' remains constant from trial to trial. Similarly, the probability
of failure remains the same.
Subject to the fulfilment of the above conditions the probability of x successes in n trials (x< n )
denoted by P(x) can be derived as:
P(X= x) = P(x) = nCx px qn-x , (x= 0, 1, 2…, n) (i)
p+q = 1, p > 0, q> 0
The binomial random variable x is defined as
x = number of successes observed when a binomial experiment is performed. This probability
distribution is called the binomial probability distribution or binomial distribution where x is
called the binomial variate.
78
ii) A fixed probability of outcome in any trial and,
iii) Independent trials.
Example: There are 20 lottery tickets with 3 prizes. Find the probability that out of 5 tickets
purchased exactly two prizes are won.
Here,
N = 20, n = 5, A = 3, x = 2
P(x) = (ACx . N-ACn-x )/NCn
= (3C2. 17C3)/ 20C5 = 5/38
The parameters of a hypergeometric distribution are N, A and n. Creating a table in
hypergeometric distribution is impossible because of the multitude of combinations possible of
these three parameters. Thus, each probability in this case has to be calculated and it makes a very
time consuming and tedious task for the researcher. Because of this reason, most researchers use
hypergeometric distribution only in cases of working binomial problems without replacement.
79
Thus, a hypergeometric distribution used be used as an alternative only in the following cases:
¡ Sampling is being done without replacement.
¡ n ³ 5% N.
Hypergeometric probabilities are calculated under the assumption of equally likely sampling of
the remaining elements of the sample space.
80
As an example,
Suppose bank customers arrive randomly on weekday afternoons at an average of 3.2 customers
every 4 minutes. What is the probability of exactly 6 customers arriving in a 5- minute interval on
a weekday afternoon? The lambda for this problem is 3.2 customers per 4 minutes. The value of x
is 6 customers per 5 minutes. The probability of 6 customers randomly arriving during a 5-
minute interval in the face of a long run average has been 3.2 customers per 4-minute interval is
(3.26) (e-4.2 ) = 1073.74 (.0408) = .0608
6! 720
If a bank averages 3.2 customers every 4 minutes, the probability of 6 customers arriving during
any one 5-minute interval is .0608.
81
3. What are the types of probability distributions?
4. When is the Poisson distribution used?
5. Who propounded the Binomial distribution?
6. List the properties of hypergeometric distribution.
4. When two unbiased coins are tossed together, the number of elements in the sample space
shall be:
a. 2
b. 3
c. 4
d. 0
83
Problem solving activities
Suppose that in the bookkeeping operation of a large corporation the probability of a recording
error on any one billing is .005. Suppose the probability of a recording error from one billing to
the next is constant, and 1,000 billings are randomly sampled by an auditor.
a. What is the probability that fewer than four billings contain a recording error?
b. What is the probability that at least 10 billings contain a billing error?
c. What is the probability that all 1,000 billings contain no recording errors?
Suggested reading:
¡ Chandan, J.S. Statistics for Business and Economics. New Delhi: Vikas Publishing House
Pvt Ltd., 1998
¡ Gupta, S.C. Fundamentals of Statistics. New Delhi: Himalaya Publoshing House, 2006.
Kothari, C.R. Quantitative Technique. New Delhi: Vikas Publishing House Pvt. Ltd.,1984
¡ Black, K. Business Statistics: Contemporary Decisions Making. Wiley, 2009.
References
¡ Black, K. (2009, December 1). Business Statistics: Contemporary Decision Making (6th
ed.). Wiley.
¡ Padmalochan, H., & Hazarika, P. (ca. 2007, April 4). A Textbook of Business Statistics (1st
ed.) [Print]. S. Chand Limited.
¡ Peck, R., Olsen, C., & Devore, J. (2022, September 30). Introduction to Statistics and Data
Analysis (AP(R) Edition) (4th ed.). Brooks/Cole, Cengage Learning.
¡ Quantitative Techniques (New Format). (2013, January 1). Vikas Publishing House.
84
MODULE - 8 CONTINUOUS PROBABILITY
DISTRIBUTIONS
STRUCTURE
¡ Introduction to continuous Probability Distribution
¡ Uniform distribution
¡ Determining probabilities in a Uniform distribution
¡ Normal distribution
¡ History of normal distribution
¡ Properties of normal distribution
¡ Standardized normal distribution
¡ Importance of normal distribution
¡ Exponential distribution
¡ Properties of exponential distribution
¡ Exponential probability density function
¡ Using the normal curve to approximate binomial distribution problems
¡ The normal curve and discrete variables
¡ Questions and Exercises
LEARNING OBJECTIVES
After going through this unit, you will be able to:
¡ Understand continuous probability function
¡ Explain uniform distribution, normal and exponential distribution
¡ Understand the standardized normal distribution
¡ Understand approximation of normal curve to discrete probability function Unit Contents
Graphically, this distribution is represented as a rectangle where b-a is the base and 1/b-a is the
height. The increase in distance between a and b proportionately decreases the density at any
particular value within the distribution boundaries. Since the probability density function
integrates to 1, the height of the probability density function decreases as the base length
increases.
The formula for Uniform distribution is given by:
In uniform distribution, the total area under the curve is 1 and is equal to the product of the length
and width of the rectangle. As the distribution lies, by definition, between the x values of a and b,
the length of the rectangle is (b-a). Combining this with the fact that the area of the curve is 1, we
can find the height of the rectangle in the following way:
Area of rectangle = (Length) (Height) = 1
But, Length = b-a
Therefore, (b-a) (Height) = 1
And Height = 1/ (b-a)
These calculations show why, between the x values of a and b, the height of the distribution
remains constant (1/b-a).
86
The mean and standard deviation of a Uniform distribution is given by:
µ= (a+b)/ 2
s = (b-a) /Ö12
As an example of this situation, suppose a production line is set up to manufacture machine
braces in lots of five per minute during a shift. When the lots are weighed, variation among the
weights is detected, with lot weights ranging from 41 to 47 grams in a uniform distribution. The
height of this distribution is
Example
Suppose the amount of time it takes to assemble a plastic module ranges from 27 to 39 seconds
and that assembly times are uniformly distributed. Describe the distribution. What is the
probability that a given assembly will take between 30 and 35 seconds? Fewer than 30 seconds?
Solution
f(x) = 1/ 39?27 = 1 /12
µ = (a+b)/2 = (39+27)/2 = 33
s = b-a/ Ö12
= 39?27 /Ö12
= 12 / Ö12
= 3.464
87
because the distribution finds its application in many problems. It fits quite a few human
characteristics such as height, weight, life expectancy, IQ, etc. Other species such as animals,
insects, and trees also have characteristics that can be attributed to be possessing the
characteristics of a normal distribution. Certain variables in industry and business, too, are seen
to be normally distributed. Besides these variables, the normal distribution is an integral part of
statistics. When large sample sizes are taken for experiment, the statistics tend to be normally
distributed regardless of the kind of distribution from which they are drawn. Karl Gauss, a
mathematician-astronomer of the 18th century is associated with the normal distribution and
hence, this distribution is also called the Gaussian distribution.
A z score signifies the number of standard deviations that a value, x, is above or below the mean.
If x is above the mean, then z score is positive, if value of x is less than the mean, we have a
negative z score and when x = z, the associated z score is zero. The z distribution is a normal
distribution with a mean of 0 and a standard deviation of 1.
89
c) It is skewed to the right.
d) The x values range from zero to infinity.
e) Its apex is always at x = 0
f) The curve steadily decreases as x gets larger.
Example:
A manufacturing firm has been involved in statistical quality control for several years. As part of
the production process, parts are randomly selected and tested. From the records of these tests, it
has been established that a defective part occurs in a pattern that is Poisson distributed on the
average of 1.38 defects every 20 minutes during production runs. Use this information to
determine the probability that less than 15 minutes will elapse between any two defects.
Solution
The value of l is 1.38 defects per 20-minute interval. The value of µ can be determined by
µ = 1/ l = 1/138 = .7246
On the average, it is .7246 of the interval, or (.7246) (20 minutes) = 14.49 minutes, between
defects. The value of x0 represents the desired number of intervals between arrivals or
occurrences for the probability question. In this problem, the probability question involves 15
minutes and the interval is 20 minutes. Thus x0 is 15 20, or .75 of an interval. The question here is
to determine the probability of there being less than 15 minutes between defects. The probability
formula always yields the right tail of the distribution-in this case, the probability of there being
15 minutes or more between arrivals. By using the value of x0 and the value of , the probability
of there being 15 minutes or more between defects can be determined.
The probability of .3552 is the probability that at least 15 minutes will elapse between defects.
90
8.5 USING THE NORMAL CURVE TO APPROXIMATE BINOMIAL
DISTRIBUTION PROBLEMS
For certain kinds of the binomial distribution, they can be approximated using normal
distribution. As a general trend, when sample sizes become large, binomial distribution
approaches a normal distribution in shape regardless of the value of p.
Whenever a continuous distribution is used to approximate a discrete distribution, the question
that naturally arises is how good is the approximation? Statistics instructors usually state that it
depends. In case of the normal approximation to the binomial, the goodness of fit depends on two
quantities which define the binomial distribution: n and p. Most statisticians go by a simple "rule
of thumb" they apply while approximating for binomial with a normal distribution, such as:
When either np < 10 or n(1 -p) <10, the binomial distribution is too skewed for the normal
approximation to give accurate results.
It can be debated that using the normal distribution to approximate another discrete distribution
which we can evaluate exactly sounds a bit too good to be true. But the point to remember is that
we cannot always find exact probabilities in other situations in statistics and we will have to rely
on approximations. This involves a tradeoff between ease of calculation and exactness of answer.
An understanding of this while using the normal approximation to binomials will give us a
clearer understanding of the issues involved.
91
Figure 1.1. In such cases, it is accustomed to say that x has approximately a normal distribution.
The normal distribution can then be used to calculate approximate probabilities of events
involving x.
Example
Premature babies are those born more than 3 weeks early. Newsweek (May 16, 1988) reported
that 10% of the live births in the United States are premature. Suppose that 250 live births are
randomly selected and that the number x of "preemies" is determined.
Solution Now, because
n p = 250(.1) = 25 ³10
n(1- p = 250 (0.9) = 225 ³ 0
x has approximately a normal distribution, with
µ= 250 (0.1) = 25
s = Ö(250(0.1)(0.9) )= 4.743
The probability that x is between 15 and 30 (inclusive) is
P( 15 £ x £ 30) = P( 14.5?25)/4.743 £z £ 30.5 - 25 / 4.743
= P (-2.21£ z £ 1.1.6)
= .8770 - .0136
= .8634
92
Short Answer Questions
1. What do you mean by continuous probability distribution?
2. Narrate a brief history of normal distribution.
3. How would you determine probability under a uniform distribution?
4. List some of the properties of normal distribution.
5. What is exponential distribution? State the properties.
6. Suppose that heights of all cakes baked with a certain mix closely follow a normal
distribution with mean 5.3 cm and standard deviation 0.75 cm. find the percentage of cakes
which have a height of 4.4 cm or less.
93
2. Following is the example of a continuous probability distribution:
a. Discrete probability distribution
b. Poisson distribution
c. Normal distribution
d. Binomial distribution
3. A variable that can assume any value between two given points is called
a. Uncertain random variable
b. Irregular random variable
c. Discrete random variable
d. Continuous random variable
6. If value of interval a is 2.5 and the value of interval b is 3.5 the value of ? for uniform
distribution is
a. 0.5
b. 3
c. 2.5
d. 2
94
Match the following
Column A Column B
1 Continuous variable A F distribution
2 Exponential distribution B =a+b
3 Continuous probability function C Abraham de Moivre
4 Normal distribution D Height of individuals
5 Uniform distribution E Normal curve
6 Approximation of discrete probability distribution F
(Answer Key: 1 = D, 2 = F, 3 = A, 4 = C, 5 = B, 6 = E)
Suggested reading:
¡ Chandan, J.S. Statistics for Business and Economics. New Delhi: Vikas Publishing House
Pvt Ltd., 1998
¡ Gupta, S.C. Fundamentals of Statistics. New Delhi: Himalaya Publoshing House, 2006.
Kothari, C.R. Quantitative Technique. New Delhi: Vikas Publishing House Pvt. Ltd., 1984
¡ Black, K. Business Statistics: Contemporary Decisions Making. Wiley, 2009.
References
¡ Black, K. (2009, December 1). Business Statistics: Contemporary Decision Making (6th
ed.). Wiley.
¡ Padmalochan, H., & Hazarika, P. (ca. 2007, April 4). A Textbook of Business Statistics (1st
ed.) [Print]. S. Chand Limited.
¡ Peck, R., Olsen, C., & Devore, J. (2022, September 30). Introduction to Statistics and Data
Analysis (AP(R) Edition) (4th ed.). Brooks/Cole, Cengage Learning.
¡ Quantitative Techniques (New Format). (2013, January 1). Vikas Publishing House.
95
SAMPLING AND DIFFERENT
MODULE - 9
SAMPLING TECHNIQUES
STRUCTURE
¡ Sampling
¡ Reasons for sampling
¡ Random Versus Non-random Sampling
¡ Types of Random Sampling
LEARNING OUTCOMES
1. To understand the meaning of sampling
2. To understand random sampling methods
3. To understand non-random sampling methods
9.1 SAMPLING
This Chapter explains the sampling and sampling distribution process of some statistics. How do
get the data used for statistical analysis? Why do the researchers usually do the selection rather
than conducting a census? What are the differences between various sampling methods? What is
random and Non-random sample?
This chapter will address all the above questions about sampling and the distribution of two
statistics: Sample means and the sample proportion. These are the basic concepts for statistical
analysis.
Sampling is a process used in statistical analysis in which several observations are taken from a
larger population. The predetermined method of sampling can be applied.
Sampling is widely used in business as a means of gathering valuable information about a
population. Data are collected from samples and afterwards, it can be analysed to make
inferences and to produce the results.
96
9.3 RANDOM VERSUS NON-RANDOM SAMPLING
Random sampling and Non random sampling are two main types of sampling. Every unit of the
population has an equal chance or probability of being selected in to the sample is called random
sampling. Random sampling implies that chance enters into the process of selection. For
example, If you would like to have the opinions of Students from the university you will take a
survey of random 100 students out of 10000 population of [Link] method works if there is
an equal chance that any of the subjects in a population will be chosen. Researchers choose
simple random sampling to make generalizations about a population.
In non random sampling not every unit of the population has the same probability of being
selected into the sample. Members of non random samples are not selected by chance. For
example, they might be selected because they are at the right place at the right time or because
they know the people conducting the research. Sometimes random sampling is called probability
sampling and non-random sampling is called nonprobability sampling. Because every unit of the
population is not equally likely to be selected, assigning a probability of occurrence in non-
random sampling is impossible. The statistical methods presented and discussed in this text are
based on the assumption that the data come from random samples. Non random sampling
methods are not appropriate techniques for gathering data to be analyzed by most of the
statistical methods presented.
97
9.4.3 CLUSTER SAMPLING:
Like stratified sampling in cluster sampling, the population is divided into groups based on
similar characteristics or particular characteristics. But, in stratified sampling, one or more
samples are chosen from each stratum while in cluster sampling the clusters are chosen at
random and then takes samples from them. It is often used for market research.
Eg. A study on the impact of natural disasters may divide a population region-wise into clusters,
and then random collection is chosen to begin the study duster's overall impact.
Sometimes the clusters are too large, and the second set of clusters is taken from each original
cluster. This technique is called two-stage sampling.
e.g., if a researcher could divide India into clusters of cities. She could then divide the cities into
clusters of Areas and randomly select individual residences from the area clusters. The ?rst stage
is selecting the test cities and the second stage is selecting the areas. Cluster or area sampling
offers several advantages. Two of the foremost advantages are convenience and cost. Clusters are
usually convenient to obtain, and the cost of sampling from the entire population is reduced
because the scope of the study is reduced to the clusters.
98
9.4.5 NON-RANDOM SAMPLING
A technique of sampling used to select elements from the population by various mechanism that
does not involve a random selection, this process is called non-random sampling . Because
chance is not used for selecting items from the samples these ,techniques are called non-
probability sampling. In these sampling techniques, sampling error is not expected objectively
for such sampling techniques.
99
9.5.3 QUOTA SAMPLING
Quota Sampling is the third sampling technique of Non-random Sampling which appears to be
similar to stratified random sampling. There are subclasses of certain population, such as gender
,age group, or geographic region, are used as strata. However, the researcher uses a nonrandom
sampling method instead of randomly sampling from each stratum, to gather data from one
stratum until the desired quota of samples is filled.
With quota controls ,Quotas are described , to set the sizes of the samples hich can be obtained
from the subgroups. Generally, on the basis of proportions of the subclasses quotas are based in
the population. In this case, the quota concept is similar to that of proportional stratified
sampling. Quotas often are filled by using recent, available and applicable elements.
For example,; if the a subclass has been represented by the respondent whose quota has been
filled, the interviewer may terminate the interview. Mostly the quota sampling is used when no
actual frame if available for the population. Quota sampling can be useful if no frame is available
for the population. Also, for quota sampling, preparatory work is minimal. researcher approach
population directly where the quota can belled. The object is to gain the benefits of stratification
without the high field costs of stratification. Ultimately, it remains a non-probability sampling
method.
E.g . Instead of randomly interviewing people to obtain a quota of African Americans, the
researcher would go to the African's community residential area of the city and interview there
until expected responses are achieved to fill the quota. In such sampling, an interviewer may
begin the interviews by applying few filter questions
100
Both are summary values that describe a group, and there's a handy mnemonic device for
remembering which group each describes. Just focus on their first letters:
¡ Parameter = Population
¡ Statistic = Sample
A population is the entire group of Objects, people, Transactions, animals etc A sample is a
portion of the population.
Estimator:
An estimator is a function of the sample, i.e., it is a rule that tells you how to calculate an estimate
of a parameter from a sample.
An estimator is a statistic that estimates some fact about the population. You can also think of an
estimator as the rule that creates an estimate. For example, the sample mean(x*) is an estimator
for the population mean, ?.
For example, let's say you wanted to know the average height of children in a school with a
population of 1500 students. You take a sample of 50 children, measure them and find that the
mean height is 58 inches. This is your sample mean, the estimator. You use the sample mean to
estimate that the population mean (estimand) is about 58 inches.
An estimate is a Palue of an estimator calculated from a sample. An estimator is a statistic that
estimates some fact about the population. You can also think of an estimator as the rule that
creates an estimate. For example, the sample mean (x*) is an estimator for the population mean,
m. The quantity that is being estimated (i.e. the one you want to know) is called the estimand.
A good estimator is one that gives UNBIASED, EFFICIENT and CONSISTENT estimates. In
this post, I will explain what these terms mean. An estimator is a formula- we input our sample
values and it gives an estimate of the statistic
101
According to the central limit theorem, the mean of a sample of data will be closer to the mean of
the overall population in question, as the sample size increases, notwithstanding the actual
distribution of the data. In other words, the data is accurate whether the distribution is normal or
aberrant.
As a general rule, sample sizes of around 30-50 are assumed sufficient for the Central Limit
Theorem to hold, it means that the distribution of the sample means is fairly normally distributed.
Therefore, the more samples one takes, the more the graphed results take the shape of a normal
distribution. Note, however, that the central limit theorem will still be approximated in many
cases for much smaller sample sizes, such as n=8 or n=5.
The central limit theorem is often used in conjunction with the law of large numbers, which states
that the average of the sample means and standard deviations will come closer to equalling the
population mean and standard deviation as the sample size grows, which is extremely useful in
accurately predicting the characteristics of populations.
102
A larger standard error indicates that the means are more spread out, and thus it is more likely that
your sample mean is an inaccurate representation of the true population mean.
On the other hand, a smaller standard error indicates that the means are closer together, and thus it
is more likely that your sample mean is an accurate representation of the true population mean.
Standard error increases when standard deviation increases. Standard error decreases when
sample size increases because having more data yields less variation in your results.
SE = s / Ö n
s = the population standard deviation
Ön = the square root of the sample size
Example: The values in your sample are 52, 55, 60, and 65.
103
The result is the following pairs of data.
(53,53) (55,53) (59,53) (63,53)
(53,55) (55,55) (59,55) (63,55)
(53,59) (55,59) (59,59) (63,59)
(53,63) (55,63) (59,63) (63,63)
(53,64) (55,64) (59,64) (63,64)
(53,68) (55,68) (59,68) (63,68)
(53,69) (55,69) (59,69) (63,69)
(53,70) (55,70) (59,70) (63,70)
(64,53) (68,53) (69,53) (70,53)
(64,55) (68,55) (69,55) (70,55)
(64,59) (68,59) (69,59) (70,59)
(64,63) (68,63) (69,63) (70,63)
(64,64) (68,64) (69,64) (70,64)
(64,68) (68,68) (69,68) (70,68)
(64,69) (68,69) (69,69) (70,69)
(64,70) (68,70) (69,70) (70,70)
104
9.10 RELATIONSHIP BETWEEN SAMPLING SIZE AND SAMPLING
DISTRIBUTION
The variability of each sampling distribution decreases as the sample size increases. The range of
the sampling distribution is smaller than the range of the original population.
There is an inverse relationship between standard error and Sampling size. The variability of
sampling distribution decreases as the sample size increases.
Also, as the sample size increases the shape of the sampling distribution becomes more similar to
a normal distribution regardless of the shape of the population.
Problem:
A population has a mean of 400 and a standard deviation of 100. A simple random sample of size
200 will be taken and the sample mean will be used to estimate the population mean.
a. What is the expected value of x?
b. What is the standard deviation of x?
c. Show the sampling distribution of x.
d. What does the sampling distribution of x show?
105
9.12 PROPERTIES OF POINT ESTIMATOR
Following are the main properties of Point Estimator:
1. Bias
It is the difference between the expected value of the estimator and the value of the parameter
which is estimated. When the estimated value of the parameter and the value of the parameter
being estimated are equal, the estimator is considered unbiased.
Also, the closer the expected value of a parameter is to the value of the parameter being
measured, the lesser the bias is.
2. Consistency
Consistency tells us how close the point estimator stays to the value of the parameter as it
increases in size. The point estimator requires a large sample size for it to be more consistent and
accurate.
You can also check if a point estimator is consistent by looking at its corresponding expected
value and variance. For the point estimator to be consistent, the expected value should move
toward the true value of the parameter
106
the random number list to select 100 people randomly from your list. How representative
of the population is the sample? Find the proportion of men and women in your population
and in your sample. How do the proportions compare? Find the proportion of 25-50 -year-
olds in your sample and the proportion in the population. How do they compare?
C. Interviewing passengers of an Indian Airline about its food and other services provided to
passengers.
Studying the employee green behaviour of the manufacturing sector MSMEs (Micro,
Small and Medium Enterprises) on the basis of Investment and turnover. How they follow
the organizational green policy.
D. A city's telephone book lists 10,00,000 people. If the telephone book is the frame for a
study, how large would the sample size be if systematic sampling were done on every 200th
person?
107
Short answer type questions
1. What do you mean by quota sampling?
2. What do you mean by purposive sampling?
3. What do you mean by stratified sampling?
4. What do you mean by population and sample?
REFERENCES:
¡ Black, K. (2019). Business statistics: for contemporary decision making. John Wiley &
Sons.
¡ Anderson, D. R., Sweeney, D. J., Williams, T. A., Camm, J. D., & Cochran, J. J. (2012).
Quantitative Methods for Business (Book Only). Cengage Learning.
¡ Everitt, B. S.; Skrondal, A. (2010), The Cambridge Dictionary of Statistics, Cambridge
University Press.
¡ Gonick, L. (1993). The Cartoon Guide to Statistics. HarperPerennial.
¡ Kotz, S.; et al., eds. (2006), Encyclopedia of Statistical Sciences, Wiley.
¡ Levine, D. (2014). Even You Can Learn Statistics and Analytics: An Easy to Understand
Guide to Statistics and Analytics 3rd Edition. Pearson FT Press
108
MODULE - 10 HYPOTHESIS TESTING
STRUCTURE
¡ Research hypothesis
¡ Definition of hypothesis
¡ Types of Hypotheses
¡ Research Hypotheses
¡ Statistical Hypotheses
¡ Differences between null and alternative hypotheses in summary
¡ Substantive hypotheses
¡ Type I and Type II errors
¡ Power of Test
¡ Normal Distribution Hypothesis Test
¡ The t distribution
109
10.2.1 DEFINITION OF HYPOTHESIS
"Any supposition which we make in order to deduce conclusions in accordance with facts which
are known to be real". - Mill
"A hypothesis is an attempt at explanation : a provisional supposition made in order to explain
scientifically some facts or phenomenon." - Coffey
"A proposition which can be put to test to determine validity." - Goode and Hatt
110
workers have to score on the "loyalty" instrument (assuming higher scores indicate more loyal)
than younger workers to prove the research hypothesis? What is the "proof threshold"? Instead of
attempting to prove or disprove research hypotheses directly in this manner, business researchers
convert their research hypotheses to statistical hypotheses and then test the statistical hypotheses
using standard procedures.
A null hypothesis plus an alternate hypothesis make up statistical hypotheses. These two sections
are designed to include every outcome that could possibly result from the experiment or study.
¡ Null Hypothesis
Generally, the null hypothesis asserts that no statistical significance can be found in a collection
of provided observations, that is, the old theory is still true, the old standard is correct, and the
system is in control. Null hypothesis is represented as H0.
We can reject the null hypothesis if the sample contains sufficient data to refute the assertion that
there is no effect in the population (p £ a). If not, we are unable to rule out the null hypothesis.
Although it may sound odd, statisticians only accept the phrase "fail to reject." Avoid using
words like "prove" or "accept" when referring to the null hypothesis. Often, null hypotheses use
words like "no effect," "no difference," "no association" or "no relationship." They are always
expressed mathematically with an equality (usually =, but sometimes ³ or £).
Examples of alternative hypotheses:
Does the amount of text highlighted The amount of text highlighted in the
2 in the textbook affect exam scores? textbook has no effect on exam scores.
Does daily meditation decrease the Daily meditation does not decrease the
3 incidence of depression? incidence of depression.
¡ Alternative hypothesis
The alternative hypothesis, on the other hand, states that there are changes in the old theory and
the new theory is true, there are new standards, the system is out of control, and/or something is
happening. The alternate response to your research question is the alternative hypothesis. It
asserts that the populace is affected. Generally, the alternative hypothesis is represented as Ha.
The complement of the null hypothesis is the alternative hypothesis. The extensive nature of null
and alternative hypotheses ensures that they account for all potential outcomes. Additionally,
they are mutually exclusive, thus only one of them may be true at once.
Phrases like "an effect," "a difference," "an association" or "a relationship" are frequently used in
alternative hypotheses. In mathematics, null hypotheses are always expressed as an inequality
(usually ?, but sometimes < or >). Alternative hypotheses can be expressed in a variety of ways,
just like null hypotheses.
111
Examples of alternative hypotheses:
Sr no Research question Alternative hypothesis (Ha)
1 Does tooth flossing affect the number Tooth flossing has an effect on the
of cavities? number of cavities.
2 Does the amount of text highlighted The amount of text highlighted in the
in a textbook affect exam scores? textbook has an effect on exam scores.
3 Does daily meditation decrease the Daily meditation decreases the incidence
incidence of depression? of depression.
112
10.4 TYPE I AND TYPE II ERRORS
Statistical Principles of Hypothesis Testing
Researcher can use null and alternative hypotheses with hypothesis testing to see whether the
data confirm or disprove the study expectations; however, one can never completely confirm (or
refute) the idea, regardless how many facts one gathers. Drawing conclusions about phenomena
in the population from events observed in the sample will always be necessary (Hulley et al.,
2001).
The null hypothesis, which forms the basis of hypothesis testing, holds that there is no difference
between groups or correlation between variables in the population. It is always accompanied by a
counterargument, which is your research's forecast of a real difference between groups or a real
correlation between variables.
Let us consider: Researcher tests whether a modified Car model can satisfy the demands of
economic class consumers.
In this case:
¡ The null hypothesis (H0): Modified Car model cannot satisfy the demands of economic
class consumers.
¡ The alternative hypothesis (H1): Modified Car model can satisfy the demands of economic
class consumers.
A researcher can use HTAB System to test the Hypotheses, which involves four major tasks
(Figure 1):
Task 1. Establishing the Hypotheses
Task 2. Conducting the Test
Task 3. Taking statistical Action
Task 4. Determining the Business implications
113
Post to applying the HTAB system, the possible statistical outcomes of a study can be divided
into two groups:
1. Those that cause the rejection of the null hypothesis
2. Those that do not cause the rejection of the null hypothesis
The rejection region refers to the conceptual and visual area where statistical results that lead to
the rejection of the null hypothesis are located. The nonrejection region refers to statistical
outcomes that do not lead to the null hypothesis being rejected. (Figure 2) The researcher will
reject the null hypothesis if any sample mean falls within that range. The business researcher will
opt not to reject the null hypothesis if the sample means that fall between the two crucial values
are sufficiently similar to the population mean. These methods fall into the category of
nonrejection. Researcher decides whether the null hypothesis can be rejected based on the data
and the results of a statistical test. Since these decisions are based on probabilities, there is always
a risk of making the wrong conclusion.
If the results are statistically significant, the null hypothesis must be incorrect in order for them to
hold true. Researcher would then reject the null hypothesis in this situation. But occasionally,
this might be a Type I error.
If the results are not statistically significant, the null hypothesis is likely to be correct and they
have a high probability of occurring. As a result, null hypothesis is not rejected. But occasionally,
this might be a Type II error.
114
¡ A Type II error happens when you get false negative results: Researcher conclude that the
modified car model cannot satisfy the demands of economic class consumers, when, in
fact, it did. Your study might have overlooked important signs of progress or mistakenly
attributed any advancements to unrelated sources.
115
and represents the chances of a true positive detection conditional on the actual existence of an
effect to detect. Statistical power ranges from 0 to 1, and as the power of a test increases, the
probability b of making a type II error by wrongly failing to reject the null hypothesis decreases.
The formula for power is 1 - b. A hypothesis test's power ranges from 0 to 1, and if it's close to 1,
it's very effective in identifying a false null hypothesis. Beta is typically set at 0.2 but may be set
smaller by the researchers.
Greater sample size, effect sizes, and significance levels all result in increased power for the
researcher. Although there are other factors, such as variance (s2).
Figure 3: Illustration of the power and the significance level of a statistical test, given the
null hypothesis (sampling distribution 1) and the alternative hypothesis (sampling
distribution 2)
Source: Wikipdia.org_Power of test
116
performed by switching the test statistic. These tests are valuable because they enable us to verify
statements made about usually dispersed things. We consider looking at the mean of a sample
from a population when we hypothesise test for the mean of a normal distribution.
Steps to follow in Normal Distribution Hypothesis Testing:
1. Define the parameter in the context of the question - for a normal hypothesis test the
parameter is µ which is always the mean of something.
2. Write down the null hypothesis and the alternate hypothesis.
3. Define the test statistic X in the context of the question.
4. Write down the distribution of X under the null hypothesis.
5. State the significance level a - even though you are likely given it in the question, not
stating it risks losing a mark.
6. Test for significance or find the critical region.
7. Write a concluding sentence, linking the acceptance or rejection of H0 to the context.
117
test. Gosset made an important contribution since it paved the way for more precise statistical
tests, which some experts claim signalled the start of the contemporary age in mathematical
statistics**.
The formula for the t statistic is:
118
SELF ASSESSMENT QUESTIONS
Long answer questions
1. What do you mean by hypothesis testing
2. Explain different kinds of hypothesis.
3. Define with the help of examples two major research hypothesis
4. A study was made to compare the costs of supporting a family of four Americans for a year
in different foreign cities. The lifestyle of living in the United States on an annual income
of $75,000 was the standard against which living in foreign cities was compared. A
comparable living standard in Toronto and Mexico City was attained for about $64,000.
Suppose an executive wants to determine whether there is any difference in the average
annual cost of supporting her family of four in the manner to which they are accustomed
between Toronto and Mexico City. She uses the following data, randomly gathered from
11 families in each city, and an alpha of .01 to test this difference. She assumes the annual
cost is normally distributed and the population variances are equal. What does the
executive find?
Toronto Mexico City
$69,000 $65,000
64,500 64,000
67,500 66,000
64,500 64,900
66,700 62,000
68,000 60,500
65,000 62,500
69,000 63,000
71,000 64,500
68,500 63,500
67,500 62,400
5. A company's auditor believes the per diem cost in Nashville, Tennessee, rose significantly
between 1999 and 2009. To test this belief, the auditor samples 51 business trips from the
company's records for 1999; the sample average was $190 per day, with a population
standard deviation of $18.50. The auditor selects a second random sample of 47 business
trips from the company's records for 2009; the sample average was $198 per day, with a
population standard deviation of $15.60. If he uses the risk of committing a Type I error of
.01, does the auditor find that the per diem average expense in Nashville has gone up
significantly?
6. Employee suggestions can provide useful and insightful ideas for management. Some
companies solicit and receive employee suggestions more than others, and company
culture influences the use of employee suggestions. Suppose a study is conducted to
determine whether there is a significant difference in a mean number of suggestions a
119
month per employee between the Canon Corporation and the Pioneer Electronic
Corporation. The study shows that the average number of suggestions per month is 5.8 at
Canon and 5.0 at Pioneer. Suppose these figures were obtained from random samples of 36
and 45 employees, respectively. If the population standard deviations of suggestions per
employee are 1.7 and 1.4 for Canon and Pioneer, respectively, is there a significant
difference in the population means? Use = .05.
REFERENCES:
¡ Hulley S. B, Cummings S. R, Browner W. S, Grady D, Hearst N, Newman T. B. 2nd ed.
Philadelphia: Lippincott Williams and Wilkins; 2001. Getting ready to estimate sample
size: Hypothesis and underlying principles In: Designing Clinical Research-An
epidemiologic approach; pp. 51-63
¡ Bullard, F. A. (2009). Exoplanet detection: A comparison of three statistics or how long
should it take to find a small planet? (Doctoral dissertation).
¡ Piaw, C. Y. (2013). Mastering research statistics. Malaysia: McGraw Hill Education, New
York, United States.
¡ Wasserman, L. (2004). All of statistics: a concise course in statistical inference (Vol. 26).
New York: Springer.
¡ Keener, R. W. (2010). Theoretical statistics: Topics for a core course. New York: Springer.
¡ [Link]
120
MODULE - 11 NON-PARAMETRIC TESTS
STRUCTURE
¡ Introduction To Non-parametric Test
¡ Chi-square goodness of - fit test
¡ Chi-square test of independence
LEARNING OUTCOMES
¡ To understand the meaning of non-parametric tests
¡ To understand the use of chi square test as goodness of fit
¡ To understand the use of chi square test as test of independence
Where,
fo = frequency of observed values
fe = frequency of expected values
k = number of categories
121
c = number of parameters being estimated from the sample data
Across the distribution, this formula compares the frequency of observed values to the frequency
of expected values. Since the observed total drawn from the sample is utilised as the total for the
expected frequencies, the test loses one degree of freedom because the total number of expected
frequencies must equal the total number of observed frequencies.
To establish the frequency distribution of predicted values, a population parameter, such as, or, is
occasionally computed from the sample data. This estimation loses a degree of freedom every
time it happens. In general, k - 1 degrees of freedom are employed in the test when a uniform
distribution is used as the expected distribution or when an expected distribution of values is
provided. The degrees of freedom are k - 2 since estimating ? eliminates a third degree of freedom
when determining whether a distribution is Poisson or not. In testing to determine whether an
observed distribution is normal, the degrees of freedom are k - 3 because two additional degrees
of freedom are lost in estimating both µ and ? from the observed sample data.
In 1900, Karl Pearson developed the chi-square test. The chi-square distribution has an infinitely
long positive tail since it is the sum of the squares of k independent random variables, which
means it can never be smaller than zero. The degrees of freedom (df) associated with each
distribution in the family of chi-square distributions determine how they behave. The chi-square
distribution is noticeably tilted to the right for tiny df values (positive values). The chi-square
distribution starts to resemble the normal curve as the df rises.
Let us discuss how can the chi-square goodness-of-fit test be applied to business situations.
Example: In a survey conducted by National Research Agency on organised retail sector
customers were asked: "In general, how would you rate the level of service provided by
organised retail outlets in India?" The distribution of responses to this question was as follows:
Excellent 8%
Pretty good 47%
Only fair 34%
Poor 11%
Assume that Big bazar store manager wants to find out whether the results of this consumer
survey apply to their customers. Hence, store manager conducted 207 interviews of their
customers. The customers were questioned about how they would rank the quality of service at
the supermarket they had just left. The response categories were kept similar which are excellent,
pretty good, only fair, and poor. The observed responses from this study are:
Excellent 21%
Pretty good 109%
Only fair 62%
Poor 15%
The store manager can now apply a chi-square goodness-of-fit test to check whether the
observed response frequencies from this survey match those that would be anticipated based on
the results of the national survey.
Solution:
STEP 1. The hypotheses for this example follows:
Ho: The observed distribution is the same as the expected distribution.
122
Ha: The observed distribution is not the same as the expected distribution
STEP 2. The statistical test being used is:
STEP 6. The chi-square goodness-of-fit can then be calculated as shown in below table
123
STEP 7. As the observed value of the chi-square of 6.25 is not greater than the critical table value
of 7.8147, the store manager will not reject the null hypothesis.
Practical Implications:
STEP 8. The data gathered in the sample of 207 supermarket shoppers indicate that
the distribution of responses from supermarket shoppers in the manager's city is not significantly
different from the distribution of responses to the national survey.
The store manager may conclude that their customers do not appear to have attitudes
different from those people who took the survey.
124
statistical terms, null hypothesis) is that the type of movie and whether or not people purchased
snacks have no relationship. The owner of the movie theatre wants to know how many snacks to
purchase. When movie types and snack purchases are unrelated, estimating is easier than when
movie types influence snack sales.
¡ A list of dog breeds seen as patients at a veterinary clinic. The second variable is whether
owners feed dry food, canned food, or a combination of the two. Our theory is that dog breeds and
food types are unrelated. If this is correct, the clinic can order food based solely on the total
number of dogs, with no regard for breeds.
Let us elaborate the 1st example: Assume we collect information for 600 people in our theatre.
We know what kind of movie each person saw and whether or not they bought snacks.
To begin, consider whether the Chi-square test of independence is an appropriate method for
evaluating the relationship between movie type and snack purchases.
¡ We have a simple random sample of 600 people who saw a movie at our theatre. We meet
this requirement.
¡ Our variables are the movie type and whether or not snacks were purchased. Both variables
are categorical. We meet this requirement.
¡ The last requirement is for more than five expected values for each combination of the two
variables. To confirm this, we need to know the total counts for each type of movie and the
total counts for whether snacks were bought or not. For now, we assume we meet this
requirement and will check it later.
It appears we have indeed selected a valid method. (We still need to check that more than five
values are expected for each combination.)
Here is our data summarized in a contingency table:
Before we proceed, let's double-check the assumption of five expected values in each category.
There are more than five counts in each combination of Movie Type and Snacks in the data. But,
what are the expected counts if movie and snack purchases are made separately?
125
Table 4: Contingency table for movie snacks data with row and column totals
Type of Movie Snacks No Snacks Row totals
Action 50 75 125
Family 90 30 120
Horror 45 10 55
The row and column totals are used to calculate the expected counts for each Movie-Snack
combination. We divide the grand total by the sum of the row and column totals. This calculates
the expected number of cells in the table. For the Action-Snacks cell, for example, we have:
125*310 / 600 = 65
We rounded the answer to the nearest whole number. If there is not a relationship between movie
type and snack purchasing we would expect 65 people to have watched an action film with
snacks. Here are the actual and expected counts for each Movie-Snack combination. In each cell
of Table 5 below, the expected count appears in bold beneath the actual count. The expected
counts are rounded to the nearest whole number.
Table 5: Contingency table for movie snacks data showing actual count vs. expected count
Contingency table for movie snacks data showing actual count vs. expected count
Type of No
Snacks Row totals
Movie Snacks
50 75
Action 125
65 60
125 175
Comedy 300
155 145
90 30
Family 120
62 58
45 10
Horror 55
28 27
These calculated values will be labelled as "expected values," "expected cell counts," or some
other similar term when using software.
126
Because all of the expected counts for our data are greater than five, we can apply the
independence test.
Let's look at the contingency table before calculating the test statistic. The expected counts are
calculated using the row and column totals. Looking at each cell, we can see that some expected
counts are close to the actual counts, but the majority are not. If there is no correlation between
the type of movie and snack purchases, the actual and expected counts will be comparable. If a
relationship exists, the actual and expected counts will differ.
Actual: 50 Actual: 75
Expected: 64.58 Expected: 60.42
Actual: 90 Actual: 30
Expected: 62 Expected 58
Family
Difference: 90 – 62 = 28 Difference: 30 – 58 = -28
Squared Difference: 784 Squared Difference: 784
Divide by Expected: 784/62 = 12.65 Divide by Expected: 784/58 = 13.52
Actual: 45 Actual: 10
Expected 28.42 Expected 26.58
127
Lastly, to get our test statistic, we add the numbers in the final row for each cell:
3.29+3.52+5.81+6.21+12.65+13.52+9.68+10.35=65.033.29+3.52+5.81+6.21+12.65+13.52+
9.68+10.35=65.03
To make our decision, we compare the test statistic to a value from the Chi-square distribution.
This activity involves five steps:
1. We decide how much risk we're willing to take in assuming that the two variables aren't
independent when they are. Prior to collecting the movie data, we decided that we are
willing to take a 5% risk of claiming that the two variables - Movie Type and Snack
Purchase - are not independent when they are. We set the significance level,, to 0.05 in
statistics speak.
2. We calculate a test statistic. As shown above, our test statistic is 65.03.
3. We find the critical value from the Chi-square distribution based on our degrees of freedom
and our significance level. This is the value we expect if the two variables are independent.
4. The degrees of freedom depend on how many rows and how many columns we have. The
degrees of freedom (df) are calculated as:
df=(r-1)×(c-1)df=(r-1)×(c-1)
In the formula, r is the number of rows, and c is the number of columns in our contingency
table. From our example, with Movie Type as the rows and Snack Purchase as the columns,
we have:
df=(4-1)×(2-1)=3×1=3df=(4-1)×(2-1)=3×1=3
The Chi-square value with a = 0.05 and three degrees of freedom is 7.815.
5. We compare the value of our test statistic (65.03) to the Chi-square value. Since 65.03 >
7.815, we reject the idea that movie type and snack purchases are independent.
We conclude that there is a link between movie genre and snack purchases. Regardless of the
type of movie being shown, the owner of the movie theatre cannot estimate how many snacks to
purchase. Instead, when estimating snack purchases, the owner must consider the type of movies
being shown.
It's important to note that we can't infer that the type of movie influences snack purchases. The
independence test only tells us whether or not there is a relationship; it does not tell us which
variable causes which.
128
these data in an effort to determine whether they are Poisson distributed.
Number of Arrivals Observed Frequencies
0 7
1 18
2 25
3 17
4 12
³5 5
4. Use a chi-square goodness-of-fit test to determine whether the observed frequencies are
distributed the same as the expected frequencies (a = .05).
Category fo fe
1 53 68
2 37 42
3 32 33
4 28 22
5 18 10
6 15 8
5. Use the following data and a = .01 to determine whether the observed frequencies
represent a uniform distribution.
Category fo
1 19
2 17
3 14
4 18
5 19
6 21
7 18
8 18
6. According to an extensive survey conducted for Business Marketing by Leo J. Shapiro &
Associates, 66% of all computer companies are going to spend more on marketing this
year than in previous years. Only 33% of other information technology companies and
28% of non-information technology companies are going to spend more. Suppose a
researcher wanted to conduct a survey of her own to test the claim that 28% of all non-
information technology companies are spending more on marketing next year than this
129
year. She randomly selects 270 companies and determines that 62 of the companies do
plan to spend more on marketing next year. Use a = .05, the chi-square goodness-of-fit test,
and the sample data to test to determine whether the 28% figure holds for all non-
information technology companies.
8. Use the following contingency table to test whether variable 1 is independent of variable 2.
Let a = .01
Variable 2
201 325
Variable 1
68 110
9. Is the transportation mode used to ship goods independent of type of industry? Suppose the
following contingency table represents frequency counts of types of transportation used by
the publishing and computer hardware industries. Analyze the data by using the chi-square
test of independence to determine whether the type of industry is independent of
transportation mode. Let a = .05.
Transportation mode
Air Train Truck
Industry Publishing 32 12 41
Computer 5 6 23
hardware
1.8 REFERENCES
¡ Black, K. (2019). Business statistics: for contemporary decision making. John Wiley &
Sons.
¡ Anderson, D. R., Sweeney, D. J., Williams, T. A., Camm, J. D., & Cochran, J. J. (2020).
Modern business statistics with Microsoft Excel. Cengage Learning.
130
MODULE - 12 ANALYSIS OF VARIANCE (ANOVA)
STRUCTURE
¡ Introduction to anova
¡ Types of anova
¡ Single factor Anova
¡ Two factor Anova
131
12.4 TYPES OF ANOVA
One-way or two-way ANOVA refers to the number of independent variables in the analysis of the
variance test.
With a one-way ANOVA, we have one independent variable affecting a dependent variable. It is
used to search for statistically significant differences between two or more independent variables
A two-way ANOVA is an extension of the one-way ANOVA. With a two-way ANOVA, two
independents are affecting a dependent variable. For example, a two-way ANOVA allows
comparing productivity based on two independent variables, such as salary and skill set. It is
utilized to observe the interaction between the two factors and test the effect of two factors at the
same time. Thereby the potential interaction of two independent variables on one dependent
variable is revealed.
A three-way ANOVA, also known as three-factor ANOVA, is a statistical means of determining
the effect of three factors on an outcome.
There is a variation of ANOVA for example, MANOVA (multivariate ANOVA). It differs from
ANOVA as the former tests for multiple dependent variables simultaneously while the latter
assesses only one dependent variable at a time.
ANOVA has many applications in finance, economics, science, medicine, and social science.
ANOVA is used in finance in several different ways, such as to forecast the movements of
security prices by first determining which factors influence stock fluctuations. This analysis can
provide valuable insight into the behavior of a security or market index under various conditions.
A researcher might, for example, test students from multiple colleges to see if students from one
of the colleges consistently outperform students from the other colleges. In a business
application, an R&D researcher might test two different processes of creating a product to see if
one process is better than the other in terms of cost efficiency. In medical sciences, to compare the
effects of different treatment protocols on patient outcomes; in social science research (for
instance to assess the effects of gender and class on specified variables), in software engineering
(for instance to evaluate database management systems), in manufacturing (to assess product
and process quality metrics), and industrial design among other fields.
With ANOVA, a researcher can determine whether the variability of the outcomes is due to
chance or the factors in the analysis. ANOVA analysis is considered to be accurate than t testing
because it is flexible and requires fewer observations. It is better suited for use in complex
analyses than those that can be assessed by conducting tests. ANOVA testing allows researchers
to uncover relationship.
132
Table 1: ANOVA single factor table
We are interested to know whether a difference exists somewhere between the three different
year levels. Now we will analyze based on the one-way ANOVA technique (also known as Single
factor Anova). First, we will undertake a hand calculation. This will enable us to realize from
where the numbers originate. Then we will compare the hand-calculated numbers with that of the
inbuilt excel output
Figure 12.1
But if we remove that last step of finding the part of the average, then we are left with just the sum
of the squares. We will take the distance of each data point from the mean, square each distance,
133
and then add them together. If we stop there, that is the sum of squares (refer to Figure 1). The
'sum of squares' is a foundational component of ANOVA and Regression.
The overall sum of squares ie the Sum of Squares Total (SST) is partitioned into two components.
The first component is SSC (sum of the squares of column) and the second one is the SSE (sum of
the squares of the errors). The sum of squares of the columns (SSC), is between the columns, and
the sum of squares of the error (SSE), is within each column. This is actually about the individual
distribution around each column's mean.
Thus, SST = SSC + SSE (refer to Figure 2)
Kindly make a note:
N = Total number of observations, in this case, it is 21
C = Columns = 3
Figure 12.2
Table 2
134
01
which is 74.52. We would then square this difference and sum it up. It totals 2901.24
2. If we consider a normal distribution, the mean will be in the centre, and some of the data
points will be to the right, which is higher than the mean, and those which are lower than
the mean will be to the left.
3. The degree of freedom for SST is dftotal = N-1 = 21-1=20. (refer to Table 2)
135
Table 3
You will observe that the hand calculation output match with that of excel's ANOVA output. Thus
now you are in a position to relate to the excel's inbuilt ANOVA output.
1. The squared sum between columns (SSC) is 88.66666, which matches the hand calculation
viz 88.67.
2. The squared sum of totals (SST) is 2901.238095; this, too matches the hand calculation viz
2901.24.
3. The squared sum of errors (SSE) within columns is 2812.5714. This, too, matches the hand
calculation viz 2812.57
4. The degree of freedom, the MSC, MSE, Fstatistic, and the Fcritic value also match
perfectly.
The p-value is 0.756278 which is more than alpha 0.05 and is thus not significant.
The null hypothesis is that the means are equal to each other. Thus, we fail to reject our null
hypothesis that these three means are different. The means of the first-year students, the second-
year students, and the third-year students on their study skills exam did not differ significantly.
136
power of linear regression is based on ANOVA.
Now let us consider one more fictitious problem to understand the practical implication of 'Two-
way Anova without replication'. Like before, we will once more undertake a hand calculation
(excel) so that you know and understand the source of the numbers ie from where the numbers
originate and their importance. We will follow it up with excel's inbuilt ANOVA output to
compare and reconfirm whether the hand-calculated numbers match the excel's inbuilt ANOVA
output.
Let's assume that the 'LifeStyle Coffee' chain uses secret shoppers who appear as customers to
enter their store and document their experience in terms of customer service, cleanliness, and the
quality of their product - coffee. They give a score out of a maximum of 100. The secret shoppers
receive standardized training by 'LifeStyle Coffee' to ensure consistency and objectivity in their
store reviews. For its locations in the cities of Nagpur, Mumbai, and Pune, 'LifeStyle Coffee'
trained six secret shoppers. Each of the six secret shoppers is assigned to visit the store once in
each of the three cities. As the visit sequence will be assigned randomly, hence the name
'randomized block design'. It is called 'without replication' because each shopper is only going to
each city once (In ANOVA with replication each shopper will visit the city more than once. There
will be multiple measurements. Thus, more data will be generated). The two factors here are the
city ( columns) and the shopper (rows).
We would like to know if a difference exists in secret shopper ratings among the cities…
Are all the cities about the same in their ratings? Is one significantly higher than the other two? Or
are all three different from each other?
This is testing for differences among the cities. It's not testing whether they are good or bad,
which is a subjective experience or rating. The heart of the problem as to what makes this
problem fit for Two-way Anova is that the secret shoppers (ie rows) themselves will have their
natural variation to review their experiences. Two-way ANOVA allows accounting for the
shopper variation to determine if a difference exists among the cities without the shopper
variation clouding or masking any of the city differences. It is like untangling all the sources of
variation before we can get down to looking at any differences that might exist between the cities.
Table 4
137
To calculate the mean, the '=Average' excel function in the cell is used
1. The mean for each column is calculated
2. The mean for each row is calculated
3. The overall mean is calculated
In a one-way ANOVA:
SST (sum of squares total) = SSC (sum of squares column (or treatment or groups) + SSE (sum of
squares within or error)
Whereas, in a two-way ANOVA:
We are interested in the differences between the cities (ie the columns). By introducing the
blocking variable, we are further trying to reduce the original SSE into SSB and the remaining
will be again the new SSE, ie further splitting the SSE and attributing it to SSB and that which
cannot be attributed will be SSE. Thereby now the SSE is smaller which is the unexplained
source of error; the unexplained variance will always be there. In the end, SSC will be compared
to SSE, and SSC claims a larger part of the total variance.
SST (sum of squares total) = SSC (sum of squares column) + SSB (sum of squares block) + SSE
(sum of squares error or within).
The extent SSB accounts for SSE can be known from the ratio 'SSB divided by SSE'. In the real
sense by dividing 'MSB by MSE'.
Table 12.5
Table 12. 6
139
Table 7
140
Table 8
You will observe that the excel output and the hand calculation match.
1. The column means for the city's matches.
2. The squared sum for the rows, columns, and the error matches
3. The degree of freedom, the MSC, MSB, and MSE also match.
4. The F statistic for rows and columns and the F critical value also match.
Thus, our hand calculation is right and it has provided a sense and an understanding to know the
source of numbers.
The P-value for the columns (in cities as per the example) is 0.024. It is less than 0.05 so it is
significant whereas the P-value for the rows (in shoppers as per the example) is 0.057 which is
barely equal to or pretty close to 0.05. Thus, it is not significant. This indicates that the column
scores do have legitimate differences even after accounting for variation in shoppers' scores.
Now if you check the F table or calculate in a cell in excel as per the equation = [Link](0.05,
2,10), then the value that you will obtain is = 4.10, is the Fcritic value.
The null hypothesis is that there are no significant differences in the city (remember that it always
assumes that there is no difference).
We reject the null hypothesis as the F statistic for columns (cities) is 5.52 is larger than the
Fcritic. A significant difference in the mean quality score is present in the columns (cities).
141
To understand this better let us take an example of a plant food company. They are trying to find
the effectiveness of three different plant foods named AA, BB, and CC with one feeding per day
on eight respective plants. That is eight plants will be given the food AA, another eight plants will
be given BB and the third set of eight plants will be given the food CC. The height of the plant will
be tested before the food is given and after 75 days. The same type of seeds will be used for the
entire experiment to control for the type of seed.
Now let's add one more factor. That is increase the feeding frequency from once a day to twice a
day with another set of eight similar plants each for the same food ie AA, BB, and CC. This is a
balanced design because two factors and two, multiple, equal numbers of measurements are
present for each factor combination.
Now we have six sets of eight plants in this experiment. We have a two-factor or two-way
ANOVA with replication. It's with replication because each food and feeding frequency has eight
plants.
So now to understand how different it is when we do replication here, take note that each of the
eight plants with respective food AA, BB, and CC for one feeding has its mean & variation, and
similarly for two feedings it has its mean & variation. This is shown in the figure there are 6
means and each has its variation from its mean. This is the fundamental concept of two-way
Anova with a variation.
Table 9
Browse through the table and keep the following at the back of your mind:
1. The two row means 60 and 58.1 are pretty close to each other
2. The two-row means of 60 and 58.1 are quite close to that of the overall mean of 59.1
3. The two column means 63,2 and 64.6 are a bit above the overall mean of 59.1
142
4. The column means 49.6 for CC is way off from the overall mean of 59.1
With the help of the excel tool, I have taken the two-column means (one feeding mean & two
feeding mean) and plotted a graph [ path: excel - charts - insert line chart]. (refer to Graph1)
This graph is known as the interaction graph or the graph of marginal means. I have titled it as
'Marginal means of plant height (ignore the word - Estimated in the graph)'.
It allows us to visualize the characteristics of each factor and any interaction that may be
occurring between them. So, in a marginal means graph, as a general rule, we look to see if the
lines cross or would cross because that expresses that the factors change, and their values change
across the groups.
The factor of interest or the factor with the most levels, ie the three types of plant food is plotted
on the X-axis and the dependent variable, what is being measured is plotted on the Y ie the plant
growth.
When the plant food AA and BB is fed twice (red line) the plant grows whereas with the food CC
the growth goes down. Thus, two feedings do not produce consistent growth across all the plant
food types. This type of situation is called an 'interaction'.
An interaction occurs when the effect of one factor changes for different levels of the other factor.
In this case, the most effective feeding frequency changes across plant food types. For AA and
BB two feedings are the best but when we get to CC, it changes, one feeding is the best. That's
what we mean by an interaction. If the lines cross on a marginal means graph, it means that there
is an interaction. Non-parallel lines can also mean there is a significant interaction but the
crossing of lines is more indicative than non-parallel lines.
Graph 1
While interpreting a two-way Anova one should always look for a significant interaction first. If
the exchange is significant, then we need not interpret it further. Because it means that the two
individual factors are too intertwined and tied together to look at them individually, always go for
the interaction term first when you're looking at your F-ratio and p-value or significance.
In this example, we will avoid hand calculation. This is because by now you may have
understood as to how to do it to know how the numbers arise.
143
12.8 EXCELS BUILT-IN ANOVA - TWO-WAY ANOVA WITH REPLICATION
To obtain it the following path has been followed (path: excel - data - data - data analysis -
window - Anova two factor with replication - input the range with labels - rows per sample should
be 8 here - alpha to be maintained at 0.05 - tick the new worksheet or identify/output range where
it should appear). The output will be similar to Table 10 below. It's only that I have colored and
rearranged for readability. Let's try to understand it.
Table 10
Under source of variation, there is an item 'sample' it is the 'feeding frequency', i.e.s the rows, The
item 'columns' represent the 'plant food', and 'interaction' represent the interaction between
'feedings*plantfood'.
Now the p-value for the sample ie feeding frequency, is 0.45. It is not significant as it is more than
0.05. Secondl,y also have a look at the rows means of 60 and 58.1 for feed 1 and feed 2
respectively. They are pretty close to each other and more relative to the overall mean. Now if
you visualize and try to plot these two numbers on the graph above, you will notice that they are
close to each other and have an overall mean of 59.1. There is hardly any variability. This gives us
an idea that they are not significant.
Now have a look at the p-value of plant food ie columns. It is 0.0000 (this has been obtained after
decreasing decimals). It is significant because it is less than 0.05. Now also look at the column
means for the three plant foods it is 63.5, 64.75, and 49. There's a difference. The mean 49 is quite
a way off! If we try to plot these three numbers on the graph above, you will realize that the three
two plots are quite a way off from the overall mean. There is a lot of variability in the column
means that in the type of food So there's a significant difference
Now we will turn our attention to the item - 'interaction'. It is the interaction between
'feeding*plantfood'. The p-value is 0.0000 (this has been obtained after decreasing decimals). It
144
is lower than 0.05. Secondly, we have seen that two feedings are better for plant food AA and BB,
and in the case of CC one feeding is better. Moreover, the lines in the graph cross each other. The
crossed-row lines indicate, usually, an interaction.
Thus, the column means spread far apart from the overall mean, indicate a significant column
factor and the row means spread far apart from the overall mean usually indicate a significant
row factor.
So, just because each effect is significant, does not mean there's a significant interaction. But if an
interaction exists, and it is significant then the row and column effects cannot be evaluated
individually. The main effects are too intertwined and are too confounded together to look at
individually because the values change across. And it cannot be untangled.
Table 3
If ANOVA is statistically significant, then we're going to follow up with a post hoc test to
determine where those differences reside.
We wouldn't do a post hoc test if our ANOVA were non-significant. We would only do the post
hoc follow-up if the initial ANOVA told us that there are differences there somewhere, and now
we have to find them.
The posthoc is only necessary when you reject the null hypothesis when you say there is a
statistically significant difference between groups and when there are three or more groups.
If there are only two groups, well then the solution is easy you just look at the means whichever
group has the higher mean that's statistically significantly different from the other group.
But in the case of ANOVA, we could have three means is the first one different than the second or
different from the third is the third just different from the first we have to follow up in a way to
determine where those differences lie.
There are multiple ways in which we could conduct a post hoc test the simplest one is called a
Bonferroni correction. This is where you take your alpha level typically 0.05 and just divide it by
145
the number of tests. If we were running 5 tests, we would divide 0.05 so each test would have a
significance level of 0.01 to be considered statistically significant. This is the simplest method
and the most conservative method but not necessarily the best method in that it can increase the
chance for type two errors.
There are other ways of approaching post hoc testing that can give us a nice balance between not
inflating the type two error rate and also making sure that we're only finding differences where
they truly exist.
There are 18 different types of post hoc tests as per the SPSS software. Which one to choose
depends upon the nature of the data.
If all the assumptions are true and in case of equal sample size there's the Tukey HSD (Honestly
Significant Difference) but in case of unequal sample sizes there are Gabriel's tests and in the
case of smaller sample size Hochberg GT 2 and for the larger sample size unequal variances there
are the Gains Howl and there's the Fisher LSD test, Scheffe's test and more.
Let us take an example and apply the Tukey HSD posthoc test
This test is very commonly used by statisticians. Let's take an example and use this test for the
example we already had from Excel's built-in Anova: single Factor test (or a One-way Anova).
The formula to be used for Tukey's Posthoc analysis test is as follows:
wherein…
q = is the constant to be obtained from the Studentized Range q table (based on 'dfw' and 'k'; the
number of treatment groups)
MSw is the mean square within
nk is the number in each category.
Now let's use this above formula for the excels output above (refer to Table 3):
'dfw'= 18 and 'k' the number of treatment groups = 3. Now using the Studentized Range q table,
the corresponding number across this intersection is 3.609 (refer to Figure 3)
MSw = 156.2539 and nk = 7 ( refer to Table 3)
Now inserting these numbers in their appropriate positions, the formula looks like this :
HSD = 3.609 * square root of 156.25 divided by 7
HSD = 3.609 * 4.72
HSD = 17.03448
If the means differ by more than this HSD value which is 17.03448 then they are statistically
significantly different.
The three means are 71.714, 75.285, and 76.571. None of them differ by more than 17.03448.
The means are not statistically different.
As such in this example before we rejected the null hypothesis because the means did not differ
significantly. The null hypothesis was that the means are equal to each other. That is what is
confirmed by Tukey's Posthoc analysis. Infact, the post hoc analysis should only be done if there
146
is a statistical significance! This was just an example for you. You can undertake a similar
exercise for Table 8 and Table 10.
SUMMARY
¡ Analysis of variance, or ANOVA, is a statistical method that separates observed variance
data into different components to use for additional tests.
¡ A one-way ANOVA is used for three or more groups of data, to gain information about the
relationship between the dependent and independent variables.
¡ If no true variance exists between the groups, the ANOVA's F-ratio should equal close to 1.
¡ Analysis of variances (ANOVA) is a statistical method that analyzes the influence of one or
more independent variables on a dependent variable of interest.
¡ ANOVA is used in various applications, including in finance and financial markets to find
and confirm correlations and associations between various factors.
¡ There are a variety of ANOVA techniques, including one-way, two-way, and factor models
¡ A two-way ANOVA is an extension of the one-way ANOVA (analysis of variances) that
reveals the results of two independent variables on a dependent variable.
¡ A two-way ANOVA test is a statistical technique that analyzes the effect of the independent
variables on the expected outcome and their relationship to the outcome itself.
¡ ANOVA has many applications in finance, economics, science, medicine, and social
science.
¡ There are several post hoc analysis methods to be deployed after ANOVA. It should be
undertaken only if statistical significance exists.
147
SELF ASSESSMENT QUESTION
1. Suppose an ANOVA has been performed on a completely randomized design containing
six treatment levels. The mean for group 3 is 15.85, and the sample size for group 3 is eight.
The mean for group 6 is 17.21, and the sample size for group 6 is seven. MSE is .3352. The
total number of observations is 46. Compute the significant difference for the means of
these two groups by using the Tukey-Kramer procedure. Let ? = 0.05
2. A completely randomized design has been analyzed by using a one-way ANOVA. There
are four treatment groups in the design, and each sample size is six. MSE is equal to 2.389.
Using compute Tukey's HSD for this ANOVA.
3. Using the results of problem 11.5, compute a critical value by using the Tukey-Kramer
procedure for groups 1 and 2. Use Determine whether there is a significant difference
between these two groups.
4. In recent years, the debate over the U.S. economy has been constant. The electorate seems
somewhat divided as to whether the economy is in recovery or not. Suppose a survey was
undertaken to ascertain whether the perception of economic recovery differs according to
political affiliation. People were selected for the survey from the Democratic Party, the
Republican Party, and those classifying themselves as independents. A 25-point scale was
developed in which respondents gave a score of 25 if they felt the economy was in
complete recovery, a 0 if the economy was not in a recovery and some value in between for
more uncertain responses. To control for differences in socioeconomic class, a blocking
variable was maintained using five different socioeconomic categories. The data are given
here in the form of a randomized block design. Use to determine whether there is a
significant difference in mean responses according to political affiliation.
Political Affiliation
Socioeconomic Class Democrat Republican Independent
Upper 11 5 8
Upper middle 15 9 8
Middle 19 14 15
Lower middle 16 12 10
Lower 9 8 7
5. A randomized block design has a treatment variable with six levels and a blocking variable
with 10 blocks. Using this information and complete the following table and conclude the
null hypothesis.
Source of Variance SS df MS F
Treatment 2,477.53
Blocks 3,180.48
Error 11,661.38
Total
148
REFERENCES:
1. [Link]
2. Black Ken; Business Statistics for contemporary decision making; 6th ed.
3. h t t p s : / / w w w. y o u t u b e . c o m / w a t c h ? v = Z k j P 5 R J L Q F 4 & l i s t = P L I e G t x p v y G -
LoKUpV0fSY8BGKIMIdmfCi&index=1, Foltz Brandon Statistics 101 Linear
Regression
4. [Link]
149
MODULE - 13 SIMPLE LINEAR REGRESSION
STRUCTURE
¡ What Is a Regression?
¡ Why is it called Regression?
¡ What is the purpose of Regression
¡ Understanding Regression
¡ What is a Variable?
¡ What is a Covariance?
¡ What is a correlation coefficient 'r'?
¡ What are the Correlation caveats?
¡ Coefficient of determination (R-squared)
¡ What is a Linear Relationship?
¡ Understanding regression equation
¡ Calculating regressions: manual and excel
¡ Regression statistics table - Excel output
¡ Summary
13.2 INTRODUCTION
If you wish to know how two or more pieces of data relate to each other, for example how the total
food bill in a restaurant impacts tips (two pieces of data) or how tips are impacted by restaurant
ambiance and the price on the menu card (three pieces of data), or if you wish to create a forecast
or analyze predictions based on the relationships between the pieces of data, then the regression
is helpful. This is a tool commonly used for forecasting and statistical analysis.
This course familiarises you briefly with the underlying concepts, principles, and mechanics
related to Regression.
150
13.2.1 WHAT IS A REGRESSION?
Regression is a method that determines the strength and character of the relationship between
one dependent variable (usually denoted by Y) and a series of other variables known as
independent variables (usually denoted by X).
Regression analysis is the process of constructing a mathematical model or function that can be
used to predict or determine one variable by another variable or other variables. The most
elementary regression model is called simple regression or bivariate regression involving two
variables in which one variable is predicted by another variable. In simple regression, the
variable to be predicted is called the dependent variable and is designated as Y. The predictor is
called the independent variable, or explanatory variable, and is designated as X.
Regression is also known as 'Ordinary least squares' (OLS), or 'least squares regression' or
simple linear regression. Linear regression establishes the linear relationship between two
variables. Linear regression is graphically depicted using a straight line with the slope defining
how the change in one variable impacts a change in the other. A brief on non-linear regression is
provided at the end.
Let's take an example, to study the relationship, between the number of workers, x, and the tables,
y, produced by them. In Table 1 Given below is the sample of 10 with a duration of one hour
each. The standard deviation calculated from the table given below is Sx = 6.48 and Sy = 16.69.
Given below is the methodology to obtain covariance as per the formula in Figure 1.
Covariance is 962.4/ n-1 = 962.4 /9 = 106.93.
The sign is positive, i.e., + 106.93. The graph depicts linearity.
The scatter plot in Figure 2 represents that there exists a positive linear relationship between the
number of workers, x, and the tables, y, produced by them.
152
Table 1: Covariance calculation
153
The correlation calculation:
In the above example of 'workers and tables produced' Sx is 6.48 and Sy is 16.69 and covariance
is 106.93.
Correlation =
106.93 divided by the product of (6.48 x 16.69)
= 0.989
=r
The rule of thumb to know whether a relationship exists between the two variables is to
determine whether the correlation | r | ³ 2; divided by the square root of n (where n is the sample
size).
= | r | ³ 2 divided by the square root of 10
= | r | should be³ 0.632 then the relationship exists.
In our calculations,
| r | is 0.989 which is greater than 0.632, thus the relationship exists.
154
Figure 3: Calculation of slope
In this equation, "X" and "Y" are two variables that are related by the parameters "m" and "b".
Graphically, Y = mX + b plots in the X-Y plane as a line with slope "m" and Y-intercept "b." The
Y-intercept "b" is simply the value of "Y" when X=0.
The slope "m" is calculated from any two individual points (X1, Y1) and (X2, Y2) as shown in
Figure 3
155
b = The y-intercept or "b" is the value of Y (dependent variable) if the value of X (independent
variable) is zero, and so is sometimes simply referred to as the 'constant.'
m = is the slope of the regression line
u = The regression residual or error term
How do you interpret a regression model?
A simple regression model output may be in the form of:
Y = 3 + (6.4) X + 0.56
We would interpret the model as the value of Y changes by 6.4 X for every one unit change in X. If
X goes up by 3, Y goes up by 19.2. The Y intercept is 3 when X is zero. The regression residual or
error term is 0.56.
A multiple regression model output may be in the form of:
Y = 1.0 + (3.2) X1 - 2.0(X2) + 0.21.
Here we have a multiple linear regression that relates a dependent variable Y with two
independent variables X1 and X2.
Multiple linear regression (MLR), also known simply as multiple regression, is a statistical
technique that uses several independent variables to predict the outcome of a dependent variable.
It extends to several independent variables. Whereas Simple linear regression is a function that
allows making predictions about one variable based on the information that is known about
another variable.
We would interpret the model as the value of Y changes by 3.2X1 for every one unit change in X1
(if X1 goes up by 2, Y goes up by 6.4, etc.) holding all else constant. That means controlling for
X2, X1 has this observed relationship. Likewise, holding X1 constant, every one unit increase in
X2 is associated with a 2X decrease in Y. Note the negative sign here.
We can also note the Y-intercept of 1.0, meaning that Y = 1 when X1 and X2 are both zero. The
regression residual or error term is 0.21.
What are the assumptions that must hold for regression models?
To properly interpret the output of a regression model, the following main assumptions about the
underlying data process of what you analyzing must hold:
1. The relationship between variables is linear
2. That the variance of the variables and error term must remain constant (Homoskedasticity)
3. All independent variables are independent of one another
4. All variables are normally-distributed
156
Table 2: calculating the mathematical regression line
The tip amount increases as the total bill increases. There seems to be linearity and positivity.
The mean of the total bill is 74 whereas that of the tip amount is 10. Now it is important to
remember that these two numbers (74, 10) are known as centroids. The regression line passes
through the centroid.
157
1. Columns 3, 4, 5, and 6 have been created to calculate the slope 'm' of the linear regression
equation: Y = mX + b.
2. Column 3 denotes the deviation of each amount of the total bill from the mean (74) whereas
column 4 denotes the deviation of each amount of the tip from the mean (10).
3. Column 5 is the product of columns 3 & 4 and column 6 is the square of column 3.
4. The summation (ie total) of column 5 is 615 and that of column 6 is 4206.
5. The slope 'm 'as per the formulae is 615/4206 = 0.146219686. Thus, the linear regression
equation: Y = mX + b will read as Y = 0.146219686 X + b.
6. To calculate 'b' substitute the centroids (74, 10) as X and Y in the linear regression equation.
The calculation yields the value of 'b' as - 0.819. 'b' here is the Y-intercept (or considered as
the constant). Now we have calculated the value of slope 'm' and that of 'b'. So the linear
regression equation (or the regression line) Y = mX + b reads thus: Y = 0.146219686 X -
0.819.
7. The equation can be interpreted as follows…for every increase in the bill amount by Re 1/-,
the tip is expected (or forecasted) to increase by 0.146219686 paise. And if the bill amount
is zero, the tip amount will be - 0.819 paise i.e. negative 0.819 paise (Well, at times the y-
intercept may not make any sense in the real world).
8. Now we will use the linear regression equation (or the regression line) Y = 0.146219686 X
- 0.819 to calculate the predicted (or forecasted) tip amount based on the total bill. Kindly
remember here that we already have the observed or the actual tip amount. But now we are
calculating the predicted (or the forecast amount). It has been done in the table pasted
below:
158
6. In column 7, we square each error and sum it up. The summation obtained here is
30.07489464. This is known as the Sum of Squared Error i.e., SSE (some refer to it as the
Residual Sum of Squares {RSS}). The lower the SSE, the better it is because it means that
it fits the data, quite well.
7. The difference between the two summations ie SST and SSE is the Sum of Squared
Regression (ie SSR).
8. So, Sum of the Squared Total (SST) = Sum of Squared Regression (SSR) + Sum of Squared
Error (SSE), ie put simply it is SST= SSR + SSE. Thus, inputting SST and SSE we get SSR
which is 89.92510536.
9. With the help of the linear regression equation (or the regression line) Y= mX + b ie Y =
0.146219686 X - 0.819; we have been able to reduce the SST (some may also refer to it as
the previous error) from 120 to 30.07489464. That is the error has been reduced by
89.92510536. In other words, it means that the error has been regressed or reduced! Thus,
the use of the term 'regression'
10. The coefficient of determination ie r2 can be calculated here to understand the 'fit'. r2 =
SSR divided by SST = 89.92510536 divided by 120 = 0.7493 = 74.93%. It means that
74.93 % of the error (which in this case is 89.92510536) can be explained by using the
estimated regression equation (which is Y = 0.146219686 X - 0.819.) to predict the tip
amount. The balance of 25.07% (which is 30.07489464) that is SSE remains unexplained!
To create the scatter diagram (refer to Figure 4) with the help of Excel, you may follow the path
described below (also depends on the excel version that you may use):
Open excel - click insert - Scatter - pick up the first plot - you will get a window on the excel sheet
- right click on the window - add data - select the X values and the Y values which have been
plotted on the excel sheet - okay. Now you will get the (x, y) plots on the chart.
To obtain the Regression line: click chart design - add chart element- trendline. And you will
obtain the trend line.
Follow the same procedure - More trend line option - tick the display equation on the chart and
tick display the R squared value on the chart. Now the Figure that you obtain will resemble
Figure 4:
Figure 4: The scatter plot, the regression line, the regression equation, and R2
159
13.11 REGRESSION STATISTICS TABLE - EXCEL OUTPUT
To obtain the Regression Statistics table (Refer to Table 4), you may follow the path: - Data - Data
Analysis - Regression - Input the Y range and the X range with the labels - labels - confidence
interval at 95%. The following table will be displayed.
You can now compare Table 4 obtained from Excel with the manual calculation.
It has been labelled for your easy understanding.
The following have been labelled: Coefficient of determination ie R2, the SSR, the SSE, and the
SST, the value of the y-intercept ie b, the slope, 95% confidence interval between 0.02883093 to
0.26361, and the Mean Square Error (MSE) - s2.
The confidence interval can be interpreted as "I am 95% confident that the interval (0.02883093
to 0.26361) contains the true slope of the regression line. Because the interval does not contain
zero so I reject the null hypothesis that the slope is zero.
Mean Square Error (MSE), s2 is an estimate of s2 the variance of the error. In other words, it
means, how spread out the data points are from the regression line. MSE is SSE divided by the
product of (sample numbers minus 2) = 30.074 divided by (6-2 = 4) = 7.5187.
The standard error of the estimate or sigma, or just standard error - s, is the standard deviation of
the error term. So, all of the errors we are dealing with are standard deviation. It's just the average
distance an observation falls from the regression line in units of the dependent variable. So, since
MSE (ie s2) is squared, the standard error 's' is just the square root of it. 's' is the square root of
MSE.
s = Square root of MSE
i.e., square root of 7.51872325249643 which is 2.74202903932406.
You will find this number in Table 4 at the top LHS.
160
Nonlinear Regression
Nonlinear regression is a form of regression analysis in which data is fit to a model and then
expressed as a mathematical function. Simple linear regression relates two variables (X and Y)
with a straight line (Y = mX + b), while nonlinear regression relates the two variables in a
nonlinear (curved) relationship. Nonlinear regression modelling is similar to linear regression
modelling in that both seek to track a particular response from a set of variables graphically.
SUMMARY
1. A linear relationship (or linear association) is a term used to describe a straight-line
relationship between two variables. Linear relationships can be expressed either in a
graphical format or as a mathematical equation of the form Y = mX + b.
2. A regression is a technique that relates a dependent variable to one or more independent
variables. Simple Regression or Multiple Regression
3. A regression model can show whether changes observed in the dependent variable are
associated with changes in one or more of the independent variables. It does this by
essentially fitting a best-fit line and seeing how the data is dispersed around this line.
4. The sum of squares measures the deviation of data points away from the mean value. A
higher sum of squares indicates higher variability while a lower result indicates low
variability from the mean. There are three types of sum of squares: SST, SSE, and SSR.
5. A line of best fit is a straight line that minimizes the distance between it and some of the
data points
6. The least squares method is a procedure to find the best fit for a set of data points by
minimizing the sum of the offsets or residuals (ie error) of points from the line.
7. The least squares method provides the overall rationale for the placement of the line of best
fit among the data points being studied.
161
119 16 4 7
93 18 31 16
108 27 12 10
117 31 3 8
2. Use the following data to determine the equation of the multiple regression model.
Comment on the regression coefficients.
Predictor Coefficient
Constant 31,409.5
x1 .08425
x2 289.62
x3 -.0947
3. Develop a multiple regression model to predict y from x1, x2, and x3 using the following
data. Discuss the values of F and t.
Y x1 x2 x3
5.3 44 11 401
3.6 24 40 219
5.1 46 13 394
4.9 38 18 362
7.0 61 3 453
6.4 58 5 468
5.2 47 14 386
4.6 36 24 357
2.9 19 52 206
4.0 31 29 301
3.8 24 37 243
3.8 27 36 228
4.8 36 21 342
5.4 50 11 421
5.8 55 9 445
162
REFERENCES:
1. [Link]
2. Black Ken; Business Statistics for contemporary decision making; 6th ed.
3. h t t p s : / / w w w. y o u t u b e . c o m / w a t c h ? v = Z k j P 5 R J L Q F 4 & l i s t = P L I e G t x p v y G -
LoKUpV0fSY8BGKIMIdmfCi&index=1, Foltz Brandon Statistics 101 Linear
Regression
163
MODULE - 14 SPSS AND DATA ANALYSIS
Structure
¡ Introduction
¡ Benefits of spss
¡ Spss for data analysis
¡ Variables
¡ Variable types
¡ Conclusion
LEARNING OUTCOMES
This chapter will enable the reader to understand
1. Meaning and relevance of SPSS
2. History of SPSS
3. Fundamentals of SPSS
4. How to effectively use it for research and data anlysis.
14.1 INTRODUCTION
"The world's leading statistical software for business, government, research and academic
organizations" - IBM SPSS
SPSS refers to "Statistical Package for the Social Sciences". It is a comprehensive, predictive
analytical computerized software package which enables ease to use set of data to users ranging
from business to statistical programmers and academicians to researchers. SPSS is used for
logical batched and non-batched statistical analytical tool.
Initially SPSS was evolved by SPSS Inc. which has a mission to 'drive the widespread use of data
in decision-making' derives directly from these two themes." The two themes are- firstly, to make
difficult analytical tasks easier for every user via improvement in ability to use and data access
and enable more number of people to attain benefits from the usage of quantitative techniques for
decision making. Secondly, to focus on analyzing data regarding people their viewpoints,
attitudes and behavior.
Later, IBM acquired SPSS Inc. in 2009. Due to this acquisition, the current version of SPSS is
named as IBM SPSS Statistics. Now it has been diversified into IBM SPSS Data collection used
for survey authoring and deployment, collecting and summarizing of data, and IBM SPSS
Modeler used for data mining, text analytics and collaboration and deployment in batch and
automated scoring services.
164
14.2 BENEFITS OF SPSS
As per the details given on the website of IBM SPSS (i.e. //[Link]/analytics/data-
science/predictive-analytics/spss-statistical-software), the company focus on making
researchers more confident about their results at every level of the process of analysis. SPSS
provides versatile method and an ease to analyze or predict the data. The SPSS product family
increases the pace and simplifies the complete analysis starting from data access and preparation
to final analysis, disposition of results and ultimately reportage. The statistical capabilities of
SPSS range from simple percentages to complex examines of variance, general linear models
and multiple regression. One can use data ranging from simple integers or binary variables to
multiple response or logarithmic variables2.
Unlike conventional statistical software, IBM SPSS enables superior analytical abilities,
flexibility and usability, to provide better quality user experience, better productivity and
performance, powerful and leading edge analytics and extensibility for information technology
infrastructure.
Following are few key aspects of IBM SPSS -
¡ The ability to quickly analyze large datasets within pivot tables Seamless integration with
Microsoft® Office applications
¡ Access to a newly improved Syntax Editor, with auto-completion, auto indentation, color-
coding and other features to make it easier to automate analytics production jobs
¡ An Interactive Model Viewer
¡ Access to multiple interface languages for global teams that may be working on the same
project
¡ Quickly prepare data in just a single step with Automated Data Preparation
¡ View significance tests in the main results table
¡ Fast performance on procedures for Frequencies, Descriptive and Crosstabs
¡ Manage and analyze business datasets
165
¡ Create customized, user-defined interfaces for existing procedures and user-defined
procedures
¡ Multithreaded procedures that improve performance and scalability
¡ Direct marketing functionality that allows business users to run their own analyses
¡ Bootstrapping capabilities that improve the stability of models
¡ Nonparametric testing procedures
¡ Regularization methods including Ridge regression, Lasso and Elastic Net, that improve
predictive models by reducing coefficient variability
¡ Multithreaded algorithms including SORT, correlation, partial correlation linear
regression, multinomial linear regression, factor analysis
¡ Nearest Neighbor analysis for prediction or classification
¡ Non-linear data modeling procedures to discover more complex relationships in your data.
¡ Cox Regression that enables survival analysis for samples drawn by complex sampling
methods
¡ Support for 64-bit hardware on desktop for Windows and Mac
¡ Support for Snow Leopard™ on Mac OS® X 10.6
¡ Support for IBM System z servers running Linux®
¡ Mac and Linux users can connect clients to IBM SPSS Statistics Server
¡ Support for Python as a "front-end" cross-platform scripting language and support for R
algorithms
¡ Collaboration capabilities boost the productivity of analysts using IBM® SPSS®
Statistics, and server-based options increase scalability and performance
¡ IBM SPSS Statistics Server makes working with large data faster and more scalable, and
improves overall stability.
¡ Improved security enables it to run as non-root on Unix/Linux.
¡ Client and server software can be on different release levels (for example, client V21 and
server V20), simplifying administration
166
Figure 2: opening a new file
2. Introducing the interface
For this demonstration, we have saved the SPSS file as [Link].
a. THE DATA VIEW
The data file is showed in the IBM SPSS Statistics Data Editor. In the Data Editor, if you put the
167
mouse cursor on a variable name (the column headings), a more descriptive variable label is
displayed (if a label has been defined for that variable).
Further, to move to the first cell of the data view Press Ctrl-Home and to move to the last cell of
the data view Press Ctrl-End.
b. THE VARIABLE VIEW
To visit the variable view "Click" the Variable View tab given on the down left corner.
Review the information in the rows for each variable. In the variable view the variables are listed
in rows, with each column containing a specific kind of information of the variable. This is
different from the data view in which the variables are listed in columns (see fig 1)
To give definition of the variable double click on label id at the top of the id column. On double
clicking the names of the variable i.e. labels, in the data view variable view window will open.
Otherwise you can click on variable view and start filling each variable details.
To return back to data view click on 'data view' tab.
168
c. THE OUTPUT VIEW
For the output view one need to conduct an analysis.
So we start with simple frequency table (table to count the frequency of responses for a variable).
Figure 6: Way to compute frequencies
To do the same you can follow below given steps -
Go to MENU bar and click on the 'analyze' tab. Then go to 'Descriptive statistics' in dropdown
and click on 'Frequencies' (See Fig 6)
On clicking this tab, the dialogue will appear indicating a with variable list
169
As shown in figure 8, each variable has an icon next to them. Every icon indicates the data type
and measurement level of each variable.
Then click on the any one option (variable) from the list. if in case, the variable label and/or name
looks shortened in the list, you can position the cursor on that label to view the complete
label/name. The variable name for EDU is shown in the square brackets after the variable label
which describes the detailed name the label. EDUCATION is the variable label. If in case there is
any variable without a variable name, it would just appear in the list box.
The dialogue box can be resized depending upon the requirement, just like windows, by clicking
and dragging the outside border or the corners. As the dialogue box becomes wider, the variable
list would also become wider.
In order to analyze the variable frequency, you can chose any from the list and click on it and then
click on icon to shift the variable from left to right. Other way to do it, you can click and drag the
variable from left to right box and then click OK to run the analysis. To enable the OK button, you
need to shift at least one variable in the Variable(s) list. In the given example, the EDU variable is
shifted to Variable(s) list and then OK is clicked.
170
In SPSS, each window handles a separate task.
The results of the analysis run on SPSS is shown in viewers' window in output view. You can also
go to the any item in the viewer by clicking on the list in outline pane.
171
Further to analyze the data and see the output- Go to 'analyze' and click on 'descriptive statistics'
and 'crosstabs'. Here you can notice a dialogue box similar to previous selected options. To
finalize the analysis, Click on OK. From the output view you can select the charts or tables, copy
them, and paste them into other applications like spreadsheets or word processors.
Note: If you want to maintain the correct spacing of the tables, use a non-proportional
font like Courier New.
The syntax view
The syntax view is fundamentally the computer code that leads to a specific output. From time to
time graphical interface is preferred to pursue the daily routine work. Though at certain times
there is a need to reproduce the steps to come to a particular conclusion. This is requird to
replicate the analysis. For this SPSS syntax view is the best one.
CROSSTABS
/TABLES=TEACHER BY AGE
/FORMAT= AVALUE TABLES
/CELLS= COUNT
/BARCHART.
In the above given code, the SPSS is tutored to make crosstabs by utilizing teacher sorting the
crosstabs by age through using a specific format. Further it puts a count onto each cell and thus
making a bar chart
You should keep in mind to preserve the syntax code ones used., especially in case you are
calculating results for writing papers, reports and like.
172
Following are the steps to prepare charts and frequency distribution by executing saved syntax.
d. Go to menu bar
e. Click on analyze > Descriptive statistics > Crosstabs (keep the previous selection of
syntax)
f. Instead of pressing OK, click on Paste
g. SPSS will automatically bring the Syntax Editor with the code you just have pasted.
h. Run the syntax and get the output.
Crosstab
Crosstab refers to a short form of Cross tabulation which represents a summary table
emphasizing on the summary. The crosstabs are used for categorical data or discrete data like
gender or employment status. The crosstabs cannot be used for data which is continuous in nature
like income, dosage etc. such data can only be entered in crosstab by converting the continuous
data into groups like less than Rs 12000, between Rs.12000 to Rs.25000 and Rs. 25000 and
above.
173
VARIABLES:
Variables can be defined as a particular kind of information. Income, gender, or temperature can
be considered as variable. A few confuse in the terms like "concepts" and "variables". However,
concepts are the mental images or perceptions. Its meanings vary evidently from individual to
individual. Whereas variables are measurable, of course with varying degrees of accuracy.
VARIABLE TYPES
According to SPSS Step-by-Step Tutorial: Part 1, SPSS uses (and insists upon) what are called
strongly typed variables. Strongly typed means that you must define your variables according to
the type of data they will contain. A user can use any of the variable types, as defined by the SPSS
Help file. The SPSS data editor ACCEPTS following forms of numeric strongly types variables-
¡ The basic form of variable is the standard numeric format which is required to be provided
in scientific notation or standard format.
¡ Comma is the another numeric variable, in which value are indicated with commas to
delimit every three places and with period as decimal delimiter.
¡ Dot is the variable which is exhibited with periods delimiting every three places and
comma as a decimal delimiter.
¡ Scientific notation is a numeric variable whose values in shown with embedded E and a
signed power of ten exponents. Such variable can be either preceded by E or D with an
optional sign or a sign alone.
¡ Date is another numeric variable whose values are displayed in one of several calendar date
or clock-time formats. A user can select the format like dates with slashes, hyphens,
periods, commas, or blank spaces as delimiters and enter the data.
¡ Custom currency are the variables which are displayed in one of the custom currency
formats that you have defined in the Currency tab of the Options dialog box.
¡ String are the values of the string variables which is neither numeric nor used for
calculations. Also known as alphanumeric variables, such variables consist of characters
up to a certain defined length, which distinct uppercase and lowercase
174
Missing values
If you do not enter any data in a field, it will be considered as missing and SPSS will enter a period
for you.
175
Step 3
176
Step 2
Correlations
tech_led is
Pearson
1 .502**
Correlation
tech_led
Sig. (2-tailed) .000
N 498 498
Pearson
.502** 1
Correlation
is
Sig. (2-tailed) .000
N 498 498
**. Correlation is significant at the 0.01 level (2 -
tailed).
177
The window below will pop up, and ask you to choose where to save it ("Browse…").
The default location will be shown as in the hard drive. You need to select the location carefully
like in the desktop or the Documents folder).
You can also mention the SPSS that in which category of file you want to save the file.
178
You will usually want to select "All Visible Objects" and export it as a Word/RTF (.doc) file. This
is the easiest way to save all your work in useful format (RTF is Rich Text Format, which can be
read in nearly any text application on any platform
SPSS Functionality
SPSS has a very flexible data handling capability. SPSS provides has huge range of statistical and
mathematical functions, statistical procedures and can read data in almost any format (e.g.,
numeric, alphanumeric, binary, dollar, date, time formats) as explained above. To the help of
researcher SPSS also has excellent data manipulation utilities.
The following is a brief overview of some of the functionalities of SPSS:
¡ Data transformations
¡ Descriptive Statistics
¡ Data Examination
¡ Reliability tests
¡ Contingency tables
¡ Correlation
¡ T-tests
¡ ANOVA
¡ MANOVA
¡ General Linear Model (Release 7.0 and higher)
¡ Regression
¡ Logistic Regression
179
¡ Nonlinear Regression
¡ Loglinear Regression
¡ Factor Analysis
¡ Discriminant Analysis
¡ Cluster anlaysis
¡ Probit analysis
¡ Multidimensional scaling
¡ Survival analysis
¡ Forecasting/Time Series
¡ Graphics and graphical interface.
¡ Nonparametric analysis
CONCLUSION
This chapter aims to give basic information about the most sorted statistical software for data
analysis in social sciences and management. This software eases the data entry, processing and
output derivation, which eliminates time taking manual process of analysis. Fow effective usage
of SPSS, it is very important to understand varied aspect of SPSS and then start working for your
research. There are many more information related to SPSS which is provided in SPSS tutorials
provided from time to time by the IBM with every upgraded version of SPSS.
REFERENCES
¡ Arkkelin, D. (2014). Using SPSS to understand research and data analysis.
[Link]
¡ Dan Flynn (n.d.), Guide to SPSS, Barnard College
[Link]
¡ Landau, S. (2004). A handbook of statistical analyses using SPSS. CRC.
[Link]
nalyses_using_SPSS.pdf
¡ Garth, Andrew (2008), Analysing data using SPSS, Sheffield Hallam University.
[Link]
Webpages
¡ [Link]
¡ [Link]
[Link]
[Link]
¡ [Link]
¡ [Link]
¡ [Link]
180
T2217
BUSINESS STATISTICS
MBA I SEM I
ISBN: 978-93-95877-05-3