0% found this document useful (0 votes)
13 views185 pages

MBA Business Statistics E-Content Guide

The document outlines the e-content for a Business Statistics course for MBA students, detailing its establishment under the UGC Act and the academic structure, including modules on various statistical concepts. It emphasizes the importance of statistics in research, defining it comprehensively and outlining its characteristics, methods, and stages of statistical investigation. The content is intended for internal academic use, with strict copyright restrictions.

Uploaded by

Ajay Khedkar
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
13 views185 pages

MBA Business Statistics E-Content Guide

The document outlines the e-content for a Business Statistics course for MBA students, detailing its establishment under the UGC Act and the academic structure, including modules on various statistical concepts. It emphasizes the importance of statistics in research, defining it comprehensively and outlining its characteristics, methods, and stages of statistical investigation. The content is intended for internal academic use, with strict copyright restrictions.

Uploaded by

Ajay Khedkar
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Established under Section 3 of the UGC Act.

1956 I Awarded Category - I by UGC

E-CONTENT
BUSINESS STATISTICS
MBA I SEM I
Dr. Anugamini Priya Srivastava
Established under Section 3 of the UGC Act. 1956 I Awarded Category - I by UGC

Gram: Lavale, Tal: Mulshi, Dist: Pune, Maharashtra, India Pin: 412115

E-CONTENT
BUSINESS STATISTICS
MBA I SEM I
Internal Advisory Board (Self-Learning Material)

Chancellor : Prof. Dr. S B Mujumdar ([Link]. Ph.D.)


Distinguished Academician &
Educationist (Awarded Padma
Bhushan and Padma Shri by President
of India)
Pro-Chancellor, Symbiosis International : Dr. Vidya Yeravdekar
(Deemed University) & Principal Director, Symbiosis
Vice-Chancellor : Dr. Rajani [Link]
Provost, Faculty of Health Sciences : Dr. Rajiv Yeravdekar
Dean-Academics & Administration : Dr. Bhama Venkataramani
Dean, Faculty of Law : Dr. Shashikala Gurpur
Dean, Faculty of Management : Dr. R. Raman
Dean, Faculty of Computer Studies : Dr. Dhanya Pramod
Dean, Faculty of Media and Communication : Dr. Ruchi Jaggi
Dean , Faculty of Humanities & Social Sciences : Dr. Jyoti Chandiramani
Dean, Faculty of Engineering : Dr. Ketan Kotecha
Dean, Faculty of Architecture & Design : Dr. Sanjeevani Ayachit
Director, Symbiosis School for Online & Digital Learning : Dr. Raju Ganesh Sunder

Programme Coordinator - Management Programmes


Dr. Pravin Narayan Mahamuni

Self-Learning Material: E-Content


T2217 Business Statistics (MBA I SEM I)

Authors : Dr. Anugamini Priya Srivastava

Editor & Reviewer : Dr. Akriti Chaubey

ISBN: 978-93-95877-05-3

Copy Right © Registrar, Symbiosis International (Deemed University), Pune

All rights reserved. No part of this content may be reproduced or used in any form or by any means without
prior permission from the publisher. This content is produced only for academic purpose and is for internal
circulation only.

Acknowledgement : Every attempt has been made to trace the copyright holders of material reproduced in
this content. Should any infringement have occurred. We apologize for the same and would be pleased to make
necessary corrections in future editions of this content.

Published by : Symbiosis School for Online and Digital Learning, SIU, Lavale, Pune

Layout and Designing : Neha Creations, Pune 411030


CONTENT

MODULE - 1 : Introduction to Business Statistics 01

MODULE - 2 : Data Presentation 12

MODULE - 3 : Measures of Central Tendency 25

MODULE - 4 : Measures of Dispersion 41

MODULE - 5 : Measures of Dispersion 53

MODULE - 6 : Introduction to Probability 61

MODULE - 7 : Discrete Probability Distribution 72

MODULE - 8 : Continuous Probability Distributions 85

MODULE - 9 : Sampling and Different Sampling Techniques 96

MODULE - 10 : Hypothesis Testing 109

MODULE - 11 : Non-Parametric Tests 121

MODULE - 12 : Analysis Of Variance (ANOVA) 131

MODULE - 13 : Simple Linear Regression 150

MODULE - 14 : Spss And Data Analysis 164


INTRODUCTION TO
MODULE - 1
BUSINESS STATISTICS

STRUCTURE
¡ Statistics and research
¡ Statistics definition
¡ Characteristics of statistics as statistical data
¡ Statistical methods
¡ Stages of statistical investigation
¡ Functions of statistics
¡ Scope of the statistics
¡ Limitations of statistics
¡ Summary

1.1 LEARNING OBJECTIVES


¡ To understand what is business statistics
¡ To understand the characteristics of business statistics
¡ To understand the uses and limitations of business statistics
¡ To understand the process of statistics

1.2 STATISTICS AND RESEARCH


Research is all about looking for something which remains to be searched. It is a process to
explore different topics and detect certain conclusions. Over time, analysis has been utilised to
bring groundbreaking theories and principles to guide other disciplines and subject areas.
However, research and exploration require empirical testing. Where on the one hand, theoretical
models are derived from literature and exploration, and statistical methods are utilised to test
statistically and verify the theoretical model. Researchers have acknowledged the importance of
statistics in research. The use of scientific methods has displayed their relevance in scientific
research. The complicated yet straightforward process of research starts from planning to
drawing meaningful conclusions and can be sorted to be dominated by statistical methods. The
research steps comprise planning, designing, collecting, and analysing data, followed by making
valid inferences and reporting the final results. All of these steps of research require statistical
understanding.
Where planning and designing require specific research design, statistics help to understand the
conceptual framework and enable effective execution of the next steps. The clarity of the
conceptual framework can help decide the method to select for data collection. Data collection
involves sample selection and data collection based on the research design. Statistical
understanding helps explore the variety of methods for sample selection and collection of data.

01
From probability sampling to non-probability sampling, statistical knowledge helps researchers
choose the best possible sampling method.
Similarly, statistical understanding helps identify and explore different modes of collection of
data and choose the best possible one from primary and secondary data sources. Along the same
line, statistics provides a wide range of methods to explore the data duly collected. It enables the
recording and cleaning of the data based on missing values and outliers. Further, it has in
identifying the most effective methodology to understand the characteristics of the data per
variable, explore the average or dispersion in the data, the fitness of the model and the
predictability of variables on others. Statistics also support making valuable inferences based on
the statistical thresholds. Using different references and guidelines, researchers can make useful
inferences based on their research ideas and hypothesis. And lastly, in research, statistics can
provide proper formats of data presentation - tabular or diagrammatic to enable researchers to
represent their creativity in report writing.
Thus, statistical analysis in research can prove to be the lifeline.

1.3 STATISTICS DEFINITION


many scholars and researchers defined statistics over some time. On the one hand, statistics was
defined as statistical data. This definition was taken as understanding in the plural sense of
statistics. According to Webster, statistics refers to "the classified facts respecting the condition
of people in a state, especially those facts which can be stated in numbers or tables of numbers or
any tabular or classified arrangement". The definition provided by Webster is too narrow. It can
find the scope of statistics to only such facts and figures related to the people's condition in a
state.
Yule and Kendall defined statistics as "quantitative data affected to a marked extent by a
multiplicity of causes". In this definition, again, the author restricted the meaning of statistics to
only quantitative data and varied factors affecting it. However, statistics can also comprise some
part of qualitative data analysis.
Considering the limitation of these definitions, professor Horace Secrist provided the most
comprehensive explanation for statistics as statistical data. Professor defined statistics as
"aggregates of facts affected to a marked extent by a multiplicity of causes, numerically
expressed, enumerated or estimated according to reasonable standards of accuracy, collected in a
systematic manner for a predetermined purpose and placed in relation to each other”
The definition provided explains the different dimensions and characteristics of statistics as a
statistical method. So, let's understand every segment of this definition explained in 1.4.

1.4 CHARACTERISTICS OF STATISTICS AS STATISTICAL DATA


1.4.1. Aggregates of facts
the first characteristic of statistics elaborates that statistics do not work with single and isolated
numbers of figures. It always includes aggregated facts and figures. The aggregated facts and
figures insurance effective comparability and establishment of relationships. For example, data
on the salary of an individual, individual purchase decision, birth, death etc., does not comprise
statistics. However, aggregated data on these factors can be taken into consideration in statistics.

02
1.4.2. Affected to a marked extent by a multiplicity of causes
generally, the facts and figures are influenced by the number of forces operating together for a
considerable time. Assuming that a particular point or figure works in isolation will be utterly
wrong in statistics. For example, suppose you talk about the rice production quality in a year. In
that case, we need to consider the data on the quantity of rainfall, the quality of soil seeds, and the
cultivation methods. So, in statistics, every piece of information collected is affected by multiple
causes. Different ways and means have been identified for segregating the effects of other forces
on a process. Similarly, statistics also considered the difficulty involved in measuring the
complex impact variety of factors which are not measurable.

1.4.3. Numerically expressed


all the facts and figures are numerically expressed in statistics. All the figures need to be
presented as numbers; only then can they be processed through statistical methods. For example,
a statement like "the babies learn faster in the first year after birth". This statement is not
numerically expressed. While a word like "the babies learn 39% of their senses in the first year
after birth" can be considered under statistics.

1.4.4. Enumerated or estimated according to reasonable standards of accuracy-


in statistics, the data is either estimated or enumerated. In situations where actual data is
unavailable, the overall statistics listed-like in election rallies - will be difficult to count the
number of individuals in recovery. Thus, enumeration is a better way to identify the statistical
data. While in a classroom, we can easily count and estimate the number of students taking
sessions. Here we can easily say that 45 students in a school are attending lectures.

1.4.5. Collected systematically for a predetermined purpose -


in statistics, to get accurate answers, the data must be collected in a systematized order keeping
the pre-determined goal in mind. As already mentioned, research and statistics go hand in hand.
So, to have reliable and accurate data collection for the nuanced analysis, research objectives
have to be kept in mind. Along the same line, the order of collection of data should not be done in
a haphazard manner, i.e. It should be done in a scientific way following every step of research.
For example, if the research problem is identifying the impact of leadership on employee
behaviour, the researcher needs to understand and clarify the research questions first, identify the
research design based on propositions or hypotheses, and then start the data collection. In the
same line, the data collection should start by identifying the respondents who can give the
answers accurately. So, in this example researcher need select leaders and their immediate
employees to collect the data. Further, after putting them, the data collection method should be
identified and finalized. It can be a questionnaire method, an interview method, or so on. Based
on the research problem; the researcher will have more clarity about the research design and the
prospective respondents, which will thus help in the systematic collection of data

1.4.6. Placed in relation to each other


The statistics are collected and analyzed to understand future trends and patterns. Similarly,
statistics help us understand how things have changed over time. This makes us conclude that
statistics need to have comparable statistical data. The data has to be collected in such a manner

03
that enables chronological or region-wise comparison for geographical comparisons, for
example, statistics collected on the per capita income of employees working in developed
nations. All of them can be compared when converted to dollars as the currency all these
characteristics make statistics very crucial in the field of research. Without following these
characteristics, data cannot be called statistics. Like it is well said all statistics are numerical
statements of facts, but all numerical statements of facts and not statistics.

1.5 STATISTICAL METHODS


Many researchers also define statistics as statistical methods. Very often, it is known as a
singular sense of defining statistics. A few definitions given by researchers on defining statistical
methods are given below.
According to professor A.L. Bowley, "statistics may be called the science of counting". On
another occasion, professor Bowley said, " statistics may be called the science of averages".
Another definition was given as statistics is the science of measurement of the social organism,
regarded as a whole in all its manifestations". In the first definition by professor Bowley, the
emphasis was on collecting data, including counting. However, the other aspects of statistics,
like presentation analysis interpretations, are entirely ignored. In the second definition, the main
emphasis was on averages considering statistics as only enabling the calculation of averages.
However, statistics involves a large number of statistical methods and the is premises. It is not
only about averages. Statistics do include tests like dispersion, skewness, correlation, regression
etc., although all of these tests for wholly ignored from the second definition. Talking about the
third definition given again, statistics was restricted to the scope of sociology, including a man
and his activities. With criticism about these definitions, professor Bowley recommended that
"statistics cannot be confined to any scope".
Boddington also defines statistics as statistical methods as "the science of estimates and
probabilities". Although again in, this definition is only on estimates and probabilities, which
are just a tiny part of statistical methods.
The most comprehensive definition given to define statistical method was given by Croxton and
Cowden. They defined statistics as "statistics may be defined as a science of collection,
presentation, analysis and interpretation of data".
Croxton and Cowden's small yet comprehensive definition provided the significant steps
involved in statistical methods. It has already been emphasised that statistics have to have data
collected systematically. Thus, this definition clarifies the four stages of statistical investigation
given below.

1.6 STAGES OF STATISTICAL INVESTIGATION


The definition provided four different stages of statistical investigation. However, one more day
is added to business statistics - an organisation of data. So, let's understand the 5 stages of
statistical investigation.
1.6.1 Collection of data: the first step of the statistical investigation is the collection of data.
Based on the pre-determined purpose, the data is supposed to be collected in a systematized
manner. Since it is the first step, proper care must be exercised. This is important because it
creates the foundation for further analysis. If the data is not collected correctly, then that data
will not be considered reliable. The data can be collected from two different sources- namely,
primary source and secondary source the primary source includes first-hand data collection

04
where a statistician contains the data based on the research objectives. This data is collected
based explicitly on the purpose of the research projects. Another way to gather information is the
secondary sources. The secondary sources are the published or unpublished sources from where
the investigator collects the data and use it in their projects. However, they need to be very
cautious regarding the use of secondary sources of data. The reliability and suitability of such
data should be tested and only utilised in this study.

1.6.2 Organization of data:


the next step is to organise the data. This step includes editing and classification of data. When
the data is collected, they can be specific errors in the data, like missing values and irrelevant
answers. Similarly, in the case of open-ended questions, the responses can be different for each
respondent. This step of the data organisation helps investigators edit the collected data carefully
to avoid missing values, inconsistencies, irrelevant answers, vague answers and wrong
calculations. Once the data is clarified and revised, the information is classified based on
common characteristics. And finally, based on the classification and tabulations, the present
data can be more refined.

1.6.3 Presentation of data:


once the data is organised in a tabular form, we can present the data through diagrammatic
formats the most common way of giving data are bar graphs pie charts and histograms. We will
discuss all these methods of data presentation in the upcoming chapters. The data submitted
using the most appropriate diagrams can provide more order and understanding of the data.

1.6.4 Analysis of the data:


depending upon the purpose of the study, tabulations are done for the data. As mentioned, the
tabular presentation of data enhances the clarity of the characteristics of the data. Once all three
steps are done - collection, organisation and presentation of data, the next step is the data
analysis. During the data analysis, the investigator needs to identify the most suitable data
analysis method to conclude. For example, if the investigator wants to explain the effectiveness
of leaders in an organisation, a questionnaire method can be utilised for data collection, which
can then be organised in a table, avoiding the missing values and outliers and then descriptive
analytics, including measures of location and measures of dispersion, can be initially utilised to
serve the purpose. Further, if a statistician wants to establish a relationship between leadership
and employee stress, then the Pearson correlation test can be used.

1.6.5 Interpretation of data:


interpretation of the data is the last step of statistical investigation; however not the simplest one.
This step calls for a high degree of skill and experience among researchers to provide practical
conclusions. If the data is not adequately investigated, it cannot allow proper interpretation and
will give defeated and vague deductions. The correct interpretation is believed to be a valid
conclusion of the study and, thus, in decision-making.

05
1.7 FUNCTIONS OF STATISTICS
So now, let's understand the functions of statistics
1.7.1. Definiteness: the statistics specifically require a clear and definite form of general
statements. The numbers and details must be precisely defined and mentioned for accurate
results. For example, saying statements like the GDP of the country has increased. This statement
is neither specific nor definite in any sense. While rewriting the same statement as the country's
GDP has risen from 11% in 2020 to 15% in 2022, this statement provides more clarity to the
readers with a definite meaning.

1.7.2. Simplified mass of figure: along with being definite, statistics helps condense the
gathering of data into a few significant figures with proper and more precise meanings. In a raw
data set, multiple columns and rows of data exist. Using statistical methods, one can condense the
data based on common characteristics and give proper meaning for better interpretations. For
example, by reading the census reports on individual salary structures in a country, a researcher
might not understand the income of the entire population. However, a simplified value of per
capita income can be easily remembered and understood by everyone.

1.7.3. Facilitate comparison: another function of statistics is to facilitate comparison. When


collected systematically, statistics can ensure comparison. Unless the figures can be compared
with another figure of the same kind, they are devoid of any meaning in statistics. For example,
data on rainfall is collected for one year while the unit of measurement differs in every country. In
such a situation, the comparison will not be possible until all data is transformed into the same
division. Similarly, when car sales data indicate sales in rs crore for two different years, like
20,000 cr in 2020, increased to 25,000 cr in 2022, it can easily ensure comparisons. So, statistics
can provide a more straightforward comparison between data

1.7.4. Formulating and testing hypothesis: statistical methods can help investigators to use
different tools and measurements to formulate and test the hypothesis. For example, if the
investigator aims to test the effect of increased dopamine on students' performance in the
classroom, then using statistical understanding, null and alternative hypotheses can be
formulated. If the alternative hypothesis is that dopamine affects students' performance, then a
simple linear regression analysis can be used to test the theory.

1.7.5. Helps in prediction - statistics also enables effective prediction for the future. Using
statistics, researchers can evaluate the overall trend and utilise the information to forecast future
events. For example, data on employee turnover from the last 5 years can be used to predict the
prospective number of employees that might leave the organisation in the next 6 months. Based
on this prediction, good employee retention policies can be developed well in advance.

1.8 SCOPE OF THE STATISTICS


Statistics can be applied in different fundamental and multi-disciplinary areas. For example,
statistics can be utilized in marketing to design and re-designing advertisements for products,
placements of products on shelves in retail stores, bringing realistic ideas into ads to attract more
customers and so on. The list is on and on. The scope of statistics is so vast that it not only helps in

06
proper data collection, summarization and presentation, but it also helps marketers to develop
new strategies and deal with problems.

1.8.1 Operations
In the field of operations management and supply chain management, statistics support quality
management. On the one hand, it helps manufacture the products with ideal size and
composition; on the other hand, it helps quality check through random selection of items from the
lots. Through a random sample of articles from conveyor belts, managers can check whether the
quality of the product is acceptable or rejected. Similarly, on delivery of raw materials, managers
and executives can verify if the materials delivered are of prescribed quality or not.

1.8.2 Accounting
In accounting, statistics are utilized explicitly in the auditing process. In the auditing process,
auditors cross-verify vouchers or bills with transactions shown in ledgers and journals, also
known as bookkeeping. For cross-verification, the auditor selects vouchers and statements
randomly from the registers and checks them in correspondence to transaction entries. Through
this method of random sampling, there are higher chances that frauds and mistakes can be
detected, if any. We will discuss random samples in upcoming chapters.

1.8.3 Human resource management


In human resource management, information systems use statistics to shortlist candidates for
interviews, select employees for training based on training needs, ensure effective performance
appraisals, predict future employee turnover and absenteeism, and so on. The advanced usage of
statistics is done in human resource analytics.

1.8.4 Banking
Similarly, in banking, statistics are used to collect and analyze information to understand future
economic conditions and evaluate other external needs to understand every line of business in
which they might be directly or indirectly interested.

1.8.5 Others
Similarly, statistics can be utilized in investment, purchase, credit and control, personnel
management, and research and development. Thus, we can conclude that the scope of the
statistics is vast and continuously increasing. Due to this, it is tough to define statistics. At the
same time, it is unwise to explain the exact usage of statistics as it is permeated almost every
aspect of our lives.

1.9 LIMITATIONS OF STATISTICS


From the above discussion, it is pretty clear that statistics is way too much value in our lives.
However, like every concept has some limitations, so is the case with statistics. It is not a magical
device which provides a correct solution to problems. It requires a proper data collection
structure for critical analysis to attain valid conclusions. Without a systematic process, statistics
can lead investigators to draw wrong conclusions.
So, let's see and discuss the limitations of statistics.
07
1.9.1. Does not deals with isolated measurements - statistics do not deal with individual or
isolated measurements. It requires aggregated or average estimates for analysis and decision-
making. For example, to make policies for employee retention, hr managers need to take
information on compensation and rewards, the experience of employees, job satisfaction and
other related factors. But the observations on each element have to be taken as the average of all
responses of individuals. No policies can be developed based on only individual reactions.

1.9.2. Deals with numbers- statistics involves numeric facts and figures. It does not enable the
usage of qualitative or text data. For example, data on males and females, usually considered
nominal data, can be used for making a diagrammatic presentation. However, it cannot be used
for analysis. For analysis, these labels must be transformed into codes, like male-0 and female-1.

1.9.3. Are true only on an average- the conclusions in statistics are based on the purpose of
research, hypothesis developed, research design adopted and disciplinary background of
researcher. So, it is not true that the conclusion made in one analysis is universally true. A
generalisation of results is required to verify the application of the decision in another context.

1.9.4. Only a means- statistical methods furnish only one way of studying a problem. They may
not be the best under all circumstances. So, it should be carefully noted that statistics is only a
means, not an end. It analyses the facts and throws light on the actual situation. Complete
dependence on statistics male lead to fallacious conclusions in many cases. Proper consideration
of factors affecting and cautious organization and analysis is required for accurate conclusions.

1.9.5. Can be misused- the most significant limitation of statistics is that it is liable to be
misused. It is essential to understand that statistics are likely, and they can be moulded in any
manner to establish right or wrong conclusions if the statistical findings are based on incomplete
information, one may arrive at false conclusions. That's why it requires experience and skills to
draw sensible conclusions from the data. Otherwise, there is every likelihood of a wrong
interpretation.

1.10 SUMMARY
The chapter involved a basic understanding of the term business statistics. Below given is the
summary of this chapter.
¡ Statistics and research: research and statistics go hand-in-hand. However, research can
be conducted without statistics, but statistics without research in the business world will
not help in attaining practical conclusions and decisions.
¡ Statistics definition: statistics can be defined as both statistical methods and data.
Different scholars provided different meanings to explain the crux of statistics and its
characteristics.
¡ Characteristics of statistics as statistical data: significant features of statistics as
statistical data are that it comprises aggregates of facts affected to a marked extent by a
multiplicity of causes, numerically expressed, enumerated or estimated according to

08
reasonable standards of accuracy, collected in a systematic manner for a predetermined
purpose and placed with each other.
¡ Statistical methods: in the singular sense, statistics is also defined as statistical methods.
¡ Stages of statistical investigation: statistics involves 5 stages- collection, organisation,
presentation, analysis and interpretation.
¡ Scope of statistics: statistics are helpful in every discipline, from operations management
to production management to human resource management and marketing. It is also
beneficial in fundamental grounds of accounting and banking.
¡ Limitations of statistics: statistics have certain limitations which restrict their usage in
research work. It only uses aggregates and numerical data for analysis. It is determined to
be correct on average and can lead to misleading results if not conducted with diligence
and understanding of the research purpose.

1.11 SELF ASSESSMENT QUESTIONS


Long questions
1. Explain the evolution of statistics
2. Explain how statistics and research are interconnected.
3. Explain the limitations of statistics
4. Discuss the functions of statistics
5. Elaborate on the scope of statistics
6. Explain the characteristics of statistics

Short questions
1. How statistics can be helpful in the field of banking
2. How statistics can be misused
3. Define statistics as statistical methods
4. Define statistics as statistical data

Fill in the blanks


1. Yule and Kendall defined statistics as "quantitative data affected to a marked extent by a
……………………………………". (multiplicity of causes)
2. All the facts and figures are ……………………….. Expressed in statistics.(numerically)
3. Boddington defined statistical methods as "the science of …………………………".
(estimates and probabilities)
4. In statistics, to get accurate answers, the data must be collected in a systematized order
keeping the …………………………. in mind. (purpose)

09
True/ False
1. Statistics specifically require a clear and definite form of general statements. (True)
2. Statistics help us understand how things have changed over time. (True)
3. To have reliable and accurate data collection for the nuanced analysis, research objectives
need not be kept in mind. (False)
4. In statistics, unless the figures can be compared with another figure of the same kind, they
are devoid of any meaning. (True)

Multiple choice question


1. Statistical investigation includes 5 steps process. Choose the correct process from the
below -
a. Collection, organisation, presentation, analysis and interpretation
[Link], presentation, organisation, analysis and interpretation
c. Collection, organisation, analysis, presentation and interpretation
[Link], presentation, analysis, organisation and interpretation

2. In accounting, statistics is used explicitly for


a. Quality assurance
b. Auditing process
c. Recruitment
[Link] of the given options

3. Statistics are prone to be misused because


a. They can be moulded to establish right or wrong conclusions
b. It helps condense the collected data into a few significant figures with proper and more
precise meanings
c. It is affected to a marked extent by a multiplicity of causes
[Link] of the given options

4. Choose which of the following are not a function of statistics


a. Facilitate comparison
b. Formulating and testing hypothesis
c. Helps in prediction
d. Includes only qualitative data

10
1.12 REFERENCES
¡ Black, K. (2019). Business statistics: for contemporary decision making. John Wiley &
Sons.
¡ Anderson, D. R., Sweeney, D. J., Williams, T. A., Camm, J. D., & Cochran, J. J. (2020).
Modern business statistics with Microsoft Excel. Cengage Learning.

11
MODULE - 2 DATA PRESENTATION

STRUCTURE
¡ Introduction
¡ Frequency distributions
¡ Steps to develop frequency distributions
¡ Class midpoint
¡ Relative frequency
¡ Cumulative frequency
¡ Quantitative data graphs
¡ Qualitative data presentations

Learning Objectives
¡ To understand the meaning of diagrammatic presentations
¡ To understand the frequency distributions
¡ To understand quantitative modes of data presentation
¡ To understand the qualitative modes of data presentation.

2.1 INTRODUCTION
In the era of data analytics and data science, it is essential to understand how these data can be
summarized and presented to ensure effective communication of results. Thus, in this chapter,
we will discuss different modes of diagrammatic presentations using qualitative and quantitative
data with the help of examples.
The first step toward analyzing the data is to explore the data using tabular or diagrammatic
presentations.

2.2 FREQUENCY DISTRIBUTIONS:


The frequency distributions refer to a summary of data presented in class intervals and
frequencies. The frequency distributions are relatively more accessible to construct. However, to
make it properly, standard rules and guidelines are to be followed. Depending upon the nature of
the data, different shapes and designs of frequency distribution tables can be developed. Even if
the data are identical, various forms can be created, depending upon the taste of individual
researchers.

12
Table 2.1: Data on the number of employees fallen sick in 12 months in the year 2021
34 12
23 34
20 23
56 16
67 11
20 10

STEPS TO DEVELOP FREQUENCY DISTRIBUTIONS


Determine the range of the raw data: The range often is defined as the difference between the
largest and smallest numbers. So before moving ahead, researchers need to identify the range of
numbers. For example, in the data given in table 2.1, the range of the data is 67-10 = 57.
Determine how many classes it will contain: next step is to identify how many classes one needs
to present the data in distribution effectively. A standard rule for it is to select between 5 and 15
classes. If the number of class are too few, then the data summary may be too general to be of any
use to researchers. While if the number of classes is very high, then the frequency distribution
will not aggregate the data correctly and will not be helpful to researchers. Although, the decision
of a number of classes is arbitrarily taken by the researcher's diligence. See table 2.2; the data
given in table 2.1 was divided into 6 classes with the frequency of values within each classes.

Table 2.2: Data on the number of employees fallen sick in 12 months in the year 2021
(grouped)
Classes Frequency
10-20 4
20-30 4
30-40 2
40-50 0
50-60 1
60-70 1

Determine the width of the class interval the next step is to identify the class width. The formula
to calculate the class width = range of data by a number of classes. For the data given in table 1.1,
the approximation would be =
57 = 9.5
6
So the class width can be considered as 10, rounding off to the following whole number, which is
10. The frequency distribution must start at a value equal to or lower than the lowest number of
the ungrouped data and end at a value equal to or higher than the highest number. So, based on the
minimum days of absenteeism as 10 and the highest at 67, the frequency distribution can start at
10 and end at 70. (see table 2.2)

2.3 CLASS MIDPOINT


The midpoint of each class interval is known as the class midpoint. It is also called a class mark.

13
For data summarization and presentation, class midpoints are essential. In other words, a class
midpoint can be defined as a value halfway across the class interval and calculated as the average
of the two class endpoints. We can calculate midpoints using the given formula =

Where,
LL represents the lower limit of the class interval
UL represents the upper limit of the class interval
Thus, using this formula, for the data in table 1.3, midpoints can be calculated as

(10+20) / 2 = 15

2.4 RELATIVE FREQUENCY


The following form of frequency which are used for depicting table and diagram indicating
relative proportions are related to frequency. It can be defined as the proportion of the total
frequency in any given class interval in a frequency distribution. Table 2.3 lists the relative
frequency for all classes. The relative frequency can be calculated by -
class frequency
total frequency

i.e. for example, for the class 10-20, the relative frequency would be = 4/12
= 0.333333
Table 2.3: Class Midpoints, Relative Frequencies, and Cumulative Frequencies for
absenteeism

Class Frequency Mid Relative Cumulative


interval value frequency frequency
10-20 4 15 0.33 4
20-30 4 25 0.33 8
30-40 2 35 0.17 10
40-50 0 45 0 10
50-60 1 55 0.08 11
60-70 1 65 0.08 12
Total 12

2.5 CUMULATIVE FREQUENCY


The cumulative frequency refers to running the total of frequencies through the classes of a
frequency distribution. In other words, the cumulative frequency for a class interval can be
calculated by adding the frequency estimate of the class interval with the prior cumulative total.
As shown in Table 2.3, class 10-20 is the same as its frequency, i.e. 4, while for class interval 20-
30, the cumulative frequency will be 4+4 =8.

14
Similarly, for other classes, the cumulative frequency can be calculated. This process continues
through the last interval, at which point the cumulative total equals the sum of the frequencies. In
other words, to verify the estimates in cumulative frequency, we can cross-check that the final
number is the cumulative frequency column with the total frequency column. In example 2.3, the
final cumulative value is 12, equal to the total frequency.

2.6 QUANTITATIVE DATA GRAPHS


Quantitative data are the data which are numerically expressed and are measured by interval and
ratio scale of measurement. To effectively present any data, it is suggested to transform it into
graphs and plots. There are five major types of quantitative data graphs -
1) histogram.
2) frequency polygons
3) Ogives
4) dot plot
5) Stem and leaf plots

2.6.1 HISTOGRAM
The histogram is a series of contiguous bars or rectangles that denote the frequency of data in
given class intervals. If the class interval is equal, then the frequency of the values in each class
interval is represented through the height of the bars. While on the other hand, if the class
intervals are unequal, then the relative comparisons of class frequencies are depicted in the areas
of the bars (rectangles). As shown in figure 2.1, a histogram involves an x-axis labelled with class
endpoints and a y-axis marked with their respective frequencies. Here, the x-axis is also known
as abscissa, while the y-axis is named ordinates. Thus, in the histogram are drawn drawing a
horizontal line from the frequency value of one class endpoint to another class endpoint and
interlinking each one vertically from the frequency value to the x-axis to form a series of bars, as
shown in figure 2.1.

Figure 2.1: Histogram depicting the data on absenteeism

15
2.6.2 FREQUENCY POLYGONS
Frequency polygons are another way of presenting a quantitative data set. Theoretically, it is
constructed by scaling class midpoints along the horizontal axis and the frequency scale along
the vertical axis. In other words, it is similar to a histogram. But instead of bars or rectangles, it is
based on plotting a dot at the class midpoint and then connecting each dot by a series of line
segments. As seen in figure 2.2, the dots are plotted and then a line segment connecting each dot
is drawn.

Figure 2.2: Excel produced a frequency polygon for the days of absenteeism

2.6.3 OGIVES
An ogive is another way of presenting quantitative data. However, unlike histograms and
polygons, it is made based on cumulative frequencies. So, a graphical presentation of cumulative
frequency values. In the construction of ogives, the following steps are taken
¡ Labelling x-axis with class endpoints and y-axis with frequencies
¡ Scaling y-axis enough to include the total of frequencies
¡ Starting with plotting a 0 at the beginning of the first class
¡ Preceding with marking each dot at the end of each class interval for the cumulative values.
¡ Lastly, connect all dots to complete the presentation.

The diagrammatic presentation of ogive is most useful when the researcher aims to see running
totals. For example, if researchers are interested in controlling the overall production costs, an
ogive can depict the cumulative costs of a financial year. As shown in figure 2.3, a particularly
steep slope occurs in the 20-30 class interval, signifying a significant jump in class frequency
totals.

16
Figure 2.3: Ogive using excel

2.6.4 DOT PLOTS


Dot plots are simple charts generally made to display continuous, quantitative data. Each data
value is plotted along the horizontal axis using a dot. If multiple data points have the same values,
the dots will stack up vertically. If there are many close points, it may not be possible to display
all data values along the horizontal axis. The dot plots are helpful mainly for observing the
overall shape of the data distribution points along with intervals with identifying data
values or intervals for which there are groupings and gaps in the data. To simplify, dot plots are
made using the following steps-

Table 2.4: Re-organize the data for dot plots

Value Frequency Values Frequency


10 5 10 0
20 2 10 1
30 4 10 2
40 3 10 3
50 1 10 4
60 1 20 0
70 0 20 1
40 0
40 1
40 2
40 3
50 0
50 1
50 2
60 0
70 0

17
¡ First, reorganize the data into a "long" format (as shown in table 2.4)
¡ Step 2: Create a dot plot using the "scatterplot" option in excel as shown in figure 2.4.

Figure 2.4: Excel steps for scatterplot.

¡ Step 3: Customize the chart as shown in figure 2.4


- Delete the gridlines.
- Delete the title.
- Increase the size of the individual dots.
- Change the x-axis to only span from 1 to 7.

Figure 2.5: Dot plot for data on absenteeism using excel.

2.6.5 STEM AND LEAF PLOTS


Another way of organising and presenting quantitative data using the frequency distribution is
the stem and leaf plot. This plot is developed by separating the left and right digits for each
number of the data into a stem and a leaf. The process has three steps
¡ Left most digits are considered as stem and consist of higher valued digits

18
¡ Right-most digits are considered leaves and consist of lower values.
¡ If a set of data has only two digits, the stem is the value on the left, and the leaf is the value
on the right
Let's develop a stem and leaf plot using the example in table 2.1. The stem and leaf plot are
depicted in Figure 2.6

Figure 2.6: Stem and leaf plot

1 0
1 2 6
2 0 0 3 3
3 4 4
5 6
6 7

The primary benefit of using stem and leaf plots is that the investigator can readily see whether
the scores are in the upper or lower end of each bracket and also determine the spread of the
scores. Further, this plot can help determine the spread of the scores. Another advantage of using
this plot is that the values of the original raw data are retained. In contrast, other frequency
distributions and graphic presentations use the class midpoints to depict the values in a class
interval.

2.7 QUALITATIVE DATA PRESENTATIONS


Now let's discuss the qualitative graphs for data presentations. We are going to discuss two types
of qualitative data graphs 1. Pie charts, 2. Bar charts,

2.7.1 PIE CHARTS


A pie chart is a circular depiction of data. Under this diagrammatic presentation, the area under
the whole pie represents 100% of the data and slices of the pie represent a percentage breakdown
of the sublevels. As shown in Figure 2.7, the data on the gender of respondents is depicted in a pie
chart based on the values given in Table 2.5.

Table 2.5: Data on gender of respondents

FEMALE 13
MALE 7

19
Figure 2.7: Pie chart using excel

Here we can see the total area under the pie is 100%, and the angle is 360 degrees. So, to identify
the proportion, we can use the formula as female = 13/20*100 = 65 % & male = 7/20* 100 = 35%.
Thus, based on the example shown in figure 1.7, we can conclude that the respondents in the
survey were majorly females (65%) while males were only 35%.
Now converting this into degrees, we can calculate the angle as 13/20*360 = 234 degrees for
females. Similarly, for males = 7/20 * 360 = 126 degrees. Thus, the pie chart shows the relative
magnitude of the part to the whole. Pie charts are widely used to depict variables like market
share, resource allocations, investment patterns etc.

2.7.2 BAR GRAPHS


Another simple way of presenting qualitative data is a bar graph or bar chart. In a bar chart, we
can show two or more categories along with one axis and a series of bars along the other axis, one
for each category. Usually, the length of the bar indicates the magnitude of the measure. As
shown in figure 2.8, the data given in table 1.5 is shown as bar chart. The bar graph is qualitative
because the categories are non-numerical, and it may be either horizontal or vertical.

Figure 2.8: Bar chart using excel

20
2.8 COUNTRY MAPS
Apart from the above-given diagrams for qualitative data, we can also use country or state-based
data and use country maps in excel to depict. The steps to develop a country map is stated below
1. Summarize the data in excel
2. Go to insert tab and click on maps option under the charts tab as shown in figure 2.9

Figure 2.9: steps to choose country maps in excel

Click on map options and the diagram will automatically shown as shown in figure 2.10.

Figure 2.10: Country maps using excel

21
However, there are a few points that are needed to be kept in mind.
a. Maps charts can only plot high-level geographic details only. No latitude or longitude, or
street address can be mapped in it.
b. Maps charts also support one-dimensional display only.
c. Online connections are required to create new maps in excel or append data to existing
maps.
d. Existing maps can be viewed offline.

2.10 SUMMARY
In the era of data analytics and data science, it is essential to understand how these data can be
summarized and presented to ensure effective communication of results
Frequency distributions: The frequency distributions refer to a summary of data presented in
class intervals and frequencies.
Class midpoint: a class midpoint can be defined as a value halfway across the class interval and
calculated as the average of the two class endpoints
Relative frequency: It can be defined as the proportion of the total frequency in any given class
interval in a frequency distribution
Cumulative frequency: refers to running the total of frequencies through the classes of a
frequency distribution.
Quantitative data graphs: Quantitative data are the data which are numerically expressed and are
measured by interval and ratio scale of measurement. To effectively present any data, it is
suggested to transform it into graphs and plots. There are five major types of quantitative data
graphs - 1) histogram, 2) frequency polygons, 3) Ogives, 4) dot plot, 5) Stem and leaf plots
Qualitative data presentations can be done using two types of graphs 1. Pie charts, 2. Bar charts,

2.11 SELF ASSESSMENT QUESTIONS


Long questions
1. According to T-100 Domestic Market, the top seven airlines in the United States by
domestic boardings in a recent year were Southwest Airlines with 81.1 million, Delta Airlines
with 79.4 million, American Airlines with 72.6 million, United Airlines with 56.3 million,
Northwest Airlines with 43.3 million, US Airways with 37.8 million, and Continental Airlines
with 31.5 million. Construct a pie chart and a bar graph to depict this information.

2. According to the National Retail Federation and Center for Retailing Education at
the University of Florida, the four main sources of inventory shrinkage are employee theft,
shoplifting, administrative error, and vendor fraud. The estimated annual dollar amount in
shrinkage ($ millions) associated with each of these sources follows:
Employee theft $17,918.6
Shoplifting 15,191.9
Administrative error 7,617.6

22
Vendor fraud 2,553.6
Total $43,281.7
Construct a pie chart and a bar chart to depict these data.

3. The following data represent the number of passengers per flight in a sample of 50 flights
from Wichita, Kansas, to Kansas City, Missouri.
23 46 66 67 13 58 19 17 65 17
25 20 47 28 16 38 44 29 48 29
69 34 35 60 37 52 80 59 51 33
48 46 23 38 52 50 17 57 41 77
45 47 49 19 32 64 27 61 70 19
a. Construct a dot plot for these data.
b. Construct a stem-and-leaf plot for these data. What does the stem-and-leaf plot tell you
about the number of passengers per flight?

4. Construct a histogram and a frequency polygon for the following data.


Class Interval Frequency
30-under 32 5
32-under 34 7
34-under 36 15
36-under 38 21
38-under 40 34
40-under 42 24
42-under 44 17
44-under 46 8

Short questions
1. Define pie charts.
2. Define Bar charts
3. Define the advantages of stem and leaf plots
4. Explain the uses of country maps.

True and false


1. Pie charts are used for depicting absolute frequencies (False)
2. Bar charts are similar to histogram (False)

23
3. Ogive are developed based on cumulative frequency (True)
4. Country maps can be used to depict street addresses (False)

Fill in the blanks


1. The midpoint of each class interval is also known as ……………………. (class mark)
2. An ogive is another way of presenting ………………….. data. (quantitative)
3. The ……………………… frequency refers to running the total of frequencies through
the classes of a frequency distribution. (cumulative frequency)
4. In a stem and leaf plot, ……………….. digits are used to depict as leaf. (right)

2.12 REFERENCES
¡ Black, K. (2019). Business statistics: for contemporary decision making. John Wiley &
Sons.
¡ Anderson, D. R., Sweeney, D. J., Williams, T. A., Camm, J. D., & Cochran, J. J. (2020).
Modern business statistics with Microsoft Excel. Cengage Learning.

24
MODULE - 3 MEASURES OF CENTRAL TENDENCY

STRUCTURE
¡ Introduction to measures of central tendency
¡ Characteristics of a good average
¡ S Symbol
¡ Different types of measures of central tendency
¡ A.M. of Grouped frequency distribution
¡ Composite A.M.
¡ Advantages and disadvantages of A.M.
¡ Geometric Mean
¡ Advantages and disadvantages of G.M.
¡ Uses of G.M.
¡ Harmonic Mean
¡ Relationship among A.M., G.M. and H.M.
¡ Median
¡ Advantages and Limitations of Median
¡ Mode
¡ Relationship between Mean, Median and Mode
¡ Quartiles, Deciles and Percentiles

3.1 LEARNING OBJECTIVES


After going through this unit, you will be able to:
1. Understand what measures of central tendency are
2. Learn the different measures of central tendency
3. Explain the different types of means
4. Understand the relationship between mean, median and mode
5. Explain Quartiles, deciles and percentiles

3.2 INTRODUCTION
While working with data, we may sometimes need such a numerical expression which can tell us
certain characteristics of the whole data set, such as the central most value, the lowest value, the
highest value or the value which appears the most frequently. One such type of measure which
can tell us the average value of a distribution are known as the 'measures of central tendency' or
25
also known as 'Averages', more popularly. Such a measure which represents the middle most
value should obviously be greater than the smallest value and less than the highest value. A
measure of central tendency or an average of a certain distribution is nothing but a representative
value of that distribution which enables us to comprehend in a single effort the significance of the
whole. It should be a value which lies between the two limits, i.e, the highest and lowest points in
the data set, possibly at the centre, where most of the values of the series cluster.
Measures of central tendency or averages are additionally, also called as measures of central
location. These are arithmetical measures intended to represent the central value of a data set. We
can say that an average of a distribution (of the values) of any variable (say weight of some
students in a class in cms) is a representative value of that variable.
In any observation set, the representative value of a distribution usually lies at or near the centre
of the distribution. This happens due to the inherent tendency of a distribution of data of any kind
that the major part of the values gets concentrated at the centre. Since this average is reflective of
this tendency of the data, hence, the average is called a measure of central tendency.
Depending upon the nature of a distribution, different methods of obtaining the representative
value have been evolved and as a result, we have several averages or measures of central
tendency.

3.3 CHARACTERISTICS OF A GOOD AVERAGE


According to statisticians Yale and Kendall, an average will be considered good or efficient if it
possesses the following characteristics:
¡ An average should be easily understandable
¡ It should be rigidly defined. The definition should be such that it's interpretation is not
subjective in nature.
¡ The average should be such that it can be calculated easily.
¡ The average of a variable should be based on all values of the variable.
¡ The average should not reflect significant change in it's value if there is a change in sample.
In simple terms, an average should possess sampling stability.
¡ Such an average should be such that it is not unduly affected by extreme values, i,e, the
formula for average should be such that it does not show undue large change due to the
presence of one or two very large or very small values in the distribution set.

3.4 S SYMBOL
In order to denote sum (i.e, a total of certain quantities), the Greek letter S (capital sigma) is used.
For example, if a variable x takes the values x1, x2, x3…xn, then the sum of these values of the
n n
variable x i.e., (x1 + x2 + x3+… + xn) is denoted by S t =1 xi or S x. The symbol S t = 1xi means
that the lower limit of i is 1 and the upper limit of i is n, i.e., i takes the value 1, 2, 3, …, n and the
symbol S means that all the values of xi for i = 1, 2, 3,…, n are to be added. Again the symbol S x
implies 'sum of the values of x'.
Illustration: Express with the help of S symbol:
a. x4 + x5 + x6 + x7 + x8 + x9 + x10
Solution: x4 + x5 + x6 + x7 + x8 + x9 + x10 = S10i=4 xi

26
3.5 DIFFERENT TYPES OF MEASURES OF CENTRAL TENDENCY
The following three types of averages or measures of central tendency are used:
a) Mean b) Median c) Mode

a) Mean:
Arithmetic mean is the average of a group of observations and is calculated by adding all the
numbers and then dividing the sum so obtained by the number of observations. Because it is the
arithmetic mean out of the three types of means that is most commonly used by statisticians, the
arithmetic mean is more commonly known as only 'mean'.
Here, the mean for the population is represented by the Greek letter mu ( ). And the mean for the
sample is denoted by x ?. While talking of mean, there are three types of means, namely:
i) Arithmetic mean: (AM)
ii) Geometric mean: (GM)
ii) Harmonic mean: (HM)

i) The A.M. of a variable x is denoted by the symbol x ? and is defined to be the sum of the
values of x divided by the number of values of x.
Formula of AM in case of individual series:
x: x1, x2, x3, …, xn, `x will be:
t
`x = (x1+x2+x3+?+xn) = S i=1 xi = Sx
n n n

Formula in case of ungrouped frequency distribution:

x x1 x2 x3 … xn
f f1 f2 f3 … fn

3.6 A.M. OF GROUPED FREQUENCY DISTRIBUTION


In discussing grouped frequency distribution, while calculating the A.M., we make an
assumption that the observations included in a class represented by a class interval are
concentrated around the center of the class interval. To obtain the A.M. of a grouped frequency
distribution, we consider the frequencies of the class intervals to be the frequencies of the mid
values of the corresponding classes. By this process, we successfully convert a grouped
frequency distribution to a discrete (or ungrouped) frequency distribution. Hence by applying

27
the A.M. formula for discrete or ungrouped frequency distributions we can find the arithmetic
mean of a grouped frequency distribution.

Illustrative examples (1):


i. Find the A.M. of the following numbers
5, 8, 10, 15, 24 and 28.
ii. Find the A.M. of the following series:
x: 4, -2, 7, 0 and -1.
Solution:
5+8+10+15+24+28
i. The required A.M. = = 90/6= 15.
6
4+ −2 +7+0+(−1)
ii. The required A.M. = = 8/5 = 1.6
5
Example 2: Find A.M. of the following frequency distribution:
x: 1 2 3 4 5 6 7 8 9
f: 7 11 16 17 26 31 11 1 1

Solution: First of all we shall prepare the following frequency table:


x f fx
1 7 7
2 11 22
3 16 48
4 17 68
5 26 130
6 31 186
7 11 77
8 1 8
9 1 9
N= 121 åfx = 555

åfx
\ A.M = = 555/121 = 4.59 (approx.)
N

Note: All measures of central tendency or averages of a distribution will possess the same
unit of the distribution.

3.7 COMPOSITE A.M.


The A.M. of two or more distributions is called a combined or composite A.M.
If there are n1 values in a distribution (d1) where the A.M. is x ?1 and similarly the A.M. of n1
values of yet another set of distribution (d2) is x ?2 then the A.M. of the combined distributions
(d1 and d2) is given by:

Example: The average marks obtained by two groups of students in an examination are 75 and
85. If the average marks of all the students is 80, find the ratio of students in the two groups.
Solution: Let x denote marks of all the students, x1 denote marks of the first group, x2 denote
marks of the second group, n1 denote no. of students of the first group and n2 denote number of
students of the second group.

28
3.8 ADVANTAGES AND DISADVANTAGES OF A.M.
Advantages:
¡ A.M is easy to determine and understand.
¡ The A.M. is based on all values of the distribution.
¡ It can be used for further algebraic treatment.
¡ The formula for A.M is rigidly defined implying that for a given series, the value of A.M.
remains unique.
¡ It provides a good basis for comparison.
¡ The values of a series need not be arranged in any order for calculating the A.M.
¡ If the A.M. and the number of observations in the series are known, then we can also find
out the sum of the distribution.

Disadvantages:
¡ The A.M. is unduly affected by extreme (i.e, very large or small) values.
¡ The A.M. cannot be computed even if one of the values in the series is missing.
¡ The determination of the A.M. in case of a grouped frequency distribution can be
misleading as it is based on an unrealistic assumption that the observations of each class is
concentrated around the centre of that class.

3.9 GEOMETRIC MEAN


If a variable x takes the values x1, x2, x3,..., xn, then the nth root of the product of these n values is
called the geometric mean of the variable x and is denoted by G.
1/n
Thus, G = (x1 x2 x3 … xn)
Again if the frequencies of x1, x2, x3, …, xn are f1, f2, f3,…, fn respectively then the GM of x
will be defined as:
G = ( x1f1 x2f2 x3f3 … xnfn) 1/N, N = Sf

29
1.7.1 Advantages and disadvantages of G.M. :
Advantages:
¡ G.M. is rigidly defined.
¡ G.M. is based on all values of the distribution.
¡ It is possible to do further mathematical treatment in case of G.M.
¡ In comparison to A.M., G.M. is less affected by extreme values.
Disadvantages:
¡ G.M. is not that easy to determine and understand.
¡ G.M. of a distribution cannot be determined if there is even one negative value in the series.
Also, if there is at least one zero value, then the G.M. will be zero.

3.9.1 Uses of G.M.


¡ G.M. finds extensive use in averaging ratios, rates and percentages.
¡ As population increases in geometric progression, in determining the average rate of
increase, G.M. is used.
¡ G.M. is considered to be the top average in the construction of index numbers.

3.10 HARMONIC MEAN


Harmonic mean is the reciprocal of the A.M.s of the reciprocals of values in a distribution. If a
variable takes on the values of x1, x2, x3, …, xn, then the harmonic mean of x which is denoted
by H is given by:

Example: Determine HM of the following numbers:


46.1, 21, 127, 202
Solution:
Calculations for H.M.

30
The H.M. is given by:

= 4/ 0.0429 = 93.24.

3.10.1 Advantages and disadvantages of H.M.:


Advantages:
¡ The H.M. of a distribution is based on all the observations.
¡ It is rigidly defined.
¡ It is capable of further mathematical treatment.
¡ It is suitable in case of those series that have a wide dispersion.

Disadvantages:
¡ It is difficult to calculate and understand.
¡ If even a single value in the distribution is zero, then the H.M. cannot be computed.
¡ It gives more importance to smaller values.

3.10.2 Uses of H.M.:


Harmonic mean is used most frequently in cases of finding out the average speed of an object that
has traversed equal distances in different times with different speeds. To find mean mileage of a
car that traverses equal distances with different mileage, the H.M. is obtained.

3.11 RELATIONSHIP AMONG A.M. , G.M. AND H.M.


a. For any finite number of positive values, A.M ³ G.M. ³H.M.
b. For any two positive numbers, A.M. x H.M. = (G.M.)2

Note:
¡ Although A.M. is for all values, i.e., positive, negative and zero, G.M. is defined for positive
values only and H.M. is defined for non-zero values.
¡ If x1 = x2, then A.M. = G.M. = H.M. If x1 x2, then A.M > G.M. > H.M.
¡ The above property holds true for any finite number of positive values.

3.12 MEDIAN
Median is the second type of average that is used. Median is the middle most value in a
distribution when the said distribution is arranged either in ascending or descending order. It

31
divides the distribution into two equal parts. Thus there are equal number of observations on the
right and left of the median value, i.e, the number of observations greater than and less than the
median are equal.
For determining the median of an individual series, we have to make sure that all the values in a
distribution are arranged in a definite order, i.e, whether those values are in ascending or
descending order or not. If the values are not in a definite order, then these values have to be
arranged in either an ascending or descending order. For a distribution with odd number of terms
(value), the median is the middlemost value. If there are even number of terms in the array, then
the median is the average of the two middle numbers.
Symbolically, for odd number of terms in the distribution,
Median = ((n+1)/2)th value from the beginning or the end
For even number of terms in the distribution,
Median = average of the n/2th value and the ( n/2 +1 )th value
Example: Determine median for the following series:
i. 77, 73, 72, 70, 75, 79, 78
ii 94, 33, 86, 68, 32, 80, 48, 70
Solution:
i. Arranging the values of the series in ascending order, we get
70, 72, 73, 75, 77, 78, 79
No. of terms in the series = 7
The required median = (7+1)/2 = 4th term = 75.

ii. Arranging the series in ascending order, we get


32, 33, 48, 68, 70, 80, 86, 94
No. of terms in the series = 8 = even number
The required median = average of the n/2th value and the ( n/2 +1 )th value
n/2th value = 68, ( n/2 +1 )th value = 70
The required median = (68+70)/2 = 4th term = 69.

Note: By arranging the terms in descending order, the same value of median will be obtained.

3.12.1 Median of an ungrouped frequency distribution:


To determine the median of an ungrouped frequency distribution, we first have to arrange the
distribution in a definite order and then form a cumulative frequency table. If the number of
observations is odd, then the (N+1)/2 th (N is the total frequency) tern will be the median. If it is
so that (N+1)/2 is greater than a term x but less than or equal to another term y, where both x and
y are two consecutive values in the cumulative frequency column, then the observation whose
cumulative frequency is y shall be the median. However, if there are even number of terms, then
the A.M. (average) of the N/2 th and the (N/2 +1 ) th terms will be the median.

32
3.12.2 Median of a grouped frequency distribution:
While computing the median of a grouped frequency distribution, the cumulative frequencies of
the various class intervals has to be found out first. Then we have find out the median class. The
median class is the class which contains the median value. We find out the median value by
applying the same principles we did in case of the ungrouped frequency distribution (for odd
number terms, N/2th term is median and for even number of terms, average of the N/2 and [N/2
+1] the term shall be the median). After detecting the median class, the particular median value is
determined by using the following formula:

Where, L = Lower class limit (lower class boundary)


f = Frequency i.e., simple frequency of the median class
fc = Cumulative frequency of the class preceding the median class
N = Total frequency
I = Length of the median class (h maybe used in place of I)
Note: This formula for obtaining the median in case of a grouped frequency distribution holds
true only when the distribution is in ascending order. If the distribution is in a descending order, it
has to be arranged in an ascending order first, in order to apply the above formula.

3.12.3 Advantages and Limitations of Median:


Advantages:
¡ The median is not affected by extreme values.
¡ It is easily determined and understood.
¡ Median can be determined graphically.
¡ Medians of individual distributions and ungrouped frequency distributions can be
determined by mere observations in most cases.

Disadvantages:
¡ Unlike the other measures of central tendency, determination of median requires the
distribution to be arranged in a definite order if it is not in any order.
¡ Median is not based on all observations of the distribution.
¡ In comparison to mean, it is more affected by fluctuations in sampling.

Use of median:
In order to determine the average in case of distributions having open-end class intervals, median
is the best measure of central tendency. In case of income distribution, median would yield better
results.

33
Example: Determine median for the following distribution:

Daily 50-55 55-60 60-65 65-70 70-75 75-80 80-85


wages
(Rs)
No of 6 10 22 30 16 12 15
workers

Solution:
Table for determining median
Weekly wages No. of workers (f) Cumulative frequency ( fc)
50-55 6 6
55-60 10 16
60-65 22 38
65-70 30 68
70-75 16 84
75-80 12 96
80-85 15 111
N = 111

Since no. of classes is 7 (odd), thus median will be ( (N+1)/2)th term. Thus Median will be
(111+1)/2 = 56th term. From the cumulative frequency table, we find that the 56th term lies in the
class 60-70. Therefore 60-70 is the median class.

3.13 MODE
The mode of a distribution is that value which occurs the most frequently in the distribution. It is
that distribution whose frequency is the maximum. It should be noted that mode is not unique
which means that a distribution may have more than one mode. Distributions that have more than
one mode are called bimodal distributions and those that have more than two modes are termed
as multimodal.
Thus, an individual distribution does not have a mode. Even in case of a discrete frequency
distribution, each observation has the same frequency and thus has no mode. In case of an
ungrouped frequency distribution, mode can be determined by observation, in most cases. In
case of a grouped frequency distribution, mode is obtained by using the following formula:

34
f1 = frequency of the modal class
f0 = frequency of the class preceding the modal class
f2 = frequency of the class succeeding the modal class
I = length of the modal class. (Instead of I, the symbol h may also be used.)
Note:
¡ The class (specified by a class interval) whose frequency is the maximum is called the
modal class.
¡ The above formula is applicable when all the classes are of equal length.

Example:
Determine the mode/modes of the following series, if any
i. 3, 4, 5, 2, 3, 4, 1, 6, 4;
ii. 7, 9, 11, 7, 6, 5, 9, 13;
iii. 3, 5, 6, 7, 9, 12, 3, 6, 5, 9, 12, 7
Solution:
i. The number 4 is repeated the maximum number of times (3 times). Hence the mode of the
distribution is 4.
ii. Here we see that both the numbers 7 and 9, appear twice in the distribution. Since the
frequency of these numbers is the highest (2), thus 7 and 9 are the two modes of these
series.
iii. In this series, the frequency of each observation is the same and hence this series has no
mode.

3.13.1 Advantages and Limitations of Mode:


Advantages:
¡ The mode of an ungrouped frequency distribution can be determined by observation alone.
¡ Mode is not affected by extreme values.
¡ It can be determined graphically.
¡ It is also easy to understand.

Disadvantages:
¡ Mode is not based on all observations.
¡ Unlike the other averages, it is not capable of further mathematical treatment.

3.13.2 Use of Mode


The usefulness of mode is found in industries and business. A shoe maker can make use of mode
by being interested in the modal size of shoes and manufacturing them in larger quantities.

35
Weather forecasts are also based on mode. The mode is an appropriate measure of central
tendency for nominal-level data.

3.14 RELATIONSHIP BETWEEN MEAN, MEDIAN AND MODE


An experimental or empirical relationship between mean, median and mode has been established
by Prof. Karl Pearson. For any distribution the following relationship approximately holds:

Mean - Mode = 3 (Mean - Median)

3.15 QUARTILES, DECILES AND PERCENTILES


We have learnt so far that mean, median and mode are the measures of central tendency. In
addition to this, there are few other measures of central tendency which have been found to be
similar to the median and hence these too are studied along with the measures of central
tendency. These three measures are Quartiles, Deciles and Percentiles. They indicate quantities
at some specific places of distribution.

Quartiles
Quartiles are those measures of central tendency which divide the distribution into four equal
parts when the distribution is arranged in an ascending order. There are three quartiles in a
distribution namely Q1, Q2 and Q3.

Deciles
The nine quantities that divide a distribution into ten equal parts are called the deciles of the
distribution. The distribution needs to be arranged in an ascending order. These are denoted by
D1, D2, D3, …, D9.

Percentiles
Percentiles divide a distribution into hundred equal parts. There are 99 percentiles because it
takes 99 dividers to separate a group of data into 100 parts. The nth percentile is the value such
that at least n percent of the data are below that value and at most (100 - n) percent are above that
value.

SELF ASSESSMENT Questions


Long Answer Questions
1. What do you mean by measures of central tendency? Discuss briefly the methods of
measuring averages.
2. What are the different types of mean? Define each of them and state their relative merits,
demerits.
3. Define median and mode. Explain how these measures are calculated in case of grouped
and ungrouped data.

36
4. Find the missing frequency if the arithmetic mean is Rs 33 thousand.
Loss of sales (Rs in thousand) 0-10 10-20 20-30 30-40 40-50 50-60
No. of families 10 15 30 - 25 20

5. Calculate A.M. and median of the distribution. Hence calculate mode using empirical
relation between the three.
Class intervals 59-61 61-63 63-65 65-67 67-69
Frequency 4 30 45 15 6

6. The following list shows the 15 largest banks in the world by assets according to Standard
and Poor's. Compute the median and the mean assets from this group. Which of these two
measures do you think is most appropriate for summarizing these data, and why? What is the
value of Q2? Determine the 63rd percentile for the data. How could such information on
percentiles potentially help banking decision-makers?
Bank Assets (US$ millions)
Industrial & Commercial Bank of China 4,009
China Construction Bank Corp. 3,400
Agricultural Bank of China 3,236
Bank of China 2,992
Mitsubishi UFJ Financial group 2,785
JP Morgan Chase & Co. 2,534
HSBC Holdings 2,522
BNP Paribas 2,357
Bank of America 2,281
Credit Agricole 2,117
Wells Fargo & Co. 1,952
Japan Post Bank 1,874
Citigroup Inc. 1,842
Sumitomo Mitsui Financial Group 1,175
Deutsche Bank 1,166

Short Answer Questions


1. Mention the desirable characteristics of a good average.
2. Define quartiles, deciles and percentiles.
3. Explain composite mean.
4. State the empirical relation between mean, median and mode.
5. The average weight of the following distribution is 58.5 kg.

Weight (kg) 50 55 60 x + 12.5 70 Total

No. of men 1 4 2 2 1 10

Find the value of x.


37
6. Suppose the average salaries of elementary school teachers in three cities A, B and C are Rs
13,300, Rs 14,500 and Rs 21,000. Given that there are 13,000, 17,200 and 2,400 elementary
school teachers in these cities, find the average salary of all these elementary school teachers in
these cities.

Fill in the blanks


1. A.M. is very much affected by _______________ (Extreme values)
2. An average___________the given data. (Summarises)
3. For an open ended distribution, _________cannot be determined. (Mean)
4. The arithmetic mean of -1, 0 and 1 is __________ (0)
5. A distribution which has two modes is called ____________ (Bi modal)
6. Median is more suited average for grouped data with ___________classes. (Open end)

True and False


1. Mean is one of the measure of central tendency. (True)
2. Median can be computed without arranging the distribution in any definite order. (True)
3. A distribution can only have one mode. (False)
4. Geometric Mean (G.M.) is used to calculate the average speed of an object. (False)
5. For any finite distribution, A.M. ³ G.M. ³ H.M. (True)
6. Mean usually implies A.M. (True)

Multiple Choice Questions


1. Which of the following represents median?
a. First quartile
b. Fourth decile
c. Second quartile
d. None of the above

2. Which of the following relations among the location parameters does not hold?
a. Q2 = median
b. P50 = median
c. D5 = median
d. D6 = median

38
3. Extreme values have no effect on:
a. A.M.
b. Median
c. G.M.
d. H.M.

4. If the average of 7, 9, 12, x, 5, 4 and 11 is 9 then x is:


a. 13
b. 14
c. 15
d. 8

5. The mean of 8 numbers is 15. After a new number 24 is added, the new mean shall be:
a. 8
b. 16
c. 12
d. 10

6. The mode of the distribution of values 5, 9, 7, 7, 5, 9, 6, 7, 5, 4, 3, 4, 1, 5 is


a. 5
b. 9
c. 7
d. 3

Problem solving activities


A research agency administers a demographic survey to 90 telemarketing companies to
determine the size of their operations. When asked to report how many employees now work in
their telemarketing operation, the companies gave responses ranging from 1 to 100. The agency's
analyst organizes the figures into a frequency distribution.

Number of employees working in Telemarketing Number of companies


0- Under 20 32
20- under 40 16

40- under 60 13
60 – under 80 10
80 – under 100 19

Compute the mean, median, and mode for this distribution.

39
Suggested reading:
¡ Chandan, J.S. Statistics for Business and Economics. New Delhi: Vikas Publishing House
Pvt Ltd., 1998
¡ Gupta, S.C. Fundamentals of Statistics. New Delhi: Himalaya Publoshing House, 2006.
¡ Kothari, C.R. Quantitative Technique. New Delhi: Vikas Publishing House Pvt. Ltd., 1984
¡ Black, K. Business Statistics: Contemporary Decisions Making. Wiley, 2009.

References
¡ Black, K. (2009, December 1). Business Statistics: Contemporary Decision Making (6th
ed.). Wiley.
¡ Padmalochan, H., & Hazarika, P. (ca. 2007, April 4). A Textbook of Business Statistics (1st
ed.) [Print]. S. Chand Limited.
¡ Peck, R., Olsen, C., & Devore, J. (2022, September 30). Introduction to Statistics and Data
Analysis (AP(R) Edition) (4th ed.). Brooks/Cole, Cengage Learning.
¡ Quantitative Techniques (New Format). (2013, January 1). Vikas Publishing House.

40
MODULE - 4 MEASURES OF DISPERSION

STRUCTURE
¡ Measures of dispersion
¡ Different measures of dispersion
¡ Absolute and Relative measures of dispersion
¡ Different measures of dispersion
¡ Range
¡ Interquartile range and Quartile Deviation (Q.D.)
¡ Co-efficient of Quartile Deviation (Q.D.)
¡ Mean Deviation (M.D.)
¡ Standard Deviation
¡ Empirical relationship between Q.D, M.D. and S.D.
¡ Coefficient of Variation

4.1 LEARNING OBJECTIVES


After going through this unit, you will be able to:
¡ Understand what measures of dispersion are
¡ Learn the different measures of dispersion
¡ Explain the difference between absolute and relative measures of dispersion

4.2 MEASURES OF DISPERSION


While we know that the measures of central tendency provide a representative value of a single
set of data. They determine the middle most values of a distribution or that value which occurs
most frequently in the distribution. However, sometimes these measures of central tendency are
not fully representative of a given set of data. This happens in the case of distributions where the
extent of variation in individual values in relation to the average, or in relation to the other values
is large. As an illustration, let us observe the following three series:
Series A 40 40 40 40 40
Series B 35 39 41 42 43
Series C 10 18 35 57 80

We observe that in the first series, all the values are 40. Thus, the A.M. which is 40, fully
represents the series as well as the individual items. In the second series, although the mean is 40,
the values are not scattered much as the minimum value in the series is 35 and the maximum is 43.
Thus, for the second series too, we can say that the mean is a good representative of the series.

41
Here, the discrepancy between the mean and other values isn't very high.
In case of the third series, we observe that all the values are very different. Though the mean is
same like the other two series, i.e., 40, the values are widely scattered, with 10 being the
minimum value and 80 being the maximum. Clearly, in this case, the mean neither satisfactorily
represents the series generally, nor the individual values of the series in particular.
Thus, measures of central tendency lose their effectiveness and cannot be the representative
when the extent of variation (dispersion or scatteredness) of the individual values of a
distribution in relation to their average or in relation to the other values becomes large. Hence it is
important for a statistician to not only know the average of any type, but also the scatteredness of
a distribution.
Scatteredness of data about an average is termed as dispersion or variation. Quoting Spiegel,
"The degree to which numerical data tend to spread about an average value is called variation or
dispersion."
The study of average alone without knowledge of dispersion may lead to an erroneous
conclusion. A person with the knowledge of average and without the knowledge of variability
was once travelling with his family where he had to cross a river without a boat to reach the
destination. He knew that the average depth of the river was 100cm and the average height of his
family was 130cm. So he and his family decided to swim across the river on foot. But it so
happened that the maximum depth of the river was 150cm and the height of his youngest son was
60cm. We can only fathom what must have happened to the family.
Depending upon the nature of a distribution, different methods of obtaining the representative
value have been evolved and as a result, we have several averages or measures of central
tendency.

4.2.1 Objectives of measuring variability


Following are the objectives of studying variability:
¡ To study the reliability of average: The study of variability helps in assessing the
reliability of the average by determining the extent to which the data under study are
homogenous
¡ To control variation: The study of variation is done to see if the variation among data is
significant and if it so then such study may help in suggesting measures to control the
variation.
¡ To make comparison among series: Measures of variability are useful in comparing two
or more series with regards to disparity or differences.
¡ To make further statistical analysis: Standard deviation which is a measure of variability
is useful for study of higher measures such as skewness, kurtosis, regression, correlation,
etc.

4.2.2 Absolute and Relative measures of dispersion


The various types of measures of dispersion can be divided into a) Absolute measures b) Relative
measures
a. Absolute measures: Absolute measures of dispersion are those which are expressed in the
same units of the distributions for which these measures are obtained. For instance, if the

42
original distribution is in kilograms, an absolute measure will also be in kilograms. For this
reason an absolute measure of dispersion cannot be used to compare the variability
between two or more distributions.
b. Relative measures: A relative measure of dispersion is one which is calculated as a
percentage or coefficient of an absolute measure. A relative measure of dispersion is also
known as the coefficient of dispersion.
A relative measure of dispersion is free from any unit.

4.3 DIFFERENT MEASURES OF DISPERSION


There are two types of measures of dispersion, namely a) Absolute measures and b) Relative
measures.
i. Absolute measures: There are four types of absolute measures of dispersion:
¡ Range
¡ Interquartile Range and Quartile Deviation (Q.D.)
¡ Mean Deviation (M.D.)
¡ Standard Deviation (S.D.) and variance.
ii. Relative measures: These are the following:
¡ Coefficient of Quartile deviation
¡ Coefficient of mean deviation
¡ Coefficient of standard deviation
¡ Coefficient of variation

4.5 RANGE
The range in a distribution is the difference between the smallest and the largest value of that
distribution. Thus if L denotes the largest observation and S denotes the smallest observation,
then Range R = L - S.

4.5.1 Advantages and Limitation of Range:


Range is easy to be calculated and understood. Range is used in quality assurance, where the
range is used to prepare control charts. However, the disadvantages arise in the form of range
being dependent only on the two extreme values. This may, very often, lead us to a wrong
conclusion. Also, range cannot be computed in case of one or both the first and last class intervals
being open.
Example: Determine range for the following distribution:
Weight (kgs) 40 47 56 62 70
No. of students 4 7 11 3 1
Solution: Here, L = 70, S = 40
Range = L - S = 70 - 40 = 30kg.

43
4.6 INTERQUARTILE RANGE AND QUARTILE DEVIATION (Q.D.)
Interquartile range is the difference between the third quartile (Q3) and the first quartile, Q1 of
the distribution.
Quartile deviation, Q.D. is half of the interquartile range of a distribution.
Thus, Interquartile range = Q3- Q1
Quartile deviation = Q3 - Q1 ) / 2

4.6.1 Advantages:
¡ Q.D. is a better measure of dispersion than range as unlike range that takes into account
only two of the values, Q.D. involves 50% of the values.
¡ Q.D. is not affected by extreme values as the lowest 25% and highest 25% of the
observations are not considered while calculating the Q.D.
¡ It is the only measure which can be used for open ended class intervals.

4.6.2 Disadvantages:
¡ Since Q.D. is based on only 50% of the observations, it disregards the other half of the
observations.
¡ It is not amenable for further mathematical treatment.
Interquartile range is specifically useful for those data users who are more interested in values
towards the middle and less interested in the extremes.

4.6.3 Co-efficient of Quartile Deviation (Q.D.)


Co-efficient of Q.D. is a relative measure of dispersion and is defined as follows:

4.7 MEAN DEVIATION (M.D.)


Mean deviation is the third type of measure of dispersion and is the arithmetic mean (A.M.) of the
absolute deviations of the observations in a distribution from the average (usually mean or
median). For applying mean deviation, variance and standard deviation, the data should be at
least an interval level data.

44
Here A = mean or median of x and d = x-A
Note:
i. In case of mean deviation from mean, 'A' may be AM, GM or HM but is usually taken as
AM.
ii. |x-A| is called as the absolute value of the deviation x - A. By absolute value we mean the
magnitude of the value without considering the sign.
iii. Since the sum of the deviations measured from the bus is zero, hence in case of mean
deviation we always take absolute deviations.

4.7.1 Advantages and disadvantages of Mean Deviation:


Advantages:
¡ The M.D of a distribution is based on all the observations.
¡ It is less affected by extreme values.
¡ Since deviations are taken from average (mean, median, mode), thus, mean deviation is
considered to be a good measure for comparing the variability among two or more
distributions.

Disadvantages:
¡ In case of M.D. absolute values are taken and the actual signs of deviations are discarded.
¡ Mean deviation from mode is not considered to be a good measure of dispersion.
¡ For a grouped frequency distribution containing open- end class intervals, one cannot
determine mean deviation.

4.7.2 Coefficient of Mean Deviation:


While mean deviation is an absolute measure of dispersion, coefficient of mean deviation is a
relative measure of dispersion. The formula is given below:
Coefficient of mean deviation = (Mean deviation)/(The average from which mean deviation is
taken)
Thus,

45
Example: For the following distribution determine the mean deviation (M.D.) from mean and its
coefficient.

4.8 STANDARD DEVIATION


Standard deviation is a popular measure of variability. It is used as an independent measure of
analysis as well as used as a part of other analyses, such as computing confidence intervals and in
hypothesis testing. It is the positive square root of the arithmetic mean of the squares of the
deviations of the values of a variable from its arithmetic mean.
If the variable x takes n values x1, x2,… , xn and if x ? be the arithmetic mean of these values, then

4.8.1 Advantages and limitations of S.D.


Advantages:
¡ Standard deviation is considered to be the best measure of dispersion.
¡ It is rigidly defined and is based on all observations.
¡ Standard deviation possesses the highest sampling stability as compared to the other
measures of dispersion.
¡ The main disadvantage that is found in case of mean deviation, i.e., it disregards the
algebraic signs of deviations, standard deviation faces no such limitation.
¡ The normal curve can also be analysed with the help of S.D.

46
¡ It is standard deviation which can be considered as the basis of sampling theory and
correlation analysis.
Limitations:
¡ The biggest disadvantage is that S.D. is difficult to calculate in comparison to the other
measures of variability.
¡ Also, as compared to M.D., it is more affected by extreme values.
Example: Find standard deviation of the following observation:
8, 10, 12, 14, 16, 18, 20, 22, 24, 26
Solution:

4.8 VARIANCE
Variance is the square of standard deviation. Thus, we can say that standard deviation is the
positive square root of variance. Thus, variance = 2 where implies standard deviation and

4.9 EMPIRICAL RELATIONSHIP BETWEEN Q.D, M.D. AND S.D.


It can be verified empirically that for a distribution the following relationship among Q.D., M.D.,
and S.D. holds true. However, in case of a normal distribution, this relationship is exact.

47
4.10 COEFFICIENT OF VARIATION
Coefficient of variation is a relative measure of dispersion developed by Professor Karl Pearson.
Abbreviated as C.V., it is useful in comparing the variability of two or more sets of data especially
if they are expressed in different units of measurement. C.V. is expressed as a percentage. The
formula for C.V. is given as follows:

The coefficient of variation is the most popular relative measure of dispersion. While comparing
the variability between two distributions, the distribution having the minimum C.V., is
considered to less variable, more stable, more consistent, uniform or more homogenous. On the
other hand, the distribution which has higher C.V. is said to be more variable, less stable, less
consistent or homogenous.
Coefficient of variation can be sued in another situation, say, for comparing the relative
consistency of the prices of shares of two companies. It will help an investor to decide which
company's share prices are relatively stable and he can choose to invest in the same. The shares
which are more stable or consistent in the fluctuation of prices will be preferred by the investor.
As an example, suppose the scores of a cricket player A in three matches are 0, 10 and 80. The
scores of another cricket player B in these matches are 28, 30 and 32. Although the mean scores
of both the players are same, both being 30, we can say that B is more consistent than A. even
people without having the knowledge of variability or dispersion will say this. If we calculate the
C.V. for both the series of scores, we shall find that C.V. score of B is less than the C.V. score of A.
In fact, the magnitude of any measure of dispersion will be less in case of scores attained by B
than the scores attained by A.

SELF ASSESSMENT Questions


Long Answer Questions
1. What do you mean by dispersion? Mention the various measures of dispersion.
2. What are absolute and relative measures of dispersion? Write in what situation relative
measures are used.
3. Define standard deviation and explain why standard deviation is most widely used in
statistical studies as a measure of variation.
4. Why do you need to study both measures of central tendency and measures of dispersion
together?
5. Determine QD and coeff. of QD for the following distribution:

Height (inches) 60-62 62-64 64-66 66-68 68-70 70-72


No. of boys 5 18 42 20 8 7

6. The following frequency distribution gives the height (in inches) of 100students selected at
random from a college having 3000 students. Calculate standard deviation.
Weight (kg) 30-34 35-39 40-44 45-49 50-54
No. of boys 5 11 26 10 8

48
Short Answer Questions
1. Discuss any two objectives of measuring variability.
2. What is variance?
3. What do you meant by coefficient of variation?
4. Explain in what situations range serves as a useful measure of variability.
5. What is the empirical relationship between QD, MD and SD?
6. For a distribution, the coefficient of variation is 22.5% and the value of arithmetic average
is 7.5. Find out the value of standard deviation.

Fill in the blanks


1. Measures of variability tells how much the data in a distribution is ________ (Scattered)
2. The difference between the smallest and largest observation in the distribution is termed
______ (Range)
3. There are mainly___________types of measures of variability. (Two)
4. Relative measures of dispersion are __________ (Unitless)
5. The relationship between MD and SD is _________ (5 M.D = 4 S.D)
6. The numerical value of a standard deviation can never be __________ (Negative)

True and False


1. The degree to which numerical data tend to spread about an average value is called the
variation or dispersion of data. (True)
2. Standard deviation as a measure of dispersion, does not consider all observations in the set.
(False)
3. The range only considers two values, the highest and lowest of the data set.(True)
4. Relative measure of dispersion is presented in the same way in which the unit of
distribution is given. (False)
5. The absolute measure of dispersion for range is called the coefficient of range. (True)
6. Interquartile range is the difference between the third quartile and the first quartile (True)

Multiple Choice Questions


1. Which of the following are methods under measures of dispersion?
a. Standard deviation
b. Mean deviation
c. Range
d. All of the above

49
2. Which of the following are characteristics of a good measure of dispersion?
a. It should be easy to calculate
b. It should be based on all the observations within a series
c. It should not be affected by the fluctuations within the sampling
d. All of the above

3. The coefficient of variation is a percentage expression for


a. Standard deviation
b. Quartile deviation
c. Mean deviation
d. None of the above

4. While calculating the standard deviation, the deviations are only taken from
a. The mode value of a series
b. The median value of a series
c. The quartile value of a series
d. The mean value of a series

5. Which of the following cannot be calculated for open-ended distributions?


a. Standard deviation
b. Mean deviation
c. Range
d. None of the above

6. If all the observations within a series are multiplied by five, then


a. The new standard deviation would be decreased by five
b. The new standard deviation would be increased by five
c. The new standard deviation would be half of the previous standard deviation
d. The new standard deviation would be multiplied by five

50
Match the following
Column A Column B
1 Measures of dispersion A Open ended class interval
2 Absolute measure B Scatteredness
3 Coefficient of variation C Less consistency
4 Quartile deviation D Karl Pearson
5 Easiest measure of dispersion E Mean deviation
6 Variability F Range
(Answer Key: 1 = B, 2 = E, 3 = D, 4 = A, 5 = F, 6 = C)

Problem solving activities


Shown below are the top food and drug stores in the United States in a recent year according to
Fortune magazine.
Company Revenues ($ billions)
Kroger 66.11
Walgreen 47.41
CVS/Caremark 43.81
Safeway 40.19
Publix Super Markets 21.82
Supervalu 19.86
Rite Aid 17.27
Winn-Dixie Stores 7.88

Assume that the data represent a population.


1. Find the range.
2. Find the population variance.
3. Find the population standard deviation.
4. Find the coefficient of variation.

51
Suggested reading:
¡ Chandan, J.S. Statistics for Business and Economics. New Delhi: Vikas Publishing House
Pvt Ltd., 1998
¡ Gupta, S.C. Fundamentals of Statistics. New Delhi: Himalaya Publishing House, 2006.
¡ Kothari, C.R. Quantitative Technique. New Delhi: Vikas Publishing House Pvt. Ltd., 1984
¡ Black, K. Business Statistics: Contemporary Decisions Making. Wiley, 2009.

References
¡ Black, K. (2009, December 1). Business Statistics: Contemporary Decision Making (6th
ed.). Wiley.
¡ Padmalochan, H., & Hazarika, P. (ca. 2007, April 4). A Textbook of Business Statistics (1st
ed.) [Print]. S. Chand Limited.
¡ Peck, R., Olsen, C., & Devore, J. (2022, September 30). Introduction to Statistics and Data
Analysis (AP(R) Edition) (4th ed.). Brooks/Cole, Cengage Learning.
¡ Quantitative Techniques (New Format). (2013, January 1). Vikas Publishing House.

52
MODULE - 5 MEASURES OF DISPERSION

STRUCTURE
¡ Z score
¡ Skewness
¡ Kurtosis
¡ Empirical Rule
¡ Chebyshev's Theorem

LEARNING OBJECTIVES
After going through this unit, you will be able to:
¡ Understand a few other characteristics of a distribution
¡ Explain what z score is
¡ Learn the measures of shape of a distribution
¡ Understand the ways of applying standard deviation

5.1 Z-SCORE
A z score represents how far a data point is from the mean. How many standard deviations above
or below the population mean a raw value (x) is when the data is normally distributed is shown
by the z score. The use of z scores enables translating a value's raw distance from the mean into
units of standard deviations.

A negative z score implies that the raw value of x lies below the mean whereas a positive z score
implies the contrary, i.e., the raw value (x) is above the mean. For example, for a data set that is
normally distributed with a mean of 50 and a standard deviation of 10, suppose a statistician
wants to determine the z score for a value of 70. This value (x = 70) is 20 units above the mean, so
the z value is

Here, the z score signifies that the raw value of 70 is two standard deviations above the mean. The
z score is interpreted using the empirical rule which states that 95% of all values in a distribution

53
are within two standard deviations of the mean if the data set follows a normal distribution.
As a z score is the numerical score of standard deviations an individual data point is from the
mean, the empirical rule can also be stated in terms of z score.
Between z = -1.00 and z = +1.00 are approximately 68% of the values.
Between z = -2.00 and z = +2.00 are approximately 95% of the values.
Between z = -3.00 and z = +3.00 are approximately 99.7% of the values.

5.2 MEASURES OF SHAPE


Measures of shape are tools that can be used to describe the shape of a distribution of data. In this
section, we examine two measures of shape, skewness and kurtosis.

5.3 SKEWNESS
The concept of skewness can be acquired from the concept of symmetricity. If a distribution of
data is such that the right half is a mirror image of the left half, such a distribution is said to be
symmetrical. Statistically, a distribution is said to be symmetric about it's mean (A.M.) if the
observations of the distribution equidistant from the A.M. have equal frequencies.
Let us consider the following distribution:

It can be observed that the values of x, which are equidistant from 6 are 4 and 8, 2 and 10. 4 and 8
have the same frequency 3, whereas 2 and 10 have the same frequency of 1.
A distribution may also be said to be symmetric about it's A.M. if the deviations of the values
from their A.M. are such that corresponding to each positive deviation there is a negative
deviation of equal magnitude.

A symmetric distribution
54
Coming to skewness, a distribution which is asymmetrical or lacks symmetry is called a skewed
distribution. Thus, in case of a skewed distribution, the magnitudes of the positive and negative
deviations of the values from the mean are unequal or do not balance.
The figure on the left represents a negatively skewed distribution whereas the figure on the right
represents a positively skewed distribution.

Now for a positively skewed distribution which has a longer tail towards the right, like the figure
on the right, A.M. Median Mode. Whereas for a negatively skewed distribution where the tail
lengthens towards the left, A.M. Median Mode.
For a symmetric distribution, A.M. = Median = Mode.

5.3.1 Measures of skewness


Since for a skewed distribution, A.M. > Median > Mode or A.M. < Median < Mode, hence the
difference between A.M. and Mode i.,e, AM- Mode can also be taken as a measure of skewness.
Statistician Karl Pearson is credited with developing at least two coefficients of skewness that
can be used to determine the degree of skewness in a distribution. In this regard, Karl Pearson has
regarded (A.M. - Mode) / (S.D.) as a measure of skewness. It is, in fact called Pearson's first
measure of skewness (Sk). Apart from this, we also know that (A.M. - Mode) is approximately
equal to 3 (A.M. - Median). Hence, (3 (A.M.-Median))/? is also taken as a measure of skewness
by Karl Pearson and is known as Pearson's second measure of skewness (Sk).
Thus we have,
Pearson's first measure of skewness (Sk) = (A.M. - Mode) / (S.D.)
Pearson's second measure of skewness (Sk) = (3 (A.M. - Median)) / s
Example: Given below are the A.M., the median and the S.D. of two distributions. Determine
which distribution is more skewed.
i. A.M. = 22 ; Median = 24 ; S.D. = 10
ii. A.M. = 22 ; Median = 25 ; S.D. = 12
Solution: Here Pearson's second measure of skewness will be applicable

55
Since the absolute value of skewness of (ii) is greater than the absolute value of (I), hence
distribution (ii) is more skewed.

5.4 KURTOSIS
Kurtosis is concerned with the flatness or peakedness of frequency curve - the graphical
representation of a frequency distribution.
From the standpoint of Kurtosis, the normal curve is termed as mesokurtic, i.e, of intermediate
peakedness. A curve which is more peaked than the normal curve is called leptokurtic and a curve
flatter than the normal curve is called platykurtic.

5.5 EMPIRICAL RULE


The empirical rule is an important rule of thumb that states the approximate percentage of values
that lie within the given number of standard deviations from the mean of a set of observations if
the data follows a normal observation. The empirical rule is used only for three numbers of
standard deviations: l , 2 , and 3 . The empirical rule usually applies as long as the distribution is
mound shaped, i.e, it follows a normal distribution.
The empirical rule:

56
Example: A company produces a lightweight valve that is specified to weigh 1365 grams.
Unfortunately, because of imperfections in the manufacturing process not all of the valves
produced weigh exactly 1365 grams. In fact, the weights of the valves produced are normally
distributed with a mean weight of 1365 grams and a standard deviation of 294 grams. Within
what range of weights would approximately 95% of the valve weights fall? Approximately 16%
of the weights would be more than what value? Approximately 0.15% of the weights would be
less than what value?
Solution: Because the valve weights are normally distributed, the empirical rule applies.
According to the empirical rule, approximately 95% of the weights should fall within m ± 2s
=1365 ± 2(294) = 1365 ± 588. Thus, approximately 95% should fall between 777 and 1953.
Approximately 68% of the weights should fall within m ± 1s, and 32% should fall outside this
interval. Because the normal distribution is symmetrical, approximately 16% should lie above
m± 1s = 1365 + 294 = 1659. Approximately 99.7% of the weights should fall within ± 3s, and
.3% should fall outside this interval. Half of these, or .15%, should lie below m - 3s = 1365 -
3(294) = 1365 - 882 = 483.

5.6 CHEBYSHEV'S THEOREM


Chebyshev's theorem comes handy when the distribution is not normally distributed and the
Empirical rule cannot be applied. The empirical rule applies only when data are known to be
approximately normally distributed. This theorem applies to all distributions regardless of their
shape and thus can be used even if the shape of the distribution is unknown. Though Chebyshev's
theorem, too, can be applied, in theory, to normal distributions, it is usually empirical rule that is
preferred in case of normal distributions. This theorem, unlike the empirical rule, is not a rule of
thumb and is given in a formula format which makes it's application more wide. Chebyshev's
2
theorem states that at least 1-1/ k values will fall within k standard deviations of the mean
regardless of the shape of the distribution. The formula is given as:
Within k standard deviations of the mean, m ± ks , lie at least
1- 1
k2
proportion of the values.
Assumption: k > 1
Example: In the computing industry the average age of professional employees tends to be
younger than in many other business professions. Suppose the average age of a professional
employed by a particular computer firm is 28 with a standard deviation of 6 years. A histogram of
professional employee ages with this firm reveals that the data are not normally distributed but
rather are amassed in the 20s and that few workers are over 40. Apply Chebyshev's theorem to
determine within what range of ages would at least 80% of the workers' ages fall.
Solution: Because the ages are not normally distributed, it is not appropriate to apply the
empirical rule; and therefore Chebyshev's theorem must be applied to answer the question.
2
Chebyshev's theorem states that at least 1 - 1 /k proportion of the values are within m ± ks .
Because 80% of the values are within this range,
Let 1 - 1/k2 = 0.80
Solving for k yields

57
2
.20 = 1/ k
2
=>k = 5.000
k = 2.24
Chebyshev's theorem says that at least .80 of the values are within ±2.24 of the mean. For m = 28
and s = 6, at least .80, or 80%, of the values are within 28 ±2.24(6) = 28 ± 13.4 years of age or
between 14.6 and 41.4 years old.

SELF ASSESSMENT QUESTIONS


Long Answer Questions
1. What do you mean by measures of shape?
2. Explain skewness. What are the two types of skewness distributions?
3. Explain the z score.
4. Write a note on Chebyshev's Theorem.
5. For a distribution, the mean is 4.75, median is 4.72 and the S.D. is 0.92. Find the coefficient
of skewness and comment on the nature and shape of the distribution.
6. Of a certain distribution, the Karl Pearson's coefficient of skewness is 0.32, S.D. is 6.5 and
mean is 29.6. Find the mode and the median of the distribution.
Height (inches) 60-62 62-64 64-66 66-68 68-70 70-72
No. of boys 5 18 42 20 8 7

Short Answer Questions


1. What is a mesokurtic distribution?
2. What is the position of A.M., Median and Mode when a distribution is negatively skewed?
3. What does the empirical rule signify?
4. Why do we need to study the measures of shape?
5. A local hotel offers ballroom dancing on Friday nights. A researcher observes the
customers and estimates their ages. Discuss the skewness of the distribution of ages if the
mean age is 51, the median age is 54, and the modal age is 59.
6. On a certain day the average closing price of a group of stocks on the New York Stock
Exchange is $35 (to the nearest dollar). If the median value is $33 and the mode is $21, is
the distribution of these stock prices skewed? If so, how?

Fill in the blanks


1. ______________ represents how far a data point is from the mean. (z score)
2. In a symmetric distribution, the right half is the _________ of the left half. (mirror image)
3. Kurtosis shows the___________of a distribution. (Peakedness)
4. Skewness is when the distribution is ____________ (Asymmetric)

58
5. For a symmetric distribution, _________ (A.M. = Median = Mode)
6. A curve flatter than the normal curve is called __________ (Platykurtic)

True and False


1. Measures of shape show the curve or peakedness of a distibution. (True)
2. For symmetric distribution, A.M.>Median>Mode. (False)
3. For a positively skewed distribution A.M.>Median>Mode. (True)
4. Mesokurtic is when the curve of the distribution is very peaked. (False)
5. Coefficient of skewness is given by Karl Pearson. (True)
6. Chebyshev's Theorem is used when the data is not normally distributed. (True)

Multiple Choice Questions


1. The values of mean, median and mode can be
a. Some time equal
b. Never equal
c. Always equal
d. None of these

2. If mean, median, and mode are all equal then distribution will be
a. Positive Skewed
b. Negative Skewed
c. Symmetrical
d. None of these

3. A symmetrical distribution has mean equal to 4. Its mode will be


a. Equal to 4
b. Less than 4
c. Greater than 4
d. Not equal to 4

4. The shape of symmetrical distribution is


a. U shaped
b. Bell Shaped
c. J Shaped
d. None of these

59
5. A curve whose tail is longer to the right is called
a. Negatively skewed
b. Positively skewed
c. Symmetrical
d. None of these option

6. For a positively skewed distribution


a. Mean > Mode
b. Mode > Mean
c. Median > Mean
d. None of these options

Suggested reading:
¡ Chandan, J.S. Statistics for Business and Economics. New Delhi: Vikas Publishing House
Pvt Ltd., 1998
¡ Gupta, S.C. Fundamentals of Statistics. New Delhi: Himalaya Publishing House, 2006.
¡ Kothari, C.R. Quantitative Technique. New Delhi: Vikas Publishing House Pvt. Ltd., 1984
¡ Black, K. Business Statistics: Contemporary Decisions Making. Wiley, 2009.

References
¡ Black, K. (2009, December 1). Business Statistics: Contemporary Decision Making (6th
ed.). Wiley.
¡ Padmalochan, H., & Hazarika, P. (ca. 2007, April 4). A Textbook of Business Statistics (1st
ed.) [Print]. S. Chand Limited.
¡ Peck, R., Olsen, C., & Devore, J. (2022, September 30). Introduction to Statistics and Data
Analysis (AP(R) Edition) (4th ed.). Brooks/Cole, Cengage Learning.
¡ Quantitative Techniques (New Format). (2013, January 1). Vikas Publishing House.

60
MODULE - 6 INTRODUCTION TO PROBABILITY

STRUCTURE
¡ Introduction
¡ Usefulness of the theory
¡ Definition of useful terms
¡ Formula
¡ Dependent and Independent events
¡ Approaches
¡ Classical or Mathematical Approach to Probability
¡ Relative frequency of Occurrence
¡ Subjective Probability
¡ Types of probabilities
¡ Conditional Probability
¡ Symbols associated with probability
¡ Addition and Multiplication Theorem
¡ Bayes' Theorem
¡ Questions and Exercises

LEARNING OBJECTIVES
After going through this unit, you will be able to:
¡ Explain what the meaning of probability is
¡ Understand the usefulness of probability
¡ Describe events and their role in probability
¡ Explain the different approaches to probability
¡ Explain marginal, join, union probabilities
¡ Explain conditional probability
¡ Solve probability problems using addition and multiplication rules
¡ Describe Bayes' Theorem Unit Contents

6.1 INTRODUCTION TO PROBABILITY THEORY


In our daily lives, we come across certain situations. Sometimes, the answer to those situations is
definite and known to us. At other times, the outcome of the situation is not known to us with

61
certainty.
Consider the following experiments:
a. Zinc is mixed with dilute sulphuric acid.
b. A uniform coin is tossed upwards.
c. A dice is rolled.
The result of experiment (a) is zinc sulphate and hydrogen produced from the chemical reaction
between zinc and sulphuric acid. However, for the experiment (b), the result, i.e., whether head
or tail will turn up, is uncertain. Experiment (c), too, might give any result between 1-6.
Situations such as (b) and (c) are termed as random experiments and the result of such an
experiment is called an event. The likelihood of a particular event occurring is defined as
Probability. It is the measure of certainty of an event. The numerical probability value lies
between 0 and 1, where 0 denotes the impossibility of the event occurring and 1 denotes
assurance of the same event happening.

6.2 USEFULNESS OF THE THEORY


Probability is indispensable in the branch of statistics. It is the basis of Inferential statistics-
where probability is used to obtain the results, in case, the characteristics of the population are
not known. Quoting Prof. Ya-Lin-Chou, "Statistics is the science of decision making with
calculated risks in the face of uncertainty." The use of probability has been extended to various
branches of Science, Arts and Commerce. In industries of insurance and gaming, probability is
used to determine the likelihood of certain events occurring for determining specific rates and
charges, payoffs, etc.

6.3 DEFINITION OF IMPORTANT TERMS


Random Experiment: Random experiment is an experiment that, when performed repeatedly,
under the same conditions, may result in any of the given possible outcomes associated with the
experiment. For example, when an unbiased die is rolled repeatedly, it may show any of the six
faces every time it is rolled.
Trial: The act of performing a random experiment is called a trial.
Event: An event is an outcome or a set of outcomes of a trial. For example, the outcome of getting
a 'head' while tossing a coin is called an event and is usually denoted in capital letters like A, B, C
etc.
Elementary event: Elementary event, also called a single event is a single possible outcome of
an experiment. The event of getting a 5 while rolling a die is an elementary event.
Composite event: Composite or compound event is the combination of one or more single
events in a random experiment. The event of getting a sum of 6 while rolling two dice
simultaneously can be termed as a composite event.
Favourable cases: The outcomes which indicate the occurrence of an event are termed as
favourable cases to the event. For example, the number of favourable cases to the event of getting
a 'queen card' when a card is drawn at random from a pack of full cards is 4, as there are 4 queen
cards in a full stack.
Equally likely events: Equally likely events are said to be two or more events associated with a

62
random experiment where each event has got an equal chance of occurring.
Mutually exclusive events: Two or more events are said to be mutually exclusive if one event
prevents the occurrence of the other event. While tossing a coin, the occurrence of either head or
tail is an example of mutually exclusive event.
Complementary event: The complementary to an event A, is denoted by A . All the events of
an experiment not in A are considered its complement. By formula, A  = 1 - A.
Exhaustive events: Exhaustive events are those which include all the possible outcomes of a
random experiment. If a coin is tossed, the two events, namely head (H) and tail (T) will be the set
of exhaustive events in this situation.
Sample space: The set of all possible outcomes of a random experiment is called the sample
space. When we roll a die, any of the six faces between 1-6 can turn up and hence the sample
space will be S= {1, 2, 3, 4, 5, 6}.

6.4 FORMULA FOR PROBABILITY


Probability is given by:
P(E) = Number of Favourable Outcomes/Number of total outcomes
P(E) = n(E)/n(S)
Here,
n(E) = Number of event favourable to event E
n(S) = Total number of outcomes

6.5 DEPENDENT AND INDEPENDENT EVENTS


Statistical dependence is a relation between different characteristics measured on the same units.
Two events are said to be independent if the occurrence or non-occurrence of one does not affect
the probability of occurrence or non-occurrence of the other.
¡ If a die is rolled twice, then the event of getting a 2 both times are independent.
¡ If a coin is tossed twice, the event of getting a head in the first toss and a head in the second
toss are independent events.
Two events are said to be dependent if the occurrence of one event affects the occurrence or non-
occurrence of the other event. In case of two dependent events A and B, B can occur only when A
is known to have occurred and vice versa.
For example, the probability of drawing a Queen from a pack of 52 cards is 4/52 or 1/13; if a
Queen can be drawn and has not been replaced in the pack, the probability of drawing a Queen
again is 3/51 or 1/17. Thus the occurrence of the first event has affected the probability of the
occurrence of the second event.

6.6 APPROACHES TO PROBABILITY


6.6.1 Classical or Mathematical Approach to Probability
This approach determines the probability of an event in the form of a ratio m/n, where n are the
equally likely, mutually exclusive and exhaustive events and if m of them are favourable to an
63
event A, then probability of event A is given by:
Symbolically, P(A) = m/n = Total no of cases favourable to A
Total no of all possible cases
Illustrative examples
Let us consider the events A, B and C with rolling of a die A= {1, 2, 3, 4, 5, 6}, B= {7}, C= {1, 5}
Event A's favourable cases = 1, 2, 3, 4, 5, 6 P(A)= (Total no of cases favourable to A/ Total no of
all possible cases)
No of favourable cases = 6 = 1
No of total outcomes 6
Therefore, Event A's favourable cases = 6/6 = 1
Event B's favourable cases= 0, as 7 is not on the face of a die
P(B)= 0/6= 0.
Coming to event C, favourable cases = 1, 5
P(c)= 2/6=1/3
Range of possible probabilities 0 £ P(E) ³ 1

6.6.2 Relative frequency of Occurrence


In this approach, the probability is measured by the number of times the event has happened in
the past divided by the total number of opportunities for the event to occur.
Symbolically, P(A) = (No of times an event has occurred / Total no of opportunities for the event
to occur)
Relative frequency is used when the theoretical probabilities of an event cannot be determined.
For example, in a soccer game, it is difficult to know exactly how likely is it that one team shall
score goals against another; we can never compute the theoretical probability of soccer events. In
such cases, we can use the relative frequency to estimate the theoretical probability.

Subjective Probability
This approach to probability is based on a measure of the feelings and insights of the person
determining the probability. It measures the degree of a person's belief or subjective assessment
that the event will occur. The subjective approach is used when the other probability approaches
cannot be used-examples of situations like electoral outcomes, results of cricket matches, etc. In
a business situation, it can be the likelihood of a successful product launch, investment outcomes
etc. Though it is not a mathematical or scientific approach, this method uses the information,
experience, knowledge, stored and processed in the human mind. Sometimes, it can be merely an
estimate. In other cases, human experience and judgement can yield accurate probabilities.

6.7 MARGINAL, UNION, JOINT PROBABILITIES


Marginal probability is given as P(E), where E is some event. This probability is computed by
dividing a subtotal by the whole. For example, the marginal probability is the probability that a

64
person owns an Audi car. This probability will be obtained by dividing the number of Audi
owners by the total number of car owners. Another example could be the probability of a person
who wears spectacles which can be obtained by dividing the number of spectacle wearers to the
total number of people.
Symbolically, Marginal probability is given by P(E)= n(E)/ n(S)
where n(E) = Number of event favourable to event E
n(S) = Total number of outcomes
Another type of probability is Union probability which denotes the union of two events. The
union probability of two events, A and B, is given by P(A È B), which is the probability that either
A will occur or B will occur or both A and B will occur. An example of union probability is the
probability that a person owns an Audi or a Ford. The condition for union probability is that the
person has to own at least an Audi or a Ford. Another situation could be the probability of
someone having red hair or a tattoo. In this union, all people who have red hair are included,
along with people who have a tattoo and all redheads who have a tattoo. In an organization, the
probability that an employee is a clerical worker or a male is union probability.
Joint probability is the third type of probability which is the intersection of events denoted as
P(AÇ B), for two events A and B. To qualify for a joint probability, both events must occur. An
example of joint probability will be the probability of a person owning both an Audi and a Ford.
Merely owning either an Audi or a Ford will not be sufficient for joint probability. The
probability of a person having red hair and a tattoo is an example of joint probability.

6.8 CONDITIONAL PROBABILITY


In certain cases, a manager might want to know the outcome of an event that has already occurred
and may want to know the chances of a second event occurring based upon the knowledge of the
earlier event. For example, let us assume that a new brand of deodorant is being introduced in the
market. Based on previous market research, the manufacturer has some idea about the chances of
its success. Now, he introduces the product in a few selected stores in certain areas before
marketing it nationally. A highly positive response from the test-market area will improve his
confidence about the success of his brand nationally. Accordingly, the manufacturer's
assessment of high probability of sales for his brand would be conditional upon positive response
from the test market.
Let A and B be two dependent events. The probability of occurrence of event B given that event A
has already occurred is called the conditional probability of B and is given as:

Note: If B does not depend on A, then the symbol used is P(B) and not P(B|A). In this case, the
probability of event B occurring is not dependent on occurrence of event A. P ( B|A) is read as
'probability of B given A' or 'probability B post A'.
As an example, let us suppose that we roll a die and it is known that the number that came up is
greater than 3. We want to find out the probability that the outcome is an even number greater
than 3.
Let Event A = even
And Event B = larger than 3

65
Then P( even| greater than 4) = [P (even and greater than 3)/ P(greater than 3)]
Or
P (A/B) = P (AB)/ P(B)
= (1/6) (2/6) = ½.

Figure 6.1: Types of probabilities

6.9 CERTAIN SYMBOLS ASSOCIATED WITH PROBABILITY


I. If A and B are two mutually exclusive events, then the symbol P(A+B) or P (AÈB) denotes
the probability of occurrence of either A or B. Id A and B are two non-mutually exclusive
events then the symbol P (A+B) or P ( AÈB) denotes the probability of occurrence of at
least one event, i.e., either only A or only B or both.
ii. The symbol P(AB) or P (A ? B) denotes the probability of simultaneous occurrence of both
the events A and B.
iii. P(AC) or P(A') denotes the probability of non-occurrence of the event A. It denotes the
complement of the event A and it contains all the outcomes of the associated random
experiment which are not in A.
iv. P(`A`B ) or P (`A Ç`B ) denotes the probability of non-occurrence of both A and B.
v. P (`AB) or P (`A ÇB) denotes the probability of non-occurrence of A and occurrence of B.
vi. P (A`B) or P (AÇ`B ) denotes the probability of occurrence of A and non-occurrence of B.

6.10 ADDITION AND MULTIPLICATION THEOREM OF PROBABILITY


Law of Addition
The general law of addition determines the probability of union of two events, A and B. If the two
events are mutually exclusive, then the probability that either of the two events will occur is
given by the sum of their probabilities. It is given by:

66
P (A or B) = P(A) + P (B)
However, if the two events are not mutually exclusive, then the probability of either event A or B
occurring is given by the probability that event A occurs plus the probability that event B occurs,
subtracted by the common probability of both A and B occurring.
Symbolically, it can be written as:
P( AðB) = P(A) + P(B) - P(AðB)

6.10. LAW OF MULTIPLICATION


The law of Multiplication is used to find out the joint probability of two events. Similar to the
conditions stated in the law of addition, if A and B are two mutually exclusive events, then the
joint probability of their occurrence is given by the product of their separate probabilities. It is
given as:

P( AB) = P(A) * P(B) or P(A) P(B)


To understand this better, if we toss a coin twice, then the probability that the first toss lands in a
head and the second toss lands in a tail is given by,
P( HT) = P(H) * P(T) = ½ * ½ = ¼
As against this, if A and B are not independent, which means that the occurrence of either event is
dependent on the occurrence of the other event, then the probability that they both will occur is
given by
P( AÇB) = P(A) . P (B|A) = P (B) . P (A|B)

6.11 BAYES' THEOREM


Bayes' Theorem or Rule is an extension of the law of Conditional Probability. It was developed
by Rev. Thomas Bayes. Thomas Bayes being a religious preacher himself, was motivated to
establish this theorem to prove the existence of God by looking at the world itself, which was
created by God. This theorem contributes to the statistical decision-making theory by relating
posteriori probability to a priori probability. This formula extends the use of
Conditional probabilities to allow revision of original or prior probabilities of outcomes of
events based upon observation and analysis of new information.
This theorem uses the conditional probability formula by describing the condition in terms of the
additional information, which leads to the revised probability of the outcome of an event.
Let us suppose that there are 60 students in our statistics class, out of which 40 are female
students and 20 are male students. Out of the 40 female, 20 are Indian students and the remaining
20 are foreign students. Out of the 20 male students, 15 are Indians and the other 5 are foreign
students, so that out of all the 60 students, 35 are Indians and 25 are foreigners. This data can be
presented in a tabular form as follows:

Indian Foreigner Total


Male 15 5 20
Female 20 20 40
Total 35 25 60

67
Based on this information, the probability that a student picked up at random will be female is
40/60 or 0.67 since there are 40 females out of 60 students. Now let us assume that we are given
additional information that the person picked up at random is Indian, then what is the probability
that this person is a female? This additional information will result in the revised probability or
posterior probability in that it is assigned to the outcome after the additional information is made
available.
Since we want to determine the revised probability of picking a female student at random,
provided we know that the student is an Indian, so let A1 be the event female, A2 be the event
male, and B be the event Indian. Thus, based on our knowledge of conditional probability, Bayes'
Theorem shall be as follows:
P(A1 |B) = {P(A1) P (B| A1)}/{ P(A1) P(B|A1) + P(A2) P (B|A2)}
In the above example, there are 2 basic events which are A1 (female) and A2 (male). However, if
there are n basic events, A, A2, …, An, then Bayes' Theorem can be generalized as,
P(A1|B) = P(A1) P(B|A1)}/{P(A1) P (B|A1) + P(A2) P(B|A2) + … + P(An) P(B|An)}
Solving the previous two events, P(A1|B) =
= {(40/60) (20/40)}/{(40/60) (20/40) + (20/60) (15/20)}
= 20/35 = 4/7 = 0.57
This example demonstrated clearly how the probability of picking up a female student was 0.67
however, after the additional information received that the student is a foreigner, the posterior
probability becomes 0.57.

SELF ASSESSMENT QUESTIONS


Long Answer Questions
1. Define probability. State the various approaches to probability.
2. Explain conditional probability.
3. State the additive and multiplicative law of probability.
4. In a single throw with two dice, find the chance of throwing (i) two four, (ii) four-five (i.e,
one die shows four and the other die shows five points) and (iii) doublets (i.e., both dice
show the same figure)
5. An urn contains 5 white and 7 red balls. Two successive draws if one ball are made. Find
the probability that the first drawn ball is white and the second drawn ball is red if (i) the
first drawn ball is replaced before the second draw, (ii) the first drawn ball is not replaced.
6. The probability that a contractor will get a plumbing contract is 2/3, and the probability that
he will not get an electric contract is 5/9. If the probability of getting any one contract is 4/5,
what is the probability that he will get both the contracts?

Short Answer Questions


1. State the use of probability in everyday life.
2. What is a mutually exclusive event?
3. What do you mean by the complementary event?

68
4. Explain Bayes' Theorem.
5. Explain subjective probability.
6. What do you mean by dependent events?

Fill in the blanks


1. An…………is an outcome or a set of outcomes of an activity or a result of a trial. (Event)
2. If A and B are mutually exclusive events, then P (AB) = ……………. (0)
3. If a card is drawn at random from a pack of cards, the probability of getting either a king or
queen is ……………. (2/13)
4. The probability of getting more than 2 points when a die is thrown is ………….. (2/3)
5. If events are…………. exclusive, then the occurrence of any one of the events prevents
any of the other events from occurring. ((Mutually exclusive)
6. ………….is applied when it is necessary to compute the probability if both events A and B
occur at the same time. (Multiplication rule)

True and False


1. Probability ranges from -1 to 1.(False)
2. If an event cannot take place, then the probability of its occurrence is 0. (True)
3. If events A and B are dependent, then the probability that they both will occur is the product
of their separate probabilities. (False)
4. Anchoring is the practice of assessing the probability of an event based on past experiences
and factoring in new information to assign a probability value.(True)
5. Bayes' theorem makes use of conditional probability formula where the condition can be
described in terms of the additional information which would result in the revised
probability of the outcome of an event. (True)
6. The probability that a die will show eight points when it is thrown is 1/6. (False)

MCQs
1. If A is an uncertain event, then
a. P(A)³0
b. 0£P(A)£1
c. 0<P(A) < 1
d. None of the above

2. If P(A)= P(B), then two events A and B are


a. Mutually exclusive
b. Exhaustive

69
c. Equally likely
d. Independent

3. If P(A?B) [or P(AB)] is zero then the two events A and B are:
a. Mutually exclusive
b. Exhaustive
c. Equally likely
d. Independent

4. If two events A and B, P(A?B) =1, then A and B are:


a. Mutually exclusive
b. Equally likely
c. Dependent
d. Exhaustive

5. If P(A/B) =P(A), then


a. B is independent of A
b. A is independent of B
c. B is dependent on A
d. Both (a) and (b)

6. If two events A and B are independent, then


a. They can be mutually exclusive
b. They cannot be mutually exclusive
c. They cannot be exhaustive
d. Both (b) and (c)

Problem solving activities


In a recent year, business failures in the United States numbered 83,384, according to Dun &
Bradstreet. The construction industry accounted for 10,867 of these business failures. The South
Atlantic states accounted for 8,010 of the business failures. Suppose that 1,258 of all business
failures were construction businesses located in the South Atlantic states. A failed business is
randomly selected from this list of business failures.
a. What is the probability that the business is located in the South Atlantic states?
b. What is the probability that the business is in the construction industry or located in the
South Atlantic states?

70
c. What is the probability that the business is in the construction industry if it is known that
the business is located in the South Atlantic states?
d. What is the probability that the business is located in the South Atlantic states if it is known
to be a construction business?
e. What is the probability that the business is not located in the South Atlantic states if it is
known that it is not a construction business?

Suggested reading:
¡ Chandan, J.S. Statistics for Business and Economics. New Delhi: Vikas Publishing House
Pvt Ltd., 1998
¡ Gupta, S.C. Fundamentals of Statistics. New Delhi: Himalaya Publishing House, 2006.
Kothari, C.R. Quantitative Technique. New Delhi: Vikas Publishing House Pvt. Ltd.,
1984
¡ Black, K. Business Statistics: Contemporary Decisions Making. Wiley, 2009.

References
¡ Black, K. (2009, December 1). Business Statistics: Contemporary Decision Making
(6th ed.). Wiley.
¡ Padmalochan, H., & Hazarika, P. (ca. 2007, April 4). A Textbook of Business Statistics (1st
ed.) [Print]. S. Chand Limited.
¡ Peck, R., Olsen, C., & Devore, J. (2022, September 30). Introduction to Statistics and Data
Analysis (AP(R) Edition) (4th ed.). Brooks/Cole, Cengage Learning.
¡ Quantitative Techniques (New Format). (2013, January 1). Vikas Publishing House.

71
DISCRETE PROBABILITY
MODULE - 7
DISTRIBUTION

STRUCTURE
¡ Introduction to a random variable
¡ Discrete Probability distribution
¡ Expected Value
¡ Variance
¡ Binomial Distribution
¡ Bernoulli Trials
¡ Additional properties of binomial distribution
¡ Hypergeometric Distribution
¡ Poisson Probability Distribution
¡ Questions and Exercises

7.1 LEARNING OBJECTIVES


After going through this unit, you will be able to:
¡ Explain the functionality of random variable
¡ Understand the difference between a discrete and continuous variable
¡ Describe the expected value
¡ List the discrete probability distributions
¡ Explain binomial, hypergeometric and Poisson distributions

7.2 RANDOM VARIABLE


The variable whose numerical values are associated with the outcomes of a random experiment
is called a random variable. Its value can be related to the sample space of a random experiment.
For example, while tossing two unbiased coins together, the total set of outcomes can be
represented as, S = {HH, HT, TH, TT}. We can find out a number associated with each sample
point. If we let X denote the number of heads in this example, then we arrive at the following:

Sample Point HH HT TH TT
X 2 1 1 0

Corresponding to the sample point HH, we get X=2, to each of HT and TH we get X=2 and for the
sample point TT, X=0.

72
(T, T)

0
(H, T)

(T, H)

(H, H) 2

Through this example, we learn that each sample point may be assigned a numerical value of X,
and more than one sample point can be given the same value. Yet another observation is that
variable X assumes each value with a definite probability, and the sum of all the probabilities
results in unity (1).

In the above example, P(X=2) + P(X=1) + P(X=1) + P(X=0)


= 2/4+1/4+1/4+0/4
= 4/4=1.

We can thereby say that, a random variable is a numerical valued function defined on a sample
space. A random variable can be divided into discrete and continuous random variables.
The result of a random experiment is associated with the random variable, and depending upon
different trials, the value changes and hence is termed a random variable. Say the number of girls
in a three-child family, the number of light bulbs that will turn out defective in a bundle of 150
bulbs, etc. Such a value cannot be predicted with certainty and thus is a random variable.
When a random variable takes on a finite number of values or a countably infinite number of
variables, such a variable is called a discrete random variable. In the majority of statistical
situations, discrete variables take non-negative whole numbers.
When a random variable takes on an uncountably infinite number of values, it is called a
continuous random variable. Such variables take on matters at every point in a given range; the
main distinction between discrete and continuous random is that the former involves counting,
while the latter involves measuring.
A discrete variable may take the form of events like
¡ number of broken eggs in a carton,
¡ the total number of items purchased,
¡ age of individuals
¡ number of defects in a batch of 60 items Continuous variables will include measurements
like

73
¡ Height and weight of individuals
¡ amount of time spent in the store
¡ measuring the time between customer arrivals at a retail outlet

7.3 DISCRETE PROBABILITY DISTRIBUTION


Let X be a discrete random variable and let P(x) = P(X=x) such that 0 £P(x) ³1 and SP(x)=1
summation being taken over the various values of the variable.
Then the function pi = P(X=xi) or P(x)= P(X=x) is called the probability mass function or simply
the probability function of the random variable X and all the set of all the possible ordered pairs
{x, P(x)}is termed as discrete probability distribution or the discrete theoretical distribution of
the random variable X. Each probability is the limiting relative frequency of occurrence of the
corresponding x value when the chance experiment is repeatedly performed. In a frequency
distribution the total frequency is distributed among different values of (classes) of the variable
and in probability distribution the total probability 1 is distributed among the various possible
values taken by the random variable.

Note: Random variables are also called variates and are denoted by capital letters X, Y…,
whereas their specific values are denoted respectively by small letters x, y…

Example
i) Discrete distribution of the number of power cuts in a day (Table 4.1)

Table 7.1: Number of power cuts in day

Number of power cuts Probability


0 .37
1 .31
2 .18
3 .09
4 .04

An executive is considering taking leave on a Friday and holding a Zoom meeting online for her
clients' presentation at home. She recognizes that the area of her residence is prone to power cuts
and is concerned about the possibility of the same happening on the said day. The above table
shows a discrete distribution that contains the number of power cuts that could occur on the day
of her presentation and the probability that each number of power cuts will occur. It is apparent
from the table that the most likely number of power cuts is 0 and, 1 with a probability of .37 and
.31, respectively.

74
ii) If two unbiased coins are thrown and if X denotes the number of heads turning up, then we
have
P (HH) = ¼, P(HT) = ¼, P(TH) = ¼, P(TT) = ¼ Then P(X=0) = P(TT) = ¼,
P(X=1) = P(HT+TH) = P(HT) + P(TH) = ¼ + ¼ ½
P(X=2) = P(HH) = ¼
Consequently, we have the following probability distribution of the discrete random variable X:
X=x 0 1 2
P(x) 1/4 1/2 1/4

iii) If three unbiased coins are thrown, then the probability distribution of the discrete random
variable X denoting the number of heads will be as follows

X=x 0 1 2 3
P(x) 1/8 2/8 3/8 1/8

iv) If X denotes the sum of points shown by a pair of dice when they are thrown, then P(X=x) =
P(x) can be displayed as:
X=x 2 3 4 5 6 7 8 9 10 11 12
P(x) 1/36 2/36 3/36 4/36 5/36 6/36 5/36 4/36 3/36 2/36 1/36

The probability values and the different values of X can be represented graphically to obtain the
probability histogram, probability polygon and probability curve.

Expected Value
In the above example, we may observe that the probability of 6/36, where X takes the value of
7, is the highest. But 7 is also the mean or average of the values of X. Thus, the mean or the
average value is the most probable or the most likely. This mean or average value of the random
variable X is the expected value of the random variable X.
This expected value of a discrete distribution is the long-run occurrence of averages. It is also the
mean of the discrete random variable. If a trial is held only once, then a discrete random variable
yield result only once. However, if the trial is repeated long enough, the average of the outcomes
is more likely to yield a mean or an expected value. The mean or expected value is computed by
first multiplying each possible x value by the probability of observing that value and then adding
up the resulting quantities.
Symbolically,
µ = E(x) = S [x . P(x)]
Where
E = long run average
x = an outcome
P(x) = probability of that outcome
75
Computing the expected value from Table 4.1

x P(x) x P(x)
0 .37 0
1 .31 .31
2 .18 .36
3 .09 .27
4 .04 .16
= 1 .1

S [ x. P(x)] = 1.1 power cuts


In the long run, the mean or expected number of power cuts on the given day for the lady is 1.1
power cuts. Of course, the number of power cuts will never be 1.1.
Importance of Expectation: Expected value is of great significance in arriving at a decision
under conditions of uncertainty when the probabilities of choice of different acts are known. The
expected monetary value indicates the average profit that would be gained if a particular act was
selected while the expected opportunity loss would give us the difference between the highest
possible profit for an event and the accrual profit obtained for the particular action taken.

7.4 VARIANCE
The variance of a discrete random variable is calculated by first subtracting the mean from each
possible value of x to obtain the deviations, then squaring each deviation and multiplying the
result by the probability of the corresponding x value, and then adding all those quantities
together.
Symbolically,
2 2
s x=S(x-m) p(x)
Where x = an outcome
p(x) = probability of a given outcome
m = mean
The standard deviation of x, denoted by ðx, is the square root of the variance. The formula for
standard deviation is given by:

x P(x) (x  ) 2 (x-ð) 2  P(x)


0 .37 (0-1.1) 2 = 1.21 (1.21) (.37) = .44
1 .31 (1-1.1)2 = .01 (.01) (.31) = .00
2 .18 (2-1.1)2 = .81 (.81) (.18)= .14
3 .09 (3-1.1)2 = 3.61 (3.61) (.09) = .32
4 .04 (4-1.1)2 = 8.41 (8.41) (.04) = .33
(x-μ)2 . P(x) = 1.23
Thus the variance of s2 = 1.23 and
Standard Deviation is s =Ö(1.23) = 1.11 power cuts.

76
The variance s 2x and standard deviation s x (of x) occurs when the probability distribution
describes how x values are distributed among members of a population (so that the probabilities
are population relative frequencies).
Example 1
A person plays a game of throwing a die under the condition that he could get as many rupees as
the number of points on the upper most face. Find the expectation and variance of his winnings.
Solution: Let X denote the amount received by the person. Then X will be a random variable
taking the values Rs 1, Rs 2, Rs 3, Rs 4, Rs 5, Rs 6 with probabilities 1/6 each.
Now the expectation of the person's winnings, i.e, E(X) is given by

7.5 BINOMIAL DISTRIBUTION


Binomial probability distribution is one of the more frequently encountered discrete probability
distributions. It is the most fundamental discrete probability distribution in Statistics. Binomial
distribution, also known as Bernoulli distribution, was named after Swiss mathematician James
Bernoulli (1654-1705), who innovated it in 1700 and the theory was first published in 1713, eight
years after his death. Such distribution arises when the experiment of interests consists of making
a sequence of dichotomous observations (those that have two possible values in the observation).
The process of making such an observation is called a trial. As an example, one characteristic of
blood type in humans is the Rh factor, which can be either positive or negative (dichotomous).
We could conduct such an experiment that consists of noting the Rh factor for each of 30 blood
donors as a sequence of 30 dichotomous trials, where each trial consists of observing the Rh
factor (positive or negative) of a single donor.
The properties of a Binomial distribution are as follows:
1. A trial is repeated under the same conditions for a fixed and finite number of times, say, n.
2. Each such trial results only in two possible outcomes (success or failure) which are
mutually exclusive. (Trial means an attempt to produce an uncertain event A. The outcome
in which we are interested is called success. 'p' denotes the probability of success in each
trial and q (= 1- p) is the constant probability of failure in each trial.

77
3. Each trial is independent of other trials.
4. The probability of success 'p' remains constant from trial to trial. Similarly, the probability
of failure remains the same.
Subject to the fulfilment of the above conditions the probability of x successes in n trials (x< n )
denoted by P(x) can be derived as:
P(X= x) = P(x) = nCx px qn-x , (x= 0, 1, 2…, n) (i)
p+q = 1, p > 0, q> 0
The binomial random variable x is defined as
x = number of successes observed when a binomial experiment is performed. This probability
distribution is called the binomial probability distribution or binomial distribution where x is
called the binomial variate.

7.6 BERNOULLI TRIALS


Repeated independent trials are called Bernoulli trials or Bernoullian series of trials if there are
only two possible outcomes (success or failure) for each trial and the probabilities of success and
failure in each remain the same throughout the trials.
Using the term 'Bernoulli trials', binomial distribution can precisely be defined as follows:
Let the random variable X denote the number of successes in n Bernoulli trials. Then the
probability of x successes i.e, P(X= x) or P(x) is given by:
P(X= x) = P(x) = nCx px qn-x , (x= 0, 1, 2…, n) p+q = 1, p > 0, q> 0
This is the binomial probability distribution.

7.7 ADDITIONAL PROPERTIES OF BINOMIAL DISTRIBUTION


¡ Binomial distribution is a discrete probability distribution.
¡ n and p are the two parameters of binomial distribution.
¡ The sum of all the probabilities is unity, i.e., P(0) + P(1) + P(2) +… + P(n) = 1.
¡ For a binomial distribution, the mean ð = np; standard deviation, ð = Önpq
¡ Binomial distribution tends to normal distribution as n increases.
¡ Binomial distribution is symmetrical if p = q = ½. It is positively skewed if p <½,
negatively skewed if p > ½.

7.8 IMPORTANCE OR UTILITY OF BINOMIAL DISTRIBUTION


Theoretical distributions, such as Binomial distributions help by providing data on the basis of
which observed (empirical or experimental) results can be assessed. When theoretical
distributions are available, then observed distributions serve no purpose. Theoretical
distributions provide the decision maker with a sound basis for taking rational and dependable
decisions. However, in order to apply binomial distribution the conditions for it have to hold true-
i) Two mutually exclusive outcomes,

78
ii) A fixed probability of outcome in any trial and,
iii) Independent trials.

7.9 HYPERGEOMETRIC DISTRIBUTION


The second discrete statistical distribution of our discussion is the hypergeometric distribution.
Hypergeometric distribution is used to complement analyses that can be done by using the
binomial distribution however, the hypergeometric distribution finds its place only in those trials
which are done without replacement. The hypergeometric distribution, similar to the binomial
distribution, consists of two possible outcomes- success and failure. However, to use the
hypergeometric distribution, one must know the size of the population and the proportion of
successes and failures in it. As sampling is done without replacement in the case of
hypergeometric distribution, so events cannot be considered independent and hence information
about the population build up must be known in order to re determine the probability of a success
in each successive trial.
Properties of hypergeometric distribution are as follows:
¡ It is a discrete probability distribution.
¡ Each trial has two outcomes- success or failure.
¡ Sampling is done without replacement.
¡ The population (N), is finite and known.
¡ The number of successes (A) in the population, is known. The formula for hypergeometric
distribution is given as follows:
P(x) = (ACx . N-ACn-x )/NCn
Where
N= size of the population
n= sample size
A = number of successes in the population
x = number of successes in the sample; sampling is done without replacement

Example: There are 20 lottery tickets with 3 prizes. Find the probability that out of 5 tickets
purchased exactly two prizes are won.
Here,
N = 20, n = 5, A = 3, x = 2
P(x) = (ACx . N-ACn-x )/NCn
= (3C2. 17C3)/ 20C5 = 5/38
The parameters of a hypergeometric distribution are N, A and n. Creating a table in
hypergeometric distribution is impossible because of the multitude of combinations possible of
these three parameters. Thus, each probability in this case has to be calculated and it makes a very
time consuming and tedious task for the researcher. Because of this reason, most researchers use
hypergeometric distribution only in cases of working binomial problems without replacement.

79
Thus, a hypergeometric distribution used be used as an alternative only in the following cases:
¡ Sampling is being done without replacement.
¡ n ³ 5% N.
Hypergeometric probabilities are calculated under the assumption of equally likely sampling of
the remaining elements of the sample space.

7.9 POISSON PROBABILITY DISTRIBUTION


Poisson distribution was developed by Siemon Denis Poisson (1781-1840), a French
mathematician and is named after him. He developed this theory during the later part of his life.
Whereas binomial distribution looks at the dichotomous outcomes of a given definite number of
trials, Poisson distribution focuses on only the finite number of occurrences over some interval
or continuum. An example might of this sort might focus on the number of cars that randomly
arrive at a car service facility during a 20 minute interval. The Poisson distribution describes rare
events.
Following are the properties of Poisson distribution:
¡ It is a discrete probability distribution.
¡ It describes rare or improbable events.
¡ The trials in this distribution are independent of each other.
¡ It describes such discrete occurrences over an interval.
¡ The occurrences in each interval can range from 0 to infinity.
¡ The number of occurrences that is expected should be held constant during the experiment.
¡ The sum of all the probabilities of successes is unity, i.e, P(0) + P() + P(2) +… = 1

Examples of Poisson distribution are as follows:


¡ Number of telephone calls per minute at a telephone booth
¡ Number of sewing flaws per pair of jeans during production
¡ Number of hazardous waste sites per state in India
¡ The number of mistakes committed by a good typist per 400 words typed The formula for
Poisson distribution is as follows:
P(x) = (lx e-l)/ x!
where, x = 0, 1, 2, 3…
l = long-run average
e = 2.718282
Note: The value of ? must be held constant throughout a Poisson experiment. One should be very
specific in describing the interval for which a particular ? is used. It can be so that during a 5
minute interval, the number of customers that arrive at a Mercedes showroom may vary from
hour to hour, day to day and month to month. Thus, a researcher must not apply a given lambda to
intervals for which lambda changes.

80
As an example,
Suppose bank customers arrive randomly on weekday afternoons at an average of 3.2 customers
every 4 minutes. What is the probability of exactly 6 customers arriving in a 5- minute interval on
a weekday afternoon? The lambda for this problem is 3.2 customers per 4 minutes. The value of x
is 6 customers per 5 minutes. The probability of 6 customers randomly arriving during a 5-
minute interval in the face of a long run average has been 3.2 customers per 4-minute interval is
(3.26) (e-4.2 ) = 1073.74 (.0408) = .0608
6! 720
If a bank averages 3.2 customers every 4 minutes, the probability of 6 customers arriving during
any one 5-minute interval is .0608.

7.10 IMPORTANCE OF POISSON DISTRIBUTION


Poisson distribution is used in those cases where the probability of occurrence of the event is
small and the number of trials are large. There are varieties of such events in Insurance, Physics,
Economics, Biology, Business, where the probability of occurrence of an event is small whereas
the number of trials are large. Poisson distribution finds its application in all such cases.

SELF ASSESSMENT QUESTIONS


Long Answer Questions
1. Explain expected value with the help of examples.
2. State the distinctive features of the Binomial, Hypergeometric and Poisson distribution.
3. Explain the circumstances when the following probability distributions are used:
a. Binomial distribution
b. Poisson distribution
c. Hypergeometric distribution
4. Six coins are tossed simultaneously. What is the probability of getting (i) two heads,
(ii) at least two heads (iii) more than two heads?
5. If 5% of the electric bulbs manufactured by a company are defective, use Poisson
distribution to find the probability that in a sample of 100 bulbs (i) none is defective,
(ii) 5 bulbs are defective, (Given e5 = 0.007)
6. Suppose 20 major computer companies operate in the United States and that 14 are located
in California's Silicon Valley. If three computer companies are selected randomly from the
entire list, what is the probability that one or more of the selected companies are located in
the Silicon Valley?

Short Answer Questions


1. What do you mean by a random variable?
2. Define the term 'expected value'.

81
3. What are the types of probability distributions?
4. When is the Poisson distribution used?
5. Who propounded the Binomial distribution?
6. List the properties of hypergeometric distribution.

Fill in the blanks


1. When p is small (say 0.1) the binomial distribution is skewed to the ________(Right)
2. __________describes the distribution of probabilities where there are only two mutually
exclusive outcomes for each trial of an experiment. (Binomial distribution)
3. Random variables are called _________ (Variate)
4. Hypergeometric distribution is used in cases of trials without ________ (Replacement)
5. Poisson distribution was developed by_________ (Siemon Denis Poisson)
6. When a random variable takes on a finite number of values it is called __________
(Discrete)

True and False


1. Random variable can take any value depending upon trial to trial. (True)
2. Number of telephone calls made per minute at a small business is an example of Poisson
distribution. (True)
3. Binomial distribution is a dichotomous distribution. (True)
4. For a binomial distribution, mean = mnp. (False)
5. Poisson distribution is used for all common events. (False)
6. Hypergeometric distribution is used in case of infinite population. (False)

Multiple Choice Questions


1. The sum of the values of the random variable weighted by the probability that the random
variable will take on the value:
a. Variance of random variable
b. Expected value of random variable
c. Mean of random variable
d. Standard value of random variable

2. In a Binomial Distribution, if p, q and n are probability of success, failure and number of


trials respectively then variance is given by
a. np
b. npq
c. np2q
d. npq2
82
3. When a random variate can take any value in the given interval a ? x ? b, the distribution is
a:
a. Poisson distribution
b. Binomial distribution
c. Geometric distribution
d. Continuous probability distribution

4. When two unbiased coins are tossed together, the number of elements in the sample space
shall be:
a. 2
b. 3
c. 4
d. 0

5. The certainty of an event in probability is given by:


a. 0
b. 1
c. 2
d. 1.5

6. A binomial distribution has a mean of 5 and a variance 4. The number of trials is


a. 18
b. 17
c. 25
d. None of these

Match the following


Column A Column B
1 Random variable (C) A Finite number of values
2 Binomial distribution (E) B Önpq
3 Discrete variable (A) C Takes on a numerical value
4 Continuous variable (D) D Height of individuals
5 Expected value (F) E Dichotomous distribution
6 Standard deviation (B) F Mean value

83
Problem solving activities
Suppose that in the bookkeeping operation of a large corporation the probability of a recording
error on any one billing is .005. Suppose the probability of a recording error from one billing to
the next is constant, and 1,000 billings are randomly sampled by an auditor.
a. What is the probability that fewer than four billings contain a recording error?
b. What is the probability that at least 10 billings contain a billing error?
c. What is the probability that all 1,000 billings contain no recording errors?

Suggested reading:
¡ Chandan, J.S. Statistics for Business and Economics. New Delhi: Vikas Publishing House
Pvt Ltd., 1998
¡ Gupta, S.C. Fundamentals of Statistics. New Delhi: Himalaya Publoshing House, 2006.
Kothari, C.R. Quantitative Technique. New Delhi: Vikas Publishing House Pvt. Ltd.,1984
¡ Black, K. Business Statistics: Contemporary Decisions Making. Wiley, 2009.

References
¡ Black, K. (2009, December 1). Business Statistics: Contemporary Decision Making (6th
ed.). Wiley.
¡ Padmalochan, H., & Hazarika, P. (ca. 2007, April 4). A Textbook of Business Statistics (1st
ed.) [Print]. S. Chand Limited.
¡ Peck, R., Olsen, C., & Devore, J. (2022, September 30). Introduction to Statistics and Data
Analysis (AP(R) Edition) (4th ed.). Brooks/Cole, Cengage Learning.
¡ Quantitative Techniques (New Format). (2013, January 1). Vikas Publishing House.

84
MODULE - 8 CONTINUOUS PROBABILITY
DISTRIBUTIONS

STRUCTURE
¡ Introduction to continuous Probability Distribution
¡ Uniform distribution
¡ Determining probabilities in a Uniform distribution
¡ Normal distribution
¡ History of normal distribution
¡ Properties of normal distribution
¡ Standardized normal distribution
¡ Importance of normal distribution
¡ Exponential distribution
¡ Properties of exponential distribution
¡ Exponential probability density function
¡ Using the normal curve to approximate binomial distribution problems
¡ The normal curve and discrete variables
¡ Questions and Exercises

LEARNING OBJECTIVES
After going through this unit, you will be able to:
¡ Understand continuous probability function
¡ Explain uniform distribution, normal and exponential distribution
¡ Understand the standardized normal distribution
¡ Understand approximation of normal curve to discrete probability function Unit Contents

8.1 CONTINUOUS PROBABILITY DISTRIBUTION


In the previous unit, we learnt about random variables and how continuous variable is one of the
types of it. A continuous sample space will take as many sample points as there are in some given
range. A continuous probability distribution shall take values from continuous variable which
shall be taken on for every point in a given interval.
In continuous distributions, the probabilities of outcomes occurring between the start point and
the end point of the range is determined by finding out the area under the curve between these two
points. There many forms of continuous distributions such as uniform distribution, normal
distribution, chi- square distribution, exponential distribution, t and F distribution.
We shall be discussing the uniform, normal and exponential distributions in this module.
85
8.2 UNIFORM DISTRIBUTION
In statistics, uniform continuous distribution, also known as the rectangular distribution, is a
simple continuous distribution in which the same height, or f(x) is obtained over a range of
values.
A uniform distribution is defined as follows:

Graphically, this distribution is represented as a rectangle where b-a is the base and 1/b-a is the
height. The increase in distance between a and b proportionately decreases the density at any
particular value within the distribution boundaries. Since the probability density function
integrates to 1, the height of the probability density function decreases as the base length
increases.
The formula for Uniform distribution is given by:

0 for all other values

In uniform distribution, the total area under the curve is 1 and is equal to the product of the length
and width of the rectangle. As the distribution lies, by definition, between the x values of a and b,
the length of the rectangle is (b-a). Combining this with the fact that the area of the curve is 1, we
can find the height of the rectangle in the following way:
Area of rectangle = (Length) (Height) = 1
But, Length = b-a
Therefore, (b-a) (Height) = 1
And Height = 1/ (b-a)
These calculations show why, between the x values of a and b, the height of the distribution
remains constant (1/b-a).

86
The mean and standard deviation of a Uniform distribution is given by:
µ= (a+b)/ 2
s = (b-a) /Ö12
As an example of this situation, suppose a production line is set up to manufacture machine
braces in lots of five per minute during a shift. When the lots are weighed, variation among the
weights is detected, with lot weights ranging from 41 to 47 grams in a uniform distribution. The
height of this distribution is

f(x) = 1/(b-a) = 1/ (47-41)


= 1/ 6

8.2.1 Determining probabilities in a Uniform distribution


While in case of discrete probability distributions, the probability function yields the value of the
probability, however, in case of continuous distribution, the probability is calculated by finding
out the area over an interval of the function.
Probabilities in a Uniform distribution is given by:
P(x) = x2 - x1 / b - a
Where,
a £ x1£x2 £ b
We should remember that the area between a and b is 1. The probability for any interval that
includes a and b shall be 1. The probability of x ³ b or of x £ a is zero because there is no area
above b or below a.

Example
Suppose the amount of time it takes to assemble a plastic module ranges from 27 to 39 seconds
and that assembly times are uniformly distributed. Describe the distribution. What is the
probability that a given assembly will take between 30 and 35 seconds? Fewer than 30 seconds?
Solution
f(x) = 1/ 39?27 = 1 /12
µ = (a+b)/2 = (39+27)/2 = 33
s = b-a/ Ö12
= 39?27 /Ö12
= 12 / Ö12
= 3.464

8.3 NORMAL DISTRIBUTION


Normal distribution is one of the commonly used continuous probability distribution. This is

87
because the distribution finds its application in many problems. It fits quite a few human
characteristics such as height, weight, life expectancy, IQ, etc. Other species such as animals,
insects, and trees also have characteristics that can be attributed to be possessing the
characteristics of a normal distribution. Certain variables in industry and business, too, are seen
to be normally distributed. Besides these variables, the normal distribution is an integral part of
statistics. When large sample sizes are taken for experiment, the statistics tend to be normally
distributed regardless of the kind of distribution from which they are drawn. Karl Gauss, a
mathematician-astronomer of the 18th century is associated with the normal distribution and
hence, this distribution is also called the Gaussian distribution.

The above image depicts the normal distribution.

8.3.1 History of the Normal Distribution


The first person to be associated with the normal distribution is Karl Gauss who recognized that
the errors of repeated measurement of objects are often normally distributed. Thus, it is known as
the Gaussian distribution or the normal curve of error. At a lesser extent, some credit is also given
to Pierre-Simon de Laplace (1749-1827) for discovering the normal distribution. However,
many people now believe that Abraham de Moivre (1667-1754), a French mathematician, first
understood the normal distribution. De Moivre determined that the binomial distribution
approached the normal distribution as a limit.

8.3.2 Properties of normal distribution


i. It is a continuous probability distribution.
ii. The normal curve is bell shaped and perfectly symmetrical about the line X = µ.
iii. The normal distribution is described or characterized by two parameters: the mean, µ , and
the standard deviation?.
iv. It is asymptotic to the horizontal axis.
v. It is unimodal.
vi. It is a family of curves.
vii. Area under the curve is 1.
viii. The area under the two tails extend to infinity.
88
The normal distribution is symmetrical. Each half of the distribution is a mirror image of the
other half. In theory, the normal distribution is asymptotic to the horizontal axis, as it does not
touch the x-axis, and it goes forever in each direction; however, the reality is that most
applications of the normal curve are experiments that have finite limits of potential outcomes.
The normal curve sometimes is referred to as the bell-shaped curve. It is unimodal because
values build up in only one portion of the graph-the center of the curve. The area under the curve
yields the probabilities, so the total of all probabilities for a normal distribution is 1. Because the
distribution is symmetric, the area of the distribution on each side of the mean is 0.5.

8.3.3 Standardized normal distribution


In normal distribution, for every pair of m & s we have a different normal distribution. This
characteristic of the normal curve (a family of curves) gives rise to difficulty. It so happens that as
a result of other normal distributions coming up, analysis of each combination of  and € could
become tedious. Fortunately, though, a mechanism has been developed to deal away with this
and with the help of which all normal distributions can be converted into a single distribution: the
z distribution. This process yields the standardized normal distribution.
The conversion formula for any x value of a given normal distribution is given as:
z = x-m/s, s# 0

A z score signifies the number of standard deviations that a value, x, is above or below the mean.
If x is above the mean, then z score is positive, if value of x is less than the mean, we have a
negative z score and when x = z, the associated z score is zero. The z distribution is a normal
distribution with a mean of 0 and a standard deviation of 1.

8.3.4 Importance of normal distribution


i. Data obtained from psychological, physical and biological measurements approximately
follow normal distribution.
ii. Distributions like binomial, Poisson etc. can be approximated to normal distribution.
iii. Normal curve is used to find confidence limits of the population parameters.
iv. The theory of errors of observations in physical measurements are based on normal
distribution.

8.4 EXPONENTIAL DISTRIBUTION


The third continuous distribution that we shall discuss is the exponential distribution. It is closely
related to the Poisson distribution. While the Poisson is a discrete distribution and explains
random occurrences over some interval, the exponential distribution is a continuous probability
distribution and it explains the time interval between random occurrences.

8.4.1 Properties of exponential distribution


a) It is a continuous distribution.
b) It is a family of distributions.

89
c) It is skewed to the right.
d) The x values range from zero to infinity.
e) Its apex is always at x = 0
f) The curve steadily decreases as x gets larger.

8.4.2 Exponential probability density function


-lx)
f (x) =le
Where x ³0
l>0
And e = 2.71828

The defining parameter of an exponential distribution is l. Each unique value of l determines a


different exponential distribution, resulting in a family of exponential distributions. The mean of
an exponential distribution is s= 1/ l , and the standard deviation of an exponential distribution is
s = 1/ l.

Example:
A manufacturing firm has been involved in statistical quality control for several years. As part of
the production process, parts are randomly selected and tested. From the records of these tests, it
has been established that a defective part occurs in a pattern that is Poisson distributed on the
average of 1.38 defects every 20 minutes during production runs. Use this information to
determine the probability that less than 15 minutes will elapse between any two defects.
Solution
The value of l is 1.38 defects per 20-minute interval. The value of µ can be determined by
µ = 1/ l = 1/138 = .7246

On the average, it is .7246 of the interval, or (.7246) (20 minutes) = 14.49 minutes, between
defects. The value of x0 represents the desired number of intervals between arrivals or
occurrences for the probability question. In this problem, the probability question involves 15
minutes and the interval is 20 minutes. Thus x0 is 15 20, or .75 of an interval. The question here is
to determine the probability of there being less than 15 minutes between defects. The probability
formula always yields the right tail of the distribution-in this case, the probability of there being
15 minutes or more between arrivals. By using the value of x0 and the value of  , the probability
of there being 15 minutes or more between defects can be determined.

P(x ³ x0) = P(x ³ .75) = elx0= e(-1.38)(.75) = e-1.035 = .3552

The probability of .3552 is the probability that at least 15 minutes will elapse between defects.

90
8.5 USING THE NORMAL CURVE TO APPROXIMATE BINOMIAL
DISTRIBUTION PROBLEMS
For certain kinds of the binomial distribution, they can be approximated using normal
distribution. As a general trend, when sample sizes become large, binomial distribution
approaches a normal distribution in shape regardless of the value of p.
Whenever a continuous distribution is used to approximate a discrete distribution, the question
that naturally arises is how good is the approximation? Statistics instructors usually state that it
depends. In case of the normal approximation to the binomial, the goodness of fit depends on two
quantities which define the binomial distribution: n and p. Most statisticians go by a simple "rule
of thumb" they apply while approximating for binomial with a normal distribution, such as:
When either np < 10 or n(1 -p) <10, the binomial distribution is too skewed for the normal
approximation to give accurate results.
It can be debated that using the normal distribution to approximate another discrete distribution
which we can evaluate exactly sounds a bit too good to be true. But the point to remember is that
we cannot always find exact probabilities in other situations in statistics and we will have to rely
on approximations. This involves a tradeoff between ease of calculation and exactness of answer.
An understanding of this while using the normal approximation to binomials will give us a
clearer understanding of the issues involved.

8.6 THE NORMAL CURVE AND DISCRETE VARIABLES

Fig 8.1 A normal curve approximation to a probability histogram

The probability distribution of a discrete variable x is represented pictorially by a probability


histogram. The probability of a particular value is represented by the area of the rectangle
centered around that particular value. x can have different values which are usually isolated
points on the number line, bearing whole numbers. For example, if x = the IQ of a randomly
selected 9-year-old child, then x is a discrete random variable, because an IQ score must be a
whole number.
A normal curve has been found to well approximate a probability distribution, as illustrated in

91
Figure 1.1. In such cases, it is accustomed to say that x has approximately a normal distribution.
The normal distribution can then be used to calculate approximate probabilities of events
involving x.

Example
Premature babies are those born more than 3 weeks early. Newsweek (May 16, 1988) reported
that 10% of the live births in the United States are premature. Suppose that 250 live births are
randomly selected and that the number x of "preemies" is determined.
Solution Now, because
n p = 250(.1) = 25 ³10
n(1- p = 250 (0.9) = 225 ³ 0
x has approximately a normal distribution, with
µ= 250 (0.1) = 25
s = Ö(250(0.1)(0.9) )= 4.743
The probability that x is between 15 and 30 (inclusive) is
P( 15 £ x £ 30) = P( 14.5?25)/4.743 £z £ 30.5 - 25 / 4.743
= P (-2.21£ z £ 1.1.6)
= .8770 - .0136
= .8634

8.8 SELF ASSESSMENT QUESTIONS


Long Answer Questions
1. Distinguish between discrete and continuous probability function.
2. Explain uniform distribution. How would you determine probabilities in a uniform
distribution?
3. What is normal distribution? Explain standardized normal distribution.
4. Explain how normal curve is used to approximate binomial distribution.
5. A financial analyst computed the return on stockholder's equity for all companies listed on
the BSE. She found that the mean of this distribution was 10 %, with a standard deviation of
5 %. She was interested in examining further those companies whose return on
stockholder's equity was between 16 % and 22 %. Of the approximately 1300 companies
listed on the exchange, how many were of interest to her?
6. Suppose the average speeds of passenger trains traveling from Newark, New Jersey, to
Philadelphia, Pennsylvania, are normally distributed, with a mean average speed of 88
miles per hour and a standard deviation of 6.4 miles per hour. What is the probability that a
train will average less than 70 miles per hour?

92
Short Answer Questions
1. What do you mean by continuous probability distribution?
2. Narrate a brief history of normal distribution.
3. How would you determine probability under a uniform distribution?
4. List some of the properties of normal distribution.
5. What is exponential distribution? State the properties.
6. Suppose that heights of all cakes baked with a certain mix closely follow a normal
distribution with mean 5.3 cm and standard deviation 0.75 cm. find the percentage of cakes
which have a height of 4.4 cm or less.

Fill in the blanks


1. Uniform distribution is a form of________distribution. (Continuous)
2. _______ is the height of the rectangle in an uniform distribution. (1/b-a)
3. __________is associated with normal distribution. (Karl Gauss)
4. The curve of a normal distribution is _________ (Bell shaped)
5. Exponential distribution is skewed to the ____________ (Right)
6. Discrete distributions are approximated using the________curve. (Normal)

True and False


1. Continuous random variable are counted and not measured. (False)
2. Binomial distribution can be approximated by a normal curve. (True)
3. Height of individuals is normally distributed. (True)
4. x values range from zero to infinity in case of an exponential distribution. (True)
5. Uniform distribution is a type of discrete probability distribution. (False)
6. Normal distribution is one of the most commonly used distributions. (True)

Multiple Choice Questions


1. Which one of these variables is a continuous random variable?
a. The time it takes a randomly selected student to complete an exam.
b. The number of tattoos a randomly selected person has.
c. The number of women taller than 68 inches in a random sample of 5 women.
d. D. The number of correct guesses on a multiple choice test.

93
2. Following is the example of a continuous probability distribution:
a. Discrete probability distribution
b. Poisson distribution
c. Normal distribution
d. Binomial distribution

3. A variable that can assume any value between two given points is called
a. Uncertain random variable
b. Irregular random variable
c. Discrete random variable
d. Continuous random variable

4. Normal distribution is symmetric is about


a. Variance
b. Mean
c. Standard deviation
d. Covariance

5. The area under a standard normal curve is :


a. 0
b. 1
c.
d. Not defined

6. If value of interval a is 2.5 and the value of interval b is 3.5 the value of ? for uniform
distribution is
a. 0.5
b. 3
c. 2.5
d. 2

94
Match the following
Column A Column B
1 Continuous variable A F distribution
2 Exponential distribution B =a+b
3 Continuous probability function C Abraham de Moivre
4 Normal distribution D Height of individuals
5 Uniform distribution E Normal curve
6 Approximation of discrete probability distribution F
(Answer Key: 1 = D, 2 = F, 3 = A, 4 = C, 5 = B, 6 = E)

Problem solving activities


According to the Internal Revenue Service, income tax returns one year averaged $1,332 in
refunds for taxpayers. One explanation of this figure is that taxpayers would rather have the
government keep back too much money during the year than to owe it money at the end of the
year. Suppose the average amount of tax at the end of a year is a refund of $1,332, with a standard
deviation of $725. Assume that amounts owed or due on tax returns are normally distributed.
a. What proportion of tax returns show a refund greater than $2,000?
b. What proportion of the tax returns show that the taxpayer owes money to the government?
c. What proportion of the tax returns show a refund between $100 and $700?

Suggested reading:
¡ Chandan, J.S. Statistics for Business and Economics. New Delhi: Vikas Publishing House
Pvt Ltd., 1998
¡ Gupta, S.C. Fundamentals of Statistics. New Delhi: Himalaya Publoshing House, 2006.
Kothari, C.R. Quantitative Technique. New Delhi: Vikas Publishing House Pvt. Ltd., 1984
¡ Black, K. Business Statistics: Contemporary Decisions Making. Wiley, 2009.

References
¡ Black, K. (2009, December 1). Business Statistics: Contemporary Decision Making (6th
ed.). Wiley.
¡ Padmalochan, H., & Hazarika, P. (ca. 2007, April 4). A Textbook of Business Statistics (1st
ed.) [Print]. S. Chand Limited.
¡ Peck, R., Olsen, C., & Devore, J. (2022, September 30). Introduction to Statistics and Data
Analysis (AP(R) Edition) (4th ed.). Brooks/Cole, Cengage Learning.
¡ Quantitative Techniques (New Format). (2013, January 1). Vikas Publishing House.

95
SAMPLING AND DIFFERENT
MODULE - 9
SAMPLING TECHNIQUES

STRUCTURE
¡ Sampling
¡ Reasons for sampling
¡ Random Versus Non-random Sampling
¡ Types of Random Sampling

LEARNING OUTCOMES
1. To understand the meaning of sampling
2. To understand random sampling methods
3. To understand non-random sampling methods

9.1 SAMPLING
This Chapter explains the sampling and sampling distribution process of some statistics. How do
get the data used for statistical analysis? Why do the researchers usually do the selection rather
than conducting a census? What are the differences between various sampling methods? What is
random and Non-random sample?
This chapter will address all the above questions about sampling and the distribution of two
statistics: Sample means and the sample proportion. These are the basic concepts for statistical
analysis.
Sampling is a process used in statistical analysis in which several observations are taken from a
larger population. The predetermined method of sampling can be applied.
Sampling is widely used in business as a means of gathering valuable information about a
population. Data are collected from samples and afterwards, it can be analysed to make
inferences and to produce the results.

9.2 REASONS FOR SAMPLING


Instead of conducting a census taking a sample offers several advantages.
1. The sample can save Time & money.
2. The scope of the study can be broadened through sampling with given resources
3. A sample can save the product because sometimes the research process is destructive
4. In case when accessing the population is difficult to sample is the only option

96
9.3 RANDOM VERSUS NON-RANDOM SAMPLING
Random sampling and Non random sampling are two main types of sampling. Every unit of the
population has an equal chance or probability of being selected in to the sample is called random
sampling. Random sampling implies that chance enters into the process of selection. For
example, If you would like to have the opinions of Students from the university you will take a
survey of random 100 students out of 10000 population of [Link] method works if there is
an equal chance that any of the subjects in a population will be chosen. Researchers choose
simple random sampling to make generalizations about a population.
In non random sampling not every unit of the population has the same probability of being
selected into the sample. Members of non random samples are not selected by chance. For
example, they might be selected because they are at the right place at the right time or because
they know the people conducting the research. Sometimes random sampling is called probability
sampling and non-random sampling is called nonprobability sampling. Because every unit of the
population is not equally likely to be selected, assigning a probability of occurrence in non-
random sampling is impossible. The statistical methods presented and discussed in this text are
based on the assumption that the data come from random samples. Non random sampling
methods are not appropriate techniques for gathering data to be analyzed by most of the
statistical methods presented.

9.4 TYPES OF RANDOM SAMPLING


The following are commonly used random sampling methods:
¡ Simple Random Sampling
¡ Cluster Sampling
¡ Stratified random Sampling
¡ Systematic Sampling

9.4.1 SIMPLE RANDOM SAMPLING:


Simple Random sampling is very common approach of Random Sampling. It means that every
single member of the population is put in to a one big group and then choosing it randomly.
E.g. restaurant keeps the fishbowl on the counter and asks customers to put their visiting cards.
Once a month, a single Business card is withdrawn to award one lucky diner with a one meal free
offer.
A pharmaceutical company wants to test the effectiveness of a new drug. Volunteers are taken out
from various age groups and then the effectiveness of drug is observed for various age groups.

9.4.2 STRATIFIED RANDOM SAMPLING


In this sampling method, the population is divided into groups based on similar characteristics.
Each group will be called a stratum (A plural form of Strata) and those one or many choices are
made randomly from each stratum.
eg. A study on tax reform stratified a population according to income level and them random
sampling is done from each level.

97
9.4.3 CLUSTER SAMPLING:
Like stratified sampling in cluster sampling, the population is divided into groups based on
similar characteristics or particular characteristics. But, in stratified sampling, one or more
samples are chosen from each stratum while in cluster sampling the clusters are chosen at
random and then takes samples from them. It is often used for market research.
Eg. A study on the impact of natural disasters may divide a population region-wise into clusters,
and then random collection is chosen to begin the study duster's overall impact.
Sometimes the clusters are too large, and the second set of clusters is taken from each original
cluster. This technique is called two-stage sampling.
e.g., if a researcher could divide India into clusters of cities. She could then divide the cities into
clusters of Areas and randomly select individual residences from the area clusters. The ?rst stage
is selecting the test cities and the second stage is selecting the areas. Cluster or area sampling
offers several advantages. Two of the foremost advantages are convenience and cost. Clusters are
usually convenient to obtain, and the cost of sampling from the entire population is reduced
because the scope of the study is reduced to the clusters.

9.4.4 SYSTEMATIC SAMPLING


Systematic sampling is a forth random sampling [Link] sampling is not done in an
attempt to reduce sampling error. But, It is used because for convenience and relative ease of
administration. With systematic sampling, every k th item is selected to produce a sample of size
n from a population of size N. The value of k, sometimes called the sampling cycle, can be
determined by the following formula. If k is not an integer value, the whole-number value should
be used.

[Link] DETERMINING THE VALUE OF K


where
n = sample size N = population size k = size of interval for selection
k=N/n
Example of systematic sampling, a MIS researcher wanted to sample the Medium enterprises in
Maharashtra . She had enough ?nancial support to sample 1,000 companies (n). The Directory of
Maharashtra Manufacturers listed approximately 10,000 total manufacturers in Maharashtra
(N) in alphabetical order. The value of k was 10 (10,000/1,000) and the researcher selected every
10th company in the directory for his sample. Did the researcher begin with the ?rst company
listed or the 10th or one somewhere between? In selecting every k th value, a simple random
number table should be used to determine a value between 1 and k inclusive as a starting point.
The second element for the sample is the starting point plus k. In the example, k = 10, the
researcher would have gone to a table of random numbers to determine a starting point between 1
and 10. Suppose he selected the number [Link] would have started with the 5th company, then
selected the 15nth (5 + 10), and then the 25th, and so on.
With other advantages like Systematic sampling is evenly distributed across the frame. A
researcher can judge if the sampling plan is being followed or not.

98
9.4.5 NON-RANDOM SAMPLING
A technique of sampling used to select elements from the population by various mechanism that
does not involve a random selection, this process is called non-random sampling . Because
chance is not used for selecting items from the samples these ,techniques are called non-
probability sampling. In these sampling techniques, sampling error is not expected objectively
for such sampling techniques.

9.5 TYPES OF NON-RANDOM SAMPLING


The following are commonly used Non- random sampling methods:
¡ Convenience Sampling
¡ Judgement Sampling
¡ Quota Sampling
¡ Snowball Sampling

9.5.1 CONVENIENCE SAMPLING


In Convenience Sampling, elements for the sample are selected as per the convenience of the
researcher. The researcher typically chooses elements that are nearby, readily available, or those
who are ready to participate. The sample tends to be less variable than the population because in
many environments the extreme elements of the population are not readily available.
E.g. a convenience sample of homes where interviews are taken by visiting door to door which
includes the houses which are on ground floor or first floor, the houses which doesn't have dogs,
Apartment near the street and the houses where people are willing to speak friendly.

9.5.2 JUDGMENT SAMPLING


When the elements for the samples are chosen by the judgement of researchers it is called
Judgement sampling. Researchers often think that they can have sound judgement ability which
can save time and Money.
If judgement sampling is done , calculating a probability that if element is going to be selected in
to sampling is not possible sometime .As probabilities are based on non -random selection the
sampling errors are not determined objectively. Other disadvantage associated with judgement
sampling is the researcher tends to make errors of selecting the elements in the sample with
judgement in one direction. These systematic errors lead to biases. The researcher may not
include the extreme elements. This kind of sampling does not provide any authentic objective
method for determining whether researcher's judgement is better than or weaker than others.
Judgement sampling is used in three cases:
i) To select unique respondents who can give special information or when respondent is
especially informative
ii) When respondent is difficult to reach ,respondents from special population
iii) When the purpose is in-depth investigation particular type of respondents are identified
Eg. A researcher is interested in studying the reason, why people wear eyeglasses to read books.
Common judgement of researcher may be focussed entirely on the population who indeed wear
eyeglasses. This is the judgement researcher has made in sampling

99
9.5.3 QUOTA SAMPLING
Quota Sampling is the third sampling technique of Non-random Sampling which appears to be
similar to stratified random sampling. There are subclasses of certain population, such as gender
,age group, or geographic region, are used as strata. However, the researcher uses a nonrandom
sampling method instead of randomly sampling from each stratum, to gather data from one
stratum until the desired quota of samples is filled.
With quota controls ,Quotas are described , to set the sizes of the samples hich can be obtained
from the subgroups. Generally, on the basis of proportions of the subclasses quotas are based in
the population. In this case, the quota concept is similar to that of proportional stratified
sampling. Quotas often are filled by using recent, available and applicable elements.
For example,; if the a subclass has been represented by the respondent whose quota has been
filled, the interviewer may terminate the interview. Mostly the quota sampling is used when no
actual frame if available for the population. Quota sampling can be useful if no frame is available
for the population. Also, for quota sampling, preparatory work is minimal. researcher approach
population directly where the quota can belled. The object is to gain the benefits of stratification
without the high field costs of stratification. Ultimately, it remains a non-probability sampling
method.
E.g . Instead of randomly interviewing people to obtain a quota of African Americans, the
researcher would go to the African's community residential area of the city and interview there
until expected responses are achieved to fill the quota. In such sampling, an interviewer may
begin the interviews by applying few filter questions

9.5.4 SNOWBALL SAMPLING


Snowball sampling is forth sampling technique of Non-random sampling , in which survey
subjects are selected and on the basis of their refeerals the sampling is done .The researcher
identifies a person who fits the study profile. The researcher then contacts the subject for the
elements who match the same profile and checks if others can also fit in the profile This is very
cost-effective and efficient method of conducting survey which is particularly useful when
survey subjects are difficult to visit personally .
E.g If survey is based on Information Technology sectors . Since IT employees are working from
home and belonging to different cities and states , in the pandemic , snowball sampling technique
is applicable to collect the responses as IT Employees work in teams and can refer the teammates
to fill up the survey.

9.6 CONCEPT OF PARAMETER AND ESTIMATOR


A parameter is used to describe the entire population being studied. For example, we want to
know the average length of a fish. This is a parameter because it is states something about the
entire population of fish.
the properties of entire population are the parameters which described in numbers
Statistics are numbers that describe the properties of samples.
For example, the average income for INDIA is a population parameter. Conversely, the average
income for a sample drawn from INDIA. is a sample statistic. Both values represent the mean
income, but one is a parameter vs a statistic.

100
Both are summary values that describe a group, and there's a handy mnemonic device for
remembering which group each describes. Just focus on their first letters:
¡ Parameter = Population
¡ Statistic = Sample
A population is the entire group of Objects, people, Transactions, animals etc A sample is a
portion of the population.
Estimator:
An estimator is a function of the sample, i.e., it is a rule that tells you how to calculate an estimate
of a parameter from a sample.
An estimator is a statistic that estimates some fact about the population. You can also think of an
estimator as the rule that creates an estimate. For example, the sample mean(x*) is an estimator
for the population mean, ?.
For example, let's say you wanted to know the average height of children in a school with a
population of 1500 students. You take a sample of 50 children, measure them and find that the
mean height is 58 inches. This is your sample mean, the estimator. You use the sample mean to
estimate that the population mean (estimand) is about 58 inches.
An estimate is a Palue of an estimator calculated from a sample. An estimator is a statistic that
estimates some fact about the population. You can also think of an estimator as the rule that
creates an estimate. For example, the sample mean (x*) is an estimator for the population mean,
m. The quantity that is being estimated (i.e. the one you want to know) is called the estimand.
A good estimator is one that gives UNBIASED, EFFICIENT and CONSISTENT estimates. In
this post, I will explain what these terms mean. An estimator is a formula- we input our sample
values and it gives an estimate of the statistic

(The sample mean is an estimator for the population mean)

9.7 CENTRAL LIMIT THEOREM (CLT)


The central limit theorem (CLT) states that the distribution of sample means approximates a
normal distribution as the sample size gets larger, regardless of the population's distribution.
Sample sizes equal to or greater than 30 are often considered sufficient for the Central limit
Theorem to hold.
Sample sizes equal to or greater than 30 are often considered sufficient for the Central Limit
Theorem to hold.
A key aspect of Central Limit Theorem is that the average of the sample means and standard
deviations will equal the population mean and standard deviation.
A sufficiently large sample size can predict the characteristics of a population more accurately.
Central Limit Theorem is useful in finance when analysing a large collection of securities to
estimate portfolio distributions and traits for returns, risk, and correlation.

101
According to the central limit theorem, the mean of a sample of data will be closer to the mean of
the overall population in question, as the sample size increases, notwithstanding the actual
distribution of the data. In other words, the data is accurate whether the distribution is normal or
aberrant.
As a general rule, sample sizes of around 30-50 are assumed sufficient for the Central Limit
Theorem to hold, it means that the distribution of the sample means is fairly normally distributed.
Therefore, the more samples one takes, the more the graphed results take the shape of a normal
distribution. Note, however, that the central limit theorem will still be approximated in many
cases for much smaller sample sizes, such as n=8 or n=5.
The central limit theorem is often used in conjunction with the law of large numbers, which states
that the average of the sample means and standard deviations will come closer to equalling the
population mean and standard deviation as the sample size grows, which is extremely useful in
accurately predicting the characteristics of populations.

9.8 USEFULNESS OF CENTRAL LIMIT THEOREM


The central limit theorem is useful when one analyses large data sets as it allows one to assume
that the sampling distribution of the mean will be normally-distributed in most of the cases. This
allows for easy statistical analysis and inference.
For example, investors can use central limit theorem to aggregate data of individual's security
performance and generate distribution of sample means which represents a larger population
distribution for security returns over a period of time.

9.8.1 Standard Error:


The standard error of the regression (S), also known as the standard error of the estimate,
represents the average distance that the observed values fall from the regression line.
Conveniently, it tells you how wrong the regression model is on average using the units of the
response variable. Standard error represents how population mean is different from a sample
mean. It tells you how much the sample mean would vary if you were to repeat a study using new
samples from within a single population.
Standard Error represents the average distance that observed values fall from the regression line.
Standard error of regression (s) is also known as standard error of Estimate.
SE is calculated by dividing the standard deviation of the sample by the square root of the sample
size. The mean of the total population has to be calculated. then calculate each measurements
deviation from the mean. Each deviation will be squared from the mean.

102
A larger standard error indicates that the means are more spread out, and thus it is more likely that
your sample mean is an inaccurate representation of the true population mean.
On the other hand, a smaller standard error indicates that the means are closer together, and thus it
is more likely that your sample mean is an accurate representation of the true population mean.
Standard error increases when standard deviation increases. Standard error decreases when
sample size increases because having more data yields less variation in your results.
SE = s / Ö n
s = the population standard deviation
Ön = the square root of the sample size
Example: The values in your sample are 52, 55, 60, and 65.

9.8.2 Calculate the mean by adding these 4 samples and divide it by 4.


(52 + 55+ 60 + 65)/4 = 58 (Step 1).
calculate the sum of the squared deviations of each sample value from the mean (Steps 2-4).
¡ Using the values in this example, the squared deviations are (58 - 52) ^2= 36, (58 - 55) ^2=
9, (58 - 60) ^2=4 and (58 - 65) ^2=49. Therefore, the sum of the squared deviations is 98 (36
+ 4 + 9 + 49).
¡ Next, divide the sum of the squared deviations by the sample size minus one and take the
square root (Steps 5-6). The standard deviation in this example is the square root of [98 / (4
- 1)], which is about 5.72.
¡ Lastly, divide the standard deviation, 5.72, by the square root of the sample size, 4 (Step 7).
¡ The resulting value is 2.86 which is the standard error (SE) of the values in this example.

9.9 SAMPLING DISTRIBUTION


When the researcher reaches to the conclusion about the population parameter from the statistics
while the selection of random sample from the population computes the statistic on the sample. It
is essential to know the distribution of the statistics in attempting to analyse the sample statistics.
So far, we have seen various distributions like the hypergeometric distribution, the binomial
distribution, the exponential distribution, the uniform distribution, the Poisson distribution, the
normal distribution. In this section we are going to explore the sample mean which is one of the
most common statistics used in inferential process. To assign and compute the probability of
occurrence of a particular value of the sample mean the researcher must know sample means
distribution. The distribution possibilities can be examined by taking a population with the
particular distribution, selecting the given sample randomly, computing the sample means, and
attempt to determine how the means are distributed.
e.g. The small finite population consist of only N=8 numbers:53 55 59 63 64 69 70.
from this population suppose we take all possible samples of size n = 2
Using an Excel-produced histogram, we can see the shape of the distribution of this population of
data.

103
The result is the following pairs of data.
(53,53) (55,53) (59,53) (63,53)
(53,55) (55,55) (59,55) (63,55)
(53,59) (55,59) (59,59) (63,59)
(53,63) (55,63) (59,63) (63,63)
(53,64) (55,64) (59,64) (63,64)
(53,68) (55,68) (59,68) (63,68)
(53,69) (55,69) (59,69) (63,69)
(53,70) (55,70) (59,70) (63,70)
(64,53) (68,53) (69,53) (70,53)
(64,55) (68,55) (69,55) (70,55)
(64,59) (68,59) (69,59) (70,59)
(64,63) (68,63) (69,63) (70,63)
(64,64) (68,64) (69,64) (70,64)
(64,68) (68,68) (69,68) (70,68)
(64,69) (68,69) (69,69) (70,69)
(64,70) (68,70) (69,70) (70,70)

Mean of each Samples


53 53.5 54.5 57.5 58.5 60 60.5 61.5 62 62.5 56.5 57 59 61 61.5 63.5 64 64.5 58.5 59 61 63
63.5 65.5 66 66.5 59 5 9.5 61.5 63.5 64 66 66.5 67 60 61.5 63.5 65.5 66 68 68.5 69
61.5 62 64 66 66.5 68.5 69 69.5 62 62.5 64.5 66.5 67 69 69.5 70

104
9.10 RELATIONSHIP BETWEEN SAMPLING SIZE AND SAMPLING
DISTRIBUTION
The variability of each sampling distribution decreases as the sample size increases. The range of
the sampling distribution is smaller than the range of the original population.
There is an inverse relationship between standard error and Sampling size. The variability of
sampling distribution decreases as the sample size increases.
Also, as the sample size increases the shape of the sampling distribution becomes more similar to
a normal distribution regardless of the shape of the population.

Problem:
A population has a mean of 400 and a standard deviation of 100. A simple random sample of size
200 will be taken and the sample mean will be used to estimate the population mean.
a. What is the expected value of x?
b. What is the standard deviation of x?
c. Show the sampling distribution of x.
d. What does the sampling distribution of x show?

9.11 POINT ESTIMATOR AND PROPERTIES OF POINT ESTIMATORS


Point estimators are functions that are used to find an approximate value of a population
parameter from random samples of the population. They use the sample data of a population to
calculate a point estimate or a statistic that serves as the best estimate of an unknown parameter of
a population.
Point estimation involves the use of sample data to calculate a single value (known as a point
estimate since it identifies a point in some parameter space) which is to serve as a "best guess" or
"best estimate" of an unknown population parameter (for example, the population mean). More
formally, it is the application of a point estimator to the data to obtain a point estimate.
Point estimation can be contrasted with interval estimation: such interval estimates are typically
either confidence intervals, in the case of frequentist inference, or credible intervals, in the case
of Bayesian inference. More generally, a point estimator can be contrasted with a set estimator.
Examples are given by confidence sets or credible sets. A point estimator can also be contrasted
with a distribution estimator. Examples are given by confidence distributions, randomized
estimators, and Bayesian posteriors.
Eg. The sample standard deviation (s) is a point estimate of the population standard deviation (?).
The sample mean (sx) is a point estimate of the population mean, m.
2 2
The sample variance (s ) is a point estimate of the population variance (s ).

105
9.12 PROPERTIES OF POINT ESTIMATOR
Following are the main properties of Point Estimator:
1. Bias
It is the difference between the expected value of the estimator and the value of the parameter
which is estimated. When the estimated value of the parameter and the value of the parameter
being estimated are equal, the estimator is considered unbiased.
Also, the closer the expected value of a parameter is to the value of the parameter being
measured, the lesser the bias is.

2. Consistency
Consistency tells us how close the point estimator stays to the value of the parameter as it
increases in size. The point estimator requires a large sample size for it to be more consistent and
accurate.
You can also check if a point estimator is consistent by looking at its corresponding expected
value and variance. For the point estimator to be consistent, the expected value should move
toward the true value of the parameter

3. Most efficient or unbiased


The most efficient point estimator is the one with the smallest variance of all the unbiased and
consistent estimators. The variance measures the level of dispersion from the estimate, and the
smallest variance should vary the least from one sample to the other.
Generally, the efficiency of the estimator depends on the distribution of the population. For
example, in a normal distribution, the mean is considered more efficient than the median, but the
same does not apply in asymmetrical distributions
A point estimator is a statistic used to estimate the value of an unknown parameter of a
population. It uses sample data when calculating a single statistic that will be the best estimate of
the unknown parameter of the population.
On the other hand, interval estimation uses sample data to calculate the interval of the possible
values of an unknown parameter of a population. The interval of the parameter is selected in a
way that it falls within a 95% or higher probability, also known as the confidence interval.
The confidence interval is used to indicate how reliable an estimate is, and it is calculated from
the observed data. The endpoints of the intervals are referred to as the upper and lower
confidence limits.
EXAMPLES:
Develop a frame for the population of each of the following research projects.
A. Measuring the Emotional Intelligence and wellbeing of CMM 5 level and highly CRYSIL
rated 15 IT company's employees.
B. Conducting a telephone survey in the capitals of various states of India, to determine the
famous tourist places they are interested in.
Make a list of 500 people from the above cities through the telephone Include men and
women, of various ages, various educational levels, and so on. Number the list and then use

106
the random number list to select 100 people randomly from your list. How representative
of the population is the sample? Find the proportion of men and women in your population
and in your sample. How do the proportions compare? Find the proportion of 25-50 -year-
olds in your sample and the proportion in the population. How do they compare?

C. Interviewing passengers of an Indian Airline about its food and other services provided to
passengers.
Studying the employee green behaviour of the manufacturing sector MSMEs (Micro,
Small and Medium Enterprises) on the basis of Investment and turnover. How they follow
the organizational green policy.
D. A city's telephone book lists 10,00,000 people. If the telephone book is the frame for a
study, how large would the sample size be if systematic sampling were done on every 200th
person?

SELF ASSESSMENT QUESTIONS


Long answer questions
1. Explain the meaning of non-random sampling methods.
2. What do you mean by sampling. Explain with the help of examples
3. Explain different types of random sampling methods

107
Short answer type questions
1. What do you mean by quota sampling?
2. What do you mean by purposive sampling?
3. What do you mean by stratified sampling?
4. What do you mean by population and sample?

Fill in the blanks


1. In ……………….. Sampling, elements for the sample are selected at the convenience of
the researcher. (Convenience)
2. Like ……………….. sampling in cluster sampling, the population is divided into groups
based on similar characteristics or particular characteristics. (Stratified)
3. ……………… sampling implies that chance enters into the process of selection.
(Random)

REFERENCES:
¡ Black, K. (2019). Business statistics: for contemporary decision making. John Wiley &
Sons.
¡ Anderson, D. R., Sweeney, D. J., Williams, T. A., Camm, J. D., & Cochran, J. J. (2012).
Quantitative Methods for Business (Book Only). Cengage Learning.
¡ Everitt, B. S.; Skrondal, A. (2010), The Cambridge Dictionary of Statistics, Cambridge
University Press.
¡ Gonick, L. (1993). The Cartoon Guide to Statistics. HarperPerennial.
¡ Kotz, S.; et al., eds. (2006), Encyclopedia of Statistical Sciences, Wiley.
¡ Levine, D. (2014). Even You Can Learn Statistics and Analytics: An Easy to Understand
Guide to Statistics and Analytics 3rd Edition. Pearson FT Press

108
MODULE - 10 HYPOTHESIS TESTING

STRUCTURE
¡ Research hypothesis
¡ Definition of hypothesis
¡ Types of Hypotheses
¡ Research Hypotheses
¡ Statistical Hypotheses
¡ Differences between null and alternative hypotheses in summary
¡ Substantive hypotheses
¡ Type I and Type II errors
¡ Power of Test
¡ Normal Distribution Hypothesis Test
¡ The t distribution

10.1 LEARNING OUTCOMES


This chapter content will help you understand
1. Null and Alternative Hypothesis,
2. Type I and TypeII errors.
3. Power of a test.
4. Hypothesis testing using the normal distribution.
5. The t distribution.
6. One sample, paired and independent samples t-tests.

10.2 RESEARCH HYPOTHESES


In the field of research or business, searching for answers to questions is routine. To reach
towards their final aim researchers begin with developing tentative answers of their questions, in
research it is termed as "Hypothesis". By beginning with theories that are already widely
accepted and understood, hypotheses aim to clarify the issue and direct attention toward finding
a solution. In the modern world, theory and facts are inextricably interwoven, necessitating
continuous stimulation of facts by theory and theory by realities. Existing ideas are rejected or
reformulated in response to facts, which prompts the development of new theories. When facts
are included, theories might be redefined or clarified.

109
10.2.1 DEFINITION OF HYPOTHESIS
"Any supposition which we make in order to deduce conclusions in accordance with facts which
are known to be real". - Mill
"A hypothesis is an attempt at explanation : a provisional supposition made in order to explain
scientifically some facts or phenomenon." - Coffey
"A proposition which can be put to test to determine validity." - Goode and Hatt

10.3 TYPES OF HYPOTHESES


Three types of hypotheses that will be explored here:
1. Research hypotheses
2. Statistical hypotheses
3. Substantive hypotheses
Although much of the focus will be on testing statistical hypotheses, it is also important
for business decision makers to have an understanding of both research and substantive
hypotheses.

10.3.1 RESEARCH HYPOTHESES


Business researchers frequently have a notion or theory about how the study will turn out before
they start it based on experience or prior work. Research hypotheses are concepts that have been
established prior to an experiment or study being carried out. Research hypotheses are most
similar to the earlier-described theories. It is a prediction of what the researcher expects a study
or experiment to reveal.
Some examples of research hypotheses in business might include:
¡ Older workers are more loyal to a company.
¡ Companies with more than $1 billion in assets spend a higher percentage of their annual
budget on advertising than do companies with less than $1 billion in assets.
¡ The implementation of a Six Sigma quality approach in manufacturing will result in
greater productivity.
¡ The price of scrap metal is a good indicator of the industrial production index six months
later.
¡ Airline company stock prices are positively correlated with the volume of OPEC oil
production.

10.3.2 STATISTICAL HYPOTHESES


To scientifically test research hypotheses, a more formal hypothesis structure needs to be set up
using statistical inferences.
Suppose business researchers want to "prove" the research hypothesis that older workers are
more loyal to a company. A "loyalty" survey instrument is either developed or obtained. If this
instrument is administered to both older and younger workers, how much higher do older

110
workers have to score on the "loyalty" instrument (assuming higher scores indicate more loyal)
than younger workers to prove the research hypothesis? What is the "proof threshold"? Instead of
attempting to prove or disprove research hypotheses directly in this manner, business researchers
convert their research hypotheses to statistical hypotheses and then test the statistical hypotheses
using standard procedures.
A null hypothesis plus an alternate hypothesis make up statistical hypotheses. These two sections
are designed to include every outcome that could possibly result from the experiment or study.
¡ Null Hypothesis
Generally, the null hypothesis asserts that no statistical significance can be found in a collection
of provided observations, that is, the old theory is still true, the old standard is correct, and the
system is in control. Null hypothesis is represented as H0.
We can reject the null hypothesis if the sample contains sufficient data to refute the assertion that
there is no effect in the population (p £ a). If not, we are unable to rule out the null hypothesis.
Although it may sound odd, statisticians only accept the phrase "fail to reject." Avoid using
words like "prove" or "accept" when referring to the null hypothesis. Often, null hypotheses use
words like "no effect," "no difference," "no association" or "no relationship." They are always
expressed mathematically with an equality (usually =, but sometimes ³ or £).
Examples of alternative hypotheses:

Sr no Research question Null hypothesis (H0)


Does tooth flossing affect the number Tooth flossing has no effect on the number
1 of cavities? of cavities.

Does the amount of text highlighted The amount of text highlighted in the
2 in the textbook affect exam scores? textbook has no effect on exam scores.

Does daily meditation decrease the Daily meditation does not decrease the
3 incidence of depression? incidence of depression.

¡ Alternative hypothesis
The alternative hypothesis, on the other hand, states that there are changes in the old theory and
the new theory is true, there are new standards, the system is out of control, and/or something is
happening. The alternate response to your research question is the alternative hypothesis. It
asserts that the populace is affected. Generally, the alternative hypothesis is represented as Ha.
The complement of the null hypothesis is the alternative hypothesis. The extensive nature of null
and alternative hypotheses ensures that they account for all potential outcomes. Additionally,
they are mutually exclusive, thus only one of them may be true at once.
Phrases like "an effect," "a difference," "an association" or "a relationship" are frequently used in
alternative hypotheses. In mathematics, null hypotheses are always expressed as an inequality
(usually ?, but sometimes < or >). Alternative hypotheses can be expressed in a variety of ways,
just like null hypotheses.

111
Examples of alternative hypotheses:
Sr no Research question Alternative hypothesis (Ha)
1 Does tooth flossing affect the number Tooth flossing has an effect on the
of cavities? number of cavities.

2 Does the amount of text highlighted The amount of text highlighted in the
in a textbook affect exam scores? textbook has an effect on exam scores.
3 Does daily meditation decrease the Daily meditation decreases the incidence
incidence of depression? of depression.

10.3.4 DIFFERENCES BETWEEN NULL AND ALTERNATIVE HYPOTHESES IN


SUMMARY

Null hypotheses (H0) Alternative hypotheses (Ha)


Definition A claim that there is no effect in A claim that there is an effect in
the population. the population.
Also known as H0 Ha
H1
Typical No effect An effect
phrases used No difference A difference
No association An association
No relationship A relationship
No change A change
Does not increase Increases
Does not decrease Decreases
Symbols used Equality symbol (=, ≥, or ≤) Inequality symbol (≠, <, or >)

p≤α Rejected Supported


p>α Failed to reject Not supported

10.3.5 SUBSTANTIVE HYPOTHESES


A business researcher concludes based on the data gathered throughout the study when testing a
statistical hypothesis. It is typical to refer to a result as statistically significant if the null
hypothesis is rejected and as a result, the alternative hypothesis is accepted. For statisticians and
business researchers, the term "significant" simply denotes that the null hypothesis has been
chosen to be rejected and that the experiment's outcome is unlikely to have been the result of
chance. However, the word "significant" is more frequently used in common business contexts to
mean "important" or "a substantial amount." One issue that might occur when testing statistical
hypotheses is that some data qualities may provide a conclusion that is statistically significant
but has no bearing on the business. Business decision-makers must establish what, to them, is a
substantive result in addition to knowing a statistically significant result. When a statistical
analysis yields conclusions that are significant to the decision maker, it is said to have produced a
'substantive result'. Decision-makers and business researchers alike should be mindful that
results that are statistically significant are not always conclusive results.

112
10.4 TYPE I AND TYPE II ERRORS
Statistical Principles of Hypothesis Testing
Researcher can use null and alternative hypotheses with hypothesis testing to see whether the
data confirm or disprove the study expectations; however, one can never completely confirm (or
refute) the idea, regardless how many facts one gathers. Drawing conclusions about phenomena
in the population from events observed in the sample will always be necessary (Hulley et al.,
2001).
The null hypothesis, which forms the basis of hypothesis testing, holds that there is no difference
between groups or correlation between variables in the population. It is always accompanied by a
counterargument, which is your research's forecast of a real difference between groups or a real
correlation between variables.
Let us consider: Researcher tests whether a modified Car model can satisfy the demands of
economic class consumers.
In this case:
¡ The null hypothesis (H0): Modified Car model cannot satisfy the demands of economic
class consumers.
¡ The alternative hypothesis (H1): Modified Car model can satisfy the demands of economic
class consumers.
A researcher can use HTAB System to test the Hypotheses, which involves four major tasks
(Figure 1):
Task 1. Establishing the Hypotheses
Task 2. Conducting the Test
Task 3. Taking statistical Action
Task 4. Determining the Business implications

Figure 10.1 : HTAB System of hypothesis Testing


Source: Book name, 296

113
Post to applying the HTAB system, the possible statistical outcomes of a study can be divided
into two groups:
1. Those that cause the rejection of the null hypothesis
2. Those that do not cause the rejection of the null hypothesis
The rejection region refers to the conceptual and visual area where statistical results that lead to
the rejection of the null hypothesis are located. The nonrejection region refers to statistical
outcomes that do not lead to the null hypothesis being rejected. (Figure 2) The researcher will
reject the null hypothesis if any sample mean falls within that range. The business researcher will
opt not to reject the null hypothesis if the sample means that fall between the two crucial values
are sufficiently similar to the population mean. These methods fall into the category of
nonrejection. Researcher decides whether the null hypothesis can be rejected based on the data
and the results of a statistical test. Since these decisions are based on probabilities, there is always
a risk of making the wrong conclusion.

Figure 10.2 : Rejection and Nonrejection Regions


Source: Book name, 297

If the results are statistically significant, the null hypothesis must be incorrect in order for them to
hold true. Researcher would then reject the null hypothesis in this situation. But occasionally,
this might be a Type I error.
If the results are not statistically significant, the null hypothesis is likely to be correct and they
have a high probability of occurring. As a result, null hypothesis is not rejected. But occasionally,
this might be a Type II error.

Considering earlier example: Type I and Type II errors


¡ A Type I error happens when you get false positive results: Researcher conclude that the
modified car model can satisfy the demands of economic class consumers, when, in fact, it
did not. These improvements could have been the result of measurement errors or other
random variables.

114
¡ A Type II error happens when you get false negative results: Researcher conclude that the
modified car model cannot satisfy the demands of economic class consumers, when, in
fact, it did. Your study might have overlooked important signs of progress or mistakenly
attributed any advancements to unrelated sources.

10.4.1 TYPE I ERROR


Rejecting the null hypothesis when it is in fact true is a Type I error. It entails drawing conclusions
about outcomes that are statistically significant when, in fact, they were just the consequence of
chance or unrelated causes.
The significance level you select (alpha or a) determines the likelihood that you will make this
mistake. You determined that figure at the start of your research to determine the statistical
likelihood of getting your results (p value).
The typical significance level is 0.05 or 5%. If the null hypothesis is accurate, your results only
have a 5% chance or less of occurring. The findings of your test are statistically significant and
compatible with the alternative hypothesis if the p value is less than the significance level. Your
results are deemed statistically non-significant if your p value is larger than the significance
level.

10.4.2 TYPE II ERROR


When the null hypothesis is ignored even when it is untrue, this is known as a Type II error.
Because hypothesis testing can only tell you whether you reject the null hypothesis, this is not
precisely the same as "accepting" the null hypothesis.
An alternative definition of a Type II error is failing to recognise an effect when there actually
was one. In truth, it's possible that your study lacked the statistical power to identify an effect of a
particular size.
Power measures how well a test can identify an actual effect when it exists. Typically, a power
level of 80% or above is regarded as appropriate. The statistical power of a study is inversely
correlated with the likelihood of a Type II error. The likelihood of making a Type II error
decreases with increasing statistical power.

[Link] POWER OF TEST


Bullard F (2009) explains interpretation of power in multiple ways:
¡ Power is the probability of rejecting the null hypothesis when, in fact, it is false.
¡ Power is the probability of making a correct decision (to reject the null hypothesis) when
the null hypothesis is false.
¡ Power is the probability that a test of significance will pick up on an effect that is present.
¡ Power is the probability that a test of significance will detect a deviation from the null
hypothesis, should such a deviation exist.
¡ Power is the probability of avoiding a Type II error.
A binary hypothesis test's power in statistics refers to the likelihood that it will correctly reject the
null hypothesis when a particular alternative hypothesis is true. It is commonly denoted by 1 - b,

115
and represents the chances of a true positive detection conditional on the actual existence of an
effect to detect. Statistical power ranges from 0 to 1, and as the power of a test increases, the
probability b of making a type II error by wrongly failing to reject the null hypothesis decreases.
The formula for power is 1 - b. A hypothesis test's power ranges from 0 to 1, and if it's close to 1,
it's very effective in identifying a false null hypothesis. Beta is typically set at 0.2 but may be set
smaller by the researchers.
Greater sample size, effect sizes, and significance levels all result in increased power for the
researcher. Although there are other factors, such as variance (s2).

Figure 3: Illustration of the power and the significance level of a statistical test, given the
null hypothesis (sampling distribution 1) and the alternative hypothesis (sampling
distribution 2)
Source: Wikipdia.org_Power of test

10.5 NORMAL DISTRIBUTION HYPOTHESIS TEST


The binomial distribution and normal distribution hypothesis tests, which are the two main types
of hypothesis testing. In binomial hypothesis tests, the probability parameter pp is being
examined. In typical hypothesis tests, µ, the mean parameter, is being tested. This provides us
with a crucial distinction that we can utilise to decide what test to run when. Words like mean,
average, and overall are all indicators that you should employ a normal hypothesis test because
they all test a mean parameter. If you would model it with a normal distribution, you wish to
perform a normal hypothesis test. The scenario and context should also be taken into
consideration.
Similar to binomial distribution hypothesis tests, normal distribution hypothesis tests can also be

116
performed by switching the test statistic. These tests are valuable because they enable us to verify
statements made about usually dispersed things. We consider looking at the mean of a sample
from a population when we hypothesise test for the mean of a normal distribution.
Steps to follow in Normal Distribution Hypothesis Testing:
1. Define the parameter in the context of the question - for a normal hypothesis test the
parameter is µ which is always the mean of something.
2. Write down the null hypothesis and the alternate hypothesis.
3. Define the test statistic X in the context of the question.
4. Write down the distribution of X under the null hypothesis.
5. State the significance level a - even though you are likely given it in the question, not
stating it risks losing a mark.
6. Test for significance or find the critical region.
7. Write a concluding sentence, linking the acceptance or rejection of H0 to the context.

10.6 IN CASE OF MULTIPLE OBSERVATIONS


We can generate a more accurate result and have a broader critical region if we are provided with
numerous observations to base our hypothesis test on.
If we have a normally distributed variable X ˜ N (µ,z2), the average of n observations of X has the
distribution X ˜ (µ, z2/n).
This implies that, in the case of several observations, the hypothesis test can be performed using
the average of the observations rather than just one. Since this has a bigger critical region than
just a single observation, it is in fact much superior.

10.7 THE T DISTRIBUTION


The sample standard deviation must be employed in the estimate process when the population
standard deviation is unknown. In this section, a statistical method for estimating a population
mean using the sample mean is described when it is unknown what the population standard
deviation will be.
Assume a business researcher wants to calculate the average flight time of a 767 jet from New
York to Los Angeles. The business researcher is unlikely to know the population standard
deviation because she does not know the population mean or average time. The researcher can
construct the estimate by taking a random sample of flights and computing a sample mean and
standard deviation. Another business researcher is using a random sample of people to study the
impact of movie video advertisements on consumers.
The researcher wants to estimate the population mean response but has no idea what the
population standard deviation is. To perform this analysis, he will have access to the sample
mean and standard deviation. When the population standard deviation is unknown, the z
formulas are inapplicable (and is replaced by the sample standard deviation). Instead, a British
statistician named William S. Gosset devised another mechanism to deal with such cases. Gosset
was Karl Pearson's student and a close personal friend. Gosset published his initial research on
the t test under the alias "Student." Because of this, the t test is sometimes known as the student's t

117
test. Gosset made an important contribution since it paved the way for more precise statistical
tests, which some experts claim signalled the start of the contemporary age in mathematical
statistics**.
The formula for the t statistic is:

Footnote: ** Adapted from "Behavioral Statistics: An Historical Perspective," by Arthur L.


Dudycha and Linda W. Dudycha, in "Statistical Issues: A Reader for the Behavioral Sciences,"
edited by Roger Kirk and published by Brooks/Cole in Monterey, California, in 1972.
Because every sample size has a different distribution, the possibility of numerous t tables exists,
the t distribution is actually a series of distributions. Only a few critical numbers are shown, and
each line in the table includes values from a distinct t distribution in order to make these t values
easier to handle. The t statistic is used under the presumption that the population is regularly
distributed. Nonparametric methods should be utilised if the population distribution is either
abnormal or unknown.
While testing hypothesis you know you need to use a t-test, but you're not sure which one to use.
In such case check for followings: Is the objective to compare the means of two groups, or is it
limited to how the mean of a single group compares to some number?
Use a one-sample t-test if you are only interested in how the mean of a single group compares to a
single number. If one is testing whether the average student consumes significantly more than
2000 calories per day, a one-sample t-test is appropriate (e.g., you are comparing the mean
number of calories consumed to see whether it is significantly greater than the number 2000). A
one-sample t-test is used to compare a single population to a standard value (for example, to
determine whether the average lifespan of a specific town is different from the country average).
If you're comparing the means of two groups, ask yourself: Did the two groups of numbers come
from the same people? If this is the case, we must employ a paired-samples t-test (also known as a
repeated-samples t-test). A paired t-test is used to compare a single population before and after
some experimental intervention or at two different points in time (for example, measuring
student performance on a test before and after being taught the material). A paired-samples t-test
is also appropriate if the researcher is conducting a "matched" design in which they purposefully
chose pairs of subjects that are similar in various characteristics (e.g., age, gender, medical
history, etc.) A paired-samples t-test is appropriate whenever numbers in the first and second
groups are paired and there is a meaningful relationship between a value in the first group of
scores and the corresponding value in the second group of scores.
In all other cases where a t-test is appropriate, an independent-samples t-test is preferable. This is
suitable for "between-subjects" designs in which two groups of subjects are expected to differ on
a critical manipulation. For example, if you were testing the effect of caffeine on plant growth,
you could have two groups: one control group given water and one experimental group of plants
given a caffeine solution. You should use an independent-samples t-test because there is no
meaningful pairing between the scores in the two groups because you are using completely
different plants in each group.

118
SELF ASSESSMENT QUESTIONS
Long answer questions
1. What do you mean by hypothesis testing
2. Explain different kinds of hypothesis.
3. Define with the help of examples two major research hypothesis
4. A study was made to compare the costs of supporting a family of four Americans for a year
in different foreign cities. The lifestyle of living in the United States on an annual income
of $75,000 was the standard against which living in foreign cities was compared. A
comparable living standard in Toronto and Mexico City was attained for about $64,000.
Suppose an executive wants to determine whether there is any difference in the average
annual cost of supporting her family of four in the manner to which they are accustomed
between Toronto and Mexico City. She uses the following data, randomly gathered from
11 families in each city, and an alpha of .01 to test this difference. She assumes the annual
cost is normally distributed and the population variances are equal. What does the
executive find?
Toronto Mexico City
$69,000 $65,000
64,500 64,000
67,500 66,000
64,500 64,900
66,700 62,000
68,000 60,500
65,000 62,500
69,000 63,000
71,000 64,500
68,500 63,500
67,500 62,400

5. A company's auditor believes the per diem cost in Nashville, Tennessee, rose significantly
between 1999 and 2009. To test this belief, the auditor samples 51 business trips from the
company's records for 1999; the sample average was $190 per day, with a population
standard deviation of $18.50. The auditor selects a second random sample of 47 business
trips from the company's records for 2009; the sample average was $198 per day, with a
population standard deviation of $15.60. If he uses the risk of committing a Type I error of
.01, does the auditor find that the per diem average expense in Nashville has gone up
significantly?

6. Employee suggestions can provide useful and insightful ideas for management. Some
companies solicit and receive employee suggestions more than others, and company
culture influences the use of employee suggestions. Suppose a study is conducted to
determine whether there is a significant difference in a mean number of suggestions a
119
month per employee between the Canon Corporation and the Pioneer Electronic
Corporation. The study shows that the average number of suggestions per month is 5.8 at
Canon and 5.0 at Pioneer. Suppose these figures were obtained from random samples of 36
and 45 employees, respectively. If the population standard deviations of suggestions per
employee are 1.7 and 1.4 for Canon and Pioneer, respectively, is there a significant
difference in the population means? Use = .05.

Short answer questions


1. What do you mean by hypothesis
2. Explain null hypothesis
3. Explain alternative hypothesis
4. Explain type 1 errors
5. Explain type 2 errors

Fill in the blanks


1. When the null hypothesis is ignored even when it is untrue, this is known as a ……….
error. (Type II)
2. The ……………… hypothesis asserts that no statistical significance can be found in a
collection of provided observations. (Null hypothesis)
3. A technique of sampling used to select elements from the population by various
mechanism that does not involve a random selection, this process is ……………………
sampling. (Non-probability)

REFERENCES:
¡ Hulley S. B, Cummings S. R, Browner W. S, Grady D, Hearst N, Newman T. B. 2nd ed.
Philadelphia: Lippincott Williams and Wilkins; 2001. Getting ready to estimate sample
size: Hypothesis and underlying principles In: Designing Clinical Research-An
epidemiologic approach; pp. 51-63
¡ Bullard, F. A. (2009). Exoplanet detection: A comparison of three statistics or how long
should it take to find a small planet? (Doctoral dissertation).
¡ Piaw, C. Y. (2013). Mastering research statistics. Malaysia: McGraw Hill Education, New
York, United States.
¡ Wasserman, L. (2004). All of statistics: a concise course in statistical inference (Vol. 26).
New York: Springer.
¡ Keener, R. W. (2010). Theoretical statistics: Topics for a core course. New York: Springer.
¡ [Link]

120
MODULE - 11 NON-PARAMETRIC TESTS

STRUCTURE
¡ Introduction To Non-parametric Test
¡ Chi-square goodness of - fit test
¡ Chi-square test of independence

LEARNING OUTCOMES
¡ To understand the meaning of non-parametric tests
¡ To understand the use of chi square test as goodness of fit
¡ To understand the use of chi square test as test of independence

11.1 INTRODUCTION TO NON-PARAMETRIC TEST


To examine the probability of multinomial distribution trials along a single dimension, apply the
chi-square goodness-of-fit test. For instance, if economic class is the variable under study and
there are three possible outcomes-lower income class, middle income class, and upper income
class-then economic class is the single dimension and the three outcomes are the three classes.
One and only one of the possible outcomes can happen on each trial. In other words, a family unit
can only belong to one income class-lower income, middle income, or upper income-and cannot
belong to more than one.
In order to assess whether there is a discrepancy between what was predicted and what was
observed, the chi-square goodness-of-fit test compares the expected, or theoretical, frequencies
of categories from a population distribution to the observed, or actual, frequencies from the
distribution. Officials from the airline business, for instance, may hypothesise that the age
distribution of those buying airline tickets is a certain manner. An actual sample of ticket buyers'
ages can be randomly selected in order to confirm or refute this expected distribution, and the
observed results can then be compared to the predicted results using the chi-square goodness-of-
fit test. The observed arrivals at bank teller windows can also be tested to see if they follow the
expected Poisson distribution.

11.2 FORMULA FOR CHI-SQUARE GOODNESS OF - FIT TEST

Where,
fo = frequency of observed values
fe = frequency of expected values
k = number of categories

121
c = number of parameters being estimated from the sample data
Across the distribution, this formula compares the frequency of observed values to the frequency
of expected values. Since the observed total drawn from the sample is utilised as the total for the
expected frequencies, the test loses one degree of freedom because the total number of expected
frequencies must equal the total number of observed frequencies.
To establish the frequency distribution of predicted values, a population parameter, such as, or, is
occasionally computed from the sample data. This estimation loses a degree of freedom every
time it happens. In general, k - 1 degrees of freedom are employed in the test when a uniform
distribution is used as the expected distribution or when an expected distribution of values is
provided. The degrees of freedom are k - 2 since estimating ? eliminates a third degree of freedom
when determining whether a distribution is Poisson or not. In testing to determine whether an
observed distribution is normal, the degrees of freedom are k - 3 because two additional degrees
of freedom are lost in estimating both µ and ? from the observed sample data.
In 1900, Karl Pearson developed the chi-square test. The chi-square distribution has an infinitely
long positive tail since it is the sum of the squares of k independent random variables, which
means it can never be smaller than zero. The degrees of freedom (df) associated with each
distribution in the family of chi-square distributions determine how they behave. The chi-square
distribution is noticeably tilted to the right for tiny df values (positive values). The chi-square
distribution starts to resemble the normal curve as the df rises.
Let us discuss how can the chi-square goodness-of-fit test be applied to business situations.
Example: In a survey conducted by National Research Agency on organised retail sector
customers were asked: "In general, how would you rate the level of service provided by
organised retail outlets in India?" The distribution of responses to this question was as follows:
Excellent 8%
Pretty good 47%
Only fair 34%
Poor 11%
Assume that Big bazar store manager wants to find out whether the results of this consumer
survey apply to their customers. Hence, store manager conducted 207 interviews of their
customers. The customers were questioned about how they would rank the quality of service at
the supermarket they had just left. The response categories were kept similar which are excellent,
pretty good, only fair, and poor. The observed responses from this study are:
Excellent 21%
Pretty good 109%
Only fair 62%
Poor 15%
The store manager can now apply a chi-square goodness-of-fit test to check whether the
observed response frequencies from this survey match those that would be anticipated based on
the results of the national survey.
Solution:
STEP 1. The hypotheses for this example follows:
Ho: The observed distribution is the same as the expected distribution.

122
Ha: The observed distribution is not the same as the expected distribution
STEP 2. The statistical test being used is:

STEP 3. Let a = .05


STEP 4: Because a chi-square of 0 denotes perfect agreement between distributions, chi-square
goodness-of-fit tests are one-tailed. Due to the fact that chi-square can never be negative and is
always determined by the sum of squared values, any departure from zero difference always goes
in the positive direction.
With four categories in this example (excellent, pretty good, only fair, and poor), k = 4
The degrees of freedom: k - 1 = 4 - 1 = 3
For a = .05 and df = 3, the critical chi-square value is:
2
X 0.5,3 = 7.8147
After the data are analyzed, an observed chi-square greater than 7.8147 must be computed
in order to reject the null hypothesis.
STEP 5. The observed values gathered in the sample data n = 207. The expected proportions are
given, but the expected frequencies must be calculated by multiplying the expected proportions
by the sample total of the observed frequencies, as shown below:

Table 1: Construction of Expected Values for Service Satisfaction Study

Expected Frequency (fe)


Response Expected proportion (proportion * sample total)
Excellent 0.08 .08*207 16.56
Pretty
good 0.47 .47*207 97.29
Only fair 0.34 .34*207 70.38
Poor 0.11 .11*207 22.77

STEP 6. The chi-square goodness-of-fit can then be calculated as shown in below table

Table 2: Calculation of Chi-Square


(fo - fe)2 / fe
Response fo fe
Excellent 21 16.56 1.19
Pretty good 109 97.29 1.41
Only fair 62 70.38 1
Poor 15 22.77 2.65
207 207 6.25

123
STEP 7. As the observed value of the chi-square of 6.25 is not greater than the critical table value
of 7.8147, the store manager will not reject the null hypothesis.
Practical Implications:
STEP 8. The data gathered in the sample of 207 supermarket shoppers indicate that
the distribution of responses from supermarket shoppers in the manager's city is not significantly
different from the distribution of responses to the national survey.
The store manager may conclude that their customers do not appear to have attitudes
different from those people who took the survey.

11.3 CHI SQUARE -TEST OF INDEPENDENCE


The chi-square goodness-of-fit test is used to determine whether the distribution of frequencies
for categories of one variable, such as age or number of bank arrivals, is the same as some
hypothesised or expected distribution. The goodness-of-fit test, on the other hand, cannot be
used to analyse two variables at the same time. The Chi-square test of independence determines
whether or not two variables are likely to be related. We have counts for two nominal or
categorical variables. We also believe that the two variables are unrelated. The test allows us to
determine whether or not our idea is plausible.
To determine whether two variables are independent, a different chi-square test, the chi-square
test of independence, can be used to analyse the frequencies of two variables with multiple
categories. This type of analysis is frequently desirable. A market researcher, for example, might
want to know whether the type of soft drink preferred by a consumer is unrelated to the
consumer's age.
The chi-square test of independence can be used to analyze any level of data measurement, but it
is particularly useful in analyzing nominal data. Following are the types of questions that may get
asked in business research:
¡ In which region of the country do you reside?
A. Northeast B. Midwest C. South D. West
¡ Which type of financial investment are you most likely to make today?
A. Stocks B. Bonds C. Treasury Bills
On a questionnaire, the following two questions might be used to measure geographic region and
type of financial investment. The business researcher would tally the frequencies of responses to
these two questions into a two-way table called a contingency table. Because the chi-square test
of independence uses a contingency table, this test is sometimes referred to as contingency
analysis.
If the two variables are independent, they are not related. The chi-square test of independence is,
in some ways, a test of whether the variables are related. A chisquare test of independence's null
hypothesis is that the two variables are independent (not related). If the null hypothesis is
rejected, the two variables are not independent and are related.
Two variables are required for the Chi-square test of independence. Our assumption is that the
variables are unrelated. Here are several examples:
¡ Consider, we have a list of movie genres which is our first variable. Our second variable is
whether or not those genres' patrons purchased snacks at the theatre. Our hypothesis (or, in

124
statistical terms, null hypothesis) is that the type of movie and whether or not people purchased
snacks have no relationship. The owner of the movie theatre wants to know how many snacks to
purchase. When movie types and snack purchases are unrelated, estimating is easier than when
movie types influence snack sales.
¡ A list of dog breeds seen as patients at a veterinary clinic. The second variable is whether
owners feed dry food, canned food, or a combination of the two. Our theory is that dog breeds and
food types are unrelated. If this is correct, the clinic can order food based solely on the total
number of dogs, with no regard for breeds.
Let us elaborate the 1st example: Assume we collect information for 600 people in our theatre.
We know what kind of movie each person saw and whether or not they bought snacks.
To begin, consider whether the Chi-square test of independence is an appropriate method for
evaluating the relationship between movie type and snack purchases.
¡ We have a simple random sample of 600 people who saw a movie at our theatre. We meet
this requirement.
¡ Our variables are the movie type and whether or not snacks were purchased. Both variables
are categorical. We meet this requirement.
¡ The last requirement is for more than five expected values for each combination of the two
variables. To confirm this, we need to know the total counts for each type of movie and the
total counts for whether snacks were bought or not. For now, we assume we meet this
requirement and will check it later.
It appears we have indeed selected a valid method. (We still need to check that more than five
values are expected for each combination.)
Here is our data summarized in a contingency table:

Table 3: Contingency table for movie snacks data

Type of Movie Snacks No Snacks


Action 50 75
Comedy 125 175
Family 90 30
Horror 45 10

Before we proceed, let's double-check the assumption of five expected values in each category.
There are more than five counts in each combination of Movie Type and Snacks in the data. But,
what are the expected counts if movie and snack purchases are made separately?

11.3.1 Finding expected counts


To find expected counts for each Movie-Snack combination, we first need the row and column
totals, which are shown below:

125
Table 4: Contingency table for movie snacks data with row and column totals
Type of Movie Snacks No Snacks Row totals

Action 50 75 125

Comedy 125 175 300

Family 90 30 120

Horror 45 10 55

Column totals 310 290 GRAND TOTAL = 600

The row and column totals are used to calculate the expected counts for each Movie-Snack
combination. We divide the grand total by the sum of the row and column totals. This calculates
the expected number of cells in the table. For the Action-Snacks cell, for example, we have:
125*310 / 600 = 65
We rounded the answer to the nearest whole number. If there is not a relationship between movie
type and snack purchasing we would expect 65 people to have watched an action film with
snacks. Here are the actual and expected counts for each Movie-Snack combination. In each cell
of Table 5 below, the expected count appears in bold beneath the actual count. The expected
counts are rounded to the nearest whole number.

Table 5: Contingency table for movie snacks data showing actual count vs. expected count
Contingency table for movie snacks data showing actual count vs. expected count

Type of No
Snacks Row totals
Movie Snacks

50 75
Action 125
65 60

125 175
Comedy 300
155 145

90 30
Family 120
62 58

45 10
Horror 55
28 27

Column totals 310 290 GRAND TOTAL = 600

These calculated values will be labelled as "expected values," "expected cell counts," or some
other similar term when using software.

126
Because all of the expected counts for our data are greater than five, we can apply the
independence test.
Let's look at the contingency table before calculating the test statistic. The expected counts are
calculated using the row and column totals. Looking at each cell, we can see that some expected
counts are close to the actual counts, but the majority are not. If there is no correlation between
the type of movie and snack purchases, the actual and expected counts will be comparable. If a
relationship exists, the actual and expected counts will differ.

11.3.2 Performing the test


The basic idea behind calculating the test statistic is to compare actual and expected values given
the data's row and column totals. First, we compute the difference between the actual and
expected values for each Movie-Snacks combination. The difference is then squared.
Combinations with fewer actual values than expected are given the same weight as combinations
with more actual values. Then we divide by the combination's expected value. These values are
added up for each Movie-Snacks combination. This provides us with our test statistic.
Table 6 below shows the calculations for each Movie-Snacks combination to two decimal places.

Table 6: Preparing to calculate our test statistic


Type of
Snack No Snacks
Movie

Actual: 50 Actual: 75
Expected: 64.58 Expected: 60.42

Action Difference: 50 – 64.58 = -14.58 Difference: 75 – 60.42 = 14.58


Squared Difference: 212.67 Squared Difference: 212.67
Divide by Expected: 212.67/64.58 = Divide by Expected: 212.67/60.42 =
3.29 3.52

Actual: 125 Actual 175


Expected 155 Expected 145
Comedy
Difference: 125 – 155 = -30 Difference: 175 – 145 = 30
Squared Difference: 900 Squared Difference: 900
Divide by Expected: 900/155 = 5.81 Divide by Expected: 900/145 = 6.21

Actual: 90 Actual: 30
Expected: 62 Expected 58
Family
Difference: 90 – 62 = 28 Difference: 30 – 58 = -28
Squared Difference: 784 Squared Difference: 784
Divide by Expected: 784/62 = 12.65 Divide by Expected: 784/58 = 13.52

Actual: 45 Actual: 10
Expected 28.42 Expected 26.58

Horror Difference: 45 – 28.42 = 16.58 Difference: 10 – 26.58 = -16.58


Squared Difference: 275.01 Squared Difference: 275.01
Divide by Expected: 275.01/28.42 = Divide by Expected: 275.01/26.58 =
9.68 10.35

127
Lastly, to get our test statistic, we add the numbers in the final row for each cell:
3.29+3.52+5.81+6.21+12.65+13.52+9.68+10.35=65.033.29+3.52+5.81+6.21+12.65+13.52+
9.68+10.35=65.03

To make our decision, we compare the test statistic to a value from the Chi-square distribution.
This activity involves five steps:
1. We decide how much risk we're willing to take in assuming that the two variables aren't
independent when they are. Prior to collecting the movie data, we decided that we are
willing to take a 5% risk of claiming that the two variables - Movie Type and Snack
Purchase - are not independent when they are. We set the significance level,, to 0.05 in
statistics speak.
2. We calculate a test statistic. As shown above, our test statistic is 65.03.
3. We find the critical value from the Chi-square distribution based on our degrees of freedom
and our significance level. This is the value we expect if the two variables are independent.
4. The degrees of freedom depend on how many rows and how many columns we have. The
degrees of freedom (df) are calculated as:
df=(r-1)×(c-1)df=(r-1)×(c-1)
In the formula, r is the number of rows, and c is the number of columns in our contingency
table. From our example, with Movie Type as the rows and Snack Purchase as the columns,
we have:
df=(4-1)×(2-1)=3×1=3df=(4-1)×(2-1)=3×1=3
The Chi-square value with a = 0.05 and three degrees of freedom is 7.815.
5. We compare the value of our test statistic (65.03) to the Chi-square value. Since 65.03 >
7.815, we reject the idea that movie type and snack purchases are independent.
We conclude that there is a link between movie genre and snack purchases. Regardless of the
type of movie being shown, the owner of the movie theatre cannot estimate how many snacks to
purchase. Instead, when estimating snack purchases, the owner must consider the type of movies
being shown.
It's important to note that we can't infer that the type of movie influences snack purchases. The
independence test only tells us whether or not there is a relationship; it does not tell us which
variable causes which.

SELF ASSESSMENT QUESTIONS


Long answer questions
1. explain the chi-square test of independence with the help of examples
2. explain the meaning of non-parametric tests
3. Quite often in the business world, random arrivals are Poisson distributed. This
distribution is characterized by an average arrival rate, per some interval. Suppose a teller
supervisor believes the distribution of random arrivals at a local bank is Poisson and sets
out to test this hypothesis by gathering information. The following data represent a
distribution of frequency of arrivals during1-minute intervals at the bank. Use .05 to test

128
these data in an effort to determine whether they are Poisson distributed.
Number of Arrivals Observed Frequencies
0 7
1 18
2 25
3 17
4 12
³5 5

4. Use a chi-square goodness-of-fit test to determine whether the observed frequencies are
distributed the same as the expected frequencies (a = .05).
Category fo fe
1 53 68
2 37 42
3 32 33
4 28 22
5 18 10
6 15 8

5. Use the following data and a = .01 to determine whether the observed frequencies
represent a uniform distribution.
Category fo
1 19
2 17
3 14
4 18
5 19
6 21
7 18
8 18

6. According to an extensive survey conducted for Business Marketing by Leo J. Shapiro &
Associates, 66% of all computer companies are going to spend more on marketing this
year than in previous years. Only 33% of other information technology companies and
28% of non-information technology companies are going to spend more. Suppose a
researcher wanted to conduct a survey of her own to test the claim that 28% of all non-
information technology companies are spending more on marketing next year than this

129
year. She randomly selects 270 companies and determines that 62 of the companies do
plan to spend more on marketing next year. Use a = .05, the chi-square goodness-of-fit test,
and the sample data to test to determine whether the 28% figure holds for all non-
information technology companies.

7. Cross-cultural training is rapidly becoming a popular way to prepare executives for


foreign management positions within their company. This training includes such aspects
as foreign language, revisit orientations, meetings with former expatriates, and cultural
background information on the country. According to Runzheimer International, 30% of
all major companies provide formal cross-cultural programs to their executives being
relocated in foreign countries. Suppose a researcher wants to test this figure for companies
in the communications industry to determine whether the figure is too high for that
industry. In a random sample, 180 communications firms are contacted; 42 provide such a
program. Let a = .05 and use the chi-squaregoodness-of-fit test to determine whether the
.30 proportion for all major companies is too high for this industry.

8. Use the following contingency table to test whether variable 1 is independent of variable 2.
Let a = .01

Variable 2
201 325
Variable 1
68 110

9. Is the transportation mode used to ship goods independent of type of industry? Suppose the
following contingency table represents frequency counts of types of transportation used by
the publishing and computer hardware industries. Analyze the data by using the chi-square
test of independence to determine whether the type of industry is independent of
transportation mode. Let a = .05.

Transportation mode
Air Train Truck
Industry Publishing 32 12 41
Computer 5 6 23
hardware

1.8 REFERENCES
¡ Black, K. (2019). Business statistics: for contemporary decision making. John Wiley &
Sons.
¡ Anderson, D. R., Sweeney, D. J., Williams, T. A., Camm, J. D., & Cochran, J. J. (2020).
Modern business statistics with Microsoft Excel. Cengage Learning.

130
MODULE - 12 ANALYSIS OF VARIANCE (ANOVA)

STRUCTURE
¡ Introduction to anova
¡ Types of anova
¡ Single factor Anova
¡ Two factor Anova

12.1 LEARNING OUTCOMES


1. To understand ANOVA method
2. To understand how to conduct single factor Anova
3. To understand how to conduct two factor Anova

12.2 INTRODUCTION TO ANOVA


Analysis of variance (ANOVA) is a statistical tool that splits the total variability found inside a
data set into two parts: the factors that influence and those which do not, The former are known as
the systematic factors and the latter as random factors. The ANOVA test determines independent
variables' influence on the dependent variable in a regression study.
Ronal Fisher created ANOVA in the 1920s. Till then the t- and z-test methods were used for
statistical analysis. Thus, ANOVA is also called the Fisher analysis of variance and is considered
an extension of the t- and z-tests. The Formula for ANOVA is:

F = MSE divided by MST


where:
F=ANOVA coefficient
MST=Mean sum of squares due to treatment
MSE=Mean sum of squares due to error

12.3 WHAT DOES ANOVA DO?


ANOVA and the F statistic (also called the F-ratio), determine the comparison and relationship
between two and more than two groups simultaneously. It determines the variability between
samples and within samples.
The result of the ANOVA's F-ratio statistic will be close to 1 if no real difference exists between
the tested groups, called the null hypothesis. The distribution of all possible values of the F
statistic is known as F-distribution. It is a group of distribution functions, with two characteristic
numbers, called the numerator degrees of freedom and the denominator degrees of freedom.

131
12.4 TYPES OF ANOVA
One-way or two-way ANOVA refers to the number of independent variables in the analysis of the
variance test.
With a one-way ANOVA, we have one independent variable affecting a dependent variable. It is
used to search for statistically significant differences between two or more independent variables
A two-way ANOVA is an extension of the one-way ANOVA. With a two-way ANOVA, two
independents are affecting a dependent variable. For example, a two-way ANOVA allows
comparing productivity based on two independent variables, such as salary and skill set. It is
utilized to observe the interaction between the two factors and test the effect of two factors at the
same time. Thereby the potential interaction of two independent variables on one dependent
variable is revealed.
A three-way ANOVA, also known as three-factor ANOVA, is a statistical means of determining
the effect of three factors on an outcome.
There is a variation of ANOVA for example, MANOVA (multivariate ANOVA). It differs from
ANOVA as the former tests for multiple dependent variables simultaneously while the latter
assesses only one dependent variable at a time.
ANOVA has many applications in finance, economics, science, medicine, and social science.
ANOVA is used in finance in several different ways, such as to forecast the movements of
security prices by first determining which factors influence stock fluctuations. This analysis can
provide valuable insight into the behavior of a security or market index under various conditions.
A researcher might, for example, test students from multiple colleges to see if students from one
of the colleges consistently outperform students from the other colleges. In a business
application, an R&D researcher might test two different processes of creating a product to see if
one process is better than the other in terms of cost efficiency. In medical sciences, to compare the
effects of different treatment protocols on patient outcomes; in social science research (for
instance to assess the effects of gender and class on specified variables), in software engineering
(for instance to evaluate database management systems), in manufacturing (to assess product
and process quality metrics), and industrial design among other fields.
With ANOVA, a researcher can determine whether the variability of the outcomes is due to
chance or the factors in the analysis. ANOVA analysis is considered to be accurate than t testing
because it is flexible and requires fewer observations. It is better suited for use in complex
analyses than those that can be assessed by conducting tests. ANOVA testing allows researchers
to uncover relationship.

12.5 SINGLE FACTOR ANOVA (OR ONE-WAY ANOVA)


Let's consider a fictitious problem (refer to Table1): 21 students at the autonomous university of
Delhi, Kolkata, and Chennai are selected for a skill assessment test. Seven first-year, seven
second-year, and seven third-year undergraduate students are randomly selected. The students
are given a skill assessment test, with a maximum score of 100.

132
Table 1: ANOVA single factor table

We are interested to know whether a difference exists somewhere between the three different
year levels. Now we will analyze based on the one-way ANOVA technique (also known as Single
factor Anova). First, we will undertake a hand calculation. This will enable us to realize from
where the numbers originate. Then we will compare the hand-calculated numbers with that of the
inbuilt excel output

12.5.1 HAND CALCULATION IN THE EXCEL SHEET.


It is called a single factor because each year there are several levels. In this case, within each year
there are randomly selected seven students (refer to Table1). The scores have been keyed in each
column (columns are also known as a group or treatment). In experimental contexts, they are also
called 'completely randomized design'. The mean for each column is keyed, it will have its own
distribution and variance. The overall variance, ie the mean of all the 21 scores taken together
(sometimes called the grand mean) is 74.52. A quick look at the three-column means indicates
that the mean of year 1(Delhi) is slightly off the mark as it is three points lower than the overall
mean. That's the reason we are doing ANOVA
ANOVA as you are aware is the analysis of variation. Variance is the average squared deviation,
or the average squared difference, of a data point from the mean. So, we take the distance of each
data point from the mean, square that distance, add those together, and then find the average. That
is variance.

Figure 12.1
But if we remove that last step of finding the part of the average, then we are left with just the sum
of the squares. We will take the distance of each data point from the mean, square each distance,
133
and then add them together. If we stop there, that is the sum of squares (refer to Figure 1). The
'sum of squares' is a foundational component of ANOVA and Regression.
The overall sum of squares ie the Sum of Squares Total (SST) is partitioned into two components.
The first component is SSC (sum of the squares of column) and the second one is the SSE (sum of
the squares of the errors). The sum of squares of the columns (SSC), is between the columns, and
the sum of squares of the error (SSE), is within each column. This is actually about the individual
distribution around each column's mean.
Thus, SST = SSC + SSE (refer to Figure 2)
Kindly make a note:
N = Total number of observations, in this case, it is 21
C = Columns = 3

Figure 12.2

12.5.2 SSC (SUM OF SQUARES OF COLUMNS)


1. It is the level of the single factor we are looking at i.e. years 1, 2, and 3 which are the
columns.
2. The difference between each column's mean and the overall mean. Square those deviations
and add them up. This sums to 88.67
3. The degree of freedom for SSC is dfcolumns = C-1 =3-1= 2
4. MSC = mean square of the columns = SSC divided by dfcolumns = SSC divided by 2 = 88.67
divided by 2 = 44.33 (refer to Table 2)

Table 2

12.5.3 SST (SUM OF SQUARES OF TOTAL).


1. It is the difference between each data point (so all 21 data points) and the overall mean,

134
01
which is 74.52. We would then square this difference and sum it up. It totals 2901.24
2. If we consider a normal distribution, the mean will be in the centre, and some of the data
points will be to the right, which is higher than the mean, and those which are lower than
the mean will be to the left.
3. The degree of freedom for SST is dftotal = N-1 = 21-1=20. (refer to Table 2)

12.5.4 SSE (SUM OF SQUARES OF ERROR)


1. SSE is an error within each column
2. SST = SSC + SSE. Now that we have calculated SST and SSC, by keying in this equation,
we get SST = 2901.24. Alternatively, it can be calculated by finding the difference between
each score within the column and its mean, squaring it, and then summing it up for all three.
The result will match.
3. The degree of freedom for SEE is dferror is N-C = 21-3=18
4. MSE, the mean squared error = SSE divided by dferror = SSE divided by 18 = 2812.57
divided by 18 = 156.25 (refer to Table 2)
F statistics = MSC divided by MSE = 44.33 divided by 156.25 = 0.283726.
0.283726 is the ratio of two variances. It is the variance ratio between the columns divided by the
variance within the column. So that's the F ratio: 'between' divided by 'within'.
Now, 0.283726, the F value obtained, should be compared with the standard Fcritical value
available in the F distribution chart based on the alpha value of 0.05 and the degrees of freedom of
2 and 18. It can also be calculated in excel by keying the formula:
= F distribution, Inverse, Right Tailed
=[Link] (0.05,2,18)
= 3.55455
= F critical value
The null hypothesis is 'means are equal to each other.
The F-statistic value of 0.283726 is not more significant than the F-critical value of 3.55455. It
means that there is no significant difference in mean test scores by year of students. Or it's
another way of saying that the means, the three standards come from a common population.
Thus, we fail to reject our null hypothesis that these three means are different. Therefore, the
means of the first-year students, the second-year students, and the third-year stuents, on their
study skills exam did not differ significantly.
Now let us compare the above output, i.e., the numbers obtained by hand calculation with excel's
built-in Anova table/output.

12.6 EXCELS IN-BUILT ANOVA - SINGLE FACTOR.


To obtain it the following path has been followed (path: excel - data - data - data analysis -
window - Anova single factor - input the range with labels - select labels - alpha to be maintained
at 0.05 - tick the new worksheet or identify/output range where it should appear).
The output will be similar to table 3 below. It's only that I have colored and rearranged for
readability. Let's observe it and compare it to the hand calculations.

135
Table 3
You will observe that the hand calculation output match with that of excel's ANOVA output. Thus
now you are in a position to relate to the excel's inbuilt ANOVA output.
1. The squared sum between columns (SSC) is 88.66666, which matches the hand calculation
viz 88.67.
2. The squared sum of totals (SST) is 2901.238095; this, too matches the hand calculation viz
2901.24.
3. The squared sum of errors (SSE) within columns is 2812.5714. This, too, matches the hand
calculation viz 2812.57
4. The degree of freedom, the MSC, MSE, Fstatistic, and the Fcritic value also match
perfectly.
The p-value is 0.756278 which is more than alpha 0.05 and is thus not significant.
The null hypothesis is that the means are equal to each other. Thus, we fail to reject our null
hypothesis that these three means are different. The means of the first-year students, the second-
year students, and the third-year students on their study skills exam did not differ significantly.

12.7 TWO-WAY ANOVA WITHOUT REPLICATION


Two-way Anova without replication is also known as a 'completely randomized block design'
(blocks are referred to as subgroups).
In a One-way Anova, we selected a random sample for each column (columns are also known as
a group or treatment). We had the variance of the columns (SSC) and then we had the error
variance (SSE), and then we added those together, and we got the overall variance (SST).
Now a Two-way Anova allows us to account for variation at the row level. Thus, we are adding a
new dimension. By adding factors to the rows (or adding blocks) we can extract the row variance
from the overall error variance. This is because some of the error variances are due to the variance
in the rows. The objective is to reduce unexplained, unknown errors.
So, now there are four types of 'sums of squares'. The total sums of squares (SST), are made up of
the column variance (SSC), the row/block variance (SBR), and the error variance (SSE). A
similar concept of assigning (or allocating or partitioning) a certain proportion of variance, to
specific factors or variables, is at the heart of simple and multiple regression. The explanatory

136
power of linear regression is based on ANOVA.
Now let us consider one more fictitious problem to understand the practical implication of 'Two-
way Anova without replication'. Like before, we will once more undertake a hand calculation
(excel) so that you know and understand the source of the numbers ie from where the numbers
originate and their importance. We will follow it up with excel's inbuilt ANOVA output to
compare and reconfirm whether the hand-calculated numbers match the excel's inbuilt ANOVA
output.
Let's assume that the 'LifeStyle Coffee' chain uses secret shoppers who appear as customers to
enter their store and document their experience in terms of customer service, cleanliness, and the
quality of their product - coffee. They give a score out of a maximum of 100. The secret shoppers
receive standardized training by 'LifeStyle Coffee' to ensure consistency and objectivity in their
store reviews. For its locations in the cities of Nagpur, Mumbai, and Pune, 'LifeStyle Coffee'
trained six secret shoppers. Each of the six secret shoppers is assigned to visit the store once in
each of the three cities. As the visit sequence will be assigned randomly, hence the name
'randomized block design'. It is called 'without replication' because each shopper is only going to
each city once (In ANOVA with replication each shopper will visit the city more than once. There
will be multiple measurements. Thus, more data will be generated). The two factors here are the
city ( columns) and the shopper (rows).
We would like to know if a difference exists in secret shopper ratings among the cities…
Are all the cities about the same in their ratings? Is one significantly higher than the other two? Or
are all three different from each other?
This is testing for differences among the cities. It's not testing whether they are good or bad,
which is a subjective experience or rating. The heart of the problem as to what makes this
problem fit for Two-way Anova is that the secret shoppers (ie rows) themselves will have their
natural variation to review their experiences. Two-way ANOVA allows accounting for the
shopper variation to determine if a difference exists among the cities without the shopper
variation clouding or masking any of the city differences. It is like untangling all the sources of
variation before we can get down to looking at any differences that might exist between the cities.

12.7.1 HAND CALCULATION FOR MEANS (REFER TO TABLE4)


Kindly take note of the following:
¡ n = N = 18
¡ Columns (Groups, Treatments) = C = 3
¡ Rows (Blocks) = B = 6

Table 4

137
To calculate the mean, the '=Average' excel function in the cell is used
1. The mean for each column is calculated
2. The mean for each row is calculated
3. The overall mean is calculated
In a one-way ANOVA:
SST (sum of squares total) = SSC (sum of squares column (or treatment or groups) + SSE (sum of
squares within or error)
Whereas, in a two-way ANOVA:
We are interested in the differences between the cities (ie the columns). By introducing the
blocking variable, we are further trying to reduce the original SSE into SSB and the remaining
will be again the new SSE, ie further splitting the SSE and attributing it to SSB and that which
cannot be attributed will be SSE. Thereby now the SSE is smaller which is the unexplained
source of error; the unexplained variance will always be there. In the end, SSC will be compared
to SSE, and SSC claims a larger part of the total variance.
SST (sum of squares total) = SSC (sum of squares column) + SSB (sum of squares block) + SSE
(sum of squares error or within).
The extent SSB accounts for SSE can be known from the ratio 'SSB divided by SSE'. In the real
sense by dividing 'MSB by MSE'.

12.7.2 HAND CALCULATION FOR SSC (REFER TO TABLE 5)


1. The column mean subtracted from the overall mean
2. The difference squared
3. The squares summed/added.
4. dfcolumns = Degree of freedom for SSC = C - 1 = 3-1 = 2
5. 87.5 is the squared difference for one column. It has to be multiplied by the number of rows
which is 6. Thus, SSC is 525
6. MSC = mean square of the columns = SSC divided by dfcolumns = SSC divided by 2
7. MSC = 525 divided by 2 = 262.5

Table 12.5

12.7.3 HAND CALCULATION FOR SSB (REFER TO TABLE 6)


The row (i.e., blocks) mean subtracted from the overall mean
1. The difference squared.
138
2. The squares summed/added.
3. dfblocks = The degree of freedom for SSB = B - 1 = 6-1= 5
4. 250 is the squared difference for one row. It has to be multiplied by the number of columns
which is 3. SSB is 750.
5. MSB = mean square of the blocks error = SSB divided by the dfblocks = SSB divided by 5
6. MSB = 750 divided by 5 = 150.

Table 12. 6

12.7.4 HAND CALCULATION FOR SST AND SSE (REFER TO TABLE 7)


1. Each score by the shopper is subtracted from the overall mean
2. The difference is squared
3. The squares summed/added
4. SST is 1750.
5. dftotal = degree of freedom for SST = N-1 = 18-1 = 17?
6. SST = SSC + SSB + SSE
7. SSE = SST - SSC - SSB
8. SSE = 1750 - 525 - 750 = 475
9. dferror = The degree of freedom for SEE = (C-1) *(B-1) = (2) *(5) = 10
10. MSE = mean squared error = SSE divided by the dferror = SSE divided by 10
11. MSE = 475 divided by 10 = 47.5

139
Table 7

F statistics = MSC divided by MSE = 262.5/ 47.5 = 5.53.


5.53 is the F ratio of two variances, 'between the columns' divided by variance 'inside the
column'.
[ For your information…MSN divided by MSE = 150 / 47.5 = 3.157 is the F ratio of variances
'between the rows' divided by variance 'inside the rows'].

12.7.5 EXCELS BUILT-IN ANOVA - TWO-WAY ANOVA WITHOUT REPLICATION


(REFER TO TABLE 8)
To obtain it, the following path has been followed (path: excel - data - data - data analysis -
window - Anova two factor without replication - input the range with labels - select labels - alpha
to be maintained at 0.05 - tick the new worksheet or identify/output range where it should
appear). The output will be similar to table x pasted below. It's only that I have colored and
rearranged for readability. Let's observe it and compare it to the hand calculations.

140
Table 8

You will observe that the excel output and the hand calculation match.
1. The column means for the city's matches.
2. The squared sum for the rows, columns, and the error matches
3. The degree of freedom, the MSC, MSB, and MSE also match.
4. The F statistic for rows and columns and the F critical value also match.
Thus, our hand calculation is right and it has provided a sense and an understanding to know the
source of numbers.
The P-value for the columns (in cities as per the example) is 0.024. It is less than 0.05 so it is
significant whereas the P-value for the rows (in shoppers as per the example) is 0.057 which is
barely equal to or pretty close to 0.05. Thus, it is not significant. This indicates that the column
scores do have legitimate differences even after accounting for variation in shoppers' scores.

F Critical value ie F critic = FAlpha, dfC, dfE = F0.05.2,10

Now if you check the F table or calculate in a cell in excel as per the equation = [Link](0.05,
2,10), then the value that you will obtain is = 4.10, is the Fcritic value.
The null hypothesis is that there are no significant differences in the city (remember that it always
assumes that there is no difference).
We reject the null hypothesis as the F statistic for columns (cities) is 5.52 is larger than the
Fcritic. A significant difference in the mean quality score is present in the columns (cities).

12.7.6 TWO-WAY ANOVA WITH REPLICATION


In a two-way ANOVA with replication, we have multiple measurements, not just a single one.
This allows for a new type of measurement, known as the interaction between two factors.

141
To understand this better let us take an example of a plant food company. They are trying to find
the effectiveness of three different plant foods named AA, BB, and CC with one feeding per day
on eight respective plants. That is eight plants will be given the food AA, another eight plants will
be given BB and the third set of eight plants will be given the food CC. The height of the plant will
be tested before the food is given and after 75 days. The same type of seeds will be used for the
entire experiment to control for the type of seed.
Now let's add one more factor. That is increase the feeding frequency from once a day to twice a
day with another set of eight similar plants each for the same food ie AA, BB, and CC. This is a
balanced design because two factors and two, multiple, equal numbers of measurements are
present for each factor combination.
Now we have six sets of eight plants in this experiment. We have a two-factor or two-way
ANOVA with replication. It's with replication because each food and feeding frequency has eight
plants.
So now to understand how different it is when we do replication here, take note that each of the
eight plants with respective food AA, BB, and CC for one feeding has its mean & variation, and
similarly for two feedings it has its mean & variation. This is shown in the figure there are 6
means and each has its variation from its mean. This is the fundamental concept of two-way
Anova with a variation.

Table 9

Browse through the table and keep the following at the back of your mind:
1. The two row means 60 and 58.1 are pretty close to each other
2. The two-row means of 60 and 58.1 are quite close to that of the overall mean of 59.1
3. The two column means 63,2 and 64.6 are a bit above the overall mean of 59.1

142
4. The column means 49.6 for CC is way off from the overall mean of 59.1
With the help of the excel tool, I have taken the two-column means (one feeding mean & two
feeding mean) and plotted a graph [ path: excel - charts - insert line chart]. (refer to Graph1)
This graph is known as the interaction graph or the graph of marginal means. I have titled it as
'Marginal means of plant height (ignore the word - Estimated in the graph)'.
It allows us to visualize the characteristics of each factor and any interaction that may be
occurring between them. So, in a marginal means graph, as a general rule, we look to see if the
lines cross or would cross because that expresses that the factors change, and their values change
across the groups.
The factor of interest or the factor with the most levels, ie the three types of plant food is plotted
on the X-axis and the dependent variable, what is being measured is plotted on the Y ie the plant
growth.
When the plant food AA and BB is fed twice (red line) the plant grows whereas with the food CC
the growth goes down. Thus, two feedings do not produce consistent growth across all the plant
food types. This type of situation is called an 'interaction'.
An interaction occurs when the effect of one factor changes for different levels of the other factor.
In this case, the most effective feeding frequency changes across plant food types. For AA and
BB two feedings are the best but when we get to CC, it changes, one feeding is the best. That's
what we mean by an interaction. If the lines cross on a marginal means graph, it means that there
is an interaction. Non-parallel lines can also mean there is a significant interaction but the
crossing of lines is more indicative than non-parallel lines.

Graph 1

While interpreting a two-way Anova one should always look for a significant interaction first. If
the exchange is significant, then we need not interpret it further. Because it means that the two
individual factors are too intertwined and tied together to look at them individually, always go for
the interaction term first when you're looking at your F-ratio and p-value or significance.
In this example, we will avoid hand calculation. This is because by now you may have
understood as to how to do it to know how the numbers arise.

143
12.8 EXCELS BUILT-IN ANOVA - TWO-WAY ANOVA WITH REPLICATION
To obtain it the following path has been followed (path: excel - data - data - data analysis -
window - Anova two factor with replication - input the range with labels - rows per sample should
be 8 here - alpha to be maintained at 0.05 - tick the new worksheet or identify/output range where
it should appear). The output will be similar to Table 10 below. It's only that I have colored and
rearranged for readability. Let's try to understand it.

Table 10
Under source of variation, there is an item 'sample' it is the 'feeding frequency', i.e.s the rows, The
item 'columns' represent the 'plant food', and 'interaction' represent the interaction between
'feedings*plantfood'.
Now the p-value for the sample ie feeding frequency, is 0.45. It is not significant as it is more than
0.05. Secondl,y also have a look at the rows means of 60 and 58.1 for feed 1 and feed 2
respectively. They are pretty close to each other and more relative to the overall mean. Now if
you visualize and try to plot these two numbers on the graph above, you will notice that they are
close to each other and have an overall mean of 59.1. There is hardly any variability. This gives us
an idea that they are not significant.
Now have a look at the p-value of plant food ie columns. It is 0.0000 (this has been obtained after
decreasing decimals). It is significant because it is less than 0.05. Now also look at the column
means for the three plant foods it is 63.5, 64.75, and 49. There's a difference. The mean 49 is quite
a way off! If we try to plot these three numbers on the graph above, you will realize that the three
two plots are quite a way off from the overall mean. There is a lot of variability in the column
means that in the type of food So there's a significant difference
Now we will turn our attention to the item - 'interaction'. It is the interaction between
'feeding*plantfood'. The p-value is 0.0000 (this has been obtained after decreasing decimals). It

144
is lower than 0.05. Secondly, we have seen that two feedings are better for plant food AA and BB,
and in the case of CC one feeding is better. Moreover, the lines in the graph cross each other. The
crossed-row lines indicate, usually, an interaction.
Thus, the column means spread far apart from the overall mean, indicate a significant column
factor and the row means spread far apart from the overall mean usually indicate a significant
row factor.
So, just because each effect is significant, does not mean there's a significant interaction. But if an
interaction exists, and it is significant then the row and column effects cannot be evaluated
individually. The main effects are too intertwined and are too confounded together to look at
individually because the values change across. And it cannot be untangled.

12.9 POST HOC ANALYSIS


Post hoc means afterward this is how we follow up on statistically significant results with
ANOVA (also called an omnibus test because it tests overall differences among a variety of
groups).

Table 3

If ANOVA is statistically significant, then we're going to follow up with a post hoc test to
determine where those differences reside.
We wouldn't do a post hoc test if our ANOVA were non-significant. We would only do the post
hoc follow-up if the initial ANOVA told us that there are differences there somewhere, and now
we have to find them.
The posthoc is only necessary when you reject the null hypothesis when you say there is a
statistically significant difference between groups and when there are three or more groups.
If there are only two groups, well then the solution is easy you just look at the means whichever
group has the higher mean that's statistically significantly different from the other group.
But in the case of ANOVA, we could have three means is the first one different than the second or
different from the third is the third just different from the first we have to follow up in a way to
determine where those differences lie.
There are multiple ways in which we could conduct a post hoc test the simplest one is called a
Bonferroni correction. This is where you take your alpha level typically 0.05 and just divide it by

145
the number of tests. If we were running 5 tests, we would divide 0.05 so each test would have a
significance level of 0.01 to be considered statistically significant. This is the simplest method
and the most conservative method but not necessarily the best method in that it can increase the
chance for type two errors.
There are other ways of approaching post hoc testing that can give us a nice balance between not
inflating the type two error rate and also making sure that we're only finding differences where
they truly exist.
There are 18 different types of post hoc tests as per the SPSS software. Which one to choose
depends upon the nature of the data.
If all the assumptions are true and in case of equal sample size there's the Tukey HSD (Honestly
Significant Difference) but in case of unequal sample sizes there are Gabriel's tests and in the
case of smaller sample size Hochberg GT 2 and for the larger sample size unequal variances there
are the Gains Howl and there's the Fisher LSD test, Scheffe's test and more.
Let us take an example and apply the Tukey HSD posthoc test
This test is very commonly used by statisticians. Let's take an example and use this test for the
example we already had from Excel's built-in Anova: single Factor test (or a One-way Anova).
The formula to be used for Tukey's Posthoc analysis test is as follows:

wherein…
q = is the constant to be obtained from the Studentized Range q table (based on 'dfw' and 'k'; the
number of treatment groups)
MSw is the mean square within
nk is the number in each category.
Now let's use this above formula for the excels output above (refer to Table 3):
'dfw'= 18 and 'k' the number of treatment groups = 3. Now using the Studentized Range q table,
the corresponding number across this intersection is 3.609 (refer to Figure 3)
MSw = 156.2539 and nk = 7 ( refer to Table 3)
Now inserting these numbers in their appropriate positions, the formula looks like this :
HSD = 3.609 * square root of 156.25 divided by 7
HSD = 3.609 * 4.72
HSD = 17.03448
If the means differ by more than this HSD value which is 17.03448 then they are statistically
significantly different.
The three means are 71.714, 75.285, and 76.571. None of them differ by more than 17.03448.
The means are not statistically different.
As such in this example before we rejected the null hypothesis because the means did not differ
significantly. The null hypothesis was that the means are equal to each other. That is what is
confirmed by Tukey's Posthoc analysis. Infact, the post hoc analysis should only be done if there

146
is a statistical significance! This was just an example for you. You can undertake a similar
exercise for Table 8 and Table 10.

SUMMARY
¡ Analysis of variance, or ANOVA, is a statistical method that separates observed variance
data into different components to use for additional tests.
¡ A one-way ANOVA is used for three or more groups of data, to gain information about the
relationship between the dependent and independent variables.
¡ If no true variance exists between the groups, the ANOVA's F-ratio should equal close to 1.
¡ Analysis of variances (ANOVA) is a statistical method that analyzes the influence of one or
more independent variables on a dependent variable of interest.
¡ ANOVA is used in various applications, including in finance and financial markets to find
and confirm correlations and associations between various factors.
¡ There are a variety of ANOVA techniques, including one-way, two-way, and factor models
¡ A two-way ANOVA is an extension of the one-way ANOVA (analysis of variances) that
reveals the results of two independent variables on a dependent variable.
¡ A two-way ANOVA test is a statistical technique that analyzes the effect of the independent
variables on the expected outcome and their relationship to the outcome itself.
¡ ANOVA has many applications in finance, economics, science, medicine, and social
science.
¡ There are several post hoc analysis methods to be deployed after ANOVA. It should be
undertaken only if statistical significance exists.
147
SELF ASSESSMENT QUESTION
1. Suppose an ANOVA has been performed on a completely randomized design containing
six treatment levels. The mean for group 3 is 15.85, and the sample size for group 3 is eight.
The mean for group 6 is 17.21, and the sample size for group 6 is seven. MSE is .3352. The
total number of observations is 46. Compute the significant difference for the means of
these two groups by using the Tukey-Kramer procedure. Let ? = 0.05
2. A completely randomized design has been analyzed by using a one-way ANOVA. There
are four treatment groups in the design, and each sample size is six. MSE is equal to 2.389.
Using compute Tukey's HSD for this ANOVA.
3. Using the results of problem 11.5, compute a critical value by using the Tukey-Kramer
procedure for groups 1 and 2. Use Determine whether there is a significant difference
between these two groups.
4. In recent years, the debate over the U.S. economy has been constant. The electorate seems
somewhat divided as to whether the economy is in recovery or not. Suppose a survey was
undertaken to ascertain whether the perception of economic recovery differs according to
political affiliation. People were selected for the survey from the Democratic Party, the
Republican Party, and those classifying themselves as independents. A 25-point scale was
developed in which respondents gave a score of 25 if they felt the economy was in
complete recovery, a 0 if the economy was not in a recovery and some value in between for
more uncertain responses. To control for differences in socioeconomic class, a blocking
variable was maintained using five different socioeconomic categories. The data are given
here in the form of a randomized block design. Use to determine whether there is a
significant difference in mean responses according to political affiliation.

Political Affiliation
Socioeconomic Class Democrat Republican Independent
Upper 11 5 8
Upper middle 15 9 8
Middle 19 14 15
Lower middle 16 12 10
Lower 9 8 7

5. A randomized block design has a treatment variable with six levels and a blocking variable
with 10 blocks. Using this information and complete the following table and conclude the
null hypothesis.
Source of Variance SS df MS F
Treatment 2,477.53
Blocks 3,180.48
Error 11,661.38
Total

148
REFERENCES:
1. [Link]
2. Black Ken; Business Statistics for contemporary decision making; 6th ed.
3. h t t p s : / / w w w. y o u t u b e . c o m / w a t c h ? v = Z k j P 5 R J L Q F 4 & l i s t = P L I e G t x p v y G -
LoKUpV0fSY8BGKIMIdmfCi&index=1, Foltz Brandon Statistics 101 Linear
Regression
4. [Link]

149
MODULE - 13 SIMPLE LINEAR REGRESSION

STRUCTURE
¡ What Is a Regression?
¡ Why is it called Regression?
¡ What is the purpose of Regression
¡ Understanding Regression
¡ What is a Variable?
¡ What is a Covariance?
¡ What is a correlation coefficient 'r'?
¡ What are the Correlation caveats?
¡ Coefficient of determination (R-squared)
¡ What is a Linear Relationship?
¡ Understanding regression equation
¡ Calculating regressions: manual and excel
¡ Regression statistics table - Excel output
¡ Summary

13.1 LEARNING OUTCOMES


1 By the end of this chapter you will be able to understand
2 The meaning of regression
3 Meaning of variables and concepts
4 Meaning of correlation and covariance
5 Meaning of linear regression
6 Using excel to compute regression analysis

13.2 INTRODUCTION
If you wish to know how two or more pieces of data relate to each other, for example how the total
food bill in a restaurant impacts tips (two pieces of data) or how tips are impacted by restaurant
ambiance and the price on the menu card (three pieces of data), or if you wish to create a forecast
or analyze predictions based on the relationships between the pieces of data, then the regression
is helpful. This is a tool commonly used for forecasting and statistical analysis.
This course familiarises you briefly with the underlying concepts, principles, and mechanics
related to Regression.

150
13.2.1 WHAT IS A REGRESSION?
Regression is a method that determines the strength and character of the relationship between
one dependent variable (usually denoted by Y) and a series of other variables known as
independent variables (usually denoted by X).
Regression analysis is the process of constructing a mathematical model or function that can be
used to predict or determine one variable by another variable or other variables. The most
elementary regression model is called simple regression or bivariate regression involving two
variables in which one variable is predicted by another variable. In simple regression, the
variable to be predicted is called the dependent variable and is designated as Y. The predictor is
called the independent variable, or explanatory variable, and is designated as X.
Regression is also known as 'Ordinary least squares' (OLS), or 'least squares regression' or
simple linear regression. Linear regression establishes the linear relationship between two
variables. Linear regression is graphically depicted using a straight line with the slope defining
how the change in one variable impacts a change in the other. A brief on non-linear regression is
provided at the end.

13.2.2 WHY IS IT CALLED REGRESSION?


The statistical technique most likely was termed "regression" by Sir Francis Galton in the 19th
century to describe the statistical feature of biological data (such as heights of people in a
population) to regress to some mean level. In other words, while there are shorter and taller
people, only outliers are very tall or short, and most people cluster somewhere around (or
"regress" to) the average. The word 'Regress' can be considered as 'retreat' or 'revert' to the mean.

13.2.3 WHAT IS THE PURPOSE OF REGRESSION?


Regression is a powerful statistical inference tool used to predict future outcomes based on past
observations (Regression cannot indicate the cause). It is used in several contexts in business,
finance, and economics. For instance, it is used to help investment managers value assets and
understand the relationships between factors such as commodity prices and the stocks of
businesses dealing in those commodities. Regression can also help predict sales for a company
based on weather, previous sales, GDP growth, or other types of conditions. Regression is used to
determine how many specific factors such as the price of a commodity, interest rates, particular
industries, or sectors influence the price movement of an asset.

13.2.4 UNDERSTANDING REGRESSION


Regression captures the correlation (i.e., association or connection or relationship) between
variables observed in a data set and quantifies whether those correlations are statistically
significant.
The two basic types of regression are simple linear regression and multiple linear regression.
Simple linear regression uses one independent variable X to explain or predict the outcome of the
dependent variable Y, while multiple linear regression uses two or more independent variables to
predict the outcome (while holding all others constant).
Now let us review the concepts which are closely interrelated with Regression. We will review-
Variables, Covariance, Correlation coefficient, and Linear relationship and then create a
regression equation, and a regression line and interpret the output.
151
13.3 WHAT IS A VARIABLE?
A Variable is a quantity that may assume any one of a set of values that is alterable, adjustable, or
changeable. The core of a regression model is the relationship between two different variables,
called the dependent and independent variables. The dependent variable is dependent on the
independent variable. For example, the sales of luxury wristwatches depend on the disposable
income of the customer. The independent variable is illustrated on the X-axis and the dependent
variable on the Y-axis. It is also important to determine the direction and strength of the
relationship between these two variables to forecast wristwatch sales. If disposable income
increases/decreases by 1%, how much will the sales of luxury wristwatches increase or
decrease?

13.4 WHAT IS A COVARIANCE?


The formula to calculate the relationship between two variables is called covariance. Covariance
is a descriptive measure of the linear association between two variables. The calculation
indicates to you the direction of the relationship. It does not indicate the strength of the
relationship.
A positive value indicates a direct or increasing linear relationship. If one variable increases and
the other variable tends to also increase, the covariance would be positive. A negative value
indicates a decreasing relationship. If one variable goes up and the other tends to go down, then
the covariance would be negative.
They follow a linear pattern. That is they Covary.

Figure 1: Covariance formula

Let's take an example, to study the relationship, between the number of workers, x, and the tables,
y, produced by them. In Table 1 Given below is the sample of 10 with a duration of one hour
each. The standard deviation calculated from the table given below is Sx = 6.48 and Sy = 16.69.
Given below is the methodology to obtain covariance as per the formula in Figure 1.
Covariance is 962.4/ n-1 = 962.4 /9 = 106.93.
The sign is positive, i.e., + 106.93. The graph depicts linearity.
The scatter plot in Figure 2 represents that there exists a positive linear relationship between the
number of workers, x, and the tables, y, produced by them.

152
Table 1: Covariance calculation

13.5 WHAT IS A CORRELATION COEFFICIENT 'R'?


Correlation is a measure of the degree of relatedness of variables. Correlation shows the strength
of a relationship between two variables and is expressed numerically by the correlation
coefficient. The correlation coefficient's values range between -1.0 and +1.0.
A perfect positive correlation means that the correlation coefficient is exactly 1. This implies that
as one variable move, either up or down, the other variable moves in lockstep, in the same
direction. A perfect negative correlation means that two assets move in opposite directions, while
a zero correlation implies no linear relationship at all.

Figure 2: Scatter plot for covariance depicts linearity and positivity

Pearson correlation coefficient (r) = rxy


= Covariance (x, y) divided by [standard deviation (x); ie Sx * standard deviation (y); ie Sy]
= Cov (x, y) divided by the product of (Sx*Sy)
In one of our previous examples, if the correlation is +1 and the disposable income increases by
1%, then wristwatch sales would increase by 1%. If the correlation is -1, a 1% increase in
disposable income would result in a 1% decrease in wristwatch sales - the exact opposite.

153
The correlation calculation:
In the above example of 'workers and tables produced' Sx is 6.48 and Sy is 16.69 and covariance
is 106.93.
Correlation =
106.93 divided by the product of (6.48 x 16.69)
= 0.989
=r
The rule of thumb to know whether a relationship exists between the two variables is to
determine whether the correlation | r | ³ 2; divided by the square root of n (where n is the sample
size).
= | r | ³ 2 divided by the square root of 10
= | r | should be³ 0.632 then the relationship exists.
In our calculations,
| r | is 0.989 which is greater than 0.632, thus the relationship exists.

13.6 WHAT ARE THE CORRELATION CAVEATS?


Covariance provides the 'direction' ( positive, negative, near zero) of the linear relationship of the
two variables, whereas correlation provides 'direction and strength'. The covariance result has no
upper or lower bound and its size is dependent on the scale of the variables. While correlation is
always between -1 and +1 and its scale is independent of the scale of the variables themselves.
Covariance is not standardized while correlation is standardized. Correlation is only applicable
to linear relationships. Correlation strength does not necessarily mean that the correlation is
statistically significant; related to the sample size.

13.7 COEFFICIENT OF DETERMINATION (R-SQUARED)


The coefficient of determination (R-squared) is a metric (ie a measure) that is used to measure
how much of the variation in outcome can be explained by the variation in the independent
variables. R2 can only be between 0 and 1, where 0 indicates that the outcome cannot be
predicted by any of the independent variables and 1 indicates that the outcome can be predicted
without error from the independent variables.

13.8 WHAT IS A LINEAR RELATIONSHIP?


A linear relationship (or linear association) is a statistical term used to describe a straight-line
relationship between two variables. Linear relationships can be expressed either in a graphical
format where the variable and the constant are connected via a straight line or in a mathematical
format where the independent variable is multiplied by the slope coefficient, and added by a
constant, which determines the dependent variable. A linear relationship satisfies the equation:
Y = mX + b where
m = slope
b = y-intercept or a constant

154
Figure 3: Calculation of slope

In this equation, "X" and "Y" are two variables that are related by the parameters "m" and "b".
Graphically, Y = mX + b plots in the X-Y plane as a line with slope "m" and Y-intercept "b." The
Y-intercept "b" is simply the value of "Y" when X=0.
The slope "m" is calculated from any two individual points (X1, Y1) and (X2, Y2) as shown in
Figure 3

13.9 UNDERSTANDING REGRESSION EQUATION


Linear regression models often use a least-squares approach to determine the line of best fit. It
provides the overall rationale for the placement of the line of best fit among the data points or
scatter plots. It is also known as the least squares regression line. By definition, a line is always
straight, so the best fit line is linear.
It minimizes the vertical distance from the data points to the regression line. That is, it minimizes
the distance between the regression line and where observations fall in the data sets. A square is,
in turn, determined by squaring the distance between a data point and the regression line (or mean
value of the data set) and then adding them together. It is also known as variation. Variation refers
to the difference between each data set from the mean. The line of best fit will minimize this
value. A low sum of squares indicates little variation between data sets while a higher one
indicates more variation.
The term "least squares" is used because it is the smallest sum of squares of errors. The least
squares approach is a popular method for determining regression equations, and it tells you about
the relationship between the independent and the dependent variable. Once this process has been
completed, a regression model is constructed. The regression equation describes the relationship
between the dependent variable (Y) and the independent variable (X).
The general form of the regression model is:
Simple linear regression equation:
Y = b + mX + u
Multiple linear regression equation:
Y = b + m_1X_1 + m_2X_2 + m_3X_3 + ... + m_tX_t + u
Where…
Y = The dependent variable you are trying to predict or explain or forecast.
X = The independent variable(s) you are using to associate with Y.

155
b = The y-intercept or "b" is the value of Y (dependent variable) if the value of X (independent
variable) is zero, and so is sometimes simply referred to as the 'constant.'
m = is the slope of the regression line
u = The regression residual or error term
How do you interpret a regression model?
A simple regression model output may be in the form of:
Y = 3 + (6.4) X + 0.56
We would interpret the model as the value of Y changes by 6.4 X for every one unit change in X. If
X goes up by 3, Y goes up by 19.2. The Y intercept is 3 when X is zero. The regression residual or
error term is 0.56.
A multiple regression model output may be in the form of:
Y = 1.0 + (3.2) X1 - 2.0(X2) + 0.21.
Here we have a multiple linear regression that relates a dependent variable Y with two
independent variables X1 and X2.
Multiple linear regression (MLR), also known simply as multiple regression, is a statistical
technique that uses several independent variables to predict the outcome of a dependent variable.
It extends to several independent variables. Whereas Simple linear regression is a function that
allows making predictions about one variable based on the information that is known about
another variable.
We would interpret the model as the value of Y changes by 3.2X1 for every one unit change in X1
(if X1 goes up by 2, Y goes up by 6.4, etc.) holding all else constant. That means controlling for
X2, X1 has this observed relationship. Likewise, holding X1 constant, every one unit increase in
X2 is associated with a 2X decrease in Y. Note the negative sign here.
We can also note the Y-intercept of 1.0, meaning that Y = 1 when X1 and X2 are both zero. The
regression residual or error term is 0.21.
What are the assumptions that must hold for regression models?
To properly interpret the output of a regression model, the following main assumptions about the
underlying data process of what you analyzing must hold:
1. The relationship between variables is linear
2. That the variance of the variables and error term must remain constant (Homoskedasticity)
3. All independent variables are independent of one another
4. All variables are normally-distributed

13.10 CALCULATING REGRESSIONS: MANUAL AND EXCEL


Now that some concepts are clear let's do a simple exercise using manual and excel regression
tools.
Manual calculation:
We are taking an example of a tip received by the restaurant staff from customers. Refer to Table
2. Just eyeballing columns 1 and 2, you can see that there will be a positive correlation between
the tip amount and the total bill.

156
Table 2: calculating the mathematical regression line

Table 3: Calculating SSE (ie RSS) and SSR

The tip amount increases as the total bill increases. There seems to be linearity and positivity.
The mean of the total bill is 74 whereas that of the tip amount is 10. Now it is important to
remember that these two numbers (74, 10) are known as centroids. The regression line passes
through the centroid.

157
1. Columns 3, 4, 5, and 6 have been created to calculate the slope 'm' of the linear regression
equation: Y = mX + b.
2. Column 3 denotes the deviation of each amount of the total bill from the mean (74) whereas
column 4 denotes the deviation of each amount of the tip from the mean (10).
3. Column 5 is the product of columns 3 & 4 and column 6 is the square of column 3.
4. The summation (ie total) of column 5 is 615 and that of column 6 is 4206.
5. The slope 'm 'as per the formulae is 615/4206 = 0.146219686. Thus, the linear regression
equation: Y = mX + b will read as Y = 0.146219686 X + b.
6. To calculate 'b' substitute the centroids (74, 10) as X and Y in the linear regression equation.
The calculation yields the value of 'b' as - 0.819. 'b' here is the Y-intercept (or considered as
the constant). Now we have calculated the value of slope 'm' and that of 'b'. So the linear
regression equation (or the regression line) Y = mX + b reads thus: Y = 0.146219686 X -
0.819.
7. The equation can be interpreted as follows…for every increase in the bill amount by Re 1/-,
the tip is expected (or forecasted) to increase by 0.146219686 paise. And if the bill amount
is zero, the tip amount will be - 0.819 paise i.e. negative 0.819 paise (Well, at times the y-
intercept may not make any sense in the real world).
8. Now we will use the linear regression equation (or the regression line) Y = 0.146219686 X
- 0.819 to calculate the predicted (or forecasted) tip amount based on the total bill. Kindly
remember here that we already have the observed or the actual tip amount. But now we are
calculating the predicted (or the forecast amount). It has been done in the table pasted
below:

Let's understand Table 3:


1. Column 1 is the independent variable, which is the total bill whereas column 2 is the tip
amount which is the dependent variable. This is because the tip amount depends on the
total bill. As per the convention, the independent variable is plotted on the X-axis and the
dependent on the Y-axis.
2. Column 3 is the difference between the observed(actual) tip amount and the 'mean' of the
tips which is 10.
3. Column 4 is the square of each difference of the tip obtained in column 3. The summation
of the squares obtained here is 120. This is known as the Sum of the Squared Total i.e., SST.
If we consider only the dependent variable, the sum of squares is due to error (from the
mean which is 10 in this case). Therefore, it is also the total and maximum sum of squares
for the data.
4. In column 5, we use the linear regression equation obtained earlier ie Y = 0.146219686 X -
0.819. In this equation, we input each value of X to obtain the estimated (forecasted) value
of the tip amount Y. For example, Y = 0.146219686 * 34 - 0.819 = 4.152. Now, this is the
predicted amount as per the linear regression equation. We do this for each of the next 5
observations.
5. In column 6 we state the difference between the observed (actual) tip amount and that of
the estimated amount calculated in the previous step. This is the 'error' when compared
with the predicted amount.

158
6. In column 7, we square each error and sum it up. The summation obtained here is
30.07489464. This is known as the Sum of Squared Error i.e., SSE (some refer to it as the
Residual Sum of Squares {RSS}). The lower the SSE, the better it is because it means that
it fits the data, quite well.
7. The difference between the two summations ie SST and SSE is the Sum of Squared
Regression (ie SSR).
8. So, Sum of the Squared Total (SST) = Sum of Squared Regression (SSR) + Sum of Squared
Error (SSE), ie put simply it is SST= SSR + SSE. Thus, inputting SST and SSE we get SSR
which is 89.92510536.
9. With the help of the linear regression equation (or the regression line) Y= mX + b ie Y =
0.146219686 X - 0.819; we have been able to reduce the SST (some may also refer to it as
the previous error) from 120 to 30.07489464. That is the error has been reduced by
89.92510536. In other words, it means that the error has been regressed or reduced! Thus,
the use of the term 'regression'
10. The coefficient of determination ie r2 can be calculated here to understand the 'fit'. r2 =
SSR divided by SST = 89.92510536 divided by 120 = 0.7493 = 74.93%. It means that
74.93 % of the error (which in this case is 89.92510536) can be explained by using the
estimated regression equation (which is Y = 0.146219686 X - 0.819.) to predict the tip
amount. The balance of 25.07% (which is 30.07489464) that is SSE remains unexplained!
To create the scatter diagram (refer to Figure 4) with the help of Excel, you may follow the path
described below (also depends on the excel version that you may use):
Open excel - click insert - Scatter - pick up the first plot - you will get a window on the excel sheet
- right click on the window - add data - select the X values and the Y values which have been
plotted on the excel sheet - okay. Now you will get the (x, y) plots on the chart.
To obtain the Regression line: click chart design - add chart element- trendline. And you will
obtain the trend line.
Follow the same procedure - More trend line option - tick the display equation on the chart and
tick display the R squared value on the chart. Now the Figure that you obtain will resemble
Figure 4:

Figure 4: The scatter plot, the regression line, the regression equation, and R2
159
13.11 REGRESSION STATISTICS TABLE - EXCEL OUTPUT
To obtain the Regression Statistics table (Refer to Table 4), you may follow the path: - Data - Data
Analysis - Regression - Input the Y range and the X range with the labels - labels - confidence
interval at 95%. The following table will be displayed.

Table 4: Regression statistic output from excel

You can now compare Table 4 obtained from Excel with the manual calculation.
It has been labelled for your easy understanding.
The following have been labelled: Coefficient of determination ie R2, the SSR, the SSE, and the
SST, the value of the y-intercept ie b, the slope, 95% confidence interval between 0.02883093 to
0.26361, and the Mean Square Error (MSE) - s2.
The confidence interval can be interpreted as "I am 95% confident that the interval (0.02883093
to 0.26361) contains the true slope of the regression line. Because the interval does not contain
zero so I reject the null hypothesis that the slope is zero.
Mean Square Error (MSE), s2 is an estimate of s2 the variance of the error. In other words, it
means, how spread out the data points are from the regression line. MSE is SSE divided by the
product of (sample numbers minus 2) = 30.074 divided by (6-2 = 4) = 7.5187.
The standard error of the estimate or sigma, or just standard error - s, is the standard deviation of
the error term. So, all of the errors we are dealing with are standard deviation. It's just the average
distance an observation falls from the regression line in units of the dependent variable. So, since
MSE (ie s2) is squared, the standard error 's' is just the square root of it. 's' is the square root of
MSE.
s = Square root of MSE
i.e., square root of 7.51872325249643 which is 2.74202903932406.
You will find this number in Table 4 at the top LHS.

160
Nonlinear Regression
Nonlinear regression is a form of regression analysis in which data is fit to a model and then
expressed as a mathematical function. Simple linear regression relates two variables (X and Y)
with a straight line (Y = mX + b), while nonlinear regression relates the two variables in a
nonlinear (curved) relationship. Nonlinear regression modelling is similar to linear regression
modelling in that both seek to track a particular response from a set of variables graphically.

SUMMARY
1. A linear relationship (or linear association) is a term used to describe a straight-line
relationship between two variables. Linear relationships can be expressed either in a
graphical format or as a mathematical equation of the form Y = mX + b.
2. A regression is a technique that relates a dependent variable to one or more independent
variables. Simple Regression or Multiple Regression
3. A regression model can show whether changes observed in the dependent variable are
associated with changes in one or more of the independent variables. It does this by
essentially fitting a best-fit line and seeing how the data is dispersed around this line.
4. The sum of squares measures the deviation of data points away from the mean value. A
higher sum of squares indicates higher variability while a lower result indicates low
variability from the mean. There are three types of sum of squares: SST, SSE, and SSR.
5. A line of best fit is a straight line that minimizes the distance between it and some of the
data points
6. The least squares method is a procedure to find the best fit for a set of data points by
minimizing the sum of the offsets or residuals (ie error) of points from the line.
7. The least squares method provides the overall rationale for the placement of the line of best
fit among the data points being studied.

SELF ASSESSMENT QUESTIONS


1. Use a computer to develop the equation of the regression model for the following data.
Comment on the regression coefficients. Determine the predicted value of y for
x1 = 33, x2 = 29, and x3 = 13.
Y x1 x2 x3
114 21 6 5
94 43 25 8
87 56 42 25
98 19 27 9
101 29 20 12
85 34 45 21
94 40 33 14
107 32 14 11

161
119 16 4 7
93 18 31 16
108 27 12 10
117 31 3 8

2. Use the following data to determine the equation of the multiple regression model.
Comment on the regression coefficients.
Predictor Coefficient
Constant 31,409.5
x1 .08425
x2 289.62
x3 -.0947

3. Develop a multiple regression model to predict y from x1, x2, and x3 using the following
data. Discuss the values of F and t.
Y x1 x2 x3
5.3 44 11 401
3.6 24 40 219
5.1 46 13 394
4.9 38 18 362
7.0 61 3 453
6.4 58 5 468
5.2 47 14 386
4.6 36 24 357
2.9 19 52 206
4.0 31 29 301
3.8 24 37 243
3.8 27 36 228
4.8 36 21 342
5.4 50 11 421
5.8 55 9 445

162
REFERENCES:
1. [Link]
2. Black Ken; Business Statistics for contemporary decision making; 6th ed.
3. h t t p s : / / w w w. y o u t u b e . c o m / w a t c h ? v = Z k j P 5 R J L Q F 4 & l i s t = P L I e G t x p v y G -
LoKUpV0fSY8BGKIMIdmfCi&index=1, Foltz Brandon Statistics 101 Linear
Regression

163
MODULE - 14 SPSS AND DATA ANALYSIS

Structure
¡ Introduction
¡ Benefits of spss
¡ Spss for data analysis
¡ Variables
¡ Variable types
¡ Conclusion

LEARNING OUTCOMES
This chapter will enable the reader to understand
1. Meaning and relevance of SPSS
2. History of SPSS
3. Fundamentals of SPSS
4. How to effectively use it for research and data anlysis.

14.1 INTRODUCTION
"The world's leading statistical software for business, government, research and academic
organizations" - IBM SPSS
SPSS refers to "Statistical Package for the Social Sciences". It is a comprehensive, predictive
analytical computerized software package which enables ease to use set of data to users ranging
from business to statistical programmers and academicians to researchers. SPSS is used for
logical batched and non-batched statistical analytical tool.
Initially SPSS was evolved by SPSS Inc. which has a mission to 'drive the widespread use of data
in decision-making' derives directly from these two themes." The two themes are- firstly, to make
difficult analytical tasks easier for every user via improvement in ability to use and data access
and enable more number of people to attain benefits from the usage of quantitative techniques for
decision making. Secondly, to focus on analyzing data regarding people their viewpoints,
attitudes and behavior.
Later, IBM acquired SPSS Inc. in 2009. Due to this acquisition, the current version of SPSS is
named as IBM SPSS Statistics. Now it has been diversified into IBM SPSS Data collection used
for survey authoring and deployment, collecting and summarizing of data, and IBM SPSS
Modeler used for data mining, text analytics and collaboration and deployment in batch and
automated scoring services.

164
14.2 BENEFITS OF SPSS
As per the details given on the website of IBM SPSS (i.e. //[Link]/analytics/data-
science/predictive-analytics/spss-statistical-software), the company focus on making
researchers more confident about their results at every level of the process of analysis. SPSS
provides versatile method and an ease to analyze or predict the data. The SPSS product family
increases the pace and simplifies the complete analysis starting from data access and preparation
to final analysis, disposition of results and ultimately reportage. The statistical capabilities of
SPSS range from simple percentages to complex examines of variance, general linear models
and multiple regression. One can use data ranging from simple integers or binary variables to
multiple response or logarithmic variables2.

Figure 1: IBM Statistics 21


Source: [Link]

Unlike conventional statistical software, IBM SPSS enables superior analytical abilities,
flexibility and usability, to provide better quality user experience, better productivity and
performance, powerful and leading edge analytics and extensibility for information technology
infrastructure.
Following are few key aspects of IBM SPSS -
¡ The ability to quickly analyze large datasets within pivot tables Seamless integration with
Microsoft® Office applications
¡ Access to a newly improved Syntax Editor, with auto-completion, auto indentation, color-
coding and other features to make it easier to automate analytics production jobs
¡ An Interactive Model Viewer
¡ Access to multiple interface languages for global teams that may be working on the same
project
¡ Quickly prepare data in just a single step with Automated Data Preparation
¡ View significance tests in the main results table
¡ Fast performance on procedures for Frequencies, Descriptive and Crosstabs
¡ Manage and analyze business datasets

165
¡ Create customized, user-defined interfaces for existing procedures and user-defined
procedures
¡ Multithreaded procedures that improve performance and scalability
¡ Direct marketing functionality that allows business users to run their own analyses
¡ Bootstrapping capabilities that improve the stability of models
¡ Nonparametric testing procedures
¡ Regularization methods including Ridge regression, Lasso and Elastic Net, that improve
predictive models by reducing coefficient variability
¡ Multithreaded algorithms including SORT, correlation, partial correlation linear
regression, multinomial linear regression, factor analysis
¡ Nearest Neighbor analysis for prediction or classification
¡ Non-linear data modeling procedures to discover more complex relationships in your data.
¡ Cox Regression that enables survival analysis for samples drawn by complex sampling
methods
¡ Support for 64-bit hardware on desktop for Windows and Mac
¡ Support for Snow Leopard™ on Mac OS® X 10.6
¡ Support for IBM System z servers running Linux®
¡ Mac and Linux users can connect clients to IBM SPSS Statistics Server
¡ Support for Python as a "front-end" cross-platform scripting language and support for R
algorithms
¡ Collaboration capabilities boost the productivity of analysts using IBM® SPSS®
Statistics, and server-based options increase scalability and performance
¡ IBM SPSS Statistics Server makes working with large data faster and more scalable, and
improves overall stability.
¡ Improved security enables it to run as non-root on Unix/Linux.
¡ Client and server software can be on different release levels (for example, client V21 and
server V20), simplifying administration

14.3 SPSS FOR DATA ANALYSIS


1. Opening the data file
As mentioned in SPSS brief guide for SPSS 21 version, in order to open a new SPSS file one need
to follow below given steps
¡ Go to MENU option and choose:
File > Open > Data...
One can also Open the data file through the FILE button on the toolbar as shown in figure 1
Figure 1: Diagram showing the process to open a new SPSS data file

166
Figure 2: opening a new file
2. Introducing the interface
For this demonstration, we have saved the SPSS file as [Link].
a. THE DATA VIEW

Figure 3: the SPSS data view

The data file is showed in the IBM SPSS Statistics Data Editor. In the Data Editor, if you put the

167
mouse cursor on a variable name (the column headings), a more descriptive variable label is
displayed (if a label has been defined for that variable).

Figure 4: Actual data labels are displayed.

Further, to move to the first cell of the data view Press Ctrl-Home and to move to the last cell of
the data view Press Ctrl-End.
b. THE VARIABLE VIEW
To visit the variable view "Click" the Variable View tab given on the down left corner.
Review the information in the rows for each variable. In the variable view the variables are listed
in rows, with each column containing a specific kind of information of the variable. This is
different from the data view in which the variables are listed in columns (see fig 1)
To give definition of the variable double click on label id at the top of the id column. On double
clicking the names of the variable i.e. labels, in the data view variable view window will open.
Otherwise you can click on variable view and start filling each variable details.
To return back to data view click on 'data view' tab.

Figure 5: The variable view

168
c. THE OUTPUT VIEW
For the output view one need to conduct an analysis.
So we start with simple frequency table (table to count the frequency of responses for a variable).
Figure 6: Way to compute frequencies
To do the same you can follow below given steps -
Go to MENU bar and click on the 'analyze' tab. Then go to 'Descriptive statistics' in dropdown
and click on 'Frequencies' (See Fig 6)
On clicking this tab, the dialogue will appear indicating a with variable list

Figure 7: View with variable name

Figure 8: View with the variable labels

169
As shown in figure 8, each variable has an icon next to them. Every icon indicates the data type
and measurement level of each variable.

Then click on the any one option (variable) from the list. if in case, the variable label and/or name
looks shortened in the list, you can position the cursor on that label to view the complete
label/name. The variable name for EDU is shown in the square brackets after the variable label
which describes the detailed name the label. EDUCATION is the variable label. If in case there is
any variable without a variable name, it would just appear in the list box.
The dialogue box can be resized depending upon the requirement, just like windows, by clicking
and dragging the outside border or the corners. As the dialogue box becomes wider, the variable
list would also become wider.
In order to analyze the variable frequency, you can chose any from the list and click on it and then
click on icon to shift the variable from left to right. Other way to do it, you can click and drag the
variable from left to right box and then click OK to run the analysis. To enable the OK button, you
need to shift at least one variable in the Variable(s) list. In the given example, the EDU variable is
shifted to Variable(s) list and then OK is clicked.

THE OUTPUT VIEW


The output window is the window which shows the outcomes of your various enquiries such as
frequency distributions, statistical tests, cross-tabs and charts. Hence, each window in SPSS is
assigned with separate tasks.

Figure 9:The output view

170
In SPSS, each window handles a separate task.
The results of the analysis run on SPSS is shown in viewers' window in output view. You can also
go to the any item in the viewer by clicking on the list in outline pane.

THE DRAFT VIEW


The next screen that has important role in SPSS is the draft view as shown below. The draft view
shows the output of the analysis executed on SPSS. The same view can also be printed. It does not
include content pane or on the output pane notations.
To open this following are the steps
1. Go to menu bar
2. select file option
3. go to new and click on output option

Figure 10: Process to open output file


The output view has its own window which contains its own menu separate from the main SPSS
window.

Figure 11: The output view

171
Further to analyze the data and see the output- Go to 'analyze' and click on 'descriptive statistics'
and 'crosstabs'. Here you can notice a dialogue box similar to previous selected options. To
finalize the analysis, Click on OK. From the output view you can select the charts or tables, copy
them, and paste them into other applications like spreadsheets or word processors.
Note: If you want to maintain the correct spacing of the tables, use a non-proportional
font like Courier New.
The syntax view
The syntax view is fundamentally the computer code that leads to a specific output. From time to
time graphical interface is preferred to pursue the daily routine work. Though at certain times
there is a need to reproduce the steps to come to a particular conclusion. This is requird to
replicate the analysis. For this SPSS syntax view is the best one.

CROSSTABS
/TABLES=TEACHER BY AGE
/FORMAT= AVALUE TABLES
/CELLS= COUNT
/BARCHART.
In the above given code, the SPSS is tutored to make crosstabs by utilizing teacher sorting the
crosstabs by age through using a specific format. Further it puts a count onto each cell and thus
making a bar chart

Figure 12: the syntax view

You should keep in mind to preserve the syntax code ones used., especially in case you are
calculating results for writing papers, reports and like.

172
Following are the steps to prepare charts and frequency distribution by executing saved syntax.
d. Go to menu bar
e. Click on analyze > Descriptive statistics > Crosstabs (keep the previous selection of
syntax)
f. Instead of pressing OK, click on Paste
g. SPSS will automatically bring the Syntax Editor with the code you just have pasted.
h. Run the syntax and get the output.

Crosstab
Crosstab refers to a short form of Cross tabulation which represents a summary table
emphasizing on the summary. The crosstabs are used for categorical data or discrete data like
gender or employment status. The crosstabs cannot be used for data which is continuous in nature
like income, dosage etc. such data can only be entered in crosstab by converting the continuous
data into groups like less than Rs 12000, between Rs.12000 to Rs.25000 and Rs. 25000 and
above.

Figure 13: Process to open crosstabs

3. ENTERING AND MODIFYING DATA


Crafting the definitions of the data
Data refers to a set of values of quantitative or qualitative variables. It can be either nominal (data
to which numbers are given as code/ label to define attributes), Ordinal (data can be put in order
without having numerical meaning beyond the order e.g. in 5 point Likert scale 1 denotes
strongly disagree while 5 denotes strongly agree); Interval (numerical data with meaningful
distance between numbers, except zero e.g. 60°c hot tea is cannot be considered as twice hot to
one at 30°c), and Ratio (numerical data with meaningful distance between numbers including
zero e.g. weight of A is 90 Kgs can be considered as twice heavier to B with weight of 45 Kgs).
Thus when quantities are defined for a variable, it is considered as relevant data for analysis.
Though before understanding the further analytical process, it is very important to understand
the types of variables available in statistics.

173
VARIABLES:
Variables can be defined as a particular kind of information. Income, gender, or temperature can
be considered as variable. A few confuse in the terms like "concepts" and "variables". However,
concepts are the mental images or perceptions. Its meanings vary evidently from individual to
individual. Whereas variables are measurable, of course with varying degrees of accuracy.

VARIABLE TYPES
According to SPSS Step-by-Step Tutorial: Part 1, SPSS uses (and insists upon) what are called
strongly typed variables. Strongly typed means that you must define your variables according to
the type of data they will contain. A user can use any of the variable types, as defined by the SPSS
Help file. The SPSS data editor ACCEPTS following forms of numeric strongly types variables-
¡ The basic form of variable is the standard numeric format which is required to be provided
in scientific notation or standard format.
¡ Comma is the another numeric variable, in which value are indicated with commas to
delimit every three places and with period as decimal delimiter.
¡ Dot is the variable which is exhibited with periods delimiting every three places and
comma as a decimal delimiter.
¡ Scientific notation is a numeric variable whose values in shown with embedded E and a
signed power of ten exponents. Such variable can be either preceded by E or D with an
optional sign or a sign alone.
¡ Date is another numeric variable whose values are displayed in one of several calendar date
or clock-time formats. A user can select the format like dates with slashes, hyphens,
periods, commas, or blank spaces as delimiters and enter the data.
¡ Custom currency are the variables which are displayed in one of the custom currency
formats that you have defined in the Currency tab of the Options dialog box.
¡ String are the values of the string variables which is neither numeric nor used for
calculations. Also known as alphanumeric variables, such variables consist of characters
up to a certain defined length, which distinct uppercase and lowercase

Variable names and labels


In SPSS, no funny characters like spaces or hyphens are used. Only variable names which are
within the limit of eight characters, period are used. Following are the related rules
¡ Variable names must not end with a period
¡ Variable names must begin with a letter.
¡ Variable names must be no longer than eight characters.
¡ Variable names cannot contain blanks or special characters.
¡ Variable names must be unique.
¡ Variable names are not case sensitive

174
Missing values
If you do not enter any data in a field, it will be considered as missing and SPSS will enter a period
for you.

Save the work


As indicated above, SPSS acts differently from normal windows. So does the saving process of
SPSS. In order to save the work done by the researcher on SPSS, they need to follow below given
process
Step 1: File -> Save As
Step 2: Locate the location where you want to save the file.
Step 3: Give name to your file
Step 4: Save
The Data Editor files are saved as .sav, while the output files (from the SPSS Viewer) are saved as
.spo. You can also select the format in which you wish to save
Step 1

Figure14: Go to save as to save the file


Step 2

Figure 14: location search

175
Step 3

Figure 15: File types


Cutting and pasting
Every time an output is derived, it provides an option to the researcher to copy and paste that table
or diagram to word doc, with all the formatting preserved. This is useful when you want to
prepare a lab report (or paper) and want to insert a graph or table.
You can also right-click the name of an object in the left-hand pane of the SPSS Viewer and do the
same.
Step 1

176
Step 2

Correlations
tech_led is
Pearson
1 .502**
Correlation
tech_led
Sig. (2-tailed) .000
N 498 498
Pearson
.502** 1
Correlation
is
Sig. (2-tailed) .000
N 498 498
**. Correlation is significant at the 0.01 level (2 -
tailed).

EXPORTING THE OUTPUT


If in case, while pasting the graph in word doc, the images gets automatically cropped, in that
case it is better to export the output to word doc. To do this go to File, then Export.

177
The window below will pop up, and ask you to choose where to save it ("Browse…").

The default location will be shown as in the hard drive. You need to select the location carefully
like in the desktop or the Documents folder).

You can also mention the SPSS that in which category of file you want to save the file.

178
You will usually want to select "All Visible Objects" and export it as a Word/RTF (.doc) file. This
is the easiest way to save all your work in useful format (RTF is Rich Text Format, which can be
read in nearly any text application on any platform

SPSS Functionality
SPSS has a very flexible data handling capability. SPSS provides has huge range of statistical and
mathematical functions, statistical procedures and can read data in almost any format (e.g.,
numeric, alphanumeric, binary, dollar, date, time formats) as explained above. To the help of
researcher SPSS also has excellent data manipulation utilities.
The following is a brief overview of some of the functionalities of SPSS:
¡ Data transformations
¡ Descriptive Statistics
¡ Data Examination
¡ Reliability tests
¡ Contingency tables
¡ Correlation
¡ T-tests
¡ ANOVA
¡ MANOVA
¡ General Linear Model (Release 7.0 and higher)
¡ Regression
¡ Logistic Regression

179
¡ Nonlinear Regression
¡ Loglinear Regression
¡ Factor Analysis
¡ Discriminant Analysis
¡ Cluster anlaysis
¡ Probit analysis
¡ Multidimensional scaling
¡ Survival analysis
¡ Forecasting/Time Series
¡ Graphics and graphical interface.
¡ Nonparametric analysis

CONCLUSION
This chapter aims to give basic information about the most sorted statistical software for data
analysis in social sciences and management. This software eases the data entry, processing and
output derivation, which eliminates time taking manual process of analysis. Fow effective usage
of SPSS, it is very important to understand varied aspect of SPSS and then start working for your
research. There are many more information related to SPSS which is provided in SPSS tutorials
provided from time to time by the IBM with every upgraded version of SPSS.

REFERENCES
¡ Arkkelin, D. (2014). Using SPSS to understand research and data analysis.
[Link]
¡ Dan Flynn (n.d.), Guide to SPSS, Barnard College
[Link]
¡ Landau, S. (2004). A handbook of statistical analyses using SPSS. CRC.
[Link]
nalyses_using_SPSS.pdf
¡ Garth, Andrew (2008), Analysing data using SPSS, Sheffield Hallam University.
[Link]

Webpages
¡ [Link]
¡ [Link]
[Link]
[Link]
¡ [Link]
¡ [Link]
¡ [Link]

180
T2217
BUSINESS STATISTICS
MBA I SEM I

ISBN: 978-93-95877-05-3

SYMBIOSIS INTERNATIONAL (DEEMED UNIVERSITY)


Gram: Lavale, Tal: Mulshi, Dist: Pune, Maharashtra, India Pin: 412115

You might also like