Decision Science
Decision Science
1
Course Code: 22ONMBT***
Course Objective:
1. The objective of this course is to acquaint the learners with various statistical tools
2. The course aims at providing fundamental knowledge and exposure to the learners to
use various statistical methods to understand, analyze and interpret data for decision
making.
3. The course has universal utility in making learners ready for corporate jobs and
entrepreneurial ventures.
Course Description:
This course begins with the basics of Statistics. It further elaborates on various sources of
data, depiction of data and applicability of the measure of the central tendencies and
dispersion. Later on the course emphasizes on the knowledge and applications of the
correlation, regression and time series data in real time situations. Finally, the learners got
a chance to acquaint themselves with the basics of Hypothesis Testing and applications
thereof.
2
Table of Contents
Module 1 .............................................................................................................................................................. 4
1.1: Basics of Statistics ...................................................................................................................... 4
[Link] Collection Sources ............................................................................................................ 19
1.3 Descriptive Statistical Methods................................................................................................ 39
Module 2 .......................................................................................................................................................... 100
2.1 Correlation Analysis ............................................................................................................... 100
2.2 Regression Analysis ................................................................................................................ 126
2.3. Time Series Analysis .............................................................................................................. 155
Module 3 .......................................................................................................................................................... 200
3.1. Types of Analysis.................................................................................................................... 200
3.2. Basics of Hypothesis Testing ................................................................................................. 229
3.3 Applications of Hypothesis Testing ....................................................................................... 251
3
Module 1
Learning outcomes
1.1.1 Meaning
The collection, description, analysis, and drawing of conclusions from quantitative data are
all part of the field of statistics, which is applied mathematics. Calculus of differential and
integrals, linear algebra, and probability theory are key mathematical concepts in statistics.
The study of statistics involves the manipulation of data, including methods for data
statistics are the two main subfields of statistics. There are various ways to communicate
point (ratio-level). Simple random, systematic, stratified, or cluster sampling are just a few
sampling methods employed to gather statistical data. Almost every division of every
The scope of Statistics can be very well appreciated as the researcher follows:
4
Statistics help in economic planning
To solve economic issues, statistical data and other statistical analysis approaches are
extremely helpful. Economic plans are formed on the basis of statistical information.
Statistics are also used to assess the plan's effectiveness. Numbers can be used to express
poverty, etc.
corporate policies and predicting future trends. The foundation of modern business is the
precision and accuracy of estimates and statistical forecasting of future product demand,
market trends, and other factors. Businessmen's ability to accurately estimate and make
Statistics in administration
they have been used to gather data on governmental fiscal and military policy.
The state compiles a vast amount of data on many characteristics of the populace and uses it
Statistics in research
Every form of research project makes heavy use of statistical methodologies. The use of
statistics in doing various sorts of research is helpful in all fields, including social science,
5
With the use of the proper statistical techniques, the yield of a crop grown with various seeds,
fertilisers, etc., may be studied. In a similar vein, statistics are very useful to the insurance
industry as well as to other sectors, including astronomy, banking, railroads, and the physical
sciences.
Every form of research project makes heavy use of statistical methodologies. The use of
statistics in doing various sorts of research is helpful in all fields, including social science,
With the use of the proper statistical techniques, the yield of a crop grown with various seeds,
fertilisers, etc., may be studied. Similarly, statistics are very useful to the insurance industry
In the contemporary era, often known as the age of planning, statistics are essential to
Statistical data and techniques of statistical analysis are the two terms that have to be
immensely useful in involving economic problems, such as wages, price, time series analysis,
Statistics is an irresponsible device for production control. Business executives are counting
on more and more on statistical techniques for studying the much and desires of the important
customers.
6
4. Statistics and industry
In industry, statistics are widely used in quality control. In production engineering, to find out
whether the product is conforming to the specifications or not, various statistical tools, such
Statistics are intimately related to recent advancements in statistical techniques and are the
In medical science, the statistical tools for collection, presentation, and analysis of observed
facts relating to causes and incidence of disease and the result of the application of various
In education and physiology, statistics have found wide applications such as determining or
determining the reliability and validity of a test, factor analysis, etc. There are many
Branches of Statistics
Data collection, descriptive statistics, and inferential statistics are the three main subfields of
7
1. Data Collection- The process of gathering data is what matters most. In terms of
mathematics, we often don't need to worry too much about this because we merely use the
data that is provided, but there are important considerations to make when gathering data.
This is quite simple for data like test scores from a class. Since each student has a specific
mark assigned to them, the data set is simply made up of all the marks. Data collection is
occasionally a challenge. Because bees move and fly around, it can be difficult to count them;
in these situations, you might have to estimate. Additionally, you should be mindful about
where you receive your data from if you are collecting it. Assume, for instance, that you want
to survey people about their election preferences. You must select a representative sample of
people because it is impractical to ask the entire nation (the population). It's not as simple as
it seems. For instance, polls were occasionally conducted in the middle of the 20th century by
This seems representative, but only the wealthy had telephones in those days, so you were
only polling a small portion of society—a group that would be more likely to support one
party than the other. When doing a survey by email, the same problem could arise today.
As a result, there are problems with data collection, and you must do prior to dealing, be
certain that the info has been gathered fairly. with it, attempt to communicate it, and draw
inferences.
When doing a survey by email, the same problem could arise today. As a result, there are
problems with data collection, and you must do prior to dealing, be certain that the info has
been gathered fairly. with it, attempt to communicate it, and draw inferences.
2. Descriptive Statistics- As previously said, it's frequently impractical to gather data from
the complete population because it would be very costly and time-consuming and the
8
population might change while you collect the data. Instead, we frequently need to take a
sample.
Every bit of information needs to be summarised because if you just list it out in its entirety,
no one will understand it. Imagine if every single individual surveyed by a polling
organisation had their votes posted on the TV news; it would be a massive list of parties, and
you couldn't draw any conclusions. Instead, you are given visual representations of the data
(a bar chart, for instance) that may indicate the proportion of votes each party received. In the
general election of 2010, close to 30 million people cast ballots. You would be completely
bewildered if every vote were simply listed and displayed one after the other; instead, a
summary of votes is given (for example, as percentages: Conservative 36%, Labour 29%,
Liberal Democrat 23%, Others 12%). This is an illustration of descriptive statistics, which
variability, also referred to as measures of dispersion, are the foundation of all descriptive
statistics.
Frequency Distribution- The frequency distribution shows the frequency or count of the
various outcomes in a data collection or sample, and is used for both quantitative and
qualitative data. Typically, the frequency distribution is displayed as a table or graph. The
count or frequency of the values' occurrences within an interval, range, or particular group
A presentation or summary of grouped data that has been divided into mutually exclusive
classes and the number of occurrences in each class is called a frequency distribution. It
9
enables the presentation of raw data in a more structured and orderly manner. Bar charts,
histograms, pie charts, and line graphs are examples of common charts and graphs used in
Central Tendency- A dataset's descriptive summary utilising a single value that represents
the centre of the data distribution is referred to as having a central tendency. Measures of
central location are another name for measures of central tendency. The measurements of
The average or most frequent number in a data set is known as the mean, which is regarded
as the most widely used measure of central tendency. The middle score for a data collection
in ascending order is referred to as the median. The score or value that appears most
Variability- A summary statistic that reflects the level of sample dispersion is known as a
measure of variability. How far away the data points appear to fall from the centre is
The range and width of the distribution of values in a data set are indicated by the terms
dispersion, spread, and variability. The spread's many elements and features are represented
10
The distance between the greatest and lowest values within a data collection is idealised as
the range, which shows the degree of dispersion. The standard deviation is used to calculate
the average variance in a set of data and gives information about how far a value from a data
set is from the mean value of the same data set. The variance, which is just an average of the
3. Inferential Statistics- The part of statistics that deals with drawing inferences from the
data is called inferential statistics. You are simply asking, "What is this data saying to us, and
After a series of incidents, a council, for instance, might be considering lowering the speed
limit on a major road. In order to determine whether the speed limit should be lowered (for
instance, if a lot of automobiles are going too quickly), they might survey the speeds of cars
(gather data) first. Be aware, however, that this may not be the case; everyone may be
travelling at a pace that is totally acceptable, and the incidents may not be related to speed at
all. It's called inferential statistics when you take your available data and draw a "inference"
or "conclusion" from it. When we talk about topics like hypothesis testing, where we check to
see whether the data backs up a claim we make, we'll see much more of this in the future.
11
BASIS FOR
DESCRIPTIVE STATISTICS INFERENTIAL STATISTICS
COMPARISON
Operations Organize, analyze and present Compares, test and predicts data.
data in a meaningful way.
Statistics is designed to study a group of values instead of single observation studies the mass
In general, statistics studies only the quantitative characters of the given problems instead of
qualitative characters. The problems which cannot be studied quantitatively (i.e. in numerical
12
form), such as poverty, leadership, beauty, intelligence, honesty, etc., are not directly studied
in statistics.
There are only a few results that are accurately correct in statistics, and almost all are only
Statistics deals with averages that are obtained from different individual items. The laws of
The best solution under all conditions of the given problem is not p by the statistical
methods.
Statistics cannot be of much help in studying the provided problem, like a country's culture,
Mean, median, mode, variance, and other descriptive statistics are used to describe the key
aspects of the data. Through numbers and graphs, it summarizes the data.
Inferential statistics allows us to infer information about the population from a sample.
Inferential statistics' major objective is to extrapolate findings from the sample to the
population data. For instance, we need to determine the typical data analyst salary in India.
13
1. The first option is to study the data of data analysts across India and ask them about their
2. The second option is to use a sample of data analysts from the central IT cities in India
The first alternative is not feasible since it is difficult to compile data from all data analysts in
India. Both time and money are spent on it. To solve this problem, we will consider the
second alternative: to gather a small sample of data analyst salaries and use the average of
those prices as the Indian average. Inferential statistics are those that allow us to infer
Key Terms- Statistics, Data, Decision making, Inferential Statistics, Descriptive Statistics,
Summary- In this unit, we have discussed the basic idea of statistics, applications, and
limitations. Also students can get the idea of Descriptive and Inferential Statistics in this unit
Case Study:
A portfolio manager closely monitors the price-earnings ratios (defined as the current
market price divided by earnings over the previous four quarters) of 200 common
equities. He believes that it is appropriate to become an active buyer when the majority
norms. Low P-E ratios may indicate that investors in general are overly pessimistic.
14
Furthermore, equities with low P-E ratios gain from rising earnings in two ways: (a) higher
earnings multiplied by a stable P-E ratio equals a higher market price, and (b) rising
(c) Display the frequency distribution as a histogram or frequency polygon and comment
on the pattern.
15
Exercise: Choose the correct option
1. The arithmetic mean of the scores of a group of students in a test was 52. The
brightest 20% of them secured a mean score of 80 and the dullest 25% a mean score of
A. 45
B. 50
C. 51.4 approx
D. 54.6 approx.
2. Let A1, A2, be two AM’s and G1, G2 be two GM’s between a and b,then (A1 + A2) /
a) (a+b) / 2ab
b) 2ab/(a+b)
c) (a+b)/(ab)
a) AP
b) GP
c) HP
a) AP
b) GP
c) HP
16
5. If A and G be the A.M and G.M between two positive numbers then the numbers are
a) True
b) False
6. If one geometric mean G and two airthmetic mean A1, A2 are inserted between two
a) 2G
b) G
c) G2
AM ≤ GM.
a) True
b) False
a) 4
b) 5
c) 6
d) 16
9. The kurtosis defines the peak of the curve in the region which is
17
d. around the variance
10. In kurtosis, the beta is greater than three and quartile range is preferred for
a. mesokurtic distribution
c. leptokurtic distribution
d. platykurtic distribution
1. What is statistics?
Answers
Exercise: MCQ
1. (C)
2. (C)
3. (B)
18
4. (D)
5. (D)
6. (B)
7. (A)
8. (C)
9. (D)
10. (B)
A collection of measurements and facts is referred to as data, and it can be utilised to provide
an individual or organisation with the knowledge they need to make a well-informed choice.
It helps the analyst understand, examine, and evaluate a variety of socioeconomic problems,
determining their sources so that potential solutions can be developed. Along with theoretical
information, data may also contain specific numerical facts that lend weight to the claim.
Data collection is the initial stage of performing a statistical study, and it can be done using
Primary Source
19
It is a synthesis of data from the primary source. It provides the researcher with raw,
quantitative data that is directly related to the statistical investigation. In other words, the
primary sources of information give the researcher immediate access to the research issue.
Secondary Sources
already acquired data from primary sources. Quantitative and unprocessed data that are
pertinent to the investigation are not made available to the researcher directly. In turn, the
illustration, reviews, scholarly books and journals, articles, and government websites that
A solid study will need both primary and secondary sources of data collection, even though
primary sources provide the information more credibility because they include supporting
evidence.
It's important to first define data before discussing collection. The short answer is that data is
the act of gathering, gauging, and analysing precise data from a range of pertinent sources in
order to address issues, provide answers, assess results, and predict trends and possibilities.
Data collection is crucial because of how heavily dependent our society is on it. To assure
quality assurance, keep research integrity, and make educated business decisions, accurate
The researchers must specify the data sources, data types, and methodologies used during
data gathering. We'll quickly find that there are numerous approaches to gathering data.
There is heavy reliance on data collection in research, commercial, and government fields.
20
Before an analyst begins collecting data, they must answer three questions first:
What methods and procedures will be used to collect, store, and process the information?
Data can also be divided into qualitative and quantitative categories. Descriptions like colour,
size, quality, and appearance are all included in qualitative data. Unsurprisingly, quantitative
data involves numbers. Examples include statistics, poll results, percentages, etc.
Let’s get into specifics. Using the primary/secondary methods mentioned above, here is a
Interviews
The researcher surveyed a sizable sample of people through in-person interviews or other
forms of mass contact like phone or mail. This approach is by far the most typical way to
collect data.
Projective Technique
When prospective respondents are aware of the purpose of the interview and are reluctant to
reply, projective data collection is performed. For instance, if a representative from a cell
phone company asks them, a person might be hesitant to answer inquiries concerning their
21
Delphi Technique
Greek mythology describes the Oracle at Delphi as the high priestess of Apollo's temple who
offered counsel, prophesies, and wisdom. Researchers utilise the Delphi method to acquire
data by asking a group of experts for their opinions. Each expert responds to inquiries in their
area of expertise, and the responses are combined to form a single view.
Focus Groups
Focus groups and interviews are both frequently employed methods. A moderator gathers the
group of anywhere between six and twelve participants and serves as its facilitator.
Questionnaires.
of open-ended or closed-ended questions about the topic at hand are presented to the
respondents.
Unlike primary data collection, there are no specific collection methods. Instead, since the
information has already been collected, the researcher consults various data sources, such as:
Financial Statements
Sales Reports
Retailer/Distributor/Deal Feedback
Business Journals
22
Government Records (e.g., census, tax records, Social Security info)
Trade/Business Magazines
The Internet
The data must be structured and presented in a coherent and understandable manner in order
to assist in further statistical analysis because the acquired data, also known as raw data or
reduce a large amount of data into ever-smaller, more digestible chunks. Classification is the
process of dividing data into distinct classes or subclasses based on specific criteria, whereas
tabulation deals with the orderly arrangement and presentation of classified data. Therefore,
the initial stage in the tabulation is classification. For Example, letters in the post office are
classified according to their destinations, viz., Delhi, Madurai, Bangalore, Mumbai etc.,
Objects of Classification: The following are the main objectives of classifying the data:
4. It enables one to get a mental picture of the information and helps in drawing inferences.
23
Types of classification: Statistical data are classified in respect of their characteristics.
a) Chronological classification
b) Geographical classification
c) Qualitative classification
d) Quantitative classification
arranged according to the order of time expressed in years, months, weeks, etc., The data is
Example-
according to geographical region or place. For instance, the production of paddy in different
Example-
the same qualities or traits, such as sex, literacy, religion, employment, etc. Such qualities
cannot be quantified using a scale. For instance, if the population needs to be categorised
24
according to one characteristic, let's say sex, we can divide them into two groups, males and
females. On the basis of other criteria, "marriage status," they can also be categorised as
"married" or "single."
nature, two classes are created, one that possesses the attribute and the other that does not.
Population
Male Female
A manifold categorization is one that takes into account two or more attributes and creates
many classifications.
sex and marital status, the population is first divided into males and females. On the basis of
the attribute "employment," each of these classes may then be further divided into "married"
and "single," and as a result, the population is divided into four classes.
25
Still the classification may be further extended by considering other attributes like marital
height, weight, etc., data is categorised quantitatively. For instance, the group of kids might
(i) the variable (i.e.) the weight in the above example, and
There are 50 children with weights ranging from 5 to 10 kg, 200 children ranging between 10
to 15 kg and so on.
Types of Tables
26
There are different types of tables based on a different basis. On the basis of the objective or
purpose of the study, the tables are classified into two types, viz.
On the basis of the nature of the data, the tables are classified into two types, viz.
Further, on the basis of the elements or characteristics covered, tables are broadly classified
Tabulation is the process of condensing categorised or grouped data into a table format for
27
organised set of categorised data in columns and rows is called a table. A statistical table
enables the researcher to present a vast amount of data in a detailed and organised manner. It
makes comparison easier and frequently identifies patterns in data that would not be apparent
otherwise. In actuality, classification and "tabulation" are not two separate procedures. In
reality, they work in tandem. Data are categorised before being tabulated, and they are then
Depiction of Data
The relationship between facts, ideas, information, and concepts is shown in a diagram
lines that cross the stated coordinate point on the surface. By enabling one to assess how
the amount of one variable has changed with respect to another over time, it facilitates the
research of a relationship between two variables. It is useful for analysing series and
Definition: The investigator must tabulate the data after gathering them in order to
evaluate their prominent characteristics. The presentation of data is the term used for such
a setup.
Any obtained data may be arranged in a frequency distribution table and then displayed
using pictographs or bar graphs. The lengths of the equally wide bars that make up a bar
28
graph, which is a depiction of numbers, depend on the frequency and scale you select. The
collected raw data can be placed in any one of the given ways:
2. Ascending order
3. Descending order
Example: Let the marks obtained by 30 students of class VIII in a class test, out of 50
39,25,5,33,19,21,12,41,12,21,19,1,10,8,1239,25,5,33,19,21,12,41,12,21,19,1,10,8,12
17,19,17,17,41,40,12,41,33,19,21,33,5,1,2117,19,17,17,41,40,12,41,33,19,21,33,5,1,21
The data in the given form is known as raw data or ungrouped data. The above-given data
1. Bar chart
3. Histogram
29
4. Pie chart
5. Line graph
The qualitative data is represented visually in the bar graph. The data is presented either
frequencies.
The bars are organised in frequency order, emphasising the more important groups. It is
simple to determine which sorts of data in a collection outweigh the others by looking at
all the bars. Bar graphs come in a variety of forms, including single, stacked, and grouped.
A frequency table or frequency distribution is a way to present raw data that makes it
30
Using the tally marks, the frequency distribution table is created. A type of numerical
system that uses vertical lines for counting is called a tally mark. To get a total of 55, the
Example:
Consider a jar containing the different colours of pieces of bread as shown below:
31
Another type of graph that displays data using bars is the histogram. The histogram is used
for numerical data, and classes, or ranges of values, are presented at the bottom. The
categories with higher frequencies have thicker bars than the other types.
Although a histogram and a bar graph appear quite similar, they differ due to the data
level. The frequency of the categorical data is represented in bar graphs. A categorical
The numerical distributions of a dataset are depicted using a pie chart. In this graph, a
circle is divided into various sectors, where each sector reflects the percentage of a
32
Graphical Representation of Data: Line Graph
A line graph is a type of graph that use points and lines to show change over time. In other
words, the chart that depicts a line connecting several points or a line illustrating the
The straight line or curve that connects a number of succeeding data points in the graphic
serves as an illustration of the quantitative data between two changing variables. Two
variables are compared on the vertical and horizontal axes in linear charts.
33
Key Terms- Graphs, Tables, Data, Depiction of Data, Bar graphs, Pie Charts, Histogram,
Frequency.
Summary
This unit represents the various data representation methods in statistics such as bar graph,
Case Study:
The welfare committee of a big housing complex is investigating the prospect of hiring
private security guards to patrol the complex's front gate 24 hours a day, seven days a
week. The housing complex has 810 flats, and owners were asked to vote for or against
Yes 194
No 121
Not sure 73
No response 422
(a) Convert the data to percentages and create a bar chart and a pie chart. Which of these
34
(a) After removing the 'no response' group, convert the remaining 388 replies to
Assume you have been designated as a poll officer; what would you like to advise to the
A. 11
B. 13
C. 17
D. 23
2. Find the median of the call received on 7 consecutive days 11,13, 17, 13, 23,25,19
A. 13
B. 23
C. 25
D. 17
A. 12,9
B. 7,9
C. 7,12
D. 11,9
4. When the Mean of a number is 18, what is the Mean of the sampling distribution?
A. 21
B. 18
C. 27
D. 23
35
A. Frequency distribution
C. Tabulation
D. Classification
6. When the quantitative and qualitative data are arranged according to a single
A. One-way
B. Bivariate
C. Manifold division
D. Dichotomy
A. Frequency distribution
B. Chronological distribution
C. Ordinal distribution
D. Nominal distribution
36
9. What are the general tables of data used to show data in an orderly manner known
as?
B. Manifold tables
C. Repository tables
10. What is the table where the variables are subdivided with interrelated features
known as?
B. Sub-parts of a table
C. One-way table
D. Two-way table
37
3. How data can be represented?
Answers
Exercise: MCQ
1. (B)
2. (D)
3. (C)
4. (B)
5. (C)
6. (A)
7. (A)
8. (A)
9. (C)
10. (D)
38
1.3 Descriptive Statistical Methods
Depending on the nature of the data and the goal for which it was obtained, the description of
statistical data may be highly detailed or quite succinct. One must watch out for being neither
too brief nor too long when describing data verbally or statistically. We can compare two or
more distributions for the same time period or within the same distribution over time using
the measures of central tendency. For instance, using an average, it is possible to compare the
average tea consumption in two separate regions for the same time period or in one region
over a period of two years, such as 2003 and 2004. The calculation of arithmetic mean by the
Adding all the observations and dividing the sum by the number of observations results the
It may be noted that the Greek letter μ denotes the mean of the population and n denotes the
39
μ = ∑x/n. The formula given above is the basic formula that forms the definition of
arithmetic mean and is used in the case of ungrouped data where weights are not involved.
For grouped data, arithmetic mean may be calculated by applying any of the following
methods: (i) Direct method, (ii) Short-cut method, (iii) Step-deviation method 24. In the case
of the direct method, the formula x = ∑fm/n is used. Here m is the mid-point of various
classes, f is the frequency of each class, and n is the total number of frequencies. The
Example- The following table gives the marks of 58 students in Statistics. Calculate the
Solution-
40
It may be noted that the mid-point of each class is taken as a good approximation of the true
mean of the class. This is based on the assumption that the values are distributed fairly evenly
throughout the interval. When large numbers of frequencies occur, this assumption is usually
accepted.
In the case of the short-cut method, the concept of the arbitrary mean is followed. The
formula for calculation of the arithmetic Mean by the short-cut method is given below:
f = frequency
The use of the direct method would be quite laborious when the numbers are really big and/or
in fractions. The expedient approach is preferable in such circumstances. This is due to the
fact that the computation effort required by the shortcut technique is significantly decreased,
41
especially when it comes to calculating the product of values and their related frequencies.
However, suppose calculations are done automatically with a calculator rather than manually.
In that case, it might not be essential to employ the shortcut approach as using the direct way
The shortcut technique employs an arbitrary or assumed mean, as is evident from the
formula. The correction factor for the discrepancy between the actual mean and the assumed
mean is represented by the second component in the formula (∑fd ÷ n) . (∑fd ÷ n) will be 0
if the assumed mean proves to be the same as the real mean. The short-cut method's
application is predicated on the idea that the sum of deviations from the true mean is equal to
zero. As a result, how the assumed Mean is related to the actual mean will determine how
deviations are calculated from any other number. For the figures given earlier pertaining to
marks obtained by 58 students, we calculate the average marks by using the short-cut method
Example-
It may be noted that we have taken arbitrary mean as 35 and deviations from midpoints. In
other words, the arbitrary mean has been subtracted from each value of the mid-point and the
42
Now we take up the calculation of arithmetic mean for the same set of data using the step-
deviation method.
It will be clear that the solution is the same in all three situations. Due to its easier
computations, the step deviation approach is the most practical. It should be highlighted that
the result would remain the same even if we chose a different arbitrary mean and recalculated
departures from that value. Since we now understand the various ways the arithmetic mean
can be determined, we are equipped to deal with any situation where the calculation of the
43
1.3.2 Weighted Average MEAN
The average of the provided data set is the weighted mean, sometimes referred to as weighted
average. It is an average that was determined by giving various weights to various individual
variables. The weighted mean and the arithmetic mean are equivalent if all of the values are
the same.
The average mean or arithmetic mean is the same as the weighted mean. When data is
presented in various forms, it is calculated and compared to the sample mean or arithmetic
mean. Although weighted mean and average mean typically behave similarly, they do have
some divergent characteristics. Higher weighted data values than lower weighted data values
contribute to a greater weighted mean. Negative weights cannot exist; some of them may be
zero, but not all of them, because division by zero is forbidden. Compared to weights with a
In the case of ungrouped data where weights are involved, our approach for calculating the
The average of all values that are prioritised is known as a weighted average. The weighted
average of values is calculated by dividing the total weight by the total value.
The following weighted mean formula can be used to calculate the weighted mean for a given
44
Example 1: Suppose a student has secured the following marks in three tests:
Laboratory -25
Final -20
collection of numbers' values, denotes the central tendency of that set of numbers. Measures
of central tendencies in mathematics and statistics are used to summarise the values of the
entire data set. Mean, median, mode, and range are the key metrics for identifying central
tendencies. One of these gives a general understanding of the data, which is the data set's
mean. The data set's mean indicates what the data set's average number is. The various mean
types are arithmetic mean (AM), geometric mean (GM), and harmonic mean (HM).
The Geometric Mean (GM) is the average value or mean that, by taking the product of a set
of numbers' values as its root, indicates the set's central tendency. In essence, where n is the
45
total number of values, we multiply all 'n' values and subtract the nth root of the numbers. For
instance, the geometric mean for a pair of numbers, such as 8 and 1, is equivalent to √(8×1) =
√8 = 2√2.
As a result, the geometric mean is also described as the product of n numbers at the nth root.
Keep in mind that this is not the same as the arithmetic mean. Data values are summed, then
divided by the total number of values to determine the arithmetic mean. In contrast, the given
data values are multiplied in a geometric mean, and the final product of the data values is
obtained by taking the root of the radical index. Take the square root, for instance, if you
have two data values. If you have three data values, take the cube root; if you have four, take
There are two additional methods that are occasionally employed in business and economics
in addition to the three central tendency measures already mentioned. The geometric mean
and the harmonic mean are these. The harmonic mean is less significant than the geometric
mean. Below, we go over both of these methods. We start with the geometric mean. The
geometric mean of a distribution is determined by taking the nth root of the product of its n
observations. Symbolically,
46
= (1/n) log (x₁ · x₂ · ... · xₙ)
= (∑ log xᵢ) / n
Similar calculations must be made if there are three observations: we must determine the
cube root of the product of these three observations, for example, and so on. The more
elements there are, the more challenging it is to multiply the integers and determine the root.
For example, the given data sets are: For example, for data set, 4, 10, 16, 24
10, 15 and 20 Here n = 4
Here, the number of data points = 3 Therefore, the G.M = (4 ×10 ×16 ×
Arithmetic mean or mean = 24)1/4
(10+15+20)/3 = 153601/4
47
Before learning the relationship between the AM, GM, and HM, it is necessary to understand
their respective formulas. Assuming "a" and "b" are the two numbers and that there are 2
values,
AM = (a+b)/2
GM = √(ab)
⇒HM = 2/[(a+b)/ab
HM = GM2 /AM
⇒GM2 = AM × HM
Or else,
GM = √[ AM × HM]
As a result, 𝐺𝑀2 = AM × HM is the relationship between AM, GM, and HM. As a result,
the geometric mean's square is equal to the sum of the harmonic and arithmetic means.
Let us also see why the G.M for the given data set is always less than the arithmetic mean for
So,
48
Now let’s subtract the two equations
A−G ≥ 0
The geometric mean is utilised in many fields and has numerous advantages over the
Since many of the value line indexes used by financial departments employ G.M. to
determine the annual return on the investment portfolio, it is used in stock indexes.
In finance, the geometric mean is used to determine average growth rates, commonly
Additionally, studies on biological processes like bacterial growth and cell division use
geometric means.
Median
When the data are ordered in an ascending or descending order of magnitude, the median is
defined as the value of the middle item (or the mean of the values of the two middle items).
In an ungrouped frequency distribution, the median is the middle value if n is odd, and the n
values are sorted in ascending or descending order of magnitude. The median is the mean of
49
We have to first arrange it in either ascending or descending order. These figures are arranged
5,7,10,15,18,19,21,25,33
Now, as the series consists of an odd number of items, to find out the value of the middle
n+1
𝑚𝑒𝑑𝑖𝑎𝑛 =
2
n+1
=5
2
That is, the size of the 5th item is the median. This happens to be 18.
Assume there are 23 total items in the series. In order to include 23 in the above sequence at
the proper location, that is, between 21 and 25, we may have to do so. The series now
consists of 5, 7, 10, 15, 18, 19, 21, 23, 25 and 33. Using the calculation above, the median
size is the 5.5th item. Here, we must average the values of the fifth and sixth items. With an
n+1
It may be noted that the formula itself is not the formula for the median; it merely
2
indicates the position of the median, namely, the number of items we have to count until we
arrive at the item whose value is the median. In the case of the even number of items in the
series, we identify the two items whose values have to be averaged to obtain the median. In
50
the case of a grouped series, the median is calculated by linear interpolation with the help of
l2 + l1
𝑀 = l1 (𝑚 − 𝑐)
f
m = the middle item or (n + 1)/2nd, where n stands for the total number of items
c = the cumulative frequency of the class preceding the one in which the median lies
Mode
The mode is a different way to quantify central tendency, and it is the value at the location
where things are concentrated most intensively. Consider the following sequence as an
illustration: 8,9, 11, 15, 16, 12, 15,3, 7, 15 37 The greatest number of times figure 15 appears
in a sequence of ten observations is three. Therefore, the mode is 15. Because the series listed
above is discontinuous, the variable cannot be in the fractional form. If the series were
continuous, we could state the mode to be around 15 without performing any more
calculations. The following formula is used to identify the mode for grouped data:
f1 − f0
𝑀𝑜𝑑𝑒 = l1 𝑋(𝑖)
(f1 − f0) + (f1 − f2)
Where, l1 = the lower value of the class in which the mode lies
51
f1 = the frequency of the class in which the mode lies
The class intervals should be consistent throughout when using the aforementioned formula.
On the grounds that the frequencies are dispersed equally over the class, the class intervals
should be made uniform if they are not already. The application of the aforementioned
formula in the case of unequal class intervals will produce false results.
Dispersion-
The numerous central value measures provide us with a single number that sums up all the
data. But until all of the observations are identical, the average by itself cannot effectively
described. Even if the central value in two or more distributions may be the same, the way the
distributions are formed can vary greatly. We may analyse this crucial aspect of distribution
2. "The degree to which numerical data tend to spread about an average value is called the
3. Dispersion or spread is the degree of the scatter or variation of the variable about a
52
4. "The measurement of the scatterness of the mass of figures in a series about an average is
Measure of Dispersion
the Lorenz Curve are the five measurements of dispersion. The first four of them are
1.3.5 Range
The range, or difference between a data set's maximum value and minimum value, is the
R=H-L
R = range
H = highest value
L = lowest value
The range is the easiest measure of variability to calculate. To find the range, follow these
steps:
53
2. Subtract the lowest value from the highest value.
This process is the same regardless of whether your values are positive or negative, or whole
numbers or fractions.
Participant 1 2 3 4 5 6 7 8
Age 37 19 31 29 21 26 33 36
First, order the values from low to high to identify the lowest value (L) and the highest value
(H).
Age 19 21 26 29 31 33 36 37
R=H–L
R = 37 – 19 = 18
The quartile deviation or interquartile range is a more accurate indicator of variation within a
distribution than the range. Here, the center 50% of the distribution is used in order to avoid
using the 25% at either end of the distribution. The difference between the third and the first
54
Symbolically, interquartile range = Q3- Q1
Many times the interquartile range is reduced in the form of semi-interquartile range or
When the quartile deviation is low, the items that make up the middle 50% of the distribution
have a low variance. If the quartile deviation is high, on the other hand, it means that the
items in the middle 50% of the distribution have a wide range. The two quartiles, Q3 and QI,
are equally spaced from the median in a symmetrical distribution; it should be observed.
Symbolically,
M-Q1 = Q3-M
The majority of business and economic statistics are asymmetrical; thus, this is rarely the
case. However, it is reasonable to suppose that the interquartile range contains about 50% of
the observations. It should be highlighted that the quartile deviation or the interquartile range
is an exact indicator of dispersion. It can be transformed into the following relative measure
of dispersion:
Q3 − Q1
𝐶𝑜𝑒𝑓𝑓𝑖𝑐𝑖𝑒𝑛𝑡 𝑜𝑓 𝑄𝐷 =
𝑄3 + 𝑄1
The computation of a quartile deviation is very simple, involving the computation of upper
Example 1- Take the following data set into consideration: 22, 12, 14, 7, 18, 16, 11, 15,
Solution-
55
First, we need to arrange data in ascending order to find Q3 and Q1 and avoid any duplicates.
Q1 = ¼ (9 + 1)
=¼ (10)
Q1=2.5 Term
Q3=¾ (9 + 1)
=¾ (10)
The average for Q1 is 2nd, which is 11, plus the difference between 3rd and 4th,
Q3 is the seventh term and product of 0.5; the eighth term's difference from the
Q.D. = Q3 – Q1 / 2
=5.5/2
Q.D.=2.75.
56
Example 2- The textile maker Harry Ltd. is developing a compensation plan. The
management is debating the launch of a new venture, but they want to first determine
The management has compiled its 10 most recent days' worth of average daily
155, 169, 188, 150, 177, 145, 140, 190, 175, 156.
Solution- Here, there are 10 observations, therefore our initial step would be to arrange the
140, 145, 150, 155, 156, 169, 175, 177, 188, 190
=¼ (10+1)
=¼ (11)
=¾ (11)
57
The second term is 145, and by adding 0.75 * (150 - 145) which is 3.75, we get
148.75.
The eighth term is 177, thus by adding 0.25 * (188 - 177), which is 2.75, the answer is
179.75.
Q.D. = Q3 – Q1 / 2
=31/2
Q.D.=15.50.
58
Use the Quartile Deviation formula to find out the dispersion in % marks.
Solution- In this case, there are 25 observations, therefore, our first step would be to arrange
59
Calculation of Q1 can be done as follows,
=¼ (25+1)
=¼ (26)
60
Q3=¾ (n+1)th term
=¾ (26)
Q3 = 19.50 Term
range:
The outcome of adding 0.50 * (156 - 154) which is 1 to the sixth term, which is 154,
is 155.00.
The outcome of adding 0.50 * (177 - 177) to the 19th term, which is 177, is 177.
Q.D. = Q3 – Q1 / 2
=22/2
Q.D.= 11.
Semi-interquartile range, also known as quartile deviation. Again, the interquartile range
represents the difference in variance between the third and first quartiles. The interquartile
range shows how far apart from the mean or average the observations or values in the
provided dataset are. When attempting to understand or conduct a study on the dispersion of
the observations or samples from the given data sets that are found in the main or middle
body of the given series, the quartile deviation or semi-interquartile range is frequently used.
61
This situation would typically occur in a distribution where the data or the observations have
a tendency to lie heavily in the main body or middle of the given set of data, or the series, and
the distribution or the values do not lie towards the extremes, or if they do, they are not of
The average deviation is another name for the mean deviation. The absolute amounts by
which the various items depart from the mean are averaged, as the name suggests. We ignore
positive and negative signs when computing the mean deviation since positive and negative
𝛴|𝑥|
𝑀𝐷 =
𝑛
Where MD = mean deviation, |x| = deviation of an item from the mean ignoring positive and
Example-
Advantages of Mean
Deviation-
62
1. The fact that mean deviation is based on all observations gives it an advantage as a
measure of dispersion. Contrast this with conventional metrics of dispersion like range and
2. The calculation is easy. This is so because the calculation only requires basic, simple
procedures like adding the absolute difference and average, then dividing the result by the
3. The process of averaging the differences eliminates all outliers and paints a clear picture of
4. It measures dispersion more accurately than the standard deviation. This is so because the
standard deviation squares deviations rather than measuring their actual values. Additionally,
this increases the likelihood that the occurrence of extreme values will have an impact on
standard deviation.
using it. When comparing two sets of data values, such relative metrics of dispersion are
useful.
1. When calculated about the median, mean deviation has the lowest value. As a result, the
mean deviation from the median is never greater than the mean deviation from the mean or
2. In the case of a symmetric distribution, the mean departure from the median generally
3. The change in scale has no impact on mean deviation. This indicates that the value of the
mean deviation remains the same if we add a single fixed value to all the observations.
63
4. On the other hand, a change in scale has an impact on mean deviation. The mean deviation
is multiplied by the same positive amount when we multiply each value by that number.
1. Because of its precision and ease of usage, economists frequently employ it.
2. When determining how much wealth is distributed within a society, it is helpful. This is so
that all values, including those of exceedingly wealthy or extremely impoverished people, are
3. Since it is the most accurate measure of variability for this usage, it is used to forecast
business cycles.
2. It might occasionally produce inaccurate results. When deviations are taken from the
median rather than the mean, the mean deviation produces the best results. However, median
3. The method is incorrect just mathematically since it ignores the algebraic signs when
The standard deviation and mean deviation are similar in that they both quantify deviations
from the mean. However, due to its advantageous mathematical characteristics, the standard
64
deviation is favoured over the mean deviation, quartile deviation, and range. We introduce
Example-
Solution-
108
𝑀𝑒𝑎𝑛 = = 18
6
The second column shows the deviations from the mean. The third or the last column shows
the squared deviations, the sum of which is 70. The arithmetic mean of the squared deviations
is:
(𝑥 − 𝜇)2 70
∑ = = 11.67 𝑎𝑝𝑝𝑟𝑜𝑥
𝑁 6
The variance is the average of the squared deviations. It should be noted that many words that
are used interchangeably to express this variance include "the variance of the distribution X,"
"the variance of X," "the variance of the distribution," and "only the variance."
Symbolically,
(𝑥 − 𝜇)2
𝑉𝑎𝑟 𝑋 = ∑
𝑁
It is also written as
65
2
(𝑥𝑖 − 𝜇)2
𝜎 =∑
𝑁
Where 𝜎 2 (called sigma squared) is used to denote the variance. Although the variance is a
measure of dispersion, the unit of its measurement is (points). If a distribution relates to the
income of families, then the variance is (𝑅𝑠)2 and not rupees. Similarly, if another
distribution pertains to marks of students, then the unit of variance is (𝑚𝑎𝑟𝑘𝑠)2 . To address
produced by taking the square root of variance. Using the example of a single observation
In applied Statistics, the standard deviation is more frequently used than the variance. This
2 (∑𝑥𝑖 )2
∑𝑥
√ 𝑖 −
𝜎= 𝑁
𝑁
We use this formula to calculate the standard deviation from the individual observations
given earlier.
66
In a symmetrical, bell-shaped curve:
(i) About 68 percent of the values in the population fall within + 1 standard deviation from
the mean.
(ii) About 95 percent of the values will fall within +2 standard deviations from the mean.
(iii) About 99 percent of the values will fall within + 3 standard deviations from the mean.
Given that it represents the variance in the same units as the original data, the standard
or more distributions. We ought to employ a relative measure of dispersion for this objective.
The coefficient of variation, which connects the standard deviation and the mean such that
the standard deviation is reported as a percentage of the mean, is one such indicator of
relative dispersion. As a result, the measurement unit for the standard deviation is no longer
𝜎
Symbolically, 𝐶𝑉(𝐶𝑜𝑒𝑓𝑓𝑖𝑐𝑖𝑒𝑛𝑡 𝑜𝑓 𝑉𝑎𝑟𝑖𝑎𝑡𝑖𝑜𝑛) = 𝜇 ∗ 100
1. The most used measure of dispersion, it can be used in a wide range of circumstances.
2. All observations are taken into account when calculating the standard deviation. Other
metrics of dispersion, such as range, on the other hand, are not dependent on all observations.
3. A fair estimate of the population standard deviation is provided by the sample standard
deviation. Since the standard deviation is unaffected by sample fluctuations, it can be used to
4. The standard deviation is used to compute the data's skewness and kurtosis, which informs
67
5. Using the formula for the combined standard deviation, we may determine the standard
deviation of two data sets if we are given their individual standard deviations. Such formulas
1. The square of the differences of the observations from the mean, rather than the actual
distance of each observation from the mean, is used to calculate the standard deviation.
2. Outliers will add a significant amount to the numerator when the differences are squared
because doing so makes large values even larger. This indicates that the standard deviation
gives extreme values more weight. The standard deviation is therefore susceptible to the
impact of outliers.
3. Since the method requires extracting square roots, performing it by hand would be highly
One form of measure of dispersion is the coefficient of variation. A number called a measure
of dispersion is used to assess the degree of data variability. As a result, the coefficient of
variation is used to assess how much data deviate from the mean or average value. The
68
For calculating the coefficient of variation, there are two formulas. The sample coefficient of
variance and the population coefficient of variation are these. In statistics, the term
"population" refers to the entire group being studied. In other terms, the population refers to
the entire set of data. The sample is the particular portion of the population that has been
selected. The sample is utilised to reflect the study's overall population. There is no
difference between the sample mean and the population mean. There are two coefficient of
variation calculations, though, because the standard deviation values vary. Here are some of
them:
𝜎
Population Coefficient of Variation = 𝜇 ∗ 100.
𝑠
Sample Coefficent of Variation = 𝜇 ∗ 100
∑(𝑥𝑖 −𝜇)2
σ is the standard deviation of the population. It is given by 𝜎 = √ 𝑁
∑(𝑥𝑖 −𝜇)2
s is the standard deviation of the sample. It is given by 𝜎 = √ 𝑁−1
Example: Two plants C and D of a factory show the following results about the number of
69
Average monthly wages $2500 $2500
Standard deviation 9 10
Using coefficient of variation formulas, find in which plant, C or D is there greater variability
in individual wages.
Solution-
To Find: Which plant has greater variability? For this, we need to find the coefficient of
variation. The plant that has a higher coefficient of variation will have greater variability.
CV = (9/2500) × 100
CV = 0.36%
CV = (σ/μ) × 100
CV = (10/2500) × 100
CV = 0.4%
70
It is readily evident that this "explained variation"
We should be aware that if there is a significant correlation between our x and y, the stronger
that correlation is, the more we can explain the "spread" of our y values by merely admitting
that they are near the best-fit line, which, due to the significant correlation, has a non-zero
As a result, we can query what percentage of the total variation is still unknown and what
To find a nice, tight expression for the "unexplained variation", consider the following:
Remembering
71
The first term on the right is referred to as the unexplained variation since it is obvious that
the total variation should be the sum of the explained variation and the unexplained variance.
Seasonal Variations can be measured by the method of simple average. The data ought to be
Seasonal Variations can be measured by the method of simple average. The data should be
72
Method of Simple Averages:
This is the simplest and easiest method for studying Seasonal Variations. The procedure of
Procedure:
(i) Arrange the data by months, quarters or years according to the data given.
(iv) Find the average of averages, and it is called Grand Average (G)
(v) Compute Seasonal Index for every season (i.e) months, quarters or year is given by
73
1.3.10 Skewness and Kurtosis
Skewness
Depending on the model, skewness might lower the interpretation of feature relevance or
break model assumptions if the values of a particular independent variable (feature) are
skewed.
Skewness, which differs from the symmetrical normal distribution (bell curve), is a measure
To determine skewness, one can use the normal distribution. Data are symmetrically
distributed when we discuss the normal distribution. Due to the fact that all measurements
with a central tendency fall in the middle, the symmetrical distribution has no skewness.
74
The left and right sides of data that are symmetrically distributed have an equal number of
observations. (For example, if the dataset contains 90 values, the left side will have 45
observations, and the right side will contain 45 observations.) But what if the distribution is
not symmetrical? This type of data is referred to as asymmetrical data, and time skewness
Types of skewness
contrast to symmetrically distributed data, where all measures of the central tendency (mean,
median, and mode) are equal to each other, with positively skewed data, the measures are
dispersing. Positively skewed distributions are thus types of distributions where the mean,
median, and mode of the distribution are positive rather than negative or zero.
Positively skewed statistics have a mean that is higher than their median (a large number of
data-pushed on the right-hand side). In other words, the outcomes are skewed to the negative.
Since the median is the middle value and the mean is always the highest value, the mean will
75
Extremely positive skewness is undesirable for distribution since it might lead to inaccurate
findings when present in high concentrations. The skewed data is being brought closer to a
normal distribution with the aid of data transformation technologies. The most well-known
transformation for favourably skewed distributions is the log transformation. The natural
data are plotted on the graph's right side while the distribution's tail spreads out to the left.
When data is negatively skewed, the mean is lower than the median (a large number of data-
distribution's mean, median, and mode are negative rather than positive or zero.
The Median is the middle value, and mode is the highest value, and due to an unbalanced
76
Subtract a mode from a mean, then divides the difference by standard deviation.
The value is scaled down to a narrow range of -1 to +1 when we divide the covariance values
by the standard deviation because Pearson's correlation coefficient ranges from -1 (perfectly
denoting no linear relationship. That describes the correlation values' range quite accurately.
If the data show a high mode, Pearson's initial skewness coefficient is helpful. However,
Pearson's first coefficient is not favoured if the data contain a low mode or multiple modes; in
this case, Pearson's second coefficient may be preferable because it does not depend on the
mode.
The data are significantly skewed if the skewness is between -1 and -0.5 (negative skewed)
77
The data are considered to be highly skewed if the skewness is less than -1 (negative
The data's quartiles serve as the foundation for Bowley's coefficient of skewness. It is based
on the data set's middle 50% of observations. This indicates that each tail of the data set
In the case of a symmetric distribution, the Q1 and Q3 quartiles are equally spaced from the
78
Step 2 - Enter the Range or classes (X) seperated by comma (,)
Deciles or percentiles of the data serve as the foundation for Kelly's coefficient of skewness.
The middle 50% of the data set's observations form the basis of the Bowley's coefficient of
skewness. This indicates that the 25 percent of observations in each tail of the data set are left
Kelly proposed a skewness index based on the middle 80% of the data set's observations.
For a symmetric distribution, the first decile namely D1 and ninth decile D9 are equidistant
79
Interpretations-
Kurtosis
80
Kurtosis can be used to detect outliers in our data. It provides the overall level of outliers that
are present.
The peak and tail of the data can be heavy-tailed and flat, resembling punching or squashing
the distribution. The term for this is negative kurtosis (Platykurtic). Positive Kurtosis is a
term used to describe a distribution that has a light tail and a steeper top curve (Leptokurtic).
kurtosis is indicated by a kurtosis value larger than three. The range of kurtosis value in this
kurtosis fewer than three. The range of values for a negative kurtosis is from -2 to infinity.
Excess Kurtosis
In probability and statistics, the excess kurtosis is used to compare the kurtosis coefficient to
the normal distribution. Leptokurtic distributions have positive excess kurtosis, platykurtic
distributions have negative excess kurtosis, and zero excess kurtosis (Mesokurtic
81
Leptokurtic (kurtosis > 3)
Leptokurtic has extremely long and thin tails, which increases the likelihood of outliers.
Positive values of kurtosis suggest a peaked distribution with fat tails. A distribution where
more of the numbers are distributed away from the mean and in the tails is said to have an
The majority of the data points are present and close to the mean because the platykurtic
distribution has a lower tail and elongated tails around the centre. Comparing a platykurtic
Mesokurtic (kurtosis = 3)
Mesokurtic is the same as the normal distribution, which means kurtosis is near to 0. In
Mesokurtic, distributions are moderate in breadth, and curves are a medium peaked height.
82
Examples 1- Suppose we have the following observations:
{12 13 54 56 25}
Solution
First, we must determine the sample mean and the sample standard deviation:
83
Skewness is positive. Hence, the data has a positively skewed distribution.
Example 2- Using the data from the example above (12 13 54 56 25), determine the
Solution-
84
The excess kurtosis is then obtained by taking 3 out of the sample kurtosis.
Moments
When choosing the probability distribution we will use, statistical moments are essential
since they allow us to characterise the characteristics of the statistical distribution. They are
Moments are popularly used to describe the characteristic of a distribution. Let’s say the
random variable of our interest is X then, moments are defined as the X’s expected values.
85
For Example, E(X), E(X²), E(X³), E(X⁴),…, etc.
Skewness
Kurtosis
Raw Moments
The expected value of Xn is the raw moment, or the n-th moment, of a probability density
Centered Moments
The predicted value of a given integer power of the random variable's deviation from the
mean is what is known as the central moment, which is a moment of a probability distribution
Standardized Moments
often normalised by dividing the standard deviation, making the moment scale-invariant.
A sample is the subset of the population. The process of selecting a sample is known as
86
Sampling is a strategy for choosing specific individuals or a subset of the population in order
to draw conclusions from them statistically and estimate the characteristics of the entire
they do not have to study the full community in order to gather useful information.
It serves as the foundation of any research design because it is also a time- and money-
efficient strategy. For the best derivation, sampling techniques can be utilised in research
survey software.
For example, if a drug manufacturer would like to research the adverse side effects of a drug
on the country’s population, it is almost impossible to conduct a research study that involves
everyone. In this case, the researcher decides a sample of people from each demographic and
then researches them, giving him/her indicative feedback on the drug’s behavior.
There are two forms of sampling used in market action research: probability sampling and
non-probability sampling. Let's examine these two sampling techniques in more detail.
selects a few criteria and randomly selects individuals of a population. With the use of this
selection parameter, each member has an equal chance of being included in the sample.
87
Simple Random Sampling
Stratified sampling
Systematic sampling
Cluster Sampling
for the study. This type of sampling is not a set or predetermined selection procedure. Due of
this, it is challenging to ensure that every component of a population has an equal chance of
Convenience Sampling
Purposive Sampling
Quota Sampling
This essay is an overview of research that Jared Wadley conducted on October 1st, 2008. The
study was conducted to look into the consequences of gun exhibitions, which were on the rise
88
not only in the USA but also internationally. In order to conduct the research, demographic
samples that were supposed to represent different regions of America were sampled.
These inquiries, in my opinion, were used by the researcher to direct him appropriately on the
type of sampling process, data collection, data interpretation, data analysis, and data
In order to conduct the research, Wadley combined both dependent and independent factors.
The variables that can be measured, have an impact on the research, and depend on the
independent variables are referred to in this context as the independent variables. The number
of deaths, suicides, and homicides were some of the dependent variables employed in this
study. He was able to deliver more accurate results thanks to this understanding. Dependent
variables, on the other hand, are variables that do not depend on one another and are
unaffected by study.
No matter what experimental conditions are applied, they don't change. The number of gun
exhibitions held during this time period is thus one of the independent variables in our study
(Maxim, L.W. 2010). Even if the number of deaths was increasing as a result of viewers'
exposure to these shows, the number of shows was unaffected by this unfavourable trend
since viewers never linked these deaths to these programmes, especially prior to the
The larger USA is where this study was conducted. The researcher had to develop a better
method of identifying and choosing the population to utilise as the participants of this
research because the USA is such a large and populous country. The sampling method was
89
acceptable because it was difficult to involve the full American population in the
investigation. These sample populations were chosen to represent the entire American
population.
Simple random, in which the population was deliberately chosen using random number
tables, and systematic sample, in which the population was chosen based on numerous
characteristics deemed pertinent to this research, were some of the sampling techniques used.
For instance, the two most populous states in the USA, Texas and California, provided the
greatest number of samples. In addition, the researcher used a stratified quota sampling
approach, in which larger populations were separated into smaller ones and later given
preferential treatment.
The population was split into men, women, children, and the elderly, according to this
statement. This is as a result of how each group interpreted the gun exhibitions. Gun
exhibitions are more popular among young people, while older people find them to be
terribly damaging and dishonourable. This indicates that more youth were expected to
participate in this study than in any other demographic segment of the American population.
However, many biases were found during the course of this research in the population that
was chosen, how the data were measured, and how the study's subject was handled.
As a result, when the researcher disregarded the sampling's general rules, selection biases
became evident.
There were underrepresented areas and overrepresented ones. Bias in the measurements
would result from this. To gather, analyse, interpret, and present the data for this study, a
variety of methods and approaches were used. These included surveys, interviews, in-person
90
Many ethical decisions had to be taken when doing this investigation. The study's participants
were not to be forced to provide information or see such "traumatising" programmes. Their
According to this study, there is no connection between watching these episodes and an
increase in suicide and homicide rates. Additionally, it was discovered that watching such
programmes does not necessarily change how youngsters behave because such movies
frequently include warnings like "do not try this at home!" Finally, it was discovered that the
strict restrictions' enforcement does not necessarily result in a sharp decline in the number of
gun exhibitions.
Finally, it may be assumed that this research was properly conducted precisely at the precise
moment that this kind of information needed to be made public. Even if several obstacles had
Glossary
statistics, which are more predictive in nature, descriptive statistics strive to summarise.
Sample- An excerpt taken from a bigger population is known as a sample. The outcome is
referred to be a random sample if the drawing is carried out in a way that gives every
91
Parameter- An value that is derived from a population is referred to as a parameter. This
figure would be a parameter if I had access to all of the data for all people on Earth and
Statistics- A statistic is a number produced from a sample. This value would be a statistic
if I determined the average age of a sample of humanity living on Earth (far more
from a sample and generalise them to the characteristics of the population as a whole.
This ability is not a given and greatly depends on the type of sample used, the size of the
Distribution- The arranging of data by a single variable's values in ascending order, from
low to high, is known as a distribution. This configuration, as well as its features like
Mean- One of the three main measures of central tendency, along with median and mode,
that jointly assess a crucial and fundamental element of distribution is the mean. The
mean is perhaps the statistic that researchers generally use the most. The sample mean is
Median- The median is the score in a distribution that lies between the top and bottom
50% of scores and is located at the 50% percentile. The median can divide a set of
Mode- The score that appears in the distribution the most frequently is the mode. A
distribution with more than one mode is said to be multimodal, and one with two modes
is referred to as bimodal.
92
Skew- Skew occurs when there are more scores at one end of the distribution than the
other. The situation is known as negative skew when a distribution's scores are more
concentrated at the high end, and a tail is produced by the relatively smaller number of
low-end values. When a distribution has a tail at the high end, it has positive skew.
In general, we would anticipate that the mean in a negatively skewed distribution would
be lower than the median and that the mean in a positively skewed distribution would be
Range- The range, one of the most crucial measurements of dispersion, is the distinction
variance. Despite not being used frequently on its own, variance can be a helpful
calculation when heading toward a more descriptive statistical measurement like standard
deviation.
Standard Deviation- The average difference between each distribution score and its mean
is what is known as the distribution's standard deviation. The standard deviation gives a
These 2 measurements, coupled with the mean, give a clear picture of how the scores are
distributed.
Interquartile Range-The IQR is the difference between the score defining the third and
Key Terms- Arithmetic Mean, Geometric Mean, Mean, Median, Mode, Statistics, Kurtosis,
93
Summary- This unit deals with descriptive and inferential statistical methods like mean,
median, mode, standard deviation, skewness, various methods of skewness, and kurtosis,
Important Formula-
94
Range=Maximumvalue–Minimumvalue
95
Exercise: Choose the correct option
B. Normal
D. Symmetrical
A. AM
B. AM > Median
C. AM > Mode
A. 0
B. 4
C. 8
D. 6
4. The degree to which numerical data tend to spread out about an average value is
called
A. Variation
96
B. Skewness
C. Flatness
D. Constant
5. When a distribution is symmetrical and has one mode, the highest point on the curve
is called the
B. Median
C. Mean
D. Mode
A. Platykurtic
B. Mesokurtic
C. Positively skewed
D. Symmetrical
A. 15
B. 20
C. 25
D. 5
A. Dispersion
97
B. Skewness
C. Symmetry
D. Kurtosis
A. Symetrical
B. Skewed
C. Both A and B
D. None of these
10. In Uni model distribution if mode is less than mean, then skewness will be
A. Symmetrical
B. Normal
C. Positively Skewed
D. Negatively Skewed
1. What is skewness?
3. What is kurtosis?
98
3. Explain Standard deviation and Kurtosis in detail?
Answers
Exercise: MCQ
1. (D)
2. (D)
3. (C)
4. (B)
5. (D)
6. (C)
7. (B)
8. (D)
9. (D)
10. (C)
99
Module 2
Learning Outcomes
Introduction
For the comparison and analysis of distributions containing only one variable, or univariate
kurtosis are useful. However, another crucial component of statistics is expressing the
Understanding the correlations between two or more variables is essential for decision-
making in many business research scenarios. For instance, knowing if the interest rate on
bonds is linked to the prime interest rate might assist a broker forecast how the bond market
would behave. An account executive may find it useful to know whether there is a significant
correlation between advertising spending and sales spending for a company while researching
The statistical techniques of correlation and regression aid in understanding the relationship
between two or more variables that may be connected in a similar manner, such as the bond
interest rate and prime interest rate, advertising expenditure and sales, income and
consumption, crop yield and fertiliser use, height and weight, and so forth.
In all these cases involving two or more variables, we may be interested in seeing:
100
if there is any association between the variables;
if so, what form the relationship between the two variables takes;
how we can make use of that relationship for predictive purposes, that is, forecasting;
and
Correlation analysis and regression analysis are two approaches to examining the relationship
between two or more variables because these issues are interconnected. Let's say there is a
correlation between two or more variables. In this situation, regression analysis can be used
to estimate the inaccuracy of estimations and predict the value of the other variable(s) using
What is Correlation?
one variable tends to be accompanied by matching signals in the other(s), then two or more
variables are said to be highly linked. It has a range of -1 to +1. It aids in comprehending the
“The nature and strength of the link between the variables are measured by the correlation
between them.”
When the goal of exploratory research is to find variables that might be related in some
manner to the variable of interest, correlation is frequently utilised as a measure of the degree
101
The notion of variables and the distinction between dependent and independent variables
make it evident that variables may be related to one another. For instance, demand and supply
are tied to commodity price, and agricultural output is influenced by rainfall, student grades
A measure of the type of link between two or more variables is called correlation. It
fluctuates from -1 to +1. It aids in comprehending the degree and direction of the link
In the words of Croxton and Cowden, “When the relationship is of a quantitative nature,
the appropriate statistical tool for discovering and measuring the relationship and
Correlation measures the strength of the relationship between two or more variables. For
example, the relationship between income and consumption expenditure, price and
When the relationship between variables is known, it is easy to predict the value of one
It helps understand the behaviour of various economic variables like demand, supply,
Correlation can be classified in several ways. The important ways of classifying correlation
are:
102
(ii) Linear and non-linear (curvilinear) and
Positive and Negative Correlation- Positive correlation is defined as the movement of both
variables in the same direction, i.e., when one variable rises, the other variable rises on
average as well, or when one variable falls, the other variable falls on average. Conversely, if
the variables are moving in the opposite way, we say that there is a negative correlation, as in
Linear and Non-linear (Curvilinear)- Correlation A case of linear correlation occurs when
changes in one variable are accompanied by similar changes in the other variable in a fixed
X: 10 20 30 40 50
Y: 25 50 75 100 125
In the aforementioned case, the ratio of change is the same. Therefore, it is an instance of
linear correlation. All of the points will lie on a single straight line if we plot these variables
on graph paper. On the other hand, non-linear or curvilinear correlation occurs when the
amount of change in one variable does not follow a constant ratio with the change in another
variable. A non-linear connection would result from changing a few numbers in series X or
series Y.
Simple, Partial, and Multiple Correlation- The number of variables included in a study
determines how these three types of correlation are distinguished from one another. A
correlation is referred to as a simple correlation if there are just two variables present in the
are present in a study. Three or more variables are analysed simultaneously in multiple
103
correlations. However, with partial correlation, we just take into account the interaction of
two variables while holding the impact of the remaining variable(s) constant.
Consider a situation where there are three variables: X, Y, and Z. X represents the number of
hours studied, Y represents intelligence, and Z represents the total number of exam points
earned. We will investigate the relationships between the marks attained (Z) and the two
variables, the quantity of study time (X) and I.Q. (Y). In contrast, a study utilising partial
correlation is one in which the link between X and Z is examined while maintaining a
The commonly used methods for studying linear relationships between two variables involve
both graphic and algebraic methods. Some of the widely used methods include:
1. Scatter Diagram
2. Correlation Graph
Scatter Diagram
The Dotogram or Dot Diagram is another name for this approach. One of the simplest ways
approach, dots are used to plot both variables on graph paper. "Scatter Diagram" is the name
given to the resulting diagram. By looking at the diagram, we may get a general notion of the
type and strength of the link between the two variables. The spreading of dots across the
104
graph is referred to as scatter. When analysing correlation, the following things should be
kept in mind:
If the plotted points are very close to each other, it indicates high degree of correlation. If
the plotted points are away from each other, it indicates low degree of correlation.
If the points on the diagram reveal any trend (either upward or downward), the variables
are said to be correlated, and if no trend is revealed, the variables are uncorrelated.
The correlation is positive if there is an upward movement from the lower left hand
corner to the higher right-hand corner, showing that the values of the two variables move
in the same direction. In contrast, if the points show a downward trend from the upper left
to the lower right, the correlation is negative since the values of the two variables in this
in particular, if all the points lie on a straight line starting from the left bottom and going
up towards the right top, the correlation is perfect and positive, and if all the points like
105
on a straight line starting from the left top and coming down to the right bottom, the
Example-
Given the following data on sales (in thousand units) and expenses (in thousand rupees) of a
Months: J F M A M J J A S O
Sales: 50 50 55 60 62 65 68 60 60 50
Expenses:11 13 14 16 16 15 15 14 13 13
b) Do you think that there is a correlation between sales and expenses of the firm? Is it
Solution-
(b) Figure shows that the plotted points are close to each other and reveal an upward trend. So
there is a high degree of positive correlation between sales and expenses of the firm.
106
Graphical Method
This Correlogram approach is quite straightforward. Two series' worth of data are plotted on
a graph sheet. By comparing the direction and proximity of two curves, we may determine
the correlation. A positive correlation is present when both of the curves on the graph are
going in the same direction. On the other hand, correlation is considered to be negative if
both curves are travelling in the opposite direction. A lack of connection can be seen in a
Example- Find out graphically, if there is any correlation between price yield per plot (qtls);
Plot no.: 1 2 3 4 5 6 7 8 9 10
X: 3.5 4.3 5.2 5.8 6.4 7.3 7.2 7.5 7.8 8.3
Y: 6 8 9 12 10 15 17 20 18 24
107
The above graph demonstrates that the two curves move in the same direction and are also
quite near to one another, indicating that the price yield per plot (qtls) and amount of fertiliser
Remark: By visualising the relationship between the variables, the scatter diagram and
correlation graph, two graphic tools, provide the reader a sense of the data. These are easily
understood and help us develop a reasonable, if rough, understanding of the type and strength
of the link between the two variables. These techniques, however, are unable to measure the
connection between them. With the use of algebraic techniques, which compute the
The coefficient of correlation rxy between two variables x and y, for the bivariate dataset (xi,yi)
r(x,y)=cov(x,y)/σxσy
where,
Here, 𝑥̅ and 𝑦̅ are simply the respective means of the distributions of x and y.
Alternate Formula
108
You can apply the following formulas if any data is presented as a class-distributed frequency
distribution:
where,
109
2.1.5 Properties of the Pearson’s Correlation Coefficient
different bivariate distributions. For instance, you can examine the Pearson's correlation
coefficients from both scenarios to see how much of your decision to skip the movie is
connected to your friends not joining you and to your own lack of interest in the film. This
distinct numbers in economics, where the cost price or the market shares depend on a variety of
different elements.
⇒ The value of r always lies between +1 and -1. Depending on its exact value, we see the
r value variation:
110
A number higher than 0 denotes a positive connection, meaning that as one variable's value
rises, the value of the other variable also rises. If the value of one variable is greater than 0, the
⇒ The Pearson product-moment correlation does not take into consideration whether a variable
has been classified as a dependent or independent variable. It treats all variables equally.
⇒ A change of origin of the system, or any scaling of the variables doesn’t affect the value
A nonparametric measurement of the strength and direction of the link between two ranking
monotonicity of a connection, i.e., whether the link between two continuous or ordered variables
does not actually require monotonicity, but if we already know there isn't a monotonic
relationship between the two variables, it makes no sense to use Spearman's correlation to find
On the other hand, one would use a Pearson's correlation to determine the strength and direction
of any linear relationship if, for instance, the relationship appears linear (as shown by a
scatterplot). Monotony –
111
Spearman Ranking of the Data
Before performing the Spearman's Rank Correlation analysis, we must rank the data under
consideration. This is essential because we must compare if, when one variable is increased, the
other follows a monotonic relation (regularly increases or decreases) with respect to it.
We must therefore compare the values of the two variables at each level. Such "levels" are
assigned to each value in the dataset by the ranking algorithm so that we can quickly compare
them.
Assign number 1 to n (the number of data points) corresponding to the variable values in
In the case of two or more values being identical, assign to them the arithmetic mean of the
Examples of values for the selling price are: 28.2, 32.8, 19.4, 22.5, 20.0, and 22.5 These are the
matching ranks: 2, 1, 5, 3.5, 4, 3.5 The highest value (32.8) is ranked first, and 28.2 is ranked
second. When two numbers (22.5) are similar, it is necessary to calculate the arithmetic mean of
112
Spearman Rank Correlation formula-
𝟔𝜮𝒊 ⅆ𝟐𝒊
𝒓𝑹 = 𝟏 −
𝒏(𝒏𝟐 − 𝟏)
where n is the number of data points of the two variables and di is the difference in the ranks of
the ith element of each random variable considered. The Spearman correlation coefficient, ρ, can
The closer ρ is to zero, the weaker the association between the ranks.
Correlation and partial correlation are related ideas. It demonstrates the concept that just
because two variables show a correlation doesn't mean that they are causally related. When
two variables are conditional on one or more other factors, partial correlation measures the
correlation between the two variables. This suggests that when there is a connection between
two variables, the confounder (or controlling variable), a common cause of the misleading
association, may contribute to the explanation of the correlation. This portion is removed,
To gauge the strength of the association between two variables while taking into
consideration the influence of one or more additional factors, partial correlation is used. Your
study's key variables should be continuous, regularly distributed, logically connected, and
113
devoid of outliers. Your variables should also have a comparable dispersion across each of
The partial correlation formula between random variables X and Y with Z factored out is
given by when there is only one confounding factor (then it is referred to as a first-order
partial correlation).
Assumptions are a part of all statistical methods. Your data must meet specific criteria in
order for the results of a statistical method to be accurate, which is what assumptions mean.
Continuous
Normally Distributed
Linearity
No Outliers
Covariate(s)
Continuous
114
You must care about a continuous variable. Continuous refers to a variable that can have any
conceivable value. Age, height, weight, test scores, survey results, yearly salary, etc. are a
Normally Distributed
The variable that matters to you ought to be dispersed normally. This is referred to as being
regularly distributed in statistics (aka it must look like a bell curve when you graph the data).
Only if the variable you care about is normally distributed should you conduct an
115
Linearity
The factors that matter to you must be correlated linearly. This means that if the variables are
No Outliers
There must be no outliers in the variables that you care about. When it comes to outliers, or
data points with extremely high or low values, Pearson's correlation is sensitive. When you
plot your variables, look for any points that are far from the other points to determine if there
Making ensuring the variables have a similar dispersion across their ranges is known as
homoscedasticity in statistics.
Covariate(s)
If you have one or more covariates, you should only do partial correlation. When analysing
the variable connection of interest, a covariate is a variable whose effects you want to
eliminate. For instance, you might want to take education level into account while evaluating
the link between age and memory function. In this manner, you may be certain that the results
116
Variable 1: Height
Variable 2: Weight
Covariate: Age
In this illustration, we are interested in the correlation between height and weight while
taking age into consideration. As a result, we start by gathering data on a set of people's
First, we make sure the variables that are important to us conform to the partial correlation's
presumptions. We proceed with the study after establishing that height and weight are
normally distributed, devoid of outliers, scattered similarly over their ranges, and linearly
The analysis will yield a p-value and a correlation coefficient, or "r." R values are between -1
and 1. When r is negative, the variables are said to be inversely connected (i.e. when one
Positive numbers, on the other hand, show that as one variable rise, the other rises as well.
When age effects are taken into account, the p-value shows the likelihood that our results
would not have been seen if there was no real association between height and weight. If the
p-value is less than or equal to 0.05, our result is considered statistically significant, and we
may be confident that the observed difference is not the result of random chance.
117
The effect of the independent variables must be additively and not jointly related.
Summary- This unit deals with data sources, data collections, data tables, correlation, types
Case Study:
Certain rocket motors are created by fusing two types of fuel, an igniter and a sustainer.
The goal of this research is to look at the relationship between the strength of this
Empirical evidence implies that this relationship is linear, and the major goal of this
study is to evaluate this claim and create the best model feasible, based on a data set of
118
1 2158.7 15.5
2 1678.15 23.75
3 2316 8
4 2061.3 17
5 2207.5 5
6 1708.3 19
7 1784.7 24
8 2575 2.5
9 2357.9 7.5
10 2277.7 11
11 2165.2 13
12 2399.55 3.75
13 1779.8 25
14 2336.75 9.75
15 1765.3 22
16 2053.5 18
17 2414.4 6
18 2200 12.5
19 2654.2 2
20 1753.7 21.5
119
Procedures for Analyzing:
The first section of the analysis attempts to determine whether or not there is a
significant linear association between the variables Strength and Age. A scatterplot will
be created for this purpose, and the correlation coefficient will be calculated. If the
FOUND RESULTS
120
The data clearly shows a negative linear trend. A linear regression analysis of the data
makes logical.
negative linear relationship between Strength and Age. Also from the table, it can be
121
Answers-
Strength and age were found to have a significant association Based on the
scatterplot and the correlation, there is a clear negative linear link between them.
Strength=2625.355−36.961Age
122
C. The coefficient of correlation is not dependent on both the change of scale and change
of origin
A. It is a bivariate analysis
B. It is a multivariate analysis
C. It is a univariate analysis
D. Both a and c
A. Standard error
123
B. Correlation
C. Regression
7. Which one of the following statements about the correlation coefficient is correct?
B. Both the change of scale and the change of origin have no effect on the correlation
coefficient.
8. Choose the correct option concerning the correlation analysis between 2 sets of data.
10. The correlation for the values of two variables moving in the same direction is
A. Perfect positive
124
B. Negative
C. Positive
D. No correlation.
1. What is correlation?
Answers
Exercise: MCQ
1. (D)
2. (C)
3. (C)
125
4. (D)
5. (D)
6. (C)
7. (A)
8. (A)
9. (C)
10. (D)
In the corporate world, it frequently becomes important to make a forecast in order for
circumstance.
For instance, a business is curious about how much the demand for In the coming five years,
the number of televisions will rise while taking population growth into account in a specific
city. Here, it is assumed categorically that an increase in population will result in a growing
interest in televisions. Consequently, to ascertain the type and degree of association. The
relationship between these two factors becomes crucial for the business.
2.2.1 Introduction
126
A study on heredity titled "Natural Inheritance" was published in 1889 by Sir Francis Galton,
a cousin of Charles Darwin. He shared his discovery that sweet pea plant seed sizes appeared
to "revert" or "regress" to the mean size over time. He also shared the findings of a study on
the correlation between dads' heights and their sons' heights. Height of father versus height of
son data pairs were fitted with a straight line. He discovered a "regression to mediocrity" here
as well. The sons' heights showed a shift away from their fathers and toward the norm in
height.
The concept of statistical regression is attributed to Sir Galton. The name "regression" still
exists even if the majority of applications of regression analysis may have little in common
with Galton's "regression to the mean." Today, it describes the statistical method of
simulating the interaction between two or more variables. Regression analysis, in its broadest
meaning, refers to the estimation or prediction of an unknown value for one variable using
known values for the other variable (s). It is one of the most significant and often applied
statistical approaches in nearly all natural, social, and physical disciplines. We will just
discuss simple regression in this course, which is linear regression with just two variables: a
Multiple regressions are regression analyses that look at more than two variables at once.
Simple regression involves only two variables; one variable is predicted by another variable.
Only two variables are used in simple regression, where one variable predicts the other. The
dependent variable is the one that needs to be predicted. The independent variable, also
referred to as the explanatory variable, is the predictor. For instance, when attempting to
forecast television set demand based on population growth, the demand for television sets is
127
used as the dependent variable and population growth is used as the independent or predictor
variable.
Choosing which variable is which can occasionally lead to issues. Often, the decision is clear-
cut, as in the case of population increase and TV demand, as it would be absurd to assume
that the latter might be influenced by the former. The dependent variable must be TV
Linear Regression
Creating techniques for fitting a straight line, or a regression line as it is frequently known, to
data on two variables is the process of revealing a linear relationship. The best estimate of
one variable for any given value of the other variable is represented graphically or as a
relationship by the line of regression. The independent and dependent factors affect how the
line is referred to. The Regression line of Y on X is a line that provides the best estimate of Y
for any value of X when X and Y are two variables whose relationship has to be shown.
If the dependent variable changes to X, then the best estimate of X by any value of Y is called
Regression line of X on Y.
1. Linear Regression
2. Logistic Regression
3. Polynomial Regression
4. Ridge Regression
5. Lasso Regression
6. Quantile Regression
128
9. Partial Least Squares Regression
1. Linear Regression
The modelling method that is most frequently employed assumes a linear relationship
between an independent variable (V) and a dependent variable (Y) (X). It uses a best-fit line,
commonly referred to as a regression line. The equation for the linear relationship is Y =
c+m*X + e, where c stands for the intercept, m for the slope, and e for the error term.
The simple and complex versions of the linear regression model, respectively, have different
numbers of dependent and independent variables (with one dependent variable and more than
2. Logistic Regression
The logistic regression method is appropriate when the dependent variable is discrete. In
other words, this method is used to determine the likelihood of events that are mutually
129
exclusive, such as pass/fail, true/false, 0/1, and so on. Thus, the probability has a value
between 0 and 1, the target variable has a range of two possible values, and its relationship to
3. Polynomial Regression
polynomial regression analysis is performed. The best fit line is curved instead of straight in
4. Ridge Regression
130
The ridge regression technique is used when the independent variables are highly correlated
and the data shows multicollinearity. Even if least squares estimates are impartial in
multicollinearity, their variances are high enough to induce a difference between the observed
value and the true value. By inflating the regression estimates, ridge regression lowers
standard errors.
The multicollinearity issue in the ridge regression equation is solved by the lambda (λ)
variable.
5. Lasso Regression
The lasso (Least Absolute Shrinkage and Selection Operator) method penalises the absolute
magnitude of the regression coefficient, just like ridge regression does. The lasso regression
method also uses variable selection, which causes the coefficient values to zero off
completely.
131
6. Quantile Regression
A part of the linear regression method is the quantile regression methodology. When the
conditions for linear regression are not met or when there are outliers in the data, it is used.
132
7. Bayesian Linear Regression
The Bayes theorem is utilised in Bayesian linear regression, a type of regression analysis
method used in machine learning to determine the values of the regression coefficients. This
method calculates the posterior distribution of the features rather than the least-squares. As a
result, the method performs better in terms of stability than standard linear regression.
regression data. By biassing the regression estimates, the significant components regression
approach, like ridge regression, lowers standard errors. The training data are first modified
using principal component analysis (PCA), and the changed samples are then utilised to train
the regressors.
133
A quick and effective method for covariance-based regression analysis is partial least squares
multicollinearity between the variables. Regression is used once the procedure reduces the
When working with highly correlated data, elastic net regression combines the ridge and
lasso regression techniques. By leveraging the penalties connected to the ridge and lasso
Degree of Correlation
134
The coefficient of correlation measures the strength of the association between two variables.
1. Perfect correlation: Two variables have a perfect correlation if they vary in the same
proportion (increase or decrease). An ideal correlation in this case could be either a positive
or negative correlation.
Coefficient of correlation (r) = −1: If there is a perfect negative relationship between two
2. Zero correlation: The correlation between two variables is zero if there is no relationship
between them. It suggests that a change in one variable's value has no bearing on the change
Correlation coefficient (r) = 0: The value of correlation will be zero if there is no correlation
between the two variables. It does not necessarily follow that these two factors are
independent, though. It merely shows that the two variables don't have a linear connection.
3. Limited degree of correlation: Between perfect correlation and zero correlation, there is a
limited degree of correlation, meaning that the coefficient of correlation is between +1 and 1.
High degree of correlation: The correlation between two data series is close to one. The
correlation between two sets of data is neither very high nor very low. Low degree of
135
1. Regression coefficient y on x
byx=n⋅∑fdxdy-∑fdx⋅∑fdyn⋅∑fdx2-(∑fdx)2⋅hyhx
2. Regression coefficient x on y
bxy=n⋅∑fdxdy-∑fdx⋅∑fdyn⋅∑fdy2-(∑fdy)2⋅hxhy
3. Regression Line y on x
y-ˉy=byx(x-ˉx)
4. Regression Line x on y
x-ˉx=bxy(y-ˉy)
Regression analysis is mostly used to carry out the financial procedure. Therefore, the
1. Forecasting:
dangers. For instance, demand analysis predicts how many items a buyer is likely to
purchase.
Demand, however, is not the only dependent variable when it comes to business. Much
more than just direct income can be predicted using regression analysis.
By estimating the amount of people who would pass in front of a certain billboard, for
instance, we could forecast the highest bid for an advertisement. Regression analysis is a
key tool used by insurance companies to predict the creditworthiness of policyholders and
the number of claims that might be made during a specific time period.
2. CAPM:
136
The linear regression model is a key component of the Capital Asset Pricing Model
(CAPM), which determines the relationship between an asset's expected return and the
Regression analysis is used to calculate a stock's beta coefficient. Beta is a metric for
We can quickly calculate it in Excel using the SLOPE tool because it reflects the slope of
certain rival.
It can also be used to figure out how the stock prices of two companies are related to one
It might help the company identify the factors affecting their sales in contrast to the
comparable company. These methods can help small businesses succeed quickly and
4. Identifying problems:
For instance, a manager of a retail store might believe that extending the hours of
However, RA can contend that the higher income is insufficient to offset rising operating
costs brought by extended working hours (such as additional employee labour charges).
137
This research may therefore provide quantitative support for decisions and assist
managers in avoiding errors based solely on intuition.
Reliable source
Regression analysis (and other types of statistical analysis) are now being used by many
companies and their top executives to make better business decisions and cut down on
Regression analysis is a tool that managers can use to sift through data and select the
Correlation Regression
In Correlation, both the independent and dependent However, in Regression, both the
different.
The primary objective of Correlation is, to find out a When it comes to regression, its primary
138
association between the values. haphazard variable based on the values
Correlation stipulates the degree to which both of the However, regression specifies the effect
variables can move together. of the change in the unit in the known
value.
Procedures for forecasting can benefit from using regression lines. Its goal is to describe how the
dependent variable (y variable) and one or more independent variables are related (x variable).
By entering various values for the independent variables, an analyst can predict future
behaviours of the dependent variables by using the equation derived from the regression line.
139
Regression Line Formula: y = a + bx + u
The straight line is typically found through linear regression. The least squares regression line
is another name for it. In a bivariate dataset, it is represented. Let's assume that the dependent
Y = a0 + a1x
Where a0 is the constant and a1 is the regression coefficient and x is the value of the
independent variable. If you are given a random sample of observation, the population
140
Y’ = a0 + a1x, where a0 is the constant and b1 is the regression coefficient. Here you will, ‘x’
is the value of the independent variable, and y’ is the predicted value of the dependent
variable.
Regression coefficient values remain the same. Since shifting of origin takes place because
If the variables x and y are changed to u and v, respectively u= (x-a)/p v=(y-c) /q, Here p
If there are two lines of regression. Both of these lines intersect at a specific point [x’, y’].
Variables x and y are taken into consideration. According to the property, the intersection
of both the lines of regression, i.e. y on x and y, is [x’, y’]. This is the solution for both of
You will discover that the geometric mean of the two coefficients represents the
correlation coefficient between the two variables, x and y. Additionally, the common sign
of the two correlation coefficients will be indicated by the sign over the values of the
coefficients. So, if, according to the property, regression coefficients are byx= (b) and bxy=
(b’) then the correlation coefficient is r=+-sqrt (byx + bxy) so, in some cases, both the
coefficients give a negative value, and r is also negative. If both the values of coefficients
The regression constant (a0) is equal to the y-intercept of the regression line. Where a0 and
141
The standard error of estimation is also known as the standard error of the regression (s). It is
typical separation between the observed values and the regression line. The standard error of
regression gives an indication of how closely the observations are fitting the regression line
when it has smaller values. If the standard error is "0," then the correlation is flawless, and
The standard error of estimate measures the difference between the actual values of Y and the
anticipated (calculated) values of Y on the regression line, just as the standard deviation
measures the variation in a set of data from its mean. Both linear regression models and non-
linear regression models can use the standard error of the estimate. It is crucial for the
computation of prediction intervals and confidence intervals. The following formula can be
where,
Similarly,
where,
142
xe = Estimated value of x for a given value of y
The large the value of Syx or Sxy the greater the scatter on the line of regression. In such a
case the degree of correlation series is poor. The error of an estimate is an absolute measure
and is given by the ratio S/σ. This ratio is also used for finding the value of the coefficient of
correlation.
change in the dependent variable's value in relation to the unit change in the independent
variable.
There will be two regression coefficients if there are two regression equations:
change in X for the unit change in Y and is denoted by the symbol bxy. In a symbolic
When the deviations from the real means of X and Y are taken into consideration, the
143
The following formula is applied when deviations from the assumed mean are
obtained:
Regression Coefficient of Y on X: The symbol byx is used that measures the change
In case, the deviations are taken from the actual means; the following formula is used:
The bxy can be calculated by using the following formula when the deviations are taken
The Regression Coefficient is also called a slope coefficient because it determines the slope
of the line, i.e. the change in the independent variable for the unit change in the independent
variable.
144
The link between a predictor variable and the responder is described by regression
coefficients, which are estimations of the unknowable population parameters. Coefficients are
the numbers that multiply the predictor values in linear regression. Let's say you have the
regression formula y = 3X + 5. In this equation, the predictor is X, the constant is +5, and the
coefficient is +3.
The direction of the link between a predictor variable and a responder variable is shown by
A positive sign indicates that as the predictor variable increases, the response variable
also increases.
A negative sign indicates that as the predictor variable increases, the response variable
decreases.
When the predictor is changed by one unit, the coefficient value shows the average change in
the response. If a coefficient is +3, for instance, the mean response value rises by 3 for each
Let's start by examining the equation for linear regression in its general form:
y=B*x+A
Here, the coefficients A and B determine the equation, with x serving as the independent
variable and y serving as the dependent variable. The distinction between the equations for
linear regression and multiple regression is that the multiple regression equation must be able
to handle several inputs, as opposed to just the single input required by the linear regression
equation. The equation for multiple regression uses this change to account for:
145
y = B 1*x 1, B 2*x 2, B 3*x 3,..., B n*x n, and A
For instance, the first independent variable's value is x 1, the second independent variable's
value is x 2, and so on. It continues as we continue to include independent variables until the
Note: any number, n, of independent variables may be used in this multiple regression
The same subscripts are used by the B coefficients, indicating that they are the coefficients
associated with each independent variable. As previously, A is only a constant that indicates
what the dependent variable, y, is worth when all of the independent variables, the xs, are
equal to zero.
Here's an illustration of multiple regression: Imagine that you are responsible for traffic
planning in your city and that you must determine the typical commute time for automobiles
travelling from the east to the west side of the city. Although you are unsure of the typical
time it takes, you are aware that it will vary depending on a number of variables, including
the distance travelled, the quantity of stoplights encountered, and the amount of other
The following two goals are given in a case study scenario where you are the Chief Analytics
146
Objective 1: Improve the campaigns' conversion rates or the percentage of customers who
The first goal was accomplished in the earlier sections of this case study example. Clients’
propensities to respond to campaigns were estimated using the classification models (Parts 5,
Parts 6, Parts 7 & Parts 8). The second goal is now left up to you: calculate the estimated
earnings from each consumer, assuming the customer reacts to the ad. This is a common
regression issue. You will use the data for 4200 consumers, out of 100,000 solicited
customers, who reacted to the prior campaigns, to create a regression model. All of these
4200 clients reside in various communities that can be divided into the three categories listed
below.
Greater Cities
Major Cities
Little Towns
By the way, there are 1400 clients in each of these three groups, which are distributed equally
among the customers. The average profit made by these three types of cities was the first
thing you looked at. The average profit figures for these categories are different, as you can
see in the chart below. Remember these average values; they will be useful when we create
147
The second question is whether or not these average profitability figures differ significantly.
The distributions of all 4200 clients, broken out by location category, provide the answer to
this query. These distributions are depicted in the above figure (towards right). The following
table shows the density distribution for all 4200 of our original data's clients, broken down by
geography category. As you can see, certain situations in this distribution of profits are
1. As a result of their citizens' stronger earning potential and disposable money, large cities
148
2. Due to their larger socio-economic diversity, large metropolitan areas also have a wider
Let's construct a straightforward regression model using these categories as the predictor
variables, keeping the information above in mind. The outcomes of our regression model are
as follows:
Std. t
Coefficients: Estimate Pr(>|t|)
Error value
Intercept 46 0.4691 98.06 <2e-16
Mid Sized
8 0.6635 12.06 <2e-16
Cities
Large Cities 22 0.6635 33.16 <2e-16
Multiple R-
0.2069
squared:
Adjusted R-
0.2065
squared:
F-statistic (P
2.20E-16
Value)
Keep in mind that the model's only predictor variables are mid-sized and major cities. The
intercept component incorporates the knowledge of tiny towns. Additionally, because these
predictor variables are dummy variables, their only two possible values are 0 or 1. For
example, if the location is a small town, mid-sized cities are equal to zero,
Recall the above average figures, this is the same average value for small towns. Now, if the
149
Again this is the same as the average value for mid-sized cities. Finally, the estimated profit
Now the next question is : how good is this model? For this we will have to scroll up to the
P values for certain coefficients: The value in the coefficients' right-most column, 2e-16, is
extremely low. This indicates that the coefficients won't go zero practically certainly,
according to the model. This is comparable to your odds of defeating Usain Bolt, which are
For our model, the adjusted R-squared value is 0.2065. This indicates that only the location
category accounts for 20% of the variation in earnings. This is not bad for a single category
variable, and if we continue to include more significant variables in the model described
Key Terms- Regression, Regression lines, correlation, linear regression, Standard error,
Summary- In this unit , we have covered topics related to correlation, correlation analysis,
various methods associated with it, regressions, and standard error, etc.
150
Exercise: Choose the correct option
A. Karl Pearson
B. R.A Fischer
D. Francis Galton.
2. Choose the least likely assumption of a classic normal linear regression model?
A. The independent variable and the dependent variable have a linear relationship.
151
D. None of the preceding.
3. Which one of the below statements regarding the regression line is correct?
A. Regression coefficient of X on Y
C. Regression coefficient of Y on X
D. Correlation coefficient of Y on X.
5. Which of the following statements is true about the arithmetic mean of two regression
coefficients?
152
C. A regression line is also known as the prediction equation
A. by 1
B. no change
C. by intercept
D. by its slope
A. lm(formula, data)
B. lr(formula, data)
C. lrm(formula, data)
D. [Link](formula, data)
153
A. Supervised Learning
B. Unsupervised Learning
C. Semi-Supervised Learning
2. What is regression?
Answers
Exercise: MCQ
1. (D)
2. (B)
3. (D)
4. (C)
154
5. (D)
6. (D)
7. (D)
8. (A)
9. (A&C)
10. (A)
A time series is a collection of observations on a single variable that are made at regular
intervals of time. The succeeding intervals are typically separated by equal amounts of time,
such as 10 years, one year, one quarter, one month, one week, one day, and one hour, etc. The
population statistics for India is a time series, with a 10-year lag between each succeeding
figure. Similar annual data are provided for national income, agricultural and industrial
production, etc.
The analysis of time series entails its breakdown into distinct elements that have an impact on
the variable's value over a specific period. It is a quantitative and objective assessment of the
impact of different variables on the activity in question. The analysis of any time series data
155
Because it enables us to understand the effects of diverse pressures, the study of historical
behaviour is crucial. This can make it easier to anticipate how events will develop in the
future, forecast the value of the variable, and make future plans.
There are three key elements in a typical time series that appear to be independent of one
Trend: The long-term, overall trend of either an increase or decrease in the forecast variable
y's average (or mean) value over time. Over time, the trend growth rate typically fluctuates.
Cycles- A cycle is defined as an upward and downward oscillation of unclear duration and
magnitude about the trend line caused by seasonal effect, with either a long time and irregular
swings. The average length of a business cycle is larger than one year but less than five to
seven years. Four phases make up the movement: peak (prosperity), contradiction (recession),
Seasonal: This is a specific instance of a cycle component of a time series in which the
cycle's size and length are constant and occur at regular intervals throughout the year. For
instance, festival seasons may see a significant boost in a retail store's average sales.
Irregular- A short-term unforeseen and non-recurring set of events can create erratic or
The principal methods of measuring trend fall into below mentioned categories:
2. Method of Averages
156
3. Method of least squares
The goal of the time series methods is to use a mathematical formula to predict the future of
an observable historical trend for a particular variable. These approaches make no attempt to
explain why the variable under research will change in the future. The use of a causal
mechanism overcomes this drawback of the time series approach. The causal approach looks
for variables that affect the variable in some way or cause it to fluctuate in a predictable way.
Regression analysis and correlation analysis are the two causal techniques that have already
been covered. Some time series techniques, such as freehand curves and moving averages,
only describe the values of the input data, but semi-average and least squares techniques
assist in finding a trend equation that may be used to characterise the input data values.
Freehand Method
The data can often be easily and possibly adequately represented by a freehand curve that is
drawn smoothly over the data values. By simply extending the trend line, the forecast may be
produced. The following prerequisites should be met by a trend line fitted by the freehand
method:
The following prerequisites must be met for a trend line to be fit by hand:
(i) The trend line should be straight or be a combination of long, progressive curves.
(ii) The total vertical deviation of the observations above the trend line should be equal to the
(iii) It is best to have a minimal sum of squares for the vertical deviations of the observations
157
(iv) The trend line should cut through the cycles so that, not only for the entire series, but
ideally for each complete cycle as well, the area above the trend line and the area below the
Example- Fit a trend line to the following data by using the freehand method.
Solution- presents the freehand graph of sales turnover (Rs. in lakh) from 1991 to 1998. The
so what works well for one person might not work well for another.
Building a freehand trend takes a lot of time if a cautious and meticulous job is to be
done.
158
Methods of Averages
The goal of smoothing techniques is to eliminate the random fluctuations resulting from the
time series' irregular components and, in doing so, give us a general sense of how the data are
moving over time. Three smoothing techniques will be covered in this section.
(iii) Semi-averages
The data requirements for the techniques to be discussed in this section are minimal and these
Moving Averages
The moving Averages Method gives a trend with a fair degree of accuracy. In this method,
we take the arithmetic mean of the values for a certain time span. The time span can be three
years, four -years, five- years and so on, depending on the data set and our interest. We will
It is crucial to first smooth out the irregular pattern in the historical values of the variable
before using this as the foundation for a future projection if we are watching the movement of
some variable values over time and trying to project this movement into the future. A series
of moving average calculations can be used to accomplish this. This method is arbitrary and
is reliant on the duration of the period used to calculate moving averages. The period should
159
cycle in the series in order to eliminate the impact of cyclical changes. The moving averages,
which serve as an estimate of the next period’s value of a variable given a period of length n
Moving average,
the term ‘moving’ is used because it is obtained by summing and averaging the values from a
given number of periods, each time deleting the oldest value and adding a new value.
Procedure:
(i) Select the moving averages' timeframe (three- years, four -years).
(iii) If the moving average is an odd number, centering it is not an issue; the average value
will be centred every three years, with the exception of the second year.
160
(v) If the moving average is an even number, the first four values' average will be positioned
between the second and third year, and the second four values' average will be positioned
between the third and fourth year. The third year will see a new average of these two
averages. This holds true for the remaining values in the issue. The centering of the averages
The limitation of this method is that it is highly subjective and dependent on the length of
(i) As the size of n (the number of periods averaged) rises, the approach becomes less
sensitive to actual changes in the data while also smoothing out variances better.
(ii) Moving averages struggle to detect patterns. Since these are averages, it won't predict a
change to either a higher or lower level because it will always remain within previous levels.
Example- Use the data below to calculate the number of students enrolled in a higher
Solution:
161
Example- Use the information below to calculate the number of pupils enrolled in a
Solution:
162
Weighted Moving Averages
weighted average of the most recent n values, different values could be used. Since there is
The most recent observation is typically given the most weight, whereas previous data values
As more recent data points are more pertinent than those from the distant past, weighted
moving averages give more weight to more recent data items. The weights should total 1 (or
163
Weighted moving average = Σ(Weight for period n) (Data value in period n) /ΣWeights
Example-
The specified price is multiplied by the corresponding weighting before the values are added
The WMA's denominator is the sum of the price periods expressed as a triangular number.
The weighted five-day moving average in the aforementioned example from the table would
be $22.65:
164
In this illustration, the most recent data point received the highest weighting out of a random
total of 15. Any value's values can be weighed however you see fit. The weighted average's
lower value in comparison to the simple average shows that recent selling pressure might be
stronger than some traders think. When utilising weighted moving averages, the most
Semi-Average Method
If a linear function can accurately describe the data, we may estimate the slope and intercept
of the trend using the semi-average method. Simply dividing the data into two sections and
calculating their individual arithmetic means is the technique. These two points are plotted
corresponding to the middle of the class interval that each portion covers, and a straight line
connecting these two points create the necessary trend line. The slope is determined by the
ratio of the difference in the arithmetic means of the number of years between them or the
change per unit of time, and the intercept value is the arithmetic mean of the first section.
A time series using the formula y = a + bx is the outcome. A and b are the intercept and slope
values, and y is the estimated trend value. Always include a reference to the year where x = 0
and a description of the units of x and y in your equation's full formulation. If the trend is
linear, the semi-average method of creating a trend equation may be acceptable and relatively
simple to commute. The forecast will be skewed and less accurate if the data diverge
The semi-averages are computed using this method to determine the trend values. We'll
Procedure:
(i) The information is split into two equal portions. If the number of data points is odd, two
equal sections can be created by simply leaving out the middle year.
165
(ii) Each component's average is calculated, giving us two points.
(iii) The midpoint (year) of each half is where each point is plotted.
(vi) According to semi-averages' methodology, this line represents the trend line.
Example 1- Fit a trend line by the method of semi-averages for the given data.
Solution-
Due to the odd number of years (seven), we will omit the production value of the middle year
and instead calculate the averages of the first three and last three years.
166
Example-2 Fit a trend line by the method of semi-averages for the given data.
Solution-
Since there are eight even years, we can divide the provided data in half and get the averages
for the first four years and the final four years.
167
Methods of Least Square
For medium- to long-term projections, the trend project approach involves fitting a trend line
to a set of historical data points and then projecting the line into the future. Depending on the
movement of time-series data, various mathematical trend equations (such as exponential and
1. The study of trends allows us to describe a historical pattern so that we may evaluate the
2. The study enables us to make future intermediate- and long-term forecasting projections by
3. Using trends as a reference for short-term (one-year or less) forecasting of general business
cycle conditions allows us to isolate and then minimise its influencing impacts on the time-
series model. Model for Linear Trend The least squares method can be used if we desire to
168
A least squares line's slope and y-intercept, or the height at which it intersects the y-axis, are
used to define it (the angle of the line). The following equation can be used to represent the
line if we can determine the y-intercept and slope. y = anticipated value of the dependent
variable, where y = a + bx an is the y-axis intercept. b = slope of the regression line (or the
rate of change in y for a given change in x) x = unrelated variable (which is time in this case)
Because it produces what mathematics refers to as a "line of best fit," least squares is one of
the most popular techniques for detecting trends in data. The characteristics of this trend line
include
(i) the summation of all vertical deviations about it is zero, that is, Σ(y- yˆ ) = 0,
(ii) the summation f all vertical deviations squared is a minimum, that is, Σ(y- yˆ ) is least,
and
(iii) the line goes through the mean values of variables x and y.
It is determined for linear equations by the simultaneous solutions of the two normal
equations, Σy = na + bx and xy = aΣx + bΣx2. When two terms in three equations can be
removed by coding the data so that ∑x = 0, we get ∑y = na and ∑xy = b∑x2 instead. When
working with time-series data, coding is simple. We chose x = 0 for the time period's centre
when coding the data, and we have an equal number of plus and minus periods on either side
of the trend line that add to zero. The values of constants a and b can also be determined as
̅̅̅̅
∑𝒙𝒚 − 𝒏𝒙𝒚
𝒃= 𝟐
̅ − 𝒃𝒙
,𝒂 = 𝒚 ̅
̅)𝟐
𝜮𝒙 − 𝒏(𝒙
169
When the distribution of the deviations is roughly normal, the method of least squares
provides the most accurate measurement of the secular trend in a time series.
The approach can be applied when the trend is quadratic, exponential, or linear.
Extremely big deviations from the trend are given too much weight by the least-
squares method.
Only during the period, it refers to the least-squares line is the best.
Its position could be altered by the removal or addition for one or more time periods.
The majority of financial, investment, and commercial choices are based on predictions of
Forecasting and time series analysis are crucial steps in understanding how financial markets
behave in a dynamic and powerful way. An expert can foresee the necessary projections for
crucial financial applications in a variety of sectors, such as risk evolution, option pricing &
Time series analysis, which may be used to forecast interest rates, foreign exchange risk,
stock market volatility, and many other things, has evolved into an integral aspect of financial
170
This study is used in investments to monitor price swings and a security's price evolution. For
For the short term, such as the observation per hour for a business day, and
For the long term, such as observation at the month end for five years
To track how a specific asset, security, or economic variable behaves or changes over time,
time series analysis is incredibly helpful. For instance, it can be used to assess how certain
underlying changes react when applied to other data observations made within the same time
period.
A data-driven industry, medicine has developed and is still making significant advances in
Think about the scenario where time series and a medical approach are combined. Data
mining and CBR (case-based reasoning) work in synergy to pre-process time series data for
feature mining, which can be used to track patients' development over time.
In the field of medicine, it is crucial to look at how behaviour changes over time rather than
drawing conclusions based just on the time series' absolute values. The typical demonstration
of linking time series with case-based monitoring is to diagnose heart rate variability in
However, time series in the context of the epidemiology domain has only lately and slowly
records should be connected over time and collected precisely at regular intervals.
171
Healthcare applications utilising time series analysis have produced significant
prognostication for the industry as well as for individual patients' health diagnoses once the
government has installed enough scientific devices to collect good and lengthy temporal
data.
Medical Instruments
Time series analysis has made its way into medicine with the advent of medical devices such
as
Medical professionals now have more opportunities to use time series for medical diagnostics
As a result of the development of wearable sensors and smart electronic healthcare devices,
people may now take routine measurements automatically and with little input, leading to a
reliable collection of longitudinal medical data for both ill and healthy people.
Different fields of astronomy and astrophysics are among the present and modern
Astronomical specialists are skilled in time series for calibrating devices and researching
things of their interest because astronomy, being specific in its field, heavily depends on
172
Time series data has a long history in the astronomy field; for instance, sunspot time series
were recorded in China in 800 BC, making sunspot data collecting as well-recorded natural
phenomenon. Time series data have an essential impact on knowing and quantifying anything
To discover variable stars that are used to surmise stellar distances, and
These systems, which depend on the wavelengths and light intensities of light to transmit
time series data in real-time, enable astronomers to observe phenomena as they happen.
Astroinformatics and astrostatistics are new fields of study that have emerged in recent
decades as a result of data-driven astronomy; these paradigms integrate key fields including
statistics, data mining, machine learning, and artificial intelligence. Here, time series analysis
would play a role in the quick detection and classification of astronomical objects as well as
Aristotle, a Greek philosopher, conducted research on weather events in antiquity with the
goal of determining the origins and consequences of weather variations. Later, scientists
began to compile weather-related data, recording it on an hourly or daily basis and storing it
in various locations, using the instrument "barometer" to calculate the status of atmospheric
conditions.
173
Newspapers started printing personalised weather forecasts over time, and as technology
Many countries have set up tens of thousands of weather forecasting stations all around the
world in order to conduct atmospheric measurements using computer methods for quick
compilations.
These stations are outfitted with highly functioning equipment and are connected to one
another in order to gather weather data from various locations and anticipate weather
As the process examines prior data trends, time series forecasting assists organisations in
making wise business decisions. It can be helpful in predicting future possibilities and events
Reliability: Time series forecasting is very trustworthy when the data has a wide range of
time intervals in the form of numerous observations over a longer time horizon. By
Growth: Time series is the best asset to use when evaluating endogenous as well as
improvement of internal human capital within firms that leads to economic growth. Time
series forecasting, for instance, can be used to analyse the effects of any policy variable.
Trend estimation: To find trends, time series methods can be used. For instance, these
Seasonal patterns: Variations in recorded data points may reveal seasonal patterns and
oscillations that serve as the foundation for data forecasting. The information gathered is
174
important for markets whose products vary seasonally and helps businesses manage their
Secular Trends
It speaks of the data's propensity to trend upward or downward over the long run. Examples
of secular trends that have an upward orientation include changes in productivity, an increase
in the rate of capital formation, population expansion, etc. Conversely, deaths brought on by
better medical care and cleanliness show a downward trend. All of these factors work slowly
1. Graphical method
1. Graphical Method
By placing the time variable on the X-axis and the value variable on the Y-axis, the values of
a time series are plotted using this method on graph paper. After that, a smooth curve is free-
hand drawn through the points that were plotted. To forecast the values, an extension of the
175
trend line shown above can be used. When sketching the smooth curve freehand, the
(ii) There should be roughly the same number of points above and below the line or curve.
(iii) The vertical deviation of the points above and below the smoothed line, when added
Merits
Demerits
It is a subjective method
The values of trend obtained by different statisticians would be different and hence not
reliable.
Example- Annual power consumption per household in a certain locality was reported
below.
176
Solution:
2. Semi-Average Method
The series is split into two equal halves using this manner, and the average of each portion is
(i) In the event that the series has an even number of years, it can be divided in half. Place the
values at the midpoint of each of the two series' respective durations by calculating the
(ii) It is impossible to divide a series into two equal halves if the number of years in the series
is odd. There won't be a midway year. Find the arithmetic mean for each segment of the data
(iii) The trend values for other years can be calculated by adding or subtracting successively
from each year that comes before or after any given year.
Merits
177
It does not require many calculations.
Demerits
It is used for the calculation of averages, and they are affected by extreme values.
Example- Calculate the trend values using semi-averages methods for the income from the
Solution:
178
Trend values for the previous and successive years of the central years can be calculated by
A series of arithmetic means of the variance values in a sequence make up a moving average.
Another method of creating a smooth curve for a time series of data is as shown here.
The seasonal variations are more typically removed using moving averages. The moving
average method, even when used to estimate trend values, aids in the establishment of a trend
line by removing the time series' cyclical, seasonal, and random changes. The length of the
When utilising this strategy, selecting the moving average's length is crucial.
The smoothing of variances for a moving average is greatly influenced by the selection of the
right length.
In general, if the number of years for the moving average is more, then the curve becomes
smooth.
Merits
It can be used to find the figures on either extreme; that is, for the past and future years.
Demerits
179
In non-periodic data, this method is less effective.
Selection of proper ‘period’ or ‘time interval’ for computing moving average is difficult.
Values for the first few years and as well as for the last few years cannot be found.
The following steps must be taken into account in order to determine the trend values
Add up the first three years' numbers, then compare the yearly total to the median year.
Keep the first year's value, add the values of the following three, and compare them to the
median year.
This procedure must be carried out again until all data values needed for calculations have
been obtained.
To obtain the 3-year moving averages, which serve as the trend values we need, each 3-
Example- Calculate the 3-year moving averages for the loans issued by co-operative banks
for non-farm sector/small scale industries based on the values given below.
180
Solution: The three-year moving averages are shown in the last column.
Summarize the first four years' values and arrange them between the second and third
years.
Leave the first year value alone and add the next four values starting with the second
This process must be repeated until the final item's value is considered.
Add the first two 4-year moving totals, then record the result next to the third year.
Keep the initial 4-year moving total and add the following two, then position them against
This procedure must be carried out repeatedly until all 4-yearly moving totals have been
Divide the 4-years moving total by 8 to get the moving averages which are our required
trend values.
181
4. Method of least squares
A line from which the sum of all deviations from various locations is zero is known as the
line of best fit. This is the most effective way to get trend values. It provides a practical
foundation for figuring out the time series' line of best fit. It is a formula for calculating
trends. In addition, as compared to other fitting techniques, the sum of the squares of these
variances would be the smallest. As a result, this technique is called the Method of Least
(i) The sum of the deviations of the actual values of Y and Ŷ (estimated value of Y) is Zero.
that is Σ(Y–Ŷ) = 0.
(ii) The sum of squares of the deviations of the actual values of Y and Ŷ (estimated value
Procedure:
(ii) The constants ‘a’ and ‘b’ are estimated by solving the following two normal
Equations ΣY = n a + b ΣX ...(2)
182
The constant ‘a’ gives the mean of Y and ‘b’ gives the rate of change (slope).
(v) By substituting the values of ‘a’ and ‘b’ in the trend equation (1), we get the Line of Best
Fit.
Secular trend, one of the time series' four elements, shows the direction the data will take
over the long run. The least squares approach is one mathematical methodology that can be
used to determine the trend values. The line created using this method is referred to as the
line of best fit because it is the most frequently utilised in practice and produces the least sum
It aids in value projections for the future. It is crucial for determining the trend values of time
Solved Examples-
Given below are the data relating to the production of sugarcane in a district.
Fit a straight line trend by the method of least squares and tabulate the trend values.
Solution:
183
Computation of trend values by the method of least squares (ODD Years).
Y = a+bX;
Given below are the data relating to the sales of a product in a district.
Fit a straight line trend by the method of least squares and tabulate the trend values.
184
Solution:
185
similarly other values can be obtained
The method of least squares is a device for finding the equation which best fits a given set of
observations.
Suppose we are given n pairs of observations, and it is required to fit a straight line to these
y = a + bx
Any value of a and b would give a straight line, and once these values are obtained, an
estimate of y can be obtained by substituting the observed values of y. In order for the
is desirable that the estimated values of yi, say y^ i on the whole close enough to the
observed values yi, i = 1, 2, …, n. According to the principle of least squares, the best fitting
186
is minimum. This leads us to two normal equations.
Solving these two equations, we get the vales for a and b and the fit of the trend equation
(line of best):
y = a + bx
Substituting the observed values xi in the above equation, we get the trend values yi, i = 1, 2,
…, n.
Merits
Trend values can be provided for each of the specified time periods.
187
Demerits
In comparison to the other ways, the calculations for this method are challenging.
necessitates recalculations.
Only the near future, not the far future, can be predicted for the trend.
(i) Subtract the first year from all the years (x)
(ii) Find ui = xi – A
y = a + b u = a + b (x – A)
Example
188
Fit a straight line trend by the method of least squares for the following consumer price index
Solution:
= a + bu where u = X – 2
y = 197.4 + 16.2 (X – 2)
189
= 197.4 + 16.2 X – 32.4
= 16.2 X + 165
X = 0, y = 165 + 0 = 165
Hence, the trend values for 2010, 2011, 2012, 2013 and 2014 are 165, 181.2, 197.4, 213.6
ii) Find ui = 2X – (n – 1)
Then follow the same procedure used in the previous method for odd years
190
2.3.6 Case Study (Time Series Analysis)
A client (a Multistore Retailer) had witnessed unusual fluctuation in demand for certain
SKUs from a particular product category. The client wanted to create a forecasting
model to estimate demand for certain SKUs for 1 to 12 months in the future. A time-
series dataset with monthly data for pricing, sales, and around 50 current-period or
lagged potential predictor factors was created. To forecast future demand, an ensemble
of LSTM and Autoregressive Time-Series Model was built. Forecasted demand was
used by the client company to better control production and inventory costs and boost
profitability.
Strategic Challenge:
Rapid expansion in Indian middle-class wealth causes periodic excess demand and big
spikes in the demand for specific SKUs in the client's manufacturing process.
Furthermore, fresh supplies of identical products on the market were rapidly arising,
The goal of the research was to create a forecasting model of the demand for a specific
data.
Build a robust forecasting model of demand for a given SKU using time-series
191
Create a web-based forecast simulation tool that allows clients to input updated
Analytical Design:
The Advanced Analytics team built a time-series analysis dataset, modifying all series to
be monthly and addressing missing values, holidays, and sales seasons properly. More
to identify the most promising linear and nonlinear predictors, lagged predictors, and
predictor combinations. The mean absolute prediction error was used to calculate
predictive power. More than ten distinct models were researched and analysed in order
to find the top five models for each required forecast time range (1,2,3,6,9 & 12
future, while accounting for serial correlation (the correlation over time of the impact of
unobserved variables on the variable being predicted—in this case, demand). The
accuracy.
Result:
The result was the creation of a Web-based forecasting tool that allowed the client's
management team to enter updated values of predictor variables each month and
forecast future demand for a certain SKU. Following that, the client organisation
evaluated the model by comparing forecasted vs. actual results for the first several
months. The resulting forecast accuracy was outstanding, prompting the client
organisation to:
192
Conduct a follow-up study utilising the forecasting method in another product
category.
Glossary
Regression: Sir Francis Galton's research on the heights of brothers through generations is
where the term regression first appeared. Children with unusually tall (or short) parents
explanatory variables X1 through Xk. The explanatory variables are multiplied by the
it relates to Y) and the interdependence of the explanatory variables, use the estimations
Time series are collections of statistical data that are organised and displayed according to
time. Based on the historical data in the time series, time series analysis predicts the data
When growth has ups and downs within the same year, we are witnessing seasonal
variance.
Cyclic patterns in the data across time are known as cyclical variation.
Free hand curves are rather straightforward, easy to understand, and uncomplicated.
Simply plot the curve using the data points that are available, then extend the trend line to
193
The moving average method bases its operation on arithmetic mean calculations of a
fixed number of readings OR observations over a certain period (3 years, 4 years etc.).
Least Squares: This technique employs regression analysis to identify the time series data'
trend line.
Important formula-
Y = a + bX
Where,
Y = Dependent variable
X = Independent variable
194
b = [n(∑xy) – (∑x)(∑y)]/ [n(∑x2) – (∑x)2]
Regression Coefficient
In the linear regression line, the equation is given by:
Y = b0 + b1X
byx = r(σy/σx)
bxy = r(σx/σy)
Where,
σx = Standard deviation of x
σy = Standard deviation of y
Key Terms- Time Series Analysis, Decision making, Business, Trends, Regression
Summary- In this unit, students have gone through, Methods of Least Squares, moving
averages, trend lines, cyclic patterns, time series, regression and regression analysis in details.
195
Exercise: Choose the correct option
A) Only 3
B) 1 and 2
C) 2 and 3
D) 1 and 3
E) 1,2 and 3
A) Naive approach
B) Exponential smoothing
C) Moving Average
A) Seasonality
B) Trend
C) Cyclical
D) Noise
196
A) Seasonality
B) Cyclical
5) The below time series plot contains both Cyclical and Seasonality component.
A) TRUE
B) FALSE
6) Adjacent observations in time series data (excluding white noise) are significantly
A) TRUE
B) FALSE
A) TRUE
B) FALSE
8) Which of the following is not a necessary condition for weakly stationary time series?
197
A) Mean is constant and does not depend on time
B) Autocovariance function depends on s and t only through their difference |s-t| (where t and
D) Smoothing Splines
10) If the demand is 100 during October 2016, 200 in November 2016, 300 in December
2016, 400 in January 2017. What is the 3-month simple moving average for February
2017?
A) 300
B) 350
C) 400
2. What is regression?
198
1. What are some real-world applications of Time-Series Forecasting?
2. The following figures relates to the profits of a commercial concern for 8 years
Find the trend of profits by the method of three yearly moving averages.
Answers
Exercise: MCQ
1. (E)
2. (D)
3. (E)
4. (A)
5. (B)
6. (B)
7. (A)
8. (D)
9. (C)
10. (A)
199
Module 3
Learning outcomes
Exploratory data analysis is the initial examination of data to identify links between variables
in the data and to acquire understanding of trends, patterns, and relationships between
different entities in the data set using statistics and visualisation tools (EDA).
Exploratory data analysis is divided into two categories, each of which is categorised as
Uni-Variate Analysis
Analyzing just one variable is referred to as univariate analysis. This is easy to remember
200
To comprehend the distribution of values for a single variable, use univariate analysis.
To learn more about the distribution of values in the dataset, we might opt to run
"household size":
201
There is only one reliable variable in a univariate analysis because uni means one and variate
means variable. Univariate analysis aims to derive the data, characterise and summarise it,
and examine any patterns that may be there. It investigates every variable in a dataset
The Central Tendency (mean, mode, and median), Dispersion (range, variance), Quartiles
(interquartile range), and Standard deviation are a few patterns that are simple to spot using
univariate analysis.
There are three typical methods for carrying out univariate analysis:
1. Summary Statistics
The most typical application of univariate analysis is the use of summary statistics to describe
a variable.
Measures of central tendency: They are numerical values used to pinpoint where a dataset's
Measures of dispersion: these figures show how dispersed the values in the dataset are.
Examples include the variance, standard deviation, interquartile range, and range.
2. Frequency Distributions
The creation of a frequency distribution, which details how frequently various values appear
3. Charts
Making charts to show the value distribution for a particular variable is yet another technique
202
Common examples include:
Boxplots
Histograms
Density Curves
Pie Charts
The Household Size variable from our previously described dataset is used in the examples
that follow to demonstrate how to carry out each type of univariate analysis:
We can calculate the following measures of central tendency for Household Size:
These values give us an idea of how spread out the values are for this variable.
203
More methods for Uni Variate Analysis-
The frequency distribution table displays the frequency of each occurrence in the data.
Example:
The list of IQ scores is: 118, 139, 124, 125, 127, 128, 129, 130, 130, 133, 136, 138, 141, 142,
IQ Range Number
118-125 3
126-133 7
134-141 4
142-149 2
150-157 1
Bar Charts
When comparing several categories of data or distinct groupings of data, the bar graph is
highly useful. Monitoring alterations over time is useful. For showing discrete data, it works
best.
Histograms
Histograms show the same categorical variables against the category of data as bar charts do.
These categories are shown in histograms as bins that represent the number of data points in a
204
Pie Charts
Pie charts are typically used to see how a group is divided into more manageable parts. The
slices of the pie indicate the relative size of that particular category, and the entire pie
indicates 100%.
205
Frequency Polygons
There are two variables in this situation since bi means two and variate imply variable. The
analysis focuses on the relationship between the two variables and the reason. The bivariate
1. Scatterplots.
2. Correlation Coefficients.
206
Scatter Plot
Dots are used in a scatter plot to symbolise distinct data points. These charts make it simpler
to determine whether two variables are connected. The pattern that emerges reveals the nature
(linear or non-linear) and intensity of the link between the two variables.
Linear Correlation
The degree of a linear link between two numerical variables is shown by linear correlation.
There is no propensity to change in accordance with the values of the second quantity if there
207
Here, r measures the strength of a linear relationship and is always between -1 and 1 where -1
denotes perfect negative linear correlation and +1 denotes perfect positive linear correlation
3. Linear Regression
With straightforward linear regression, bivariate analysis can be carried out in a third
approach.
We select one variable to serve as an explanatory variable and the other variable to serve as a
response variable using this approach. Next, we identify the line that most closely "fits" the
dataset so that we can determine the precise relationship between the two variables.
For instance, the dataset's line of best fit looks like this:
This indicates that an average exam score increase of 3.85 is correlated with each additional
hour of study. We can determine the precise correlation between study time and exam score
208
3.1.2 Bivariate Analysis of two categorical Variables (Categorical-Categorical)
Chi-square Test
To ascertain the correlation between categorical variables, utilise the chi-square test. It is
more frequency table categories. Two categorical variables are totally dependent on one
another when the likelihood is zero and completely independent when the probability is one.
Here, O stands for the observed value, E for the expected value, and subscript c denotes the
degrees of freedom.
Example- Let's imagine you want to determine if gender influences political party
preference in any way. To find out which political party respondents prefer, you
conduct a basic random sample poll of 440 voters. The table below displays the survey's
results:
Solution-
209
Step 1: Define the Hypothesis
Similarly, you can calculate the expected value for each of the cells.
Now you will calculate the (O - E)2 / E for each cell in the table.
Where
O = Observed Value
210
E = Expected Value
= 9.837
The crucial statistic must be identified before drawing any conclusions, which necessitates
knowing our degrees of freedom. The degrees of freedom in this situation are equal to the
product of the number of rows minus one and the number of columns minus one in the table,
The key statistic from the chi-square table is the last statistic you compare our acquired
statistic to. As you can see, the critical statistic is less than our actual statistic of 9.83 and is
5.991 for an alpha level of 0.05 and two degrees of freedom. Because the crucial statistic is
higher than your obtained statistic, you can reject our null hypothesis.
This indicates that there is enough evidence to support your claim that political party
211
3.1.3 Bivariate Analysis of one numerical and one categorical variable (Numerical-
Categorical)
A statistical test called the Z test is run on data that roughly fits the normal distribution. For
assessing hypotheses, the z test can be applied to proportions, two samples, or one sample.
When the population variance is known, it determines whether or not the means of two big
samples differ.
Calculating if there is a significant difference between a sample and the population requires
If the probability of Z is small, the difference between the two averages is more significant.
Z Test Formula
212
In order to determine whether there is a difference between the means of two populations, the
z test formula compares the z statistic with the z critical value. The z critical value separates
the acceptance and rejection sections of the distribution graph in hypothesis testing. The null
hypothesis can be rejected if the test statistic is within the rejection region; otherwise, it
cannot be rejected. Below is the z-test formula for setting up the necessary hypothesis tests
One-Sample Z Test
When the population standard deviation is known, a one-sample z test is performed to
determine whether there is a discrepancy between the sample mean and the population mean.
The following algorithm is provided to set a one sample z test based on the z test statistic:
Decision criteria: Reject the null hypothesis if the z statistic exceeds the z critical value.
Right-Tailed Test:
213
Alternative Hypothesis: H1: μ > μ0
Decision Criteria: Reject the null hypothesis if the z statistic is greater than the z critical
value.
Two-Tailed Test:
Decision Criteria: Reject the null hypothesis if the z statistic is greater than the z critical
value.
A two sample z test is used to check if there is a difference between the means of two
Similar to the one-sample test, the two-sample z test can be set up. The means of the two
samples will be compared using this test, nevertheless. The null hypothesis, for instance, is
stated as H 0: μ 1 = μ 2.
214
Example- A gym trainer claimed that all the new boys in the gym are above average
weight. A random sample of thirty boys weight have a mean score of 112.5 kg and the
population mean weight is 100 kg and the standard deviation is 15. Is there sufficient
Solution-
215
T-Tests
The t-test formula enables us to compare the average values of two data sets and ascertain
whether or not they represent the same population. The critical value from the t-table is used
to compare the t-score against. If the t-score is large, the groups are dissimilar, and if it is
The sample population is subjected to the t-test formula. The mean, variance, and standard
deviation of the data under comparison all affect the t-test formula. On the n number of
samples that were gathered, three different sorts of t-tests may be run.
One-sample test,
216
Independent sample t-test and
The degree of freedom (df = n-1) and the accompanying value are found using the t-table to
determine the critical value (usually 0.05 or 0.1). The initial premise is incorrect, and we infer
that the results are significantly different if the t-test yielded statistically > CV.
If the sample size is large enough, then we use a Z-test, and for small sample size, we use a
T-test.
T Test Formula-
217
Example- Find the t-test value for the following two sets of values: 7, 2, 9, 8 and 1, 2, 3,
4?
Solution-
218
Standard deviation for the first set of data: S1 = 3.11
When more than two groups' averages are statistically different from one another, the
ANOVA test is performed to assess whether there is a significant difference between them.
219
This comparison of averages of a numerical variable for more than two categories of a
Example of Anova
The following data is given:
Standard
Types of Animals Number of animals Average Domestic animals
Deviation
Dogs 5 12 2
Cats 5 16 1
Hamsters 5 20 4
Solution:
Construct the following table:
Animal name n x s s2
Dogs 5 12 2 4
Cats 5 16 1 1
Hamster 5 20 4 16
p=3
n=5
220
N = 15
x̄ = 16
SST = ∑n (x−x̄)2
SST= 5(12−16)2+5(16−16)2+11(20−16)2
= 160
MST=SST/p-1
MST=160/3-1
MST=80
SSE = ∑ (n−1) s2
SSE = 4 × 4 + 4 × 1 + 4 × 16
SSE = 84
MSE=SSE/N-p
MSE=8415/38415-3
MSE=7
F=MST/MSE
F=80/7
F=11.429
221
3.1.5 Multivariate Analysis
When more than two variables must be studied at once, multivariate analysis is necessary.
extremely difficult for the human brain to visualise a relationship among four variables on a
graph. Cluster analysis, factor analysis, multiple regression analysis, principal component
analysis, etc. are examples of multivariate analysis types. There are more than 20 alternative
approaches to multivariate analysis, and which one to choose depends on the data set and the
Cluster Analysis
Different objects are grouped together using cluster analysis so that there is a maximum
similarity between objects belonging to the same group and a minimum similarity in all other
cases. When the measure is distance or similarity and the rows and columns of the data table
A data table with several interconnected metrics is reduced in dimension using principal
component analysis, or PCA. The original variables in this case are changed into a fresh set
The dataset that demonstrates multicollinearity is analysed using PCA. The gap between
variances and their true value can be very wide, despite the bias in least squares estimations.
As a result, PCA increases some bias and decreases the regression model's standard error.
222
Correspondence Analysis
Using information from a contingency table, correspondence analysis can be used to reveal
relative relationships between and among two different groupings of variables. A contingency
analysis, it is more illuminating. But it's also more intricate than a single-variate study.
Multivariate statistics analyse three or more variables simultaneously to look for any potential
interactions.
Application areas
223
Climatology: (minimum temperature, maximum temperature. rainfall, humadity) on a day
Key Terms- Statistics, Hypothesis, Significance, Tests, Uni variate, Bi variate, Multi Variate,
Summary- In this unit, areas related to testing, types of testing hypothesis testing,
Univariate, Multivariate and Bivaraiate analysis, ANOVA, etc, have been covered in detail.
Case Study:
A FMCG firm wanted to investigate the effects of four different training programmes
separated into four equal-sized groups and then subjected to the various sales training
programmes. Because some students dropped out throughout the training programmes
owing to illness, vacations, and other reasons, the number of trainees who completed the
programmes varied by group. At the end of the training programmes, each salesperson
was assigned a sales area at random from a group of sales areas judged to have similar
sales potential. The table shows the sales made by each of the four groups of salespeople
224
Questions for Discussion-
the researcher.
a) Statistic
b) Hypothesis
c) Level of Significance
d) Test-Statistic
a) Null Hypothesis
b) Statistical Hypothesis
c) Simple Hypothesis
d) Composite Hypothesis
a) Null Hypothesis
b) Statistical Hypothesis
c) Simple Hypothesis
d) Composite Hypothesis
225
4. A hypothesis which defines the population distribution is called?
a) Null Hypothesis
b) Statistical Hypothesis
c) Simple Hypothesis
d) Composite Hypothesis
a) Null Hypothesis
b) Positive Hypothesis
c) Negative Hypothesis
d) Alternative Hypothesis.
7 What is an outlier?
226
a) It shows the results you would expect to find by chance
d) It compares the results you might get from various statistical tests
10. What is the name of the test that is used to assess the relationship between two
ordinal variables?
a) Spearman's rho
b) Phi
c) Cramer's V
d) Chi square
227
3. What is Uniariate analysis?
3. What are Z test and T test? Explain in detail with suitable example.
Answers
Exercise: MCQ
1. (B)
2. (A)
3. (B)
4. (C)
5. (D)
6. (C)
7. (C)
8. (B)
9. (D)
10. (A)
228
3.2. Basics of Hypothesis Testing
variables are connected to different elements of the study question. A testable prediction is
what a hypothesis is. A statement is examined in the research to see if it is true or incorrect in
order to determine its veracity. Diverse facets of the research issue must be investigated by
between variables relevant to each part of the research issue. The researcher might investigate
several areas of the research by testing these linkages. A hypothesis is a potential explanation
without a solid foundation. As a result, the researcher establishes logical connections between
or among the research's variables. The correlations between these variables, which are
connected by a common theme, provide the research's framework. These logical connections
or falsifiable presumptions provide the researcher with a starting point for the inquiry.
The decision-maker is given this tool through hypothesis testing. The operations manager
would take a sample of filled bottles from the ongoing bottling process if he were to employ
this instrument. The strength of the evidence the sample of bottles produced will be
A hypothesis that needs to be tested is the implicit assertion (μ = 1,000 cm 3), and the
hypotheses.
229
The following hypotheses, for instance, would be developed by a researcher investigating
The more patriarchy there is in a society, the more discrimination against women
The more traditional traditions present in a culture, the more discrimination against
Use of Family-Planning Practice in an Area": The higher the standard of education, the
The higher the availability of family-planning services, the higher the use of family
The higher the standards of living, the higher the use of the family-planning practice
will be.
Characteristics of Hypothesis
1. Empirically Testable
4. Predictable
5. Manageable
Importance of Hypothesis
230
1. It provides the research with a focus.
3. It identifies the researcher's emphasis because, in the absence of hypotheses, the research
7. It saves time, money, and energy because the researcher wouldn't have to focus these
1. Simple Hypothesis
2. Complex Hypothesis
4. Null Hypothesis
5. Alternative Hypothesis
6. Logical Hypothesis
7. Statistical Hypothesis
1. Simple Hypothesis
231
Any hypothesis that indicates a link between two variables—the independent and dependent
2. Complex Hypothesis
A hypothesis that indicates a relationship between more than two variables is referred to as a
complex hypothesis.
Examples:
1. The rate of crime will increase as poverty and illiteracy in society rise (three variables -
2. The agricultural productivity will increase if fertilisers, better seeds, and modern
equipment are used more frequently (Four variable-three independent variables and one
dependent variable).
3. Poverty and crime rates increase with the level of illiteracy in a culture. (Two dependent
3. Working Hypothesis.
A working hypothesis is one that has been approved for testing and development during the
investigation. It is a theory that is presumptively appropriate to explain certain facts and the
connections between various occurrences. This hypothesis is accepted to be tested for study
Any hypothesis that is originally accepted for consideration in the research is acceptable.
232
4. Alternative Hypothesis
A new hypothesis (to replace the working hypothesis) is produced and tested to examine the
desired feature of the research if the working hypothesis is incorrect or rejected. This new
As implied by the name, it is a different hypothesis (or connection) that is used when the
hypothesis.
5 Null Hypothesis
The aim of a null hypothesis is clear. Making a null hypothesis with the purpose to
disapprove, reject, or nullify it allows the researcher to confirm the existence of a relationship
between the variables. In order to establish that there is a relationship between the variables, a
null hypothesis is typically created as a reverse technique. 𝐻0 denotes the Null Hypothesis.
6 Statistical Hypothesis
A statistical hypothesis is a hypothesis that can be statistically tested. Any theory that has the
quantitative methods. A statistical hypothesis can also be stated to have measurable variables
7. Logical Hypothesis
233
A logical hypothesis is one that can be substantiated rationally. It is a relationship that may be
expressed as a hypothesis, and its veracity can be established by connecting its interlinks
using logical justifications. It can be supported by logical evidence to prove it. It doesn't
always follow that statistical methods cannot be used to verify a logical hypothesis. It might
or might not be statistically verifiable, but in light of the logical arguments, it looks so
The next stage is to acquire data from a representative sample of the population after
outlining the null and alternative hypotheses. The fact that we cannot be certain of our
samples cannot be completely eliminated until the sample size equals that of the population,
it is possible that the result reached is flawed and causes an error. There are two different
Type I Error
234
A Type I error occurs when a true null hypothesis is incorrectly rejected during statistical
testing. The operations manager would be making a type I error if he rejected 𝐻0 and
concluded that the process had gotten out of hand when in fact it had not.
When a null hypothesis is disregarded during the hypothesis testing procedure even though it
A null hypothesis is established prior to the start of a test in hypothesis testing. In some
circumstances, the null hypothesis makes the assumption that there is no causal connection
between the test item and the stimuli being given to the test subject in order to cause an
However, mistakes can happen where the null hypothesis is rejected, indicating that a cause-
and-effect link exists between the testing variables when in fact, a false positive occurred.
A hypothesis is tested using sample data in a process known as hypothesis testing. The goal
of the test is to demonstrate that the conjecture or hypothesis is supported by the inputted
data. The idea that there is no statistically significant relationship between the two data sets,
Consider the case where the null hypothesis holds that an investment plan doesn't outperform
a market index like the S&P 500. To find out if the investment strategy performed better than
the S&P, the researcher would test the historical performance of the method using samples of
235
data. The null hypothesis would be disproved if the test's outcomes revealed that the approach
n=0 is used to indicate this circumstance. The null hypothesis, which states that the stimuli do
not impact the test subject, would then need to be rejected if, after the test is completed, the
results appear to indicate that the stimuli administered to the test subject induced a reaction.
untrue, it should always be rejected. Errors can, however, arise under some circumstances.
It is occasionally wrong to reject the null hypothesis, which states that there is no connection
between the test subject, the stimuli, and the result. A "false positive" result occurs when it
appears that the stimuli had an effect on the subject but the outcome was random and
something other than the stimuli was responsible. A type I error is what is referred to as this
"false positive," which results in an inaccurate rejection of the null hypothesis. A type I error
Let's take the trial of a criminal suspect as an example. The person's innocence is the null
hypothesis, and guilt is the alternative. In this instance, a type I error would result in the
person being found guilty even though they were innocent and being imprisoned.
In medical testing, a type I error could provide the impression that treatment is lessening the
severity of the condition when it actually has no such impact. The null hypothesis in the
testing of a novel drug is that the drug has no effect on how the disease develops. Say a lab is
looking into a brand-new cancer medication. The medicine does not impact the rate at which
236
The cancer cells cease growing after the medicine has been applied to them. Thus, the null
hypothesis that the medicine would have no impact would be rejected by the researchers. In
this instance, rejecting the null would be the correct conclusion if the medicine was the
reason of the growth halt. However, this would be an example of an inaccurate rejection of
the null hypothesis if anything else during the test caused the growth slowdown rather than
Type II Error
When one fails to reject a null hypothesis that is actually wrong, this error is known
statistically as a type II error. This term is used in the context of hypothesis testing. A type II
error, often called an error of omission, results in a false negative. When the patient is
infected, a disease test, for instance, can return a negative result. This is a type II error
A type I error in statistical analysis is when a genuine null hypothesis is rejected, whereas a
type II error is when a false null hypothesis is not correctly rejected. Despite the fact that the
A type II error, often referred to as a second-kind error or a beta error, validates a hypothesis
that ought to have been disproven, such as the assertion that two observations are the same
despite the fact that they are not. Even when the alternative hypothesis represents the actual
state of nature, a type II error does not reject the null hypothesis. To put it another way, an
237
By establishing more strict standards for rejecting a null hypothesis, type II errors can be
minimised. If an analyst, for instance, considers anything that falls within the +/- bounds of a
95% confidence interval to be statistically insignificant (a negative result), then lowering that
tolerance to +/- 90% and then narrowing the bounds will result in fewer negative results and
lower the likelihood of a false negative. However, following these instructions tends to
increase the likelihood of running into a type I error—a false-positive outcome. The
likelihood or risk of committing a type I error or type II error should be taken into account
Type II error is the incorrect choice to accept (rather than reject, to be more precise) an
erroneous null hypothesis. The operations manager would be making a type II error if he did
not reject 𝐻0 and assumed that the process was under control when it had actually gotten out
of control. The number of type I and type II errors should be kept to a minimum because they
are both undesirable. Let's examine how to reduce the likelihood of type I and type II errors.
It may be clear that it is possible to completely eliminate the likelihood of type I mistake,
even with faulty sample evidence. Regardless of the evidence, simply accept the null
hypothesis.
We will never make a type I error since we will never reject any null hypotheses, including a
genuine null hypothesis. It is clear that this would be unwise, though. If we always accept a
null hypothesis, we will undoubtedly accept any false null hypothesis that is presented,
regardless of how absurd it may be. In other words, the likelihood that we will make a type II
lower the likelihood of type II error to zero because this would mean rejecting every genuine
null hypothesis, regardless of how accurate it is. We will have a type I error probability of 1.
238
Therefore, we cannot and should not try to completely avoid either type of error. We should
plan, organize, and settle for some small, optimal probability of each type of error.
Let's say a biotechnology company wishes to assess the effectiveness of two of its diabetes
medications. The two medicines are equally effective, according to the null hypothesis. The
claim that the corporation seeks to disprove with the one-tailed test is a null hypothesis, H0.
The counterargument, Ha, claims that the two medications are not equally effective. The
natural condition that is supported by rejecting the null hypothesis is represented by the
To compare the therapies, the biotech business conducts a significant clinical trial involving
3,000 diabetic patients. The 3,000 patients are randomly split into two groups of equal size,
with one group receiving one treatment while the other receives the other treatment. It
that it will reject the null hypothesis even if it is true or a 5% chance that it will make a type I
error.
Assume that 2.5%, or 0.025, is the beta value. Consequently, the likelihood of making a type
II error is 97.5%. The null hypothesis should be disproved if the two drugs are not equivalent.
However, a type II error happens if the biotech company does not reject the null hypothesis
A type I error rejects the null hypothesis even when it is true, in contrast to a type II error,
which does not (i.e., a false positive). The level of significance chosen for the hypothesis test
239
The probability of making a type II error, commonly known as beta, is one minus the test's
power. Increasing the sample size would boost the test's power while lowering the likelihood
The most common policy in statistical hypothesis testing is to establish a significance level,
denoted by α, and to reject 𝐻0 when the p-value falls below it. When this policy is followed,
In other words, we can say that the rejection region for 𝐻0 is the area under the curve where
the p-value is less than α. This region is also called critical region.
The standard values for α are 10%, 5%, and 1%. Suppose α is set at 5%. In the preceding
example, for a sample mean of 1,000.5, the p-value was 16%, and 𝐻0 will not be rejected. For
a sample mean of 1001, the p-value will be 2.28%, which is below α = 5%. Hence 𝐻0 will be
rejected.
Let us analyze in some detail the implications of using a significance level α for rejecting a
null hypothesis.
The first thing to note is that if we do not reject 𝐻0 , this does not prove that 𝐻0 is true. For
example, if α = 5% and the p-value = 6%, we will not reject 𝐻0 . But there is only about
6% chance that 𝐻0 is true, which is hardly proof that 𝐻0 is true. It may be possible that 𝐻0
is false and by not rejecting it, we are committing a type II error. For this reason, we
should say "We cannot reject 𝐻0 at an α of 5%" rather than "We accept 𝐻0 ."
240
The second thing to note is that α is the maximum probability of type I error we set for
probability of committing a type I error. In other words, setting α = 5% means that we are
The third thing to note is that the selected value of α indirectly determines the probability
of type II error as well. In general, other things remaining the same, increasing the value
of α will decrease the probability of type II error. This should be intuitively obvious. For
example, increasing α from 5% to 10% means that in those instances with a p-value in the
range 5% to 10% the 𝐻0 that would not have been rejected before would now be rejected.
Thus, some cases of false 𝐻0 that escaped rejection before may not escape now. As a
The fourth thing to note about α is the meaning of (1 - α). If we set α = 5%, then (1 - α) =
95% is the minimum confidence level that we set in order to reject 𝐻0 . In other words, we
The interval "inside the sample distribution of the test statistic that is consistent with the null
Let's put it another way: Suppose you perform a z-test-style hypothesis test. The test's results
are presented as a z-value, which has a wide range of potential values. Some values will fall
within an interval that shows the null hypothesis is true within that range of values. The
241
Due to the fact that a hypothesis test cannot tell you which hypothesis is true (the alternate or
null hypothesis) or even which is most likely true, you must provisionally accept the null
hypothesis.
It just examines if your data provide enough support to reject the null hypothesis. The null
Let's imagine if our "experiment" had you catching a kid in the act of stealing a cookie:
Null hypothesis (H0): The child didn’t steal the cookie (innocent until proven guilty!).
You have a good feeling the kid took the cookie. But even with all the evidence gathered, you
can't be certain that the youngster is guilty. The alternative theory that the youngster is
culpable is therefore unsupported by sufficient data. To put it another way, you can't accept
the null hypothesis that the child is guilty and reject the null hypothesis that the child is
innocent. This does not imply that the youngster is a victim. Simply put, you lacked sufficient
proof to hold them accountable. The null hypothesis of innocence is not actually "accepted,"
despite the fact that your result fell within the acceptance range. Simply accept it on a
Later on you might find crumbs in their bed, leading you to revisit your findings.
Rejection Region
A rejection region, also known as a critical region, is a region of the parameter space where a
null hypothesis will be rejected if a result is seen that falls inside it. Typically, a significance
test is used in hypothesis testing, and the rejection region is provided as a statistic, such as a t
score or a z score. The value of the real (non-standardized) parameter can be specified just as
simply.
242
Since a rejection zone is just a more precise way of expressing the significance threshold,
they are identical. A standardised score directly specifies the area under the parameter
distribution that will result in rejection because it is expressed in terms of standard deviations
(e.g., 1.644 SD or.1.96 SD). This information can be easily converted into alpha (α) during
the planning stage or a p-value after the test is finished by computing the cumulative
distribution function.
The observed p-value is compared to the crucial Z score, which is located at the edge of the
rejection region, once the test is finished, and if it falls within the region, the null hypothesis
is rejected.
Notably, if a value is outside the rejection zone, it does not always follow that the null
hypothesis can be accepted; rather, it only indicates that there is insufficient data to support
its rejection.
The observed p-value is compared to the crucial Z score, which is located at the edge of the
rejection region once the test is finished, and if it falls within the region, the null hypothesis is
rejected.
If your test findings fall into a particular section of a graph known as a rejection region (also
known as a critical region), you would reject the null hypothesis. In other words, your results
243
Testing ideas or experimental results is the main goal of statistics. For instance, you might
have created a brand-new fertiliser that you believe causes plants to grow 50% more quickly.
1. Your experiment must be able to be repeated in order to demonstrate that your theory is
accurate.
2. Be compared to a well-known plant fact (in this example, probably the average growth rate
A hypothesis test is the name given to this kind of statistical analysis. The testing procedure
includes the rejection region. In particular, it is a branch of probability that indicates the
The alpha level you decide to accept as a researcher is up to you. For instance, if you wanted
to have a 95% confidence level in the significance of your data, you would select a 5% alpha
level (100% - 95%). The rejection zone is between those 5% levels. The 5% in a one-tailed
test would be in one tail. The rejection region for a two-tailed test would be in both tails.
244
3.2.6 Procedure of Hypothesis Testing
A number of crucial ideas about hypothesis testing have been introduced to us. We can now
establish a standard testing approach in a more organised manner. By this point, it should be
evident that the process of testing a hypothesis essentially consists of two stages. In the first
stage, we design the experiment and establish the criteria for rejecting the null hypothesis. In
the second stage, we determine if the null hypothesis can be rejected using the sample data.
Step 1: State the Null and the Alternate Hypotheses. i.e. H0 and H1
Step 3: Choose the test statistic and define the critical region in terms of the test statistic
by comparing the observed value of the test statistic with the cut- off value or the critical
245
Key Terms- Hypothesis Testing, Type 1 and Type 2, Level of Significance, Region,
Summary- This unit deals with the hypothesis testing and its types, Null Hypothesis,
Alternate Hypothesis, and its procedures in details, level of significance, acceptance and
Case Study:
The slogan "made in China" has become a source of anxiety in recent years, as Indian
businesses seek to shield their products from foreign competition. In recent years, India
has seen a significant trade imbalance as a result of a flood of imported goods that enter
the nation and are offered at lower prices than comparable Indian-made items. One
major source of concern is electronic products, with total imported items continuously
increasing from the 1990s to the 2004s. Concerned about product quality concerns,
worker layoffs, and excessive costs, Indian corporations have spent millions on
simplify the analysis, we have coded the year using the coded variable x = Year 1989.
246
Questions for Discussion-
1. Determine the least-squares line for estimating import volume as a function of year
3. Predict the volume of commodities imported in each of the years 2002, 2003, and 2004
4. Are the forecasts made in Step 4 realistic approximations of the actual values
5. Enter the 1989-2004 data into your database and compute the regression line. What
influence did the new data points have on the slope> What effect does SSE have?
6. Does a straight line appear to be an accurate model for the data given the form of the
scattered diagram for the years 1989-2004? What other model would be more
appropriate?
247
Exercise: Choose the correct option
a) Level of Confidence
b) Level of Significance
c) Level of Margin
d) Level of Rejection
2. The point where the Null Hypothesis gets rejected is called as?
a) Significant Value
b) Rejection Value
c) Acceptance Value
d) Critical Value
3. If the Critical region is evenly distributed then the test is referred as?
a) Two tailed
b) One tailed
c) Three tailed
d) Zero tailed
a) Null Hypothesis
b) Simple Hypothesis
c) Alternative Hypothesis
d) Composite Hypothesis
5. Which of the following is defined as the rule or formula to test a Null Hypothesis?
a) Test statistic
248
b) Population statistic
c) Variance statistic
d) Null statistic
a) Right tailed
b) Left tailed
c) Center tailed
d) Cross tailed
7. Consider a hypothesis where H0 where ϕ0 = 23 against H1 where ϕ1 < 23. The test is?
a) Right tailed
b) Left tailed
c) Center tailed
d) Cross tailed
a) 1-α
b) β
c) α
d) 1-β
249
10. Alternative Hypothesis is also called as?
a) Composite hypothesis
b) Research Hypothesis
c) Simple Hypothesis
d) Null Hypothesis
Answers
Exercise: MCQ
1. (B)
2. (D)
250
3. (A)
4. (C)
5. (A)
6. (A)
7. (B)
8. (A)
9. (C)
10. (B)
251
The use of hypothesis testing in manufacturing processes includes establishing if the
implementation of a new technique or procedure in the manufacturing facility was the source
of any anomalies in the product's quality or not. Suppose manufacturing plant X checks
whether a certain method increases the number of defective products produced each quarter;
say this number is 200. To confirm this, the researcher must now compute the mean of the
number of defective items created before the start and end of the quarter.
Null Hypothesis (Ho): Before and after using the new production process, the average
Alternative Hypothesis (Ha): Before and after the new manufacturing procedure was
implemented, the average quantity of defective goods produced varied, i.e. μ after ≠ μ before.
The null hypothesis is disapproved and it may be concluded that changes in the method of
production cause a rise in the number of defective items produced each quarter if the
resulting p-value of the hypothesis testing is less than the significant value, i.e. α =.05.
252
To Plan the Marketing Strategies
strategies on the sales of the product, many organisations frequently use hypothesis testing.
For instance, the company's marketing division believed that increasing their investment in
digital advertisements would increase sales. The marketing department may increase the
budget for digital advertisements for a specific time period in order to test this assumption
and then at the conclusion of that time period, analyse the data that was gathered. To validate
Null Hypothesis (Ho): The average sales are the same before and after the increase in the
Alternative Hypothesis (Ha): After an increase in the budget for digital advertisements, i.e.,
The marketing department can reject the null hypothesis and come to the conclusion that
increasing the budget for digital advertising may increase product sales if the P-value is less
253
In clinical Trials
Hypothesis testing is used to analyse the effects of new therapeutic techniques, medications,
or procedures on patient health. For instance, a pharmacist thinks that the new medication is
causing diabetic patients' blood pressure to increase. The researcher must check the sample
patients' blood pressure before and after taking the new medication for roughly a specific
time period, say one month, in order to test this hypothesis. The hypothesis testing process is
Null Hypothesis (Ho): The average blood pressure is the same before and after taking the
Alternative Hypothesis (Ha): The average blood pressure before and after taking the
respectively.
The null hypothesis is rejected if the p-value of the hypothesis test is less than the
significance value (let's say.o5), at which point it can be deduced that the new medication is
254
In Testing Effectiveness of Essential Oils
Due to their multiple advantages, essential oils are becoming more and more popular
nowadays. Many essential oils, including ylang-ylang, lavender, and chamomile, make this
claim. You might want to investigate the real healing potential of each of these oils. Let's say
you believe that lavender essential oil can help you feel less anxious and stressed. You can
conduct hypothesis testing by restating the hypothesis as follows to examine this supposition:
Group A, or the experimental group, in this experiment receives lavender oil, whereas group
B, or the control group, receives a placebo. The data is then gathered using different
statistical tools, and both the experimental group and the control group's stress levels are then
analysed. Following the calculation, the p-value and significance threshold are discovered to
be 0.25 and 0.05, respectively. The null hypothesis is rejected because the p values are less
than the significance values, and it is therefore clear that lavender oil is effective at lowering
255
In Testing Fertilizer’s Impact on Plants
The effect of pesticides, fertilisers, and other substances on the growth of plants or animals is
now also investigated using hypothesis testing. Let's say a researcher wishes to confirm his
hypothesis that a certain fertiliser may cause a plant to grow more quickly in a month than its
typical growth of 10 inches. He constantly applied that fertiliser to the plant for over a month
to confirm this presumption. The mathematical process for testing the hypothesis in this
situation is as follows:
Null Hypothesis (Ho): There is no relationship between the fertiliser and plant growth. that
Alternative Hypothesis (Ha): The fertiliser causes the plant to grow more quickly,
The null hypothesis can be rejected if the p-value for the hypothesis test is less than the level
of significance, let's say.05, at which point you can draw the conclusion that the specific
256
In Testing the Effectiveness of Vitamin E
Let's say the researcher believes that vitamin E contributes to hair growing more quickly. In
an experiment he conducted, the experimental group received vitamin E for three months,
whereas the control group received a placebo. After three months have passed, the results are
then analysed. He restates his theory as follows to support his initial assumption:
The null hypothesis (Ho): It states that there is no correlation between vitamin E and the
Alternative Hypothesis (Ha): Under the assumption that all other factors remain constant,
the group of individuals who received vitamin E experiences faster hair growth than the
group as a whole did before doing so. In this case, μafter > before
Following statistical analysis, the significance level and the p-value, in this case, are,
respectively, o.o5 and 0.20. The researcher can therefore draw the conclusion that consuming
257
In Testing the Teaching Strategy
Let's imagine that Mr. X and Mr. Y, the two teachers, disagree on the ideal teaching
approach. Mr. Y contends that the weekly test is a waste of time and will not affect the
performance of the students in the yearly examinations, contrary to Mr. X's claim that giving
the children the weekly tests will improve their performance in the annual exams. Now we
may test our hypotheses to see which of the two is correct. The researcher could put forward
The null hypothesis (Ho): The average marks earned by the kids when they took the weekly
examinations and when they didn't were the same, proving the null hypothesis (Ho) that there
is no correlation between the weekly tests and how well the kids perform in the annual
Alternative Hypothesis (Ha): The children will score better on the yearly examinations if
they are required to take the weekly tests in addition to the annual exams or if μafter is
followed by μbefore.
The researcher can draw the conclusion that if the weekly assessment system is applied, the
children will score better on their annual tests if the p-value of the hypothesis testing is less
258
When examining the underlying premise of intelligence
Let's say a principal claims that the IQ level of the pupils attending her school is above
average. The researcher might select a sample of about 50 randomly chosen pupils from that
institution to back up her claim. Let's assume that those kids have an average IQ of roughly
110 and that the mean population as a whole has an average IQ of 100 with a standard
The null hypothesis (Ho): The population mean IQ score of 100 is a known fact, hence the
Alternative Hypothesis (Ha): The students' average IQ score is above average, or μ> 100.
We are going for the "greater than" assumption, thus it is a one-tailed test. Assume for the
moment that the significance level in this situation, or alpha level, is 5%, or 0.05, which
equates to a Z score of 1.645. The statistical calculation (112.5 - 100) / (15/√30) = 4.56 yields
the Z score. The last step is to compare the values of the computed and expected z scores.
The null hypothesis, which states that the average IQ score of the students attending that
school is above average, is rejected in this case since the calculated Z score is lower than the
expected Z score.
1. Conduct a hypothesis test to see if your decision and conclusion would change if your
belief were that the brown trout’s Mean I.Q. is not four.
Solution
a. H0:μ=4H0:μ=4
b. Ha:μ≠4Ha:μ≠4
259
c. Let X¯ X¯ the
average
e. t=1.95t=1.95
f. p-value=0.076p-value=0.076
h.
i. α:0.05α:0.05
hypothesis
average
i. (3.8865,5.9468)
seen or felt the presence of an angel. Some others question if the percentage is actually
so high. It carries out its own research. Only two of the 76 Americans who were polled
had really seen or felt the presence of an angel. Would you concur with the Newsweek
poll as a consequence of the contingent's poll? Give three reasons why the findings of
260
Solution
a. H0:p≥0.13H0:p≥0.13
b. Ha:p<0.13Ha:p<0.13
proportion
proportion
e. –2.688
f. p-value=0.0036p-value=0.0036
h.
i. alpha: 0.05
hypothesis
i. (0,0.0623)(0,0.0623).
The“plus-4s” confidence
interval
is (0.0022, 0.0978)
261
3. "Untitled," by Stephen Chen
I've often wondered how software is released and sold to the public. Ironically, I work
for a company that sells products with known problems. Unfortunately, most of the
problems are difficult to create, which makes them difficult to fix. I usually use the test
program X, which tests the product, to try to create a specific problem. When the test
program is run to make an error occur, the likelihood of generating an error is 1%.
So, armed with this knowledge, I wrote a new test program Y, that will generate the
same error that test program X creates, but more often. To find out if my test program
is better than the original so that I can convince the management that I'm right, I ran
my test program to find out how often I can generate the same error. When I ran my
test program 50 times, I generated the error twice. While this may not seem much
better, I think that I can convince the management to use my test program instead of
Solution
a. H0:p=0.01H0:p=0.01
b. Ha:p>0.01Ha:p>0.01
proportion
of errors generated
proportion
262
e. 2.13
f. 0.0165
h.
i. α:0.05α:0.05
hypothesis
the proportion
i. Confidence
interval
: (0,0.094)(0,0.094).
The“plus-4s” confidence
interval
is (0.004,0.144)(0.004,0.144).
263
The t-test serves as the big sample equivalent of the large sample z test. Generally speaking, a
sample with a n<30 is considered tiny. Due to the non-normal distribution of tiny samples, a
t-test is required.
One of the three different types of T-tests is the "One sample T Test." It is employed to
determine if the population's mean, from which the sample was selected, corresponds to
a given value.
The One Sample T Test is used to examine if a sample of observations might have
originated from a process that adheres to a particular parameter (like the mean).
For instance, you might want to determine whether a sample mean of 15 items matches a
postulated mean (population). In essence, you want to determine whether or not the
sample represents the given population. Let’s suppose you want to test if the mean
weight of a manufactured component (from a sample size 15) is of a particular value (55
The null hypothesis typically presupposes that there is no difference between the
hypothesised mean and the sample means (comparison mean). The T-Test is used to
The alternate hypothesis could be one of the following three scenarios, depending on
Case 1: H1 : x̅ != µ. used when the comparison Mean and the genuine sample mean are
264
Case 2: H1 : x̅ > µ. used when the comparison Mean is higher than the genuine sample
Case 3: H1 : x̅ < µ. used when the comparison Mean is greater than the genuine sample
Where x̅ is the sample mean and µ is the population mean for comparison.
Example- A customer service company wants to know if their support agents are
A report states that each ticket typically takes 20 minutes to be resolved. The sample
group's mean ticket purchase time is 21 minutes, with a 7-minute standard deviation. Can
you determine whether or not the company's assistance performance exceeds the industry
norm?
Step 1: Define the Null Hypothesis (H0) and Alternate Hypothesis (H1)
Example:
The alternate hypothesis can also state that the sample mean is greater than or less than
265
Step 2: Compute the test statistic (T)
𝑍 𝑥̅ − 𝜇
𝑡= =
𝑠 𝜎̂
√𝑛
Use the degree of freedom and the alpha level (0.05) to find the T-critical.
Step 4: Determine if the computed test statistic falls in the rejection region.
Alternately, simply compute the P-value. If it is less than the significance level (0.05 or
Example-
Problem Statement:
We have the potato yield from 12 different farms. We know that the standard potato
x = [21.5, 24.5, 18.5, 17.2, 14.5, 23.2, 22.1, 20.5, 19.4, 18.1, 24.1, 18.5]
Test if the potato yield from these farms is significantly better than the standard yield.
Solution:
266
Step 1: Define the Null and Alternate Hypothesis
H0: x̅ = 20
H1: x̅ > 20
n = 12. Since this is one sample T-test, the degree of freedom = n-1 = 12-1 = 11.
𝑥1 + 𝑥2 + 𝑥3 + ⋯ + 𝑥𝑛
𝑥̅ =
𝑛
𝑥̅ = 20.175
σ=3.0211
𝑥̅ − 𝜇 𝑥̅ − 𝜇
𝑇= =
𝑠𝑒 𝜎̂
√𝑛
267
20.175 − 20
𝑇= = 0.2006
3.0211
√12
Confidence level = 0.95, alpha=0.05. For one tailed test, look under 0.05 column. For
Because of how we specify the alternative hypothesis, a one-tailed test is used in this
case. We would have used a two-tailed test if the null hypothesis had been as simple as
268
Since the computed T Statistic is less than the T-critical, it does not fall in the rejection
region.
To establish whether or not two population means are equal, a two sample t-test is employed.
Let's say we wish to determine whether the mean weight of two different turtle species is
equal. Weighing each individual turtle would take too much time and money because there
Instead, we might choose 15 turtles at random from each group and use the average weight of
each sample to assess whether the mean weights of the two populations are equal:
269
The mean weight between the two samples will, however, almost certainly differ by at least a
(x̅1 -x̅ 2 )
Test statistic: t=
1 1
SP (√n +n )
1 2
where x
̅ 1 and x̅ 2 are the sample means, n1 and n2 are the sample sizes, and where sp is
calculated as:
You can reject the null hypothesis if the p-value for the test statistic t with (n1+n2-1) degrees
of freedom is less than your selected level of significance (popular options are 0.10, 0.05, and
0.01).
270
Example- Let's say we wish to determine whether the mean weight of two different turtle
species is equal. The following procedures will be used to conduct a two sample t-test with a
Suppose we collect a random sample of turtles from each population with the following
information:
Sample 1:
Sample size n1 = 40
Sample 2:
Sample size n2 = 38
We will perform the two sample t-test with the following hypotheses:
271
(𝑛1 −1)𝑠12 +(𝑛2 −1)𝑠22
𝑠𝑃 = √ 𝑛1 +𝑛2 −2
(40−1)18.52 +(38−1)16.72
= 𝑠𝑃 = √ 40+38−2
= 17.647
(x̅1 -x̅2 ) 1 1
t= = (300-305) / 17.647(√ + ) = -1.2508
1 1
SP (√ + ) 40 38
n n1 2
According to the T Score to P Value Calculator, the p-value associated with t = -1.2508 and
We are unable to reject the null hypothesis because this p-value is more than our level of
significance α, which is 0.05. The difference in mean turtle weight between these two
The estimated T statistic clearly does not fall into the rejection region. Therefore, the null
272
As long as the data has a normal distribution, a z test can be used to determine if the means of
two populations differ or not. It is necessary to set up the null hypothesis, the alternative
hypothesis, and calculate the value of the z test statistic for this reason. The z critical value
A normal distribution population with independent data points and a sample size higher than
If the z test statistic is statistically significant when contrasted with the crucial value, the null
In order to determine whether there is a difference between the means of two populations, the
z test formula compares the z statistic with the z critical value. The z critical value separates
the acceptance and rejection sections of the distribution graph in hypothesis testing. The null
hypothesis can be rejected if the test statistic is within the rejection region; otherwise, it
cannot be rejected. Below is the z test formula for setting up the necessary hypothesis tests
determine whether there is a discrepancy between the sample mean and the population mean.
𝑥̅ − 𝜇
𝑧= 𝜎
√𝑛
273
𝑥̅ is the sample mean, 𝜇 is the population mean, 𝜎 is the population standard deviation and n
The algorithm to set a one sample z test based on the z test statistic is given as follows:
Decision Criteria: If the z statistic < z critical value then reject the null hypothesis.
Decision Criteria: If the z statistic > z critical value then reject the null hypothesis.
Decision Criteria: If the z statistic > z critical value then reject the null hypothesis.
To determine whether there is a difference between the means of two samples, a two sample
z test is utilised. The following is the formula for the z test statistic:
(𝑥
̅̅̅1 − 𝑥̅2 ) − (𝜇1 − 𝜇2 )
𝑧=
𝜎12 𝜎22
√
𝑛1 + 𝑛2
274
𝑥1 𝜇1 , 𝑎𝑛𝑑 𝜎12 are the sample mean, population mean and population variance respectively
̅̅̅,
𝑥2 𝜇2 , 𝑎𝑛𝑑 𝜎22 are the sample mean, population mean and population
for the first sample. ̅̅̅,
Similar to the one-sample test, the two-sample z test can be set up. The means of the two
samples will be compared using this test, nevertheless. For example, the null hypothesis is
given as H0 : μ1=μ2.
275
For categorical data, a statistical test called Pearson's chi-square test is used. It is employed to
assess whether your data significantly depart from your expectations. The Pearson's chi-
To determine if two categorical variables are connected to one another, utilise the chi-
Among the most popular nonparametric tests are Pearson's chi-square (Χ2) tests, sometimes
known as chi-square tests. For data that do not adhere to the assumptions of parametric tests,
Use a chi-square test or equivalent nonparametric test if you want to test a hypothesis
groupings like animals or countries, can be nominal or ordinal. They cannot have a normal
There are two different kinds of Pearson's chi-square tests, but they all determine whether a
categorical variable's observed frequency distribution differs significantly from its predicted
a frequency distribution.
276
Frequency distribution tables are frequently used to depict frequency distributions. The
particular kind of frequency distribution table called a contingency table can be used to
display the number of observations in each combination of groups when there are two
categorical variables.
House sparrow 15
House finch 12
Black-capped chickadee 9
Common grackle 8
European starling 8
Mourning dove 6
Both of Pearson’s chi-square tests use the same formula to calculate the test statistic, chi-
square (Χ2):
277
Where:
The larger the chi-square, the greater the discrepancy between the observations and the
expectations (O E in the equation). You evaluate the chi-square value against a crucial value
These tests are actually the same in terms of mathematics. However, because they are
The Chi-square goodness of fit test can be used when there is only one categorical variable. It
significantly deviates from your expectations. The idea is that the categories will have equal
278
Example: Hypotheses for the chi-square goodness of fit test
Null hypothesis (H0): The bird species visit the bird feeder in equal proportions.
Alternative hypothesis (HA): The bird species visit the bird feeder
in different proportions.
Null hypothesis (H0): The bird species visit the bird feeder in the same proportions
Alternative hypothesis (HA): The bird species visit the bird feeder
in different proportions from the average over the past five years.
You may do an independence test using the chi-square formula when you have two
categorical variables. You can use it to see if there is a relationship between the two
variables. When two variables are independent of one another, they do not affect each other's
Null hypothesis (H0): The proportion of people who are left-handed is the same for
279
If there is a statistically significant difference between the means of three or more
The one-way and two-way ANOVAs are the two forms of ANOVAs that are used the most
frequently.
One-way ANOVA: Used to examine the influence of a single factor on a response variable.
Two-way ANOVA: used to ascertain the effects of two factors on a response variable and to
ascertain whether or not the two factors interact with the answer variable.
The following examples provide an example of how to perform each type of ANOVA.
Consider a professor who wants to discover if using three distinct study methods will result in
different exam results. He enlists 30students to take part in the study and randomly assigns
280
each participant to use one of the three methods to study for an exam in order to evaluate this.
All of the students take the same exam at the end of the month.
The professor performs a one-way ANOVA and gets the following results:
281
The p-value for the F test is 0.1138, and the F test statistic is 2.3575. We lack adequate
information to conclude that the three studying methods result in different mean exam scores
Consider a botanist who is curious about the effects of sunlight exposure and watering
frequency on plant growth. She sows 40 seeds and gives them two months to grow while
providing them with various amounts of sunlight exposure and hydration schedules. She
notes the height of each plant after two months. The outcomes are displayed below:
The professor performs a two-way ANOVA and gets the following results:
282
The interaction between watering frequency and sun exposure had a p-value of 0.310898.
For watering frequency, the p-value was 0.975975. At an alpha level of 0.05, this is not
statistically significant.
The p-value for exposure to sunlight was 0.000003. This has a 0.05 alpha level
significance statistically.
These findings suggest that the only variable with a statistically significant impact on plant
height is sunshine exposure. And because there is no interaction effect, the effect of sunlight
exposure is consistent across each level of watering frequency. That is, whether a plant is
watered daily or weekly has no impact on how sunlight exposure affects a plant.
283
Case Study (Hypothesis Testing)
For their lift policy salesforce, The Titan Insurance Company has just implemented a new
incentive payment programme. It wants to get a head start on determining whether the new
plan will work or not. There are signs that the sales force is selling more insurance, but since
The total sum guaranteed for the policies sold by a salesperson during the month is often how
life insurance firms gauge their monthly performance. Consider the scenario where salesman
X sold seven policies with the following sums assured: £1000, £2500, £3000, £5000, £10000,
and £35000. The total of these amounts assured, or £61,500, represents X's output for the
month.
With Titan's new strategy, salespeople receive little regular pay but are compensated
significantly with bonuses based on their performance (i.e. to the total sum assured of policies
sold by them). The plan is costly for the business, but sales growth is expected to more than
284
makeup for it. According to the agreement with the sales force, the plan would be scrapped
after six months if it does not at least break even for the business.
The programme has been running for four months at this point. After varying throughout the
Titan has selected 30 random salespeople to test the efficacy of the plan by measuring their
output in the penultimate month prior to transition and again in the fourth month following
the changeover (they have deliberately chosen months not too close to the changeover). The
1 57 62
2 103 122
3 59 54
4 75 82
5 84 84
6 73 86
7 35 32
8 110 104
9 44 38
10 82 107
11 67 84
12 64 85
285
13 78 99
14 53 39
15 41 34
16 39 58
17 80 73
18 87 53
19 73 66
20 65 78
21 28 41
22 62 71
23 49 38
24 84 95
25 63 81
26 77 58
27 67 75
28 101 94
29 91 100
30 50 68
Data Preparation-
286
It will be best to transform the given data to thousands as they are currently expressed
in 000.
Problem 1
Describe the five per cent significance test you would apply to these data to determine
whether the new scheme has significantly raised outputs. What conclusion does the test
lead to?
Solution:
It is asked whether the new scheme has significantly raised the output. It is an example of
Note: Two-tailed test could have been used if it was asked “new scheme has
Let,
H0: μ1 = μ2 ; μ2 – μ1 = 0
Since the population standard deviation is unknown, paired sample t-test will be used.
Since the p-value (=0.06529) is higher than 0.05, we accept (fail to reject) the NULL
287
Problem 2
Let's say it has been determined that Titan has to increase its average output by £5000 in
order to break even. What is the alternative hypothesis, if this figure is it:
(b) What is the p-value of the hypothesis test if we test for a difference of $5000?
Solution:
2.b. What is the p-value of the hypothesis test if we test for a difference of $5000?
Solution:
P-value = 0.6499
Solution:
HA: μd > 0
With α = 0.05 and df = 29, critical value for t statistic (or t_critical ) will be 1.699127.
288
Hence, H0 will be rejected for test statistics ≥ 1.699127.
Graphically,
289
Probability (type II error) is P(Do not reject H0 | H0 is false)
= P (t < | μd = 5000)
290
= P (t < -0.245766)
= 0.4037973
Now, β=0.5934752,
= 0.4065248
Since we only have evidence from the sample(s), hypotheses cannot be proved or
most.
Even if the test statistic is within the Acceptance Region or has a p-value of , using
the phrase "accept H0" in place of "do not reject" should be avoided. This merely
means that there is insufficient statistical support for rejecting the H0 from the
sample. H0 will not be accepted since we have attempted to negate (reject) it but have
The interval estimation technique known as the confidence interval can also be used
outside the confidence interval, which means that the confidence interval does not
Important formula-
291
𝑍 𝑥̅ − 𝜇
𝑡= =
𝑠 𝜎̂
√𝑛
2. sample mean
𝑥1 + 𝑥2 + 𝑥3 + ⋯ + 𝑥𝑛
𝑥̅ =
𝑛
4. T Statistic formula
𝑥̅ − 𝜇 𝑥̅ − 𝜇
𝑇= =
𝑠𝑒 𝜎̂
√𝑛
(x̅1 -x̅ 2 )
5. Test statistic: t=
1 1
SP (√n +n )
1 2
7. z test statistic:
𝑥̅ − 𝜇
𝑧= 𝜎
√𝑛
8. Two sample z test statistic:
292
(𝑥
̅̅̅1 − 𝑥̅2 ) − (𝜇1 − 𝜇2 )
𝑧=
𝜎12 𝜎22
√
𝑛1 + 𝑛2
9. chi-square (Χ2):
Key Terms- Sample Tests, Z Tests, Hypothesis Testing, Anova, Sample standard deviation.
Summary- In the given unit, students can learn the applications of hypothesis testing, and
sample t tests.
Glossary
Alpha Level- Alpha risk is another name for it. It's the likelihood that rejecting your
always a number between 0 and 1, with 0.05 being the most popular choice. When
your test is finished, and the data has been processed statistically, you will get a p-
Alternate Hypothesis- A hypothesis that disagrees with the null hypothesis; the two
Beta level- also known as the beta risk. It is the risk that you are willing to take to
make a Type B error, which is not to reject your null hypothesis when it is actually
false.
293
Conclusion- A declaration that details the strength of the evidence (sufficient or
insufficient), the significance level, and whether the original claim is confirmed (null)
Confidence level- called the confidence interval as well. This indicates how certain
you can be that your conclusion is accurate. Because the alpha and confidence levels
1 – α = confidence level
Critical region- All values that might lead us to reject the null hypothesis H0 are
Critical value(s)- The value(s) that define the boundary between the critical and non-
critical regions. The sample statistics are not used to determine the critical values.
The rejection region and the non-rejection region are divided by a critical value.
Critical value(s)- The value(s) that define the boundary between the critical and non-
critical regions. The sample statistics are not used to determine the critical values.
The rejection region and the non-rejection region are divided by the letter A.
Decision- a claim supported by the null hypothesis. Either the "null hypothesis" is
rejected, or the "null hypothesis is failed to be rejected." The null hypothesis won't
A p-value is a probability of getting a test statistic that is at least as extreme as the one
Error- In hypothesis testing, there are two main sorts of errors: type A errors, where a
valid hypothesis is disproved, and type B errors, where a false hypothesis is accepted.
294
Left-tailed test- The hypothesis test is a left-tailed test if the alternative hypothesis H1
Null hypothesis- The claim you're attempting to refute. This is the standard
P-value- A essential component of any hypothesis test results is the p-value. It's a
number between 0 and 1 and measures the likelihood that random fluctuations caused
any data that could lead you to reject the null hypothesis. It is determined by applying
a statistical significance test to test results. You reject the null hypothesis if the p-
value is less than your alpha threshold, and you do not reject the null hypothesis if the
When determining terms for a hypothesis test, the standard error is computed in a
manner known as "pooling." The two proportions are averaged in the pooled form,
and the standard error is calculated using just one proportion. Pooled computations
are preferred by ASQ, Villanova, and most other organisations. The two proportions
general.
References-
Press
295
Suggested Readings-
1. [Link]
testing/anova/
2. [Link]
significance-types-and-measures-statistics/15249
Study Tips-
1. Instead of mass practise, use distributive practise. To study statistics, set aside one to two
hours per day at the same time for six days of the week (leave the seventh day off). Avoid
cramming four or five hours of study into one or two sessions each week. This is a
fundamental idea.
2. At least once a week, study in student groups of three or four. A deeper level of learning is
actually cemented through verbal exchange and interpretation of concepts and skills with
other pupils.
3. Avoid attempting to retain formulas (A good instructor will never ask you to do this).
Concepts, concepts, concepts: study. Remember that you can always look up the formula in a
4. Complete as many and different of the exercises and problems as you can. Ideally, a
workbook is included with your textbook. Simply reading about statistics won't teach you
anything about it. Pushing the pencil and continually honing your techniques are required.
5. In statistics, look for recurring themes. There are probably only a few number of crucial
abilities that come up repeatedly. If necessary, request that your instructor stress these.
296
6. Become a Gestalt therapist! Recognize that statistics as a whole is larger than the sum of its
parts. It is very simple to become preoccupied with minor details and lose sight of the bigger
picture.
7. Take action if you suffer from math or statistics anxiety, which affects roughly 70% of
people in general. Most institutions offer top-notch counselling programmes to lessen this
impairment because they recognise how crippling this issue is. Get assistance for yourself.
This could end up being the wisest choice you make during your college career.
297
Exercise: State the Type I and Type II errors in complete sentences given the following
statements.
c. The mean starting salary for San Jose State University graduates is at least $100,000
per year.
f. The mean number of cars a person owns in his or her lifetime is not more than ten.
g. About half of Americans prefer to live away from cities, given the choice.
j. Private universities mean tuition cost is more than $20,000 per year.
3. What is data?
2. A particular brand of tires claims that its deluxe tire averages at least 50,000 miles before it
needs to be replaced. From past studies of this tire, the standard deviation is known to be
298
8,000. A survey of owners of that tire design is conducted. From the 28 tires surveyed,
the mean lifespan was 46,500 miles with a standard deviation of 9,800 miles.
3. A professor wants to know if her introductory statistics class has a good grasp of basic
math. Six students are chosen at random from the class and given a math proficiency test. The
professor wants the class to be able to score above 70 on the test. The six students get the
Can the professor have 90% confidence that the mean score for the class on the test would be
above 70?
Answers
a. Type I error: We conclude that the Mean is not 34 years, when it really is 34 years.
Type II error: We conclude that the mean is 34 years, when in fact it really is not 34
years.
b. Type I error: We conclude that more than 60% of Americans vote in presidential
elections, when the actual percentage is at most 60%.Type II error: We conclude that
at most 60% of Americans vote in presidential elections when, in fact, more than 60%
do.
299
c. Type I error: We conclude that the Mean starting salary is less than $100,000 when it
really is at least $100,000. Type II error: We conclude that the Mean starting salary is
d. Type I error: We conclude that the proportion of high school seniors who get drunk
each month is not 29% when it really is 29%. Type II error: We conclude that
the proportion of high school seniors who get drunk each month is 29% when, in fact,
it is not 29%.
e. Type I error: We conclude that fewer than 5% of adults ride the bus to work in Los
Angeles when the percentage that does is really 5% or more. Type II error: We
conclude that 5% or more adults ride the bus to work in Los Angeles when, in fact,
f. Type I error: We conclude that the Mean number of cars a person owns in his or her
lifetime is more than 10, when in reality it is not more than 10. Type II error: We
conclude that the mean number of cars a person owns in his or her lifetime is not
g. Type I error: We conclude that the Proportion of Americans who prefer to live away
from cities is not about half, though the actual proportion is about half. Type II error:
We conclude that the proportion of Americans who prefer to live away from cities is
h. Type I error: We conclude that the duration of paid vacations each year for Europeans
is not six weeks, when in fact it is six weeks. Type II error: We conclude that the
duration of paid vacations each year for Europeans is six weeks when, in fact, it is
not.
300
i. Type I error: We conclude that the Proportion is less than 11%, when it is really at
least 11%. Type II error: We conclude that the proportion of women who develop
j. Type I error: We conclude that the Average tuition cost at private universities is more
than $20,000, though in reality it is at most $20,000. Type II error: We conclude that
the average tuition cost at private universities is at most $20,000 when, in fact, it is
301
Suggested Reading
9788189611330
2. Levine, D., Sazbat, K. and Stephan, D. 2013. Business Statistics, 7thEdition, Pearson
4. Croucher, J. 2011. Statistics: Making Business Decisions, 13thEdition, Tata McGraw Hill,
ISBN: 9780074710419.
5. Gupta, S. 2011. Statistical Methods, 4thEdition, Sultan Chand & Sons, ISBN: 8180548627
302