Il 0% ha trovato utile questo documento (0 voti)
13 visualizzazioni302 pagine

Decision Science

Cu MBA notes

Caricato da

Bashidh Basheer
Copyright
© All Rights Reserved
Per noi i diritti sui contenuti sono una cosa seria. Se sospetti che questo contenuto sia tuo, rivendicalo qui.
Formati disponibili
Scarica in formato PDF, TXT o leggi online su Scribd
Il 0% ha trovato utile questo documento (0 voti)
13 visualizzazioni302 pagine

Decision Science

Cu MBA notes

Caricato da

Bashidh Basheer
Copyright
© All Rights Reserved
Per noi i diritti sui contenuti sono una cosa seria. Se sospetti che questo contenuto sia tuo, rivendicalo qui.
Formati disponibili
Scarica in formato PDF, TXT o leggi online su Scribd

Decision Science

1
Course Code: 22ONMBT***

Course Objective:

1. The objective of this course is to acquaint the learners with various statistical tools

and techniques used in business decision-making.

2. The course aims at providing fundamental knowledge and exposure to the learners to

use various statistical methods to understand, analyze and interpret data for decision

making.

3. The course has universal utility in making learners ready for corporate jobs and

entrepreneurial ventures.

Course Description:

This course begins with the basics of Statistics. It further elaborates on various sources of

data, depiction of data and applicability of the measure of the central tendencies and

dispersion. Later on the course emphasizes on the knowledge and applications of the

correlation, regression and time series data in real time situations. Finally, the learners got

a chance to acquaint themselves with the basics of Hypothesis Testing and applications

thereof.

2
Table of Contents

Module 1 .............................................................................................................................................................. 4
1.1: Basics of Statistics ...................................................................................................................... 4
[Link] Collection Sources ............................................................................................................ 19
1.3 Descriptive Statistical Methods................................................................................................ 39
Module 2 .......................................................................................................................................................... 100
2.1 Correlation Analysis ............................................................................................................... 100
2.2 Regression Analysis ................................................................................................................ 126
2.3. Time Series Analysis .............................................................................................................. 155
Module 3 .......................................................................................................................................................... 200
3.1. Types of Analysis.................................................................................................................... 200
3.2. Basics of Hypothesis Testing ................................................................................................. 229
3.3 Applications of Hypothesis Testing ....................................................................................... 251

3
Module 1

Learning outcomes

At the end of this module, you will be able to

 Explain the meaning, importance, limitations and applications of Statistics

 Discuss the different sources of data collection

 Formulate the various forms of descriptive statistics method

1.1: Basics of Statistics

1.1.1 Meaning

The collection, description, analysis, and drawing of conclusions from quantitative data are

all part of the field of statistics, which is applied mathematics. Calculus of differential and

integrals, linear algebra, and probability theory are key mathematical concepts in statistics.

The study of statistics involves the manipulation of data, including methods for data

collection, examination, analysis, and conclusion-making. Inferential and descriptive

statistics are the two main subfields of statistics. There are various ways to communicate

statistics, from non-numerical descriptors (nominal-level) to numerical in reference to a zero-

point (ratio-level). Simple random, systematic, stratified, or cluster sampling are just a few

sampling methods employed to gather statistical data. Almost every division of every

business uses statistics, which are crucial to investment.

1.1.2 Scope of Statistics

The scope of Statistics can be very well appreciated as the researcher follows:

4
 Statistics help in economic planning

To solve economic issues, statistical data and other statistical analysis approaches are

extremely helpful. Economic plans are formed on the basis of statistical information.

Statistics are also used to assess the plan's effectiveness. Numbers can be used to express

economic issues, including production, consumption, wages, price profits, unemployment,

poverty, etc.

 Statistics in business and management

For businesspeople, statistics are incredibly useful. It aids entrepreneurs in developing

corporate policies and predicting future trends. The foundation of modern business is the

precision and accuracy of estimates and statistical forecasting of future product demand,

market trends, and other factors. Businessmen's ability to accurately estimate and make

decisions depends on their experience and effective application of statistical techniques.

 Statistics in administration

Without numbers, efficient administration cannot be understood. Since statistics' inception,

they have been used to gather data on governmental fiscal and military policy.

The state compiles a vast amount of data on many characteristics of the populace and uses it

to create various policies for their welfare.

 Statistics in research

Every form of research project makes heavy use of statistical methodologies. The use of

statistics in doing various sorts of research is helpful in all fields, including social science,

health, and agriculture.

5
With the use of the proper statistical techniques, the yield of a crop grown with various seeds,

fertilisers, etc., may be studied. In a similar vein, statistics are very useful to the insurance

industry as well as to other sectors, including astronomy, banking, railroads, and the physical

sciences.

1.1.3 Importance and Limitations of Statistics

Every form of research project makes heavy use of statistical methodologies. The use of

statistics in doing various sorts of research is helpful in all fields, including social science,

health, and agriculture.

With the use of the proper statistical techniques, the yield of a crop grown with various seeds,

fertilisers, etc., may be studied. Similarly, statistics are very useful to the insurance industry

and various sectors. Below are the importance of statistics:

1. Statistics and planning

In the contemporary era, often known as the age of planning, statistics are essential to

planning. The government is returning to planning for economic development almost

everywhere in the world.

2. Statistics and economics

Statistical data and techniques of statistical analysis are the two terms that have to be

immensely useful in involving economic problems, such as wages, price, time series analysis,

and demand analysis.

3. Statistics and business

Statistics is an irresponsible device for production control. Business executives are counting

on more and more on statistical techniques for studying the much and desires of the important

customers.

6
4. Statistics and industry

In industry, statistics are widely used in quality control. In production engineering, to find out

whether the product is conforming to the specifications or not, various statistical tools, such

as inspection plans, control charts etc., are used.

5. Statistics and Mathematics

Statistics are intimately related to recent advancements in statistical techniques and are the

effect of wide applications of mathematics.

6. Statistics and modern science

In medical science, the statistical tools for collection, presentation, and analysis of observed

facts relating to causes and incidence of disease and the result of the application of various

drugs and medicine are of great importance.

7. Statistics, psychology, and education

In education and physiology, statistics have found wide applications such as determining or

determining the reliability and validity of a test, factor analysis, etc. There are many

applications of statistics in every field.

Branches of Statistics

Data collection, descriptive statistics, and inferential statistics are the three main subfields of

statistics. Let's examine these ideas in greater detail.

The primary Limitations of Statistics are the following:

7
1. Data Collection- The process of gathering data is what matters most. In terms of

mathematics, we often don't need to worry too much about this because we merely use the

data that is provided, but there are important considerations to make when gathering data.

This is quite simple for data like test scores from a class. Since each student has a specific

mark assigned to them, the data set is simply made up of all the marks. Data collection is

occasionally a challenge. Because bees move and fly around, it can be difficult to count them;

in these situations, you might have to estimate. Additionally, you should be mindful about

where you receive your data from if you are collecting it. Assume, for instance, that you want

to survey people about their election preferences. You must select a representative sample of

people because it is impractical to ask the entire nation (the population). It's not as simple as

it seems. For instance, polls were occasionally conducted in the middle of the 20th century by

phoning random numbers in a telephone directory.

This seems representative, but only the wealthy had telephones in those days, so you were

only polling a small portion of society—a group that would be more likely to support one

party than the other. When doing a survey by email, the same problem could arise today.

As a result, there are problems with data collection, and you must do prior to dealing, be

certain that the info has been gathered fairly. with it, attempt to communicate it, and draw

inferences.

When doing a survey by email, the same problem could arise today. As a result, there are

problems with data collection, and you must do prior to dealing, be certain that the info has

been gathered fairly. with it, attempt to communicate it, and draw inferences.

2. Descriptive Statistics- As previously said, it's frequently impractical to gather data from

the complete population because it would be very costly and time-consuming and the

8
population might change while you collect the data. Instead, we frequently need to take a

sample.

Descriptive statistics' fundamental goal is to "display the data" in an understandable manner.

Every bit of information needs to be summarised because if you just list it out in its entirety,

no one will understand it. Imagine if every single individual surveyed by a polling

organisation had their votes posted on the TV news; it would be a massive list of parties, and

you couldn't draw any conclusions. Instead, you are given visual representations of the data

(a bar chart, for instance) that may indicate the proportion of votes each party received. In the

general election of 2010, close to 30 million people cast ballots. You would be completely

bewildered if every vote were simply listed and displayed one after the other; instead, a

summary of votes is given (for example, as percentages: Conservative 36%, Labour 29%,

Liberal Democrat 23%, Others 12%). This is an illustration of descriptive statistics, which

"describe" or "summarise" the total data such that it is understandable to individuals.

Types of Descriptive Statistics- Either measures of central tendency or measurements of

variability, also referred to as measures of dispersion, are the foundation of all descriptive

statistics.

Frequency Distribution- The frequency distribution shows the frequency or count of the

various outcomes in a data collection or sample, and is used for both quantitative and

qualitative data. Typically, the frequency distribution is displayed as a table or graph. The

count or frequency of the values' occurrences within an interval, range, or particular group

are provided alongside each entry in the table or graph.

A presentation or summary of grouped data that has been divided into mutually exclusive

classes and the number of occurrences in each class is called a frequency distribution. It

9
enables the presentation of raw data in a more structured and orderly manner. Bar charts,

histograms, pie charts, and line graphs are examples of common charts and graphs used in

frequency distribution presentation and visualisation.

Central Tendency- A dataset's descriptive summary utilising a single value that represents

the centre of the data distribution is referred to as having a central tendency. Measures of

central location are another name for measures of central tendency. The measurements of

central tendency are the mean, median, and mode.

The average or most frequent number in a data set is known as the mean, which is regarded

as the most widely used measure of central tendency. The middle score for a data collection

in ascending order is referred to as the median. The score or value that appears most

frequently in a data set is referred to as the mode.

Variability- A summary statistic that reflects the level of sample dispersion is known as a

measure of variability. How far away the data points appear to fall from the centre is

determined by the variability measures.

The range and width of the distribution of values in a data set are indicated by the terms

dispersion, spread, and variability. The spread's many elements and features are represented

by the range, standard deviation, and variance, respectively.

10
The distance between the greatest and lowest values within a data collection is idealised as

the range, which shows the degree of dispersion. The standard deviation is used to calculate

the average variance in a set of data and gives information about how far a value from a data

set is from the mean value of the same data set. The variance, which is just an average of the

squared deviations, represents the extent of the dispersion.

3. Inferential Statistics- The part of statistics that deals with drawing inferences from the

data is called inferential statistics. You are simply asking, "What is this data saying to us, and

what should we do?" This is a rather broad area.

After a series of incidents, a council, for instance, might be considering lowering the speed

limit on a major road. In order to determine whether the speed limit should be lowered (for

instance, if a lot of automobiles are going too quickly), they might survey the speeds of cars

(gather data) first. Be aware, however, that this may not be the case; everyone may be

travelling at a pace that is totally acceptable, and the incidents may not be related to speed at

all. It's called inferential statistics when you take your available data and draw a "inference"

or "conclusion" from it. When we talk about topics like hypothesis testing, where we check to

see whether the data backs up a claim we make, we'll see much more of this in the future.

Difference between Descriptive and Inferential Statistics

11
BASIS FOR
DESCRIPTIVE STATISTICS INFERENTIAL STATISTICS
COMPARISON

Meaning Descriptive Statistics is A type of statistics known as


concerned with describing the inferential statistics concentrates on
population under study. inferring information about the
population from sample analysis
and observation.

Operations Organize, analyze and present Compares, test and predicts data.
data in a meaningful way.

Form of the final Charts, Tables and Graphs Probability


Result

Usage To describe a situation. To explain the chances of


occurrence of an event.

Function It explains the data, which is It aims to arrive at a conclusion


already known, to summarize about the population that goes
sample. beyond what is known from the
statistics.

The primary Limitations of Statistics are the following:

1. Statistics is incapable to explain individual items

Statistics is designed to study a group of values instead of single observation studies the mass

of phenomena and the conclusion on certain characteristics obtained.

2. Statistics are unable to study qualitative characters

In general, statistics studies only the quantitative characters of the given problems instead of

qualitative characters. The problems which cannot be studied quantitatively (i.e. in numerical

12
form), such as poverty, leadership, beauty, intelligence, honesty, etc., are not directly studied

in statistics.

3. Statistical results are not accurately correct

There are only a few results that are accurately correct in statistics, and almost all are only

approximately correct. In other words, a hundred percent accuracy is impossible in statistical

work because statistical laws are based on average.

4. Statistics deal with average

Statistics deals with averages that are obtained from different individual items. The laws of

statistics are true only on average.

5. Statistics is only one of the methods of studying a given problem

The best solution under all conditions of the given problem is not p by the statistical

methods.

Statistics cannot be of much help in studying the provided problem, like a country's culture,

religion, and philosophy unless they are supplemented by other evidence.

1.1.4 Applications of inferential statistics in managerial decision-making

Mean, median, mode, variance, and other descriptive statistics are used to describe the key

aspects of the data. Through numbers and graphs, it summarizes the data.

Inferential statistics allows us to infer information about the population from a sample.

Inferential statistics' major objective is to extrapolate findings from the sample to the

population data. For instance, we need to determine the typical data analyst salary in India.

There are two alternatives.

13
1. The first option is to study the data of data analysts across India and ask them about their

salaries and take an average.

2. The second option is to use a sample of data analysts from the central IT cities in India

and take their average and consider that across India.

The first alternative is not feasible since it is difficult to compile data from all data analysts in

India. Both time and money are spent on it. To solve this problem, we will consider the

second alternative: to gather a small sample of data analyst salaries and use the average of

those prices as the Indian average. Inferential statistics are those that allow us to infer

information about the population from a piece of data.

Key Terms- Statistics, Data, Decision making, Inferential Statistics, Descriptive Statistics,

Tables, Mean, Median, Mode

Summary- In this unit, we have discussed the basic idea of statistics, applications, and

limitations. Also students can get the idea of Descriptive and Inferential Statistics in this unit

Case Study:

A portfolio manager closely monitors the price-earnings ratios (defined as the current

market price divided by earnings over the previous four quarters) of 200 common

equities. He believes that it is appropriate to become an active buyer when the majority

of stocks in a representative sample have low price-earnings (P-E) ratios by historical

norms. Low P-E ratios may indicate that investors in general are overly pessimistic.

14
Furthermore, equities with low P-E ratios gain from rising earnings in two ways: (a) higher

earnings multiplied by a stable P-E ratio equals a higher market price, and (b) rising

earnings are frequently followed by rising P-E ratios.

Questions for Discussion-

(a) Arrange the variable's values in an array.

(b) Create a frequency distribution table.

(c) Display the frequency distribution as a histogram or frequency polygon and comment

on the pattern.

(d) Create an ogive and a cumulative frequency distribution.

15
Exercise: Choose the correct option

1. The arithmetic mean of the scores of a group of students in a test was 52. The

brightest 20% of them secured a mean score of 80 and the dullest 25% a mean score of

31. The mean score of remaining 55% is_________?

A. 45

B. 50

C. 51.4 approx

D. 54.6 approx.

2. Let A1, A2, be two AM’s and G1, G2 be two GM’s between a and b,then (A1 + A2) /

G1G2 is equal to _______

a) (a+b) / 2ab

b) 2ab/(a+b)

c) (a+b)/(ab)

3. The series a,(a+b)/2, b is in _______

a) AP

b) GP

c) HP

d) None of the mentioned

4. The series a, (ab)1/2, b is in _______

a) AP

b) GP

c) HP

d) None of the mentioned

16
5. If A and G be the A.M and G.M between two positive numbers then the numbers are

A + (A2 – G2)1/2, A – (A2 – G2)1/2.

a) True

b) False

6. If one geometric mean G and two airthmetic mean A1, A2 are inserted between two

numbers, then (2A1 – A2) (2A2 – A1) is equal to _______

a) 2G

b) G

c) G2

d) None of the mentioned

7. State whether the given statement is true or false.

AM ≤ GM.

a) True

b) False

8. If between two numbers which are root of given equation. x2 – 18x + 16 = 0, a GM is

inserted then the value of that GM is?

a) 4

b) 5

c) 6

d) 16

9. The kurtosis defines the peak of the curve in the region which is

a. around the mode

b. around the mean

c. around the median

17
d. around the variance

10. In kurtosis, the beta is greater than three and quartile range is preferred for

a. mesokurtic distribution

b. mega curve distribution

c. leptokurtic distribution

d. platykurtic distribution

Short Answer Type Questions-

1. What is statistics?

2. What is the significance of statistics?

3. Write down the branches of statistics.

Long Answer Type Questions-

1. Write down the applications of statistics in detail.

2. What are the Scope of statistics?

3. What are the limitations of statistics?

Answers

Exercise: MCQ

1. (C)

2. (C)

3. (B)

18
4. (D)

5. (D)

6. (B)

7. (A)

8. (C)

9. (D)

10. (B)

[Link] Collection Sources

A collection of measurements and facts is referred to as data, and it can be utilised to provide

an individual or organisation with the knowledge they need to make a well-informed choice.

It helps the analyst understand, examine, and evaluate a variety of socioeconomic problems,

including unemployment, poverty, inflation, etc. Understanding the difficulties helps in

determining their sources so that potential solutions can be developed. Along with theoretical

information, data may also contain specific numerical facts that lend weight to the claim.

Data collection is the initial stage of performing a statistical study, and it can be done using

either primary sources or secondary ones.

1.2.1 Sources of Data

 Primary Source

19
It is a synthesis of data from the primary source. It provides the researcher with raw,

quantitative data that is directly related to the statistical investigation. In other words, the

primary sources of information give the researcher immediate access to the research issue.

For example, statistical data, artistic efforts, interview transcripts, etc.

 Secondary Sources

It is a compilation of information drawn from several institutions or organisations that have

already acquired data from primary sources. Quantitative and unprocessed data that are

pertinent to the investigation are not made available to the researcher directly. In turn, the

secondary source of data analysis summarises or analyses the primary sources. As an

illustration, reviews, scholarly books and journals, articles, and government websites that

host polls or data

A solid study will need both primary and secondary sources of data collection, even though

primary sources provide the information more credibility because they include supporting

evidence.

1.2.2 Collection of Data

It's important to first define data before discussing collection. The short answer is that data is

a variety of types of information formatted in a specific way. As a result, data collecting is

the act of gathering, gauging, and analysing precise data from a range of pertinent sources in

order to address issues, provide answers, assess results, and predict trends and possibilities.

Data collection is crucial because of how heavily dependent our society is on it. To assure

quality assurance, keep research integrity, and make educated business decisions, accurate

data collecting is required.

The researchers must specify the data sources, data types, and methodologies used during

data gathering. We'll quickly find that there are numerous approaches to gathering data.

There is heavy reliance on data collection in research, commercial, and government fields.

20
Before an analyst begins collecting data, they must answer three questions first:

 What’s the goal or purpose of this research?

 What kinds of data are they planning on gathering?

 What methods and procedures will be used to collect, store, and process the information?

Data can also be divided into qualitative and quantitative categories. Descriptions like colour,

size, quality, and appearance are all included in qualitative data. Unsurprisingly, quantitative

data involves numbers. Examples include statistics, poll results, percentages, etc.

1.2.3 Data Collection Techniques

Let’s get into specifics. Using the primary/secondary methods mentioned above, here is a

breakdown of specific techniques.

1. Primary Data Collection

Interviews

The researcher surveyed a sizable sample of people through in-person interviews or other

forms of mass contact like phone or mail. This approach is by far the most typical way to

collect data.

Projective Technique

When prospective respondents are aware of the purpose of the interview and are reluctant to

reply, projective data collection is performed. For instance, if a representative from a cell

phone company asks them, a person might be hesitant to answer inquiries concerning their

phone service. Projective data collection requires interviewees to finish an incomplete

question with their thoughts, feelings, and attitudes.

21
Delphi Technique

Greek mythology describes the Oracle at Delphi as the high priestess of Apollo's temple who

offered counsel, prophesies, and wisdom. Researchers utilise the Delphi method to acquire

data by asking a group of experts for their opinions. Each expert responds to inquiries in their

area of expertise, and the responses are combined to form a single view.

Focus Groups

Focus groups and interviews are both frequently employed methods. A moderator gathers the

group of anywhere between six and twelve participants and serves as its facilitator.

Questionnaires.

A straightforward, easy-to-use approach to gathering data is through questionnaires. A series

of open-ended or closed-ended questions about the topic at hand are presented to the

respondents.

2. Secondary Data Collection

Unlike primary data collection, there are no specific collection methods. Instead, since the

information has already been collected, the researcher consults various data sources, such as:

 Financial Statements

 Sales Reports

 Retailer/Distributor/Deal Feedback

 Customer Personal Information (e.g., name, address, age, contact info)

 Business Journals

22
 Government Records (e.g., census, tax records, Social Security info)

 Trade/Business Magazines

 The Internet

1.2.4 Classification of Data

The data must be structured and presented in a coherent and understandable manner in order

to assist in further statistical analysis because the acquired data, also known as raw data or

ungrouped data, is always in an unorganised state. Thus, it is crucial for an investigator to

reduce a large amount of data into ever-smaller, more digestible chunks. Classification is the

process of dividing data into distinct classes or subclasses based on specific criteria, whereas

tabulation deals with the orderly arrangement and presentation of classified data. Therefore,

the initial stage in the tabulation is classification. For Example, letters in the post office are

classified according to their destinations, viz., Delhi, Madurai, Bangalore, Mumbai etc.,

Objects of Classification: The following are the main objectives of classifying the data:

1. It condenses the mass of data in an easily assimilable form.

2. It eliminates unnecessary details.

3. It facilitates comparison and highlights the significant aspect of data.

4. It enables one to get a mental picture of the information and helps in drawing inferences.

5. It helps in the statistical treatment of the information collected.

23
Types of classification: Statistical data are classified in respect of their characteristics.

Broadly there are four basic types of classification, namely

a) Chronological classification

b) Geographical classification

c) Qualitative classification

d) Quantitative classification

a) Chronological classification: In chronological classification, the collected data are

arranged according to the order of time expressed in years, months, weeks, etc., The data is

generally classified in ascending order of time.

Example-

b) Geographical classification: In this type of classification, the data are classified

according to geographical region or place. For instance, the production of paddy in different

states in Iraq, the production of wheat in different countries, etc.,

Example-

c) Qualitative classification: In this sort of categorization, data are categorised according to

the same qualities or traits, such as sex, literacy, religion, employment, etc. Such qualities

cannot be quantified using a scale. For instance, if the population needs to be categorised

24
according to one characteristic, let's say sex, we can divide them into two groups, males and

females. On the basis of other criteria, "marriage status," they can also be categorised as

"married" or "single."

As a result, when categorization is done based on a single attribute, which is dichotomous in

nature, two classes are created, one that possesses the attribute and the other that does not.

Simple or dichotomous classification is the name for this kind of classification.

A simple classification may be shown as under

Population

Male Female

A manifold categorization is one that takes into account two or more attributes and creates

many classifications.

For instance, if we categorise a population concurrently based on two characteristics, such as

sex and marital status, the population is first divided into males and females. On the basis of

the attribute "employment," each of these classes may then be further divided into "married"

and "single," and as a result, the population is divided into four classes.

(i) Male married

(ii) Male single

(iii) Female married

(iv) Female single

25
Still the classification may be further extended by considering other attributes like marital

status etc. This can be explained by the following chart

d) Quantitative classification: According to certain attributes that may be measured, such as

height, weight, etc., data is categorised quantitatively. For instance, the group of kids might

be divided into the following categories based on their weight.

In this type of classification, there are two elements, namely

(i) the variable (i.e.) the weight in the above example, and

(ii) the frequency in the number of children.

There are 50 children with weights ranging from 5 to 10 kg, 200 children ranging between 10

to 15 kg and so on.

Types of Tables

26
There are different types of tables based on a different basis. On the basis of the objective or

purpose of the study, the tables are classified into two types, viz.

(1) General purpose table and

(2) Specific purpose table or summary table.

On the basis of the nature of the data, the tables are classified into two types, viz.

(1) Primary, or original table and

(2) Secondary or derived table.

Further, on the basis of the elements or characteristics covered, tables are broadly classified

into two types, viz.

(1) Simple and

(2) Complex table.

Thus, the different types of tables can be diagrammatical as under:

1.2.5 Tabulation of Data

Tabulation is the process of condensing categorised or grouped data into a table format for

ease of comprehension and quick access to the needed information by an investigator. An

27
organised set of categorised data in columns and rows is called a table. A statistical table

enables the researcher to present a vast amount of data in a detailed and organised manner. It

makes comparison easier and frequently identifies patterns in data that would not be apparent

otherwise. In actuality, classification and "tabulation" are not two separate procedures. In

reality, they work in tandem. Data are categorised before being tabulated, and they are then

shown in various table columns and rows.

Depiction of Data

Data Representation: Analyzing numerical data is a technique called data representation.

The relationship between facts, ideas, information, and concepts is shown in a diagram

using data representation. It is a fundamental learning technique that is straightforward and

simple to comprehend. It is consistently based on the data type in a certain domain.

Graphical representations come in a wide variety of sizes and shapes.

A graph is a chart used in mathematics where statistical data is represented by c urves or

lines that cross the stated coordinate point on the surface. By enabling one to assess how

the amount of one variable has changed with respect to another over time, it facilitates the

research of a relationship between two variables. It is useful for analysing series and

frequency distributions in a given context.

Data Representation in Maths

Definition: The investigator must tabulate the data after gathering them in order to

evaluate their prominent characteristics. The presentation of data is the term used for such

a setup.

Any obtained data may be arranged in a frequency distribution table and then displayed

using pictographs or bar graphs. The lengths of the equally wide bars that make up a bar

28
graph, which is a depiction of numbers, depend on the frequency and scale you select. The

collected raw data can be placed in any one of the given ways:

1. Serial order of alphabetical order

2. Ascending order

3. Descending order

Data Representation Example

Example: Let the marks obtained by 30 students of class VIII in a class test, out of 50

according to their roll numbers, be:

39,25,5,33,19,21,12,41,12,21,19,1,10,8,1239,25,5,33,19,21,12,41,12,21,19,1,10,8,12

17,19,17,17,41,40,12,41,33,19,21,33,5,1,2117,19,17,17,41,40,12,41,33,19,21,33,5,1,21

The data in the given form is known as raw data or ungrouped data. The above-given data

can be placed in the serial order as shown below:

A few of the graphical representation of data is given below:

1. Bar chart

2. Frequency distribution table

3. Histogram

29
4. Pie chart

5. Line graph

Pictorial Representation of Data: Bar Chart

The qualitative data is represented visually in the bar graph. The data is presented either

horizontally or vertically and compares different quantities, traits, timings, and

frequencies.

The bars are organised in frequency order, emphasising the more important groups. It is

simple to determine which sorts of data in a collection outweigh the others by looking at

all the bars. Bar graphs come in a variety of forms, including single, stacked, and grouped.

Graphical Representation of Data: Frequency Distribution Table

A frequency table or frequency distribution is a way to present raw data that makes it

simple to comprehend the information it contains.

30
Using the tally marks, the frequency distribution table is created. A type of numerical

system that uses vertical lines for counting is called a tally mark. To get a total of 55, the

cross line is laid over the four lines.

Example:

Consider a jar containing the different colours of pieces of bread as shown below:

Construct a frequency distribution table for the data mentioned above.


Answer:

Graphical Representation of Data: Histogram

31
Another type of graph that displays data using bars is the histogram. The histogram is used

for numerical data, and classes, or ranges of values, are presented at the bottom. The

categories with higher frequencies have thicker bars than the other types.

Although a histogram and a bar graph appear quite similar, they differ due to the data

level. The frequency of the categorical data is represented in bar graphs. A categorical

variable, like gender or hair colour, has two or more categories.

Graphical Representation of Data: Pie Chart

The numerical distributions of a dataset are depicted using a pie chart. In this graph, a

circle is divided into various sectors, where each sector reflects the percentage of a

specific element as a whole. As a result, it is also referred to as a circular graph or chart.

32
Graphical Representation of Data: Line Graph

A line graph is a type of graph that use points and lines to show change over time. In other

words, the chart that depicts a line connecting several points or a line illustrating the

connection between the points is what we mean by this.

The straight line or curve that connects a number of succeeding data points in the graphic

serves as an illustration of the quantitative data between two changing variables. Two

variables are compared on the vertical and horizontal axes in linear charts.

33
Key Terms- Graphs, Tables, Data, Depiction of Data, Bar graphs, Pie Charts, Histogram,

Frequency.

Summary

This unit represents the various data representation methods in statistics such as bar graph,

histogram, pie charts etc, deptiction and tabulation of data.

Case Study:

The welfare committee of a big housing complex is investigating the prospect of hiring

private security guards to patrol the complex's front gate 24 hours a day, seven days a

week. The housing complex has 810 flats, and owners were asked to vote for or against

the project. The following information was gathered:

Should the Guards be appointed

Yes 194

No 121

Not sure 73

No response 422

Questions for Discussion

(a) Convert the data to percentages and create a bar chart and a pie chart. Which of these

graphs do you prefer? Why?

34
(a) After removing the 'no response' group, convert the remaining 388 replies to

percentages and create bar and pie charts again.

Assume you have been designated as a poll officer; what would you like to advise to the

president of the welfare committee based on your data analysis?

Exercise: Choose the correct option

1. Find the mode of the call received on 7 consecutive day 11,13,13,17,19,23,25

A. 11
B. 13
C. 17
D. 23

2. Find the median of the call received on 7 consecutive days 11,13, 17, 13, 23,25,19

A. 13
B. 23
C. 25
D. 17

3. Find the mode and median of the 9 consecutive number 12,7,8,14,21,23,27,7,11

A. 12,9
B. 7,9
C. 7,12
D. 11,9

4. When the Mean of a number is 18, what is the Mean of the sampling distribution?

A. 21
B. 18
C. 27
D. 23

5. What is the arrangement of data in rows and columns known as?

35
A. Frequency distribution

B. Cumulative frequency distribution

C. Tabulation

D. Classification

6. When the quantitative and qualitative data are arranged according to a single

feature, what is the tabulation known as?

A. One-way

B. Bivariate

C. Manifold division

D. Dichotomy

7. Which function does the tabulation origin spot specify?

A. The list of integers

B. The list of maxterms

C. The list of minterms

D. None of the above

8. In a tabular presentation, what is the summary and presentation of data with

different non-overlapping classes defined as?

A. Frequency distribution

B. Chronological distribution

C. Ordinal distribution

D. Nominal distribution

36
9. What are the general tables of data used to show data in an orderly manner known

as?

A. Double characteristic tables

B. Manifold tables

C. Repository tables

D. Single characteristics tables

10. What is the table where the variables are subdivided with interrelated features

known as?

A. Order level table

B. Sub-parts of a table

C. One-way table

D. Two-way table

Short Answer Type Questions-

1. What are Mean, Median, and Mode?

2. What data collection?

3. What are the sources of data.

Long Answer Type Questions-

1. Write down the data collection techniques in detail.

2. Write down the types of data tables in detail.

37
3. How data can be represented?

Answers

Exercise: MCQ

1. (B)

2. (D)

3. (C)

4. (B)

5. (C)

6. (A)

7. (A)

8. (A)

9. (C)

10. (D)

38
1.3 Descriptive Statistical Methods

Depending on the nature of the data and the goal for which it was obtained, the description of

statistical data may be highly detailed or quite succinct. One must watch out for being neither

too brief nor too long when describing data verbally or statistically. We can compare two or

more distributions for the same time period or within the same distribution over time using

the measures of central tendency. For instance, using an average, it is possible to compare the

average tea consumption in two separate regions for the same time period or in one region

over a period of two years, such as 2003 and 2004. The calculation of arithmetic mean by the

direct method is shown below.

1.3.1 Arithmetic Mean

Adding all the observations and dividing the sum by the number of observations results the

arithmetic mean. Suppose we have the following observations:

10, 15,30, 7, 42, 79 and 83

It may be noted that the Greek letter μ denotes the mean of the population and n denotes the

total number of observations in a population. Thus the population mean

39
μ = ∑x/n. The formula given above is the basic formula that forms the definition of

arithmetic mean and is used in the case of ungrouped data where weights are not involved.

Grouped data Arithmetic Mean

For grouped data, arithmetic mean may be calculated by applying any of the following

methods: (i) Direct method, (ii) Short-cut method, (iii) Step-deviation method 24. In the case

of the direct method, the formula x = ∑fm/n is used. Here m is the mid-point of various

classes, f is the frequency of each class, and n is the total number of frequencies. The

calculation of arithmetic mean by the direct method is shown below.

Example- The following table gives the marks of 58 students in Statistics. Calculate the

average marks of this group.

Solution-

40
It may be noted that the mid-point of each class is taken as a good approximation of the true

mean of the class. This is based on the assumption that the values are distributed fairly evenly

throughout the interval. When large numbers of frequencies occur, this assumption is usually

accepted.

In the case of the short-cut method, the concept of the arbitrary mean is followed. The

formula for calculation of the arithmetic Mean by the short-cut method is given below:

Where A = arbitrary or assumed Mean

f = frequency

d = deviation from the arbitrary or assumed mean

The use of the direct method would be quite laborious when the numbers are really big and/or

in fractions. The expedient approach is preferable in such circumstances. This is due to the

fact that the computation effort required by the shortcut technique is significantly decreased,

41
especially when it comes to calculating the product of values and their related frequencies.

However, suppose calculations are done automatically with a calculator rather than manually.

In that case, it might not be essential to employ the shortcut approach as using the direct way

might not present any issues.

The shortcut technique employs an arbitrary or assumed mean, as is evident from the

formula. The correction factor for the discrepancy between the actual mean and the assumed

mean is represented by the second component in the formula (∑fd ÷ n) . (∑fd ÷ n) will be 0

if the assumed mean proves to be the same as the real mean. The short-cut method's

application is predicated on the idea that the sum of deviations from the true mean is equal to

zero. As a result, how the assumed Mean is related to the actual mean will determine how

deviations are calculated from any other number. For the figures given earlier pertaining to

marks obtained by 58 students, we calculate the average marks by using the short-cut method

Example-

It may be noted that we have taken arbitrary mean as 35 and deviations from midpoints. In

other words, the arbitrary mean has been subtracted from each value of the mid-point and the

resultant figure is shown in column d.

42
Now we take up the calculation of arithmetic mean for the same set of data using the step-

deviation method.

It will be clear that the solution is the same in all three situations. Due to its easier

computations, the step deviation approach is the most practical. It should be highlighted that

the result would remain the same even if we chose a different arbitrary mean and recalculated

departures from that value. Since we now understand the various ways the arithmetic mean

can be determined, we are equipped to deal with any situation where the calculation of the

arithmetic mean is necessary.

43
1.3.2 Weighted Average MEAN
The average of the provided data set is the weighted mean, sometimes referred to as weighted

average. It is an average that was determined by giving various weights to various individual

variables. The weighted mean and the arithmetic mean are equivalent if all of the values are

the same.

The average mean or arithmetic mean is the same as the weighted mean. When data is

presented in various forms, it is calculated and compared to the sample mean or arithmetic

mean. Although weighted mean and average mean typically behave similarly, they do have

some divergent characteristics. Higher weighted data values than lower weighted data values

contribute to a greater weighted mean. Negative weights cannot exist; some of them may be

zero, but not all of them, because division by zero is forbidden. Compared to weights with a

lower weighted mean, contribute to the higher weighted mean

In the case of ungrouped data where weights are involved, our approach for calculating the

arithmetic mean will be different from the one used earlier.

Weighted Average Meaning

The average of all values that are prioritised is known as a weighted average. The weighted

average of values is calculated by dividing the total weight by the total value.

Formula for Weighted Mean

The following weighted mean formula can be used to calculate the weighted mean for a given

non-negative set of data with non-negative weights:

44
Example 1: Suppose a student has secured the following marks in three tests:

Mid-term test -30

Laboratory -25

Final -20

The simple arithmetic mean will be

1.3.3 Geometric Mean


The Geometric Mean (GM) is an average value or mean that, by calculating the product of a

collection of numbers' values, denotes the central tendency of that set of numbers. Measures

of central tendencies in mathematics and statistics are used to summarise the values of the

entire data set. Mean, median, mode, and range are the key metrics for identifying central

tendencies. One of these gives a general understanding of the data, which is the data set's

mean. The data set's mean indicates what the data set's average number is. The various mean

types are arithmetic mean (AM), geometric mean (GM), and harmonic mean (HM).

What is Geometric Mean?

The Geometric Mean (GM) is the average value or mean that, by taking the product of a set

of numbers' values as its root, indicates the set's central tendency. In essence, where n is the

45
total number of values, we multiply all 'n' values and subtract the nth root of the numbers. For

instance, the geometric mean for a pair of numbers, such as 8 and 1, is equivalent to √(8×1) =

√8 = 2√2.

As a result, the geometric mean is also described as the product of n numbers at the nth root.

Keep in mind that this is not the same as the arithmetic mean. Data values are summed, then

divided by the total number of values to determine the arithmetic mean. In contrast, the given

data values are multiplied in a geometric mean, and the final product of the data values is

obtained by taking the root of the radical index. Take the square root, for instance, if you

have two data values. If you have three data values, take the cube root; if you have four, take

the fourth root; and so on.

Geometric Mean Formula

There are two additional methods that are occasionally employed in business and economics

in addition to the three central tendency measures already mentioned. The geometric mean

and the harmonic mean are these. The harmonic mean is less significant than the geometric

mean. Below, we go over both of these methods. We start with the geometric mean. The

geometric mean of a distribution is determined by taking the nth root of the product of its n

observations. Symbolically,

If we have only two observations, say, 4 and 16, then

Taking logarithm on both sides,

log GM = log (x₁ · x₂ · ... · xₙ)1/n

46
= (1/n) log (x₁ · x₂ · ... · xₙ)

= (1/n) [log x₁ +log x₂ + ... + log xₙ]

= (∑ log xᵢ) / n

So, the geometric mean is defined as GM = Antilog (∑ log xᵢ) / n.

This is a different version of GM's formula.

Similar calculations must be made if there are three observations: we must determine the

cube root of the product of these three observations, for example, and so on. The more

elements there are, the more challenging it is to multiply the integers and determine the root.

Logarithms are employed in calculations to make them simpler.

Difference between Arithmetic Mean and Geometric Mean

Arithmetic Mean Geometric Mean

The geometric mean can be found by


In the arithmetic mean, we add data
multiplying all the numbers in the given
values and then divide it by the total
data set and take the nth root for the
number of values.
obtained result.

For example, the given data sets are: For example, for data set, 4, 10, 16, 24
10, 15 and 20 Here n = 4
Here, the number of data points = 3 Therefore, the G.M = (4 ×10 ×16 ×
Arithmetic mean or mean = 24)1/4
(10+15+20)/3 = 153601/4

Mean = 45/3 =15 G.M = 11.13

Relation between AM, GM and HM

47
Before learning the relationship between the AM, GM, and HM, it is necessary to understand

their respective formulas. Assuming "a" and "b" are the two numbers and that there are 2

values,

AM = (a+b)/2

⇒ 1/AM = 2/(a+b) ……. (I)

GM = √(ab)

⇒GM2 = ab ……. (II)

HM= 2/[(1/a) + (1/b)]

⇒HM = 2/[(a+b)/ab

⇒ HM = 2ab/(a+b) ….. (III)

Now, substitute (I) and (II) in (III), we get

HM = GM2 /AM

⇒GM2 = AM × HM

Or else,

GM = √[ AM × HM]

As a result, 𝐺𝑀2 = AM × HM is the relationship between AM, GM, and HM. As a result,

the geometric mean's square is equal to the sum of the harmonic and arithmetic means.

Let us also see why the G.M for the given data set is always less than the arithmetic mean for

the data set. Let A and G be A.M. and G.M.

So,

A = (a+b)/2 and G=√ab

48
Now let’s subtract the two equations

A−G = (a+b)/2 − √ab = (a+b−2√ab)/2 = (√a−√b)2/2 ≥ 0

A−G ≥ 0

This indicates that A ≥ G

Applications of Geometric Mean

The geometric mean is utilised in many fields and has numerous advantages over the

arithmetic mean. The following are a few examples of applications:

 Since many of the value line indexes used by financial departments employ G.M. to

determine the annual return on the investment portfolio, it is used in stock indexes.

 In finance, the geometric mean is used to determine average growth rates, commonly

referred to as the compounded annual growth rate (CAGR).

 Additionally, studies on biological processes like bacterial growth and cell division use

geometric means.

1.3.4 MEDIAN and MODE

Median
When the data are ordered in an ascending or descending order of magnitude, the median is

defined as the value of the middle item (or the mean of the values of the two middle items).

In an ungrouped frequency distribution, the median is the middle value if n is odd, and the n

values are sorted in ascending or descending order of magnitude. The median is the mean of

the two middle values when n is even.

Suppose we have the following series:

15, 19,21,7, 10,33,25,18 and 5

49
We have to first arrange it in either ascending or descending order. These figures are arranged

in ascending order as follows:

5,7,10,15,18,19,21,25,33

Now, as the series consists of an odd number of items, to find out the value of the middle

item, we use the formula

n+1
𝑚𝑒𝑑𝑖𝑎𝑛 =
2

Where n is the number of items. In this case, n is 9; as such,

n+1
=5
2

That is, the size of the 5th item is the median. This happens to be 18.

Assume there are 23 total items in the series. In order to include 23 in the above sequence at

the proper location, that is, between 21 and 25, we may have to do so. The series now

consists of 5, 7, 10, 15, 18, 19, 21, 23, 25 and 33. Using the calculation above, the median

size is the 5.5th item. Here, we must average the values of the fifth and sixth items. With an

average of 18 and 19, the median is 18.5, according to this.

n+1
It may be noted that the formula itself is not the formula for the median; it merely
2

indicates the position of the median, namely, the number of items we have to count until we

arrive at the item whose value is the median. In the case of the even number of items in the

series, we identify the two items whose values have to be averaged to obtain the median. In

50
the case of a grouped series, the median is calculated by linear interpolation with the help of

the following formula:

l2 + l1
𝑀 = l1 (𝑚 − 𝑐)
f

Where M = the median X

l1 = the lower limit of the class in which the median lies

l2 = the upper limit of the class in which the median lies

f = the frequency of the class in which the median lies

m = the middle item or (n + 1)/2nd, where n stands for the total number of items

c = the cumulative frequency of the class preceding the one in which the median lies

Mode

The mode is a different way to quantify central tendency, and it is the value at the location

where things are concentrated most intensively. Consider the following sequence as an

illustration: 8,9, 11, 15, 16, 12, 15,3, 7, 15 37 The greatest number of times figure 15 appears

in a sequence of ten observations is three. Therefore, the mode is 15. Because the series listed

above is discontinuous, the variable cannot be in the fractional form. If the series were

continuous, we could state the mode to be around 15 without performing any more

calculations. The following formula is used to identify the mode for grouped data:

f1 − f0
𝑀𝑜𝑑𝑒 = l1 𝑋(𝑖)
(f1 − f0) + (f1 − f2)

Where, l1 = the lower value of the class in which the mode lies

51
f1 = the frequency of the class in which the mode lies

fo = the frequency of the class preceding the modal class

f2 = the frequency of the class succeeding the modal class

i = the class-interval of the modal class

The class intervals should be consistent throughout when using the aforementioned formula.

On the grounds that the frequencies are dispersed equally over the class, the class intervals

should be made uniform if they are not already. The application of the aforementioned

formula in the case of unequal class intervals will produce false results.

Dispersion-

The numerous central value measures provide us with a single number that sums up all the

data. But until all of the observations are identical, the average by itself cannot effectively

explain a group of observations. The variability or dispersion of the observations must be

described. Even if the central value in two or more distributions may be the same, the way the

distributions are formed can vary greatly. We may analyse this crucial aspect of distribution

using measures of dispersion.

Following are some key definitions of dispersion:

1. "Dispersion is the measure of the variation of the items." -A.L. Bowley

2. "The degree to which numerical data tend to spread about an average value is called the

variation of dispersion of the data." -Spiegel

3. Dispersion or spread is the degree of the scatter or variation of the variable about a

central value." -Brooks & Dick

52
4. "The measurement of the scatterness of the mass of figures in a series about an average is

called measure of variation or dispersion." -Simpson & Kajka

Measure of Dispersion

Inter-quartile Range or Quartile Deviation,Range, Standard Deviation, Mean Deviation, and

the Lorenz Curve are the five measurements of dispersion. The first four of them are

mathematical techniques, and the final one is a graphical technique.

1.3.5 Range

The range, or difference between a data set's maximum value and minimum value, is the

easiest way to assess dispersion.

The formula to calculate the range is:

R=H-L

 R = range

 H = highest value

 L = lowest value

The range is the easiest measure of variability to calculate. To find the range, follow these

steps:

1. Order all values in your data set from low to high.

53
2. Subtract the lowest value from the highest value.

This process is the same regardless of whether your values are positive or negative, or whole

numbers or fractions.

Range example- Your data set is the ages of 8 participants.

Participant 1 2 3 4 5 6 7 8

Age 37 19 31 29 21 26 33 36

First, order the values from low to high to identify the lowest value (L) and the highest value
(H).

Age 19 21 26 29 31 33 36 37

Then subtract the lowest from the highest value.

R=H–L

R = 37 – 19 = 18

The range of our data set is 18 years.

1.3.6 Quartile Deviation

The quartile deviation or interquartile range is a more accurate indicator of variation within a

distribution than the range. Here, the center 50% of the distribution is used in order to avoid

using the 25% at either end of the distribution. The difference between the third and the first

quartile is represented by the interquartile range, to put it another way.

54
Symbolically, interquartile range = Q3- Q1

Many times the interquartile range is reduced in the form of semi-interquartile range or

quartile deviation as shown below:

Semi interquartile range or Quartile deviation = (Q3 – Q1)/2

When the quartile deviation is low, the items that make up the middle 50% of the distribution

have a low variance. If the quartile deviation is high, on the other hand, it means that the

items in the middle 50% of the distribution have a wide range. The two quartiles, Q3 and QI,

are equally spaced from the median in a symmetrical distribution; it should be observed.

Symbolically,

M-Q1 = Q3-M

The majority of business and economic statistics are asymmetrical; thus, this is rarely the

case. However, it is reasonable to suppose that the interquartile range contains about 50% of

the observations. It should be highlighted that the quartile deviation or the interquartile range

is an exact indicator of dispersion. It can be transformed into the following relative measure

of dispersion:

Q3 − Q1
𝐶𝑜𝑒𝑓𝑓𝑖𝑐𝑖𝑒𝑛𝑡 𝑜𝑓 𝑄𝐷 =
𝑄3 + 𝑄1

The computation of a quartile deviation is very simple, involving the computation of upper

and lower quartiles.

Example 1- Take the following data set into consideration: 22, 12, 14, 7, 18, 16, 11, 15,

12. You need to figure out the quartile deviation.

Solution-

55
First, we need to arrange data in ascending order to find Q3 and Q1 and avoid any duplicates.

7, 11, 12, 13, 14, 15, 16, 18, 22

Calculation of Q1 can be done as follows,

Q1 = ¼ (9 + 1)

=¼ (10)

Q1=2.5 Term

Calculation of Q3 can be done as follows,

Q3=¾ (9 + 1)

=¾ (10)

Q3= 7.5 Term

The quartile deviation can be calculated using the formulas below.

 The average for Q1 is 2nd, which is 11, plus the difference between 3rd and 4th,

which is 0.5, for a total of (12-11)*0.5 = 11.50.

 Q3 is the seventh term and product of 0.5; the eighth term's difference from the

seventh term is (18–16)*0.5, and the outcome is 16 + 1 = 17.

Q.D. = Q3 – Q1 / 2

Using the quartile deviation formula, we have (17-11.50) / 2

=5.5/2

Q.D.=2.75.

56
Example 2- The textile maker Harry Ltd. is developing a compensation plan. The

management is debating the launch of a new venture, but they want to first determine

the extent of their production spread.

The management has compiled its 10 most recent days' worth of average daily

production data for each (typical) employee.

155, 169, 188, 150, 177, 145, 140, 190, 175, 156.

Use the Quartile Deviation formula to help management find dispersion.

Solution- Here, there are 10 observations, therefore our initial step would be to arrange the

data in ascending order.

140, 145, 150, 155, 156, 169, 175, 177, 188, 190

The following formulas can be used to calculate Q1:

Q1= ¼ (n+1)th term

=¼ (10+1)

=¼ (11)

Q1= 2.75th Term

Calculation of Q3 can be done as follows,

Q3= ¾ (n+1)th term

=¾ (11)

Q3= 8.25 Term

The quartile deviation can be calculated using the formulas below.

57
 The second term is 145, and by adding 0.75 * (150 - 145) which is 3.75, we get

148.75.

 The eighth term is 177, thus by adding 0.25 * (188 - 177), which is 2.75, the answer is

179.75.

Q.D. = Q3 – Q1 / 2

Using the quartile deviation formula, we have (179.75-148.75) / 2

=31/2

Q.D.=15.50.

Example 3- The percentage score distribution of Ryan's International Academy's

students will be examined.

The data is for the 25 students.

58
Use the Quartile Deviation formula to find out the dispersion in % marks.

Solution- In this case, there are 25 observations, therefore, our first step would be to arrange

the data in ascending order.

59
Calculation of Q1 can be done as follows,

Q1= ¼ (n+1)th term

=¼ (25+1)

=¼ (26)

Q1= 6.5th Term

Calculation of Q3 can be done as follows,

60
Q3=¾ (n+1)th term

=¾ (26)

Q3 = 19.50 Term

Following are some examples of how to calculate quartile deviation or semi-interquartile

range:

 The outcome of adding 0.50 * (156 - 154) which is 1 to the sixth term, which is 154,

is 155.00.

 The outcome of adding 0.50 * (177 - 177) to the 19th term, which is 177, is 177.

Q.D. = Q3 – Q1 / 2

Using the quartile deviation formula, we have (177-155) / 2

=22/2

Q.D.= 11.

Uses of Quartile Deviation-

Semi-interquartile range, also known as quartile deviation. Again, the interquartile range

represents the difference in variance between the third and first quartiles. The interquartile

range shows how far apart from the mean or average the observations or values in the

provided dataset are. When attempting to understand or conduct a study on the dispersion of

the observations or samples from the given data sets that are found in the main or middle

body of the given series, the quartile deviation or semi-interquartile range is frequently used.

61
This situation would typically occur in a distribution where the data or the observations have

a tendency to lie heavily in the main body or middle of the given set of data, or the series, and

the distribution or the values do not lie towards the extremes, or if they do, they are not of

great significance for the calculation.

1.3.7 Mean Deviation

The average deviation is another name for the mean deviation. The absolute amounts by

which the various items depart from the mean are averaged, as the name suggests. We ignore

positive and negative signs when computing the mean deviation since positive and negative

deviations from the mean are equivalent. Symbolically,

𝛴|𝑥|
𝑀𝐷 =
𝑛

Where MD = mean deviation, |x| = deviation of an item from the mean ignoring positive and

negative signs, n = the total number of observations.

Example-

Advantages of Mean

Deviation-

62
1. The fact that mean deviation is based on all observations gives it an advantage as a

measure of dispersion. Contrast this with conventional metrics of dispersion like range and

quartile deviation, which do not account for all possible values.

2. The calculation is easy. This is so because the calculation only requires basic, simple

procedures like adding the absolute difference and average, then dividing the result by the

total number of observations.

3. The process of averaging the differences eliminates all outliers and paints a clear picture of

the level of data variability.

4. It measures dispersion more accurately than the standard deviation. This is so because the

standard deviation squares deviations rather than measuring their actual values. Additionally,

this increases the likelihood that the occurrence of extreme values will have an impact on

standard deviation.

5. The coefficient of mean deviation, a relative indicator of dispersion, can be computed

using it. When comparing two sets of data values, such relative metrics of dispersion are

useful.

Properties of Mean Deviation

1. When calculated about the median, mean deviation has the lowest value. As a result, the

mean deviation from the median is never greater than the mean deviation from the mean or

the mean deviation from the mode.

2. In the case of a symmetric distribution, the mean departure from the median generally

corresponds to 57.5% of the data values.

3. The change in scale has no impact on mean deviation. This indicates that the value of the

mean deviation remains the same if we add a single fixed value to all the observations.

63
4. On the other hand, a change in scale has an impact on mean deviation. The mean deviation

is multiplied by the same positive amount when we multiply each value by that number.

Uses of Mean Deviation

1. Because of its precision and ease of usage, economists frequently employ it.

2. When determining how much wealth is distributed within a society, it is helpful. This is so

that all values, including those of exceedingly wealthy or extremely impoverished people, are

taken into account.

3. Since it is the most accurate measure of variability for this usage, it is used to forecast

business cycles.

Limitations of Mean Deviations

1. It cannot be subjected to more algebraic analysis.

2. It might occasionally produce inaccurate results. When deviations are taken from the

median rather than the mean, the mean deviation produces the best results. However, median

is not a good indicator in a series with wide variability in the items.

3. The method is incorrect just mathematically since it ignores the algebraic signs when

deviating from the mean.

1.3.8 Standard Deviation

The standard deviation and mean deviation are similar in that they both quantify deviations

from the mean. However, due to its advantageous mathematical characteristics, the standard

64
deviation is favoured over the mean deviation, quartile deviation, and range. We introduce

another notion, variance, before defining the standard deviation.

Example-

Solution-
108
𝑀𝑒𝑎𝑛 = = 18
6

The second column shows the deviations from the mean. The third or the last column shows

the squared deviations, the sum of which is 70. The arithmetic mean of the squared deviations

is:

(𝑥 − 𝜇)2 70
∑ = = 11.67 𝑎𝑝𝑝𝑟𝑜𝑥
𝑁 6

The variance is the average of the squared deviations. It should be noted that many words that

are used interchangeably to express this variance include "the variance of the distribution X,"

"the variance of X," "the variance of the distribution," and "only the variance."

Symbolically,

(𝑥 − 𝜇)2
𝑉𝑎𝑟 𝑋 = ∑
𝑁

It is also written as

65
2
(𝑥𝑖 − 𝜇)2
𝜎 =∑
𝑁

Where 𝜎 2 (called sigma squared) is used to denote the variance. Although the variance is a

measure of dispersion, the unit of its measurement is (points). If a distribution relates to the

income of families, then the variance is (𝑅𝑠)2 and not rupees. Similarly, if another

distribution pertains to marks of students, then the unit of variance is (𝑚𝑎𝑟𝑘𝑠)2 . To address

this shortcoming, the standard deviation—a more accurate indicator of dispersion—is

produced by taking the square root of variance. Using the example of a single observation

from previously, we calculate the square root of the variance.

In applied Statistics, the standard deviation is more frequently used than the variance. This

can also be written as:

2 (∑𝑥𝑖 )2
∑𝑥
√ 𝑖 −
𝜎= 𝑁
𝑁

We use this formula to calculate the standard deviation from the individual observations

given earlier.

Uses of the Standard Deviation

The standard deviation is a frequently used measure of dispersion. It enables us to determine

as to how far individual items in a distribution deviate from its mean.

66
In a symmetrical, bell-shaped curve:

(i) About 68 percent of the values in the population fall within + 1 standard deviation from

the mean.

(ii) About 95 percent of the values will fall within +2 standard deviations from the mean.

(iii) About 99 percent of the values will fall within + 3 standard deviations from the mean.

Given that it represents the variance in the same units as the original data, the standard

deviation is an absolute measure of dispersion. As a result, it is inapplicable to compare two

or more distributions. We ought to employ a relative measure of dispersion for this objective.

The coefficient of variation, which connects the standard deviation and the mean such that

the standard deviation is reported as a percentage of the mean, is one such indicator of

relative dispersion. As a result, the measurement unit for the standard deviation is no longer

used, and the new unit is changed to percent.

𝜎
Symbolically, 𝐶𝑉(𝐶𝑜𝑒𝑓𝑓𝑖𝑐𝑖𝑒𝑛𝑡 𝑜𝑓 𝑉𝑎𝑟𝑖𝑎𝑡𝑖𝑜𝑛) = 𝜇 ∗ 100

Advantages of Mean Standard Deviation

1. The most used measure of dispersion, it can be used in a wide range of circumstances.

2. All observations are taken into account when calculating the standard deviation. Other

metrics of dispersion, such as range, on the other hand, are not dependent on all observations.

3. A fair estimate of the population standard deviation is provided by the sample standard

deviation. Since the standard deviation is unaffected by sample fluctuations, it can be used to

calculate the population's standard deviation.

4. The standard deviation is used to compute the data's skewness and kurtosis, which informs

us about the data's symmetry and shape.

67
5. Using the formula for the combined standard deviation, we may determine the standard

deviation of two data sets if we are given their individual standard deviations. Such formulas

do not provide combined values for the other dispersion measures.

Disadvantages of Standard Deviation

1. The square of the differences of the observations from the mean, rather than the actual

distance of each observation from the mean, is used to calculate the standard deviation.

2. Outliers will add a significant amount to the numerator when the differences are squared

because doing so makes large values even larger. This indicates that the standard deviation

gives extreme values more weight. The standard deviation is therefore susceptible to the

impact of outliers.

3. Since the method requires extracting square roots, performing it by hand would be highly

time-consuming. Calculators, however, may easily address this issue.

1.3.9 Coefficient of Variation

One form of measure of dispersion is the coefficient of variation. A number called a measure

of dispersion is used to assess the degree of data variability. As a result, the coefficient of

variation is used to assess how much data deviate from the mean or average value. The

acronym CV stands for coefficient of variation.

Coefficient of Variation formula

68
For calculating the coefficient of variation, there are two formulas. The sample coefficient of

variance and the population coefficient of variation are these. In statistics, the term

"population" refers to the entire group being studied. In other terms, the population refers to

the entire set of data. The sample is the particular portion of the population that has been

selected. The sample is utilised to reflect the study's overall population. There is no

difference between the sample mean and the population mean. There are two coefficient of

variation calculations, though, because the standard deviation values vary. Here are some of

them:
𝜎
 Population Coefficient of Variation = 𝜇 ∗ 100.

𝑠
 Sample Coefficent of Variation = 𝜇 ∗ 100

∑(𝑥𝑖 −𝜇)2
σ is the standard deviation of the population. It is given by 𝜎 = √ 𝑁

∑(𝑥𝑖 −𝜇)2
s is the standard deviation of the sample. It is given by 𝜎 = √ 𝑁−1

Example: Two plants C and D of a factory show the following results about the number of

workers and the wages paid to them.

No. of workers 5000 6000

69
Average monthly wages $2500 $2500

Standard deviation 9 10

Using coefficient of variation formulas, find in which plant, C or D is there greater variability

in individual wages.

Solution-

To Find: Which plant has greater variability? For this, we need to find the coefficient of

variation. The plant that has a higher coefficient of variation will have greater variability.

Coefficient of variation for plant C.

Using the coefficient of variation formula,

CV = (σ/μ) × 100, μ≠0

CV = (9/2500) × 100

CV = 0.36%

Now, CV for plant D

CV = (σ/μ) × 100

CV = (10/2500) × 100

CV = 0.4%

Plant C has CV = 0.36 and plant D has CV = 0.4

Answer: Hence plant D has greater variability in individual wages.

Explained and Unexplained Variation

As we evaluate the scatterplot's total fluctuation of y values, given by

70
It is readily evident that this "explained variation"

We should be aware that if there is a significant correlation between our x and y, the stronger

that correlation is, the more we can explain the "spread" of our y values by merely admitting

that they are near the best-fit line, which, due to the significant correlation, has a non-zero

slope and thereby "spreads the y-values out."

As a result, we can query what percentage of the total variation is still unknown and what

percentage could be described by the linear model itself.

To find a nice, tight expression for the "unexplained variation", consider the following:

Clearly, this is true:

Squaring both sides and summing over all i, we then have:

We claim the last term is zero. Here's why:

Remembering

71
The first term on the right is referred to as the unexplained variation since it is obvious that

the total variation should be the sum of the explained variation and the unexplained variance.

Measurement of Seasonal Variation

Seasonal Variations can be measured by the method of simple average. The data ought to be

accessible in seasonally appropriate weeks, months, and quarters.

Methods of measuring Seasonal Variations By Simple Averages :

Seasonal Variations can be measured by the method of simple average. The data should be

available in season wise likely weeks, months, quarters.

72
Method of Simple Averages:

This is the simplest and easiest method for studying Seasonal Variations. The procedure of

simple average method is outlined below.

Procedure:

(i) Arrange the data by months, quarters or years according to the data given.

(ii) Find the sum of the each months, quarters or year.

(iii) Find the average of each months, quarters or year.

(iv) Find the average of averages, and it is called Grand Average (G)

(v) Compute Seasonal Index for every season (i.e) months, quarters or year is given by

(vi) If the data is given in months

Similarly we can calculate SI for all other months.

(vii) If the data is given in quarter

73
1.3.10 Skewness and Kurtosis

Skewness
Depending on the model, skewness might lower the interpretation of feature relevance or

break model assumptions if the values of a particular independent variable (feature) are

skewed.

Skewness, which differs from the symmetrical normal distribution (bell curve), is a measure

of asymmetry found in a probability distribution that is observed in statistics.

To determine skewness, one can use the normal distribution. Data are symmetrically

distributed when we discuss the normal distribution. Due to the fact that all measurements

with a central tendency fall in the middle, the symmetrical distribution has no skewness.

74
The left and right sides of data that are symmetrically distributed have an equal number of

observations. (For example, if the dataset contains 90 values, the left side will have 45

observations, and the right side will contain 45 observations.) But what if the distribution is

not symmetrical? This type of data is referred to as asymmetrical data, and time skewness

enters the picture.

Types of skewness

1. Positive skewed or right-skewed

In statistics, a positively skewed distribution is a particular type of distribution where, in

contrast to symmetrically distributed data, where all measures of the central tendency (mean,

median, and mode) are equal to each other, with positively skewed data, the measures are

dispersing. Positively skewed distributions are thus types of distributions where the mean,

median, and mode of the distribution are positive rather than negative or zero.

Positively skewed statistics have a mean that is higher than their median (a large number of

data-pushed on the right-hand side). In other words, the outcomes are skewed to the negative.

Since the median is the middle value and the mean is always the highest value, the mean will

be greater than the median.

75
Extremely positive skewness is undesirable for distribution since it might lead to inaccurate

findings when present in high concentrations. The skewed data is being brought closer to a

normal distribution with the aid of data transformation technologies. The most well-known

transformation for favourably skewed distributions is the log transformation. The natural

logarithm of each value in the dataset is proposed by the log transformation.

2. Negative skewed or left-skewed

The direct opposite of a positively skewed distribution is a negatively skewed distribution. In

statistics, a negatively skewed distribution is a distribution model in which the majority of

data are plotted on the graph's right side while the distribution's tail spreads out to the left.

When data is negatively skewed, the mean is lower than the median (a large number of data-

pushed on the left-hand side). A distribution is said to be negatively skewed if the

distribution's mean, median, and mode are negative rather than positive or zero.

The Median is the middle value, and mode is the highest value, and due to an unbalanced

distribution median will be higher than the mean.

Calculate the skewness coefficient of the sample

Pearson’s first coefficient of skewness

76
Subtract a mode from a mean, then divides the difference by standard deviation.

The value is scaled down to a narrow range of -1 to +1 when we divide the covariance values

by the standard deviation because Pearson's correlation coefficient ranges from -1 (perfectly

negative linear relationship) to +1 (perfectly positive linear relationship), with a value of 0

denoting no linear relationship. That describes the correlation values' range quite accurately.

If the data show a high mode, Pearson's initial skewness coefficient is helpful. However,

Pearson's first coefficient is not favoured if the data contain a low mode or multiple modes; in

this case, Pearson's second coefficient may be preferable because it does not depend on the

mode.

Pearson’s second coefficient of skewness

Multiply the difference by 3, and divide the product by standard deviation.

The skewness is almost symmetrical if it is between -0.5 and 0.5.

The data are significantly skewed if the skewness is between -1 and -0.5 (negative skewed)

or between 0.5 and 1 (positive skewed).

77
The data are considered to be highly skewed if the skewness is less than -1 (negative

skewed) or higher than 1 (positive skewed).

Bowley's coefficient of skewness

The data's quartiles serve as the foundation for Bowley's coefficient of skewness. It is based

on the data set's middle 50% of observations. This indicates that each tail of the data set

contains 25% of the observations according to Bowley's coefficient of skewness.

In the case of a symmetric distribution, the Q1 and Q3 quartiles are equally spaced from the

median. Q2. Thus, Q3−Q2=Q2−Q1.

It is specified that the Bowley's coefficient of skewness is-

Steps to find Bowley's coefficient of skewness for grouped data-

Step 1 - Select type of frequency distribution (Discrete or continuous)

78
Step 2 - Enter the Range or classes (X) seperated by comma (,)

Step 3 - Enter the Frequencies (f) seperated by comma

Step 4 - Click on "Calculate" button for decile calculation

Step 5 - Gives output as number of observation (N)

Step 6 - Gives output as Q1, Q2 and Q3

Step 7 - Gives output as Bowley's Coefficient of Skewness

Kelly’s coefficient of skewness

Deciles or percentiles of the data serve as the foundation for Kelly's coefficient of skewness.

The middle 50% of the data set's observations form the basis of the Bowley's coefficient of

skewness. This indicates that the 25 percent of observations in each tail of the data set are left

by Bowley's coefficient of skewness.

Kelly proposed a skewness index based on the middle 80% of the data set's observations.

For a symmetric distribution, the first decile namely D1 and ninth decile D9 are equidistant

from the median i.e. D5. Thus, D9−D5=D5−D1.

The Kelley's coefficient of skewness based is defined as-

79
Interpretations-

Steps to find Kelly's coefficient of skewness for grouped data-

Step 1 - Select type of frequency distribution (Discrete or continuous)

Step 2 - Enter the Range or classes (X) seperated by comma (,)

Step 3 - Enter the Frequencies (f) seperated by comma

Step 4 - Click on "Calculate" button for decile calculation

Step 5 - Gives output as number of observation (N)

Step 6 - Gives output as D1, D5 and D9

Step 7 - Gives output as Kelly's Coefficient of Skewness

Kurtosis

80
Kurtosis can be used to detect outliers in our data. It provides the overall level of outliers that

are present.

The peak and tail of the data can be heavy-tailed and flat, resembling punching or squashing

the distribution. The term for this is negative kurtosis (Platykurtic). Positive Kurtosis is a

term used to describe a distribution that has a light tail and a steeper top curve (Leptokurtic).

Kurtosis is predicted to have a value of 3. In a symmetric distribution, this is seen. Positive

kurtosis is indicated by a kurtosis value larger than three. The range of kurtosis value in this

situation ranges from 1 to infinity. Additionally, a negative kurtosis will be indicated by a

kurtosis fewer than three. The range of values for a negative kurtosis is from -2 to infinity.

The greater the value of kurtosis, the higher the peak.

Excess Kurtosis

In probability and statistics, the excess kurtosis is used to compare the kurtosis coefficient to

the normal distribution. Leptokurtic distributions have positive excess kurtosis, platykurtic

distributions have negative excess kurtosis, and zero excess kurtosis (Mesokurtic

distribution). Since kurtosis in normal distributions is 3, excess kurtosis is calculated by

subtraction of kurtosis from 3.

Excess kurtosis = Kurt – 3

Types of excess kurtosis:

1. Leptokurtic or heavy-tailed distribution (kurtosis more than normal distribution)

2. Mesokurtic (kurtosis same as the normal distribution)

3. Platykurtic or short-tailed distribution (kurtosis less than normal distribution)

81
Leptokurtic (kurtosis > 3)

Leptokurtic has extremely long and thin tails, which increases the likelihood of outliers.

Positive values of kurtosis suggest a peaked distribution with fat tails. A distribution where

more of the numbers are distributed away from the mean and in the tails is said to have an

extreme positive kurtosis.

platykurtic (kurtosis < 3)

The majority of the data points are present and close to the mean because the platykurtic

distribution has a lower tail and elongated tails around the centre. Comparing a platykurtic

distribution to a normal distribution, the former is flatter (less peaked).

Mesokurtic (kurtosis = 3)

Mesokurtic is the same as the normal distribution, which means kurtosis is near to 0. In

Mesokurtic, distributions are moderate in breadth, and curves are a medium peaked height.

82
Examples 1- Suppose we have the following observations:

{12 13 54 56 25}

Determine the skewness of the data.

Solution

First, we must determine the sample mean and the sample standard deviation:

83
Skewness is positive. Hence, the data has a positively skewed distribution.

Example 2- Using the data from the example above (12 13 54 56 25), determine the

type of kurtosis present.

Solution-

84
The excess kurtosis is then obtained by taking 3 out of the sample kurtosis.

As a result, extra kurtosis = 0.7861 - 3 = 2.2139.

We have a platykurtic distribution since the excess kurtosis is negative.

Moments

When choosing the probability distribution we will use, statistical moments are essential

since they allow us to characterise the characteristics of the statistical distribution. They are

useful in describing the distribution as a result.

Moments are popularly used to describe the characteristic of a distribution. Let’s say the

random variable of our interest is X then, moments are defined as the X’s expected values.

85
For Example, E(X), E(X²), E(X³), E(X⁴),…, etc.

Four Statistical Moments

 The Expected Value or Mean

 Variance and Standard Deviation

 Skewness

 Kurtosis

All the above have been discussed in this module

Different Types of Moments

Raw Moments

The expected value of Xn is the raw moment, or the n-th moment, of a probability density

function f(x). The Crude moment is another name for it.

Centered Moments

The predicted value of a given integer power of the random variable's deviation from the

mean is what is known as the central moment, which is a moment of a probability distribution

of a random variable defined about the mean of the random variable.

Standardized Moments

A probability distribution's standardised moment is a higher degree central moment that is

often normalised by dividing the standard deviation, making the moment scale-invariant.

Sample Sampling and Sampling Techniques

A sample is the subset of the population. The process of selecting a sample is known as

samplingThe number of elements in the sample is the sample size

86
Sampling is a strategy for choosing specific individuals or a subset of the population in order

to draw conclusions from them statistically and estimate the characteristics of the entire

population. Researchers frequently utilise various sampling techniques in market research so

they do not have to study the full community in order to gather useful information.

It serves as the foundation of any research design because it is also a time- and money-

efficient strategy. For the best derivation, sampling techniques can be utilised in research

survey software.

For example, if a drug manufacturer would like to research the adverse side effects of a drug

on the country’s population, it is almost impossible to conduct a research study that involves

everyone. In this case, the researcher decides a sample of people from each demographic and

then researches them, giving him/her indicative feedback on the drug’s behavior.

Types of sampling: sampling methods

There are two forms of sampling used in market action research: probability sampling and

non-probability sampling. Let's examine these two sampling techniques in more detail.

Probability sampling: Probability sampling is a sampling approach in which a researcher

selects a few criteria and randomly selects individuals of a population. With the use of this

selection parameter, each member has an equal chance of being included in the sample.

87
 Simple Random Sampling

 Stratified sampling

 Systematic sampling

 Cluster Sampling

 Multi stage Sampling

Non-probability sampling: In non-probability sampling, participants are chosen at random

for the study. This type of sampling is not a set or predetermined selection procedure. Due of

this, it is challenging to ensure that every component of a population has an equal chance of

being represented in a sample.

 Convenience Sampling

 Purposive Sampling

 Quota Sampling

 Referral /Snowball Sampling

1.3.11 Case Study (Inferential Statistics)

This essay is an overview of research that Jared Wadley conducted on October 1st, 2008. The

study was conducted to look into the consequences of gun exhibitions, which were on the rise

88
not only in the USA but also internationally. In order to conduct the research, demographic

samples that were supposed to represent different regions of America were sampled.

The study's hypothesis were:

1. Do gun shows increase the incidence of killings using firearms?

2. Do gun shows in the USA result in an increase in suicidal cases?

3. Do strict laws result in fewer deaths from firearms?

These inquiries, in my opinion, were used by the researcher to direct him appropriately on the

type of sampling process, data collection, data interpretation, data analysis, and data

presentation (Maxim, L.W. 2010).

In order to conduct the research, Wadley combined both dependent and independent factors.

The variables that can be measured, have an impact on the research, and depend on the

independent variables are referred to in this context as the independent variables. The number

of deaths, suicides, and homicides were some of the dependent variables employed in this

study. He was able to deliver more accurate results thanks to this understanding. Dependent

variables, on the other hand, are variables that do not depend on one another and are

unaffected by study.

No matter what experimental conditions are applied, they don't change. The number of gun

exhibitions held during this time period is thus one of the independent variables in our study

(Maxim, L.W. 2010). Even if the number of deaths was increasing as a result of viewers'

exposure to these shows, the number of shows was unaffected by this unfavourable trend

since viewers never linked these deaths to these programmes, especially prior to the

completion of a scientific study.

The larger USA is where this study was conducted. The researcher had to develop a better

method of identifying and choosing the population to utilise as the participants of this

research because the USA is such a large and populous country. The sampling method was

89
acceptable because it was difficult to involve the full American population in the

investigation. These sample populations were chosen to represent the entire American

population.

Simple random, in which the population was deliberately chosen using random number

tables, and systematic sample, in which the population was chosen based on numerous

characteristics deemed pertinent to this research, were some of the sampling techniques used.

For instance, the two most populous states in the USA, Texas and California, provided the

greatest number of samples. In addition, the researcher used a stratified quota sampling

approach, in which larger populations were separated into smaller ones and later given

preferential treatment.

The population was split into men, women, children, and the elderly, according to this

statement. This is as a result of how each group interpreted the gun exhibitions. Gun

exhibitions are more popular among young people, while older people find them to be

terribly damaging and dishonourable. This indicates that more youth were expected to

participate in this study than in any other demographic segment of the American population.

However, many biases were found during the course of this research in the population that

was chosen, how the data were measured, and how the study's subject was handled.

As a result, when the researcher disregarded the sampling's general rules, selection biases

became evident.

There were underrepresented areas and overrepresented ones. Bias in the measurements

would result from this. To gather, analyse, interpret, and present the data for this study, a

variety of methods and approaches were used. These included surveys, interviews, in-person

observations, charts, and pie charts, to name a few.

90
Many ethical decisions had to be taken when doing this investigation. The study's participants

were not to be forced to provide information or see such "traumatising" programmes. Their

legal and fundamental human right is violated by this.

According to this study, there is no connection between watching these episodes and an

increase in suicide and homicide rates. Additionally, it was discovered that watching such

programmes does not necessarily change how youngsters behave because such movies

frequently include warnings like "do not try this at home!" Finally, it was discovered that the

strict restrictions' enforcement does not necessarily result in a sharp decline in the number of

gun exhibitions.

Finally, it may be assumed that this research was properly conducted precisely at the precise

moment that this kind of information needed to be made public. Even if several obstacles had

to be overcome, this research ended well.

Glossary

 Descriptive Statistics- A group of statistical techniques known as descriptive statistics are

used to quantitatively characterise or sum up a set of data. Comparable to inferential

statistics, which are more predictive in nature, descriptive statistics strive to summarise.

 Population- A population is a chosen individual or group that represents all of the

participants in a particular group of interest.

 Sample- An excerpt taken from a bigger population is known as a sample. The outcome is

referred to be a random sample if the drawing is carried out in a way that gives every

member of the population an equal probability of selection.

91
 Parameter- An value that is derived from a population is referred to as a parameter. This

figure would be a parameter if I had access to all of the data for all people on Earth and

calculated the mean age.

 Statistics- A statistic is a number produced from a sample. This value would be a statistic

if I determined the average age of a sample of humanity living on Earth (far more

realistic). Hence, statistics as a field of study.

 Generalizability- Generalizability is the ability to extrapolate findings from data obtained

from a sample and generalise them to the characteristics of the population as a whole.

This ability is not a given and greatly depends on the type of sample used, the size of the

sample, and a number of other elements.

 Distribution- The arranging of data by a single variable's values in ascending order, from

low to high, is known as a distribution. This configuration, as well as its features like

shape and dispersion, reveal details about the underlying sample.

 Mean- One of the three main measures of central tendency, along with median and mode,

that jointly assess a crucial and fundamental element of distribution is the mean. The

mean is a single, condensed numerical summary of a distribution and is the

straightforward arithmetic average of a distribution of variable values (or scores). The

mean is perhaps the statistic that researchers generally use the most. The sample mean is

written as x̄, while the population mean is written as.

 Median- The median is the score in a distribution that lies between the top and bottom

50% of scores and is located at the 50% percentile. The median can divide a set of

distribution scores in half and determine a distribution's skew.

 Mode- The score that appears in the distribution the most frequently is the mode. A

distribution with more than one mode is said to be multimodal, and one with two modes

is referred to as bimodal.

92
 Skew- Skew occurs when there are more scores at one end of the distribution than the

other. The situation is known as negative skew when a distribution's scores are more

concentrated at the high end, and a tail is produced by the relatively smaller number of

low-end values. When a distribution has a tail at the high end, it has positive skew.

 In general, we would anticipate that the mean in a negatively skewed distribution would

be lower than the median and that the mean in a positively skewed distribution would be

higher than the median.

 Range- The range, one of the most crucial measurements of dispersion, is the distinction

between a distribution's highest and minimum values.

 Variance- The statistical average of the score distribution's dispersion is known as a

variance. Despite not being used frequently on its own, variance can be a helpful

calculation when heading toward a more descriptive statistical measurement like standard

deviation.

 Standard Deviation- The average difference between each distribution score and its mean

is what is known as the distribution's standard deviation. The standard deviation gives a

decent indication of how dispersed a disquisition's scores are on an individual basis.

These 2 measurements, coupled with the mean, give a clear picture of how the scores are

distributed.

 Interquartile Range-The IQR is the difference between the score defining the third and

first quartiles, or the 75th percentile and 25th percentile, respectively.

Key Terms- Arithmetic Mean, Geometric Mean, Mean, Median, Mode, Statistics, Kurtosis,

Skewness, Moments, Standard Deviation, Range.

93
Summary- This unit deals with descriptive and inferential statistical methods like mean,

median, mode, standard deviation, skewness, various methods of skewness, and kurtosis,

sample, sampling, and sampling techniques.

Important Formula-

94
Range=Maximumvalue–Minimumvalue

95
Exercise: Choose the correct option

1. The first three moments of a distribution about the mean

are 1, 4, and 0. The distribution is

A. Skewed to the left

B. Normal

C. Skewed to the right

D. Symmetrical

2. The distribution is positively skewed if

A. AM

B. AM > Median

C. AM > Mode

D. Both ÄM > Mode’ and ‘AM > Median’

3. In symmetrical distribution if Q1=4, Q3=12, then median is

A. 0

B. 4

C. 8

D. 6

4. The degree to which numerical data tend to spread out about an average value is

called

A. Variation

96
B. Skewness

C. Flatness

D. Constant

5. When a distribution is symmetrical and has one mode, the highest point on the curve

is called the

A. All of the options

B. Median

C. Mean

D. Mode

6. If the moment Ratio β2= 3 then the distribution is

A. Platykurtic

B. Mesokurtic

C. Positively skewed

D. Symmetrical

7. In Symmetrical distribution Q3-Q1=20 , Median = 15, Q3 is equal to

A. 15

B. 20

C. 25

D. 5

8. The degree of peakedness is called-

A. Dispersion

97
B. Skewness

C. Symmetry

D. Kurtosis

9. The coefficient of skewness is always zero for

A. Symetrical

B. Skewed

C. Both A and B

D. None of these

10. In Uni model distribution if mode is less than mean, then skewness will be

A. Symmetrical

B. Normal

C. Positively Skewed

D. Negatively Skewed

Short Answer Type Questions-

1. What is skewness?

2. What is the significance of skewness?

3. What is kurtosis?

Long Answer Type Questions-

1. Write down the sampling techniques in detail.

2. Write down the meaning and limitations of Mean deviation.

98
3. Explain Standard deviation and Kurtosis in detail?

Answers

Exercise: MCQ

1. (D)

2. (D)

3. (C)

4. (B)

5. (D)

6. (C)

7. (B)

8. (D)

9. (D)

10. (C)

99
Module 2
Learning Outcomes

At the end of this unit, you will be able to

 Discuss the different methods of correlation analysis

 Explain the difference between correlation & regression analysis

 Discuss the role of time series in business

2.1 Correlation Analysis

Introduction

For the comparison and analysis of distributions containing only one variable, or univariate

distributions, statistical methods of measures of central tendency, dispersion, skewness, and

kurtosis are useful. However, another crucial component of statistics is expressing the

relationship between two or more variables.

Understanding the correlations between two or more variables is essential for decision-

making in many business research scenarios. For instance, knowing if the interest rate on

bonds is linked to the prime interest rate might assist a broker forecast how the bond market

would behave. An account executive may find it useful to know whether there is a significant

correlation between advertising spending and sales spending for a company while researching

the impact of advertising on sales.

The statistical techniques of correlation and regression aid in understanding the relationship

between two or more variables that may be connected in a similar manner, such as the bond

interest rate and prime interest rate, advertising expenditure and sales, income and

consumption, crop yield and fertiliser use, height and weight, and so forth.

 In all these cases involving two or more variables, we may be interested in seeing:

100
 if there is any association between the variables;

 if there is an association, is it strong enough to be useful;

 if so, what form the relationship between the two variables takes;

 how we can make use of that relationship for predictive purposes, that is, forecasting;

and

 how good such predictions will be.

Correlation analysis and regression analysis are two approaches to examining the relationship

between two or more variables because these issues are interconnected. Let's say there is a

correlation between two or more variables. In this situation, regression analysis can be used

to estimate the inaccuracy of estimations and predict the value of the other variable(s) using

data from one (or more) variable(s).

What is Correlation?

A measure of association between two or more variables is a correlation. When movement in

one variable tends to be accompanied by matching signals in the other(s), then two or more

variables are said to be highly linked. It has a range of -1 to +1. It aids in comprehending the

degree and direction of the link between two variables.

“The nature and strength of the link between the variables are measured by the correlation

between them.”

When the goal of exploratory research is to find variables that might be related in some

manner to the variable of interest, correlation is frequently utilised as a measure of the degree

of relatedness of two variables.

2.1.1 Significance of Correlation

101
The notion of variables and the distinction between dependent and independent variables

make it evident that variables may be related to one another. For instance, demand and supply

are tied to commodity price, and agricultural output is influenced by rainfall, student grades

are influenced by study time, quantity requested may be influenced by advertisement

spending, consumption is influenced by income, and so on.

A measure of the type of link between two or more variables is called correlation. It

fluctuates from -1 to +1. It aids in comprehending the degree and direction of the link

between two variables.

In the words of Croxton and Cowden, “When the relationship is of a quantitative nature,

the appropriate statistical tool for discovering and measuring the relationship and

expressing it in brief formula is known as correlation.”

 Correlation measures the strength of the relationship between two or more variables. For

example, the relationship between income and consumption expenditure, price and

quantity demanded etc.

 When the relationship between variables is known, it is easy to predict the value of one

variable when the other variable is known.

 It helps understand the behaviour of various economic variables like demand, supply,

GDP, interest, money supply, inflation, income and expenditure.

2.1.2 Types of Correlation Analysis

Correlation can be classified in several ways. The important ways of classifying correlation

are:

(i) Positive and negative,

102
(ii) Linear and non-linear (curvilinear) and

(iii) Simple, partial and multiple.

Positive and Negative Correlation- Positive correlation is defined as the movement of both

variables in the same direction, i.e., when one variable rises, the other variable rises on

average as well, or when one variable falls, the other variable falls on average. Conversely, if

the variables are moving in the opposite way, we say that there is a negative correlation, as in

the case of supply and demand fluctuations.

Linear and Non-linear (Curvilinear)- Correlation A case of linear correlation occurs when

changes in one variable are accompanied by similar changes in the other variable in a fixed

ratio. Look at the information below:

X: 10 20 30 40 50

Y: 25 50 75 100 125

In the aforementioned case, the ratio of change is the same. Therefore, it is an instance of

linear correlation. All of the points will lie on a single straight line if we plot these variables

on graph paper. On the other hand, non-linear or curvilinear correlation occurs when the

amount of change in one variable does not follow a constant ratio with the change in another

variable. A non-linear connection would result from changing a few numbers in series X or

series Y.

Simple, Partial, and Multiple Correlation- The number of variables included in a study

determines how these three types of correlation are distinguished from one another. A

correlation is referred to as a simple correlation if there are just two variables present in the

investigation. It is a problem of partial or multiple correlations when three or more variables

are present in a study. Three or more variables are analysed simultaneously in multiple

103
correlations. However, with partial correlation, we just take into account the interaction of

two variables while holding the impact of the remaining variable(s) constant.

Consider a situation where there are three variables: X, Y, and Z. X represents the number of

hours studied, Y represents intelligence, and Z represents the total number of exam points

earned. We will investigate the relationships between the marks attained (Z) and the two

variables, the quantity of study time (X) and I.Q. (Y). In contrast, a study utilising partial

correlation is one in which the link between X and Z is examined while maintaining a

constant average I.Q. (Y).

2.1.3 Methods of Correlation Analysis

The commonly used methods for studying linear relationships between two variables involve

both graphic and algebraic methods. Some of the widely used methods include:

1. Scatter Diagram

2. Correlation Graph

3. Pearson’s Coefficient of Correlation

4. Spearman’s Rank Correlation

5. Concurrent Deviation Method

Scatter Diagram

The Dotogram or Dot Diagram is another name for this approach. One of the simplest ways

to diagrammatically portray a bivariate distribution is with a scatter diagram. In this

approach, dots are used to plot both variables on graph paper. "Scatter Diagram" is the name

given to the resulting diagram. By looking at the diagram, we may get a general notion of the

type and strength of the link between the two variables. The spreading of dots across the

104
graph is referred to as scatter. When analysing correlation, the following things should be

kept in mind:

 If the plotted points are very close to each other, it indicates high degree of correlation. If

the plotted points are away from each other, it indicates low degree of correlation.

 If the points on the diagram reveal any trend (either upward or downward), the variables

are said to be correlated, and if no trend is revealed, the variables are uncorrelated.

 The correlation is positive if there is an upward movement from the lower left hand
corner to the higher right-hand corner, showing that the values of the two variables move

in the same direction. In contrast, if the points show a downward trend from the upper left

to the lower right, the correlation is negative since the values of the two variables in this

scenario travel in opposite directions.

 in particular, if all the points lie on a straight line starting from the left bottom and going
up towards the right top, the correlation is perfect and positive, and if all the points like

105
on a straight line starting from the left top and coming down to the right bottom, the

correlation is perfect and negative.

Example-

Given the following data on sales (in thousand units) and expenses (in thousand rupees) of a

firm for 10 month:

Months: J F M A M J J A S O

Sales: 50 50 55 60 62 65 68 60 60 50

Expenses:11 13 14 16 16 15 15 14 13 13

a) Make a Scatter Diagram

b) Do you think that there is a correlation between sales and expenses of the firm? Is it

positive or negative? Is it high or low?

Solution-

(a) The Scatter Diagram of the given data is shown below

(b) Figure shows that the plotted points are close to each other and reveal an upward trend. So

there is a high degree of positive correlation between sales and expenses of the firm.

106
Graphical Method

This Correlogram approach is quite straightforward. Two series' worth of data are plotted on

a graph sheet. By comparing the direction and proximity of two curves, we may determine

the correlation. A positive correlation is present when both of the curves on the graph are

going in the same direction. On the other hand, correlation is considered to be negative if

both curves are travelling in the opposite direction. A lack of connection can be seen in a

graph if there is no clear pattern due to irregular swings in the curves.

Example- Find out graphically, if there is any correlation between price yield per plot (qtls);

denoted by Y and quantity of fertilizer used (kg); denote by X.

Plot no.: 1 2 3 4 5 6 7 8 9 10

X: 3.5 4.3 5.2 5.8 6.4 7.3 7.2 7.5 7.8 8.3

Y: 6 8 9 12 10 15 17 20 18 24

Solution- The Correlogram of the given data is shown

107
The above graph demonstrates that the two curves move in the same direction and are also

quite near to one another, indicating that the price yield per plot (qtls) and amount of fertiliser

used to have a close relationship (kg).

Remark: By visualising the relationship between the variables, the scatter diagram and

correlation graph, two graphic tools, provide the reader a sense of the data. These are easily

understood and help us develop a reasonable, if rough, understanding of the type and strength

of the link between the two variables. These techniques, however, are unable to measure the

connection between them. With the use of algebraic techniques, which compute the

correlation coefficient, we may measure the degree of correlation.

2.1.4 Karl Pearson’s correlation co-efficient

The coefficient of correlation rxy between two variables x and y, for the bivariate dataset (xi,yi)

where i = 1,2,3…..N; is given by –

r(x,y)=cov(x,y)/σxσy

where,

⇒ cov(x,y): the covariance between x and y

Here, 𝑥̅ and 𝑦̅ are simply the respective means of the distributions of x and y.

⇒ σx and σy are the standard deviations of the distributions x and y.

Alternate Formula

108
You can apply the following formulas if any data is presented as a class-distributed frequency

distribution:

⇒ cov(x,y): the covariance between x and y

where,

xi: The central value of the i’th class of x

yj: The central value of the j’th class of y

fio,fij: Marginal Frequencies of x and y

fij: Frequency of the (i,j)th cell

In any case, the following equality must always hold:

Total frequency = N = Σi,jfij = Σifio = Σjfjo

A single formula for discrete datasets

109
2.1.5 Properties of the Pearson’s Correlation Coefficient

⇒ r is unit-less. As a result, it can also be used to compare associations between completely

different bivariate distributions. For instance, you can examine the Pearson's correlation

coefficients from both scenarios to see how much of your decision to skip the movie is

connected to your friends not joining you and to your own lack of interest in the film. This

parameter is consequently of the highest importance in determining the relationship between

distinct numbers in economics, where the cost price or the market shares depend on a variety of

different elements.

⇒ The value of r always lies between +1 and -1. Depending on its exact value, we see the

following degrees of association between the variables-

r value variation:

110
A number higher than 0 denotes a positive connection, meaning that as one variable's value

rises, the value of the other variable also rises. If the value of one variable is greater than 0, the

value of the other variable is lower, indicating a negative connection.

⇒ The Pearson product-moment correlation does not take into consideration whether a variable

has been classified as a dependent or independent variable. It treats all variables equally.

⇒ A change of origin of the system, or any scaling of the variables doesn’t affect the value

of r. The sign might change depending on the sign of scaling done.

2.1.6 Rank Correlation Coefficient

A nonparametric measurement of the strength and direction of the link between two ranking

variables is the Spearman's correlation coefficient, denoted by or rR. It establishes the

monotonicity of a connection, i.e., whether the link between two continuous or ordered variables

contains a monotonic component.

Comparatively to a linear relationship, monotonicity is "less restricting." Spearman's correlation

does not actually require monotonicity, but if we already know there isn't a monotonic

relationship between the two variables, it makes no sense to use Spearman's correlation to find

the strength and direction of that relationship.

On the other hand, one would use a Pearson's correlation to determine the strength and direction

of any linear relationship if, for instance, the relationship appears linear (as shown by a

scatterplot). Monotony –

111
Spearman Ranking of the Data

Before performing the Spearman's Rank Correlation analysis, we must rank the data under

consideration. This is essential because we must compare if, when one variable is increased, the

other follows a monotonic relation (regularly increases or decreases) with respect to it.

We must therefore compare the values of the two variables at each level. Such "levels" are

assigned to each value in the dataset by the ranking algorithm so that we can quickly compare

them.

 Assign number 1 to n (the number of data points) corresponding to the variable values in

the order highest to lowest.

 In the case of two or more values being identical, assign to them the arithmetic mean of the

ranks that they would have otherwise occupied.

Examples of values for the selling price are: 28.2, 32.8, 19.4, 22.5, 20.0, and 22.5 These are the

matching ranks: 2, 1, 5, 3.5, 4, 3.5 The highest value (32.8) is ranked first, and 28.2 is ranked

second. When two numbers (22.5) are similar, it is necessary to calculate the arithmetic mean of

the ranks they would have otherwise occupied (3+4/2).

112
Spearman Rank Correlation formula-

𝟔𝜮𝒊 ⅆ𝟐𝒊
𝒓𝑹 = 𝟏 −
𝒏(𝒏𝟐 − 𝟏)

where n is the number of data points of the two variables and di is the difference in the ranks of

the ith element of each random variable considered. The Spearman correlation coefficient, ρ, can

take values from +1 to -1.

 A ρ of +1 indicates a perfect association of ranks

 A ρ of zero indicates no association between ranks and

 Ρ of -1 indicates a perfect negative association of ranks.

The closer ρ is to zero, the weaker the association between the ranks.

2.1.7 Partial Correlation

Correlation and partial correlation are related ideas. It demonstrates the concept that just

because two variables show a correlation doesn't mean that they are causally related. When

two variables are conditional on one or more other factors, partial correlation measures the

correlation between the two variables. This suggests that when there is a connection between

two variables, the confounder (or controlling variable), a common cause of the misleading

association, may contribute to the explanation of the correlation. This portion is removed,

leaving only the partial correlation between the two variables.

To gauge the strength of the association between two variables while taking into

consideration the influence of one or more additional factors, partial correlation is used. Your

study's key variables should be continuous, regularly distributed, logically connected, and

113
devoid of outliers. Your variables should also have a comparable dispersion across each of

their respective ranges.

The partial correlation formula between random variables X and Y with Z factored out is

given by when there is only one confounding factor (then it is referred to as a first-order

partial correlation).

Assumptions for Partial Correlation

Assumptions are a part of all statistical methods. Your data must meet specific criteria in

order for the results of a statistical method to be accurate, which is what assumptions mean.

The assumptions for Pearson Correlation include:

 Continuous

 Normally Distributed

 Linearity

 No Outliers

 Similar Spread Across Range

 Covariate(s)

Let’s dive in to each one of these separately.

Continuous

114
You must care about a continuous variable. Continuous refers to a variable that can have any

conceivable value. Age, height, weight, test scores, survey results, yearly salary, etc. are a

few excellent examples of continuous variables.

Normally Distributed

The variable that matters to you ought to be dispersed normally. This is referred to as being

regularly distributed in statistics (aka it must look like a bell curve when you graph the data).

Only if the variable you care about is normally distributed should you conduct an

independent samples t-test on your data.

115
Linearity

The factors that matter to you must be correlated linearly. This means that if the variables are

plotted, a straight line can be drawn to represent the data's shape.

No Outliers

There must be no outliers in the variables that you care about. When it comes to outliers, or

data points with extremely high or low values, Pearson's correlation is sensitive. When you

plot your variables, look for any points that are far from the other points to determine if there

are any outliers.

Similar Spread Across Range

Making ensuring the variables have a similar dispersion across their ranges is known as

homoscedasticity in statistics.

Covariate(s)

If you have one or more covariates, you should only do partial correlation. When analysing

the variable connection of interest, a covariate is a variable whose effects you want to

eliminate. For instance, you might want to take education level into account while evaluating

the link between age and memory function. In this manner, you may be certain that the results

aren't influenced by education level.

Partial Correlation Example

116
Variable 1: Height

Variable 2: Weight

Covariate: Age

In this illustration, we are interested in the correlation between height and weight while

taking age into consideration. As a result, we start by gathering data on a set of people's

height, weight, and ages.

First, we make sure the variables that are important to us conform to the partial correlation's

presumptions. We proceed with the study after establishing that height and weight are

normally distributed, devoid of outliers, scattered similarly over their ranges, and linearly

connected (see above for details).

The analysis will yield a p-value and a correlation coefficient, or "r." R values are between -1

and 1. When r is negative, the variables are said to be inversely connected (i.e. when one

variable increases, the other decreases).

Positive numbers, on the other hand, show that as one variable rise, the other rises as well.

When age effects are taken into account, the p-value shows the likelihood that our results

would not have been seen if there was no real association between height and weight. If the

p-value is less than or equal to 0.05, our result is considered statistically significant, and we

may be confident that the observed difference is not the result of random chance.

Problems related to Partial Correlation-

 The gross or zero order correlation must have linear regressions

117
 The effect of the independent variables must be additively and not jointly related.

Key Terms- Significance, Correlation Analysis, Correlation coefficient, Partial correlation,

statistics, r value, variables.

Summary- This unit deals with data sources, data collections, data tables, correlation, types

of correlation, partial correlation, mean, median, mode, covariates, etc.

Case Study:

Certain rocket motors are created by fusing two types of fuel, an igniter and a sustainer.

The goal of this research is to look at the relationship between the strength of this

binding and the age of the propellant.

Empirical evidence implies that this relationship is linear, and the major goal of this

study is to evaluate this claim and create the best model feasible, based on a data set of

20 observations from different planes.

The obtained data set is displayed below:

Observation Strength Age

118
1 2158.7 15.5

2 1678.15 23.75

3 2316 8

4 2061.3 17

5 2207.5 5

6 1708.3 19

7 1784.7 24

8 2575 2.5

9 2357.9 7.5

10 2277.7 11

11 2165.2 13

12 2399.55 3.75

13 1779.8 25

14 2336.75 9.75

15 1765.3 22

16 2053.5 18

17 2414.4 6

18 2200 12.5

19 2654.2 2

20 1753.7 21.5

119
Procedures for Analyzing:

The first section of the analysis attempts to determine whether or not there is a

significant linear association between the variables Strength and Age. A scatterplot will

be created for this purpose, and the correlation coefficient will be calculated. If the

results of these techniques indicate that a linear relationship is plausible, a full

regression model will be computed.

FOUND RESULTS

First, a scatterplot is displayed:

120
The data clearly shows a negative linear trend. A linear regression analysis of the data

makes logical.

The following is a correlation:

The correlation is R = -0.947, which suggests that there is a considerable degree of

negative linear relationship between Strength and Age. Also from the table, it can be

determined that this association is significant at the 0.01 significance level.

The factor of determination is

This suggests that age explains 89.61% of the variation in Strength.

Questions for discussion-

(a) Calculate the regression coefficients, then form a regression equation.

(b) Draw a histogram and scatterplot diagram.

121
Answers-

 Strength and age were found to have a significant association Based on the

scatterplot and the correlation, there is a clear negative linear link between them.

 The linear regression equation found is

Strength=2625.355−36.961Age

 The data do not appear to contradict any of the regression assumptions.

Exercise: Choose the correct option

1. Which of the following are types of correlation?

A. Positive and Negative

B. Simple, Partial and Multiple

C. Linear and Nonlinear

D. All of the above

2. Which of the following is true for the coefficient of correlation?

A. The coefficient of correlation is not dependent on the change of scale

B. The coefficient of correlation is not dependent on the change of origin

122
C. The coefficient of correlation is not dependent on both the change of scale and change

of origin

D. None of the above

3 Which of the following statements is true for correlation analysis?

A. It is a bivariate analysis

B. It is a multivariate analysis

C. It is a univariate analysis

D. Both a and c

4. If the values of two variables move in the same direction, ___________

A. The correlation is said to be non-linear

B. The correlation is said to be linear

C. The correlation is said to be negative

D. The correlation is said to be positive

5. If the values of two variables move in the opposite direction, ___________

A. The correlation is said to be linear

B. The correlation is said to be non-linear

C. The correlation is said to be positive

D. The correlation is said to be negative

6. Which of the following techniques is an analysis of the relationship between two

variables to help provide the prediction mechanism?

A. Standard error

123
B. Correlation

C. Regression

D. None of the above

7. Which one of the following statements about the correlation coefficient is correct?

A. The correlation coefficient is unaffected by scale changes.

B. Both the change of scale and the change of origin have no effect on the correlation

coefficient.

C. The correlation coefficient is unaffected by the change of origin.

D. The correlation coefficient is affected by changes of origin and scale.

8. Choose the correct option concerning the correlation analysis between 2 sets of data.

A. Multiple correlations is a correlational analysis comparing two sets of data.

B. A partial correlation is a correlational analysis comparing two sets of data.

C. A simple correlation is a correlational analysis comparing two sets of data.

D. None of the preceding

9. The correlation coefficient is?

A. The square of the coefficient of determination

B. Can never be negative

C. The square root of the coefficient of determination.

D. The same as r square

10. The correlation for the values of two variables moving in the same direction is

A. Perfect positive

124
B. Negative

C. Positive

D. No correlation.

Short Answer Type Questions-

1. What is correlation?

2. What down the types of correaltion?

3. What is the significance of correlation?

Long Answer Type Questions-

1. Explain Scatter diagram in detail.

2. Explain Spearman’s Rank Coeffcient.

3. Explain Partial correlation in depth

Answers

Exercise: MCQ

1. (D)

2. (C)

3. (C)

125
4. (D)

5. (D)

6. (C)

7. (A)

8. (A)

9. (C)

10. (D)

2.2 Regression Analysis

In the corporate world, it frequently becomes important to make a forecast in order for

management to decide on a product or a specific course of action. A forecast requires the

identification of a relationship between two or more important variables in a unique

circumstance.

For instance, a business is curious about how much the demand for In the coming five years,

the number of televisions will rise while taking population growth into account in a specific

city. Here, it is assumed categorically that an increase in population will result in a growing

interest in televisions. Consequently, to ascertain the type and degree of association. The

relationship between these two factors becomes crucial for the business.

2.2.1 Introduction

126
A study on heredity titled "Natural Inheritance" was published in 1889 by Sir Francis Galton,

a cousin of Charles Darwin. He shared his discovery that sweet pea plant seed sizes appeared

to "revert" or "regress" to the mean size over time. He also shared the findings of a study on

the correlation between dads' heights and their sons' heights. Height of father versus height of

son data pairs were fitted with a straight line. He discovered a "regression to mediocrity" here

as well. The sons' heights showed a shift away from their fathers and toward the norm in

height.

The concept of statistical regression is attributed to Sir Galton. The name "regression" still

exists even if the majority of applications of regression analysis may have little in common

with Galton's "regression to the mean." Today, it describes the statistical method of

simulating the interaction between two or more variables. Regression analysis, in its broadest

meaning, refers to the estimation or prediction of an unknown value for one variable using

known values for the other variable (s). It is one of the most significant and often applied

statistical approaches in nearly all natural, social, and physical disciplines. We will just

discuss simple regression in this course, which is linear regression with just two variables: a

dependent variable and an independent variable.

Multiple regressions are regression analyses that look at more than two variables at once.

Independent and Dependent Variables

Simple regression involves only two variables; one variable is predicted by another variable.

Only two variables are used in simple regression, where one variable predicts the other. The

dependent variable is the one that needs to be predicted. The independent variable, also

referred to as the explanatory variable, is the predictor. For instance, when attempting to

forecast television set demand based on population growth, the demand for television sets is

127
used as the dependent variable and population growth is used as the independent or predictor

variable.

Choosing which variable is which can occasionally lead to issues. Often, the decision is clear-

cut, as in the case of population increase and TV demand, as it would be absurd to assume

that the latter might be influenced by the former. The dependent variable must be TV

demand, and the independent variable must be population growth.

Linear Regression

Creating techniques for fitting a straight line, or a regression line as it is frequently known, to

data on two variables is the process of revealing a linear relationship. The best estimate of

one variable for any given value of the other variable is represented graphically or as a

relationship by the line of regression. The independent and dependent factors affect how the

line is referred to. The Regression line of Y on X is a line that provides the best estimate of Y

for any value of X when X and Y are two variables whose relationship has to be shown.

If the dependent variable changes to X, then the best estimate of X by any value of Y is called

Regression line of X on Y.

Type of Regression Analysis

1. Linear Regression

2. Logistic Regression

3. Polynomial Regression

4. Ridge Regression

5. Lasso Regression

6. Quantile Regression

7. Bayesian Linear Regression

8. Principal Components Regression

128
9. Partial Least Squares Regression

10. Elastic Net Regression

1. Linear Regression

The modelling method that is most frequently employed assumes a linear relationship

between an independent variable (V) and a dependent variable (Y) (X). It uses a best-fit line,

commonly referred to as a regression line. The equation for the linear relationship is Y =

c+m*X + e, where c stands for the intercept, m for the slope, and e for the error term.

The simple and complex versions of the linear regression model, respectively, have different

numbers of dependent and independent variables (with one dependent variable and more than

one independent variable).

2. Logistic Regression

The logistic regression method is appropriate when the dependent variable is discrete. In

other words, this method is used to determine the likelihood of events that are mutually

129
exclusive, such as pass/fail, true/false, 0/1, and so on. Thus, the probability has a value

between 0 and 1, the target variable has a range of two possible values, and its relationship to

the independent variable is depicted by a sigmoid curve.

3. Polynomial Regression

In order to depict a non-linear relationship between dependent and independent variables,

polynomial regression analysis is performed. The best fit line is curved instead of straight in

this variation of the multiple linear regression model.

4. Ridge Regression

130
The ridge regression technique is used when the independent variables are highly correlated

and the data shows multicollinearity. Even if least squares estimates are impartial in

multicollinearity, their variances are high enough to induce a difference between the observed

value and the true value. By inflating the regression estimates, ridge regression lowers

standard errors.

The multicollinearity issue in the ridge regression equation is solved by the lambda (λ)

variable.

5. Lasso Regression

The lasso (Least Absolute Shrinkage and Selection Operator) method penalises the absolute

magnitude of the regression coefficient, just like ridge regression does. The lasso regression

method also uses variable selection, which causes the coefficient values to zero off

completely.

131
6. Quantile Regression

A part of the linear regression method is the quantile regression methodology. When the

conditions for linear regression are not met or when there are outliers in the data, it is used.

Quantile regression is used in statistics and econometrics.

132
7. Bayesian Linear Regression

The Bayes theorem is utilised in Bayesian linear regression, a type of regression analysis

method used in machine learning to determine the values of the regression coefficients. This

method calculates the posterior distribution of the features rather than the least-squares. As a

result, the method performs better in terms of stability than standard linear regression.

8. Principal Components Regression

The principle components regression approach is frequently used to evaluate multicollinear

regression data. By biassing the regression estimates, the significant components regression

approach, like ridge regression, lowers standard errors. The training data are first modified

using principal component analysis (PCA), and the changed samples are then utilised to train

the regressors.

9. Partial Least Squares Regression

133
A quick and effective method for covariance-based regression analysis is partial least squares

regression. It is beneficial for regression issues where there is a high likelihood of

multicollinearity between the variables. Regression is used once the procedure reduces the

number of variables to a reasonable number of predictors.

10. Elastic Net Regression

When working with highly correlated data, elastic net regression combines the ridge and

lasso regression techniques. By leveraging the penalties connected to the ridge and lasso

regression techniques, it regularises regression models.

Degree of Correlation

134
The coefficient of correlation measures the strength of the association between two variables.

These are what they are:

1. Perfect correlation: Two variables have a perfect correlation if they vary in the same

proportion (increase or decrease). An ideal correlation in this case could be either a positive

or negative correlation.

Coefficient of correlation (r) = 1: If there is a perfect positive association between two

variables, then the coefficient of correlation (r) will equal 1.

Coefficient of correlation (r) = −1: If there is a perfect negative relationship between two

variables, the coefficient of correlation (r) will equal one.

2. Zero correlation: The correlation between two variables is zero if there is no relationship

between them. It suggests that a change in one variable's value has no bearing on the change

in the other variable's value.

Correlation coefficient (r) = 0: The value of correlation will be zero if there is no correlation

between the two variables. It does not necessarily follow that these two factors are

independent, though. It merely shows that the two variables don't have a linear connection.

3. Limited degree of correlation: Between perfect correlation and zero correlation, there is a

limited degree of correlation, meaning that the coefficient of correlation is between +1 and 1.

This comparatively low correlation strength could be strong, medium, or low.

High degree of correlation: The correlation between two data series is close to one. The

correlation between two sets of data is neither very high nor very low. Low degree of

correlation: There is little correlation between the two data series.

Regression Equation formula for Grouped Data

135
1. Regression coefficient y on x
byx=n⋅∑fdxdy-∑fdx⋅∑fdyn⋅∑fdx2-(∑fdx)2⋅hyhx
2. Regression coefficient x on y
bxy=n⋅∑fdxdy-∑fdx⋅∑fdyn⋅∑fdy2-(∑fdy)2⋅hxhy
3. Regression Line y on x
y-ˉy=byx(x-ˉx)
4. Regression Line x on y
x-ˉx=bxy(y-ˉy)

2.2.2 Applications of Regression Analysis

Regression analysis is mostly used to carry out the financial procedure. Therefore, the

following list comprises 5 financial and related uses of regression analysis.

1. Forecasting:

 Regression analysis is frequently used in business to predict potential opportunities and

dangers. For instance, demand analysis predicts how many items a buyer is likely to

purchase.

 Demand, however, is not the only dependent variable when it comes to business. Much

more than just direct income can be predicted using regression analysis.

 By estimating the amount of people who would pass in front of a certain billboard, for

instance, we could forecast the highest bid for an advertisement. Regression analysis is a

key tool used by insurance companies to predict the creditworthiness of policyholders and
the number of claims that might be made during a specific time period.

2. CAPM:

136
 The linear regression model is a key component of the Capital Asset Pricing Model

(CAPM), which determines the relationship between an asset's expected return and the

associated market risk premium.

 Financial analysts typically use it to forecast corporate returns and operational

performance in their financial research.

 Regression analysis is used to calculate a stock's beta coefficient. Beta is a metric for

comparing return volatility to overall market risk.

 We can quickly calculate it in Excel using the SLOPE tool because it reflects the slope of

the CAPM regression.

3. Comparing with competition:

 It can be used to assess how well a business is doing financially in comparison to a

certain rival.

 It can also be used to figure out how the stock prices of two companies are related to one

another (this can be extended to find a correlation between 2 competing companies, 2

companies operating in an unrelated industry etc).

 It might help the company identify the factors affecting their sales in contrast to the
comparable company. These methods can help small businesses succeed quickly and

within a short period of time.

4. Identifying problems:

 Regression is beneficial for identifying erroneous judgments as well as for providing

factual support for management decisions.

 For instance, a manager of a retail store might believe that extending the hours of

operation will greatly increase sales.

 However, RA can contend that the higher income is insufficient to offset rising operating
costs brought by extended working hours (such as additional employee labour charges).

137
 This research may therefore provide quantitative support for decisions and assist
managers in avoiding errors based solely on intuition.

Reliable source

 Regression analysis (and other types of statistical analysis) are now being used by many

companies and their top executives to make better business decisions and cut down on

guessing and gut feelings.

 Regression enables businesses to approach management scientifically. Large and small

businesses are regularly overrun with enormous amounts of data.

 Regression analysis is a tool that managers can use to sift through data and select the

pertinent variables to help them reach the best conclusions.

2.2.3 Difference between Correlation & Regression

Correlation Regression

‘Correlation’ as the name says it determines the ‘Regression’ explains how an

interconnection or a co-relationship between the independent variable is numerically

variables. associated with the dependent variable.

In Correlation, both the independent and dependent However, in Regression, both the

values have no difference. dependent and independent variable are

different.

The primary objective of Correlation is, to find out a When it comes to regression, its primary

quantitative/numerical value expressing the intent is, to reckon the values of a

138
association between the values. haphazard variable based on the values

of the fixed variable.

Correlation stipulates the degree to which both of the However, regression specifies the effect

variables can move together. of the change in the unit in the known

variable(p) on the evaluated variable (q).

Correlation helps to constitute the connection Regression helps in estimating a

between the two variables. variable’s value based on another given

value.

2.2.4 Regression Lines

Procedures for forecasting can benefit from using regression lines. Its goal is to describe how the

dependent variable (y variable) and one or more independent variables are related (x variable).

By entering various values for the independent variables, an analyst can predict future

behaviours of the dependent variables by using the equation derived from the regression line.

139
Regression Line Formula: y = a + bx + u

Multiple Regression Line Formula: y= a + b1x1 +b2x2 + b3x3 +…+ btxt + u

The Least Square Regression

The straight line is typically found through linear regression. The least squares regression line

is another name for it. In a bivariate dataset, it is represented. Let's assume that the dependent

variable is y. X is the unrelated factor. The line of population regression is

Y = a0 + a1x

Where a0 is the constant and a1 is the regression coefficient and x is the value of the

independent variable. If you are given a random sample of observation, the population

regression line is estimated by

140
Y’ = a0 + a1x, where a0 is the constant and b1 is the regression coefficient. Here you will, ‘x’

is the value of the independent variable, and y’ is the predicted value of the dependent

variable.

Properties of the Regression Lines

 Regression coefficient values remain the same. Since shifting of origin takes place because

of the change of scale. The property says

 If the variables x and y are changed to u and v, respectively u= (x-a)/p v=(y-c) /q, Here p

and q are the constants. Byz =q/p*bvu Bxy=p/q*buv.

 If there are two lines of regression. Both of these lines intersect at a specific point [x’, y’].

Variables x and y are taken into consideration. According to the property, the intersection

of both the lines of regression, i.e. y on x and y, is [x’, y’]. This is the solution for both of

the equations of variables x and y.

 You will discover that the geometric mean of the two coefficients represents the

correlation coefficient between the two variables, x and y. Additionally, the common sign

of the two correlation coefficients will be indicated by the sign over the values of the

coefficients. So, if, according to the property, regression coefficients are byx= (b) and bxy=

(b’) then the correlation coefficient is r=+-sqrt (byx + bxy) so, in some cases, both the

coefficients give a negative value, and r is also negative. If both the values of coefficients

are positive, the r will be positive.

 The regression constant (a0) is equal to the y-intercept of the regression line. Where a0 and

a1 are the regression parameters.

2.2.5 Standard Error

141
The standard error of estimation is also known as the standard error of the regression (s). It is

a measurement of an observation's variance around a calculated regression line. It shows the

typical separation between the observed values and the regression line. The standard error of

regression gives an indication of how closely the observations are fitting the regression line

when it has smaller values. If the standard error is "0," then the correlation is flawless, and

there is no fluctuation that corresponds to the computed line.

The standard error of estimate measures the difference between the actual values of Y and the

anticipated (calculated) values of Y on the regression line, just as the standard deviation

measures the variation in a set of data from its mean. Both linear regression models and non-

linear regression models can use the standard error of the estimate. It is crucial for the

computation of prediction intervals and confidence intervals. The following formula can be

used to compute the Standard Error:

where,

 Syx = Standard error of estimate of y on x

 ye = Estimated value of y for a given value of x

Similarly,

where,

 Sxy = Standard error of estimate of x on y

142
 xe = Estimated value of x for a given value of y

The large the value of Syx or Sxy the greater the scatter on the line of regression. In such a

case the degree of correlation series is poor. The error of an estimate is an absolute measure

and is given by the ratio S/σ. This ratio is also used for finding the value of the coefficient of

correlation.

2.2.6 Regression coefficient


The constant "b" in the regression equation known as the regression coefficient indicates the

change in the dependent variable's value in relation to the unit change in the independent

variable.

There will be two regression coefficients if there are two regression equations:

 Regression Coefficient of X on Y: The regression coefficient of X on Y quantifies the

change in X for the unit change in Y and is denoted by the symbol bxy. In a symbolic

sense, it can be expressed as:

When the deviations from the real means of X and Y are taken into consideration, the

following formula can be used to get the bxy:

143
The following formula is applied when deviations from the assumed mean are

obtained:

 Regression Coefficient of Y on X: The symbol byx is used that measures the change

in Y corresponding to the unit change in X. Symbolically, it can be represented as:

In case, the deviations are taken from the actual means; the following formula is used:

The bxy can be calculated by using the following formula when the deviations are taken

from the assumed means:

The Regression Coefficient is also called a slope coefficient because it determines the slope

of the line, i.e. the change in the independent variable for the unit change in the independent

variable.

144
The link between a predictor variable and the responder is described by regression

coefficients, which are estimations of the unknowable population parameters. Coefficients are

the numbers that multiply the predictor values in linear regression. Let's say you have the

regression formula y = 3X + 5. In this equation, the predictor is X, the constant is +5, and the

coefficient is +3.

The direction of the link between a predictor variable and a responder variable is shown by

the sign of each coefficient.

 A positive sign indicates that as the predictor variable increases, the response variable

also increases.

 A negative sign indicates that as the predictor variable increases, the response variable

decreases.

When the predictor is changed by one unit, the coefficient value shows the average change in

the response. If a coefficient is +3, for instance, the mean response value rises by 3 for each

unit increase in the predictor.

2.2.7 Introduction to Multiple Regression

Let's start by examining the equation for linear regression in its general form:

y=B*x+A

Here, the coefficients A and B determine the equation, with x serving as the independent

variable and y serving as the dependent variable. The distinction between the equations for

linear regression and multiple regression is that the multiple regression equation must be able

to handle several inputs, as opposed to just the single input required by the linear regression

equation. The equation for multiple regression uses this change to account for:

145
y = B 1*x 1, B 2*x 2, B 3*x 3,..., B n*x n, and A

Subscripts in this equation stand in for many independent variables.

For instance, the first independent variable's value is x 1, the second independent variable's

value is x 2, and so on. It continues as we continue to include independent variables until the

final independent variable, x n, is included in the equation.

Note: any number, n, of independent variables may be used in this multiple regression

model, and additional terms may be added as necessary.

The same subscripts are used by the B coefficients, indicating that they are the coefficients

associated with each independent variable. As previously, A is only a constant that indicates

what the dependent variable, y, is worth when all of the independent variables, the xs, are

equal to zero.

Here's an illustration of multiple regression: Imagine that you are responsible for traffic

planning in your city and that you must determine the typical commute time for automobiles

travelling from the east to the west side of the city. Although you are unsure of the typical

time it takes, you are aware that it will vary depending on a number of variables, including

the distance travelled, the quantity of stoplights encountered, and the amount of other

vehicles on the road.

Case Study (Regression Analysis)

The following two goals are given in a case study scenario where you are the Chief Analytics

Officer and Head of Business Strategy at DresSMart Inc.

146
Objective 1: Improve the campaigns' conversion rates or the percentage of customers who

purchase items from the marketing product catalog.

Objective 2: Increase revenue from consumers who are converted.

The first goal was accomplished in the earlier sections of this case study example. Clients’

propensities to respond to campaigns were estimated using the classification models (Parts 5,

Parts 6, Parts 7 & Parts 8). The second goal is now left up to you: calculate the estimated

earnings from each consumer, assuming the customer reacts to the ad. This is a common

regression issue. You will use the data for 4200 consumers, out of 100,000 solicited

customers, who reacted to the prior campaigns, to create a regression model. All of these

4200 clients reside in various communities that can be divided into the three categories listed

below.

 Greater Cities

 Major Cities

 Little Towns

By the way, there are 1400 clients in each of these three groups, which are distributed equally

among the customers. The average profit made by these three types of cities was the first

thing you looked at. The average profit figures for these categories are different, as you can

see in the chart below. Remember these average values; they will be useful when we create

our regression model.

147
The second question is whether or not these average profitability figures differ significantly.

The distributions of all 4200 clients, broken out by location category, provide the answer to

this query. These distributions are depicted in the above figure (towards right). The following

table shows the density distribution for all 4200 of our original data's clients, broken down by

geography category. As you can see, certain situations in this distribution of profits are

negative due to consumer returns of goods and other losses.

The charts above reveal a few intuitive conclusions:

1. As a result of their citizens' stronger earning potential and disposable money, large cities

have higher average profit values than the other cities.

148
2. Due to their larger socio-economic diversity, large metropolitan areas also have a wider

distribution of profit than the other two categories.

Let's construct a straightforward regression model using these categories as the predictor

variables, keeping the information above in mind. The outcomes of our regression model are

as follows:

Std. t
Coefficients: Estimate Pr(>|t|)
Error value
Intercept 46 0.4691 98.06 <2e-16
Mid Sized
8 0.6635 12.06 <2e-16
Cities
Large Cities 22 0.6635 33.16 <2e-16
Multiple R-
0.2069
squared:
Adjusted R-
0.2065
squared:
F-statistic (P
2.20E-16
Value)

The following is the linear equation for this regression model

Keep in mind that the model's only predictor variables are mid-sized and major cities. The

intercept component incorporates the knowledge of tiny towns. Additionally, because these

predictor variables are dummy variables, their only two possible values are 0 or 1. For

example, if the location is a small town, mid-sized cities are equal to zero,

Recall the above average figures, this is the same average value for small towns. Now, if the

location is a mid-sized city then

149
Again this is the same as the average value for mid-sized cities. Finally, the estimated profit

through the resident customer of a large city is:

Now the next question is : how good is this model? For this we will have to scroll up to the

regression model results and look at the following two things:

P values for certain coefficients: The value in the coefficients' right-most column, 2e-16, is

extremely low. This indicates that the coefficients won't go zero practically certainly,

according to the model. This is comparable to your odds of defeating Usain Bolt, which are

incredibly slim but not zero.

For our model, the adjusted R-squared value is 0.2065. This indicates that only the location

category accounts for 20% of the variation in earnings. This is not bad for a single category

variable, and if we continue to include more significant variables in the model described

above, the value of the adjusted R-squared will continue to rise.

Key Terms- Regression, Regression lines, correlation, linear regression, Standard error,

multiple regression, application, statistics.

Summary- In this unit , we have covered topics related to correlation, correlation analysis,

various methods associated with it, regressions, and standard error, etc.

150
Exercise: Choose the correct option

1. Who introduced the term ‘regression’?

A. Karl Pearson

B. R.A Fischer

C. Croxton and Cowden

D. Francis Galton.

2. Choose the least likely assumption of a classic normal linear regression model?

A. The independent variable and the dependent variable have a linear relationship.

B. The independent variable is normally distributed.

C. There is no randomness in the independent variable.

151
D. None of the preceding.

3. Which one of the below statements regarding the regression line is correct?

A. The prediction equation is another name for a regression line.

B. A regression line is also referred to as the line of the average relationship.

C. The estimating equation is another name for a regression line.

D. All of the above.

E. None of the preceding.

4. The slope of the regression line of Y on X is also referred to as the:

A. Regression coefficient of X on Y

B. The correlation coefficient of X on Y

C. Regression coefficient of Y on X

D. Correlation coefficient of Y on X.

5. Which of the following statements is true about the arithmetic mean of two regression

coefficients?

A. It is less than the correlation coefficient

B. It is equal to the correlation coefficient

C. It is greater than or equal to the correlation coefficient

D. It is greater than the correlation coefficient

6. Which of the following statements is true about the regression line?

A. A regression line is also known as the line of the average relationship

B. A regression line is also known as the estimating equation

152
C. A regression line is also known as the prediction equation

D. All of the above

7. In a simple linear regression model (One independent variable), If we change the

input variable by 1 unit. How much output variable will change?

A. by 1

B. no change

C. by intercept

D. by its slope

8. Function used for linear regression in R is __________

A. lm(formula, data)

B. lr(formula, data)

C. lrm(formula, data)

D. [Link](formula, data)

9. What does a simple linear regression analysis examine?

A. The relationship between only two variables

B. The relationship between one dependent and one independent variable

C. The relationship between many variables

D. The relationship between two dependent and one independent variable

10. Linear Regression is an Example of-

153
A. Supervised Learning

B. Unsupervised Learning

C. Semi-Supervised Learning

Short Answer Type Questions-

1. What is the correlation of the coefficient?

2. What is regression?

3. What is Linear Regression?

Long Answer Type Questions-

1. Explain Multiple Regression.

2. What is Standard Error?

3. Write down the difference between correlation and regression?

Answers

Exercise: MCQ

1. (D)

2. (B)

3. (D)

4. (C)

154
5. (D)

6. (D)

7. (D)

8. (A)

9. (A&C)

10. (A)

2.3. Time Series Analysis

A time series is a collection of observations on a single variable that are made at regular

intervals of time. The succeeding intervals are typically separated by equal amounts of time,

such as 10 years, one year, one quarter, one month, one week, one day, and one hour, etc. The

population statistics for India is a time series, with a 10-year lag between each succeeding

figure. Similar annual data are provided for national income, agricultural and industrial

production, etc.

2.3.1 Objectives of Time Series Analysis

The analysis of time series entails its breakdown into distinct elements that have an impact on

the variable's value over a specific period. It is a quantitative and objective assessment of the

impact of different variables on the activity in question. The analysis of any time series data

has two primary goals:

(i) Analyzing the historical behaviour of data

(ii) To create future projections.

155
Because it enables us to understand the effects of diverse pressures, the study of historical

behaviour is crucial. This can make it easier to anticipate how events will develop in the

future, forecast the value of the variable, and make future plans.

2.3.2 Components of a Time Series

There are three key elements in a typical time series that appear to be independent of one

another and appear to be impacting time-series data.

Trend: The long-term, overall trend of either an increase or decrease in the forecast variable

y's average (or mean) value over time. Over time, the trend growth rate typically fluctuates.

Cycles- A cycle is defined as an upward and downward oscillation of unclear duration and

magnitude about the trend line caused by seasonal effect, with either a long time and irregular

swings. The average length of a business cycle is larger than one year but less than five to

seven years. Four phases make up the movement: peak (prosperity), contradiction (recession),

trough (depression), and expansion (recovery or growth).

Seasonal: This is a specific instance of a cycle component of a time series in which the

cycle's size and length are constant and occur at regular intervals throughout the year. For

instance, festival seasons may see a significant boost in a retail store's average sales.

Irregular- A short-term unforeseen and non-recurring set of events can create erratic or

irregular movements in a time series. There is no set pattern for these.

2.3.3 Methods of Time Series

The principal methods of measuring trend fall into below mentioned categories:

1. Free Hand Curve methods

2. Method of Averages

156
3. Method of least squares

The goal of the time series methods is to use a mathematical formula to predict the future of

an observable historical trend for a particular variable. These approaches make no attempt to

explain why the variable under research will change in the future. The use of a causal

mechanism overcomes this drawback of the time series approach. The causal approach looks

for variables that affect the variable in some way or cause it to fluctuate in a predictable way.

Regression analysis and correlation analysis are the two causal techniques that have already

been covered. Some time series techniques, such as freehand curves and moving averages,

only describe the values of the input data, but semi-average and least squares techniques

assist in finding a trend equation that may be used to characterise the input data values.

 Freehand Method

The data can often be easily and possibly adequately represented by a freehand curve that is

drawn smoothly over the data values. By simply extending the trend line, the forecast may be

produced. The following prerequisites should be met by a trend line fitted by the freehand

method:

The following prerequisites must be met for a trend line to be fit by hand:

(i) The trend line should be straight or be a combination of long, progressive curves.

(ii) The total vertical deviation of the observations above the trend line should be equal to the

total vertical deviation of the observations below the trend line.

(iii) It is best to have a minimal sum of squares for the vertical deviations of the observations

from the trend line.

157
(iv) The trend line should cut through the cycles so that, not only for the entire series, but

ideally for each complete cycle as well, the area above the trend line and the area below the

trend line are equal.

Example- Fit a trend line to the following data by using the freehand method.

Year 1991 1992 1993 1994 1995 1996 1997 1998

Sales turnover : (Rs. in lakh) 80 90 92 83 94 99 92 104

Solution- presents the freehand graph of sales turnover (Rs. in lakh) from 1991 to 1998. The

forecast can be obtained simply by extending the trend line.

Limitations of the freehand method


 This method is quite subjective because the trend line depends on individual judgement,

so what works well for one person might not work well for another.

 If the trend line is utilised to make predictions, it won't be very useful.

 Building a freehand trend takes a lot of time if a cautious and meticulous job is to be

done.

158
Methods of Averages

The goal of smoothing techniques is to eliminate the random fluctuations resulting from the

time series' irregular components and, in doing so, give us a general sense of how the data are

moving over time. Three smoothing techniques will be covered in this section.

(i) Moving averages

(ii) Weighted moving averages

(iii) Semi-averages

The data requirements for the techniques to be discussed in this section are minimal and these

techniques are easy to use and understand.

Moving Averages

The moving Averages Method gives a trend with a fair degree of accuracy. In this method,

we take the arithmetic mean of the values for a certain time span. The time span can be three

years, four -years, five- years and so on, depending on the data set and our interest. We will

see the working procedure of this method.

It is crucial to first smooth out the irregular pattern in the historical values of the variable

before using this as the foundation for a future projection if we are watching the movement of

some variable values over time and trying to project this movement into the future. A series

of moving average calculations can be used to accomplish this. This method is arbitrary and

is reliant on the duration of the period used to calculate moving averages. The period should

be an integer value that corresponds to or is a multiple of the expected average duration of a

159
cycle in the series in order to eliminate the impact of cyclical changes. The moving averages,

which serve as an estimate of the next period’s value of a variable given a period of length n

are expressed as:

Moving average,

where t = current time period

D = actual data which is exchanged each period

n = length of time period In this method,

the term ‘moving’ is used because it is obtained by summing and averaging the values from a

given number of periods, each time deleting the oldest value and adding a new value.

Procedure:

(i) Select the moving averages' timeframe (three- years, four -years).

(ii) Averages for odd-numbered years can be calculated by:

(iii) If the moving average is an odd number, centering it is not an issue; the average value

will be centred every three years, with the exception of the second year.

(iv) In the case of even years, averages can be obtained by calculating,

160
(v) If the moving average is an even number, the first four values' average will be positioned

between the second and third year, and the second four values' average will be positioned

between the third and fourth year. The third year will see a new average of these two

averages. This holds true for the remaining values in the issue. The centering of the averages

is the name given to this technique.

The limitation of this method is that it is highly subjective and dependent on the length of

period chosen for constructing the averages.

There are three drawbacks to moving averages:

(i) As the size of n (the number of periods averaged) rises, the approach becomes less

sensitive to actual changes in the data while also smoothing out variances better.

(ii) Moving averages struggle to detect patterns. Since these are averages, it won't predict a

change to either a higher or lower level because it will always remain within previous levels.

(iii) Moving average requires huge archives of historical data.

Example- Use the data below to calculate the number of students enrolled in a higher

secondary school in a specific hamlet over a three-year period.

Solution:

Computation of three- yearly moving averages.

161
Example- Use the information below to calculate the number of pupils enrolled in a

higher secondary school in a specific city on a four-year moving average basis.

Solution:

Computation of four- yearly moving averages.

162
Weighted Moving Averages

Each observation in moving averages is given equal weight (weight). To determine a

weighted average of the most recent n values, different values could be used. Since there is

no established formula to determine them, the choice of weights is somewhat discretionary.

The most recent observation is typically given the most weight, whereas previous data values

receive less weight.

As more recent data points are more pertinent than those from the distant past, weighted

moving averages give more weight to more recent data items. The weights should total 1 (or

100%) when added together.

Mathematically, a weighted moving average can be written as

163
Weighted moving average = Σ(Weight for period n) (Data value in period n) /ΣWeights

Example-

Date Closing Price of AAPL Weighting

June 26 $22.72 5/15

June 25 $22.59 4/15

June 24 $22.57 3/15

June 23 $22.71 2/15

June 20 $22.73 1/15

The specified price is multiplied by the corresponding weighting before the values are added

up to determine the weighted average. The following is the WMA formula:

The WMA's denominator is the sum of the price periods expressed as a triangular number.

The weighted five-day moving average in the aforementioned example from the table would

be $22.65:

164
In this illustration, the most recent data point received the highest weighting out of a random

total of 15. Any value's values can be weighed however you see fit. The weighted average's

lower value in comparison to the simple average shows that recent selling pressure might be

stronger than some traders think. When utilising weighted moving averages, the most

common option for traders is to give recent values more weight.

Semi-Average Method

If a linear function can accurately describe the data, we may estimate the slope and intercept

of the trend using the semi-average method. Simply dividing the data into two sections and

calculating their individual arithmetic means is the technique. These two points are plotted

corresponding to the middle of the class interval that each portion covers, and a straight line

connecting these two points create the necessary trend line. The slope is determined by the

ratio of the difference in the arithmetic means of the number of years between them or the

change per unit of time, and the intercept value is the arithmetic mean of the first section.

A time series using the formula y = a + bx is the outcome. A and b are the intercept and slope

values, and y is the estimated trend value. Always include a reference to the year where x = 0

and a description of the units of x and y in your equation's full formulation. If the trend is

linear, the semi-average method of creating a trend equation may be acceptable and relatively

simple to commute. The forecast will be skewed and less accurate if the data diverge

significantly from linearity.

The semi-averages are computed using this method to determine the trend values. We'll

examine how this strategy operates right now.

Procedure:

(i) The information is split into two equal portions. If the number of data points is odd, two

equal sections can be created by simply leaving out the middle year.

165
(ii) Each component's average is calculated, giving us two points.

(iii) The midpoint (year) of each half is where each point is plotted.

(iv) Draw a straight line connecting the two spots.

(v) Either side of the straight line can be expanded.

(vi) According to semi-averages' methodology, this line represents the trend line.

Example 1- Fit a trend line by the method of semi-averages for the given data.

Solution-

Due to the odd number of years (seven), we will omit the production value of the middle year

and instead calculate the averages of the first three and last three years.

166
Example-2 Fit a trend line by the method of semi-averages for the given data.

Solution-

Since there are eight even years, we can divide the provided data in half and get the averages

for the first four years and the final four years.

167
Methods of Least Square

For medium- to long-term projections, the trend project approach involves fitting a trend line

to a set of historical data points and then projecting the line into the future. Depending on the

movement of time-series data, various mathematical trend equations (such as exponential and

quadratic) can be created.

Reasons to study trend:

A few reasons to study trends are as follows:

1. The study of trends allows us to describe a historical pattern so that we may evaluate the

success of the prior policy.

2. The study enables us to make future intermediate- and long-term forecasting projections by

allowing us to use trends as a tool.

3. Using trends as a reference for short-term (one-year or less) forecasting of general business

cycle conditions allows us to isolate and then minimise its influencing impacts on the time-

series model. Model for Linear Trend The least squares method can be used if we desire to

create a linear trend line using a precise statistical technique.

168
A least squares line's slope and y-intercept, or the height at which it intersects the y-axis, are

used to define it (the angle of the line). The following equation can be used to represent the

line if we can determine the y-intercept and slope. y = anticipated value of the dependent

variable, where y = a + bx an is the y-axis intercept. b = slope of the regression line (or the

rate of change in y for a given change in x) x = unrelated variable (which is time in this case)

Because it produces what mathematics refers to as a "line of best fit," least squares is one of

the most popular techniques for detecting trends in data. The characteristics of this trend line

include

(i) the summation of all vertical deviations about it is zero, that is, Σ(y- yˆ ) = 0,

(ii) the summation f all vertical deviations squared is a minimum, that is, Σ(y- yˆ ) is least,

and

(iii) the line goes through the mean values of variables x and y.

It is determined for linear equations by the simultaneous solutions of the two normal

equations, Σy = na + bx and xy = aΣx + bΣx2. When two terms in three equations can be

removed by coding the data so that ∑x = 0, we get ∑y = na and ∑xy = b∑x2 instead. When

working with time-series data, coding is simple. We chose x = 0 for the time period's centre

when coding the data, and we have an equal number of plus and minus periods on either side

of the trend line that add to zero. The values of constants a and b can also be determined as

follows for any regression line:

̅̅̅̅
∑𝒙𝒚 − 𝒏𝒙𝒚
𝒃= 𝟐
̅ − 𝒃𝒙
,𝒂 = 𝒚 ̅
̅)𝟐
𝜮𝒙 − 𝒏(𝒙

Merits of Least Square Methods

169
 When the distribution of the deviations is roughly normal, the method of least squares

provides the most accurate measurement of the secular trend in a time series.

 Unbiased estimates of the parameters can be found in the least-squares estimations.

 The approach can be applied when the trend is quadratic, exponential, or linear.

Demerits of Least Square Methods

 Extremely big deviations from the trend are given too much weight by the least-

squares method.

 Only during the period, it refers to the least-squares line is the best.

 Its position could be altered by the removal or addition for one or more time periods.

2.3.4 Applications of Time series in business decision-making problems

Time series in Financial and Business Domain

The majority of financial, investment, and commercial choices are based on predictions of

future changes and demand in the financial sector.

Forecasting and time series analysis are crucial steps in understanding how financial markets

behave in a dynamic and powerful way. An expert can foresee the necessary projections for

crucial financial applications in a variety of sectors, such as risk evolution, option pricing &

trading, portfolio design, etc., by looking at financial data.

Time series analysis, which may be used to forecast interest rates, foreign exchange risk,

stock market volatility, and many other things, has evolved into an integral aspect of financial

analysis. Policymakers and business professionals use financial forecasting to decide on

production, purchasing, market sustainability, resource allocation, etc.

170
This study is used in investments to monitor price swings and a security's price evolution. For

instance, it is possible to record a security's price;

 For the short term, such as the observation per hour for a business day, and

 For the long term, such as observation at the month end for five years

To track how a specific asset, security, or economic variable behaves or changes over time,

time series analysis is incredibly helpful. For instance, it can be used to assess how certain

underlying changes react when applied to other data observations made within the same time

period.

Time series in Medical Domain

A data-driven industry, medicine has developed and is still making significant advances in

time series analysis of human knowledge.

Think about the scenario where time series and a medical approach are combined. Data

mining and CBR (case-based reasoning) work in synergy to pre-process time series data for

feature mining, which can be used to track patients' development over time.

In the field of medicine, it is crucial to look at how behaviour changes over time rather than

drawing conclusions based just on the time series' absolute values. The typical demonstration

of linking time series with case-based monitoring is to diagnose heart rate variability in

conjunction with respiration based on the sensor readings.

However, time series in the context of the epidemiology domain has only lately and slowly

emerged as approaches to time series analysis necessitate recordkeeping systems so that

records should be connected over time and collected precisely at regular intervals.

171
Healthcare applications utilising time series analysis have produced significant

prognostication for the industry as well as for individual patients' health diagnoses once the

government has installed enough scientific devices to collect good and lengthy temporal

data.

Medical Instruments

Time series analysis has made its way into medicine with the advent of medical devices such

as

 Electrocardiograms (ECGs) were invented in 1901: For diagnosing cardiac conditions by

recording the electrical pulses passing through the heart.

 Electroencephalogram (EEG) was invented in 1924: For measuring electrical

activity/impulses in the brain.

Medical professionals now have more opportunities to use time series for medical diagnostics

because to these advancements.

As a result of the development of wearable sensors and smart electronic healthcare devices,

people may now take routine measurements automatically and with little input, leading to a

reliable collection of longitudinal medical data for both ill and healthy people.

Time Series in Astronomy

Different fields of astronomy and astrophysics are among the present and modern

applications where time series plays a key role,

Astronomical specialists are skilled in time series for calibrating devices and researching

things of their interest because astronomy, being specific in its field, heavily depends on

graphing objects, trajectories, and exact measurements.

172
Time series data has a long history in the astronomy field; for instance, sunspot time series

were recorded in China in 800 BC, making sunspot data collecting as well-recorded natural

phenomenon. Time series data have an essential impact on knowing and quantifying anything

about the cosmos.

Similarly, time series analysis was utilised in earlier eras.

 To discover variable stars that are used to surmise stellar distances, and

 To observe transitory events such as supernovae to understand the mechanism of the

changing of the universe with time.

These systems, which depend on the wavelengths and light intensities of light to transmit

time series data in real-time, enable astronomers to observe phenomena as they happen.

Astroinformatics and astrostatistics are new fields of study that have emerged in recent

decades as a result of data-driven astronomy; these paradigms integrate key fields including

statistics, data mining, machine learning, and artificial intelligence. Here, time series analysis

would play a role in the quick detection and classification of astronomical objects as well as

the independent characterisation of unique events.

Time Series in Forecasting Weather

Aristotle, a Greek philosopher, conducted research on weather events in antiquity with the

goal of determining the origins and consequences of weather variations. Later, scientists

began to compile weather-related data, recording it on an hourly or daily basis and storing it

in various locations, using the instrument "barometer" to calculate the status of atmospheric

conditions.

173
Newspapers started printing personalised weather forecasts over time, and as technology

developed, forecasts eventually went beyond just basic weather conditions.

Many countries have set up tens of thousands of weather forecasting stations all around the

world in order to conduct atmospheric measurements using computer methods for quick

compilations.

These stations are outfitted with highly functioning equipment and are connected to one

another in order to gather weather data from various locations and anticipate weather

conditions at all times according to requirements.

Time Series in Business Development

As the process examines prior data trends, time series forecasting assists organisations in

making wise business decisions. It can be helpful in predicting future possibilities and events

in the following ways.

 Reliability: Time series forecasting is very trustworthy when the data has a wide range of

time intervals in the form of numerous observations over a longer time horizon. By

utilising data observations at varied time periods, it delivers illuminating information.

 Growth: Time series is the best asset to use when evaluating endogenous as well as

overall financial performance and growth. Endogenous growth is essentially the

improvement of internal human capital within firms that leads to economic growth. Time

series forecasting, for instance, can be used to analyse the effects of any policy variable.

 Trend estimation: To find trends, time series methods can be used. For instance, these

methods examine data observations to determine when measurements show a decline or

increase in sales of a specific product.

 Seasonal patterns: Variations in recorded data points may reveal seasonal patterns and

oscillations that serve as the foundation for data forecasting. The information gathered is

174
important for markets whose products vary seasonally and helps businesses manage their

product development and delivery needs.

2.3.5 Methods to measure secular trends

Secular Trends

It speaks of the data's propensity to trend upward or downward over the long run. Examples

of secular trends that have an upward orientation include changes in productivity, an increase

in the rate of capital formation, population expansion, etc. Conversely, deaths brought on by

better medical care and cleanliness show a downward trend. All of these factors work slowly

and gradually to affect the time series variable.

Methods of Measuring Trend

Trend is measured using by the following methods:

1. Graphical method

2. Semi averages method

3. Moving averages method

4. Method of least squares

1. Graphical Method

By placing the time variable on the X-axis and the value variable on the Y-axis, the values of

a time series are plotted using this method on graph paper. After that, a smooth curve is free-

hand drawn through the points that were plotted. To forecast the values, an extension of the

175
trend line shown above can be used. When sketching the smooth curve freehand, the

following considerations must be made.

(i) The line or curve should be smooth

(ii) There should be roughly the same number of points above and below the line or curve.

(iii) The vertical deviation of the points above and below the smoothed line, when added

together, is the sum of their squared vertical deviations.

Merits

 It is a simple method of estimating trends.

 It requires no mathematical computations.

 This method can be used even if the trend is not linear.

Demerits

 It is a subjective method

 The values of trend obtained by different statisticians would be different and hence not

reliable.

Example- Annual power consumption per household in a certain locality was reported

below.

Draw a free hand curve for the above data.

176
Solution:

2. Semi-Average Method

The series is split into two equal halves using this manner, and the average of each portion is

plotted at the halfway point of its duration.

(i) In the event that the series has an even number of years, it can be divided in half. Place the

values at the midpoint of each of the two series' respective durations by calculating the

average of the two portions of the series.

(ii) It is impossible to divide a series into two equal halves if the number of years in the series

is odd. There won't be a midway year. Find the arithmetic mean for each segment of the data

after splitting it into two pieces. So, we obtain semi-averages.

(iii) The trend values for other years can be calculated by adding or subtracting successively

from each year that comes before or after any given year.

Merits

 This method is very simple and easy to understand

177
 It does not require many calculations.

Demerits

 This method is used only when the trend is linear.

 It is used for the calculation of averages, and they are affected by extreme values.

Example- Calculate the trend values using semi-averages methods for the income from the

forest department. Find the yearly increase.

Source: The Principal Chief conservator of forests, Chennai-15.

Solution:

Difference between the central years = 2012 – 2009 = 3

Difference between the semi-averages = 82.513 – 53.877 = 28.636

Increase in trend value for one year = 28.636 /3 = 9.545

178
Trend values for the previous and successive years of the central years can be calculated by

subtracting and adding, respectively, the increase in annual trend value.

3. Moving Averages Method

A series of arithmetic means of the variance values in a sequence make up a moving average.

Another method of creating a smooth curve for a time series of data is as shown here.

The seasonal variations are more typically removed using moving averages. The moving

average method, even when used to estimate trend values, aids in the establishment of a trend

line by removing the time series' cyclical, seasonal, and random changes. The length of the

time series data determines the moving average's period.

When utilising this strategy, selecting the moving average's length is crucial.

The smoothing of variances for a moving average is greatly influenced by the selection of the

right length.

In general, if the number of years for the moving average is more, then the curve becomes

smooth.

Merits

 It can be easily applied

 It is useful in the case of series with periodic fluctuations.

 It does not show different results when used by different persons

 It can be used to find the figures on either extreme; that is, for the past and future years.

Demerits

179
 In non-periodic data, this method is less effective.

 Selection of proper ‘period’ or ‘time interval’ for computing moving average is difficult.

 Values for the first few years and as well as for the last few years cannot be found.

Moving averages odd number of years (3 years)

 The following steps must be taken into account in order to determine the trend values

using the three-yearly moving average approach.

 Add up the first three years' numbers, then compare the yearly total to the median year.

[This amount is known as the moving total]

 Keep the first year's value, add the values of the following three, and compare them to the

median year.

 This procedure must be carried out again until all data values needed for calculations have

been obtained.

 To obtain the 3-year moving averages, which serve as the trend values we need, each 3-

yearly moving total must be divided by three.

Example- Calculate the 3-year moving averages for the loans issued by co-operative banks

for non-farm sector/small scale industries based on the values given below.

180
Solution: The three-year moving averages are shown in the last column.

Moving averages - even number of years (4 years)

 Summarize the first four years' values and arrange them between the second and third

years.

 Leave the first year value alone and add the next four values starting with the second

year. Write the sum next to the centre position.

 This process must be repeated until the final item's value is considered.

 Add the first two 4-year moving totals, then record the result next to the third year.

 Keep the initial 4-year moving total and add the following two, then position them against

the fourth year.

 This procedure must be carried out repeatedly until all 4-yearly moving totals have been

added up and centred.

 Divide the 4-years moving total by 8 to get the moving averages which are our required

trend values.

181
4. Method of least squares

A line from which the sum of all deviations from various locations is zero is known as the

line of best fit. This is the most effective way to get trend values. It provides a practical

foundation for figuring out the time series' line of best fit. It is a formula for calculating

trends. In addition, as compared to other fitting techniques, the sum of the squares of these

variances would be the smallest. As a result, this technique is called the Method of Least

Squares and meets the criteria listed below:

(i) The sum of the deviations of the actual values of Y and Ŷ (estimated value of Y) is Zero.

that is Σ(Y–Ŷ) = 0.

(ii) The sum of squares of the deviations of the actual values of Y and Ŷ (estimated value

of Y) is the least, that is, Σ(Y–Ŷ)2 is the least ;

Procedure:

(i) The straight line trend is represented by the equation Y = a + bX …(1)

where Y is the actual value, X is time, a, and b are constants

(ii) The constants ‘a’ and ‘b’ are estimated by solving the following two normal

Equations ΣY = n a + b ΣX ...(2)

ΣXY = a ΣX + b ΣX2 ...(3)

Where ‘n’ = the number of years given in the data.

(iii) By taking the mid-point of the time as the origin, we get ΣX = 0

(iv) When ΣX = 0, the two normal equations reduce to

182
The constant ‘a’ gives the mean of Y and ‘b’ gives the rate of change (slope).

(v) By substituting the values of ‘a’ and ‘b’ in the trend equation (1), we get the Line of Best

Fit.

Secular trend, one of the time series' four elements, shows the direction the data will take

over the long run. The least squares approach is one mathematical methodology that can be

used to determine the trend values. The line created using this method is referred to as the

line of best fit because it is the most frequently utilised in practice and produces the least sum

of squared variances between the actual and computed values.

It aids in value projections for the future. It is crucial for determining the trend values of time

series data in the economy and in business.

Solved Examples-

Given below are the data relating to the production of sugarcane in a district.

Fit a straight line trend by the method of least squares and tabulate the trend values.

Solution:

183
Computation of trend values by the method of least squares (ODD Years).

Therefore, the required equation of the straight line trend is given by

Y = a+bX;

Y = 45.143 + 1.036 (x-2003)

The trend values can be obtained by

When X = 2000 , Yt = 45.143 + 1.036(2000–2003) = 42.035

When X = 2001, Yt = 45.143 + 1.036(2001–2003) = 43.071,

similarly other values can be obtained.

Given below are the data relating to the sales of a product in a district.

Fit a straight line trend by the method of least squares and tabulate the trend values.

184
Solution:

Computation of trend values by the method of least squares.

In case of EVEN number of years, let us consider

185
similarly other values can be obtained

Computation of Trend using Method of Least squares

The method of least squares is a device for finding the equation which best fits a given set of

observations.

Suppose we are given n pairs of observations, and it is required to fit a straight line to these

data. The general equation of the straight line is:

y = a + bx

where a and b are constants.

Any value of a and b would give a straight line, and once these values are obtained, an

estimate of y can be obtained by substituting the observed values of y. In order for the

equation y = a + b x gives a good representation of the linear relationship between x and y, it

is desirable that the estimated values of yi, say y^ i on the whole close enough to the

observed values yi, i = 1, 2, …, n. According to the principle of least squares, the best fitting

equation is obtained by minimizing the sum of squares of differences

186
is minimum. This leads us to two normal equations.

Solving these two equations, we get the vales for a and b and the fit of the trend equation

(line of best):

y = a + bx

Substituting the observed values xi in the above equation, we get the trend values yi, i = 1, 2,

…, n.

Merits

 Personal prejudice is totally eliminated by the least squares method.

 Trend values can be provided for each of the specified time periods.

 Using this technique, we may predict future values.

187
Demerits

 In comparison to the other ways, the calculations for this method are challenging.

 It overlooks cyclical, seasonal, and irregular changes; adding new observations

necessitates recalculations.

 Only the near future, not the far future, can be predicted for the trend.

Steps for calculating trend values when n is odd:

(i) Subtract the first year from all the years (x)

(ii) Take the middle value (A)

(ii) Find ui = xi – A

(iv) Find ui2 and uiyi

Then use the normal equations:

Then the estimated equation of straight line is:

y = a + b u = a + b (x – A)

Example

188
Fit a straight line trend by the method of least squares for the following consumer price index

numbers of the industrial workers.

Solution:

The equation of the straight line is y = a + bx

= a + bu where u = X – 2

The normal equations give:

y = 197.4 + 16.2 (X – 2)

189
= 197.4 + 16.2 X – 32.4

= 16.2 X + 165

That is, y = 165 + 16.2X

To get the required trend values, put X = 0, 1, 2, 3, 4 in the estimated equation.

X = 0, y = 165 + 0 = 165

X = 1, y = 165 + 16.2 = 181.2

X = 2, y = 165 + 32.4 = 197.4

X = 3, y = 165 + 48.6 = 213.6

X = 4, y = 165 + 64.8 = 229.8

Hence, the trend values for 2010, 2011, 2012, 2013 and 2014 are 165, 181.2, 197.4, 213.6

and 229.8 respectively.

Steps for calculating trend values when n is even:

i) Subtract the first year from all the years (x)

ii) Find ui = 2X – (n – 1)

iii) Find ui2 and ui yi

Then follow the same procedure used in the previous method for odd years

190
2.3.6 Case Study (Time Series Analysis)

A client (a Multistore Retailer) had witnessed unusual fluctuation in demand for certain

SKUs from a particular product category. The client wanted to create a forecasting

model to estimate demand for certain SKUs for 1 to 12 months in the future. A time-

series dataset with monthly data for pricing, sales, and around 50 current-period or

lagged potential predictor factors was created. To forecast future demand, an ensemble

of LSTM and Autoregressive Time-Series Model was built. Forecasted demand was

used by the client company to better control production and inventory costs and boost

profitability.

Strategic Challenge:

Rapid expansion in Indian middle-class wealth causes periodic excess demand and big

spikes in the demand for specific SKUs in the client's manufacturing process.

Furthermore, fresh supplies of identical products on the market were rapidly arising,

resulting in its periodic oversupply and subsequent price reduction owing to

competition. Our customer sought to anticipate future changes in demand in order to

better manage product stocks and improve profitability.

The goal of the research was to create a forecasting model of the demand for a specific

product category SKUs. The specific objectives were to:

 Compile a database of microeconomic and macroeconomic time series variables

that could be used as possible predictors of demand, as well as transactional

data.

 Build a robust forecasting model of demand for a given SKU using time-series

regression and deep learning approaches.

191
 Create a web-based forecast simulation tool that allows clients to input updated

predictor factors and monitor updated projections of SKU demand.

Analytical Design:

The Advanced Analytics team built a time-series analysis dataset, modifying all series to

be monthly and addressing missing values, holidays, and sales seasons properly. More

than 20 macroeconomic variables were subjected to variable selection methods in order

to identify the most promising linear and nonlinear predictors, lagged predictors, and

predictor combinations. The mean absolute prediction error was used to calculate

predictive power. More than ten distinct models were researched and analysed in order

to find the top five models for each required forecast time range (1,2,3,6,9 & 12

months). A forecast simulator based on an ensemble of LSTM and Autoregressive

Time-Series Models was developed to forecast 1, 2, 3, 6, 9, and 12 months into the

future, while accounting for serial correlation (the correlation over time of the impact of

unobserved variables on the variable being predicted—in this case, demand). The

ensemble technique aggregated forecasts from numerous models, enhancing forecast

accuracy.

Result:

The result was the creation of a Web-based forecasting tool that allowed the client's

management team to enter updated values of predictor variables each month and

forecast future demand for a certain SKU. Following that, the client organisation

evaluated the model by comparing forecasted vs. actual results for the first several

months. The resulting forecast accuracy was outstanding, prompting the client

organisation to:

 Use the model projections as an input to business operations; and

192
 Conduct a follow-up study utilising the forecasting method in another product

category.

Glossary

 Regression: Sir Francis Galton's research on the heights of brothers through generations is

where the term regression first appeared. Children with unusually tall (or short) parents

"regress" to the demographic mean in terms of height.

 Regression analysis- Today, any modelling of a forecast variable Y as a function of a

group of explanatory variables X1 through Xk is referred to as regression analysis. ratios

of regression Regression involves modelling the forecast variable Y as a function of the

explanatory variables X1 through Xk. The explanatory variables are multiplied by the

regression coefficients. To comprehend the significance of each explanatory variable (as

it relates to Y) and the interdependence of the explanatory variables, use the estimations

of these regression coefficients (as they relate to Y).

 Time series are collections of statistical data that are organised and displayed according to

time. Based on the historical data in the time series, time series analysis predicts the data

for the future.

 When growth has ups and downs within the same year, we are witnessing seasonal

variance.

 Cyclic patterns in the data across time are known as cyclical variation.

 Free hand curves are rather straightforward, easy to understand, and uncomplicated.

Simply plot the curve using the data points that are available, then extend the trend line to

anticipate OR predict the future.

193
 The moving average method bases its operation on arithmetic mean calculations of a

fixed number of readings OR observations over a certain period (3 years, 4 years etc.).

 Least Squares: This technique employs regression analysis to identify the time series data'

trend line.

Important formula-

Correlation Coefficient Formula

The above formulas can also be written as:

The sample correlation coefficient formula is:

Simple Linear Regression Equation

Y = a + bX

Where,

Y = Dependent variable

X = Independent variable

a = [(∑y)(∑x2) – (∑x)(∑xy)]/ [n(∑x2) – (∑x)2]

194
b = [n(∑xy) – (∑x)(∑y)]/ [n(∑x2) – (∑x)2]

Regression Coefficient
In the linear regression line, the equation is given by:

Y = b0 + b1X

Here b0 is a constant and b1 is the regression coefficient.

The formula for the regression coefficient is given below.

b1 = ∑[(xi – x)(yi – y)]/ ∑[(xi – x)2]

The regression coefficient of y and x formula is:

byx = r(σy/σx)

The regression coefficient of x on y formula is:

bxy = r(σx/σy)

Where,

σx = Standard deviation of x

σy = Standard deviation of y

Key Terms- Time Series Analysis, Decision making, Business, Trends, Regression

coefficients, R squared Value.

Summary- In this unit, students have gone through, Methods of Least Squares, moving

averages, trend lines, cyclic patterns, time series, regression and regression analysis in details.

195
Exercise: Choose the correct option

1) Which of the following is an example of time series problem?

1. Estimating number of hotel rooms booking in next 6 months.

2. Estimating the total sales in next 3 years of an insurance company.

3. Estimating the number of calls for the next one week.

A) Only 3

B) 1 and 2

C) 2 and 3

D) 1 and 3

E) 1,2 and 3

2) Which of the following is not an example of a time series model?

A) Naive approach

B) Exponential smoothing

C) Moving Average

D)None of the above

3) Which of the following can’t be a component for a time series plot?

A) Seasonality

B) Trend

C) Cyclical

D) Noise

E) None of the above

4) Which of the following is relatively easier to estimate in time series modeling?

196
A) Seasonality

B) Cyclical

C) No difference between Seasonality and Cyclical

5) The below time series plot contains both Cyclical and Seasonality component.

A) TRUE

B) FALSE

6) Adjacent observations in time series data (excluding white noise) are significantly

independent and identically distributed (IID).

A) TRUE

B) FALSE

7) Smoothing parameter close to one gives more weight or influence to recent

observations over the forecast.

A) TRUE

B) FALSE

8) Which of the following is not a necessary condition for weakly stationary time series?

197
A) Mean is constant and does not depend on time

B) Autocovariance function depends on s and t only through their difference |s-t| (where t and

s are moments in time)

C) The time series under considerations is a finite variance process

D) Time series is Gaussian

9) Which of the following is not a technique used in smoothing time series?

A) Nearest Neighbour Regression

B) Locally weighted scatter plot smoothing

C) Tree based models like (CART)

D) Smoothing Splines

10) If the demand is 100 during October 2016, 200 in November 2016, 300 in December

2016, 400 in January 2017. What is the 3-month simple moving average for February

2017?

A) 300

B) 350

C) 400

D) Need more information

Short Answer Type Questions-

1. What is Moving Average?

2. What is regression?

3. What is Linear Regression?

Long Answer Type Questions-

198
1. What are some real-world applications of Time-Series Forecasting?

2. The following figures relates to the profits of a commercial concern for 8 years

Find the trend of profits by the method of three yearly moving averages.

3. Explain Method of Least Squares?

Answers

Exercise: MCQ

1. (E)

2. (D)

3. (E)

4. (A)

5. (B)

6. (B)

7. (A)

8. (D)

9. (C)

10. (A)

199
Module 3

Learning outcomes

At the end of this module, you will be able to

 Discuss the different types of analysis and application

 Formulate the basis of hypothesis testing

 Explain the application of hypothesis testing

3.1. Types of Analysis

3.1.1 Introduction to Uni-variate, Bi-variate, and Multi-Variate Analysis of Data

Exploratory data analysis is the initial examination of data to identify links between variables

in the data and to acquire understanding of trends, patterns, and relationships between

different entities in the data set using statistics and visualisation tools (EDA).

Exploratory data analysis is divided into two categories, each of which is categorised as

either graphical or non-graphical. Each approach is either univariate, bivariate, or

multivariate after that.

Uni-Variate Analysis

Analyzing just one variable is referred to as univariate analysis. This is easy to remember

because the word "uni" denotes "one."

200
To comprehend the distribution of values for a single variable, use univariate analysis.

Comparing this kind of analysis to the following

 Analysis of two variables is known as bivariate analysis.

 Analysis of two or more variables is known as multivariate analysis.

Example- suppose we have the following dataset:

To learn more about the distribution of values in the dataset, we might opt to run

univariate analysis on any of the individual variables.

As an illustration, we may decide to run a univariate analysis on the variable

"household size":

201
There is only one reliable variable in a univariate analysis because uni means one and variate

means variable. Univariate analysis aims to derive the data, characterise and summarise it,

and examine any patterns that may be there. It investigates every variable in a dataset

independently, and categorical and numerical variables are both acceptable.

The Central Tendency (mean, mode, and median), Dispersion (range, variance), Quartiles

(interquartile range), and Standard deviation are a few patterns that are simple to spot using

univariate analysis.

There are three typical methods for carrying out univariate analysis:

1. Summary Statistics

The most typical application of univariate analysis is the use of summary statistics to describe

a variable.

Two common categories of summary statistics are:

Measures of central tendency: They are numerical values used to pinpoint where a dataset's

centre resides. The median and mean are two examples.

Measures of dispersion: these figures show how dispersed the values in the dataset are.

Examples include the variance, standard deviation, interquartile range, and range.

2. Frequency Distributions

The creation of a frequency distribution, which details how frequently various values appear

in a dataset, is another method of carrying out univariate analysis.

3. Charts
Making charts to show the value distribution for a particular variable is yet another technique

to undertake univariate analysis.

202
Common examples include:

 Boxplots

 Histograms

 Density Curves

 Pie Charts

The Household Size variable from our previously described dataset is used in the examples

that follow to demonstrate how to carry out each type of univariate analysis:

We can calculate the following measures of central tendency for Household Size:

 Mean (the average value): 3.8

 Median (the middle value): 4

These values give us an idea of where the “center” value is located.

We can also calculate the following measures of dispersion:

 Range (the difference between the max and min): 6

 Interquartile Range (the spread of the middle 50% of values): 2.5

 Standard Deviation (an average measure of spread): 1.87

These values give us an idea of how spread out the values are for this variable.

203
More methods for Uni Variate Analysis-

Frequency Distribution Tables

The frequency distribution table displays the frequency of each occurrence in the data.

Finding patterns is made simpler as the data is presented in a concise manner.

Example:

The list of IQ scores is: 118, 139, 124, 125, 127, 128, 129, 130, 130, 133, 136, 138, 141, 142,

149, 130, 154.

IQ Range Number

118-125 3

126-133 7

134-141 4

142-149 2

150-157 1

Bar Charts

When comparing several categories of data or distinct groupings of data, the bar graph is

highly useful. Monitoring alterations over time is useful. For showing discrete data, it works

best.

Histograms

Histograms show the same categorical variables against the category of data as bar charts do.

These categories are shown in histograms as bins that represent the number of data points in a

range. For displaying continuous data, it works well.

204
Pie Charts

Pie charts are typically used to see how a group is divided into more manageable parts. The

slices of the pie indicate the relative size of that particular category, and the entire pie

indicates 100%.

205
Frequency Polygons

A frequency polygon is used to compare datasets or show the cumulative frequency

distribution, much like histograms.

Bi- Variate Analysis

There are two variables in this situation since bi means two and variate imply variable. The

analysis focuses on the relationship between the two variables and the reason. The bivariate

analysis comes in three different flavours.

There are three common ways to perform the bivariate analysis:

1. Scatterplots.

2. Correlation Coefficients.

3. Simple Linear Regression.

206
Scatter Plot

Dots are used in a scatter plot to symbolise distinct data points. These charts make it simpler

to determine whether two variables are connected. The pattern that emerges reveals the nature

(linear or non-linear) and intensity of the link between the two variables.

Linear Correlation

The degree of a linear link between two numerical variables is shown by linear correlation.

There is no propensity to change in accordance with the values of the second quantity if there

is no correlation between the two variables.

207
Here, r measures the strength of a linear relationship and is always between -1 and 1 where -1

denotes perfect negative linear correlation and +1 denotes perfect positive linear correlation

and zero denotes no linear correlation.

3. Linear Regression

With straightforward linear regression, bivariate analysis can be carried out in a third

approach.

We select one variable to serve as an explanatory variable and the other variable to serve as a

response variable using this approach. Next, we identify the line that most closely "fits" the

dataset so that we can determine the precise relationship between the two variables.

For instance, the dataset's line of best fit looks like this:

Exam grade: 69.07 +3.85 (hours studied)

This indicates that an average exam score increase of 3.85 is correlated with each additional

hour of study. We can determine the precise correlation between study time and exam score

by applying this linear regression model.

208
3.1.2 Bivariate Analysis of two categorical Variables (Categorical-Categorical)

Chi-square Test

To ascertain the correlation between categorical variables, utilise the chi-square test. It is

determined by comparing the measured frequencies to the expected frequencies in one or

more frequency table categories. Two categorical variables are totally dependent on one

another when the likelihood is zero and completely independent when the probability is one.

Here, O stands for the observed value, E for the expected value, and subscript c denotes the

degrees of freedom.

Example- Let's imagine you want to determine if gender influences political party

preference in any way. To find out which political party respondents prefer, you

conduct a basic random sample poll of 440 voters. The table below displays the survey's

results:

Use the following instructions to conduct a Chi-Square test of independence to determine

whether gender is associated with political party preference.

Solution-

209
Step 1: Define the Hypothesis

H0: There is no link between gender and political party preference.

H1: There is a link between gender and political party preference.

Step 2: Calculate the Expected Values

Now you will calculate the expected frequency.

For example, the expected value for Male Republicans is:

Similarly, you can calculate the expected value for each of the cells.

Step 3: Calculate (O-E)2 / E for Each Cell in the Table

Now you will calculate the (O - E)2 / E for each cell in the table.

Where

O = Observed Value

210
E = Expected Value

Step 4: Calculate the Test Statistic X2

X2 is the sum of all the values in the last table

= 0.743 + 2.05 + 2.33 + 3.33 + 0.384 + 1

= 9.837

The crucial statistic must be identified before drawing any conclusions, which necessitates

knowing our degrees of freedom. The degrees of freedom in this situation are equal to the

product of the number of rows minus one and the number of columns minus one in the table,

or (r-1) (c-1). We have (3-1) (2-1) = 2

The key statistic from the chi-square table is the last statistic you compare our acquired

statistic to. As you can see, the critical statistic is less than our actual statistic of 9.83 and is

5.991 for an alpha level of 0.05 and two degrees of freedom. Because the crucial statistic is

higher than your obtained statistic, you can reject our null hypothesis.

This indicates that there is enough evidence to support your claim that political party

preference and gender are related.

211
3.1.3 Bivariate Analysis of one numerical and one categorical variable (Numerical-

Categorical)

Z-test and t-test

A statistical test called the Z test is run on data that roughly fits the normal distribution. For

assessing hypotheses, the z test can be applied to proportions, two samples, or one sample.

When the population variance is known, it determines whether or not the means of two big

samples differ.

Calculating if there is a significant difference between a sample and the population requires

the use of Z and T-tests.

If the probability of Z is small, the difference between the two averages is more significant.

Z Test Formula

212
In order to determine whether there is a difference between the means of two populations, the

z test formula compares the z statistic with the z critical value. The z critical value separates

the acceptance and rejection sections of the distribution graph in hypothesis testing. The null

hypothesis can be rejected if the test statistic is within the rejection region; otherwise, it

cannot be rejected. Below is the z-test formula for setting up the necessary hypothesis tests

for a one-sample and two-sample z-test.

One-Sample Z Test
When the population standard deviation is known, a one-sample z test is performed to

determine whether there is a discrepancy between the sample mean and the population mean.

The formula for the z-test statistic is given as follows:

The following algorithm is provided to set a one sample z test based on the z test statistic:

Left Tailed Test:

H0: The null hypothesis is that μ= μ0

Alternative Hypothesis: H1: μ < μ0.

Decision criteria: Reject the null hypothesis if the z statistic exceeds the z critical value.

Right-Tailed Test:

H 0: The null hypothesis is that μ= μ0.

213
Alternative Hypothesis: H1: μ > μ0

Decision Criteria: Reject the null hypothesis if the z statistic is greater than the z critical

value.

Two-Tailed Test:

H 0: The null hypothesis is that μ = μ0.

Alternative Hypothesis: H1: μ≠μ0

Decision Criteria: Reject the null hypothesis if the z statistic is greater than the z critical

value.

Two Sample Z Tests

A two sample z test is used to check if there is a difference between the means of two

samples. The z test statistic formula is given as follows:

Similar to the one-sample test, the two-sample z test can be set up. The means of the two

samples will be compared using this test, nevertheless. The null hypothesis, for instance, is

stated as H 0: μ 1 = μ 2.

Rejection Region for Null Hypothesis

214
Example- A gym trainer claimed that all the new boys in the gym are above average

weight. A random sample of thirty boys weight have a mean score of 112.5 kg and the

population mean weight is 100 kg and the standard deviation is 15. Is there sufficient

evidence to support the claim of the gym trainer?

Solution-

215
T-Tests

The t-test formula enables us to compare the average values of two data sets and ascertain

whether or not they represent the same population. The critical value from the t-table is used

to compare the t-score against. If the t-score is large, the groups are dissimilar, and if it is

small, the groups are similar.

The sample population is subjected to the t-test formula. The mean, variance, and standard

deviation of the data under comparison all affect the t-test formula. On the n number of

samples that were gathered, three different sorts of t-tests may be run.

 One-sample test,

216
 Independent sample t-test and

 Paired samples t-test

The degree of freedom (df = n-1) and the accompanying value are found using the t-table to

determine the critical value (usually 0.05 or 0.1). The initial premise is incorrect, and we infer

that the results are significantly different if the t-test yielded statistically > CV.

If the sample size is large enough, then we use a Z-test, and for small sample size, we use a

T-test.

T Test Formula-

217
Example- Find the t-test value for the following two sets of values: 7, 2, 9, 8 and 1, 2, 3,

4?

Solution-

Formula for mean-

Formula for standard deviation-

Construct the following table for standard deviation-

218
Standard deviation for the first set of data: S1 = 3.11

Number of terms in second set: n2 = 4

Mean for second set of data:

3.1.4 ANALYSIS OF VARIANCE (ANOVA)

When more than two groups' averages are statistically different from one another, the

ANOVA test is performed to assess whether there is a significant difference between them.

219
This comparison of averages of a numerical variable for more than two categories of a

categorical variable is appropriate.

Example of Anova
The following data is given:

Standard
Types of Animals Number of animals Average Domestic animals
Deviation

Dogs 5 12 2

Cats 5 16 1

Hamsters 5 20 4

Calculate the Anova coefficient.

Solution:
Construct the following table:

Animal name n x s s2

Dogs 5 12 2 4

Cats 5 16 1 1

Hamster 5 20 4 16

p=3
n=5

220
N = 15
x̄ = 16
SST = ∑n (x−x̄)2

SST= 5(12−16)2+5(16−16)2+11(20−16)2

= 160

MST=SST/p-1

MST=160/3-1

MST=80

SSE = ∑ (n−1) s2

SSE = 4 × 4 + 4 × 1 + 4 × 16

SSE = 84

MSE=SSE/N-p

MSE=8415/38415-3

MSE=7

F=MST/MSE

F=80/7

F=11.429

221
3.1.5 Multivariate Analysis

When more than two variables must be studied at once, multivariate analysis is necessary.

Multivariate analysis is used to examine more complicated sets of data because it is

extremely difficult for the human brain to visualise a relationship among four variables on a

graph. Cluster analysis, factor analysis, multiple regression analysis, principal component

analysis, etc. are examples of multivariate analysis types. There are more than 20 alternative

approaches to multivariate analysis, and which one to choose depends on the data set and the

desired outcome. The most typical methods include:

Cluster Analysis

Different objects are grouped together using cluster analysis so that there is a maximum

similarity between objects belonging to the same group and a minimum similarity in all other

cases. When the measure is distance or similarity and the rows and columns of the data table

represent the same units, it is employed.

Principal Component Analysis (PCA)

A data table with several interconnected metrics is reduced in dimension using principal

component analysis, or PCA. The original variables in this case are changed into a fresh set

of variables called the "Principal Components" in Principal Component Analysis.

The dataset that demonstrates multicollinearity is analysed using PCA. The gap between

variances and their true value can be very wide, despite the bias in least squares estimations.

As a result, PCA increases some bias and decreases the regression model's standard error.

222
Correspondence Analysis

Using information from a contingency table, correspondence analysis can be used to reveal

relative relationships between and among two different groupings of variables. A contingency

table is a 2D table comprising groups of variables in the rows and columns.

3.1.6 Application of Uni-Variate, Bi-Variate, and Multi-Variate Analysis-

Multivariate analysis is a simultaneous study of several variables. Compared to univariate

analysis, it is more illuminating. But it's also more intricate than a single-variate study.

Multivariate statistics analyse three or more variables simultaneously to look for any potential

interactions.

Application areas

Social science: (gender, age, Nationality) of an individual

223
Climatology: (minimum temperature, maximum temperature. rainfall, humadity) on a day

Econometrics: (input costs, production, profit) of a firm

Key Terms- Statistics, Hypothesis, Significance, Tests, Uni variate, Bi variate, Multi Variate,

Chi square tests.

Summary- In this unit, areas related to testing, types of testing hypothesis testing,

Univariate, Multivariate and Bivaraiate analysis, ANOVA, etc, have been covered in detail.

Case Study:

A FMCG firm wanted to investigate the effects of four different training programmes

on the sales ability of their salespeople. Thirty-two participants were randomly

separated into four equal-sized groups and then subjected to the various sales training

programmes. Because some students dropped out throughout the training programmes

owing to illness, vacations, and other reasons, the number of trainees who completed the

programmes varied by group. At the end of the training programmes, each salesperson

was assigned a sales area at random from a group of sales areas judged to have similar

sales potential. The table shows the sales made by each of the four groups of salespeople

during the first week after finishing the training programme:

224
Questions for Discussion-

1. Use the proper approach to analyse the experiment.

2. Identify and investigate any noteworthy effects of treatments or factors of interest to

the researcher.

3. What practical implications does this experiment have?

4. Write a paragraph describing the findings of your analysis.

Exercise: Choose the correct option

1. A statement made about a population for testing purpose is called?

a) Statistic

b) Hypothesis

c) Level of Significance

d) Test-Statistic

2. If the assumed hypothesis is tested for rejection considering it to be true is called?

a) Null Hypothesis

b) Statistical Hypothesis

c) Simple Hypothesis

d) Composite Hypothesis

3. A statement whose validity is tested on the basis of a sample is called?

a) Null Hypothesis

b) Statistical Hypothesis

c) Simple Hypothesis

d) Composite Hypothesis

225
4. A hypothesis which defines the population distribution is called?

a) Null Hypothesis

b) Statistical Hypothesis

c) Simple Hypothesis

d) Composite Hypothesis

5. If the null hypothesis is false then which of the following is accepted?

a) Null Hypothesis

b) Positive Hypothesis

c) Negative Hypothesis

d) Alternative Hypothesis.

6. What is the difference between a bar chart and a histogram?

a) A histogram does not show the entire range of scores in a distribution

b) Bar charts are circular, whereas histograms are square

c) There are no gaps between the bars on a histogram

d) Bar charts represents numbers, whereas histograms represent percentages

7 What is an outlier?

a) A type of variable that cannot be quantified

b) A score that is left out of the analysis because of missing data

c) An extreme value at either end of a distribution

8. What is the function of a contingency table, in the context of bivariate analysis?

226
a) It shows the results you would expect to find by chance

b) It summarises the frequencies of two variables so that they can be compared

c) It lists the different levels of p value for tests of significance

d) It compares the results you might get from various statistical tests

9. When might it be appropriate to conduct a multivariate analysis test?

a) If the relationship between two variables might be spurious

b) If there could be an intervening variable

c) If a third variable might be moderating the relationship

d) All of the above

10. What is the name of the test that is used to assess the relationship between two

ordinal variables?

a) Spearman's rho

b) Phi

c) Cramer's V

d) Chi square

Short Answer Type Questions-

1. What is the multivariate analysis?

2. How many types of analysis are there?

227
3. What is Uniariate analysis?

Long Answer Type Questions-

1. Explain Anova using an example.

2. With the help of suitable example, explain Chi-Square Test?

3. What are Z test and T test? Explain in detail with suitable example.

Answers

Exercise: MCQ

1. (B)

2. (A)

3. (B)

4. (C)

5. (D)

6. (C)

7. (C)

8. (B)

9. (D)

10. (A)

228
3.2. Basics of Hypothesis Testing

3.2.1 Introduction to Hypothesis and its types

A preliminary correlation between two or more variables is called a hypothesis. These

variables are connected to different elements of the study question. A testable prediction is

what a hypothesis is. A statement is examined in the research to see if it is true or incorrect in

order to determine its veracity. Diverse facets of the research issue must be investigated by

the researcher. As a result, a researcher makes an assumption about a potential correlation

between variables relevant to each part of the research issue. The researcher might investigate

several areas of the research by testing these linkages. A hypothesis is a potential explanation

for the relationship between the variables.

A researcher intentionally develops hypotheses since it is challenging to begin research

without a solid foundation. As a result, the researcher establishes logical connections between

or among the research's variables. The correlations between these variables, which are

connected by a common theme, provide the research's framework. These logical connections

or falsifiable presumptions provide the researcher with a starting point for the inquiry.

The decision-maker is given this tool through hypothesis testing. The operations manager

would take a sample of filled bottles from the ongoing bottling process if he were to employ

this instrument. The strength of the evidence the sample of bottles produced will be

considered in the evaluation;

A hypothesis that needs to be tested is the implicit assertion (μ = 1,000 cm 3), and the

statistical method that enables us to do so is known as hypothesis testing or testing of

hypotheses.

229
The following hypotheses, for instance, would be developed by a researcher investigating

"Discrimination Against Women in a Rural Society"

 Discrimination against women will increase in illiterate societies.

 The more patriarchy there is in a society, the more discrimination against women

there will be.

 The more traditional traditions present in a culture, the more discrimination against

women there will be.

The following hypotheses might be developed by a researcher whose focus is "Extent of

Use of Family-Planning Practice in an Area": The higher the standard of education, the

higher the use of the family-planning practice.

 The higher the availability of family-planning services, the higher the use of family

planning practice will be.

 The higher the standards of living, the higher the use of the family-planning practice

will be.

Characteristics of Hypothesis

1. Empirically Testable

2. Simple and Clear

3. Specific and relevant to the theme of research

4. Predictable

5. Manageable

Importance of Hypothesis

230
1. It provides the research with a focus.

2. It facilitates exploring different facets of the research.

3. It identifies the researcher's emphasis because, in the absence of hypotheses, the research

may potentially concentrate on unimportant and undesired components of the study.

4. It aids in the creation of research methodology.

5. It stops research from being done in a blind manner.

6. It guarantees the correctness and accuracy of the research's findings.

7. It saves time, money, and energy because the researcher wouldn't have to focus these

resources on extraneous aspects of the study if there were no hypotheses.

3.2.2 Types of Hypothesis

The types of hypotheses are as follows:

1. Simple Hypothesis

2. Complex Hypothesis

3. Working or Research Hypothesis

4. Null Hypothesis

5. Alternative Hypothesis

6. Logical Hypothesis

7. Statistical Hypothesis

1. Simple Hypothesis

231
Any hypothesis that indicates a link between two variables—the independent and dependent

variables—is referred to as a simple hypothesis.

Examples: The rate of crime in society would increase as unemployment increases.

Poorer fertiliser use would result in lower agricultural productivity.

A society's crime rate would be higher the poorer it was.

2. Complex Hypothesis

A hypothesis that indicates a relationship between more than two variables is referred to as a

complex hypothesis.

Examples:

1. The rate of crime will increase as poverty and illiteracy in society rise (three variables -

two independent variables and one dependent variable).

2. The agricultural productivity will increase if fertilisers, better seeds, and modern

equipment are used more frequently (Four variable-three independent variables and one

dependent variable).

3. Poverty and crime rates increase with the level of illiteracy in a culture. (Two dependent

variables and one independent variable total three variables)

3. Working Hypothesis.

A working hypothesis is one that has been approved for testing and development during the

investigation. It is a theory that is presumptively appropriate to explain certain facts and the

connections between various occurrences. This hypothesis is accepted to be tested for study

since it is anticipated that it will lead to a useful theory.

Any hypothesis that is originally accepted for consideration in the research is acceptable.

232
4. Alternative Hypothesis

A new hypothesis (to replace the working hypothesis) is produced and tested to examine the

desired feature of the research if the working hypothesis is incorrect or rejected. This new

hypothesis is referred to as an alternate hypothesis.

As implied by the name, it is a different hypothesis (or connection) that is used when the

working hypothesis is unable to produce the necessary theory. H₁ is the alternative

hypothesis.

5 Null Hypothesis

A hypothesis that expresses no link between variables is known as a null hypothesis. It

disproves the relationship between the variables.

Examples: The amount of crime in a society is unrelated to poverty.

In society, the rate of unemployment has little to do with illiteracy.

The aim of a null hypothesis is clear. Making a null hypothesis with the purpose to

disapprove, reject, or nullify it allows the researcher to confirm the existence of a relationship

between the variables. In order to establish that there is a relationship between the variables, a

null hypothesis is typically created as a reverse technique. 𝐻0 denotes the Null Hypothesis.

6 Statistical Hypothesis

A statistical hypothesis is a hypothesis that can be statistically tested. Any theory that has the

ability to be statistically validated is acceptable. It implies that it can be evaluated using

quantitative methods. A statistical hypothesis can also be stated to have measurable variables

or to be capable of being turned into quantifiable indications for statistical testing.

7. Logical Hypothesis

233
A logical hypothesis is one that can be substantiated rationally. It is a relationship that may be

expressed as a hypothesis, and its veracity can be established by connecting its interlinks

using logical justifications. It can be supported by logical evidence to prove it. It doesn't

always follow that statistical methods cannot be used to verify a logical hypothesis. It might

or might not be statistically verifiable, but in light of the logical arguments, it looks so

probable that these logical factors are sufficient to verify it.

3.2.3 Type I & Type II Errors

The next stage is to acquire data from a representative sample of the population after

outlining the null and alternative hypotheses. The fact that we cannot be certain of our

interferences with 100% certainty is a significant constraint. Since differences between

samples cannot be completely eliminated until the sample size equals that of the population,

it is possible that the result reached is flawed and causes an error. There are two different

kinds of errors, as may be seen in the Table below.

Type I and Type II Errors of Hypothesis Testing

Type I Error

234
A Type I error occurs when a true null hypothesis is incorrectly rejected during statistical

testing. The operations manager would be making a type I error if he rejected 𝐻0 and

concluded that the process had gotten out of hand when in fact it had not.

When a null hypothesis is disregarded during the hypothesis testing procedure even though it

is true and should not be disregarded, this is known as a type I error.

A null hypothesis is established prior to the start of a test in hypothesis testing. In some

circumstances, the null hypothesis makes the assumption that there is no causal connection

between the test item and the stimuli being given to the test subject in order to cause an

outcome to the test.

However, mistakes can happen where the null hypothesis is rejected, indicating that a cause-

and-effect link exists between the testing variables when in fact, a false positive occurred.

Type I errors are what are known as these false positives.

Understanding a Type I Error

A hypothesis is tested using sample data in a process known as hypothesis testing. The goal

of the test is to demonstrate that the conjecture or hypothesis is supported by the inputted

data. The idea that there is no statistically significant relationship between the two data sets,

variables, or populations under consideration in the hypothesis is known as a null hypothesis.

A researcher would typically strive to refute the null hypothesis.

Consider the case where the null hypothesis holds that an investment plan doesn't outperform

a market index like the S&P 500. To find out if the investment strategy performed better than

the S&P, the researcher would test the historical performance of the method using samples of

235
data. The null hypothesis would be disproved if the test's outcomes revealed that the approach

outperformed the index.

n=0 is used to indicate this circumstance. The null hypothesis, which states that the stimuli do

not impact the test subject, would then need to be rejected if, after the test is completed, the

results appear to indicate that the stimuli administered to the test subject induced a reaction.

If a null hypothesis is discovered to be true, it should never be rejected, and if it is found to be

untrue, it should always be rejected. Errors can, however, arise under some circumstances.

False Positive Type I Error

It is occasionally wrong to reject the null hypothesis, which states that there is no connection

between the test subject, the stimuli, and the result. A "false positive" result occurs when it

appears that the stimuli had an effect on the subject but the outcome was random and

something other than the stimuli was responsible. A type I error is what is referred to as this

"false positive," which results in an inaccurate rejection of the null hypothesis. A type I error

results in the rejection of a proposition that wasn't warranted.

Examples of Type I Error

Let's take the trial of a criminal suspect as an example. The person's innocence is the null

hypothesis, and guilt is the alternative. In this instance, a type I error would result in the

person being found guilty even though they were innocent and being imprisoned.

In medical testing, a type I error could provide the impression that treatment is lessening the

severity of the condition when it actually has no such impact. The null hypothesis in the

testing of a novel drug is that the drug has no effect on how the disease develops. Say a lab is

looking into a brand-new cancer medication. The medicine does not impact the rate at which

cancer cells grow, according to their null hypothesis.

236
The cancer cells cease growing after the medicine has been applied to them. Thus, the null

hypothesis that the medicine would have no impact would be rejected by the researchers. In

this instance, rejecting the null would be the correct conclusion if the medicine was the

reason of the growth halt. However, this would be an example of an inaccurate rejection of

the null hypothesis if anything else during the test caused the growth slowdown rather than

the medicine supplied (i.e., a type I error).

Type II Error

When one fails to reject a null hypothesis that is actually wrong, this error is known

statistically as a type II error. This term is used in the context of hypothesis testing. A type II

error, often called an error of omission, results in a false negative. When the patient is

infected, a disease test, for instance, can return a negative result. This is a type II error

because, despite being wrong, we accept the test's negative conclusion.

A type I error in statistical analysis is when a genuine null hypothesis is rejected, whereas a

type II error is when a false null hypothesis is not correctly rejected. Despite the fact that the

error does not happen by accident, it rejects the alternative theory.

Understanding a Type II Error

A type II error, often referred to as a second-kind error or a beta error, validates a hypothesis

that ought to have been disproven, such as the assertion that two observations are the same

despite the fact that they are not. Even when the alternative hypothesis represents the actual

state of nature, a type II error does not reject the null hypothesis. To put it another way, an

erroneous conclusion is accepted as fact.

237
By establishing more strict standards for rejecting a null hypothesis, type II errors can be

minimised. If an analyst, for instance, considers anything that falls within the +/- bounds of a

95% confidence interval to be statistically insignificant (a negative result), then lowering that

tolerance to +/- 90% and then narrowing the bounds will result in fewer negative results and

lower the likelihood of a false negative. However, following these instructions tends to

increase the likelihood of running into a type I error—a false-positive outcome. The

likelihood or risk of committing a type I error or type II error should be taken into account

while conducting a hypothesis test.

Type II error is the incorrect choice to accept (rather than reject, to be more precise) an

erroneous null hypothesis. The operations manager would be making a type II error if he did

not reject 𝐻0 and assumed that the process was under control when it had actually gotten out

of control. The number of type I and type II errors should be kept to a minimum because they

are both undesirable. Let's examine how to reduce the likelihood of type I and type II errors.

It may be clear that it is possible to completely eliminate the likelihood of type I mistake,

even with faulty sample evidence. Regardless of the evidence, simply accept the null

hypothesis.

We will never make a type I error since we will never reject any null hypotheses, including a

genuine null hypothesis. It is clear that this would be unwise, though. If we always accept a

null hypothesis, we will undoubtedly accept any false null hypothesis that is presented,

regardless of how absurd it may be. In other words, the likelihood that we will make a type II

error is 1. Similarly, we believe it would be unwise to constantly reject a null hypothesis to

lower the likelihood of type II error to zero because this would mean rejecting every genuine

null hypothesis, regardless of how accurate it is. We will have a type I error probability of 1.

238
Therefore, we cannot and should not try to completely avoid either type of error. We should

plan, organize, and settle for some small, optimal probability of each type of error.

Examples of Type II Error

Let's say a biotechnology company wishes to assess the effectiveness of two of its diabetes

medications. The two medicines are equally effective, according to the null hypothesis. The

claim that the corporation seeks to disprove with the one-tailed test is a null hypothesis, H0.

The counterargument, Ha, claims that the two medications are not equally effective. The

natural condition that is supported by rejecting the null hypothesis is represented by the

alternative hypothesis, Ha.

To compare the therapies, the biotech business conducts a significant clinical trial involving

3,000 diabetic patients. The 3,000 patients are randomly split into two groups of equal size,

with one group receiving one treatment while the other receives the other treatment. It

chooses a significance level of 0.05, indicating that it is prepared to accept a 5% probability

that it will reject the null hypothesis even if it is true or a 5% chance that it will make a type I

error.

Assume that 2.5%, or 0.025, is the beta value. Consequently, the likelihood of making a type

II error is 97.5%. The null hypothesis should be disproved if the two drugs are not equivalent.

However, a type II error happens if the biotech company does not reject the null hypothesis

when the medications are not equally successful.

Type I Errors vs. Type II Errors

A type I error rejects the null hypothesis even when it is true, in contrast to a type II error,

which does not (i.e., a false positive). The level of significance chosen for the hypothesis test

is equivalent to the likelihood of making a type I error. Therefore, there is a 5% probability

that a type I mistake could happen if the threshold of significance is 0.05.

239
The probability of making a type II error, commonly known as beta, is one minus the test's

power. Increasing the sample size would boost the test's power while lowering the likelihood

that a type II mistake would be made.

3.2.4 Level of Significance

The most common policy in statistical hypothesis testing is to establish a significance level,

denoted by α, and to reject 𝐻0 when the p-value falls below it. When this policy is followed,

one can be sure that the maximum probability of type I error is α.

Policy: When the p-value is less than α, reject 𝐻0 .

In other words, we can say that the rejection region for 𝐻0 is the area under the curve where

the p-value is less than α. This region is also called critical region.

The standard values for α are 10%, 5%, and 1%. Suppose α is set at 5%. In the preceding

example, for a sample mean of 1,000.5, the p-value was 16%, and 𝐻0 will not be rejected. For

a sample mean of 1001, the p-value will be 2.28%, which is below α = 5%. Hence 𝐻0 will be

rejected.

Let us analyze in some detail the implications of using a significance level α for rejecting a

null hypothesis.

 The first thing to note is that if we do not reject 𝐻0 , this does not prove that 𝐻0 is true. For
example, if α = 5% and the p-value = 6%, we will not reject 𝐻0 . But there is only about

6% chance that 𝐻0 is true, which is hardly proof that 𝐻0 is true. It may be possible that 𝐻0

is false and by not rejecting it, we are committing a type II error. For this reason, we

should say "We cannot reject 𝐻0 at an α of 5%" rather than "We accept 𝐻0 ."

240
 The second thing to note is that α is the maximum probability of type I error we set for

ourselves. Since α is the maximum p-value at which we reject 𝐻0 , it is the maximum

probability of committing a type I error. In other words, setting α = 5% means that we are

willing to put up with up to 5% chance of committing a type I error.

 The third thing to note is that the selected value of α indirectly determines the probability

of type II error as well. In general, other things remaining the same, increasing the value

of α will decrease the probability of type II error. This should be intuitively obvious. For

example, increasing α from 5% to 10% means that in those instances with a p-value in the

range 5% to 10% the 𝐻0 that would not have been rejected before would now be rejected.

Thus, some cases of false 𝐻0 that escaped rejection before may not escape now. As a

result, the probability of type II error will decrease.

 The fourth thing to note about α is the meaning of (1 - α). If we set α = 5%, then (1 - α) =
95% is the minimum confidence level that we set in order to reject 𝐻0 . In other words, we

want to be at least 95% confident that 𝐻0 is false before we reject it.

3.2.5 Acceptance and Rejection Region

The interval "inside the sample distribution of the test statistic that is consistent with the null

hypothesis 𝐻0 from hypothesis testing" is known as the acceptance area.

Let's put it another way: Suppose you perform a z-test-style hypothesis test. The test's results

are presented as a z-value, which has a wide range of potential values. Some values will fall

within an interval that shows the null hypothesis is true within that range of values. The

acceptance region is that space.

241
Due to the fact that a hypothesis test cannot tell you which hypothesis is true (the alternate or

null hypothesis) or even which is most likely true, you must provisionally accept the null

hypothesis.

It just examines if your data provide enough support to reject the null hypothesis. The null

hypothesis is not true if the alternative hypothesis is not accepted.

Let's imagine if our "experiment" had you catching a kid in the act of stealing a cookie:

 Null hypothesis (H0): The child didn’t steal the cookie (innocent until proven guilty!).

 Alternate hypothesis (H1): The child did steal the cookie.

You have a good feeling the kid took the cookie. But even with all the evidence gathered, you

can't be certain that the youngster is guilty. The alternative theory that the youngster is

culpable is therefore unsupported by sufficient data. To put it another way, you can't accept

the null hypothesis that the child is guilty and reject the null hypothesis that the child is

innocent. This does not imply that the youngster is a victim. Simply put, you lacked sufficient

proof to hold them accountable. The null hypothesis of innocence is not actually "accepted,"

despite the fact that your result fell within the acceptance range. Simply accept it on a

temporary basis—possibly grudgingly—and release the child from custody.

Later on you might find crumbs in their bed, leading you to revisit your findings.

Rejection Region

A rejection region, also known as a critical region, is a region of the parameter space where a

null hypothesis will be rejected if a result is seen that falls inside it. Typically, a significance

test is used in hypothesis testing, and the rejection region is provided as a statistic, such as a t

score or a z score. The value of the real (non-standardized) parameter can be specified just as

simply.

242
Since a rejection zone is just a more precise way of expressing the significance threshold,

they are identical. A standardised score directly specifies the area under the parameter

distribution that will result in rejection because it is expressed in terms of standard deviations

(e.g., 1.644 SD or.1.96 SD). This information can be easily converted into alpha (α) during

the planning stage or a p-value after the test is finished by computing the cumulative

distribution function.

The observed p-value is compared to the crucial Z score, which is located at the edge of the

rejection region, once the test is finished, and if it falls within the region, the null hypothesis

is rejected.

Notably, if a value is outside the rejection zone, it does not always follow that the null

hypothesis can be accepted; rather, it only indicates that there is insufficient data to support

its rejection.

The observed p-value is compared to the crucial Z score, which is located at the edge of the

rejection region once the test is finished, and if it falls within the region, the null hypothesis is

rejected.

If your test findings fall into a particular section of a graph known as a rejection region (also

known as a critical region), you would reject the null hypothesis. In other words, your results

are statistically significant if they fall within that range.

The rejection regions in a two-tailed t-distribution.

243
Testing ideas or experimental results is the main goal of statistics. For instance, you might

have created a brand-new fertiliser that you believe causes plants to grow 50% more quickly.

1. Your experiment must be able to be repeated in order to demonstrate that your theory is

accurate.

2. Be compared to a well-known plant fact (in this example, probably the average growth rate

of plants without fertilizer).

A hypothesis test is the name given to this kind of statistical analysis. The testing procedure

includes the rejection region. In particular, it is a branch of probability that indicates the

likelihood that your theory (or "hypothesis") is correct.

The alpha level you decide to accept as a researcher is up to you. For instance, if you wanted

to have a 95% confidence level in the significance of your data, you would select a 5% alpha

level (100% - 95%). The rejection zone is between those 5% levels. The 5% in a one-tailed

test would be in one tail. The rejection region for a two-tailed test would be in both tails.

A one tailed test with the rejection region in one tail.

244
3.2.6 Procedure of Hypothesis Testing

A number of crucial ideas about hypothesis testing have been introduced to us. We can now

establish a standard testing approach in a more organised manner. By this point, it should be

evident that the process of testing a hypothesis essentially consists of two stages. In the first

stage, we design the experiment and establish the criteria for rejecting the null hypothesis. In

the second stage, we determine if the null hypothesis can be rejected using the sample data.

The specific procedures are as follows:

Step 1: State the Null and the Alternate Hypotheses. i.e. H0 and H1

Step 2: Specify a level of significance α

Step 3: Choose the test statistic and define the critical region in terms of the test statistic

Step 4: Make necessary computations

calculate the observed value of the test statistic

find the p- value of the test

Step 5: Decide to accept or reject the null hypothesis either

by comparing the p- value with α or

by comparing the observed value of the test statistic with the cut- off value or the critical

value of the test statistic.

245
Key Terms- Hypothesis Testing, Type 1 and Type 2, Level of Significance, Region,

Statistics, Range, Procedure of hypothesis.

Summary- This unit deals with the hypothesis testing and its types, Null Hypothesis,

Alternate Hypothesis, and its procedures in details, level of significance, acceptance and

rejection region as well.

Case Study:

The slogan "made in China" has become a source of anxiety in recent years, as Indian

businesses seek to shield their products from foreign competition. In recent years, India

has seen a significant trade imbalance as a result of a flood of imported goods that enter

the nation and are offered at lower prices than comparable Indian-made items. One

major source of concern is electronic products, with total imported items continuously

increasing from the 1990s to the 2004s. Concerned about product quality concerns,

worker layoffs, and excessive costs, Indian corporations have spent millions on

advertising to manufacture electronic goods that will meet consumer desires. To

simplify the analysis, we have coded the year using the coded variable x = Year 1989.

246
Questions for Discussion-

1. Determine the least-squares line for estimating import volume as a function of year

for the years 1989-2000.

2. Is there a significant linear relationship between import volume and year?

3. Predict the volume of commodities imported in each of the years 2002, 2003, and 2004

using 95% prediction ranges.

4. Are the forecasts made in Step 4 realistic approximations of the actual values

observed in these years? Explain.

5. Enter the 1989-2004 data into your database and compute the regression line. What

influence did the new data points have on the slope> What effect does SSE have?

6. Does a straight line appear to be an accurate model for the data given the form of the

scattered diagram for the years 1989-2004? What other model would be more

appropriate?

247
Exercise: Choose the correct option

1. The rejection probability of Null Hypothesis when it is true is called as?

a) Level of Confidence

b) Level of Significance

c) Level of Margin

d) Level of Rejection

2. The point where the Null Hypothesis gets rejected is called as?

a) Significant Value

b) Rejection Value

c) Acceptance Value

d) Critical Value

3. If the Critical region is evenly distributed then the test is referred as?

a) Two tailed

b) One tailed

c) Three tailed

d) Zero tailed

4. The type of test is defined by which of the following?

a) Null Hypothesis

b) Simple Hypothesis

c) Alternative Hypothesis

d) Composite Hypothesis

5. Which of the following is defined as the rule or formula to test a Null Hypothesis?

a) Test statistic

248
b) Population statistic

c) Variance statistic

d) Null statistic

6. Consider a hypothesis H0 where ϕ0 = 5 against H1 where ϕ1 > 5. The test is?

a) Right tailed

b) Left tailed

c) Center tailed

d) Cross tailed

7. Consider a hypothesis where H0 where ϕ0 = 23 against H1 where ϕ1 < 23. The test is?

a) Right tailed

b) Left tailed

c) Center tailed

d) Cross tailed

8. Type 1 error occurs when?

a) An individual rejects H0 if it is True

b) An individual rejects H0 if it is False

c) An individual accepts H0 if it is True

d) An individual accepts H0 if it is False

9. The probability of Type 1 error is referred as?

a) 1-α

b) β

c) α

d) 1-β

249
10. Alternative Hypothesis is also called as?

a) Composite hypothesis

b) Research Hypothesis

c) Simple Hypothesis

d) Null Hypothesis

Short Answer Type Questions-

1. What is Hypothesis testing?

2. Write down the types and importance of Hypothesis?

3. What is Acceptance and Rejection Region?

Long Answer Type Questions-

1. Explain Type-I and Type-II error with an appropriate example.

2. Explain Level of Significance in detail?

3. Write down the procedure of Hypothesis.

Answers

Exercise: MCQ

1. (B)

2. (D)

250
3. (A)

4. (C)

5. (A)

6. (A)

7. (B)

8. (A)

9. (C)

10. (B)

3.3 Applications of Hypothesis Testing

3.3.1 Real-life applications of Hypothesis testing

Here are some real-life applications of Hypothesis testing

To examine the production procedures:

251
The use of hypothesis testing in manufacturing processes includes establishing if the

implementation of a new technique or procedure in the manufacturing facility was the source

of any anomalies in the product's quality or not. Suppose manufacturing plant X checks

whether a certain method increases the number of defective products produced each quarter;

say this number is 200. To confirm this, the researcher must now compute the mean of the

number of defective items created before the start and end of the quarter.

The depiction of this example's hypothesis testing is provided below.

Null Hypothesis (Ho): Before and after using the new production process, the average

number of defective goods produced is equal, or μ after = μ before.

Alternative Hypothesis (Ha): Before and after the new manufacturing procedure was

implemented, the average quantity of defective goods produced varied, i.e. μ after ≠ μ before.

The null hypothesis is disapproved and it may be concluded that changes in the method of

production cause a rise in the number of defective items produced each quarter if the

resulting p-value of the hypothesis testing is less than the significant value, i.e. α =.05.

252
To Plan the Marketing Strategies

To ascertain the effect of recently implemented marketing strategies, campaigns, or other

strategies on the sales of the product, many organisations frequently use hypothesis testing.

For instance, the company's marketing division believed that increasing their investment in

digital advertisements would increase sales. The marketing department may increase the

budget for digital advertisements for a specific time period in order to test this assumption

and then at the conclusion of that time period, analyse the data that was gathered. To validate

their presumption, they must do hypothesis testing. Here,

Null Hypothesis (Ho): The average sales are the same before and after the increase in the

digital advertisement budget, or μafter = μbefore.

Alternative Hypothesis (Ha): After an increase in the budget for digital advertisements, i.e.,

"μafter >μbefore," average sales rise.

The marketing department can reject the null hypothesis and come to the conclusion that

increasing the budget for digital advertising may increase product sales if the P-value is less

than the significant value, such as.05.

253
In clinical Trials

Hypothesis testing is frequently used by physicians and pharmacists in therapeutic trials.

Hypothesis testing is used to analyse the effects of new therapeutic techniques, medications,

or procedures on patient health. For instance, a pharmacist thinks that the new medication is

causing diabetic patients' blood pressure to increase. The researcher must check the sample

patients' blood pressure before and after taking the new medication for roughly a specific

time period, say one month, in order to test this hypothesis. The hypothesis testing process is

then carried out as follows:

Null Hypothesis (Ho): The average blood pressure is the same before and after taking the

medication, or μafter = μbefore

Alternative Hypothesis (Ha): The average blood pressure before and after taking the

medication is lower than the average blood pressure afterwards, or μafter<μbefore

respectively.

The null hypothesis is rejected if the p-value of the hypothesis test is less than the

significance value (let's say.o5), at which point it can be deduced that the new medication is

to blame for the rise in diabetic patients' blood pressure.

254
In Testing Effectiveness of Essential Oils

Due to their multiple advantages, essential oils are becoming more and more popular

nowadays. Many essential oils, including ylang-ylang, lavender, and chamomile, make this

claim. You might want to investigate the real healing potential of each of these oils. Let's say

you believe that lavender essential oil can help you feel less anxious and stressed. You can

conduct hypothesis testing by restating the hypothesis as follows to examine this supposition:

Null Hypothesis (Ho): No effect of lavender essential oil on lowering anxiety

Alternative Hypothesis (Ha): Lavender oil aids in lowering anxiety

Group A, or the experimental group, in this experiment receives lavender oil, whereas group

B, or the control group, receives a placebo. The data is then gathered using different

statistical tools, and both the experimental group and the control group's stress levels are then

analysed. Following the calculation, the p-value and significance threshold are discovered to

be 0.25 and 0.05, respectively. The null hypothesis is rejected because the p values are less

than the significance values, and it is therefore clear that lavender oil is effective at lowering

people's levels of stress.

255
In Testing Fertilizer’s Impact on Plants

The effect of pesticides, fertilisers, and other substances on the growth of plants or animals is

now also investigated using hypothesis testing. Let's say a researcher wishes to confirm his

hypothesis that a certain fertiliser may cause a plant to grow more quickly in a month than its

typical growth of 10 inches. He constantly applied that fertiliser to the plant for over a month

to confirm this presumption. The mathematical process for testing the hypothesis in this

situation is as follows:

Null Hypothesis (Ho): There is no relationship between the fertiliser and plant growth. that

is, Equals 20 inches

Alternative Hypothesis (Ha): The fertiliser causes the plant to grow more quickly,

measuring > 20 inches.

The null hypothesis can be rejected if the p-value for the hypothesis test is less than the level

of significance, let's say.05, at which point you can draw the conclusion that the specific

fertiliser is what's causing the faster plant growth.

256
In Testing the Effectiveness of Vitamin E

Let's say the researcher believes that vitamin E contributes to hair growing more quickly. In

an experiment he conducted, the experimental group received vitamin E for three months,

whereas the control group received a placebo. After three months have passed, the results are

then analysed. He restates his theory as follows to support his initial assumption:

The null hypothesis (Ho): It states that there is no correlation between vitamin E and the

sample group's hair growth, i.e., μafter = μbefore

Alternative Hypothesis (Ha): Under the assumption that all other factors remain constant,

the group of individuals who received vitamin E experiences faster hair growth than the

group as a whole did before doing so. In this case, μafter > before

Following statistical analysis, the significance level and the p-value, in this case, are,

respectively, o.o5 and 0.20. The researcher can therefore draw the conclusion that consuming

vitamin E causes quicker hair growth.

257
In Testing the Teaching Strategy

Let's imagine that Mr. X and Mr. Y, the two teachers, disagree on the ideal teaching

approach. Mr. Y contends that the weekly test is a waste of time and will not affect the

performance of the students in the yearly examinations, contrary to Mr. X's claim that giving

the children the weekly tests will improve their performance in the annual exams. Now we

may test our hypotheses to see which of the two is correct. The researcher could put forward

the following hypothesis:

The null hypothesis (Ho): The average marks earned by the kids when they took the weekly

examinations and when they didn't were the same, proving the null hypothesis (Ho) that there

is no correlation between the weekly tests and how well the kids perform in the annual

exams. (μBefore = μAfter)

Alternative Hypothesis (Ha): The children will score better on the yearly examinations if

they are required to take the weekly tests in addition to the annual exams or if μafter is

followed by μbefore.

The researcher can draw the conclusion that if the weekly assessment system is applied, the

children will score better on their annual tests if the p-value of the hypothesis testing is less

than the level of significance, say.05.

258
When examining the underlying premise of intelligence

Let's say a principal claims that the IQ level of the pupils attending her school is above

average. The researcher might select a sample of about 50 randomly chosen pupils from that

institution to back up her claim. Let's assume that those kids have an average IQ of roughly

110 and that the mean population as a whole has an average IQ of 100 with a standard

deviation of 15. The results of the hypothesis testing are as follows:

The null hypothesis (Ho): The population mean IQ score of 100 is a known fact, hence the

null hypothesis is that μ= 100.

Alternative Hypothesis (Ha): The students' average IQ score is above average, or μ> 100.

We are going for the "greater than" assumption, thus it is a one-tailed test. Assume for the

moment that the significance level in this situation, or alpha level, is 5%, or 0.05, which

equates to a Z score of 1.645. The statistical calculation (112.5 - 100) / (15/√30) = 4.56 yields

the Z score. The last step is to compare the values of the computed and expected z scores.

The null hypothesis, which states that the average IQ score of the students attending that

school is above average, is rejected in this case since the calculated Z score is lower than the

expected Z score.

Solved Examples on Hypothesis Testing

1. Conduct a hypothesis test to see if your decision and conclusion would change if your

belief were that the brown trout’s Mean I.Q. is not four.

Solution

a. H0:μ=4H0:μ=4

b. Ha:μ≠4Ha:μ≠4

259
c. Let X¯ X¯ the

average

I.Q. of a set of brown trout.

d. two-tailed Student's t-test

e. t=1.95t=1.95

f. p-value=0.076p-value=0.076

g. Check student’s solution.

h.

i. α:0.05α:0.05

ii. Decision: Reject the null

hypothesis

iii. Reason for decision: The p-valuep-value is greater than 0.05

iv. Conclusion: There is insufficient evidence to conclude that the

average

IQ of brown trout is not four.

i. (3.8865,5.9468)

2. According to a survey conducted for Newsweek, 13% of Americans claim to have

seen or felt the presence of an angel. Some others question if the percentage is actually

so high. It carries out its own research. Only two of the 76 Americans who were polled

had really seen or felt the presence of an angel. Would you concur with the Newsweek

poll as a consequence of the contingent's poll? Give three reasons why the findings of

the two polls could differ, using complete phrases.

260
Solution

a. H0:p≥0.13H0:p≥0.13

b. Ha:p<0.13Ha:p<0.13

c. Let P'=P′= the

proportion

of Americans who have seen or sensed angels

d. normal for a single

proportion

e. –2.688

f. p-value=0.0036p-value=0.0036

g. Check student’s solution.

h.

i. alpha: 0.05

ii. Decision: Reject the null

hypothesis

iii. Reason for decision: The p-value is less than 0.05.

iv. Conclusion: There is sufficient evidence to conclude that the percentage of

Americans who have seen or sensed an angel is less than 13%.

i. (0,0.0623)(0,0.0623).

The“plus-4s” confidence

interval

is (0.0022, 0.0978)

261
3. "Untitled," by Stephen Chen

I've often wondered how software is released and sold to the public. Ironically, I work

for a company that sells products with known problems. Unfortunately, most of the

problems are difficult to create, which makes them difficult to fix. I usually use the test

program X, which tests the product, to try to create a specific problem. When the test

program is run to make an error occur, the likelihood of generating an error is 1%.

So, armed with this knowledge, I wrote a new test program Y, that will generate the

same error that test program X creates, but more often. To find out if my test program

is better than the original so that I can convince the management that I'm right, I ran

my test program to find out how often I can generate the same error. When I ran my

test program 50 times, I generated the error twice. While this may not seem much

better, I think that I can convince the management to use my test program instead of

the original test program. Am I right?

Solution

a. H0:p=0.01H0:p=0.01

b. Ha:p>0.01Ha:p>0.01

c. Let P'=P′= the

proportion

of errors generated

d. Normal for a single

proportion

262
e. 2.13

f. 0.0165

g. Check student’s solution.

h.

i. α:0.05α:0.05

ii. Decision: Reject the null

hypothesis

iii. Reason for decision: The p-value is less than 0.05.

iv. Conclusion: At the 5% significance level, there is sufficient evidence to

conclude that the

the proportion

of errors generated is more than 0.01.

i. Confidence

interval

: (0,0.094)(0,0.094).

The“plus-4s” confidence

interval

is (0.004,0.144)(0.004,0.144).

3.3.2 Small Sample t-Tests

263
The t-test serves as the big sample equivalent of the large sample z test. Generally speaking, a

sample with a n<30 is considered tiny. Due to the non-normal distribution of tiny samples, a

t-test is required.

a) Single Sample T-test

One of the three different types of T-tests is the "One sample T Test." It is employed to

determine if the population's mean, from which the sample was selected, corresponds to

a given value.

The One Sample T Test is used to examine if a sample of observations might have

originated from a process that adheres to a particular parameter (like the mean).

Usually, it is used with little samples.

For instance, you might want to determine whether a sample mean of 15 items matches a

postulated mean (population). In essence, you want to determine whether or not the

sample represents the given population. Let’s suppose you want to test if the mean

weight of a manufactured component (from a sample size 15) is of a particular value (55

grams), with a 99% confidence.

The null hypothesis typically presupposes that there is no difference between the

hypothesised mean and the sample means (comparison mean). The T-Test is used to

determine whether or not the null hypothesis can be rejected.

The alternate hypothesis could be one of the following three scenarios, depending on

how the problem is phrased:

Case 1: H1 : x̅ != µ. used when the comparison Mean and the genuine sample mean are

not comparable. Employ the two-tailed T-test.

264
Case 2: H1 : x̅ > µ. used when the comparison Mean is higher than the genuine sample

mean. Utilize the upper-tailed T-test.

Case 3: H1 : x̅ < µ. used when the comparison Mean is greater than the genuine sample

mean. Utilize the lower-tailed T-test.

Where x̅ is the sample mean and µ is the population mean for comparison.

Example- A customer service company wants to know if their support agents are

performing on par with industry standards.

A report states that each ticket typically takes 20 minutes to be resolved. The sample

group's mean ticket purchase time is 21 minutes, with a 7-minute standard deviation. Can

you determine whether or not the company's assistance performance exceeds the industry

norm?

The procedure of doing Single Sample Test-

Step 1: Define the Null Hypothesis (H0) and Alternate Hypothesis (H1)

Example:

H0: Sample mean (x̅ ) = Hypothesized Population mean (µ)

H1: Sample mean (x̅ ) != Hypothesized Population mean (µ)

The alternate hypothesis can also state that the sample mean is greater than or less than

the comparison mean.

265
Step 2: Compute the test statistic (T)

𝑍 𝑥̅ − 𝜇
𝑡= =
𝑠 𝜎̂
√𝑛

where s is the standard error.

Step 3: Find the T-critical from the T-Table

Use the degree of freedom and the alpha level (0.05) to find the T-critical.

Step 4: Determine if the computed test statistic falls in the rejection region.

Alternately, simply compute the P-value. If it is less than the significance level (0.05 or

0.01), reject the null hypothesis.

Example-

Problem Statement:

We have the potato yield from 12 different farms. We know that the standard potato

yield for the given variety is µ=20.

x = [21.5, 24.5, 18.5, 17.2, 14.5, 23.2, 22.1, 20.5, 19.4, 18.1, 24.1, 18.5]

Test if the potato yield from these farms is significantly better than the standard yield.

Solution:

266
Step 1: Define the Null and Alternate Hypothesis

H0: x̅ = 20

H1: x̅ > 20

n = 12. Since this is one sample T-test, the degree of freedom = n-1 = 12-1 = 11.

Let’s set alpha = 0.05, to meet 95% confidence level.

Step2:Calculate the Test Statistic (T)

1. Calculate the sample mean

𝑥1 + 𝑥2 + 𝑥3 + ⋯ + 𝑥𝑛
𝑥̅ =
𝑛

𝑥̅ = 20.175

2. Calculate sample standard deviation

(𝑥1 − 𝑥̅ )2 + (𝑥2 − 𝑥̅ )2 + ⋯ + (𝑥3 − 𝑥̅ )2


𝜎̅ =
𝑛−1

σ=3.0211

3. Substitute in the T Statistic formula

𝑥̅ − 𝜇 𝑥̅ − 𝜇
𝑇= =
𝑠𝑒 𝜎̂
√𝑛

267
20.175 − 20
𝑇= = 0.2006
3.0211
√12

Step 3: Find the T-Critical

Confidence level = 0.95, alpha=0.05. For one tailed test, look under 0.05 column. For

d.o.f = 12 – 1 = 11, T-Critical = 1.796.

Because of how we specify the alternative hypothesis, a one-tailed test is used in this

case. We would have used a two-tailed test if the null hypothesis had been as simple as

"the sample means are not equal to 20."

Step 4: Does it fall in rejection region?

268
Since the computed T Statistic is less than the T-critical, it does not fall in the rejection

region.

b) Two Sample T-Test

To establish whether or not two population means are equal, a two sample t-test is employed.

Let's say we wish to determine whether the mean weight of two different turtle species is

equal. Weighing each individual turtle would take too much time and money because there

are thousands of turtles in each colony.

Instead, we might choose 15 turtles at random from each group and use the average weight of

each sample to assess whether the mean weights of the two populations are equal:

269
The mean weight between the two samples will, however, almost certainly differ by at least a

small amount. Whether or if this difference is statistically significant is the question.

Fortunately, we can respond to this query using a two sample t-test.

A two-sample t-test always uses the following null hypothesis:

 H0: μ1 = μ2 (the two population means are equal)

The alternative hypothesis can be either two-tailed, left-tailed, or right-tailed:

 H1 (two-tailed): μ1 ≠ μ2 (the two population means are not equal)

 H1 (left-tailed): μ1 < μ2 (population 1 mean is less than population 2 mean)

 H1 (right-tailed): μ1> μ2 (population 1 mean is greater than population 2 mean)

We use the following formula to calculate the test statistic t:

(x̅1 -x̅ 2 )
Test statistic: t=
1 1
SP (√n +n )
1 2

where x
̅ 1 and x̅ 2 are the sample means, n1 and n2 are the sample sizes, and where sp is

calculated as:

(𝑛1 − 1)𝑠12 + (𝑛2 − 1)𝑠22


𝑠𝑃 = √
𝑛1 + 𝑛2 − 2

where s12 and s22 are the sample variances.

You can reject the null hypothesis if the p-value for the test statistic t with (n1+n2-1) degrees

of freedom is less than your selected level of significance (popular options are 0.10, 0.05, and

0.01).

270
Example- Let's say we wish to determine whether the mean weight of two different turtle

species is equal. The following procedures will be used to conduct a two sample t-test with a

significance level α of 0.05 to test this:

Step 1: Gather the sample data.

Suppose we collect a random sample of turtles from each population with the following

information:

Sample 1:

 Sample size n1 = 40

 Sample mean weight x1 = 300

 Sample standard deviation s1 = 18.5

Sample 2:

 Sample size n2 = 38

 Sample mean weight x2 = 305

 Sample standard deviation s2 = 16.7

Step 2: Define the hypotheses.

We will perform the two sample t-test with the following hypotheses:

 H0: μ1 = μ2 (the two population means are equal)

 H1: μ1 ≠ μ2 (the two population means are not equal)

Step 3: Calculate the test statistic t.

First, we will calculate the pooled standard deviation sp:

271
(𝑛1 −1)𝑠12 +(𝑛2 −1)𝑠22
𝑠𝑃 = √ 𝑛1 +𝑛2 −2

(40−1)18.52 +(38−1)16.72
= 𝑠𝑃 = √ 40+38−2

= 17.647

Next, we will calculate the test statistic t:

(x̅1 -x̅2 ) 1 1
t= = (300-305) / 17.647(√ + ) = -1.2508
1 1
SP (√ + ) 40 38
n n1 2

Step 4: Calculate the p-value of the test statistic t.

According to the T Score to P Value Calculator, the p-value associated with t = -1.2508 and

degrees of freedom = n1+n2-2 = 40+38-2 = 76 is 0.21484.

Step 5: Draw a conclusion.

We are unable to reject the null hypothesis because this p-value is more than our level of

significance α, which is 0.05. The difference in mean turtle weight between these two

populations cannot be inferred from the available data.

The estimated T statistic clearly does not fall into the rejection region. Therefore, the null

hypothesis is not rejected.

3.3.3 Z- Test for Single & Two Samples

272
As long as the data has a normal distribution, a z test can be used to determine if the means of

two populations differ or not. It is necessary to set up the null hypothesis, the alternative

hypothesis, and calculate the value of the z test statistic for this reason. The z critical value

serves as the basis for the decision criterion.

A normal distribution population with independent data points and a sample size higher than

or equal to 30 is subjected to a z test. When the population variance is known, it is used to

determine if the means of two populations are equal to one another.

If the z test statistic is statistically significant when contrasted with the crucial value, the null

hypothesis can be rejected.

In order to determine whether there is a difference between the means of two populations, the

z test formula compares the z statistic with the z critical value. The z critical value separates

the acceptance and rejection sections of the distribution graph in hypothesis testing. The null

hypothesis can be rejected if the test statistic is within the rejection region; otherwise, it

cannot be rejected. Below is the z test formula for setting up the necessary hypothesis tests

for a one-sample and two-sample z test.

One Sample Z-Test


When the population standard deviation is known, a one-sample z test is performed to

determine whether there is a discrepancy between the sample mean and the population mean.

The following is the formula for the z test statistic:

𝑥̅ − 𝜇
𝑧= 𝜎
√𝑛

273
𝑥̅ is the sample mean, 𝜇 is the population mean, 𝜎 is the population standard deviation and n

is the sample size.

The algorithm to set a one sample z test based on the z test statistic is given as follows:

Left Tailed Test:

Null Hypothesis: H0: μ=μ0

Alternate Hypothesis: H1 : μ<μ0

Decision Criteria: If the z statistic < z critical value then reject the null hypothesis.

Right Tailed Test:

Null Hypothesis: H0H0 : μ=μ0μ=μ0

Alternate Hypothesis: H1H1 : μ>μ0μ>μ0

Decision Criteria: If the z statistic > z critical value then reject the null hypothesis.

Two Tailed Test:

Null Hypothesis: H0H0 : μ=μ0μ=μ0

Alternate Hypothesis: H1H1 : μ≠μ0μ≠μ0

Decision Criteria: If the z statistic > z critical value then reject the null hypothesis.

Two Sample Z Test

To determine whether there is a difference between the means of two samples, a two sample

z test is utilised. The following is the formula for the z test statistic:

(𝑥
̅̅̅1 − 𝑥̅2 ) − (𝜇1 − 𝜇2 )
𝑧=
𝜎12 𝜎22

𝑛1 + 𝑛2

274
𝑥1 𝜇1 , 𝑎𝑛𝑑 𝜎12 are the sample mean, population mean and population variance respectively
̅̅̅,

𝑥2 𝜇2 , 𝑎𝑛𝑑 𝜎22 are the sample mean, population mean and population
for the first sample. ̅̅̅,

variance respectively for the second sample.

Similar to the one-sample test, the two-sample z test can be set up. The means of the two

samples will be compared using this test, nevertheless. For example, the null hypothesis is

given as H0 : μ1=μ2.

3.3.4 Chi-square (goodness of fit & association of attributes)

275
For categorical data, a statistical test called Pearson's chi-square test is used. It is employed to

assess whether your data significantly depart from your expectations. The Pearson's chi-

square test comes in two varieties:

 To determine whether the frequency distribution of a categorical variable deviates

from your expectations, apply the chi-square goodness of fit test.

 To determine if two categorical variables are connected to one another, utilise the chi-

square test of independence.

Chi-square is often written as Χ2 and is pronounced “kai-square” (rhymes with “eye-square”).

It is also called chi-squared.

Among the most popular nonparametric tests are Pearson's chi-square (Χ2) tests, sometimes

known as chi-square tests. For data that do not adhere to the assumptions of parametric tests,

particularly the assumption of a normal distribution, nonparametric tests are used.

Use a chi-square test or equivalent nonparametric test if you want to test a hypothesis

regarding the distribution of a categorical variable. Categorical variables, which indicate

groupings like animals or countries, can be nominal or ordinal. They cannot have a normal

distribution since they can only have a few limited values.

There are two different kinds of Pearson's chi-square tests, but they all determine whether a

categorical variable's observed frequency distribution differs significantly from its predicted

frequency distribution. The distribution of observations among several groups is described by

a frequency distribution.

276
Frequency distribution tables are frequently used to depict frequency distributions. The

number of observations in each category is displayed in a frequency distribution table. A

particular kind of frequency distribution table called a contingency table can be used to

display the number of observations in each combination of groups when there are two

categorical variables.

Frequency of visits by bird species at a bird feeder during a 24-hour period

Bird species Frequency

House sparrow 15

House finch 12

Black-capped chickadee 9

Common grackle 8

European starling 8

Mourning dove 6

Both of Pearson’s chi-square tests use the same formula to calculate the test statistic, chi-

square (Χ2):

277
Where:

 Χ2 is the chi-square test statistic

 Σ is the summation operator (it means “take the sum of”)

 O is the observed frequency

 E is the expected frequency

The larger the chi-square, the greater the discrepancy between the observations and the

expectations (O E in the equation). You evaluate the chi-square value against a crucial value

to see whether the difference is large enough to be statistically significant.

Types of chi-square test

The two types of Pearson’s chi-square tests are:

 Chi-square of goodness of fit test

 Chi-square test of independence

These tests are actually the same in terms of mathematics. However, because they are

employed for various objectives, we frequently consider them to be distinct tests.

the goodness of fit test using chi-square

The Chi-square goodness of fit test can be used when there is only one categorical variable. It

enables you to determine whether the categorical variable's frequency distribution

significantly deviates from your expectations. The idea is that the categories will have equal

proportions, albeit this is not always the case.

278
Example: Hypotheses for the chi-square goodness of fit test

Expectation of equal proportions

 Null hypothesis (H0): The bird species visit the bird feeder in equal proportions.

 Alternative hypothesis (HA): The bird species visit the bird feeder

in different proportions.

Expectation of different proportions

 Null hypothesis (H0): The bird species visit the bird feeder in the same proportions

as the average over the past five years.

 Alternative hypothesis (HA): The bird species visit the bird feeder

in different proportions from the average over the past five years.

Chi-square test of independence

You may do an independence test using the chi-square formula when you have two

categorical variables. You can use it to see if there is a relationship between the two

variables. When two variables are independent of one another, they do not affect each other's

likelihood of belonging to a particular group.

Example: Chi-square test of independence

 Null hypothesis (H0): The proportion of people who are left-handed is the same for

Americans and Canadians.

 Alternative hypothesis (HA): The proportion of people who are left-

handed differs between nationalities.

3.3.5 One Way & Two Way Anova

279
If there is a statistically significant difference between the means of three or more

independent groups, it can be ascertained using an ANOVA or analysis of variance.

The one-way and two-way ANOVAs are the two forms of ANOVAs that are used the most

frequently.

One-way ANOVA: Used to examine the influence of a single factor on a response variable.

Two-way ANOVA: used to ascertain the effects of two factors on a response variable and to

ascertain whether or not the two factors interact with the answer variable.

The following examples provide an example of how to perform each type of ANOVA.

Example: One-Way ANOVA

Consider a professor who wants to discover if using three distinct study methods will result in

different exam results. He enlists 30students to take part in the study and randomly assigns

280
each participant to use one of the three methods to study for an exam in order to evaluate this.

All of the students take the same exam at the end of the month.

The test scores for each student are shown below:

The professor performs a one-way ANOVA and gets the following results:

281
The p-value for the F test is 0.1138, and the F test statistic is 2.3575. We lack adequate

information to conclude that the three studying methods result in different mean exam scores

because this p-value is not smaller than.05.

Example: Two-Way ANOVA

Consider a botanist who is curious about the effects of sunlight exposure and watering

frequency on plant growth. She sows 40 seeds and gives them two months to grow while

providing them with various amounts of sunlight exposure and hydration schedules. She

notes the height of each plant after two months. The outcomes are displayed below:

The professor performs a two-way ANOVA and gets the following results:

Here’s how to interpret the results:

282
 The interaction between watering frequency and sun exposure had a p-value of 0.310898.

At an alpha level of 0.05, this is not statistically significant.

 For watering frequency, the p-value was 0.975975. At an alpha level of 0.05, this is not

statistically significant.

 The p-value for exposure to sunlight was 0.000003. This has a 0.05 alpha level

significance statistically.

These findings suggest that the only variable with a statistically significant impact on plant

height is sunshine exposure. And because there is no interaction effect, the effect of sunlight

exposure is consistent across each level of watering frequency. That is, whether a plant is

watered daily or weekly has no impact on how sunlight exposure affects a plant.

283
Case Study (Hypothesis Testing)

For their lift policy salesforce, The Titan Insurance Company has just implemented a new

incentive payment programme. It wants to get a head start on determining whether the new

plan will work or not. There are signs that the sales force is selling more insurance, but since

sales constantly fluctuate from month to month in an unpredictable fashion, it is unclear

whether the strategy has had a substantial impact.

The total sum guaranteed for the policies sold by a salesperson during the month is often how

life insurance firms gauge their monthly performance. Consider the scenario where salesman

X sold seven policies with the following sums assured: £1000, £2500, £3000, £5000, £10000,

and £35000. The total of these amounts assured, or £61,500, represents X's output for the

month.

With Titan's new strategy, salespeople receive little regular pay but are compensated

significantly with bonuses based on their performance (i.e. to the total sum assured of policies

sold by them). The plan is costly for the business, but sales growth is expected to more than

284
makeup for it. According to the agreement with the sales force, the plan would be scrapped

after six months if it does not at least break even for the business.

The programme has been running for four months at this point. After varying throughout the

first two months of the switchover, it has stabilised.

Titan has selected 30 random salespeople to test the efficacy of the plan by measuring their

output in the penultimate month prior to transition and again in the fourth month following

the changeover (they have deliberately chosen months not too close to the changeover). The

outputs of the salesmen are displayed in Table-

SALESPERSON Old_Scheme New_Scheme

1 57 62

2 103 122

3 59 54

4 75 82

5 84 84

6 73 86

7 35 32

8 110 104

9 44 38

10 82 107

11 67 84

12 64 85

285
13 78 99

14 53 39

15 41 34

16 39 58

17 80 73

18 87 53

19 73 66

20 65 78

21 28 41

22 62 71

23 49 38

24 84 95

25 63 81

26 77 58

27 67 75

28 101 94

29 91 100

30 50 68

Data Preparation-

286
It will be best to transform the given data to thousands as they are currently expressed

in 000.

Problem 1

Describe the five per cent significance test you would apply to these data to determine

whether the new scheme has significantly raised outputs. What conclusion does the test

lead to?

Solution:

It is asked whether the new scheme has significantly raised the output. It is an example of

the one-tailed t-test.

Note: Two-tailed test could have been used if it was asked “new scheme has

significantly changed the output”

Mean of amount assured before the introduction of scheme = 68450

Mean of amount assured after the introduction of scheme = 72000

Difference in mean = 72000 – 68450 = 3550

Let,

μ1 = Average sums assured by salesperson BEFORE changeover. μ2 = Average sums

assured by salesperson AFTER changeover.

H0: μ1 = μ2 ; μ2 – μ1 = 0

HA: μ1 < μ2 ; μ2 – μ1 > 0 ; true difference of means is greater than zero.

Since the population standard deviation is unknown, paired sample t-test will be used.

Since the p-value (=0.06529) is higher than 0.05, we accept (fail to reject) the NULL

hypothesis. The new scheme has NOT significantly raised outputs.

287
Problem 2

Let's say it has been determined that Titan has to increase its average output by £5000 in

order to break even. What is the alternative hypothesis, if this figure is it:

(a) The probability of a type 1 error?

(b) What is the p-value of the hypothesis test if we test for a difference of $5000?

(c) Power of the test:

Solution:

2.a. The probability of a type 1 error?

Solution: Probability of Type I error = significant level = 0.05 or 5%

2.b. What is the p-value of the hypothesis test if we test for a difference of $5000?

Solution:

Let μ2 = Average sums assured by salesperson AFTER changeover.

μ1 = Average sums assured by salesperson BEFORE changeover.

μd = μ2 – μ1 H0: μd ≤ 5000 HA: μd > 5000

This is a right tail test.

P-value = 0.6499

2.c. Power of the test.

Solution:

Let μ2 = Average sums assured by salesperson AFTER changeover. μ1 = Average sums

assured by salesperson BEFORE changeover. μd = μ2 – μ1 H0: μd = 4000

HA: μd > 0

H0 will be rejected if test statistics > t_critical.

With α = 0.05 and df = 29, critical value for t statistic (or t_critical ) will be 1.699127.

288
Hence, H0 will be rejected for test statistics ≥ 1.699127.

Hence, H0 will be rejected if for ̅ ≥ 4368.176

Graphically,

289
Probability (type II error) is P(Do not reject H0 | H0 is false)

Our NULL hypothesis is TRUE at μd = 0 so that H0: μd = 0 ; HA: μd > 0

Probability of type II error at μd = 5000

= P (Do not reject H0 | H0 is false)

= P (Do not reject H0 | μd = 5000) = P (𝑥̅ < 4368.176 | μd = 5000)

= P (t < | μd = 5000)

290
= P (t < -0.245766)

= 0.4037973

Now, β=0.5934752,

Power of test = 1- β = 1- 0.5934752

= 0.4065248

Points to Keep In mind-

 Since we only have evidence from the sample(s), hypotheses cannot be proved or

refuted during hypothesis testing. A hypothesis can only be accepted or rejected at

most.

 Even if the test statistic is within the Acceptance Region or has a p-value of , using

the phrase "accept H0" in place of "do not reject" should be avoided. This merely

means that there is insufficient statistical support for rejecting the H0 from the

sample. H0 will not be accepted since we have attempted to negate (reject) it but have

not discovered enough evidence to do so.

 The interval estimation technique known as the confidence interval can also be used

to test hypotheses. We do not reject H0 if the confidence interval encompasses the

hypothesis parameter. Otherwise, we reject H0 if the hypothesised parameter is

outside the confidence interval, which means that the confidence interval does not

contain the parameter.

Important formula-

1. test statistic (T)

291
𝑍 𝑥̅ − 𝜇
𝑡= =
𝑠 𝜎̂
√𝑛

2. sample mean

𝑥1 + 𝑥2 + 𝑥3 + ⋯ + 𝑥𝑛
𝑥̅ =
𝑛

3. sample standard deviation

(𝑥1 − 𝑥̅ )2 + (𝑥2 − 𝑥̅ )2 + ⋯ + (𝑥3 − 𝑥̅ )2


𝜎̅ =
𝑛−1

4. T Statistic formula

𝑥̅ − 𝜇 𝑥̅ − 𝜇
𝑇= =
𝑠𝑒 𝜎̂
√𝑛

(x̅1 -x̅ 2 )
5. Test statistic: t=
1 1
SP (√n +n )
1 2

6. pooled standard deviation sp:

(𝑛1 −1)𝑠12 +(𝑛2 −1)𝑠22


𝑠𝑃 = √ 𝑛1 +𝑛2 −2

7. z test statistic:

𝑥̅ − 𝜇
𝑧= 𝜎
√𝑛
8. Two sample z test statistic:

292
(𝑥
̅̅̅1 − 𝑥̅2 ) − (𝜇1 − 𝜇2 )
𝑧=
𝜎12 𝜎22

𝑛1 + 𝑛2

9. chi-square (Χ2):

Key Terms- Sample Tests, Z Tests, Hypothesis Testing, Anova, Sample standard deviation.

Summary- In the given unit, students can learn the applications of hypothesis testing, and

sample t tests.

Glossary

 Alpha Level- Alpha risk is another name for it. It's the likelihood that rejecting your

null hypothesis mistakenly or making a Type A error is acceptable. Alpha level is

always a number between 0 and 1, with 0.05 being the most popular choice. When

your test is finished, and the data has been processed statistically, you will get a p-

value to compare to your alpha level.

 Alternate Hypothesis- A hypothesis that disagrees with the null hypothesis; the two

are mutually exclusive.

 Beta level- also known as the beta risk. It is the risk that you are willing to take to

make a Type B error, which is not to reject your null hypothesis when it is actually

false.

293
 Conclusion- A declaration that details the strength of the evidence (sufficient or

insufficient), the significance level, and whether the original claim is confirmed (null)

or rejected (sufficient) (alternative).

 Confidence level- called the confidence interval as well. This indicates how certain

you can be that your conclusion is accurate. Because the alpha and confidence levels

always sum up to one, calculating the confidence level is simple. i.e.:

1 – α = confidence level

 Critical region- All values that might lead us to reject the null hypothesis H0 are

collected in this set—also referred to as a rejection zone.

 Critical value(s)- The value(s) that define the boundary between the critical and non-

critical regions. The sample statistics are not used to determine the critical values.

 The rejection region and the non-rejection region are divided by a critical value.

 Critical value(s)- The value(s) that define the boundary between the critical and non-

critical regions. The sample statistics are not used to determine the critical values.

 The rejection region and the non-rejection region are divided by the letter A.

 Decision- a claim supported by the null hypothesis. Either the "null hypothesis" is

rejected, or the "null hypothesis is failed to be rejected." The null hypothesis won't

ever be accepted by us.

 A p-value is a probability of getting a test statistic that is at least as extreme as the one

found from the sample data.

 Error- In hypothesis testing, there are two main sorts of errors: type A errors, where a

valid hypothesis is disproved, and type B errors, where a false hypothesis is accepted.

Learn more about mistakes.

 Ho- Also referred to as the null hypothesis.

 H1- The alternative hypothesis, or H1, is also referred to as H(a).

294
 Left-tailed test- The hypothesis test is a left-tailed test if the alternative hypothesis H1

contains the less-than inequality sign ().

 Null hypothesis- The claim you're attempting to refute. This is the standard

presumption, according to which there was no outside influence on the experimental

outcomes and they were solely the consequence of chance.

 P-value- A essential component of any hypothesis test results is the p-value. It's a

number between 0 and 1 and measures the likelihood that random fluctuations caused

any data that could lead you to reject the null hypothesis. It is determined by applying

a statistical significance test to test results. You reject the null hypothesis if the p-

value is less than your alpha threshold, and you do not reject the null hypothesis if the

number is larger. Study up on p-values.

 When determining terms for a hypothesis test, the standard error is computed in a

manner known as "pooling." The two proportions are averaged in the pooled form,

and the standard error is calculated using just one proportion. Pooled computations

are preferred by ASQ, Villanova, and most other organisations. The two proportions

are used independently in the unpooled variant. IASSC recommends unpooled in

general.

References-

1. Higher Engineering Mathematics By [Link], Khanna Publishers

2. Probability and Statistics for Engineers and Scientists by Sheldon [Link],Academic

Press

3. Bhattacharyya, G. K., and R. A. Johnson, (1997). Statistical Concepts and Methods,

John Wiley and Sons, New York.

295
Suggested Readings-

1. [Link]

testing/anova/

2. [Link]

significance-types-and-measures-statistics/15249

Study Tips-

1. Instead of mass practise, use distributive practise. To study statistics, set aside one to two

hours per day at the same time for six days of the week (leave the seventh day off). Avoid

cramming four or five hours of study into one or two sessions each week. This is a

fundamental idea.

2. At least once a week, study in student groups of three or four. A deeper level of learning is

actually cemented through verbal exchange and interpretation of concepts and skills with

other pupils.

3. Avoid attempting to retain formulas (A good instructor will never ask you to do this).

Concepts, concepts, concepts: study. Remember that you can always look up the formula in a

textbook later in life when you need to utilise a statistical technique.

4. Complete as many and different of the exercises and problems as you can. Ideally, a

workbook is included with your textbook. Simply reading about statistics won't teach you

anything about it. Pushing the pencil and continually honing your techniques are required.

5. In statistics, look for recurring themes. There are probably only a few number of crucial

abilities that come up repeatedly. If necessary, request that your instructor stress these.

296
6. Become a Gestalt therapist! Recognize that statistics as a whole is larger than the sum of its

parts. It is very simple to become preoccupied with minor details and lose sight of the bigger

picture.

7. Take action if you suffer from math or statistics anxiety, which affects roughly 70% of

people in general. Most institutions offer top-notch counselling programmes to lessen this

impairment because they recognise how crippling this issue is. Get assistance for yourself.

This could end up being the wisest choice you make during your college career.

297
Exercise: State the Type I and Type II errors in complete sentences given the following

statements.

a. The Mean number of years Americans work before retiring is 34.

b. At most 60% of Americans vote in presidential elections.

c. The mean starting salary for San Jose State University graduates is at least $100,000

per year.

d. Twenty-nine percent of high school seniors get drunk each month.

e. Fewer than 5% of adults ride the bus to work in Los Angeles.

f. The mean number of cars a person owns in his or her lifetime is not more than ten.

g. About half of Americans prefer to live away from cities, given the choice.

h. Europeans have a mean paid vacation each year of six weeks.

i. The chance of developing breast cancer is under 11% for women.

j. Private universities mean tuition cost is more than $20,000 per year.

Short Answer Type Questions-

1. What type of graphs are used to depict the bivariate analysis?

2. What do you mean by bivariate frequency distribution?

3. What is data?

Long Answer Type Questions-

1. Explain the real life applications of Hypothesis testing.

2. A particular brand of tires claims that its deluxe tire averages at least 50,000 miles before it

needs to be replaced. From past studies of this tire, the standard deviation is known to be

298
8,000. A survey of owners of that tire design is conducted. From the 28 tires surveyed,

the mean lifespan was 46,500 miles with a standard deviation of 9,800 miles.

Using α=0.05α=0.05 is the data highly inconsistent with the claim??

3. A professor wants to know if her introductory statistics class has a good grasp of basic

math. Six students are chosen at random from the class and given a math proficiency test. The

professor wants the class to be able to score above 70 on the test. The six students get the

following scores: 62, 92, 75, 68, 83, 95.

Can the professor have 90% confidence that the mean score for the class on the test would be

above 70?

Answers

Exercise: Type I and Type II error:

a. Type I error: We conclude that the Mean is not 34 years, when it really is 34 years.

Type II error: We conclude that the mean is 34 years, when in fact it really is not 34

years.

b. Type I error: We conclude that more than 60% of Americans vote in presidential

elections, when the actual percentage is at most 60%.Type II error: We conclude that

at most 60% of Americans vote in presidential elections when, in fact, more than 60%

do.

299
c. Type I error: We conclude that the Mean starting salary is less than $100,000 when it

really is at least $100,000. Type II error: We conclude that the Mean starting salary is

at least $100,000 when, in fact, it is less than $100,000.

d. Type I error: We conclude that the proportion of high school seniors who get drunk

each month is not 29% when it really is 29%. Type II error: We conclude that

the proportion of high school seniors who get drunk each month is 29% when, in fact,

it is not 29%.

e. Type I error: We conclude that fewer than 5% of adults ride the bus to work in Los

Angeles when the percentage that does is really 5% or more. Type II error: We

conclude that 5% or more adults ride the bus to work in Los Angeles when, in fact,

fewer than 5% do.

f. Type I error: We conclude that the Mean number of cars a person owns in his or her

lifetime is more than 10, when in reality it is not more than 10. Type II error: We

conclude that the mean number of cars a person owns in his or her lifetime is not

more than 10 when, in fact,it is more than 10.

g. Type I error: We conclude that the Proportion of Americans who prefer to live away

from cities is not about half, though the actual proportion is about half. Type II error:

We conclude that the proportion of Americans who prefer to live away from cities is

half when, in fact, it is not half.

h. Type I error: We conclude that the duration of paid vacations each year for Europeans

is not six weeks, when in fact it is six weeks. Type II error: We conclude that the

duration of paid vacations each year for Europeans is six weeks when, in fact, it is

not.

300
i. Type I error: We conclude that the Proportion is less than 11%, when it is really at

least 11%. Type II error: We conclude that the proportion of women who develop

breast cancer is at least 11%, when in fact it is less than 11%.

j. Type I error: We conclude that the Average tuition cost at private universities is more

than $20,000, though in reality it is at most $20,000. Type II error: We conclude that

the average tuition cost at private universities is at most $20,000 when, in fact, it is

more than $20,000.

301
Suggested Reading

1. T R Jain & S C Aggarwal, Statistics for MBA, VK Publications, ISBN 818961133X,

9788189611330

2. Levine, D., Sazbat, K. and Stephan, D. 2013. Business Statistics, 7thEdition, Pearson

Education, India, ISBN: 9780132807265.

3. Gupta, C. and Gupta, V. 2004. An Introduction to Statistical Methods, 23rdEdition, Vikas

Publications, India, ISBN: 9788125916543.

4. Croucher, J. 2011. Statistics: Making Business Decisions, 13thEdition, Tata McGraw Hill,

ISBN: 9780074710419.

5. Gupta, S. 2011. Statistical Methods, 4thEdition, Sultan Chand & Sons, ISBN: 8180548627

302

Potrebbero piacerti anche