Business Statistics
Module 2
MEASURES OF CENTRAL TENDENCY
INTRODUCTION
The values of the extent of the observations are not equal, but we notice a general tendency of
such observations to cluster around a particular level. In this situation it may be preferable to
characterise each group of observations by such a level, which is called the central tendency of
that group. This single value for each group of observations serves as a representative of that
group. This level around which the observations tend to cluster may vary from group to group.
One of the most important objectives of this biostatistical analysis is to get a single value that
describes the characteristic of the entire mass of data. Such a value is called an “average”.
The objective here is to find one representative value which can-be used to locate and summarise
the entire set of varying values. This one value can be used to make many decisions concerning
the entire set. We can define measures of central tendency (or location) to find some central value
around which the data tend to cluster.
SIGNIFICANCE OF MEASURES OF CENTRAI TENDENCY
Measures of central tendency i.e. condensing the mass of data in one single value , enable us to get
an idea of the entire data. For example, it is impossible to remember the individual incomes of
millions of earning people of India. But if the average income is obtained, we get one single value
that represents the entire population.
For quantitative data it is observed that there is a tendency of the data to be distributed about a
central value which is a typical value and is called a measure of central tendency. It is also called a
measure of location because it gives the position of the distribution on the axis of the variable.
There are three commonly used measures of central tendency, viz., Mean, Median
and Mode. The mean again may be of three types, viz. Arithmetic Mean (A.M.), Geometric Mean
(G.M.) and Harmonic Mean (H.M.).Measures of central tendency also enable us to compare two
or more sets of data to facilitate comparison.
For example, the average sales figures of April may be compared with the sales figures of
previous months.
PROPERTIES OF A GOOD MEASURE OF CENTRAL TENDENCY
[1] It should be easy to understand and calculate.
[2] It should be rigidly defined.
[3] It should be based on all observations.
[4] It should be least affected by sampling fluctuation.
[5] It should be capable of further algebraic treatment.
[6] It should be least affected by extreme values.
[7] It should be calculated in case of open end interval.
Following are some of the important measures of central tendency which are commonly used in
business
and industry.
• Arithmetic Mean
• Weighted Arithmetic Mean
• Median
• Quantiles(quartiles, deciles and percentiles)
• Mode
• Geometric Mean
• Harmonic Mean
ARITHMETIC MEAN
1) The arithmetic mean (or mean or average) is the most commonly used and readily
understood measure of central tendency. In statistics, the term average refers to any of the
measures of central [Link] new roman
Ungrouped data/Raw data
2) The arithmetic mean is defined as being equal to the sum of the numerical values of each
and every observation divided by the total number of observations.
Discrete data
3) When the observations are classified into a frequency distribution.
Merits and Demerits of Arithmetic Mean
The following are the merits and demerits of arithmetic mean:
(a) Merits of Arithmetic Mean
(i) It is commonly understood and most widely used.
(ii) It is simple and easy to calculate.
(iii) It is based on all the observations.
(iv) It is a good measure for comparison.
(v) It is adaptable to arithmetic and algebraic treatment.
(b) Demerits of Arithmetic Mean
(i) The value of mean is highly affected by abnormal and extreme values.
(ii) It may not be actually present in the series. For example, the average of 2,
3 and 10 is 5, which is not an observation of the series.
(iii) It can be calculated if certain item is missing. Further, in case of open-end interval, it
is calculated on certain assumption.
(iv) It cannot be located by mere observation.
Calculation of Arithmetic Mean
Mainly three forms of data are available, which are given below:
(i) Individual series or ungrouped data
(ii) Discrete series
(iii) Continuous series
MEDIAN
The ‘median’ is another important and widely used measure of central tendency. Median of a
distribution is the value of the variable which divides it into two equal parts, i.e., median is the
value such that the number of observations above it is equal to the number of observations below
it. The median is thus a positional average.
In case of ungrouped data, if the number of observations is odd, then median is the middle value
after the values have been arranged in ascending or descending order of magnitude. In case of
even number of observation, there are two middle terms and median is obtained by taking the
arithmetic mean of the two middle terms.
MODE
The ‘Mode’ is another measure of central tendency which is conceptually very useful. It is derived
from the French word “La mode” which means fashion.
“Mode is the value which occurs most frequently in a set of observations and around which the
other items of the set cluster densely.” In other words, mode is the value of the variable which is
predominant in the series. For example, the mode of a series 3, 5, 8, 5, 4, 5, 9, 3 would be 5, since
the value 5 occurs most frequently than any of the others. “The value of the variable at which the
curve reaches a maximum is called the mode”.There are many situations in which arithmetic mean
and median fail to reveal the true characteristic of data. For example, when we talk of most
common wage, most common income, most common height, most common size of sole or ready-
made garments.
Merits and Demerits of Mode
Merits
(i) It is easy to calculate, sometimes it is found only by inspection.
(ii) It is not affected by extreme values.
(iii) It can be calculated from open end classes.
(iv) It is simple and precise.
(v) Mode is that point where there is more concentration of frequencies. Hence,
it is the best representative of the data.
Demerits
(i) It is not based on all the items of the distribution.
(ii) It cannot be treated algebraically.
(iii) Equal intervals are needed for the calculation of mode, which is a draw-
back.
(iv) In certain situations, it is not clearly defined. Also in case of bi-modal or
multimodal distribution, it is not defined.
Quartiles
The quartiles divides the series in four equal parts. There are three quartiles name
Q1, Q2 and Q3, Q1 is called the first (lower) quartile and Q3 is called the third (upper) quartile Q2
is called the second quartile which divides the series into two equal parts. Second quartile Q2
coincides with median i.e., the value of Q2 and median is same. Twenty five per cent value are
less than Q1 and twenty five percent values are greater than Q3 and the rest fifty per cent values
lie between Q1 and Q3. Quartiles are widely used in economics and business.
MEASURES OF DISPERSION
Dispersion is important not only as merely supplementary to the average, but because of the
scatter distribution. According the Spurr and Bonini, “In matters of health, variations in body
temperature, pulse beat and blood pressure are basic guides to diagnosis. Prescribed treatment is
designed to control their variation.”
A measure of dispersion describes the degree of scatter shown by the observations
and is usually measured as an average deviation about some central value. Measures
of dispersion gives us additional information that enables us to judge the reliability
of our measure of central value. It makes possible to compare two series of data in
respect of their variability.
Following are the most commonly used measures of dispersion:
1. Range
2. Interquartile range and quartile deviation
3. Mean deviation
4. Standard deviation
5. Coefficient of variation.
Properties of a Good Measure of Dispersion
Characteristics of an ideal measure of dispersion are the same that of average, viz. in
brief :
(i) It should be simple to understand and easy to compute.
(ii) It should be rigidly defined.
(iii) It should be based on all the observations.
(iv) It should be amenable to further algebraic treatment.
(v) It must have sampling stability.
(vi) It should not be affected by extreme observations.
Absolute and Relative Measure of Dispersion
Measures of dispersion may be either absolute or relative. Absolute measure of dispersion cannot
be used for comparison purposes if expressed in different units.
The absolute measure of dispersion can be compared with another, only if the two
belong to the same population. For instance, when we measure the height and weight of the
students they may be in metres and kilograms respectively–two different units.
These different units cannot be measured through absolute method, to know the variability.
Therefore, two series cannot be compared if the absolute measure of dispersion of each series is
expressed as a ratio or percentage of the average. For comparing the variability, even if the
distributions are in the same units, the relative measure of dispersion is computed. In other words,
the coefficient of dispersion of each group should be calculated in order to compare two series.
RANGE
The difference between the largest and smallest value of the variates is called its
range.
Range = L – S,
where L = Largest value
S = Smallest value
Coefficient of range = L-S
L+S
Advantages of Range
1, It is simple to understand and easy to calculate.
2. It is used to study variations in the prices of commodities and movement in
the prices of securities.
3. It is used in weather forcasting e.g., It gives an idea of the variation between
maximum and minimum levels of temperature.
4. It is used in quality control for drawing R-charts.
Disadvantages of Range
1. It depends only on two values (largest and smallest) and ignores all other
values, it is highly misleading.
2. It is not useful for frequency distribution.
3. It is affected by sampling fluctuation.
4. It cannot be computed if the distribution is open-ended.
5. It is not suitable for further mathematical treatment.
Uses of Range
(i) Range is used in industries for the statistical quality control of manufactured
product by the construction of control chart.
(ii) The meteorological department uses the range for weather forecasts since
public is interested to know the limits within which the temperature is
likely to vary on a particular day.
(iii) Range is useful in studying the variations in the prices of stock, shares, and
other commodities that are sensitive to price changes from one period to
another period.
QUARTILE DEVIATION OR SEMI- INTERQUARTILE RANGE
Quartile deviation or semi-interquartile range (Q.D.) is given by
Q.D. = Q3 - Q1
2
where Q3 is the third quartile and
Q1 is the first quartile.
Quartile deviation is defined as half the distance between the third and the first
quartile.
Quartile deviation is an absolute measure of dispersion. The relative measure of dispersion, known
as coefficient of quartile deviation, is calculated as follows :
Coefficient of Q.D. = Q3 - Q1
Q3 + Q1
Interquartile Range
The difference between the third quartile and first quartile is known as interquartile range. It is
given by
Interquartile range = Q3 – Q1,
where Q3 is the third quartile and Q1 is the first quartile.
Merits and Demerits of Quartile Deviation
Merits
(i) It is simple to understand and easy to compute.
(ii) It is not influenced by the extreme values.
(iii) It can be found out with open end distribution.
(iv) It is not affected by the presence of extreme values.
Demerits
(i) It ignores the first 25% of the items and the last 25% of the items.
(ii) It is a positional average; hence not amenable to further mathematical treatment.
(iii) Its value is affected by sampling fluctuations.
(iv) It gives only a rough measure.
MEAN DEVIATION
Mean deviation (M.D.) is defined as the average of the absolute deviations taken from an average
usually, the mean, median or mode. Mean deviation is a measure of dispersion which is based on
all values of a set of data.
Merits and Demerits of Mean Deviation
Merits
(i) It is simple to understand and easy to compute.
(ii) It is based on all the observations.
(iii) It is not much affected by the fluctuations of sampling.
(iv) It is less affected by the extreme values.
(v) It is rigidly defined.
(vi) It is better measure for comparison.
(vii) It is flexible, because it can be calculated from any measure of central
tendency.
Demerits
(i) It is a non-algebraic treatment.
(ii) Algebraic positive and negative signs are ignored. In mean deviation + 5
and – 5 have the same meaning. It is mathematically unsound and illogical.
(iii) It is not a very accurate measure of dispersion.
(iv) It is not suitable for further mathematical calculation.
(v) It is rarely used. It is not as popular as standard deviation.
Uses of Mean Deviation
(i) It will help to understand the standard deviation.
(ii) It is useful in marketing problems.
(iii) It is useful while using small samples.
(iv) It is used in statistical analysis of economic business and social phenomena.
(v) It is useful in calculating the distribution of wealth in a community or a
nation.
(vi) It is useful in forecasting business cycles.
STANDARD DEVIATION
Standard deviation is the most commonly used absolute measure of dispersion. The concept of
standard deviation was first introduced by Karl Pearson in 1893. The ‘Standard’ is assigned to this
measure of variation is probably because it is the most commonly used and is the most flexible in
terms of variety of applications of all the measures of dispersions. It is clear that standard
deviation is a measure of the spread in a set of observations. In this method, the drawback of
ignoring the algebraic sign as in mean deviation is overcome by taking the square of deviations,
thereby, making all the deviations positive.
“Standard deviation is positive square root of the arithmetic mean of the squares of the deviations
taken from arithmetic mean”.
Merits and Demerits of Standard Deviation
Merits
(i) It is rigidly defined and its value is always definite.
(ii) It is based on all the observations and the actual signs of deviations are
used.
(iii) It is less affected by sampling fluctuations.
(iv) It is possible for further algebraic treatment.
Demerits
(i) It is not easy to understand and to calculate.
(ii) It gives more weight to extreme values, because the values are squared up.
(iii) It is affected by the value of every item in the series.
(iv) As it is an absolute measure of variability, it cannot be used for the purpose
of comparison.
(v) It has not found favour with the economists and businessmen.
Uses of Standard Deviation
(i) Standard deviation is the best measure of dispersion.
(ii) It is widely used in statistics because it possesses most of the characteris-
tics of an ideal measure of dispersion.
(iii) It is widely used in sampling theory and by biologists.
(iv) It is used in coefficient of correlation and in the study of symmetrical fre-
quency distribution.
COEFFICIENT OF VARIATION
The standard deviation is an absolute measure of dispersion. It is expressed in terms of units in
which the original data are collected. The standard deviation of the height of plants cannot be
compared with the standard deviation of the weight of plants because they are expressed in
different units, i.e., height (cm) and weight (gm). Therefore, the standard deviation must be
converted into a relative measure of dispersion for the purpose of comparison. The relative
measure of dispersion is known as the coefficient of variation.
Coefficient of standard deviation will be in fraction and as such not very good for comparison.
Therefore, the coefficient of standard deviation is multiplied by 100 gives the coefficient of
variations.
COMPARISON BETWEEN MEAN DEVIATION AND STANDARD DEVIATION
Mean deviation Standard deviation
Deviations are calculated from mean, medianor mode. Deviations are calculated only from
mean.
Algebraic signs are ignored while calculating mean deviations Algebraic signs are taken into account.
t is simple to calculate It is difficult to calculate.
It lacks mathematical properties, because It is mathematically sound, because
algebraic signs are ignored. algebraic signs are taken into account.
MEASURES OF SKEWNESS, KURTOSIS AND MOMENTS
Skewness
“Skewness means lack of symmetry or lopsidedness in a frequency distribution”. The object of
measuring skewness is to estimate the extent to which a distribution is
distorted from a perfectly symmetrical distribution. Skewness indicates whether the curve is
turned more to one side than to other, i.e., whether the curve has a longer tail on-one side.
Skewness can be positive as well as negative. Skewness is positive if the longer tail of the
distribution lies towards the right and negative if it lies towards the left.
Symmetrical Distribution
A frequency distribution is called symmetric if the frequencies are symmetrically distributed on
both sides of the centre point of the frequency curve. or A frequency distribution, in which the
values of mean, median and mode are equal, is called symmetrical distribution.
Skewed Distribution
frequency distribution which is not symmetrical is called skewed distribution. It
is of two types:
(a) Positively skewed
(b) Negatively skewed
Positively Skewed Distribution: A frequency distribution is said to be +ve skewed if the
frequency curve, gives a longer tail to the right hand side.
or
In a +ve skewed distribution, the value of mean is maximum and that of mode
is least, the median lies between mean and mode, i.e., Mode Mean Median +ve skewed
distribution
Mean > Median > Mode.
Negatively Skewed Distribution: A frequency distribution is said to be negatively skewed if the
frequency curve gives a longer tail on the left hand side. or In a –ve skewed distribution, the value
of the mode is maximum and that of mean is least, the median lies between mode and mean, i.e.,
Mode > Median > Mean
Moderately Symmetrical Distribution: If in a frequency distribution, the interval between the
mean and median is approximately one-third of the interval between the mean and mode, then the
distribution is called moderately symmetrical distribution.
DEFINING THE TERM ‘KURTOSIS’
Given two frequency distributions, which have the same variability as measured by
the standard deviation, they may be relatively more or less flat topped than the “Normal curve”. A
frequency curve may be symmetrical but it may not be equally
flat topped with the Normal curve. The relative flatness of the top is called Kurtosis
or convexity of the frequency [Link] enables us to have an idea about the “flatness or
peakedness” of the frequency curve.
Normal Curves or Mesokurtic Curves
The curves which are neither flat, nor sharply peaked, are known as Normal curves or mesokurtic
curves.
Platykurtic Curves
The curves which are flatter than the Normal curves, are known as platykurtic curves.
Leptokurtic Curves
The curves which are more sharply peaked than the normal curve, are known as leptokurtic
curves.
DEFINING ‘MOMENTS’
Business Statistics
Module 3
Fundamentals of sampling, Testing of Hypothesis,
Time series Analysis
Population: the elements about which we wish to make some inferences • Census: a census
involves complete details of the elements of a population
• Universe: the universe is the entire group of items the researcher wish to study and about which
they wish to generalize.
• Population element: the individual participant or object on which the measurement is taken
• Population Parameter: A parameter is a summary description of a fixed characteristic or measure
of the target population. A parameter denotes the true value which would be obtained if a census
rather than a sample was undertaken.
• Target population: the collection of elements or objects that possess the information about which
inferences are to be made. • Sample: a group of cases, participants, events, or records consisting of
a portion of the target population, carefully selected to represent that population
• Sample Statistic: A statistic is a summary description of a characteristic or measure of the
sample. The sample statistic is used as an estimate of the population parameter.
• Sampling unit: the basic unit containing the elements of the population to be sampled. •
Sampling frame: a representation of the elements of the target population.
Basics of Sampling Theory
Definition of sampling
According
Define
Pop to Levin and Rubin, Sam to refer not only to people,
statisticians use the word, population,
ulati Ele d
but, plinsample, to describe a portion
to all items that have been chosen for study. They use the word,
me target
onfrom the population.
chosen g
nt popula
fra
Sampling is the process of selecting a small number of elements from a larger defined target
tion me
group of elements such that the information gathered from the small group will allow judgments
to be made about the larger groups.
Fundamental of sampling
Sampling means obtaining information from a portion of a larger group or from universe.
Elements are selected in a manner that they yield all most all information about the whole
universe, if and when selected according to some scientific principles and procedures.
A sample refers to a smaller, manageable version of a larger group. It is a subset containing the
characteristics of a larger population. in other words, sampling is a process, which allows us to
study a small group of people from the large group to derive inferences that are likely to be
applicable to all the people of the large group. This is done as When a researcher conducts a
research, it’s rarely possible to collect data from every person in that group or to study the whole
population. Instead, the researcher selects a sample. Hence, a sample is the group of individuals
who will actually participate in the research. In order to get a clearer picture we must first
differentiate between population and a sample.
The population is the entire group that a researcher wants to draw conclusions about.
The sample is the specific group of individuals that a researcher will collect data from.
The population can be defined in terms of geographical location, age, income, and many other
characteristics.
Sampling Design
A sample design is a technique or the procedure that the researcher would adopt in selecting items
for the sample from a given population. Sample design is determined before data are collected.
Researcher must select/prepare a sample design which should be reliable and appropriate for his
research study.
STEPS IN SAMPLE DESIGN
[Link] of universe can be finite (number of items is certain, eg : number of workers in a factory)
or infinite (number of items is infinite, eg : number of stars in the sky).
2. Sampling unit • It may be a geographical one (state, district, village, etc.,) or a construction unit
(house, flat, etc.,) or a social unit (family, club, school, etc.,) or an individual.
3. Source list • It contains the names of all items of a universe.
4. Size of sample • It refers to the number of items to be selected from the universe to constitute a
sample
5. Parameters of interest : one must consider the specific population parameters which are of
interest (proportion of persons with some characteristic in the population)
6. Budgetary constraint • Cost considerations
7. Sampling procedure: The researcher must decide the type of sample he will use i.e., he must
decide about the technique to be used in selecting the items for the sample.
Developing a Sampling Plan
Define the Population of Interest
• Identify a Sampling Frame
• Select a Sampling Method
• Determine Sample Size
• Execute the Sampling Plan
•
Defining Population of Interest
Population of interest is entirely dependent on research design.
Some Bases for Defining Population:
• Geographic Area
• Demographics
• Usage/Lifestyle
• Awareness
Sampling Frame
A list of population elements (people, companies, houses, cities, etc.) from which units to be
sampled can be selected.
Difficult to get an accurate list.
Sample frame error occurs when certain elements of the population are accidentally omitted or
not included on the list.
sampling Errors
Sampling error is any type of bias that is attributable to mistakes in either drawing a sample or
determining the sample size.
By definition, when you have collected a sample from a population, you have less than complete
information about the population. This, in turn, means that there is a chance that the sample
statistics you calculate, (for example, the mean of a variable, a frequency distribution, etc.) may
not be an unbiased estimate of the population parameter.
Sampling error decreases rapidly as the sample size increases from a few hundred to about 1000
respondents. However, there is rarely any reason to select larger samples while comparing the
increased cost of survey with reduction in sampling error.
Non-Sampling Error
Before determining the sample size, we will briefly review other sources of error in surveys. When
you read a a news article that reports the results of a national poll, the error in the estimates is
always listed. However, experienced survey researchers know that errors due to other sources are
typically greater than the error due to sampling alone. Following are some other types of errors. •
Measurement errors, caused by poorly written questions, poorly designed questionnaires,
respondent errors in completing questionnaires, and so on.
• Non-response errors, caused because the respondents are not a representative subset of the
population.
• Data coding errors, caused, by errors in coding and entering the data. Of these error sources, the
first two are typically more severe. In mail surveys, non-response error is often the most serious
problem.
Types of sampling methods
Sampling are basically of two types – probability sampling and non-probability sampling.
Probability Nonprobability
sampling sampling
1. Probability sampling/ Random sampling: Probability sampling is defined as a sampling
technique in which the researcher chooses samples from a larger population using a method based
on the theory of probability. The researcher sets a few criteria and chooses members of a
population randomly. This is done so that all the members have an equal opportunity to be a part
of the sample with this selection parameter.
2. Non-probability sampling/ Non-random sampling: It is a sampling technique where the
samples are chosen deliberately and not randomly. Non-probability sampling is defined as a
sampling technique in which the researcher selects samples based on the subjective judgment of
the researcher rather than random selection. It is a less stringent method. This sampling method
depends heavily on the expertise of the researchers. It is carried out by observation, and
researchers use it widely for qualitative research. Non-probability sampling is a sampling method
in which not all members of the population have an equal chance of participating in the study,
unlike probability sampling. Each member of the population has a known chance of being
selected. Non-probability sampling is most useful for exploratory studies like a pilot survey
(deploying a survey to a smaller sample compared to pre-determined sample size). Researchers
use this method in studies where it is impossible to draw random probability sampling due to time
or cost considerations.
Types of probability sampling with examples:
There are four types of probability sampling techniques.
They are -simple random sampling, stratified random sampling, cluster sampling, and systematic
sampling Simple random sampling:
Simple random sampling
As the name suggests, is an entirely random method of selecting the sample. A simple random
sample is a is a randomly selected subset of a population in which each member of the subset has
an equal probability of being chosen to be a part of a sample. As such, a simple random sample is
an unbiased surveying technique. It is one of the best probability sampling techniques that helps in
saving time and resource. It is a reliable method of obtaining information where every single
member of a population is chosen randomly, merely by [Link] could be more accurately called
a randomly chosen sample.
Random samples are used to avoid bias and other unwanted effects. However, it isn’t quite as
simple as it seems; choosing a random sample isn’t as simple as just picking 100 people from
10,000 people. One has to be sure that one’s random sample is truly random and fairly
homogenous. ‘Lottery method’ is one example of random sample where the selection of items
entirely depends on luck or probability, and therefore this sampling technique is also sometimes
known as a method of chances. Another example of simple random sampling is the ‘use of random
numbers.’ The use of random numbers is an alternative method that also involves numbering the
population. The use of a number table similar to the one below can help with this sampling
technique.
Stratified random sampling:
Stratified random sampling is a method of sampling that involves dividing a population into
smaller sub-groups called strata. The groups or strata are organized based on the shared
characteristics or attributes such as gender, age, income range, job profile, or educational
attainment, etc. of the members in the group. While sampling, a researcher organises these groups
and then sample is drawn from each group separately using simple random sampling method.
Stratified random sampling is also known as quota random sampling and proportional random
sampling. For example, the company has 800 female employees and 200 male employees. You
want to ensure that the sample reflects the gender balance of the company, so you sort the
population into two strata based on gender. Then you use random sampling on each group,
selecting 80 women and 20 men, which gives you a representative sample of 100 people.
Types of non-probability sampling with examples
There are four types of non-probability sampling; they are- Convenience sampling, Judgmental or
purposive sampling, Quota sampling and Snowball sampling.
Convenience sampling ▫
It is used in exploratory research where the investigator is interested in getting an inexpensive
approximation of the fact. As the name implies, the sample is selected because it is convenient.
Also called haphazard or accidental sampling
Judgmental or purposive sampling:
Judgemental or purposive sampling method, researchers select the samples based purely on the
researcher’s discretion, knowledge and credibility. In other words, researchers choose only those
people who they deem fit to participate in the research study. In purposive sampling, we sample
with a purpose in mind. We usually would have one or more specific predefined groups we are
seeking. Judgmental or purposive sampling is not a scientific method of sampling, and the
downside to this sampling technique is that the preconceived notions of a researcher can influence
the results. Thus, this research technique involves a high amount of ambiguity. The purposive
sampling technique is most effective when one needs to study a certain cultural domain with
knowledgeable experts within. Purposive sampling may also be used with both qualitative and
quantitative research techniques.
Snowball sampling:
Snowball sampling is a sampling method that researchers apply when the subjects are inaccessible
or hard to find. In snowball sampling, the researcher begins by identifying someone who meets the
criteria for inclusion in the study at hand and the researcher asks them to recommend others who
they may know who also meet the criteria. This sampling system works like the referral program.
Snowball sampling is especially useful when the sample size is small and we are trying to reach
populations that are difficult to trace. Researchers also implement this sampling method in
situations where the topic is highly sensitive and not openly discussed—for example, surveys to
gather information about individuals with HIV/ AIDS. Not many victims will readily respond to
the questions.
Quota Testing of Hypothesis sampling:
In quota sampling, a quota is assigned Quota sampling is a non-probabilistic form of stratified
sampling. In this sampling method, the population is divided into strata or into mutually exclusive
sub-groups that are similar or homogenous and from which the sample items are selected on the
basis of a given quota or proportion. In quota sampling, care is taken to maintain the correct
proportions representative of the population. For example, if the population consists of 45%
female and 55% male, the sample should reflect those percentages.
Sampling Techniques
Probability
• Simple random sampling
• Systematic random sampling
• Stratified random sampling
• Cluster sampling
Nonprobability
• Convenience sampling
• Judgment sampling
• Quota sampling
• Snowball sampling
Sample size determined using the formula
The calculation of the sample size is concerned with the number of respondents required. To
determine the number to select for the sample drawn from the sampling frame, you must estimate
the non-response rate. The actual sample size to be drawn is: So, if any survey organization
decides that they need 700 respondents, and the expected response rate from the population is
50%, then 700/0.50, or 1400, customers must be drawn from the sampling [Link] of
Testing of Hypothesis:
In hypothesis testing, we decide whether to accept or reject a particular value of a set, of particular
values of a parameter or those of several parameters. It is seen that, although the exact value of a
parameter may be unknown, there is often same idea about the true value. The data collected from
samples helps us in rejecting or accepting our hypothesis. In other words, in dealing with
problems of hypothesis testing, we try to arrive at a right decision about a pre-stated hypothesis.
Definition:
A test of a statistical hypothesis is a two action decision problem after the experimental sample
values have been obtained, the two–actions being the acceptance or rejection of the hypothesis.
Statistical Hypothesis:
If the hypothesis is stated in terms of population parameters (such as mean and variance), the
hypothesis is called statistical hypothesis.
• Example: To determine whether the wages of men and women are equal.
• A product in the market is of standard quality.
• Whether a particular medicine is effective to cure a disease.
Hypothesis Testing Procedures
The following are the five steps in testing of hypothesis
1. Specify H0 and H1 , the null and alternative hypothesis and an acceptable level of signifance.
2. Determine an apporiate sample based test statistics and the rejection region for the specified H0.
3. Collect the sample data and calculate the test statistics.
4. Make a decision to either reject or fail to reject H0.
5. Interpret the results .
Types of hypothesis
Parametric Hypothesis:
A statistical hypothesis which refers only the value of unknown parameters of probability
distribution whose form is known is called a parametric hypothesis.
The null hypothesis0 (denoted by 0H0) is a statement that the value of a population parameter
(such as proportion, mean, or standard deviation) is equal to some claimed value.
We test the null hypothesis directly. Either reject H0 or fail to reject [Link] alternative hypothesis
(denoted by H1 or Ha or HA) is the statement that the parameter has a value that somehow differs
from the Null [Link] symbolic form of the alternative hypothesis must use one of these
symbols: , <, >.We have two kinds of alternative hypothesis:-
(a) One sided alternative hypothesis
(b) Two sided alternative hypothesis
The test related to (a) is called as ‘one – tailed’ test and those related to (b) are called as ‘two
tailed’ tests.
Critical region (or rejection region)
The critical region (or rejection region) is the set of all values of the test statistic that cause us to
reject the null hypothesis.
Acceptance and rejection regions in case of a two-tailed test with 5% significance level.
Significance Level
The significance level is the 0probability that the test statistic will fall in the 0critical region when
the null hypothesis is 0actually true. Common choices for are 0.05, 00.01, and 0.10.
Critical Value
A critical value is any value that separates the critical region (where we reject the null hypothesis)
from the values of the test statistic that do not lead to rejection of the null hypothesis. The critical
values depend on the nature of the null hypothesis, the sampling distribution that applies, and the
significance level .
Two-tailed, Right-tailed,Left-tailed Tests
The tails in a distribution are the extreme regions bounded by critical values.
Two TAILED Test
H0: = is divided equally between the two tails of the critical region
H1:
Right-tailed Test
H0: =
H1: >
Left tailed test
H0: =
H1: <
Type of Errors
Type I Error
• A type I error is the mistake of rejecting the null hypothsis when it is true.
• The symbol (alpha) is used to represent the probability of a type I error.
• A Type I Error is the mistake of rejecting the null hypothesis when it is true.
Type II Error : A Type II error is the mistake of failing to reject the null hypothesis when it is
[Link] symbol (Beta ) is used to represent the probability of a type IIerror.
There may be four possible situations that arise in any test procedure which have been summaries
are given below:
Decison
Accept H0 Reject H0
H0 is true
Correct Decision Type I Error
H0 is false Type II Error Correct Decision
Time series analysis
The term ‘Time Series’ consists of quantitative data which are arranged in the order of their
occurrence. For example, when we collect the data regarding population per capita income, sales
and prices of a commodity, etc. for a particular time period and arrange the data so obtained in a
series, called time series. Thus, according to ‘Spiegel’.
‘A time series is a set of observations taken at specified times, usually at equal intervals.’ In the
analysis of time series, time is the most important variable which may be either year, month, week,
day, hour or even minutes or seconds.
SIGNIFICANCE OF TIME SERIES ANALYSIS
The analysis of time series is of great importance because of the following reasons.
(a) To understand past behaviour: With the help of time series analysis, we can observe data over a
period of time and easily understand the changes which have taken place in the past. This helps us
in predicting the future behaviour.
(b) To predict future behaviour: With the help of time series analysis, we can plan for the future
business activities. There are so many statistical techniques by means of which we can analyse the
time series so obtained regarding the prediction of various future variations in business and
economics.
(c) To evaluate current accomplishments: By using time series analysis, we can investigate the
cause for the growth and decay of the achievements in business activities.
COMPONENTS OF TIME SERIES
The term ‘Trend’, is the basic tendency of production, sales and income etc. to grow or decline
over a period of time. Trend does not include short-range oscillations, it includes steady
movements over a long period of time. In business activities, we come across many economic
time series. Some series increase slowly and some increase fast. Some others series decrease at
varying rates, some remain constant for a long period of time. That is to say, we
have two types of trends which are:
(i) Linear or straight line trend
(ii) Non-linear Trend
There are mainly four types of components or patterns, or movements of a time series, given as
(a) Secular trend
(b) Seasonal variations
(c) Cyclical variations
(d) Irregular variations
(a) Secular Trend: The trends that occur as a result of general tendency of the data to increase or
decrease, over a long period of time, are known as secular [Link] we say that secular trend
refers to the general tendency of the data to increase or decrease over a long period of time.
(b) Seasonal Variations: The trends that take place during a period of 12 months as a result of
change in climate, weather conditions etc. are called seasonal variations or season variations are
those periodic movements in business activity which occur regularly every year and have their
origin in the nature of the year itself. There are some factors causing seasonal variation which are
as under:
(i) Weather and climate changes: It is the most important factor causing seasonal variations. The
change in the climate and weather conditions such as rainfall, humidity, etc. effects the different
products differently. For example, there is a greater demand of soft drinks and cotton clothes in
summer whereas, there is a greater demand of hot drinks and wollen clothes in winter.
(ii) Traditions and habits of a culture: For seasonal variations in time series, customs, traditions
and habits are important factor. For example, on certain occasions like Deepawali, there is a big
demand for sweets and also there is a great demand for cash before the festivals. Similarly, most of
the students buy their books in the first few months after the opening of schools and colleges.
Hence the sales of books, sweets etc. show seasonal variations.
(c) Cyclic Variations: Cyclic variations refer to the oscillatory variations in a time series which
have a duration anywhere between 2 to 10 years. These variations arise due to trade cycles or
business cycles. A business cycle has four phases namely
(i) Prosperity, (ii) Recession, (iii) Depression, (iv) Recovery.
(d) Irregular Variations: The variations in business activity which do not repeat in a definite
pattern are called irregular variations. They are also called erratic variations or accidental
variations or random [Link] variations take place due to special causes like floods,
earth quakes, strikes and wars etc. The declination in industrial output due to the strike in a factory
is an example of irregular variations.
METHODS OF MEASURING TREND
There are various methods that can be used for determining trend which are listed
below:
1. Free hand or graphic method
2. Semi-average method
3. Moving average method
4. Method of least square.
Freehand or Graphic Method: This is the simplest method of studying trend. The various steps
involved are given below.
Step I. Plot the given time series on a graph and examine the direction of the trend based on the
plotted information.
Step II. Draw a straight line (dotted) which will best fit to the data. The curve so obtained is called
freehand curve and this method is also called trend fitting by inspection.
Merits and Demerits of Freehand Curve Method
Merits
(i) It is the simplest method to measure trend.
(ii) It is flexible to use whether the trend is a straight line or a curve.
(iii) It requires no mathematical computations
(iv) It can be used to predict the future behaviour of business activity.
Demerits
(i) This method is subjective in nature as the trend line depends on the personal
judgement. It means different persons may draw different trend lines or
curves from the same set of given information.
(ii) It is time consuming.
(iii) It leaks accuracy.
Semi-average Method: In this method, the given data is divided into two
parts. The various steps involved are
Step I. After dividing the data into two parts, find A.M. of each part.
Step II. Plot the two values of A.M. obtained in step. I on the graph corresponding to the time
periods. Join these two points by a straight line, the straight line so obtained is the required trend
line.
‘Semi-average method’ can be applied in two situations given below:
(a) When the number of years given is even
(b) When the number of years given is odd.
Merits and Demerits of ‘Semi–Average Method’
Merits
1. It is the simplest method as compared to other methods like Moving Average
2. It is objective in nature as everyone will get the same trend for a given data.
3. It is time saving.
Demerits
1. It is based on straight line relationship between the plotted points. But this
relationship may or may not exist.
2. It is affected by extreme values. It means that if there are extremes in either
half or both halves of the series, the trend line so obtained will not give a
true picture of the growth factor.
3. Moving Average Method: In this method, we compute moving averages such as 3-yearly
moving average, 4-yearly moving average, 5-yearly moving average, etc. The period of moving
average is decided by keeping in mind the periodicity of data. The period is determined by plotting
the data on the graph paper and the average time interval of successive peaks or toughs are
noticed. While selecting the period of moving average, it is necessary to consider that after how
many years most of the fluctuations occur in the data.
Moving Average Method is studied in two different situations.
(a) Odd Period Moving Average
(b) Even Period Moving Average.
Odd Period Moving Average
When the period of moving average is odd, say, K years, where K is odd. The various
steps involved are:
Step I. Add all the values corresponding to first K years in the time series
Step II. Leave the first year, apply step I again. Continue this process further
till we reach the last value of the series.
Step III. Divide the moving totals obtained in step I and step II by the periods
of the moving average.
Step IV. The trend values (or moving averages) of different years.
Even Period Moving Average
If the moving average is an even period say, four yearly or six yearly, the moving averages are
placed at the centre of the time span. This can be done by the following two methods:
(a) Moving average by centering the totals
(b) Moving average by centering the averages.
Period of Moving Average
In time series analysis, we may come across some situations when the period of moving average is
not known. The period of the moving average for determining the trend values may be either 3
yearly, or 4 yearly of 5 yearly etc. or some other period. The basic principle in determining the
period of moving average is that it should be equal to the period of cyclic variations so that all
type of cyclical fluctuation are either eliminated or reduced to minimum.
In some other cases, the period of moving average in a series is not uniform. The cycle may
complete in five years or in seven years or in eight or nine years. Under such circumstances, the
average duration of the cycle is calculated and this calculated average is taken as the period of
moving average.
In most of the cases, the duration of the cycle is found out by plotting original data on a graph
paper and reading the time distances between various peaks or troughs. The average of these time
distances would give the average duration of the cycle and this is taken as the period of moving
average.
MEASUREMENT OF SHORT-TERM FLUCTUATIONS
‘Short-term fluctuations’ may be defined as the difference of trend values and the original data.
Merits and Demerits of Moving Average Method
Merits
The method of moving average has the following advantages.
1. It is a simple method if compared to the method of least squares.
2. It is a flexible method. It means, if some more figures are added to the data, the previous
calculations will not change. But we will get some more trend values.
3. It is most suitable method in eliminating cyclic fluctuations. If the period of moving average
happens to coincide with the period of cyclic fluctuations in the given data, cyclic fluctuations are
automatically eliminated.
Demerits
There are some limitations of this method which are given below.
1. Trend values cannot be obtained for all the years. For example, in a 3 yearly moving average,
trend values cannot be obtained for the first year and the last
year.
2. There is no hard and fast rule in selecting the period of moving average. One has to select the
period of moving average based on his own judgement.
3. It cannot be used in forecasting as this method is not represented by a math- ematical function.
4. It gives no appropriate computations when the trend is a straight line. The moving average lies
either above or below the true sweep of the data.
Method of Least Squares
One of the best method of trend fitting in a time series analysis is the method of least squares. This
method is widely used in practice. To fit a trend line, consider the
following conditions.
(i) Sum of the deviations of the actual value and computed trend value is zero i.e., Σ(Y – Yc) =
0, where
Y = the actual values
Yc = the trend values
(ii) Sum of the squares of the deviations of the actual values Y and computed trend values Yc is
least from this line, i.e., Σ(Y – Yc)2 = 0.
Business Statistics
Module 4
Correlation and Regression Analysis
DEFINITION OF CORRELATION
According to Ya Lun Chou, “correlation analysis attempts to determine the degree of relationship
between variables”.
According to W.I. King, “correlation means that between two series or group of data there exists
some casual connection”.
According to Croxton and Cowden, “the relationship of quantitative nature, the appropriate
statistical tool for discovering and measuring the relationship and expressing it in brief formula is
known as correlation”.
According to A.M. Tuttle, “correlation is an analysis of the covariation between two or more
variables”.Thus, the association of any two variables is known as correlation. In otherwords, we
say that corresponding to a change in one variable there is a change in another variable, they are
said to be correlated. This change may be in either direction. If one variable increases (or
decreases) the other may also increases (or decreases). The above definitions make it clear that the
term “correlation refers to the study of relationship between two variables”.
TYPES OF CORRELATION
In a bivariate distribution, correlation is classified into many types, but the important are:
1. Positive and negative correlation
2. Simple and multiple correlation
3. Partial and total correlation
4. Linear and non-linear correlation.
Positive and Negative Correlation
Positive and negative correlation depend upon the direction of the change of the variables. If two
variables tend to move together in the same direction i.e., an increase or decrease in the value of
one variable is accompanied by an increase or decrease inthe value of other variable, then the
correlation is called positive or direct correlation.
Height and weight, age of wife and husband, intake of calories and proteins, rainfall and yield of
crops, price and supply are examples of positive correlation.
If two variables tend to move together in opposite directions i.e., an increase or decrease in the
value of one variable is accompanied by a decrease or increase in the value of other variable, then
the correlation is called negative or inverse correlation. Price and demand, yield of crops and
price, literacy status and total fertility rates among adult female population, rise in prices and
consumption of qualitative food (milk or eggs) etc. are the examples of negative correlation.
Simple and Multiple Correlation
When we study only two variables, the relationship is described as simple correlation. For
examples, the yield of wheat and use of fertilizers, plant yield and number of tillers, number of
pods and number of clusters, quantity of money and price level, demand and price etc. But in a
multiple correlation, we study more than two variables simultaneously.
However, multiple correlation consists of the measurement of the relationship between a
dependable variable and two or more independent variables.
For examples, when we study the relationship between plant yield with that of a number of pods
and a number of clusters in pulses and if we study the relationship between agricultural
production, rainfall and quantity of fertilizers used, it will be a multiple correlation.
Partial and Total Correlation
To study of two variables excluding some other variables is called partial correlation. For
examples, the correlation between yield of maize and fertilizers excluding, the effect of pesticides
and manures is called partial correlation, we study price and demand, eliminating the supply side
is called partial correlation. In total correlation, all the facts are taken into account.
Linear and Non-linear Correlation
If the ratio of change between two variables is uniform, then there will be linear correlation
between [Link] ratio of change between the variables is the same. If we plot these on the graph,
we get a straight line.
In a curvilinear or non-linear correlation, the amount of change in one variable
does not bear a constant ratio of the amount of change in the other variable. The graph of non-
linear or curvilinear relationship will form a curve.
METHODS OF STUDYING CORRELATION
The different methods of finding out the relationship between two variables are:
[Link] Method
1. Scatter diagram or scattergram or dot diagram,
2. Simple graph or correlation graph.
[Link] Method
3. Karl Pearson’s coefficient of correlation,
4. Spearman’s coefficient of rank correlation,
5. Coefficient of concurrent deviation,
6. Method of least squares.
Scatter Diagram Method
This is the simplest method of finding out whether there is any relationship present between two
variables by plotting the values on a chart, known as scatter diagram.
By this method a rough idea about the correlation of two variables can be judged. In this method,
the given data are plotted on a graph paper in the form of dots. X variables are plotted on the
horizontal axis and Y variables on the vertical axis. Thus, we have the dots and we can know the
scatter or concentration of the various points.
If the plotted points form a straight line running from the lower left-hand corner to
the upper righthand corner, then there is a perfect positive correlation (i.e., r = +1). On the other
hand, if the points are in a straight line, having a falling trend from the upper left-hand corner to
the lower right-hand corner, it reveals that there is a perfect negative or inverse correlation (i.e., r
= – 1). If the plotted points fall in a narrow band, and the points are rising from lower left-hand
corner to the upper right-hand corner, there will be a high degree of positive correlation between
the two variables. If the plotted points fall in a narrow band from the upper left-hand corner to the
lower right-hand corner, there will be a high degree of negative correlation. If the plotted points lie
scatter all over the diagram, there is no correlation between the two variables.
Merits and Demerits of Scatter Diagram
Merits
(i) It is a simple and attractive method of finding out the nature of correlation between two
variables.
(ii) It is a non-mathematical method of studying correlation. It is easy to understand.
(iii) We can get a rough idea at a glance whether it is positive or negative correlation.
(iv) It is not affected by the extreme values.
(v) The correlation between two variables can be known only on the basis of
diagram.
Demerits
(i) It gives only a rough idea about the correlation.
(ii) It does not give the degree or extent of relationship between two variables.
Simple Graph
The values of two variables are plotted on a graph paper, we get two curves, one for X variable
and another for Y variable. These two curves reveal the direction and closeness of the two curves
and reveal whether or not the variables are related. If both the curves move in the same direction
i.e., parallel to each other, either upward or downward, correlation is said to be positive.
Karl Pearson’s Coefficient of Correlation
Karl Pearson, a great biometrician and statistician, suggested a mathematical method for
measuring the magnitude of linear relationship between two variables. Karl Pearson’s method is
the most widely used method in practice and is known as Pearsonian coefficient of correlation. It
gives information about the direction as well as the magnitude of the relationship between two
variables. It is denoted by the symbol ‘r.
Merits and Demerits
Merits
(i) It is the best measure as far as the algebraic point of view is concerned because it is based on
all the observations of both the series.
(ii) It is the most popular mathematical method used for measuring the degree of relationship.
(iii) It is an ideal measure as far as the biostatistical point of view is concerned because it is based
on arithmetic mean and standard deviation, which are the best measures of central tendency and
dispersion respectively.
(iv) The main features of this coefficient is that it gives information about the direction as well as
the magnitude of the relationship between the two variables.
Demerits
(i) It assumes linear relationship between two variables but in practice it is not always possible.
(ii) It lies between + 1 and –1 need a very careful interpretation, otherwise it will be
misinterpreted. Careless interpretation will be fallacious.
(iii) It is affected by extreme values.
Degree of Correlation
The degree of correlation between two variables can be ascertained by the quantitative value of
coefficient of correlation which can be found out by calculation, Karl Pearson has given a formula
for measuring correlation coefficient (r). However, the results of this formula varies between + 1
and – 1. In case of perfect positive correlation, the result will be r = + 1 and in case of perfect
negative correlation, the result will be r = – 1. However, in the absence of correlation, the result
will be r = 0. It indicates that the degree of scattering is very large. In experimental research, it is
very difficult to find such values of r as + 1, – 1 and 0.
Coefficient of Correlation and Probable Error
To find out the reliability or the significance of the value of Karl Pearson’s coefficient of
correlation, probable error is used. With the help of probable error the limits of coefficient of
correlation are obtained and the reliability of the value of the coefficient is assessed.
Spearman’s Coefficient of Rank Correlation
In 1904, Charles Edward Spearman, a British psychologist found out the method ofascertaining
the coefficient of correlation by ranks. This method is based on rank.
This measure is useful in dealing with qualitative characteristics, such as intelligence, beauty,
morality, character etc. It cannot be measured quantitatively, as in the case of Karl Pearson’s
coefficient of correlation, but it is based on the ranks given to the observations. It can be used
when the data are irregular or extreme items are erratic or inaccurate, because coefficient of rank
correlation is not based on the assumption of formality of data.
MULTIPLE AND PARTIAL CORRELATION
When the values of one variable are associated with or influenced by other variable, i.e., the age of
husband and wife, the height of father and son, the supply and demand of a commodity and so on,
Karl Pearson’s coefficient of correlation can be used as a measure of linear relationship between
them. But sometimes there is interrelation between many variables and the value of one variable
may be influenced by many others, e.g., the yield of crop per acre say (x1) depends upon quality
of seed (x2),fertility of soil (x3), fertilizer used (x4), irrigation facilities (x5), weather conditions
(x6) and so on. Whenever we are interested in studying the joint effect of a group of variables
upon a variable not included in that group, our study is that of multiple correlation and multiple
regression.
Coefficient of Multiple Correlation
In a trivariate distribution in which each of the variables x1, x2 and x3 has N observations, the
coefficient of multiple correlation of x1 on x2 and x3 usually denoted by R1.23.
REGRESSION ANALYSIS AND PROPERTIES OF REGRESSION
COEFFICIENTS
In regression, we intend to describe the dependence of a variable on an independent variable. We
employ regression equations to lend support to hypothesis regarding the possible causation of
changes in y by changes in x, for purposes of prediction of y in terms of x, and for purposes of
explaining some of the variations of y by x, by using the latter variable as a statistical control.
Studies of the effects of temperature on heartbeat rate, protein intake on growth rate in a child, age
of person on blood pressure, and dose of an insecticide on mortality of the insect population are all
typical examples of regression.
In regression analysis, we can predict or estimate the value of one variable from the
given value of the other variable. Regression explains the functional form of two variables one as
dependent variable and other as independent variable. For examples, “temperature and oxygen
content of water are correlated. We can find out the expected amount of the dissolved oxygen for a
given temperature”. “Age and blood pressure are correlated, we may find the expected amount of
systolic blood pressure for a given age. Thus the regression of the systolic blood pressure readings
(y) on the age of subjects (x)”. “The regression of gain in height or weight (y) on the levels of
protein or calorie intake (x)”.
Uses of Regression Analysis
Regression analysis is useful in many scientific studies.
(i) Regression analysis is used in biostatistics in all those fields where two or more relative
variables are having the tendency to go back to the average.
(ii) Regression analysis predicts the value of dependent variables from the values of independent
variables.
(iii) Regression analysis is highly useful and the regression line or equation helps to estimate the
value of dependent variable, when the values of independent variables are used in the equation.
(iv) We can calculate the coefficient of correlation (r) with the help of regression coefficients.
(v) Regression analysis in statistical estimation of demand curves, supply curves, production
function, cost function, consumption function, etc. can be predicted.
DIFFERENCE BETWEEN CORRELATION AND REGRESSION
CORRELATION REGRESSION
Correlation is the relationship between two or Regression means going back and it is a
more variables, which vary in sympathy with the mathematical measure showing the average
other in the same or the opposite relationship between two variables.
direction.
Both the variables x and y are random x is a random variable and y is a fixed variable.
variables. Sometimes both the variables may be random
variables.
It find out the degree of relationship between It indicates the cause and effect relationship between
two variables and not the cause and effect of the the variables and establishes a function relationship.
variable.
t is used for testing and verifying the relation Besides verification, it is used for the prediction of
between two variables and gives limited one value, in relationship to the other given value.
information.
The coefficient of correlation is a relative Regression analysis is an absolute measure. If we
measure. The range of relationship lies know the value of independent variable, we can find
between +_ 1. the value of the dependent variable.
There may be nonsense correlation between two In regression, there is no such nonsense regression.
variables.
It has limited application, because it is It has wider application, as it studies linear
confined only to linear relationship between the and non-linear relationship between the variables.
variables.
If the coefficient of correlation is positive, The regression coefficient explains that the
then the two variables are positively decrease in one variable is associated with
correlated and vice versa. the increase in the other variable.
REGRESSION EQUATIONS
If the bivariate data are plotted on a graph paper, a scatter diagram is obtained which
indicates some relationship between two variables. The dots of scatter diagram tend to concentrate
around a curve. This curve is known as regression curve and its functional form is called
regression equation. When this curve is a straight line, it is
called regression line and regression is said to be linear and if it is a curve, it is called non-linear
regression.
In bivariate data we have two variables and therefore, we have two regression lines
as follows:
(i) Regression line of y on x, it is denoted by
y = a + bx
where x is independent variable and y is dependent variable.
(ii) Regression line of x on y, it is denoted by
x = a + by
where y is independent variable and x is dependent variable.
When the regression lines show some trend upward or downward, we say that there is some
correlation between two variables. But if both the lines of regression are perpendicular to each
other, we say that both the variables are uncorrelated or r = 0. Further, if both the lines of
regression coincide, we say that there is a perfect correlation between the variables or r = 1.
METHODS OF FITTING REGRESSION LINES
There are mainly three methods for fitting of regression lines:
(i) Method of least squares (normal equations)
(ii) Deviation from actual mean method.
(iii) Deviation from assumed mean method.
(i) Method of Least Squares: If the set of paired data gives the indication that regression is linear,
then we can fit two regression lines (a) y on x (b) x on y.
(ii) Deviation from Actual Mean Method: If the arithmetic mean of both the series x and y are not
in fraction, this method is suitable for fitting regression lines. This method is easier and simpler to
calculate than the previous method, which is a tedious one. We can find out the deviations of x and
y series from their respective means.
(iii) Deviation Taken from the Assumed Mean Method: The difference between the above said
method and this is that instead of taking deviations from the arithmetic mean, we take deviations
from the assumed mean. If the actual mean is in fraction, this method can be used.
MULTIPLE REGRESSION
The concepts and techniques for analysing the association among three or more variables are
natural extensions of those explored in the bivariate situation discussed so far. In multiple
regression model, we assume that a linear relationship exists between some variable y, which we
call the dependent variable, and k independent variables x1, x2, ..., xk. The independent variables
are sometimes referred to as explanatory variables, because of their use in explaining the variation
in y; or as predictor variables, because of their use in predicting y. In general, we ought to be able
to improve our predicting ability by including more independent variables in such an
[Link] Regression Equation: Regression equation of x1 on x2 and x3.
*************************************************************