0% found this document useful (0 votes)
22 views125 pages

MBA Business Statistics Syllabus

The document outlines the syllabus for a Business Statistics course. It covers four units: (1) measures of central tendency and dispersion, time series analysis, and correlation; (2) index numbers, regression, probability distributions, and hypothesis testing; (3) estimation theory and hypothesis testing techniques; (4) suggested readings. The course provides an introduction to statistical concepts and methods used in business decision making and research.

Uploaded by

Bablu Prasad
Copyright
© Attribution Non-Commercial (BY-NC)
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
22 views125 pages

MBA Business Statistics Syllabus

The document outlines the syllabus for a Business Statistics course. It covers four units: (1) measures of central tendency and dispersion, time series analysis, and correlation; (2) index numbers, regression, probability distributions, and hypothesis testing; (3) estimation theory and hypothesis testing techniques; (4) suggested readings. The course provides an introduction to statistical concepts and methods used in business decision making and research.

Uploaded by

Bablu Prasad
Copyright
© Attribution Non-Commercial (BY-NC)
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

MBA 015

BUSINESS STATISTICS
Syllabus
Unit I
Role of statistics: Applications of inferential statistics in managerial decision-making; Measures of central tendency: Mean, Median and Mode and their implications; Measures of Dispersion: Range, Mean deviation, Standard deviation, Coefficient of Variation (CV), Skewness, Kurtosis.

Unit II
Time series analysis: Concept, Additive and Multiplicative models, Components of time series, Trend analysis: Least Square method Linear and Non-Linear equations, Applications in business decisionmaking. Index Numbers: Meaning, Types of index numbers, uses of index numbers, Construction of Price, Quantity and Volume indices: Fixed base and Chain base methods. Correlation: Meaning and types of correlation, Karl Pearson and Spearman rank correlation. Regression: Meaning, Regression equations and their application, Partial and Multiple correlation and regression: An overview.

Unit III
Probability: Concept of probability and its uses in business decision-making; Addition and multiplication theorems; Bayes Theorem, and its applications. Probability Theoretical Distributions: Concept and application of Binomial; Poisson and Normal distributions.

Unit IV
Estimation Theory and Hypothesis Testing: Sampling theory; Formulation of Hypotheses; Application of z-test, t-test, F-test and 2 (chi square) test. Techniques of association of Attributes and Testing. Suggested Reading 1. Beri Statistics for Management (Tata McGraw-Hill, 2nd edition). 2. Chandan J S Statistics for Business and Economics (Vikas, 1998, 1st edition). 3. Render and Stair Jr Quantitative Analysis for Management (Prentice-Hall, 7th edition). 4. Sharma J K Business Statistics (Pearson Education, 2nd edition). 5. Gupta C B, Gupta V An Introduction to Statistical Methods (Vikas, 1995, 23rd edition). 6. Levin Rubin Statistics for Management (Pearson 2000, New Delhi, 7th edition). 7. Earshot L- Essential Quantitative Methods for Business Management and Finance (Palgrave, 2001).
MBA-015 Business Statistics Page 1

THE GREEK ALPHABET


Upper Case Lower Case Pronunciation Upper Case Lower Case Pronunciation

, ,

Alpha Beta Gamma Delta Epsilon Zeta Eta Theta Iota Kappa Lambda Mu

, , , ,

Nu Xi Omicron Pi Rho Sigma Tau Upsilon Phi Chi Psi Omega

MBA-015

Business Statistics

Page 2

UNIT I Introduction
The science of statistics has assumed great importance in recent years. Today it has become an allimportant science, without which no other science can progress. Modern age is the age of statistics and it can be correctly said that the extent of economic development of a country can be best known by finding out the extent to which statistical organization has developed there. In India, the importance of statistics increased considerably after independence, with the start of the era of economic planning. In fact, economic planning cannot be imagined in the absence of statistical data. Today, statistics is a very commonly used word but surprisingly enough, it is understood by different people in different senses. To some, the word statistics connotes just a set of figures like the number of children born in India, in various months of a particular year and their sex. Some people think of statistics as something used to reinforce a qualitative statement. Thus, when someone says that the economic condition of the Indian masses has improved during the last five years, and gives the figures of rising per capita income during this period, he is using statistics to support a qualitative statement. To many others statistics means the representation of a phenomenon with the help of figures, charts, diagrams and pictograms, etc. There are many who think of statistics as a complex set of relationships between a number of variables and also as a technique by which they can reduce the element of uncertainty in their decision process. It is seen by many as a device to achieve a degree of precision in the concept and theories of social sciences which by nature are inexact. However, if we analyse the way the word statistics is looked at, we find that broadly speaking there are two categories one, in which the word refers to a set of figures and other in which it refers to a set of techniques and methods. Therefore, in common parlance, the word statistics denotes some sort of numerical data. If, for example, somebody says that he has studied the statistics of man-hours lost by the Indian cotton mills due to strikes in the year 2002, or that he has seen the statistics of automobile accidents in the USA, he refers to the numerical figures or data relating to these phenomena. In this sense, statistics are numerical descriptions of the quantitative aspects of things. They take the form of counts or measurements. Statistics about the membership of a certain hostel, for example, include a count of the number of members of various kinds, as postgraduate or undergraduate or over and under 21 years of age. They might include such measurements as the weights and heights of the members, etc., or the ratio between weights and heights. The use of the word statistics in this sense is always in plural. However, any figure or set of figures cannot be called statistics irrespective of any other
MBA-015 Business Statistics Page 3

consideration. Many things have to be taken into account before using the word statistics for any group of figures. However, instead of using the word statistics in the above sense for indicating numerical facts, data is a more appropriate term. The second sense in which the word statistics is used refers to the techniques and methods used in the collection, analysis and interpretation of data. In this sense the word is used in singular. Thus Statistics are numerical facts, but Statistics is also a body of methods for making decisions when there is uncertainty arising from the incompleteness or the instability of the information available. The decisions may be required to be taken either for the practical purpose of selecting a course of action or for the scientific purpose of gaining knowledge.

Definition:
The term statistics has been defined differently by different authors. Some authors have defined the word as used in the first sense (numerical data) while others have defined it as used in the second sense (the science). Of the first type of definitions, the one given by Horace Secrist is the most exhaustive: By statistics we mean aggregates of facts affected to a marked extent by multiplicity of causes numerically expressed, enumerated or estimated according to reasonable standards of accuracy, collected in a systematic manner for a pre-determined purpose and placed in relation to each other. The above definition makes it clear that statistics (as numerical data) should possess the following characteristics: (i) They should be aggregates of facts.

(ii) They should be affected to a marked extent by multiplicity of causes. (iii) They should be numerically expressed. (iv) They should be enumerated or estimated according to reasonable standards of accuracy. (v) They should be collected in a systematic manner. (vi) They should be collected for a predetermined purpose. (vii) They should be placed in relation to each other. Another definition of statistics states: Statistics are the classified facts representing the conditions of the people in a State ... specially those facts which can be stated in numbers or in tables of numbers or in any tabular or classified arrangement. This definition is rather narrow, but it gives an idea about the origin of the term statistics.

MBA-015

Business Statistics

Page 4

Of the second type of definitions of the tern statistics (as statistical methods or the science of statistics), the one given by Seligman is very short and simple and yet quite comprehensive: Statistics is the science which deals with the methods of collecting, classifying, presenting, comparing and interpreting numerical data collected to throw some light on any sphere of enquiry. According to Wallis and Roberts: Statistics is a body of methods for making decisions in the face of uncertainty. On the basis of the above definitions, we may evolve a comprehensive definition of statistics on our own: Statistics (as used in the sense of data) are numerical statements of facts capable of analysis and interpretation and the science of statistics is a study of the principles and methods used in the collection, presentation, analysis and interpretation of numerical data in any sphere of enquiry.

Role of Statistics
Statistical methods are applicable to a very large number of fields. They are used in social sciences like economics, anthropology, sociology and psychology. Even professional disciplines like medicine, engineering and business rely heavily on statistical methods. Statistical methods are applied, to a limited extent, even in the realm of physical sciences which are by nature exact and experimental in character. In fact, statistical methods provide a set of tools that can be profitably used by professionals in different disciplines in the manner in which they deem fit. Statistics is related to various other sciences and disciplines such as economics, mathematics, astronomy, biology, meteorology, etc. in fact the science of statistics is associated with all the other important sciences both physical as well as social. Bowley has rightly said, A knowledge of statistics is like a knowledge of a foreign language it may prove of use at any time under any circumstances. Statistics finds use in economics, planning, commerce, business management, matters of governance, research, war, etc. It is useful for bankers, brokers and insurance companies also.

APPLICATION OF INFERENTIAL STATISTICS IN MANAGERIAL DECISION-MAKING


Inferential Statistics: Inferential statistics is concerned with making inferences from samples about the populations from which they have been drawn. In other words, if we find a difference between two samples, we would like to know, is this a "real" difference (that is, is it present in the population) or just a "chance" difference (that is, it could just be the result of random sampling error). That's what tests of statistical significance are all about. Any inferred conclusion from a sample data to the
MBA-015 Business Statistics Page 5

population from which the sample is drawn must be expressed in a probabilistic term. Probability is the language and a measuring tool for uncertainty in our statistical conclusions. Decision-making implies selecting the best alternative from among those available. Statistics can play a major role in analyzing the information available, on whose basis the various alternatives have to be compared. In order to bring about uniformity and consistency in their decisions, organizations need to remove subjectivity and introduce objectivity in the process of decision-making. Quantitative techniques in general, and among them, statistics can help in accomplishing this. Since the complexity of business environment makes the process of decision-making difficult, the decision-maker cannot rely entirely upon his observation, experience or evaluation to make decision. Decisions have to be based upon data which show relationship, indicate trends, and show rates of change in various relevant variables. The field of statistics provides methods for collecting, presenting, analyzing and meaningfully interpreting data. Thus, the statistical methodology in collection, analysis and interpretation of data for better decision-making is a sine qua non for managerial decision-making and research in both physical and social sciences, particularly in business and economics.

STATISTICS

Descriptive Statistics

Inductive Statistics

Statistical Decision Theory

Data Collection and Presentation

Statistical Inference (Hypothesis Testing) and Estimation (Regression and Correlation)

Decision Problems, Alternatives, Uncertainties, Consequences, Criterion of Choice

Statistical data constitute the basic raw material of statistical methods. These data are either readily available or collected by the analyst. The manager may face four types of situations: (i) When data need to be presented in a form which helps in easy grasping;

(ii) Where no specific action is contemplated but it is intended to test some hypotheses and draw inferences;
MBA-015 Business Statistics Page 6

(iii) When some unknown quantities have to be estimated or relationships established through observed data; and (iv) When a decision has to be made under uncertainty regarding a course of action to be followed. While the first situation falls in the realm of descriptive statistics, second and third situations fall in the area of inductive statistics, and the last situation is dealt by statistical decision theory. Thus, descriptive statistics refers to the analysis and synthesis of data. Statistical decision theory concerns itself with establishment of rules and procedures for choosing one course from the alternative courses of action under situations of uncertainty. Inductive statistics is concerned with the development of scientific criteria so that values of a group may be meaningfully estimated by examining only a small portion of that group. The group is known as population or universe and the portion is known as sample. Further, values in the sample are known as statistics and values in the population are known as parameters. Thus, inductive statistics is concerned with estimating universe parameters from the sample statistics. The term inductive statistics is derived from the inductive process which tries to arrive at information about the general (the universe) from the knowledge of the particular (the sample).

Limitations of Statistics:
1. Does not study qualitative phenomena. 2. Does not reveal the entire story. 3. Statistical laws not universally applicable. 4. Does not study individuals. 5. Is liable to be misused.

MBA-015

Business Statistics

Page 7

MEASURES OF CENTRAL TENDENCY


Condensation of data is necessary in statistical analysis because a large number of big figures are not only confusing to mind but difficult to analyse also. In order to reduce the complexity of data and to make them comparable it is essential that the various phenomena which are being compared are reduced to one figure each. If, for example, a comparison is made between the marks obtained by a group of 200 students belonging to a university and marks obtained by another group of 200 students belonging to another university, it would be impossible to arrive at any conclusion if the two series relating to these marks are directly compared. On the other hand, if each of these series is represented by one figure, comparison would be an extremely easy affair. It is obvious that a figure which is used to represent a whole series should neither have the lowest value in the series nor the highest value but a value somewhere between these two limits, possibly in the centre, where most of the items of the series cluster. Such figures are called the measures of central tendency or averages. The average represents a whole series and as such, its value always lies between the minimum and maximum values and generally it is located near the centre or middle of the distribution. The word average or the term measure of central tendency can be defined as a typical value around which other figures congregate.

Objectives of Averaging
(i) To get a single value which is representative of the characteristics of the entire mass of data. (ii) To facilitate comparison.

Characteristics of a Good Average


(i) It should be rigidly defined.

(ii) It should be based on all the observations of the series. (iii) It should be capable of further algebraic treatment. (iv) It should be easy to calculate and simple to follow. (v) It should not be affected by fluctuations of sampling.

Measures of Various Orders


Statistical series may differ from each other in the following three ways: (i) They may differ in the values of the variable around which most of the items cluster.

(ii) They may differ in the extent to which items are dispersed around the central value. (iii) They may differ in the extent of departure from a normal distribution.
MBA-015 Business Statistics Page 8

Accordingly, there are three measures designed to study the above differences. They are respectively known as: (i) Measures of the first order, or measures of central tendency.

(ii) Measures of the second order, or measures of dispersion. (iii) Measures of the third order, that is, skewness, kurtosis, etc.

Types of Averages
(a) Mathematical Averages: 1. Arithmetic Average or Mean. 2. Geometric Mean. 3. Harmonic Mean. (b) Averages of Position: 1. Median. 2. Mode. Besides these there are some less important averages like the quadratic mean. There are also some averages which are mostly calculated by using the technique of the arithmetic average in a modified form. Examples of such averages are weighted average, moving average and progressive average. These find use in the analysis of commercial statistics and the utility of moving average is very great in the analysis of time series. Of the above, mean, median and mode are the most popular ones, and we shall limit our discussion to these only.

MEAN (ARITHMETIC AVERAGE)


Arithmetic average or mean of a series is the figure obtained by dividing the total values of the various items by their number. Calculation of Mean in a Series of Individual Observations: Direct Method:

Where X = arithmetic average, X = values of the variable, = summation or total,


MBA-015 Business Statistics Page 9

N = number of items. Short-cut Method:

Where A = assumed mean, dx = sum of deviations from assumed mean. Calculation of Mean in a Discrete Series: Direct Method:

Where fX = sum of products of X and corresponding frequencies f = sum of frequencies Short-cut Method:

Where fdx = sum of product of deviations from assumed mean and respective frequencies. Step-Deviation Method:

Where dx = deviations from assumed mean divided by a common factor, i = the common factor. Calculation of Mean in a Continuous Series: The process is the same as in case of a discrete series. The same formulae can be used after finding the mid-values of various classes. The method remains the same for both the exclusive as well as inclusive classes. Mean can also be calculated if class intervals are unequal. However, mean cannot

MBA-015

Business Statistics

Page 10

be calculated in case of open-ended classes. (In such a situation, mean may be calculated by assuming the missing class limits.) Direct method:

Shortcut method:

Step-deviation method:

Mathematical Properties of Mean (i) The sum of the deviations of the items from the mean is always equal to zero.

(ii) The sum of squared deviations of the items from the mean is less than the sum of squared deviations of items from any other value. (iii) If the mean values of two or more series are known individually, then the mean of the combined series formed by taking the observations of all these series together can be easily expressed in terms of the weighted average of the means of the component series. (iv) The mean of all the sums (or differences) of corresponding observations in two series (with equal number of observations) is equal to the sum (or difference) of the means of the two series. Merits of Mean (i) It is rigidly defined.

(ii) It is based on all the observations of the series. (iii) It is capable of further algebraic treatment. (iv) It is simple to follow, and, arguably, it is also easy to calculate. (v) It is least affected by fluctuations of sampling. Drawbacks of Mean (i) Sometimes abnormal items may considerably affect the value of mean.

(ii) It cannot be calculated if the exact values of all the observations are not known. (iii) It is more difficult to calculate than median and mode. (iv) It can have a value that does not even exist in the series.

MBA-015

Business Statistics

Page 11

(v) As a corollary to the above point, mean may sometimes give results that appear absurd. For instance, we may find that the average number of children in the families of a certain region is 3.4! (vi) It may give fallacious conclusions. In such situations, measures of the second and third order may be needed to be calculated. (vii) Mean has an upward bias, that is, it gives more importance to greater items of a series and less importance to smaller items.

MEDIAN
Median is the value of the middle item of a series arranged in ascending or descending order of magnitude. As opposed to mean which is the mathematical average, median is the positional average. It divides the series into two equal parts, one containing items smaller than the median value and the other containing items greater than it. The value of the median is given by the value of the middle item, irrespective of all other values. Calculation of Median The calculation of median involves two steps: (i) Finding the location of the middle item; and

(ii) Finding the value of this middle item. The middle item in a series of individual observations and also in a discrete series is where n is the total number of observations. In case of a continuous series term of the series. Once the middle item is located its value has to be found out. In a series of individual observations, if the total number of items is an odd figure, the value of the middle item is the median value. On the other hand, if the number of items is even, the median value is the average of the two items in the centre of the distribution. Calculation of Median in a Series of Individual Observations/Discrete Series: (i) Arrange the values in ascending or descending order of magnitude. (ii) M = value of item
th th

item,

item is the middle

Calculation of Median in a Continuous Series: (i) Arrange the values in ascending or descending order of magnitude.

MBA-015

Business Statistics

Page 12

(ii) Find the value of (iii) a.

to identify the median class.

b. where M = median, L1 = lower class limit of median class, L2 = upper class limit of median class, N = total number of observations, cf = cumulative frequency of the class preceding the median class, f = frequency of the median class, i = class interval of the median class. Merits of Median (i) It is rigidly defined.

(ii) It is easily calculated (sometimes it can be located by mere visual inspection) and understood. (iii) It is not affected by the values of the extreme items. (iv) Even if all values are not known, median can be calculated. (v) It gives best results in a study of those phenomena which are incapable of direct quantitative measurement, for example, intelligence. Drawbacks of Median (i) It may not be representative of the series in certain cases. This is specially so when there are wide variations between the values of different items. (ii) It is not suitable for further algebraic treatment. (iii) In case of continuous series, calculation of median requires interpolation which assumes that all the values of a class are evenly spread over the values in that particular class. (However, this drawback is also present in case of mean and mode.) (iv) Median does not consider the values of all the items. Even if the values of some of the items are changed, the median may remain unchanged. (This is a drawback of mode also.) (v) It is more likely to be affected by the fluctuations of sampling. (vi) It cannot be calculated without rearranging the data in ascending or descending order.
MBA-015 Business Statistics Page 13

Fractiles (Quartiles, Deciles, Percentiles)


Just as median is appositional average and partitions the series into two equal parts, similarly there are other positional values which divide a series into different number of equal parts. The most common of these partition or positional values (fractiles) are quartiles, deciles and percentiles.

Quartiles
The values which divide the given data into four equal parts are known as quartiles. Obviously, there will be three such values. The first quartile, which is also called the lower quartile and is represented by Q1, covers the first 25% items of the series so that 75% of the items are greater than it. It divides the first half of the series into two equal parts. Similarly, the second quartile (Q2) covers the first 50% items of the series and divides it into two equal parts (It is the same as median). The third quartile, also called the upper quartile (Q3) covers the first 75% of the items of the series and divides the second half of the series into two equal parts. Before calculation of quartiles, as in the case of median, the series has to be arranged in ascending or descending order. Q1 < Q2 < Q 3

Where Qn = nth quartile, cf = cumulative frequency of the class preceding nth quartile class, f = frequency of the nth quartile class, i = class interval of the nth quartile class.

Deciles
The values which divide the given data into ten equal parts are known as deciles. There are nine deciles denoted by D1, D2, D3, D4, D5, D6, D7, D8, and D9. The fifth decile (D5)is equal to the median.

MBA-015

Business Statistics

Page 14

Where Dn = nth decile, cf = cumulative frequency of the class preceding nth decile class, f = frequency of the nth decile class, i = class interval of the nth decile class.

Percentiles
The values that divide the given data into hundred equal parts are known as percentiles. There will therefore be ninety nine deciles denoted by P1, P2, P3 P99. The fiftieth percentile (P50) is equal to the median.

Where Pn = nth percentile, cf = cumulative frequency of the class preceding nth percentile class, f = frequency of the nth percentile class, i = class interval of the nth percentile class. Thus M = Q2 = D5 = P50 Q1 = P25 Q3 = P75

Ogive (Locating Median Graphically)


Median, quartiles, deciles, percentiles, etc., can be obtained graphically in any one of the following two ways: First Method: (i) Draw an ogive (cumulative frequency curve) by less than method.

(ii) Plot the values of the variable on the x-axis and the cumulated values (less than) on the y-axis. (iii) Determine the position of the median value by the formula .

MBA-015

Business Statistics

Page 15

(iv) Locate this value on y-axis and from this draw a line parallel to the base cutting the curve at a point. From this point of intersection draw a line parallel to the y-axis and the point where it cuts the x-axis or the base line is the median value. (v) The values of quartiles, deciles and percentiles, etc., can also be located in the same manner. Second Method: (i) Draw two ogives (one less than and the other more than) from the data given. (ii) From the point of intersection of these ogives draw a line parallel to y-axis. It would be a perpendicular on x-axis. The point where this perpendicular touches the x-axis or the base line gives the value of the median.

MODE
Mode is the most typical or fashionable value of the series. It is the most common item of a series. It is generally the value which occurs the largest number of times in a series, that is, the value with the highest frequency. The mode of a distribution is the value at the point around which the items tend to be most heavily concentrated. At times, therefore, mode may not necessarily be the value which occurs the largest number of times in a series, as, in some cases, the point of maximum concentration may be around some other value. In some cases there may be more than one points of concentration of values and the series may be bi-modal or multi-modal. Calculation of Mode in a Series of Individual Observations: A series of individual observations has to be converted into a discrete or continuous series for computation of mode. Calculation of Mode in a Discrete Series: (i) Visual Inspection: The value with highest frequency is the mode. (ii) Grouping Method: This method is adopted when a series is either bi-modal or multi-modal, or if high frequencies are concentrated around more than one item. This method helps in identifying the point of maximum concentration. a. Arrange the values in ascending (or descending) order. b. Add frequencies in twos. i. Add frequencies of item numbers 1 and 2, 3 and 4, . ii. Add frequencies of item numbers 2 and 3, 4 and 5, . c. Add frequencies in threes. i. Add frequencies of item numbers 1, 2 and 3, 4, 5 and 6, . ii. Add frequencies of item numbers 2, 3 and 4, 5, 6 and 7, .
MBA-015 Business Statistics Page 16

iii. Add frequencies of item numbers 3, 4 and 5, 6, 7 and 8, . d. If necessary, add frequencies in fours and fives also. e. The size of items containing the maximum frequencies in each case are noted down. f. The item that has the maximum frequency the largest number of times is called the mode.

N.B.: In case of continuous series, grouping helps in identifying the modal class. Calculation of Mode in a Continuous Series:

Where Z = mode, L1 = lower limit of modal class, L2 = upper limit of modal class, f0 = frequency of the class preceding modal class, f1 = frequency of the modal class, f2 = frequency of the class succeeding modal class, i = class interval.

Histograms (Locating Mode Graphically)


The value of mode can be graphically determined in a frequency distribution. For this, the series has to be first represented by a histogram. After this, two lines are drawn diagonally inside the modal class in such a manner that they touch the upper corner of the modal bar and the upper corner of the adjacent bar. Then a perpendicular is drawn from the intersection of these lines to the x-axis. The point at which the perpendicular touches the x-axis gives the modal value. Merits of Mode (i) It is a very simple measure of central tendency. Often it can be calculated by mere visual inspection of the data. (ii) It is commonly understood. (iii) The values of mean and median can sometimes be not even be present in the series. But mode will always be present in the series.
MBA-015 Business Statistics Page 17

(iv) Mode is normally not affected by extreme values. (v) For determination of mode, it is not even necessary to know the values of all the items of the series. Drawbacks of Mode (i) It is ill-defined, indeterminate, indefinite.

(ii) It is not based on all the observations of the series. (iii) It is not capable of further mathematical treatment. (iv) It may be unrepresentative in many cases. (v) In many cases, it may be impossible to get a definite value of mode.

EMPIRICAL RELATIONSHIP BETWEEN MEAN, MEDIAN AND MODE


In a symmetrical distribution, mean, median and mode are identical, that is, they have the same value. However, in actual life, most distributions are not symmetrical they are skewed. In case of moderately skewed (or moderately asymmetrical) distributions, the value of the mean, median and mode have the following empirical relationship (given by Karl Pearson):

Z = 3M2X Thus XM = (XZ) We can see that the difference between mean and mode is three times the difference between mean and median. In other words, in moderately skewed distributions, median is closer to mean than mode. If we plot any frequency distribution in the form of a curve, mode (Z) touches the peak of the curve indicating maximum frequency, median (M) divides the area of the curve into two equal halves, and mean (x) is the centre of gravity. This relationship between the values of mean, median and mode is of great value. In cases where mode is ill-defined, or the series is bi-modal and mode cannot be calculated through the formula, this relationship can give us an empirical value of the mode.

IMPLICATIONS OF MEAN, MEDIAN AND MODE


Mean. Arithmetic average or mean of a series is the figure obtained by dividing the total values of the various items by their number.

MBA-015

Business Statistics

Page 18

Median. As the name itself suggests, median is the value of the middle item of a series. It divides the series into two equal groups, one containing values less than the median and the other part containing values greater than it. It is a positional average. Mode. It is the most fashionable value, or, more accurately, the value around which there is maximum concentration of items. Choice of an Average The choice of an average should be made very cautiously, as, if a wrong average has been chosen, inaccurate conclusions are likely to follow. Selection of a particular average should be done after giving consideration to a number of points noted below: (i) The object or purpose for which the average is being calculated.

(ii) Whether the average would be used for further statistical computations. (iii) The nature of data available. a. If the distribution is more or less symmetrical, any one of the three measures, that is, mean median or mode may be used. Their values would not differ much. b. If the data is very asymmetrical, median or mode should be used instead of mean. c. If the distribution has unequal class intervals, median should be used as it considers only the class interval of the median class. Mean can also be used in such cases. If it is necessary to calculate mode in such a situation, the series should be regrouped to get classes with equal intervals. d. In case of any open ended classes, mean cannot be calculated, therefore median or mode should be used. At times, it may be worthwhile to calculate more than one average and draw conclusions based on their characteristics. To sum up, it can be said that: Arithmetic Mean should be used when (a) The distribution is not very skewed; (b) The distribution does not have open-ended classes; (c) When the distribution does not have very large and very small items. Median should be used when we have open ended distribution and, more particularly, when plotted as a curve, the distribution is found to be of J-shape or reverse J-shape. In case of price or income distribution, median is the best average.

MBA-015

Business Statistics

Page 19

Mode should be used when we are dealing with qualitative data, and where we have to find the preferences of people. Consumer preferences are best studied with the help of mode. In some cases of discrete series like the average number of rooms in households or average size of shoe, mode is the only appropriate average. Limitations of Averages (1) Even the most judiciously chosen average will have its limitations. An average is a single figure representing a series, and no single figure can condense in itself all the properties of the items it represents. For instance, if the average height of a group of women is less than the average height of a group of men, it does not mean that no woman is taller than any man. (2) An average may have a value which is not even present in the series. (3) As a corollary to the above point, mean may sometimes give results that appear absurd. For instance, we may find that the average number of children in the families of a certain region is 3.4! (4) Averages do not throw light on the formation of the series. The mean of 2, 3 and 55 is 20, and the mean of 19, 20 and 21 is also 20. However, if wrong conclusions are drawn by the use of averages, it is not the fault of the averages. The fault lies with the person drawing the conclusions. The inherent limitations of averages should always be kept in mind and they should not be expected to reveal more than what they can.

MBA-015

Business Statistics

Page 20

DISPERSION
We have seen that the measures of central tendency have their own limitations. Often, they fail to reveal the entire story of a phenomenon. There may be a dozen series whose averages may be identical but which may differ from each other in several ways. Obviously, in such cases further statistical analysis of the data is necessary so that these differences between various series may also be studied. Consider the following three series: Series A Series B Series C 40 40 40 40 40 40 40 40 40 Total Mean 360 40 36 37 38 39 40 41 42 43 44 360 40 1 9 20 30 40 50 60 70 80 360 40

In the first series the mean is 40 and the values of all the items are identical. The items are not at all scattered and the mean fully discloses the characteristics of this distribution. In the second case, though the mean is again 40, yet all the items have different values. But the items are not very scattered. In this case also the mean value is a good representative of the series. However, in the third series, although the mean is still 40, the values are very widely scattered. The mean is 40 times the smallest value and half of the maximum value. It therefore does not satisfactorily represent the individual items in this group. More importantly, no differentiation can be made between the above three series on the basis of mean alone. In order to have a deeper analysis of these three series, it is essential that we study something more than their central tendencies. We find that the scatter in the first case is nil, in the second case it is small, while in the third series the items are very scattered. It is evident from this example that a study of the extent of the scatter around an average should also be studied to throw more light on the composition of a series. The name given to this scatter is dispersion.

MBA-015

Business Statistics

Page 21

Dispersion Defined Dispersion or spread is the degree of the scatter or variation of the variable about a central value. Measures of variability are usually used to indicate how tightly bunched the sample values are around the mean (or any other measure of central tendency). Since for a precise study of dispersion we have to average deviations of the values of the various items, from their average, the various measures of dispersion are called averages of the second order. Dispersion or variation can be expressed either in terms of the original unit of a series or as an abstract figure like a ration or percentage. These are known as absolute and relative dispersion respectively. In a comparison of variability of two or more series, relative dispersion should be taken into account. Objects of Measuring Dispersion (1) To judge the reliability of measures of central tendency. (2) To make a comparative study of the variability of two series. (3) To identify the causes of variability with a view to control it. (4) To serve as a basis for further statistical analysis. Measures of Dispersion (1) Range; (2) Inter Quartile Range; (3) Semi-Inter Quartile Range (Quartile Deviation); (4) Average Deviation (Mean Deviation); (5) Standard Deviation or Root Mean Square (rms) Deviation; (6) Lorenz Curve. Of the above, the first three are positional measures. Fourth and fifth are algebraic measures or calculation measures. The last one is a graphic method. We shall limit our discussion to range, mean deviation and standard deviation. Properties of a Good Measure of Dispersion These are the same as the properties of a good measure of central tendency. These are: (i) A good measure of dispersion should be rigidly defined.

(ii) It should be based on all the observations of the series.


MBA-015 Business Statistics Page 22

(iii) It should be capable of further algebraic treatment. (iv) It should be simple to calculate and easy to follow. (v) It should not be affected by fluctuations of sampling.

Range
Range is the simplest possible measure of dispersion. It is the difference between the values of the extreme items of a series. R = LS Where L = value of the greatest item of the series, S = value of the lowest item of the series.

In case of continuous series range is calculated by any one of the following two methods: (i) By subtracting the lower limit of the lowest class from the upper limit of the highest class. (ii) By subtracting the mid-value of the lowest class from mid-value of the highest class. Merits (i) It is rigidly defined. (ii) It is easy to calculate and simple to follow. Demerits (i) It is greatly affected by the fluctuations of sampling.

(ii) It is not based on all the observations of the series. (iii) It cannot be used in case of open-ended distributions. Uses of Range (i) Quality control.

(ii) Variation in money rates, share values, exchange rates and gold prices, etc. (iii) Weather forecasting.

MBA-015

Business Statistics

Page 23

MEAN DEVIATION
The range suffers from a major defect; that it is calculated by taking into account only two values of the series. This method of studying dispersion (by location of limits) is also called the method of limits. Because of the demerits of range, we shall now consider another two more measures of dispersion that take into account all the observations of a series, and is calculated in relation to a central value. This method of calculating dispersion is called the method of averaging deviations. The first among these two measures is mean deviation. Mean deviation of a series is the arithmetic average of the deviations of various items from a measure of central tendency (mean, median or mode). Theoretically, deviations can be taken from any of the three averages mentioned above, but in actual practice mean deviation is calculated either from mean or from median. Mode is usually not considered, as its value is very often not well-defined. Between mean and median, the latter is supposed to be better than the former because the sum of deviations from median is less than the sum of the deviations from mean. Therefore the value of mean deviation from median is always less than the value calculated from mean. Since the purpose of a measure of dispersion is to study the variation of items from a central value, while aggregating deviations, the algebraic signs are not taken into account. This step of ignoring the signs renders mean deviation incapable of further algebraic treatment. Mean deviation is also known as the first moment of dispersion. Mean Deviation from Mean

Mean Deviation from Median

Mean Deviation from Mode

MBA-015

Business Statistics

Page 24

Mean Coefficient of Dispersion From Mean

From Median

From Mode

Calculation of Mean Deviation in a Series of Individual Observations Direct Method: The above-discussed formulae give the direct method of calculation of mean deviation (in a series of individual observations. Short-cut Method: Median (or mean, or mode) is calculated and the total of the values of the items below the median (or mean, or mode) and above it are found out. The former is subtracted from the latter and divided by the number of items. The resulting figure is the mean deviation from median (or mean, or mode). Symbolically:

Calculation of Mean Deviation in a Discrete Series Direct Method (i) Calculate median (or mean, or mode).

(ii) Find the deviations from the median (or mean, or mode), ignoring sign. (iii) Multiply these deviations with the respective frequencies and total these products. (iv) Divide the total (fdM or fdx or fdZ) by the number of observations to get mean deviation.

MBA-015

Business Statistics

Page 25

Symbolically:

Short-cut Method 1. Deviations are taken from an assumed median (or mean, or mode), multiplied by their respective frequencies, and the products are totaled. 2. Number of items less than the actual median (or actual mean, or actual mode) are multiplied by the difference between the actual and assumed mean. 3. Similarly, the number of items greater than the actual median (or actual mean, or actual mode) are multiplied by the difference between the actual and assumed mean. 4. The latter (No. 3) is deducted from the former (No. 2) and the balance is added to the sum of the products of deviations from the assumed median (or mean, or mode) and their frequencies (No. 1). 5. To resulting figure when divided by the number of items gives the value of mean deviation. Calculation of Mean Deviation in a Continuous Series Direct Method Same procedure as in case of discrete series. Short-cut Method A. If the actual and assumed median (or mean, or mode) lie in the same class, the procedure is the same as in case of discrete series. B. If the actual and assumed median (or mean, or mode) lie in different classes, following adjustments are required: 1. In such a situation, the frequency of the class in which the actual median (or actual mean, or actual mode) lies is treated separately. It is multiplied by the difference of the deviation of the mid-value from the actual and assumed medians (or mean, or mode). 2. This product is subtracted from the total deviations from the assumed median (or mean, or mode). 3. The rest of the procedure is the same as in the case of discrete series.

MBA-015

Business Statistics

Page 26

Merits of Mean Deviation 1. It is rigidly defined. 2. Its calculation is not very difficult. 3. It is readily understood. 4. It is based on all observations 5. It is not affected much by the values of the extreme items. Demerits of Mean Deviation 1. It is not capable of further mathematical treatment. 2. At times, mean deviation may not be a very accurate measure of dispersion. a. If it is calculated from mode, mode may be ill-defined. b. If it is calculated from median, it may not be very reliable if the degree of variability is high in the series. c. If it is calculated from mean, it is not very scientific because the sum of deviations from mean (ignoring the signs) is greater than the sum of deviations from median. Utility of Mean Deviation 1. Finds favour with economists and businessmen due to simplicity in calculation. 2. Has an advantage over standard deviations which gives greater importance to the deviations of extreme values. 3. This method has been found to be more useful than others in forecasting business cycles. 4. It is useful for small sample studies where elaborate statistical analysis is not needed.

STANDARD DEVIATION
The concept of standard deviation was first used by Karl Pearson in 1893. It is the most commonly used measure of dispersion. It removes the drawback in the technique of the calculation of mean deviation, where algebraic signs are ignored. One of the easiest ways of doing away with algebraic signs is to square the figures, and this process is adopted in the calculation of standard deviation. Standard deviation is the square root of the arithmetic average of the squares of the deviations measured from the mean. It is therefore also known as the root-mean-square (rms) deviation from mean. Some other terms like mean square error and error of mean square are also used to denote standard deviation.

MBA-015

Business Statistics

Page 27

Thus, in the calculation of standard deviation, first mean is calculated, and the deviations of various items from this mean value are squared. The squared items are totaled and the sum is divided by the number of items. The square root of the resulting figure is the standard deviation of the series. Symbolically,

Difference between Mean Deviation and Standard Deviation (i) Mean deviation does not take into account the algebraic signs. Standard deviation does not ignore them but nullifies their effect by squaring the deviations. (ii) Mean deviation can be calculated from mean, median or mode. Standard deviation is calculated from mean only. (iii) Standard deviation possesses many more mathematical properties than mean deviation. Calculation of Standard Deviation in a Series of Individual Observations Direct Method (No. 1) The formula discussed above. Direct Method (No. 2)

Shortcut Method

MBA-015

Business Statistics

Page 28

Standard Coefficient of Dispersion/Coefficient of Standard Deviation Standard deviation is an absolute measure of dispersion. For the purpose of comparison, a relative measure of dispersion is calculated by dividing the standard deviation by mean. Thus, the standard coefficient of dispersion = Calculation of Standard Deviation in a Discrete Series Direct Method

Shortcut Method

Step Deviation Method

Calculation of Standard Deviation in a Continuous Series

Mathematical Properties of Standard Deviation (1) The standard deviation of a series can be found out from the standard deviations of its component parts and their means.

(2) The standard deviation of first N natural numbers is

(3) The sum of the squares of deviations taken from the mean is minimum. Thus, standard deviation is calculated on the basis of minimum deviations from the measure of central tendency. (4) In a normal distribution the following area relationship holds good:
MBA-015 Business Statistics Page 29

Mean 1 covers 68.27% of the items Mean 2 covers 95.45% of the items Mean 3 covers 99.73% of the items Further, it should be noted that in symmetrical or moderately asymmetrical series a range which is 6 times of the standard deviation usually covers at least 99% of the observations. This property gives a concrete meaning of standard deviation. (In a normal distribution, mean deviation is 0.7979 of standard deviation. Further, arithmetic mean mean deviation would cover 57.51% of the items.) Coefficient of Standard Deviation A relative measure of dispersion, coefficient of standard deviation =

Coefficient of Variation

Variance Variance = 2 Thus, = Merits, Demerits and Uses of Standard Deviation Merits: (1) It is rigidly defined. (2) It is based on all the observations. (3) It is amenable to further algebraic treatment and possesses many mathematical properties. (4) It is less affected by fluctuations of sampling than most other measures of dispersion. (5) The squaring of deviations makes them positive and the difficulty about algebraic signs which was experienced in case of mean deviation is removed. Demerits: (1) Standard deviation is not easy to calculate, nor is it easily understood. (2) It gives more weight to extreme values, because the squares of the deviations that are big in size, would be proportionately greater than the squares of those deviations that are comparatively small.

MBA-015

Business Statistics

Page 30

Uses: Despite the drawbacks mentioned above, the standard deviation is the best measure of dispersion and should be used wherever possible. Just as the mean is the best measure of central tendency (leaving some exceptional cases), standard deviation is the best measure of dispersion. However, since standard deviation gives greater weight to extreme items, it does not find much favour with economists and businessmen who are more interested in the results of the modal class.

SKEWNESS
So far, we have discussed the methods of measuring the central tendency of a frequency distribution and the methods of studying the concentration of items round the central value. However, these measures do not reveal whether the dispersal of values on either side of an average is symmetrical or not. If observations are arranged in a symmetrical order round a measure of central tendency, we get a symmetrical distribution. When plotted, such a distribution gives a normal or ideal curve. In a perfectly symmetrical distribution, the values of mean, median and mode coincide and the quartiles are equidistant from the median. A normal curve is a bell-shaped frequency curve in which the values on either side of a measure of central tendency are symmetrical. In order to study a frequency distribution, it would be of great use to know whether it would give a normal curve, and if not, to what extent it would deviate from a normal distribution. In fact measures of central tendency and measures of dispersion should always be supplemented by the measures of the third order, for instance the measure of skewness. Skewness is the opposite of symmetry and its presence tells us that a particular distribution is not symmetrical, or, in other words, it is skew. Definition A distribution is said to be skewed when the mean and the median fall at different points in the distribution, and the balance of the curve (or its centre of gravity) is shifted to one side or to the other (left or right). Symmetrical Distribution: It is bell-shaped and in it, there is no skewness. The value of mean, median and mode is identical. Positively Skewed Distribution: It is skewed to the right. In it the value of mean is more than the value of median and median is greater than mode.

MBA-015

Business Statistics

Page 31

Negatively Skewed Distribution: It is skewed to the left. In it the value of mode is more than the values of median and median is greater than mean. Thus: (i) In a symmetrical distribution, x = M = Z

(ii) In a positively skewed distribution, x > M > Z (iii) In a negatively skew distribution, x < M < Z Difference between Measures of Dispersion and Skewness (1) Dispersion deals with the spread of values around central value. Skewness on the other hand deals with the symmetry of distribution around a central value. (2) Dispersion deals with the amount of variation and skewness deals with the direction of variation. (3) Dispersion helps in finding out the extent to which a central value is a representative of the whole distribution. Skewness deals with the nature of variations on either side of a central value. Test of Skewness (a) In a skew distribution values of mean, median and mode would not coincide. The mean and mode would be pulled wide apart and median would usually lie between them. (b) In a skew distribution the two quartiles would not be equidistant from the median. (c) A skew distribution would not give a bell-shaped curve. (d) The sum of positive deviations from median is not equal to the sum of negative deviations. (e) Frequencies are not equally distributed at various points which are equidistant from mode on either side. Measures of Skewness First Measures of Skewness: (i) Mean Mode (x - Z)

(ii) Mean Median (x - M) (iii) Median Mode (M - Z) These measures suffer from some drawbacks. These are: (i) These measures are expressed in the units of a distribution and therefore cannot be compared with measures expressed in another distribution with a different unit.
MBA-015 Business Statistics Page 32

(ii) There may be considerable variation in different distributions. In one distribution the difference between mean and median may be very large as compared to similar difference in another distribution and yet the two distribution may give similar curves. For example, the difference between mean and median in case of weights may be much larger than a similar difference in heights yet the two curves representing weight and height may look similar because relative variations are not studied by these absolute measures. Relative Measures of Skewness Relative measures of skewness are known as coefficients of skewness. Coefficient of skewness

However, these measures of skewness are not as popular as those given by Karl Pearson, Bowley, Kelly. Thus the important measures of relative skewness are: (1) Karl Pearsons Coefficient of Skewness:

If mode is ill-defined:

(2) Bowleys Coefficient of Skewness:

(3) Kellys Coefficient of Skewness:

Pearsons formula is the most popular.


MBA-015 Business Statistics Page 33

MOMENTS
Moment is a familiar mechanical term for the measure of a force with reference to its tendency to produce rotation. The strength of this tendency depends on the amount of the force and the distance from the origin of the point at which the force is exerted.

This is the same formula as used for calculation of mean. It is for this reason that the arithmetic average is also called the first moment about origin. If the origin is not zero but the mean value then the first moment about mean would be indicates the difference between the values of a variable (X) and the mean (X). In case of grouped data, the first moment about mean would be , where d

Moments about Mean Series of Individual Observations

Grouped Data

MBA-015

Business Statistics

Page 34

Moments can be extended to 5 or 6 or more, but in actual practice the first four moments are enough to describe a series. Objective Moments are calculated to study the nature of a distribution. They tell us whether a distribution is symmetrical or not. They also tell us about the nature of symmetry whether the symmetrical curve is (i) Normal

(ii) More flat than a normal curve (iii) More peaked than a normal curve Thus we can study the skewness of a distribution with the help of moments. To study the nature of a distribution, two constants are calculated from the moments. They are

1 tells us whether a distribution is skewed or whether it is symmetrical. 2 tells us the difference between a symmetrical curve and a normal curve. This is known as Kurtosis.

KURTOSIS
Kurtosis is another measure which tells us about the form of a distribution. It tells us whether the distribution, if plotted on a graph, would give us a normal curve, a curve more flat than the normal curve or a curve more peaked than the normal curve. The word kurtosis in Greek language means bulginess Definition The degree of kurtosis of a distribution is measured relative to the peakedness of a normal curve. If a distribution is more peaked than the normal distribution, it is called leptokurtic. If the distribution is more flat than the normal distribution, it is called platykurtic. The normal distribution is known as mesokurtic. (Platykurtic curves are like platypus, squat with short tails, leptokurtic curves are like kangaroos, high with long tails, noted for leaping.) Measures of Kurtosis Kurtosis is measured by coefficient 2 or its derivation 2.
MBA-015 Business Statistics Page 35

The standard value of 2 is taken as 3 and curves with 2 less than 3 are called platykurtic while curves with 2 more than 3 are called leptokurtic. In a normal or mesokurtic curve, 2 is equal to 3. Therefore, for a normal curve 2 is 0, in a platykurtic curve, 2 is less than 0 and in a leptokurtic curve, 2 is greater than 0. Coefficients of Skewness based on Moments

If there is no skewness in a distribution the value of 1 = 0. If the value of 1 is more than zero it is an indication of skewness and the greater the value of 1, greater is the extent of skewness. However, a drawback of 1 is that it cannot assume a negative value (32 will be positive, as it is a square; 2 is also positive as it is square of deviations). This drawback is removed by 1.

MBA-015

Business Statistics

Page 36

UNIT II TIME SERIES ANALYSIS


Introduction and Application Time series refers to such a series in which one variable is time. In other words, if we have a chronologically arranged values of a variable over successive time periods, it would be a time series. Examples of time series would be the figures of national income, or production or costs or prices spread over a period of time. The analysis of such figures is called the Analysis of Time Series. Time series analysis is done primarily for the purpose of making forecasts for the future and also for the purpose of evaluating past performances. An economist or a businessman is very naturally interested in estimating the future figures of national income, population, prices and wages, etc. In fact, the success or failure of an economist or a businessman depends to a large extent on the accuracy of his future forecasts. Forecasting is done by analyzing the past behavior of the variable under study. Hence the analysis of time series assumes great importance in the study of all economic problems. However, it should not be concluded from the above discussion that analysis of time series is useful only to economists and businessmen. Analysis of past data is done in a variety of other fields also. A sociologist may find the analysis of past data relating to crimes very useful for studying the crime situation in a country and suggesting appropriate measures for controlling it. Similarly, in a study of weather conditions, or in studying the incidence of cancer or in studying the body temperature of a patient, an analysis of past data is very useful. Thus we can conclude by saying that the analysis of time series is helpful in studying any phenomenon whose values are or can be chronologically arranged over successive time periods. Definition A time series may be defined as a collection of readings belonging to different time periods, of some economic variable or composite of variables such as production of steel, per capita income, gross national product, price of tobacco or index of industrial production. Utility of Time Series Analysis (i) (ii) (iii) (iv) It helps in the analysis of past behavior of a variable. It helps in forecasting. It helps in evaluation of current achievement. It helps in making comparative studies.

MBA-015

Business Statistics

Page 37

Components of a Time Series 1. Secular Trend or Long Term Movements (T) The general tendency of the time series data to increase or to decrease or to remain segregated during a long period of time is called secular trend. (i) The secular trend is generally either upward or downward.

(ii) General trend is the result of such forces which are more or less constant for a long time or which change very gradually. (iii) The term long period of time is a relative term. (iv) The secular trend may be linear or non-linear. (v) It is not necessary that a time series should always have a rising or falling tendency. 2. Seasonal Variations (S) Seasonal variations refer to such movements in a time series which are due to forces which are rhythmic in nature and which repeat themselves periodically every season. The seasonal variations may be attributed to the following causes: (i) Causes resulting from natural forces. (ii) Causes resulting from local customs and traditions. 3. Cyclical Variations (C) The cyclical variations in a series are the recurrent variations whose duration is more than one year. The difficulty associated with business cycles is that their period is not uniform and further, cyclical variations are mixed up with erratic and irregular variations which are difficult to be isolated. 4. Irregular Variations (I) Irregular variations are the effect of random factors. These are generally mixed up with seasonal and cyclical variations and are caused by irregular and accidental factors like floods, famines, wars, strikes, lockouts, etc. Time Series Decomposition Models The analysis of time series consists of two major steps: 1. Identifying the various factors or influences which produce the variations in the time series, and 2. Isolating, analyzing and measuring the effects of these factors independently, by holding other things constant.

MBA-015

Business Statistics

Page 38

The purpose of decomposition models is to break a time series into its components: Trend (T), Cyclical (C), Seasonality (S), and Irregularity (I). Decomposition of time series aims to isolate influence of each of the four components on the actual series so as to provide a basis for forecasting. There are many models by which a time series can be analysed; two models commonly used for decomposition of a time series are: Multiplicative Model. The actual values of a time series, represented by Y can be found by multiplying four components at a particular time period. The effect of four components on the time series is interdependent. The multiplicative time series model is defined as: Y=TxCxSxI The multiplicative model is appropriate in situations where the effect of C, S and I is measured in relative sense and not in absolute sense. The geometric mean of C, S and I is assumed to be less than one. Additive Model. In this model, it is assumed that the effect of various components can be estimated by adding the various components of a time series. It is stated as: Y=T+C+S+I Here C, S and I are absolute quantities and can have positive or negative values. It is assumed that these four components are independent of each other. However, in real-life time series data this assumption does not hold good. Measurement of Trend (various methods) The various methods by which trend values can be determined are: 1. Freehand or Graphic Method. This is the simplest of all the methods of finding out the trend values. First of all the values of a time series are plotted on a graph paper in the form of a histogram. After this a freehand smoothed curve is drawn through these points in such a way that the curve represents the general tendency of the data. The trend values can then be read for various time periods by locating them on the trend line against each time period. 2. Method of Semi-Averages. In this method, as the name suggests, semi-averages are calculated to find out the trend values. By semi-averages is meant the averages of the two halves of a series. Thus, if we have a time series running from 1994 to 2003, it would be divided into two halves one half from 1994 to 1998 and the other half from 1999 to 2003. The averages of the values of these
MBA-015 Business Statistics Page 39

two halves would be found out. These averages would be plotted against the mid-value of each half. The two points are then joined by a straight line which can be extended on either side. This is the trend line. 3. Method of Moving Averages. Moving average method is a simple device of reducing fluctuations and obtaining trend values with a fair degree of accuracy. In this method the average value of a number of years is taken as the trend value for the middle point of the period of moving average. The process of averaging smoothes the curve and reduces the fluctuations. The first thing to be decided in this method is the period of the moving average that is, to decide about the number of consecutive items whose average would be calculated each time. Suppose it has been decided that the period of the moving average would be 5 years, then the arithmetic average of the first 5 items would be placed against the item no. 3 and then the average of item numbers 2, 3, 4, 5 and 6 would be placed against item no. 4. This process would be repeated till the average of the last five items has been calculated. 4. Method of Least Squares (Fitting Straight Line Trend). One of the best ways of obtaining trend values is the method of least squares. With this method, a straight line trend is obtained, this line is called the Line of the Best Fit. It is a line from which the sum of the deviations of various points on either side is equal to zero. Also, the sum of the squares of these deviations would be the least as compared to the sums of squares of the deviations obtained by using other lines. That is why this method is known as the method of least squares. (i) Linear Trend

(ii) Non-Linear Trend

MBA-015

Business Statistics

Page 40

If X = 0, then

MBA-015

Business Statistics

Page 41

CONSTRUCTION OF INDEX NUMBERS AND THEIR USES


Introduction Index numbers are devices that measure the change in the level of a phenomenon with respect to time, geographical location or some other characteristic. Initially, index numbers were designed to study the change in the price level or the purchasing power of money. Today there is hardly any phenomena which does not make use of the device of index numbers. Wherever a comparative study has to be made, index numbers are very handy. In present day situation changes in production, consumption, exports, imports, national income, cost of living, incidence of crime, number of road accidents, business failures and a very wide variety of other phenomena are studied with the help of index numbers. Index numbers are supposed to be barometers which measure the change in the level of a phenomenon. Their importance in the realm of quantitative techniques cannot be over emphasized. Definition An index number is a statistical measure designed to show changes in variable or a group of related variables with respect to time, geographical location or other characteristic. Characteristics of Index Numbers 1. Index numbers are a specialized type of average. 2. Index numbers study the effects of such factors which cannot be measured directly. 3. Index numbers bring out the common characteristics of a group of items. 4. The changes measured by index numbers can be either in relation to time or in relation to place. Uses of Index Numbers 1. Help in studying trends. 2. Help in policy formulation. 3. Help in measuring the purchasing power of money. 4. Help in deflating various values. 5. Act as economic barometers. Types of Index Numbers (a) Price Index Numbers (Wholesale and Retail). (b) Quantity Index Numbers. (c) Value Index Numbers.

MBA-015

Business Statistics

Page 42

(d) Special Purpose Index Numbers. Selection of the Base Index numbers measure the relative changes in the level of a phenomenon as compared to the level of the same phenomenon on a previous date. This previous date or the period on which the current variations are based is known as the Base Period of Index Number. In the construction of index numbers the selection of the base period is a very important step. There are two methods by which the base period can be selected: (i) Fixed Base Method As the name suggests, in this method the base period is fixed. A particular year is generally chosen arbitrarily and the prices of the subsequent years are expressed as relatives of the prices of the base year. Sometimes instead of choosing a single year as the base, a period of a few years is chosen and the average price of this period is taken as the base year price. Fixed base can be used for an indefinite period. The year which is selected as a base should be a normal year. Or, in other words, the price level in this year should neither be abnormally low nor abnormally high. If an abnormal year is chosen as the base, the price relatives of the current year calculated on its basis would give misleading conclusions. In order to remove this difficulty associated with the selection of a normal year, the average price of a few years is sometimes taken as the base price, as stated earlier. The idea is that if a few years average is taken, abnormalities in one direction would be set off against abnormalities in another direction. Sometimes when an index number is being constructed for a past period, the average of the whole period may be taken as the base. But this is possible only if the index number relates to the past. (ii) Chain Base Method In this method, there is no fixed base period. The year immediately preceding the one for which price relatives have to be calculated is assumed as the base year. In this way there is no fixed base. It goes on changing. The chief advantage of this method is that the price relatives of a year can be compared with the price levels of the immediately preceding year. Businessmen and others are more interested in comparison of this type rather than in comparisons relating to distant past. Yet another advantage of the chain base method is that under it, it is possible to include new items in an index number or to delete old items which are no more important. In fixed base method, this is not possible. But chain base method has a drawback and it is that with it, comparisons cannot be made over a long period.

MBA-015

Business Statistics

Page 43

Methods of Constructing Index Numbers A. Un-weighted Index Numbers. (each item is supposed to have the same weight.) B. Weighted Index Numbers. (weights are assigned to various items in accordance with their importance.) Un-weighted index numbers can be further divided in two categories: (i) Simple Aggregative Method.

(ii) Simple Average of Relatives Method.

Similarly, weighted index numbers can also be divided in two categories: (i) Weighted Aggregative Method. 1. Laspeyres Method.

2. Paasches Method.

3. Drobish and Bowleys Method.

4. Fishers Ideal Method.

5. Marshall-Edgeworth Method.

MBA-015

Business Statistics

Page 44

6. Walsh Method.

7. Kellys Method.

(ii) Weighted Average of Relatives Method. (beyond the scope of our discussion.) Quantity or Volume Index Numbers The quantity index numbers measure average change in quantities and enable us to compare changes in physical quantity of goods produced or sold. Laspeyres Method

Paasches Formula

Fishers Formula

Value Index Numbers Value is the product of price and quantity. A simple value ration is equal to the value of the current year divided by the value of the base year. If this ration is multiplied by 100, we get the value index number.

MBA-015

Business Statistics

Page 45

Tests of Adequacy of Index Number Formulae 1. Unit Test This test requires that the formula for the construction of index number should be such which is not affected by the unit in which the prices or quantities have been quoted. This test is satisfied by all the formulae discussed above except the simple (un-weighted) aggregative index formula. 2. Time Reversal Test If the formula for calculating an index number is such that it will give the same ratio between one point of comparison and the other, no matter which of the two is taken as base, it satisfies the time reversal test. That is P01 x P10 = 1 This test is not satisfied by Laspeyres and Paasches formulae. It is satisfied by: (1) Fishers Ideal Formula (2) Kellys Formula (3) Marshall-Edgeworth Method (4) Walsch Formula 3. Factor Reversal Test A formula that permits inter-changing the price and quantities without giving inconsistent result, that is, the two results multiplied together should give the true value ratio. That is

This test is satisfied only by the Fisher's Ideal Index Number. 4. Circular Test It is a sort of extension of the time reversal test. Suppose an index number is constructed for the year 2003 with the base of 2002 and another index number for 2002 on the base of 2001, then it should be possible for us to directly get an index number for 2003 on the base of 2001. If this is possible, then the circular test is satisfied. If P01 represents the price change of the current year on the base year and P12 the price change of the base year on some other base, and P20 the price change of the current year on this second base, then the following equation should be satisfied: P01 x P12 x P20 = 1 This test is fulfilled only by un-weighted index numbers.
MBA-015 Business Statistics Page 46

Cost of Living Index Numbers These index numbers are also called consumer price index numbers, retail price index numbers, cost of living price index numbers and price of living index numbers. Cost of living index numbers do not actually measure the actual cost of living nor the fluctuations in the cost of living due to causes other than the changes in the price level. Cost of living index numbers only tell us how much the consumers of a particular class have to pay to get a basket of goods and services at a particular point of time in comparison to what they paid for this basket in the base year. The necessity of the construction of cost of living index numbers arises on account of the fact that wholesale price index numbers measure the variations only in the general level of prices. These variations do not throw light on the effects of rise and fall of prices on the cost of living of different classes of people in a society. Different groups of people consume different types of commodities and even the same type of commodities are not consumed in the same proportion by different classes of people. The relative importance of different commodities is therefore different in case of different types of people. In order to measure the effects of the rise and fall in the prices of various commodities on the cost of living of different classes of people, separate index numbers are constructed for different groups. It should be remembered that a cost of living index number tells us about the variations in the cost of living of only one group of persons living in a particular region. By region, we mean an area within which retail process are almost equal and by groups we mean classes distinguished from each other on the basis of incomes. Thus, there cannot be one cost of living index number for textile workers of the whole country, because retail process in different places differ and the pattern of consumption is also not alike in different localities. Similarly, we cannot have a cost of living index number for the whole population of a particular town because groups of persons with different incomes spend their income on various commodities in different ways and the relative importance of various commodities to all persons is not identical. Limitations of Index Numbers (1) They are only approximate indicators of the relative level of a phenomenon. (2) Index numbers use only limited number of items in their calculation. (3) The quality of the product is very difficult to be accounted for in the construction of index numbers. (4) Index numbers are liable to be misused. (5) Indices constructed for one purpose cannot be used for another purpose. Unit III
MBA-015 Business Statistics Page 47

CORRELATION AND REGRESSION Correlation


Introduction In various types of analysis discussed so far, we have confined ourselves to such series where various items assumed different values of one variable. There can be, however, such series also where each item assumes the values of two or more variables. For example, if the heights and weights of a group of persons are measured we shall get such series where each member of the group would assume two values one relating to height and another relating to weight. If, besides heights and weights, the chest measurements were also taken, each member of the group would assume three values relating to three different variables. In such situations, sometimes it appears that the values of the various variables so obtained are inter-related. It is likely that such relationship may be obtained in two series relating to the heights and weights of a group of persons. It may be observed that weights increase with increase in heights so that tall people are heavier than short sized people. Similarly, if data are collected about the prices of a commodity and the quantities sold at different prices, two series would be obtained. In these two series we are again likely to find some relationship. With an increase in the price of the commodity, the quantity sold is bound to decrease. Such relationships can be find in many other types of series also. The term correlation (or co-variation) indicates the relationship between two such variables in which with changes in the value of one variable, the value of the other variable also changes. Definition If two or more quantities vary in sympathy so that movements in the one tend to be accompanied by corresponding movements in the other(s) then they are said to be correlated. Utility (1) The study of correlation reduces the range of uncertainty associated with decision making. (2) Correlation analysis is very helpful in understanding economic behavior. (3) Correlation study helps in identifying such factors which can stabilize a disturbed economic situation. (4) Correlation study helps to estimate the likely change in a variable with a particular amount of change in a related variable. (Here we take help of regression analysis.)

MBA-015

Business Statistics

Page 48

(5) Inter-relationship studies between different variables are very helpful tools in promoting research and opening new frontiers of knowledge. There can be correlation between two variables due to any one or more of the following reasons: (1) There is a cause-effect relationship between the two variables. (2) Both the correlated variables are being affected by a third variable. (3) Related variables might be mutually affecting each other so that neither of them could be designated as a cause or effect. (4) The correlation may be due to random or chance factors. (5) There might be a situation of nonsense or spurious correlation between the two variables under study. Types of Correlation (i) Positive or Negative Correlation

(ii) Simple, Multiple or Partial Correlation In simple correlation we study only two variables say price and demand. In multiple correlation, we study together the relationship between three or more factors like production, rainfall and use of fertilizers. In partial correlation, though more than two factors are involved, but correlation is studied only between two factors and the other factors are assumed to be constant. (iii) Linear or Non-Linear (Curvilinear) Correlation Methods of Studying Correlation (i) Scatter Diagram

(ii) Correlation Graph (iii) Coefficient of Correlation (iv) Coefficient of Correlation by Rank Differences (v) Coefficient of Concurrent Deviation

Karl Pearsons Coefficient of Correlation

MBA-015

Business Statistics

Page 49

When deviations are taken from assumed mean

When deviations are taken from actual mean

In case of Grouped Data

Mathematical Properties of Coefficient of Correlation (1) It lies between -1 and +1. It cannot exceed unity. (2) It is not affected by change of scale or origin.

MBA-015

Business Statistics

Page 50

Coefficient of Correlation by the Method of Least Squares Will be discussed while studying Regression Analysis. Spearmans Rank Correlation Coefficient

In case of equal ranks

Coefficient of Concurrent Deviation

Regression
The dictionary meaning of the word regression is stepping back or going back. Regression implies going back or returning towards the mean. Regression lines study the average relationship between two series, and throw light on their covariance. For example, if the coefficient of correlation between the heights of fathers and sons is +0.7, it means that if a group of fathers have heights which are more than the average by x inches, their sons could have heights which would be more than average by 0.7x inches. Thus, the heights of the sons regress towards the mean. The study of this tendency is the subject matter of regression. Definition Regression analysis attempts to establish the nature of the relationship between variables that is, to study the functional relationship between the variables and thereby provide a mechanism for predicting, or forecasting. Utility The definition makes it clear that regression analysis is done for estimating or predicting the unknown value of one variable from the known value of the other variable. This is a very useful statistical tool which is used both in natural as well as social sciences.

MBA-015

Business Statistics

Page 51

In the field of business this tool of statistical analysis is very widely used. Businessmen are interested in predicting future production, consumption, investment, prices, profits, sales, etc. In fact the success of a businessman depends on the correctness of the various estimates that he is required to make. In sociological studies and in the field of economic planning, projections of population, birth rates, death rates and other similar variables are of great use. With the help of regression analysis we can estimate or predict the effect of one variable on the other. However, in social sciences there often is multiple causation which means that a large number of factors affect various variables. The regression study which confines itself to a study of two variables is called simple regression. The regression analysis which studies more than two variables at a time is called multiple regression. In simple regression analysis there are two variables one of which is known as an independent variable or regresser or predictor or explanator. On the basis of the values of this variable the values of the other variable are predicted. This other variable is called the dependent or regressed or explained variable. With the help of regression studies we can also calculate the value of the coefficient of correlation. The coefficient of determination (square of coefficient of correlation), which measures the effect of the independent variable on the dependent variable gives us an indication about the predictive value of the regression studies. Comparison of Correlation and Regression Studies (1) Correlation studies are meant for studying the co-variation of two variables. Regression analysis tells us about the relative movement in the variables under study and with its help we can also predict the value of one variable by taking into account the value of the other variable. (2) Correlation between two series is not necessarily a cause-effect relationship. Regression, on the other hand, presumes one variable as a cause and the other as its effect. (3) The coefficient of correlation varies between 1. The regression coefficients have the same sign as the correlation coefficient. Further, regression coefficients can have a value higher than unity but the product of the two regression coefficients can never exceed unity because the correlation coefficient is the square root of the product of the two regression coefficients. For regression analysis, a graph is needed to be drawn. For this, first we have to draw a scatter diagram. When all related pairs or values of X and Y have been plotted, we draw two regression lines to predict the values of X and Y variables. The regression line which is used to predict the values of Y for corresponding values of X is called the Regression Line of Y on X. similarly the regression line used to predict a value of X for a value of Y is called the Regression Line of X on Y. If the coefficient of
MBA-015 Business Statistics Page 52

correlation between X and Y is perfect, that is, its value is either +1 or -1, the two regression lines will coincide; there will be only one regression line, as the variations in the two series in such cases always increases or decreases by a constant figure. Regression can be linear or non-linear. We shall limit our discussion to linear regression only. Regression lines can be drawn by: (a) Free Hand Curve Method In the free hand curve method, we first plot the pairs of the values of X and Y in the form of a scatter diagram one point for one pair of values. After this we draw two free hand straight lines. One of these lines is drawn in such a way that the positive deviations of Y-series from its mean are cancelled by the negative deviations. That is, the sum of the deviations on one side of the line shall be equal to the sum of the deviations on the other side. This will be the regression line of Y on X. The other regression line would be drawn in such a way that the positive deviations of X-series from its mean would cancel the negative deviations. This is the regression line of X on Y. The two regression line cut each other at the point, the coordinates of which give the mean values of the two series. However, it is very difficult to draw regression lines by freehand curve method. Usually a piece of thread is repeatedly adjusted in such a manner in the scatter diagram that the positive and negative deviations cancel each other. Once these lines are drawn, we can predict or estimate the values of Y from the regression line of Y on X, and similarly the values of X can be predicted from the regression line of X on Y. (b) Method of Least Squares In order to avoid the difficulties associated with the drawing of regression lines by the free hand curve method, a mathematical relationship is established between the movements of X and Y series and algebraic equations are obtained to represent the relative movements of X and Y series. One such method is the method of least squares. In this method we minimize the sum of squares of the deviations between the given values of a variable and its estimated values given by the line of best fit. Line of regression of Y on X is the line which gives the best estimate for the value of Y for a specified value of X. similarly the line of regression of X on Y is the line which gives the best estimate for the value of X for a specified value of Y. If the values of Y are plotted on the Y-axis (that is, the vertical axis) then the regression line of Y on X will be such that it minimizes the total of the squares of the vertical deviations. Similarly, if
MBA-015 Business Statistics Page 53

the values of X are plotted on the X-axis (the horizontal axis) the regression line of X on Y will be such which minimizes the total of the squares of the horizontal deviations. The line of best fit is obtained by the equation of straight line

And that in the method of least squares this line is obtained with the help of the following two normal equations:

If the values of X and Y variables are substituted in the above equations we get the values of a and b and thus get the regression line of Y on X. Here Y is the dependent variable and X the independent variable. To get the regression line of X on Y, we will have to assume X as the dependent variable and Y as the independent variable. We will then get the two equations for the two regression lines. Regression line of X on Y

Will be given by the normal equations:

After obtaining the two regression equations, from the first equation (Y on X) we obtain any two values of Y for some values of X, plot them, and join them with a straight line to get the regression line of Y on X. Similarly, from the regression equation of X on Y, we get two values of X for any two values of Y, plot and join them with a straight line to obtain the regression line of X on Y. The two lines cut each other at the point which gives the average values of X and Y. After drawing the regression lines, in order to find any value of Y for a given value of X we will draw a perpendicular from the X-series (for the given value of X) and the point at which it cuts the regression line of Y on X will indicate the coupled value of Y which can be read on the Yscale (by drawing a line parallel to X-axis form the point where the perpendicular joins the regression line of Y on X). Similarly values of X-series can be found by using the regression line of X on Y.

MBA-015

Business Statistics

Page 54

Why Two Regression Lines? The answer to the above question is that one regression line cannot minimize the sum of squares of deviations for both the X and Y series unless the relationship between them indicates perfect positive or negative correlation. In case of perfect correlation, one line is enough because X and Y series have the same type of deviations. Ordinarily in social sciences perfect correlation is very rarely found. For this reason one regression line minimizes the sum of the squares of deviations of the Yseries and the other regression line takes care of the deviations of the X-series. Regression Equations Regression equation of Y on X:

Regression equation of X on Y:

In these equations a and b are constants which determine the positions of the lines of regression. The parameter a indicates the level of the line of regression (the distance of the line above or below the origin). The parameter b determines the slope of the line, that is, the corresponding change in Y in relation to per unit change in X or vice versa. The values of a and b in these equations will be given by the normal equations discussed above. Method of Deviations from the Means This method is much simpler than the method of least squares. Here, instead of taking the actual values of X and Y, we take their deviations from the mean values. According to this method, the regression equation of Y on X is given by

is called the Regression Coefficient of Y on X and is denoted by byx. Regression equation of X on Y is given by

is called the Regression Coefficient of X on Y and is denoted by bxy.

MBA-015

Business Statistics

Page 55

Thus, the Regression Coefficient of Y on X

Similarly, the Regression Coefficient of X on Y

Thus, the two regression equations can be rewritten as Regression equation of Y on X

Regression equation of Y on X

Properties of Regression Coefficients (1) The geometric mean of byx and bxy gives the value of the coefficient of correlation. That is

The sign of r will be the same as the sign of the two regression coefficients. (2) Both the regression coefficients will have the same algebraic sign. Either both will be positive, or both will be negative.

MBA-015

Business Statistics

Page 56

(3) The values of both the regression coefficients cannot be greater than one. At most, only one of them can be greater than one, not both. When deviations are taken from assumed mean

In case of grouped data

Partial and Multiple Correlation and Regression


Social and natural phenomena are generally affected simultaneously by a large number of factors. The effect of these factors on one another is studied by correlation and regression studies. In simple correlation between two variables, as we have studied earlier, our assumption was that the effect of other factors on the phenomenon under study is ignored. For example, when we study correlation between price and demand and treat price as a dependent factor and demand as an independent factor, we ignore the other independent factors like money in circulation, exports, and imports, which also affect price. Another assumption that we make in a study of simple correlation or regression is that the various independent variables affecting a dependent variable are not interrelated. It was assumed that they are mutually independent. However, in actual practice there is an association between different independent variables affecting a dependent variable. Thus, our study of simple correlation and regression makes such assumptions which are not true and to this extent the relationship studied is not absolutely dependable. It is necessary to study the effect of all the factors. Partial and multiple correlation and regression analysis is done to achieve this objective.

MBA-015

Business Statistics

Page 57

In partial correlation we study the effect of one independent variable on a dependent variable by excluding the effect of other independent factors. Thus if price (a dependent variable) is being affected by (i) demand,

(ii) money in circulation, and (iii) exports, we can study the relationship between price and demand by excluding the effects of money supply and exports. This will be called a study of partial correlation between price and demand. In simple correlation the effect of these factors was not excluded. It was just ignored. In multiple correlation we study the effects of all the independent variables simultaneously on a dependent variable. Thus, if we study the effects of demand, money supple and exports (all independent variables) on price (dependent variable) at the same time or if we study their combined effect on price it would be called multiple correlation study. Multiple and partial correlation and regression studies are very complex in nature if many independent variables are involved. The calculations in such cases are possible only through the use of computers. Partial Correlation Partial correlation is also called net correlation. It is a study of the relationship between one dependent variable and one independent variable by keeping the other independent variables constant. Simple correlation between two variables is called zero order coefficient as no factor is held constant. If a partial correlation is studied between two variables by keeping a third variable constant, it would be called a first order coefficient, as one variable is kept constant. Similarly, if two variables are kept constant, we get the second order coefficient, and so on. A first order coefficient may be indicated by r12.3 (or r21.3), which means that a partial correlation is being studied between variables (1) and (2), while variable (3) is kept constant.

Partial Correlation Analysis: Utility and Limitations The analysis of partial correlation is of great significance particularly in the field of social sciences, where most of the phenomena have multiple causation. A number of variables are in operation at
MBA-015 Business Statistics Page 58

the same time and their effects cannot be studied in isolation as is done in physical and experimental sciences where the variables can be controlled and the effect of each variable can be studied separately. The partial correlation analysis is one such technique where it is possible to keep some variables constant and study the effects of only one independent variable on a dependent variable. In partial correlation analysis the effects of independent variables other than the one which is being studied are not ignored as is done in the case of simple correlation. The utility of partial correlation is great in interrelated series an din various experimental designs where interrelated phenomena are to be studied. However, partial correlation analysis has some limitations: (1) It is presumed in the calculation of partial correlation coefficient that the simple correlations or zero order correlations from which partial correlation is computed have linear relationships between the variables. In actual practice, particularly in social sciences, this assumption may lead to misleading results as a linear relationship generally does not exist in such phenomena. (2) The effects of the independent variables are studied additively not jointly. It means that it is presumed that the various independent variables are independent of each other. In actual practice this may not be true and there may be an interaction among the factors. (3) The reliability of the partial correlation coefficient decreases as their order gores up. Therefore it is necessary that the size of the items in the gross correlation should be large. (4) It involves a lot of calculation and its analysis is not easy. Multiple Correlation In multiple correlation we study three or more variables at a time. The effect of all the independent factors on a dependent factor is studied simultaneously. If it is desired to study the relationship between output, soil fertility, chemical manures and rainfall we will have to pick up the variable whose value has to be estimated. It will be called the dependent variable. If, for example, we wish to know how output is related to soil fertility, chemical manures and rainfall, then output would be the dependent factor and soil fertility, chemical manures and rainfall would be the independent factors. We shall then study the effect of all these three independent factors on output which is the dependent factor. Which factor should be chosen as the dependent factor would depend on the purpose of the enquiry. A dependent variable is indicated by X1 and the independent variables by X2, X3, X4, . The Coefficient of Multiple Linear Correlation is represented by R and the necessary subscripts are added to it. Thus

MBA-015

Business Statistics

Page 59

R1.23 = Multiple Correlation Coefficient with X1 as dependent variable and X2 and X3 as independent variables.

The following points should be noted about the multiple correlation coefficient: (i) It is a non-negative coefficient. Its value ranges between 0 and 1.

(ii) R1.23 r12. R1.23 r13. (iii) If R1.23 = 0 then r12 = 0 and r13 = 0. (iv) R1.23 is the same thing as R1.32. the position of the subscript to the right of the dot does not make a difference. (v) By squaring R1.23 we get the coefficient of multiple determination. Utility and Limitations The utility and the limitations of the multiple correlation coefficient are the same as those of partial correlation coefficient. Multiple correlation coefficient gives the effects of all independent variables on a dependent variable. This is very useful in understanding the movements in the dependent variable. If multiple and partial correlation coefficients are studied together, a very useful analysis of the relationship between the various variables is possible. The chief limitation of the multiple correlation coefficient is that like partial correlation coefficient it also presumes that a linear relationship exists between the zero order coefficients. It also assumes that the independent variables affect the dependent variable in an independent manner and have an additive property. If, however, there is inter-relation between independent variables their effects cannot be distinct nor can they be additive. The multiple correlation coefficient is cumbersome in calculation and difficult in interpretation. As such, it should be calculated with care and interpreted with caution. Multiple Regression Analysis In case of a simple regression study between two variables X and Y we find out the relative movement of Y-series for a unit movement of X-series and vice-versa. We therefore have two regression equations one of Y on X and the other of X on Y. In multiple regression analysis there are three or more variables say X1, X2 and X3. We now take X1 as the dependent variable and try to find out its relative movement for movements in both X 2 and X3

MBA-015

Business Statistics

Page 60

which are independent variables. Thus in multiple regression analysis the effect of two or more independent variables on one dependent variable is studied. The procedure for studying multiple regression is similar to the one we have for simple regression, with the difference that the other variables are added in the regression equation. If there are three variables X1, X2 and X3, the multiple regression would take the following form:

In the above equation b12.3 indicates the slope of the regression line of X1 on X2 when X3 is held constant. Similarly b13.2 indicates the slope of the regression line of X1 on X3 when X2 is held constant. In most problems the changes in X1 are due partially to changes in X2 and partially to changes in X3. It is for this reason that b12.3 is called the partial regression coefficient of X1 on X2, keeping X3 constant and b13.2 is called the partial regression coefficient of X1 on X3, keeping X2 constant. Assumptions of Linear Multiple Regression Analysis (i) The dependent variable is a random variable. The independent variable need not necessarily be random. (ii) The relationship between independent and dependent variables is linear.

MBA-015

Business Statistics

Page 61

UNIT IV PROBABILITY
Introduction In common parlance, the term probability refers to the chance of happening or not happening of an event. The theory of probability provides a numerical measure to the element of chance or uncertainty. It enables us to take decision under conditions of uncertainty with a calculated risk. Today the theory of probability has been very extensively developed and there is hardly any discipline physical or social where it is not being used. In the field of business and economics it is very widely used for quantitative analysis of various problems and it forms the very basis of the modern theory of decision making. Terminology Random [Link] is an experiment which, if conducted repeatedly under homogeneous conditions, does not give the same result. The result may be any one of the various possible outcomes. For example, if an unbiased die is thrown, it will not always fall with any particular number up. Any of the six numbers on the die may come up. Trial and Event. The performance of a random experiment is called a trial and the outcome an event. Thus, throwing of a die would be called a trial and the result (falling of any one of the six numbers on a face of the die) an event. Events can be either simple (elementary) or compound (composite). An event is called simple if it corresponds to a single possible outcome. Thus in throwing a die, the chance of getting 3 is a simple event. However, the chance of getting an odd number is a compound event. A compound event can further be decomposed into simple events. For example, in the abovementioned compound event, there are three simple events. Exhaustive Cases. All possible outcomes of an event are known as exhaustive cases. In the throw of a single die, there are six exhaustive cases. Favourable Cases. The number of outcomes which result in the happening of a desired event are called favourable cases. Thus, in a single throw of a die the number of favourable cases of getting an odd number are three, that is, 1, 3 and 5.

MBA-015

Business Statistics

Page 62

Mutually Exclusive Events. Two or more events are said to be mutually exclusive if the happening of any one of them excludes the happening of all others in single (same) experiment. Thus, in the throw of a single die, the events 5 and 6 are mutually exclusive because both cannot occur simultaneously. Equally Likely Cases. Two or more events are said to be equally likely if the chance of their happening is equal, that is, there is no preference of any one event over the other. Thus in a throw of an unbiased die, the coming up of 1, 2, 3, 4, 5 or 6 is equally likely. Independent and Dependent Events. An event is said to be independent if its happening is not affected by the happening of other events and if it itself also does not affect the happening of other events. Thus, in the throw of a die repeatedly, coming up of 5 on the first throw is independent of the coming up of 5 again in the second or subsequent throws. However, if we are successively drawing cards from a pack, without replacement, the events would be dependent. For example, the chance of getting a King on the first draw is 4/52. However, if this card is not replaced, the chance of getting a King in the second draw would depend on whether we got a King in the first draw or not. Permutations and Combinations The word permutation refers to the arrangements and the word combination refers to groups. These terms find their usage in the calculation of probability. The number of permutations of n dissimilar things taken r at a time

The number of combinations of n dissimilar things taken r at a time

Therefore

Some Basic Concepts of Set Theory Set. A set is a collection of distinct and well defined objects. The objects forming the set are called the elements or members of the set. Distinct means that no two elements in a set are the same and well defined means that on the basis of a rule, it should be absolutely clear whether a particular object belongs to a particular set or not.

MBA-015

Business Statistics

Page 63

Subset. A set A is called a subset of a set B if each element of set A also belongs to set B. Set A can be smaller or equal to set B. Equal Sets. Two sets A and B are said to be equal if all the elements of set A belong to set B and all the elements of set B belong to set A. Null/Empty/Void Set. It is a set having no element. It is denoted by { } or . Disjoint Sets. Two sets A and B are said to be disjoint if there is no element common in them, that is, if there is no element which belongs to both set A and set B. Union of Two Sets. If A and B are two given sets then their union is the set of those elements which belong either to set A or set B or to both. The union of sets A and B is denoted as AB. Intersection of Sets. If A and B are two sets then their intersection is the set of those elements which are common to both the sets A and B. it is denoted as AB. Complement of a Set. Let U be any set (the universal set) and set A be its subset. Then the complement of set A in relation to U is that set B whose elements belong to U but not to A. Complement of A is denoted by , A or Ac. Sample Space. The set of all possible outcomes of an experiment is called the sample space of the experiment and is denoted by S. Each outcome is called an element or a sample point of the sample space. Sample space is also called the universal set or event space or possibility set. An event is a subset of the sample space. An event may consist of one or more sample points. De Morgans Laws (i) (AB)c = (AcBc) (ii) (AB)c = (AcBc) Probability Ordinarily speaking, the probability of an event denotes the likelihood of its happening. The value of probability ranges between 0 and 1. If an event is certain to happen, its probability would be 1, and if it is certain that the event would not take place, then the probability of its happening is 0. Ordinarily, in social sciences, the probability of the happening of an event is neither 0 nor 1. The reason is that in social sciences we deal with situations where there is always an element of uncertainty about the happening or not happening of an event. For this reason the probability of the events is generally between 0 and 1. The general rule of the happening of an event is that if an event can happen in m ways and fail to happen in n ways, then the probability p of the happening of the event is given by
MBA-015 Business Statistics Page 64

Odds in Favour of the Occurrence of the Event

Odds Against the Occurrence of the Event

P(A) + P() = 1 Note: if n events are mutually exclusive and exhaustive, the sum of their individual probabilities = 1 Classical or a priori Probability This is the earliest approach to the theory of probability. It is based on the definition of probability given by Laplace the ratio of the number of favourable cases to the total number of equally likely cases. This theory assumes that the various outcomes of an event are equally likely and so the probability of their happening is also equal. Since this theory was one of the first attempts to determine the probability of an event even before the event has taken place, and often it tries to determine the probability of events which have never taken place before (that is, there is no past data available) it is called a priori probability. This theory is suitable for calculating the probabilities of various events in games of chance, where various events are equally likely to happen. Thus, while throwing a die there are six possible events that can happen. They are supposed to be equally likely and the probability of getting any particular number (say 5) would be 1/6. And the probability that this event will not take place = It means that P + q = 1 or 1 q = p or 1 p = q. Relative Frequency Theory of Probability The relative frequency approach to the theory of probability is not based on a priori probabilities. In many situations it is not possible to have equally likely events, on which assumption the classical theory of probability is based. For example, whether the prices of the shares of a company would go

MBA-015

Business Statistics

Page 65

up or down are not two events which may be called equally likely. No doubt there are three alternatives here: (i) The prices may remain constant.

(ii) The prices may go up. (iii) The prices may go down. But these events are not equally likely to happen. Similarly, whether a machine would turn out an acceptable article or an unacceptable article are two alternatives but they also are not equally likely. Therefore the a priori probability theory cannot be applied. In such situations the probability of the happening of an event is determined on the basis of past experience or on the basis of relative frequency of success in the past. Thus, if a machine has been turning out 10% unacceptable articles in past the relative frequency for unacceptable articles would be 10% of the total items. However, relative frequency should always be estimated on the basis of a large number of readings in the past. The larger the number of past readings, the greater will be the accuracy of the result. Since in relative frequency approach probabilities are calculated on the basis of past experience these probabilities are called posterior probabilities as contrasted with a priori probabilities calculated through the classical approach. It should be understood that a priori probabilities are generally calculated in games of chance and posterior probabilities in problems relating to various types of economic and social phenomena where prior probabilities are not constant. A priori probabilities are deductive in nature, and are based on theory instead of evidence or experience or experimentation. Posterior probabilities, also called empirical probabilities, are based on experience of the past and on experiments conducted. It is thus empirical in nature. For example, mortality tables, tables relating to expectation of life at various ages, birth rates, death rates, rates of depreciation of machinery, etc. are all based on past experience. The posterior probability of an event

Subjective Theory of Probability This theory is also known as Personalistic Theory of Probability because it presumes that any decision reflects the personality of the decision maker, and subjective elements are important in assigning a probability to an event. This is very true in actual life. Suppose we have data relating to the price of a share for the last 2 years and further supposing that out of 1000 quotations relating to

MBA-015

Business Statistics

Page 66

this share price, there was a price rise on 400 occasions, then the empirical probability of a price rise in this share is or 0.4

However, with these data, some people would buy the shares and others would sell them for the reason that their subjective estimate coupled with relative frequency approach may give them different ideas about the price of these shares in future. Here the personality of the decision maker is reflected in the ultimate decision. The decision under this theory is taken on the basis of the available data plus the effects of other factors, many of which may be subjective in nature. This theory is very commonly used in business decision making. The approach in this theory is very broad based and flexible. Axiomatic Approach to Probability This approach is entirely mathematical in character and is based on set theory. To start with, some concepts are laid down and certain properties or postulates commonly known as axioms are defined. From these axioms alone the entire theory of probability is derived by deductive logic. The axiomatic probability includes the concepts of both the classical as well as the empirical definitions of probability. In this concept, for mutually exclusive (disjoint) events, the probability of the happening of either A or B is given by P(AB) = P(A) + P(B) For events of simultaneous occurrence, or the probability of both A and B happening together is given by P(AB) = P(A) x P(B) Conditional Probability Two events A and B are said to be dependent when B can occur only when A is known to have already occurred or vice versa. The probabilities associated with such events are called conditional probabilities. The probability of the occurrence of event A when event B has already occurred is called the conditional probability of the occurrence of A given that event B has already occurred and is denoted by P(A/B) (read as: Probability of A, given B).

MBA-015

Business Statistics

Page 67

Theorems of Probability (i) Addition Theorem If A and B are any two events then the probability that at least one of them occurs P(AB) = P(A) + P(B) P(AB) If A and B are mutually exclusive events, then AB = , P(AB) = 0, therefore P(AB) = P(A) + P(B) If there are three events A, B and C, the probability of the occurrence of at least one of them P(ABC) = P(A) + P(B) + P(C) P(AB) P(BC) P(CA) + P(ABC) If the events are mutually exclusive P(ABC) = P(A) + P(B) + P(C) (ii) Multiplication Theorem The probability of the simultaneous occurrence of the events A and B P(AB) = P(AB) = P(A) x P(B/A) = P(B) x P(A/B) If the events A and B are independent P(B/A) = P(B), P(A/B) = P(A) Therefore, for independent events P(AB) = P(A) x P(B) Bayes Theorem and Inverse Probability Bayes theorem is based on the proposition that probabilities should be revised when new information is available. The idea of revising probabilities is used by us in our daily life. Probabilities are often revised as soon as some new information is available about the problem concerned. The need for revising probabilities arises from a need to make better use of available information, and thereby reduce the element of risk involved in decision making. Bayes theorem is very widely used in decision theory. On the basis of this theorem, probabilities of various outcomes are revised upwards or downwards depending on the evidence obtained. The probabilities before revision are called prior probabilities and those after revision, posterior probabilities. Theorem Imagine a situation where two uncertain events (A and Not A) are possible. Suppose we know their probability, that is, we know the probability of As happening and also the probability of As not happening. These probabilities are prior probabilities because they are probabilities before any
MBA-015 Business Statistics Page 68

further information is available. Suppose an investigation is now conducted. The investigation may have several outcomes which would be dependent on event A. For any particular outcome (which may be called B) the conditional probabilities P(B/A) and P(B Not A) are available. The result itself serves to revise the probabilities for event A and event Not A. The resulting values would be posterior probabilities since they have been obtained after the results of the investigation. The posterior probabilities are actually conditional probabilities of the form P(A/B) or P(Not A/B). Thus, according to Bayes Theorem, the posterior probability of event A for a particular result of an investigation B may be found as

Inverse Probability It is the probability of the happening of an event as a result of factor (a) if the event could have happened as a result of factors (a) or (b) or (c) or or (n). With this concept of inverse probability, the Bayes theorem can be stated as: An event A can occur only if any one of the set of exhaustive and mutually exclusive events B1, B2, B3, , Bn occurs. The probabilities P(B1), P(B2), P(B3), , P(Bn) and the conditional probabilities P(A/Bi) where i = 1, 2, 3, , n are known. Then the conditional probability P(Bi/A) when A has actually occurred is given by

Where i = 1, 2, 3, , n Joint and Marginal Probabilities Joint probabilities are arrived at by multiplying two or more probabilities depending on the number of events involved. Marginal probabilities are the sum of probabilities (prior or posterior) of two or more events. Both joint and marginal probabilities are used in the Bayes theorem. Mathematical Expectation Mathematical expectation is the weighted mean of expected values of a variable. The mean is weighted in the sense that the various values of a variable are multiplied by their probabilities. If X is a random variable which can assume any one of the values x1, x2, x3, , xn, with respective probabilities p1, p2, p3, , pn, then the mathematical expectation of X (or the expected value of X) which is denoted by E(x) would be
MBA-015 Business Statistics Page 69

E(x) = p1x1 + p2x2 + p3x3 + + pnxn In a game of chance if a player would gain a sum a if he wins and would lose a sum b if he loses then the mathematical expectation would be (a x p) + (-b x q) = ap - bq Here loss is regarded as a negative gain. If the mathematical expectation of a game is zero, it is a fair game. If it is more than zero, it is biased to the player, that is, the game is in favour of the player. If the mathematical expectation is negative, the game is biased against the player. At most gambling places, the mathematical expectation for the players is negative. Addition and Multiplication Laws of Expectation If X and Y are two random variables then the expected value of X and Y together, that is E(X+Y) = E(X) + E(Y) If X and Y are two independent random variables then the expected value of XY, that is E(XY) = E(X) E(Y) Variance of X in terms of expectation

Since the expected value of X is the arithmetic mean of X series over a period of time

Some more results E(ax + b) = aE(x) + b Var(ax) = a2Var(x) Var(a + bx) = b2Var(x) Var(ax + by) = a2Var(x) + b2Var(y) Var(a) = 0
MBA-015 Business Statistics Page 70

Where a and b are constants.

MBA-015

Business Statistics

Page 71

PROBABILITY THEORETICAL DISTRIBUTIONS


Introduction Frequency distribution can be obtained in two ways, namely, (i) By compiling actual frequencies through collection of data or by conducting experiments, (ii) By deriving the frequencies on the basis of some mathematical relationships. The actual or observed frequency distributions (also called empirical or experimental distribution) are used to estimate certain values in the universe on the basis of sample studies. Thus, when we are finding the average height of adult males through a sample study our purpose is to estimate the average height of the adult males in the universe from which the sample has been taken. However, there are certain situations where we can derive expected values on the basis of some mathematical relationship. For example, if we toss a coin we know that the probability of heads is and of tails is also . If we toss a coin 100 times the expected number of heads is 50. This is the theoretical or derived frequency of heads. In actual practice, if we toss a coin 100 times and heads come up 60 times, then this is the observed frequency of heads. Thus, in such cases, we have an observed frequency of 60 and an expected frequency of 50. If we toss 10 coins 256 times we can have a set of observed frequencies by conducting the experiment and getting observed frequencies. We can also have theoretical frequencies by finding out the value of would be called a theoretical frequency distribution or a probability distribution. Thus, we can define theoretical or probability distributions as such distributions which are not obtained by actual observations or experiments but are mathematically deduced on certain assumptions. The importance of theoretical distributions cannot be overemphasized. They provide us data on the basis of which the results of actual observations or experiments can be assessed. In fact, where theoretical distributions are available, there is no need for having observed distributions. Theoretical distributions provide the decision maker with a sound basis for taking rational and dependable decisions. Theoretical distributions are of many types. We will, however, discuss only three such distributions which are more popular than others and are widely used. These are (i) Binomial distribution, . This distribution

(ii) Poisson distribution, and (iii) Normal distribution.


MBA-015 Business Statistics Page 72

Out of the above, the first two, namely, the binomial and the Poisson distributions are discrete distributions and the normal distribution is a continuous distribution.

Binomial Distribution
It is a particular case of multinomial distribution and is of very great importance in research and problems connected with probability and sampling. Binomial distribution studies such experiments which can have only two possible outcomes. Let us consider the tossing of a coin. Suppose we toss two coins (a and b) simultaneously. The possible outcomes shall be Coins a b H H H T T H T T

H = Heads T = Tails

Thus the coins can fall in any one of four different ways, as shown above. If p stands for the probability of a coin falling heads and q for falling tails, then The probability of two heads = p x p = p2 The probability of one heads and one tails = (p x q) + (q x p) = pq + pq = 2pq The probability of two tails = q x q = q2 The above results are the expansion of (p + q)2 (p + q)2 = p2 + 2pq + q2 Let us now analyse the sample of three coins. Ways of Arising a b c Three heads H H H H H T Two heads, one tails H T H T H H H T T One heads, two tails T H T T T H Three tails T T T Type of Event Probability of Way Probability of the Type of Result p3 p2q pqp qp2 pq2 qpq q2p q3 p3 3p2q

3pq2 q3

These are the terms in the expansion of (p + q)3 In case of an unbiased coin, p = q = . Therefore p3 = ()3 = 1/8 3p2q = 3 x ()2 x = 3/8
MBA-015 Business Statistics Page 73

3pq2 = 3 x x ()2 = 3/8 q3 = ()3 = 1/8 Similarly, the terms of the binomial expansion of

Now, probability of n successes

Probability of zero successes

Probability of at least one success = P(1) + P(2) + P(3) + + P(n) = 1 P(0)

Probability of exactly r successes

Probability of at least r successes = P(r) + P(r + 1) + P(r + 2) + + P(n)

Probability of at most r successes = P(0) + P(1) + P(2) + + P(r)

If these n trials are repeated N times, the number of exactly r, (r-1), (r-2), successes can be obtained by the expansion of N(p + q)n In this manner we can obtain the theoretical or expected frequencies when, say, N = 256. The actual frequencies when n coins are tossed simultaneously 256 times shall most probably be different from these. However, if the value of N is very large (say, 2560 or 25600) then the difference between the expected and observed frequencies would not be significant.

MBA-015

Business Statistics

Page 74

Assumptions of Binomial Distribution (i) The number of trials or n is finite and fixed. The performance must be repeated a fixed number of times. (ii) The experiment must result in either of the two events happening or, in other words, there must be only two possible outcomes of the event which are mutually exclusive and exhaustive. (iii) The probability of the happening of the event (or success) in any trial of the event is constant for all trials. Similarly, the probability of not happening of the event (or failure) is also constant. In other words p and q shall remain constant in all trials. (iv) All trials must be independent of each other. General Form of Binomial Distribution The general form of binomial distribution depends on the following two things: (1) The values of p and q, and (2) The value of the exponent n. If p = q = , the distribution would be symmetrical and the expected frequencies on either side of the central value would be identical. Thus, in the case where four coins are tossed 256 times, the frequencies for 0, 1, 2, 3 and 4 heads would be 16, 64, 96, 64 and 16, respectively. The distribution is therefore symmetrical. If, however, p is not equal to q, the distribution would not be symmetrical but skewed. If p < q, the distribution will be positively skewed (skewed towards right) and if p > q, the distribution will be negatively skewed (towards left). If p = q, the effect of increasing the value of the exponent n would be that the value of mean would increase and so would the value of measure of dispersion. If p q, an increase in the value of n would raise the value of the mean and dispersion, but would reduce the asymmetry of the distribution. It means that even if p and q are not equal, even then we can obtain a more or less symmetrical distribution if the value of n is large. Thus we can conclude that the type of distribution that we shall obtain would depend on the values of p and q and also on the value of n. Constants in a Binomial Distribution The following relationships hold good in a binomial distribution: (i) The mean of a binomial distribution is np.

(ii) The standard deviation of a binomial distribution is (npq) and variance is npq. (iii) The moments in a binomial distribution are as follows:

MBA-015

Business Statistics

Page 75

1 = 0 2 = npq 3 = npq(q p) 4 = 3n2p2q2 + npq(1 6pq) (iv) The 1 and 2 (on the basis of the above moments) are

If 1 = 0, the distribution is symmetrical. (v) The values of 1 and 2 are

If the distribution is normal, both 1 and 2 would be zero. Note: In a binomial distribution, variance would always be less than mean as mean is np and variance is npq (neither p nor q = 0 and both p and q are less than 1 as p + q = 1).

Poisson Distribution
This is also a discrete probability distribution ands is useful in such cases where the value of p is very small and the value of n is very large. In such cases the binomial distribution does not give appropriate theoretical frequencies. The Poisson distribution in such situations has been found to be quite suitable. Poisson distribution is a limiting form of binomial distribution as n moves towards infinity and p moves towards zero but np or mean remains constant and finite. This distribution is used very widely both in physical sciences as well as social sciences. This distribution studies the probabilities of rare events. Rare events are common in all sciences. In physics the emission of radioactive substance is a rare event. In biology, the number of bacteria in a unit may be very small. The number of defective articles produced by a high quality machine may be very small. The number of accidental deaths by falling from a roof would again be a small number. These phenomena are rare in the sense that the probabilities of their happening is very small. In fact

MBA-015

Business Statistics

Page 76

we do not know the probabilities of their not happening. For example, we do not know the number of people who did not die by accidentally falling from a roof. All we know is the number of a few people who died in this manner. In all such cases the Poisson distribution is used to give theoretical or expected values. Conditions of Poisson Distribution (i) (ii) (iii) (iv) (v) The variable is discrete. The number of trials, that is, n is very large. The probability of success, that is, p is very small. The probability of success in each trial is constant. np is constant and finite.

Form of Poisson Distribution (i) It is a discrete distribution.

(ii) It has a single parameter the mean of the distribution denoted by m. (iii) The probability of exactly 0, 1, 2, 3, , n successes is found by the successive term of the expansion

Where e = 2.71828 When x = 0, probability of exactly zero success

Similarly, probability of 1 success

Probability of 2 successes

Probability of r successes

MBA-015

Business Statistics

Page 77

Constants of Poisson Distribution 1. Mean of Poisson distribution is m Mean = rp(r)

2. Variance of Poisson distribution is also m Thus Mean = Variance = m, = m

3. The moments (about mean) of the Poisson distribution are 1 = 0 2 = variance = m 3 = m 4 = m + 3m2

Fitting of a Poisson Distribution Finding out the theoretical values in a Poisson distribution is very easy. We have to first find out the value of mean and calculate the frequency of 0 success. Once this is done the other frequencies can be obtained very easily as shown below When r = 0

MBA-015

Business Statistics

Page 78

If this probability is multiplied by N (total number of observations) we get the frequency for 0 success

Normal Distribution
Normal distribution is a continuous probability distribution and is most important of all the theoretical distributions. Many physical measurements and natural phenomena have observed frequency distributions which very closely resemble the normal distribution. These include measurements of heights and weights of both persons and things, characteristics like IQ, etc. A more important reason for the importance of this distribution lies in the fact that a theoretical property of the sample mean allows us to use the normal distribution to find out the probabilities of various sample results. This distribution is extremely useful in the analysis of data concerning economic and business problems. This distribution plays a basic role in situations where inferences are made regarding the value of the population mean when only the sample mean can be calculated directly. Such situations are very common in the field of social sciences. The normal distribution is a limiting case of binomial distribution if (i) The number of trials or the value of n is very large (nearly infinity) (ii) If neither p nor q is very small. If p and q are nearly equal, the normal distribution is very close to binomial distribution. The normal distribution can be regarded as a limiting case of Poisson distribution if the value of mean, that is, m is very large (nearly infinity). Importance of Normal Curve 1. If a random sample is taken from any universe then as the sample size increases the mean of the sample approaches the normal distribution (with mean and standard deviation /n). this is the central limit theorem. This property of normal distribution is very important and it
MBA-015 Business Statistics Page 79

enables us to draw inferences about the universe by making sample studies. According to this theorem we can estimate the upper and lower limits within which a value in the universe would lie, by conducting sample studies. 2. Even when the assumptions of a normal distribution are not satisfied, the results given by a normal distribution study, in many cases, are found to be highly satisfactory. However, theoretically the results hold good under the assumption that the properties of the normal distribution are applicable to the problem under study. 3. There are many mathematical properties possessed by this distribution which make its application possible in a wide variety of situations and for making varied types of studies. 4. It is very useful in statistical quality control where the control limits are set by using this distribution. Properties of Normal Distribution 1. The normal curve is symmetrical about the mean, that is, there is no skewness in it. 2. The mean, median and mode have the same value. 3. The height of the normal curve is maximum at the mean value. This ordinate divides the curve into two equal and identical parts. 4. Since the curve is symmetrical, the first and third quartiles are equidistant from the median. 5. Since there is only one point of maximum frequency, the normal distribution is uni-modal. 6. The mean deviation is 0.7979 of the standard deviation. 7. Semi-inter-quartile range is equal to the probable error which is 0.6745 of standard deviation. 8. The points of inflection (the points at which the curve changes direction) are each at a distance of one standard deviation from the mean. 9. The curve is asymptotic to the base line. It continues to approach but never touches the base line. 10. The various ordinates at different distances from mean ordinate (in terms of standard deviation) stand in a fixed proportion to the height of the mean ordinate. Thus the height of an ordinate at (one standard deviation) distance on either side of the mean ordinate is 60.653%

of the height of the mean ordinate. 11. The most important relationship in the normal curve is the area relationship. Since the ordinate at a given distance from the mean has always the same relationship with the mean ordinate,

it follows that the area of the curve enclosed between the mean ordinate and an ordinate at distance from mean would always be the same proportion of the total area of the curve. Thus, the area enclosed between the mean ordinate and an ordinate at a distance of one standard deviation ( ) from the mean is always 34.134% of the total area of the curve. It means that the
MBA-015 Business Statistics Page 80

area enclosed between the two ordinates at

distance from mean on either side would always

be 68.268% of the total area. Similarly the area between the two ordinates at 2 distance from mean on either side would be 95.45% of the total area. Ordinates at 3 distance from mean on either side would enclose 99.73% of the total area. Area of the Normal Curve between Mean Ordinate and the Ordinates at various Sigma Distances from the Mean as percentage of the Total Area Distance from the Mean Ordinate 0.5 1.0 1.5 1.96 2.0 2.5 2.5758 3.0 Percentage of Total Area at one side of mean at both sides of mean 19.146 38.292 34.134 68.268 43.319 86.638 47.500 95.000 47.725 95.450 49.379 98.758 49.500 99.000 49.865 99.730

We have specifically mentioned the area enclosed between the ordinates at 1.96 distance, 2.578 distance and 3 distance because in various tests of significance these figures are most commonly used. Constants of Normal Distribution (i) The mean of the normal distribution is X and the standard deviation is .
2

(ii) 2 = 3 = 0 3 = 3 (iii) (iv)

(normal distribution has no skewness) (normal distribution has no kurtosis)

The Standard Normal Curve A normal curve in which mean is zero and standard deviation is unity is known as the standard normal curve. Its utility is very great because curves with mean X and standard deviation can be

converted into a standard normal curve by change of origin and scale. This becomes necessary as otherwise in different distributions with different values of mean and standard deviation, it would be very difficult to find out the area between various ordinates. Once they are converted into standard normal curve, this problem is solved.

MBA-015

Business Statistics

Page 81

In the original scale the mean and standard deviation are X and

but in the new scale which is called

the z-scale, the mean is zero and standard deviation is unity. This process of changing the X-scale into z-scale is called z-transformation. z is called the standard normal variate and is given by

Where X is the mean and

is the standard deviation of the given normal distribution, X is the

observed value at which we want to find the value of z. The values of z for different values of X, define a normal distribution with mean 0 and standard deviation equal to unity. This new distribution is called standard normal distribution or unit normal distribution. The curve drawn with this distribution is called standard normal curve. One should not that (i) The area under the standard normal curve is unity.

(ii) The mean of the standard normal curve is 0. (iii) The standard deviation of the standard normal curve is equal to unity. Equation of the normal curve:

where y = computed height of an ordinate at distance x from mean y0 = height of the maximum ordinate at the mean. It is constant in the equation. e = 2.71828 = standard deviation x = any given value of the dependent variable expressed as deviation from mean, i.e., X = (X X) The maximum ordinate

where N = total number of items in the sample i = class interval = 3.1416, (2) = 2.5066

MBA-015

Business Statistics

Page 82

Thus

The equation of the normal curve can now be written as

Equation of the Standard Normal Distribution

where

when we have to find the area under the standard normal curve we only change the value of X into value of z. After this we can find out the area enclosed between various values of z form the tables.

MBA-015

Business Statistics

Page 83

UNIT V SAMPLING THEORY


Introduction Statistics which are collected and analysed are either descriptive or inductive in character. Descriptive statistics are those which describe some characteristics of a set of figures. Inductive statistics refer to drawing inference about a universe on the basis of the examination of a part of the universe only. In other words, inductive statistics refer to estimating values of the universe on the basis of sample study. In modern decision-making process in various fields of human activities, most of our decisions are based on the examination of a few objects only or in other words they are based on sample studies. This process of drawing inferences about a universe on the basis of sample studies naturally involves an element of risk, because one may draw wrong inferences about a universe by studying a sample. Modern statistical theory makes an attempt to evaluate this risk in terms of probability. Types of Universe Finite and Infinite Universe By finite universe we mean such populations which contain a definite number of units. As against this, an infinite universe is one in which the number of units is infinite. For example, the number of leaves in a tree is an infinite universe. Even though it may be possible to count the number of leaves in a tree, since the number is quite large, it is considered an infinite universe. Infinite populations are better for sampling studies as the probabilities of various events can be better estimated of the universe is infinite. Hypothetical and Existent Universe Hypothetical universe is one which does not consist of concrete objects. Such universes consist of an infinite number of items. Existent universe, as the name suggests, refers to a population of concrete objects. In hypothetical universe the values of p and q remain constant in various trials. This is a very important property of such populations. It enables us to fit a particular curve to such data with a high degree of accuracy. Objects of Sampling The most important aim of sampling studies is to obtain maximum information about the phenomenon under study with the least sacrifice of money, time and energy. Another aim of sampling studies is to obtain the best possible values of the parameters.
MBA-015 Business Statistics Page 84

Principles of Sampling (i) Law of Statistical Regularity. A group of objects chosen at random from a particular group tends to possess the characteristics of that group (universe). (ii) Principle of Inertia of Large Numbers. As sample size increases, results would be more reliable. (iii) Principle of Persistence. If some items of the universe possess some specific characteristics, these characteristics would be found in the sample also and even if the sample size is increased or population is increased, these characteristics would be reflected. (iv) Principle of Optimisation. Effort should be made to get the best possible or optimum results both in terms of cost as well as efficiency. (v) Principle of Validity. A sample design is called valid only when the inferences drawn from it are valid for the universe from which the sample has been taken. Census versus Sample Enumeration Census Method. In census type of enquiry there is complete enumeration of all the items of the universe and the question of taking a sample does not arise. Normally, census method should give exact and accurate results, but it suffers from certain limitations: (i) It is very expensive, particularly if the size of universe is large.

(ii) It requires more time fro completion. (iii) It needs substantial manpower and administrative control. The census method is however used in certain cases where (a) Complete information is needed about the entire universe. (b) The size of universe is not big and the need for accurate result is great. (c) The occurrence of a defect (during the manufacture of an item) may cause damage to the machine itself or to the life of the person(s) operating it. Sample Method. Sample method is one where we study some selected items from the entire universe for drawing general inferences. This method has many advantages over census such as speed, economy, adaptability and scientific approach. The merits of sampling method are: (i) It takes less time.

(ii) It is less expensive. (iii) It is more dependable. Even though the sampling method is not free from sampling and nonsampling errors, yet if it is properly designed, results would be more reliable and dependable than the results of census method. The reason is that in case of sampling it is possible to find out the degree of reliability of the result in terms of probability. Further, non-sampling errors
MBA-015 Business Statistics Page 85

due to wrong recording of observations or non-response or incomplete response or improper training of investigators, etc., which are substantial in case of census enquiry, can be easily controlled if the enquiry uses sampling where the numbers involved are less and one can have better trained personnel and can also use sophisticated statistical techniques of processing and analysis of data. (iv) In some cases only sampling method is possible. (v) This method has to be used in such cases also where the item itself is destroyed in the course of inspection. Limitations of Sampling Sampling studies can give better results only if the sample are drawn systematically, their size is adequate and an appropriate sample design is used. Other limitations are that if, for example, some selected units of the sample do not respond, or if the personnel conducting the survey are not qualified, the sample results may be highly misleading. Sometimes, if the sample size is large, sample surveys may need more time, money and labour also. Precision in Sampling Since the main aim of studies is to obtain information about the problem under study in the universe at large, and since sampling studies are made only from a few units collected out of a large number constituting the universe, the question of dependability often arises. It is clear that if a sample fails to reveal the main characteristics of the universe, it does not serve its purpose. Theory of sampling makes an attempt to indicate the degree of reliability that can be placed on various estimates obtained from sampling studies. This is done by assigning limits within which the estimate is expected to vary. These limits vary with the degree of confidence which we wish to achieve in our assertions. Thus, if we want to assert a fact with a very high degree of confidence, the limits which can be placed, will be wide so that the chance of the estimate going beyond them is minimum. It means that the degree of confidence which can be put in any estimate is expressed in terms of probability. The accuracy or precision of estimates depends on a variety of factors. The first is the manner in which the estimate is made from the sample data. This leads us to the theory of estimation. The second is the manner in which the sample was obtained. This leads us to the study of the technique of sampling. A third factor is the size of the sample. If the size of the sample is small, much reliance cannot be placed on the estimate.

MBA-015

Business Statistics

Page 86

Errors in Sampling The word error is used in a specialized sense in statistics. It means the difference between the true value and the estimated value. Statistical errors may arise due to inappropriate definitions of statistical units, bias of the investigator or the inherent instability of the collected data. Such errors are called errors of origin. Errors may also arise on account of manipulation in counting, measurement, description or approximation. Such errors are known as errors of manipulation. Yet another cause of statistical errors may be the use of incomplete data, errors may also arise on account of inadequacy of the size of the sample and all such errors are called errors of inadequacy. Sampling and Non-Sampling Errors Errors in statistics are classified into two categories: (1) Sampling Errors Sampling errors have their origin in sampling and they arise on account of the fact that sample has been used to estimate parameters or population values. Sampling errors are attributed to fluctuations in sampling. These arise due to the following reasons: (i) Improper selection of the sample.

(ii) Substitution. (iii) Faulty demarcation of statistical units. (iv) Errors due to variability of population and wrong method of estimation. Measurement of Sampling Error. A measure of sampling error is provided by the standard of the estimate. Estimation of sampling error can reduce the element of uncertainty associated with interpretation of data. In most cases the degree of precision or the opposite of error, would depend on the size of the sample. The standard error of estimate is inversely proportional to the square root of the sample size. In other words, as the sample size increases, element of error is reduced. (2) Non-Sampling Errors Non-sampling errors generally arise when data are not properly observed, approximated and processed. These arise due to the following reasons: (i) Improper or ambiguous definition of various terms. (ii) Incomplete questionnaire and defective methods of interviewing. (iii) Personal bias of the investigator. (iv) Lack of trained and qualified investigators and failure of respondents to give correct answers. (v) Improper coverage and inadequate or incomplete response. (vi) Errors in compilation and tabulation.

MBA-015

Business Statistics

Page 87

Measurement of Errors Statistical errors can be measured either: (a) Absolutely or (b) Relatively Absolute and Relative Errors. If the error is measured absolutely it is called an absolute error and if it is measured relatively it is called relative error. Absolute error is the difference between the true value and the estimate. Relative error is the ratio of the absolute error to the estimate. If U stands for the actual value, U for the estimated value, Ue for the absolute error and e for the relative error, then

Positive and Negative Errors. Absolute and relative errors can be either positive or negative. If the true value exceeds the estimate, the error is said to be positive and, on the other hand, if the estimate exceeds the true value the error is called negative. Classes of Errors Broadly speaking, errors may be either (a) Biased, or (b) Unbiased. Biased Errors. Biased errors are those which arise on account of some bias in the mind of the investigator or the informant or in the instruments of measurement. Biased errors are cumulative. Unbiased Errors. Unbiased errors are those which arise just on account of chance. Unbiased errors are generally compensating. Types of Sampling Samples can be selected from a universe in any one of the following three manners: (i) Purposive or subjective or judgement sampling.

(ii) Probability sampling. (iii) Mixed sampling.

MBA-015

Business Statistics

Page 88

Some of the important sub-types of sampling schemes under probability and mixed sampling are: (1) Simple random sampling. (2) Stratified random sampling. (3) Systematic sampling. (4) Multi-stage sampling. (5) Quasi-random sampling. (6) Area sampling. (7) Simple Cluster sampling. (8) Multi-stage Cluster sampling. (9) Quota sampling. Choice of Sampling Techniques It is very difficult to say that any one of the techniques listed above would always be better than the rest. Factors like nature of the problem, size of the universe, size of the sample, availability of resources and time affect the choice of an appropriate method. For example, in studies where the size of the sample is small in relation to the size of the universe, purposive sampling would be better than random sampling, but as the size of the sample increases, the importance of random sampling also goes up. Judging Reliability of a Sample The reliability of a sample can be determined in the following ways: (i) A number of samples may be taken from the same universe and the results of various samples compared. If there is not much variation in the results of the different samples, it is a measure of its reliability. (ii) Sub-samples may be taken from the main sample and studied. If the results of the subsamples are similar to those given by the main sample, it is a measure of its reliability. (iii) If some mathematical properties are found in the distribution under study, the sample result can be compared with expected values obtained on the basis of mathematical relationship and if the difference between them is not significant, the sample has given dependable results. In distributions where binomial, normal, Poisson or any other theoretical distribution is applicable, sample results can be compared with the expected values and they would give an idea about the reliability of the sample.

MBA-015

Business Statistics

Page 89

ESTIMATION THEORY AND HYPOTHESIS TESTING


The aim of all sample studies is to obtain the best possible parameter values. Sample studies by themselves have no significance unless they enable us to have some idea about the universe from which the samples have been drawn. Therefore it is necessary that samples are representative of the universe and contain its main characteristics. We know that statistical data can be collected either through a census or through a sample. In case of a census study, the question of estimating parameter values does not arise as the study itself covers the entire universe. However, when sample studies are conducted, parameter values have to be estimated. In this section we shall discuss the techniques which enable us to make generalizations on the basis of sample studies. We will also discuss the extent to which such generalizations are valid and, if they are valid, the degree of confidence with which such a statement can be made. In sample studies, the following problems are encountered: (i) How to estimate parameter values or how to generalize the result of the sample to the entire population? (ii) How to find out whether the generalizations made are valid or not? (iii) What is the level of confidence with which a particular generalization can be made? The answers to these problems are provided by an important branch of statistics statistical inference, which can be broadly divided into the following heads: (i) Estimation Theory, and (ii) Hypothesis Testing.

Estimation Theory
Estimation theory refers to the techniques and methods by which population parameters are estimated from sample studies. Estimation of parameter is absolutely essential whenever a sample study has been conducted as people are generally interested in parameter values only. The theory of estimation has been grouped in two classes: (i) Point Estimation a single statistic is used to estimate the parameter value. (ii) Interval Estimation a probable range is determined within which the real value of the parameter is expected to lie.

MBA-015

Business Statistics

Page 90

Point Estimation A particular value of the sample (or statistic) which is used to estimate the parameter value is called the point estimate or estimator of the parameter. A good estimator would give us a parameter value which is as close to the real value as possible. It should possess the following characteristics: (i) It should be unbiased.

(ii) It should be consistent. (iii) It should be efficient. (iv) It should be sufficient. Interval Estimate Even the best possible point estimate may differ from the population value and make the estimate unsatisfactory. In such a situation, the question arises that can we make some generalizations about the universe values if not precisely, then within a certain range? This is what we aim at in interval estimate. If sample size is large, we try and estimate the parameter values within certain limits. We can also determine the probability of our being wrong while making a particular statement about a parameter. The limits within which a parameter value is estimated is called the fiducial limits or confidence interval or confidence limits. These limits would vary on the basis of the degree of precision which is desired to be achieved. We know that in a normal distribution mean 1.96 includes 95% of the items of the distribution. It means that if our limits are within mean 1.96 , the probability of our going wrong is 0.05.

Hypothesis Testing
A hypothesis in statistics is simply a quantitative statement about a population. It is an assumption that is made about parameter values and then its validity is tested by statistical techniques which ultimately tell us whether the hypothesis is correct and is sustained or whether it is false and is to be rejected. To important concepts associated with hypothesis testing and theory of estimation are sampling distribution and standard error. Sampling Distribution and Standard Error A sampling distribution is an array of sample studies relating to a universe. If we calculate the mean of the sampling distribution it could be deemed to be the mean of the universe. Similarly the standard deviation of the sampling distribution would be called the standard error.

MBA-015

Business Statistics

Page 91

If n samples are taken from a universe then the mean of their mean values would be close to the mean of the entire universe and the standard deviation of these mean values would be called the standard error. The concept of standard error is very useful in testing statistical hypothesis. If all possible samples of a certain size are taken from the universe then the mean of the sampling distribution of means would be equal to the mean of the universe. That is E(X) = where E(X) stands for the mean of sampling distribution and stands for the mean of the universe. further, variance of X is nearly equal to . Variance of X depends on the sample size the bigger the size of the sample, the smaller would be the variance. If the variance is small the standard error would also be small and the degree of precision would be greater. According to Central Limit Theorem, if X1, X2, X3, , Xn is a random sample from a normal population with mean and variance variance
2

, then the sample mean X is also normally distributed with mean and

. This is true even if the universe from which the samples have been taken is not normal,

provided the sample size is large (greater than 30). Standard Error Utility: (i) It provides and idea about the degree of precision of a sample or in other words, it tells us about the extent to which a sample is reliable. The greater the standard error, the greater is the element of unreliability of the sample and lesser is the degree of precision. Precision is the reciprocal of standard error . The reliability of a sample would vary as the square root of

the number of items in the sample. Therefore, if the degree precision has to be doubled or the degree of uncertainty has to be halved, the size of the sample would have to be increased fourfold .

(ii) The standard error helps in determining the limits within which a parameter value is expected to vary. This is so because, as we have already seen, the sampling distribution tends to be a normal distribution if the sample size is large. In a normal distribution we know that 95% of the items have a value within mean 1.96 (where denotes standard error). This means

that 95% of the sample means will range within the mean of the universe 1.96 . similarly, in a normal distribution mean 3 covers 99.73% of the items. The chance of sample mean going outside 3(standard error) is very small (p = 0.0027). it means that only 27 times out of 10,000 the sample mean can cross these limits.
MBA-015 Business Statistics Page 92

(iii) The standard error helps in testing a hypothesis. Usually hypothesis is tested at 95% level of significance which means that we can afford to be wrong only in 5 cases out of 100. In a normal curve 1.96 covers 95% of the values. Therefore, if our sample has given us a result which falls outside the range of 1.96 it is outside the level of precision desired. On the other hand, if the difference between the observed and expected values is within this range, we can say that the difference is on account of fluctuations of sampling and can be ignored. We will be wrong in our assertion only 5 times out of 100 or the probability of our being wrong is 0.05. If we want greater caution, we can test the hypothesis at a higher level of precision. We can set the range at 2.58 . this range covers 99% of the values. If the difference between observed and expected values falls within this range it means that the difference is not significant. We will be wrong in our assertion only once in a 100 cases (p = 0.01). Difference between Standard Deviation and Standard Error The difference between standard deviation and standard error is that the former concerns original values and the latter concerns the statistic computed from samples of original values. The standard deviation of the distribution of sample means is called the standard error of the mean. The standard deviation applies to the distribution of items around their average while the standard error of the mean applies to the distribution of averages of samples around the true average of the universe. Hypothesis Testing Hypothesis testing involves finding out whether the difference in the estimated value of the parameter and the true value, if known, is significant or it could have arisen due to fluctuations of sampling. Many times we may be interested in knowing whether the results given to us by two different samples drawn from the same universe are consistent with each other or whether the difference between their values is significant and could not have arisen due to fluctuations of sampling. In all such cases we formulate a hypothesis and then test its validity. A hypothesis is a quantitative statement about a population. It may or may not be true. By testing the hypothesis we can find out whether it deserves acceptance or rejection. Procedure for Hypothesis Testing (i) Set up a hypothesis. Generally a single hypothesis is not set. Two complimentary hypotheses are set at the same time. If one hypothesis is accepted then the other is automatically rejected and vice versa. These hypotheses are known as a. Null hypothesis or hypothesis of no difference (denoted by H0).

MBA-015

Business Statistics

Page 93

b. Alternative hypothesis (denoted by Hi). The concept of null hypothesis is very important in the theory of statistical inference and tests of significance. This hypothesis presumes that there is no difference between any two values which are being compared and any difference is on account of fluctuations of sampling. As against null hypothesis, the alternate hypothesis assumes that the differences are not due to fluctuations of sampling; they are real. An example: H0: = 120kg Hi: 120kg or Hi: > 120kg or < 120kg (ii) Determine the level of significance. It means that we have to determine the level of confidence with which a particular hypothesis is accepted or rejected. (iii) Decide the test statistic. A test statistic is a statistic computed from a sample. Test statistics are generally based on some probability distribution. Some of the common probability distributions which are used in hypothesis testing are z, t, 2 and F, etc. The choice of the test statistic would depend upon the nature of the distribution and the size of the sample. In case of small samples the properties of a normal distribution are not applicable and z-test has to be replaced by t-test. (iv) Draw conclusions. The last step in hypothesis is to draw conclusions about accepting or rejecting the hypothesis. A null hypothesis will be rejected if the sample statistic falls in the rejection region. For example, at 5% level of significance, the mean of a sample would not represent the mean of the universe if its value is not within 95% value enclosed in a normal distribution. The rejection area is 2.5% on either side of the normal curve. In such a case we reject the null hypothesis which means that we conclude that the difference between the sample mean and the population mean could not be due to chance. The alternative hypothesis is accepted. On the other hand, if the sample result is within the 95% level of acceptance, we conclude that the difference between sample mean and population mean could be due to chance and hence the sample mean can be used as a parameter value. The null hypothesis is accepted and the alternative hypothesis (of difference) stands rejected.

MBA-015

Business Statistics

Page 94

Errors in Hypothesis Testing Two types of errors are generally committed in hypothesis testing. They are Type I Error and Type II Error. Type I error is committed when we reject a correct or true hypothesis. Type II error is committed when we accept a wrong or incorrect hypothesis. Type I error (of rejecting a null hypothesis when it is true) is denoted by . = Probability of Type I Error = Probability of rejecting H0 when it is true. Type II error (of accepting a null hypothesis when it is not true) is denoted by . = Probability of Type II Error = Probability of accepting H0 when it is not true. If the difference between two means is zero and if our test indicated rejection of the null hypothesis, we commit type I error. If the difference between the two means is not zero but our test suggests acceptance of null hypothesis, we commit type II error. Accept H0 Reject H0 H0 is true No error Type I Error H0 is false Type II Error No error Though efforts are made to reduce both type I and type II errors, it is not possible to reduce both at the same time. The value of can be reduced only by increasing the value of . A business executive will have to compare the payoff of with the payoff of and find out if it is worthwhile to have a larger probability of type I error or a larger probability of type II error. However, in general, it is more risky to accept a false hypothesis (type II error) than to reject a true hypothesis (type I error). The probability of committing a type I error is kept at 5% level, that is, the probability is generally fixed at 0.05. in other words, the critical region is 5% and acceptance region is 95%. Two Tailed and One Tailed Test of Hypothesis The critical region (or the region of rejection) which is generally 5%, is kept on both sides of the normal distribution in a two tailed test. It means that 2.5% of the critical region is on the extreme left of the normal curve and 2.5% is on the extreme right. The remaining 95% area would be the acceptance region. Two tailed test is applied when, for example, the null hypothesis is that mean weight of the undergraduate male students is 120 pounds and the alternate hypothesis is that it is not 120 pounds. In such a situation the actual value of the population mean may be 120 pounds or more than 120 pounds or less than 120 pounds.

MBA-015

Business Statistics

Page 95

On the other hand, if the null hypothesis is that the mean weight of the male undergraduate students is 120 pounds and the alternate hypothesis is that it is more than 120 pounds then the one tailed test (right tailed test) would be applied. If the alternate hypothesis is that it is less than 120 pounds, then again we will use the one tailed test (left tailed test) Sign of Alternative Hypothesis > = < Type of Test to be Used Right Tailed Test Two Tailed Test Left Tailed Test

The Critical Value of z for Two Tailed and One Tailed Tests (taken from the Area Table of the Normal Curve) Critical Values of z Level of Significance 1% 5% 10% z Critical Value for One Tailed Test +2.33 or -2.33 +1.645 or -1.645 +1.28 or -1.28 z Critical Value for Two Tailed Test +2.58 and -2.58 +1.96 and -1.96 +1645 and -1.645+

MBA-015

Business Statistics

Page 96

TECHNIQUES OF ASSOCIATION OF ATTRIBUTES & TESTING Association of Attributes


Introduction Statistics deal with quantitative data. Quantitative data may arise in one of the following ways: (a) An investigator may measure the actual magnitude of some variable. In such a situation a fairly accurate quantitative measurement is possible. (b) However, at times the data might be such that it may not be possible for an investigator to measure their magnitude. In such cases the observer can only study the presence or absence of a particular quality in a group of individuals. He then has to take decision on the basis of some standard definition of the term in question. Such data in which the quantitative measurement of the magnitude is not possible and in which only the presence or absence of an attribute can be studied, are called the statistics of attributes. Classification of Data In the analysis of statistics relating to attributes, the first thing is the classification of data. Here data are classified on the basis of the presence or absence of a particular attribute. If only one attribute is being studied, the population would be classified into two groups one comprising of those units in which this attribute is present and the other consisting of those in which it is not present. If more than one attributes are taken into account, the number of classes would be more than two. The following points should be noted about the classification of data according to attributes: (1) Classification is Arbitrary and Vague. It should be clearly understood that when the universe is divided in, say, two classes, there may not be any clear-cut line of demarcation between them. One attribute may gradually transform into the other attribute and there may be many cases on the borderline between the two attributes. In all types of analysis relating to statistics of attributes, this point should always be kept in mind. (2) Classification by Dichotomy. If only one attribute is being studied the universe is divided in two parts one in which the attribute is present and the other in which it is not present. These classes are mutually exclusive. Such a classification where the universe is divided in two parts is called classification by dichotomy. In actual analysis there are more than two classes in which the universe is divided and such classification is called manifold classification. Notation and Terminology Usually capital letters A, B, C, are used to denote the presence of attributes and the Greek lower case letters , , , are used to denote the absence of these attributes respectively. The number of
MBA-015 Business Statistics Page 97

units possessing a particular attribute represented by A would be termed as belonging to Class A. Similarly, those in which this attribute is absent would be termed as belonging to Class . If two attributes are being studied, their combination can be represented by the combination of the letters representing the two attributes. Thus, if blindness is represented by A and deafness by B then AB would represent blindness and deafness, A would represent blindness and absence of deafness, B would represent the absence of blindness and presence of deafness and would represent the absence of both blindness as well as deafness. Class Frequencies. The number of units in different classes are called class frequencies. Class frequencies are represented by enclosing the class symbols by braces. Thus (AB) would represent the frequency of the class AB. Number of Classes. If there is one attribute represented by A, the total number of classes is 3, namely, A, and N (if the total or N is also taken as a class). In general the number of classes is equal to 3n, where n stands for the number of attributes. Order of Classes. The various classes and their frequencies are demarcated on the basis of an order. Thus N is a class of Zero Order. are classes of the First Order. are classes of the Second Order. Similarly the frequencies of these classes are also known as frequencies of the zero, first or second order. If there are only two attributes under study, then the second order classes and frequencies are called the classes or the frequencies of the ultimate order. Since these are the last set of classes and frequencies, the number of classes of the ultimate order is equal to 2n, where n stands for the number of attributes. Positive and Negative Classes. The classes which represent the presence of an attribute or attributes are called positive classes. The classes which represent the absence of an attribute or attributes are called negative classes. The classes in which one attribute is present and the other is absent are called pairs of contraries. Thus N, A, B and AB are positive classes. and are negative classes. A and B are pairs of contraries.

MBA-015

Business Statistics

Page 98

Relationship between Classes of Various Order In classifying statistical data according to attributes the following simple rule should be kept in mind. Any class frequencies can always be expressed in terms of class frequencies of higher order. On the basis of this rule we can set up various types of relationship[s between the frequencies of different orders. If there is one attribute only, represented by A, the frequency of the universe or N can be divided into two classes A and . Thus (N) = (A) + () Now, if one more attribute B is taken into account, the first order classes, that is, A and can each be divided into two classes one in which attribute B is present and the other in which it is absent. Thus (A) = (AB) + (A) () + (B) + () Similarly (B) = (AB) + (B) () + (A) + () Consistency of Data Meaning. It is obvious that in statistics of attributes no class frequency can be negative. If any class frequency is negative the data are said to be inconsistent. Such inconsistency may be due to wrong counting or inaccurate calculations. In order to test whether a set of figures is consistent or not, various class frequencies should be found and if no class frequency is negative, apparently the data are consistent. However, consistence of data is no proof of their accuracy. The easiest way to check for consistency is to find the ultimate class frequencies because if there is any inconsistency, one or more of these ultimate class frequencies would most probably be negative. Some Rules for Testing the Consistency of Data (I) In case of single attribute: (1) (A) 0, otherwise (A) will be negative. (2) (A) (N), otherwise () will be negative since (N) = (A) + (). (II) Two attributes: (3) (AB) 0, otherwise (AB) will be negative. (4) (AB) (A) + (B) (N), otherwise () will be negative.

MBA-015

Business Statistics

Page 99

(5) (AB) (A), otherwise (A) will be negative since (A) = (AB) + (A) (6) (AB) (B), otherwise (B) will be negative since (B) = (AB) + (B) Incomplete Data. The above rules are also used to fill the gaps if the data are incomplete. With the help of these rules it is possible to lay down the maximum and minimum limits of a particular class frequency. Methods of Studying Association Meaning of Association. Association of attributes refers to such techniques by which we can measure the relationship between two such phenomena whose size cannot be measured and where we can only find out the presence or absence of an attribute. Just as in case of correlation analysis we study the relationship between two variables whose value can be measured, similarly in case of association we study the relationship between two attributes which are not capable of quantitative measurement. The study of association can be done by any of the following methods: I. Comparison of Actual and Observed Frequencies. II. Comparison of Various Proportions and Products. III. Calculation of Yules Coefficient of Association. IV. Calculation of Coefficient of Collignation. V. Calculation of Coefficient of Contingency. Comparison of Actual and Observed Frequencies Whenever we want to study association between two attributes A and B we try to find out whether attribute A is more commonly found with attribute B than is ordinarily expected. Thus, in a study of association, the first thing to be calculated is the expected value of (AB). This value is calculated on the basis of simple rules of probability. Thus, if in a population of 100 students, 20 are married, the probability of coming across a married student is .

If two attributes A and B are studied in a universe and if the frequency of A is represented by (A) and of B by (B), then

MBA-015

Business Statistics

Page 100

The combine probability of two independent events is equal to the product of their individual probabilities.

Therefore, if attributes A and B are independent,

Criterion of Independence. If there is no kind of relationship between the attributes A and B we may expect to find the same proportion of As in Bs as in s. If this is not the case, it indicates an association between the attributes A and B. Two attributes A and B are said to be independent if the observed frequency of (AB) is equal to its expected frequency, i.e., .

Complete Association and Dissociation. There is complete association (perfect positive association) between two attributes A and B if all As are Bs or all Bs are As. This would be possible in any one of the following three cases. (i) (A) = (B) =(AB).

(ii) (A) = (AB). (iii) (B) = (AB). There would be complete dissociation (perfect negative association) between A and B if none of the As is B. Alternatively, if N contains no such unit which possesses neither of the two attributes, it is also taken to be a case of complete dissociation. Therefore, there is complete dissociation between attributes A and B if (i) (AB) = 0 or (ii) () = 0 Limitation. The main limitation is that this method only determines the nature of association between A and B, that is, whether the association is positive or negative. It does not tell us about the degree of association. While studying association using this method, it must be kept in mind that if the value of (AB) is found to be greater than the value of , it should not be at once concluded that there is

MBA-015

Business Statistics

Page 101

positive association between the two attributes. It is quite possible, particularly when the difference between observed and the expected values is not much, that the association may be the result of sampling fluctuations and the true association may be zero. As such, unless the difference the observed and the expected values is very significant, we should not conclude that there is any association or dissociation between the two attributes. Comparison of Various Proportions and Products This method is slightly better than the previous method as it compares the proportion between the various classes. Thus attributes A and B are (i) Independent if

(ii) Positively associated if

(iii) Negatively associated if

If the first relation holds good, the corresponding relationships

would also hold good further it can also be concluded that A and B would be independent if

It can also be found out easily that A and B would be independent if

Yules Coefficient of Association Yules coefficient aims at finding out the extent of association between two attributes. It is given by

MBA-015

Business Statistics

Page 102

If Q = 0, the two attributes are independent. If Q = +1, there is perfect association between the two attributes. If Q = -1, there is perfect dissociation between the two attributes. The chief characteristic of this coefficient is that it is independent of the relative proportions of As, Bs, s and s in the data. That is, if all terms containing A are multiplied by a constant, the value of Q would not be affected. The same would hold good for all values containing B, and .

MBA-015

Business Statistics

Page 103

SIGNIFICANCE TESTS IN VARIABLES (LARGE SAMPLES)


Introduction The aims of sampling studies in statistics of variables are the same as in case of statistics of attributes. Here also we compare the actual or observed frequencies with those expected under certain assumptions and try to find out whether the difference can be attributed to chance. As in sampling of attributes here too we try to obtain one or two constants for the universe mean or standard deviation because these can give an idea about the parent distribution. In sampling studies relating to variables, as in sampling studies of attributes, our third aim is to assess the reliability of our estimates. Difference between Large and Small Samples Though no hard and fast line can be drawn between large and small samples, yet, conventionally if a sample consists of more than 30 items, it is considered to be large. A sample consisting of up to 30 items is deemed to be a small sample. The tests of significance used in large samples are different from those used in small samples because small samples fail to satisfy the assumptions under which large sample analysis is done. The assumptions under which significance tests are applied in case of large samples are: (i) The random sampling distribution of statistic has the properties of the normal curve. In case of small samples it may not be so. (ii) Sample values are close to parameter values and can be used for analysis in the absence of parameter values. For example, for calculating the standard error, if the standard deviation of the universe is not known, it can be substituted by the standard deviation of the sample. This is not possible in case of small samples unless necessary adjustments are made. Significance Tests The significance tests used in statistics of variables are based on the calculation of standard error. The method of calculating standard error of various measures associated with variables are given below: 1. 2. 3. SE of Mean (when population SE of Mean (when population SE of Median

is known) is not known)

MBA-015

Business Statistics

Page 104

4. 5. 6. 7. 8.

SE of Quartiles SE of Quartile Deviation SE of Mean Deviation SE of Standard Deviation SE of Variance

9.

SE of Coefficient of Variation

10. SE of Coefficient of Skewness 11. SE of Pearsonian Coefficient of Correlation 12. SE of Spearmans Rank Correlation 13. SE of Regression Coefficient of Y on X (byx) 14. SE of Regression Coefficient of X on Y (bxy) 15. SE of Regression Estimate of Y on X 16. SE of Regression Estimate of X on Y 17. SE of Coefficient of Association Standard Error of the Difference of Sample Means Suppose two samples have given us two mean values and we have to find out whether there is a significant difference between the two values, or whether they could have come from one universe or universes having the same mean and standard deviation. Here we shall calculate the standard error of the difference of the two sample means and then we shall find out whether the difference between them is more than, say, three times the standard error of the difference. If the actual difference between the two means is more than thrice the standard error of the difference, it is said to be significant, otherwise the difference can be due to fluctuations of sampling. There is no hard and fast rule of taking the criterion of three standard errors only. We know that mean 3 standard error would cover 99.73% of the cases and the probability of a difference equal
MBA-015 Business Statistics Page 105

to or more than three standard errors arising out of chance would only be 10.9973 or 0.0027. we can test the results at 1.96 standard error in which case the probability of getting a difference equal to or more than 1.96 standard error due to sampling fluctuations would be 0.05. The formulae for the standard error of the difference of two sample means are as follows: 1. Difference of two Sample Means (If p is known) 2. Difference of two Sample Means (If p is not known) ( 1 = S1, 2 = S2) 3. Difference of two Sample Means (when correlation exists) 4. Difference between first Sample Mean and Combined Mean p = 12 (Combined) 5. Difference between second Sample Mean and Combined Mean p = 12 (Combined) Comparing Mean of a Sample with the Pooled Mean of two Samples Sometimes the mean of a sample is compared with the combined mean of the two samples taken together. In such cases the standard error of the difference of first sample and the combined sample is calculated as

Comparing Means of Correlated Samples If the two samples whose mean values are being compared are not independent but are correlated the formula for the calculation of standard error of the difference between two such means is modified as follows

Standard Error of the Difference between two Sample Medians The standard error of the difference of two sample medians

MBA-015

Business Statistics

Page 106

Standard Error of the difference between two Sample Standard Deviations 1. Difference of two sample standard deviations (when p is known) 2. Difference of two sample standard deviations (when p is not known) ( 12 = S12, 22 = S22) 3. Difference of deviation ( p = 4. Difference of deviation ( p =
1 12)

and combined standard

2 12)

and combined standard

MBA-015

Business Statistics

Page 107

SIGNIFICANCE TESTS IN SMALL SAMPLES


The analysis of small samples has to be done by techniques which are different from those applicable in case of large samples. Size of Samples and Standard Measures. Some effects of the smallness of the sample size on various standard measures: (1) The smaller the sample the greater would be the variation in the value of mean of different samples. (2) The standard deviation of the smaller samples tends to be smaller than true standard deviation of the universe; the smaller the sample, the smaller would be the standard deviation. (3) Since the standard deviation of small samples tends to be smaller than the standard deviation of the universe, all errors base on these smaller standard deviations also tend to be small. As such, unless corrections are made for the errors of small samples, they would tend to underestimate the actual error.

Students t Distribution
This distribution is applicable to small samples and was developed by W. S. Gossett. The Students t Distribution is used when the sample size is less than 30. The Students t Distribution is

Where the sample standard deviation

If the sample SD is given without using n1 as denominator

It should be noted that the only difference in the calculation of S in large and small samples is that whereas in the former case the sum of the squares of deviations of various items from mean, that is, (XX)2 is divided by n (the number of items), in case of small samples it is divided by n1, which are the degrees of freedom. The degree of freedom in such problems is n1 because one has the

MBA-015

Business Statistics

Page 108

freedom to change only n1 items as the last item has to be the difference between X and the sum of n1 items. The degrees of freedom are indicated by (Greek small letter nu). Once the value of t has been calculated, it is compared with its table value. Form of t-Distribution (1) Like normal distribution, it also has a symmetrical bell shaped frequency curve. The only parameter which determines that shape of the curve is the degree of freedom. With a change in the degrees of freedom, the shape of the curve will also change. (2) The mean of the t-distribution is zero, as in the case of normal distribution. (3) The variance of t-distribution is greater than one and as the sample size increases, it tends to move towards unity. It is for this reason that it has been mentioned above that the single parameter which changes the form of t-distribution is the degrees of freedom, or the size of the sample. If the sample size becomes very large, there is no difference between t-distribution and normal distribution. (4) t-distribution can be used even in case of large samples but the large sample theory cannot be used for small samples. (5) The value of t ranges from - to +. Application of t-Distribution (i) t-test for significance of single mean, population variance being unknown.

(ii) t-test for the significance of the difference between two sample means, the population variance being equal but unknown. (iii) t-test for significance of an observed correlation coefficient. Test Statistic for t-Distribution

Where

And

MBA-015

Business Statistics

Page 109

Or

If the value of the standard deviation of the sample is given without using n1 as denominator, then

Confidence or Fiducial Limits for For given degrees of freedom v and for a given level of confidence , the fiducial limits for are given by

Assumptions for t-Test (i) The parent population from which the samples are drawn is normal.

(ii) The given sample is random. (iii) The population standard deviation is not known. t-Test for Difference of Means Null Hypothesis:

(i)

The test statistic in such cases

Where a. X1 and X2 are the two sample means, b. n1 and n2 are the two sample sizes, c. S is the pooled or combined standard deviation of the two samples. (ii) The value of S or the combined standard deviation

MBA-015

Business Statistics

Page 110

When the deviations have been taken from assumed averages A1 and A2

(iii) When the standard deviations of the two samples are given, the pooled or combined standard deviation

(iv) If the calculated value of t is more than the critical value for the given degrees of freedom and at a given level of significance, the difference in the two samples is significant and the null hypothesis is rejected. Assumptions for Testing the Difference of Means (i) The universe from which the samples are drawn are normally distributed.

(ii) The two samples are random samples and have been drawn independently. (iii) The variances of the two populations are equal. It follows that the standard deviations of the two universes (
1

and

2)

are also equal.

(iv) Population variances are not known. Significance Test for Dependent Samples or Paired Observations In the t-test for difference of small sample means, we had presumed that the samples are independent of each other. However, there might be situations where the samples are not independent and the values of the second sample depend on the values of the first sample. Such a situation may arise in businesses where, for example, the effect of advertising on sales may have to be analysed. In this type of an experiment, n1 = n2 and every item has two values. The significance test applied in such cases for small samples is

Where d is the mean of the differences of the paired values. S is the standard deviation of the difference, given by

MBA-015

Business Statistics

Page 111

t-Test for Significance of an Observed Sample Correlation Coefficient The standard error of a small sample coefficient of correlation is given by

And the value of t is the ratio which r bears to the standard error:

The calculated value of t is compared with the critical value for n 2 degrees of freedom to find out whether the value of r is significant or not. Z-Test of Significance of Correlation Coefficient Z-Transformation According to this method of testing the significance of the coefficient of correlation in small samples, given by Fisher, the coefficient of correlation in the sample, r is transformed into Z, where

The standard error of Z

Similarly, the coefficient of correlation in the universe, P can be transformed into (the Greek small letter xi). The value of is calculated in exactly the same manner as the value of Z. If the coefficient of correlation of the universe is not known, it is supposed to be 0 and then the value of also becomes 0. To test the significance of a coefficient of correlation the difference between Z and is calculated and the relationship of the difference and the standard error of Z is then found out. If this difference is more than twice (or thrice, depending upon the level of significance), the standard error is supposed to be significant. Variance Ratio test (F-Test) One important property of F-test is that it can be used to determine whether the two independent estimates of population variance significantly differ between themselves, or whether they establish the fact that both the samples have come from the same universe and have a common variance. For

MBA-015

Business Statistics

Page 112

this purpose the ratio of variance given by the two samples is obtained. This ratio is called F-ratio, given by

Where

And

In calculating F the numerator is the greater variance and the denominator is the smaller variance. Therefore its value can never be less than 1.

The calculated value of F is compared with the critical value of F for the given level of significance and given degrees of freedom. If the calculated value is greater than the critical value, the difference is significant, otherwise it is insignificant and could have arisen due to fluctuations of sampling. The degrees of freedom are indicated by 1 and 1 and are (n11) and (n21) respectively. The F-test shall be discussed in detail in the section on Analysis of Variance (ANOVA).

MBA-015

Business Statistics

Page 113

2 TEST
In Z-test, t-test and F-test we make assumptions about the population values or parameters. Therefore these tests are called parametric tests. However, in many situations it is not possible to make dependable assumptions about the form of the parent distribution (that is, whether it is normal or not) from which the samples have been drawn. Z-, t- and F-tests are not applicable in such cases. To study the problem connected with such situations various tests have been evolved which do not need any assumptions about the parameters. Such tests are distribution-free tests and are called non-parametric tests. 2 (chi-square) test is a non-parametric test. Such non-parametric tests have assumed great importance in statistical analysis and statistical inference because they are easy to compute and can be used without making assumptions about parameters as they are distribution-free tests. These tests are not as reliable as parametric tests but where parameters cannot be rigidly assumed, such tests have to be used. 2-test is a test which describes the magnitude of difference between observed frequencies and the frequencies expected under certain assumptions. With the help of 2-test, it is possible to find out whether such differences are significant or are insignificant and could have arisen due to fluctuations of sampling. In 2-test the only problem is to decide how the expected frequencies have to be arrived. Once the expected values have been arrived at the calculation of 2 and its interpretation are very easy and involve the following steps: (i) Calculate the expected frequencies (denoted by E).

(ii) Find out the difference between observed frequencies (denoted by O) and expected frequencies (OE). (iii) Square up the various values of (OE) [that is, find out (OE)2] and divide each value of (OE)2 by the respective value of E [ (iv) Total all values of ].

and this gives the value of 2. That is

(v) Compare the calculated value of 2 with the independent value of 2 (from tables) for the given degrees of freedom and at the desired level of significance. (vi) If the calculated value of 2 is more than the relevant table value, the difference between observed and expected values is significant. If the observed value of 2 is less than the table

MBA-015

Business Statistics

Page 114

value, the differences between the observed and expected frequencies is not significant and could have arisen due to fluctuations of sampling. Certain Terms used in 2-Test Degrees of Freedom The term refers to the number of independent constraints in a set of data. Suppose there is a 2x2 association table and the actual frequencies of the various classes are as follows: A AB B B 60 22 38 A 40 8 32 30 70 100 Suppose that we presume that the two attributes A and B are independent then the expected frequency of the class AB would be or 18. Now, once we decide the expected frequency of the

class AB, the expected frequencies of the remaining three classes are automatically fixed. Thus, for the class B, expected frequency must be 6018 or 42. Similarly, for A the frequency will be 3018 or 12. For it will be 7042 or 28. It means that so far as this table is concerned, we have only one choice of our own and in the remaining three classes we have no freedom to fill in the frequencies as we like. It means that we have only one degree of freedom in this case. There is one independent constraint and three constraints are dependent. In such tables, the degrees of freedom are calculated by the formula

Where stands for the degrees of freedom, c for the number of columns and r for the number of rows. Thus in a 2x2 table, there will be (21)x(21) or 1 degree of freedom. Similarly, in a 2x3 table there will be (21)x(31), that is 2 degrees of freedom. If the data are not given in the shape of a contingency table as above but are given in the shape of a series of individual observations or discrete or continuous series, then the degrees of freedom are calculated in a different way. Take the following illustration relating to 1024 throws of ten coins each, which gives the following distribution: Number of Heads Actual Frequencies 0 2 1 10 2 38
MBA-015 Business Statistics Page 115

Number of Heads Actual Frequencies 3 106 4 188 5 257 6 226 7 128 8 59 9 7 10 3 Total 1024 In this illustration, if we write down the expected frequencies we have the freedom to write any ten figures we may choose, but the eleventh figure must be equal to the difference between 1024 and the total of the ten figures that we have written. Thus, there are 10 degrees of freedom in this case. In such cases, therefore, the degrees of freedom would be equal to n1, where n is the number of frequencies (or values in case of a series of independent observations. Thus, in the above illustration where the binomial distribution was applicable, we lost one degree of freedom, when a poisson distribution is fit, two degrees of freedom are lost, as in its calculation there are two constraints (total frequency and arithmetic average). In case of a normal distribution, three degrees of freedom are lost as there are three constraints, viz., total frequency, mean and standard deviation. Form of 2 Distribution The 2-distribution has only one parameter and it is or the degree of freedom. Like t-distribution, the 2-distribution also varies as the degrees of freedom change. If the degrees of freedom are low, the 2-distribution is highly skewed to the right. As the degrees of freedom increase the degree of skewness decreases, and if the degrees of freedom are substantial, normal distribution with a mean Constants of 2 Distribution The constants of a 2-distributionwith degrees of freedom are (i) Mean = ; Mode = 2; Variance = 2 assumes the shape of a

and standard deviation equal to unity.

(ii) Moments: 1 = 0; 2 = 2; 3 = 8; 4 = 48 + 122 (iii)

MBA-015

Business Statistics

Page 116

hence

hence

Application of 2 Distribution (i) 2-test for goodness of fit.

(ii) 2-test for association of attributes or for independence of attributes. (iii) 2-test to determine if the universe has a specified value of variance. Conditions for Validity of 2 Test The 2-test should be used only when (i) The total frequency is fairly large. Conventionally this test is not used if n < 50.

(ii) No expected frequency should be small (say, less than 5). If an expected frequency is less than 5, it should be pooled with the frequency of the neighbouring class. However, this reduces the degree of freedom. (iii) The distribution should not be of proportions or percentages, etc. It should be of original units. Note: If an expected frequency is less than 5 in case of a 2x2 table, it cannot be pooled with the frequency of the adjoining class. In such cases Yates Correction is used. In Yates correction, 0.5 is added to the observed frequency which is less than 5. The remaining frequencies are adjusted so as to keep the sub-totals and the grand total unchanged. For example A A B 2 10 12 B 2.5 9.5 12 would become 6 6 12 5.5 6.5 12 8 16 24 8 16 24 2 Test for Goodness of Fit if we have a set of observations through an experiment and we are interested in knowing whether these values are consistent with the values which may be obtained under some hypothesis, then the test is said to be test for goodness of fit. If the observed values are close to the hypothesized or

MBA-015

Business Statistics

Page 117

expected values, the fit is said to be good. If however, the difference between the two set of figures are found to be significant, the fit is not good. 2 can also be calculated by the formula

This formula is helpful if there are fractions in the deviations, and the application of the earlier formula becomes tedious. 2 Test for Independence of Attributes in testing association of attributes or their independence the expected values are calculated on the criterion of independence. Thus, the independence values of class . Similarly, other

expected values are calculated. The same principle which applies in a 2x2 table is applied for n x n and other forms of contingency table also. The rest of the procedure of calculating the value of 2 is the same as in case of testing the goodness of fit. The interpretation is also done in the same way. Alternate Method In 2x2 tables it is not necessary to calculate the expected values and the value of 2 can be obtained directly. This method was first used by Brandt and Snedecar. b1 B2 Tb c1 c2 T c T1 T2 T

Where b1, b2, c1, c2 are observed values. 2 for Population Variance 2 may also be used to ascertain whether the variance in the universe could be a specified figure. For this we find the value of 2 by the following rule:

Where n denoted the size of sample, s2 denotes the variance of the sample and variance of the universe.

denotes the

MBA-015

Business Statistics

Page 118

The Additive Property of 2 2 has a very useful property of addition. If a number of sample studies have been conducted in the same field then the results can be pooled together for obtaining an accurate idea about the real position. Suppose ten experiments have been conducted to test whether a particular vaccine is effective against a particular disease. Now here we shall have ten different values of 2 and ten different values of . We can add the ten values of 2 to obtain one value and similarly the ten values of can also be added together. Now we can test the results of all these ten experiments combined together. Precautions in the Use of 2 Test (i) Frequencies should not be small. Normally no frequency should be less than 10 but generally 2 test can be used if no frequency is less than 5. (ii) Frequency of non-occurrence should never be omitted. For example, if we know the number of cases which have been cured by, say, 4 different medicines, we should not use 2 test unless we also know the number of cases not cured by these medicines. (iii) 2 test should not be used if we do not have the original data. It is not used on proportions, percentages, or rates. (iv) If repeated measurements are made on the same units, 2 test should not be used. For example, if we have obtained the marks of 10 students before and after coaching, we should not use 2 test to find out if coaching has done any good. In such cases the difference of means should be tested. 2 test cannot be used because the table we will get in such cases would not be a contingency tables as the same persons would be counted twice. (v) Care should be taken that a. The degrees of freedom are correctly found out; b. The critical values of 2 are correctly used; c. The hypothesis is properly set and tested; d. The sum of observed and expected frequencies in various sub-totals and in the grand total are the same; and e. The expected values have been calculated on a rational basis.

MBA-015

Business Statistics

Page 119

ANALYSIS OF VARIANCE (ANOVA)


It was R. A. Fisher who introduced the term Variance in the analysis of statistical data. Fisher developed an elaborate technique of analysing the variance of two or more series for the purpose of studying their characteristic. Variance is a measure of the variability of the values of various items of a series from the mean. The technique of the analysis of variance (ANOVA) as developed by Fisher is capable of fruitful application in a variety of problems.

F-ratio is therefore the ratio of two independent 2 variates divided by their respective degrees of freedom. The F-test tells us whether the variances of the two universes from which the samples have been taken are equal. We set up a null hypothesis , that is, the population

variances are the same. In other words, the assumption we make is that the two independent estimates of the population variances do not differ significantly and we test this hypothesis by calculating F-ratio and comparing the calculated value of F with the critical value of F from the table. If the calculated value of F is more than the critical value for the given degrees of freedom and the given significance level, the null hypothesis is rejected and the inference is that the difference is significant. The samples in such cases do not come from the same universe or if they have come from two universes, the variance of these two universes are not the same. In ANOVA, an attempt is made to find out whether the means given by a number of samples are significantly different from one another. In case of just two samples, Z-test or t-test may be used for the purpose, but these tests cannot be used if the number of samples is more than two. The technique of ANOVA is very useful in such cases and with its help it is possible to study the significance of the difference of mean values of a large number of samples at the same time. For example, it can be used in a research where the effects of one or two variables has to be determined on the basis of a number of experiments conducted simultaneously. Variance analysis can tell us whether different sample data classified on the basis of one or two variables are meaningful or not. ANOVA studies the significance of the difference in means by analysing variance. The variances would differ only when the means are significantly different.

MBA-015

Business Statistics

Page 120

Assumptions in F-Test (i) The universe from which the samples are drawn follows the properties of the standard normal curve. (ii) The samples drawn form the universe are random and independent of each other. (iii) The variances of the populations from which the samples have been taken do not significantly differ from each other. We start with the null hypothesis, In real life problems, these assumptions may or may not hold good. However, unless the universes are highly skewed, minor variations in the assumptions do not affect the validity of the F-test. If the means of the samples are not significantly different form the mean of all the samples taken together, then it can be safely assumed that they have come from the same or similar population. Technique of Analysing Variance The technique of ANOVA in case of single variable and in case of two variables is similar. In both cases a comparison is made between the variance of sample means with the residual variance. However, in case of a single variable (one-way classification), the total variance is divided into two parts: (i) Variance between the samples; and (ii) Variance within the samples. The latter variance is the residual variance. In case of two variables (two-way classification) the total variance is divided into three parts: (i) Variance due to the first variable;

(ii) Variance due to the second variable; and (iii) Residual variance. Variance Analysis in One-Way Classification If there are, say, four samples, the null hypothesis would be that their means are not different, that is . The alternate hypothesis shall be . With these

hypotheses, we perform the following steps for ANOVA in one-way classification: (i) Calculate the sum of the squares of variation between samples.

(ii) Calculate the sum of the squares of variation within the samples. (iii) Calculate the total sum of the squares of variation. This will be the sum of (i) and (ii) and should be calculated directly also for verification of calculations. (iv) Calculate the F-ratio.

MBA-015

Business Statistics

Page 121

(v) Compare the calculated F-ratio with the critical value of F-ratio from the table. (vi) Draw inference whether the null hypothesis is accepted or rejected. These steps are discussed below: (i) Variation between Samples. The variation between sample means can be either on account of difference in variable or due to chance. On the other hand, the difference between the values of items in any single sample can be on account of chance only as the variable remains uniform throughout the sample. The results of different samples are shown in columns. The sum of squares of variation between the samples is called SSC (sum of squares in column). If the number of items is equal in all samples, it is calculated as follows: a. Find out the mean of each sample ( ). .

b. Find the mean of the sample means or the grand mean

c. Find the deviation of sample means from the grand mean; square these deviations and multiply by the number of items in samples; find the sum of these figures. This is SSC. d. Divide the sum of squares of deviations from previous step by the degrees of freedom, given by (c1), where c denotes the number of samples. This is denoted as MSC (mean square of columns). (ii) Variation within the Samples. Any variation within a sample could be on account of chance only. The difference in the values of various items within a sample, which is due to chance, is called an estimate of the error. It is denoted by SSE (sum of the squares of the deviations from the mean of the series). It is calculated as follows: a. Find the mean of the sample. b. Find the deviation of various items of the sample from the mean value (of the sample). Square these deviations and obtain their total. c. Repeat this process for all samples, and total the sum of the squares of the deviation of the various samples from their respective means. This would give the value of SSE. d. Divide SSE by the degrees of freedom which would be or , where is the total number of items in all samples and is the number of

samples. It could be also calculated as

, as the number of columns would be equal to

the number of samples. This would be the variance within the sample (or variance due to chance). This is denoted by MSE (means square of the error). (iii) Total sum of Squares. This is the total of SSR and SSE. In other words, it is found by adding the sum of squares of deviations between the samples and the sum of squares of deviations

MBA-015

Business Statistics

Page 122

within the samples. However, for the sake of verification of calculations, it should be calculated as follows: a. Find the grand mean. b. Find the deviations of all the items of all the samples from the grand mean. Square these deviations and obtain the total of these squares (SST, that is the total sum of squares). c. Divide SST by the degrees of freedom ( ). This is the total variance.

(iv) Calculation of Variance Ratio (F-Ratio), comparison with tabulated value and inference. F-ratio (variance ratio) is the ratio between greater variance and smaller variance. Generally variance between the samples is more than the variance within the samples and

If, however, the variance within the samples is more than the variance between the samples, the numerator and denominator should be interchanged and the degrees of freedom adjusted accordingly. The calculated F-ratio is compared with the critical value of F to draw inferences. If the calculated value of F is more than the critical value, the null hypothesis is rejected and the differences between means is said to be significant. It means that we accept the alternate hypothesis that .

The above points for analysing variance in one-way classification can be analysed in the following ANOVA table: Analysis of Variance (ANOVA) Table (One-Way Classification) Source of Variation Sum of Squares Degrees of Freedom Mean Squares Between Samples SSC C1 Within Samples Total Where SSC = Sum of squares between samples SSE = Sum of squares within samples SST = Total sum of squares MSC = Mean square between samples MSE = Mean square within samples F = Ratio of MSE to MSC.
MBA-015 Business Statistics Page 123

F-Ratio

SSE SST

NC n1

Coding of Data (Change of Scale or Origin) The F ratio has a very important property that its value remains unchanged if all the figures are multiplied or divided by a common factor or if a common factor is added to or subtracted from each figure. The above property is of very great help in calculating the F ratio when the figures are large, and the calculations become tedious. In such cases the figures may be divided or multiplied to get a convenient set of figures. This property may also be used in case of fractional figures. When such simplification is done, it is said that the data have been coded to simplify calculations. The value of F remains unchanged when coded data are used for calculations. Variance Analysis in Two-Way Classification Here we study the effect of two variables. The procedure of analysis in two-way classification is: total columns as well as rows. The effect of one factor is studied through the column-wise figures and totals and the effect of the other factor is studied through the row-wise figures and totals. The variances are calculated for both the columns and rows and they are compared with the residual variance or error. The table of variance analysis in a two-way classification takes the following form: ANOVA Table (Two-Way Classification) Source of Variation Sum of Squares Degrees of Freedom Mean Square (i) (ii) (iii) (iv) Between Columns SSC (c1) Between Rows Residual or Error Total Where SSC = Sum of squares between columns SSR = Sum of squares between rows SSE = Sum of squares due to errors SST = Total sum of squares. The SSE or residual sum of squares is found out by subtracting the total of SSC and SSR from the total sum of squares. In a two-way classification one should be careful in finding out the degrees of freedom. For column figures it is (c1), for row it is (r-1) and for the residual it is (c1)(r1). Therefore, while calculating SSR SSE SST (r1) (c1)(r1) n1 (r1)

Variance Ratio (v)

MBA-015

Business Statistics

Page 124

F-ratio for columns and for rows the degrees of freedom for the numerator may not be the same. In the first case it is (c1) and in the second case it is (r1). However, the degrees of freedom for the denominator in both cases would be the same, that is, (c1)(r1). Assumption in Analysis of Variance of Two Variables When we analyse the variance in a two-way classification, we assume that the factor effects are additive so that the values are affected by a constant amount on account of each variable. But this may not necessarily be true. There may be interaction between the two factors. One factor might be affecting the other in some way. However, in the analysis of variance in a two-way classification, this effect is ignored, and the inference drawn by variance analysis is subject to this limitation.

MBA-015

Business Statistics

Page 125

You might also like