0% found this document useful (0 votes)
13 views86 pages

Fisheries Statistics: Concepts & Applications

Uploaded by

sushamasethu293
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
13 views86 pages

Fisheries Statistics: Concepts & Applications

Uploaded by

sushamasethu293
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

STATISTICS

SYLLABUS - THEORY

 Definition of statistics, scope of fisheries statistics.


 Basic concepts of population and sample, random sampling.
 Collection of data; census enumeration and sample surveys, their advantages and
disadvantages, preparation of schedules and questionnaires;
 Classification of data, frequency and cumulative frequency table.
 Diagrammatic and graphical representation of data - bar diagrams, pie-diagram,
histogram, frequency polygon, frequency curve and Ogive;
 Important measures of central tendency - arithmetic mean, media ;and mode, relative
merits and demerits of these measures.
 Important measures of dispersion -range, mean deviation, variance and standard
deviation, relative merits and demerits of these measures; relative measures of
dispersion - coefficient of variation;
 Measures of skewness and kurtosis.
Unit 1:Definition of statistics

Learning objective
The reader is going to be enlighted about statistics and its uses
Fisheries statistics is defined in this unit. The usage of statistics in the field of fisheries are
explained.
Chapter 1: Definition of statistics
1.1 Definition of statistics

Statistics may be defined as the science of collection, organization, presentation, analysis and
interpretation of numerical data.

According to the above definition, there are five stages in a statistical investigation.
[Link] Collection

Collection of data constitutes the first step in a statistical investigation. Utmost care must be
exercised in collecting data because they form the foundation of statistical analysis. If data are
faulty, the conclusions drawn can never be reliable. The data may be available from existing
published or unpublished sources or else may be collected by the investigator himself. The
firsthand collection of data is one of the most difficult and important tasks faced by a
statistician. Therefore, like all scientific pursuits, the investigator must take into account
whatever data have already been collected by others. This would save the investigator from
foreseeable pitfalls, unnecessary labour and duplication of efforts.
[Link] Organization

1
Data collected from published sources are generally in organized form. However, a large mass
of figures that the collected from a survey frequently needs organization. The first step in
organizing a mass of data is editing. The collected data must be edited very carefully so that the
omissions, inconsistencies, irrelevant answers and wrong computation in the returns from a
survey may be corrected or adjusted. After the data have been edited the next step is to classify
them. Classification is the process of arranging the data according to some common
characteristics possessed by the items constituting the data. The last step in organization is
tabulation. The object of tabulation is to arrange the data in columns and rows so that there is
absolute clarity in the data presented.
[Link] Presentation

After the data have been collected and organized they are ready for presentation. Data
presented in an orderly manner facilitates statistical analysis.
[Link] Analysis

After collection, organization and presentation the next step is that of analysis. A major part of
this text is devoted to the methods used in analyzing the presented data are numerous ranging
from simple observation of data to complicated, sophisticated and highly mathematical
techniques: However, in this text only the most commonly used methods of statistical analysis
are included.
Chapter 2: Functions of statistics
1.2.1 Functions of Statistics

It presents facts in a definite form.

It simplifies mass of figures.

It facilitates comparison.

It helps in formulating and testing hypothesis.

It helps in prediction.

It helps in the formulation of suitable policies.


1.2.2 Definiteness

Numerical expressions are convincing and, therefore, one of the most important functions of
statistics is to present general statements in a precise and definite form. Statements or facts
conveyed in exact quantitative terms are always more convincing than vague utterances.
Statistics present facts in a precise and definite form and thus help proper comprehension of
what is stated. Consider, for example, a statement sex ratio (i.e., number of females per 1000
males) is going down in India. The reader would not have a clear idea of the situation from this
statement. But if we say the sex ratio has gone down from 934 in 2000 to 929 in 2010, it

2
conveys a definite meaning. Similarly, statements like ‘there is a lot of unemployment in India’,
‘the population of India is growing at a very fast rate’, ‘the prices of various commodities are
rising’, ’the number of students seeking admission to professional courses is increasing’, etc.,
hardly convey any worthwhile information as they do not specify the numerical dimensions
involved.
1.2.3 Condensation

Not only does Statistics present facts in a definite form but it also helps in condensing mass of
data into a few significant figures. In a way, statistical methods present a meaningful overall
information from the mass of data. Thus, it is impossible for one to form a precise idea about
the income position of the people of India form a record of individual incomes of the entire
population. However, the figure of per capital income can be easily remembered by every one.
1.2.4 Comparison

Unless figures are compared with others of the same kind they are often devoid of any
meaning. For example, if we say that the production of Maruti Udyog Ltd. has increased
considerably shall not be meaningful unless some comparison of figures is made. But the
statement there has been an increase from 200 cars a day in Sept. 1985 to 1000 cars a day in
April 2000 definitely indicates the increasing trend in production.
1.2.5 Formulating and Testing Hypothesis

Statistical methods are extremely useful in formulating and testing hypothesis and to develop
new theories. For example, hypothesis like whether chloromycetin is effective in preventing
typhoid, whether the credit squeeze is effective in checking price increase, whether students
have benefited from the extra coaching, etc., can be tested by appropriate statistical tools.
1.2.6 Prediction

Plans and policies of organizations are invariable formulated well in advance of the time of their
implementation. A knowledge of future trends is very helpful in framing suitable policies and
plans. Statistical methods provide helpful means of forecasting future events. For example, if a
businessman has to decide how much he should produce in the next year, he would like to
know the expected sales for that year. He may use his subjective judgement and make a guess.
However, a better method for him would be to analyse the sales data of the past years or
arrange a statistical survey of the market to obtain necessary data for estimating the sales
volume for the next year.
1.2.7 Formulation of policies

Statistics provide the basic material for framing sutiable policies. For example, it may be
necessary to decide how much oil a nation should import in the next year, the decision would
depend upon the expected internal production and the likely demand for oil in the next year. In
the absence of information regarding the estimated domestic output and demand for oil the
decision on imports cannot be made with reasonable accuracy.
Chapter 3: Fisheries statistics

3
1.3.1 Fisheries statistics

‘Biostatistics’ or ‘biometry’ is the science of statistics as applied to quantitative study of the


biological phenomena. Fisheries statistics relates to the branch of biostatistics applied to the
study of fish and fisheries as well as to the study of socio-economic aspects of fisheries as a
resource wealth utilized by man for avocation and food. It also encompasses fisheries data.
Chapter 4: Scope of fisheries statistics
1.4.1 Inventory of potential resources

Estimation of total water area available for exploitation, area actually exploited, manpower
employed in fishing and allied activities, the type of craft and gear employed for fishing etc.
1.4.2. Production

Estimation of inland and marine fish landings disaggregated according to mechanized and non-
mechanised crafts, species, sizes etc. It also encompasses fish production through aquaculture,
fish feed production through hatcheries etc.
1.4.3 Fish stock assessment

Growth of fish populations, their size (length of weight) and age structure, natality, recruitment
and mortality, estimation of stock, optimum yield etc.
1.4.4 Morphometric and meristic analysis

Measurement of various body proportions such as total length, standard length, fork length,
head length etc. for the purpose of statistical comparison with similar measurements for a sub
species or a closely related species and establishing the levels of significance at which
differences occur. In other words it serves the purpose of establishing the variations and
relationships between different quantitative morphological characters of two or more closely
related species.
Counts of spines and rays of fins, scales, vertebrae etc., constitute meristic characteristics of
fish. These characteristics form one of the important taxonomic tools for differentiating closely
related species
1.4.5 Designing experiments for quantitative inferences

Designing field and laboratory experiments to quantify various biotic and abiotic phenomena in
aquatic environment and interpret casual relationships in quantitative terms of these variables
on fish behavior, growth, production, survival, spawning, etc.
1.4.6 Genetic studies

Study of various fish characters as regards to their heritable and non-heritable properties and
patterns in different generations, efficiency of different selection procedures for improving fish
stocks etc.
1.4.7 Quality control

4
Checking the quality of frozen fish and fish products and ascertaining whether the product is
conforming to specifications or not, using inspection plans, control charts etc.
1.4.8 Market research and business

Estimation of cost of production, price and price spread, estimation of supply and demand,
consumption and distribution pattern, income, its distribution, investment and returns,
production and inventory control, managerial decision making etc.
Chapter 5: Sources of fisheries data
1.5.1 Handbook on fisheries statistics

National fisheries data are periodically released by the Government of India, Fisheries Division,
Ministy of Agriculture and Cooperation as an official document for Central Board of Fisheries
meeting. It contains information on production, exports, fishing harbours, training in fisheries,
outlays and expenditure, prices, fishing resources and other aspects of fisheries.
1.5.2 Statistics of marine products exports

It is published annually by the Marine Products Export Development Authority, Kochi. It


contains information on country-wise exports, region-wise exports, item-wise exports, average
unit value, world markets, prices, marine fish landings etc.
1.5.3 FAO yearbook of fishery statistics

The yearbook of fishery statistics is published annually by the Food and Agricultural
Organisation (FAO) of the United Nations in different volumes. The volume on ‘capture
production’ provides annual data for countries or areas by species groups, major fishing areas,
continents etc. The volume on ‘aquaculture production’ contains data for countries or area by
culture environment/species groups. The volume on ‘commodities’ includes disposition of
world fishery production, estimated total international trade in fishery commodities, imports,
exports, etc.
Chapter 6: Computer and fisheries statistics
1.6.1 Computer and fisheries statistics

The wide spread use of computers has had a tremendous impact on statistical analysis of
fisheries data. Computers can perform calculations much faster and accurately. The use of
computers makes it possible to devote more time for improvement of quality of data collected
and interpretation of results. Many number of computer software packages are available for
carrying out most of the descriptive and inferential statistical analysis. Some of the statistical
packages that are widely used include Statistical Packages for the Social Sciences (SPSS), SAS
and SYSTAT. These packages differ with respect to their input requirements, their output
formats and the specific calculations they perform. Therefore, before using a particular package
or a program, input requirements should be studied carefully and program’s output format to
be studied for proper interpretation of results. Spread sheet software Microsoft Excel is also
used for conducting statistical analysis using Data analysis / Analysis Tool pak and Tool bar
function fx.

5
Length based Fish Stock Assessment (LFSA) and Electronic Length Frequency Analysis (ELEFAN)
packages developed during eighties by Food and Agricultural Organisation (FAO) and
International Centre for Living Aquatic Resources Management (ICLARM) respectively have
been very popular in fish stock assessment studies. In nineties FAO and ICLARM joined together
and developed a new program package for length based fish stock assessment called FISAT
(FAO – ICLARM Stock Assessment Tools) incorporating the positive aspects of the parent
packages LFSA and ELEFAN. Other important fish stock assessment programmes produced by
FAO include BEAM1, 2, 3 and 4 for bio-economic modeling and CLIMPROD dealing with surplus
production models incorporating environmental variables. The computer software packages
mentioned here are only indicative and not exhaustive.
Unit 2: Collection of data

Learnin objective
While going for sampling, the preparation of questionnaire is explained. The already available
since of information are also covered.
2.1.1 Introduction

The investigator is faced with one of the most difficult problems of obtaining or gathering the
desired information or data. Utmost care must be exercised while collecting data because they
constitute the foundation on which the superstructure of statistical analysis is built. The results
obtained from the analysis are properly interpreted and policy decisions are taken. Hence, if
the data are inaccurate and inadequate the whole analysis may be faulty and the decisions
taken misleading.

Data may be obtained either from the primary source or the secondary source. A primary
source is one that itself collects the data; a secondary source is one that makes available data
which were collected by some other agency. For example, the data collected by the Ministry of
Industries and made available through various publications constitute primary source.
However, if the Ministry of Industries uses data collected by some other organization, say,
National Sample Survery Organisation, this will constitute secondary source for the Ministry. A
primary source usually has more detailed information particularly on the procedures followed
in collecting and compiling the data. It may be noted that a given source may be partly primary
and partly secondary. The journal ‘Agricultural Situation in India’ gives data compiled by the
Ministry of Agriculture and Irrigation and may also contain related information collected by
other Ministries.

It is preferable to make use of the primary source wherever possible for the following reasons:

(i) The secondary source may contain mistakes due to errors in transcription made when the
figures were copied from the primary source

(ii) The primary source frequently includes definitions of terms and units used.

6
(iii) The primary source often includes a copy of the schedule and a description of the
procedure used in selecting the sample and in collecting the data.

(iv) Primary source usually shows data in greater detail.

Depending on the source, statistical data are classified under two categories:

(i) Primary data, and

(ii) Secondary data.


2.1.2 Primary data

Primary data are obtained by a study specifically designed to fulfill needs of the problem at
hand. Such data are original in character and are generated in large number of surveys
conducted mostly by Government and also by some individuals, institutions and research
bodies. For example, data obtained in a population census by the Office of the Registrar
General and Census Commissioner, Ministry of Home Affairs, are primary data.
2.1.3 Secondary Data

Data which are not originally collected but rather obtained from published or unpublished
sources are known as secondary data. For example, for the Office of the Registrar General and
Census Commissioner the census data are primary whereas for all others, who use such data,
they are secondary. The secondary data constitute the chief material on the basis of which
statistical work is carried out in many investigations.
2.1.4 Points for Primary data

In fact, before collecting primary data it is desirable that one should go through the existing
literature and learn what is already known of the general area in which the specific problem
falls and all surrounding information that may give us leads and lessons. This can help in getting
an idea about the possible pitfalls, avoiding duplication of efforts and waste of resources. It
should be noted that it is the process of assembling primary data which is called ‘collection’ of
statistics and is different from the process of ‘compiling’ statistics (i.e secondary data) from
various published sources. To quote Crum, Patton and Tebbutt, ‘Collection means the
assembling, for the purpose of particular investigation of entirely new data, presumably not
already available in published sources’. We have used the term ‘collection’ in this book strictly
in the narrow sense defined above.
2.1.5 Difference between primary and secondary data

The difference between primary and secondary data is only of degree – data which are primary
in the hands of one become secondary in the hands of another. Data are primary for the
individual agency or institution collecting them whereas for the rest of the world they are
secondary. A few examples would clarify the distinction. Suppose an investigator wants data
about the spending habits of the students of Delhi University. If he collects the data himself or

7
through his agents adopting any suitable method such as contacting and interviewing students
or circulating a questionnaire, the data would constitute primary data for him. On the other
hand, if the students union has already made a similar survey and the investigator obtained
data from union office, such data would constitute secondary data for him. Similarly, statistics
collected by various departments of the Government such as Labour Bureau and Central
Statistical Orgainisation are primary for the respective departments whereas for all others
they are secondary data.
2.1.6 Advantages of secondary data

Secondary data offers the following advantages:

(1) It is highly convenient to use information which someone else has compiled. There is no
need for printing data collection forms, hiring enumerators, editing and tabulating the results,
etc. Researchers alone or with some clerical assistance may obtain information from published
records compiled by somebody else.

(2) If secondary data are available they are much quicker to obtain than primary data.

(3) Secondary data may be available on some subjects where it would be impossible to collect
primary data. For example, census data cannot be collected by an individual or research
organization, but can only be obtained from Government publications.
2.1.7 Disadvantage of secondary data

However, two major problems are encountered in using secondary data:

(1) The first is the difficulty of finding data which exactly fit is to the need of the present study.

(2) The second problem is finding data which are sufficiently accurate.
2.1.8 Choice between Primary and Secondary data

The investigator must decide at the outset whether he will use primary data or secondary data
in an investigation. The choice between the two depends mainly on the following
considerations:

(i) Nature and scope of the enquiry,

(ii) Availability of financial resources,

(iii) Availability of time,

(iv) Degree of accuracy desired, and

(v) The collecting agency, i.e., whether an individual, an institution or a Government body.

8
It may be pointed out that most statistical analysis rests upon secondary data. Primary data are
generally used in those cases where the secondary data do not provide an adequate basis for
analysis. In certain cases, both primary as well as secondary data may be employed. The reason
why secondary data are being increasingly used is that published statistics are now available
covering diverse fields so that an investigator finds required data readily available to him in
many cases.
2.1.9 Methods of collecting primary data

Primary data may be obtained by applying any of the following methods.

I. Direct personal interviews.

II. Indirect oral interviews.

III. Information from correspondentes.

IV. Mailed questionnaire method.

V. Schedules sent through enumerators.

These methods are discussed below:


[Link] Direct Personal Interviews

Under this method of collecting of data, there is a face-to-face contact with the persons from
whom the information is to be obtained (known as informants). The interviewer asks them
questions pertaining to the survey and collects the desired information. Thus, if a person wants
to collect data about the working conditions of the workers of a fish processing Industry, he
would go to the Industry, contact the workers and obtain the desired information. The
information thus obtained is first hand or original in character.

Merits . The advantages of personal interview are:

1. Response is more encouraging as most people are willing to supply information when
approached personally.

2. The information obtained by this method is likely to be more accurate because the
interviewer can clarify the doubts of the informants about certain questions and thus obtain
correct information. In case the interviewer apprehends that the informant is not giving
accurate information, he may cross-examine him and thereby try to obtain the information.

3. It is also possible through personal interview to collect supplementary information about the
informant’s personal characteristics and environment and such information often proves very
useful while interpreting results.

9
4. Some questions about which the informant may likely to be sensitive which can be carefully
combined with other questions by the interviewer. He can twist the questions keeping in mind
the informant’s reactions. He can change the subject, if necessary, or explain the survey
problem further if it appears that the informant is not inclined to supply any information. In
other words, a delicate situation can usually be handled more effectively by a personal
interview than by other survey techniques.

5. The language of communication can be adjusted to the status and educational level of the
person interviewed, thus inconvenience and misinterpretation on the part of the informant can
be avoided.

Limitations. Important limitations of the personal interview method are:

1. It may be very costly where the number of persons to be interviewed is large and they are
spread over a wide area.

2. The chances of personal prejudice and bias are greater under this method as compared to
other methods.

3. The interviewers have to be thoroughly trained and supervised, otherwise they may not be
able to obtain the desired information. Untrained or poorly trained people may spoil the entire
work.

4. More time is required for collecting information by this method as compared to others. This
is because interviews can be held only at the convenience of the informants. Thus, if
information is required to be obtained from the working members of households, interviews
will have to be held in the evening or on weekend. Since only an hour or two can be used for
interviews in the evening, the work may have to be continued for a long time, or a large staff
may have to be employed involving huge expenditure.

Suitability . This method is suitable for intensive rather than extensive field surveys. Hence, it
should be used only in those cases where intensive study of a limited field is desired.

It may be noted that when personal interview method is adopted, the investigator instead of
going personally and conducting a face-to-face interview may also obtain information on
telephone. For example, the television viewers may be asked to comment on certain
programmes on phone. The method is less expensive. However, this method suffers from some
serious defects like: (i) not every one owns a phone and hence only a very limited group can be
approached by this method, (ii) very few questions can be asked on phone, (iii) since telephone
interview has to be conducted very quickly, the respondents may give vague and reckless
answers, and (iv) there may be serious errors of communication on telephone.

Because of these reasons, telephone interviews are not very commonly used.
[Link] Indirect Oral Interviews

10
Under this method of collecting data, the investigator contacts third parties called witnesses
capable of supplying the necessary information. The method is generally adopted in those cases
where the information to be obtained is of a complex nature and the informants are not
inclined to respond if approached directly. For example, in an enquiry regarding addiction to
drugs, alcohol, etc., people may be reluctant to supply information about their own habits. It
would be necessary in that case to get the desired information from those dealing in drugs,
liquor or other people who may be knowing them, for example, their neighbours, friends, etc.
Similarly, if a fire has broken out at a certain place the cause of the fire may be traced by
contacting persons living in the neighbourhood of that area. In a similar manner, clues about
thefts or murders are obtained by the police by interrogating third parties who are supposed to
have knowledge about the case under investigation. Enquiry Committees and Commissions
appointed by the Government generally adopt this method to get people’s views and all
possible details of facts relating to the enquiry.

This method is very popular in practice. However, the correctness of information obtained
depends upon a number of factors, such as:

1. The type of persons whose evidence is being recorded. If the people do not know the full
facts of the problem under investigation or if they are prejudiced it will not be possible to arrive
at correct conclusions.

2. The ability of the interviewers to draw out the information from witness by means of
appropriate questions and cross-examination.

3. The honesty of interviewers who are collecting the information. It might happen that
because of bribery, nepotism or certain other reasons those who are collecting the information
give it such a twist that correct conclusions are not arrived at.

For the success of this method it is necessary that the evidence of one person alone is not relied
upon; the views of a number of persons should be ascertained to find the real position. Utmost
care must be exercised in the selection of these persons because it is on their views that the
final conclusions are reached.

Suitability. This method is suitable in such cases where indirect sources of information are
required to be tapped either because direct sources do not exist or cannot be relied upon or
would be reluctant to part with the information.
[Link] Information from Correspondents

Under this method, the investigator appoints local agents or correspondents in different places
to collect information.
These correspondents collect and transmit the information to the central office where the data
are processed. Newspaper agencies generally adopt this method. Correspondents in different
places supply information relating to such events as accidents, riots, strikes, etc., to the head
office. The correspondents may be paid or honorary persons but generally they are paid. This

11
method is also adopted by various departments of government in such cases where regular
information is to be collected from a wide area. For example, in the construction of wholesale
price index numbers regular information is obtained from correspondents appointed in
different areas. This method is particularly suitable in case of crop estimates. The special
advantage of this method is that it is cheap and appropriate for extensive investigation.
However, it may not always ensure accurate results because of the personal prejudice and bias
of the correspondents.

Suitability. As stated above, this method is generally adopted in those cases where the
information is to be obtained at regular intervals from a wide area.
[Link] Mailed Questionnaire Method

Under this method, a set of questions pertaining to the survey (known as questionnaire) is
prepared and is sent to the informants by post. The questionnaire contains questions and
provides space for recording answers. A request is made to the informants through a covering
letter to fill up the questionnaire and send it back within a specified time.

The questionnaire studies can be classified on the basis of:

(i) The degree to which the questionnaire is formalized/structured,

(ii) This disguise or lack of disguise of the questionnaire, and

(iii) The communication method used.

When no formal questionnaire is in use, interviewer adapt their questioning to each interview
as it progresses or perhaps elicit responses by indirect methods such as showing pictures on
which the respondent comments. When a prescribed sequence of question is followed, it is
referred to as structured study. On the other hand, when no prescribed sequence of questions
exists, the study is non-structured.

When questionnaires are constructed so that the objective is clear to the respondents, they are
non-disguised; on the other hand, when the objective is not clear, the questionnaire is a
disguised one. Using these two bases of classification, four types of studies can be
distinguished:

1. Non-disguised structured

2. Non-disguised non-structured

3. Disguised structured, and

4. Disguised non-structured.

12
Merits.

1. This method of collecting data can be easily adopted where the field of investigation is very
vast and the informants are spread over a wide geographical area.

2. It is also relatively cheap and expeditious provided the informants respond in time.

3. On questions of personal nature or questions requiring reaction by the family, this method is
generally superior to either personal interviews or telephone method.

Limitations.

1. This method can be adopted only where the informants are literate people so that they can
understand written questions and send the answers in writing.

2. It involves some uncertainty about the response. Co-operation on the part of informants may
be difficult to presume.

3. The information supplied by the informants may not be correct and it may be difficult to
verify the accuracy.

The success of this method depends upon the skill with which the questionnaire is drafted and
the extent to which willing co-operation of the informants is secured. Since the advantages of
the personal contact are lost in the mailed questionnaire, the form and tone of the
questionnaire must be designed to supply as far as possible the missing personal element.
Where the information is required by a government department, it is generally available on
account of legal or administrative sanctions. In other cases, it is necessary to take informants
into confidence so that they furnish correct information.

To make this method work effectively the following suggestions are made:

1. The questionnaire should be so framed that it does not become as undue burden on the
respondents, otherwise they may not return them back.

2. Prepaid postage stamp should be affixed.

3. The sample should be large.

4. It should be adopted in such enquiries where it is expected that the respondents would
return the questionnaire because of their own interest in the enquiry.

5. Its use should be preferred in such enquiries where there could be a legal compulsion to
supply the information so that the risk of non-response is eliminated.

13
Suitability.

This method is appropriate in cases where informants are spread over a wide area, i.e., in case
of extensive surveys.
[Link] Schedules sent through Enumerators

Yet another method of collecting information is that of sending schedules through the
enumerators or interviewers. The enumerators contact the informants, get replies to the
questions contained in a schedule and fill them in their own handwriting in the questionnaire
form. The essential difference between the mailed questionnaire method and this method is
that the questionnaire is sent to the informants by post in the former, in the latter the
enumerators carry the schedule personally to the informants. The method is free from most of
the Limitations of the mailed questionnaire method.

Merits. The main advantages of the method are:

1. It can be adopted in those cases where informants are illiterate.

2. There can be very little chance for non-response as the enumerators go personally to obtain
the information.

3. The information received is more reliable as the accuracy of statements can be checked by
supplementary questions wherever necessary.

Limitations. 1. Among various methods of collecting primary data, this method is quite costly as
enumerators are generally paid persons.

2. The success of the method depends largely upon the training imparted to the enumerators.

3. Skilled interviewing requires experience and training, but there is a tendency for statisticians
to neglect this extremely important part of the data collecting process. Without good
interviewing most of the information collected is of doubtful value.

4. The way in which the enumerators conduct the interview would affect the data collected.
When questions are asked by a number of different interviewers, it is possible that variations in
the personalities of the interviewers will cause variation in the answers obtained. This variation
will not be obvious. Hence every effort must be made to remove as much of variation as
possible due to different interviewers.

Suitability . This method is quite popularly used in practice. The main reason for this is a very
high rate of response because of the personal contact of the enumerators.
2.1.10 Designing the Questionnaire

14
Before framing the questionnaire it is essential to set out in detail the data which we desire
from the answers to questionnaire. It shall be wise if we can construct the type of tables which
we would like to obtain from the enquiry. It may not always be possible to set out all the
possible data we would like, in advance, since many things may be learnt in the course of
enquiry and one may find that what he believed to be ideal was not in fact ideal. For this reason
those who are likely to be concerned with analyzing the results should be called in at the very
early stage. For example, it may not be very appropriate for a government statistician to collect
some data on, say, unemployment, and then hand them over to an economist to analyse. The
wise thing would have been to consult the economist first on what data were desirable.

The success of the questionnaire method of collecting information depends largely on the
proper drafting of the questionnaire. Drafting questionnaire is a highly specialized job and
requires a great deal of skill and experience. It is difficult to lay down any hard and fast rules to
be followed in this connection. However, the following general principles may be helpful in
framing a questionnaire:

1. Covering letter. The person conducting the survey must introduce himself and state the
objective of the survey. It is desirable that –

(i) A short letter is enclosed. The letter should state in as few a words as possible the purpose of
the survey and how the informant would tend to benefit from it.

(ii) Enclose a self-addressed stamped envelope for the respondent’s convenience in returning
the questionnaire.

(iii) Assure the respondent that his answers will be kept in strictest confidence.

(iv) Promise the respondent that he will not be solicited after he fills up the questionnaire.

(v) If possible, offer special inducements (free gifts, concession coupons, etc.) to return the
questionnaire.

(vi) If the respondent is interested, promise a copy of the results of the survey to him.

2. Number of questions should be small. The number of questions should be kept to the
minimum. The precise number of questions to be included would naturally depend on the
object and scope of the investigation. Fifteen to twenty-five may be regarded as a fair number.
If a lengthy questionnaire is unavoidable, it should preferably be divided into two or more
parts.

It should be noted that there is an inverse relationship between the length of a questionnaire
and the rate of response to the survey. That is, the longer the questionnaire, the lower will be
the rate of response, the shorter the questionnaire, the higher will be the rate of response.
Therefore, each question must be clearly presented in as few a words as possible and each

15
question should be deemed essential to the survey. In addition questions must be free from
ambiguities.

3. Questions should be arranged logically. The questions must be arranged in a logical order so
that a natural and spontaneous reply to each is induced. They should not skip back and forth
from one topic to another. Thus, it is undesirable to ask a man how many children he has
before asking whether he is married or not. Similarly, it would be illogical to ask a man his
income before asking him whether he is employed or not. Thus the sequence of the questions
should be considered carefully in terms of the purpose of the study and the persons who will
supply the information. Questions applying identification and description of the respondent
should come first followed by major information questions. If opinions are requested, such
questions should usually be placed at the end of the list. Two different questions worded
differently be included on the same subject to provide cross-check on important points.

4. Questions should be short and simple to understand. Unless the person being interviewed is
technically trained, technical terms should be avoided. Words such as “capital” or “income”
that have different meanings for different persons should not be used unless a clarification is
included in the question.

5. Ambiguous questions ought to be avoided. ‘Ambiguous questions’ means different things to


different people. It will not be possible to obtain comparable replies from correspondents who
take a question to mean different things. For example,

Consider the following question:

Do you smoke? Yes / No

There are several ambiguities in this question. It is not clear whether the desired response
pertains to cigars, cigarattes, pipes or combinations thereof. Also it is not clear whether
occasional smoking or habitual smoking was the primary concern of the question. If we are
interested only in current cigarette consumption, it would be better to ask:

How many cigarettes do you currently smoke each day?

Less than 5

5 to 9

10 to 14

15 to 19

20 and above

16
6. Personal questions should be avoided. As far as possible, questions of a personal and
peculiar nature should not be asked. For example, questions about income, Sales-tax paid, etc.,
may not be willingly answered in writing. Where such information is essential, it should be
obtained by personal interviews. Even then, such questions should be asked only at the end of
the interview, when the informants feel more at ease with the interviewer.

7. Instructions to the informants. The questionnaire should provide necessary instructions to


the informants. For example, the questionnaire should specify the time within which it should
be sent back and the address at which it should be sent. Instructions about units of
measurements, etc., should also be given. For instance, if there is a question on weight, it
should be specified as to whether weight is to be expressed in pounds or kilograms or in some
other units.

8. Objective type Questions. Avoid questions of opinion and keep to questions of fact. In
factual studies, it is highly desirable that questions are so designed that objective answer may
be forthcoming. For example, instead of asking the condition of a building, allow the informant
or enumerator to state the condition in his own words. It is desirable to ask if a structure was in
good condition, needed minor repairs, needed structural repairs or was unfit for use. No doubt,
answer to such questions may not be completely objective but they can be readily tabulated.
Similarly, while asking students how do they normaly travel to college, frame a question of the
type:

How do you normally travel to college?

i) By bus (ii) By your own car (iii) By your own scooter

(iv) By taxi (v) On foot (vi) Any other

The respondent will tick mark the particular alternative applicable to him.

This type of question is known as multiple-choice question. It suggests several answer among
which the respondent may choose. If a multiple choice question is used, all alternatives should
be stated and a ‘don’t know’category be left in the questionnaire. Such questions not only
facilitate tabulation but will take very little time of the respondent to fill the questionnaire.
However, this type of question is excellent if most of the possible answers are both known and
few in number. When the possible answers are numerous, a limited list – even if accompanied
by "any other” category – may elicit response different from that which otherwise would be
forthcoming. Multiple choice questions tend to bias result by the order in which alternative
answers are given. When ideas are involved, the first item in the list of alternative has a
favourable bias. The use of multiple- choice question is indicated only when the investigator is
confident of the existence of a limited group of important alternatives and it should be avoided
when the there are many possible responses of relatively equal significance.

17
9. “Yes” or “No” question. As far as possible the questions should be of such a nature that they
can be answered easily in ‘Yes’ or ‘No’. Such questions pose a simple alternative to the
respondent. This is an excellent technique if applied to situations where a clear-cut alternative
exists. The questions “Do you own a car?”, “Are you married?”, “Did you vote in last election?”
can easily be answered with a “yes” or “no”. However, when the alternatives is not clear-cut,
the “yes” or “no” type question should be avoided. A question such as “Do you favour the
Government policies?” usually cannot be answered with a simple reply. The Government has so
many policies and only the most radical or partisam would favour or oppose them all. A typical
citizen may endorse many, have no opinion on some and reject others. The “Yes” or “no”
question in this case compels him to compress a variety of opinions into a simple alternative
which may, in reality, not exist.

Sometimes a respondent cannot give a simple “yes” or “no” answer either because he has not
yet made up his mind or because the lacks information on the topic. For example, the answer
to the question “Are you in favour of public schools?” may not always be in ‘yes’ or ‘no’
because the respondent has not thought over it. In such cases additional alternative such as ‘do
not know’, ‘undecided’; no opinion’ should be included.

10. Specific information questions and open-end questions. Specific information questions call
for a specific item of information. For example, “What is your age?”, “How many children do
you have ?”, etc. These questions are simple and direct and are well adapted to securing
information of this type. Care should be taken to use this type of question only where the
respondent can answer correctly. The open question does not pose alternatives or request
specific information. It leaves the respondent free to make whatever reply he chooses. For
example, the question, what should be done to enhance the practical utility of [Link] course?
Why do you use Colgate toothpaste or ‘Lux’ soap are open-end questions. In many ways open
question is superior to other types—there is no danger of being unduly restrictive suggesting
answers, posing false alternatives, and introducing some bias. It also may serve to interest the
respondent in the interview itself, especially if he is asked his opinion at the outset. However,
open questions are difficult to tabulate. Since no restriction is placed upon the variety of
answers, many will often be forthcoming. This not only increased the labour involved but
frequently leads to improper tabulation. Hence every effort must be made to minimize open
questions in the questionnaire.

11. Questionnaire should look attractive. A questionnaire should be made to look as attractrive
as possible. The printing and the paper used, etc., should be good and plenty of space should be
left for answers depending upon the type of questions.

12. Questions requiring calculations should be avoided. Questions should not require
calculcations to be made. For example, informants should not be asked yearly income, for in
most cases they are paid monthly. Similarly, questions necessitating calculation of ratios and
percentages, etc., should not be asked as it may take much time and the informant may not
send back the questionnaire.

18
13. Pre-testing the questionnaire. The questionnaire should be pre-tested with a group before
mailing it out. The advantage of pre-testing is that the shortcomings of the questionnaire can
be discovered and it can be revised in the light of the tryout.

14. Cross-checks. If possible, one or more change the serial should be incorporated into the
questionnaire to determine whether the respondent is answering at least the important
questions correctly.

Method of tabulation. The method to be used for tabulating the results should be determined
before the final draft of the questionnaire is made. If the results of the questionnaire are to be
computerized, it is desirable to consult the computer experts before making a final draft.
2.1.11 Pre-testing the questionnaire (or pilot survey)

Before final form of the questionnaire is adopted it is desirable to carry out a preliminary
experiment on a sample basis. When questionnaires are to be distributed on a large scale, it is
absolutely essential to pre-test them. There are many avantages of pre-testing the
questionnaire, such as:

1. The investigator can find out what are the drawbacks of the questionnaire, i.e., which
questions ought to be edited/revise and which more ought to be added.

2. An idea can be formed about the extent of non-response likely to be expected.

3. Greater co-operation of the informants can be secured. Even persons most allergic to write
can, with proper inducement, be prevailed upon to answer the questions. It is the surveyor’s
job to find out what these appeals are.

While pre-testing the questionnaire, it is important always to cover a cross-section of the


population eventually to be surveyed. When the sample is drawn, it should be broken down
into various sub-samples by taking, for instance, every tenth or every hundredth case from the
entire list.

The work of pre-testing the questionnaire must be done with utmost care and caution
otherwise unnecessary and unwanted changes have to introduced. Proper testing, revising and
re-testing the questionnaire would yield high dividends.

If time and budget permit, a second pilot study should be undertaken on a fresh sampling of
respondents to further improve the final document.
2.1.12 Sources of secondary data

In most of the studies the investigator may find it difficult to collect primary data on all related
issues and as such he makes use of the data collected by others. There is a vast amount of
published information from which statistical studies may be made and fresh statistics are

19
constantly in a state of production. The sources of secondary data can broadly be classified
under two heads:

(1) Published sources, and (2) Unpublished sources


[Link] Published Sources

The various sources of published data are:

1. Reports and official publications of

(a) International bodies such as the ‘World Bank’, ‘International Labour Organisation’,
‘Statistical Office of the United Nations’.

(b) Central and State Governments such as Abstract of the Indian Union, Economic Survey of
India, India 1988-89, etc.

(c) Reports of the Ad-hoc Committees and Commissions appointed by the Government such as
Sarkaria Committee, Mehrotra Committee, Shah Commission, Fourth Pay Commission, etc.

2. Semi-official publications of various local bodies such Municipal Corporations and District
Boards.

3. Publications of autonomous and private institutes; such as

(a) Trade and Professional bodies, such as , the Federation of Indian Chambers of Commerce
and Industry, the Institute of Chartered Accountants, the Institute of Foreign Trade. The
prestigious journals of these institutes are respectively “Economic Trends’, ‘The Chartered
Accountant’, ‘Foreign Trade Review’.

(b) Financial and economic journals such as ‘Indian Economic Review’, ‘Reserve Bank of India
Bulletin’, ‘ Indian Finance’.

(c) Annual Reports of Joint Stock Companies and Corporations.

(d) Publications brought out by various autonomous Research Institutes and Scholars such as
Institute of Economic,Pune.

It should be noted that the publications mentioned above vary with regard to the periodicity of
publication. Some are published at regular intervals(yearly, monthly, weekly, etc.) whereas
others are ad hoc publications, i.e., with no regularity about periodicity of publication.
[Link] Unpublished Sources

20
There are various sources of unpublished data such as records maintained by various
Government and private offices, studies made by research institutions, scholars, etc. Such
sources can be used wherever necessary.
2.1.13 Editing primary and secondary data

Once data have been obtained either from primary or secondary source, the next step in a
statistical investigation is to edit the data, i.e., to scrutiny the same. The main objective of
editing is to detect possible errors and irregularities. The task of editing is a highly specialized
one and requires great care and attention. Negligence in this respect may render useless the
findings of an otherwise valuable study. However, it should be noted that the work of editing
data collected from internal records and published sources is relatively simple – it is the data
collected from a survey that need extensive editing.

While editing primary data the following considerations need attention:

1. The data should be complete.

2. The data should be consistent,

3. The data should be accurate, and

4. The data should be homogeneous.

1. Editing for completeness. The editior should see that each schedule and questionnaire is
complete in all respects, i.e., answer to each and every question has been furnished. If some
questions have not been answered and those questions are of vital importance the informants
should be contacted again either personally or through correspondence. It may happen that in
spite of best efforts a few questions remain unanswered. In such questions, the editor should
mark ‘No answer ’ in the space provided for answers and if the questions are of vital
importance then the schedule or questionnaire should be dropped.

2. Editing for consistency. While editing the data for consistency, the editor should see that the
answers to questions are not contradictory in nature. If there are mutually contradictory
answers, he should try to obtain the correct answers either by referring back the questionnaire
or by contacting, wherever possible, the informant in person. For example, if amongst others,
reply to the questions: (a) Are you married? (b) Mention the number of children you have, and
the are respectively ‘no’ and to ‘three’, then there is a contradiction and it should be clarified.

3. Editing for accuracy. The reliability of conclusions depends basically on the correctness of
information. If the information supplied is wrong, conclusions can never be valid. It is,
therefore, necessary for the investigators to see that the information is accurate in all respects.
However, this is one of the most difficult tasks of the investigators. If the inaccuracy is due to
arithmetic errors, it can be easily detected and corrected. But if the cause of inaccuracy is faulty

21
information supplied, it may be difficult to verify it, e.g., information relating to income, age,
etc.

4. Editing for uniformity. By homogeneity we mean the condition in which all the questions
have been understood in the same sense. The investigators must check all the questions for
uniform interpretation. For example, as to the question of income, if some informants have
given monthly income, others annual income and still others weekly income or even daily
income, no comparison can be made. Similarly, if some persons have given the basic income
whereas others the total income, no comparison is possible. The investigators should check up
that the information supplied by the various people is homogeneous and uniform.
2.1.14 Precautions in the use of secondary data

Since secondary data have already been obtained it is highly desirable that a proper scrutiny of
such data is made before they are used by the investigator. In fact, the user has to be extra-
cautious while using secondary data. In this context, Prof. Bowely rightly points out that
“secondary data should not be accepted at their face value”. The reason is that such data may
be erroneous in many respects due to bias, inadequate size of the sample, substitution, errors
of definition, arithmetical errors, etc. Even if there is no error such data may not be suitable and
adequate for the purpose of the enquiry. Hence before using such data, the investigator should
consider the following aspects:

[Link]

Whether the data are suitable for the purpose of investigation in view. Before using secondary
data the investigator must ensure that the data are suitable for the purpose of the enquiry. The
suitability of data can be judged in the light of the nature and scope of investigation. For
example, if the object of enquiry is to study the wage levels including allowances of workers
and the data relate to basic wages alone, such data would not be suitable for the immediate
purpose. It may be difficult to find data which exactly fit to the needs of the present project.

Quit often secondary data do not satisfy immediate needs because they have been compiled
for other purposes. Even when directly pertinent to the subject under study, secondary data
may be just enough off the point to make them of little or no use. The value of secondary data
is frequently impaired by:

(i) Variation in the units of measurement: consumer income, for example, may be measured by
individual, family household, spending units or tax return.

(ii) Definition of classes may be different: for example, definition of literate, educated, poor may
vary from researcher to researcher.

(iii)Variation in the period to which the data is related to: the data available may relate to a
different time period and not serve the purpose of researcher. Data published to promote the
interests of a particular group whether it is political, social or commercial are suspicious.

22
2. Adequacy

Whether the data are adequate for the investigation. If it is found that the data are suitable for
the purpose of investigation, they should be tested for adequacy. Adequacy of the data is to be
judged in the light of the requirements of the survey and the geographical area covered by the
available data. For example, in the illustration given above, if our objective is to study the wage
rates of the workers in fish processing industry in India and if the available data cover only the
State of Tamil Nadu., it would not serve the purpose. The question of adequacy may also be
considered in the light of the time period for which the data are available. For example, for
studying trend of prices we may use data for the last 8-10 years but from the source known to
us data may be available for the last 2-3 years only which would not serve the purpose.

3. Reliability

Whether the data are reliable. It is very difficult to find out whether the secondary data are
reliable or not. The following tests, if applied, may be helpful to determine how far the given
data are reliable:

(i) Which specific method of data collection was used? If a source fails to give a detailed
description of its method of data collection, researchers should be hesitant about using the
information provided. When the methodology is described, researchers should subject it to a
painstaking examination.

Data published to promote the interests of a particular group whether it is political, commercial
or social are suspect.

(ii) Was the collecting agency unbiased or did it “have an axe to grind”?

(iii) If the enumeration was based on a sample, was the sample representative?

(iv) Were the enumerators capable and properly trained? Incompetent or poorly trained
enumerators cannot be depended upon to produce useful result.

(v) Was there a proper check on the accuracy of field work?

(vi) Was the editing, tabulating and analysis carefully and conscientiously done? Carelessness in
either one or more of these functions can render of little value the findings of an otherwise
valuable study.

(vii) What degree of accuracy was desired by the compiler? How far was it achieved?
Unit 3: Sampling and sample method

Learning objective

23
For getting information, sampling is necessary. The different sampling methods are dealt with
advantages and disadvantages.
3.1.1 Introduction

When secondary data are not available for the problem under study, a decision may be taken to
collect primary data by using any of the methods discussed in the previous chapter. The
required information may be obtained by following either the census methods or the sample
method.
3.1.2 Census and sample method

Under the census or complete enumeration survey method, data are collected from each and
every unit (person, household, field, shop, factory, etc., as the case may be) of the population
or universe which is the complete set of items which are of interest in any particular situation.
For example, if the average wage of workers working in fish processing industry in India is to be
calculated, then wage figures would be obtained from each and every worker working in the
industry and by dividing the total wages which all these workers received by the number of
workers working in the industry, we would get the figure of average wage. Some of the merits
of the census method are:

(i) Data are obtained from each and every unit of the population.

(ii) the results obtained are likely to be more representative, accurate and reliable.

(iii) It is an appropriate method of obtaining information on rare events such as areas under
some crops and yield thereof, the number of persons of certain age groups, their distribution by
sex, educational level of people, etc. This is the reason why throughout the world the
population data are obtained by conducting a census generally every 10 years by the census
method.

(iv) Data of complete enumeration census can be widely used as a basis for various surveys

However, despite these advantages the census method is not very popularly used in practice.
The effort, money and time required for carrying out complete enumeration will generally be
very large and in many cases cost may be so prohibitive that the very idea of collecting
information may have to be dropped. This is more true of underdeveloped countries where
resources constitute a big constraint. Also if the population is infinite or the evaluation process
destroys the population unit, the method cannot be adopted.

Sampling is simply the process of drawing a sample the population under study. Thus, using the
sampling technique instead of every unit of the universe only a part of the universe is studied
and the conclusions are drawn on that basis for the entire universe. A sample is a subset of
population units. The process of sampling involves three elements:

a. determining the sample size

24
b. Selecting the sample units,

c. collecting obtaining information from sampled units

The three elements cannot generally be considered in isolation from one another. Sample
selection, data collection, and estimation are all interwoven and each has an impact on the
others. Sampling is not haphazard selection—it embodies definite rules for selecting the
sample. But having followed a set of rules for sample selection, we cannot consider the
estimation process independent of it—estimation is guided by the manner in which the sample
has been selected.

Although much of the development in the theory of sampling has taken place only in recent
years, the idea of sampling is pretty old. Since times immemorial people have examined a
handful of grains to ascertain the equality of the entire lot. A housewife examines only two or
three grains of boiling rice to know whether the pot of rice is ready or not. A doctor examines a
few drops of blood and draws conclusion about the blood constitution of the whole body. A
businessman places orders for material by examining only a small sample of the same. A
teacher may put questions to one or two students and find out whether the class as a whole is
following the lesson. In fact there is hardly any field where the technique of sampling is not
used.

It should be noted that a sample is not studied for its own sake. The basic objective of its study
is to draw inference about the population. In other words, sampling is a tool which helps to
know the characteristics of the universe or population by examining only a small part of it. The
values obtained from the study of sample, such as the average and dispersion, are known as
‘statistics’. On the other hand, such values for the population are called ‘parameters’.
3.1.3 Theoretical basis of sampling

On the basis of sample study we can predict and generalize the behaviour of mass phenomena.
This is possible because there is no statistical population whose elements would vary from each
other without limit. For example, wheat varies to a limited extent in colour, protein content,
length, weight, etc., it can always be identified as wheat. Similarly, apples of the same tree may
vary in size, colour, taste, weight, etc., but they can always be identified as apples. Thus we find
that although diversity is a universal quality of mass data, every population has characteristic
properties with limited variation. This makes possible to select a relatively small unbiased
random sample that can portray fairly well the traits of the population.

There are two important laws on which the theory of sampling is based:

1. Law of ‘Statistical Regularity’, and

2. Law of ‘Inertia of Large Numbers’,


3.1.4 Law of Statistical Regularity

25
This law is derived from the mathematical theory of probability. In the words of King: "the law
of statistical regularity lays down that a moderately large number of items chosen at random
from a large group are almost sure on the average to possess the characteristics of the large
group”. In other words, this law points out that if a sample is taken at random from a
population, it is likely to possess almost the same characteristics as that of the population. This
law directs our attention to an important point, that is, the desirability of choosing the sample
at random.

By random selection we mean a selection where each and every item of the population has an
equal chance of being selected in the sample. In other words, the selection must not be made
by deliberate exercise of one’s discretion. A sample selected in this manner would be a
representative of the population. If this condition is satisfied it is possible for one to depict fairly
accurately the characteristics of the population by studying only a part of it. Hence, this law is
of great practical significance because it makes possible a considerable reduction of the work
necessary before any conclusion is drawn regarding a large universe. For example, if one
intends to make a study of the average height of the students of an University it is not
necessary to measure the heights of each and every student. A few students may be selected at
random from each college, their heights may be measured and the average height of university
students in general may be inferred.

It should be noted that the results derived from sample data may be different from that of the
population. This is for the simple reason that the sample is only a part of the whole universe.
For example, the average height of students of the University may come out to be 160cm. By
census method whereas it may be 159cm. or 161cm. for the sample taken. It should be just a
coincidence if the height comes out to be exactly 160cm. under both the methods. However,
there would not be much difference in the results derived if the sample is a representative of
the population.
3.1.5 Law of Inertia of Large Numbers

This law is a corollary of the law of statistical regularity. It is of great significance in the theory of
sampling. It states that, other things being equal, larger the size of the sample, more accurate
the results are likely to be. This is because large numbers are more stable as compared to small
ones. The difference in the aggregate result is likely to be insignificant, when the number in the
sample is large, because when large numbers are considered the variations in the component
parts tend to balance each other and, therefore, the variation in the aggregate is insignificant.
For example, if a coin is tossed 10 times we should expect equal number of heads and tails, i.e.,
5 each. But since the experiment is tried a small number of times it is likely that we may not get
exactly 5 heads and 5 tails. The result may be a combination of 9 heads and 1 tail, or 8 heads
and 2 tails, or 7 heads and 3 tails. If the same experiment is carried out 1,000 times the chance
of 500 heads and 500 tails would be very high, i.e., the result would be very near to 50% heads
and 50% tails. The basic reason for such likelihood is that the experiment has been carried out
sufficiently large number of times and possibility of variation in one direction compensating
others in a different direction is greater. If at one time we get continuously 5 heads, it is likely
that at other time we may get continuously 5 tails, and so on, and for the experiment as a

26
whole the number of heads and tails may be more or less equal. Similarly, if it is intended to
study the variation in the production of rice over a number of years and data are collected from
one or two States only, the result would reflect large variations in production due to the
favourable factors in operation. If, on the other hand, figures of production are collected for all
the States in India, it is quite likely that we find little variation in the aggregate. This does not
mean that the production would remain constant for all the years. It only implies that the
changes in the production of the individual States will be counterbalanced so as to reflect
smaller variations in production for the country as a whole.
3.1.6 Essentials of sampling

If the sample results are to have any worthwhile meaning, it is necessary that a sample
possesses the following essentials:

i. Representativeness. A sample should be selected such that it truly represents the universe
otherwise the results obtained may be misleading. To ensure representativeness the random
method of selection should be used.

ii. Adequacy. The size of sample should be adequate, otherwise characteristic of the universe
may be obtained is to less accuracy.

iii. Independence. All items of the sample should be selected independently of one another and
all items of the universe should have the same chance of being selected in the sample. By
independence of selection we mean that the selection of a particular item in one draw has
influence on the probabilities of selection in any other draw.
3.1.7 Methods of sampling

Various methods of sampling can be grouped under two broad heads: probability sampling
(also known as random sampling) and non-probability (or non-random) sampling.
[Link] Probability sampling

Probability sampling methods are those in which every item in the universe has a known
chance, or probability, of being chosen for the sample. This implies that the selection of sample
items is independent of the person making the study-that is, the sampling operation is
controlled so objectively that the items will be chosen strictly at random.
[Link] Non-probability Sampling

Non-probability sampling methods are those which do not provide every item in the universe
with a known chance of being included in the sample. The selection process is, at least, partially
subjective.

It may be noted that the term random sample is not used to describe the data in the sample
but the process employed to select the sample. Randomness is thus a property of the sampling

27
procedure instead of an individual sample. As such, randomness can enter processed sampling
in a number of ways and hence random samples may be of many kinds.
[Link] Advantages of Probability Sampling

The following are the basic advantages of probability sampling methods:

1. Probability sampling does not depend upon the existence of detailed information about the
universe for its effectiveness.

2. Probability sampling provides estimates which have measurable precision

3. it is possible to evaluate the relative efficiency of various sample designs only when
probability sampling is used.
[Link] Limitations of Probability Sampling

Despite the great advantages of probability sampling techniques mentioned above it has
certain limitations because of which also non-probability sampling is used in practice. These
limitations are

Theoretical Knowledge

1. Probability sampling requires a very high level of skill and experience for its use

2. It requires a lot of time to plan and execute a probability sample

3. The costs involved in probability sampling are generally large as compared to non-probability
sampling

Non probability sampling

Non–random sampling is a process of sample selection without the application of


randomization. In other words, a non-random sample is selected on a basis other than the
probability consideration such as convenience, judgment etc.

The most important difference between random and non-random sampling is that the pattern
of sampling variability can be ascertained in case of random sampling, whereas there is no way
of knowing the pattern of variability in the process in non-random sampling.

Some of the important popularly used sampling methods that are are discussed in the following
sections.
3.1.8 Non-probability sampling Methods

Judgment Sampling

28
In this method of sampling the selectuon of sample items depends exclusively on the
judgement of the investigator. In other words, the investigator exercises his judgement in the
selection and includes those items in the sample which he thinks are most typical of the
universe with regard to the characteristics under investigation. For example, if sample of ten
students is to be selected from a class of sixty for analyzing the spending habits of students, the
investigator would select 10 students who, in his opinion, are representative of the class.

Merits: Though the principles of sampling theory are not applicable to judgment sampling, the
method is sometimes used in solving many types of economic and business problems. The use
of judgment sampling is justified under a variety of circumstances.

i. when the number of units in the population, simple random selection may miss to include the
more important elements, whereas judgment sampling would certainly include them in the
sample.

ii. When some unknown traits of a population are to be studied and some of whose
characteristics are known, we may then stratify the population according to these known
properties and select sampling units from each stratum on the basis of judgement. This method
is used to obtain a more representative sample.

iii. In solving everyday business problems and making public policy decisions, executives and
public officials are often pressed for time and cannot wait for probability sample designs.
Judgment sampling is then the only practical method to arrive at solutions to their urgent
problems.

Limitations. i. This method is not scientific because the population units to be sampled may be
affected by the personal prejudice or bias of the investigator. Thus, judgment sampling involves
the risk that the investigator may establish foregone conclusions by including those items in the
sample which conform to his preconceived notions. For example, if an investigator holds the
view that the wages of workers in a certain establishment are very low, and if he adopts the
judgment sampling method, he may include only those workers in the sample whose wages are
low and thereby establish his point of view which may be far from the truth. Since an element
of subjectiveness is possible, this method cannot be recommended for general use.

ii. There is no objective way of evaluating the reliability of sample results

The success of this method depends upon the excellence in judgment. If the individual making
decisions is knowledgeable about the population and has good judgment, then the resulting
sample may be representative, otherwise the inferences based on the sample may be
erroneous. It may be noted that even if a judgment sample is reasonably representative, there
is no objective method for determining the size or likelihood of sampling error. This is a big
disadvantage of this method.
[Link] Convenience Sampling

29
A convenience sample is obtained by selecting ‘convenient’ population units.

The method of convenience sampling is also called the chunk. A chunk refers to that fraction of
the population being investigated which is selected neither by probability nor by judgment but
by convenience. A sample obtained from readily available lists such as automobile registrations;
telephone directories, etc., is a convenience sample and not a random sample even if the
sample is drawn at random from the lists. If a person is to submit a project report on labour-
management relations in processing industry and he takes a processing industry close to his
office and interviews some people over there, he is following the convenience sampling
method. Convenience samples are prone to bias by their very nature—selecting population
elements which are convenient to choose almost always make them special or different from
the best of the elements in the population in some way.

Hence the sample obtained by the convenience sampling method can hardly be representative
of the population—they are generally biased and unsatisfactory. However, convenience
sampling is often used for making pilot studies. Questions may be tested and preliminary
information may be obtained by the chunk before the final sampling design is decided upon.
[Link] Quota sampling

Quota sampling is a type of judgment sampling and is perhaps the most commonly used
sampling technique in non-probability category. In a quota sample, quotas are set up according
to some specified characteristics such as income groups, age, political or religious affiliations,
and so on. Each interviewer is then told to interview the population units which constitute his
quota. Within the quota, the selection of sample items depends on personal judgment. For
example, in a radio listening survey, the interviewers may be told to interview 500 people living
in a certain area and that out of every 100 persons interviewed 60 are to be housewives, 25
farmers and 15 children under the age of 15. Within these quotas the interviewer is free to
select the people to be interviewed. The cost of conducting interview per person interviewed
may be relatively small for a quota sample. However, there are numerous opportunities for bias
which may invalidate the results. For example, interviewers may miss farmers working in the
fields or talk with those housewives who are at home. If a person refuses to respond, the
interviewer simply selects someone else. Because of the risk of personal prejudice and bias
entering the process of selection, the quota sampling is not widely used in practical work.

Quota sampling and stratified random sampling are similar in as much as in both methods the
universe is divided into parts and the total sample is selected from all the parts. However, the
two procedures diverge radically. In stratified random sampling the sample within each stratum
is chosen at random. In quota sampling, the sampling within each cell is not done at random;
the field representatives are given wide latitude in the selection of respondents to meet their
quotas.

Quota sampling if often used in public opinion studies. It occasionally provides satisfactory
results if the interviewers are carefully trained and if they follow their instructions closely. It is
often found that since the choice of respondents within a cell is left to the field representatives,

30
the more accessible and articulate people within a cell will usually be the ones who are
interviewed. Slight negligence on the part of interviewers may lead to interviewing ineligible
respondents. Even with alert and conscientious field representatives it is often difficult to
determine such control category as age, income, educational qualifications, etc.
3.1.9 Probability sampling methods

Simple Random Sampling

Simple random sampling refers to that sampling technique in which each and every unit of the
population has an equal opportunity of being selected in the sample. In simple random
sampling inclusion of items in the sample is simply a matter of chance—personal bias of the
investigator does not influence the selection. It should be noted that the word ‘random’ does
not mean ‘haphazard’ or ‘hit-or-miss’- it rather means that the selection process is such that
the chance only determines which items shall be included in the sample. As pointed out by
Chou, when a sample of size n is drawn from a population with N elements the sample is
‘simple random sample’ if any of the following is true. And, if any one of the following is true, so
are the other two.

1. All n items of the sample are selected independently of one another and all N items in the
population have the same chance of being included in the sample. By independent of selection
we mean that the selection of particular item in one draw has no influence on the probabilities
of selection in any other draw.

2. At each selection, all remaining items in the population have the same chance of being
drawn. If sampling is made with replacement, i.e., when each unit drawn from the population is
returned prior to drawing the next unit, each item has a probability of 1/N of being drawn at
each selection. If sampling is without replacement, i.e., when each unit drawn from the
population is not returned prior to drawing the next unit, the probability of selection of each
item remaining in the population at the first draw is 1/N, at the second draw 1/(N-1), at the
third draw is 1/(N-2), and so on. It should be noted that sampling with replacements has very
limited and special use in statistics—we are mostly concerned with sampling without
replacement.

3. All the possible samples of a given size n are equally likely to be selected.

To ensure randomness of selection one may adopt either the Lottery Method or use table of
random numbers.
[Link] Lottery Method

This is a very popular method of taking a random sample. Under this method, all items of the
universe are numbered or named on separate slips of paper of identical size and shape. These
slips are then folded and mixed up in a container or drum. A blindfold selection is then made of
the number of slips required to constitute the desired sample size. The selection of items thus
depends entirely on chance. The method would be quite clear with the help of an example. If

31
we want to take a sample of 10 persons out of a population of 100, the procedure is to write
the names of all the 100 persons on separate slips of paper, fold these slips, mix them
thoroughly and then make a blindfold selection of 10 slips.

The above methods is very popular in lottery draws where a decision about prizes is to be
made. However, while adopting lottery method it is absolutely essential to see that the slips are
of identical size, shape and colour, otherwise there is a lot of possibility of personal prejudice
and bias affecting the results.
[Link] Table of random Numbers

The lottery method discussed above becomes quite cumbersome as the size of population
increases. An alternative method of random selection is that of using the table of random
numbers.

The random numbers are generally obtained by some mechanism which, when repeated a large
number of times, ensured approximately equal frequencies for the number from 0 to 9 and also
proper frequencies for various combinations of numbers (such as 00,01,…………..99; 000, 001,
………….999; etc) that could be expected in a random sequence of the digits 0 to 9.

Several standard table of random numbers are available, among which the following may be
specially mentioned, as they have been tested extensively for randomness:

1. Tippett’s (1927) random number tables consisting of 41,600 random digits grouped into
10,400 sets of four –digit random numbers:

2. Fisher and Yates (1938) table of random numbers with 15,000 random digits arranged into
1,500 sets of ten-digit random numbers:

3. Kendall and Babington Smith (1939) table of random numbers consisting of 1,00,000 random
digits grouped into 25,000 sets of four-digit random numbers:

4. Rand Corporation (1955) table of random numbers consisting of 1,00,000 random digits
grouped into 20,000 sets of five-digit random numbers; and

5. C.R. Rao, Mitra and Mathai (1966) table of random numbers

Trippett’s table of random numbers is most popularly used in practice. We give below the first
forty sets from Tippett’s table as an illustration of the general appearance of random numbers:
2952 6641 3992 9792 7969 5911 3170 5624
4167 9524 1545 1396 7203 5356 1300 2693
2670 7483 3408 2762 3563 1089 6913 7691
0560 5246 1112 6107 6008 8125 4233 8776
2754 9143 1405 9025 7002 6111 8816 6446

32
It is important that the starting point in the table of random numbers be selected in some
random fashion so that every units has an equal chance of being selected.

One may question, and quite rightly, as to how it is ensured that these digits are random. It may
be pointed out that the digits in the table were chosen haphazardly but the real guarantee of
their randomness lies in practical tests. Tippett’s numbers have been subjected to numerous
tests and used in many purposes. An example to illustrate how Tippett’s table of random
numbers may be used is given below. Suppose we have to select 20 items out of 6,000. The
procedure is to number all the 6,000 items from 1 to 6,000. A page from Tippett’s table may
then be consulted and the first twenty numbers up to 6,000 are noted down. Items bearing
those numbers will be included in the sample. Making use of the portion of the table the
required numbers are noted as:
2952 3992 5911 3170 5624 4167
1545 1396 5356 1300 2693 2370
3408 2762 3563 1089 0560 5246
1112 4233

The items which bear the above numbers constitute the sample.

If the size of universe is less than 1,000 the procedure will be different, as Tippett’s numbers are
available only in four figures. Thus, for example, if it is desired to take a sample of 10 items out
of 400, all items from 1 to 400 should be numbered as 0001 to 0400. We may now select 10
numbers from the table which are up to 0400.

If the size of universe is less than 100, the table is used as follows: Suppose ten numbers from 0
to 80 are required. We start anywhere in the table and write down the numbers in pairs. The
table can be read horizontally, vertically,diagonally or in any methodical way. Starting with the
first and reading horizontally first we obtain 29, 52, 66, 41, 39,92,97,79,69,59,11,31, 70, 56, 24,
70,56, 24,41, 67 and so on. Ignoring the numbers greater than 80, we obtain for our purpose
ten random numbers, namely, 29, 52, 66,41, 39, 73, 69, 59, 11 and 31.

Fishers and Yates tables consist of 15,000 numbers. These have been arranged in two digits in
300 blocks, each block consisting of 5 rows and 5 columns. Kendall and Smith also constructed
random numbers (10,000 in all) by using a randomizing machine. However, this method of
random selection cannot be followed in case of articles like ghee, oil, petrol, wheat, etc.

Merits. 1. Since the selection of items to the sample depends entirely on chance there is no
possibility of personal bias affecting the results.

[Link] to judgment sampling a random sample represent the universe in a better way. As
the size of the sample increases, it become increasingly representative of the population.

3. The analyst can easily assess the accuracy of the estimate because sampling errors follow the
principles of chance. The theory of random sampling is further developed than that of any

33
other type of sampling which enables the analyst to provide the most reliable information at
the least cost.

Limitations. 1. The use of simple random sampling necessitates a completely catalogued


universe from which draw the sample. But it is often difficult for the investigator to have up-to-
date list of all the items of the population to be sampled. This restricts the use of this method in
economic and business data where very often we have to employ restricted random sampling
designs.

2. The size of the sample required to ensure statistical reliability is usually larger under random
sampling than non-probability sampling.

3. From the point of view of field survey it has been claimed that cases selected by random
sampling tend to be too widely dispersed geographically and that the time and cost of
collecting data become too large.

4. Random sampling may produce the most non-random-looking results. For example, thirteen
cards from a well-shuffled pack of playing cards may consist of one suit. But the probability of
this type of occurrence is very, very low.
3.1.10 Restricted Random Sampling

Stratified Sampling

Stratified random sampling or simply stratified sampling is one of the random methods which,
by using the available information concerning the population, attempts to design a more
efficient sample than obtained by the simple random procedure.

While applying stratified random sampling technique, the procedure followed is given below.

1. The universe to be sampled is subdivided (or stratified) into groups which are mutually
exclusive and include all items in the universe.

2. A simple random sample is then chosen independently from each group.

This sampling procedure differs from simple random sampling, where in the latter the sample
items are chosen at random from the entire universe. In stratified random sampling the
sampling is designed so that a designated number of items is chosen from each stratum. In
simple random sampling the distribution of the sample among strata is left entirely to chance.
[Link] How to Select Stratified Random Sample?

Some of the issues involved in setting up stratified random sample are;

1. Base of Stratification

34
What characteristic should be used to subdivide the universe into different strata? As a general
rule, strata are created on the basis of a variable known to be correlated with the variable of
interest and for which information on each universe element is known. Strata should be
constructed in a way which will minimize differences among sampling units within strata, and
maximize difference among strata.

For example, if we are interested in studying the consumption pattern of the people of Delhi,
the city of Delhi may be divided into various parts (such as zones or wards) and from each part
a sample may be taken at random. Before deciding on stratification we must have knowledge of
the traits of the population. Such knowledge may be based upon expert judgment, past data,
preliminary observations from pilot studies, etc.

The purpose of stratification is to increase the efficiency of sampling by dividing a


heterogeneous universe in such a way that (i) there is as great a homogeneity as possible within
each stratum and (ii) a marked difference is possible between the strata.

2. Number of strata

How many strata should be constructed? The practical consideration limit the number of strata
that is feasible, costs of adding more strata may soon outrun benefits. As a generalization more
than six strata may be undersirable.

3. Sample size within strata

How many observations should be taken from each stratum? When deciding this question we
can use either a proportional or a disproportional allocation. In proportional allocation, one
samples each stratum in proportion to its relative weight. In disproportional allocation this is
not the case. It may be pointed out that proportional allocation approach is simple and if all one
knows about each stratum is the number of items in that stratum, the different strata are
sampled at different rates. As a general rule when variability among observations within a
stratum is high, one samples that stratum at a higher rate than for strata with less internal
variation.
[Link] Proportional and Disproportional Stratified Sample

In a proportional stratified sampling plan, the number of items drawn from each stratum is
proportional to the size of the stratum. For example, if the population is divided into five
groups, their respective sizes being 10,15,20,30 and 25 per cent of the population and a sample
of 5,000 is drawn, the desired proportional sample may be obtained in the following manner.
From stratum one 5,000 (0.10) = 500 items
From stratum two 5,000 (0.15) = 750 items
From stratum three 5,000 (0.20) = 1,000 items
From stratum four 5,000 (0.30) = 1,500 items
From stratum five 5,000 (0.25) = 1,250 items
Total = 5,000 items

35
Proportional stratification yields a sample that represents the universe with respect to the
proportion in each stratum in the population. This procedure is satisfactory if there is no great
difference in dispersion from stratum to stratum. But it is certainly not the most efficient
procedure, especially when there is considerable variation in different strata. This indicates that
in order to obtain maximum efficiency in stratification, we should assign greater representation
to a stratum with a large dispersion and smaller representations to one with small variation.

In disproportional stratified sampling equal number of cases is taken from each stratum
regardless of how the stratum is represented in the universe. Thus, in the above example, an
equal number of items (1,000) from each stratum may be drawn. In practice disproportional
sampling is common when sampling from a highly variable universe, wherein the variation of
the measurements differs greatly from stratum to stratum.

The following is that information about data the number of lecturers, readers and professors in
a University
Length of No of No. of No. of Total
service [Link]. Assoc. prof. Professors
Less than 5 2,000 250 50 2,300
yrs.
5-10 yrs 3,000 220 80 3,300
10-15 yrs 1,500 170 30 1,700
More than 880 80 40 1,000
15 yrs
Total 7,380 720 200 8300

Work out how many lecturers, readers and Professors would be selected from each category if
(i) we follow stratified proportionate sampling method and take 10% of the universe equivalent
to the sample size. (ii) if the size of the sample is 10% of the universe but the lecturers, readers
and professors are to be in the ratio of 5:3:2 and weight age of the length of service is to be in
the ratio of 4:3:2:1.

Solution. (i) The sample size is 10% of the universe hence 830 persons would be selected in the
sample. Since 12 strata are formed and we want to follow proportionate stratified sampling
method, we will take 10% from each stratum. The number of persons selected shall be as
follows:
Length of Lecturers Readers Professors Total
service
Less than 200 25 5 230
yrs.
5-10 yrs 300 22 8 330
10-15 yrs 150 17 3 170
More than 88 8 4 100
yrs

36
Total 738 72 20 830

(ii) In the second case also the size of sample is 830 but the lecturers, readers and professors
are to be in the ratio of 5:3:2 of the sample, i.e., we take 415 lectures, 249 readers and 166
professer. Then the length of service weightage in the ratio of 4:3:2:1 and to be taken as
follows.
Length of Lecturers Readers Professors Total
service
Less than 332
166 100 66
5 years
5-10 249
125 74* 50
10-15 166
83 50 33
Above 15 83
41 25 17
Total 415 249 166 830

Merits. 1. More representative. Since the population is first divided into various strata and then
a sample is drawn from each stratum there is a little possibility of any essential group of the
population being completely excluded. A more representative sample is thus secured. C.J.
Grohman has rightly pointed out that this type of sampling balances the uncertainty of random
sampling against the bias of deliberate selection.

2. Greater accuracy. Stratified sampling ensures greater accuracy. The accuracy is maximum if
each stratum is formed so that it consists of uniform or homogenous items.

3. Greater geographical concentration. As compared with random sample, stratified samples


can be more concentrated geographically, i.e. , the units from different strata may be selected
in such a way that all of them are localized in one geographical area. This would greatly reduce
the time and expenses of interviewing.

Limitations. 1. Utmost care must be exercised in dividing the population into various strata.
Each stratum must contain, as far as possible, homogenous items as otherwise the results may
not be reliable. If proper stratification of the population is not done, the sample may have the
effect of bias.

2. The items from each stratum should be selected at random. But this may be difficult to
achieve in the absence of skilled sampling supervisors and a random selection within each
stratum may not be ensured.

37
3. Because of the likelihood that a stratified sample will be more widely distributed
geographically than a simple random sample cost per observation may be quite high.
[Link] Systematic sampling

A systematic sample if formed by selecting one unit at random and then selecting additional
units at evenly spaced intervals until the sample has been formed. This method is popularly
used in those cases where a complete list of the population from which sample is to be drawn is
available. The list may be prepared in alphabetical, geographical, numerical or some other
order. The items are serially numbered. The first item is selected at random generally by
following the Lottery method. Subsequent items are selected by taking every kth item from the
list where ‘k’ refers to the sampling interval or sampling ratio, i.e., the ratio of population size to
the size of the sample. Symbolically:

where k= Sampling interval, N=Universe size, and n=Sample size.

While calculating k, it is possible that we get a fraction value. In such case, we should use
approximation procedure, i.e., if the fraction is less than 0.5, it should be omitted and if it is
more than 0.5 should be taken as 1. If it is exactly 0.5 it should be omitted if the number is even
and should be taken as 1, if the number is odd. This is based on the principle that the number
after approximation should preferably be even. For example if the number of students is
respectively 1020, 1150 and 1100 and we want to take a sample of 200, k shall be:

(i) =

(ii) =

(iii) =

Illustration 2. In a class there are 96 students with Roll Nos. from 1 to 96. It is desired to take
sample of 10 students. Use the systematic sampling method to determine the sample.

Solution

= =

From 1 to 96 Roll Nos. the first student between 1 and k, i.e., 1 and 10, will be selected at
random and then we will go on taking even kth student. Suppose the first student comes out to
be 4th in the enrolement. The sample would then consist of the following Roll No.

38
4,14,24,34,44,54,64,74,84,94

Systematic sampling is relatively a simple technique and may be more efficient statistically than
simple random sampling, provided the lists are arranged wholly at random. However, it is rarely
that this requirement is fulfilled. The nearest approach to randomness is provided by
alphabetical lists such as are found in telephone directory although even these may have
certain non-random characteristics.

Merits . The systematic sampling design is simple and convenient to adopt. The time and work
involved in sampling by this method are relatively less. The results obtained are also found to
be generally satisfactory, provided care is taken to see that there are no periodic features
associated with the sampling interval. If populations are sufficiently large, systematic sampling
can often be expected to yield results similar to those obtained by proportional stratified
sampling.

Limitations . The main limitation of the method is that it becomes less representative if we are
dealing with populations having “hidden periodicities”. Also if the population is ordered in a
systematic way with respect to the characteristics the investigator is interested in, then it is
possible that only certain types of items will be included in the population, or at least more of
certain types than others. For instance, in a study of workers’ wages the list may be such that
every tenth worker on the list gets wages above Rs. 750 per month.
[Link] Multi-stage sampling or Cluster Sampling

Under this method, the random selection is made of primary, intermediate and final (or the
ultimate) units from a given population or stratum. There are several stages in which the
sampling process is carried out. At first, the first stage units are sampled by some suitable
method, such as simple random sampling. Then, a sample of second stage units is selected from
each of the selected first stage units, again by some suitable method which may be the same as
or different from the method employed for the first stage units. Further stages may be added as
required. The procedure may be illustrated as follows;

Suppose we want to take a sample of 5,000 households from the State of Tamil Nadu. At the
first stage, the State may be divided into a number of districts and a few districts selected at
random. At the second stage, each district may be subdivided into a number of villages and a
sample of villages may be taken at random. At the third stage, a number of households may be
selected from each of the villages selected at the second stage. To take another example
suppose in a particular survey, we wish to take a sample of 10,000 students from Delhi
University. We may take colleges—primary units—as the first stage, then draw departments as
the second stage, and choose students as the third and last stage.

Merits. Multi-stage sampling introduces flexibility in the sampling method which is lacking in
the other methods. It enables existing divisions and sub-divisions of the population to be used
as units at various stges, and permits the field work to be concentrated and yet large area to be
covered. Another advantage of the method is that subdivision into second stage units(i.e., the

39
construction of the second stage frame) need be carried out for only those first stage units
which are included in the sample. It is, therefore, particularly valuable in surveys of under-
developed areas where no frame is generally sufficiently detailed and accurate for subdivision
of the material into reasonably small sampling units.

Limitations. A multi-stage sample is less accurate than a sample containing the same number
of final stage units which have been selected by some suitable stage process.

We have discussed above the various random procedures in independent designs. In practice
we often combine two or more of these methods into a single design.
3.1.11 Selection of Appropriate Method of sampling

Having discussed various methods of sampling, the question now arises as to which method
should be adopted in a particular situation. It should be noted that each method has its own
speciality and none of them no one method can be regarded as the best under all
circumstances— A number of factors such as nature of the problem, size of universe, size of
sample, availability of finance, time, etc., would affect the choice of a particular method of
sampling.
3.1.12 Size of sample

An important decision has to be taken in adopting a sampling technique is about the size of the
sample. Size of sample means the number of unit to be sampled from the population for
investigation. Different opinions have been expressed by experts on this point. For example,
some have suggested that the sample size should be 5% of the size of population while others
are of the opinion that sample size should be at least 10%. However, these views are of little
use in practice because no hard and fast rule can be laid down that sample size should be 5%,
10% or 25% of the universe size. It may be pointed out that mere size alone does not ensure
representativeness. A smaller sample but well selected sample is small it may not represent the
universe and the inference drawn about the population may be misleading. On the other hand,
if the size of sample is very large, it may be too burdensome financially, require a lot of time
and may have serious problems of managing it. Hence the sample size should neither be too
small nor too large. It should be ‘optimum’. Optimum size, according to Parten, is one that
fulfils the requirements of efficiency representativeness, reliability and flexibility.

The following factors should be considered while deciding the sample size:

(i) The size of the universe—the large the size of the universe, the bigger should be the sample
size.

(ii) The resource available – If the resources available are vast a larger sample size could be
taken. However, in most cases resources constitute a big constraint on sample size.

(iii) The degree of accuracy or precision desired—The greater the degree of accuracy desired
the larger should be sample size. However, it does not necessarily mean that bigger samples

40
always ensure greater accuracy. If a sample is selected by experts by following scientific
method, it may ensure better results even when it is small compared to a situation in which a
large sample size is selected by inexperienced people.

(iv) Homogeneity or Heterogeneity of the Universe—if the universe consists of homogenous


units a small sample serve the purpose but if the universe consists of heterogenous units a
large sample may be inevitable.

(v) nature of study—For an intensive and continuous study a small sample may be suitable. But
for studies which are not likely to be repeated and are quite extensive in nature, it may be
necessary to take a larger sample size.

(vi)Method of sampling adopted—the size of sample is also influenced by the type of sampling
plan adopted. For example, if the sample is a simple random sample it may necessitate a bigger
sample size. However, in a properly drawn stratified sampling plan, even a small sample may
give better results.

(vii) Nature of respondents—Where it is expected a large number of respondents will not co-
operate and send back the questionnaires, a large sample should be selected.

The above factors have to be properly weighted before arriving at the sample size. However,
the selection of optimum sample size is not that simple as it might seem to be. If the sample is
used which is larger than necessary, resources are wasted, if the sample is smaller than
required the objective of the analysis may not be achieved.

If, for example, p= 0.5 and q = 0.5, p = 0.005, then n shall be determined as follows:

n= = 10,000.
[Link] Determination of sample size

A number of formulae have been devised for determining the sample size depending upon the
availability of information. A few formulae are given below:

where

n= Sample size

Z=Value at a specified level of confidence or desired degree of precision.

= Standard deviation of the population

41
d = Difference between population mean and sample mean.

The steps in computing the sample from the above formula are:

(i) Select the desired degree of precision, i.e., specified level of confidence and designate it as
small ‘z’ (at 1% level of significance or 99% confidence level the value of ‘z’ is 2-58 and at 5%
level of significance or 95% confidence level 1.96)

(ii) Multiply the ‘z’ selected in step 1 by the standard deviation of the universe which may be
assumed.

(iii) Divide the product of the preceding step by the standard error of mean or difference
between population and sample mean. Square the resultant quotient. The result is the size of
sample required.

Example. Determine the sample size if population mean=25, Sample means=23 and the
desired degree of precision is 99%.

Solution

Z= 2.576 (at 15 level the Z value is 2.576)

Substituting the values

n= = (7.728)2 = 59.72 or 60

Similarly, the sample size can be determined from the formula for determining the standard
error of mean, i.e.,

or n =

If is 10 and = 2.5, n shall be

n= = (4)2 =16

also from the formula for calculating standard error proportion, the sample size can be
determined:

42
n=

If, for example, p=0.5 and q=0.5, =0.005, then n shall be determined as follows:

n= =10,000
3.1.13 Merits and limitations of sampling

Merits. The sampling technique has the following merits over the complete enumeration
survey:

1. Less time consuming. Since the sample is a study of a part of the population, considerable
time and labour are saved when a sample survey is carried out. For these reasons a sample
provides more timely data in practice then a census.

2. Less cost. Although the amount of effort and expense involved in collecting information is
always greater per unit of the sample than a complete census, the total financial burden of a
sample survey is generally less than that of a complete census. This is because of the fact that in
sampling, we study only a part of population and the total expense of collecting data is less
than that required when the census method is adopted. This is a great advantage particularly in
an underdeveloped economy where much of the information would be difficult to collect by the
census method for lack of adequate resources.

3. More reliable results . Although sampling involves certain inaccuracies owing to sampling
errors, the result obtained is generally more reliable than obtained from a complete count.
There are several reasons for this. First, it is always possible to determine the extent of
sampling errors. Secondly, other types of errors to which a survey is subject, such as inaccuracy
of information, incompleteness of returns, etc., are likely to be more serious in a complete
census than in sample survey. This is because more effective precautions can be taken in a
sample survey to ensure that information is accurate and complete. For these reasons not only
may the total error be expected to be smaller in a sample survey but sample result can also be
used with a greater degree of confidence because of our knowledge of the probable size of
error. Thirdly, it is possible to avail the services of experts and to impart thorough training to
the investigators in a sample survey which further reduces the possibility of errors. Follow-up
work can also be undertaken much more effectively in the sampling method. Indeed, even a
complete census can only be tested for accuracy.

4. More detailed information. Since sampling saves time and money, it is possible to collect
more detailed information in a sample survey. For example, if the population consists of 1,000
persons in a survey of the consumption pattern of the people, the two alternative techniques
available are as follows:

43
(a) We may collect the necessary data from each one of the 1,000 people through a
questionnaire containing, say, 100 questions (census method), or

(b) We may take a sample of 100 persons (i.e., 10% of population) and prepare questionnaire
containing as many as 10 questions. The expenses involved in the latter case would almost be
the same as in the former but it will enable nine times more information to be obtained.

5. Sampling method is the only method that can, be used in certain cases. There are some cases
in which the census method is inapplicable and the only practicable means is provided by the
sample method. For example, if one is interested in testing the breaking strength of chalks
manufactured in a factory under the census method all the chalks would be broken in the
process of testing. Hence, census method is impracticable and resort must be had to the sample
method. Similarly, if the producer wants to find out whether the tensile strength of a lot of
steel wires meets the specified standard, he must resort to sample method because census
would mean complete destruction of all the wires. Also if the population under investigation is
infinite, sampling is the only possible solution.

6. The sample method is often used to judge the accuracy of the information obtained on a
census basis. For example, in the population census which is conducted very often (10 years in
our country) the field officers employ the sample method to determine the accuracy of
information obtained by the enumerators on the census basis.

Limitations. Despite various advantages, sampling is not completely free form limitations. Some
of the difficulties involved in sampling are stated below:

1. A sample survey must be carefully planned and executed, otherwise the results obtained may
be inaccurate and misleading. Of course, serious errors may arise in sampling, if the sampling
procedure is not perfect.

2. Sampling generally require the service of experts, if only for consultation purposes. In the
absence of qualified and experienced persons, the information obtained from sample surveys
cannot be relied upon.

3. At times the sampling plan may be so complicated which may require more time, labour and
money than census. This is so if the size of the sample is a large proportion of the total
population and if complicated weighted procedures are used. With each additional
complication in the survey, the chances of error multiply and greater care has to be taken
which, in turn, means more time and labour.

4. If the information is required for each and every unit in the domain of study, a complete
enumeration survey is necessary.
3.1.14 Sampling and non-sampling errors

44
To appreciate the need for sample surveys, it is necessary to understand clearly the role of
sampling and non-sampling errors in complete enumeration and sample surveys. The errors
arising due to drawing inferences about the population on the basis of few observations
(sampling) is termed sampling error. Clearly, the sampling error in this sense is non-existent in
complete enumeration survey, since the whole population is surveyed. However, the errors
arising at the stage of ascertainment and processing of data, which are termed non-sampling
errors, are common both in complete enumeration and sample surveys.
[Link] Sampling errors

Even if utmost care has been taken in selecting a sample, the results derived from a sample
study may not be exactly equal to the true value in the population. The reason is that estimates
of are based on a part of population and not on the whole and samples are seldom, if ever,
perfect miniature of the population. Hence sampling gives rise to certain errors known as
sampling errors (or sampling fluctuations). These errors would not be present in a complete
enumeration survey. However, the errors can be controlled. The modern sampling theory helps
in designing the survey in such a manner that the sampling errors can be made small.

Sampling errors are two types: biased and unbiased

1. Biased errors: These errors arise from any bias in selection, estimation, etc. For example, if in
place of simple random sampling, deliberate sampling has been used in a particular case some
bias is introduced in the result and hence such errors are called biased sampling errors.

2. Unbiased errors. These errors arise due to chance differences between the members of
population included in the sample and those not included. An error in statistics is the difference
between the value of a statistic and that of the corresponding parameter.

Thus the total sampling error is made up of errors due to bias, if any, and the random sampling
error. The essence of bias is that it forms a constant component of error that does not decrease
in a large population as the number in the sample increased. Such error is, therefore, also
known as cumulative or non compensating error. The random sampling error, on the other
hand, decreases on an average as the size of the sample increases. Such error is, therefore, also
known as non-cumulative or compensating error.
[Link].1 Causes of Bias

Bias may arise due to:

(i) in appropriate process of selection

(ii) in appropriate work during the collection; and

(iii) in appropriate methods of analysis

45
(i) Faulty selection. Faulty selection of the sample may give rise to bias in a number of ways,
such as:

(a) Deliberate selection of a ‘representative’ sample.

(b) Conscious or unconscious bias in the selection of a ‘random’ sample. The randomness of
selection may not really exist, even though the investigator claims that he has a random sample
if he allows his desire to obtain a certain result to influence his selection.

(c) Substitution. Substitution of an item in place of one chosen in random sample sometimes
leads to bias. Thus, if it is decided to interview every 50th householder in the street, it would be
inappropriate to interview the 51st or any other number in his place as the characteristics
possessed by them differ from those who were originally to be included in the sample.

(d) Non-response. If all the items to be included in the sample are not covered there will be
bias even though no substitution has been attempted. This fault occurs particularly in mailed
questionnaires method, which are incompletely returned. Moreover, the information supplied
by the informants may also be biased.

(e) An appeal to the vanity of the person questioned may give rise to yet another kind of bias.
For example, the question ‘Are you a good student?’ is such that most of the students would
succumb to vanity and answer ‘Yes;

(ii) Bias due to Faulty Collection of Data. Any consistent error in measurement will give rise to
bias whether the measurement are carried out on a sample or on all the units of population.
The danger of error is, however, likely to be greater in sampling work, since the units measured
are often smaller. Bias may arise due to improper formulation of the decision, problem or
wrongly defining the population, specifying the wrong decision, securing an inadequate frame,
and so on. Biased observations may result from a poorly designed questionnaire, an ill-trained
interviewer, failure of a respondent’s memory, etc. Bias in the flow of data may be due to
unorganized collection procedure, faulty editing or coding of responses.

(iii) Bias in Analysis. In addition to bias which arises from faulty process of selection and faulty
collection of information, faulty methods of analysis may also introduce bias. Such bias can be
avoided by adopting the proper methods of analysis.
[Link].2 Avoidance of Bias

If possibilities of bias exist, fully objective conclusions cannot be drawn. The first essential of
any sampling or census procedure must, therefore, be the elimination of all sourced of bias. The
simplest and the only certain way of avoiding bias in the selection process is for the sample to
be drawn either entirely at random, or at random subject to restrictions which, while improving
the accuracy, are of such a nature that they do not introduce bias in the results. In certain
cases, systematic selection may also be permissible.
[Link].3 Method of Reducing Sampling Errors

46
Once the absence of bias has been ensured, attention should be given to the random sampling
errors. Such errors must be reduced to the minimum so as to attain the desired accuracy.

Apart from reducing errors of bias, the simplest way of increasing the accuracy of a sample is to
increase its size. The sampling error usually decreases with increase in sample size (number of
units selected in the sample) and in fact in many situations the decrease is inversely
proportional to the square root of the sample size as can be seen from the diagram below.

From this diagram it is clear that though the reduction in sampling error is substantial for initial
increases in sample size, it becomes marginal after a certain stage. In other words, considerably
greater effort is needed after a certain stage to decrease the sampling error than in the initial
instances. Hence after that stage sizeable reduction in cost can be achieved by lowering even
slightly the precision required. From this point of view, there is a strong case for resorting to a
sample survey to provide estimates within permissible margins of error instead of a complete
enumeration survey, as in the latter the effort and the cost needed will be substantially higher
due to the attempt to reduce the sampling error to zero.

As regards non-sampling errors they are likely to be more in case of complete enumeration
survey than in case of a sample survey, since it is possible to reduce the non-sampling errors to
a greater extent by using better organization and suitably trained personnel at the field and
tabulation stages. The behavior of the non-sampling errors with increase in sample size is likely
to increase with increase in sample size. In many situations, it is quite possible that the non-
sampling error in a complete enumeration survey is greater than both the sampling and non-
sampling errors taken together in a sample survey, and naturally in such situations the latter is
to be preferred to the former.
[Link] Non-sampling Errors

When a complete enumeration of units in the universe if made, one would expect that it would
give rise to data free from errors. However, in practice it may not be so. For example, it is
difficult to completely avoid errors of observation or ascertainment. So also in the processing of
data, tabulation errors may be committed affecting the final results. Errors arising in this
manner are termed non-sampling errors, as they are due to factors other than the inductive
process of inferring about the population from a sample. Thus, the data obtained in an
investigation by complete enumeration, although free from sampling error, would still be

47
subject to non-sampling error, whereas the results of a sample survey would be subject to
sampling error as well as non-sampling error.

Non-sampling errors can occur at every stage of planning and execution of the census or
survey. Such errors can arise due to a number of causes such as defective methods of data
collection and tabulation, faulty definition, incomplete coverage of the population or sample,
etc. More specifically, non-sampling errors may arise from one or more of the following factors:

1. Data specification being inadequate and inconsistent with respect to the objective of the
census or survey.

2. Inappropriate statistical unit.

3. Inaccurate or inappropriate methods of interview, observation or measurement with


inadequate or ambiguous schedules, definitions or instructions.

4. Lack of trained and experienced investigators.

5. Lack of adequate inspection and supervision of primary staff.

6. Errors due to non-response, i.e., incomplete coverage in respect of units.

7. Errors in data processing operations such as coding, punching, verification, etc.

8. Errors committed during presentation and printing of tabulated results.

These sources are not exhaustive, but are given to indicate some of the possible sources of
error. In a sample survey, non-sampling errors may also arise due to defective frame and faulty
selection of sampling units.
[Link].1 Control of Non-Sampling Errors

In some situations the non-sampling errors may be large and deserve greater attention than the
sampling errors. While, in general sampling errors decrease with increase in sample size, non-
sampling errors tend to increase with the sample size. In the case of complete enumeration
non-sampling errors and in the case of sample surveys both sampling and non-sampling errors
require to be controlled and reduced to a level at which their presence does not vitiate the use
of final results.
3.1.15 Reliability of samples

The reliability of samples can be tested in the following ways:

1. Different samples of the same size should be taken from the same universe and their results
be compared. If results are similar, the sample will be reliable.

48
2. If the measurements of the universe are known, then they should be compared with the
measurements of the sample. In case of similarity of measurements, the sample will be reliable.

3. Sub-samples should be taken from the sample drawn. If the results of sample and sub-
samples are similarity, the sample may be considered as reliable.
Unit 4: Classification and Tabulation of Data

Learning objective
The necessary steps to be taken for different classification for the data are illustrated with
examples.
The method of grouping the raw data into frequency tables are explained for analysis purpose.
Chapter 1: Classification of Data
4.1.1 Objectives of classification

The principal objectives of classifying the data are :

1. To condense the mass of data in such a manner that similarities and dissimilarities can
be readily apprehended. Millions of figures can thus be arranged in a few classes having
common features.
2. To facilitate comparison.
3. To pointout the most significant features of the data at a glance.
4. To focus the important information collected.
5. To enable a statistical treatment of the material collected.
4.1.2 Types of classification

Broadly, the data can be classified on the following four bases :

1. Georgraphical, i.e. area-wise, eg. Cities, districts, etc.


2. Chronological, i.e. on the basis of time.
3. Qualitiative, eg. According to some attributes.
4. Quantitative i.e. in terms of magnitudes.
Chapter 2: Tabulation of data or frequency Rules
4.2.1 Formation of a discrete frequency distribution

The process of preparing this type of distribution is very simple. We have just to count the
number of times a particular value is repeated which is called the frequency of that class. In
order to facilitate counting prepare a column of “tallies”. In another column, place all possible
values of variable from the lowest to the highest. Then put a bar (vertical line) opposite the
particular value to which it relates. To facilitate counting, blocks of five bars are prepared and
some space is left in between each block. We finally count the number of bars and get
frequency.

The process shall be clear from the following examples :

49
Illustration 1 : In a survey of 35 families in a village, the number of children per family was
recorded and the following data was obtained :
1 0 2 3 4 5 6

7 2 3 4 0 2 5

8 4 5 12 6 3 2

7 6 5 3 3 7 8

9 7 9 4 5 4 3

Represent the data in the form of a discrete frequency distribution.

FREQUENCY DISTRIBUTION OF THE NUMBER OF CHILDREN

It is clear from the table that the number of children varies from 0 to 12. There were 2 families
with no child, 5 families with 4 children each only one family with 12 children.

Illustration 2 : Count the number of letters in each word of the para given below (ignoring
comma, full-stop, etc.) and prepare a discrete frequency distribution.

Today, to a very striking degree, our culture has become a statistical culture. Even a person who
may never have heard of an index number, is attached in an intimate fashion by the gyrations
of those index numbers which describe the cost of living.
No. of Letters Tallies Frequency
1 III 3

2 IIII 9

3 I 6

4 IIII 4

5 II 7

6 IIII 5

7 IIII 4

8 I 4

9 I 1

50
10 1

This method of classifying helps in condensing the data only where values are largely repeated,
otherwise hardly any condensation will be done.
4.2.2 Formation of continuous frequency distribution

This type of classification is most popular in practice. The following technical terms are
important when a continuous frequency distribution is formed or data are classified according
to class-intervals
[Link] Class limits

The class limits are the lowest and the highest values that can be included in the class. For
example, take the class 20-40. The lowest value of the class is 20 and the highest is 40. The two
boundaries of class are known as the lower limit and the upper limit of the class. The lower limit
of a class is the value below which there can be no item the class. The upper limit of a class is
the value above which no item can belong to that class. Of the class 70-89, 70 is the lower limit
and 89 is the upper limit, i.e. in this class there can be no value which is less than 70 or more
than 89. Similarly, if we take the class 90 – 109, there can be no value in that class which is less
than 90 or more than 109.
[Link] Class intervals
The difference between the Upper and Lower limit of a class is known as class interval of that
class. For example, in the class 100-200, the class interval is 200-100 = 100. An important
decision while constructing a frequency distribution is about the width of the class interval i.e.
whether it should be 10,20,50,100,500 etc. The decision would depend upon a number of
factors such as the range in the data, i.e. the difference between the smallest and largest item,
the details required and number of classes to be formed, etc. A simple formula to obtain the
estimate of appropriate class interval, i.e. is

i=
where, L = largest item
S = smallest item
K = the number of classes.
For example, if the salary of 100 employees in a commercial undertaking varied between Rs.500
and Rs.5,500 and we want to form 10 classes, then the class interval would be

i=
L = 5,500, S = 500, k = 10

i= = = 500
The starting class would be 500 – 1000, the next 1000-1500 and so on.

51
The question now is how to fix the number of classes, i.e. The number can be either fixed
arbitrarily keeping in view the nature of problem under study or it can be decided using
Sturges’ Rule. According to this rule, number of classes can be determined by the formula :
k = 1 + 3.322 log N
where N = total number of observations and
Log = logarithms of the number
Thus, if 10 observations are being studied, the number of classes shall be :
k = 1 + (3.322 x 1) = 4.322 or 4.
and if 100 observations are being studied, the number of classes shall be
k = 1 + (3.322 x 2) = 1 + 6.644 = 7.644 or 8
It should be noted that since log is used in the formula, the number of classes shall generally be
between 4 and 20 – it cannot be less than 4 even if N is less than 10 and if N is 10 lakh k will be
1+(3.322x6) = 19.932 or 20.
Sturges suggested the following formula for determining the magnitude of class interval :

i=
Where Range is the difference between the largest and smallest items.
For example, if in the above illustration we apply this formula the magnitude of class interval
shall be :

i= = =654.1 or 650
If we take a class interval of 650, the number of classes formed would be 5000/650, i.e. 7.69 or
8.
It may be noted that the application of above formula may give a value involving fractions and
odd intervals. For example, we got i=6.541. In such cases suitable approximation should be
made.
[Link] Class Frequency

The number of observations corresponding to a particular class is known as the frequency of


that class or the class frequency. In the following illustration, the frequency of the class 1000-
1100 is 50 which implies that there are 50 persons having income between Rs.1000 and
Rs.1100. If we add together the frequencies of all individual classes, we obtain the total
frequency. Thus, in the same problem, the total frequency of the six classes is 550 which means
that in all there are 550 persons whose income has been studied.

Class Mid-point or Class Mark : It is the value lying half-way between the lower and upper class
limits of a class-interval. Mid-point of a class is ascertained as follows :

Mid-point of a class=

For the purpose of further calculations in statistical work the mid-point of each class is taken to
represent that class.

52
There are two methods of classifying the data according to classes, namely (i) ‘exclusive’
method and (ii) ‘inclusive’ method.
[Link] ‘Exclusive’ Method

When the class intervals are so fixed that the upper limit of one class is the lower limit of the
next class it is known as the exclusive method of classification. The following data are classified
on this basis :
Income (Rs.) No. of Persons
1000 – 1100 50

1100 – 1200 100

1200 – 1300 200

1300 – 1400 150

1400 – 1500 40

1500 – 1600 10
TOTAL 550

It is clear that the ‘exclusive’ method ensures continuity of data in as much as the upper limit of
one class is the lower limit of the next class. Thus, in the above example, there are 50 persons
whose income is between Rs.1000 and Rs.1099.99. A person whose income is Rs.1100 would be
included in the class 1100-1200. This method is widely followed in practice. However, it is
confusing to a layman who has no knowledge of statistics. For example, if a questionnaire
includes an item asking the respondent the number of times he visits the Super Bazar in a
month and he is required to tick one of the categories : 5-10, 10-15 and 15-20, a person who
visits the Super Bazar 10 times may 5-10 or 10-15. In the absence of any specific instructions,
some people may tick the class 5-10 while others 10-15. Hence whenever this method is used it
is necessary to give clear instructions in the questionnaire. However, the reader should note
that if class intervals are given like 0-10, 10-15, etc. it is always presumed that upper limit is
exclusive, i.e. the item of that value is not included in that class.

A better way of expressing the classes when exclusive method is followed is :


Income (Rs.) No. of Persons
1000 but under 1100 50

1100 but under 1200 100

1200 but under 1300 200

1300 but under 1400 150

53
1400 but under 1500 40

1500 but under 1600 10


TOTAL 550

It avoids confusion of the type when classes are expressed 1000-1100, 1100-1200, 1200-1300,
etc. It is suggested that in practice this approach should be preferred over the previous one.
[Link] ‘Inclusive’ Method

Under the ‘inclusive’ method of classification, the upper limit of one class is included in that
class itself. The following example, illustrates the method :
Income (Rs.) No. of Persons
1000-1099 50

1100-1199 100

1200-1299 200

1300-1399 150

1400-1499 40

1500-1599 10
TOTAL 550

In the class 1000-1099 we include persons whose income is between Rs.1000 and Rs.1099. If
the income of a person is exactly Rs.1100 he is included in the next class. The above example
makes it clear that there is no confusion here of the type we find under the ‘exclusive’ method.
We may have classes like 1000-1099.5 or 1000-1099.99, and so on.

To decide whether to use the inclusive or the exclusive method it is important to determine
whether the variable under observation is a continuous or discrete one. In case of continuous
variables the upper limit exclusive method must be used. For example, the variable height
being inherently a continuous one should be stated as 60” and under 62”,62" and under 64”,
and so on. The inclusive method should, in general, be used in case of discrete variables. Thus,
in classifying factories according to number of workers, the limits should be stated as, for
example, 100-199 employees, 200-299 employees and not 100-200, 200-300, etc.
4.2.3 Considerations in the construction of Frequency distributions
It is difficult to lay down any hard and fast rules for constructing a frequency distribution, since
it depends on the nature of the given data and the object of classification.
However, the following general considerations may be borne in mind for ensuring meaningful
classification of data :

54
(1) The number of classes should preferably be between 5 and 20. However, there is no
rigidity about it. The classes can be more than 20 depending upon the total number of items in
the series and the details required, but they should not be less than five because in that case
the classification may not reveal the essential characteristics. The choice of number of classes
basically depends upon :
(a) the number of figures to be classified.
(b) the magnitude of the figures
(c) the details required, and
(d) ease of calculation for further statistical work.
(2) As far as possible one should avoid values of class-intervals, as 3,7, 11, 26, 39, etc.
Preferably, one should have class-intervals of either five or multiples of 5 like 10,20,25, 100, etc.
The reason is that the human mind is accustomed more to think in terms of certain multiples of
5,10 and the like. However, where the data necessitate a class-interval of less than 5 it can be
any value between 1 and 4.
(3) The starting point, i.e. the lower limit of the first class, may either be zero or 5 or multiple of
5. For example, if the lowest value of the data is 63 and we have taken a class-interval of 10,
then the first class can be 60-70, instead of 63-73. Similarly, if the lowest value of the data is 76
and the class interval is 5 then the first class can be 75 to 80 rather than 76 to 81.
(4) To ensure continuity we should follow ‘exclusive’ method of classification. However, if
‘inclusive’ method has been adopted it is necessary to adjust the class limits between two
classes to have continuity. The adjustment consists of finding the difference between the lower
limit of the second class and the upper limit of the first class, dividing the difference by two,
subtracting the value so obtained from all lower limits and adding the value to all upper limits.
This can be expressed in the form of a formula as follows :
Correction factor=

How the adjustment is made when data are given by inclusive method can be seen from the
following examples :
Monthly wages (in No. of workers Monthly Wages (in No. of workers
Rs.) Rs.)
800-899 5 1100-1190 8
900-999 10 1200-1299 2
1000-1099 15
To adjust the class limits, we take here the difference between 900 and 899, which is one. By
dividing it by two we get ½ or 0.5. This (0.5) is called the correction factor. Deduct 0-5 from the
lower limits of all classes and add 0.5 to upper limits. The adjusted classes would then be as
follows :
Monthly wages (in No. of workers Monthly Wages (in No. of workers
Rs.) Rs.)
799.5-899.5 5 1099.5-1199.5 8
899-.5-999.5 10 1199.5-1299.5 2
999.5-1099.5 15

55
It should be noted that before adjustment the class-interval was 99 but after adjustment, it is
100. Observe another case :
Variable Frequency
5-9.5 8
10-14.5 10
15-19.5 2

The correction factor here is = =0.25


After adjustment the classes will be :
Variable Frequency
4.75-9.75 8
9.75-14.75 10
14.75-19.75 2
The class-interval now is 5 and not 4.5. Taking a third example, if the class limits are
Variable Frequency
5-9.99 8
10-14.99 10
15-19.99 2

The correction factor would be= = =0.005


After adjustment the classes will become :
Variable Frequency
4.995-9.995 8
9.995-14.995 10
14.995-19.995 2
(5) Wherever possible, it is desirable to use class intervals of equal sizes because comparisons
of frequencies among classes are facilitated and subsequent calculations from the distribution
are simplified. However, this is not always a practical procedure. For example, in case of data
on monthly income of families, in order to show the details for the portion of frequency
distribution where the majority of incomes lie, class intervals of 100 or 200 may be used
starting say from 600-700 onwards upto about 1200 then intervals of 400 to 500 may be used
upto 2500 or so and a final class of 2500 and above may be shown for the relatively small
number of families having these highest incomes. It is obvious that if we have an equal class
intervals were used, say 500, too many families would be lumped together in the first one or
two classes, and the information how these incomes were distributed would be lost. To resolve
this dilemma at times we use open-end interval along with equal or inequal class intervals. The
use of unequal class sizes and open-end intervals generally becomes necessary in cases where
most of data are concentrated within a certain range, where gaps appear in which relatively
few items are observed and where there are very few extremely small or large values.
(6) Open-end distribution presents problems of graphing and further analysis. When the
frequency distribution is being employed as the only technique of presentation, open-end
classes do not seriously reduce its usefulness as long as only a few items fall in these classes.
However, use of the distribution for purposes of further mathematical computation is difficult

56
because a mid-point value, which can be used to present, the class, cannot be determined for
an open-end class.
(7) In any frequency distribution the size of items or the value are indicated on the left-hand
side and the number of times the items in those sizes or values have repeated are indicated by
frequencies on the right-hand side corresponding to the respective size or values.
Unit 5: Diagrammatic and graphic presentation

Learning objectives
For easy observation, the method of drawing diagrams and graphs shown with examples.
Chapter 1 . Diagrammatic Presentation
5.1.1 Significance of diagrams and graphs
Diagrams and graphs are extremely useful because of the following reasons :
1. They give a bird’s eye view of the entire data and, therefore, the information
presented can be easily understood. It is a fact that as the number and magnitude of figures
increases they become more confusing and their analysis tends to be more strenuous. Pictorial
presentation helps in proper understanding of the data as it gives an interesting form to it. The
old saying ‘A picture is worth 10,000 words’ is very true. The mind through the eye can more
readily appreciate the significance of figures in the form of pictures than it can follow the
figures themselves.
2. They are attractive to the eye. Figures are dry but diagrams delight the eye. For this
reason diagrams create greater interest than cold figures. Thus, while going through journals
and newspapers, the readers generally skip over the figures but most of them do look at the
diagrams and graphs. Since diagrams have attraction value, they are very popular in exhibitions,
fairs, conferences, board meetings and public functions.
3. They have a great memorizing effect. The impressions created by diagrams last much
longer than those created by the figures presented in a tabular form.
4. They facilitate comparison of data relating to different periods of time of different
regions. Diagrams help one in making quick and accurate comparison of data. They bring out
hidden fact and relationship and can stimulate as well as aid analytical thinking and
investigation.

5.1.2 Comparison of Tabular and Diagrammatic Presentation


Data may be presented in the form of tables as well as diagrams and graphs. Both forms of
presentation have their own usefulness for particular purposes. Hence, the choice of the form
of presentation must be made with due thought and care. The following points may be kept in
view in this connection :
1. Tables contain precise figures whereas diagrams give only an approximate idea. Exact
values can be read from a table.
2. More information can be presented in one table than either in one graph or diagram.
3. Tables usually require much close reading and are more difficult to interpret than
diagrams.
4. Graphs and diagrams have a visual appeal and, therefore, prove to be more impressive
to laymen.

57
5.1.3 Difference between Diagrams and Graphs

Though there is no clear-cut line of demarcation between the two, following points of
difference may be noted.

1. For constructing a graph we generally make use of graph paper whereas a diagram is
generally constructed on a plain paper. In other words, a graph represents
mathematical relationship (though not necessary functional) between two variables
whereas a diagram does not.
2. Diagrams are more attractive to the eyes and as such they suitable for publicity and
propaganda. They do not add anything to the meaning of the data and, therefore, from
the point of view of a statistician or research workers they are not helpful in analysis.
Graphs, on the other hand, are very much used by the statistician and the research
worker in analysis.
3. For representing frequency distributions and time series, graphs are more appropriate
than diagrams. In fact, for presenting frequency distributions diagrams are rarely used.
5.1.4 General rules for constructing diagrams

The following general rules should be observed while constructing diagrams :

1. Title : Every diagram must be given a suitable title. The title should convey in as few a
words as possible the main idea that the diagrams intent to portray. However, the brevity
should not be secured at the cost of clarity or omission of essential details. The title may be
given either at the top of the diagram or below it.

2. Proportion between width and height. A proper proportion between the height and
width of the diagram should be maintained. If either the height and width is too short or
too long in proportion, the diagram would given an ugly look. While there are no fixed rules
about the dimensions, a convenient standard as suggested by Lutz in the book entitled
“Graphic Presentation” may be adopted for general use. It is known as “Root-two”, that is, a
ratio of 1 (short side) to 1.414 (long side). Modifications may, no doubt, be made to
accommodate a diagram in the space available.

3. Selection of scale : The scale showing the values may be in even numbers or in multiples
of five or ten, eg. 25, 50, 75, or 20, 40, 60. Odd values like 1, 3, 5, 7 may be avoided.

4. Footnotes. In order to clarify certain points about the diagram, footnote may be given at
the bottom of the diagram.

5. Index. An index illustrating different types of lines or different shades, colours should be
given so that the reader can easily make out the meaning of the diagram.

6. Neatness and cleanliness. Diagrams should be absolutely neat and clean.

58
7. Simplicity. Diagrams should be as simple as possible so that the reader can understand
their meaning clearly and easily. For the sake of simplicity, it is important that too much
material should not be loaded in a single diagram otherwise it may become too confusing
and prove worthless. Several simple charts are often better and more effective than one or
two complex ones which may present the same material in a confusing way.
5.1.4 General rules for constructing diagrams

The following general rules should be observed while constructing diagrams :

1. Title : Every diagram must be given a suitable title. The title should convey in as few a
words as possible the main idea that the diagrams intent to portray. However, the brevity
should not be secured at the cost of clarity or omission of essential details. The title may be
given either at the top of the diagram or below it.

2. Proportion between width and height. A proper proportion between the height and
width of the diagram should be maintained. If either the height and width is too short or
too long in proportion, the diagram would given an ugly look. While there are no fixed rules
about the dimensions, a convenient standard as suggested by Lutz in the book entitled
“Graphic Presentation” may be adopted for general use. It is known as “Root-two”, that is, a
ratio of 1 (short side) to 1.414 (long side). Modifications may, no doubt, be made to
accommodate a diagram in the space available.

3. Selection of scale : The scale showing the values may be in even numbers or in multiples
of five or ten, eg. 25, 50, 75, or 20, 40, 60. Odd values like 1, 3, 5, 7 may be avoided.

4. Footnotes. In order to clarify certain points about the diagram, footnote may be given at
the bottom of the diagram.

5. Index. An index illustrating different types of lines or different shades, colours should be
given so that the reader can easily make out the meaning of the diagram.

6. Neatness and cleanliness. Diagrams should be absolutely neat and clean.

7. Simplicity. Diagrams should be as simple as possible so that the reader can understand
their meaning clearly and easily. For the sake of simplicity, it is important that too much
material should not be loaded in a single diagram otherwise it may become too confusing
and prove worthless. Several simple charts are often better and more effective than one or
two complex ones which may present the same material in a confusing way.
5.1.6. One-dimensional or Bar diagrams

Bar diagrams are the most common type of diagrams used in practice. A bar is a thick line
whose width is shown merely for attention. They are called one-dimensional because it is
only the length of the bar that matters and not the width. When the number of items is

59
large, lines may be drawn instead of bars to economise space. The special merits of bar
diagrams are the following :

(i) They are readily understood even by those unaccustomed to reading charts or those who
are not chart-minded.

(ii) They posses the outstanding advantage that they are the simplest and the easiest to
make.

(iii) When a large number of items are to be compared they are the only form that can be
used effectively.

While constructing bar diagrams the following points should be kept in mind.

(i) The width of the bars should be uniform throughout the diagram.

(ii) The gap between one bar and another should be uniform throughout.

(iii) Bars may be either horizontal or vertical. The vertical bars should be preferred. Figures
at the end of each bar so that the reader can know the precise value without looking at the
scale. This is particularly so where the scale is too narrow, for example, 1” on paper may
represent 10 crore people.

Types of Bar Diagrams :

Bar diagrams are of the following type :

(a) Simple bar diagrams

(b) Sub-divided bar diagrams

(c) Multiple bar diagrams

(d) Percentage bar diagrams

(e) Deviation bars


[Link] Simple Bar Diagrams

A simple bar diagram is used to represent only one variable. For example, the figures of
sales, production, population, etc. for various years may be shown by means of a simple bar
diagram. Since these are of the same width and only the length varies, it becomes very easy
for the reader to study the relationship. Simple bar diagrams are very popular in practice.
This can be either vertical or horizontal. In practice, vertical bars are more popular.
However, an important limitation of such diagrams is that they can present only one

60
classification or one category of data. For example, while presenting the population for the
last five decades, one can only depict the total population in the simple bar diagrams, and
not its sex-wise distribution.
[Link] Sub-divided Bar Diagrams

In a sub-divided bar diagram each bar representing the magnitude of a given phenomenon is
further sub-divided according to its various components. Each component occupies a part of
the bar proportional to its share in the total.

For example, the number of students in various courses for [Link]., [Link]., B.A., M.A., in
various colleges may be represented by a sub-divided bar diagram. While constructing such a
diagram, the various components is that of presenting each bar in the same order. A common
and helpful arrangements is that of presenting each bar in the order of magnitude from the
largest component at the base of the bar to the smallest at the end. To distinguish between the
different components, it is useful to use different shades or colours. Index or key should be
given explaining these differences.

Sub-divided bar diagrams should not be used where the number of components is more than
10 or 12, for, in that case, the diagram will be overloaded with information which cannot be
easily compared and understood.

The component bar diagram can be used to represent either the absolute data or distribution
ratios such as percentage distributors. It is, in fact, an excellent method for presenting a set of
distribution ratios diagrammatically. The subdivided bar diagrams can be constructed both on
horizontal and vertical bases.
[Link] Multiple bars
In a multiple bar diagram two or more sets of interrelated data are represented. The technique
of drawing such a diagram is the same as that of simple bar diagram. The only difference is that
since more than one phenomenon is represented, different shades, dots or crosses are used to
distinguish between the bars. Wherever a comparison between two or more related variables is
to be made, multiple bar diagram may be preferred
[Link] Percentage bars

Percentage bars are particularly useful in statistical work which require the portrayal of relative
changes in data. When such diagrams are prepared, the length of the bars is kept equal to 100
units and segments are cut in these bars to represent the components (percentages) of an
aggregate.
5.1.7 Two-dimensional diagrams

As distinguished from one-dimensional diagrams in which only the length of the bar is taken
into account, in two-dimensional diagrams the length as well as the width of the bars is
considered. Thus the area of the bars represents the given two-way classified data. Two

61
dimensional diagrams are also known as surface diagrams or area diagrams. The important
types of such diagrams are :

a) Rectangles

b) Squares and

c) Circles
[Link] Rectangles

This form is quite popular. Since the area of a rectangle is equal to the product of its length and
width, while constructing such diagram both length and width are considered. When two set of
figures are to be represented by rectangles, either of the two methods may be adopted. We
may represent the figures as they are given or may convert them to percentage and then
subdivide the length into various components. The latter method is more popular than the
former as it enables comparison to be made on a percentage basis.
[Link] Squares

The rectangular method of diagrammatic presentation is difficult to use where the values of
item vary widely. For example, if in the illustration given above the number of units sold of
commodity A and B are 20 and 240 respectively, the width of the rectangles would be in the
ratio of 5:60 or 1:12. If this ratio, is taken, the diagram would look very unwidely. It is order to
overcome this difficulty that squares are used.

The method of drawing a square diagram is very simple. One has to take the square-root of the
values of various items that are to be shown in the diagrams and then select a suitable scale to
draw the squares.
[Link] Circles

Another way of preparing a two-dimensional diagram is in the form of circles. In such diagrams
both the total and the component parts or sector can be shown. The area of a circle is
proportional to the square of its radius. As in the construction of squares, the square-roots of
various figures are worked out while constructing the circle. However, in the latter case the
radii of the circles can be obtained by dividing the value of pie and taking square-root.

Circles can be used in all those cases in which squares are used. However, in both these types of
diagram it is difficult to judge the relative magnitudes with precision.

Circles are difficult to compare and as such are not very popular in statistical work. When it is
necessary to use circles, they should be compared on an area basis rather than on a diameter
basis, as the diameter basis is very misleading. Compared to rectangles, circles are more
difficult to construct and interpret.
5.1.8 Pie diagram

62
Pie diagrams are very popularly used in practice to show percentage breakdowns. For example,
with the help of a pie diagram we can show how the expenditure of the Government is
distributed over different heads like Agriculture, Irrigation, Industry, Transport, Defence, etc.
Similarly through a pie diagram we can show how the expenditures incurred by an industry are
divided under different heads like raw materials, wages and salaries selling and distribution
expenses, etc. The pie chart is so called because the entire graph looks like a pie, and the
components resemble slices cut from pie.

While making comparisons, pie diagrams should be used on a percentage basis and not on an
absolute basis, since a series of pie diagrams showing absolute figures would require that larger
totals by represented by larger circle. Such presentation involves difficulties of two-dimensional
comparisons. However, when pie diagrams are constructed on a percentage basis, percentages
can be presented by circles equal in size. It may be noted that this problem does not arise in the
use of a single pie diagram.

In laying out the sectors for pie chart, it is desirable to follow some logical arrangement or
sequence. It is common practice to begin the largest component sector of a pie diagram at 12
O’clock position on the circle. Usually the other component sectors are placed in clockwise
succession in descending order of magnitude, except for catch-all components like
‘miscellaneous’ and ‘all others’ which are shown last, contrast with adjacent sectors.

In constructing a pie chart the first step is to prepare the data so that the various component
values can be transposed into corresponding degrees on the circle. Suppose there are four
components in a series representing the following values : (i) 60 percent, (ii) 25 percent, (iii) 15
percent(360/100=3.6),(iv) 5 percent the corresponding values of the four components in the
illustration are (60) x (3.6) = 216; (25) x (3.6) = 90; (10) x (3.6) = 36; (5) x (3.6) = 18.

The second step is to draw a circle of appropriate size with a compass. The size of the radius
depends upon the available space and other factors of presentation.

The third step is to measure points on the circle representing the size of each sector with the
help of a protractor. The ordinary protractor is based upon a scale in which the total circle is
360 degrees, but it is possible to purchase a protractor in which the entire circle is divided not
into 360 but 100 equal parts so that the angle representing any desired percentage can be read
directly.

In laying out the sectors for a pie chart it is desirable to follow some logical arrangement,
pattern or sequence. For example, it is a common procedure to arrange the sectors according
to size, with the largest at the top others in sequence running clockwise. An essential feature of
the pie chart is the careful identification of each sector with some kind of explanatory or
descriptive label. If there is sufficient room, the labels can be placed inside the sectors;
otherwise the labels should be placed in contiguous positions outside the circle, usually with an
arrow pointing towards the appropriate sector.
[Link] Limitation of Pie Diagrams

63
Pie diagrams are at times less effective than bar diagrams for accurate reading and
interpretation, particularly when series is divided into a large number of components or the
difference among the components is very small. It is generally inadvisable to attempt to portray
a series of more than five or six categories by means of a pie chart. If, for example, there are
eight, ten or more categories it may be very confusing to differentiate the relative values
portrayed specially when several small sectors are of approximately the same size. This type of
diagram, although frequently used, appears upon comparison inferior to simple bar diagram,
the divided bar diagram or a group of curves.
5.1.9 Pictograms

Pictograms are very popularly used in presenting statistical data. They are not abstract
presentations such as lines or bars but really depict the kind of data we are dealing with.
Pictures are attractive and easy to comprehend and as such this method is particularly useful in
presenting statistics to the layman. When pictographs are used data is represented
through a pictorial symbol that is carefully selected.

While constructing a pictograph the following points should be kept in mind :

The pictorial symbol should be self-explanatory. If we are telling a story about aeroplane, the
symbol should clearly indicate an aeroplane. The following points should be kept in mind while
selecting a pictorial symbol :

a) A symbol must represent a general concept (like man, woman, child, bus) not an individual of
the species (not Hitler, Akbar or [Link]’s car).

b) A symbol should be clear concise and interesting.

c) A symbol must be clearly distinguishable from every other symbol.

d) A symbol should suit the size of paper, ie. It should be neither too small nor too large.

e) Last, but not the least, an artist should use the principles of a good design established by the
fine and applied arts when drawing a pictorial symbol.

2) Changes in numbers are shown by more or fewer symbols, not by larger or smaller ones.

3) Pictographs should be simple to understand and convey essential facts.

Merit :

1) Compared with other types of diagrams, pictographs have greater attraction value
and, When the attention of masses is to be drawn such as in exhibitions, fairs, they are very
popularly used. They stimulate interest in the information being represented.

64
2) Facts portrayed in pictorial form are generally remembered longer than facts presented in
tables or in non-pictorial charts.

Limitations :

1) They are difficult to construct. Besides, it is necessary to use one symbol to represent a fixed
number of units which may create difficulties. Thus if one symbol is representing 10 crore
persons, the question is how to represent a population of 31.85 crores. In such a case either the
symbol should be proportionately smaller or the figure approximated to 30 crores. In either
case, error is introduced.

2) Pictographs give only an overall picture, they do not give minute details. For greater accuracy
we should write actual figure under, above or on one side of the symbol.
5.1.10 Cartograms

Cartograms or statistical maps are used to give quantitative information on a geographical


basis. They are thus used to represent spatial distributions. The quantities on the map can be
shown in many ways, such as through shades or colour, by dots, by placing pictograms in each
geographical unit and by placing the appropriate numerical figure in each geographical unit.

Statistical maps should be used only where geographic comparisons are of primary importance
and where approximate measures will suffice. For more accurate representation of size, bar
charts are preferable. To be sure, maps are sometimes combined, which are drawn in the
appropriate areas.
5.1.11 Choice of a Suitable diagram

Which diagram out of several ones to select in a given situation is a ticklish problem. The choice
would primarily depend upon two factors, namely : (i) the nature of the data; (ii) the type of
people for whom the diagram is meant. On the nature of the data would depend whether to
use one-dimensional, two-dimensional or three-dimensional diagram, and if it is one-
dimensional, whether to adopt the simple bar or subdivided bar, multiple bar or some other
type. As already stated, a cubic diagram would be preferred to a bar if the magnitudes of the
figures are very wide apart. The type of people for whom the diagram is intended must also be
considered. For example, for drawing attention of an uneducated mass, pictographs and
cartograms are more effective than cubes, circles, etc. Different types of diagrams such as bars,
rectangles, cubes, pictographs, pie charts, have specialised uses. However, bar diagrams are
the most popular in practice. There are different types of bars and the appropriate type of bar
chart can be choosen on the following basis :

a) Simple bar charts should be used where changes in totals are required to be conveyed.

b) Component bar charts are more useful where changes in totals as well as in the size of
components figures (absolute ones) are required to be displayed.

65
c) Percentage composition bar charts are better suited where changes in the relative size of
component figures are to be exhibited.

d) Multiple bar charts should be used where changes in the absolute values of the components
figures are to be emphasised and the overall total is of no importance.

However, multiple and component bar charts should be used only when there are not more
than three or four components, as a large number of components make the bar charts too
complex to enable worthwhile visual impression to be gained. When a large number of
components have to be shown, a pie chart is more suitable.

A pie chart is particularly useful where it is desired to show the relative proportions of the
figures that go to make up a single overall total. Unlike bar charts it is not restricted to three or
four component figures although its effectiveness tends to dwindle with more than seven or
eight components.

However, pie charts cannot be used effectively where a series of figures is involved, as a
number of different pie charts are not easy to compare. Nor should changes in the overall total
be shown by changing the size of the ‘pie’.

Occasionally, circles are used to represent size. But it is difficult to compare them and they
should not be used when it is possible to use charts. This is because it is easier to compare the
lengths of lines or bars then to compare areas or volumes.

Cubes should be used in those cases where the difference between the smallest and largest
values to be represented is very large. In other cases cubes should not be used because
comparison is too difficult with the help of cubes.

Pictographs and cartograms are very elementary form of visual presentation. However, they are
more informative and more effective than other forms for presenting data to the general public
who, by and large, neither possess much ability to understand nor take interest in the less
attractive forms of presentation. The pictograph is admirably suited to the illustrations of
exhibits or articles in newspapers and magazines or for dressing up annual reports. Cartograms
or statistical maps are particularly effective in bringing out the geographical pattern that may be
concealed to the data.
Chapter 2: Graphical Presentation of Data
5.2.1 Graphs

A large variety of graphs are used in practice. However, here we shall discuss only some
important types of graphs which are more popular. Broadly various graphs can be divided
under the following two heads :

1. Graphs of time series.


2. Graphs of frequency distributions.

66
Constructing charts and graphs is an art which can be acquired through practice. There are a
number of simple rules, adoption of which leads to the effectiveness of the graphs. However,
before discussing these rules, the elementary procedure of constructing a graph is considered.
5.2.2 Technique of Constructing Graphs

For constructing graphs, we make use of graph paper. Two simple lines are first drawn which
intersect each other at right angles. The lines are known as coordinate axes. The point of
intersection is known as the point of origin or the ‘zero’ point. The horizontal line is called the
axis of X or ‘abscissa’ and the vertical line the axis of Y or ‘ordinate’. The alternative appellations
are X-axis and Y-axis respectively.

O is the point of origin, XOX’ is the axis of X or the ‘abscissa’ and YOY’ the axis of Y or the
‘ordinate’. Both positive as well as negative values can be shown on the graphs. Distances
measured towards the right or upward from the origin are positive and those measured
towards the left or downwards are negative.

The whole plotting area is divided into four quadrants as shown above. In quadrant I, both the
values of X and Y are positive. In quadrant II, Y is positive, X is negative; in quadrant III, both X as
well as Y are negative and in quadrant IV, X is positive whereas Y is negative. Since most
business data are positive quadrant I is most frequently used.

It is conventional to take the independent variable on the horizontal scale and the dependent
on the vertical scale. In case of time series, time is represented on the horizontal sale and the
variable on the vertical scale. For each axis a convenient scale is chosen which represents the
unit of a variable. The choice is made in such a manner that the entire data is accommodated in
the space available. The scale on X-axis and Y-axis need not be identical.

On the arithmetic line graph, the Y scale must begin at zero as origin. Thus the X-axis always
runs through this zero origin. The zero line is the base line and the curve is interpreted in terms
of distance from this base line. In one special case, when we present graphically a series of
changes from a norm of 100% then the 100% line is considered the base line.

Once the scale is chosen equal space would represent equal amounts in case of natural scale.
However, in case of ratio scale, it is not so. No hard and fast rule can be laid down about the
ratio of the scale on the abscissa and on the ordinate because much would depend upon the
given data and the size of the paper. However, conventionally X-axis is taken 1 ½ times as long
as Y-axis. But there is no rigidity about it.

After the choice of the scale is made the last step in constructing a graph is to plot the given
data by taking the corresponding values of X and Y. The various points so obtained are then
joined by line segments.
5.2.3 Graphs of Time Series or Line Graphs

67
When we observe the values of a variable at different points of time, the series so formed is
known as time series. The technique of graphic presentation is extremely helpful in analysing
change at different points of time. On the X-axis we generally take the time and on the Y-axis
the value of the variable and join the various points by straight lines. The graph so formed is
known as the line graph. Such graphs are most widely used in practice. They are the simplest to
understand, easiest to make and most adaptable to many uses. They require the least technical
skill and at the same time enable one to present more information of a complex nature in a
perfectly understandable form than any other kind of chart. Many variables can be shown on
the same graph and a comparison can be made.

Graphs of time series can be constructed either on a natural scale or on a ratio scale. In natural
arithmetic scale absolute change from one period to another area shown whereas in a ratio
scale the rates of change or the relative changes are shown. First of all, we take up line graphs
on a natural scale and then study such graphs on ratio scale.
[Link] Rules for constructing the line graphs on Natural Scale

In constructing a graph of time series on natural scale the following points should be kept in
mind :

 Take the time on the X-axis (horizontal) and the variable on the Y-axis (vertical). The unit
of time in which the variable under consideration is measured should be clearly stated
in the title eg. an indication should be given as to whether the years are calendar or
financial or whether the variable is measured as at a date.
 Begin Y-axis with zero and select a suitable scale so that the entire data is
accommodated in the space available. On the arthmetic scale equal magnitude must be
represented by equal distances. This requirement is true for both the X-axis as well as
the Y-axis but for each separately. For example, I cm on Y-axis may represent 1,000 units
whereas 1” on X-axis may represent gap between 1993 and 1994. The scale should be so
chosen that the horizontal axis is longer than the vertical one. If the fluctuations in the
variable are too small or if the lowest value of the variable is large, the false base should
be used.
 Corresponding to the time factor plot the value of the variable and join the various
points by straight lines (and not with curves). The points on the graph should not be
indicated by circles or crosses rather dots should be used so that they disappear into
line.
 Join the various points with straight lines, not curves.
 If on one graph more than one variable is shown, they should be distinguished by the
use of thick, thin, dotted lines, etc. or different colours be used. Every graph should be
given a suitable title. The unit of time in which the variable under consideration is
measured should be clearly stated in the title, i.e. an indication should be given as to
whether the years are calendar or financial or whether the variable is measured as at a
date.

68
 Lettering on the graph, i.e. indication of years, units, etc. should be done horizontally
and not vertically so that in order to read what is written it is not necessary to turn the
graph from one side to another.

[Link] False Base Line

One of the fundamental rules while constructing graphs is that the scale on the Y-axis should
begin from zero. Where the lowest value to be plotted on the Y scale is relatively high and a
detailed scale is required to bring out he variations in all the data, starting the Y scale with zero
introduces difficulties. For example, if we have a series of production figures over a number of
years ranging from 15000 units to 25000 units, then starting with a zero origin would have one
of two undesirable consequences: either (i) the necessarily large intervals (say 5000 units) on
the Y scale would make us lose sight of the extent of fluctuations in the curve : (ii) a necessarily
large graph to permit small intervals (say 1000 units) would entail a waste of a large part of the
graph, in addition to poor visual communication.

The solution is to break the Y scale : If the zero origin is the shown then the scale is broken by
drawing a horizontal wavy line (also called kinked or zig-zag line) or a vertical wavy line
between zero and the first unit on the Y scale which in our illustration would be 15000 units.
These lines are drawn to make the reader aware of the fact that false base has been used.
Three important objects of false base line are :

1. Variations in the data are clearly shown

2. A large part of the graph is not wasted or space is saved by using false base.

3. The graph provides a better visual communication.


[Link] Graphs of one variable

When only one variable is to be represented, on the X-axis measure time and on the Y-axis the
value of the variable and plot the various points and join them by straight lines. The fluctuation
of this line shows the variation in the variable, and the distance of the plotting from the base
line of the graph indicates the magnitude
[Link] Graph of Two or More variables

If the unit of measurement is the same, we can represent two or more variables on the same
graph. This facilitaties comparison. However, when the number of variables is very large (say,
exceeding five or six) and they are all shown on the same graph, the chart becomes quite
confusing because different lines may cut each other and make it difficult to understand the
behaviour of the variables. Therefore, for the sake of clarity we should not represent more than
5 or 6 variables on the same graph. When two or more variables are shown on the same graph
it is desirable to use thick, thin, broken, dotted lines, etc. to distinguish between the various
variables.

69
[Link] Graphs having two scales

If two variables are expressed in two different units, then we will have two scales – one on the
left and the other on the right. To facilitate comparison, each scale is made proportional to the
respective average of each. The average values of both the variables are kept in the middle of
the graph, and then scales are determined.
5.2.4 Range chart

It is a very good method of showing the range of variation, ie. The minimum and maximum
values of a variable. For example, if we are interested in showing the minimum and maximum
price of a commodity for different periods of time or the minimum and maximum
temperatures, or the minimum and maximum price of shares of some company for different
periods, the range chart would be very appropriate. The above data can be best represented
through a range chart, the following are the steps in constructing such a chart : 1. Take time on
the X-axis and the variable on the Y-axis. 2. Draw two curves by plotting the given data – one
curve representing the highest values and the other one the lowest values. Curve A represents
lowest prices, whereas curve B highest prices. The gap between curve A and curve B represents
the range of variation. 3. For emphasing difference between the lowest and higher values the
use of colour or some shade, etc. should be made.
5.2.5 Graphs of frequency distributions

A frequency distribution can be presented graphically in any of the following ways :

1. Histogram
2. Frequency polygon
3. Smoothed frequency curve
4. Ogives or cumulative frequency curves.
[Link]. Histogram

Out of several methods of presenting a frequency distribution graphically, histogram is the


most popular and widely used in practice. A histogram is a set of vertical bars whose areas
are proportional to the frequencies represented.

While constructing histogram the variable is always taken on the X-axis and the frequencies
depending on it on the Y-axis. Each class is then represented by a distance on the scale that
is proportional to its class-interval. The distance for each rectangle on the X-axis shall
remain the same in case the class-intervals are uniform throughout. If they are different the
width of the rectangles shall also vary. The Y-axis represents the frequencies of each class
which constitute the height of its rectangle. In this manner we get a series of rectangles
each having a class-interval distance as its width and the frequency distance is its height.
The area of the histogram represents the total frequency as distributed throughout the
classes.

70
The histogram should be clearly distinguished from a bar diagram. The distinction lies in the
fact that wheres a bar diagram is one dimensional. Ie. Only the length of the bar is material
and not the width, a histogram is two-dimensional that is, in a histogram both the length as
well as the width are important.

The histogram is most widely used for graphical presentation of a frequency distribution.
However, we cannot construct a histogram for distribution with open-end classes.
Moreover, a histogram can be quite misleading if the distribution has unequal class-
intervals and suitable adjustments in frequencies are not made.

The technique of constructing histogram is given below (i) for distributions have equal class-
intervals, and (ii) for distributions having unequal class-intervals.

When class-intervals are equal, take frequency on the Y-axis, the variable on the X-axis and
construct adjacent rectangles. In such a case the height of the rectangles will be
proportional to the frequencies.

When class-intervals are unequal, a correction for unequal class intervals must be made.
The correction consists of finding for each class the frequency density or the relative
frequency density. The frequency density is the frequency for that class divided by the
width of that class. A histogram or frequency density polygon constructed from these
density values would have the same general appearance as the corresponding graphical
display developed from equal class-intervals.

For making the adjustment we take that class which has lowest class-interval and adjust the
frequencies of other classes in the following manner. If one class-interval is twice as wide as
the one having lowest class-interval we divided the height of its rectangle by two, if it is
three times more we divide the height of its rectangle by three, etc. i.e. the heights will be
proportional to the ratio of the frequencies of the width of the class.

Construction of Histogram when only Mid-points are given. When only mid-points are given,
ascertain the upper and lower limits of the various classes and then construct the histogram
in the same manner.
[Link]. Frequency polygon

A frequency polygon is a graph of frequency distribution. It has more than four sides. It is
particularly effective in comparing two or more frequency distributions. There are two ways
in which a frequency polygon may be constructed.

1. We may draw a histogram of the given data and then join by straight lines the mid-points
of the upper horizontal side of each rectangle with the adjacent ones. The figure so formed
is called frequency polygon. It is an accepted practice to close the polygon at both ends of
the distribution by extending them to the base line. When this is done two hypothetical
classes at each end would have to be included – each with a frequency of zero. This

71
extension is made with the objective of making the area under polygon equal to the area
under the corresponding histogram. The readers are advised to follow this practice.

2. Another method of constructing frequency polygon is to take the mid-points of the


various class-intervals and then plot the frequency corresponding to each point and to join
all these points by straight lines. The figures obtained would exactly be the same as
obtained by method No. 1. The only difference is that here we have not to construct a
histogram.

By constructing a frequency polygon the value of mode can be easily ascertained. If from
the apex of the polygon a perpendicular is drawn on the X-axis, we get the value of mode.
Moreover, frequency polygons facilitate comparison of two or more frequency distributions
on the same graph.
[Link].1 Frequency polygon has certain advantages over the histogram
1. The frequency polygons of several distributions may be plotted on the same graph,
thereby making certain comparisons possible, whereas histograms cannot be usually
employed in the same way. To compare histograms we must have a separate graph for
each distribution. Because of this limitation for purposes of making a graphic
comparison of frequency distributions, frequency polygons are preferred.
2. The frequency polygon is simpler than its histogram counterpart.
3. It sketches an outline of the data pattern more clearly.
4. The polygon becomes increasingly smooth and curve-like as we increase the number of
classes and the number of observations.
[Link]. Smoothed Frequency Curve

A smoothed frequency curve can be drawn through the various points of the polygon. The
curve is drawn freehand in such a manner that the area included under the curve is
approximately the same as that of the polygon. The object of drawing a smoothed frequency
curve is to eliminate as far as possible accidental variations that might be present in the data.
While smoothing a frequency polygon the fact that it is really derived from the histogram
should always be kept in mind. This would imply that the top of the curve would overtop the
highest point of the polygon particularly when the magnitude of class-interval is large. The
curve should look as regular as possible and sudden turns should be avoided. The extent of
smoothing would, however, depend upon the nature of the data. If it is a natural phenomenon
like tossing of coin, smoothing may be freely resorted to as such phenomenon normally has
symmetrical curves, but if the phenomenon is social or economic the curve is generally skewed
and as such smoothing cannot be carried too far.

For drawing a smoothed frequency curve it is necessary to first draw the polygon and then
smooth it out. As discussed earlier, the polygon can be constructed even without first
constructing a histogram by plotting the frequencies at the mid-points of class-intervals. This
may save some time but the smoothing of the polygon cannot be done properly without a
histogram. Hence, it is desirable to proceed in a sequence, i.e. first draw a histogram than a
polygon and lastly smooth it to obtain the smoothed frequency curve. This curve should begin
72
and end at the base line and as a general rule it may be extended to the mid-points of the class-
intervals just outside the histogram. The area under the curve should represent the total
number of frequencies in the entire distribution.

The following points should be kept in mind while smoothing a frequency curve :

1. Only frequency distribution based on samples should be smoothed.


2. Only continuous series should be smoothed.
3. The total area under the curve should be equal to the area under the original histogram
or polygon.
[Link]. Ogives or Cumulative Frequency curves

4. At times we are interested in knowing ‘how many workers of a factory earn less than
Rs.700 per month’ or ‘how many workers earn more than Rs.1000 per month’
‘percentage of students who have failed, etc. To answer these questions, it is necessary
to add the frequencies. When frequencies are added, they are called cumulative
frequencies. These frequencies are then listed in a table called a cumulative frequency
table. The curve obtained by plotting cumulative frequencies is called a cumulative
frequency curve or an Ogive (pronounced Ojive).
5. There are two methods of constructing Ogive, namely:
6. (a) The ‘less than’ method.
7. (b) The ‘more than’ method.
8. (a) ‘Less than’ method. In the ‘less than’ method we start with the upper limits of the
classes and go on adding the frequencies are plotted we get a rising curve.
9. (b) ‘More than’ method. In the ‘more than’ method we start with the lower limits of the
classes and from the frequencies we subtract the frequency of each class. When these
frequencies are plotted we get a declining curve.
[Link].1 Utility of Ogives

From the point of graphic presentation, the ogive is especially used for the following purposes :

1. To determine as well as to portray the number or proportion of cases above or below a


given value.
2. To compare two or more frequency distributions. Generally there is less overlapping
when comparing several ogives on the same grid than when comparing several simple
frequency curves in this manner.
3. Ogives are also drawn for determining certain value graphically such as median,
quartiles, deciles, etc.

Despite the great significance of ogives, it should be noted that they are not as simple to
interpret as one may feel and hence the reader must be careful while using them.

[Link] Limitations of Diagrams and Graphs

73
Although diagrams and graphs are powerful and effective media for presenting statistical data,
they are not under all circumstances and for all purposes complete substitute for tabular and
other forms of presentation. The well trained specialist in this field is one who recognise not
only the advantages but also the limitations of these techniques. He knows when to use and
when not to use these methods and from its repertoire is able to select the most appropriate
form for every purpose. Julin has beautifully said, “Graphic statistics has a role to play of its
own; it is not the servant of numberical statistics, but it cannot pretend, on the other hand, to
precede or displace the latter”.

The main limitations of diagrams and graphs are

1. They can present only approximate values.

2. They can approximately represent only limited amount of information.

3. They are intended mostly to explain quantitative facts to the general public. From the point
of view of the statistician, they are not of much help in analysing data.

4. They can be easily misinterpreted and, therefore, can be used for grinding one’s axe during
advertisement, propaganda and electioneering. As such as diagrams should never be accepted
without a close inspection of the bonafides because things are very often not what they appear
to be.

5. The two-dimensional diagrams and the three-dimensional diagrams cannot be accurately


appraised visually, and, therefore, as far as possible their use should be avoided.
Unit 6: Averages or Measures of Central Tendency or Measures of Location

Learning objectives
The reader is taught to calculate the different averages for the data. The merits and demerits
of averages and its uses are explained.
Chapter 1: Averages or Measures of Central Tendency or Measures of Location
6.1.1 Averages or Measures of Central Tendency or Measures of Location

According to Professor Bowley, averages are “statistical constants which enable us to


comprehend in a single effort the significance of the whole”. They give us an idea about the
concentration of the values in the central part of the distribution. Plainly speaking, an average
of a statistical series is the value of the variable which is representative of the entire
distribution. The following are the five measures of central tendency that are in common use:

i) Arithmetic Mean or simply Mean Iii) Median,

ii) Mode (iv) Geometric Mean, and Harmonic Mean


6.1.2 Requisites for an Ideal Measure or Central Tendency

74
According to Professor Yule, the following are the characteristics to be satisfied by an ideal
measure of central tendency:

i) It should be rigidly defined

ii) It should be readily comprehensible and easy to calculate

iii) It should be based upon all the observations

iv) It should be suitable for further mathematical treatment. By this we mean that if we are
given the averages and sizes of a number of series, we should be able to calculate the average
of the composite series obtained on combining the given series.

v) It should be affected as little as possible by fluctuations of sampling.

In addition to the above criteria, we may add the following (which is not due to Prof. Yule):

vi) It should not be affected much by extreme values


Chapter 2: Arithmetic Mean
6.2.1 Arithmetic Mean

Arithmetic mean of a set of observations is defined as their total divided by the number of
observations, e.g. the arithmetic mean of n observation x1, x . . . . xn is given by

In case of frequency distribution xi/fi, i= 1, 2, ….. n where fi is the frequency of the variable xi,

= = =
In case of grouped or continuous frequency distribution, xi is taken to be the mid-point of the
class-interval.
6.2.2 Merits and Demerits of Arithmetic Mean

Merits

i) It is rigidly defined

ii) It is easy to understand and easy to calculate

iii) It is based upon all the observations

75
iv) It is amenable to algebraic treatment. The mean of the composite series in terms of the
means and sizes of the component series is given by

v) Of all the averages, arithmetic mean is affected least by fluctuations of

sampling. This property is sometimes described by saying that mean is a stable

average.

Thus, we see that arithmetic mean satisfies all the properties laid down by Prof. Yule for an
ideal average.

Demerits

i) Arithmetic mean is affected very much by extreme values. In case of extreme items,
arithmetic mean gives a distorted picture of the distribution and no longer remains
representative of the distribution.

ii) Arithmetic mean may lead to wrong conclusions if the details of the data from which it is
computed are not given. Let us consider the following marks obtained by two students A and B
in three tests, viz, terminal test, half-yearly examination and annual examination respectively.

Marks in : I Test II Test III Test Average marks

A 50% 60% 70%

B 70% 60% 50%

Thus average marks obtained by each of the two students at the end of the year are 60%. If we
are given the average marks alone we conclude that the level of intelligence of both the
students at the end of the year is same. This is a fallacious conclusion since we find from the
data that student A has improved consistently while student B has deteriorated consistently.

iii) Arithmetic mean cannot be calculated if the extreme class is open, e.g. below 10 or above
70. Moreover, even if a single observation is missing mean cannot be calculated.
6.2.3 Weighted Mean

In calculating arithmetic mean we suppose that all the items in the distribution have equal
importance. But in practice this may not be so. If some items in a distribution are more
important than others, then this point must be borne in mind, in order that average computed

76
is representative of the distribution. In such cases, proper weightage is to be given to various
items – the weights attached to each item being proportional to the importance of the item in
the distribution. For example, if we want to have an idea about the change in cost of living of a
certain group of people, then the simple mean of the prices of the commodities consumed by
them will not do, since all the commodities are not equally important e.g. wheat, rice and
pulses are more important than cigarettes, tea, confectionery, etc.

Let wi be the weights attached to the ith item having the value xi, i = 1, 2 …….n. Then, we define.

Weighted arithmetic mean or weighted mean = ∑ i wi xi / ∑ i wi …..…..

It may be observed that the formula for weighted mean is the similar to the formula for simple
mean with fi, (i = 1,2……n) the frequencies replaced by wi, (i = 1,2,……n), the weights.

Weighted mean gives the result equal to the simple mean if the weights assigned to each of the
variate values are equal. It results in higher value than the simple mean if smaller weights are
given to smaller items and larger weights to larger items. If the weights attached to larger items
are smaller and those should items attached to highted mean, weighted mean results in
smaller value than the simple mean.
Chapter 3: Median
6.3.1 Median

Median of a data is the value which divides the data it into two equal parts. It is the value which
exceeds and is exceeded by the same number of observations, i.e. it is the value such that the
number of observations above it is equal to the number of observations below it. The median is
thus a positional average.

In case of ungrouped data, if the number of observations is odd then median is the middle
most value arranged in the ascending or descending order of magnitude. In case of even
number of observations, there are two middle terms and median is obtained by taking the
arithmetic mean of the middle most terms. For example, the median of the values 25, 20, 15,
35, 18 i.e. 15, 18, 20, 25, 35,25,20 and the median of 8, 20, 50, 25, 15, 30 i.e. of 8,15, 20, 25, 30,
50 is ½ (20+25) = 22.5.

Remark: In case of even number of observations, in fact any value lying between the two
middle values can be taken as median but conventionally we take it to be the mean of the
middle terms.

In case of discrete frequency distribution median is obtained by considering the cumulative


frequencies. The steps for calculating median are given below:

i) Find N/2, where N= ∑ fi

i) See the cumulative frequency (c.f) is either equal to just greater than N/2.

77
ii) The corresponding value of x is median

Example: Obtain the median for the following frequency distribution:

x:123456789

f : 8 10 11 16 20 25 15 9 6

Solution
x f c.f.
1 8 8
2 10 18
3 11 29
4 16 45
5 20 65
6 25 90
7 15 105
8 9 114
9 6 120
120

Here N = 120 and have N/2 = 60

Cumulative frequency (c.f.) just greater than N/2, is 65 and the value of x corresponding to 65 is
5. Therefore median is 5.

In the case of continuous frequency distribution, the class corresponding to the c.f. just greater
than N/2 is called the median class and the value of median is obtained by the following
formula:

Median = l +

l is the lower limit of the median class

f is the frequency of the median class

h is the magnitude of the median class

c is the c.f. of the class preceding the median class

and N= ∑ f
6.3.2 Merits and Demerits of Median

78
Merits:

i) It is rigidly defined

ii) It is easily understood and is easy to calculate. In some cases it can be located merely by
inspection.

iii) It is not at all affected by extreme values

iv) It can be calculated for distributions with open-end classes.

Demerits:

i) In case of even number of observations median cannot be determined exactly. We merely


estimate it by taking the mean of the two middle most terms.

ii) It is not based on all the observations. For example, the median of 10, 25, 50, 60 and 65 is 50.
We can replace the observations 10 and 25 by any two values which are smaller than 50 and
the observations 60 and 65 by any two values greater than 50, without affecting the value of
median. This property is sometimes described by saying that median is insensitive.

iii) It is not amenable to algebraic treatment

iv) As compared with mean, it is affected much by fluctuations of sampling.

Uses:

i) Median is the only average to be used while dealing with qualitative data which cannot be
measured quantitatively but still can be arranged in ascending or descending order of
magnitude, e.g. to find the average intelligence or average honesty among a group of people.

ii) It is to be used for determining the typical value in problems concerning wages, distribution
of wealth, etc.
Chapter 4: Mode
6.4.1 Mode
Mode is the value which occurs most frequently in a set of observations and around which the
other items of the set cluster densely. In other words, mode is the value of the variable which is
predominant in the series. Thus in the case of discrete frequency distribution, mode is the value
of x corresponding to maximum frequency.
In case of continuous frequency distribution, mode is given by the formula:
Mode=

l+ =I+

79
where l is the lower limit, h the magnitude and f1 the frequency of the modal class, fo and f2 are
the frequencies of the classes preceding and succeeding the modal class respectively.
6.4.2 Merits and Demerits of Mode

Merits

i) Mode is readily comprehensible and easy to calculate. Like median, mode can be found in
some cases merely by inspection.

ii) Mode is not at all affected by extreme values

iii) Mode can be conveniently located even if the frequency distribution has classes of unequal
magnitude provided the modal class and the classes preceding and succeeding it are of the
same magnitude. Open-end classes also do not pose any problem in the determination of
mode.

Demerits

i) Mode is ill-defined. It is not always possible to find a clearly defined mode. In some cases, we
may come across distributions with two modes. Such distributions are called bi-modal. If a
distribution has more than two modes, it is said to be multimodal.

ii) It is not based upon all the observations

iii) It is not capable of further mathematical treatment

iv) As compared with mean, mode is affected to a greater extent by fluctuations of sampling

Uses

Mode is the average to be used to find the ideal size, e.g. in business forecasting, in the
manufacture of ready-made garments, shoes, etc.
Chapter 5: Geometric and Harmonic mean
6.5.1 Geometric Mean
Geometric mean of a set of n observations is the nth root of their product. Thus the geometric
mean, G, of n observations xi, i=1,2 …. n is
G = (x1 x2 …… xn) 1/n
The computation is facilitated by the use of logarithms. Taking logarithm on both sides, we get

log G= =

G=antilog
In case of frequency distribution xi/fi, (I = 1, 2 ……n) geometric mean, G is given by

80
G=
Taking logarithm of both sides, we get
l

log G =

=
Thus we see that logarithm of G is the arithmetic mean of the logarithms of the given values.
We get

G=antilog
In the case of grouped or continuous frequency distribution, x is taken to be the value
corresponding to the mid-points of the classes.
6.5.2 Merits and Demerits of Geometric Mean
Merits:
i) It is rigidly defined
ii) It is based upon all the observations
iii) It is suitable for further mathematical treatment
iv) It is not affected much by fluctuations of sampling
v) It gives comparatively more weight to small items
Demerits:
i) Because of its abstract mathematical character, geometric mean is not easy to understand
and to calculate for a non-mathematics student.
ii) If any one of the observations is zero, geometric mean becomes zero and if any one of
the observations is negative, geometric mean becomes imaginary regardless of the
magnitude of the other items.
Uses
Geometric mean is used
i) To find the rate of population growth and the rate of interest
ii) In the construction of index numbers
6.5.3 Harmonic Mean
Harmonic mean of a number of observations is the reciprocal of the arithmetic mean of the
reciprocals of the given values. Thus, harmonic mean, H of n observation x, i=1, 2, ……n is

H=
In case of frequency distribution xi/fi (I = 1, 2 …..n)

H= , N=
6.5.4 Merits and Demerits of Harmonic Mean

81
Merits

Harmonic mean is rigidly defined, based upon all the observations and is suitable for further
mathematical treatment. Like geometric mean it is not affected much by fluctuations of
sampling. It gives greater importance to small items and is useful only when small items have to
be given a very high weightage.

Demerits: Harmonic mean is not easily understood and is difficult to compute.


Unit 7: Relative measures of dispersion
Learning objectives
The importance of the measures of dispersion and its uses are dealt.
Chapter 1: Relative measures of dispersion
7.1.1 DISPERSION

Averages or the measures of central tendency give us an idea of the concentration of the
observations about the central part of the distribution. If we know the average alone we cannot
form a complete idea about the distribution as well be clear from the following example:

Consider the series (i) 7, 8, 9, 10, 11 (ii) 3, 6, 9, 12, 15 (iii) 1, 5, 9, 13, 17. In all these cases we
see that n, the number of observations is 5 and the mean is 9. If we are given that the mean
of 5 observations is 9, we cannot form an idea as to whether it is the average of first series or
second series or third series or of any other series of 5 observations sum is 45. Thus we see
that the measures of central tendency are inadequate to give us a complete idea of the
distribution. They must be supported and supplemented by some other measures. One such
measure is Dispersion.

Literal meaning of dispersion is ‘scatteredness’. We study dispersion to have an idea about the
homogeneity or heterogeneity of the distribution. In the above case we say that series (i) is
more homogeneous (less dispersed) than the series (ii) or (iii) or we say that series (iii) is more
heterogeneous (more scattered) than the series (i) or (ii).

i) It should be rigidly defined

ii) It should be easy to calculate and easy to understand

iii) It should be based on all the observations

iv) It should be amenable to further mathematical treatment

v) It should be affected as little as possible by fluctuations of sampling


7.1.2 Measures of Dispersion

The following are the measures of dispersion.

82
i) Range

ii) Quartile deviation or semi-interquartile range

iii) Mean deviation and

iv) variance and Standard deviation


7.1.3 Range

The range is the difference between two extreme observations of the data. If A and B are the
greatest and the smallest observations respectively in a data, then its range is A-B.

Range is the simplest but a crude measure of dispersion. Since it is based on two extreme
observations which themselves are subject to chance fluctuations, it is not at all a reliable
measure of dispersion.
7.1.4 Quartile Deviation

Quartile deviation or semi-interquartile range Q is given by

Q = ½ (Q3-Q1)

where Q1 and Q3 are the first and third quartiles of distribution respectively.

Quartile deviation is definitely a better measure than the range as it makes use of 50% of the
data. But since it ignores the other 50% of the data, it cannot be regarded as a reliable
measure.
7.1.5 Mean Deviation

If xi/fi, I = 1,2, …..n is the frequency distribution, then mean deviation from the average A
(usually mean, median or mode) is given by

Mean deviation = =N

Where represents the modulus or the absolute value of the deviation (xi – A), when
the –ve sign is ignored.

Since mean deviation is based on all the observations, it is a better measure of dispersion than
range or quartile deviation. But the step of ignoring the signs of the deviation (x i-A) creates
artificiality and renders it useless for further mathematical treatment.

It may be pointed out here that mean deviation is least when taken from median.
7.1.6 Standard Deviation and Root Mean Square Deviation

83
Standard deviation usually denoted by the Greek letter small sigma ( ) is the positive square
root of the arithmetic mean of the squares of the deviations of the given values from their
arithmetic mean. For the frequency distribution xi/fi – 1,2, …..n,

Where is the arithmetic mean of the distribution and =N.

The step of squaring the deviations (xi – ) overcomes the drawback of ignoring the signs in
mean deviation. Standard deviation is also suitable for further mathematical treatment.
Moreover of all the measures, standard deviation is affected least by fluctuations of sampling.
Unit 8 : Measures of skewness and kurtosis

Learning objectives
The importance of the measures of dispersion and into uses are dealt.
The symmetry and peakedness of the available data are explained.
Chapter 1: Skewness
8.1.1 SKEWNESS

Literally, skewness means lack of symmetry. We study skewness to have an idea about the
shape of the curve which we can draw with the help of the given data. A distribution is said to
be skewed if

i) Mean, median and mode fall at different points

i.e. Mean Median Mode

ii) Quartiles are not equidistant from median, and

iii) The curve drawn with the help of the given data is not symmetrical but stretched more to
one side than to the other.
8.1.2 Measures of skewness

Various measures of skewness are

1) Sk = M – Md 2) Sk = M – Mo

Where M is the mean, Md, the median and Mo, the mode of the distribution.

3) Sk = (Q3 – Md) – (Md – Q1)

84
These are the absolute measures of skewness. As in dispersion, for comparing two series we do
not calculate these absolute measures but we calculate the relative measures called the co-
efficients of skewness which are pure numbers independent of units of measurement. The
following are the co-efficients of skewness:

1) Prof. Karl Pearson’s Co-efficient of Skewness

= ................................(1)

where s is the standard deviation of the distribution

It has been shown that for any distribution, (M-Md) / s lies between +1. Hence the limits for the
co-efficient of skewness are +3. In practice, these limits are rarely attained.

Skewness is positive if M> Mo or M>Md and negative if M<Mo or M<Md.

2. Prof. Bowley’s Co-efficient of Skewness

Based on quartiles,

= = .............................(2)

Thus Sk = + 1 if Md = Q1 and Sk = -1 if Q3 = Md. Bowley’s co-efficient of skewness lies between


+1

3. Based upon moments, co-efficient of skewness is

= ......................................(3)

Where symbols have their usual meanings. Thus Sk = 0 if either b 1 = 0 or b 2 = -3. But since b 2 =
m 4 / m 2 2 , cannot be negative, Sk = 0 if and only if b 1 = 0. In this respect b 1 is taken to be a
measure of skewness. The co-efficient in (3) is to be regarded as without sign.

We observe in (1) and (2) that skewness can be positive as well as negative. The skewness is
positive if the larger tail of the distribution lies towards the higher values of the variate (the
right), i.e. if the curve drawn with the help of the given data is stretched more to the right than
to the left and is negative in the contrary case.
Chapter 2: Kurtosis
8.2.1 Kurtosis

85
If we know the measures of central tendency, dispersion and skewness, we still cannot form a
complete idea about the distribution as will be clear from the above figure in which all the
three curves A, B and C are symmetrical about the mean ‘m’ and have the same range.

In addition to these measures we should know one more measure which Prof. Karl Pearson calls
as the Convexity of a curve or Kurtosis. Kurtosis enables us to have an idea about the flatness or
peakedness of the curve. It is measured by the co-efficient b 2 or its derivation g 2 given by

b 2= m 4/ m 22g 2=b 2– 3

Curve of the type ‘A’ which is neither flat nor peaked is called the normal curve or mesokurtic
curve and for such a curve b 2 = 3 i.e. g 2 =0. Curve of the type ‘B’ which is flatter than the normal
curve is known as platykurtic and for such a curve b 2 < 3, i.e. g 2 <0. Curve of the type ‘C’ which
is more peaked than the normal curve is called leptokurtic and for such a curve b 2 > 3 i.e. g 2 >0.

THE END

86

You might also like