0% found this document useful (0 votes)
18 views122 pages

Overview of Statistics and Sampling

This document provides an overview of statistics as a subject area. It begins with definitions and a brief history of statistics, noting its origins in state administration and early use in India and other countries for collecting vital records and conducting surveys. The document then discusses the scope of statistics in economics, management, and industry. It defines key terms like population and sample, explaining that a population is the whole group under study while a sample is a subset of the population. Different sampling methods are mentioned like probability sampling, judgment sampling, and mixed sampling. The document concludes by discussing data condensation and graphical methods.

Uploaded by

prajktabhalerao
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
18 views122 pages

Overview of Statistics and Sampling

This document provides an overview of statistics as a subject area. It begins with definitions and a brief history of statistics, noting its origins in state administration and early use in India and other countries for collecting vital records and conducting surveys. The document then discusses the scope of statistics in economics, management, and industry. It defines key terms like population and sample, explaining that a population is the whole group under study while a sample is a subset of the population. Different sampling methods are mentioned like probability sampling, judgment sampling, and mixed sampling. The document concludes by discussing data condensation and graphical methods.

Uploaded by

prajktabhalerao
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd

Unit 01.

Population and Sample

Contents:
Definition & History of Statistics
Scope in different areas
Population & Sample
Methods of Sampling and
Data Condensation & Graphical Methods

Definition & History of Statistics


The subject of Statistics, as it seems, is
not a new discipline but it is as old as
the human society itself.
Its origin can be traced to the old days
when it was regarded as the science of
State-craft and was the by-product of
administrative activity of the Sate.
The word Statistics seems to have
been derived from the Latin word
status or the Italian word statista
or the German word statistik each of

In India, an efficient system of collecting


official and administrative statistics
existed even more than 2,000 years ago,
in particular, during the reign of Chandra
Gupta Maurya (324-300 B.C.).
From Kautilyas Arthshastra it is known
that even before 300 B.C. a very good
system of collecting Vital Statistics and
registration of births and deaths was in
vogue.

During Akbars reign (1556-1605


A.D.), Raja Todarmal, the then land and
revenue minister, maintained good
records of land and agricultural statistics.

In Aina-e-Akbari written by Abul


Fazl
(in 1596-97), one of the nine gems of
Akbar, we find the detailed accounts of
the administrative & statistical surveys
conducted during Akbars reign.

In Germany, the systematic


collection of official statistics originated
towards the end of 18th century when, in
order to have an idea of the relative
strength of different German States,
information regarding population and
output industrial & agricultural was
collected.

In England, statistics were the


outcome of Napoleonic wars. The wars
necessitated the systematic collection
of numerical data to enable the
government to assess the revenues and
expenditures with greater precision and
then to levy new taxes in order to meet
the cost of war.

Seventeenth century saw the origin of


the Vital Statistics. Captain John Grant of
London (1620-1674), known as the
father of Vital Statistics, was the first
man to study the statistics of births and
deaths.
To name the few the following are the
giants who contributed towards modern
statistics (what we have today) which is
based on probability concept.

Casper Newman, Sir William Petty


(1623-1687), James Dodson, Dr. Price
contributed towards concept of
Insurance.
Pascal(1623-1662), P. Fermat(16011665), James Bernoulli (1654-1705), DeMoivre(1667-1754), Laplace(17491827), Gauss(1777-1855), Theory of
Probability, Principle of Least squares &
Normal Law of Errors.

Mathematicians & Statisticians from 18th,


19th & 20th centuries contributed towards
Modern theory of Probability, Regression
Analysis, Correlation Analysis, Probability
& exact sampling distributions, theory of
estimation, testing of hypothesis etc. To
name the few giants:
Sir R. A. Fisher, Francis Galton, Karl
Pearson, W. S. Gosset, Pascal, James
Bernoulli, P.C. Mahalanobis, P. V.
Sukhatme, R. C. Bose, Panse, J. N.
Shrivastava, S. N Roy, C.R. Rao,

Definition of
Statistics
By some giants
A) Statistics as numerical data:

Statistics are the classified facts


representing the conditions of the people
in a state specially those facts which
can be stated in number or in tables of
numbers or in any tabular or classified
arrangement. -- Webster
Statistics are numerical statement of
facts in any department of enquiry placed
in relation to each other. -- Bowley
By statistics we mean quantitative data
affected to a marked extent by
multiplicity of causes.
-- Yule

Statistics are measurements,


enumerations or estimates of natural
phenomenon, usually systematically
arranged, analyzed and presented as to
exhibit important inter-relationships
among them. -- A. M. Tuttle
Statistics may be defined as the
aggregate of facts to a marked extent by
multiplicity of causes, numerically
expressed, enumerated or estimated
according to a reasonable standard of
accuracy, collected in a systematic
manner, for a predetermined purpose and

B) Statistics as Statistical Methods


Statistics may be called as science of
counting.
Statistics may be rightly called the
science of averages. -- Bowley A. L.
Statistics is the science of estimates and
probabilities. -- Boddington.
Statistics is the science and art of
handling aggregate of facts observing,
enumeration, recording classifying and
otherwise systematically treating them.
-- Harlow.

Scope of Statistics
in
Economics
Management Sciences

and
Industry

Scope of Statistics in Economics:


Statistical data and technique of
statistical analysis have proved
immensely useful in solving a variety of
economic problems, such as wages,
prices, consumption, production,
distribution of income and wealth etc.
Statistical tool like Index numbers, Time
series Analysis, Demand Analysis and
Forecasting Techniques are extensively
used for efficient planning and economic

Scope of Statistics in Economics:


Empirical studies based on sound statistical
analysis have led to the formulation of many
economic lows. For example:
i. Engels Law of Consumption, (1895) was
based on detailed and systematic studies of
family budgets of a number of families.
ii. Peretos Law of Income Distribution is based
on the empirical study of the income data of
different countries of the world at different
times.
iii. Empirical studies based on the observation of
the actual behavior of the buyers in the market
led Revealed Preference Analysis of Prof.
Samuelson.

Scope of Statistics in Economics:


The extensive use of Mathematics &
Statistics in the study of economics have
led to the development of new disciplines
called Economic Statistics and
Econometrics.
These days, advance statistical
techniques are used to fit the economic
models for obtaining optimum results
subject to a number of constraints on the
resources like capital, labor, production
capacity etc.

Scope of Statistics in Management Sciences

Statistical tools & techniques are


widely used in decision making. For
efficient working of different work areas
viz. marketing, sales, production, logistics,
inventory, etc.

Index numbers, Time series Analysis,


Forecasting, SQC, etc statistical tools are
important regarding decision making.

Correlation and Regression Analysis


are such techniques which are vital
regarding decision making.

Scope of Statistics in Management Sciences

Along with these Linear


Programming, Transportation Problems,
Sequencing, PERT & CPM, Assignment
Problems, Inventory control are few
optimization techniques to find the
optimum solution.

Scope of Statistics in Industry:


In Industry, Statistics is extensively used
in Quality Control. The main objective in
any production process is to control the
quality of the manufactured product so
that it conforms to specifications. This is
called process control and is achieved
through the powerful technique of control
charts and inspection plans.

Scope of Statistics in Industry:


The discovery of the control charts was made
by a young physicist Dr. W. A. Shewhart of
the Bell Telephone Laboratories (U.S.A.) in
1924 and is based on setting 3 (3-sigma)
control limits which has its basis on the
theory of probability & normal distribution.
Now a days 6 control limits are widely used
where chance of error is almost negligible.
Inspection plans are based on special kind of
sampling techniques which are very
important aspect of statistical theory.

Population & Sample

Population
Population in general means number of
living persons in a particular geographical
area on a particular time. It is the usual
meaning and is used as population of a
country.
With reference to statistics, meaning of
population is broader sense. Here it
means Each and every, or all. The
meaning is each and every unit which
covers under a given problem is called
statistical population.

Definition:
The group of individuals under study is
called population or universe.
An aggregate of objects or individuals
under study is called Population or
Universe.
Population may contain finite or infinite
elements. Accordingly, it is called as
finite or infinite population.
e.g. Total number of people living in a
country, Total number of students in a
college, Total number of buses with PMT,
etc.

Sample

Any part of population or fraction


of population under study is known
as sample.
A finite subset of statistical
individuals in a population is called
sample and the number of
individuals in a sample is called the
sample size.

In a production process say out of


100 items manufactured & 10 are
chosen at random for testing of
quality. Then it is known as sample.

While purchasing food grains,


we inspect only a handful of
grains and draw conclusion
about the quality of the whole
lot. In this case, handful of
grains is a sample and the
whole lot is a population.

When data is collected from each


and every unit of population, it is called
census enumeration or census method.
In census, the results are more accurate
and reliable.
It requires more manpower.
It incurs huge cost and is time consuming
too.
To avoid this different sampling methods
are used.

Sampling Methods

The method by which sample is chosen


out of population is called sampling
method.
There are many sampling methods
depending on types of population,
purpose of sampling etc.
Following are types of sampling methods:

Types of sampling methods


The techniques or methods of selecting a
sample is of fundamental importance in
the theory of sampling and usually
depends upon the nature of data and
type of enquiry.
Sampling Methods may be broadly
classified under the following heads:

Subjective or judgment sampling


Probability sampling
And
Mixed sampling

Mixed sampling

If the samples are selected partly


according to some laws of chance and
partly according to a fixed sampling rule,
they are termed as mixed samples
and the technique of selecting such
samples is known as mixed sampling.

Types of mixed sampling techniques


Simple Random Sampling (SRS)
Stratified Random Sampling
Systematic Sampling
Multistage Sampling
Area Sampling
Simple Cluster Sampling
Multistage Cluster Sampling
Quota Sampling, etc.
Quasi Random Sampling

Simple Random Sampling (SRS)


It is the technique of drawing a sample in
such a way that the population has an
equal and independent chance of being
included in the sample.
In this method, an equal probability of
selection is assigned to each unit of the
population at the first draw.
It also implies an equal probability of
selecting any unit from the available units
at subsequent draws.

Simple random sampling can be


subdivided into two techniques, namely
a. Simple Random Sampling Without
Replacement (SRSWOR) and
b. Simple Random Sampling With
Replacement (SRSWR)

Simple Random Sampling With


Replacement (SRSWR)
In SRSWR, first sample is selected at random
from the universe, recorded, studied and then
replaced back in the population.
Then, similarly, second element is selected at
random. This process is continued till a sample
of required size is selected.
In this sampling technique population size
remains the same in each draw.
The main drawback here is that, the same
element may get selected more than once in
the sample.

Simple Random Sampling Without


Replacement (SRSWOR)
Here in SRSWOR, first elements is
selected at random but not replaced back
in the population. This method of
selecting sample is called as simple
random sampling without
replacement.
Here population size decreases at each
draw.
The problem of getting the same sample
more than once is solved in SRSWOR.

Selection of a Simple Random Sample


Random sample refers to that method
of sample selection in which every item
has an equal chance of being selected.
But random sample does not depend
upon the method of selection only, but
also on the size and nature of the
population.
Some procedure which is simple and
good for small population is not so for
the large population.

Generally, the method of selection


should be independent of the properties
of sampled population.
Proper care has to be taken to ensure
that selected sample is random.

Random sample can be obtained by


any of the following methods.
a. By Lottery system
b. Mechanical Randomization or
Random Numbers method.

a) Lottery System
The simplest method of selecting a random
sample is the lottery system.
Let us assume that we need to select r
candidates out of n. This consists in identifying
each and every member or unit of the population
with a distinct number, recorded on a slip or a
card say, 1 to n.
These slips should be as homogeneous as possible
in shape, size, colour, etc., to avoid the human
bias.
These slips are then put in a bag and thoroughly
shuffled and then r slips are drawn one by one.
The r candidates corresponding to numbers on
the slips drawn, will constitute a random sample.

Mechanical Randomization or
Random Numbers Method
The lottery method described above is
quite time consuming and cumbersome
to use if the population is sufficiently
large.
The most practical and inexpensive
method of selecting a random sample
consists in the use of Random Number
Tables, which have been so constructed
that each of the digits 0, 1, 2, ..., 9
appear with approximately the same
frequency and independently of each

Method of drawing random sample:


1. Identify the N units in the population
with the numbers from 1 to N.
2. Select at random, any page of the
random number tables and pick up the
numbers in any row or column or
diagonal at random.
3. The population units corresponding to
the numbers selected in step 2
constitute the random sample.

Merits & Limitations of SRS

Merits
1. Since the sample units are selected at
random giving each unit an equal chance of
being selected, the element of subjectivity or
personal bias is completely eliminated.
2. As such a simple random sample is more
representative of the population as compared
to the judgment or purposive sampling.
3. Theory of random sampling is highly
developed so that it enables us to obtain the
most reliable and maximum information at
the least cost, and results in saving time,
money and labor.

Limitations
1. Selection of a simple random sample
requires an up-to-date frame, i.e. a
completely catalogued population from
which samples are to be drawn.
Frequently, it is virtually impossible to
identify the units in the population
before the sample is drawn and this
restricts the use of SRS technique.
2. Administrative Inconvenience. A
simple random sample may result in the
selection of the sampling units which are
widely spread geographically and in such

3. At times a simple random sample might give


most non-random looking results. For
example, if we draw a random sample of size
13 from a pack of cards, we may get all the
cards of the same suit. However, the
probability of such an outcome is extremely
small.
4. For a given precision, SRS usually requires
larger sample size as compared to Stratified
random sampling.
5. If the sample is not sufficiently large, then it
may not be representative of the population
and thus may not reflect the true
characteristics of the population.

Stratified Random Sampling

Stratification means division into layers.


Auxiliary information (Past data or some
other information) related to the
character under study may be used to
divide the population into various groups
such that,
i. Units within each group are as
homogenous as possible and
ii. The group means are as widely
different as possible.

Thus, a population consisting of N


sampling units is divided into k relatively
homogenous mutually disjoint (nonoverlapping) subgroups, termed as
strata, of sizes N1, N2, . . . Nk, such that
N = N i.
If a simple random sample is of size ni,
(I = 1, 2, . . . , k) is drawn from each of
the stratum respectively such that n =
ni, the sample is termed as Stratified
Random Sample of size n and the
technique of drawing such a sample is

In stratified random sampling the two points,


viz.,
1. proper classification of the population into
various strata, and
2. a suitable sample size from each stratum,
are equally important. If the stratification is
faulty, it cannot be compensated by taking large
sample.
The criterion which enables us to classify various
sampling units into different strata is termed as
stratifying factor (s.f.).
Some of the commonly used stratifying factors
are, age, sex, educational or income level,
geographical area, economic status and so on.

A s.f. is called effective if it divides the


given population into different strata
which are homogenous (or nearly so)
within themselves and the units in
different strata are as unlike as possible.
Such an organization gives estimates
with greater precision.
In many fields of highly skewed
distributions, stratification is an
exceedingly valuable tool.

Advantages of Stratified Random Sampling

More Representative. Stratified


sampling ensures any desired
representation in the sample of the
various strata in the population.
It over-rules the possibility of any
essential group of population being
completely excluded in the sample.
Stratified sampling thus provides a more
representative cross section of the
population and is frequently regarded as
the most efficient system of sampling.

Greater Accuracy. Stratified sampling


provides estimates with increased
precision. Moreover, stratified sampling
enables us to obtain the result of known
precision for each of the stratum.
Administrative Convenience. As
compared with SRS, the stratified
samples would be more concentrated
geographically. Accordingly, the time and
money involved in collecting the data and
interviewing the individuals may be
considerably reduced and the supervision

Sometimes the sampling problems may


differ markedly in different parts of the
population, e.g. a population under study
consisting of
i) literates and illiterates or ii) people
living in institutes, hostels, hospitals, etc.,
and those living in ordinary homes.
In such cases, we can deal with the
problem through stratified sampling by
regarding the different parts of the
population as stratum and tackling the
problems of the survey within each

Systematic Sampling
Systematic sampling is a commonly
employed technique if the complete and
up-to-date list of sampling units is
available.
This consists in selecting only the first
unit at random, the rest being
automatically selected according to some
predetermined pattern involving regular
spacing of units.

Let us suppose that N sampling units


are serially numbered from 1 to N in
some order and a sample size of n is to
be drawn such that
N = n*k
k = N/n
where, k usually called the sampling
interval, is an integer.

Systematic sampling consists in drawing


a random number, say, i k and
selecting the unit corresponding to this
number and every kth unit subsequently.
Thus the systematic sample of size n will
consists of units
i, i+k, i+2k, . . . , i+(n-1)k
The random number i is called the
random start and its value determines,
as a matter of fact, the whole sample.

Merits and Demerits

Merits
Systematic sampling is operationally more
convenient than SRS or stratified random
sampling.
Time and work involved is also relatively much
less.
Systematic sampling may be more efficient than
SRS provided the frame (the list from which
sample units are drawn) is arranged wholly at
random. The most common approach to
randomness is provided by alphabetical lists
such as names in telephone directory, although
even these may have certain non-random
characteristics.

Demerits
The main disadvantage of systematic
sampling is that systematic samples are
not in general random samples, since the
requirement in merit three is rarely
fulfilled.
If N is not a multiple of n, then
i) the actual sample size is different
from that
required, and
ii) sample mean is not an unbiased
estimate of

Data Condensation Methods

Important Terms
Raw data,
Attributes,
Variables,
Classification,
Frequency distribution,
Cumulative frequency distribution.

Raw data: The data collected in any


statistical
investigation is known as raw data.
Attributes: A qualitative characteristic
like religion,
sex, blood group, nationality,
defectiveness of an item produced,
beauty, etc. are termed as
attributes.
Constant: The characteristics which
does not

Variable: A quantitative characteristic


(which
changes its value & can be
measured) like
profit, population of a country,
weight of a
person, etc, is known as variable.
A quantitative variable ca be divided into
two types, namely i) discrete variable &
ii) continuous variable.

Discrete variable: The variable which


can take only particular values is called
as discrete variable.
e.g. Number of defectives in a lot, size of
readymade garments, number of
members in a family, etc. which take
integer values.
Continuous variable: The variable
which can take all possible values in a
given specified range is called as
continuous variable.
e.g. Age, income, weight of a person,

Classification

The data collected from various sources is


not arranged systematically and it is
unprocessed data. We can not draw any
conclusions and can not interpret the
data. Classification of data is required for
drawing conclusions.

Classification is arrangement of data


in groups according to similarities or
common characteristics.
Classification is the process of
arranging data into sequences and
groups according to their common
characteristics or separating them into
different but related parts.
The entire process of making
homogenous and non-overlapping groups
of observations according to similarities is
called as classification.

Objectives
1) It condenses the data.
2) It omits unnecessary details.
3) It eases the process of data tabulation.
4) It facilitates the comparison with other
data.

Basis of Classification
Basis generally depend on the nature and
purpose of the data. To name the few:
Geographical classification
Chronological classification
Qualitative classification
Quantitative classification

Geographical classification:
This type depends upon geographical
regions. In such cases, classification may
be done by countries, states, districts,
Talukas, rural-urban, etc.
Chronological classification:
When statistical data is classified
according to the time of its occurrence it
is known as chronological classification.
For example: data regarding monthly
sales, daily rainfall, yearly production,

Qualitative classification:
When the data is classified according to some
qualitative phenomenon like beauty, honesty,
sex, grades in exam, etc. the classification is
qualitative classification. In this type the data
is classified according to the presence or
absence of the attributes in the given units.
Quantitative classification:
If the data is classified on the basis of
phenomenon which is capable of quantitative
measurement like age, height, weight,
production, income, prices, etc., it is termed as
quantitative classification. This classification is
also called as classification by variables.

Frequency distribution
A frequency distribution means the data
classified on the basis of quantitative
variable. Frequency distribution can be
classified in two parts as individual series
and frequency series.
Frequency distribution can be classified
as discrete frequency distribution and
continuous frequency distribution.
Individual series is the series in which
items are listed singly. This series may be
unorganized or organized.

When observations, discrete or


continuous, are available on a single
characteristic of a large number of
individuals, often it becomes necessary to
condense the data as far as possible
without loosing any information of
interest.

Let us consider the marks in


Mathematics obtained by 250 students of
MITSOM College selected at random from
among those appearing in an
examination.

32
54
38
44
68
41
30
43
46
41
40
31
40
40
36
46
48
32
40
17
48
47
37
52
48

47
32
26
21
41
53
33
32
50
38
33
51
43
45
32
40
50
31
50
42
50
55
52
45
44

41
31
50
45
30
48
37
24
26
40
42
45
40
19
61
32
43
42
27
57
31
57
47
23
60

51
46
40
31
52
21
35
38
15
37
36
41
34
24
30
34
55
34
47
35
58
37
46
41
38

41
15
38
37
52
28
29
38
23
40
51
50
34
34
44
44
43
34
34
38
33
41
44
47
38

30
37
42
41
60
49
37
22
42
48
42
53
44
47
43
54
39
32
44
17
44
54
50
33
44

39
32
35
44
42
42
38
41
25
45
56
50
38
37
50
35
41
33
34
33
26
42
44
42
38

18
56
22
18
38
36
40
50
52
30
44
32
58
33
31
39
48
24
33
46
29
45
38
24
43

48
42
62
37
38
41
32
17
38
28
35
45
49
37
38
31
53
43
47
36
31
47
42
48
40

53
48
51
47
34
29
49
46
46
31
38
48
28
36
45
48
34
39
42
23
37
43
19
39
48

This representation of data does not


furnish any useful information and is
rather confusing to mind. A better way
may be to express the figures in an
ascending or descending order of
magnitude, commonly termed as array.
But this does not reduce the bulk of the
data.
A much better representation is use of
tally mark.

Marks
15
17
18
19
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39

No. of Students Total


Tally Marks
frequency
||
2
|||
3
||
2
||
2
||
2
||
2
|||
3
||||
4
|
1
|||
3
|
1
|||
3
||
2
||||
5
|||| ||||
10
|||| ||||
10
|||| |||
8
|||| |||| |
11
||||
5
||||
5
|||| |||| ||
12
|||| |||| |||| ||
17
|||| |
6

Marks
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
60
61
62
68

N. of Students
Total
- Tally Marks frequency
|||| |||| |
11
|||| ||||
10
|||| |||| |||
13
|||| |||
8
|||| |||| ||
12
|||| ||
7
|||| ||
7
|||| |||
8
|||| |||| ||
12
3
|||
|||| ||||
10
||||
4
||||
5
||||
4
|||
3
||
2
||
2
||
2
||
2
|||
3
|
1
|
1
|
1

A bar (|) called tally mark is put against the


number when it occurs. Having occurred four
times, the fifth occurrence is represented by
putting a cross tally (|) on the first four
tallies. This technique facilitates the counting
of the tally marks at the end.
The representation of the data as above is
known as frequency distribution. Marks are
called the variable (x) and the number of
students against the marks is known as the
frequency (f) of the variable.
The word frequency is derived from how
frequently a variable occurs.

This representation, though better than


an array, does not condense the data
much and it is quite cumbersome to go
through this huge mass of data.
Frequency distribution is a series where
we count how many times a particular
value or a particular group is repeated
called frequency.

If the identity of the individuals about


whom a particular information is taken is
not relevant, nor the order in which the
observations arise, then the first real step
of condensation is to divide the observed
range of variable into a suitable number
of class-intervals and to record the
number of observations in each class.
For example, in the above case, the data
may be expressed as:

Marks
(x)
15-19
20-24
25-29
30-34
35-39
40-44
45-49
50-54
55-59
60-64
65-69
Total

No. of students
(f)
9
11
10
44
45
54
37
26
8
5
1
250

Such a table showing the distribution of


the frequencies in the different classes is
called a frequency table and the manner
in which the class frequencies are
distributed over the class intervals is
called the grouped frequency distribution
of the variable.
Class: It is a group of numbers in which
items are placed.
Class limit: For each group or class we
consider two numbers. These two numbers
are called class limits. The lowest number
is the lower limit of the class and the
highest number is called the upper limit.

Class mark or Midvalue: It is the midpoint of the class interval.


= (Upper limit + Lower limit)/2
= (Upper boundary + Lower
boundary)/2
When classes are 100-200, 200-300, 300400,etc, we observe that 200 is upper
class limit for 100-200 class and lower
limit for 200-300 class. Such classes are
said to be continuous.

If class limits are as seen in the previous


table, viz. 15-19, 20-24, 25-29, ....etc, we
observe that 19 is upper class limit of 15-19
class and 20 is the lower class limit of next
class. Here, class limits are not continuous,
also called as inclusive classes. Here, the
lower and upper limit of the class interval is
included. If they are not continuous, then
we have to make them continuous.
In this example we make class limits
continuous by subtracting and adding 0.5
respectively to the lower and upper limit of
each class.

So, the resultant continuous classes are:


14.5-19.5, 19.5-24.5, 24.5-29.5, etc.
These are called as exclusive classes.
Here, the upper limit of the class interval
is excluded and included in the next class
interval.
Width or Magnitude of class interval:
When class limits are continuous, then
the difference between upper class limit
and lower class limit is called as width or
magnitude or span of the classes.
In the above example, 19.5-24.5, 24.529.5,etc, width is 5 as the difference

In spite of great importance of


classification in statistics, no hard and
fast rules can be laid down for it. The
following points may be kept in mind for
classification:
These classes should be clearly defined
and should not lead to any ambiguity.
These classes should be mutually
exclusive and non overlapping.
The classes should be of equal width.
Indeterminate classes, e.g., the open-end
classes like less than a or greater than
b should be avoided as far as possible

The number of classes should be neither


be too large nor too small. It should
preferably lie between 5 and 15.
However, the number of classes may be
more than 15 depending upon the total
frequency and the details required, But it
is desirable that it is not less than 5 since
in that case classification will not reveal
the essential characteristics of the
population.
The following formula due to Struges may
be used to determine an approximate
number k of classes.

Cumulative frequency:
These are cumulative totals of
frequencies. These are of two types.
1. When cumulative frequencies are
based on upper limits of classes, it is
called below or less than type cumulative
frequencies.
2. When cumulative frequencies are
based on lower limits of classes, it is
called above or more than type
cumulative frequencies.
For example:

Less than type


Marks Frequency
cumulative
frequency

More than type


cumulative
frequency

0-10

4+4+8+12+7+1
=36

10-20

1+7=8

4+4+8+12+7=35

20-30

12

1+7+12=20

4+4+8+12=28

30-40

1+7+12+8=28

4+4+8=16

1+7+12+8+4
=32

4+4=8

1+7+12+8+4
+4=36

40-50
50-60

Ex. 1 Daily earnings of 50 doctors in a city


are as follows. Classify the data taking
classes as 40-44, 45-49, 50-54, etc. and
obtain cumulative frequency column.
68, 60, 55, 50, 40, 44, 42, 50, 50, 55,
55, 60, 60, 70, 70, 56, 50, 44, 70, 63,
52, 56, 45, 64, 70, 72, 65, 58, 53, 45,
54, 45, 58, 65, 75, 75, 65, 59, 55, 46,
60, 55, 48, 65, 76, 48, 55, 66, 60, 80.

Daily
earning
40-44
45-49
50-54
55-59
60-64
65-69
70-74
75-79
80-84

Tally
marks
||||
|||| |
|||| |||
|||| ||||
|||| ||||
|||| |
||||
|
|
Total

No. of
Doctors
4
6
8
10
9
6
5
1
1
50

C.F.
4
10
18
28
37
43
48
49
50

Exercise

Ex.1. The data given below gives number of


portable torches sold by Vijay on 25 working
days. Prepare a frequency distribution of number
of torches sold.
1, 4, 1, 1, 2, 2, 1, 2, 0, 1, 1, 3, 0,
1, 5, 4, 1, 2, 3, 1, 1, 1, 4, 1, 2.
Ex.2. Among a group of students 10% scored marks
below 20, 20% scored marks between 20 and 40,
35% scored marks between 40 and 60, 20%
scored marks between 60 and 80 and remaining
30 students scored marks between 80 and 100.
Using this information prepare a frequency
distribution. Prepare less than type and more
than type cumulative frequencies.

Ex.3. From the following observations


prepare a frequency distribution table in
ascending order starting with 5-10(Using
Exclusive method). Prepare less than type
as well as more than type cumulative
frequencies.
12, 36, 40, 30, 28, 20, 19, 19, 27, 15,
26, 20, 19, 7, 26, 37, 5, 20, 11, 17,
37, 10, 10, 16, 45, 33, 21, 30, 20, 5

Ex.4. In a sample study about tea drinking


habits in two towns A and B the following data
was obtained.
Town A:
52% of the population were males,
65% of the people were tea drinkers,
40% of the population were male tea drinkers.
Town B:
50% of the people males,
75% of the people were tea drinkers,
42% of the people were male tea drinkers.
Tabulate the above information.

Ex.5. Following is the frequency distribution


of rainfall in Mumbai for 78 years.
Rainfall in
Frequency
inches
5-9
10
10-14
17
15-19
15
20-24
18
25-29
14
30-34
0
35-39
2
40-44
2

1.
2.
3.
4.

Obtain class boundaries of 3rd class


Find class mark of 1st class
Find class width of any class
Number of years having less than 25
inches rainfall
5. Number of years having more than 29
inches rainfall.

[Link] the following distribution of age


of Life Insurance Policy holders prepare a
frequency distribution and also
cumulative frequency distribution on
more than basis.
Age (Yrs.)
No. of Policy
Holders
Less than
9
15
Less than
25
25
Less than
63
35

Graphical Representation of
Data

Histogram
Frequency Polygon
Multiple Bar Diagram
Subdivided Bar Diagram

Histogram

Frequency Polygon

Multiple Bar Diagram

Sub-divided Bar
Diagram

Pie Chart

You might also like