0% found this document useful (0 votes)
31 views4 pages

Overview of Statistics and Data Collection

The document discusses the importance and scope of statistics, detailing its functions, methods of data collection, and classification of data. It highlights the roles of various statistical organizations in India, such as the Ministry of Statistics and Programme Implementation and the Indian Statistical Institute. Additionally, it covers concepts of central tendency, measures of dispersion, and various graphical representations of data.

Uploaded by

akshys1907
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
31 views4 pages

Overview of Statistics and Data Collection

The document discusses the importance and scope of statistics, detailing its functions, methods of data collection, and classification of data. It highlights the roles of various statistical organizations in India, such as the Ministry of Statistics and Programme Implementation and the Indian Statistical Institute. Additionally, it covers concepts of central tendency, measures of dispersion, and various graphical representations of data.

Uploaded by

akshys1907
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

2

CHAPTER 1
4. Enlarges human knowledge and experience.
Official Statistics

STATISTICS- SCOPES AND DEVELOPMENT 5. Helps in formulating policies, testing hypothesis and The Ministry of Statistics and Programme Implementation
forecasting future events. (MOSPI) has two wings, Statistics and Programme
Implementation.
The Statistics Wing called National Statistical Office (NSO)
Scope of Statistics consists of the Central Statistical Office (CSO), the Computer
Centre and the National Sample Survey Of- fice (NSSO).
Statistics have a crucial role in developing the world. CSO coordinates all statistical activities in the country.
1. Planning
The word Statistics have been derived from Latin word ”Status” or NSSO is the largest organisation in India, conducting regular
the Italian word ”Statista”. The meaning of these words is socio-economic surveys. It has four divisions- SDRD, FOD, DPD
2. Economics
”Political State” or a Government. and CPD.
Sir Ronald Aylmer Fisher is known as father of modern statistics. Indian Statistical Institute (ISI)
3. Industry
Defintion of Statistics by Croxton and Cowden- ”Statistics can ISI is a unique institution devoted to the research, teaching and
be defined as the collection, presentation and interpretation of application of Statis- tics. It is founded by Prof. Prasanta
4. Mathematics Chandra Mahalanobis in [Link] is known as father of indian
numerical data.”
Functions of Statistics statistics
29th june, birth day of P C Mahalanobis, is celebrated as
5. Psychology and Education
national statistics day. ISI publishes a journal ’Sankhya’.
1. Simplifies complexity. Economics and Statistics Department
6. Management Studies
The Directorate of Economics and Statistics is the nerve centre of
2. Presents facts in a definite and precise form. the Kerala State statistical system and it is the nodal agency
Some applied areas of Statistics are coordinating all statisical activities in the state.
3. Provides comparison.
• Actural Science

1
• Biostatistics

• Agricultural Statistics

Even if there is a wide application of Statistics in day to day life,


Statistics has also some limitations. The misuse of Statistics is
the main cause of discredit to this sci- ence.

The representative part of the population is known as sample. The 4. Telephone interview
method of col- lecting data from the sample is known as sampling
or sample survey. 5. Mailed questionnaires and schedules
The statistical survey may be either by census method or by
sampling method. The factors which can vary from one object to
another are called variables. 6. Focus group discussion
The variable which can not be numerically measured is called
qualitative [Link] variable which are numerically measured
is called quantitative variable. Questionnaires and schedules are series of questions arranged in
CHAPTER 2 If the variable takes specific values only, it is called discrete a logical order so as to collect information for a specified purpose.
variable. A continuous variable takes any value within the defined A questionnaire is usually mailed by post or email to selected
range of variables. [Link] approaches personally to the informant
A nominal scale of measurement is used to name categories and collects information from them.
such as gender,nation etc. In the ordinal scale of measurement, we Sources of secondary data
COLLECTION OF DATA
can put an order to the data according to the relation among the
values of variables. The data regarding a quantitative variable is a
• Government publications
cardinal data.
Based on the sources of collection, statistical data may be classified
as primary and secondary data. • Office records in panchayats, municipalities etc.
Primary data collected by the investigator for the first time for his/her
Data means any measurement, result, fact or observation which
own purpose. Data obtained from existing sources which may be • Survey reports of various research organizations
gives information. Data collection is the systematic gathering of data
published or unpublished are known as secondary data.
for a particular purpose from various sources.
Methods of primary data collection are • Survey reports in journals, newspapers and other publications
Statistical investigation includes collection, classification,
presentation, analysis and interpretation of data according to well
defined procedures. The person authorized to make investigation 1. Direct personal interview • Websites
is known as investigator.
The investigators depute some persons to collect the data from
the field. These per- sons are known as Enumerators. The 2. Indirect oral investigation
process is known as Enumeration.
A population consists of all elements, individuals, items or objects 3. Direct observation
whose charac- teristics are being studied. If data are collected
from each and every unit of the population, the investigation is
called census.

table used for one way classification is one way table. If we If two characteristics are measured simultaneously from each
consider two characteris- tics at a time for classification for data, it unit, it is known as bi variate data. The frequency distribution of a
is termed as two way classification of data and the table is two bi variate data is called bi variate frequency table
way table.
The number of repetitions of a particular observation in a series is
called frequency of the observation.
The series of observations in which items are listed individually is
called raw data or individual series. A discrete frequency table is
that series in which data are presented in a way that exact
CHAPTER 3 measurements of units are clearly shown.
Frequency tables with classes and corresponding frequencies are
known as continuous frequency table.
If the lower limit of the first class or upper limit of the last class
CLASSIFICATION AND TABULATION is not specified, then it is called open end class.
Frequency
Relative frequency = = f
Total frequency N

Frequency
Percentage frequency
Total =
frequency ×N 100= f × 100

Sum of the relative frequencies is 1 and sum of the percentage


Classification of data is the process of grouping the data frequencies is 100 The number of observations less than or equal to
according to some charac- teristics . a particular value is called less than cumulative frequency of that
Types of classification value.
The number of observations greater than or equal to a particular
• Qualitative classification value is called greater than cumulative frequency or more than
cumulative frequency of that value. If only one characteristic of the
sampling units is measured for the study, it is called uni variate
• Quantitative classification
data.

• Chronological classification

• Geographical classification

Tabulation of data is the method of representing data with the help


of a statistical table.
In one way classification only one characteristic is considered for
classification. The
SORRY FOR BAD LAYOUT :(
7
4. Percentage Bar Diagram

In constructing a pie diagram, the first step is to prepare the data so


that the various component values can be transposed into
corresponding degrees on the circle using the formulae
Item frequency
Angle Total × 360
frequency
=

CHAPTER 4 The most commonly used graphs for representing a frequency CHAPTER 5
distribution are

1. Histogram
DIAGRAMS AND GRAPHS CENTRAL TENDENCY
2. Frequency Polygon

3. Frequency Curve

Diagrams and Graphs are the methods for simplifying the The property of the observations in a data to cluster or
complexity of quanti- tative [Link] and Graphs are more 4. Ogives (cumulative frequency curves) concentrate around a value is known as Central Tendency.
attractive and impressive. Diagrams Commonly used diagrams Measures of central tendency (averages) are the values which gives
are an idea about the concentration of observations in the central part of
5. Scatter plot
the distribution.
1. Bar diagrams

There are two types of ogives. Less than ogive and


2. Pie diagram greater than ogive. Scatter plot is used to represent a bi Desirable properties of a good average
variate data.
Bar Diagram

There are four types of bar diagram. 1. Simple and rigid definition.

1. Simple Bar Diagram


2. Simple to understand and easy to calculate.

2. Multiple Bar Diagram


3. Based on all the observations.

3. Sub divided Bar Diagram(Component Bar Diagram)

Σf
x¯ = (5.2)
x
N

4. Least affected by extreme values.

5. Least affected by fluctuations of sampling.

6. Capable of further mathematical treatment.

JOIN OUR GROUP FOR MORE STUFF ❤


The various measures of central tendencies are

1. Arithmetic Mean(AM)

2. Median

3. Mode

4. Geometric Mean(GM)

5. Harmonic Mean(HM)

Arithmetic Mean Arithmetic mean which is also known as


mean is denoted by x¯
Sum of the
Mean = observations No. of
observations

(i) For a raw


Data Σ
(5.1
x
)
x¯ =
n

where n is the no. of observations.

(ii) For a a Discrete Frequency Distribution

where N = Σf is the total frequency. Median Modal class is the class having highest frequency. Mode = l +
(f1−f0)c

(ii)For a a Continuous Frequency Distribution 2f1−f0−f2


Median is the value of middle most observation in
the data. Median for raw data Where l is the lower limit of the modal class, f1 is the frequency of
When observation are arranged in ascending or descending order the modal class,f0 is the frequency of the preceding class to the
Σf of magnitude, Me- dian is the (n+1 )th item in the data where n is the modal class, f2 is the frequency of the succeeding class to the
x¯ = (5.3)
x no. of observations.
2
modal class and c is the class interval of the modal [Link] can
N Median for discrete frequency distribution also be locate using histogram.
N+1 Empirical relationship between mean, median
Median is the observation having cumulative 2 frequency , when and mode is: Mean-Mode= 3(Mean-Median)
where N = Σf is the total frequency and midpoint of the class is the observations are arranged in ascending order.
taken as the value of x. Or
Median for continuous frequency distribution
Mathematical properties of AM
Mode= 3Median- 2Mean
Median class is the class 2where N th observation lies.
Geometric Mean (GM)
1. Σ(x − x¯) = 0 M edian =
( N −m)c
2 l
f
GM is the nth root of the product of n observations in the data set.
2. Σ(x − a) is least when a = x¯
2
where l isthe lower limit of the median class, c is the class interval
of the median class, f is the median class and m is the cumulative
frequency of the class preceding the median class. √
3. If each observation is increased by ’a’, then the mean is also GM = n
x1 × × ... × = × 1
Median can be located graphically using ogives.
increased by ’a’ and each observation is decreased by ’a’, x2 xn (x1 x2 × ... × xn)n
then the mean is also decreased by ’a’. Mode
4. If each observations is multiplied by p, p /
= 0, then the
Mode of a data is defined as the value that is repeated most often in Harmonic Mean (HM)
mean of the new observations is px¯.
the data. Mode for a raw data
Weighted AM Mode of a raw data is the observation having the maximum HM is the reciprocal of the AM of the reciprocals of the given
frequency in the data. Mode for a discrete frequency distribution observations.
x¯w =ΣwΣwx Mode is the observation having the highest frequency. Mode for a
continuous fre- quency distribution
Combined AM n
HM = 1 Σ
If x¯1 and x¯2 are the means of two groups of n1 and n2 observations x

respectively, the mean of the combined group of n1 + n2


observations are given by x¯= n1 x̄ 1 +n2 x̄ 2 n +n
1 2
Relations among AM, GM and HM

1. For a set of positive values, AM ≥ GM ≥ HM.


17
3N
( −m3)c3
When all the observations are the same ,then AM = GM = Q3 = l+3 4 f3

HM
wherel1 and l3 are the lower limits of
quartile classes, f1 and f3 are the
2
2. (GM ) = AM × HM frequencies of the quartile classes, c1 and
√ c3 are the class intervals of the quartile
or GM = AM × HM classes
and m1 and m3 are the cumulative frequencies preceding the
quartile classes.
Quartiles, deciles and percentiles are the partition values which
divide data into several equal parts. CHAPTER 6
Deciles divide the distribution into ten equal parts and there are 9
Quartiles divide a data into four equal parts. There are three deciles. Median is the 5th decile
quartiles, denoted by Percentiles divide a distribution into hundred equal parts. There are
99 percentiles. Median is the 50th percentile.
Q1, Q2 and Q3. Q2 is the A box plot is a graph of a data set that consists of a line extending DISPERSION
median. Quartiles for raw from the minimum value to the maximum value and a box with
data lines drawn at the Q1, the median, and Q3.
Arrange the n observations in ascending order of magnitude.
n+1 th
Q1 = value of
4 ( ) item in the series

Q = value of 3(n+1)th
item in the series Dispersion is the degree of scatter or variation of the variable about
3 4 a central value. Measures of dispersion are

Quartiles for discrete frequency distribution


(N+1)
Q1 = observation having cumulative frequency
4 1. Range
3(N+1)
Q3 = observation having cumulative frequency
4
2. Quartile Deviation
where N is the total frequency.
3. Mean Deviation
Quartiles for continuous frequency distribution

Prepare cumulative frequency. N be the total frequency. Locate 4. Standard Deviation


the classes having
N 3N
cumulative frequencies
4 4and . These classes are called quartile Range
classes.
Range = Highest value - Lowest value

( N − m1)c1 = H- L
Q 1 = l1 + 4

f1 19

20
SD σ
CV =Mean × 100 x̄ = × 100
Quartile Deviation(QD)
Q3−Q1
Q3−Q1 Coefficient of QD
Q +Q=
QD = 2 3 1

Q1 and Q3 are explained in the previous chapter for raw data, Covariance is a measure of strength of linear relationship
discrete frequency distribution and continuous frequency between two variables Cov(x,y) indicates whether the variables
distribution. are positively related or negatively related in a bi variate
Mean Deviation (MD) distribution.
Cov(x,y) = Σ(x−x̄)(y−ȳ)
n n =
Σxy
− x¯ × y¯
MD for raw data CHAPTER 7
MD = Σ|X−A|
n , where A is any
average. MD for discrete
frequency distribution
MD = Σf|X−A| , where A is any average and N is the total
SKEWNESS AND KURTOSIS
N
frequency. MD for continuous frequency distribution
MD = Σf|X−A|
N , where A is any average , N is the total frequency
and x is the mid value of the class.
Standard Deviation (SD)

SD for raw Skewness means the absence of symmetry in a data set. For a
data SD σ symmetric distribution Mean= Median= Mode.
Σ(x−x̄)2 n = Σx
n
2
− x¯2 There are two types of skewness ‘
=
SD for discrete frequency distribution
1. Positive skeness
SD σ = N Σfx − x¯2 , where N is the total frequency.
2

SD for continuous frequency distribution 2. Negative skewness

SD σ = Σfx − x¯2 , where N is the total frequency and x is the


2

mid valueN of the For a positively skewed data Mode <


Median < Mean For a negatively skewed
class. data Mean < Median < Mode. Measures
of skewness
Coefficient of variation(CV) and coefficient of QD are relative
measures of dispersion. CV is used to compare the consistency or 1. Karl Pearsons Coefficient of Skewness
stability between two or more sets of data.
Sk = Mean−Mode
SD

22

2. Bowleys coefficient of Skewness Sample space: The set of all possible outcome of random
experiment is called the sample space. Sample space is usually
Q3+Q1−2Median listed in curly brackets{} and is denoted by S.
SB =
Q3−Q1 Sample point: Each element in the sample space is called a
sample point. Events: An event is a set of outcomes which
3. Coefficient of Skewness based on moments have some characteristics in com- mon.
Equally likely event: Two or more events which have an
β = µ32 equally likely chance or equal probability of occurrence are
1 µ23 called equally likely events.
√ Mutually exclusive (disjoint) events: Events are said to be
γ1 = β 1 CHAPTER 8
mutually exclusive if the happening of any one of them
µ3 determines nature of skewness. If µ3 > 0, the distribution is excludes the happening of all the others. Exhaustive events:
positively skewed. If µ3 < 0, the distribution is negatively A set of events is called exhaustive,if all the events together
skewed. If µ3 = 0, the distri- bution is symmetric. consume the entire sample space.
PROBABILITY Basic properties of probability
Kurtosis

Kurtosis is the measure of peakedness or flatness of the


frequency distribution. There are three types of Kurtosis.
• The probability is always between 0 and 1.
Probability is the way of measuring the chances of something
• The probability of occurrence of an impossible event is 0.
(a) Lepto kurtic to happen. According to Ya-Lin Chou: ”Probability is the
science of decision making with calculated risks in the face of
• the probability of something to occur is 1.
(b) Meso kurtic uncertainty.”
Random experiment
• Probability can not be negative.
(c) Platty kurtic
An experiment is called random experiment if it satisfies the
following condi- tions: Mathematical or classical definition of probability
Measure of Kurtosis
(a) It has more than one outcome.
If a random experiment results in ’n’ exhaustive , mutually
β2 = µ4
exclusive and equally likely outcomes out of which ’m’ are
µ22
(b) It is not possible to predict the outcome in advance. favourable to the occurrence of an
γ2 = β 2 − 3
(c) It can be repeated any number of times under identical
If β2 = 3 (γ2 = 0) the curve is conditions.
meso kurtic. If β2 > 3 (γ2 > 0) the
curve is lepto kurtic. If β2 < 3 (γ2 < Trial: A trial is an action which results in one of several
0) the curve is platty kurtic. possible outcomes or results.
24
26
If S is the sample
Algebra of events space P(S)=1 Axiom 3:
event ’A’. Then the probability of occurrence of A, usually Additivity
denoted by P(A) is given by For two mutually exclusive events
• Complement of A:A¯ or A or A → not A
c ′
A1 and A2, P (A1 A2) = P (A 1 )+ P
Number of favourable (A2)
P (A) = • A or B → At least one of the event A or B occurs.
cases
• A and B → Simultaneous occurrence of A and B
m
=
• A and B → neither A nor B
′ ′

number ofexhaustive
cases n
• A and B → only A

or • (A and B ) or (A and B) → exactly one among A and B


′ ′

Addition rule of probability


number of outcomes N (A)
P (A) =
in A N (S) Rule 1
=
number of outcomes When two events A and B are mutually exclusive,
in S

Rule 2

When two events A and B are not mutually exclusive,


P (A or B) = P (A)+ P (B) − P (A and B)
Frequency approach of Probability

If after n repetitions of an experiment, where n is very large,


an event A is observed to occur in m of these, then the P (A)
= limn→∞ m n

Given the frequency of the distribution, the probability is computed


as
frequency of the class
P (A) = total frequency in the frequency
distribution

Axioms on
Probability Axiom 1:
Non-negativity For
any event A, P (A) ≥
0 Axiom 2: Certainty

A1 and A2, then


P (A1)P (A/A1)
CHAPTER 9 P (A1/A) = P (A1)P (A/A1)+P (A2)P (A/A2)
CHAPTER 10
P (A2/A) =
P (A2)P (A/A2) SAMPLING TECHNIQUES
P (A1)P (A/A1)+P (A2)P (A/A2)

This theorem can be extended if we have three mutually


CONDITIONAL PROBABILITY exclusive and exhaus- tive events A1, A2 and A3, then In census, data is collected from each and every unit of the
P (A1/A) = P (A1)P (A/A1) population. The method of collecting data from the sample is
P (A1)P (A/A1)+P (A2)P (A/A2)+P (A3)P (A/A3)
known as sampling or sample survey.
P (A2/A) =
P (A2)P (A/A2) Sampling errors are seen in sample surveys due to the fact
P (A and B) P (A1)P (A/A1)+P (A2)P (A/A2)+P (A3)P (A/A3)
that only a part of the population is used for enquiry.
, provided P (A) =
/0
P (A) P (A3/A) =
P (A3)P (A/A3) Sampling errors decreases as sample size increases.
P (A1)P (A/A1)+P (A2)P (A/A2)+P (A3)P (A/A3)
Multiplication Theorem Errors other than sampling errors in a survey are called non
sampling errors.
P (A and B) = P (A).P (B/A) = P (B).P
(A/B) Methods of Sampling

Independent and dependent events

Two events are said to be independent if the occurrence of one


event do not affect the occurrence of the other. Otherwise the • Non Probability Sampling
events are said to be depen- dent.
If P (A/B) = P (A) and P (B/A) = P (B), then A and B are • Probability sampling
independent. Non probability sampling
Multiplication Theorem for Independent Events
• Convenience sampling
If A and B are independent, multiplication
theorem becomes P(A and B) = P(A). P(B) • Judgement sampling
Total Probability Theorem
• Quota sampling
If A1 and A2 are two mutually exclusive and exhaustive
events and A is any other event which can occur along with
A1 and A2, then total probability the- orem states that Probability sampling
P (A) = P (A1)P (A/A1)+ P (A2)P (A/A2)
• Simple Random Sampling
If A1, A2 and A3 are three mutually exclusive and exhaustive
events and A is any other event which can occur along with
• Systematic Sampling
A1, A2 and A3, then total proba- bility theorem states that
P (A) = P (A1)P (A/A1)+ P (A2)P (A/A2)+ P (A3)P (A/A3) • Stratified Sampling

Bayes’ Theorem
• Cluster Sampling
If A1 and A2 are two mutually exclusive and exhaustive
events and A is any other event which can occur along with • Multi Stage Sampling

Simple Random Sampling(SRS)

SRS is a probability sampling in which each unit in the


population has an equal chance of being included in the
sample.
• Simple
Random Sampling Without Replacement
(SRSWOR)

• Simple Random Sampling With Replacement


(SRSWR)

If a population consists of N units and a sample of n units


to be taken, the possible number of samples in SRSWOR is
NCn and in SRSWR is N n

You might also like