Overview of Statistics and Data Collection
Overview of Statistics and Data Collection
CHAPTER 1
4. Enlarges human knowledge and experience.
Official Statistics
STATISTICS- SCOPES AND DEVELOPMENT 5. Helps in formulating policies, testing hypothesis and The Ministry of Statistics and Programme Implementation
forecasting future events. (MOSPI) has two wings, Statistics and Programme
Implementation.
The Statistics Wing called National Statistical Office (NSO)
Scope of Statistics consists of the Central Statistical Office (CSO), the Computer
Centre and the National Sample Survey Of- fice (NSSO).
Statistics have a crucial role in developing the world. CSO coordinates all statistical activities in the country.
1. Planning
The word Statistics have been derived from Latin word ”Status” or NSSO is the largest organisation in India, conducting regular
the Italian word ”Statista”. The meaning of these words is socio-economic surveys. It has four divisions- SDRD, FOD, DPD
2. Economics
”Political State” or a Government. and CPD.
Sir Ronald Aylmer Fisher is known as father of modern statistics. Indian Statistical Institute (ISI)
3. Industry
Defintion of Statistics by Croxton and Cowden- ”Statistics can ISI is a unique institution devoted to the research, teaching and
be defined as the collection, presentation and interpretation of application of Statis- tics. It is founded by Prof. Prasanta
4. Mathematics Chandra Mahalanobis in [Link] is known as father of indian
numerical data.”
Functions of Statistics statistics
29th june, birth day of P C Mahalanobis, is celebrated as
5. Psychology and Education
national statistics day. ISI publishes a journal ’Sankhya’.
1. Simplifies complexity. Economics and Statistics Department
6. Management Studies
The Directorate of Economics and Statistics is the nerve centre of
2. Presents facts in a definite and precise form. the Kerala State statistical system and it is the nodal agency
Some applied areas of Statistics are coordinating all statisical activities in the state.
3. Provides comparison.
• Actural Science
1
• Biostatistics
• Agricultural Statistics
The representative part of the population is known as sample. The 4. Telephone interview
method of col- lecting data from the sample is known as sampling
or sample survey. 5. Mailed questionnaires and schedules
The statistical survey may be either by census method or by
sampling method. The factors which can vary from one object to
another are called variables. 6. Focus group discussion
The variable which can not be numerically measured is called
qualitative [Link] variable which are numerically measured
is called quantitative variable. Questionnaires and schedules are series of questions arranged in
CHAPTER 2 If the variable takes specific values only, it is called discrete a logical order so as to collect information for a specified purpose.
variable. A continuous variable takes any value within the defined A questionnaire is usually mailed by post or email to selected
range of variables. [Link] approaches personally to the informant
A nominal scale of measurement is used to name categories and collects information from them.
such as gender,nation etc. In the ordinal scale of measurement, we Sources of secondary data
COLLECTION OF DATA
can put an order to the data according to the relation among the
values of variables. The data regarding a quantitative variable is a
• Government publications
cardinal data.
Based on the sources of collection, statistical data may be classified
as primary and secondary data. • Office records in panchayats, municipalities etc.
Primary data collected by the investigator for the first time for his/her
Data means any measurement, result, fact or observation which
own purpose. Data obtained from existing sources which may be • Survey reports of various research organizations
gives information. Data collection is the systematic gathering of data
published or unpublished are known as secondary data.
for a particular purpose from various sources.
Methods of primary data collection are • Survey reports in journals, newspapers and other publications
Statistical investigation includes collection, classification,
presentation, analysis and interpretation of data according to well
defined procedures. The person authorized to make investigation 1. Direct personal interview • Websites
is known as investigator.
The investigators depute some persons to collect the data from
the field. These per- sons are known as Enumerators. The 2. Indirect oral investigation
process is known as Enumeration.
A population consists of all elements, individuals, items or objects 3. Direct observation
whose charac- teristics are being studied. If data are collected
from each and every unit of the population, the investigation is
called census.
table used for one way classification is one way table. If we If two characteristics are measured simultaneously from each
consider two characteris- tics at a time for classification for data, it unit, it is known as bi variate data. The frequency distribution of a
is termed as two way classification of data and the table is two bi variate data is called bi variate frequency table
way table.
The number of repetitions of a particular observation in a series is
called frequency of the observation.
The series of observations in which items are listed individually is
called raw data or individual series. A discrete frequency table is
that series in which data are presented in a way that exact
CHAPTER 3 measurements of units are clearly shown.
Frequency tables with classes and corresponding frequencies are
known as continuous frequency table.
If the lower limit of the first class or upper limit of the last class
CLASSIFICATION AND TABULATION is not specified, then it is called open end class.
Frequency
Relative frequency = = f
Total frequency N
Frequency
Percentage frequency
Total =
frequency ×N 100= f × 100
• Chronological classification
• Geographical classification
CHAPTER 4 The most commonly used graphs for representing a frequency CHAPTER 5
distribution are
1. Histogram
DIAGRAMS AND GRAPHS CENTRAL TENDENCY
2. Frequency Polygon
3. Frequency Curve
Diagrams and Graphs are the methods for simplifying the The property of the observations in a data to cluster or
complexity of quanti- tative [Link] and Graphs are more 4. Ogives (cumulative frequency curves) concentrate around a value is known as Central Tendency.
attractive and impressive. Diagrams Commonly used diagrams Measures of central tendency (averages) are the values which gives
are an idea about the concentration of observations in the central part of
5. Scatter plot
the distribution.
1. Bar diagrams
There are four types of bar diagram. 1. Simple and rigid definition.
Σf
x¯ = (5.2)
x
N
1. Arithmetic Mean(AM)
2. Median
3. Mode
4. Geometric Mean(GM)
5. Harmonic Mean(HM)
where N = Σf is the total frequency. Median Modal class is the class having highest frequency. Mode = l +
(f1−f0)c
HM
wherel1 and l3 are the lower limits of
quartile classes, f1 and f3 are the
2
2. (GM ) = AM × HM frequencies of the quartile classes, c1 and
√ c3 are the class intervals of the quartile
or GM = AM × HM classes
and m1 and m3 are the cumulative frequencies preceding the
quartile classes.
Quartiles, deciles and percentiles are the partition values which
divide data into several equal parts. CHAPTER 6
Deciles divide the distribution into ten equal parts and there are 9
Quartiles divide a data into four equal parts. There are three deciles. Median is the 5th decile
quartiles, denoted by Percentiles divide a distribution into hundred equal parts. There are
99 percentiles. Median is the 50th percentile.
Q1, Q2 and Q3. Q2 is the A box plot is a graph of a data set that consists of a line extending DISPERSION
median. Quartiles for raw from the minimum value to the maximum value and a box with
data lines drawn at the Q1, the median, and Q3.
Arrange the n observations in ascending order of magnitude.
n+1 th
Q1 = value of
4 ( ) item in the series
Q = value of 3(n+1)th
item in the series Dispersion is the degree of scatter or variation of the variable about
3 4 a central value. Measures of dispersion are
( N − m1)c1 = H- L
Q 1 = l1 + 4
f1 19
20
SD σ
CV =Mean × 100 x̄ = × 100
Quartile Deviation(QD)
Q3−Q1
Q3−Q1 Coefficient of QD
Q +Q=
QD = 2 3 1
Q1 and Q3 are explained in the previous chapter for raw data, Covariance is a measure of strength of linear relationship
discrete frequency distribution and continuous frequency between two variables Cov(x,y) indicates whether the variables
distribution. are positively related or negatively related in a bi variate
Mean Deviation (MD) distribution.
Cov(x,y) = Σ(x−x̄)(y−ȳ)
n n =
Σxy
− x¯ × y¯
MD for raw data CHAPTER 7
MD = Σ|X−A|
n , where A is any
average. MD for discrete
frequency distribution
MD = Σf|X−A| , where A is any average and N is the total
SKEWNESS AND KURTOSIS
N
frequency. MD for continuous frequency distribution
MD = Σf|X−A|
N , where A is any average , N is the total frequency
and x is the mid value of the class.
Standard Deviation (SD)
SD for raw Skewness means the absence of symmetry in a data set. For a
data SD σ symmetric distribution Mean= Median= Mode.
Σ(x−x̄)2 n = Σx
n
2
− x¯2 There are two types of skewness ‘
=
SD for discrete frequency distribution
1. Positive skeness
SD σ = N Σfx − x¯2 , where N is the total frequency.
2
22
2. Bowleys coefficient of Skewness Sample space: The set of all possible outcome of random
experiment is called the sample space. Sample space is usually
Q3+Q1−2Median listed in curly brackets{} and is denoted by S.
SB =
Q3−Q1 Sample point: Each element in the sample space is called a
sample point. Events: An event is a set of outcomes which
3. Coefficient of Skewness based on moments have some characteristics in com- mon.
Equally likely event: Two or more events which have an
β = µ32 equally likely chance or equal probability of occurrence are
1 µ23 called equally likely events.
√ Mutually exclusive (disjoint) events: Events are said to be
γ1 = β 1 CHAPTER 8
mutually exclusive if the happening of any one of them
µ3 determines nature of skewness. If µ3 > 0, the distribution is excludes the happening of all the others. Exhaustive events:
positively skewed. If µ3 < 0, the distribution is negatively A set of events is called exhaustive,if all the events together
skewed. If µ3 = 0, the distri- bution is symmetric. consume the entire sample space.
PROBABILITY Basic properties of probability
Kurtosis
number ofexhaustive
cases n
• A and B → only A
′
Rule 2
Axioms on
Probability Axiom 1:
Non-negativity For
any event A, P (A) ≥
0 Axiom 2: Certainty
Bayes’ Theorem
• Cluster Sampling
If A1 and A2 are two mutually exclusive and exhaustive
events and A is any other event which can occur along with • Multi Stage Sampling