Chapter seven
7. Data processing analysis
7.1. Coding, editing and cleaning the data
7.2. Data analysis
7.3. Concept of Economic Modelling
7.4. Testing hypothesis: Falsification as a scientific
approach
Data processing and analysis
processing implies editing, coding, classification and
tabulation of collected data so that they are amenable
to analysis.
[Link]:-is a process of examining the collected raw data
(specially in surveys) to detect errors and omissions and to
correct these when possible. Can be classified into two. Field
editing and central editing
Field editing consists in the review of the reporting forms by
the investigator for completing (translating or rewriting) what
the latter has written in abbreviated and/or in illegible format
the time of recording the respondents’ responses. should be
done as soon as possible after the interview, preferably on the
very day or on the next day
Central editing
take place when all forms or schedules have
been completed and returned to the office
2. Coding:- refers to the process of assigning numerals or
other symbols to answers so that responses can be put into a
limited number of categories or classes
3. Classification: reducing a large volume of raw data into
homogeneous groups
4. Tabulation:- arranging the same data in some kind of
concise and logical order.
Thus, tabulation is the process of summarizing raw data and
displaying the same in compact form (i.e., in the form of
statistical tables) for further analysis.
What is Data Analysis?
Examining data for its relevance
Preparation of tables
Graphic display of information
Estimating the unknown
Establishing functional relationship between cause and effect
Understanding the Trends and making forecasts, Regession, Factor
Analysis, Cluster analysis … and many more!
Preparing a document stating the methodology and interpreting the
results
Meaning of data analysis
Analysis of data means to make the raw data meaningful or to draw
some results from the data after the proper treatment.
The ‘null hypotheses’ are tested with the help of analysis data so to
obtain some significant results.
Thus, the analysis of data involves estimating the values of unknown
parameters of the population and testing of hypotheses for drawing
inference
NEED FOR ANALYSIS OF DATA OR TREATMENT OF DATA
1. To make the raw data meaningful,
2. To test null hypothesis,
3. To obtain the significant results,
4. To draw some inferences or make generalization, and
5. To estimate parameters
Continued
Analysis may, therefore, be categorized as
1. Descriptive Statistical Analysis, and
2. inferential analysis (Inferential analysis is often known as
statistical analysis)
Descriptive Statistical Analysis
“Descriptive analysis is largely the study of distributions of
one variable( Unidemensional analysis)
The characteristics of location, spread, and shape describe
distributions. The common measures of location, often
called central tendency, include mean, median, and mode.
The common measures of spread, alternatively called
measures of dispersion, are variance, standard deviation,
and range. The common measures of shape are skewness
and kurtosis
Cont’d
• In respect of the measures of skewness and
kurtosis, we mostly use the first measure of
skewness based on mean and mode or on mean
and median.
• Other measures of skewness, based on quartiles
or on the methods of moments, are also used
sometimes.
• Kurtosis is also used to measure the peakedness
of the curve of the frequency distribution
• Calculating measures of relationship-coefficient of correlation,
Reliability and validity by the
• Inferential Statistical Analysis
• Inferential statistical analysis involves the process of sampling,
the selection for study of a small group that is assumed to be
related to the large group from which it is drawn.
• inferential statistics concern with the process of generalization
• is concerned with the various tests of significance for testing
hypotheses in order to determine with what validity data can be
said to indicate some conclusion or conclusions.
Cont’d
• Correlation analysis studies the joint variation of two or
more variables for determining the amount of
correlation between two or more variables.
• Causal (regression) analysis is concerned with the
study of how one or more variables affect changes in
another variable. It is thus a study of functional
relationships existing between two or more variables
What is Data Analysis?
Basing on the study of the variables the analysis is
divided into three types they are:
D a ta A n a lys is
U n ivaria te B iva ria te M u ltiva ria te
A n a lys is A n a lys is A n a lys is
X1 X1, X2 X1, X2 , …, Xn
One variable Two variables Multiple variables
at a time at a time at a time
What is Data Analysis?
Univariate Data Bivariate Data Multivariate Data
central tendency - mean, analysis of two variables Multiple Regression Analysis
mode, median simultaneously
Factor Analysis
dispersion - range, variance, correlations
Cluster Analysis
max, min, quartiles, standard
comparisons, relationships,
deviation. Discriminant Analysis
causes, explanations
frequency distributions
tables where one variable is
bar graph, histogram, pie contingent on the values of
chart, line the other variable.
graph, box-and-whisker plot
independent and dependent
variables
Descriptive Statistics
MEASURES OF CENTRALL TENDENCY
Mean
Score Score
Median (X) (Y)
Score
(Z)
Mode Average (X) 20 10 20
Geometric Mean = 20 18 15 21
Harmonic Mean
Average (Y) 19 25 19
= 20 21 20 21
MEASURES OF DISPERSION
S.D (X) = 1.56 22 30 22
Range & variance 100 100 Total 100
S.D (Y) = 7.07 Total
Quartile Deviation Mean 20 20 Mean 20
Mean Deviation
Standard Deviation
Coefficient of Variation
MEASURES OF SHAPE
Skewness = 0.85 Skewness = -0.25
Skewness = 0
Correlation Analysis
• Correlation is a statistical measure that indicates the
extent to which two or more variables fluctuate
together
• Amongst the measures of relationship, Karl
Pearson’s coefficient of correlation is the frequently
used measure in case of statistics of variables
• Multiple correlation coefficient, partial correlation
coefficient, regression analysis, etc., are other
important measures often used by a researcher
Correlation Analysis
X Y
SCATTER DIGRAM WITH POSITIVE
15 13.35 CORRELATION. r = + 0.5790
26 16.12 18
27 16.74 17
25 16 16
25.5 13.59 15
14
Y
27 15.73
13
32 15.65 12
18 13.85 11
22 16.07 10
10 20 30 40
20 12.8
26 13.65
X
24 14.42
PRICE QUANTITY SOLD
SCATTER DIAGRAM WITH NEGATIVE
4.5 125
CORRELATION r = - 0.6345
5.5 115
4.5 140
4.5 140
QUANTITY SOLD
155 4 150
135 5.5 150
115 5.5 130
95 6.5 120
75 5 130
2 4 6 8
5.5 100
PRICE 6 105
4.5 150
Concept of economic modeling
• An economic model is a simplified description of reality,
designed to yield hypotheses about economic behavior
that can be tested.
• Economic models generally consist of a set of mathematical
equations that describe a theory of economic behavior. or example,
y = f(x1, x2, ...), where y is understood to be a dependent variable,
the behavior of which is determined in some sense by one or more
independent variables, x1, x2, etc.
• It is necessarily subjective in design because
there are no objective measures of economic outcomes.
• The aim of model builders is to include enough
equations to provide useful clues about how rational agents
behave or how an economy works
• Modeling provides a logical,
abstract template to help organize the analyst's thought.
• Models constitute the primary vehicles for conducting
economic analysis.
• Economic theorists, attempting to push the frontiers of
economic knowledge, employ models to discern the
economic nature of the world.
• Shows the relationship between variables
1. Estimating a Model by Regression Analysis The simplest
form of regression model relates two variables, of which one
is described as dependent, and the other as independent. The
format for this simplest regression model may be described
in functional-
• y = a + bx;
[Link] Regression Models
• Linear with multiple variables
[Link]-linear Regression Model
• For example, a quadratic polynomial equation
includes linear and second- order (or squared)
terms in the format:
• y = a + b1x + b2x 2 .
• The general format for an k-th order polynomial
model is (5) y = a + b1x + b2x 2 + b3x 3 + ... +
bkx k ;
[Link] series models
•The End
Thank you