CHAPTER EIGHT
Data Processing and Analysis
The goal of any research is to provide information out of row data. The row data after collection
has to be processed and analyzed in line with the outline (plan) laid down for the purpose at the
time of developing the research plan. Response on measurement instruments (words, check mark
etc.) conveys little information as such. The compiled data must be classified, processed, analyzed
and interpreted carefully before their complete meanings and implications can be understood.
There are two stages of data analysis, data processing and analysis. Some authors do like to make
difference between processing and analysis. However we see them separately these terms briefly
8.1. Data processing
Data possessing implies editing, coding, classification and tabulation of collected data so that they
are amendable to analysis.
Editing: Is a process of examining the collected raw data to detect errors and omission (extreme
values) and to correct those when possible
It involves a careful scrutiny of completed questionnaires or schedules
It is done to assure that the data are
o Accurate
o Consistent with other data gathered
o Uniformly entered
o As complete as possible
o And has been well organized to facilitate coding and tabulation
Editing can be either field editing or central editing
Field editing: Consist of reviewing of the reporting forms by the investigator for completing what
has been written in abbreviation and/ or in illegible form at a time of recording the respondents’
response
This sort of editing should be done as soon as possible after the interview or observation.
1
Central editing: It will take place at the research office. Its objective is to correct errors such as
entry in the wrong place, entry recorded in month
Coding: Refers to the process of assigning numerical or other symbols to answers so that responses
can be put into a limited number of categories or classes. Such classes should be appropriate to the
research problem under consideration.
There must be a class of every data items. They must be mutually exclusive (a specific answer can
be placed in one and only one cell in a given category set)
Coding is necessary for efficient analysis and through it several replies may be reduced to a small
number of classes, which contain the critical information required for analysis
E.g., Closed end question
1 [ ] Yes
2 [ ] No
Or
Less than 200 [ ] 001
201- 699 [ ] 002
1500 and more [ ] 006
Coding is used when the researcher uses computer to analyze the data otherwise it can be avoided.
Classification: Most research studies result in a large volume of raw data, which must be reduced
into homogeneous group. Which means to classify the raw data or arranging data in-groups or
classes on the basis of common characteristics?
Data Classification implies the processes of arranging data in groups or classes on the basis of
common characteristics. Data having common characteristics placed in one class and in this way
the entire data get divided into a number of groups or classes.
Classification according to attributes: Data are classified on the basis of common
characteristics, which can either be descriptive (such as literacy, sex, honesty, etc) or numerical
(such as, weight, age height, income, expenditure, etc.). Descriptive characteristics refer to
qualitative phenomenon, which cannot be measured quantitatively: only their presence or absence
2
in an individual item can be noticed. Data obtained this way on the basis of certain attributes are
known as statistics of attributes and their classification is said to be classification according to
attributes.
Classification according to class interval: Unlike descriptive characteristics the numerical
characteristics refer to quantitative phenomenon, which can be measured through some statistical
unit. Data relating to income, production, age, weighted, come under category. Such data are
known as statistics of variables and are classified on the basis of class interval. Fore example,
individuals whose incomes, say, are within 1001-1500 Birr can form one group, those whose
incomes within 500-1000 Birr form another group and so on. In this way the entire data may be
divided into a number of groups or classes or what are usually called, class interval. Each class-
interval, thus, has an upper as well as lower limit, which is known as class limit. The difference
between the two-class limits is known as class magnitude. The number of items that fall in a given
class is known as the frequency of the given class. All the classes with their respective frequency
are taken together and put in the form of table are describing as group frequency distribution or
simply frequency distribution. Classification according to class intervals usually involves the
following problems.
How many classes should be there? What should be their class size (magnitude)? The
answer is left to the skill and experience of the researcher. However, the objective
should be to display the data in such a way as to make it meaningful to the analyst.
Concerning the class size, each group is expected to have equal size. Multiples of 2.5
and 10 are generally preferred while determining the class size. Some statistician adopts
the following formula.
(i = R/ (1+3.3 log/N)
Where, I = class size
R = Range (i.e, difference between the value of the largest item and smallest item among the items
to be grouped.
N = Number of item to grouped
Some problems in processing
3
Don’t know (DK) Responses: During data processing, the researcher often comes across some
responses that are difficult to handle. Don’t know (DK) is one example of such responses. When
the DK response group is small, it is of little significance. But when it is relatively big, it becomes
a matter of major concern.
How the DK responses are to be dealt with by researcher?
Prevention is the best!
The best way is to design better types of question. Good rapport (understanding) of interviews
with respondents will result in minimizing DK response.
But what about the DK responses that have already taken place?
One way to tackle this issue is to estimate the allocation of DK answers from other data in the
questionnaire
The other way is to keep DK responses as a separate replay category if DK response happens to
be legitimate, otherwise we should let the reader make his own decision.
8.2. Analysis
Data analysis is further transformation of the processed data to look for patterns and relations
among data groups.
By analysis we mean the computation of certain indices or measures along with searching for
patterns or relationship that exist among the data groups. Analysis particularly in case of survey
or experimental data involves estimating the values of unknown parameters of the population and
testing of hypothesis for drawing inferences.
Analysis can be categorized as
Descriptive Analysis
Inferential (Statistical) Analysis
8.2.1. Descriptive analysis:
Descriptive analysis is largely the study of distribution of one variable. Analysis begins for most
projects with some form of descriptive analysis to reduce the data into a summary format.
4
Descriptive analysis refers to the transformation of raw data into a form that will make them easy
to understand and interpret.
Descriptive response or observation is typically the first form of analysis. The calculation of
averages, frequency distribution, and percentage distribution is the most common form of
summarizing data.
The most common forms of describing the processed data are:
Tabulation
Percentage
Measurements of central tendency
Measurements of dispersion
Measurement of asymmetry
Data transformation and index number
Tabulation: Refers to the orderly arrangement of data in a table or other summary format. It
presents responses or the observations on a question-by-question or item-by-item basis and
provides the most basic form of information. It tells the researcher how frequently each response
occurs
This starting pint of analysis requires the counting of responses or observations for each of the
categories. E.g., Frequency tables,
Need for tabulation
It conserves space and reduces explanatory and descriptive statement to a minimum
It facilitate the process of comparison
It facilitate the summation of items and the detection of errors and omission
It provide basis for various statistical computation,
Percentage: Whether the data are tabulated by computer or by hand, it is useful to have
percentages and cumulative percentage. Table containing percentage and frequency distribution is
easier to interpret. Percentages are useful for comparing the trend over time or among categories
5
Measure of central tendency: Describing the central tendency of the distribution with the mean,
median or mode is another basic form of descriptive analysis.
These measures are most useful when the purpose is to identify typical values of a variable or the
most common characteristics of a group. Measure of central tendency is also known as statistical
average. Mean, median and mode are most popular averages.
Mean (arithmetic mean) is the common measure of central tendency
Mode is not commonly used but in such study like estimating the popular size of shoes it can be
used
Median is commonly used in estimating the average of qualitative phenomenon like estimating
intelligence.
Measurement of dispersion: Is a measurement how the value of an item scattered around the true
value of the average.
Average value fails to give any idea about the dispersion of the values of an item or a variable
around the true value of the average.
After identifying the typical value of a variable the researcher can measure how the value of an
item is scattered around the true value of the mean. It is a measurement of how far is the value of
the variable from the average value. It measures the variation of the value of an item. Important
measures of dispersion are:
Range: Measures the difference between the maximum and the minimum value of the observed
variable
Mean deviation: It is the average dispersion of an observation around the mean value. (Xi –
X)/n
Variance: It is mean square deviation. It measures the sample variability.
Measurement of asymmetry (skew-ness):: When the distribution of items is happen to be
perfectly symmetrical, we then have a normal curve and the relating distribution is normal
distribution. Such curve is perfectly bell shaped curve in which case the value of Mean = Median
= Mode
6
Under this condition the skew-ness is altogether absent. If the curve is distorted (whether on the
right or the left side), we have asymmetric distribution this indicates that there is a skew ness.
Z=M=X
If the curve is skewed on the right side we call it positive skew ness
Positively skewed data
XMZ
X is mean, M and Z is mode
In such case Z M X
But when the curve is skewed toward left, we call it negative skew ness.
Negatively skewed data
X M Z
And X M Z
7
Skew-ness is, thus a measurement of asymmetry and shows the manner in which the items are
clustered around the average. In a symmetric (normal distribution) the items show a perfect balance
on either side of the mode, but in a skewed distribution the balance is skewed one side or distorted.
The amount by which the balance exceeds on one side measures the skew-ness.
Knowledge about the shape of the distribution is crucial to the use of statistical measure in research
analysis. Since most method make specific assumption about the nature of distribution.
Data transformation: It is the process of changing original form of data to a form that is more
suitable to perform a data analysis that will achieve the research objective. The researcher often
modifies the value of scalar data or even create new variable
Index numbers: Most of the time, financial information (price, value of output, interest rate, and
exchange rate) will be adjusted for possible price changes by using index numbers (like CPI, PPI).
An index number is a number, which is used to measure the level of a given phenomenon at some
standard date.
Index numbers measures only the relative changes.
Different indices serve different purpose
Commodity index serves as a measure of changes in the phenomenon on that commodity
only
Some index numbers are used to measure cost of living (CPI)
In economic sphere they are often termed as economic barometer
Scores of observation are recalibrated so that they may be related to certain base period or base
number. Most commonly used index number to reduce the influence of price change on our
observation is CPI
Researcher also uses index numbers to make comparison between observations. When series (data)
are expressed in same units, we can use, averages for the purpose of comparison. But two or more
series are expressed in different units; statistical average cannot be used to compare them. By
converting numbers in to index number we can make comparison between two or more series.
8
8.2.2. Inferential Analysis
Most researcher wishes to go beyond the simple tabulation of frequency distribution and
calculation of averages and / or dispersion. They frequently conduct and seek to determine the
relationship between variables and test statistical significance. When the population is consisting
of more than one variable it is possible to measure the relationship between them.
If we have data on two variables we said to have a bivariate variable, if the data is more than two
variables then the population is known as multivariate population. If for every measure of a
variable, X, we have corresponding value of variable, Y, the resulting pairs of value are called a
bivariate population
In case of bivariate or multivariate population, we often wish to know the relationship between the
two or more variables from the data obtained.
E.g., we may like to know, “Whether the number of hours students devote for study is
somehow related to their family income, to age, to sex, or to similar other factors.
There are several methods of determining the relationship between variables.
Two questions should be answered to determine the relationship between variables.
1. Is there exist association or correlation between the two or more variables? If yes, then up to
what degree?
This will be answered by the use of correlation technique. Correlation technique can be different
In case of bivariate population correlation can be found using
Cross tabulation
Karl Pearson’s coefficient of correlation: It is simple correlation and commonly used
Charles Spearman’s coefficient of correlation
In case of multivariate population correlation can be studied through:
Coefficient of multiple correlation
Coefficient of partial correlation
9
2. Is there any cause and effect (causal relationship) between two variables or between one variable
on one side and two or more variables on the other side?
This question can be answered by the use of regression analysis. In regression analysis the
researcher tries to estimate or predict the average value of one variable on the basis of the value of
other variable. For instance a researcher estimates the average value score on statistics knowing a
student’s score on a mathematics examination.
There are different techniques of regression.
In case of bivariate population cause and effect relationship can be studied through simple
regression.
In case of multivariate population: Causal relationship can be studied through multiple
regression analysis.
Time series Analysis; Successive observations of the given phenomenon over a period of time
are analyzed through time series analysis. It measures the relationship between variables and time
(trend)
Time series will measure seasonal (seasonal fluctuation), cyclical irregular fluctuation, and Trend.
The analysis of time series is done to understand the dynamic condition of achieving the short term
and long-term goal of business firm for forecasting purpose
The past trend can be used to evaluate the success or failure of management or any other policy.
Based on past trend the future patterns can be predicted and policy may accordingly be formulated.
10
CHAPTER NINE
Interpretation and Reporting the Research Result
After collecting and analyzing the data, the researcher has to accomplish the task of drawing
inferences followed by the report writing. Interpretation has to be done carefully so that misleading
conclusion will not be drawn and the whole purpose of doing research will not be vitiated.
It is through interpretation that the researcher can expose relations and processes that underline his
findings. If hypotheses are tested and upheld (confirmed), the researcher may arrive at
generalization.
But incase the researcher had no hypothesis to start with; he would try to explain his findings on
the basis of some theory.
All the analytical information and consequential inferences may well be communicated, preferably
through research report, to the consumers of research results who may be either an individuals or
groups or some public or private organization.
9.1. Meanings and Technique of interpretation of interpretation
Interpretation refers to the task of drawing inferences from the collected facts after analytical or
experimental study.
The task of interpretation has two parts or has two major aspects
1) The effort to establish continuity in research through linking the results of a given study with
those of others.
2) The establishment of explanatory concept.
In one sense, interpretation is concerned with relationships within the collected data, partially
overlapping analysis.
Interpretation also extends beyond the data of the study to include the results of other research,
theory hypothesis.
11
Why interpretation?
Interpretation is considered as a basic component of research process because of the following
reasons: It is through interpretation that the researcher can well understand the abstract principle
that works beneath (beyond) his findings.
It will lead to the establishment of explanatory concepts that can serve as a guide for further
research study. It opens new avenues of intellectual adventure and stimulates the quest for more
knowledge.
Researcher can only be better appreciated only through interpretation why his findings are what
they are and can make others to understand the real significance of his research findings. The
interpretation of exploratory research often results into hypothesis for experimental research.
Technique of interpretation
The task of interpretation is not an easy job. Rather it requires a good skill on the part of researcher.
Interpretation is an art that one learns through practice and experience. The researcher may, at
times, seek the guidance from experts for accomplishing the task of interpretation.
There are no existing rules to guide the researcher about how to interpret the data.
However, the following suggested steps could be helpful.
1) Researcher must give reasonable explanation of the relation, which he has found and he must
interpret the lines of relationship in terms of the underlying processes and must try to find out
the thread of uniformity that lies under the surface layer of his diversified research findings.
2) Extraneous information, if collected during the study, must be considered while interpreting
the final result of research study, for it may prove to be a key factor in understanding the
problem under consideration.
3) It is advisable, before embarking upon final interpretation, to consult some one having insight
into the study and who is frank and honest and will not hesitate to point out omissions and
errors in logical argumentation. Such a consultation will result in correct interpretation and,
thus, enhance the utility of research result.
4) Researcher must accomplish the task of interpretation only after considering all relevant
factors affecting the problem to avoid false generalization.
12
He must not be in hurry while interpreting results, for quite often the conclusion, which appear to
be all right at the beginning, may not at all be accurate.
Precaution in interpretation
Researcher must pay attention to the following points for correct interpretation.
At the outset, researcher must invariably satisfy himself that: the data are appropriate, trust
worthy and adequate for drawing inferences. The data reflect good homogeneity (no
extreme) and proper analysis has been done through statistical or any other methods.
The researcher must remain cautious about the errors that can possibly arise in the process
of interpreting results. Error can arise due to
False generalization and/or due to wrong interpretation of statistical measures, such
as:
The application of findings beyond the rang of observation
Identification of correlation with causation and the like
He should be well equipped with and must know the correct use of statistical measures
for drawing inferences concerning his study.
Broad generalization must be avoided, because the coverage restricted to a particular time, a
particular area and particular condition. Such restriction, if any, must invariably be specified
and the result must be framed within their limit.
The researcher must remember that there should be constant interaction between initial
hypothesis and, empirical observation and theoretical conceptions. It is exactly in this area of
interaction between theoretical orientation and empirical observation that opportunity for
originality and creativity lies. (V. Young, 1849)
9.2. Reporting the research result
Writhing report is the last step in a research study and requires a set of skills somewhat different
from those called for in research of the earlier stages of research. This task should be accomplished
by the researcher with at most care. He may also seek the assistance and guidance of experts for
the purpose. The research task remains incomplete till the report has been presented and/or written.
13
Even the most brilliant hypothesis, well-designed and conducted research study, and the most
striking generalization and findings are of little importance unless they are effectively
communicated to others.
The purpose of research is not well served unless the findings are made known to others.
Layout of research report
Layout of the report means as to what the research report should contain and look like. A
comprehensive layout of the research report should comprise
Preliminary pages
The main text
The end matter
1) Preliminary pages
In this part the report should carry
Title
Acknowledgment (this can be in the form of preface and forward, in larger study)
Table of content
List of tables (figures) 1
2) Main text
The main text provides the complete outline of the research report along with all details. Title of
the research is repeated at the top followed by abstract and then follows the other details on pages
numbered consecutively beginning with second page. Each main section of the report should begin
on a new page.
Main text can have the following sections
Introduction
o Background of the study
1
Preliminary pages are commonly numbered by Roman numbers
14
o Rationale
Objectives
Literature Review
Material and Methodology
o Data (or material)
o Methodology used,
o Limitation of the study
Results and discussion (in some cases, Empirical Analysis)
Summary, Conclusion and Recommendation or
o Concluding comment or
Since, some of the main sections of the report have been explained in some detail in chapter four
section two, here attempts were made to explain only selected parts of the report, which need
special attentions.
Introduction: the major subdivisions of this part are generally the ones shown in the proposal:
statement of the problem, significance of the study, and the organization of the study. This part of
the study should be lucid complete and concise. It has to be written in a lively and stimulating
manner in order to arouse the interest of the reader to go through the report.
Literature Review: this is a section for documentation with insight theoretical and empirical
investigation that had been carried out as related to the study at hand
Material and Methodology or Data and Methodology: this part includes detailed description
of the manner in which decision have been made about the type of data needed for the study, the
tools and approaches used for their collection and the method by which they have been collected,
justification of the selection of the particular method of data collection. Definition of the
population, the sampling techniques used to select sample elements with its full justification, the
size of the sample and the rational for the size, statistical tools used to analyze the data the rational
for using them will be dealt in detail in this section.
15
Limitations: No report is perfect, so it is important to indicate its implications. If there were
problems with non-response errors, or sampling procedures, they should be discussed.
The discussion of limitation should avoid overemphasizing the weakness, though
Its aim should be to provide a realistic basis for assessing the results.
Result and Discussion: A detailed presentation of the findings of the study (the results of the data
analysis) with supporting data in the form of tables and charts together with a validation of results.
In other words in this section the data is presented in tables and figures followed by narrative
discussion and justifications. Two things may require special attention while writing this part of
the report.
Tables that are too lengthy may better be placed in the appendix
Tables and figures should be explained. As tables and figures are expected to be self
explanatory, the textual discussion should not be a duplicate of the table. Only
important facts that lead to generalization will be discussed.
This section generally comprises the body of the report, extending over several sub-sections.
It should contain statistical summaries and reductions of the data rather than the raw data. All
results should be presented in logical sequences and divided into readily identifiable sections. All
relevant results must find a place in the report.
Summary and Conclusion: Toward the end of this section, the researcher should again put down
the results of his research clearly and precisely. This part begins with a brief restatement of the
problem, the hypothesis, description of the problem and discussion of findings and conclusion of
the study. Most readers skip other details of the report and may prefer to read only this part in
order to get an overview of the study and judge its relevance. Thus, it should be written with
maximum diligence, clarity and brevity. Moreover, this section must focus attention to
Announce the acceptance or the rejection of the stated hypothesis.
Simply unanswered question that were raised in due course of the study and which
required further investigation in there are relevant to this part.
A researcher should also state the implication that flows from the results the study for the general
reader is interested in the implication that for understanding the human behavior.
16
Such implication may have three aspects as stated below:
A statement the inferences drawn from the present study which may be expected to apply
in similar circumstances
The condition of the present study, which may limit the extent of the legitimate
generalization of the inferences drawn from the study.
The relevant questions that still remain unanswered or new questions rose by the study
along with suggestion for the kind of research that would provide answer for them.
Generally, it is considered as a good practice to finish the report with a short conclusion, which
summarize and recapitulates the main points of the study. The conclusion drawn from the study
should be clearly related to the hypothesis or the problem that are stated in the introductory section.
At the same times, a forecast of the problem future of the subject and indication of the kind of
research, which needs to be done in those particular fields, is useful and desirable. Conclusions are
opinion based on the results, where as recommendations are suggestions for action.
Recommendation: In accordance with the result of the outcome of the research work a researcher
may forward (suggest) possible solution that may alleviate the problem in question. The
recommendation to be acceptable it should meet the following requirements;
Should be clear an unambiguous
Need to be realistic, plausible and operational
Should point out the responsible body to translate the suggested solution into practice
Should be modest than assertive
3) End matter:
Here belong sections like: References (bibliography): It should be based on alphabetical listing
of names and Appendix
17