0% found this document useful (0 votes)
381 views88 pages

Data Analysis in Research Methodology

Uploaded by

Pihu Salam
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
381 views88 pages

Data Analysis in Research Methodology

Uploaded by

Pihu Salam
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PPTX, PDF, TXT or read online on Scribd

DATA ANALYSIS

INTRODUCTION

Analysis of data is the most important phase


of the research process, which involves the
computation of the certain measures along
with searching for patterns of relationship that
exists among data groups. Analysis of data
includes compilation, editing, coding,
classification, and presentation of data.
DEFINITION

1. Analysis is the process of organizing and


synthesizing the data so as to answer research
questions and test hypothesis.
[ACC. TO S.K. SHARMA]
2. Analysis is the process of breaking a complex topic
into smaller parts to gain a better understanding of it.
[ACC. TO S.K. SHARMA]
PURPOSE OF DATA ANALYSIS

• To make the raw data meaningful.


• To test null hypotheses.
• To test the statistical significance of data .
• To draw inferences and make generalizations.
• To estimate parameters.
OBJECTIVES OF DATA ANALYSIS

• Difference between qualitative and


quantitative data and analysis.
• Analyze data gathered from questionnaires.
• Analyze data gathered from interviews.
• Analyze data gathered from observation
studies.
• Presenting your findings
IMPORTANCE OF DATA ANALYSIS
• It helps in structuring the findings from
different sources of data.
• It is very helpful in breaking a macro problem
into micro parts.
• It acts like a filter .
• It helps in keeping human bias away from the
research conclusion with the help of proper
statistical treatment.
SCALES OF MEASUREMENTS
1. NOMINAL-LEVEL MEASUREMENT

• It is the lowest of the four levels of measurements.


• It involves the assignment of number to represent
categories or classes of things.
• It consists of categories that are not more or less
than each other but are different from one another
in some way. In other words, they have no
quantitative values.
• Example of nominal-level measurement
• Gender :- Male , Female
• Habitat :- Urban , Rural
2. ORDINAL-LEVEL MEASUREMENT

• It ranks objects based on their relative standing


on a specific attribute.
• The order of ranking is imposed on categories.
• It reflects only magnitude and does not equal
intervals or an absolute zero point.
• Examples
• Health status :- Poor , Good , Excellent
• Income status :- Low income , Middle income ,
Upper income
[Link] INTERVAL-LEVEL MEASUREMENT

• In this level, there is specification of ranking


of objects on an attribute of the distance
between those objects.
• There is , more or less, equal numerical
distance between intervals.
• There is no absolute zero because there is no
zero point or absence of concept.
• Examples of Interval-Level Measurement:-
• Temperature :- 10°-80° , 40°-50°
4. RATIO-LEVEL MEASUREMENT

• It is the highest level for measurement.


• It has all the three attributes: magnitude, equal
intervals, and absolute zero point.
• It represents continuous values.
• Examples of Ratio-Level Measurement :-
• Biophysical parameters :- Weight , Height
TYPES OF DATA ANALYSIS
There are two types of data analysis.

?
?
1.) QUALITATIVE DATA ANALYSIS
• DEFINITION :- Qualitative data analysis is a
process of gathering, structuring and interpreting
qualitative data to understand what it represents.
Qualitative data is non- numerical .
• PURPOSE:-
• Any type of research that produces findings not
arrived at by statistical procedures
• Persons, lives, experiences, emotions
• Qualitative data analysis - an inductive process of
organising data into categories / themes and
identifying patterns among them.
IMPORTANCE OF QUALITATIVE DATA
ANALYSIS

• It allows researchers to explore new ideas and


hypotheses
• It allows researchers to gain a more in- depth
understanding of their subject matter.
• This can lead to the development of new
research questions that can be tested in future
studies.
TYPES OF QUALITATIVE DATA
ANALYSIS TEST
1. Chi-squaretest
2. Sign test
3. Wilcoxon signed-rank test
4. Mann-whitney u-test
1.) CHI-SQUARE TEST

Definition :- This is nonparametric test used


to find out the association between two events
in binomial or multinomial samples. It is
represented by a symbol 'x' and is used to find
out association between two discrete attributes.
For example, this test can be used if one wants
to find the association between smoking in
pregnancy and low-birth-weight babies, blood
pressure and renal diseases, and obesity and
coronary diseases, etc.
Characteristic of Chi-Square Test :-

1. It is positively skewed.
2. It is non-negative.
3. It is based on degrees of freedom.
4. When the degrees of freedom change a new
distribution is created.
Steps to Chi-SquareTest

• Identify the Hypothesis


• Calculate Chi-Square Statistic
• Determine the Level of Significance
• Determine Degrees of Freedom pital.
• Create Contingency Table
• Find the Critical Value
• Calculate Expected Frequencies
• Interpret Results
Question:-

500 elementary school boys and girls are


asked which is their favourite colour blue
green or pink.

50
30 50 120
1. Identify the Hypothesis :-
1. The null hypothesis (H0) is that there is no
association between the two variables.
2. The alternative hypothesis (H1) is that there is
an association of any kind.
Cont..
• [Link] Chi-Square Statistic :- To
calculate the Chi-square statistic, follow these
steps:
• 1. Create a contingency table of observed
frequencies for each category.
• 2. Calculate expected frequencies for each
category under the null hypothesis.
OBSERVED VALUE TABLE

50
30 50 120
130 200 170
Expected value table

Expected Blue Green Pink ROW

Boys 78 120 102 300

Girls 52 80 68 200

COLUMN 130 200 170 N=500


Cont...
3. Compute the Chi-square statistic using the
formula: Χ² = Σ [ (O – E)² / E], where O is the
observed frequency and E is the expected
frequency.
[Link] the Level of Significance :-
• The significance level, (α).
• If the significance level is 10%, or α = 0.1
Cont...
[Link] Degrees of Freedom :-
The degrees of freedom for the chi-square are
calculated using the following formula:
df = (r-1)(c-1) where r is the number of rows and c is
the number of columns.
[Link] Contingency Table :-
[Link] the Critical Value :-
Cont...
[Link] Results :- The interpretation of
results should include a discussion of the
evidence gathered. This includes aspects of
validity, strengths and weaknesses of the
evidence, possible sources of bias that may be
present in the included studies, and the
potential bias of the review.
[Link] TEST

Definition
The sign test is probably the simplest of all the
nonparametric methods. This test is utilized in
comparing a single sample with some
hypothesized value. Therefore, this test is of use
in situations in which the one-sample or paired t-
test might conventionally be applied. The name
sign test is so given as it attributes a sign, either
positive (+) or negative (-), to every observation.
Objectives

✓ Perform the sign test for a single population

✓ Carry out the sign test for matched-pair data


median.

✓ Perform the sign test for binomial data


from two dependent samples.
Steps in performing sign test
Following steps are needed while conducting
the sign test.
Step I: Mention the hypothesized value for
comparison and the null hypothesis in
particular.
Step II: Attribute a sign (+/-) to every
observation as per whether it is less or greater
than the hypothesized value.
Cont....
Step III: Find out
N+ = The number of observations that are
greater than the hypothesized value.
N- = The number of observations that are less
than the hypothesized value.
S = The smaller of N+ and N-
Step IV: Determine an appropriate p-value
(Using binomial distribution table ).
QUESTION:-
The scores obtained by employees of
private sector and public sector bank
work motivation are given as following.
Employees from private Employees from public
sector bank sector bank

12 23
13 12
13 12
15 14
17 14
18 21
19 23
21 21
22 23
21 21
11 20
23 20
[Link] SIGNED-RANK TEST

Definition:-
The Wilcoxon signed-rank test is a non-
parametric rank test for statistical hypothesis
testing used either to test the location of a
population based on a sample of data, or to
compare the locations of two populations
using two matched samples.
The Wilcoxon Signed Rank test comprises
five basic steps
Step I: Mention the hypothesized value for
comparison in particular and the null
hypothesis.
Step II:Ignoring their sign, rank all the
observations in increasing order of
[Link], ignore observations that
are similar to the hypothesized value.
If two observations have similar magnitude,
they are provided an average ranking,
irrespective of sign.
Cont...
Step III: Attribute a sign (+/-) to each observation
as per whether it is greater or less than the
hypothesized value (as in the sign test).
Step IV: Determine
a. R+ = Addition of all positive ranks
b. R- = Addition of all negative ranks
c. R = Less than R+ and R-
Step V: Determine an appropriate p-value.
[Link]-whitney u-test
Definition:-
Mann-Whitney U test is the non-parametric
alternative test to the independent sample t-
test. It is a non-parametric test that is used to
compare two sample means that come from
the same population, and used to test whether
two sample means are equal or not.
Steps mann-whitney u-test

Step I:
Give rank to all observations in increasing order of their
magnitudes by ignoring which group they come from. If
two observations are of same magnitude, irrespective of
group, they are provided an average ranking.
Step II:
Sum up the ranks in the smaller of the two groups (S). If
the two groups are similar in size then either of them can be
chosen.
Step III: Determine an appropriate p-value
2. QUANTITATIVE DATA ANALYSIS

DEFINITION :- Quantitative data are


information in numeric form. theycan either be
counted (such as the number of people
whoattend a training) or compared on a
numerical scale (suchas the number of training
participants who said that a training was “very
helpful” or “somewhat helpful”).
OBJECTIVE

• Aims to achieve an in-depth understanding of


a situation/topic/ issue.
• Classify features, count them, and construct
statistical models in an attempt to explain what
is observed.
• Researcher knows clearly in advance what
he/she is looking for.
CHARACTERISTICS

• The data is usually gathered using more


structured research instruments.
• The results are based on larger sample sizes
that are representative of the population.
• All aspects of the study are carefully designed
before data is collected.
• Data are in the form of numbers and statistics.
• Researcher uses tools, such as questionnaires
or equipment to collect numerical data.
IMPORTANCE

• Qualitative data provides the means by which


analysts can quantify the world around them.
• Quantitative research is a powerful tool for
anyone looking to learn more about their
market and customers.
• Quantitative data analysis helps to summarize
and interpret numerical results from close-
ended questions to understand what is
happening.
TYPES OF QUANTITATIVE DATA
ANALYSIS TEST
There are many types of quantitative data
analysis.
• Mean
• Median
• Mode
• Anova
• T test
• Z test
1. MEAN

Definition:-The arithmetic mean is the quantity


obtained by summing two or more numbers and
then dividing by the total number of numbers. It
is also known as average or average value.
Arithmetic mean is represented by X.
Importance of mean :-
• 1. It's the most simplest method .
• 2. In this calculation importance is given to every
individual value.
• 3. Importance is given to every data's value.
Cont...
Formula:-
Sum of the values (Σx)
X =
Number of the values (n)
Merits of Mean

Merits of Mean
• It is a simple average to understand and easiest
to compute.
• It is affected by the value item in a series.
• It is a reliable method of calculating average.
• Calculate value that is not based on the
position of series.
Demerits of Mean

• Very small and very large items usually affect


the value of average.
• In the distribution with open-ended classes,
values of mean cannot be computed without
• making assumptions.
• It is not always 'a good measure of central
tendency
2. Median
Definition
Median of a set of values is the middle-most
value when the data is arranged in ascending
order of magnitude. The middle value will
divide the number of observations in the data
into two equal parts. The median is denoted by
M. It is also called positional average
Importance

• Median is important because it gives us an


idea of where the center value is located in a
dataset.
• It refers to the middle value of distribution.
• One-half of items in the distribution have a
value larger than the median value.
• One-half of items in the distribution have a
value smaller than the median value.
Formula

Merits of Median
•It is useful in case of open-ended and unequal classes.
•Extreme values do not affect the median.
•Most appropriate average in dealing with qualitative
data.
•The value of median can be determined graphically.
Demerits of Median

• For calculating the median it is necessary to


arrange the data.
• Since it is the positional averages, the value is
not determined by all the observations.
• It is not capable of further algebraic treatment.
• Median is not calculated for quantitative data.
[Link]

Definition:-
Mode is most frequently occurring number in
a data set In other words, it is the value which
has the highest frequency in the data, In other
words, the mode of the distribution value is the
point around which the items tends to be most
heavily concentrated.
Important of mode

1. A nominal statistics.
2. An inspection average.
3. The most frequently occurring score.
4. Usually occurs near the center of the
distribution.
5. The most popular scores.
Formula of mode

( f - f1 )
Z = L+ *w
2 ( f – f1 –f2 )
Cont....
Where L = lowest limit of the modal class
f = frequency of the modal class
w = width of the modal class
f₁ = frequency of class just before the modal
class
f2 = frequency of class just after the modal
class. Modal class is the one which has
the highest frequency.
Merits of Mode

• It is not affected by extreme values.


• It can be used to describe the qualitative
phenomenon.
• Values of mode can be determined graphically.
Demerits of Mode

• The value of mode cannot always be


determined.
• It is not capable of algebraic manipulation.
• It is not based on each value.
4. ANOVA

Definition:-
When a researcher wants to compare the
difference between more than two samples
means. t-test will not be useful and an
alternative inferential statistical will be can be
fulfilled by test known as analysis of variance
(ANOVA) test.
Importance of ANOVA

• The ANOVA is an important test because it


enables us to see for example how effective
two different types of treatment are and how
durable they are.
• Effectively a ANOVA can tell us how well a
treatment work.
• To test the significance between variance of
two samples.
• It is used in testing of correlation & regression.
Formula od ANOVA

Formula od ANOVA

MST
F=
MSE
Advantage of ANOVA

■ It overcomes the limitation of T-test (not


applicable on more than two group).
■ An ANOVA control TYPE I Error remains at 5%.
■ ANOVA evaluates both between and within
variance and calculates its ratio
Disadvantages of ANOVA

* All population means from each data group


must be (roughly) equal.
* All variances from each data group must be
(roughly) equal.
* The normal subject-to-subject variation may
strongly affect the error sum of squares.
5. T- TEST

Definition:- It Is Applied To Find The


Significant Difference Between Two Means.
Importance Of T – Test:-
• The calculations of a confidence interval for a
sample mean
• To test whether a sample mean is different
from a hypothesized value
• To compare mean two samples
• To compare two sample means by group
Formula of T test

( X1 - X2 )
T - Test =
SE

Types Of T- Test:-
The t-test is further divided in two types.
[Link] test
2. Unpaired test
Cont..
1. Paired t-test: It is applied on paired data of
independent observations made on same sample
before and after the intervention, for example,
comparing mean oxygen saturation among same
group of patients before and after suctioning with
a new device. Paired test is found to be more
commonly used in nursing research studies.
Formula of paired test:-
Cont..
2. Unpaired t-test: It is also known as independent
sample t-test. It is applied when we obtain data from
subjects of two independent separate groups of people
or samples drawn from two different populations, for
example, when you are comparing post-interventional
mean systolic blood pressure of male and female
participants.
Formula of unpaired test:-
(SD1)2 (SD2)2
unpaired t- test = +
n1 n2
Merits of t – test

• It reduces the possibility of guising the correct


answer
• It covers greater amount of contents than
matching types test.
• The ability to handle small sample sizes .
Demerits of t – test

• It is only appropriate for questions that can be


answered by short responses
• There is a difficult in scoring when the
questions are not prepared properly and
clearly.
• it does not give accurate results on large
datasets .
6. Z – Test

Definition:-
When a sample is larger than 30 subjects, and
a researcher wants to compare the difference
in population mean and a sample mean or the
difference between two sample means, then Z-
test is applied. .
Uses Of Z Test

• To test the population proportion.


• To test the equality of two sample proportions.
• To test the population SD when the sample is
large.
• To test the equality of two sample standard
deviations when the samples are large or when
• population standard deviations are known.
Formula Of Z – Test

Mean (X) – Population (µ)


Z =
SE of sample mean
Advantages

• It is a straightforward and reliable test.


• AZ-score can be used for a comparison of raw
scores obtained from different tests.
• While comparing a set of raw scores, the Z-
score considers both the average value and the
variability of those scores.
Disadvantages

• Z-test requires a known standard deviation


which is not always possible.
• It cannot be conducted with a smaller sample
size (less than 30).
analysis process includes the following
four steps
1. DATA PREPARATION
2. DESCRIBING THE DATA
3. DRAWINGTHE INFERENCES OF
DATA
4. INTERPRETATION AND DISCUSSION
OF DATA
1. DATA PREPARATION

Data preparation involves logging or


checking the data in,checking the data for
correctness, entering the data into the
computer, transforming the data, and
documenting as well as developing a database
structure to integrate different measures.
Cont..
• Data preparation involves the following steps:
1. Compilation
2. Editing
3. Coding
4. Classification
5. Tabulation
Cont...
1. COMPILATION: Compilation process
includes gathering together all the collected
data in a manner that a process of analysis can
be initiated.
2. EDITING:-The process of making changes,
deciding what will be removed and what will be
kept in, in order to prepare the accurate data.
[Link]: Coding is important for analysis as
numerous replies can be reduced to a small
number of classes through coding.
Cont..
4. CLASSIFICATION:-In this process, we
divide and arrange the entire data into different
categories, classifications, groups or classes.
[Link]: Tabulation is the recording of
the classified data in accurate mathematical
terms. Rows are horizontal and columns are
vertical arrangements.
2. DESCRIBING THE DATA (DESCRIPTIVE OR
SUMMARY STATISTICS)

• Descriptive statistics is used to describe the


basic features of data and to provide simple
summaries about the sample and the measures
used in a study.
• It is also used to describe the main features of a
collection of data in quantitative terms.
• Percentages, means of central tendency (mean,
median, and mode), and means of dispersion
(standard deviation, range, and mean deviation)
are the examples of descriptive statistics.
[Link] INFERENCES OF DATA

• The most commonly used inferential statistical tests


are Z-test, t-test, ANOVA etc.
• Inferential statistics helps the researcher to determine
if the difference found between two or more groups,
such as an experimental and control group.
• The decision three for the choice of inferential
statistical tests.
1. Type-I and type-II errors
2. Level of significance
3. Confidence interval
1. Type-I and type-II errors

• Type-I error occurs where null hypothesis is


rejected, when it should have been accepted.
• It is also called alpha error.
• Type-II error occurs where null hypothesis is
accepted, when it should actually have been
rejected.
• It is also known as beta error.
Cont...
2. Level of significance

• Probability of making type-I error is called


level of significance.
• It is represented by alpha or p .
• In health sciences, we generally consider the
level of significance at either 1% (0.01) or 5%
(0.05).
[Link] interval

• Confidence interval (CI) is a range of values


that with a specified degree of probability is
thought to contain the population value.
• The researcher asserts with some degree of
confidence that the population parameter lies
within those boundaries CI contains a lower
and an upper limit
4. INTERPRETATION AND DISCUSSION
OF DATA

• After the data are analysed using descriptive


and inferential statistics, results are described
in light of the statistical findings and statistical
principles.
• The next step is the interpretation of data,
which is written in a separate section as
'Discussion
summary
Thank you

You might also like