What is data analysis?
Data analysis is a process of summarizing
trends and patterns observed in the data,
determining major differentials, or
relationship among variables used in the
study and the application of appropriate
statistical tests on a set of data to answer
the problem.
The type of data analysis to use depends on:
◦ The problem
◦ The Kind of scales of measurement of the data or
variables being dealt with
Is the process of explaining the meaning of
a data in a table with emphasis in the
highlights and trends shown by the data.
It further explains the meaning of the data
and relates the findings to result of related
studies to the theoretical framework or
conceptual framework.
Nominal Data Ordinal Data
Has no mathematical Measure in which data
value
or categories of a
Also called a categorical
variable or ordered or
scale
ranked into two or
Numbers are assigned to
categories of nominal data more levels or degree,
variables to facilitate data such as from low to
processing. high or least to most.
Examples:
Examples:
Degree of malnutrition, hororoll, level of
Sex, color and civil status
anger
Interval Data Ratio Data
Has a characteristic of
Almost like the
an ordinal scale but in interval scale except
addition, the distances that the zero scale has
between points in the a real zero point
interval scale is equal
but there is no zero
point or it may be
arbitrary.
Examples: Examples:
Body temperature in farenheit, business Monthly income, number of children,
capital hours spent in studying
THE UNIVERSE OF DATA
Description and Inference
Is used to describe the nature and
characteristic of an event or a population
under investigation.
It is used to describe the characteristic of a
variable or a set of data and/or the variance
within the data.
Characteristics When to Use
1. An interval estimate 1. An interval interpretation is
2. A calculated average for all appropriate
items 2. The values of each score is
3. Affected by extreme values desired
4. Sum of deviations about the 3. Further statistical
mean is zero computation is expected
5. Subject to numerous
mathematical computations
6. Most widely used
7. Represents average quantity
Characteristics When to Use
1. An ordinal statistics 1. An ordinal interpretation is
2. A rank or position average needed
3. Not affected by extreme 2. The middle value is
values desired
4. Absolute differences from 3. Avoid influence of extreme
median, the sum of these values
differences is at a
minimum
5. Represents typical score
Characteristics When to Use
1. A nominal statistic 1. A nominal
2. An inspection average interpretation is
3. Usually occurs near needed
the center of the 2. A quick approximation
distribution of central tendency is
4. Some distribution desired
have more than one 3. The most frequently
mode occurring value is
5. Most “popular” value needed
Range Standard Deviation
Simple measure of
Gives the average of
the distances of
variation calculated
individual observation
as the highest value from the group mean,
in the distribution, the square root of the
minus the lowest average squared
value plus one deviation of each case
from the mean
Range = HV-LV
Difference in Proportions Differences in Means
To determine whether To determine whether
the difference in the the mean grade of the
proportion of male male students
smokers and the significantly differ
proportion of female from that of the
smokers is statistically female students, the
significant, the z-test z-test difference
for difference in between means can
proportions can be be applied.
applied.
Z-test ANOVA
used to test the To analyze the
comparison of two
groups difference among
three or more
test for significance of means.
the difference
between two means
(in a large sample)
Method of analysis used in testing
hypothesis.
It is used to test for significance of observed
differences or relationship between or
among the variables.
This method is used in analytical studies.
Fisher Exact Probability Test – used to estimate
the significance probability in 2x2 contingency
table.
Kolmogorov-Smirnov Test – used to test whether
a cumulative frequency distribution differs
significanctly from chance.
Mann-Whitney Test, The Wilcoxon Test or Wald-
Wolfowitz Runs Test – used to test hypothesis for
one variable about whether two independent
groups come from the same population.
Friedman Test – used to test hypothesis about
differences between more than two groups, that is
whether the several groups come from the same
population, but requires that the several groups,
matched.
Kruskal-Walls Test – same as Friedman test is testing
hypothesis but does not require matched groups.
Moses Test of Extreme Reactions – used to test
hypotheses about differences between two groups in
range.
SODA: SODA:
A B
GEND
20 30 50
ER:
(40%) (60%) (50%)
MALE
GEND
ER: 30 20 50 GENDER SODA
FEMA (60%) (40%) (50%)
LE case 1 MALE A
case 2 FEMALE B
100 case 3 FEMALE B
50 50
(100% case 4 FEMALE A
(50%) (50%)
) case 5 MALE B
... ... ...
Measures of Relationship
Measures of Significant Difference
Measures of Relationship
1. Pearson Product Moment Coefficient of
Correlation (Pearson r)
2. Spearman Rank
Measures of Significant Difference
1. Z-test
used to test the comparison of two groups
test for significance of the difference between two
means (in a large sample)
2. t-test
to test the difference between two groups
To test the significance of the difference between two
means (in a small sample)
GENDER WCC
case 1 male 111
case 2 male 110
case 3 male 109
case 4 female 102
case 5 female 104
mean WCC in males = 110
mean WCC in females = 103
In order to perform the t-test for independent samples, one independent
(grouping) variable (e.g., Gender: male/female) and at least one
dependent variable (e.g., a test score) are required. The means of the
dependent variable will be compared between selected groups based on
the specified values (e.g., male and female) of the independent variable.
The following data set can be analyzed with a t-test comparing the
average WCC score in males and females.
In the t-test analysis, comparisons of means and measures of variation
in the two groups can be visualized in box and whisker plots. These
graphs help you to quickly evaluate and "intuitively visualize" the
strength of the relation between the grouping and the dependent
variable.