0% found this document useful (0 votes)
2 views76 pages

Data analysis

The document provides an overview of data analysis methods, focusing on both descriptive and inferential techniques. It covers various graphical and statistical methods for summarizing and interpreting quantitative data, including measures of central tendency and spread. Additionally, it discusses the importance of visual representations like bar charts and histograms in conveying data insights effectively.

Uploaded by

jogtrott2
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views76 pages

Data analysis

The document provides an overview of data analysis methods, focusing on both descriptive and inferential techniques. It covers various graphical and statistical methods for summarizing and interpreting quantitative data, including measures of central tendency and spread. Additionally, it discusses the importance of visual representations like bar charts and histograms in conveying data insights effectively.

Uploaded by

jogtrott2
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

2/9/2021 Data analysis

Data analysis

Site: Moodle Printed by: Asad Alam

Course: Learning at Work - HOME Date: Tuesday, 9 February 2021, 6:07 PM

Book: Data analysis

[Link] 1/76
2/9/2021 Data analysis

Table of contents

Analysing research data: Descriptive methods


Graphical methods
Statistical methods of data description

Analysing research data: inferential methods


Inferential tests -the example of a Z-test
Investigating the difference between two samples -the 't' test
Investigating the difference between 3 or more samples - ' Analysis of Variance'
Investigating the difference between samples of nominal or weak ordinal type data
Using inferential statistics to measure association
Pearson’s product moment correlation

Introduction to tabular analysis

Simple Linear Regression and Multiple Regression

Summary and further help

[Link] 2/76
2/9/2021 Data analysis

Analysing research data: Descriptive methods

Introduction

We have already outlined that sometimes the raw data will be immediately open to analysis. In other cases some processing
will be needed particularly when the researcher wishes to convert his/her data into a quantitative form. For data from
documents, open questions and, in many cases, observational research, it has already been noted that quotes and description
will be appropriate methods of analysis. For quantitative data, graphical, statistical, tabular and modelling methods will be more
appropriate. The following sections concentrate on these quantitative approaches.

This first section concerned with analysing quantitative data concentrates on data description which is sometimes termed data
summary. Both graphical and numerical methods of description are included.

[Link] 3/76
2/9/2021 Data analysis

Graphical methods

One of the most simplest ways of investigating and/or presenting nominal data (i.e. counts in unordered categories) is with the
use of a bar chart. Much of the contextual details collected through survey methods are often nominal in nature (e.g. sex,
occupation, area of residence etc.) and a bar chart/graph is probably one of the most effective ways of describing the counts
within these categories. Figure 2a shows an example of a bar chart giving the percentage of total respondents working in three
Police Divisions. Figure 2b also illustrates this variable but the proportion of males and females interviewed is also given (a
compound bar graph). Many computer are now available which are able to produce such diagrams. Many of these can provide
quite complicated graphs and charts. It should be remembered, however, that the most effective diagrams are almost always
the most simple ones and the main reason for using graphics is to have an immediate impact on the reader. The examples
shown in Figures 2a and 2b immediately indicate to the reader that most of the respondents were drawn from Division 1 and
that most of the respondents were male.

Figure 2a: An example bar chart

[Link] 4/76
2/9/2021 Data analysis

Figure 2b: An example compound bar chart

An alternative to the bar chart is the pie chart (a circle divided into numerous segments). Whilst this is often regarded as a
somewhat 'exciting' graphic, in reality it is very difficult to interpret and inefficient at illustrating nominal information.

In a similar way to bar charts, histograms can be used to investigate weak ordinal data. As outlined in the first section, weak
ordinal data can be regarded as counts in categories or classes, which are ordered. In essence, a histogram is the same as a bar
chart, except that the arrangement of the classes along the x-axis (ie the horizontal axis) must be in order of class value (usually
in ascending order). With bar graphs, it is of no importance how the 'bars' are arranged, although it is sometimes more pleasing
to the eye to produce a gradient based on height (ie frequency). With ordinal data, however, there may be a relationship or an
association between frequency and the class value. The presentation of data in an ordered fashion is therefore important for
exploring such relationships. Figure 3 shows an example of ordinal data illustrated as a histogram. Strictly speaking, the class
intervals should be of a consistent size and the scale should be continuous along the x-axis. Sometimes histograms are referred
to as frequency diagrams because they illustrate the frequency of occurrence within each of the class values graphically. Their
general shape - formed by the height of the individual classes along the horizontal axis - is referred to as the' frequency curve'.
There are many 'classic' frequency curves referred to in the statistical texts. These are illustrated in

Figure 4. Again, the role of such diagrams is to show insight into the research data which cannot always be gained when faced
with the raw counts.

[Link] 5/76
2/9/2021 Data analysis

Figure 3: An example of a histogram

[Link] 6/76
2/9/2021 Data analysis

Figure 4: Classic frequency curve shapes

Interval or ratio data can be converted into ordinal data by constructing class intervals to which each value can be assigned. If,
for example, actual age is collected within a survey, then class: - intervals can be constructed to summarise that information
and formal statements of the kind '50 people were aged between 21 and 30, 56 between 31 and 40 etc’. Obviously, the class
intervals must cover all possible values.

A simple line graph can be used to plot individual ratio or interval values and is most often I used to display time-series
information such as the reported number of crimes each month (see Figure 5), or the annual number of custodial sentences
delivered over the last decade.

[Link] 7/76
2/9/2021 Data analysis

Figure 5

To explore the possible association between two variables then a scatter plot can be used. For example, Figure 6a illustrates a
scatter plot which shows that there is no relationship between the x and the y variable. The plot shows that the points are
scattered fairly randomly. The " x and y variables in this case could, for example, be the number of road traffic accidents and
the number of children in an area. Figure 6b, however, shows a positive relationship between line x an y. As x increases, then
so does y. An example of a positive relationship would be the incidence of crime and the level of inner city deprivation. As inner
city deprivation increases (as measured on a suitable index) then the level of crime increases. In this plot, the points tend to
form a diagonal line rising from the origin to the top right hand corner of the plot. In contrast, Figure 6c illustrates a negative
relationship. As x increases, y decreases and as y increases, x decreases. An obvious example of a negative relationship would
be the percentage of the population classed as 'high status' (i.e. from social class land l and II) and the percentage of the
population who have criminal convictions.

[Link] 8/76
2/9/2021 Data analysis

Figure 6: Types of scatter plot

There are many more pictorial ways of exploring and describing research data but the methods outlined here provide a basic
starting point. Attention now moves on to describing datasets using basic descriptive statistics.

[Link] 9/76
2/9/2021 Data analysis

Statistical methods of data description

Descriptive statistics are a simple way, usually using just one or two measures, to describe large batches of data. The measures
or methods can be split into one of two kinds. First, descriptive statistics can be used to summarise the level, sometimes known
as the central tendency of a dataset. Second, they can be used to describe the spread of a dataset.

Measures of level

When a researcher is confronted with a large dataset, then it is almost impossible to say anything about that data, without first
computing some simple summary or descriptive statistics. A measure of level can be regarded as a single value which best
summarises the dataset as a whole. The obvious example is the arithmetic mean, otherwise known as the average. For
example, rather than list the payments made to all police surgeons in anyone police force, it is far simpler to quote the average
payment made. The arithmetic mean, however, is not the only measure of central tendency or level. The paragraphs below
outline the calculation of two other main measures as well as an ordinary average and discuss their use with different levels of
information (i.e. nominal, ordinal, interval or ratio data).

The arithmetic mean

The arithmetic mean is calculated by computing the total sum of the batch of numbers and then dividing this by the number of
values in the batch. The mean is represented in many textbooks as:

- and the formula for its calculation is given by:-

Here otherwise known as' sigma' denotes the ' sum of' and indicates that the numerator consists of summing all of the x values
or all of the numbers in the batch. The denominator: consists of 'n' which represents the number of observations in the batch.

[Link] 10/76
2/9/2021 Data analysis

Obviously, the calculation of the average depends on all data values being known and can therefore only be calculated
precisely for interval or ratio data. An estimated average can be calculated for ordinal type data represented as counts in
classes. The formula for an estimated mean of ordinal data is: -

where f represents the number of observations in each class (i.e. the class frequency) and x is the midpoint of the class interval.
An example of its calculation is shown in Table 8 below. The middle of the class interval (x) is used as an estimate of all of the
values (t) which fall within that class. Table 8: The calculation of the mean for ordinal data

Age Group Number of Personnel (f) Class mid-point (x) (fx)

18-25 6 21.5 129

26-33 10 29.5 295

34-41 12 37.5 450

42-48 9 45.5 409.5

>48 4 52* 208

[Link] 11/76
2/9/2021 Data analysis

Therefore the estimated mean = 1491.5/41

= 36.4 years

*For open-ended categories, a sensible value is assigned to represent the class value.

The median

The arithmetic mean can be an inappropriate measure for certain types of frequency distribution. If the distribution is skewed
(see Figure 4), then the arithmetic average will be affected by these extreme values and will no longer represent the central
values of the dataset. In such circumstances, a much more appropriate measure is the median. The median is simply the
middle value when the numbers are put in either ascending or descending order. If there is an even number of values in the
dataset then the middle two numbers are I averaged that is, totalled and divided by two.

Figures 7a, 7b and 7c show the arrangement of the mean and median in a positively skewed, normal and negatively skewed
distribution. The diagrams illustrate how the mean is affected by extreme values, whereas the median is not.

[Link] 12/76
2/9/2021 Data analysis

Figure 7: The position of the mean and median in normal and skewed distributions

The mode

The mean and median cannot be calculated for nominal data. Here, central tendency or 'level' is measured using the mode. The
mode is simply the most frequently occurring category .The modal category in the example shown in Figure 8 is Area D. The
mode, of course, can also be used to summarise the level of ordinal data by determining the most frequently occurring class.
Modal values can also be found for interval or ratio data.

[Link] 13/76
2/9/2021 Data analysis

Measures of spread

In addition to summarising the central tendency of a dataset, another measure is needed to: describe how the rest of the data
values are spread around that centre. Are the values similar to the central measure? - indicating a fairly compact distribution, or
is there a wide range in the values? - indicating a fairly spread distribution. Four values are commonly used to summarise
spread. They are percentage in the mode, range, midspread and standard G deviation.

Percentage in the mode

This is the only measure of spread which can be used for nominal data. Its calculation is simple and consists of expressing the
frequency found in the modal class as a percentage of the total number of observations

The range

[Link] 14/76
2/9/2021 Data analysis

For ordinal, interval and ratio data a simple measure of spread is the range. This is calculated by subtracting the smallest value
in the distribution from the largest. The main disadvantage of the range is that, by definition, it includes the extremes of the
distribution and may not necessarily describe the range of the bulk of the data. Any which combats this problem is the
interquartile range, otherwise known as the midspread.

The midspread

The midspread (dq) is the range of the middle half of the data, once the data has been : placed in rank order. This has the effect
of discounting the bottom and top quarter of the data batch, thus ignoring extremes. To calculate the midspread, the difference
between the values 1 found at the upper quartile (Qu) and lower quartile (QL) is determined. The upper and lower quartiles are
found in a similar way to the median. Rather than divide the dataset into two equal parts, quartiles divide it into four. The
example below shows the calculation of the midspread and compares it with the range (see Figure 9).

Figure 9: Calculating the Midspread

Example:

Age of Personnel

at Police Headquarters

18,20,22,24,24,27 ,27 ,29,30,31 ,33,33,33,

35 ,35 ,35,36,38,39,40,41 ,41 ,42,42,43 ,63

To calculate the midspread (dq), the quartiles must first be found. The ages have already been placed in ascending order.
The number of observations = 26. By counting in 7 values from each side of the ordered distribution, the lower and upper
quartiles will be found.

The lower quartile = 27

The upper quartile = 40

[Link] 15/76
2/9/2021 Data analysis

To calculate the midspread:-

dq = Qu -Qr.

[= 40 - 27

= 13

The range of the middle half of the data is therefore 13. This compares with a full range of 45 (i.e. 63-18) which is influenced
by the extreme values in the positively skewed distribution.

The standard deviation.

Both the range and the midspread rely on just two data values to describe the spread. The. standard deviation , in contrast,
uses all of the data values and is a measure which describes how all of the values differ from the arithmetic mean, on average.
In other words, it is away of describing the average spread of all the values around the mean. The main disadvantage of the
standard deviation is that it is affected by extreme values. It is, however, an important statistical measure and is often used in
many other statistical functions. For this reason, an understanding of the measure at this juncture is beneficial for
comprehending other statistical methods. Its calculation is relatively simple and is illustrated in the example below. It is,
however, very unlikely that a researcher would need to calculate this measure manually.

Computing a standard deviation -an example

What is the standard deviation of the following Forensic Medical Examiner's Salaries? £OOOs, 7,30,5, 7,8,10,6,6,110,54

The standard deviation measures the average difference of the observations from the overall average. It is first necessary to
calculate the overall average, where

[Link] 16/76
2/9/2021 Data analysis

= 243/10 = 24.3

It is then necessary to find out the average difference of each observation from 24.3. The observation (x) is shown in the fIrst
column of Table 9 below. The difference from the mean is shown in the second column indicated as :- .

Table 9: Calculating a standard deviation – example continued

To enable the calculation of the average difference, the To calculate the overall average difference, the
difference must first be squared – otherwise the average individual differences must be summed:
(x) (x.24.3) difference will equal zero
= 10370.1

and then divided by n

7 -17.3 299.29 = 10370.1/10

= 1037.01
30 5.7 32.49

5 -19.3 372.49 Because the differences were squared, the


square root must be taken to give the overall
standard deviation = 32.2.
7 -17.3 299.29

8 -16.3 265.69 The final formula for the calculation of a


standard deviation is therefore given by: -

[Link] 17/76
2/9/2021 Data analysis

10 -14.3 204.49

6 -18.3 334.89
This is a rather tedious way of calculating

6 -18.3 334.89

110 85.7 7344.49

54 29.7 882.09

Descriptive methods -summary and conclusion

This section has outlined several useful ways of describing quantitative data. 'Pictures of data, are useful for 'exploring' datasets
and presenting results in research reports. Statistical summary measures have been introduced which describe the two main
features of any dataset - the level and the spread. The section has illustrated that the choice of statistical measure depends
upon the nature of the data.

[Link] 18/76
2/9/2021 Data analysis

Analysing research data: inferential methods

Inferential techniques are a branch of statistical methods which use the characteristics of a sample to make inferences about a
population. Many research findings are based on data gathered from a sample of observations rather than on the population
as a whole. How can the researcher be sure that the findings from the sample mirror those of the general population? In fact,
she or he can never be 100 % sure but inferential statistics allow statements to be made which are concerned with the
likelihood of the sample findings being similar or dissimilar to the population. Inferential statistics can also be used to discuss
the likelihood that the differences found between samples are actually real and have not occurred simply by chance. In all
cases, the basic concept is the same and relies on the use of probability theory to decide whether the findings are likely to have
occurred due to chance or whether they are actually significant results.

A note on hypothesis testing

There is a specific convention which must be followed when using inferential statistics to test whether results are significant or
not. Before carrying out the statistical test, both the null hypothesis and the alternative hypothesis must be set out. The null
case (denoted by Ho) always assumes 'no difference' and any observed peculiarities in research findings are there because of
chance. The alternative hypothesis (denoted by H1 assumes that the observed findings are significant because the probability of
them occurring due to chance is so small. In inferential tests, the null hypothesis is always investigated and, based on the
results of the test, either accepted or rejected. To understand the fundamental concepts behind all inferential tests, it is first
necessary to discuss the characteristics of something called the' normal curve , and its relation to probability theory .

The characteristics of the normal curve

A previous section has shown how histograms can be used to show the frequency distribution of anyone dataset. Some of the
possible distribution shapes which might result (e.g. positively skewed, negatively skewed and normal distributions) have also
been discussed in the section describing measures of level and spread. The inferential statistics described in this unit use the
characteristics provided by the normal curve to form the basis of their testing procedure. Many natural phenomena exhibit a
belI-shaped or normal distribution. If, for example, the heights of adults are plotted on a frequency curve, there will be a small
number of observations in the tails of the curve. Most observations will be found around the average (see Figure 10)

[Link] 19/76
2/9/2021 Data analysis

Figure 10: A frequency curve of adult heights

In a perfectly normal curve, the shape is symmetrical and there is an equal number of observations either side of the average,
which equals the median, which equals the mode. It therefore follows that 50% of the observations fall either side of the
average. An alternative, way of expressing this is that the chance of obtaining an observation higher than the average: is 50%
and one which is lower than the average, also 50%. This chance of 50% can be translated into a 0.5 probability (NB probability is
always expressed as a decimal, unless it is 100% chance which is equal to a probability of I). In terms of the area under the
normal curve, 50% of the area represents observations above the mean and 50% represents observations above the mean and
40% represents the observations found under the normal curve can also be described in terms of standard deviation. As noted
earlier, the standard deviation indicates the average difference of all values around the mean. In a perfectly normal curve, 68%
of the values are found within 1 standard deviation either side of the mean. 95% of the values will be within almost 2 (1.96 to be

[Link] 20/76
2/9/2021 Data analysis

exact) standard deviations and 99% of the values will fall within 2.5 standard deviations. These statements may be a little
difficult to visualise but if the height example is used to illustrate them then it may be easier to understand the ideas being
used. If it is assumed, for example, that average height is 66 inches and that the standard deviation is 4 inches, then 68% of the
population will be between 62 "and 70" (i.e. 1 standard deviation above and below the mean). Likewise, approximately 95% of
the people will be between 58 "and 74" and 99% between 54 "and 78" (NB these are not true height figures - they are only used
as an example). This is shown diagrammatically in Figures 11a, 11b and 11c below.

Figure 11: The percentage of observations under the normal curve

Z-tables and the standard normal curve

[Link] 21/76
2/9/2021 Data analysis

Statistical tables are available which provide similar information to that listed in the previous paragraph. The area which is
found between number of values and the mean has been calculated for something which is known as the 'Standard Normal
Curve', which has a mean of 0 and a standard deviation of 1. This has been done using a branch of mathematics called calculus
and its description and understanding is unimportant here. It has already been shown that the area under the curve can be
equated with probability or percentage chance. Therefore

the likelihood of persons falling between any two heights (or the likelihood of observing any; variable for that matter!) can be
derived from these calculus-derived statistical tables if the height values are first converted into standard deviations away from
the mean. This process is known as converting the values to Z-scores or standard scores. The formula for doing this, is as
follows:-

where

x = observation

and

and

= standard deviation

The relevant table in Murdoch and Barnes (1986, p.13) gives the area between the Z-score and the tail of the curve. To find the
area, for example, between 2 Z-scores and the tail, then the line of the table with "2” if at the extreme left-hand side is
consulted. Along this line and under the .00 column, a figure of .02275 is given. This means that there is a probability of 0.02275
that a value will be greater than 2 standard deviations away from the mean. Incidentally I since l it is known that there is a
probability of 0.5 that a value will be greater than the average, then the probability of a value falling between the mean and 2
standard deviations = (0.50000- I 0.02275) = 0.47725. As with the height example, discussed previously, it was noted that
approximately 95% of the values are found within 2 standard deviations either side of the average. This is derived by summing

[Link] 22/76
2/9/2021 Data analysis

.4772 + .4772 = .9544 (i.e. 95 %). Use of the tables is further illustrated in the example below.

Z scores: a further example

If the height example is again considered and it is known that the standard deviation of height is 4” and the average height is 66
", the following question can now be answered with the use of the Z-tables:

What proportion of the people are likely to be below 74". In other words, what area of the normal curve is found within the
shaded area of Figure 12 below:-

Figure 12: a further example of the use of z scores

standard deviation = 4” mean = 66”

First, calculate the Z-score:

Z = (74-66)/4

Z = 8/4

Z=2

[Link] 23/76
2/9/2021 Data analysis

In other words, 74 is 2 standard deviations away from tile mean. From the statistical tables, it has already been shown that
47.72% of tile observations fall between the mean and 2 standard deviations. Another 50% of the observations are below the
mean and so a total of 97.72% are found below 74".

Probability, Z-scores and inferential statistics - a summary

It has been shown in tile previous paragraphs that converting observations to Z scores and re- expressing them as statements
of probability (derived from tables of standard scores) has allowed estimation of the frequency of certain phenomena
occurring. So how does all of this relate to inferential data analysis and statistical testing? In essence, these concepts are extend
to make decisions regarding tile likelihood of a particular event occurring. The normal curve and Z-tables indicate that tile
likelihood of a value falling in the tails of the curve (a very high or a very low observations -an 'extreme event') is quite small.
Inferential tests, for example: which investigate whether observed differences between samples are significant, use these basic
concepts to carry out the test. If the observed difference is found in the 'tail end' of all the possible differences which could have
occurred, then the difference is unlikely to have happened by chance and tile results are said to be' significant'. This should
become more clear when the following example is considered.

[Link] 24/76
2/9/2021 Data analysis

Inferential tests -the example of a Z-test

100 police officers have undergone a community liaison course designed to increase awareness and empathy of social
problems found in certain communities. Two years previous, a national research project surveyed police officers regarding this
subject. The survey included information used to compile an empathy score of community problems. At this time, it was found
that the national average empathy score of police officers was 100 with a standard deviation of 14. The scores were distributed
normally. The present-day course employed the same survey and showed that the 100 officers attending the course had an
average score of 107. The question which needs investigating here is whether the score of 107 is significant IX larger than the
national average score of 100? The null hypothesis would therefore state that there was no significant difference between the
sample of 100 officers and the general population. The alternative hypothesis assumes that the difference is real and that it was
influenced by the existence of the community liaison course. Could this score easily have been derived by chance or has the
community liaison course had an effect on understanding and empathy of community issues. Any number of samples of 100
officers could have been taken and it would be relatively rare for the sample score to equal the population score of 100. From
the description of the normal curve, it is known that a value greater than 100 would be expected 50% of the time. How often
would a sample value of 107 be expected? Statistical research has shown that if an infinite number of samples of 100 are taken
from such a population then the sample means will be normally distributed. Furthermore, the standard deviation (otherwise
known as the standard error) of these means will be:-

In this example the standard error =

14/10 = 1.4

The standard error can be regarded as the standard deviation of all the possible sample means of size 100 which could have
been taken from the population of all police officers. So how many standard deviations (standard errors) is the value of 107
away from 100. Here the Z- score formula can be employed:-

[Link] 25/76
2/9/2021 Data analysis

= (107- 100) / 1.4

= 7/1.4

=5

In other words, the sample value of 107 is 5 standard deviations away from the population value of 100. If the Z-tables are
examined, then the value of 5 is not even reported in the tables because it is such an extreme value. It is therefore very unlikely
that 107 was gathered; from the general population by pure chance. It is more likely that this sample value has been affected by
the existence of the community liaison course. The null hypothesis of 'no

significant difference' would therefore be rejected.

A note regarding critical values and p values.

There obviously has to be a point when the researcher has to make a decision as to whether a particular value is likely to have
happened by chance or whether the value is significant. In terms of Z-scores this relates to how far into the tails of the normal
curve the ‘critical' value of Z is derived. Much published social research uses cut-off values of 95% or 99% .In other words if a
sample value could have been obtained by chance 95% or 99% of the time, then it is not regarded as significant. Obviously the
use of a 99% critical level is more strict than a 95% level. In the literature, such levels may be indicated by a p value. 95% and
99% levels equate with p=0.05 and p=0.01, respectively. Erickson and Nosanchuk (1979) provide an excellent account of critical
values and of levels of confidence in statistical testing. They also relate these concepts to the dangers of falsely accepting or
rejecting a null hypothesis and compare it to the courtroom situation of falsely finding the accused innocent or guilty.

This Z-test example has outlined the general steps which are employed in all inferential tests. The null and alternative
hypotheses are stated. A test statistic is calculated and compared with a critical value from statistical tables. If the test statistic is
greater than the table value (at a specified value of p), then the null hypothesis is rejected. If it is smaller, then the null
hypothesis is accepted. The use of a Z-test (employing a critical values of Z to help determine whether a particular result is likely

[Link] 26/76
2/9/2021 Data analysis

to have occurred by chance or not) is just one of many inferential test available , to the social researcher. There are different
statistical tests for the different levels of data (ie nominal, ordinal, interval or ratio data) and different tests for different
research scenarios (eg testing the significance of a result from one sample, testing the difference between two samples . and
testing the difference between many samples). The following sections cover some of the major statistical tests used in the
research literature and those which are likely to be most useful for project/dissertation work. It should be emphasised here,
however, that there are many others and advice should be sought before employing statistical tests in actual research.

[Link] 27/76
2/9/2021 Data analysis

Investigating the difference between two samples -the 't' test

Z-tests are used to compare a sample mean with a population mean. If the problem is one of (determining whether there is a
difference between two samples, then a 't' test should be employed. Incidentally, 't' tests should, strictly speaking, also be used
in place of Z-tests if the sample size is small, say less than 30.

In a two-sample 't' test, the difference between the means of two samples is investigated. If the difference is 'significant' then it
is assumed that the samples have been drawn from different populations. If not, then the samples are assumed to be drawn
from the same population. To carry out a two-sample ‘t’ test, a calculated 't' (based on the sample values) is compared with a
critical value found in 't' tables (NE This process is similar to comparing a calculated 'Z' with a table 'Z' in the Z-test discussed in
the previous section). If the calculated value is greater than the table value, then the null hypothesis is rejected - the null
hypothesis; being 'no significant difference'. Alternatively, if the calculated value is smaller than the table value, the null
hypothesis must be accepted.

The formula for calculating 't' is shown below:-

where x and y refer to the values of the two different samples. The Ix-yI term simply refers to the absolute value of x-y (i.e. the
sign is ignored). Again, most statistical computer packages contain the function to carry out a t-test and the calculation process
is therefore relatively unimportant. In essence, it is similar to a Z statistic. The formula calculates the difference between the
two sample means as a proportion of the standard deviation of all the possible differences (termed the standard error). A high
value indicates that the observation is found in the tail end of the distribution and is unlikely to have happened by chance. In
order to determine whether this is large enough to reject the null hypothesis, the table of critical 't' values is observed (see
Murdoch and Barnes, 1986, p. 16). Before working through an example of a t-test, some additional statistical-terms need
explanation. First, when reading the table of critical values, the degrees of freedom must be known. The degrees of freedom

[Link] 28/76
2/9/2021 Data analysis

take account of the level of uncertainty in a statistical formula by noting the size of the sample. In essence, as the sample size
gets larger, the critical values get smaller and the null hypothesis can be rejected more easily. For the t-test, the degrees of
freedom are calculated using the following formula:-

ie the number in the first sample plus the number in the second sample minus 2 (NB the degrees of freedom are often denoted
by "v" in statistical tables). Second, the test can either be or two-tailed. One or two tailed tests are relatively simple to explain: if
the alternative hypothesis states that one particular sample mean is thought to be greater or less than another, then the test is
one-tailed. If the test is simply investigating 'difference' as stated in the alternative hypothesis, then a two-tailed situation exists.
The option to choose a critical value for a one - or two-tailed test is available in the statistical tables.

The two-sample t-test: an example

Empathy scores of male and female Police Surgeons to Victims of Rape.

Female Male

80 68

79 71

78 58

[Link] 29/76
2/9/2021 Data analysis

69 62

68 52

78 67

75 63

74 70

73 59

81 61

The problem here is to determine whether female Police Surgeons have a significantly higher empathy score than their male
counterparts. From the scores above, an average empathy score of 75.5 is obtained for the females. This compares with 63.1
for the males. The t-test is used to determine whether this difference could easily have occurred by chance or not. The stages to
carrying out the statistical test are as follows:

(i) State the Null and Alternative Hypotheses:-

Ho: There is no significant gender-based difference in empathy scores.

H1: Empathy scores of female Police Surgeons are significantly greater than their male counterparts.

[Link] 30/76
2/9/2021 Data analysis

Because the alternative hypothesis has stated that female Police Surgeons have a significantly" higher empathy scores, rather
than simply significantly different, then the test is one-tailed and not two-tailed.

A confidence level of 95% will be employed in the test (i.e. p = 0.05).

(ii) Calculate t:-

Using a pocket calculator, the standard deviation and mean of each sample can be determined. The relevant values for each
sample needed to calculate the formula shown above are listed below:-

Substituting these values into the formula:-

[Link] 31/76
2/9/2021 Data analysis

(iii) Find the critical value of t from the statistical tables with 18 degrees of freedom, p=O.5 for a one-tailed test (Murdoch and
Barnes, 1986, p. 16):-

Table t = 1.73

(iv) The calculated value of 4.98 is greater than the table value. The null hypothesis is therefore rejected. It can be concluded
that female Police Surgeons have a significantly higher empathy score than their male counterparts.

[Link] 32/76
2/9/2021 Data analysis

Investigating the difference between 3 or more samples - ' Analysis of


Variance'

The analysis of variance (ANOVA) test - sometimes known as an 'F-test'- is similar to t-test in that it investigates the difference
between sample means for ratio or interval data. An F-test is used when there are three or more samples, whilst a t-test is used
for investigating the difference between two samples. Basically, an analysis of variance test compares the variance within each
sample (NB the variance is the standard deviation squared) with the variance

between them. If the samples have been drawn from the same population, then the 'within' variance will be approximately the
same as the 'between' variance - because both will reflect the overall population variance. If the samples are significantly
different and have been drawn from different populations, then the variance between the samples will be significantly larger
than the variance within.

The steps to ANOVA are similar to the statistical tests which have previously been described in this unit. A statistic which
provides the ratio of the between variance to the within variance is calculated using the sample values (an F value). This
calculated F is then compared with a critical value of F provided in statistical tables. If the calculated value is greater than the
table value then the null hypothesis is rejected. If it is smaller, then the null hypothesis is accepted. F-tests, however, can only be
carried out on parametric (normally distributed) data which is either interval or ratio in nature.

The formula for calculating F is shown below:-

F= between sample variance

within sample variance

Where the between sample variance is calculated using:-

[Link] 33/76
2/9/2021 Data analysis

where = between sample variance and k = number of samples, n is the number of individuals in each sample. The
expression:-

indicates that, for each sample, the difference between the mean of the sample and the overall mean (grand mean) of all the
sample values should be derived, squared and multiplied by the number of individuals in the sample. The summation sign
indicates that these values (for each sample) are added together.

The within sample variance is given by:-

where a2w = the within sample variance and N is the total number of individuals (in all samples). In this formula, the difference
between the individual observations and the sample mean is found, squared and summed. These sums of squared deviations
for each sample are then added together and divided by the total number of observations minus the number of samples.

The F-test formulae look very complicated but are really quite easy to work through. Again, statistical computer packages are
able to provide ANOVA tests and it is unlikely that the value would have to be calculated manually. The following example,
however, illustrates how the formulae are used in the test.

Table 10: Example of an F test (Ana1ysis of Variance (ANOVA) test)

[Link] 34/76
2/9/2021 Data analysis

Salaries for sample of FMEs in four police areas (£000s)

Area A Area B Area C Area D

Salary Salary Salary Salary

(x) (x-mean)2 (x) (x-mean)2 (x) (x-mean)2 (x) (x-mean)2

112 169 100 9 120 372.5 78 100

100 1 83 196 93 59.3 78 100

113 196 103 36 86 216.1 91 9

106 49 102 25 103 5.3 96 64

76 529 97 0 90 114.5 93 25

87 144 112 127.7 92 16

∑x = 604
∑x = 594 ∑x = 485 ∑x = 528
n=6
n=6 n=5 n=6
mean x = 100.7
mean x = 99 mean x = 97 mean x = 88
∑(x-mean)2 = 895.4
∑(x-mean)2 = 1088 ∑(x-mean)2 = 266 ∑(x-mean)2 = 314

[Link] 35/76
2/9/2021 Data analysis

Number of samples (k) = 4

Total number of individuals (N) = 23

Grand mean = (594 + 485 +604 +528)/23 = 96.1

(i) State the null and alternative hypotheses:-

Ho: There is no difference between FME salaries in Areas A, B, C, D

H1: There is a significant difference between FME salaries in Areas A, B, C, D

(ii) Calculate F

Between sample variance:-

For Area A,

and therefore:-

[Link] 36/76
2/9/2021 Data analysis

Similarly for Area B,

Similarly for Area C

and for Area D

Therefore the between sample variance =

Within sample variance


[Link] 37/76
2/9/2021 Data analysis

Calculated F:-

F = between sample variance = 191.71 = 1.42

Within sample variance 134.92

Degrees of freedom for the F test

The degrees of freedom for the between sample variance are the number of samples minus one (k-1). In this example this is
calculated as 3. The degrees of freedom for the within sample variance are the total number of individuals in the data minus
the number of samples (N-k). Again, in this example it is calculated as 19. The table value of F can be derived using Murdoch
and Barnes (1986, pp. 18-19). Within this version of critical values, the degrees of freedom for the between sample variance are
found along the horizontal axis; and the degrees of freedom for the within sample variance are presented vertically. Not all
possible values for the degrees of freedom are given. The nearest values should, obviously be used. The table J value off is
therefore given as 3.16 (p=O.O5). Since the calculated value is smaller than the f critical value, the null hypothesis is accepted
and it can be concluded that there is no significant difference in FME salaries between Areas A, B, C and D.

[Link] 38/76
2/9/2021 Data analysis

Investigating the difference between samples of nominal or weak


ordinal type data

The t-test and F-test are used to investigate the difference between samples, based on sample means. The sample means are
derived from ratio or interval type data. Much survey data, however, is not in this form and much of the resultant data is
comprised of counts in categories (ordered or unordered). A sample mean cannot be derived for nominal data and can only be
estimated for weak ordinal data. To investigate the difference between samples containing categorical data, a Chi2 test can be
used.

In essence, a Chi2 test determines whether the observed values are significantly different from those expected. The expected
values are determined by assuming that the phenomenon in question is distributed equally. As in all previous tests, a calculated
value of Chi2 is compared with a critical value derived from statistical tables and the Null hypothesis either accepted or rejected.

The formula for calculating Chi2 (often written as X2) is as follows:-

where,

d=( observed – expected )

and e = expected value.

The expected value is calculated in two different ways, depending on whether the test is for one sample or for two or more
samples

Calculating the expected value for a one-sample case

[Link] 39/76
2/9/2021 Data analysis

To calculate the expected value for a one-sample case, the total number of observations is divided by the number of categories.
The degrees of freedom in this type of Chi2 test are the number of categories minus one.

Calculating the expected values in two or more samples

With two or more samples, the expected values are found by:

Row Total x Column Total

Grand Total

and the degrees of freedom are found by subtracting one from the number of rows and multiplying this by the result of
subtracting one from the number of columns.

Chi2 - restrictions associated with its use

As previously stated, Chi2 can only be used with nominal or weak ordinal data (ie counts in categories). In particular, the counts
should be raw numbers and not rates or percentages. The data, however, does not have to be normal - Chi2 is a non-
parametric test. If the number of categories is 2, then the expected frequencies should be 5 or larger. If the number of
categories is greater than 2, no more than 1/5th of the expected frequencies should be less than 5. None should be less than 1.
If these conditions are not met, then the Chi2 test is unreliable. The two examples shown below present a one sample case and
a three sample case.

Chi2 example 1 - one sample case

Is there a significant difference between job satisfaction among a sample of 45-64 year olds?

Table 11: Job Satisfaction

[Link] 40/76
2/9/2021 Data analysis

JOB SATISFACTION Persons aged 45-64

Satisfied 35

Neither satisfied/dissatisfied 50

Not satisfied 40

(i)State the null and alternative hypotheses:

Ho: there is no significant difference in attitudes to capital punishment between I those who are employed and those who
are unemployed
H1: here is a significant difference in attitudes to capital punishment between those who are employed and those who are
unemployed

(ii) Calculate Chi2

For a one sample case, the expected value is found by dividing the total number of observations by the number of categories:-

E= (35+50+40)/3

[Link] 41/76
2/9/2021 Data analysis

E= 125/3

E= 41.7

Substituting the expected values into the formula:

(iii) Compare the calculated value (2.9) with the critical value of Chi2 derived from the tables (Murdoch and Barnes, 1986, p.17)
with 2 degrees of freedom at the 95% confidence level (i.e. p=0.05):-

Critical value of Chi2 = 5.99

Since the calculated value is smaller than the table value, the null hypothesis must therefore be accepted and it can be
concluded that there is no significant difference in job satisfaction between 45-64 year-olds.

Chi2 example 2- two sample case

Is there a significant difference between attitudes to capital punishment by employment status?

Table 12: Attitudes to Capital Punishment

[Link] 42/76
2/9/2021 Data analysis

Attitudes to Capital Punishment Employed Unemployed

Agree 60 9

Neither agree/disagree 30 11

Disagree 60 60

(i) State the null and alternative hypotheses:

Ho: there is no significant difference in attitudes to capital punishment between those who are employed and those who
are unemployed

Hi: there is a significant difference in attitudes to capital punishment between those who are employed and those who are
unemployed

(ii) Calculate Chi2:

For a two or more sample case, the calculation of the expected values is found by multiplying the row total by the column total
and dividing the result by the grand total. In this way, an expected value is calculated for each cell of the table. The row, column
and grand totals have been calculated for the given example and are shown in the table below:-

Table 13: Row, Column and Grand totals

[Link] 43/76
2/9/2021 Data analysis

Row Totals |
Attitudes to Capital Punishment Employed Unemployed
V

Agree 60 9 69

Neither agree/disagree 30 11 41

Disagree 60 60 120

Row Totals → 150 80 GRAND TOTAL 230

The calculation of the expected value for the top left-hand cell of the table (ie the employed persons who agree) is as follows:

e= Row Total x Column Total

Grand Total

= 69 x 150

230

= 45

[Link] 44/76
2/9/2021 Data analysis

This calculation, for the two or more sample case, takes into account the different number of persons interviewed or surveyed
in each of the groups. In the above example, it makes little sense to say that more of the employed group agree with capital
punishment simply because 60 is greater than 9. Far more employed people were actually included in the survey (150,
compared with 80) and so some of the difference can be accounted for by this fact. The expected calculation takes this into
account and computes the expected frequency based on the differing proportions in each employment group. Again, referring
to the top left-hand cell of [the table, slightly more of the employed group actually agree with capital punishment than was
expected (ie 60 compared with 45). The expected values for all of the cells are shown in Table 14 below:-

Table 14: The expected frequencies

Attitudes to Capital Punishment Employed Unemployed

Agree 45 24

Neither agree/disagree 26.7 14.3

Disagree 78.2 41.7

Therefore,

[Link] 45/76
2/9/2021 Data analysis

(iii) The critical value of Chi2, derived from the tables, with 2 degrees of freedom (p = 0. 05) is 5.99. Since the calculated value
(27.8) is greater than the table value, then the null hypothesis can be rejected and it can be concluded that there is a significant
difference in attitudes to capital punishment between those who are unemployed and those who are employed.

[Link] 46/76
2/9/2021 Data analysis

Using inferential statistics to measure association

Attention now focuses on the use of [Link] to investigate the association between two variables rather than
difference. The section which covered 'pictures of data' illustrated how two variables could either have a positive, a negative or
no association between each other (see Figure 6). This concept is extended here to measure the strength of that association.
Although the association between the two variables shown in both Figure 13a and 13b is positive (ie as x increases, so does y),
the strength of the association in Figure 13a is much stronger than the association in Figure 13b.

Figure 13

This section discusses the use of correlation coefficients to indicate the direction and strength of an association. A coefficient of
+1 indicates a perfect positive relationship, 0 indicates no relationship and –1 indicates a perfect negative relationship (see
Figures 14a, 14b and 14c). Two methods are described. First, Person’s Product Moment Correlation is used for investigating the
[Link] 47/76
2/9/2021 Data analysis

association between interval or ratio data which is parametric in nature (ie a histogram of the data is approximately normal in
shape). Second, Spearman’s rank correlation is used to measure association in non-parametric data or in strong ordinal data (ie
data that is ranked).

Figure 14: A strong positive relationship, no relationship and a strong negative relationship

[Link] 48/76
2/9/2021 Data analysis

Pearson’s product moment correlation

As already stated, Pearson’s correlation can be used for interval or ratio data which is parametric in nature. Further more, the
plot of the relationship between the two variables must approximate to a straight line relationship. There must be no curves.

As with all other inferential tests, a null hypothesis and an alternative hypothesis is put forward and a correlation coefficient is
calculated. Its value is then compared with a critical value provided in statistical tables and then, on the basis of the
comparison, the null hypothesis is either accepted or rejected. The formula for calculating Pearson' s correlation coefficient (r)
is as follows:-

where x refers to the first set of observations and y, the second. The number of pairs of observations is represented by n and
refers to the standard deviation of the first set of observations and refers to the standard deviation of the second.

Consider the following example:-

Table 15 shows the number of hours spent attending lectures or talks concerned with community liaison issues for ten
members of police staff. For each person, a prejudice score relating to several minority groups is also given. Is there a
relationship between community liaison education and tolerance of social minority groups?

Table 15: Pearsons correlation coefficient: an example

Hours of lectures/talks concerned with Community Liaison Prejudice Score (Low score indicates low level of prejudice)

10 1

[Link] 49/76
2/9/2021 Data analysis

3 7

12 2

11 3

6 5

8 4

14 1

9 2

10 3

2 10

(i) State the null and alternative hypotheses:-

Ho: there is no significant association between community-based education and prejudice


H1: there is a significant association between community-based education and prejudice

The alternative hypothesis could be extended to say that the association is of a negative nature (ie as the educational
experience increases, the level of prejudice decreases). In this respect, the test can be regarded as one-tailed. A simple plot of
the data illustrates this negative nature.

[Link] 50/76
2/9/2021 Data analysis

(ii) Calculate the correlation coefficient:-

The values of xy, ∑xy, σx and σy are shown in Table 16, along with the calculation of r. The correlation coefficient of -0.93
indicates a fairly strong negative association between these two variables. Next, however, the significance of this observed
association should be tested using, statistical tables to investigate whether such an association could easily have happened by
chance for this number of observations or whether the value is so large that it is probably significant.

(iii) Compare the calculated value (-0.93) with the critical values of r (Murdoch and Barnes, 1986, p. 20), using n-2 degrees of
freedom (i.e. 8) and p=0.05. The table value is given as 0.549. Since the calculated value is greater than the critical value shown
in the tables, the null hypothesis of no association can be rejected (NB the absolute value of calculated r is used. The sign simply
indicates the direction of the association). In other words, the observed association is unlikely to have happened by chance.

Table 16: the calculation of Pearsons correlation coefficient

Hours (x) Score (y) (xy)

10 1 10

3 7 21

12 2 24

11 3 33

6 5 30

[Link] 51/76
2/9/2021 Data analysis

8 4 32

14 1 14

9 2 18

10 3 30

2 10 20

mean x = 8.5 mean y = 3.8 ∑

[Link] 52/76
2/9/2021 Data analysis

Given, σx = 3.6

σy = 2.7

n = 10

Therefore,

r = 232/10 – (8.5 x 3.8)

3.6 x 2.7

r = -0.93

regarded as a modelling technique. By calculating the average (say) income of a group of individuals, the data analyst is 'fitting'
a model to the data. The average is being used to summarise what is happening, in reality, to the whole group and any data
value within the set can now be redefined as comprising two parts:

Data = Fit + Residual

(NB this is often termed the DFR equation)

[Link] 53/76
2/9/2021 Data analysis

Although an average income has been calculated from the batch, very few of the actual numbers (if any) will mirror this
average. Instead they will be made up of the average +, or -, the actual value (i.e. fit + residual). An example of a slightly more
complex model would be a set of explanations for the existence of juvenile crime:

Juvenile crime = youth unemployment + educational achievement

Here, it is assumed that the juvenile crime rate can be explained by youth unemployment and the level of educational
achievement. Obviously, this is still a relatively simple representation of reality; there are many other known and unknown
explanations for juvenile crime. It is suspected, however, that these two are the most important. The researcher would then
collect data to investigate this model and using various techniques determine how much variance in the crime rate is explained
by these two factors (i.e. investigate the DFR equation).

In summary, therefore, a model is an idealised representation of reality. Often data analysis techniques are concerned with
fitting a model to a set of data and studying it in terms of the of the DFR equation. Furthermore, if it is found that the model fits
well, then it can be used for prediction purposes. If it does not, then the model must be re-calibrated or redefined.

There are many quantitative-based techniques which are concerned with 'data modelling'. A small selection of these are
outlined here. The section begins with a very simple form of modelling and includes the straightforward description of tabular
data. This introductory part could have been inserted earlier in the Unit under the heading of 'descriptive' methods. It is placed
here, however, because it is used in the process of modelling tabular data. Likewise, the paragraphs which deal with the
analysis of tables are related to the concepts described earlier under the heading of Chi2. Such techniques are often used in
conjunction with data modelling and do not necessarily imply a different type of analysis or approach. In other words, if a
researcher has data in the form of tables, then all approaches to the analysis of tables should be considered.

[Link] 54/76
2/9/2021 Data analysis

Introduction to tabular analysis

This section begins by returning to relatively simple methods of summarising and describing the contents of tables. These form
the foundation of later elements which introduce basic data modelling. Tabular analysis is often useful in social research. It
lends itself quite readily to the results derived from survey research which are often reduced to table format. Again, many of
the research methods texts deal with the analysis of tables. Not all, however, extend the analysis to concepts of modelling. One
which does - and is also quite readable - is that written by Gilbert (1993). This text, however, does go beyond the material given
here. Another useful introduction to table description techniques is provided by Marsh (1988). She also covers some of the
material relating to data modelling. Although Hagan (1993) does not cover modelling, simple description techniques relating to
tables are provided which use examples from criminology and criminal justice.

Describing and reading tabular data

Consider the information provided in Table 18 which indicates the class background (based on father's class) of persons
working in the Magistrate's Service. The information covers all types of workers (e.g. clerks, secretaries, JP's etc) which have
been classed as either 'high', 'middle' or 'low' job-status, depending on their employment activity within the Service. The raw
information has been collected via questionnaire. The results have been coded, analysed and with the aid of a computer, the
table has been produced.

Table 18: Job status by class background: frequencies

Father’s class
Job status in Magistrate Service
Service Class Intermediate Class Working Class Total

Low 42 429 882 1353

Medium 123 101 146 370

High 90 93 70 253

[Link] 55/76
2/9/2021 Data analysis

Total 255 623 1098 1976

As it stands, the raw information within this table is difficult to interpret. It is difficult to disclose any underlying patterns which
might suggest data models. This type of table is termed a contingency table because it " shows the distribution of each variable
conditional upon each category of the other" (Marsh, 1993, p.130). The most effective way to read such tables is to re-express
the raw data as percentages. There are three alternative ways of producing percentages for the data:

Percentages based on the total number of observations


Percentages based on the row totals
Percentages based on the column totals

Each of these percentage tables provides a different type of understanding of the data in the table. Tables 19, 20 and 21 show
the different types of percentage analysis2. Table 19, which shows the information expressed as percentages of the grand total,
simply illustrates that the 42 individuals from the service class who are ranked as low job status comprise 2.1 of the total
(number surveyed etc. This type of analysis is not very useful and says little about the relationship between class background
and job status in the Magistrate Service. This type of percentage breakdown is therefore not used very often.

Father’s class
Job status in Magistrate Service
Service Class Intermediate Class Working Class Total

Low 2.1 21.7 44.6 68.4

Medium 6.2 5.1 7.4 18.7

[Link] 56/76
2/9/2021 Data analysis

High 4.6 4.7 3.5 12.8

Total 12.9 31.5 55.5 100.0 (N=1976)

NB Total may not equal 100 because of decimal rounding.

Table 20 however, allows the researcher to make statements about the percentage of people from each social class whose job
is (say) categorised as low status. For example, 86% of the intermediate class are categorised as low status, compared with 16%
of the service class. Similarly, 35% of the service class are carrying out high status jobs, in contrast to 6% of the working class.
This transformation to percentages based on the column total is now revealing some of the more interesting patterns in the
data.

Table 20: Job status by class background: column percentages

Father’s class
Job status in Magistrate Service
Service Class Intermediate Class Working Class

Low 16.5 68.9 80.3

Medium 48.2 16.2 13.3

High 35.2 14.9 6.4

100 100 100


Total
(N=225) (N=623) (N=1098)

[Link] 57/76
2/9/2021 Data analysis

Similarly, Table 21 can be useful to make statements about the class composition of each category of job status. 65% of the low
status workers have working class fathers, whereas 131 % are classed as intermediate and 3% service class. In this table, the
raw data has been converted to percentages of the row totals.

Table 21: Job status by class background: row percentages

Father’s class
Job status in Magistrate Service
Service Class Intermediate Class Working Class Total

Low 3.1 31.7 65.2 100 (N=1353)

Medium 33.2 27.3 39.5 100 (N=370)

High 35.6 36.8 27.2 100 (N=253)

Obviously the choice of percentage breakdown (i.e. whether row or column) will be determined by the research in question.
Here, for example, the research emphasis is more likely to be placed upon using class background as a causal variable in
explaining current job status and the table showing column percentages (i.e. Table 20) will be more useful in investigating this
relationship. This table indicates to the researcher, for example, what the chances are of an individual (say) from a working class
background having a 'high status' job in the Magistrate service.

Collapsing categories

It may sometimes be useful to simplify the data contained within tables to aid in the modelling process. Although the number of
categories shown in the previous set of tables was not really high, some data table can contain many categories and 'exploring’
for patterns becomes quite difficult and complex. Using again the example relating to job status and class, the job status;

[Link] 58/76
2/9/2021 Data analysis

categories could be reduced to just two. Most of the variation seems to be between low status and the other two categories
viewed as a whole and so an obvious way to reduce this information would be to merge the medium and high job status
categories. The column percentages for this new, simplified table are shown in Table 22 below:

Table 22: Reduced category table: column percentages

Father’s class
Job status in Magistrate Service
Service Class Intermediate Class Working Class

16.5 68.9 80.3


Low (Medium and High
(83.5) (31.1) (19.7)

100 100 100


Total
(N=225) (N=623) (N=1098)

The resultant table is now very simple and straightforward to interpret. The job status category has been reduced to a
dichotomy (ie two categories) and because column percentage figures are given there is really no need to even show the figures
in brackets .A single row, illustrating the percentages from each class background who are in low status jobs in the, Magistrate
Service is sufficient to describe the data. For example, over 60% more people with a working class background are categorised
as having low status jobs than those from the service class.

Using proportions to analyse and model tables

So far, the raw data have been converted to percentages to make statements about the relationships found in tables. A more
powerful alternative is the use of proportions because, unlike percentages, they can be multiplied together. Again, the
researcher has to make a decision regarding the way in which proportions are calculated. Marsh (1988, p.143) provides a rule
for dealing with proportions in contingency tables and instructs the researcher to " construct the proportions so that they sum

[Link] 59/76
2/9/2021 Data analysis

to one within the categories of the explanatory variable”. The explanatory variable is the variable which is thought to be causing
the other variable. So in the class and job status example, class of father would be regarded as the explanatory or causal
variable. It is assumed that there is some type of causal relationship between class background and current job status.

Using the rule that proportions should sum to one within the categories of the explanatory variable, a new table can now be
constructed. This is shown below in Table 23. Note that the explanatory variable has now been placed in the row position as it
is more usual to have the explanatory variable in this part of the table.

Table 23: Class background and job status: raw number (N) and proportions (p)

Low job status High or Medium job status Total

Class of Father N p N p N p

Service 42 0.16 213 0.84 255 1.00

Intermediate 429 0.69 194 0.31 623 1.00

Working 882 0.80 216 0.2 1098 1.00

1353 623 1976

Using proportions to model the tabular data

The table can now be used for the basis of modelling and construction of a DFR equation. To do this a base for comparison
must be decided upon. For instance, to determine whether 0.16 of service class people in low status jobs is high or low, it must
be compared with the other figures in that status group (ie 0.69 of intermediate class and 0.8 of working class).

[Link] 60/76
2/9/2021 Data analysis

One number is therefore used as the yardstick with which to compare all other values. It I represents the fit in the DFR equation
and by comparing all values with this base, the causal nature of the model can be investigated. Which category should be used
as the base? Marsh (1988, p.145) lists some guidelines:

The base should include a large number of observations in order to be a reliable comparison with all other values
It should be a category that is of particular interest
If there is one category which is very different to the others, then this should be used because attention will centre on the
difference

A suitable choice of category as the base for comparison in this example would be the working r class in low job status. This
category is a fairly large group and will therefore be a fairly reliable value to use in comparisons.

Figure 18 illustrates the influence of the explanatory variable (ie father's social class) on the response variable -job status.
There is a three-fold classification of the explanatory variable. The base comparison class (ie working class) is shown
underneath the horizontal line. The explanatory effect of the other two classes - the service class and the intermediate class - is
denoted by b1 and b2, respectively. The path line denoted by ‘a’ represents those people whose father was working class and
who currently hold high or medium status jobs. This latter path represents the element which is not explained by this particular
model. The model is attempting to quantify the other two paths: b1 and b2 whereby b1 represents the effect of being in the
service class as opposed to being working class on the likelihood of holding at high/medium status job. Path b2 represents the
effect of being in the intermediate class on the chances of being in such a category of job status.

Figure 18: The relationship between father's class and job status

[Link] 61/76
2/9/2021 Data analysis

One method in which these two elements can be quantified is through the use of proportion differences (Davis, 1976). A d
measure is calculated which represents the effect of the explanatory variable (i.e. father's class) on the response variable (i.e.
high status jobs). So for b1, it is derived by subtracting the proportion of the working class in high medium status jobs from the
proportion in the non-base category (i.e. service class, high medium status).

So, the effect b1 can be calculated as:

0.84 - 0.2 = 0.64

This means that people with a service class background are more likely to have a high/medium status job in the Magistrate
Service than persons with a working class background. The magnitude of this difference is 0.64. In other words, 0.64 more
service class than working class have such jobs.

Similarly, b2 can be calculated:

0.31- 0.2 = 0.11

0.31 represents the proportion of people with an intermediate class background in high/medium status jobs and 0.2
represents, again, the proportion with a working class background in such jobs.

Both of the d values are positive and represent a greater chance of being in such jobs. If, for, example, a different base category
was chosen and we were trying to explain the effect of class on being in a low status job then b1 would result in a negative value
of the same magnitude as the value shown above (ie - 0.64). This is found by subtracting the proportion of individuals who have
a working class background and low status job (0.80) from the proportion of service class with low status jobs (0.16).

The value of path a is 0.2 and represents the proportion in the base category of the explanatory variable (ie working class) who
fall in the non-base category of the response variable (ie high medium job status). As already stated, this value represents the
fitted value to which b1 and b2 are each added, and it is not prefixed by a plus or negative sign.

Attention can now return to the DFR equation:

Data = fit + residual

and the proportion of people with a working class background who are found in high status jobs (i.e. 0.84) can be reduced into
a 'fit' component (0.2) and an effect (0.64). From Table: 23, it can be deduced that the overall proportion who are in high status
jobs is 623/1976 = 0.32. Later this section, an equation of the form:

[Link] 62/76
2/9/2021 Data analysis

is used to describe the relationship between a response or dependent variable (Y) and : explanatory variables

(X 1X2 etc.). In this example, there are two explanatory variables – X1 and X2 representing the proportion in the service class and
intermediate class, respectively and so the equation for the overall proportion in high/medium status jobs (Y) (as calculated f
above at 0.32) can be written as:

Y = 0.2 + (0.64 x 0.129) + (0.11 x 0.315)

= 0.32

Modelling tabular data using proportions: summary

The discussion above has illustrated how data expressed in tabular form can be analysed to investigate and model the
relationship between causal and explanatory variables using difference in proportions (d). The overall proportion in the
response variable (i.e. in the high/medium status jobs) has been decomposed into a 'fit' element and an 'effect' element, the fit
being represented by the proportion of people with working class backgrounds in high/medium status jobs and the effect
element represents the influence of being service or intermediate class on the overall proportion. The effect of father's class on
an individual's current Magistrate Service job status has therefore; been modelled using data, originally; presented in nominal
form (i.e. counts in categories).

Attention now moves on to modelling interval or ratio level data. Again there are many techniques for modelling such data, just
as there are many other ways of modelling nominal data. The techniques of simple and multiple regression, outlined below, are
commonly used: in social research and are discussed in many texts. An understanding of their conceptual base i will help the
researcher - if so desired -to study and understand, independently, many of the other common modelling techniques which, for
various reasons have not been covered in this.

[Link] 63/76
2/9/2021 Data analysis

Simple Linear Regression and Multiple Regression

Simple linear regression is concerned with modelling the association between two variables: measured as ratio data and is
included in the widely-used family of statistical techniques known as General Linear Models (see Rose and Sullivan, 1993;
O'Brien, 1992). It extends beyond simple correlation because it assumes that one variable is causally linked to another. Whilst
correlation produced an indication of the strength and direction of the association between two variables, regression allows
investigation into the form of the relationship. In other words, questions can now be asked such as:

What affect does increasing the causal variable by a fixed amount have on the response variable.

As with other statistical methods, the aim is to simplify and summarise complex datasets to discover underlying patterns.

In regression, it is assumed that one of the variables (the 'X' variable) is affecting (i.e. causing) the values of the other variable
(the 'Y' variable). The form describes the nature of this causality and allows the researcher to predict the likely value of Y at
given values of X. This is achieved by producing a line of 'best fit' through the scatter of plots produced by the X, Y pairs using
the ordinary least squares method (OLS regression).

OLS regression can only be used to calculate the line of best fit if it is assumed that X is -1 causing Y. In this sense, X is referred
to as the independent or explanatory variable and Y, the dependent or response variable. The whole concept is illustrated in
the hypothetical' example shown in Figure 19.

Figure 19: Regression - finding a line of best fit


[Link] 64/76
2/9/2021 Data analysis

Before detailing the calculation of an OLS regression line, it is first necessary to understand the nature of the equation which
describes a straight line. A straight line is always represented by the following equation:-

where Y is the dependent variable, X the independent variable, a is where the line intercepts the Y axis and b indicates the
slope. The examples below illustrate these terms.

Example 1 - In Figure 20 below, the independent variable (X) equals the dependent variable (Y). The slope is therefore equal to
1. The slope measures the steepness of the line and is the change in Y expressed in terms of the change in X In this instance, if X
increases or decreases by (say) 1, then so does Y. Furthermore the line cuts the Y axis at 0. If all of these are put into the
equation of a straight line, then the line can be expressed as:-

Figure 20: a straight line where y = x

[Link] 65/76
2/9/2021 Data analysis

Incidentally, the formula of the line shown in Figure 22 is:

Y = 1 -0.33 (X)

and it slopes downwards from the Y axis. This is indicated by the negative sign in the equation. The correlation coefficient of the
relationship would, of course, also be negative.

Figure 22: A straight line where y = 1 + O.33x

[Link] 66/76
2/9/2021 Data analysis

As previously stated, the OLS method is used to fit a regression line to a scatter of points between an X and Y variable. Only one
explanatory or independent variable is used to explain Y and this type of regression is known as ‘Simple Linear Regression’.
Other techniques are used for ‘multiple regression’ , where more than one explanatory variable is involved (see later).
Furthermore, the OLS method can only be used for straight lines. If plots are curvilinear, they must just be transformed to a
straight line relationship before OLS can be used (see Erickson and Nosanchuk, 1979, pp.100-119). Finally, with this method, the
Y variable must be the phenomenon which is being explained by the X variable (e.g. unemployment rates (X) explain suicide
rates (Y)).

Example 2 - In Figure 21 a much more shallow line cuts the Y axis at 1. The values of X and Y show that the slope of the line is
1/3 (0.33). The equation of the line is therefore:-

Y = 1 + 0.33 (X)

Figure 21: a straight line where y = 1 + O.33x

The least squares method of producing a regression line works on the principle of minimising the sum of the squared
differences between the scatter of points and the line. In other words, the sum of these squared deviations is smaller than the
sum of squared deviations produced by any other line. In the calculation of these deviations, the vertical deviations are used
(i.e. measured along the dependent variable axis (Y)) because the analysis is also concerned with the residuals - the variation in
y not explained by this average linear relationship with X. As already outlined, the equation of a straight line can be represented
by:-
[Link] 67/76
2/9/2021 Data analysis

The OLS method provides an equation to calculate the slope (b) and the intercept (a) of the regression line. These are shown
below:-

To calculate the slope:-

To calculate the intercept:-

The use of these equations is illustrated in the example shown in Table 24 where unemployment is thought to cause suicide.
The dependent or Y variable is therefore suicide, and the independent or explanatory variable (X) is unemployment. The table
also shows some of the calculation used in the regression formulae.

Table 24: Suicide rates and Unemployment

Suicide – rate per 100,000 (Y) Unemployment % (X) (XY) (X2)

20 8 160 64

16 7 112 49

18 6 108 36

[Link] 68/76
2/9/2021 Data analysis

14 5 70 25

12 5 60 25

14 4 56 16

8 3 24 9

10 1 10 1

12 7 84 49

15 1 15 1

17 6 102 36

21 9 189 81

14 6 84 36

8 4 32 16

16 6 96 36

18 6 108 36

[Link] 69/76
2/9/2021 Data analysis

17 9 153 81

18 7 126 49

6 3 18 9

9 4 36 16

Mean X = 5.35
Mean Y = 14.13 ∑XY = 1643 ∑X2 =671
2
(X) = 28.62

To calculate the slope:-

b = 1.3 (= SLOPE)

To calculate the intercept:-

a= 14.15 – 1.3 (5.35)

[Link] 70/76
2/9/2021 Data analysis

=7.2 (= INTERCEPT)

Therefore the equation of the line of best fit is:-

Y = 7.2 + 1.3X

The equation indicates that the regression line will meet the Y axis at 7.2 (i.e. the intercept) and it has a slope of 1.3. The slope
simply indicates the amount by which Y increases for each unit increase of X. The' + ' sign indicates that the relationship is
positive (i.e. as X increases, so does Y - as X decreases, so does Y).

[Link] 71/76
2/9/2021 Data analysis

Summary and further help

In summary, this lesson has provided an insight into some of the practical issues regarding research. In particular, it has
covered some of the more popular methods of gathering primary and secondary sources of research data. It has briefly
introduced the reader to a selection of methods for dealing with qualitative data and has concentrated on a selection of
techniques for describing and analysing quantitative information. It must be stressed however that the lesson has been unable
to discuss all types of research method. An individual student may stumble across a more appropriate technique for her or his
research in relevant literature. If this happens, then do not be afraid to use these other approaches. Remember to look at as
many general research texts as possible in order to gain a wider view than the account given here. Again, do not choose a
particular data collection strategy or data analysis technique without first discussing this with tutors or supervisors and finally,
Good Luck!

Some further videos to help:

Tutorial 1 – Internal reliability (questionnaire)

[Link] 72/76
2/9/2021 Data analysis

Tutorial 2 – Correlation

[Link] 73/76
2/9/2021 Data analysis

Tutorial 3 – T-Tests

[Link] 74/76
2/9/2021 Data analysis

Tutorial 4 - Chi-squared

[Link] 75/76
2/9/2021 Data analysis

Programme Title: Chi-squared

[Link] 76/76

You might also like