0% found this document useful (0 votes)
7 views25 pages

Module2 - Introduction To Quantitative Data Analysis

Module 2 introduces key principles of quantitative data analysis, focusing on summarizing data to identify patterns and relationships between variables. It covers univariate data summaries, frequency distributions, and the role of variability in comparative analysis, along with essential statistical concepts such as estimation and hypothesis testing. The module emphasizes the importance of understanding the nature of data variation and the appropriate methods for analysis.

Uploaded by

renzo_ortiz
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
7 views25 pages

Module2 - Introduction To Quantitative Data Analysis

Module 2 introduces key principles of quantitative data analysis, focusing on summarizing data to identify patterns and relationships between variables. It covers univariate data summaries, frequency distributions, and the role of variability in comparative analysis, along with essential statistical concepts such as estimation and hypothesis testing. The module emphasizes the importance of understanding the nature of data variation and the appropriate methods for analysis.

Uploaded by

renzo_ortiz
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Module 2: Introduction to Quantitative Data Analysis

Module 2: Introduction to Quantitative


Data Analysis
Antony Fielding1
University of Birmingham & Centre for Multilevel Modelling

Rebecca Pillinger
Centre for Multilevel Modelling
Contents
Introduction.............................................................................................. 2

C2.1 Univariate Data Summary .................................................................... 3

C2.1.1 Frequency distributions .................................................................. 3


C2.1.2 Summary statistics of key features of distributions ................................10

C2.2 Comparisons and relationships: the role of variability............................... 13

C2.2.1 Comparing subgroups as a form of relationship.....................................13


C2.2.2 Relationships and their direction......................................................14
C2.2.3 The role of variability ...................................................................16
C2.2.4 Homoscedasticity and heteroscedasticity............................................18

C2.3 Some other examples of relationships and the role of explained variability ... 19

C2.3.1 Extending to two categorical explanators for a continuous response: the


possibility of interactions...........................................................................19
C2.3.2 Variability of categorical responses...................................................24
C2.3.3 Where both response and explanatory variables are continuous.................26
C2.3.4 Patterns other than straight lines .....................................................29
C2.3.5 Relationship between a continuous response and combinations of categorical
and continuous explanators: Progress of students in schools.................................31

C2.4 Working towards the idea of a formal statistical model ............................. 38

C2.5 Comments on statistical inference; uncertainty, estimation and hypothesis


testing 42

C2.5.1 Estimation.................................................................................42
C2.5.2 Confidence Intervals.....................................................................45
C2.5.3 Testing hypotheses ......................................................................46

1
With contributed material from Kelvyn Jones and Fiona Steele and extensive comment by Harvey
Goldstein.

Centre for Multilevel Modelling, 2008 0 Centre for Multilevel Modelling, 2008 1
Module 2: Introduction to Quantitative Data Analysis
C2.1.1 Frequency distributions
C2.1 Univariate Data Summary
Some of the sections within this module have online quizzes for you Before we decide on appropriate methods of analysis to address our research
question we will usually want to ‘get to know’ our data. At the very minimum we
to test your understanding. To find the quizzes: will want to establish that the variables we plan to use do indeed vary. We will
also want to establish an initial picture of the nature of that variation, since it is
the study of such variation that will form the basis of further analysis. Thus the
EXAMPLE first step in any quantitative investigation is to separately summarise each
particular variable in various informative ways. Since these initial summaries deal
From within the LEMMA learning environment with only one variable they will be called univariate analyses. For instance in our
• Go down to the Lesson for Module 2: Introduction to Quantitative Data research we might be contemplating using the variable Years of Education, from
Analysis the European Social Survey data introduced in Module 1. It will be initially useful
• Click"2.1 Univariate Data Summary" to know something about the distribution of respondents over the values of this
to open Lesson 2.1 variable. What is the general level of education? How widely dispersed are the
• Click to open the first question values? Is there a concentration of a large number of units (individuals) at
Q1(i)
particular points? Are there units with extreme values in our data set? Are there
curious or inexplicable values which may lead us to suspect errors? In the case of
nominal variables, such as Marital Status or Ethnic Group, or of ordinal variables,
Introduction such as Income measured by twelve ranges, how many units fall into each
category? The answers to these questions will affect the analyses we perform.
The aim of this module is to give an overall view of some the principles of
effective data analysis. The focus is on how we summarise data to uncover We should also re-emphasise the major point discussed in Module 1 that the level
patterns and relationships between variables, and how these relationships can of measurement of a variable may restrict what summaries are appropriate.
begin to explain the values of the variables that we observe. Some key statistical
vocabulary is introduced and concepts are illustrated by example. You will learn
about the following: C2.1.1 Frequency distributions

• Ways of summarising the shape and pattern of the values of a single Frequency distributions are perhaps the most commonly used initial summaries.
variable, at different levels of measurement. They count how many in the set of units under consideration have different values,
or groups of values, of the variable. Frequency distributions can usually be
• Understanding the role of variability in comparative analysis and the study employed whatever the level of measurement. However, there are a number of
of relationships. conventions which influence different ways of grouping values.

• The elaboration of relationships, the importance of control variables, and C2.1.1.1 Categorical data
the concepts of confounding, suppression and interaction.
Table 2.1 and Table 2.2 represent typical summaries for the categories of a
• The essential parts of a statistical model: pattern and residual variation. nominal variable, Marital Status, and an ordered variable, Education Level (defined
as the highest level of education attained).
• Key concepts in inference from samples.

Centre for Multilevel Modelling, 2008 2 Centre for Multilevel Modelling, 2008 3
Module 2: Introduction to Quantitative Data Analysis
C2.1.2 Summary statistics of key features of distributions
Table 2.1. Frequency distribution of Marital Status
These tables show the number of respondents in a subsample from the European
Number % of all cases % of valid cases Social Survey (ESS) data that fall into each of the categories of the variable.2 There
Married 22974 54.2 54.5 are also a number of categories relating to different sorts of ‘missing data’.
Separated 627 1.5 1.5 Handling such missing data is a complex problem beyond the scope of this module,
but it is important to recognise their existence. When they are removed we have
Divorced 2758 6.5 6.5
what are usually termed valid cases for further analysis. For the rest of this section
Widowed 3836 9.1 9.1 we will focus only on summarising valid cases. The tables also show percentage
Never married 11946 28.2 28.3 distributions. Percentages are usually easier to interpret because of their familiar
Refusal 102 0.2 - feel, and because they show us the relative frequency of the values of the variable
in the data. This enables us to better judge the relative importance of the values.
Don't know 32 0.1 - In Table 2.1 we can see, for instance, that over half of our subsample is married.
No answer 84 0.2 - In Table 2.2 we can see that there are very few cases in the first and last valid
Total 42359 100.0 100.0 categories. Since the ordering of the categories has meaning here, we note that
the extremes of Education Level contain relatively small numbers. The most
frequently occurring categories, ‘married’ in the first case and ‘upper secondary’
in the second, are called the modes of the distributions. There is an extra column
called the cumulative frequency in Table 2.2 that is appropriate for ordinal (or
grouped ratio or interval level) variables, and which gives the percentage of units
Table 2.2 Frequency distribution of Education Level with a value of the variable less than or equal to the value of the row in question.
So, for example, 72.8% achieve secondary education or lower or, to look at it
another way, 27.2% go beyond secondary education. A cumulative summary in the
Number % of all % of valid Cumulative %
opposite direction would produce figures like the latter3, showing the percentage
cases cases of valid cases
of units with a value of the variable greater than or equal to the value of the row
Not completed 1,684 4.0 4.0 4.0 in question.
primary
education It is often easier to pick out key features of the data by visualising them. Figure
Primary or first 5,634 13.3 13.4 17.4 2.1 is a pie-chart for the distribution of marital status where the angles (and thus
stage of basic the sizes) of the slices of the pie are in proportion to the relative frequencies in
Lower secondary 9,610 22.7 22.9 40.1 the last column of Table 2.1. The pie chart is perhaps the most common way of
or second stage visualising distributions for categorical data but bar charts are sometimes also
of basic used.
Upper secondary 13,730 32.4 32.7 72.8
Post secondary, 3,517 8.3 8.4 81.2
non-tertiary
First stage of 5,448 12.9 13.0 94.2
tertiary
Second stage of 2,419 5.7 5.8 100.0
tertiary
Refusal 43 0.1 -
Don't know 153 0.4 -
No answer 121 0.3 -
Total 42359 100.0 100.0
2
This is the subsample of cases made available on the ESS website for on-line analysis at
[Link]
3
but shifted down a row.

Centre for Multilevel Modelling, 2008 4 Centre for Multilevel Modelling, 2008 5
Module 2: Introduction to Quantitative Data Analysis Module 2: Introduction to Quantitative Data Analysis
C2.1.2 Summary statistics of key features of distributions C2.1.2 Summary statistics of key features of distributions

Table 2.3 Grouped frequency distribution of approximate Years of Residence in area at


interview (valid cases only)
Married
Whole year Number of Percentage Cumulative
Separated
range valid cases Percentage
Divorced
0-9 11606 27.8 27.8
Widowed
10-19 8667 20.7 48.5
Never married
Refusal 20-29 7470 17.9 66.4
Don't know 30-39 5400 12.9 79.3
No answer 40-49 3567 8.5 87.8
50-59 2369 5.7 93.5
60-69 1381 3.3 96.8
70-79 987 2.4 99.2
Figure 2.1: Pie chart of distribution of Marital Status
80-89 247 0.6 99.8
90-99 80 0.2 100.0
Total 41774 100
C2.1.1.2 Continuous data
There are a few new issues connected with frequency distributions for continuous
variables since to display patterns we must often group values in a meaningful
way. In the ESS there is a variable defined by answers to the question ‘How long
have you lived in this area?’, which we will call Years of Residence. Data are
recorded to the nearest whole year and the set of possible values are integers
between zero and ninety eight, so only a small percentage of units will take each
particular value. A full frequency distribution with ninety nine groups would be too
detailed, patterns would be difficult to discern and the table would not be very
informative. To avoid these problems the values must be grouped into classes or
intervals. In determining the appropriate number of intervals and their widths we
need to strike a balance between too many intervals (which leads to groups with
relatively few cases in each) or too few intervals (which results in too great a loss
of information). There are no hard and fast rules for this. One possible set of
intervals for Years of Residence is shown in Table 2.3. Here the intervals are of an
equal width of ten years, and are perhaps easier to handle when this is the case,
but as will be seen in another example later in this section there is no necessity
that this should be so.

Figure 2.2 Histogram of approximate Years of Residence in area at interview (valid cases only)

Centre for Multilevel Modelling, 2008 6 Centre for Multilevel Modelling, 2008 7
Module 2: Introduction to Quantitative Data Analysis Module 2: Introduction to Quantitative Data Analysis
C2.1.2 Summary statistics of key features of distributions C2.1.2 Summary statistics of key features of distributions

There are also useful ways of presenting such frequency distributions of continuous
variables graphically. The most common are histograms although other devices Table 2.4 Frequency distribution of households’ approximate monthly net income, all
such as stem and leaf plots (for small datasets) or box and whisker plots (also sources, in euros
called boxplots) may be employed. Figure 2.2 depicts a histogram for the
distribution of Years of Residence in Table 2.3. Each rectangle here corresponds to Income Number % of valid cases Cumulative %
one of the intervals, with the horizontal width of the rectangle proportional to the Less than 150 769 2.3 2.3
width of that interval. The areas of the rectangles are proportional to the numbers
150 to under 3005 1,988 5.9 8.2
(or percentages) in each range. It is important to note that the construction of a
histogram is governed by the areas of the rectangles, not their heights. Only if the 300 to under 500 2,996 8.9 17.1
intervals are of equal width, as they are here, are the two equivalent. The areas 500 to under 1000 5,038 15.0 32.1
give a proper visual interpretation of the relative sizes of these ranges. In the 1000 to under 1500 4,843 14.4 46.5
diagram an approximate smooth curve has also been drawn to connect the tops of
1500 to under 2000 4,115 12.2 58.7
the rectangles. This is called a frequency curve and gives us an impression of the
shape of the distribution. Here we have an example of an asymmetric distribution 2000 to under 2500 3,678 10.9 69.6
with a long tail, called a skewed distribution. Almost 50% of lengths of residence 2500 to under 3000 3,025 9.0 78.6
are under 20 years and after that point the number of units in each range 3000 to under 5000 4,539 13.5 92.1
decreases gradually as the length of residence increases.4 The long tail is therefore
5000 to under 7500 1,847 5.5 97.6
to the right and so this is an example of positive skew. If the long tail had been to
the left, i.e. towards smaller values along the continuous range, we would refer to 7500 to under 523 1.6 99.2
the distribution as negatively skewed. 10000
10000 or more 306 0.9 100.0
Sometimes the raw data provided to the analyst is already grouped into intervals Total 33667 100.0
of values, so that there is less flexibility and the only decision then might be
whether to summarise further by aggregating intervals. One such example is the
We note that almost 70% of incomes are under 2500, and there is then a gradual
Income variable, recorded from questionnaires in twelve ranges of values and
tailing off of values so that there are some very large values that are sparsely
coded 1-12 in our example dataset. The frequency distribution of valid cases is
represented. The pattern here is also of a positively skewed distribution but
displayed in Table 2.4. The widths of the ranges here are not equal but are wider
perhaps less extreme than that of the Years of Residence distribution: there is a
for ranges at the higher end of the income scale.
(short) tail at the lower end of the income scale as well as a long tail at the higher
end. Such skewed distributions are very often found with variables measuring
concepts such as income and wealth.

5
Note that for Income, we name the categories, for example, ‘150 to under 300’ whereas for Years
of Residence they were named, for example, ‘10 – 19’. This is because before the respective
categorisations Years of Residence was discrete while Income was continuous. In both cases we
want to avoid a categorisation in which some values could go into more than one category. For
Years of Residence we can do this by making the end point of each category one unit less than the
start point of the next. But if we did this for Income— by having categories ‘150 to 299’, ‘300 to
499’, and so on— then there would be nowhere to put 299.50, for example. And because this is
monthly income, so that it may be arrived at by dividing a yearly figure by 12, the values need not
be whole numbers of cents either. It therefore would not work to have categories ‘150 to 299.99’,
‘300 to 499.99’, and so on. Thus the only way to make sure that every value will fit in a category
4
More detailed examination of the ungrouped data would reveal that the 11606 cases in the first and no value can go into two different categories is to make the end points of each category ‘under
interval are almost uniformly spread over each of the separate single values in that group. [start point of the next category]’.

Centre for Multilevel Modelling, 2008 8 Centre for Multilevel Modelling, 2008 9
Module 2: Introduction to Quantitative Data Analysis Module 2: Introduction to Quantitative Data Analysis
C2.1.2 Summary statistics of key features of distributions C2.1.2 Summary statistics of key features of distributions

importance of different values or categories. Percentages are a type of summary


C2.1.1.3 Symmetric distributions for continuous data: the normal curve statistic.

For categorical data, percentages form the basis for most of the useful things one
can summarise about distributions. For continuous data and sometimes ordinal
data, mainly due to their higher level of measurement, we are interested in
specific key features of distributions and other summary statistics come into play.
We should emphasise again the point made earlier in this module that most of
these measures are only appropriate for data at these measurement levels. The
temptation to use them for nominal variables simply because their values have
numerical codes must be strenuously resisted.

First we will discuss ways of summarising central tendency. This means specifying
where along the numerical scale of the variable the set of units is located:
measurements of various kinds of average. The physics concept of centre of gravity
comes to mind as an analogy. Second, we will look at summary measures of
dispersion which give us an idea of the extent to which units are scattered or
concentrated around this central location. In Module 1 we mentioned that such
Figure 2.3 The normal distribution variation is central to most statistical analysis. As the modelling framework
develops in later modules we will begin to see that it is essentially these features
The positively skewed shapes of the two examples of frequency distributions we on which we will concentrate. In short we will see how that framework in general
have discussed are in some sense atypical of many we come across in empirical seeks explanations for why a set of cases or subset of cases is located where it is,
social science research. Far more commonly the shape of the frequency curve for relative to another set. Or we ask why individuals in one set vary away from a
a continuous variable is symmetric and has approximately the shape of the typical ‘expected’ value. Again we might be interested in why one set of cases is
distribution shown in Figure 2.3. This distribution is called a normal distribution. concentrated around a central value whilst another is more dispersed.
The normal curve has particular mathematical properties, one of which is that it is
symmetric. It plays a pivotal role in statistical analysis. It is a remarkable fact
that, across a wide range of scenarios, crucial variables do empirically seem to C2.1.2.1 Measures of central tendency
have a normal distribution.
The most commonly used statistic for measuring central tendency is the mean.
In later modules on statistical modelling, we will see that often it becomes Indeed we believe most readers of this module will already be familiar with it,
necessary or desirable to make assumptions that certain variables’ distributions6 although might have used another name for it. It is often called the average,
follow a normal shape. An examination of their histograms becomes a way of although this term can also be used as a generic term for all measures of central
checking that assumption, as we will see in the next module dealing with multiple tendency. Sometimes the term ‘mean’ is also used interchangeably with the
regression. Where we are dealing with a skewed variable from our dataset, such as concept of ‘expected value’, which will have a pivotal role in statistical modelling.
Income in Table 2.4, we have the option of transforming it (e.g. by taking From raw data on numerical values calculating the mean is straightforward
logarithms) so that it looks closer to normal. This is the most common way of enough. We sum the values of the variable over all the units and then divide by the
handling income and wealth variables in much applied economics research, since number of units. For example, when we perform this calculation for Age in the ESS
these variables are usually quite skewed. dataset, we find that it is centred at a mean of 46.7 years. However, if the raw
data only record ranges of values, as for Income in Table 2.4, then we need to
C2.1.2 Summary statistics of key features of distributions make some approximating assumptions in order to calculate the mean. What is
usually good enough for practical purposes is to assume that individuals in a range
The previous section showed ways of presenting data with the aim of making it are distributed evenly through the range. This assumption is carried forward in
easier to see its shape. Before constructing a table or chart, we summarised the calculating the mean by assuming every case has the mid-point of its range as its
data by counting the number of observations with the same value or group of value. With this convention applied to Table 2.4, the mean of the income variable
values. Percentaging was seen to be a useful way of assessing the relative is 2079 euros.

Another summary measure, which may be more useful than the mean for skewed
6
These might be variables from our original dataset, or variables that we create in our analysis.

Centre for Multilevel Modelling, 2008 10 Centre for Multilevel Modelling, 2008 11
Module 2: Introduction to Quantitative Data Analysis Module 2: Introduction to Quantitative Data Analysis
C2.1.2 Summary statistics of key features of distributions C2.1.2 Summary statistics of key features of distributions

variables such as income, is the median. This is defined to be the ‘middle’ quartile (LQ) and upper quartile (UQ) divide the set of units into four parts each
observation in the sense that half of the units in the distribution lie below it and containing 25% of the observations. Values below the LQ account for 25% of the
half above. If the distribution of units is symmetric, as in the case of variables units, whilst 25% lie above the UQ. We then define the IQR as the distance
which follow a normal distribution, then half the units will lie to the left of the between the LQ and the UQ. Although it is not sensitive to outliers, a disadvantage
mean and half to the right and so the mean and median coincide; there is nothing of the IQR is that it ignores units outside the middle 50%. We could also use ranges
to choose between them. Note that this statement is true only for the population with other percentages of units below the start and above the end point, for
from which our sample is drawn: it is unlikely that the mean and median will example 5% or 10%, but these have the same disadvantage that some units are
coincide exactly in our sample, but we will expect them to be close. being ignored.

For variables from a normal distribution the mean and median are also the same as The most common measure of dispersion is the standard deviation. This assesses
the mode, which is the peak of the frequency curve. Recall that for categorical the extent to which the observations vary around their mean; it measures the
data the mode is the category which occurs most frequently; notice how these average difference between the values of a variable and their mean. First we take
definitions are similar. the difference of the value for each unit from the mean, to give us the deviations
from the mean. Some of these will be positive and some negative. If we averaged
A difference between the mean and median occurs if the distribution is skewed these we would find they would cancel each other out and the result would be
(since it is then not symmetric). Distributions such as that for Years of Residence in zero. One way round this is to ignore the sign and average the absolute deviations.
Figure 2.2, or for Income, are positively skewed (skewed to the left) and have Indeed this measure, called the mean deviation, is occasionally used. However
large extreme positive values (extreme in the sense that there are not many units more usually we square the deviations and then average. The result is the variance
at these values). These contribute large values to the mean and pull it upwards and the standard deviation is its square root.
when the averaging is carried out. The mean will then be larger than the median,
while the mode will be smaller than the median. The reverse would be the case for When summary data are reported the standard deviation is usually given as a more
distributions which are negatively skewed, with a long tail towards smaller values. natural descriptive measure of variability than the variance. The reason it is more
Comparing the mean and median (and mode) will indicate the direction of skew natural is that its scale is the same as that of the variable it summarises. For
before the frequency distribution is plotted (but it is often hard to see from this example, if we have data on heights in centimetres and take the standard
comparison just how skewed the distribution is and whether it is far enough from deviation, and get a result of 20.5, this will be 20.5 cm. If we had measured the
symmetric to need to be treated as skewed, so it is still important to plot the same data in metres then we would get a result of 0.205 for the standard
distribution). The median for the income variable is approximated from the deviation, which will be 0.205 metres. The variance, on the other hand, will not be
frequency distribution to be 1643 euros, using the same assumptions as those we in centimetres or metres so is not so directly interpretable. However, for further
made to calculate the mean. This is considerably lower than the mean of 2079 analytical purposes and in particular when we begin to study the role of variation
euros. Thus there is a high degree of skewness and the mean is influenced by some in modelling, it becomes more natural to use the variance. In practice it does not
very large values. It could be argued that the median would be the more really matter since we can always calculate the standard deviation by taking the
appropriate indicator of average if that was taken to mean ‘typical value’. If using square root of the variance.
means or expected values is important for other, analytical, reasons, then it may
be advisable to transform the variable as suggested at the end of C [Link], so that We usually need to interpret the measure of variability, so a natural consideration
the distribution is more symmetrical. is what constitutes a ‘large’ standard deviation. This question, unless
contextualised, has little meaning. In a sense it is a ‘how long is a piece of string?’
C2.1.2.2 Measures of variation or dispersion type of query. On its own the standard deviation tells us nothing since it depends
on and must be referenced by the scale of measurement of the variable of
One obvious measure of the extent of variation is the range of values taken by the interest. For instance if we measured age in months rather than years then the
continuous variable, i.e. the difference between the highest and the lowest value standard deviation would be multiplied by 12, so that we would have a larger
amongst units being studied. Age in the ESS takes values ranging from 14 years to standard deviation but, of course, the same amount of variability. Furthermore,
98 years so the range is 84. Years of Residence goes from zero up to 98 so its range large variation is a relative concept and, as we shall explore more fully in the next
is 98. A disadvantage of the range is that it is very sensitive to extreme values or section, most summary statistics for a particular set of units can only be given a
outliers so that it may not give a clear idea of a set of values between which units sensible interpretation in a comparative setting. To illustrate these points,
generally lie. consider the following set of values for Monthly Expenditure on Magazines (£) for a
group of males: 10, 12, 15, 18, 20. The mean is £15 and the standard deviation is
An alternative measure of dispersion is the interquartile range (IQR) which is based £3.7. Is this large? Are males highly variable in their expenditure? Surely we cannot
on positional points akin to the median. Together with the median, the lower say unless we have a standard by which to judge it. Consider a group of females in

Centre for Multilevel Modelling, 2008 12 Centre for Multilevel Modelling, 2008 13
Module 2: Introduction to Quantitative Data Analysis Module 2: Introduction to Quantitative Data Analysis

a similar setting with expenditures 2, 8, 15, 22, 28. The mean of this group is the
same as for males, £15, but we might note that female expenditures are much
more variable with a standard deviation of £9.3. As evidenced by comparing their
C2.2 Comparisons and relationships: the role of
standard deviations we can now say that males are less variable than females, variability
even though on average they spend the same. We might also note that in making
our interpretation using this comparison of standard deviations we referred to the In the previous section we briefly introduced the idea of comparisons as a means of
means. This is important. What would we make of two groups which had the same contextualising data summaries. Even the most basic statistical description is
standard deviation, £5 say, but different means of £20 and £100? We will probably usually concerned with comparisons of one sort or another. Only rarely would we,
agree that on a relative basis the second group exhibits less variability. The for instance, produce a frequency distribution of a variable, calculate a median or
coefficient of variation (CV) overcomes this scale problem and may also be used to a mean, or graph a pie chart for a set of units without implicitly or explicitly
adjust for comparisons of different variables on the same group using different referring the result to another comparable situation. We may find that the mean
units of measurement. It is a scale-standardised measure defined as the standard score on an educational attainment test for a set of girls this year is 52. But of
deviation divided by the mean. In the example above, the group with a mean of what interest is this in its own right? It is useful only if, for instance, we want to
£20 has a CV of 0.25 and the group with a mean of £100 has a CV of 0.05, and see how girls fared relative to boys. Or how this set of girls compared to a similar
comparing these may give a better impression of the relative variability. set of girls last year. Or, implicitly, does this score meet certain external standards
or targets? It would not be stretching a point to say that behind most statistical
There is another useful interpretation of the standard deviation which will crop up analysis lie comparisons of one sort of another.
very frequently in the material in later modules. If the variable of interest follows
a normal distribution then approximately 95% of the units lie within 2 standard We will stress that crucial to many of these comparative exercises is the study of
deviations on either side of the mean, and 68% within one standard deviation of variability.
the mean. This idea will also be exploited in C 2.5 when it is used to assess
whether extreme sample results can have arisen by chance. C2.2.1 Comparing subgroups as a form of relationship

Don’t forget to take the online quiz! Table 2.5: Distribution of stated working hours during term of full-time secondary school
teaching staff in Pathfinder schools in 2002

From within the LEMMA learning environment Senior Staff Non- Senior Staff All teachers
• Go down to the Lesson for Module 2: Introduction to Quantitative Data (Class Teachers)
Analysis
• Click"2.1 Univariate Summary Data" Hours per Number % Number % Number %
to open Lesson 2.1 week
• Click Q 1 to open the first question 31-40 4 2.6 44 9.4 48 7.7
41-50 40 25.5 163 35.0 203 32.6
51-55 44 28.0 150 32.2 194 31.1
56-60 28 17.8 62 13.3 90 14.4
more than 41 26.1 47 10.1 14.1
60 88
Total 157 100.0 466 100.0 623 100

Mean 54.2 50.3 51.2


SD 7.3
5.3 5.1

Centre for Multilevel Modelling, 2008 1 Centre for Multilevel Modelling, 2008 13
Module 2: Introduction to Quantitative Data Analysis Module 2: Introduction to Quantitative Data Analysis
C2.2.4 Homoscedasticity and heteroscedasticity

Consider Table 2.5 above which is a typical descriptive table arising from an variables might combine together in influencing the response. We will then deal
analysis of the variable Hours Worked for teachers in the Pathfinder study7 more formally with the mechanics of building such different kinds of combined
introduced in Module 1. This variable is derived from teachers’ answers to a influence into our analysis when we discuss statistical modelling in detail from
question asking how many hours a week they worked on average. Initially a Module 3 onwards.
frequency distribution and calculation of the mean and standard deviation (sd)
might be found for all teachers as shown in bold in the last column. Whether a variable is a response or an explanatory variable is not intrinsic to that
variable but depends on the role it has in research questions that are being asked:
Analysing all teachers together, the mean of Hours Worked is 51.2 with a standard what needs to be explained in a given context and by what means. Although we
deviation of 7.3. While, as we commented above, there is little interest in this may not be using our analysis as part of an attempt to demonstrate a causal
overall summary in isolation, it will provide a useful reference point when we go process, and may not have any evidence from other sources of a causal
on to compare means and standard deviations between groups. The first two relationship, we will nevertheless often informally have the notion of a particular
columns show the distributions and summary statistics for groups of teachers causal process in mind. We will take the variable(s) measuring the cause(s) in this
formed on the basis of a binary variable Seniority (with categories Senior and Non- process as our explanatory variable(s), and the variable measuring the effect as
senior). We are now beginning to get some more interesting results; there is a our response variable.
clear difference of almost 4 hours in the mean hours worked for the two groups.
For example, in our hypothetical discrimination study, suppose we only wanted to
This is also a basic first example of the notion of a relationship. Comparing describe how equal or unequal black and white salaries are, without making any
subgroups as we do here might be viewed as an investigation of whether two claims that any inequalities we might find were caused by discrimination. It still
variables measured on teachers, Seniority and Hours Worked, are related. In other makes more sense to think of Ethnicity as a ‘cause’ and Salary as an ‘effect’ than
circumstances comparisons might be over time or place, but still at the heart of to think of Salary as a ‘cause’ and Ethnicity as an ‘effect’. So we will take Salary
these will be similar ideas about relationships. Are hours worked, for instance, as our response variable with Ethnicity as the explanatory variable. The possibility
related to whether a teacher worked in a selective or non-selective school? of discrimination is something that we have in mind and is what motivates our
decision to carry out this analysis. It thus determines our choice of response and
C2.2.2 Relationships and their direction explanatory variables, even though our research question here is ‘What, if any, is
the difference in black and white salaries?’ and not ‘Has discrimination led to a
Ideas similar and relevant to those of comparisons and relationships are expressed difference in black and white salaries?’, and even though this analysis will not
using a variety of language or terminology. Often we are concerned with one-way provide any evidence to help answer the latter question.
relationships where we might be considering a particular variable as a dependent
variable. Another name for this is the response variable. In our example this might Thus it is the nature of the underlying research relation that is the key to the
be Hours Worked, and Seniority may be ‘influencing’, ‘explaining’ or ‘affecting’ it. status of a variable. The explanatory variables Seniority in the example of
Seniority can be termed an independent variable, an explanatory variable, a Pathfinder teachers or Years of Experience in the discrimination example could be
covariate or a predictor, the usage of these names loosely relating to the purpose response variables in other contexts if we were interested in explaining such
of the analysis (although some authors use the terms interchangeably). For outcomes.
example, when we refer to Seniority as a predictor we are admitting the possibility
that knowing a person’s seniority may improve our ideas about the hours they We will also often talk about ‘the effect of the explanatory variable on the
work. response variable’. This is a loose use of language in cases where we have not
established that the relation between the two variables is causal. However the use
We should also recognise that there may be more than one explanatory variable of ‘effect’ is very widespread and can help us to express suitable questions about
under consideration. In the research example of ethnic discrimination in the legal our analysis, provided we always keep in mind that it is just a use of language and
profession that we introduced in Module 1, the dependent variable could be there may not be a causal connection between the variables in reality. It is
Income. Explanatory variables could be any or all of Years of Experience, Ethnicity, particularly important to remember this when it comes to presenting our results to
Gender, or other relevant factors. We will shortly discuss different ways such others.

For example, after performing the above analysis we may talk about being black as
7
Thomas, H., Butt, G., Fielding, A., Foster, J., Gunter, H., Lance, A., Pilkington, R., Potts, E., having a certain effect on Salary, but when we draw our conclusions we will
Powers, S., Rayner, S., Rutherford, D., Selwood, I., and Szwed, C. (2004) The Evaluation of recognise that Ethnicity may not directly cause Salary: it may be that the
difference is due to different ages or lengths of experience among black
Transforming the School Workforce Pathfinder Project. Research Report 541, Department for
employees compared to white.
Education and Skills.

Centre for Multilevel Modelling, 2008 14 Centre for Multilevel Modelling, 2008 15
Module 2: Introduction to Quantitative Data Analysis Module 2: Introduction to Quantitative Data Analysis

Worked been the same (or almost the same) for Senior and Non-senior teachers-
In particular, in this course we will not get involved in the important but complex then Seniority would have had no power to explain the variability in Hours Worked
questions of what is needed to establish causality, which are partly philosophical. because there would have been no variability in the means for different levels of
In the examples that we present we will often discuss the results of our analysis in Seniority. We could not have claimed that there was a difference in the hours
terms of ‘effects’ of the explanatory variables on the response variable. We use worked by Senior and Non-senior teachers, and thus we could not have claimed
‘effect’ interchangeably with the more neutral ‘association’, but do not use either that Seniority was responsible for some of the difference in Hours Worked across
in a causal way. In real research, there are steps that can be taken to improve our teachers as a whole.
ability to make causal inferences, although in social science it is never possible to
say with absolute certainty that X causes Y. One such step would be to collect The pattern of differing means of Hours Worked across the categories of Seniority
longitudinal data on X and Y to determine whether a change in X precedes a now constitutes the explained variability. It is the part of the variability in values
change in Y, a necessary but not sufficient condition for establishing that X causes of Hours Worked which we understand and have accounted for. But there remains
Y. some unexplained variability: notice that even after introducing Seniority to
C2.2.3 The role of variability explain the differences in Hours Worked, there are still differences between
teachers. It is not the case that every Senior teacher has a value of 54.2 for Hours
If we look again at the third column of Table 2.5, giving values and summary Worked and every Non-senior teacher has a value of 50.3 for Hours Worked.
statistics for all teachers, we may note that the value of Hours Worked is not the Rather, within the group of Senior teachers there are 4 teachers for whom Hours
same for all teachers: there are 48 teachers for whom it lies between 31 and 40 Worked lies between 31 and 40 hours, 40 teachers for whom it lies between 41 and
hours, 203 teachers for whom it lies between 41 and 50 hours, and so on. There is 50 hours, and so on, and similarly for Non-senior teachers. Once again we can
thus variability in the values of Hours Worked among teachers. measure this unexplained variability using the standard deviation; this is very
similar for the two groups at 5.3 for Senior teachers and 5.1 for Non-senior
This may seem like a trivial observation: after all, we would be very surprised (and teachers. Notice how this is reduced compared to the original standard deviation
probably somewhat suspicious) if we found that all teachers had the same value of 7.3. We have thus been able to reduce the unexplained variability by
for Hours Worked. However, it leads us to an interesting question: what is introducing Seniority.
responsible for these differences? In other words, how do the values of Hours
Worked come about? To generalise, we start by looking at our response variable alone. We notice that it
exhibits variability. All this variability is as yet unexplained since we have
We made a start on answering this question when we noted the possibility of a introduced no explanatory variables, and we can measure this unexplained
relationship between Seniority and Hours Worked. We noted that the mean of variability using the standard deviation. We then introduce an explanatory
Hours Worked was different for the two categories of Seniority: it is 4 hours higher variable, and this divides the variability into explained variability and unexplained
for Senior teachers than Non-senior teachers. Thus some of the difference in variability.
values of Hours Worked may be due to the fact that some teachers are Senior
teachers and others are Non-senior teachers, and that Senior teachers work longer The explained variability is the variability in the mean of the response across
hours in general than Non-senior teachers. values of the explanatory variable. We could measure it to get a number as we did
when we measured the unexplained variability using the standard deviation, but
What we have done here is to attempt to explain some of the variability in Hours this is of little interest. What we are interested in is describing the pattern of the
Worked. We started off by looking at the variability in our response variable means: how does the mean of the response change as we change the value of the
without taking any other variables into consideration. At this stage, all this explanatory variable? This comprises what we understand so far about how the
variability is unexplained: we have made no effort to suggest why it is that values of the response variable come about. It is a description of the relationship
different teachers have different values of Hours Worked. We can measure this between the explanatory variable and the response variable.
variability by taking the standard deviation, which we can see from Table 2.5 is 7.3
The unexplained variability is what we don't (yet) understand about the difference
We then introduced the relationship between Seniority and Hours Worked in order in values of the response across the units. We can quantify it by measuring it with
to try to explain some of the variability, that is, to begin to answer the question of the standard deviation. We can then see how much it is reduced compared to the
why different teachers have different values of Hours Worked. It seemed likely standard deviation that we measured before introducing the explanatory variable,
that Seniority might provide some explanation because the mean of Hours Worked and this will indicate how much progress we have made in explaining the
was different for Senior and Non-senior teachers: we observed a pattern in the variability of the response variable.
mean of Hours Worked across categories of Seniority. There is thus variability in
the means. Had we observed no pattern- in other words, had the mean of Hours We can, as we will see, go on to add further explanatory variables. With each

Centre for Multilevel Modelling, 2008 16 Centre for Multilevel Modelling, 2008 17
Module 2: Introduction to Quantitative Data Analysis
Module 2: Introduction to Quantitative Data Analysis
C2.3.1 Extending to two categorical explanators for a continuous response: the possibility of interactions

addition, provided that there is variability in the mean of the response across C2.3 Some other examples of relationships and the
values of the new explanatory variable even after taking into account the variables
we have already included, we will reduce the amount of unexplained variability. role of explained variability
We will also, of course, gain a greater understanding of how the values of the
response variable come about, as we add to what we can say about the pattern of C2.3.1 Extending to two categorical explanators for a continuous
the response over the values of the explanatory variables.
response: the possibility of interactions
C2.2.4 Homoscedasticity and heteroscedasticity
The summary of the situation where Seniority influences Hours Worked above
might simply be represented by Figure 2.4(a) below. The vertical lines indicate the
Returning to an examination of the standard deviations of the two groups of
mean of Hours Worked for the two different groups and clearly show the
Seniority, there is more of interest to note than simply that they are reduced
dependency of the mean on Seniority. We recognise that around these lines there
compared to the overall standard deviation. In particular, taking into account
is still some considerable variability of the units, as shown by the distribution
natural fluctuations, these data are quite compatible with the idea that Stated
curves which indicate for the two values of Seniority how many units there are
Hours Worked per Week is homoscedastic with respect to Seniority: i.e. that
with each value of Hours Worked. We can see that relatively few of the units
variability is much the same for the two values of Seniority. If variability were
actually lie on the mean line for their group of Seniority: most units lie scattered
different for each value of Seniority, Stated Hours Worked per Week would be
around this line to one side or the other. As we indicated in the previous section,
heteroscedastic with respect to Seniority.
an approach to further analysis might be to see whether another variable can now
explain this as yet unexplained variability within the seniority groups. For
It should be noted that in the above discussion about variability we were assuming
simplicity let this other variable, a hypothetical one called A, be binary (e.g.,
that the response variable was homoscedastic with respect to the explanatory
whether the school that a teacher works in is large or small) with values A1 and
variable, although this was not explicitly stated. If we had observed unequal
A2. Relevant information can be extracted from a table such as Table 2.6 which
variances across seniority groups the situation is slightly more complicated, but the
simultaneously compares the effects of Seniority and variable A on Hours Worked
total variance (across groups) would still decrease when we account for seniority
and where one possible scenario is summarised. Important information is also
effects. (Note that the total variance is not equal to the sum of the group
contained in the number of units which fall into each of the four subgroups formed
variances, but is rather a weighted average of these variances.) Heteroscedasticity
by grouping according to both Variable A and Seniority. This number appears as the
is an important aspect of analysis and will be taken up as the course proceeds.
first number in each cell of the table.

The mean hours worked for the four subgroups are also displayed as lines in panel
Don’t forget to take the online quiz for this section! (b) of Figure 2.4. What can be seen from the table and the diagram is that
Variable A has an influencing effect in addition to Seniority: within seniority groups
there are some differences in the mean of Hours Worked between the categories
From within the LEMMA learning environment of variable A. For instance there is quite a difference between the means of A1
• Go down to the Lesson for Module 2: Introduction to Quantitative Data and A2 for Senior staff. Since here the variance is relatively homogeneous across
Analysis the four subgroups, so that we again have homoscedasticity, the unexplained
• Click"2.2 Comparisons and relationships: the role of variability" variance within groups is further reduced. We have thus made more progress in
to open Lesson 2.2 explaining variation in Hours Worked by using two explanatory variables rather
• Click Q 1 to open the first question than one.

Centre for Multilevel Modelling, 2008 18 Centre for Multilevel Modelling, 2008 19
Module 2: Introduction to Quantitative Data Analysis Module 2: Introduction to Quantitative Data Analysis
C2.3.1 Extending to two categorical explanators for a continuous response: the possibility of interactions C2.3.1 Extending to two categorical explanators for a continuous response: the possibility of interactions

Figure 2.4

Table 2.6: Teacher Hours Worked per week by Seniority and binary variable A: Additive
effects with two explanatory variables.

Senior Staff Non- Senior Staff All Staff


(Class Teachers)

n: 121 n: 180 n: 301

A1 Mean: 55.6 Mean: 54.1 Mean: 54.6

SD: 4.6 SD: 4.5 SD: 6.4


n: 36 n: 286 n: 322

A2 Mean: 49.6 Mean: 47.9 Mean: 48.2

SD: 4.7 SD: 4.8 SD: 6.3


n: 157 n: 466 n: 623

Overall Mean: 54.2 Mean: 50.3 Mean: 51.2

SD: 5.3 SD: 5.1 SD: 7.3

Centre for Multilevel Modelling, 2008 20 Centre for Multilevel Modelling, 2008 21
Module 2: Introduction to Quantitative Data Analysis Module 2: Introduction to Quantitative Data Analysis
C2.3.1 Extending to two categorical explanators for a continuous response: the possibility of interactions C2.3.1 Extending to two categorical explanators for a continuous response: the possibility of interactions

Table 2.7: Teacher Hours Worked per week by Seniority and binary variable B: Interacting
There is an additional feature of this table as displayed in Figure 2.4(b) which may effects
appear contrived, but which in fact exhibits what will often be a reasonable
assumption in models as to how two effects operate simultaneously: the two Senior Staff Non- Senior Staff
variables are additive in their effects on the response. For each value of Seniority, (Class Teachers)
the effect of variable A is almost the same: within the seniority groups there are
very similar differences of 6 hours (for Senior staff) and 6.2 hours (for Non-Senior Mean: 54.6 Mean: 55.0
staff) between the means of A1 and A2. In the diagram this is evidenced by the B1
distance between the Senior-A1 and Senior-A2 lines being much the same as the
distance between the Non Senior-A1 and the Non Senior-A2 lines. Similarly the
Mean: 52.6 Mean: 47.0
difference between the seniority groups is much the same whichever value of A we
are dealing with. B2

In fact, to say that the difference between A1 and A2 in the means of Hours
Worked is equal for the two groups of Seniority is equivalent to saying that the Table 2.7 and Figure 2.4(c) demonstrate another important concept, that of
difference between the two groups of Seniority in means of Hours Worked is equal interaction. This will play an important role in the statistical modelling of later
for A1 and A2; these are just two different ways of looking at the same situation. modules. Here, when we introduce an additional explanatory binary variable B, the
So when we want to specify this kind of effect in our models in later modules, we combined effects are no longer additive. The effect of B now changes according to
only need to make one specification: we do not need to specify both that the whether we are dealing with Senior staff or not: the difference in the mean of
difference between A1 and A2 is the same for both groups of Seniority and that the Hours Worked between categories B1 and B2 is 2.0 among Senior teachers, but 8.0
difference between the groups of Seniority is the same for A1 and A2. This is a in the Non-senior group. Thus the effect of B is much smaller for Senior staff than
general principle which is true for any two explanatory variables, and any response for Non-seniors, as seen in the diagram from the much smaller distance between
variable. In particular, note that the possibility of finding additive effects is not the Senior-B1 and Senior-B2 lines than between the Non-senior-B1 and Non-senior-
confined to the special case of explanatory variables with two categories. We can B2 lines.
equally well find additive effects for categorical variables with a larger number of
categories or for continuous variables, or for any combination of these. Note that if the effect of B is different for Senior and Non Senior staff, then it
must also be true that the effect of Seniority is different for B1 and B2. Indeed we
The diagram also clearly brings out another feature which is interesting. We can can see that while for B1 Senior staff have a slightly lower mean of Hours Worked
note in Table 2.6 and Figure 2.4(a) that the gap in the overall means of the two than Non Senior staff, for B2 Senior staff have a much higher mean than Non Senior
seniority group is 3.9 hours. As displayed in Figure 2.4(b), this is larger than the staff. Again, the differing effect of B on the response for different values of
gap between the seniority groups of around 1.6 hours when we hold variable A Seniority and the differing effect of Seniority on the response for different values
constant, i.e. when we look at the difference in seniority groups for just A1 or for of B are just two different ways of looking at the same thing, so that it will only
just A2. Why should this be? The clue lies in Table 2.6, in an examination of the take one specification to include effects of this type in our models. As before, we
number of units in each of the four subgroups formed by grouping according to can find effects of this type with continuous explanatory variables or for
both Variable A and Seniority. The two explanatory variables are themselves categorical variables with more categories.
related to each other: Senior teachers are more likely to have attribute A1 than
Non-seniors (121 out of 157 Senior teachers = 77% vs. 180 out of 466 Non-senior Since we have differing effects, we say that there is an interaction between
teachers = 39%). What has happened is that before we divide up each of the variables B and Seniority in their effect on the response. In this case, conditioning
seniority groups according to variable A, part of the overall effect of Seniority on on B does not simply reduce the effect of Seniority on Hours worked while keeping
Hours Worked that we observe is due to the fact that there are relatively more of this effect similar across different values of B. In other words, the second
A1 (which has higher hours overall) in the Senior group. When we hold variable A explanatory variable does not simply reduce the effect of interest on the response
constant and effectively control for its effect, the net effect of Seniority is while keeping this effect similar across different values of the second explanatory
reduced. We could equally well argue the other way round and start with the variable. This is in contrast to the situation when we conditioned on variable A in
overall effect of variable A. Thus this is also an example of where the two Table 2.6.
explanatory variables partially confound each other and demonstrates the
importance of elaborating on relationships by conditioning on other variables. The More generally, conditioning on a second explanatory variable may produce
ability to do this is a very powerful aspect of quantitative analysis, and an entirely contrasting patterns of effects: note, for example, how in this case, when
important reason for carrying it out. we condition Seniority on B, we find that for B1 being a Senior member of staff
lowers the mean hours worked slightly while for B2 being a Senior member of staff

Centre for Multilevel Modelling, 2008 22 Centre for Multilevel Modelling, 2008 23
Module 2: Introduction to Quantitative Data Analysis
Module 2: Introduction to Quantitative Data Analysis
C2.3.1 Extending to two categorical explanators for a continuous response: the possibility of interactions
C2.3.2 Variability of categorical responses

raises the mean hours worked a lot.


Table 2.8: Seniority and Gender of full-time secondary school teaching staff in Pathfinder
An extreme situation may occur where one variable may have no effect on the schools in 2002
response at all if another variable is held at certain values. By contrast for other
values of that variable the effect may be substantial. Particular values of the
Female Male Total
second explanatory variable become necessary pre-conditions for the original
explanatory variable to influence the response. For example, imagine a trial of a % (No.) % (No.) % ( No.)
new method of teaching children to read. The response variable is the difference Senior Staff 19.4 (70) 31.7 (84) 24.6 (154)
between each child’s score on a reading test before and after the trial period
(Progress), and the main explanatory variable is whether the child has been Non-Senior 80.6 (290) 68.3 (181) 75.4 (471)
assigned to the group which receives the new method of teaching or to one which (Class teachers)
receives the standard method (Teaching Method). A further explanatory variable Total 100.0 (350) 100.0 (265) 100.0 (625)
might be Gender, and it could be that the mean value of Progress for boys in the
group receiving the new method is much greater than the mean value of Progress
for boys in the group receiving the standard method, while for girls the mean for It is, as we mentioned above, of only limited interest to know that overall 75.4% of
the group receiving the new method is barely different from the mean for the the teachers are Non-senior class teachers: this descriptive statistic is
group receiving the standard method. If this is the case, then Teaching Method has uncontextualised. But if we are interested in gender inequalities, it is of interest
no effect for girls but a substantial effect for boys. to find that 80.6% of Females are in the Non-senior group, compared to 68.3% of
Males. In other words, when we compare Males with Females we note that there is
In many contexts if interactions are present the second explanatory variable is said a heavier concentration of class teachers amongst Females than overall and
to moderate the effect of the first and is often termed a moderator variable (equivalently in this simple case) the concentration of senior staff is higher
(Baron and Kenny, 1986)8. amongst Males. There is therefore a clear relationship between Gender and
Seniority (now our response variable).
In this section, even with only basic examples in mind, we have begun to see how
quantitative analysis can uncover quite a variety of ways in which relationships can Though it might not be instantly recognisable as such, one way of thinking of Table
be more complex than the simple association between two variables which we first 2.8 is again as an examination of how one factor, Gender, might explain variability
observe. In the process, we have introduced the important concepts of control of the units over the two categories of Seniority. Whether the response variable is
variables, additive effects, interactions and confounding. These are some features continuous or categorical, the aim of our analysis is still to explain variability in
that can be built into our analysis when we are answering research questions, to values of the response, and the ideas are exactly the same: we introduce an
help us properly understand the relationship that has motivated them. explanatory variable that shows a pattern across its values in certain summary
statistic(s) for the response (in this case percentages showing the distribution of
units across categories of the response). This pattern is the explained variability,
C2.3.2 Variability of categorical responses while the difference in values of the response for units with the same value of the
explanatory variable is the unexplained variability. So here, the explained
The examples above have focused on continuously measured responses. Comparing variability consists of the pattern that 80.6% of Females are Non-senior teachers
response levels, studying relationships and analysing variability is no less an issue and 68.3% of Males are Non-senior teachers, whilst the unexplained variability
when a response is categorical. However, we can no longer conduct the discussion consists of the fact that not all Female teachers have the same value of Seniority
using means and standard deviations as summary statistics. (and not all Male teachers have the same value of Seniority): among Female
teachers there are still some who are Senior teachers and some who are Non-senior
Table 2.8 cross-tabulates some alternative descriptive summaries for the two teachers (and the same is true of Male teachers). .
binary variables Gender and Seniority for 625 secondary school teachers.
Table 2.9 is another hypothetical example, cross-classifying gender by (broad)
categories of degree subject of a group of undergraduate students. We hope to
examine the effect of gender (our explanatory variable) on the subject of degree-
note that our response now has more than two categories.
8
Baron, R.M. and Kenny, D.A. (1986) The moderator-mediator variable distinction in social
psychological research: Conceptual, strategic, and statistical considerations. Journal of Personality Table 2.9: Percentages for types of degree subject by gender: undergraduate students
and Social Psychology, 51: 1173-1182.

Centre for Multilevel Modelling, 2008 24 Centre for Multilevel Modelling, 2008 25
Module 2: Introduction to Quantitative Data Analysis Module 2: Introduction to Quantitative Data Analysis
C2.3.2 Variability of categorical responses C2.3.3 Where both response and explanatory variables are continuous

Female Male Total plotted means would lie.

% (No.) % (No.) % (No.) We can also be a bit more specific about the nature of the relationship between IQ
Science & 10.0 (22) 61.1 (110) 33.0 (132) and Reading- in other words, describe the pattern of the mean Reading score
Engineering across values of IQ. In the case of a categorical variable, being specific simply
Arts 59.1 (130) 9.4 (17) 36.8 (147) involved giving summary means of the response for each value. For reasons
explained above we cannot do this with continuous explanatory variables: there
are simply too many values. However, here it looks as though the pattern may be
Social Sciences 19.1 (42) 21.7 (39) 20.2 (81) adequately approximated by a straight line drawn through the scatter of data: the
line of plotted means that we imagined in the previous paragraph appears to be
Others 11.8 (26) 7.8 (14) 10.0 (40) more or less straight. In other words, it seems that if we increase IQ by one unit
along the horizontal axis we increase the average Reading score by the same
Total 100.0 (220) 100.0 (180) 100.0 (400) amount no matter where on the IQ scale we start. This unvarying amount
corresponds to the slope of the straight line which approximates the pattern. The
characteristics of this line (i.e. its slope and intercept) then become the summary
statistics which we use to specify the relationship.
Overall the spread over the categories is reasonably diverse, and in particular
roughly equal percentages of students study Science and Arts subjects. However
grouping by Gender reduces this spread (i.e. reduces the unexplained variability,
by explaining some of the variability) in a particular way. There is a heavy
concentration in Science and Engineering of males while females tend to opt for
Arts subjects.

C2.3.3 Where both response and explanatory variables are


continuous

Figure 2.5 is a scatter plot showing how IQ scores and Reading scores covary (vary
together) for a sample of 18 children9. The most obvious thing we can see when we
look at the plot is that there is a relationship between IQ score and Reading score:
those with higher IQ scores tend to have higher Reading scores. Just as with a
categorical explanatory variable, we can look at this as a pattern of the mean of
the response across values of the explanatory variable, and this will be the
explained variability. With continuous data there are usually not enough units
present in our dataset to be able to produce actual summary means of the
response for each and every possible value of the explanatory variable, as we did
where it was categorical. Nonetheless the graph gives us an impression of the path
these means trace as we move across the x-axis: in this case, that the higher the
IQ score the higher ‘on average’ will be the Reading score. We can imagine that if
there were many more children in the sample, so that there were many children
with each IQ score, then we could actually take the mean Reading score for each
IQ score and plot this. In reality, since we only have 18 children and in many cases
Figure 2.5: Reading and IQ measures for 18 children
there is only one child with a particular IQ score, we cannot do this directly, but
looking at the 18 points in the graph, we can see roughly where such a line of The next thing we might notice looking at the graph is that overall this group
exhibits quite a bit of variability in Reading. Values range from 50 to around 75
9
along the Reading axis, a range of about 25. However, when we focus on an
Source: Ferguson, G. A. and Takane, Y. (1989) Statistical Analysis in Psychology and Education.
examination of this variability for small ranges of IQ Score, it is much reduced (as
6th Edition. New York: McGraw–Hill (Chapter 8).

Centre for Multilevel Modelling, 2008 - 26 - Centre for Multilevel Modelling, 2008 27
Module 2: Introduction to Quantitative Data Analysis Module 2: Introduction to Quantitative Data Analysis
C2.3.3 Where both response and explanatory variables are continuous C2.3.3 Where both response and explanatory variables are continuous

can be seen if the spread of points in the 110 to 120 range of IQ score, for
example, is compared to the overall spread). The overall impression is that the To what extent can the variability in GCSE Performance over schools be explained
dispersion of Reading values is quite small around an upward trend in the general by Hours Worked by teachers? In other words, to what extent is the percentage of
level of Reading score, i.e. around the straight line pattern we noted above giving pupils getting these grades influenced by the hours worked by teachers? It seems
the relationship between the mean of Reading score and IQ score. This small from the diagram that there is very little effect and the two are unrelated: we
dispersion around the straight line is the unexplained variation after including IQ cannot see a relationship between Hours Worked and mean GCSE Performance, as
as an explanatory variable. It seems then that IQ is explaining much of the we could between IQ score and mean Reading score. There is no pattern- that is,
variation in Reading and there is quite a close relationship between the variables. no variability- of mean GCSE Performance across the values of Hours Worked. Thus
Hours Worked cannot explain any of the variability in GCSE Performance: and
We will look at this idea more formally in the context of statistical models, in this indeed looking at the graph there is an impression of as much variability in GCSE
and later modules; however, the essential ideas are evident here: we describe a Performance in the vertical direction if we hold Hours Worked constant (i.e. for
pattern (here, a linear relationship) which contains the explained variability, but particular values on the horizontal axis) as there is overall, confirming that we
we recognise that there is still unexplained variability in the response (as have not reduced the unexplained variability by the introduction of Hours Worked.
evidenced by the small residual scatter of points around the line). As we will see,
pattern and residual are the two key ingredients of many forms of statistical Based on what we see in the graph, we could report a negative finding, i.e. that
model. there is little empirical support for a hypothesised relationship. Such findings may
be as important in research as any positive ones. Usually, of course, we will go on
We might contrast this situation with the example of Figure 2.6, which is a scatter to investigate this by more formal analysis, such as will be introduced in later
plot of two variables measured on 17 schools as units in the Pathfinder study: GCSE modules, to show the apparent absence of any relationship, rather than making
Performance (percentage of pupils in the cohort getting 5 or more GCSEs at grades this claim purely on the basis of examination of the graph.
between A* and C), and Hours Worked by teaching staff.
C2.3.4 Patterns other than straight lines

In the previous section we saw one example of a pattern of the mean of a


continuous response across a continuous explanatory variable: a straight line with
a positive slope, like that shown in Figure 2.7(a) below. This is not the only
possible pattern that we can have in the case where both our explanatory and
response variables are continuous. The other panels in Figure 2.7 give some
examples of other patterns of relationship between a continuous explanatory
variable (X) and a continuous response (Y) that frequently arise in quantitative
research, while Figure 2.8 shows how data points varying around each of these
patterns might look. Panel (b) shows a linear trend with a negative slope so that in
general Y decreases in response to increases in X. Panel (c) is an extreme situation
where X appears to have no influence on Y at all. Responses vary around the same
constant value of the response whatever the value of X. Panels (d) and (e) are
examples of non-linear patterns. The first, usually called a logarithmic
relationship, shows that in general Y increases as X does but at an ever decreasing
rate. The last is an example of a quadratic relationship where Y starts off
increasing in response to increases in X but after a certain level there is a ‘turning
point’ and Y starts to decrease. We would also have a quadratic relationship if the
pattern was reflected, so that Y started off decreasing in response to increases in
X but after a certain minimum level increased in response to increases in X.

Plotting scatter diagrams as in the previous section is a very important part of


investigating what sort of pattern is appropriate. Once we have decided the
pattern, the equation for the relevant curve will form an important part of the
specification of our statistical model.
Figure 2.6: Average Weekly Hours worked by teachers and GCSE Performance in 2002 for
17 Secondary Schools in Pathfinder Study

Centre for Multilevel Modelling, 2008 28 Centre for Multilevel Modelling, 2008 29
Module 2: Introduction to Quantitative Data Analysis
Module 2: Introduction to Quantitative Data Analysis
C2.3.4 Patterns other than straight lines C2.3.5 Relationship between a continuous response and combinations of categorical & continuous
explanators: Progress of students in schools

C2.3.5 Relationship between a continuous response and


combinations of categorical and continuous explanators:
Progress of students in schools

In this section we explore some more complexities in patterns of relationship,


when we relate a continuous response to a categorical explanatory variable but
introduce another continuous variable as a control. We start by considering as a
response Achievement, consisting of scores on some test taken at the end of the
last year of secondary school by students from four schools: A, B, C and D. At this
stage we simply consider School as a nominal variable with four categories defined
over students, rather than as the higher level in a two-level hierarchical structure
(see Module 1 C 1.4 for a discussion of this distinction).

We might initially be interested in the question of whether school attended


influences Achievement. In the language of educational progress research this
would be known as a ‘raw’ outcome analysis. This initial analysis could be
undertaken in a similar way to our analysis of the relationship between Seniority
and Hours Worked in C 2.2.1: we could summarise by finding the mean
achievement and standard deviation for each school. The standard deviations
would give us an indication of the unexplained variability within schools. The
Figure 2.7 pattern of the effects of School would be summarised by the means as indicated by
the horizontal bars on the 0 - 100 scale in Figure 2.9. We might note the small gap
in achievement of 4 points between A and B and the larger gaps of 19 between B
and C and 18 between C and D. This is in line with previous research which
generally finds that some differences between schools are quite substantial.

Figure 2.8
Figure 2.9: Mean Achievement for students classified by School A: 84, B: 80, C: 61, D: 43

Centre for Multilevel Modelling, 2008 30 Centre for Multilevel Modelling, 2008 31
Module 2: Introduction to Quantitative Data Analysis Module 2: Introduction to Quantitative Data Analysis

C2.3.5 Relationship between a continuous response and combinations of categorical and continuous C2.3.5 Relationship between a continuous response and combinations of categorical & continuous
explantors: Progress of students in schools explanators: Progress of students in schools

However, further thought about the research issues might suggest that the largest
influence on students’ achievement could be their individual abilities. If students
from each school had different levels of ability then this might be responsible for
the school effects we observe, which would then be confounded with pupils’
ability. For example, some schools might select higher ability students while others
take less able children. School differences in mean Achievement may also reflect
differences in parental choice if, for example, parents of high ability children are
more proactive in sending them to the ‘best’ schools. Suppose we have
observations on a variable for each student which measured their prior ability
before they entered the school (for example, an exam taken at the end of primary
school). We might then introduce this into the analysis as a control. One pattern
we might observe is illustrated in Figure 2.10 (a), where Prior Ability is on the
horizontal axis, and the lines show the approximate relationships observed for
each school. Several pertinent features may be noted. First, this is a simple case
where the four lines are taken to have the same slope. This means that the effect
of students’ prior ability is taken to be the same for each school: a one unit change
in Prior Ability is associated with the same impact on Achievement no matter
which school we are looking at. Second, as indicated by the asterisks (*) marking
means in the figure, the higher the school mean Achievement, the higher the mean
level of Prior Ability of its intake10. Third, at any given level of Prior Ability, the
ordering of the schools on Achievement is the same as before we conditioned on
Prior Ability, as indicated by the vertical positioning of the lines11. However, lastly
and very importantly, differences between schools at each level of ability, as
indicated by the vertical distances between the lines, are now much reduced, and
may be negligible. Thus much of the effect of the schools that we initially
observed is due to the effect of the prior ability of students: the effects of School
and Prior Ability are confounded.12 This is why school effectiveness research
generally studies what are called ‘value added’ effects (i.e. the school effects
controlling for prior ability) rather than comparing raw outcomes.

10
Note that this need not be the case: with a slope for the lines identical to that in Figure 2.10 (a)
it would also be possible (if not as likely in this context) to have a situation where the higher the
school mean Achievement, the lower the school mean Prior Ability.
11
Note that it is possible, as we will see, to have a situation where the ordering of groups on the
response variable changes after conditioning on another variable.
12
An extreme case is where the lines for the schools coincide exactly. At each level of ability the
school outcomes are the same and the net effect of school is zero. This would happen if the school
Figure 2.10
means on Prior Ability were exactly linearly related to the school Achievement means, with the
slope of the line giving this relationship being equal to the slope of the school lines. Differences in It is also possible that when we control for a third variable, as we do here, the
mean raw Achievement between schools are possible but in this case they would be entirely due to
effects we initially observe are not just reduced but sometimes even entirely
changed in various ways. Figure 2.10(b) gives an instance of this: it shows a
differences in mean levels of prior ability.

Centre for Multilevel Modelling, 2008 32 Centre for Multilevel Modelling, 2008 33
Module 2: Introduction to Quantitative Data Analysis Module 2: Introduction to Quantitative Data Analysis

C2.3.5 Relationship between a continuous response and combinations of categorical & continuous C2.3.5 Relationship between a continuous response and combinations of categorical & continuous
explanators: Progress of students in schools explanators: Progress of students in schools

different possible scenario that we might find, instead of that in Figure 2.10 (a), be non-differentiated by schools, so that the lines have the same slope. It can be
for the schools of Figure 2.9 when we control for Prior Achievement. Now the line noted from the asterisks marking the means that, although achievement means are
for school B is higher than that for school A, whereas it was lower before we highly similar, the prior ability levels are quite different. Thus, although initially,
controlled for Prior Ability. Intuitively, what is happening in this case is that Prior on the Achievement levels alone, no school effects were noted, in fact real net
Ability for school B is much lower than we might expect from the small difference effects of schools on progress were being suppressed by the impact of school
in the later Achievement of schools A and B. When we control for this prior differences in prior ability that were in the opposite direction. Once these are
ability, the low prior ability of school B pupils sets their achievement in context. controlled for, it becomes apparent that at any given level of ability the net
They make more progress than students at school A in that for each level of Prior effects of the schools are quite considerable. Phenomena such as this are often
Ability students in school B achieve more on average than those in school A. The referred to as ‘true effects or relationships revealing themselves’ after introducing
line for school B is therefore at a higher level than that for school A. further variables as controls. Since in research it is very important that such
revelations take place, it should be recognised that when we have an apparent
lack of relationship at first sight, it is unwise to take this at face value, and we
would instead be well advised to see what effect including additional variables
has. This is one further reason why including control variables is a key part of
analysis strategies when we come to consider modelling frameworks in later
modules.

In C 2.3.1 we introduced the idea of interacting effects as opposed to two


variables being additive in their influence on a response variable. The schools
example in this section has so far dealt only with additive effects. This is because
we have only considered parallel lines for the schools, so for each level of Prior
Ability the differences between schools are the same, i.e. the net effect of School
is added to that of Prior Ability. However, in many real problems in educational
progress research this may be an oversimplification: the relationship between Prior
Ability and Achievement may differ across schools. We have what in educational
progress research is called ‘differential effectiveness’. The pattern could be as in
Figure 2.12, which shows how Prior Ability and School interact in their effects.
School A, for instance, has a much steeper slope so that the effect on Achievement
of a difference in Prior Ability is greater in that school than is the effect of the
same difference in any of the other schools. For school D, with the flattest slope,
the effect on Achievement of a difference in Prior Ability is less than the effect of
the same difference in any other school. Other consequences of this interaction
may be noted. Students of lower ability are expected to perform best in school D,
since the line for this school is the highest at the low end of the Prior Ability scale,
so that this school is the most effective for lower ability pupils. By contrast school
A is the most effective for students at the upper end of the ability range.

Figure 2.11

It is also entirely possible for what is referred to as a ‘suppression’ pattern of


effects to be realised when we control for a third variable. Sometimes when we
investigate a pair of variables, at first sight the relationship pattern may look
uninteresting. Continuing our example further, suppose that instead of the pattern
in Figure 2.9 the four school achievement means were A: 65.0, B: 65.0, C: 64.5, D:
64.0. These are much the same. It might be concluded that the school attended
had little effect on educational achievement. However, suppose we do as before
and contextualise by controlling for the Prior Ability variable. Suppose the
resulting pattern was as in Figure 2.11. We again take the effect of Prior Ability to

Centre for Multilevel Modelling, 2008 34 Centre for Multilevel Modelling, 2008 35
Module 2: Introduction to Quantitative Data Analysis Module 2: Introduction to Quantitative Data Analysis

C2.3.5 Relationship between a continuous response and combinations of categorical & continuous C2.3.5 Relationship between a continuous response and combinations of categorical & continuous
explanators: Progress of students in schools explanators: Progress of students in schools

Ability) can vary over those groups. Here we have treated School as a categorical
variable defined over students, but in later modules we will demonstrate another
way of examining school effects by instead taking schools as units of study at level
2 in a multilevel model (see Module 1 C 1.4 for a brief discussion of these
contrasting approaches). We will thus be able to investigate more complex
research questions concerning school effects and their interaction with other
explanatory variables, and we will be able to generalise our findings to a wider
population of schools.

Don’t forget to take the online quiz for this section!

From within the LEMMA learning environment


• Go down to the Lesson for Module 2: Introduction to Quantitative Data
Analysis
• Click "2.3 Some other examples of relationships and the role of explained
variability"
to open Lesson 2.3
• Click Q1 to open the first question

Figure 2.12 The interaction of the effects of School and Prior Ability on achievement:
differential school effectiveness.

The examples in this section should have emphasised that when we examine
relationships between variables it is of paramount importance to condition on
other variables: when this is done apparent relationships can reduce in strength,
sometimes to the point of practically vanishing; the direction of effects may be
reversed; and previously unnoticed or suppressed relationships may come to light.
Also the presence of interactions may mean that the pattern and/or strength of
the relationship between a pair of variables changes over differing values of other
variables.

Relationships between pairs of variables can thus change fundamentally when we


control for others, in ways that can lead to further, more focussed research
questions. In the next module, modelling will be introduced as a systematic
framework for taking into account many variables at once.

These examples also show quite clearly how an outcome (Achievement) can vary
over well defined groups of students (schools in this case). We have further seen
how the relationship between the outcome and an explanatory variable (Prior

Centre for Multilevel Modelling, 2008 36 Centre for Multilevel Modelling, 2008 37
Module 2: Introduction to Quantitative Data Analysis Module 2: Introduction to Quantitative Data Analysis
C2.4 Working towards the idea of forming a statistical model C2.4 Working towards the idea of forming a statistical model

For a situation like that in Table 2.7 the assumption of additivity would need to be
C2.4 Working towards the idea of a formal relaxed and we would specify a pattern of interaction in the model then evaluate
whether we do indeed have this type of effect.
statistical model
Often, instead of looking at tables like Tables 2.6 and 2.7, we just specify a model
In this module and the last we have mentioned statistical models quite a few times with additive effects and one with interaction effects, and compare the results of
as ways of performing analyses to investigate similar issues to our examples above, fitting them to see directly whether the assumption of additivity or the assumption
when we have more complex situations than those of the examples; and from of interaction is better in the situation that we have. This is partly because it can
Module 3 onwards it is models that form our framework for the quantitative often be hard to see relationships when looking at tables.
analysis of research questions. However, we have not so far given any concrete
examples of a statistical model. To do so would have involved some mathematical In Figure 2.5 the pattern of the effect of IQ on reading score was a straight line
technicalities, which might at this stage have distracted from exploring the through the data. This straight line characterisation will be the pattern part of a
fundamental ideas of relationships. We will begin to use mathematical language model for this situation. It is perhaps easy to see here where mathematical
and define particular models in the modules to come. It will then be seen that the formalism comes in, since we use equations of straight lines in this case13; and in
statistical models for relationships we introduce are essentially just mathematical fact for other patterns we will use other kinds of equation. In Module 3 we will
formalisations of the type of situations we have already considered. Models thus learn much more about how to describe even quite complicated patterns
become unifying frameworks in the sense that they are ways of putting together all mathematically, so that we can put them into the specification of a statistical
the diverse concepts that we have discussed so far (and more) into one analysis. model.
When we encounter more complex situations, this mathematical formalism will, in
addition to its main purpose of enabling the analysis to be performed, facilitate Although patterns are the first of the two key ingredients of models, the examples
summarising and understanding what is going on. of previous sections have also shown that patterns of relationships will not usually
explain all the variation in response variables. There will always be some element
In this section we aim to introduce some of the essence of what the statistical which is not explained by the pattern. In the previous examples we have called
models we will use contain, without being overly technical. In fact, many of the this the unexplained variability. In Tables 2.6 and 2.7 this was the variation in the
ingredients of statistical models have already been discussed in this module as response Hours Worked within the subgroups of the explanatory variable Seniority.
they are exemplified by the various descriptive statistical summaries we have used In Figure 2.5 it was the variation in the response around the line approximating the
when illustrating the nature of various kinds of relationships. relationship between IQ and Reading score. Although our aim in modelling will be
to reduce as much as possible such unexplained variation, by introducing more
If we review the examples, we will see that we have focused on the variability in explanatory variables or putting them together in more complex ways, we can
the values taken by different kinds of response variables. We divided this rarely eliminate it completely. We will thus never be able to perfectly predict all
variability into explained variability and unexplained variability when we values of the response variable from our explanatory variables.
introduced an explanatory variable. We have also noted that the patterns in the
mean or other summary statistic(s) constituting the explained variability arise from Ways of characterising unexplained or residual variation will then be the second
a relationship between the explanatory variable and the response variable. We key ingredient of statistical models. For instance, it is often assumed that the
have then continued this process by introducing further explanatory variables, and residual variance is the same for each value of the explanatory variables, e.g.
with each one the pattern we describe has become more complex and the amount within groups for a categorical variable- we made this assumption in the case of
of unexplained variability has been reduced further. This attempt to explain the Table 2.6, when we introduced variable A as another covariate to explain Hours
variability has been the goal of our analysis. Worked. This is the assumption of homoscedasticity which is often a feature of a
formal model. We are then effectively saying that the process by which the data
For example, Table 2.6 showed a pattern in which the four groups formed by the observations are generated was subject to the same level of variability across all
categories of variable A and Seniority had different levels on average of the values of the covariates: i.e. within each of the four subgroups for the example of
response, Hours Worked. Differences between means for the four subgroups Table 2.6. However, there may be other situations where this feature might be too
summarise the pattern. Such patterns will be essential ingredients of formal restrictive. We might then consider it more appropriate to regard some groups as
models. We commented that in this case the pattern appears to be one of additive having much larger variability than others. In C [Link] we had a simple example
effects. We might thus want to specify a model with additive effects and then use
the model to formally evaluate whether or not the explanatory variables are
working together in this way. 13
The straight line is of the form Reading= a + b × IQ score. Deciding on suitable values of a and b
from the data is part of the modelling exercise that will be introduced in Module 3.

Centre for Multilevel Modelling, 2008 38 Centre for Multilevel Modelling, 2008 39
Module 2: Introduction to Quantitative Data Analysis Module 2: Introduction to Quantitative Data Analysis
C2.4 Working towards the idea of forming a statistical model C2.4 Working towards the idea of forming a statistical model

where groups of males and females had the same average on a response, Monthly Modules 4 and 5, but note that no matter how complex the research question we
Expenditure on Magazines. In this case it would therefore be unnecessary to allow are trying to answer, the construction of the model that we use to address them
for a relationship between Gender and the response in our statistical model. will generally be achieved with specifications involving means, variances and
However the variances within these two groups, i.e. the residual variability, were covariances.
considerably different. This might be the most important feature, which we wish
to reflect in a model. Recall that the feature of differing variances is known as Much statistical investigation, particularly when we formalise statistical models,
heteroscedasticity. In this simple example we say the response is heteroscedastic will go much further than the examples we have discussed, but we have exposed
with respect to gender. most of the essentials. However complex such investigation or modelling, both
explaining variability in various ways (via patterns) and saying something about the
Similarly, in situations similar to Figure 2.5, we might desire to find an structure of the variability which cannot be explained (the residual variation) are
approximating straight line as the pattern part of our model. In Module 3 we will usually central. The technicalities of specifying models, which will be gradually
assume that the unexplained variability in the response about the line is the same introduced starting with regression models in the next module, will show how we
wherever we are on the explanatory variable X-axis. However, sometimes we may actually build in the concepts discussed above in order to do this.
alternatively allow heteroscedasticity and let the unexplained variability itself
depend in various ways on the explanatory variables. An example is where the
variance around the line is small for low values of X but increases as X increases. In
Module 5 we will give examples of modelling variance by characterising such
heteroscedastic residuals in these ways. Also in Module 5 we show that this
becomes as much an important feature of analysis as describing the pattern of
response means in relation to the values of the explanatory variables.

The structure of residual variation is in fact at the heart of multilevel modelling.


We will leave the technical details for later modules but, put simply, we are
interested in how unexplained variation can arise both from inherent variation
amongst a set of cases at level 1 and from variation amongst the higher level units
in which they are nested. Thus, for example, in examining the two-level structure
of employees (level 1) within legal firms (level 2) we should seek to characterise
both the between-employee (within-firm) and the average between-firm residual
variation in salaries.

One of the most basic questions we might be interested in is to see whether the
variation is greater at the employee level or at the firm level. If almost all the
variation is at the employee level, that implies that firms have similar levels of pay
to each other, but within each firm different employees have quite different
salaries. On the other hand, if almost all the variation is at the firm level, that
implies that some firms pay much more than others but within each firm all
employees receive roughly equal salaries. We will expect to find variances
somewhere between these two extremes, with some firms paying more than others
overall and also individuals within each firm receiving quite different salaries, so
that both the employee-level variance and the firm-level variance will be sizable.

We can also address more complicated research questions using multilevel


modelling. We can incorporate any of the concepts that can be built into single-
level models to the multilevel case - for example, all the ideas discussed above
concerning the inclusion of explanatory variables and confounding. But we can also
include in a multilevel model features that it is not possible to build into a single-
level model, thus opening up to us a wider range of research questions. We will
learn much more about what multilevel modelling can do when it is introduced in

Centre for Multilevel Modelling, 2008 40 Centre for Multilevel Modelling, 2008 41
Module 2: Introduction to Quantitative Data Analysis Module 2: Introduction to Quantitative Data Analysis
C2.5.1 Estimation C2.5 .3 Testing Hypotheses

C2.5 Comments on statistical inference; In Table 2.5 we might be interested in the difference in the population in the
means of Hours Worked of Senior and Non-senior staff. Such a population
uncertainty, estimation and hypothesis testing descriptor is known as a population parameter. Population parameters are
generally unknown and cannot be calculated directly from our sample (since we
We discussed in Module 1 issues surrounding generalisation to a population14 when have not sampled the whole population). What we do know are the realised values
we carry out quantitative analyses on a sample from that population. Statistical of sample statistics which we can use to make inferences.
approaches to this go under the general heading of statistical inference. Most
introductory courses will have concentrated on a range of inferential methods for a A sample statistic is any summary measure that can be calculated from data
variety of situations. We have assumed that there will have been previous (whether this tells us anything useful or not). For example the sum of all the
exposure to some of them. In this section we give a brief overview of some observations, the maximum value of a variable present in the sample data, the
principles which are key to understanding what inference is about. We will value of the variable for the eleventh unit sampled, or the value of the ninth unit
illustrate the ideas by reference to a simple inference problem relating to the multiplied by the value of the fifth unit and added to the minimum value present.
information on Hours Worked and Seniority presented in C 2.2.1. Specifically, we The first two examples are obviously more useful sample statistics than the last
will compare the mean Hours Worked for Senior and Non-senior staff, using the two. Among the most commonly used sample statistics are means, standard
sample data in Table 2.5 and consider what sort of generalisations we might make deviations and regression coefficients (which we will meet in Module 3).
about differences between these two groups in a wider population.
When used to make inferences, sample statistics are called estimators of
We assume an element of randomness in sample selection. There is a vast body of population parameters. The difference between sample means is an estimator of
statistical knowledge that may then be applied. However, as we discussed in the difference between population means. The particular value of the sample
Module 1, theoretical generalisation may require thought beyond such statistic that we calculate from our sample data and are using as an estimator is
considerations. Most inference methods will also require some additional called the estimate. So in the example (54.2 - 50.3) = 2.9 hours is the estimate of
assumptions to be made. For instance, it will be useful to assume that population the difference between means using this estimator.
variances of Hours Worked are the same for Senior and Non-senior staff. The
sample evidence in Table 2.5 shows that this assumption might not be Given the observations we have made about chance variation in the value of such
unreasonable. statistics due to sampling variability, it is natural to question the implications of
such estimation procedures. We expect that the estimates will not be exactly the
C2.5.1 Estimation same as the true values of the population parameters. But can we say anything
else? Fortunately we can usually elaborate on the quality of estimators. Often, for
In treating the data in Table 2.5 as a random sample, we must recognise that the example, we will be able to say that the average value of an estimator across a
values of the response, and thus of the descriptive statistics calculated from them, large number of potential sample selections will equal the corresponding
will not be the same as those which might have arisen ‘by chance’ if a different parameter value. Thus whatever the value of the sample statistic that arises on a
set of teachers had been selected. In this sense a sampling process implies a particular occasion, we are assured that its possible values are centred around the
certain degree of uncertainty about the actual outcome which will arise when it is parameter value and there is no systematic divergence from it. This property,
implemented: there will be sampling variability over the possible outcomes that which is much sought after in inferential methods, is called unbiasedness. If
could have arisen from the range of possible samples that might have been taken. instead the average value of our estimator is not equal to the corresponding
We can envisage that if we were to apply the same sample design again on parameter value, then we have bias. Whether or not we have bias will depend on
repeated occasions, different sample units would be selected each time due to what statistic we have chosen to use as an estimator for our parameter15, as well
chance. It is knowledge about this sampling variability, which statistical theory
gives us, that lies behind sound inferences.
15
Quite often there will be several different sample statistics which each appear to be a sensible
choice for an estimator of a particular parameter. However, the term ‘estimator’ need not apply
only to a sensible choice of sample statistic, which is likely to take on a value close to the true
parameter value. In fact, any statistic can be an estimator for any parameter, although in the vast
majority of cases it will be an extremely bad estimator. So even for parameters for which only one
14
As indicated in Module 1, when we talk about a population and a sample here, our remarks will sample statistic is a sensible choice, we can still talk about choosing this statistic from among all
equally well apply to a superpopulation and a population. possible estimators.

Centre for Multilevel Modelling, 2008 42 Centre for Multilevel Modelling, 2008 43
Module 2: Introduction to Quantitative Data Analysis Module 2: Introduction to Quantitative Data Analysis
C2.5.2 Confidence intervals C2.5.2 Confidence intervals

as the sample design, so that we will speak of a statistic as being a biased or formula to Table 2.5 we find that the standard error for the difference between
unbiased estimator16. the sample means for this data is 0.475 hours.

Even with an unbiased estimator, we note that the observed single value estimate We now need to use a further amazing, but well established, theoretical idea to
is subject to sampling error in the sense of being different by chance from the interpret this figure and give us a context for assessing the precision. The
population parameter value. If we can say something about the size of this error distribution of possible outcomes in this, and a wide range of other, contexts is
we can add further to our judgement of estimator quality. We cannot usually say close to normal, i.e. possible differences in the sample means are distributed
anything about the error of a particular estimate, for to do that would involve (almost) normally. The mean of this normal distribution will be the true population
knowing the value of the parameter and we are then in a circular argument. difference (because our estimator is unbiased) and its standard deviation is 0.475
However if we could measure the sampling variability of an estimator we would be (our calculation of the standard error). At the end of C [Link] we commented on
taking a step in the right direction. A natural measure, if we could find it, is the useful features of the normal distribution. These enable us to say that we have a
variance of the sampling distribution of all possible values of the estimator under 68% chance of drawing a sample for which the estimate will be within one standard
the repeated sampling we have imagined. This is called the sampling variance and error, here 0.475, of the true difference. If we wanted to be more confident, we
its square root, the standard deviation, is given the special name standard error. could say we have (approximately)20 a 95% chance of obtaining a sample statistic
The smaller is this standard error the greater the chance that the estimate that is within two standard errors, here 0.950, of the true difference.
calculated from the sample selected will be close to the corresponding parameter
value. Another term that is often used in the context of any discussion on the Standard errors (SEs), because of what they tell us about estimator precision, play
quality of designs and estimation procedures is precision. Precision is simply the such an important role in inference and the modelling of later modules that they
inverse17 of the standard error. If sampling variability is low then high precision, should routinely be reported alongside any sample estimates. In Module 3 we will
which is a desirable quality, follows. see how the standard errors of our parameter estimates are crucial when it comes
to interpreting the results.
It is an ‘amazing fact’18 that under the conditions we have discussed we are able to
derive a formula for the standard error of the difference between sample means as
an estimator for the difference between population means and then to use the C2.5.2 Confidence Intervals
sample data itself to evaluate this error, without the need to take multiple
samples and calculate the value of the estimator for each19. On applying this In the previous section we used the word ‘confident’ in saying something about the
closeness of our estimate to the true parameter value. Measuring confidence is
part of another frequently used way of reporting results. Rather than quoting the
16
It may seem that it is impossible to assess whether or not an estimator is biased, since in order to value of the estimate and a measure of precision, we specify a range of values for
the parameter together with a measure of our confidence that the true value lies
tell whether it is centred around the parameter value we need to know the parameter value, and it
in that range. The result is a confidence interval. We can surround an estimate,
is precisely because we do not know the parameter value that we need to use the estimator. for instance, by an interval consisting of 2 standard errors from the estimate in
However, fortunately, it turns out that we can in fact use statistical theory to tell us whether or each direction. Properties of the normal distribution then enable us to say that,
out of all possible samples, for 95% the confidence interval calculated using that
not an estimator is biased without needing to know the parameter value.
17
sample data will cover the true value. In our example, the 95% confidence interval
In other words, 1÷ standard error is from 2.9 – (2 × 0.475) = 1.95 up to 2.9 + (2 × 0.475) = 3.85, expressed as the
18
The book by Diamond and Jeffries referenced in the Resources section discusses three ‘Amazing range {1.95, 3.85}. We can be more confident at the cost of widening the range.
facts’ and an ‘Amazing Theorem’ (the Central Limit Theorem) whose value in inference go well The appropriate multiplier of the standard error to yield a 99% confidence interval
is 2.58, yielding the approximate range (1.67, 4.13).
beyond the basic example we are considering.
19
Under the assumption of equal population variance, value σ 2 , in the two groups, the standard
1 1
error is σ + , where n1 and n2 are the two group sample sizes. σ 2 itself is estimated by
n1 n2
taking a weighted average of the two sample variances. The basic texts listed in the Resources
section may be referred to for technical details. It might be noted that important determinants of
20
sampling error are the variability in the population and the sample size. The larger the sample size We say approximately since a more accurate multiplier of the standard error to give a 95%
(or the smaller the population variability) the more precise is the estimator. chance is 1.96 rather than 2.

Centre for Multilevel Modelling, 2008 44 Centre for Multilevel Modelling, 2008 45
Module 2: Introduction to Quantitative Data Analysis
Module 2: Introduction to Quantitative Data Analysis
C2.5.2 Confidence intervals C2.5.3 Testing hypotheses

this value of 2 is always the same: it is not special to our sample data since in
C2.5.3 Testing hypotheses using z-values we have removed the dependence on the standard deviation. Other
significance levels used in practice are 1% and 0.1% corresponding to absolute z-
Besides questions asking what the true value of a population parameter is, another values larger than 2.58 and 3.33 respectively. The z-value in our example is
common type of inference question takes the form of theoretical propositions certainly bigger than 3.33 and thus significant well beyond the 0.1% level, the most
about population parameters whose plausibility it is desired to evaluate using stringent of the conventional levels used. Compare this with what we found using
sample evidence. We might postulate, for instance, that there is no population the alternative method in the previous paragraph: the p-value was 4.7 × 10-10,
difference between the mean of Hours Worked by Senior and Non-senior teachers. which is clearly much smaller than 0.001, and thus the two methods agree as we
We then wish to see whether the sample result we observe is compatible with this would expect: they are indeed equivalent.
situation, i.e. whether it is likely to have arisen from such a population.
Alternatively, is the sample difference in means large enough to rule out this In later modules we will frequently be dealing with estimators such as the slope
possibility and lead us conclude that there is a real difference in the population coefficients in regression models. Many of the estimators we encounter will have
means? These sorts of questions are the basis for classical hypothesis testing, or approximately normal sampling distributions. Again, a common question is whether
significance testing, as it is often known. parameters are significantly different from zero. We can use the same rough and
ready rule of thumb in exploration of results to see whether anything interesting is
What are we asking here? We are asking what the chances are (what the going on: if the estimate is more than 2 times the standard error, the result is
probability is) of having a sample difference of means more extreme (larger in significant beyond the 5% level.
magnitude) than 2.9, our observed difference, if the true difference between the
population means is zero. If the population difference is indeed zero, the mean of In the discussion so far it might appear that confidence intervals and hypothesis
the (normal) sampling distribution will be zero with a standard deviation of 0.475. testing are two separable types of inference. However, they are essentially highly
To make things easier, we divide the sample statistic by the standard error to get interconnected. The information on which inference is based in both cases is the
a new statistic, z. The sampling distribution for z is, like that for the difference value of the estimator and its standard error together with some knowledge, from
between the group means, normal with a mean of 0, but, unlike that for the mean statistical theory, about the form of the sampling distribution. In addition to the
difference, it has a standard deviation of 1. The sampling distribution for z is usual interpretation, the 95% confidence interval for the difference of population
therefore the standard normal distribution. We can now compute the required means gives us a range of values which, were they to be hypothesised as the value
probabilities easily since the results for the standard normal distribution are easily for that difference, would not be questioned at 5% level of significance.
available in statistical distribution tables (or in any statistical analysis software Hypothesised values outside the confidence range would be rejected at that level.
package). Here z = 2.9/0.475 = 6.12. We can look up the probability that the In this and many other cases it is zero that is of interest as a hypothesised value so
absolute value21 of z is larger than 6.12. This is called the p-value of the sample that an implicit test at the appropriate level is achieved by examining whether the
result. It is in fact extremely low at 4.7 × 10-10. What we are saying is that if the relevant confidence interval contains zero. The 99% range (1.67, 4.13) for the two
true difference between the population means was zero, we would have about 5 means would thus certainly lead to rejection at the 1% level of the null hypothesis
chances in 10 billion of observing a sample result as extreme as the one we have of zero population mean difference, just as we rejected the hypothesis at the 1%
here. The p-value sums up the evidence against the hypothesis of no difference. level based on its p-value of less than 0.01, or on its z-value of more than 2.58.
We might regard this evidence as very conclusive and reject the hypothesis. In a
classical framework we say that we reject the null hypothesis (H0) of no difference Don’t forget to take the online quiz for this section!
between the means in favour of the alternative that there is a difference.
Sometimes also we say that the sample result is significantly different from zero.
From within the LEMMA learning environment
An alternative and more conventional way of carrying out this test procedure is to • Go down to the Lesson for Module 2: Introduction to Quantitative Data
refer the sample result to what are called significance levels. Again properties of Analysis
the normal distribution are used. For instance a z-value greater than 2 will mean a
• Click "2.5 Comments on Statistical Inference; uncertainty, estimation and
result that is statistically significant beyond the 5% level in the sense that the
hypothesis testing"
probability of such a result under the null hypothesis is less than 0.05. Note that
to open Lesson 2.5
• Click Q1 to open the first question
21
We say absolute value since we would regard a sample result of z < -6.12 as equally extreme.
The type of hypotheses we are discussing here are two tailed in the sense that population
differences in either direction are of interest. Correspondingly tests are two sided. One sided tests
Don’t forget to take the online quizzes for this module! (see page 2
for details of how to find the quizzes)
which we do not consider here are covered in pp 148-151 of the book by Diamond and Jeffries.

Centre for Multilevel Modelling, 2008 46 Centre for Multilevel Modelling, 2008 47

You might also like