Inferential Testing in Psychological Research
Inferential Testing in Psychological Research
Name:
_
Inferential testing
_______________________
Class:
_
_______________________
Date:
_
Comments:
Page 1 of 116
Q1.
A researcher carried out an overt observation study of social learning. For one week the
helping behaviour of children in a playgroup was recorded. All the children then saw a
short film in which a child was praised for tidying up toys. For the following week the
helping behaviour of the same children in the playgroup was recorded.
(a) Which of the following statements is the best description of an overt observation
study?
(b) Briefly discuss one way in which a covert observation of children might be more
beneficial than an overt observation.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(3)
(c) At the end of the observation study the researcher used a sign test to see if the
behaviour of the children was more helpful, less helpful or the same after seeing the
film than it was before they had seen the film.
Explain why the researcher decided the sign test would be an appropriate statistical
test to use on the data from this study.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
Page 2 of 116
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(4)
(Total 8 marks)
Q2.
A researcher studying depression wanted to see whether or not there was a relationship
between level of self-esteem and negative schema score. She constructed two
questionnaires and asked ten people who had been diagnosed with depression to
complete them.
One questionnaire measured the participant’s level of self-esteem. A low score (out of 50)
indicated low self-esteem.
The other questionnaire measured whether the participant showed evidence of negative
schema. A low score (out of 50) indicated evidence of negative schema. The two sets of
results for each participant are shown in the table below.
Table 1 - Self-esteem score and negative schema score for each patient
Participant 1 2 3 4 5 6 7 8 9 10
Self-esteem
8 9 9 11 13 17 18 18 20 22
score
Negative
schema 11 15 13 18 12 14 20 16 17 19
score
A Cognitive
B Emotional
C Behavioural
(1)
(b) Draw a suitable graphical display to represent the data in Table 1. Label your graph
appropriately.
Title:_______________________________________________________________
Page 3 of 116
(4)
The researcher analysed the data in Table 1 using a Spearman’s rho statistical test.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(2)
(d) Estimate the correlation coefficient most likely to result from analysis of the data in
Table 1. Shade one box only.
+0.95
+0.70
+0.30
Page 4 of 116
+0.15
(1)
(Total 8 marks)
Q3.
A psychologist decided to conduct an experiment to investigate the effect of watching
horror films before going to bed.
The 50 students were randomly split into two groups. Group 1 watched a horror film
before going to bed each night for the first week then a romantic comedy before going to
bed each night for the second week. Group 2 watched the romantic comedy in the first
week and the horror film in the second week.
When the students woke up each morning, each student received a text message that
asked if they had had a nightmare during the night. They could respond ‘yes’ or ‘no’.
(a) Write a brief consent form that would have been suitable for use in this experiment.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(6)
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
Page 5 of 116
___________________________________________________________________
___________________________________________________________________
(3)
Explain why it was important to use a repeated measures design in this case.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(2)
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(2)
(e) Explain how the psychologist could have randomly split the sample of 50 students
into the two groups.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(3)
The psychologist collated the responses of all the participants over the two-week period
and calculated the mean and standard deviation for each condition.
Mean number of nightmares reported and the standard deviation for each condition
Mean number of
nightmares in 7 Standard deviation
days
Page 6 of 116
Romantic
0.30 0.61
comedies
(f) What do the mean and standard deviation values in the table above suggest about
the effect of the type of film watched on the occurrence of nightmares? Justify your
answer.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(4)
(g) The psychologist found that the difference in the number of nightmares reported in
the two conditions was significant at p<0.05.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(2)
(h) The psychologist was concerned about the validity of the experiment.
Suggest one possible modification to the design of the experiment and explain how
this might improve validity.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(3)
(Total 25 marks)
Page 7 of 116
Q4.
In an experiment, researchers arranged for participants to complete a very personal and
embarrassing questionnaire in a room with other people. Each participant was tested
individually. The other people were confederates of the experimenter.
In condition 2: the confederates refused to complete the questionnaire and asked to leave
the experiment.
The researchers recorded the number of participants who completed the questionnaire in
each condition.
(a) Identify the type of data in this experiment. Explain your answer.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(2)
(b) Using your knowledge of social influence, explain the likely outcome of this
experiment.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(3)
(c) For this study, the researchers had to use different participants in each condition and
this could have affected the results.
Outline one way in which the researchers could have addressed this issue.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
Page 8 of 116
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(4)
(d) In order to analyse the difference in the number of participants who completed the
questionnaire in each condition, the researchers used a chi-squared test.
Apart from reference to the level of measurement, give two reasons why the
researchers used the chi-squared test.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(2)
(e) The calculated value of chi-squared in the experiment described above is 3.97
Level of significance
The calculated value of chi-squared should be equal to or greater than the critical
value to be statistically significant.
With reference to the critical values in the table above, explain whether or not the
calculated value of chi-squared is significant at the 5% level.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(2)
(Total 13 marks)
Q5.
A psychologist wanted to test whether listening to music improves running performance.
The psychologist conducted a study using 10 volunteers from a local gym. The
psychologist used a repeated measures design. Half of the participants were assigned to
condition A (without music) and half to condition B (with music).
Page 9 of 116
All participants were asked to run 400 metres as fast as they could on a treadmill in the
psychology department. All participants were given standardised instructions. All
participants wore headphones in both conditions. The psychologist recorded their running
times in seconds. The participants returned to the psychology department the following
week and repeated the test in the other condition.
(a) Identify the type of experiment used in this study. Shade one box only.
A Laboratory
B Natural
C Quasi
D Research
(1)
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(2)
Table 1: Mean number of seconds taken to complete the 400m run and the standard
deviation for both conditions
Condition A Condition B
(without music) (with music)
Explain why a histogram would not be an appropriate way of displaying the means
shown in Table 1.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(2)
(d) Name a more appropriate graph to display the means shown in Table 1. Suggest
appropriate X (horizontal) and Y (vertical) axis labels for your graph choice.
Page 10 of 116
___________________________________________________________________
X axis label:_________________________________________________________
___________________________________________________________________
Y axis label:_________________________________________________________
(3)
(e) What do the mean and standard deviation values in Table 1 suggest about the
participants’ performances with and without music? Justify your answer.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(4)
(f) Calculate the percentage decrease in the mean time it took participants to run 400
metres when listening to music. Show your workings. Give your answer to three
significant figures.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(4)
(g) The researcher used a directional hypothesis and analysed the data using a related
t-test. The calculated value of t where degrees of freedom (df) = 9 was 1.4377. He
decided to use the 5% level of significance.
Page 11 of 116
for a one-tailed test
df = 1 6.314 12.706
2 2.920 4.303
3 2.353 3.182
4 2.132 2.776
5 2.015 2.571
6 1.943 2.447
7 1.895 2.365
8 1.860 2.306
9 1.833 2.262
10 1.812 2.228
Calculated value of t must be equal to or greater than the critical value in this table for significance to
be shown.
Give three reasons why the researcher used a related t-test in this study and, using
Table 2, explain whether or not the results are significant.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(5)
(h) What is meant by a Type II error? Explain why psychologists normally use the 5%
level of significance in their research.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
Page 12 of 116
___________________________________________________________________
___________________________________________________________________
(3)
(i) Identify one extraneous variable that could have affected the results of this study.
Suggest why it would have been important to control this extraneous variable and
how it could have been controlled in this study.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(3)
(j) The report was submitted for peer review and a number of recommendations were
advised.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(6)
(Total 33 marks)
Q6.
Read the item and then answer the questions that follow.
Twenty primary school teachers were sent by their individual head teachers to attend a
training course in classroom behaviour management run by educational psychologists at a
local university. Before the training course, and again after training, the teachers were
asked to say how confident they were in managing difficult classroom behaviour.
Page 13 of 116
The researchers compared the before and after answers to see how many teachers rated
their confidence as ‘better’, ‘worse’, or ‘the same’ as it had been at the start of the course.
(a) Which of A, B, C or D best describes this study? Shade one box only.
A laboratory experiment
B pilot experiment
C natural experiment
D controlled experiment
(1)
(b) What fraction of the teachers thought that their confidence was better after the
course? Show your workings.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(2)
(c) What might the researchers conclude about the training course on the basis of the
data in the table? Explain your answer.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(2)
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(2)
(e) Which experimental design is being used in this study and why would it be an
Page 14 of 116
appropriate design in this case?
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(3)
(f) The psychologists conducting the training decided to use the Sign Test to see
whether there was a significant difference in confidence in managing difficult
classroom behaviour before and after the course.
Give the calculated value of S in this study and explain how you arrived at this
figure.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(3)
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(2)
(h) Following the training course, one of the researchers carried out an overt classroom
observation of each teacher’s primary school class. The researcher wanted to
record the frequency of difficult classroom behaviours shown by the pupils during a
normal lesson.
Suggest two behavioural categories that the researcher could record during his
observation.
Page 15 of 116
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(2)
(i) Design a tally chart/record sheet the researcher could use to record his
observations.
(3)
(j) Identify one problem that might have occurred during this observation and explain
how the observation would be improved by addressing this problem.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(4)
(Total 24 marks)
Q7.
Page 16 of 116
Read the item and then answer the questions that follow.
Mean number of verbal errors and standard deviations for both conditions
Condition A Condition B
(believed audience (believed audience
of 5 listeners) of 100 listeners)
Standard
1.30 3.54
deviation
(a) What conclusions might the psychologist draw from the data in the table? Refer to
the means and standard deviations in your answer.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(6)
(b) Read the item and then answer the question that follows.
Explain how using the standard deviation rather than the range in this situation,
would improve the study.
___________________________________________________________________
___________________________________________________________________
Page 17 of 116
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(3)
(c) Name an appropriate statistical test that could be used to analyse the number of
verbal errors in the table above. Explain why the test you have chosen would be a
suitable test in this case.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(4)
(d) The psychologist found the results were significant at p<0.05. What is meant by ‘the
results were significant at p<0.05’?
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(2)
(e) Briefly explain one method the psychologist could use to check the validity of the
data she collected in this study.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(2)
(Total 17 marks)
Q8.
Page 18 of 116
A researcher wanted to see whether cognitive behaviour therapy was an effective
treatment for depression. Twenty depressed patients who had all recently completed a
course of cognitive behaviour therapy were involved in the investigation. From their
employment records, the researcher kept a record of the number of absences from work
each patient had in the year following their treatment. This was compared with the number
of absences from work each patient had in the year prior to their treatment.
Those patients who had fewer absences from work in the year following their treatment
than in the year prior to their treatment were classified as ‘improved’ (+). Those patients
who had more absences were classified as ‘deteriorated’ (-). Those patients who had the
same number of absences were classified as ‘neither’ (0).
Table 1
1 +
2 0
3 –
4 +
5 +
6 +
7 –
8 –
9 0
10 +
11 –
12 +
13 +
14 +
15 +
16 –
17 +
18 +
19 +
Page 19 of 116
20 0
The researcher decided to use the sign test to analyse the data.
(a) Explain two factors that the researcher had to take into account when deciding to
use the sign test. Refer to the investigation above in your answer.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(4)
(b) Calculate the sign test value of s for the data in Table 1. Explain how you reached
your answer.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(2)
16 2 2 3 4
17 2 3 4 4
18 3 3 4 5
Page 20 of 116
For significance, the value of the less frequent sign is equal to, or less
than, the value of the table.
(c) With reference to the critical values in Table 2, explain whether or not the value of s
that you calculated in response to question (b) is significant at the 0.05 level for a
two tailed test.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(2)
In what ways would the use of primary data have improved this investigation?
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(3)
(e) Outline the implications of psychological research for the economy. Refer to the
investigation above in your answer.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(5)
(Total 16 marks)
Page 21 of 116
Q9.
Read the item and then answer the questions that follow.
(a) Should the hypothesis for this research be directional or non-directional? Explain
your answer.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(2)
(b) Before the observation could begin, the students needed to operationalise the
behaviour category ‘riding a bike with care’.
Explain what is meant by operationalisation and suggest two ways in which ‘riding a
bike with care’ could have been operationalised.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(4)
(c) The students thought that having a dog on a lead was a useful measure of
considerate behaviour because it had face validity. Explain what is meant by face
validity in this context.
Page 22 of 116
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(3)
(d) Identify and briefly outline two other types of validity in psychological research.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(4)
(e) Identify the behaviour sampling method used by the students. Shade one box only.
A Time sampling
B Pair sampling
C Event sampling
D Target sampling
(1)
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(3)
Page 23 of 116
(g) The data for considerate behaviours is shown in the Table 1.
Table 1
Considerate behaviours
Riding bike
Litter in bin Dog on lead
with care
Greensville 23 23 10
Browntonn 10 17 9
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(3)
Explain why the Chi-square test was an appropriate test to use in this case.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(3)
(i) In order to interpret the results of the Chi-square test the students first needed to
work out the degrees of freedom. They used the following formula.
Calculate the degrees of freedom for the data in Table 1. Show your workings.
Page 24 of 116
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(2)
(j) The calculated value of Chi-square was 6.20. Referring to the Table 2 below, state
whether or not the result of the Chi-square test is significant at the 0.05 level of
significance. Justify your answer.
Table 2
To be significant at the level shown the calculated value of Chi Square must be equal to or greater than
the critical/table value
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(3)
(k) In the discussion section of their report of the investigation the students wanted to
further discuss their results in relation to levels of significance.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
Page 25 of 116
___________________________________________________________________
(4)
(l) As a follow-up to their observation the students decided to interview some of their
peers about inconsiderate behaviours in their 6th Form Centre. The interviews were
recorded.
Explain how the students could develop their interview findings by carrying out a
content analysis and why content analysis would be appropriate in this case.
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
___________________________________________________________________
(3)
(m) Suggest one inconsiderate behaviour that the students might focus on in their
content analysis.
___________________________________________________________________
___________________________________________________________________
(1)
(Total 36 marks)
Q10.
Read the item and then answer the questions that follow.
Males Females
4 5
6 6
2 6
2 7
3 2
3 4
3 5
8 9
1 5
1 2
Median score 3 Median score 5
Page 26 of 116
(a) Explain why the data in the table is primary data and not secondary data.
(2)
(c) The researcher decided to extend the study by using an inferential test to see if
there was a significant difference between the two sets of scores.
Suggest an appropriate inferential test which the researcher could use. Justify your
choice.
(4)
(Total 9 marks)
Q11.
Read the item and then answer the questions that follow.
In a study of gender schema a researcher studied the way in which boys and
girls understood gender. An overall understanding score was calculated on the
basis of answers to a questionnaire. A high score indicated a very fixed
understanding and a low score indicated a flexible understanding.
The scores are shown in the table below.
Girls Boys
4 5
6 6
2 6
2 7
3 2
3 4
3 5
8 9
1 5
1 2
Median score 3 Median score 5
(a) Explain why the data in the table is primary data and not secondary data.
(2)
(c) The researcher decided to extend the study by using an inferential test to see if
there was a significant difference between the two sets of scores.
Suggest an appropriate inferential test which the researcher could use. Justify your
choice.
(4)
(Total 9 marks)
Q12.
Read the item and then answer the questions that follow.
Page 27 of 116
In a study of social cognition a researcher studied perspective-taking in children
aged 5 years and 9 years. An overall perspective-taking score was calculated
on the basis of answers to a questionnaire. A high score indicated good
perspective taking and a low score indicated poor perspective-taking.
The scores are shown in the table below.
5-year-olds 9-year-olds
4 5
6 6
2 6
2 7
3 2
3 4
3 5
8 9
1 5
1 2
Median score 3 Median score 5
(a) Explain why the data in the table is primary data and not secondary data.
(2)
(c) The researcher decided to extend the study by using an inferential test to see if
there was a significant difference between the two sets of scores.
Suggest an appropriate inferential test which the researcher could use. Justify your
choice.
(4)
(Total 9 marks)
Q13.
Read the item and then answer the questions that follow.
(a) Explain one advantage of using a repeated measures design in this study.
Page 28 of 116
(2)
The psychologist decides to use a sign test to see if his data are significant.
What is the calculated value of the sign test statistic ‘S’? Explain your answer.
(2)
(c) Look at the table of critical values of ‘S’ below and then answer the question that
follows.
To be significant, the calculated/observed value must be equal to or less than the critical/table value.
Using the table of critical values of ‘S’ above, state whether the findings of the study
are significant at p < 0.05. Explain your answer.
(2)
(Total 6 marks)
Q14.
Read the item and then answer the questions that follow.
(a) Explain one advantage of using a repeated measures design in this study.
(2)
Page 29 of 116
(b) The psychologist obtained the following results:
The psychologist decides to use a sign test to see if her data are significant.
What is the calculated value of the sign test statistic ‘S’? Explain your answer.
(2)
(c) Look at the table of critical values of ‘S’ below and then answer the question that
follows.
To be significant, the calculated/observed value must be equal to or less than the critical/table value.
Using the table of critical values of ‘S’ above, state whether the findings of the study
are significant at p < 0.05. Explain your answer.
(2)
(Total 6 marks)
Q15.
Read the item and then answer the questions that follow.
(a) Explain one advantage of using a repeated measures design in this study.
(2)
Page 30 of 116
• For two of the volunteers the number of cigarettes smoked stayed the same
The psychologist decides to use a sign test to see if the data are significant.
What is the calculated value of the sign test statistic ‘S’? Explain your answer.
(2)
(c) Look at the table of critical values of ‘S’ below and then answer the question that
follows.
To be significant, the calculated/observed value must be equal to or less than the critical/table value.
Using the table of critical values of ‘S’ above, state whether the findings of the study
are significant at p < 0.05. Explain your answer.
(2)
(Total 6 marks)
Q16.
In an observational study, 100 cars were fitted with video cameras to record the driver’s
behaviour. Two psychologists used content analysis to analyse the data from the films.
They found that 75% of accidents involved a lack of attention by the driver. The most
common distractions were using a hands-free phone or talking to a passenger. Other
distractions included looking at the scenery, smoking, eating, personal grooming and
trying to reach something within the car.
(b) Explain how the psychologists might have carried out content analysis to analyse
the film clips of driver behaviour.
(4)
(c) Explain how the two psychologists might have assessed the reliability of their
content analysis.
The psychologists then designed an experiment to test the effects of using a hands-
free phone on drivers’ attention. They recruited a sample of 30 experienced police
drivers and asked them to take part in two computer-simulated driving tests. Both
tests involved watching a three-minute film of a road. Participants were instructed to
click the mouse as quickly as possible, when a potential hazard (such as a car
pulling out ahead) was spotted.
Page 31 of 116
• Test A, whilst chatting with one of the psychologists on a hands-free phone
The order in which they completed the computer tests was counterbalanced.
(3)
(d) Explain why the psychologists chose to use a repeated measures design in this
experiment.
(3)
(e) Identify one possible extraneous variable in this experiment. Explain how this
variable may have influenced the results of this experiment.
(3)
(f) Explain one or more ethical issues that the psychologists should have considered in
this experiment.
(4)
(g) Write a set of standardised instructions that would be suitable to read out to
participants, before they carry out Test A, chatting on a hands-free phone.
The mean scores for each of these measures is shown in the table below.
Table to show the mean number of hazards detected and mean reaction times
in seconds for Test A and Test B
Number of hazards
26.0 23.0
detected
Reaction time in
0.45 0.27
seconds
The psychologists then used an inferential statistical test to assess whether there
was a difference in the two conditions.
(5)
(h) Identify an appropriate statistical test to analyse the difference in the number of
hazards detected in the two conditions of this experiment. Explain why this test of
difference would be appropriate.
They found no significant difference in the number of hazards detected (p > 0.05),
but there was a significant difference in reaction times (p . 0.01).
(3)
(i) Explain why the psychologists did not think that they had made a Type 1 error in
relation to the difference in reaction times.
Page 32 of 116
(2)
(j) Replication is one feature of the scientific method. The psychologists decided to
replicate this experiment using a larger sample of 250 inexperienced drivers.
Q17.
A student teacher was interested in the relationship between empathy (consideration and
feelings for others) and the time spent reading fiction. She decided to investigate whether
or not such a relationship was present in children.
The student teacher designed her own questionnaire to measure empathy in 8-year-old
children. The higher the score achieved, the greater the empathy. Twenty children, all from
one school, took part. Each child completed the questionnaire individually.
The student teacher designed another questionnaire to measure ‘time spent reading
fiction’. Each child was given this questionnaire to take home and complete with his or her
parents over a four-week period. ‘Time spent reading fiction’ included the time spent by
parents reading to the child as well as the time the child spent reading independently.
Using the responses to this questionnaire, the student teacher calculated how much time
per week, on average, each child spent reading fiction.
(a) Outline the relationship between empathy and the average number of hours spent
reading fiction per week shown in the graph above.
(1)
Page 33 of 116
(b) Name an appropriate test to determine whether or not there is a significant
relationship between the two variables in the graph above. Justify your answer with
reference to levels of measurement.
(2)
(c) Outline one way in which the student teacher could have assessed the validity of
the empathy questionnaire.
(2)
(d) Apart from the issue of validity, identify and briefly explain one methodological
limitation of the study.
(2)
(e) Explain why it was appropriate for the student teacher to use a correlation study
rather than an experiment.
(3)
(f) The student teacher noticed that some students on her course commented that they
were better able to recall information if they could read the information rather than
listen to it in lectures.
‘People who are given written information will recall more than people who hear
information in spoken form.’
In your answer, you should refer to the following and justify your design decisions:
• the sample
• relevant materials
Q18.
Some studies have suggested that there may be a relationship between intelligence and
happiness. To investigate this claim, a psychologist used a standardised test to measure
intelligence in a sample of 30 children aged 11 years, who were chosen from a local
secondary school. He also asked the children to complete a self-report questionnaire
designed to measure happiness. The score from the intelligence test was correlated with
the score from the happiness questionnaire. The psychologist used a Spearman’s rho test
to analyse the data. He found that the correlation between intelligence and happiness at
age 11 was +0.42.
(b) Identify an alternative method which could have been used to collect data about
Page 34 of 116
happiness in this study. Explain why this method might be better than using a
questionnaire.
(4)
(c) A Spearman’s rho test was used to analyse the data. Give two reasons why this test
was used.
(2)
0.10 0.05
0.05 0.025
29 0.312 0.368
30 0.306 0.362
31 0.301 0.356
Calculated rs must equal or exceed the table (critical) value for significance at the level shown.
(d) The psychologist used a non-directional hypothesis. Using the table above, state
whether or not the correlation between intelligence and happiness at age 11 (+0.42)
was significant. Explain your answer.
(3)
(e) Five years later, the same young people were asked to complete the intelligence
test and the happiness questionnaire for a second time. This time the correlation
was –0.29.
With reference to both correlation scores, outline what these findings seem to show
about the link between intelligence and happiness.
(4)
(Total 15 marks)
Q19.
A maths teacher wondered whether there was a relationship between mathematical ability
and musical ability. She decided to test this out on the GCSE students in the school. From
210 students, she randomly selected 10 and gave each of them two tests. She used part
of a GCSE exam paper to test their mathematical ability. The higher the mark, the better
the mathematical ability. She could not find a musical ability test so she devised her own.
She asked each student to sing a song of their choice. She then rated their performance
on a scale of 1–10, where 1 is completely tuneless and 10 is in perfect tune.
(b) Why might the measure of musical ability used by the teacher lack validity?
(3)
Page 35 of 116
(c) Explain how the teacher could have checked the reliability of the
mathematical ability test.
(3)
(d) Explain why the teacher chose to use a random sample in this study.
Mathematical ability test scores and musical ability ratings for 10 students
1 10 10
2 2 9
3 9 3
4 6 6
5 3 9
6 10 2
7 2 1
8 1 8
9 8 4
10 4 7
(2)
(e) In your answer book, sketch a graph to show the data in the table above.
Give the graph an appropriate title and label the axes.
(3)
(f) Discuss what the data in the table above and the graph that you have
sketched seem to show about the relationship between mathematical ability
and musical ability.
(3)
(g) The teacher noticed that most of the students who were rated highly on
musical ability were left-handed. The teacher is aware that her previous
definition of musical ability lacked validity.
You should:
Page 36 of 116
• describe the procedure that you would use, including details of how you
would assess musical ability
(h) In your answer book, draw a table to show how you would record your results.
Identify an appropriate statistical test to analyse the data that you would collect.
Justify your choice.
(3)
(Total 30 marks)
Q20.
A study was carried out to test the effectiveness of a new anger management programme.
The programme had been designed by a team of psychologists working in a young
offenders’ institution.
Fifteen male offenders aged 17– 21 years took part in the programme. An anger score for
each offender was obtained before the start of the programme. This score was based on a
questionnaire designed by the psychologists. The questionnaire had 10 items. The
maximum score was 50; the higher the score, the greater the level of anger.
Throughout the programme, the offenders were told to keep a diary of situations that
made them angry and to record their anger in these situations. After the programme had
ended, they were told to continue to keep their diary.
Two weeks later, after the programme had ended, a second anger score was obtained for
each offender. The same questionnaire was used.
Table 1: Median anger scores and the ranges before and after the programme
Before After
Median 35 24
Range 15 17
(a) Explain why measures of dispersion are often used in addition to measures of
central tendency to summarise data. Refer to the results of this study in your
answer.
(2)
(b) A Wilcoxon signed ranks test was used to test for a significant difference between
the anger scores at the start of the programme and after the programme had ended.
Page 37 of 116
test
(c) Explain why the psychologists decided to use a Wilcoxon signed ranks test to
analyse the data.
(3)
(d) Explain two possible reasons for asking each offender to keep a diary.
(4)
(e) An independent researcher reviewed the design of the study and noted that there
was no control group.
Explain how having a control group could have improved this study.
(3)
(f) The independent researcher was also concerned that the psychologists had not
checked the reliability and validity of the questionnaire used to measure the level of
anger.
Outline how the psychologists could check the reliability and the validity of the
questionnaire.
(5)
(Total 19 marks)
Q21.
Two psychologists investigated the relationship between age and recall of medical advice.
Previous research had shown that recall of medical advice tended to be poorer in older
patients. The study was conducted at a doctor's surgery and involved a sample of 30
patients aged between 18 and 78 years. They all saw the same doctor, who made notes
of the advice that she gave during the consultation.
One of the psychologists interviewed each of the patients individually, immediately after
they had seen the doctor. The psychologist asked each patient a set of questions about
what the doctor had said about their diagnosis and treatment. The patients' responses
were recorded and then typed out. Working independently the psychologists compared
each typed account with the doctor's written notes in order to rate the accuracy of the
accounts on a scale of 1 – 10. A high rating indicated that the patient's recall was very
accurate and a low rating indicated that the patient's recall was very inaccurate.
(c) The psychologists were careful to consider the issue of reliability during the
study. What is meant by reliability?
Page 38 of 116
(1)
(d) Explain how the psychologists might have assessed the reliability of their
ratings.
(3)
(e) This study collected both qualitative and quantitative data. From the
description of the study above, identify the qualitative data and the quantitative
data.
The psychologists used Spearman's rho to analyse the data from their
investigation. They chose to use the 0.05 level of significance. The result gave
a correlation coefficient of −0.52.
(2)
(f) Give two reasons why the psychologists used Spearman's rho to analyse
the data.
(2)
(g) Using the table below, state whether the result is significant or not significant
and explain why.
(2)
0.05 0.01
0.10 0.02
30 0.306 0.425
31 0.301 0.418
Calculated rs must equal or exceed the table (critical) value for significance at the level shown.
(i) Use the information in the table above to explain why the psychologists did
not think that they had made a Type 1 error in this case.
(3)
(Total 19 marks)
Q22.
The psychologists then wanted to see whether the use of diagrams in medical
consultations would affect recall of medical information.
Page 39 of 116
randomly allocated to one of two conditions. In Condition A, a doctor used diagrams to
present to each participant a series of facts about high blood pressure. In Condition B, the
same doctor presented the same series of facts about high blood pressure to each
participant but without the use of diagrams.
At the end of the consultation, participants were tested on their recall of facts about high
blood pressure. Each participant was given a score out of ten for the number of facts
recalled.
(a) In this case, the psychologists decided to use a laboratory experiment rather than a
field experiment. Discuss advantages of carrying out this experiment in a laboratory.
(4)
(b) Identify an appropriate statistical test that the psychologists could use to analyse the
data from the follow-up study. Give one reason why this test is appropriate.
(2)
(Total 6 marks)
Q23.
Psychological research suggests an association between birth order and certain abilities.
For example, first-born children are often logical in their thinking whereas later-born
children tend to be more creative. A psychologist wonders whether this might mean that
birth order is associated with different career choices. She decides to investigate and asks
50 artists and 65 lawyers whether they were the first-born child in the family or not.
(b) Identify an appropriate sampling method for this study and explain how the
psychologist might have obtained such a sample
She analysed her data using a statistical test and calculated a value of = 2.27.
She then looked at the relevant table to see whether this value was statistically
significant. An extract from the table is provided below.
Page 40 of 116
Calculated value of must be equal to or exceed the table (critical) values for significance at the level
shown
(3)
(c) Imagine that you are writing the results section of the report on this investigation.
Using information from the description of the study above and the relevant
information from the statistical table, provide contents suitable for the results
section.
Q24.
A teacher has worked in the same primary school for two years. While chatting to the
children, she is concerned to find that the majority of them come to school without having
eaten a healthy breakfast. In her opinion, children who eat ‘a decent breakfast’ learn to
read more quickly and are better behaved than children who do not. She now wants to set
up a pre-school breakfast club for the children so that they can all have this beneficial start
to the day. The local authority is not willing to spend money on this project purely on the
basis of the teacher’s opinion and insists on having scientific evidence for the claimed
benefits of eating a healthy breakfast.
(a) Explain why the teacher’s personal opinion cannot be accepted as scientific
evidence.
Refer to some of the major features of science in your answer.
A psychologist at the local university agrees to carry out a study to investigate the
claim that eating a healthy breakfast improves reading skills. He has access to 400
five-year-old children from 10 local schools, and decides to use 100 children (50 in
the experimental group and 50 in the control group). Since the children are so
young, he needs to obtain parental consent for them to take part in his study.
(6)
(b) The psychologist used a random sampling method. Explain how he could
have obtained his sample using this method.
(3)
(d) Explain why it is important to operationalise the independent variable and the
dependent variable in this study and suggest how the psychologist might do
this.
(5)
(e) The psychologist used a Mann-Whitney test to analyse the data. Give two reasons
why he chose this test.
(2)
Page 41 of 116
(f) He could have used a matched pairs design. Explain why this design would have
been more difficult to use in this study.
(2)
(g) Other than parental consent, identify one ethical issue raised in this study and
explain how the psychologist might address it.
(2)
(h) The psychologist asks some of his students to conduct a separate observational
study at the same time on the same group of children. The aim of this observational
study is to test the idea that eating a healthy breakfast affects playground behaviour.
Q25.
(a) The psychologist was also interested in the effects of a restricted diet on memory
functioning and he expected memory to become impaired. The psychologist’s
hypothesis was that participants’ scores on a memory test are lower after a
restricted diet than before a restricted diet. He gave the volunteers a memory test
when they first arrived in the research unit and a similar test at the end of the four-
week period. He recorded the memory scores on both tests and analysed them
using the Wilcoxon signed ranks test. He set his significance level at 5%.
(b) Table: Extract from table of critical values from the Wilcoxon signed ranks test
N T≤
19 53 46
20 60 52
21 67 58
22 75 65
Page 42 of 116
Calculated T must be equal to or less than the critical value (table value)
for significance at the level shown
Using the table above, state whether or not the psychologist’s result was significant.
Explain your answer.
(3)
(Total 4 marks)
Q26.
Read the text below and answer the questions that follow.
Observer A 2 5 0 6 4 3
Observer B 4 3 2 1 6 5
(a) Use the data in the Table above to sketch a scattergram. Label the axes and
give the scattergram a title.
(4)
(b) Using the data in the Table above, explain why the psychologist is concerned
about inter-rater reliability.
(4)
(c) Identify an appropriate statistical test to check the inter-rater reliability of these
two observers. Explain why this is an appropriate test.
(3)
(d) If the psychologist does find low reliability, what could she do to improve
inter-rater reliability before proceeding with the observational research?
(4)
(Total 15 marks)
Q27.
A psychologist was interested in testing a new treatment for people with eating disorders.
She put up adverts in several London clinics to recruit participants. Thirty people came
forward and they were all given a structured interview by a trained therapist. The therapist
then calculated a numerical score for each participant as a measure of their current
functioning, where 50 indicates excellent, healthy functioning and zero indicates failure to
Page 43 of 116
function adequately. The psychologist then randomly allocated half the participants to a
treatment group and half to a no-treatment group. After eight weeks, each participant was
re-assessed using a structured interview conducted by the same trained therapist, and
given a new numerical score. The trained therapist did not know which participants had
been in either group.
For each participant, the psychologist calculated an improvement score by subtracting the
score at the start of the study from the score after eight weeks. The greater the number,
the better the improvement.
(a) With reference to the data in the table above, outline what the findings of this
investigation seem to show about the effectiveness of the treatment.
(2)
(b) The psychologist used a statistical test to find out whether there was a significant
difference in improvement between the ‘treatment’ and ‘no-treatment’ groups. She
found a significant difference at the 5% level for a one-tailed test ( p ≤ 0.05).
(c) What is the likelihood of the psychologist having made a Type 1 error in this study?
Explain your answer.
(2)
(d) The psychologist assumed that improvements in the treatment group were a direct
result of the new type of treatment. Suggest two other reasons why people in the
treatment group might have improved.
(4)
(e) The psychologist could have used self-report questionnaires to assess the
participants instead of using interviews with the therapist. Explain one advantage
and one disadvantage of using self-report questionnaires in this study rather than
interviews.
(4)
(f) The psychologist needed to obtain informed consent from her participants. Write a
brief consent form which would be suitable for this study. You should include some
details of what participants could expect to happen in the study and how they would
be protected.
(5)
(g) What is meant by reliability? Explain how the reliability of the scores in this study
could be checked.
(4)
Page 44 of 116
(h) The psychologist noticed that female and male participants seemed to have
responded rather differently to the treatment.
Female patients with an eating disorder will show greater improvement in their
symptoms after treatment with the new therapy than male patients.
She used a new set of participants and, this time, used self-report questionnaires
instead of interviews with a therapist.
Imagine that you are the psychologist and are writing up the report of the study.
Write an appropriate methods section which includes reasonable detail of design,
participants, materials and procedure. Make sure that there is enough detail to allow
another researcher to carry out this study in the future.
(10)
(Total 35 marks)
Page 45 of 116
Mark schemes
Q1.
(a) [AO1 = 1]
1 mark for D
1
(b) [AO3 = 3]
3 marks for a clear and coherent discussion of why a covert observation of children
might be more beneficial than an overt observation using the detail given below.
Possible content:
(c) [AO2 = 4]
0 No relevant content.
Possible content:
• the researcher is looking for a difference in helping behaviour and doing a sign
test is one way in which the analysed data would show such a difference
• the study focused on a single group of children who were tested under both
conditions – the sign test can only be used with one group of participants /
repeated measures design
• using a sign test will allow the researcher to decide whether differences in
helping behaviour are due to chance factors or a ‘real’ effect
Page 46 of 116
• this study produces quantitative/numerical data and the sign test is one way of
analysing such data
• there are rules about when a particular test can be used and in this case the
design of the study meets the rules for using a sign test.
NOTE: reference to level of measurement is not expected, but can be credited, e.g.:
this study produces data in the form of categories / behavioural categories /
frequency counts (nominal level of measurement).
4
[8]
Q2.
(a) [AO2 = 1]
1 mark for:
A Cognitive
1
(b) [AO2 = 4]
If maximum 1 for Title, Title does not need to include ‘score’. Must include both co-
variables and reference to correlation / relationship.
Page 47 of 116
4
(c) [AO2 = 2]
PLUS
Page 48 of 116
1 mark for an explanation:
Possible content:
(d) [AO2 = 1]
+0.70.
1
[8]
Q3.
(a) [AO2 = 6]
0 No relevant content.
Possible content:
Page 49 of 116
night before going to bed for 7 nights
• a daily requirement to truthfully respond to a text message asking whether
they had experienced a nightmare
• the two-week duration of the experiment.
Ethical guidelines:
• no pressure to consent
• they can withdraw at any time
• they can withdraw their data from the experiment
• their data will be kept confidential and anonymous
• they should feel free to ask the researcher any questions at any time
• they will receive a full debrief at the end of the programme.
(b) [AO2 = 3]
2 marks for a statement with both conditions of the IV and a DV that lacks clarity
and coherence or has only one variable operationalised.
1 mark for a muddled statement with both conditions of the IV and DV present or
where neither variable is operationalised.
Possible content:
• participants will report more nightmares after watching a horror film before
bedtime than after watching a romantic comedy film before bedtime. Accept
alternative wording
• participants will report fewer nightmares after watching a horror film before
bedtime than after watching a romantic comedy film before bedtime. Accept
alternative wording.
3
(c) [AO2 = 2]
Possible explanations:
Page 50 of 116
• necessary to avoid the effects of individual differences in frequency of
nightmares
• film viewing habits, gender, hours of sleep, personality, etc, can have a big
impact on the number of nightmares recalled.
(e) [AO2 = 3]
• all 50 participants’ names / numbers are put into a hat / container / computer
• a name is drawn from the container or a random name is generated by the
computer and is assigned to the first group
• a second name is selected as before but this time goes in to the second
group; this process continues until there are 25 in each group.
OR
Mean:
1 mark – participants who watch horror films before going to bed report more
nightmares then those who watch romantic comedies before bed. Accept alternative
wording.
Plus
1 mark – mean number of nightmares reported is greater when horror films are
watched than when romantic comedies are watched. Accept alternative wording.
Standard deviation:
Page 51 of 116
Plus
1 mark – standard deviation is greater when horror films are watched before going
to bed than when romantic comedies are watched before going to bed. Accept
alternative wording.
Note - 0 marks for just stating the data from the table.
2 marks for a clear and appropriate explanation in the context of this experiment.
This means that the difference in the number of nightmares reported after watching
horror films compared to romantic comedies is significant at 0.05 level. This means
there is less than 5% (1 in 20) likelihood (probability) that the difference was due to
chance / due to something other than the IV.
1 mark for a modification to the design of the experiment that could improve validity.
Plus
2 marks for clear and coherent explanation for how the suggested modification
might improve validity of this study.
1 mark for a limited/muddled explanation for how the suggested modification might
improve validity of this study.
Possible content:
• include more than one question in the text message to the students. This
would make the aim of the experiment less obvious to guess which would in
turn reduce demand characteristics and improve the validity of the experiment
• make the conditions less obvious. Rather than having one film the students
could be asked to watch an episode from a TV series. The episodes watched
in one condition would be those which contained any scary concepts / themes
whereas episodes watched in the other condition would not contain any scary
themes / concepts
• guarantee anonymity so people will give honest answers and not feel
embarrassed
• use a broader sample, not just students. Students may be more or less
inclined to watch horror films anyway
• use a different sampling technique to avoid a self-selected sample and thus
avoid volunteer bias as that may make them more susceptible to demand
characteristics.
Page 52 of 116
study; however candidates can make the case for a matched pairs design, which
could be credit worthy.
3
[25]
Q4.
(a) [AO2 = 2]
Plus
OR
Plus
OR
Plus
If students identify more than one type of data, take the first type as the basis for
their answer. If the data is not identified or identified incorrectly, no credit can be
given for an explanation.
(b) [AO2 = 3]
1 mark for the likely outcome: more participants in condition 1 will complete the
questionnaire than in condition 2/fewer participants in condition 2 will complete the
questionnaire than in condition 1.
Plus
Page 53 of 116
Possible content:
(c) [AO3 = 4]
Plus
OR
Plus
Credit other plausible ways of addressing the question. To gain any credit answers
based on repeated measures should be appropriately detailed e.g. parallel versions
of the questionnaire, time lapse etc.
(d) [AO2 = 2]
1 mark for the researcher is looking for a difference (between two conditions/sets of
data) or an association/relationship (between two variables).
(e) [AO2 = 2]
Page 54 of 116
1 mark for stating that the value of chi squared is significant (at the 5% level).
Plus
Q5.
(a) [AO2 = 1]
1 mark
(b) [AO2 = 2]
(c) [AO2 = 2]
1 mark for explaining either you need to have continuous data or scores for each
participant in order to draw a histogram.
Plus
1 mark for identifying that the data represents two separate conditions (with
music/without music). Accept categorical/nominal.
Note: credit can be given for two separate conditions if the student explains clearly
why this would make a histogram “inappropriate”.
(d) [AO3 = 3]
Note: these are independently awarded marks, eg candidates can achieve 2 marks
for correctly labelled axes despite an incorrect graph type
Mean:
1 mark for interpreting what the mean times suggest about the effect of music on
the participants’ 400m performance – participants run faster with music (take less
time to run 400 metres) or participants run more slowly without music (take more
time to run 400 metres). Accept alternative wording.
Page 55 of 116
Plus
1 mark for an accurate justification about the difference in the mean scores in each
condition – mean time is greater in condition A than condition B (or mean time is
lower in condition B than condition A).
Standard deviation:
1 mark for an accurate comment about what the standard deviations suggest about
the spread of scores in each condition – performance is more consistent in condition
A than condition B (or performance is less consistent in condition B than condition
A). Accept alternative wording.
Plus
1 mark for a justification about the difference between the standard deviations in
each condition – standard deviation is smaller in condition A than in condition B (or
standard deviation is greater in condition B than condition A).
Note: 0 marks for just stating the data from the table, eg the mean time with music
is 117 whereas it is 123 without music.
(f) [AO2 = 4]
Marks are for calculations and/or numerical answer – no need to show unit (%)
4 marks for the correct answer given to three significant figures: 4.88 (even if no
correct workings are shown).
3 marks for correct answer not given to three significant figures eg 4.878 or 4.9.
2 marks if incorrect answer is provided even if all working is correct.
1 mark if incorrect answer and workings are partially correct eg one or two of the
correct steps.
0 marks if the incorrect answer is given to three significant figures.
Correct workings:
123 − 117 = 6
6 ÷ 123 =0.048780
0.048780 × 100 = 4.878
Answer = 4.88
(g) [AO2 = 5]
Credit other appropriate reasons e.g. reasons related to possible normal distribution,
power of the test.
Plus
Page 56 of 116
• the result is not significant (at the 5% level)
• because the calculated value of t (1.4377) is less than the critical/table value
of t, which is 1.833 (at 0.05, for a directional hypothesis where df is 9).
(h) [AO1 = 3]
A Type II error would occur where a real difference in the data is overlooked as it is
wrongly accepted as being not significant, accepting the null hypothesis in error (a
false negative).
Plus
1 mark for a reason for why the 5% level of significance is used in psychological
research.
The 5% level is used as it strikes a balance between the risk of making the Type I
and II errors (or similar).
Plus
Plus
(j) [AO1 = 6]
0 No relevant content.
Page 57 of 116
Possible content:
Process
• other psychologists check the research report before deciding whether it could
be published
• independent scrutiny by other psychologists working in a similar field
• work is considered in terms of its validity, significance and originality
• assessment of the appropriateness of the methods and designs used
• reviewer can accept the manuscript as it is, accept with revisions, suggest the
author makes revisions and re-submits or reject without the possibility of re-
submission
• editor makes the final decision whether to accept or reject the research report
based on the reviewers’ comments/recommendations
• research proposals are submitted to panel and assessed for merit.
Purposes
• to ensure quality and relevance of research, eg methodology, data analysis etc
• to ensure accuracy of findings
• to evaluate proposed designs (in terms of aims, quality and value of the
research) for research funding.
Q6.
(a) [AO2 = 1]
1 mark C
(b) [AO2 = 2]
1 mark 4/5ths
(c) [AO3 = 2]
2 marks for a clear and coherent conclusion, plus relevant explanation based on the
data.
The training course appears to have a beneficial effect on teacher confidence as the
majority of them (16 out of 20) say their confidence has improved.
(d) [AO2 = 2]
(e) [AO2 = 3]
Page 58 of 116
1 mark repeated measures design
Plus
2 marks for a clear and coherent explanation of why this design is appropriate in
this case
1 mark for a vague or muddled explanation of why this design is appropriate in this
case
Content:
(f) [AO2 = 3]
1 mark then taking the numerical value for/number of participants with the least
common/frequent sign
(g) [AO1 = 2]
Content:
(h) [AO3 = 2]
(i) [AO3 = 3]
Page 59 of 116
• Separate spaces for first and last 10 minutes
• Headed correctly with the six category spaces (may include the two used in
answer to part (h) but names of categories not essential here)
(j) [AO3 = 4]
0 No relevant content.
Relevant problems:
• limiting observations to first and last 10 minutes means the data may not be a
valid representation of disruptive behaviour in lessons. Need to carry out
observations at other times during the lesson too.
Q7.
(a) [AO2 = 2 AO3 = 4]
Page 60 of 116
justification is limited. The answer lacks clarity.
0 No relevant content.
Means
Standard deviations
(b) [AO3 = 3]
Plus
1 mark – one that takes account of the distance of all the verbal error scores
from the mean.
Plus
1 mark – not just the distance between the highest verbal error score and the
lowest verbal error score.
(c) [AO2 = 4]
1 mark for naming the t-test for independent / unrelated groups or a Mann-
Whitney test.
Plus
OR
Page 61 of 116
• data should be treated as ordinal. Cannot assume interval data because
verbal errors cannot be assumed to be of equal size (ie one verbal error
is not equivalent to any other verbal error)
• the experimental design is independent groups
• the psychologist is looking for a difference between the two conditions
• SDs are quite different.
(d) [AO1 = 2]
This means that there is a less than 5% likelihood that this difference would
occur if there is no real difference between the conditions OR the researchers
would have a 95% confidence level.
1 mark for a less clear answer which shows some understanding, eg this
means the researcher can conclude that the difference was not due to chance.
(e) [AO2 = 2]
1 mark for a partial or muddled explanation or one that is only loosely applied
to the study.
Credit answers based on any type of validity. Most answers will refer to either
face or concurrent as follows:
• asking other people if verbal errors are a good measure of verbal fluency
(face validity)
• giving participants an alternative / established verbal fluency test and
checking to see that the two sets of data are positively correlated
(concurrent validity).
Q8.
(a) [AO2 = 4]
2 marks for identifying two factors that are relevant for use of the sign test:
nominal/categorical data; test of difference; related design/repeated measures.
Plus
(b) [AO2 = 2]
Plus
Page 62 of 116
1 mark for explanation/calculation of how this was arrived at:
• The most commonly occurring sign is + (12) and the least frequently
occurring sign is – (5). The 0s are disregarded.
• The total for the least frequently occurring sign is the value of s = 5
(c) [AO2 = 2]
1 mark for stating that the value of s (5) is not significant at the 0.05 level.
Plus
(d) [AO3 = 3]
Possible points:
• Primary data are obtained ‘first-hand’ from the participants themselves
so are likely to lead to greater insight: e.g. into the patients' experience
of treatment, whether they found it beneficial, negative, etc.
• Secondary data, such as time off work, may not be a valid measure of
improvement in symptoms of depression. Primary data are more
authentic and provide more than a surface understanding: e.g.
participants may have taken time off work for reasons not related to their
depression.
• The content of the data is more likely to match the researcher’s needs
and objectives because questions, assessment tools, etc. can be
specifically tailored: e.g. an interview may produce more valid data than
a list of absences.
0 No relevant content.
Page 63 of 116
AO1 – possible content:
• Psychological research may lead to improvements in psychological
health/treatment programmes which may mean that people manage
their health better and take less time off work.
• Absence from work costs the economy an estimated 15 billion a year
annually and much of this absence is due to ‘mild’ mental illness: e.g.
stress, anxiety.
• Psychological research may lead to better ways of managing people
whilst they are at work to improve productivity: e.g. research into
motivation and workplace stress.
• ‘Cutting-edge’ scientific research may encourage investment from
overseas companies into this country.
AO2 – application
• If research (such as the investigation described) suggests that
depressives are better able to manage their condition following CBT and
return to work, then it may benefit the economy to make treatment more
widely available, improve funding, etc.
• Psychological research such as this plays an important role in sustaining
a healthy workforce and reducing absenteeism.
Q9.
(a) [AO2 = 2]
Plus
1 mark – because there is past research indicating the likely direction of the effect
(or similar)
Plus
1 mark each for two observable behaviours that could represent ‘riding a bike with
care’.
Page 64 of 116
Credit any relevant observable behaviour.
(c) [AO2 = 3]
1 mark for knowledge of the term face validity – where a behaviour appears at first
sight (on the face of it) to represent what is being measured
Plus
2 marks for clear and coherent application of the concept of face validity to the
context
1 mark for brief or muddled application of the concept of face validity to the context
(d) [AO1 = 4]
Plus
Content:
• Temporal – where findings from research that took place at a certain point in
time accurately reflect the way that behaviour would occur at a different point
in time
(e) [AO2 = 1]
(f) [AO2 = 3]
• The student pair should discuss and agree beforehand their interpretation of
the behavioural categories
• Each student should then observe the same people/space/target at the same
time but record/tally independently
Page 65 of 116
statistical test ascertain the level of agreement
(g) [AO2 = 3]
Plus
54 ÷ 9 = 6 and 36 ÷ 9 = 4
(h) [AO2 = 3]
• Data is categorical/nominal/frequency
(i) [AO2 = 2]
Plus
(2 − 1) x (3 − 1) = 1 x 2 = 2
(j) [AO2 = 3]
• Because the calculated value of Chi-square is more than the critical table
value at 0.05, either 4.60 or 5.99, depending on whether student uses one or
two-tailed values
• One-tailed is consistent with the hypothesis and Q13, but there is an argument
that the Chi-square test should always be two-tailed so either can be credited
• Where df equals 2
(k) [AO2 = 4]
Page 66 of 116
The paragraph is clear and coherent, showing sound
understanding of the concept of levels of significance and
2 3–4
effective application to the context. There is effective use of
terminology.
0 No relevant content.
Possible content:
• With the present results there is a 95% confidence in accepting the research
hypothesis/confidence that any difference/effect is due to the variables under
investigation, in this case the location of the public spaces.
• There is a 5% possibility that the same frequencies would occur if there was
no real difference between the two towns.
• The calculated value in this case well exceeds the critical value at 0.05 but
does not meet the more stringent level of significance of 0.01.
(l) [AO3 = 3]
1 mark for explaining that content analysis is suitable because the students are
analysing recordings which are a form of media.
Plus
2 marks for a clear, coherent account of how the content could be analysed
1 mark for a brief or muddled account of how the content could be analysed
Content:
They could then set up a system of categories and tally the ideas/concepts
(m) [AO3 = 1]
1 mark for any relevant inconsiderate behaviour eg leaving rubbish, leaving dirty
mugs/plates, playing music loudly, throwing books, shouting
Q10.
Page 67 of 116
(a) [AO2 = 2]
Plus
(b) [AO3 = 3]
Possible strengths
• Control issues – first hand data can be controlled whereas secondary data
may have been gathered under differing conditions
(c) [AO3 = 4]
Plus
• Data assumed to be ordinal ie not fixed intervals (also credit data is assumed
to be non-parametric)
Q11.
(a) [AO2 = 2]
Plus
(b) [AO3 = 3]
Page 68 of 116
2 marks for statement of a strength plus some elaboration
Possible strengths
(c) [AO3 = 4]
Plus
• Data assumed to be ordinal ie not fixed intervals (also credit data is assumed
to be non-parametric)
Q12.
(a) [AO2 = 2]
Plus
(b) [AO3 = 3]
Possible strengths:
• Control issues – first hand data can be controlled whereas secondary data
may have been gathered under differing conditions
Page 69 of 116
Credit other relevant strengths.
(c) [AO3 = 4]
Plus
• Data assumed to be ordinal ie not fixed intervals (also credit data is assumed
to be non-parametric)
Q13.
(a) [AO3 = 2]
Plus
Possible content:
Control for individual differences so that the researcher can be more certain that the
effect is not due to characteristics such as gender, personality etc.
(b) [AO2 = 2]
Plus
1 mark for explanation - ‘S’ is the frequency of the least common difference. There
are 8 positive differences, 2 negative differences and 2 ties.
(c) [AO2 = 2]
1 mark for identifying the correct value of N (N = 10, total number of differences)
Plus
1 mark for stating that for an N of 10, a value of ‘S’ = 2 is not significant at the p≤ .05
level.
Q14.
(a) [AO3 = 2]
Page 70 of 116
Plus
Possible content:
Control for individual differences so that the researcher can be more certain that the
effect is not due to characteristics such as gender, personality etc.
(b) [AO2 = 2]
Plus
1 mark for explanation - ‘S’ is the frequency of the least common difference. There
are 8 positive differences, 2 negative differences and 2 ties.
(c) [AO2 = 2]
1 mark for identifying the correct value of N (N = 10, total number of differences)
Plus
1 mark for stating that for an N of 10, a value of ‘S’ = 2 is not significant at the p
≤ .05 level.
Q15.
(a) [AO3 = 2]
Plus
Possible content:
Control for individual differences so that the researcher can be more certain that the
effect is not due to characteristics such as gender, personality etc.
(b) [AO2 = 2]
Plus
1 mark for explanation - ‘S’ is the frequency of the least common difference. There
are 8 positive differences, 2 negative differences and 2 ties.
(c) [AO2 = 2]
1 mark for identifying the correct value of N (N = 10, total number of differences)
Page 71 of 116
Plus
1 mark for stating that for an N of 10, a value of ‘S’ = 2 is not significant at the p
≤ .05 level.
Q16.
Please note that the AOs for the new AQA Specification (Sept 2015 onwards) have
changed. Under the new Specification the following system of AOs applies:
Although the essential content for this mark scheme remains the same, mark schemes for
the new AQA Specification (Sept 2015 onwards) take a different format as follows:
(a) AO1 = 2
Award 1 mark for a brief statement and a further mark for elaboration.
(b) AO3 = 4
• The psychologist could have begun by watching some of the film clips of
driver behaviour.
• The psychologists would then have watched the films again and counted
the number of examples which fell into each category to provide
quantitative data.
4 marks Effective
Effective explanation of the processes involved in content analysis referring to
some or all of the above points.
2 – 3 marks Reasonable
Page 72 of 116
Reasonable accurate coverage of the processes involved.
1 mark Basic
Basic identification of the processes involved in content analysis (‘watching the
films and counting’).
0 marks
No creditworthy material.
(c) AO3 = 3
2 marks for some explanation / elaboration: ‘the two psychologists could carry
out content analysis of the films separately and compare their answers’ or
‘they could re-code the films at a later date and compare the two sets of data’.
3 marks for an accurate and clear explanation which refers to deriving the
categories and checking the data. ‘The two psychologists could watch the
films separately and devise a set of categories. They could compare these and
use categories they both agreed on. They could carry out content analysis of
the films separately and compare their answers looking for agreement’.
(d) AO3 = 3
Candidates can cover one reason explained in detail here or several reasons
in less detail.
(e) AO3 = 3
• the nature and content of the conversation with the psychologist on the
hands-free phone
Page 73 of 116
Award 1 mark for basic identification of a confounding variable and a further 2
marks for elaboration of how this could have affected the dependent variable.
Example: The chat with the psychologist was not controlled (1 mark) so the
difficulty or number of questions could have varied (2 marks). This would
influence the DV as more or less attention would be required (3 marks).
(f) AO3 = 4
There are several potential ethical issues here. Candidates can focus on one
in detail or several in less detail.
4 marks Sound
An appropriate ethical issue is identified and explained in detail. Material is
accurate – or several issues are identified and discussed accurately in less detail.
2 – 3 marks Reasonable
One or more appropriate ethical issues are identified and discussed. The answer
is generally accurate.
1 mark Basic
Basic identification of an ethical issue (e.g. ‘right to withdraw’) or very brief
answers which lack detail.
0 marks
No creditworthy material.
(g) AO3 = 5
a. You will take part in a simulated driving test which will last for three
minutes.
Page 74 of 116
b. Your task will be to identify potential hazards on the road ahead.
c. When you see a hazard, you should press the mouse button as quickly
as possible.
d. Whilst you are doing the test, I will chat to you on a mobile phone and I
would like you to reply using the hands-free mobile phone headset.
For full marks, the instructions should adopt an appropriate formal tone.
Instructions which are not suitable to be read out should be awarded a
maximum mark of 2.
5 marks Effective
The standardised instructions provide accurate detail of the procedure in a clear
and concise form and participants’ understanding is checked.
4 – 3 marks Reasonable
The standardised instructions provide sufficient detail of the procedure in a
reasonably clear form.
2 marks Basic
The standardised instructions provide some details of the procedure though these
may not be clear.
1 mark Rudimentary
The standardised instructions provide few details of the procedure and may be
muddled and or inaccurate. Omissions in the instructions compromise the
procedure.
0 marks
No creditworthy material is presented.
(h) AO3 = 3
Students are required to identify an appropriate test and are asked to justify
their choice.
Award 1 mark for identification of the Wilcoxon (signed ranks) test. Candidates
could receive credit for Sign test or related t test. Note that reasons /
justification must be correct for the test supplied.
Award 1 mark for basic statement of a reason, and a further mark for
elaboration, within the context of the experiment or a further reason.
• A repeated measures design was used (1 mark) and the data can be
Page 75 of 116
treated as ordinal (1 mark).
(i) AO3 = 2
Students are told that the difference in reaction times was significant at the p ≤
0.01 level.
Award 1 mark for a basic understanding of this (‘the result is highly significant’)
and a further mark for elaboration e.g. identifying that the probability of a Type
1 error here is less than 1 / 100.
(j) AO3 = 3
Q17.
Please note that the AOs for the new AQA Specification (Sept 2015 onwards) have
changed. Under the new Specification the following system of AOs applies:
Although the essential content for this mark scheme remains the same, mark schemes for
the new AQA Specification (Sept 2015 onwards) take a different format as follows:
(a) [AO3 = 1]
(b) [AO3 = 2]
One mark for naming a test: Spearman’s rank order correlation / rho or
Pearson’s product moment correlation.
Page 76 of 116
One mark for justification. For Spearman’s rank order correlation accept: not
all data is interval – data collected for empathy test score most likely treated at
ordinal level of measurement due to self-report.
For Pearson accept: Pearson’s product moment correlation is a robust test,
even if not all data can be treated as truly interval.
Just stating ordinal / interval no credit. Accept ordinal or interval providing this
is justified with reference to at least one variable.
Unlikely but allow for an informed argument made for treating both sets of data
at interval level.
(c) [AO3 = 2]
1 mark for a knowledge of a way (not just naming a type of validity) and 2nd
mark for explaining how this would be implemented in this case. Most likely
answers will address face validity or concurrent validity, but accept any other
way such as construct validity, content validity, criterion validity and predictive
validity.
For full marks, the answer must refer to either the empathy questionnaire or
empathy test items. The ‘way’ need not be named or defined.
(d) [AO3 = 2]
(e) [AO3=3]
(f) [AO3 = 8]
Page 77 of 116
matched pairs;
• detail of sample;
• materials required for carrying out the research, eg task for assessing
levels of recall, timing device if needed;
• sufficient procedural details to carry out a replication (might include
standard instructions, ethics, etc.)
Note: standardised instructions and ethical issues are not required for full
marks.
Mark bands
Q18.
Please note that the AOs for the new AQA Specification (Sept 2015 onwards) have
changed. Under the new Specification the following system of AOs applies:
Although the essential content for this mark scheme remains the same, mark schemes for
the new AQA Specification (Sept 2015 onwards) take a different format as follows:
Page 78 of 116
• No IDA expectation in A Level essays, however, credit for references to issues,
debates and approaches where relevant.
(b) AO2/AO3 = 4
Students could also make a case for the analysis of diaries/written materials
as a way of collecting data about happiness. These would generally overcome
the problems of social desirability and demand characteristics inherent in
questionnaires. Students could also make a case for the use of observation.
(c) AO2/AO3 = 2
(d) AO2/AO3 = 3
Students should state that the obtained value of + 0.42 exceeds the critical
value for a twotailed test (.362) for N = 30. The results are therefore
statistically significant (p ≤ 0.05) Award 2 marks for a student who supplies two
pieces of information. Award 1 mark for a student who states that the results
are significant but does not provide an explanation OR the student who states
results are significant but uses incorrect values from the table. Award 0 marks
for students who argue that results are not significant.
(e) AO2/AO3 = 4
Page 79 of 116
points below:
Students may also make the point that there may be a weak tendency for
more intelligent teenagers to be less happy at 16 years of age, although this is
not statistically significant. Students may also refer to the contradiction in the
results or provide an overall conclusion.
4 marks Effective
Effective analysis and understanding.
The answer includes the findings of the two studies which are expressed clearly and
fluently with appropriate reference to intelligence and happiness. Effective use of
statistical terminology.
3 marks Reasonable
Reasonable analysis and understanding. The answer is generally focussed and
includes reference to both of the key findings which are reasonably clear. There is
reasonable use of statistical terminology.
2 marks Basic
Basic, superficial understanding. The answer is sometimes focussed OR covers
only one of the key conclusions. Expression of ideas lacks clarity. Limited use of
statistical terminology.
1 mark Rudimentary
Rudimentary with very limited understanding.
The answer is weak, muddled and may be mainly irrelevant.
Deficiency in expression of ideas results in confusion and ambiguity. The answer
lacks structure, often merely a series of unconnected assertions.
0 marks
No creditworthy material is presented.
Q19.
(a) AO2 / AO3 = 3
Page 80 of 116
The main issue is that the teacher has made up her own test:
• This involved subjective judgement on the part of the teacher who rates the
students’ musical ability. Her judgement may not reflect real differences in
musical ability and is likely to differ from other people’s judgement and / or any
absolute criteria for tunefulness.
• Lack of reliability in rating musical ability would compromise the validity of the
measure.
• As the students can choose the song they will sing, the rating of ability could
reflect the teacher liking / dislike of the song rather than the student’s ability.
• The rating may be invalid as the students selected songs which varied in
difficulty so the tunefulness reflected the difficulty of the song not the students’
ability.
In the case of the maths test candidates could refer to split half or test retest as
methods of checking reliability. They could also refer to checking the reliability of
scoring by using two separate markers for the test and comparing the scores. Credit
any other appropriate suggestion.
The teacher chose to use a random sample because it would probably be more
representative of the whole GCSE group than if she had used an opportunity or
volunteer sample. Candidates could also say that she had ready access to her
target population making it convenient for her to select a random sample.
Credit should only be awarded for scattergraphs. Other graphs gain 0 marks.
Page 81 of 116
(f) AO2 / AO3 = 3
• This means that high scorers in mathematical ability tend to achieve low
scores on musical ability and vice versa.
• The presence of two strong outliers, means that the actual correlation is very
weak and closer to zero.
• Comment on the small sample size which limits the conclusions that could be
drawn.
Page 82 of 116
• Credit can be achieved for plausible interpretations of the strength of the
correlation which are justified (ie looks moderate to strong or the outliers make
it weak in practice) or those based on rough calculations (around -0.2).
Design – 1 mark
Sampling – 2 marks
Award 1 mark for procedure, 1 mark for assessing musical ability and two further
marks for elaboration of either or both of these.
Debrief – 3 marks
• Award up to 3 marks for writing a debrief. This could include the aim of the
study, thanking participants for taking part, asking if they have any questions,
relevant ethical considerations.
Award 1 mark for a clear table appropriate for the study described in (h).
Page 83 of 116
2
Award 1 mark for the identification of an appropriate statistical test for the proposed
design.
Award 1 mark for one correct justification eg a test of difference, at least ordinal
level data.
Q20.
Please note that the AOs for the new AQA Specification (Sept 2015 onwards) have
changed. Under the new Specification the following system of AOs applies:
(a) [AO3 = 2]
(b) [AO3 = 2]
(c) [AO3 = 3]
Maximum of 3 marks can be obtained from: one mark for each reason or two
marks for each reason with explanation.
(d) [AO3 = 4]
Up to two marks for each reason and explanation. Likely points: as an aid to
memory; a qualitative measure to supplement the quantitative data collected;
to check the validity of the questionnaire; part of the therapeutic process /
Page 84 of 116
increased self-awareness.
Accept other valid reasons.
One mark for an appropriate reason and one mark for an explanation of the
reason.
(e) [AO3 = 3]
Up to three marks for outlining how a control group could have improved this
study: it is not possible to tell if the programme has caused the improvement;
improvement could have been due to the programme or due to spontaneous
recovery; by using a control group would make it more scientific; scores can
be taken at the same times (pre-programme / post-programme) as in an
experimental condition; post programme differences between the groups can
inform if programme is effective; can be more confident in inferring cause and
effect .
Allow a maximum of one mark for the general purpose of a control condition:
acts as comparison / baseline measure where nothing changes
(f) [AO3 = 5]
Up to 5 marks for addressing both reliability and validity. One of these marks
must be for reference to statistical testing.
One mark for identifying a type of validity: face validity; concurrent validity.
Accept also content validity; criterion validity; predictive validity.
Only accept identification mark if it matches how the assessment would be
carried out.
One mark for outlining how the assessment would be carried out. For example
in concurrent validity, scores from the questionnaire are compared with those
from an established but similar questionnaire known to have good validity to
see if the results are similar.
One mark for the statistical testing (checking for a positive correlation /
applying Spearman’s rank order correlation).
One mark for identifying a way of assessing reliability. Most likely is test-retest
but accept split-half reliability and item analysis.
Only accept identification mark if it matches how the assessment would be
carried out.
Do not accept inter-rated / inter-observer reliability.
One mark for outlining how the assessment would be carried out. For example
in test-retest, the same group of young offenders would be tested using the
same questionnaire at a later date to see if the findings remained consistent.
One mark for the statistical testing (checking for a positive correlation /
applying Spearman’s rank order correlation).
The one mark for statistical testing can only be credited once.
Q21.
Page 85 of 116
Please note that the AOs for the new AQA Specification (Sept 2015 onwards) have
changed. Under the new Specification the following system of AOs applies:
One mark for an accurate reason: The decision to use a directional hypothesis was
based on findings of previous research which pointed to an effect in a particular
direction ie memory is poorer with age.
• 2 marks for a directional correlational hypothesis that identifies age and recall
as the two variables but is not fully operationalised
• 1 mark for a directional hypothesis where the variables are not identified
(‘there will be a negative correlation’) or where the hypothesis lacks clarity.
(c) AO1 = 1
One mark for an accurate definition: The extent to which results or procedures are
consistent or simply 'consistency'.
One mark for identification of a way of ensuring reliability. By far the most likely
answer here is inter-rater reliability.
Two marks for some explanation/elaboration: using two separate psychologists and
comparing them.
Three marks for an accurate and clear explanation: using two separate
psychologists to rate the typed accounts for accuracy and comparing / correlating
the ratings to see how similar they are.
Candidates could make a case for test retest which would involve the same
psychologist re-examining the ratings after a period of time.
Award one mark for correct identification of one of each type of data.
• Qualitative data: the patient’s responses, the typed accounts, the doctor’s
notes.
Page 86 of 116
(f) AO2 / AO3 = 2
• the data is to be treated as ordinal because the recall accuracy is in the form
of ratings.
Second mark for explaining that -.52 exceeds .306 (p ≤ 0.05, n=30 for a one-tailed
test).
(h) AO1 = 2
One mark for a brief or muddled answer which hints at rejecting HO / accepting the
H1 in error.
Two marks for explaining the term: where the researcher rejects the null hypothesis
(or accepts the research / alternative hypothesis) when in fact the effect is due to
chance – often referred to as an error of optimists.
• This means that the researchers can be 99% certain that the results obtained
are not due to chance.
Award one mark for stating that the obtained value (-0.52) exceeds the critical value
(0.306) by a reasonable margin.
Q22.
Please note that the AOs for the new AQA Specification (Sept 2015 onwards) have
changed. Under the new Specification the following system of AOs applies:
Page 87 of 116
advantage(s) of using a laboratory experiment in this case.
The most likely advantages of the laboratory setting in this experiment include:
• Control over extraneous variables. The lab setting meant that extraneous
variables could be minimised. In this experiment, outside factors such as
waiting time, noise and stress (which would be difficult to control in a field
experiment) were removed.
• Ethical issues. In this case, the testing of memory in a field experiment would
have involved ethical issues including deception of patients or withholding of
information.
Candidates may also refer to other advantages of the laboratory setting such as
replicability. These can receive full credit if they contextualised within the scenario.
Award four marks for an answer which provides accurate and detailed discussion of
relevant advantage(s) with a clear link to the scenario.
Award two or three marks for an answer which includes discussion of relevant
advantage(s), with some reference to the scenario.
Award one mark only for an answer which merely identifies one or more relevant
advantage(s) of a laboratory experiment appropriate to this scenario.
Advantages of laboratory experiments which are not relevant to this study cannot
gain any credit eg use of technical equipment.
• One mark for correctly identifying the Mann Whitney U test or independent t
test.
• One mark awarded for an accurate reason for choice (for Mann Whitney these
are: test of difference, independent groups design / independent data or data
which can be treated at an ordinal level).
Q23.
Please note that the AOs for the new AQA Specification (Sept 2015 onwards) have
changed. Under the new Specification the following system of AOs applies:
2 marks for a clear hypothesis, 1 mark for a hypothesis which lacks clarity.
Page 88 of 116
(b) AO2 / AO3 = 3
This is a 12 mark question but marks are allocated to each of the required
components as follows:
Table: Table to show the career choices of first born and non-first born children
For 3 marks, candidates need to display the data relating to first born and non-first
born career choices on a bar chart. They should label axes correctly and draw the
columns to the correct approximate height for a sketch.
For 2 marks, candidates display data as above but labels are missing or lack clarity.
For 1 mark, candidates graph the data supplied in the question relating to first born
career choices only.
Page 89 of 116
NB Labelled axes but no bars = 0 marks.
The most likely significance level is 5% (p ≤ 0.05). Candidates are not asked to
justify their choice. Candidates who choose a more stringent level can achieve
marks but they must then follow this through when they make their statement of
results.
Candidates who erroneously report 0.05% or p = 0.5 do not gain credit for level of
significance but can achieve credit for the statement of results in relation to the
hypothesis.
For full marks, the candidate should state whether or not they can accept the
hypothesis (or they can express this in terms of rejecting the null hypothesis) at a
given significance level and refer to the observed and critical values.
Where candidates choose an inappropriate value from the table but interpret that
value correctly, they can gain 2 marks.
The critical value for x² (df =1 p 0.05 (two-tailed)) is 3.84. As the observed value of
x² 2.27 is less than the critical value, we cannot reject the null hypothesis. There is
not an association between birth order and career choice.
Page 90 of 116
Q24.
Please note that the AOs for the new AQA Specification (Sept 2015 onwards) have
changed. Under the new Specification the following system of AOs applies:
Although the essential content for this mark scheme remains the same, mark schemes for
the new AQA Specification (Sept 2015 onwards) take a different format as follows:
(a) AO2/3 = 6
Candidates need to show that they understand what differentiates opinion from
scientific evidence. They could mention some of the following:
• The teacher has only experienced one school in a particular catchment area
so she has only observed a very limited number of 5 year-olds (issues of
sampling and replicability).
• She has found out that children do not eat anything nourishing simply by
chatting with the children. She has no corroborative evidence from eg parents
(issues of objectivity).
• She uses vague phrases such as ‘decent breakfast’ without being clear what
this means (operationalisation).
• She has generated a theory and made predictions based on flimsy evidence.
• She has not used any scientific method to lead to her conclusions eg a
carefully controlled experiment, survey or observation.
• She has drawn conclusions about the effects of breakfast without considering
other variables which might affect reading skills and behaviour.
6 marks Effective
Explanation demonstrates sound understanding. Application of knowledge is
effective and shows coherent elaboration. Ideas are well structured and expressed
clearly and fluently. Consistently effective use of psychological terminology.
5 – 4 marks Reasonable
Explanation demonstrates reasonable understanding. Application of knowledge is
reasonably effective and shows some elaboration. Most ideas appropriately
structured and expressed clearly. Appropriate use of psychological terminology.
3 – 2 marks Basic
Explanation demonstrates basic, superficial understanding. Application of
knowledge is basic. Expression of ideas lacks clarity. Limited use of psychological
terminology.
Page 91 of 116
1 mark Rudimentary
Explanation is rudimentary, demonstrating very limited understanding. Application of
knowledge is weak, muddled and may be mainly irrelevant. Deficiency in expression
of ideas results in confusion and ambiguity. The answer lacks structure, often
merely a series of unconnected assertions.
0 marks
No creditworthy material is presented.
(b) AO2/3 = 3
In a random sample, every member of the identified population has an equal chance
of selection. In this case, the sampling frame consists of the 400 five-year-old
children attending ten local schools. In order to obtain a simple random sample, the
researcher has to have the names of all 400 children and can then select using one
of the following methods:
• Manual selection – Using this method, the researcher has to put each name
(or an assigned number) on a separate slip of paper and place them all in a
container. The researcher then selects 100 slips from the container. The
following conditions could apply: the container should be shaken between
each draw; the slips of paper should all be the same size and folded in the
same way so that one does not feel different from another; the selector draws
‘blind’ ie cannot see the actual slips of paper.
(c) AO2/3 = 3
Page 92 of 116
• Practical limitations eg the time and effort needed to write out 400 slips for the
manual method.
(d) AO2/3 = 5
The terms’ ‘decent breakfast’ and ‘reading skills’ are vague. It is important from the
point of view of objectivity, replicability and control of extraneous variables to make
sure that these terms are closely defined.
Suggestions as to how the psychologist might do this could include the following:
The researcher needs to specify the exact composition of the breakfast (possibly by
doing a pilot study or a literature search to identify the components of breakfast
most likely to bring about behavioural / cognitive change). He probably also needs to
specify the time at which it is consumed. The researcher needs to use a standard
reading test which should be administered to all the participants at the beginning of
the study and at the end – the dependent variable is likely to be the improvement
score.
(e) AO2/3 = 2
Reasons are:
• a test of difference
• data (scores from a reading test) are at least ordinal, this would include ordinal
/ interval and / or ratio
• independent design.
(f) AO2/3 = 2
It would have been more difficult to use a matched-pairs design because of the
number of relevant factors that would need to be controlled (eg gender, intelligence,
parental attitudes / income / education, experience of pre-school education, number
of siblings in family etc). There is a relatively small pool of children available (ie 400)
and it could be difficult to match on all these factors. It would also be very time-
consuming; it could be quite expensive to carry out the necessary surveys; it could
be quite intrusive collecting such information from parents.
(g) AO2/3 = 2
Page 93 of 116
One mark for identifying an appropriate issue and second mark for explaining how it
could be addressed.
The most likely issue is confidentiality which could be addressed by ensuring that all
scores on reading scales and all personal information are anonymised.
There are also ethical problems involved in denying the control group breakfast
although it is more difficult for candidates to suggest a way of addressing this –
perhaps to put only those children into the control group who do not eat breakfast
anyway, restricting the study length to a short period of time and, if the study results
support the hypothesis, to provide free breakfasts to these children for the rest of the
academic year.
Parental consent is excluded because it is given in the stem so answers which offer
this as an issue cannot gain credit.
(h) AO3 = 12
Design should be written clearly, succinctly and with sufficient detail for reasonable
replicability.
Candidates will not receive credit for details included in the stimulus material. These
include using a random sample of 100 children, gaining parental consent and
selection of a Mann Whitney test.
To access marks in the top band candidates must state an appropriate hypothesis in
which “playground behaviour” is clearly operationalised. The hypothesis could be
directional or non-directional.
Given the wording of the question, a correlational hypothesis is not credit-worthy,
however, the rest of the answer should be marked on its merits.
Likely aspects of “playground behaviour” would include activity levels, aggression,
cooperative play etc.
An attempt to operationalise “a healthy breakfast” should be credited. However,
candidates could assume this had already been done by the psychologist.
Page 94 of 116
9 – 7 marks Reasonable design
The design is reasonable and demonstrates knowledge and understanding of some
aspects of observational research. The selection and application of research
techniques is mostly appropriate. The description provides sufficient detail for some
aspects of the study to be implemented. Some design decisions are justified.
0 marks
No creditworthy material.
Q25.
Please note that the AOs for the new AQA Specification (Sept 2015 onwards) have
changed. Under the new Specification the following system of AOs applies:
(a) AO2 / 3 = 1
(b) AO2 / 3 = 3
If the candidate states that the result is not significant, no marks can be awarded.
Q26.
Please note that the AOs for the new AQA Specification (Sept 2015 onwards) have
changed. Under the new Specification the following system of AOs applies:
(a) AO2 / 3 = 4
Page 95 of 116
label each of the axes appropriately and plot the data accurately on the scattergram.
• the labels of the axes and the title taken together show full understanding of
the nature of the data.
Page 96 of 116
(b) AO2 / 3 = 4
For full marks, candidates should give a reasonably detailed explanation eg she is
concerned because the observers should both recognise the same types of verbal
behaviour as aggressive and you would expect their tallies to be very similar. In this
case, the observers disagree in every 10-minute time interval even though they are
both watching the same child and should be using the same criteria. In some time
slots, there is a really big difference in the number of acts.
This suggests that the observers have interpreted the criteria differently or that, at
certain times, one observer was more vigilant then the other (4 marks).
(c) AO2 / 3 = 3
1 mark for identifying the appropriate test – Spearman’s Rho or Pearson’s (with
appropriate justification).
2 further marks for explaining why it is appropriate ie the psychologist is testing for a
correlation and the data that can be treated as ordinal.
Candidates can gain no marks on this question if their choice of statistical test is
inappropriate.
(d) AO2 / 3 = 4
1 mark for a very brief answer eg ‘better training for the observers’
3 further marks for elaboration.
Page 97 of 116
There is a breadth / depth trade-off here. Candidates can elaborate on one
improvement eg explain how the training might be improved or outline several
improvements in less detail eg establish clearer criteria for categorising verbal
aggression, filming the child so that the observers can practise the categorisation.
Q27.
Please note that the AOs for the new AQA Specification (Sept 2015 onwards) have
changed. Under the new Specification the following system of AOs applies:
Although the essential content for this mark scheme remains the same, mark schemes for
the new AQA Specification (Sept 2015 onwards) take a different format as follows:
(a) AO2 / 3 = 2
One mark for one brief finding and a further mark for appropriate elaboration or for
two brief findings or one mark for a slightly muddled answer.
On average, the treatment group showed greater improvement after the treatment
than the no-treatment group. The average improvement score for the no-treatment
group was very low suggesting that the treatment gains for the treatment group were
not simply a result of the passage of time.
There was some variation in both groups as shown by the ranges but it was wider in
the treatment group. The low range in the no-treatment group suggests that most
people in this group had similar low improvement scores.
One mark for identification of a suitable test and 3 further marks for an appropriate
[Link] specification only requires knowledge of non-parametric tests.
However, if a candidate names an independent t-test and justifies its use, this is
perfectly acceptable. It is likely that most candidates will identify a non-parametric
test. The most appropriate test is the Mann-Whitney and the justifications for its use
are:
(c) AO2 / 3 = 2
One mark for correctly identifying the likelihood and one further mark for an
appropriate explanation or one mark for a slightly muddled answer.
The likelihood of making a Type 1 error is 5%. A Type 1 error occurs when a
researcher claims support for the research hypothesis with a significant statistical
test, but in fact, the variations in the scores are due to chance variables. If the level
of significance is set at 5%, there will always be a one in twenty chance or less that
Page 98 of 116
the results are due to chance rather than to the influence of the independent
variable or some other factors.
(d) AO2 / 3 = 4
Two marks for each reason. One mark for a basic identification and one further mark
for elaboration.
Expectations – the patients might expect the treatment to do them some good
and it becomes a self-fulfilling prophesy.
Other support – we do not know what other support/ treatment that the
participants might have had over the 8 week therapy period.
(e) AO2 / 3 = 4
Two marks for the advantage and two marks for the disadvantage. One mark for
simply identifying an advantage / disadvantage and the further mark for elaboration
in the context of the study. Answers which are not set in context cannot achieve full
marks.
Advantage: Much quicker to administer and to score – could all have been given out
at the same time whereas the therapist has to conduct 30 time-consuming
interviews; cheaper than interviews, ie in terms of the therapist’s time; people might
be more comfortable, and, therefore, more honest, if they have to write responses
rather than face an interviewer (could work the other way as well – see
disadvantages).
(f) AO2 / 3 = 5
• no pressure to consent – it will not affect any other aspects of their treatment if
they choose not to take part
• they can withdraw at any time
• they can withdraw their data from the study
Page 99 of 116
• their data will be kept confidential and anonymous
• they should feel free to ask the researcher any questions at any time
• they will receive a full debrief at the end of the programme.
For full marks, candidates must include a range of both procedural and ethical
points.
5 marks Effective
Consent form demonstrates sound knowledge and understanding of research
ethics.
4 – 3 marks Reasonable
Consent form demonstrates reasonable knowledge and understanding of research
ethics.
2 marks Basic
Consent form demonstrates basic, superficial knowledge and understanding of
research ethics.
1 mark Rudimentary
Consent form is rudimentary demonstrating very limited understanding of research
ethics.
0 marks
No creditworthy material is presented.
AO1: One mark for brief description, eg ’consistency’ and one further mark for
elaboration. Reliability refers to consistency over time. If a test, questionnaire, etc, is
reliable, people tend to score the same on the test if they take it again soon
afterwards.
AO2 / 3: One mark for a very brief answer, eg ‘do another test’ or ‘test them again’
or ‘use another interviewer to check’. Two marks for some elaboration.
The interviews could have been filmed and given to another trained therapist to
assess. A strong correlation between the scores given by each therapist would
demonstrate reliability.
(h) AO2 / 3 = 10
For full marks, the method section should be written clearly, succinctly and in such a
way that the study would be replicable. It should be set out in a conventional
reporting style, possibly under appropriate headings. Examiners should be mindful
that there are now different, but equally acceptable reporting styles. For example,
candidates should not be penalised for writing in the first person. The important
factor here is whether the study could be replicated.
• design
• participants
• materials
• procedures.
10 – 9 marks Effective
Effective method section that demonstrates sound knowledge and understanding of
investigation design.
The design decisions are appropriate and the description provides accurate detail of
the design, participants, materials and procedure of the study.
Effective and appropriate report style.
8 – 6 marks Reasonable
The method section demonstrates reasonable knowledge and understanding of
investigation design.
The design decisions are generally appropriate and the description provides
reasonable detail of the design, participants, materials and procedure of the study.
Generally appropriate report style.
5 – 3 marks Basic
The method section demonstrates basic knowledge and understanding of
investigation design.
Some aspects of the design are appropriate. The description provides basic detail
of some features of the study or rudimentary outline of the main features.
Expression lacks clarity.
2 – 1 mark Rudimentary
The method section demonstrates rudimentary knowledge or understanding of
research. The report is weak, muddled or incomplete.
Deficiency in expression results in confusion and ambiguity.
0 marks
No creditworthy material is presented.
Q1.
There were many correct answers to part (a) and it was apparent that the majority of
students had the knowledge to be able to identify the description of an overt observation.
In part (b) there were many correct answers and it was apparent that the majority of
students had the knowledge to be able to identify the description of an overt observation.
In part (c) most students appeared to have a good knowledge of the strengths and
limitations of the two types of observation. However, many failed to actually use this
knowledge to answer the question on discussing reasons why covert may have been
more beneficial than overt. There were many limited responses which either explained
limitations of overt observations or strengths of covert observations but without
comparison. Equally there were many responses which had implicit discussion of benefits
which was not clear and thus not sufficient for full marks.
Q2.
Part (a) was very straightforward and answered very well.
In part (b) many students were able to provide appropriate scattergrams, with accurate
title, axes, and plotting. Others missed out on one or two marks with vague titles and / or
axes. However some students, despite the questions in this section mentioning
‘relationship’, Spearman’s rho and ‘correlation’, provided completely inappropriate
graphical displays e.g. histograms and bar charts. A small minority did not attempt this
question at all.
Part (c) required reference to ‘level of measurement’. Although the majority of answers
could identify ordinal data for 1 mark, very few went on to characterise ordinal data or why
this studyproduced ordinal data, which would have fully justified the use of Spearman’s
rho.
Part (d) was very straightforward if the scattergram was plotted accurately, but a
significant minority of students were clearly unaware that 0.15 is a negligible correlation
and 0.95 a virtually perfect correlation (straight line).
Q3.
Part (a) was generally quite poorly answered. Although many students could gain some
credit for providing detail from the study and / or giving ethical issues, few were able to
demonstrate sound understanding of the requirements of a practical consent form with
many not referring to consent at all. Ethical issues were generally covered reasonably well
and were the strongest of the three elements, whilst format / style was the weakest and
was often inappropriate (eg ‘You will’, or ‘You must’, or ‘You have taken part’). Although
there were some excellent responses, generally it appeared that the students had not
really thought about the nature of giving informed consent, with sparse experimental detail
and the incorrect tone used, often reading more like a brief / debrief from or a legal
disclaimer. This is perhaps an area that schools / colleges could look at in more detail.
In part (b) despite many students showing an understanding of directional hypothesis, just
less than half failed to achieve any marks, mainly due to a focus on horror films leading to
more nightmares with no reference to the romantic comedy condition. Additionally, many
students failed to fully operationalise the variables and some struggled to write an
appropriate hypothesis for a repeated measures design, employing a writing frame for an
In part (c) most students identified individual differences as a reason to use a repeated
measures design but only about half of these students managed to clearly link their
reason to the experiment. Surprisingly, there was some confusion over the term ‘repeated
measures’, with some students focussing their answers on reliability and others on order
effects. Responses to this question again suggest that students need to gain more
practical experience of designing and conducting experiments as part of their course,
rather than focussing on theoretical learning of these concepts.
Part (e) was generally well answered with over half of the students accessing all 3 marks.
Generally, answers referring to a hat / container were more successful than those
involving a random number generator / computer, as these responses often failed to fully
describe the process.
Overall, students were able to answer part (f) well, although understanding of mean
scores was generally better than standard deviation. It appeared that students often
understood the data but failed to appropriately justify their suggestions, often simply
restating information given in the table. There were also a few costly examples of
justifications given without suggestions.
Part (g) was a good differentiator, with some excellent responses by students who clearly
explained what ‘significant at p<0.05’ meant and successfully managed to do so in the
context of the experiment. Weaker students generally gave a rote learned reason for why
psychologists generally use the 5% level or failed to contextualise their response.
Unfortunately, many students failed to explain that the ‘<0.05’ sign meant less than 5%.
Teachers should ensure that students are familiar with the mathematical symbols required
by the new specification.
Part (h) was generally not well answered. A huge range of modifications to the design of
the experiment were given but unfortunately a lot of these were inappropriate, such as
‘use a control group’, ‘use a blind procedure’, ‘use a repeated measures design’ or ‘use an
independent measures design’, suggesting either a lack of understanding of the reasoning
for experimental design or a failure to recall the experimental protocol. Some students
suggested using a matched pairs design and went on to explain effectively how this could
improve validity by reducing the likelihood of participants guessing the aim of the study
which would reduce demand characteristics. More commonly students justified the
modified design by inappropriately suggesting that it would avoid individual differences.
Changing the sample was a common answer but it was often not clear how this would
improve validity, with few being able to give more than a generic explanation. Stronger
responses generally focussed on the problem of using a text message to ascertain
whether a nightmare had been experienced, suggesting a questionnaire/interview to
reduce the chances of participants lying or to help distinguish nightmares from unpleasant
dreams.
Q4.
(a) Despite 36% of students not having the mathematical skills to answer this question
(b) Most answers provided an accurate statement of the likely outcome of the
experiment and could provide some sort of explanation. However, many students
missed the final mark by not providing accurate detail of the type of social influence
being displayed, especially in condition 2. Weaker answers were largely common
sense, lacking psychological terminology
(c) This was a demanding question as students had to explain how the researchers
might have addressed this issue in sufficient detail for 4 marks. Many chose
matched pairs as an appropriate method, but were vague about what participants
might be matched on, how matching might be carried out, and how participants
would then by distributed across conditions. There were, however, some very good
answers covering all these aspects. A significant minority chose repeated measures
as an appropriate method, but this was unlikely to be practical given the nature of
the study. Such answers earned credit only if they showed awareness of the need,
for example a very long interval between testing, the use of different but comparable
questionnaires, etc.
(d) Despite the injunction a number of students still referred to the level of
measurement as a reason for using chi-square. Otherwise this question was done
well, with many answers covering the requirement for independent data and a test of
difference (between conditions) or association (between variables). Answers
referring to correlation did not receive credit.
Q5.
(a) This question was well answered with over 90% of students correctly identifying the
type of experiment used in the study.
(b) Almost half of the students gained full credit for this question. Unfortunately, a few
gave the independent variable instead of the dependent variable, or got confused
with speed or the distance they had to run. Many students failed to state that the
running time was measured in seconds.
(c) This question was generally poorly answered, with 57% of students failing to
achieve any marks and 3% not attempting it. There was an occasional reference to
‘“the need for continuous data”’, but very few students gained the second mark.
Overall, there was a lack of understanding about what a histogram is or when they
should be used and schools/colleges need to address this. A common incorrect
answer was to assume a histogram is used for a correlation.
(d) This question was reasonably well answered. Unfortunately, a number provided a
full title instead of naming the type of graph, although when these included reference
to a bar chart they received appropriate credit. A number of the axis labels were
vague, omitting “seconds” or “mean/average”, or just writing “conditions”. There was
also some confusion regarding the type of graph with “scattergram” being the most
common incorrect response.
(e) Although there were some strong responses, generally students found this harder
than anticipated. A number failed to receive any credit due to simply defining the
(f) Despite 36% of students not having the mathematical skills to answer this question
appropriately and 3% not even attempting it, those who did have the knowledge
answered it well. Some missed or did not understand the requirement for three
significant figures.
(g) This question was generally answered very well, with the majority of students
achieving all five marks. Impressively, nearly all students identified the points in
mark scheme for choice of test and most could justify why the result was not
significant, although a few picked the wrong critical value from the table.
(h) Many students could correctly define a Type II error, although there was some
limited descriptions and some confusion between Type I and Type II errors. Few
students managed a coherent response to the second part of question, requiring an
explanation as to why the 5% level is normally used in research. Many were not
clear about the balance between making Type I and Type II errors, suggesting that a
5% level prevented these errors from occurring. Weaker students did little beyond
stating that it is used because of convention.
(i) From the stem of the question, a clearly uncontrolled variable was the type of music
participants listened to in condition B. Surprisingly few students identified this
extraneous variable but there were a variety of environmental factors given that
received credit, if they could have feasibly changed within one week and affected
the running times. Unfortunately, over half of the students appeared not to have read
the stem properly and therefore offered inappropriate extraneous variables, thus
gaining zero marks. The most common inappropriate variables were participant
variables, which would not have feasibly changed within the week, for example
running ability, fitness level, age or gender, or issues regarding order effects, which
would have been addressed by the design of the study.
(j) This question was generally quite poorly answered. Although many students could
gain some credit by defining what peer review is, very few demonstrated a detailed
understanding of theprocesses or the purposes of it, with the majority achieving
level 2. Many students made no reference to publication and saw it as a simple
checking method, often with the misconception that it involves the reviewer
repeating the study. Stronger students gave the process and then proceeded to give
its purpose, whereas weaker ones wrote all they knew, including pre-learned
evaluative points that did not gain credit. Despite some impressive answers that
showed detailed knowledge of both process and purpose, overall students had very
limited or no practical understanding of what peer review involves and teachers are
encouraged to address this.
Q16.
(a) This question required a definition of content analysis which proved challenging for
many students. Almost half of the answers achieved no marks at all. This was made
more remarkable by the fact that most were able to gain some marks on part (b)
where they were asked to explain how to carry out a content analysis for the data in
question.
(c) Most students were able to identify an appropriate method of testing the reliability of
the content analysis and collect at least one mark. The most popular answers were
test-retest and inter-rater reliability. Many failed to gain the three marks available as
their explanation of how the method of checking reliability would be carried out
lacked detail. A few students became side-tracked into improving reliability and a
small number used split half which was inappropriate in relation to content analysis
and gained no marks.
(d) This question required students to explain why a repeated measures design was
used in the experiment. Many students provided a basic answer referring to the
need for less participants or the removal of individual differences but were unable to
provide further explanation of why this would be important in this experiment.
Students who thought about the scenario and elaborated their explanation with
reference to reaction times, concentration or driving skills, achieved full marks.
(e) There was a broad range of answers to this question and about 75% of students
achieved no marks at all. Many students contradicted their previous answer to part
(d) and referred incorrectly to individual differences in reaction times and a similar
proportion referred to order effects which had been controlled by counterbalancing
or driving experience. Some students picked up on the possibility of differences in
the nature of the ‘chat’ on the phone which was encouraging. However, few students
showed any awareness of the need to match the two hazard perception tests
(stimulus materials / tasks) in this repeated measures design.
(f) Many answers to this question displayed a marked lack of common sense. Despite
referring to a simple hazard perception test, which is a key component of the driving
test, many students claimed that watching a 3-minute film of a road would be
traumatic, leading police drivers to suffer psychological harm. Others referred to
possible deception and failed to appreciate that the purpose of the experiment is
rather obvious in a repeated measures design. Better answers took issues such as
informed consent / right to withdraw and explained how these related to this
research.
(g) The question on writing instructions was answered well, with around half of students
achieving four or five marks. Some failed to gain full credit as their instructions
referenced both conditions or failed to include a check of understanding. Very weak
answers failed to refer to the conversation or made no reference to reacting as
quickly as possible.
(h) This question required students to identify an appropriate statistical test and justify
their choice. About one third of students gained the full marks for identifying the
Wilcoxon test with appropriate justification but just under half gained one mark only
for identification of the test. Common problems included justification as a test of
difference which gained no credit as it was included in the question. Other students
were confused about the type of data required for the Wilcoxon test and many
answers referred to ‘not nominal’ data.
(i) There is still evidence that few students understand the concepts of statistical error
and well over half failed to gain any marks here. Some became confused between
type 1 and type 2 errors and others referred to the number of hazards detected
rather than reaction times.
Q17.
(a) The majority of students gained the mark for this question by outlining the
relationship. A small number of students simply stated ‘positive correlation’ which did
not gain the mark as reference to the variables was required for an outline.
(b) Most students were able to name a correct test (Spearman’s or Pearson’s) although
not all could justify the choice of test with reference to the levels of measurement.
Some answers simply stated ordinal or interval, which was not enough as it was not
clear which variable was being referred to. One could argue that the empathy scale
was not equal interval and therefore should be treated as ordinal data. However, the
hours spent reading fiction per week could be treated as interval. Some suggested
an incorrect test e.g. Wilcoxon or Chi-square.
(c) Most students gained at least one mark on this question as they could outline a type
of validity. However, some just named a type which was not enough for even one
mark. Where students gained both marks, they generally referred to either face or
concurrent validity (although occasionally to predictive or criterion validity). For full
marks, answers had to explain how the type of validity would have been
implemented. Many failed to do this part of the question and thus gained only half
marks. Some students confused validity and reliability, often referring to split-half or
test-retest, and a noticeable number tried to answer the question with reference to
pilot studies.
(d) Most students were able to identify a limitation of the study and answered the
question well. The majority of the problems identified referred to the sample. A
number of answers referred to ethical issues and therefore gained no credit.
(e) Answers which simply stated that correlation looks for a relationship and
experiments investigate differences (or similar) gained only one mark. Students that
did access further marks were able to explain issues around ethics and manipulation
of variables in such a study.
(f) This question was answered well by those students who had clearly had practice at
designing and implementing their own investigations. It was pleasing to note that
some students gained full marks on this question and it was evident that some
schools and colleges had prepared students well. Students who did not gain high
marks on this question usually:
Q18.
(a) Hypothesis writing continues to be a problematic for many students, despite the
(b) Most students were able to identify an appropriate alternative method to collect data
about happiness, the most popular choices being interviews, observations and diary
studies. Some were able to provide a clear explanation of why this would be better,
but weaker students became side-tracked into describing the possible method in
detail (i.e. observational categories that might be used) and lost focus on the
question of comparison. Better students were able to refer to precise advantages of
their chosen method over questionnaires which were contextualised in relation to
measuring happiness.
(c) This question required students to give two reasons why Spearman’s rho was used
to analyse the data. Almost all students were able to accurately identify one reason,
which is encouraging!
(d) Weaker students still struggle to interpret critical and obtained values appropriately.
About a third of students managed to pick up the full 3 marks. Most of the remainder
were able to say the result was significant, gaining one mark but went on to select
the incorrect critical value from the table. A small number erroneously compared the
correlation coefficient (0.42) with 0.5 demonstrating misunderstanding throughout
the entire reasoning.
(e) This question required students to interpret a further correlation coefficient (which
was statistically insignificant) and put both pieces of information together to draw an
overall conclusion to the reported study. This proved challenging and less than 5 %
of answers achieved the full 4 marks here. Many students were able to identify that
the scores demonstrated different kinds of relationships (positive and negative) but
were unable to take this further and think about possible explanations. The better
answers focused on the inability to establish cause in correlational research and the
role played by other variables, in this case age. Some made use of their knowledge
about reliability which was creditworthy.
Q19.
(a) Hypothesis writing continues to be problematic for many students, despite the
requirement to do this at AS level. Around 40% of students achieved zero marks on
this question , having mistakenly written a directional hypothesis or one which
predicted a difference between mathematical ability and musical ability as opposed
to a relationship. Many responses lacked clarity or failed to operationalise the
variables sufficiently. The best answers were concisely and clearly worded such as
“There is a correlation (relationship) between pupils scores on a test of
mathematical ability and their scores on a test of musical ability”, which achieved the
full 3 marks.
(b) This question was answered well, with most students scoring two or all three marks.
Weaker students were able to spot the test was based on a subjective judgement
and some also made the point that singing was a poor measure of all round musical
ability. Stronger students identified the lack of control (different choices of song) and
were able to link this appropriately to investigator bias. Some students also made
the point that the test lacked validity as it had not been standardised.
The remaining 60% had some idea of ways of assessing reliability of the maths test,
the most common methods being test-retest and split-half. Some used inter-rater
reliability appropriately suggesting that two separate markers could be used for the
maths test: others became sidetracked into assuming that the study was
observational. Stronger students were able to explain two or three methods of
checking reliability in reasonable detail.
(d) This straightforward question on a random sample caught out quite a few students.
Most were able to achieve 1 mark by referring to the method as being likely to yield
a more representative sample. The weakest students simply defined random sample
and went no further.
(e) This question required students to draw a scatter graph to display the data. About
half achieved all three marks here. Many students failed to gain full marks by
inaccurate or missing labels or title. About one third of students drew an incorrect
graph, the most common error being to draw a bar chart.
(f) Most students were able to make some commentary on the generally negative
correlation shown in the graph and table. Better students noted the presence of two
outliers which weakened the overall strength of the relationship and some
commented on the impact of outliers in a small sample. A small number of students
made a rough calculation of Rs which was impressive but unnecessary to gain full
marks.
(g) This question had a range of answers from students that covered marks from 0-10.
The mark scheme allowed students to argue for different ways of designing the
experiment (independent measures or matched pairs) and of generating a sample
(volunteer or random selection from the two groups) provided these were workable
and justified. Some common errors included:
• assuming that a maths test also needed to be completed (ie incorrect IV)
Some schools and colleges had clearly prepared their students well and many
showed an impressive understanding of experimental design. Others struggled with
the question and/or, failed to read the instructions and therefore gained very few
marks. Once again, advice to teachers is: to do practical work. It was clear that
some students were very familiar with designing experiments and they had a strong
advantage here.
Q20.
(a) Students often struggled with this question. Very few understood why measures of
dispersion are used in addition to measures of central tendency and a number used
the term ‘dispersed’. Applications to the stem lacked the necessary detail to attract a
mark.
(b) Many students were able to make correct use of the table and draw an appropriate
conclusion about the statistical significance of the T value. Stronger students were
able to present this information well, with some even correctly stating that the results
were not significant at the 0.02 level and explaining why. A few less successful
students were confused about the critical and calculated values of T.
(c) It was heartening to see so many students being able to explain why this test was
used, with many students scoring the full 3 marks. A few students, however, stated
that the data was nominal or interval.
(d) Most students were able to score at least one mark on this question even if they
scored poorly on preceding and subsequent questions. Students had to think
carefully about this answer and many were able to suggest one or two sensible
reasons for the use of a diary by each offender. Elaboration and explanation of each
reason proved more of a challenge, sometimes resulting in overlap and repetition
across the two reasons offered. Some students simply repeated the stem.
(e) This question produced some answers that showed a lack of understanding of the
use of a control group with quite a few students suggesting that a control group
would consist of ‘a group of non-offenders’ or ‘normal people / people with no anger
issues from a normal population’. More informed students were able to explain why
a group of people who would not have the anger management programme would
have improved the study although some simply said that ‘it would make it more
scientific’ without explaining why.
(f) Answers to this question were most disappointing with almost two fifths of students
failing to score a single mark. What should have been a straightforward question
proved challenging for many students who seemed unfamiliar with how to establish
the reliability and validity of the questionnaire. Some students confused reliability
with validity, others simply provided definitions of each and, where an attempt was
made to apply reliability and validity to the stem, students very often referred to pilot
studies and peer review. Reliability was sometimes mistaken for replicability, the
‘split-half method’ was frequently explained as dividing the total score in half to see if
the two halves correlated and test- retest was sometimes explained as testing one
group of offenders and retesting another group of offenders. Face validity was rarely
applied to the questionnaire ie anger scores. Even students who correctly
addressed the issues of reliability and validity, failed to explain how statistical tests
of correlation would be used in this context. The value of carrying out practical
activities to enable students to ‘think like a psychologist’ and apply their knowledge
of practical activities in answer to questions such as this one cannot be overstated.
It is clear that where students had been presented with such opportunities, they
Q21.
(a) This question was answered well with most students aware that a directional
hypothesis was appropriate due to the existence of previous research. A minority of
students provided rather more detail than required for one mark.
(b) Hypothesis writing is still a problematic area for many students, despite the
requirement to do this at AS level. Many students achieved zero marks on this
question, having mistakenly written a non-directional hypothesis or one which
predicted a difference between older and younger patients. Many responses were
lacking in clarity or failed to operationalise recall adequately. The best answers were
concisely and clearly worded such as “There is a negative correlation (relationship)
between age and recall accuracy rating”, which achieved the full three marks.
(c) Although this question was worth only one mark, many students produced lengthy
answers. Some distinguished between specific types of reliability such as external or
internal. A small number of students became confused between validity and
reliability.
(d) There was a broad range of answers to this question, with students in roughly equal
measure being awarded marks across the full range. The majority had at least a
rough idea of ways of assessing reliability (the most common being inter-rater) but
found it difficult to select an appropriate method for the study detailed. The weakest
answers were those where the student focussed on reliability of the study overall,
rather than reliability of the ratings which was what the question required. Answers
that achieved the full three marks generally selected the most straightforward idea;
to take two independent psychologists who rated the typed accounts separately and
then correlated their ratings. Students who achieved only one mark suggested test
retest as a method but most were unable to carry this through and indicate that the
psychologist would need to return to the data after a suitable interval and re-rate the
accounts.
(e) This question was answered well with the majority of students achieving two marks.
There was a range of both kinds of data to draw on here including the doctor’s notes
and the patients responses (qualitative data) and ages and accuracy scores
(quantitative data).
(g) This question confused many students who were unaware that the critical value
relates to the magnitude of rho not the direction. So negative correlation drops the
minus sign when compared with the critical values. About half of students were
clearly aware of this and could compare the obtained value with the correct figure
from the table. The remainder made a number of errors, some comparing -.52 with
0.05, others claiming that the figure was smaller than .306. Some incorrectly used
the values relating to a non-directional hypothesis.
(h) Full marks were achieved by stating that the null is rejected and the experimental
hypothesis accepted, when in fact results are due to chance. Good understanding
was shown among students who referred to the level of significance being set too
leniently or the 5% likelihood of a Type 1 error occurring with the 0.05 level of
significance. In about one in three cases students confused Type 1 and Type 2
errors.
Q22.
(a) In this question, students were required to discuss the advantages of carrying out
the experiment described in the stem, in a laboratory. Fewer than half of students
made any reference to the stem and the most common mark awarded was one out
of four. Those who referred to an advantage (eg control of extraneous variables) and
linked it appropriately to the scenario (eg posters on the walls) were able to access
the full range of marks. A small but significant minority insisted on writing about
disadvantages and achieved no marks. Once again, schools and colleges should
advise students to read stems carefully and apply knowledge in Section C.
(b) Most students achieved full marks, identifying the Mann-Whitney as the appropriate
test and giving and ordinal data or independent groups as a reason. Some students
provided two or three reasons going beyond the requirements of the question. There
were a minority of cases where an incorrect answer was given, most commonly
Spearman’s rho or Wilcoxon’s signed ranks test.
Q23.
(a) Hypothesis writing is still a problematic area for many candidates – despite the
requirement to do this at AS level. Many candidates achieved zero marks on part
(a), having mistakenly written a directional or a null hypothesis. Many responses
were lacking in clarity or failed to include an operationalised DV so only achieved 1
mark. The best answers were concisely and clearly worded responses such as
“There will be an association between birth order and career choice”, which
achieved the full 2 marks.
(c) Some centres had clearly prepared their candidates very well and many showed an
impressive understanding of inferential statistics scoring 11 or 12 marks. However,
other candidates struggled with the question and collected very few marks. Some of
the most common errors were as follows.
A number of candidates did not know how to express the statistical conclusion of a
research study, by referring to observed and critical values and probability. There
were errors in correctly identifying the observed and critical values and their
relationship to the hypothesis. A large number of candidates did not label the axes of
the graph or only showed data relating to first born career choices.
Some candidates chose the wrong statistical test; some did choose the correct
statistical test but did not then state the reasons why the test was appropriate.
Yet again, advice to teachers is to do some practical work. It was clear that some
candidates were very familiar with the rationale for selecting a test and deciding if an
Q24.
(a) A challenging question because candidates needed to apply their knowledge. They
often knew about what makes something scientific (objectivity, replicability, etc) but
seemed unable to engage with the stem. There were lots of answers involving
paradigm shift, which were not relevant to this question.
(b) Most candidates had some idea about how a random sample could be obtained, but
often failed to explain the methods fully. They could suggest all the names should be
put in a hat, but did not make it clear that the names were then selected “without
looking” or “without bias”. There was some confusion with systematic sampling.
(c) Many answers displayed some confusion here, eg saying that a limitation was that it
was not representative of the whole population, when the point is that is might not
be representative of the target population of 400. Some answers referred to
problems of allocation to conditions, rather than random sampling. A good point was
made by those who said that if some parents did not give consent, the psychologist
would have to select again, and that would not be random.
(d) This question was not answered well. Most candidates seemed very unclear about
why it is important to operationalise variables. How to actually operationalise the two
variables was beyond many candidates. Some effective answers referred to food
content eg fat, sugar etc.
(f) There was some serious confusion about what exactly matched pairs design is. Few
could go beyond “it’s time consuming” or “difficult to match on all variables”. Some
referred back to the random sample and said it would not be possible; others felt
that at five-years-old children are either too similar to match or too different.
(g) Most could identify an ethical issue such as confidentiality, the right to withdraw and
protection from harm (those who did not get any breakfast or who were
embarrassed at their poor reading). Some seemed to forget that they also had to
explain how the issue would be dealt with, or they simply repeated that the right to
with draw could be dealt with by giving the right to withdraw.
(h) This question was not answered well. Many candidates failed to read the question
carefully before they attempted it. They were given the information that they were
using the same group of children (ie the 5-year olds in the previous study). Despite
the fact that the ethical issues and sampling had already been addressed in the plan
for the original study many wrote at great length about sampling and ethics. The
majority of candidates were unable to write a fully operationalised hypothesis, and
often simply restated the aim. Many seemed to think the IV was breakfast versus no
breakfast, rather than healthy versus unhealthy breakfast. Some of their ideas were
totally impractical, especially given that the children were only 5 years old. In many
answers lack of detail would have made any kind of replication very difficult.
However, some candidates did understand the need for some sort of training for the
observers, the need for clearly identified behaviour categories to record, and the
importance of being able to distinguish the two groups in the playground. Designing
a study is clearly a difficult task for candidates, and one that they need to practice.
Q25.
(b) Many candidates clearly understood how to read the table and to interpret results
and so gained the full 3 marks here. Some gained 1 mark for saying that the result
was significant but then demonstrated a complete lack of understanding in the rest
of their answer.
Q26.
(a) This question proved to be a good discriminator. Candidates who understood
scattergrams were able to make a reasonable sketch with appropriate labels and
accurately plotted data and so gained full marks. However, a disappointingly large
number of candidates clearly had no understanding of scattergrams and drew a
frequency polygon instead for which they could gain no marks. The requirement to
present and understand graphs is clearly stated on the AS specification:
‘presentation and interpretation of quantitative data including graphs, scattergrams
and tables.’
(b) Some candidates gave full answers in which they made good use of the data
contained in the table. However, fewer candidates were able to make use of the
information in the scattergram and very few referred to correlation. There were 4
marks available for this question which should have made candidates realise that
some detail was required. Answers such as ‘she was concerned because the
observers gave different ratings’ could not gain much credit. Quite a few candidates
wasted time by defining inter-rater reliability. Answered included suggestions of how
to improve reliability which, of course, was addressed in part (d).
(c) Relatively few candidates identified an appropriate test - almost every reasonably
familiar test was quoted. Experimental designs were often quoted as incorrect
reasons for test selection. Many candidates did not even suggest an inferential test
but suggested calculating the range, mean or standard deviation. Candidates who
did identify the appropriate test were usually also able to offer an appropriate
justification.
(d) This was a good discriminator. Most candidates could offer at least one solution to
this issue but many stopped after making their initial point eg ‘give them more
training’. Some were able to elaborate on this effectively to gain full marks but many
showed little understanding. Very common errors were ‘get more observers’ or
‘average the results’ or ‘only use one observer’.
Q27.
Perhaps surprisingly, many candidates achieved higher marks for this section than for
their other two questions. Most candidates attempted all parts of the question although a
significant minority did not complete part (h) suggesting that they might have run out of
time. This was a pity since this part carried 10 marks.
(a) This was a straightforward question and many candidates accessed full marks but a
surprising number were confused by median and range and some did not
understand what the range indicated about the data.
(b) Many candidates were very well prepared and got full marks here but some wrote
very confused answers showing little understanding eg ‘Spearman’s Rho because it
was nominal data and repeated measures’.
(c) There was a centre effect here. Some candidates had a good understanding of Type
(d) Candidates offered a wide range of answers, although some were a little bit too brief
or poorly explained to get both marks. There was some confusion about what is
meant by a placebo and some candidates offered two explanations which were
essentially the same as one another. Other candidates offered factors which could
apply equally to the treatment and non-treatment group. It is important in this kind of
question to read the stem carefully. Some candidates said that the therapists might
have been biased in favour of the treatment group, but the stem clearly states that
the therapist did not know who had been in which group.
(e) A lot of candidates missed the point that the advantage / disadvantage needed to be
in comparison to interviews. Many candidates gave advantages/disadvantages that
could apply equally well to both self-reports and interviews. This was acceptable
only if the candidate made it clear ie ‘People are less honest in a questionnaire’ was
not credit-worthy because it could apply to both interviews and questionnaires.
However, ‘People are less honest in a questionnaire because they are anonymous
and feel they can lie about themselves without being found out. In an interview
where they are face-to-face with the interviewer, they might find it more difficult to
lie’.
(f) Some candidates wrote excellent consent forms containing both ethical and
procedural information and expressed them in appropriate language. Some
candidates had a very vague understanding of what needed to be included here,
either only focusing on all the ethical issues (you will have the right to withdraw your
data, yourself etc) with no mention of the procedures, or vice versa. Many adopted a
rather inappropriate tone eg ‘Once you have signed this form, you are committed to
being in the study' or ‘You have to subject yourself to an interview’. It was surprising
to see that a few candidates seemed to think a consent form acted as some kind of
legal disclaimer – ‘you may suffer harm but if you sign this you can’t sue us.’
It was notable on this question that candidates who were able to express
themselves clearly and succinctly were much more likely to access full marks. Many
answers were so poorly constructed that the content was difficult to understand.
Many switched confusingly between pronouns eg They will have to have an
interview. You can withdraw at any time. I agree to be part of this study.
(g) There were some very muddled answers to this question. Candidates often didn’t
read the question carefully, and wrote something like ‘Reliability means if you do the
study again you will get similar results’ for their definition and then didn’t know what
to write for the next part of the question. Those candidates who explained it in terms
of inter-rater reliability generally gained full marks. Some candidates did not read the
question carefully and did not relate their answer to checking the scores in this
particular study. Many candidates thought incorrectly that test-retest involved using
different participants. Some candidates suggested split-half methods indicating a
lack of thought about the question. Some candidates confused reliability with
validity.
(h) Many candidates showed limited awareness of a conventional reporting style. While
it was not necessary to divide the method section into sub-sections, this strategy
might have helped candidates to include all the relevant details. Weaker answers
made no mention of gender or eating disorders and simply repeated details from the
stimulus material. Many candidates completely lost sight of the fact that gender
differences were being investigated and suggested randomly allocating participants
to groups. A lot of time was wasted in including aims / hypotheses and statistical
A Type II error occurs when a researcher fails to reject a null hypothesis that is false. Psychologists generally use a 5% significance level because it balances the risk of Type I errors (false positives) with statistical power reasonably well, providing a conventional standard for determining statistical significance .
The percentage decrease is calculated as follows: ((123 - 117) / 123) * 100 = 4.878%. Therefore, there is a 4.878% decrease in the mean time when participants listened to music .
The trade-off between breadth and depth in research involves the balance between exploring many aspects broadly or delving deeply into fewer areas. For verbal aggression, going in-depth might involve developing detailed criteria for categorisation and training observers extensively, potentially increasing validity and reliability. Alternatively, addressing several types of aggression more broadly might provide a wider understanding but with less precision and reliability in each category .
The mean values indicate that participants ran faster with music (mean = 117 seconds) compared to without music (mean = 123 seconds), suggesting that music might enhance running performance. The higher standard deviation in the music condition (14.5 seconds compared to 9.97 seconds without music) implies greater variability in how music affected different individuals' performance, which might be due to individual differences in response to music .
Students could conduct a content analysis by first listening to recordings to identify recurring themes or concepts. They would then create categories for these themes and systematically tally occurrences, ensuring to operationalise clearly defined criteria for each category. This method ensures objectivity and replicability by reducing subjective interpretation .
The operationalised dependent variable in the study is the time taken to complete the 400 meters as measured in seconds. Operationalising variables is necessary because it allows researchers to measure abstract concepts precisely and consistently across different participants, ensuring the reliability and validity of the results .
A histogram would not be appropriate for displaying the means because histograms are typically used for showing distributions of frequencies, not comparisons of central tendencies like means. A bar graph would be a more suitable choice to compare the means of the two conditions, with the conditions (with and without music) on the X-axis and mean times on the Y-axis .
The results are not statistically significant as the calculated t-value (1.4377) is less than the critical value at the 0.05 significance level for df = 9 (1.833). A related t-test was used because the same participants were measured under both conditions, allowing for control of participant-related variability, which enhances the power of the test .
The study used a laboratory experiment. This type was chosen because it allows for controlled conditions, ensuring consistent variables such as environmental settings and instructions, which helps in obtaining reliable data about the effects of listening to music on running performance .
One potential extraneous variable could be the participants' prior physical activity level or fitness. If not controlled, differences in fitness could skew the results, making it seem as though music affects performance when the effects could actually be due to differing fitness levels rather than the experimental manipulation .