0% found this document useful (0 votes)
2 views15 pages

Statistics Notes

Statistics is essential in psychology as it helps interpret research findings and enhances critical consumption of psychological information. Descriptive statistics, including measures of central tendency and variability, allow for meaningful data presentation, while understanding sample versus population is crucial for generalizing research results. Various types of variables and data, along with effective graphical representations, are important for accurately conveying and analyzing psychological data.

Uploaded by

jadie.leigh24
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
2 views15 pages

Statistics Notes

Statistics is essential in psychology as it helps interpret research findings and enhances critical consumption of psychological information. Descriptive statistics, including measures of central tendency and variability, allow for meaningful data presentation, while understanding sample versus population is crucial for generalizing research results. Various types of variables and data, along with effective graphical representations, are important for accurately conveying and analyzing psychological data.

Uploaded by

jadie.leigh24
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as DOCX, PDF, TXT or read online on Scribd

STATISTICS MODULE 2 PSYCHOLOGY

Importance of statistics in psychology


Statistics forms an integral part of the field of Psychology. Much of the work we conduct, analyse, and use
within the field of psychology is based upon research. Your foundation of statistical knowledge will allow you
to make better sense of research.
Often social media publishes stories about the latest scientific findings, self-help books make proclamations
about different ways to approach problems, and news reports interpret (or misinterpret) psychological research.
By understanding the research process, including the kinds of statistical analyses that are used, you will be able
to become a wise consumer of psychology information and make better judgments of the information you come
across.

Descriptive statistics: Central


tendency and variability
Descriptive statistics are very important because if we simply presented our raw data it would be hard to
visualize what the data was showing, especially if there is a very large dataset (as you will see in your Research
Practical Assignment). Descriptive statistics therefore enables us to present the data in a more meaningful way,
which allows easier interpretation of the data. There are two main types of descriptive statistics
1. Measures of Central Tendency
A measure of central tendency is a single value that attempts to describe a set of data by identifying the central
position within that set of data. Based on the type of data you have, you will use different measures of central
tendency.
2. Measures of dispersion
Measures of dispersion tell us how far apart our data lies (i.e., the spread of our data)/ how data is spread
 Range: This is the difference between the largest and the smallest observation in the data.
 Standard deviation/variance: This is the most commonly used measure of dispersion. It is a measure
of spread of data around the mean.

Sample Vs Population
{ When we ask a research question, often we are interested in a population (so for instance, everyone in
first-year psychology at UCT).
{ But to save resources, we don’t survey the whole population, we survey a subset, which we call a
sample.
{ From a sample – infer about broader population
{ How well you choose that sample determines how close you are to the population:
- For research design, randomly selecting a big enough sample will mean that your sample can
generalize back to (be representative of, be very close to) your population.

Vocabulary and symbols – central


tendency
• x is each score (any individual score) – a variable
• Mean = average / arithmetic average = x (with a line on top of the x). Can use for continuous data:
marks, income, GDP
• Mode: the most frequently occurring number. Used with frequencies of categorical variables. Bimodal
data = 2 modes
• Median: the number in the middle of a data set that has been ranked in ascending order/ descending
order. Most often use the median when we have SKEW DATA
• Sigma - ∑ - sum, add up
STATISTICS MODULE 2 PSYCHOLOGY

Meaning Symbols

Population Sample

Sample size N n
(number of
participants)
Mean (average) µ (mu) X (with line above x
– x bar)
IMPORTANT NOTE: Averages always reported to 2 decimal places, even i.e., 6.00

Why three measures of central


tendency (Mean, median, mode)
Median most accurate with skew data – Mode and mean either under or over represent
à Right skewed/ positive skewed = a long tail off to the right, peak to the left rather than
having the peak in the center.

à left-skewed/ negative skewed = there is a long tail off to the left.

Above graphs example


In the left-skewed distribution, the mode, the median and the mean are all quite different, and that tells us that
the distribution is left-skewed.
- The median is a better measure here, for telling us where the “centre” of the distribution is.

In the graph on the right, we have what’s called a normal curve. Here the mean, the median and the mode are
all the same.

More vocabulary – dispersion/variability


• Range: difference between highest and lowest numbers/ scope of data
• Variance: indicates how far, on average, each score is from the mean (it uses the square of the
distance).
• Standard deviation: square root of the variance
Never report one measure of central tendency alone – have to report on the spread of data.
Variance and standard deviation also give information about the shape of our data.
STATISTICS MODULE 2 PSYCHOLOGY

Adding meaning with the range


{ Suppose we have asked parents how often in the last month they have spanked their children when the
children did something wrong.
{ If we say the answer has a mean of 4.2, this has no meaning, unless you know what the range was.
{ Knowing that the range could be from 1-5, we can see that on average parents spank their children
often.
{ If we know that the actual range in the sample was 3-5, this means that all parents are spanking their
children, at least some of the time.
{ If we know that the actual range was 1-5, then we know that at least some parents were never spanking
their children

Formulae
• Population variance:
Σ ( x−μ ) 2
• σ 2=
N
• Sample variance:
Σ ( x−x ) 2
• s 2=
n−1
*NOTE – 2nd x in sample variance is mean (has line above)

{ Small SD
- Most values are close to the mean
- Mean is a very good representation of our data, most data quite close to that mean
{ Large SD
- The data is very spread out – lots of values away from the mean
- Mean no longer a good measure of central tendency
STATISTICS MODULE 2 PSYCHOLOGY

Calculating variance
x (x-mean) (x-mean)2
7 1 1
9 3 9
10 4 16
2 -4 16
3 -3 9
5 -1 1
X= 6 Σ=52

52
S2 = = 52/ 5 = 10.40
6−1
1. list your x values, and then calculate the mean. In this list of made-up values, the mean is 6.
2. subtract the mean from each individual x value; then square those values and sum them. In this case,
the sum of these values is 26.
3. Then we’ll pretend that this is a sample (not a population), and use the formula for the standard
deviation for a sample: and you can see that s = 3.2.

Box and whisker plot


{ Unlike SD which uses mean, box and whisker plot built using the median
{ Shows quite a lot of data, including the variation around the measure of central tendency.
{ The lower quartile is the median of the lower half of your data, and the upper quartile is the median
of the upper half.
{ They are called quartiles, because they divide the data set into four, that is, four quarters. The lower
quartile is also called the 25th percentile (because 25% of the data is lower than this number), and the
upper quartile is also called the 75th percentile (because 75% of the data is lower than this number).
{ The distance from the lower quartile to the upper quartile = interquartile range. (Q3 – Q1).
- Tells us where the middle 50% of data falls
- Good indication of where middle 50% is clustered – can be used instead of SD in some
situations
{ Outlier = data point very different from the rest of you data. Often calculated using the SD – certain
amount of SDs from the mean.
STATISTICS MODULE 2 PSYCHOLOGY

Here is a box and whisker plot looking at variability in hours spent sleeping during the week. In this study,
each participant wrote down how many hours they slept on Monday night, Tuesday night, etc.
{ On Monday night, the median was 7.5 hours of
sleep, with 50% of the people getting between 6 and
8 hours, and in the whole sample, between 4 and 9
hours of sleep.
{ You can also see that the median values decreased
Mon-Thursday, and that on Friday most people
slept more than any other day of the week. You can
also see that there was quite a bit of variation: on
Monday sleep times varied from 4-9 hours; on
Satureday from 6-11 hours.
{ You can also see some outliers on the graph: on
Tuesday one person got 10 hours sleep
(more than anyone else in the sample); on
Thursday one person got 3 hours’ sleep (far
less than anyone else); on Friday one
person got no sleep at all; and on Sunday,
one person got 11.5 hours sleep.

Graphs and
tables
Graphical representations allow the reader to get a
better sense of what the data means. By representing
this visually, the reader can easily see, interpret, and
understand the data being presented.

For example: Let's suppose that 200 students are


recruited from two different classes (a, b) and
randomly assigned by a researcher to two
independent experimental conditions (x, y). In each
condition, subjects' performance on a specific
experimental task is assessed. The research
hypothesis is that there is a significant difference
STATISTICS MODULE 2 PSYCHOLOGY

between the two experimental conditions. Look at the graphical presentations below. In Figure 1A, number of
subjects belonging to each class is depicted: 104 subjects belong to class a, and 96 to class b. The height of each
bar correctly represents frequency of subjects in each condition and, therefore, bar chart is informative and
pertinent. In Figure 1B, using the same graph, subjects' mean scores (and associated standard errors) in
condition x and y are depicted: 1.8 and 1.73 respectively.

Representing the same data using a box plot (Figure 1C) and a histogram (Figure 1D), we end up with a
different conclusion. Condition x shows a skewed distribution while, on the contrary, data from condition y are
more symmetrically distributed, suggesting that the two experimental conditions are not equivalent. The
research hypothesis is now supported. Without the use of graphs as reasoning tools for exploring data, it
wouldn't have been possible to detect this difference in the two experimental conditions and the
researcher would not have supported their research hypothesis.

Variables and data


Four types of variables
{ Nominal: Type of categorical data where the number is just a name, a label for something and it
doesn’t have meaning in itself (e.g., if we assign “1” to smokers and “0” to non-smokers, there’s no
real meaning to the numbers – they could just as easily be 55 and 92). Can’t calculate mean etc.
{ Ordinal: Type of categorical data where the order of the number matters, e.g., movie ratings (a 5-star
movie is better than a 1-star movie). I.e., if you rank the runners in a race, and convert seconds each
runner to finish (continuous data) into a place, you convert continuous to categorical.
{ Interval: scores like intelligence, where there is a number but “zero” is meaningless. Can add, subtract,
compare greater than and less than etc. However, can’t multiple and divide – no ratio with no true zero.
{ Ratio: numbers where zero has meaning (weight, money). Can add, subtract, compare greater than and
less than etc. Can multiple and divide
Types of data
{ Categorical data, from nominal variables and ordinal variables:
• E.g., how many people answered the questionnaire in English? In isiXhosa? In
Afrikaans?
• The data are in separate categories that don’t overlap.
{ Continuous data, from ratio variables and interval variables:
• E.g., how much do secondhand cars cost? Anything from R20,000 to R1,000,000.

Categorical data
Bar charts
• Are used with categorical data
• Can show the number of observations (frequency) in each category (Count the number of occurrences
each event happened)
• Can also show a continuous variable on the Y-axis
• The categories are shown on the x-axis

This graph shows school enrolment


in millions

This data tells a story: that there are


more than three times as many
children in primary school than in
secondary school, and half as many at
university than in secondary school.
There is an even bigger disparity
(difference) between those in primary
STATISTICS MODULE 2 PSYCHOLOGY

school and those in ECD – and this is very worrying, because ECD can really boost the ability to achieve in
education.

Shows distribution of data very clearly

Comparing two data sets


Brazil has more children in school, but the gaps
between ECD and primary school are smaller – in
Brazil, for every 3 children in primary school, there is
1 in ECD, whereas in South Africa the ratio is 1:9. In
Brazil, there is only a small difference between the
numbers of children in primary and in secondary
school.

So drawing the graph can tell a vivid story. But can


you spot a problem with this graph?

Brazil has a bigger population than South Africa, so


this is in a sense an unfair comparison.

This is fairer: it gives us the percentage of the total


population at each school level. Now you can see that we
do much better at primary school enrolment than Brazil
does, but that we are indeed lagging on ECD, secondary
school, and university.

There’s a further problem here, though:


We should really calculate these numbers as a
proportion of those eligible for each level of education,
rather than as a percentage of the total population. So,
thinking critically about each graph is important.

Important notes about drawing bar charts


STATISTICS MODULE 2 PSYCHOLOGY

These are two graphs of the same data.


F The one on the right doesn’t start at zero, and so it distorts the proportions.
F Always start your bar graph at zero.
F Graphs that are misleading in this sense can fool people – used maliciously.

Different formats

Example of a stacked bar graph, which shows you its power


This graph looks at the total health spend per person in 2018 US dollars. The data comes from 2016 (on the
left), and what is projected (on the right), if nothing changes. You can see that spending from all sources is
expected to increase, with that of government nearly doubling (increasing the most).

F Shows us that we will be spending


more on health in total in the future.
F Shows proportions of each category
of spending - tells us about four
different kinds of spending and the
overall spending.
F Breaking up one piece of information
into components
F Useful when data is mutually
exclusive
STATISTICS MODULE 2 PSYCHOLOGY

Continuous data
Organizing data
Frequency count
I.e., to count how many cars have the same price.
The things to pay attention to
- The frequency (the price R1,600 appears once in the dataset, the price R1,900 appears once, and so on)
- The percent (1 item in the dataset is 1.9% of the data, so 1.9% of cars have this value; the value R2,400
is held by 2 cars, and 2 cars are 3.8% of the data, and so on)
- Valid percent – this counts how much actual data you have (in some datasets, some data is
missing/incorrect)
- Cumulative percent – 13.2% of the data lies between R1,600 and R2,500. Shows the percent of data
that lies below a certain value. The percent up until a certain point.

Advantages of frequency counts


ü Organizes data
ü Ordered from smallest to biggest
ü Cumulative frequency shows median – median at 50%

Problems with frequency counts


û Can hardly read, because it’s such a long list.

A frequency count, while messy, can start to tell you a story.

Price
Frequency Percent Valid Percent Cumulative Percent

Valid 1600 1 1.9 1.9 1.9


1900 1 1.9 1.9 3.8
2400 2 3.8 3.8 7.5
2500 3 5.7 5.7 13.2
2600 2 3.8 3.8 17.0
2700 5 9.4 9.4 26.4
2900 2 3.8 3.8 30.2
3000 6 11.3 11.3 41.5
3100 3 5.7 5.7 47.2
3400 2 3.8 3.8 50.9
3500 1 1.9 1.9 52.8
3700 2 3.8 3.8 56.6
3900 1 1.9 1.9 58.5
4000 1 1.9 1.9 60.4
4200 1 1.9 1.9 62.3
4400 2 3.8 3.8 66.0
4500 3 5.7 5.7 71.7
4600 2 3.8 3.8 75.5
4700 3 5.7 5.7 81.1
4800 1 1.9 1.9 83.0
5200 2 3.8 3.8 86.8
5300 1 1.9 1.9 88.7
5400 1 1.9 1.9 90.6
5600 1 1.9 1.9 92.5
STATISTICS MODULE 2 PSYCHOLOGY

6600 1 1.9 1.9 94.3


7000 2 3.8 3.8 98.1
10700 1 1.9 1.9 100.0
Total 53 100.0 100.0

If you categorize your data, it’s much more helpful.


This table organizes the data into ranges/bins of R1,000. This starts to get a bit more interesting and helpful:
ü Now you can see that, in 1984, most secondhand cars cost between R2,000 and R4,999.
ü Quick and easy to interpret!
Price range Frequency Percent Cumulative percent
0-999 0 0 0
1,000-1,999 2 3.8 3.8
2,000 – 2,999 14 26.4 30.2
3,000-3,999 15 28.3 58.5
4,000 -4,999 13 24.5 83.0
5,000-5,999 5 9.4 92.4
6,000-6,999 1 1.9 94.3
7,000-7,999 2 3.8 98.1
8,000-8,999 0 - 98.1
9,000-9,999 0 - 98.1
10,000-10,999 1 1.9 100
Totals 53 100

Histograms
{ Similar to bar graph, except x-axis is continuous (no gap between categories)
{ With continuous data, we have to organize it into categories/bins before we can graph it.
{ If don’t use categories correctly – see spikes/ randomness
{ Trend lines help with ‘noisy’ data (data with a bit of randomness) – smooth out bumpiness, therefore
more accurate than graphed data.

Advantages of histograms
ü Gives a sense of spread and distribution quickly (often used to determine normal distribution) – see
skew of data.
ü Easy to spot outliers. Important because we want to understand why something is an outlier – was
something genuinely unusual or was there a mistake?
- i.e., was it an unusually valuable, or did someone make a mistake when they entered the
data? Why are there are no datapoints between 8,000 and 10,000: is data missing here?
ü See where majority of data points fell – modal class/ category.

Shapes of data/ Different types of distribution


{ Unimodal symmetric = normal distribution
{ Bimodal = two different distinct
groups/ 2 modes (common values).
I.e., height, one mode might be for
men, other for women.
STATISTICS MODULE 2 PSYCHOLOGY

Line graphs
Example
• Suppose we are testing a supplement that is supposed to enhance memory:
– We have a sample that we divide randomly into two groups:
– One group gets the medication
– The other gets a placebo
– We test their memory performance before and after administration of the drug
• Our memory scale runs from 0 (terrible) to 10 (excellent).

Average score of groups before and after drug:


Before After
Treatment group 5 10
Placebo group 1 5

What to note
F Line shows trend
F Points = data points/ means
F Direction of lines = increase/ decrease
F Put in confidence intervals/ error = points in this graph represent average. Need to show indication of
error around individual points.

Correlation
A correlation is a statistical measurement of the relationship between two variables. Correlational studies are
quite common in psychology, particularly because some things are impossible to recreate or research in a lab
setting.

Instead of performing an experiment, researchers may collect data from participants to look at relationships that
may exist between different variables. From the data and analysis they collect, researchers can then
make inferences and predictions about the nature of the relationships between different variables.

Possible correlations range from +1 to –1. A zero correlation indicates that there is no relationship between the
variables. A correlation of –1 indicates a perfect negative correlation, meaning that as one variable goes up, the
other goes down. A correlation of +1 indicates a perfect positive correlation, meaning that both variables move
in the same direction together.

• Empirical association between two variables – statistical evidence that two variables are related.
• Covariation is a statistical term for the concept that if one variable changes, the other also changes.
How the spread of two variables are related.
STATISTICS MODULE 2 PSYCHOLOGY

• Always paired data: each data point corresponds to two pieces of information. One data point = two
pieces of information.
• Tend to work with continuous numeric variables.

The correlation coefficient


Scatterplots tell us about the strength of relationship between two variables, or the correlation. Generally
the statistic we use for this is Pearson’s product-moment correlation
Pearson’s product-moment correlation coefficient:
{ Can be calculated when both variables are continuous numeric variables (not categorical variables)
{ It is designed to fall between -1 and +1. Magnitude (strength) is always between 0-1, and then it will
have a sign positive/negative (direction)
{ “No correlation” is identified by the value 0, so strong correlations are closer to either -1 or +1
{ It is identified by the letter r (the first letter of the word regression). Correlation is the foundation of
regression analysis.
{ The formula is:

How strong is strong?


Once we have a correlation, we can square it to get a sense of how much of the variance in one variable is
explained by variation in the other. (More related to regression)

You can get a sense of the relationship between two variables by calculating r2, which gives you the proportion
of variance shared between the two variables.
For instance, if the relationship between child’s and father’s education can be represented by r=0.5, then
r2=0.25; so 0.25 or 25% of the variance in child’s and father’s education is shared (the other 75% is then
affected by variables we haven’t studied).

Value of r (+ pr -) Guilford’s interpretations

<0.2 Slight, almost no relationship

0.2-0.4 Low correlation, definite but small relationship

0.4-0.7 Moderate correlation, substantial relationship

0.7-0.9 High correlation, strong relationship

0.9-1.0 Very high correlation, very dependable relationship

Not all associations are linear - Curvilinearity


If data isn’t linear, you may get a low correlation coefficient, but there might still be an association.
Correlation coefficients (Pearson product moment correlation) only work when there is a linear relationship – a
near-zero correlation coefficient only means that there is no linear relationship. There might be a curvilinear
relationship – why one should always look at the scatterplot, as well as calculating statistics. Visual
STATISTICS MODULE 2 PSYCHOLOGY

representation = clear representation. Statistic used to quantify what we see: can use a statistic to calculate
further things.
F I.e., representation of how anxiety and exam performance works
F Yerkes-Dodson Law.
Curvilinear relationship: as one variable (anxiety) increases, the
other (exam performance) goes up and then comes down.

In all these graphs: bivariate data: data where there are only two
variables involved:
- child’s and father’s education
- Alcohol consumption and eye-hand coordination
- Exam performance and anxiety

What if the variables aren’t


continuous – Ranked variables/ordinal data
• Sometimes, our variables are ranked:
– We could be looking at athletes, who are ranked first, second, third, etc. in a race
– Our data could be ranked by ourselves.
• Use Spearman’s coefficient of rank correlation, (a different correlation coefficient) or the ranked
coefficient of correlation, rs.

Assumptions about data:


 Continuous/ discrete?
 Type of data is important for the type of test.

Significance
You might find, when you are reading a paper, reports like this: (*NOTE – use 3 decimals for p-values and
record exact probability; 2 decimals for everything else)
– r=0.2, p<0.05
– r=0.5, p=0.1593
– r=0.7, p<0.01

In Psychology, p values < 0.05 (less than) are considered significant = Unlikely to happen by chance
*NOTE – p values are sensitive to sample size. If you don’t collect data from a large sample, chance of
randomness is much higher.
 “r” = The symbol for correlation
 “p” = The probability that this result has occurred by chance
- If p<0.05, that means that we are 95% certain that this result did not occur by chance.
- p<0.01 indicates that we are 99% certain that this did not occur by chance. We would then
increase our belief that this is a “true” relationship in our sample, and not something we
stumbled on by error.
 Conventionally, we pay more attention to p<0.05 and p<0.0
- The correlation of r=0.5 with a p value of 0.1593 would be one we are less certain of as being
an accurate reflection of reality.

Never look at one piece of information, but several together, to decide whether a result is worth paying
attention to – look at p & r together to determine level certainty/ significance
 Look at r value (strength of correlation): r=0.2 is a very weak correlation. Even though it is
statistically significant, it may not be worth paying attention to – whether it’s worth paying attention
to is a matter of judgement, and the size of the correlation and the p-value.
- If there is a correlation of 0.2 between anxiety and test performance, that means that anxiety
does affect test performance, but only to a small extent.
- If you want your test performance to go up, most likely you need to pay more attention to
things other than anxiety (studying more and getting enough sleep, for instance).
- If there is a correlation of 0.7 between anxiety and test performance, then you probably want
to pay a lot of attention to anxiety as affecting test performance.
 Look at the r2 value
STATISTICS MODULE 2 PSYCHOLOGY

- r2=0.04 in the first example (anxiety explains 4% of the variance in test performance) but
r2=0.49 in the second example (so anxiety explains 49%, or nearly half, of the variance – the
difference between people – in test performance).

Comparing means: The T-Test (Inferential


test)

What is a T-test?
{ Compares the mean difference between the groups - calculates the statistical significance of such a
difference.
{ Only compares 2 things
An important part of psychological research is being able to determine if the difference we are seeing between
two groups is statistically significant.
- For example, we want to test if this new drug is successful at reducing depressive symptoms. So we
look at the depression scores of two groups, one who received the new drug and one that received a
placebo.
- We see that the depression scores differ for both groups but we need to be able to determine if this
difference is due to chance, or is likely to apply to most people who would take the new drug.

Types of T-tests
• Single sample – one group = sample you collected; other group = entire population. Compares sample
to population.
• Independent sample – comparing two separate different/independent groups to each other. Assumes no
one in the one group is present in the other. Often in experiments = experimental and control group.
• Repeat measure – using the same people twice at two different points in time to see if something has
changed. (i.e., when testing an intervention)

Look at the overlap


In first graph, some people in the group who scored lower on average scored higher than some of those in the
group that scored higher on average.
• If means are more than 2 SD apart, almost always different; less than 0.5 SD apart, almost always
similar.
Significance is a way to quantify the overlap between two groups.

Independent sample T-Test


• Compares two independent groups – two groups with no overlap – no participant in both groups
• Uses the mean and variance to decide if there is a significant difference between the groups by seeing
how much overlap there is between the different scores and how much difference there is between the
two means:
- The bigger the difference between the group averages
- The less overlap there is
- = More likely there is a significant difference between those groups
• Calculates a t statistic and associated p value – need p-value to tell us about its significance
• T statistic gets bigger as your sample size increases; p value gets smaller as your sample size
increases
Degrees of Freedom (DF)
 = (number of participants in group one -1) + (number in group two -1)
 Degrees of freedom refers to the maximum number of logically independent values, which are
values that have the freedom to vary, in the data sample. Degrees of freedom is calculated by
subtracting one from the number of items within the data sample.
STATISTICS MODULE 2 PSYCHOLOGY

 Number of observations that you can lose or change without effecting the knowledge of your outcome.
Higher samples give you more degrees of freedom.
- Example: Sample size = 10 and have worked out the average, you can use 9 of the scores
along with the average to work out the value of the tenth participant.

You might also like