2.
Statistical graphs
Once the data has been summarised, the clearest way to represent it is by using a graph. The majority of
people enjoy visual representation of information! Representing data is the fifth stage of the research
process. This stage depends on the work done in the previous stages. In other words, if the data collected
is biased, the representations in this stage will be flawed, unreliable and misleading.
We represent data by using statistical graphs. Each type of graph offers a different picture of the data and
certain graphs are more appropriate for particular types of data. Let’s take a look at different types of
graphs:
● Vertical bar graphs show changes over time at discrete times, for example transit robberies
in South Africa over the past ten years.
● Horizontal bar graphs show or rank items at one point in time, for example Covid-19 deaths
during the pandemic.
● Multiple bar graphs show frequencies of related data sets at discrete times, for example the
South African population by gender over a time period of four years.
● Compound (stacked) bar graphs show compositions of related data sets and focus on the
total frequency of each data set. For example, the South African Olympic medal results since
2010.
● Pie charts show parts of the whole, for example when comparing the different allocations of
the national budget.
● Histograms show either continuous data or discrete data with many different values. For
example, the unemployment rate in South Africa by age group.
● Line and broken-line graphs show continuous data measured repeatedly over a period of
time, for example electricity usage by a high-consumption consumer in a residential area.
● Scatter plot graphs show the relationship between the data values of two sets of data. For
example, different car brands and prices.
● Box-and-whisker plots show the measures of spread of a set of data, for example a group of
learners’ test results.
● Bar graph — The graphical representation of data that uses bars to compare different
categories of data.
● Box-and-whisker plot — A diagram that shows the distribution of data along a number line
divided into quartiles.
● Broken line graph — A graph that has numbers that alternate going up and down and do not
keep to a curved or consistent line.
● Categorical data — The data that is given in the form of words, names or labels. It is generally
descriptive in nature, as data is classified and organised into categories.
● Class interval — Data that is divided into a smaller number of categories.
● Compound bar graphs — Display two or more sets of data. However, it shows a part/whole
relationship so you can easily see what amount each data group makes up of the whole.
● Continuous data — The data that is given as numbers including decimal numbers and/or
fractions. Numerical data (measurements like weight or age).
● Discrete data — Numerical data (fixed numbers like the size of a family). Data that can have
only certain values (quantities that can be counted, usually whole numbers).
● Double bar graph — The most common multiple bar graph that compares two sets of data.
● Five-number summary — Consists of the minimum and maximum values, median, lower
quartile and upper quartile.
● Frequency table — A table that lists items and uses tally marks to record the number of
times an item occurs. This is also called a frequency distribution table.
● Histogram — A graph using adjacent bars to show frequencies of continuous numerical data
with many different values.
● Interval — A range of numerical data that is part of the total range of possible values.
● Line graph — A graph that uses line segments to connect data points and shows changes in
data over time.
● Outlier — A number much smaller or bigger than the rest of the values in a set of data.
● Numerical data — Consists of numbers that are divided into discrete and continuous data.
● Pie chart — A circular diagram that is divided up into different sections or sectors. A circle is
divided into sections, illustrating the size for each category.
● Scatterplot — A graph that is made by plotting ordered pairs in a coordinate plane to show
the relationship between two sets of data, but the points are not connected by a line.
● Skewed data — Occurs when an outlier affects the measure of central tendency, as a result,
the ‘middle’ value is not accurately representative of the middle of the set of data.
PIE CHARTS
A pie chart (sometimes called a circle chart) is a circular diagram, where
each sector or section of the circle represents a data value. The pie chart
looks like a sliced pie! Each slice represents the proportion of the data in
relation to the whole. The pie chart is useful in comparing the relative size
of each data quantity. Therefore, the pie chart is often used for representing
categorical data. Each sector (or section) can be expressed as a fraction,
decimal or percentage. Let’s take a look at an example of a pie chart.
Figure 1: Music genres preferred by young people (14 to 19 years of age).
This pie chart in Figure 1 tells us that half (50%) of the young people surveyed
like rap and the other young people prefer alternative music (25%), rock and
roll (13%), country music (10%) and classical music (2%).
The method to determine the size of each sector:
Size of sector (in degrees) = Fraction of the whole x 360°
In order to draw a pie chart (as seen in Figure 1), follow this step-by-step
approach:
If 50% of the learners liked rap, then 50% of the whole circle graph (360°)
would equal 180°.
● Draw a circle with your protractor.
● Starting from the 12 o’clock position on the circle, measure an angle of
180° with your protractor. The rap component should make up half of
your circle. Mark this radius off with your ruler.
● Repeat the process for each remaining music category, drawing in the
radius according to its percentage of 360°. The final category need not
be measured as its radius is already in position.
● Bar graph — The graphical representation of data that uses bars to compare different
categories of data.
● Box-and-whisker plot — A diagram that shows the distribution of data along a
number line divided into quartiles.
● Broken line graph — A graph that has numbers that alternate going up and down and
do not keep to a curved or consistent line.
● Categorical data — The data that is given in the form of words, names or labels. It is
generally descriptive in nature, as data is classified and organised into categories.
● Class interval — Data that is divided into a smaller number of categories.
● Compound bar graphs — Display two or more sets of data. However, it shows a
part/whole relationship so you can easily see what amount each data group makes up
of the whole.
● Continuous data — The data that is given as numbers including decimal numbers
and/or fractions. Numerical data (measurements like weight or age).
● Discrete data — Numerical data (fixed numbers like the size of a family). Data that can
have only certain values (quantities that can be counted, usually whole numbers).
● Double bar graph — The most common multiple bar graph that compares two sets of
data.
● Five-number summary — Consists of the minimum and maximum values, median,
lower quartile and upper quartile.
● Frequency table — A table that lists items and uses tally marks to record the number
of times an item occurs. This is also called a frequency distribution table.
● Histogram — A graph using adjacent bars to show frequencies of continuous
numerical data with many different values.
● Interval — A range of numerical data that is part of the total range of possible values.
● Line graph — A graph that uses line segments to connect data points and shows
changes in data over time.
● Outlier — A number much smaller or bigger than the rest of the values in a set of data.
● Numerical data — Consists of numbers that are divided into discrete and continuous
data.
● Pie chart — A circular diagram that is divided up into different sections or sectors. A
circle is divided into sections, illustrating the size for each category.
● Scatterplot — A graph that is made by plotting ordered pairs in a coordinate plane to
show the relationship between two sets of data, but the points are not connected by a
line.
● Skewed data — Occurs when an outlier affects the measure of central tendency, as a
result, the ‘middle’ value is not accurately representative of the middle of the set of
data.
In statistics, the scatterplot is used to present measurements of two (or more) related data
sets. The scatter plot is particularly useful when the values of the data on the y-axis are
thought to be dependent on the values of the data on the x-axis. Therefore, one variable is
plotted against another variable to show the relationship between the two variables. The data
points, however, are not joined but form a pattern. The pattern formed illustrates the type and
the strength of the relationship between the variables. Scatterplots can illustrate various
patterns, e.g. positive (direct) relationship, negative (indirect) relationship or
no-relationship. Scatterplots can illustrate various patterns, e.g. positive (direct)
relationship, negative (indirect) relationship or no-relationship. These relationships are how
statisticians describe patterns in the data represented and are also known as correlations.
3.1 Positive correlation
When we look at the graph on the left in Figure 4 (positive correlation), the points in the
cluster run from the lower left to the upper right. The relationship between the two
variables in this cluster can be described as positive or direct. In other words, when one
variable increases, the other variable also increases.
3.2 Negative correlation
When we look at the graph in the middle in Figure 4 (negative correlation), the points in the
cluster run from the upper left to the lower right. The relationship between the two
variables in this cluster can be described as negative or indirect. In other words, when one
variable increases, the other variable decreases.
3.3 No correlation
When we look at the graph on the right in Figure 4 (no correlation), the points in the cluster
are scattered randomly without any noticeable relationship.
Interpreting and analysing data
After representing the data graphically, it is important to interpret and analyse the data. There
are a few points to take into consideration. Let’s take a look at these considerations.
● Using percentages in a table or graph is useful for comparing relationships in size, but
does not give any clear information about the actual population or sample size.
Remember, percentages are a representation of a whole. 10% could mean 1 out of
10 people or 50 out of 500 people.
● Using actual population or sample values gives a better understanding of the size, but
not the relationship between the data categories. Similar to the previous example, if
we know only 10 people participated in the research, we could draw a conclusion
that not a lot of people represent the population in comparison to a research
where 500 people participated.
● When you compare two sets of data you need to know how many items there were in
each set so that the comparison can be fair. The latter could be explained when a
survey was taken from two groups of people from two different towns. The questions
(items) are exactly the same, but the areas are different in order to compare the
findings.
● The choice of scale of the axes and the point at which the axes intersect will affect
the impression created by the graph. It is important to make sure that the graphs
that you compare have the same scale on the axes. If not, it influences the impression
that the graph illustrates.
● Tables give more information than graphs; however, graphs show trends in data
more clearly than data values in table format.
● Graphs show data clearly at a glance but the choice of graph type and the way in
which it is drawn can give you a misleading picture of the data and what it actually
illustrates.
Method of research process
It is important to question how data was collected, classified, organised, summarised and represented to
identify any errors, misinterpretations or bias. We can ask a few basic questions to double-check the
reliability and validity of the process.
● What was the size of the sample?
● Was the sample purposefully or randomly selected and representative?
● What methods were used to collect the data?
● Was the data collected facts or opinions?
● Could the sample be biased or unrepresentative?
● Is the source of the data clear?
● How was the data classified and organised?
● Which measures of central tendency and spread were used?
● Is this the best type of graph to illustrate this data
● Are the scale and the units labelled?
● Does the scale start at 0 and are the intervals the same?
● Does the title of the graph correspond to the data illustrated?
● Does the graph use 3D sections to make some parts look much bigger (or smaller) than others?
● Could the person (who drew the graph) try to give a particular message
● Average — A single value that represents a whole set of data.
● Bias — To favour one or more responses unfairly through the wording of a question or
the design of a survey.
● Data — Raw information that has been collected, without any organisation or analysis.
● Data handling — The process of collecting, organising, summarising, representing and
analysing information. The purpose of data handling is to make data more meaningful
to answer the research question(s).
● Mean — The result of dividing the sum of the data by the number of values.
● Median — The middle number in a set of data, arranged in descending order (from the
lowest to the highest).
● Mode — The number (or numbers) that occur(s) the most in a set of data.
● Range — The simplest measurement of the difference between values in a data set.
Calculate the range by subtracting the lowest value from the greatest value.