THE ANALYSIS OF BIOLOGICAL DATA
Chapter 2: Displaying Data
Complete Concepts & Definitions — Extreme Detail Guide
Every term, concept, graph type, and design principle explained in
full depth,
with real biological examples from landmark studies.
Source Text:
The Analysis of Biological Data
Whitlock & Schluter, 3rd Edition (2020)
MacMillan / W. H. Freeman
Chapter 2 · Pages 115–182
Table of Contents
1. Chapter Introduction: Why Displaying Data Matters
2. Section 2.1 — Guidelines for Effective Graphs
2.1 What Makes a Graph Good or Bad?
2.2 The Four Mistakes (Common Defects)
2.3 The Four Rules of Good Graph Design
Rule 1: Show the Data
Rule 2: Make Patterns Easy to See
Rule 3: Represent Magnitudes Honestly
Rule 4: Draw Graphical Elements Clearly
2.4 Chartjunk
2.5 Baseline Rules and Color-Blindness
3. Section 2.2 — Showing Data for One Variable
3.1 Frequency and Relative Frequency
3.2 Frequency Distribution
3.3 Frequency Table (Categorical Data)
3.4 Bar Graph
3.5 Pie Chart — and Why Bar Graphs Are Better
3.6 Histogram
Bins / Intervals
Sturges's Rule of Thumb
3.7 Shapes of Frequency Distributions
Peak and Mode
Bimodal Distribution
Symmetric Distribution
Skew (Positive and Negative)
Outliers
3.8 Violin Plot and Strip Chart (Brief Introduction)
3.8 Violin Plot and Strip Chart (Brief Introduction)
4. Section 2.3 — Showing Association Between Two Variables and Differences
Between Groups
4.1 Association Between Two Categorical Variables
Contingency Table
Grouped Bar Graph
Mosaic Plot
4.2 Association Between Two Numerical Variables
Scatter Plot
Positive, Negative, and Absent Association
4.3 Association Between a Numerical and a Categorical Variable
Strip Chart
Violin Plot (Full Explanation)
Multiple-Histogram Method
Comparing the Three Methods
5. Section 2.4 — Showing Trends in Time and Space
5.1 Line Graph
5.2 Map
6. Section 2.5 — How to Make Good Tables
6.1 Display Tables vs. Data Tables
6.2 Principles for Effective Display Tables
6.3 Worked Example: Improving a Bad Table
7. Section 2.6 — How to Make Data Files
7.1 Structure: Rows for Individuals, Columns for Variables
7.2 Human-Readable Variable Names
7.3 Auxiliary / Metadata Files
7.4 Plain Text / CSV File Format
7.5 Data Accuracy and Integrity
8. Master Summary Table of All Key Terms
9. Quick-Reference: Which Graph for Which Data?
1. Chapter Introduction: Why
Displaying Data Matters
Chapter 2 of The Analysis of Biological Data begins with a
fundamental insight that shapes everything that follows: the human
eye is a natural pattern detector. Our brains evolved over millions
of years to process visual information with extraordinary speed and
accuracy, identifying trends, clusters, outliers, and exceptions in
visual displays almost instantaneously — a task that would take
enormously longer if we tried to extract the same information from
a table of raw numbers. Biologists therefore spend significant time
creating and scrutinizing visual summaries of their data: graphs
and, to a lesser extent, tables.
The chapter opens with a reference to Florence Nightingale (1858),
one of the earliest and most influential pioneers of data
visualization. During the Crimean War, Nightingale used her
famous "wedge diagrams" (sometimes called "coxcomb" or polar
area diagrams) to display the causes of death of British troops
across different months of the war. The area of each wedge
represented the number of cases, and the color indicated the cause
of death (infectious disease, wounds, or other causes). The
diagrams revealed, with striking clarity, that the vast majority of
soldier deaths were caused by preventable infectious diseases
rather than battlefield wounds. These visualizations were so
persuasive that Nightingale successfully lobbied for public and
military health reforms that saved thousands of lives. Her story
illustrates a point central to this chapter: a well-constructed graph
does not merely describe data — it can change the world.
The authors state a prescription that they follow throughout the
entire book: the first step in any data analysis or statistical
procedure is to graph the data and look at it. Before
computing any summary statistics, before fitting any models,
before drawing any conclusions — graph the data. Visual inspection
reveals the general shape of distributions, the presence of outliers,
the nature of relationships between variables, and patterns that
bare numbers cannot convey. Graphs are thus both analytical tools
(helping the researcher understand their own data) and
communication tools (conveying results to others).
CORE PRINCIPLE OF THE CHAPTER
Always graph your data before doing anything else. The human
brain is exquisitely adapted to extract patterns from visual
information. Use this advantage at every stage of data analysis
and presentation.
The chapter is organized around two main purposes of graphs.
First, graphs are analytical tools used during the research process
to understand the data. Second, graphs are communication tools
used to present findings to a wider audience. The authors argue
that these two purposes are largely coincident: the displays that
best reveal patterns in data are almost always also the best for
communicating those patterns to others. A graph that obscures the
pattern is bad for both purposes simultaneously.
The chapter covers: (1) principles of good graph design; (2) graphs
for displaying one variable (bar graph, histogram); (3) graphs for
showing associations between two variables or differences between
groups (contingency table, mosaic plot, grouped bar graph, scatter
plot, strip chart, violin plot, multiple histograms); (4) graphs for
showing trends in time and space (line graph, map); (5) how to
make effective tables; and (6) how to organize and save data files.
2. Section 2.1 — Guidelines for Effective
Graphs
2.1 What Makes a Graph Good or Bad?
A graph is, fundamentally, a communication device. Its function is
to transmit information about patterns in data from the data
producer to the data consumer — whether that consumer is the
researcher themselves exploring their own data, or an audience of
fellow scientists reading a published paper. A good graph
accomplishes this transmission clearly, concisely, and without
distortion. A bad graph fails at one or more of these tasks — it may
mislead, confuse, hide the relevant information, or waste the
reader's cognitive effort on decorative elements that carry no
information.
To motivate the principles of good graph design, Whitlock and
Schluter present a deliberately bad graph (Figure 2.1-1 in the text)
showing mean maize plant yield under six combinations of nitrogen
fertilizer and soil water content. The data are real and interesting
— they show a clear interactive effect of nitrogen and water on
yield — but the graph renders the data nearly impossible to
interpret. The authors then systematically diagnose its four defects
and use them to derive four positive principles.
2.2 The Four Mistakes (Common Defects)
The four defects illustrated in Figure 2.1-1 of the textbook are
extremely common in published biological literature.
Understanding each in detail helps researchers avoid them in their
own work.
Mistake #1 The graph hides the data. Each bar in the maize
graph represents the average yield of four plant pots.
The individual data points — the actual observations —
are nowhere visible. This means the reader cannot see
the variation between pots, cannot assess whether any
observations are unusually high or low, cannot judge
whether the average is a good summary of the group,
and cannot evaluate whether the differences between
treatments are large relative to the natural variability
within treatments. When data are hidden behind
averages, the graph tells a simplistic and potentially
misleading story about the data.
Mistake #2 Patterns in the data are difficult to see. The
original graph uses three-dimensional bars shown in
angled perspective. The 3-D effect makes it genuinely
difficult to judge the height of the bars by eye — the
exact problem a bar graph is supposed to solve.
Because the bars appear to recede into the page, the
eye cannot reliably read off the magnitude of each bar
from the vertical axis. Any visual embellishment that
makes it harder to extract the intended information is
counterproductive. Edward Tufte (1983), whose
classic book on information graphics is cited here,
called such embellishments "chartjunk."
Mistake #3 Magnitudes are distorted. The vertical axis (plant
yield) begins at 2 rather than 0. Because the eye
instinctively reads bar height and area as proportional
to magnitude, starting the axis above zero makes the
bars appear taller relative to each other than the
actual differences in yield justify. A treatment that
produces twice the yield of another should have a bar
twice as tall, but if the axis starts at 2, even a small
difference in actual yield appears enormous. The
graph in the example visually exaggerates the
differences between treatment groups.
Mistake #4 Graphical elements are unclear. Text labels, axis
tick marks, and other figure elements are too small to
read easily. A graph that requires the reader to squint
or zoom in to read labels has failed in its
communication task. All text should be legible at the
final display size, whether in a printed journal article,
a presentation slide, or a web page.
2.3 The Four Rules of Good Graph Design
From the four mistakes above, four positive design principles are
derived. These rules are followed throughout the textbook and
should guide every graph a biologist creates.
Rule 1 Show the data.
RULE 1 — SHOW THE DATA (FULL EXPLANATION)
Show the individual data points or the full frequency
distribution whenever possible, so that the reader can see the
actual observations rather than only statistical summaries.
Showing the data allows the reader to evaluate the shape of
the distribution, judge the amount of variation within groups,
assess whether the sample size is adequate, spot potential
outliers or unusual observations, and compare the magnitude
of between-group differences with within-group variability.
The contrast between showing and hiding data is demonstrated in
Figure 2.1-2 of the text, which shows serotonin levels in desert
locusts (Schistocerca gregaria) after 0, 1, or 2 hours of crowding.
The left panel shows a strip chart: every single data point is plotted
as a dot, making visible not just the average but the entire
distribution — the spread, the skew, the outliers, the overlap
between groups, and the shift in the center of each group. The
right panel shows only bar height = average, using more ink yet
conveying far less information. From the strip chart, a reader can
clearly see that serotonin levels tend to rise with crowding
duration, but that there is substantial overlap between the
distributions and considerable variability within each group. None
of this is visible in the bar chart.
Rule 2 Make patterns in the data easy to see.
RULE 2 — MAKE PATTERNS EASY TO SEE (FULL EXPLANATION)
The central goal of a graph is to make the main pattern in the
data recognizable at a glance. This requires trying different
graph types, avoiding 3-D effects and chartjunk, not
overcrowding the graph with too much information, and
choosing the display method that best reveals the specific
pattern of interest. If the main finding is not immediately
obvious from the graph, the graph should be redesigned.
Making patterns easy to see is partly a matter of choosing the right
graph type for the data and the question (discussed in detail in
later sections), and partly a matter of keeping the graph visually
clean. Decorative elements — grid lines, drop shadows, gradients,
3-D effects, superfluous color — consume cognitive resources
without contributing information, making it harder for the brain to
perceive the underlying pattern. The principle of data-ink ratio,
articulated by Tufte, suggests that every drop of ink on the page
should be earning its place by conveying information. Anything that
uses ink without conveying information should be removed.
The authors also caution against overloading a single graph with
too much information. When a researcher has many results to
present, it is tempting to pack everything into one crowded graph.
But this usually makes the graph harder to understand. The
purpose of a graph is not to show all the data at once, but to
communicate the essential pattern clearly. Peripheral details
belong in an appendix or online supplement for the small fraction
of the audience that needs them.
Rule 3 Represent magnitudes honestly.
RULE 3 — REPRESENT MAGNITUDES HONESTLY (FULL
EXPLANATION)
The visual representation of data must be proportional to the
actual magnitudes in the data. This is especially critical for bar
graphs, where the height and area of each bar must be
proportional to the value being displayed. Bar graphs must
always begin at zero on the y-axis. Starting the axis above zero
makes differences appear larger than they truly are and
produces a dishonest graph. Other graph types (scatter plots,
strip charts, line graphs) have more flexibility on the y-axis
baseline, because they do not rely on bar area for magnitude
representation.
The authors illustrate this principle with a government data
example from British Columbia, Canada. A bar graph of education
spending per student over seven years uses a y-axis that starts
around $5,800 rather than $0. As a result, the bars appear to show
that spending increased roughly twenty-fold over the period, when
the actual increase was less than 20%. The area of each bar is
grossly out of proportion to the actual spending values. The revised
version of the graph, with the y-axis properly starting at zero,
shows the true (modest) increase and allows accurate visual
comparison. This type of distortion is alarmingly common in both
scientific and popular media.
CRITICAL RULE FOR BAR GRAPHS
A bar graph must ALWAYS have a baseline (y-axis minimum) at zero.
The human visual system instinctively interprets bar height and area
as proportional to magnitude. Starting above zero creates a
systematically misleading visual impression. This rule is non-
negotiable for bar graphs. Other graph types (line graphs, scatter
plots) are more flexible.
Rule 4 Draw graphical elements clearly.
RULE 4 — DRAW GRAPHICAL ELEMENTS CLEARLY (FULL
EXPLANATION)
All elements of a graph — axis labels, tick marks, legend text,
data point symbols — must be large enough to read
comfortably at the final display size. Axes must be labeled with
the variable name AND units of measurement. If multiple
groups are distinguished by color or symbol, the symbols must
be distinctly different from each other (not just slightly
different shades of the same color, which become
indistinguishable after printing or photocopying). Simple,
unadorned typefaces are preferred over decorative ones. The
default output of most computer graphics packages is rarely
optimal and should be edited before use.
2.4 Chartjunk
DEFINITION — CHARTJUNK
Chartjunk is a term coined by Edward Tufte (1983) to
describe visual elements in a graph that add clutter and
consume "ink" (visual processing effort) without conveying any
information about the data. Examples include three-
dimensional effects applied to 2-D data, decorative shading
and gradients, redundant grid lines, background images, drop
shadows, superfluous borders, and excessively bold colors.
Chartjunk interferes with the eye and brain's ability to "see"
patterns in the data by adding visual noise that competes with
the signal.
Chartjunk is particularly pernicious because many default graph
settings in popular computer software (spreadsheet programs,
certain statistics packages) produce it automatically. 3-D bar
charts, pie charts with decorative shadows, and graphs with heavily
colored backgrounds are all common outputs of these programs.
The researcher who accepts the default output without editing is at
risk of producing chartjunk-laden graphs. The corrective is simply
to remove every visual element that does not directly contribute to
communicating the data. When in doubt, take it out.
2.5 Color Blindness and Symbol Choice
IMPORTANT PRACTICAL NOTE — COLOR BLINDNESS
Approximately one in ten men in most populations has some
form of red-green color blindness (deuteranopia or protanopia),
making it impossible for them to distinguish red from green. If
a graph distinguishes groups using red and green alone, these
readers will see the groups as identical. To ensure your graphs
communicate to the full audience: (1) choose colors that differ
in intensity (lightness) as well as hue, so they can be
distinguished even in grayscale; and (2) use redundant coding
— distinguish groups by both color AND symbol shape (e.g.,
circles vs. triangles vs. squares) or color AND line type (solid
vs. dashed). This way, the graph is interpretable even to color-
blind readers and in black-and-white printouts.
3. Section 2.2 — Showing Data for One
Variable
When the goal is to visualize a single variable measured across a
sample of observations, the fundamental task is to show its
frequency distribution: how often each value (or range of values)
occurs in the data. Two key concepts must be defined first.
3.1 Frequency
DEFINITION — FREQUENCY
Frequency is the number of observations in a sample that
have a particular value (for a categorical variable) or that fall
within a particular interval (for a numerical variable).
Frequency is simply a count: how many times does this value
occur?
3.2 Relative Frequency
DEFINITION — RELATIVE FREQUENCY
Relative frequency is the proportion of observations in a
sample that have a particular value or fall within a particular
interval. It is calculated by dividing the frequency (count) by
the total number of observations in the sample (n).
Relative frequency = Frequency ÷ n (total number of
observations)
Relative frequency is dimensionless — it always falls between 0 and
1 (or equivalently, between 0% and 100%). It is particularly useful
when comparing distributions across samples of different sizes,
because absolute frequencies would not be directly comparable: a
frequency of 10 means something very different in a sample of 20
versus a sample of 10,000.
3.3 Frequency Distribution
DEFINITION — FREQUENCY DISTRIBUTION
The frequency distribution of a variable is a complete
description of how often each value (or each interval of values)
occurs in a sample. It lists all possible values or intervals
alongside their corresponding frequencies. The relative
frequency distribution does the same but using proportions
(relative frequencies) instead of raw counts.
Visualizing the frequency distribution is the primary goal when
displaying data for a single variable. The shape of the frequency
distribution conveys essential information about the data: where
values are concentrated, how spread out they are, whether the
distribution is symmetric or skewed, and whether there are any
unusual extreme values. This information is invisible in summary
statistics alone.
3.4 Frequency Table (Categorical Data)
DEFINITION — FREQUENCY TABLE
A frequency table is a text-based display showing the
frequency (and optionally, the relative frequency) of
occurrence of each category of a categorical variable, or each
interval of a numerical variable. Each distinct category or
interval occupies one row, and the corresponding count
occupies an adjacent column.
Frequency tables are simple and informative but lack the visual
immediacy of graphs — the eye must compare numbers rather than
perceiving lengths or areas. However, they provide the exact
numerical values that graphs do not always convey. Frequency
tables and bar graphs are complementary: the table provides
precision, the graph provides immediate visual impact.
REAL EXAMPLE — TIGER ATTACK ACTIVITIES (GURUNG ET AL.
2008)
Researchers studied 88 people killed by tigers near Chitwan National
Park, Nepal (1979–2006). They recorded the activity each person was
engaged in at the time of attack. The frequency table lists activities
(collecting grass/fodder, fishing, herding, etc.) in order from most to
least frequent (44, 11, 8, 7, 5, 5, 3, 3, 2 deaths respectively).
Ordering categories by frequency makes it immediately obvious
which activities are most dangerous — grass-collecting accounted for
half of all deaths — without requiring the reader to scan the entire
table.
BEST PRACTICE — ORDERING CATEGORIES IN A FREQUENCY
TABLE
For nominal (unordered) categorical variables: list
categories in descending order of frequency (most common
first). This makes it easy to identify the most important
categories at a glance. For ordinal (naturally ordered)
categorical variables (such as severity scores): preserve the
natural order (e.g., mild → moderate → severe). Never default
to alphabetical ordering, which conveys no information about
the relative importance of categories.
3.5 Bar Graph
DEFINITION — BAR GRAPH
A bar graph is a graphical display in which rectangular bars
of equal width are drawn for each category of a categorical
variable, with the height (and therefore area) of each bar
proportional to the frequency or relative frequency of that
category. The bars stand apart from one another (with spaces
between them), visually emphasizing that the categories are
distinct and unordered (or at least separately defined).
Bar graphs are the preferred method for displaying the frequency
distribution of a categorical variable. The key reason is that the
human visual system is highly accurate at judging the relative
lengths of parallel bars — much more accurate than it is at
comparing the angles or areas of sectors in a pie chart. A well-
drawn bar graph communicates not just the frequency of each
category but also their relative magnitudes (e.g., "this category is
four times as common as that one"), which is information that is
genuinely difficult to read from a pie chart.
Rules for Making a Good Bar Graph
Baseline at zero: The y-axis must start at zero. Bar height and
area are used by the eye to represent magnitude. Starting above
zero makes small differences look large and is dishonest.
Equal bar widths: All bars must be the same width, so that
area is proportional to frequency. Bars of unequal width would
create the false visual impression that some categories are more
"important" than others simply because their bars are wider.
Bars separated by spaces: Unlike histogram bars (which are
contiguous), bar graph bars should be separated by visible gaps,
emphasizing that the categories are discrete rather than
forming a continuum.
Ordering: For nominal categories, order bars by descending
frequency. For ordinal categories, preserve the natural order.
Report n: Include the total number of observations (n) in the
figure caption. Without knowing n, the reader cannot assess the
reliability of the frequencies shown.
No 3-D effects: Never use 3-D bars. They introduce chartjunk
and make bar heights harder to read accurately.
3.6 Pie Chart — and Why Bar Graphs Are Better
DEFINITION — PIE CHART
A pie chart is a circular graphical display in which the total
circle represents the whole sample, and each "slice" (wedge)
has an area proportional to the relative frequency of one
category. The angle of each sector at the center equals 360°
times the relative frequency of that category.
Despite their widespread use in business presentations and popular
media, pie charts have serious deficiencies as tools for scientific
data display, and are not recommended by Whitlock, Schluter, or
most experts in data visualization.
The fundamental problem is that the human visual system is far
less accurate at comparing angles and sectors than at comparing
lengths of parallel bars. When two slices in a pie chart have similar
sizes, it is genuinely difficult to tell which is larger. This problem
compounds as the number of categories increases: a pie chart with
8 or more categories becomes almost uninterpretable visually.
When exact numbers need to be communicated, labels are often
added around the perimeter — but at that point, the chart is no
better than a table and uses far more space. Side-by-side
comparison of two pie charts (e.g., comparing the frequency
distribution of an activity across two time periods) is extremely
difficult, much harder than comparing two bar charts. For all these
reasons, the bar graph should be the default choice for categorical
frequency data.
3.7 Histogram
DEFINITION — HISTOGRAM
A histogram is a graphical display of the frequency
distribution (or relative frequency distribution) of a numerical
variable, in which the data values are divided into consecutive
intervals ("bins") of equal width and the frequency of
observations in each bin is represented by the area of a
rectangular bar. Unlike a bar graph, histogram bars are
contiguous (no gaps between them), reflecting the continuous
nature of the underlying numerical variable.
The histogram is the primary recommended tool for displaying
numerical data for a single variable. Its key design principle is that
frequency is encoded as bar area, not bar height alone. This
distinction matters when bin widths differ, but since bins are
almost always of equal width in a well-drawn histogram, height and
area convey the same information. The continuous nature of the
numerical variable is reflected in the absence of gaps between
bars: a gap would imply that no values could fall in that range,
which is rarely true for numerical data.
Bins (Intervals) in a Histogram
DEFINITION — BINS / INTERVALS
Bins (also called intervals or class intervals) are the
consecutive, non-overlapping, equal-width divisions of the
range of a numerical variable used to construct a histogram.
Each observed data value is assigned to exactly one bin, and
the frequency of values falling into each bin is counted and
plotted as a bar. The choice of bin width critically determines
the appearance and informativeness of the histogram.
The choice of bin width is one of the most consequential decisions
in constructing a histogram, because the same data set can tell
very different visual stories depending on how many bins are used.
Using too narrow bins (too many of them) produces a jagged,
erratic histogram that suggests more peaks and structure than may
truly exist — the eye is drawn to random fluctuations from bin to
bin. Using too wide bins (too few) produces an oversimplified,
heavily smoothed histogram that may mask genuine features of the
distribution, such as a secondary peak. The ideal bin width
captures genuine features of the distribution without attributing
false structure to random variation.
REAL EXAMPLE — SOCKEYE SALMON BODY MASS (HENDRY ET AL.
1999)
Body mass of 228 female sockeye salmon from Pick Creek, Alaska,
was displayed with three different bin widths: 0.1 kg (narrow), 0.3 kg
(intermediate), and 0.5 kg (wide). The narrow-bin histogram
produced a bumpy distribution suggesting two or even more peaks.
The wide-bin histogram masked the genuine second peak by
aggregating too many observations into each bar. Only the
intermediate bin width (0.3 kg) clearly showed the two distinct body-
size groups in the population — a biologically meaningful pattern
corresponding to two different life-history strategies. This example
powerfully demonstrates that bin width is not merely an aesthetic
choice; it can determine whether real biological patterns are detected
or missed.
Sturges's Rule of Thumb
DEFINITION — STURGES'S RULE
Sturges's rule of thumb is a formula for choosing the
number of bins in a histogram: the suggested number of
intervals is 1 + ln(n) / ln(2), where n is the total number of
observations and ln is the natural logarithm. The result is
rounded up to the nearest integer. This formula yields
approximately 7 bins for n = 64, about 8 bins for n = 256, and
about 11 bins for n = 2048.
Number of bins ≈ ⌈ 1 + ln(n) / ln(2) ⌉
The authors note that Sturges's rule is widely regarded as overly
conservative — it tends to produce histograms with too few bins,
which can over-smooth the data. In practice, good judgment and
experimentation with several alternative bin widths is preferred
over mechanical application of any formula. The goal is always to
choose the bin width that best reveals the genuine patterns and
exceptions in the data. Additionally, break points should be chosen
at "readable" numbers (e.g., break at 0.5, not 0.483) to make the
histogram easy to interpret.
RULES FOR DRAWING A GOOD HISTOGRAM
(1) Each bar must rise from a baseline of zero — area
represents frequency. (2) Adjacent bars are contiguous (no
gaps). (3) Bins are of equal width. (4) Use "left-closed"
intervals: a value falling exactly on a boundary belongs to the
interval starting at that value (e.g., 70 goes in the 70–75 bin,
not the 65–70 bin). (5) Try several different bin widths and
choose the one that most clearly shows the genuine patterns.
(6) Use breakpoints at round, readable numbers. (7) Report n
in the caption.
3.8 Shapes of Frequency Distributions
One of the primary reasons for making a histogram is to reveal the
shape of the frequency distribution. Understanding the shape of a
distribution is critical for choosing appropriate statistical methods
(many of which assume a particular shape) and for gaining
biological insight. The following terms describe common shapes
and features.
Peak and Mode
DEFINITION — PEAK
A peak is any interval (bin) in a frequency distribution that is
noticeably more frequent than the surrounding intervals,
creating a local high point in the histogram. A peak represents
a region where values are concentrated — a common value
range in the data.
DEFINITION — MODE
The mode is the interval corresponding to the highest peak in
the frequency distribution — the most frequently occurring
value range in the data. In a histogram, the mode is simply the
tallest bar. For a bell-shaped distribution, the mode is located
at the center. For a skewed distribution, the mode is near the
tail-free end.
The mode is a measure of central tendency, but it describes the
most common value rather than the average or middle value. It is
particularly useful for distributions with obvious peaks, and less
useful for uniform or heavily skewed distributions. Unlike the mean
and median, the mode can be identified directly from a histogram
without calculation.
Bimodal Distribution
DEFINITION — BIMODAL DISTRIBUTION
A bimodal distribution is a frequency distribution that has
two distinct, separated peaks (two local maxima). A bimodal
shape in biological data often indicates the presence of two
subpopulations with different typical values — for example,
males and females if they differ substantially in size, or
animals using two alternative strategies if each strategy leads
to a different size or behavior.
Detecting bimodality is a critical benefit of using histograms rather
than summary statistics. The mean, median, and standard deviation
of a bimodal distribution are often biologically meaningless — they
describe a "typical" value that may never actually occur, falling
somewhere between the two peaks. A histogram immediately
reveals that the distribution is bimodal and suggests that the
sample may contain two distinct subgroups that should perhaps be
analyzed separately.
EXAMPLE — BIMODAL SALMON BODY MASS
The sockeye salmon body mass data, when displayed with an
appropriate bin width, showed two distinct peaks: one around 1.6–1.9
kg and another around 2.8–3.1 kg. This bimodal distribution likely
reflects two distinct life-history strategies in the population (e.g.,
early vs. late breeding fish), each associated with a characteristic
body size. This biologically important pattern would be completely
invisible from summary statistics alone.
Symmetric Distribution
DEFINITION — SYMMETRIC DISTRIBUTION
A frequency distribution is symmetric if the histogram on the
left side of the mode is a mirror image of the histogram on the
right side. The distribution is balanced around its center. Two
common examples of symmetric distributions are the uniform
distribution (in which all intervals have approximately equal
frequency) and the bell-shaped (normal) distribution (in which
frequencies are highest in the center and taper off
symmetrically in both directions).
The normal (bell-shaped, Gaussian) distribution is the most famous
symmetric distribution in statistics, partly because many biological
measurements approximate this shape (due to the Central Limit
Theorem and the fact that many traits are influenced by numerous
small, additive factors). The normal distribution is the basis of
many classical statistical tests, which assume normality of the data
or residuals. Examining a histogram for symmetry is therefore an
important step before applying normality-based statistical tests.
Skew (Positive and Negative)
DEFINITION — SKEW
Skew refers to asymmetry in the shape of a frequency
distribution for a numerical variable. A skewed distribution is
one in which the histogram is not symmetric — one tail
extends further from the mode than the other.
DEFINITION — POSITIVE (RIGHT) SKEW
Positive skew (also called right skew) describes a distribution
in which the right tail is longer than the left — most
observations are clustered at lower values, but some
observations extend far to the right (high values). In a
positively skewed histogram, the mode is to the left of center
and the right tail is stretched. Examples: income distributions
(most people earn moderate amounts, but a few earn
extraordinarily high amounts), body size of territorial animals
(most are medium-sized, a few are very large), waiting times
(most are short, but occasionally very long).
DEFINITION — NEGATIVE (LEFT) SKEW
Negative skew (also called left skew) describes a distribution
in which the left tail is longer than the right — most
observations are clustered at higher values, but some
observations extend far to the left (low values). In a negatively
skewed histogram, the mode is to the right of center and the
left tail is stretched.
REAL EXAMPLE — ZIKA-AFFECTED FETAL HEAD WIDTHS (BRASIL
ET AL. 2016)
Researchers measured the head widths (biparietal diameters in mm)
of 40 fetuses in pregnant women infected with the Zika virus, using
ultrasound at 33–36 weeks of gestational age in Rio de Janeiro. The
histogram showed that most fetuses had head widths between 80–94
mm, but three had substantially smaller heads (around 61–70 mm),
creating a long left tail. The distribution is negatively (left) skewed.
The three extreme small-headed fetuses fall well outside the range
for normal fetuses at the same gestational age and are classified as
microcephalic — a key finding of the study demonstrating the effect
of Zika infection on fetal brain development.
Skewness is biologically important because it affects the
interpretation of the mean. In a right-skewed distribution, the
mean is pulled toward the right tail and may be much higher than
the mode or median — it is not a representative "typical" value. In a
left-skewed distribution, the mean is pulled toward the left.
Understanding whether data are skewed helps the researcher
choose appropriate measures of central tendency (mean vs.
median) and appropriate statistical tests.
Outliers
DEFINITION — OUTLIER
An outlier is an observation that lies well outside the range of
values of the other observations in the data set — an extreme
value that is isolated from the main body of the data. In a
histogram, outliers appear as isolated bars separated from the
main distribution by empty space, or as a small bar on one end
of an otherwise compact distribution.
Outliers require careful attention. They have two fundamentally
different origins:
1. Data entry errors or measurement mistakes: A value of
8,500 in a variable that otherwise ranges from 80 to 95 is almost
certainly a typo (missing a decimal point). Such outliers should
be investigated, and if confirmed as errors, corrected or
removed from the data set. This is one of the main reasons to
always graph the data before doing any analysis — errors that
would be invisible in a table of numbers are immediately obvious
in a histogram.
2. Genuine extreme observations from nature: Some outliers
represent real, correctly measured values that happen to be far
from the typical range. In the Zika fetal head-width example, the
three microcephalic fetuses are genuine biological outliers —
they are not measurement errors, they are babies with
genuinely small heads, and they should not be removed from the
data. Removing genuine outliers would bias the results and
would suppress exactly the information (the existence of
microcephaly) that makes the study scientifically important.
CRITICAL WARNING — OUTLIERS MUST ALWAYS BE INVESTIGATED
Never remove an outlier simply because it is extreme. Investigate
every outlier to determine its cause. If it is a data error, correct or
remove it. If it is a genuine observation, keep it and consider its
biological meaning. Removing genuine outliers to "clean up" the data
constitutes data manipulation and can invalidate statistical results
and scientific conclusions.
3.9 Violin Plot and Strip Chart (Brief
Introduction)
While histograms are the primary tool for displaying a single
numerical variable, the violin plot and strip chart are alternatives
frequently used when comparing multiple groups. They are
described briefly here and in more detail in Section 4.3. The box
plot and cumulative frequency distribution (introduced in Chapter
3 of the textbook) are additional alternatives for displaying
numerical distributions.
4. Section 2.3 — Showing Association
Between Two Variables and Differences
Between Groups
When a study involves two variables, the central question is usually
whether the two variables are associated — whether the value of
one variable tells us something about the likely value of the other.
The appropriate graphical method depends on whether the two
variables are both categorical, both numerical, or one of each type.
DEFINITION — ASSOCIATION
Two variables are associated (or correlated) if the value of
one variable is predictably related to the value of the other.
For categorical variables, association means that the relative
frequencies of one variable differ across the categories of the
other. For numerical variables, association means that higher
(or lower) values of one variable tend to accompany higher (or
lower) values of the other.
4.1 Association Between Two Categorical
Variables
When both variables are categorical, the question is whether the
distribution of one variable changes across the categories of the
other. Three tools are available: the contingency table, the grouped
bar graph, and the mosaic plot.
Contingency Table
DEFINITION — CONTINGENCY TABLE
A contingency table is a frequency table for two (or more)
categorical variables simultaneously. It displays the frequency
(count) of observations falling into every combination of
categories of the two variables. Rows represent categories of
one variable; columns represent categories of the other. It is
called a "contingency" table because it shows how the
frequencies of one variable are "contingent upon" (depend on)
the categories of the other. The margins (row totals, column
totals, and grand total) are also typically displayed.
REAL EXAMPLE — REPRODUCTIVE EFFORT AND AVIAN MALARIA
(OPPLIGER ET AL. 1996)
Researchers investigated whether increased reproductive effort
made female great tits (Parus major) more susceptible to avian
malaria. They divided 65 nesting females into two groups: (1) an egg-
removal group (n = 30), from which two eggs were stolen, forcing the
bird to lay an extra egg and increasing reproductive stress; and (2) a
control group (n = 35) left undisturbed. Blood was taken 14 days
after hatching to test for malaria. The 2×2 contingency table showed:
Control: 7 malaria / 28 no malaria; Egg-removal: 15 malaria / 15 no
malaria. Malaria occurred in 50% of egg-removal females but only
20% of controls — suggesting that the additional reproductive effort
increased susceptibility to disease.
In a contingency table, the explanatory variable (the variable
hypothesized to cause or predict the other) is conventionally placed
in the columns, and the response variable (the variable being
predicted) is placed in the rows. This convention makes it easy to
compare the distribution of the response variable across different
values of the explanatory variable by reading down each column.
Each individual observation is counted exactly once, in exactly one
cell of the table. The total count (grand total) equals the number of
individuals in the study.
DEFINITION — 2×2 CONTINGENCY TABLE
A 2×2 contingency table is a special case of a contingency
table in which both variables have exactly two categories,
resulting in four cells plus row and column totals. It is the most
common type in biology because many experiments compare a
treatment with a control (two categories of explanatory
variable) and record a binary response (e.g., diseased vs. not
diseased). Larger contingency tables arise when one or both
variables have more than two categories.
Grouped Bar Graph
DEFINITION — GROUPED BAR GRAPH
A grouped bar graph is a bar graph for two categorical
variables simultaneously, in which bars are grouped by one
variable (the explanatory variable, shown on the x-axis) and
distinguished by color, shading, or pattern to indicate the
categories of the other variable (the response variable). Within
each group, the heights of the bars show the frequencies of
each response category. This allows visual comparison of the
response distribution across groups.
In a grouped bar graph, the spaces between bars within the same
group are narrower than the spaces between groups, visually
emphasizing that the bars within a group belong together. The
graph allows the reader to quickly compare whether the relative
heights of the response-category bars change between groups — if
they do, an association is present. All the same rules as for a simple
bar graph apply: y-axis must start at zero, bars must be of equal
width, and n must be reported.
Mosaic Plot
DEFINITION — MOSAIC PLOT
A mosaic plot is a graphical display for two categorical
variables in which rectangles are used to show the relative
frequency (proportion) of each combination of variable
categories. The rectangles within each group ("stack") are
shown proportional to relative frequency, and the width of
each stack is proportional to the number of observations in
that group. As a result, the area of each rectangle is
proportional to the overall relative frequency of that
combination of categories in the entire data set.
The mosaic plot is particularly effective at revealing associations
because the vertical position at which the colors (response
categories) meet within each stack directly shows the proportion of
each response category within that group. If the response
distribution is the same across all groups, the meeting points will
be at exactly the same vertical height in every stack — perfectly
horizontal dividing lines. If an association exists, the meeting points
will differ in height between stacks. The greater the association,
the more the heights differ.
The additional feature of variable-width stacks (proportional to
group size) means that the area of each rectangle reflects its true
overall frequency in the data — a feature not available in grouped
bar graphs. This allows the reader to simultaneously assess the
distribution of the explanatory variable (from stack widths) and the
distribution of the response variable within each group (from bar
heights within stacks).
Comparing the Three Methods: Contingency Table,
Grouped Bar Graph, Mosaic Plot
Comparison of Methods for Two Categorical Variables
Method Strengths Limitations
Contingency Provides exact numbers; Visual comparison of
table shows both absolute and patterns requires mental
relative frequencies; calculation; no graphical
necessary for statistical impression
analysis
Grouped Shows absolute frequencies Does not show relative
bar graph directly; familiar to most frequencies; comparing
readers; easy to construct proportions between groups
requires mental division;
harder to see association
clearly
Mosaic plot Shows relative frequencies Does not show absolute
visually; stack width encodes frequencies; less familiar to
group size; area encodes some readers
overall proportion; association
is immediately visually
obvious
The authors recommend trying all three and choosing whichever
most clearly communicates the pattern in the data. They note that
they often find associations are easier to see in mosaic plots than in
grouped bar graphs, because the vertical position of the dividing
line between response categories directly encodes the proportion.
4.2 Association Between Two Numerical
Variables: Scatter Plot
DEFINITION — SCATTER PLOT
A scatter plot (also called a scatterplot or scatter diagram) is
a graphical display of the relationship between two numerical
variables, in which each observation is represented as a single
point in a two-dimensional space. The x-axis (horizontal)
displays the value of the explanatory variable, and the y-axis
(vertical) displays the value of the response variable. The
pattern formed by the cloud of points reveals the nature and
strength of the association between the variables.
The scatter plot is the fundamental tool for visualizing bivariate
numerical data. It allows the reader to see not just whether an
association exists but its direction, its strength, its linearity (or
nonlinearity), and whether any unusual observations deviate from
the overall pattern.
Types of Association in a Scatter Plot
DEFINITION — POSITIVE ASSOCIATION
A positive association (positive correlation) between two
variables exists when higher values of one variable tend to co-
occur with higher values of the other — the cloud of points in a
scatter plot tends to run from the lower-left to the upper-right
of the graph. As one variable increases, the other tends to
increase as well.
DEFINITION — NEGATIVE ASSOCIATION
A negative association (negative correlation) between two
variables exists when higher values of one variable tend to co-
occur with lower values of the other — the cloud of points runs
from the upper-left to the lower-right. As one variable
increases, the other tends to decrease.
DEFINITION — ABSENT ASSOCIATION
An absent association (zero correlation) means there is no
discernible relationship between the two variables — the cloud
of points shows no directional pattern; it appears roughly
circular or random, with no trend from left to right.
REAL EXAMPLE — GUPPY ATTRACTIVENESS INHERITANCE
(BROOKS 2000)
A study examined whether attractiveness in male guppies is inherited
from father to son. The attractiveness of 36 sons (measured as a rate-
of-female-visit score relative to a standard) was plotted against their
fathers' ornamentation (a composite index of color and brightness).
The scatter plot showed a positive association: sons of highly
ornamented fathers tended to be more attractive to females, while
sons of drab fathers were less attractive. The pattern of points ran
generally from lower-left to upper-right, confirming that attractive
traits are heritable. This is a classic demonstration of how scatter
plots can reveal biological inheritance patterns that no summary
statistic alone could capture.
Each point in a scatter plot represents one observation (here, one
father-son pair). The explanatory variable (father's ornamentation
— the presumed cause) is on the x-axis, and the response variable
(son's attractiveness — the presumed effect) is on the y-axis. The
placement of variables on x and y is a deliberate and meaningful
choice: the x variable is the one we use to predict or explain the y
variable.
4.3 Association Between a Numerical and a
Categorical Variable
When one variable is numerical and the other is categorical (i.e.,
the question is whether the distribution of a numerical
measurement differs across groups), three main graphical methods
are available: the strip chart, the violin plot, and the multiple-
histogram method.
The authors specifically recommend against using bar graphs for
this situation. While bar graphs showing group means are
extremely common in published biology papers, they fail rule 1 of
good graph design: they hide the data. A bar showing only the
mean of a group conveys nothing about the spread, shape, or
outliers of the distribution within that group. Two groups could
have identical means but completely non-overlapping distributions,
a pattern that would be invisible in a bar graph but immediately
obvious in a strip chart or violin plot.
Strip Chart
DEFINITION — STRIP CHART
A strip chart (also called a dot plot) is a graphical display for
one numerical variable and one categorical variable, in which
each individual observation is represented as a dot plotted at
its numerical value on one axis (usually the y-axis), with its
group membership indicated by its position on the other axis
(usually the x-axis). It is essentially a scatter plot where the x-
axis variable is categorical rather than numerical. Points
within a group are often "jittered" (spread slightly along the x-
axis) to reduce overlap and make individual points visible.
The strip chart is the most information-rich of the three methods
because it shows every single data point. The reader can directly
see the distribution of each group — where the center is, how much
spread there is, whether the distribution is symmetric, and whether
any outliers are present. Adding a horizontal bar at the group mean
within each strip provides a summary marker without sacrificing
the individual data points.
The main limitation of strip charts is that they become difficult to
interpret when a group contains many observations (hundreds or
thousands), because the dots overlap so severely that the individual
distribution shapes cannot be discerned. In Figure 2.3-4 of the text,
the USA group (n = 1,704 males) produces such severe point
overlap in the strip chart that the distribution is essentially
invisible. In this situation, the violin plot is more appropriate.
Violin Plot
DEFINITION — VIOLIN PLOT
A violin plot is a graphical display for one numerical variable
and one categorical variable that uses a smoothed
approximation of the frequency distribution of each group,
shown symmetrically (mirrored) around a central axis. The
width of the violin at each numerical value is proportional to
the frequency (density) of observations at that value — a wide
waist means many observations near that value, a narrow neck
means few observations there. The overall shape resembles a
violin, giving the plot its name. A dot or marker typically
indicates the mean (or median) of each group.
The violin plot is a compact and visually efficient way to show the
full frequency distribution of each group when the number of
observations is large. Unlike the strip chart (which shows
individual points), the violin summarizes the data into a smoothed
shape. Unlike the bar graph (which shows only the mean), the
violin shows the location, spread, and shape of the entire
distribution. The peaks and waists of the violin directly correspond
to the peaks and valleys of the frequency distribution. Two groups
with the same mean but different spreads will look visually
different in a violin plot but identical in a bar graph.
REAL EXAMPLE — HEMOGLOBIN CONCENTRATIONS AT HIGH
ALTITUDE (BEALL ET AL. 2002)
Researchers measured hemoglobin concentrations (g/dL) in males
from four populations: high-altitude Andes (n = 71), high-altitude
Ethiopia (n = 128), high-altitude Tibet (n = 59), and sea-level USA (n
= 1,704). Both a strip chart and a violin plot were presented. For the
USA group (n = 1,704), the strip chart produced hopelessly
overlapping dots; the violin plot clearly showed a tight, roughly bell-
shaped distribution centered around 15–16 g/dL. For smaller groups,
both methods were informative. The critical biological finding was
clearly visible in both plots: only Andean males had substantially
elevated hemoglobin concentrations (center around 18–20 g/dL),
whereas Ethiopian and Tibetan high-altitude populations had
hemoglobin levels similar to sea-level Americans — suggesting that
different human populations have evolved different physiological
adaptations to high-altitude hypoxia.
Multiple-Histogram Method
The multiple-histogram method uses separate histograms for each
group, stacked vertically above one another so that the y-axis
(numerical measurement) is aligned across all histograms. This
allows direct visual comparison of the location, spread, and shape
of the distribution across groups. The key requirement is that all
histograms use the same x-axis scale, so that the positions of bars
in one histogram can be directly compared with those in another.
The multiple-histogram method provides the most detail about the
shape of each distribution — mode, skewness, bimodality, and
outliers are all clearly visible. Its main disadvantage is that it
requires considerable vertical space, making it less practical when
there are many groups to compare. It is most effective with a small
number of groups (2–4), each with a sufficient number of
observations to show a meaningful distribution shape.
Side-by-side histograms (placed next to each other horizontally
rather than stacked vertically) are not recommended, because
differences in the position of bars between histograms are much
harder to perceive when histograms are side by side than when
they share a common y-axis and are directly above one another.
Comparing Strip Chart, Violin Plot, and Multiple
Histograms
Comparison of Methods for Numerical Variable × Categorical Variable
Method Best When Shows Limitations
Strip chart Few observations Every individual Overlapping points
per group (n < data point; true obscure distribution
~50–100) distribution when n is large
Violin plot Many Smoothed Individual points not
observations per distribution visible; smoothing can
group; multiple shape, center, create artifacts
groups spread
Multiple Few groups; Detailed Requires more space;
histograms moderate-to- distribution difficult with many
large n per group shape, bins, groups
peaks
5. Section 2.4 — Showing Trends in
Time and Space
Two additional graph types are commonly used in biology to
display data collected at consecutive points in time or at multiple
locations in space: the line graph and the map.
5.1 Line Graph
DEFINITION — LINE GRAPH
A line graph uses dots connected by line segments to display
changes in a summary measurement — such as a mean, total
count, or proportion — across an ordered series of
observations, most commonly at consecutive time points. The
x-axis represents the ordering variable (usually time), and the
y-axis represents the measurement. The line segments
connecting adjacent points help the eye follow the trend,
reveal the direction and speed of change, and highlight the
overall temporal pattern.
The connecting lines in a line graph serve a specific visual function:
they direct the eye from one time point to the next, making it easy
to perceive rising trends, falling trends, cyclical patterns, sudden
spikes, and periods of relative stability. The steepness of each line
segment encodes the rate of change: a steeply rising segment
means rapid increase, a nearly flat segment means little change, a
steeply falling segment means rapid decrease. When the y-axis
baseline is at zero, the area under the curve between two time
points is proportional to the cumulative total of the measurement
over that period.
REAL EXAMPLE — MEASLES CASES IN ENGLAND AND WALES (1995–
2011)
A line graph shows the number of confirmed measles cases per
quarter from 1995 to 2011. The graph reveals several outbreak
spikes (periods of rapid increase and decrease) superimposed on a
generally low background case rate. Each spike shows a
characteristic pattern: rapid ascent as the outbreak begins, followed
by equally rapid descent as immunity spreads through the population
and the reproduction number falls below 1. The line graph makes
these epidemic dynamics immediately visible and allows comparison
of the speed and magnitude of different outbreaks across the 16-year
period.
Line graphs are appropriate when the explanatory variable is
naturally ordered (usually time) and the eye can meaningfully
follow the line from one point to the next. They should not be used
for categorical variables where the order of the groups is arbitrary
— in that case, a bar graph is more appropriate because the "line"
connecting arbitrarily ordered bars would imply a false continuity
or ordering between the categories.
5.2 Map
DEFINITION — MAP (IN DATA VISUALIZATION)
A map is the spatial analogue of a line graph: it uses a color
gradient or other visual encoding to display values of a
numerical response variable at multiple locations in physical
space. The explanatory variable is spatial location (either a
grid of coordinates, or geographic boundaries), and the
response variable is encoded by color intensity ("heat"), where
warmer or darker colors typically represent higher values.
Maps are the natural tool for biological data that vary across
geographic space: species abundance, disease incidence,
temperature, elevation, genetic diversity, or any other quantity
measured at many locations. They can summarize enormous
amounts of data into a single comprehensible image, making
spatial patterns (gradients, hotspots, barriers) immediately visible.
The term "heat map" is also used for this type of display, especially
when applied to data that are not geographic (e.g., gene expression
levels across many genes and conditions displayed on a grid).
REAL EXAMPLE — PLANT SPECIES DIVERSITY IN NORTHERN
SOUTH AMERICA
A map displays the estimated number of plant species at many points
on a fine 100 km × 100 km grid covering northern South America. A
color gradient encodes species richness — "hotter" (redder) colors
represent greater numbers of species. The map immediately reveals
the regions of peak plant biodiversity (tropical lowland Amazon basin,
Andean foothills) and areas of relatively low diversity. Summarizing
this pattern with numbers alone would require thousands of data
points, but the map communicates it in a single image.
Maps can be used to display measurements on any two-dimensional
or three-dimensional surface, not just geographic maps. A visual
representation of an MRI scan — which uses color to encode tissue
properties at each spatial position within the brain or body — is
conceptually a map. The flexibility of the map concept makes it one
of the most powerful data visualization tools in science.
6. Section 2.5 — How to Make Good
Tables
Tables serve two distinct purposes in biological research, and these
purposes require different design philosophies.
6.1 Display Tables vs. Data Tables
DEFINITION — DISPLAY TABLE
A display table is a table whose primary function is to
communicate a pattern in data to a reader — typically
appearing in the main body of a report, paper, or presentation.
Its design should prioritize clarity of the pattern over
completeness of numerical detail. Frequency tables and
compact summary tables are typical examples. The guiding
criterion is: "Does this table make the pattern easy to see?"
DEFINITION — DATA TABLE
A data table stores raw data or detailed numerical summaries
for record-keeping or specialized reference purposes. It is not
optimized for recognizing patterns and is often large. Data
tables are typically not included in the main body of a paper;
instead, they appear as appendices or online supplementary
materials for readers who need the specific numbers. Their
guiding criterion is completeness and accuracy, not visual
clarity.
6.2 Principles for Effective Display Tables
The same four principles that govern good graph design — show
the data, make patterns easy to see, represent magnitudes
honestly, draw elements clearly — apply to display tables with
appropriate modifications:
T1 Make patterns easy to see. Keep the table compact. Use
only as many significant decimal places as are needed to
communicate the pattern (no more). Don't cram too much
data into one table. Arrange rows and columns to maximize
pattern detectability. Numbers in the same column (stacked
vertically) are much easier to compare with each other than
numbers in the same row (side by side in different columns)
— arrange data accordingly.
T2 Order categories wisely. For nominal (unordered)
categories: list in order of importance or frequency, not
alphabetically. For ordinal (naturally ordered) categories
(e.g., life stages, severity scores): preserve the natural order.
Alphabetical ordering almost never helps the reader detect
patterns and should be avoided unless there is a specific
reason.
T3 Represent magnitudes honestly. When combining
numerical data into bins (as in a frequency table), use bins of
equal width so that numbers can be accurately compared.
Unequal bins distort the visual impression of relative
frequencies.
T4 Draw elements clearly. Clearly label all row and column
headers. Always provide units of measurement alongside
variable names. Use a consistent, minimal number of decimal
places. Remove unnecessary visual elements (heavy borders,
excessive shading) that add clutter.
6.3 Worked Example — Improving a Bad Table
To illustrate these principles, the text presents Table 2.5-1: a table
of inbreeding coefficients (F) and offspring survival statistics for
Spanish Habsburg king-queen pairs (Alvarez et al. 2009). The
Habsburg dynasty ruled Spain from 1516 to 1700 and is historically
famous for extensive consanguineous (within-family) marriages.
The inbreeding coefficient F measures how genetically related
parents are: F = 0 if unrelated, F = 0.25 if siblings whose own
parents were unrelated (and can be higher with multi-generational
inbreeding).
The original table (2.5-1) has several specific defects identified by
the authors:
Disordered rows: King-queen pairs are listed in historical
order, not in order of F value. A putative trend (more inbred
couples → lower offspring survival) is invisible because the rows
are not ordered to show it.
Variables of interest separated: The F column and the
survival column are separated by several intervening columns
(pregnancies, miscarriages, neonatal deaths, later deaths),
making it hard to see their relationship at a glance.
Blank lines fragmenting the table: A blank line is inserted
for each new king listed, breaking the visual continuity of the
table and making it harder to compare rows.
Excessive decimal places: F values are given to three decimal
places and survival to two or three decimal places, more
precision than needed to see the pattern.
The revised table (2.5-2) addresses all these defects:
Rows are ordered by increasing F value (from least to most
inbred).
The F column and survival column are placed adjacent to each
other.
Blank lines between kings are removed.
F values are rounded to two decimal places; survival to two
decimal places.
With these changes, the pattern becomes immediately visible:
couples with the highest F values (Philip II/Anna of Austria, F =
0.22; Philip IV/Mariana of Austria, F = 0.25) had the lowest
postnatal offspring survival (20% and 40% respectively), while
couples with the lowest F values (Philip II/Elizabeth of Valois, F =
0.01; Philip I/Joanna I, F = 0.04) had the highest survival (100% in
both cases). The pattern is not perfectly monotonic, but the overall
trend is clear in the revised table and largely invisible in the
original.
THE INBREEDING COEFFICIENT F — DEFINITION
The inbreeding coefficient F measures the probability that
two copies of a gene in an offspring are identical by descent —
that is, both copies trace back to the same ancestor. F = 0 if
the parents are completely unrelated. F = 0.25 if the parents
are full siblings (and their own parents were unrelated). Values
above 0.25 can arise from multi-generation inbreeding. High F
values are associated with inbreeding depression: reduced
fitness (survival, fertility, disease resistance) in offspring due
to increased homozygosity of deleterious recessive alleles.
7. Section 2.6 — How to Make Data
Files
Since modern graphs are almost always created with computers,
the way data are organized in computer files critically determines
how easily they can be graphed and analyzed. This section provides
practical guidelines for creating data files that are usable by all
statistics and graphing software, now and in the future.
7.1 Structure: Rows for Individuals, Columns for
Variables
FUNDAMENTAL RULE — TIDY DATA STRUCTURE
Data files should be organized so that: (1) each row contains
data for exactly one individual or sampling unit; and (2) each
column contains exactly one variable. This "one row per
individual, one column per variable" structure (sometimes
called "tidy data") is the universal standard expected by
virtually all statistics and graphing software. Every piece of
data for a given individual occupies a single row, and every
measurement of a given variable occupies a single column.
This structure may seem obvious, but violations are common. A
common mistake is to use one column per group and fill it with the
measurements from that group — which makes the grouping
information implicit and difficult for software to process. Another
mistake is to use spreadsheet formatting (merged cells, subtotals,
colored headers) that makes the file look nice visually but breaks
the structure that software expects. Data files should contain only
data — no formatting, no subtotals, no decorative elements.
For numeric variables (e.g., hemoglobin concentration), the column
should contain only numbers — no units, no symbols like "$", "%",
"#", "@", or "&". Units belong in the variable name (column
header) or in the accompanying metadata file. For categorical
variables (e.g., population group), entries can be text strings (e.g.,
"Andes", "Ethiopia") or numeric codes (e.g., 1, 2, 3, 4) — but if
numeric codes are used, their meanings must be documented.
Missing data should be left as empty (blank) cells, not coded as 0
or -999 or any other number that could be mistaken for a real
observation.
7.2 Variable Names and Readability
RULE — HUMAN-READABLE VARIABLE NAMES
Variable names (column headers) should be clear and
unambiguous enough that any reader — including the
researcher themselves after memory has faded — can
immediately understand what the variable represents. Variable
names should have no spaces or punctuation (use
"HemoglobinConcentration" or "Hemoglobin_Concentration"
rather than "Hemoglobin Concentration"). While "HC" is
shorter to type than "HemoglobinConcentration," the longer
name will be interpretable without error long after the project
is over.
The top row of the data file should contain only variable names —
no data. Each subsequent row contains one observation. An ID
column identifying each individual (or sampling unit) by a unique
number or code is strongly recommended, so that individual
records can be checked against the original source for accuracy.
7.3 Auxiliary / Metadata Files
RULE — ALWAYS CREATE AN ACCOMPANYING METADATA FILE
Every data file should be accompanied by a separate
metadata file (sometimes called a "data dictionary" or
"readme file") that documents: where the data came from; how
they were collected; the meaning of each variable (including
units); any codes used in categorical variables (e.g., "M =
male, F = female — NOT mother and father"); and any special
circumstances affecting individual observations. This metadata
file ensures that the data remain interpretable long after the
original data collection is complete and the researcher's
memory of the details has faded.
7.4 Plain Text / CSV File Format
DEFINITION — CSV FILE (COMMA-SEPARATED VALUES)
A CSV file is a plain text file format for tabular data in which:
each row of the spreadsheet corresponds to one line in the file;
entries within a row are separated by commas; and a new line
begins a new row. CSV stands for "comma-separated values."
Because it is plain text, CSV files can be read by any current or
future statistics package, spreadsheet program, or
programming language, and will never become obsolete due to
software version changes.
The authors specifically recommend saving data in plain text
formats like CSV rather than proprietary formats (such as Excel's
.xlsx format) as the primary storage format. While Excel files are
convenient for entry and viewing, they are version-dependent and
may become unreadable as software changes over the decades of a
research career. A .csv file contains only characters and will always
be readable. If a spreadsheet program is used for data entry (for
convenience), the data should be exported/saved as CSV before
analysis.
7.5 Data Accuracy and Integrity
The authors emphasize that entering data accurately and
protecting it from loss are among the most important
responsibilities of a researcher:
Double- or triple-check every data point after entry. Read
data aloud to a colleague while checking against the original
source.
Graph the data immediately after entry to detect outliers
caused by data entry errors (e.g., a value of 850 instead of 85, or
0.05 instead of 0.5) — such errors are obvious in a histogram
but invisible in a raw data table.
Store data in multiple safe locations: cloud storage, a
second computer, and/or an external drive. Loss of data due to
computer failure is a real and painful risk in research. Arrange
automatic, regular backups.
Never modify the original data file: Keep the raw data file
unchanged. All cleaning, transformations, and analysis should
be done on copies, with the cleaning process documented in
code or a separate notes file.
8. Master Summary Table of All Key
Terms and Definitions
Complete Glossary — Chapter 2: Displaying Data
Term Full Definition and Key Points
Frequency The number (count) of observations in a sample with a
particular value or within a particular interval. A raw
count.
Relative The proportion of observations with a particular value;
frequency equals frequency ÷ n. Always between 0 and 1. Allows
comparison across samples of different sizes.
Frequency A complete description of how often each value or
distribution interval occurs in a sample — the full listing of all
values/intervals with their corresponding frequencies.
Reveals the shape, center, and spread of the data.
Relative The frequency distribution expressed as proportions
frequency rather than counts. Describes what fraction of
distribution observations fall in each value or interval.
Frequency table A text-based display of the frequency distribution: one
column lists all values/categories/intervals; an adjacent
column lists the corresponding counts (and optionally
proportions). Simple but lacks visual immediacy of
graphs.
Bar graph A graphical display for the frequency distribution of a
categorical variable. Uses rectangular bars of equal
width whose heights represent frequencies or relative
frequencies. Bars are separated by gaps. Y-axis must
start at zero. Best for showing frequencies in
categorical data.
Baseline (y-axis The minimum value on the vertical axis of a bar graph.
minimum) Must always be zero for bar graphs. Starting above
zero makes differences between bars appear larger
than they truly are, creating a dishonest graph.
Pie chart A circular display using wedge areas to represent
relative frequencies of categorical data. Not
recommended by the authors: the eye is less accurate
at comparing angles and areas than bar heights, and
side-by-side comparison of multiple pie charts is
especially difficult.
Histogram A graphical display for the frequency distribution of a
numerical variable. Data are divided into equal-width
bins; bar area (and height) represents frequency in
each bin. Bars are contiguous (no gaps). Y-axis starts at
zero. Reveals distribution shape, mode, skew,
bimodality, and outliers.
Bins (intervals) Consecutive, equal-width divisions of the range of a
numerical variable used in a histogram. Each
observation falls in exactly one bin. Bin width choice
critically affects what the histogram reveals — too
narrow: jagged and misleading; too wide: over-
smoothed and hides structure. Requires good
judgment.
Sturges's rule A formula for the suggested number of histogram bins:
⌈1 + ln(n)/ln(2)⌉. Generally considered too conservative
(too few bins). Good judgment and experimentation are
preferable to mechanical application of this formula.
Peak Any interval in a frequency distribution that is
noticeably more frequent than surrounding intervals —
a local high point in the histogram. Peaks indicate
regions where data values are concentrated.
Mode The interval corresponding to the highest peak — the
most commonly occurring value range. The tallest bar
in a histogram. A measure of central tendency
describing the most frequent value.
Bimodal A frequency distribution with two distinct, separated
distribution peaks. Often indicates two distinct subpopulations or
two alternative strategies. Mean and median of a
bimodal distribution may be biologically meaningless.
Symmetric A distribution where the left half of the histogram
distribution mirrors the right half. Common examples: uniform
distribution (all values equally likely) and bell-shaped
(normal) distribution. Basis of many classical statistical
tests.
Skew Asymmetry in the shape of a frequency distribution.
The distribution has a longer "tail" on one side than the
other.
Positive (right) Right tail is longer; most values are at the low end with
skew a few extreme high values. The mean is pulled above
the median. Examples: income, waiting times, gene
expression levels.
Negative (left) Left tail is longer; most values are at the high end with
skew a few extreme low values. The mean is pulled below the
median. Examples: age at death in a well-nourished
population, human lifespan in wealthy countries.
Outlier An observation lying well outside the range of other
observations. May be a data error (investigate and
possibly correct/remove) or a genuine extreme
biological observation (keep and interpret). Always
investigate; never remove without justification.
Chartjunk Term coined by Tufte (1983): any visual element in a
graph that consumes ink without conveying information
about the data. Examples: 3-D effects, decorative
shading, redundant gridlines, background images.
Interferes with pattern recognition. Remove it.
Contingency A frequency table for two (or more) categorical
table variables, displaying frequency of every combination of
categories. Shows how the frequency distribution of
one variable depends on (is contingent upon) the
categories of another. Explanatory variable
conventionally in columns, response variable in rows.
2×2 contingency A contingency table where both variables have exactly
table two categories, resulting in four data cells. Most
common type in biology (e.g., treatment vs. control ×
diseased vs. not diseased).
Cell (in a One specific combination of row and column categories.
contingency Each individual is counted in exactly one cell. The
table) number in a cell is the frequency of that particular
combination of categories.
Explanatory The variable used to predict or explain the response
variable variable. In a scatter plot, placed on the x-axis. In a
contingency table, placed in the columns. Hypothesized
to cause or influence the response.
Response variable The variable being predicted or explained. In a scatter
plot, placed on the y-axis. In a contingency table,
placed in the rows. The outcome of interest.
Grouped bar A bar graph for two categorical variables
graph simultaneously. Bars are grouped by the explanatory
variable and coded (by color/shading) by the response
variable. Heights show frequencies. Y-axis must start at
zero. Useful for comparing frequency distributions
across groups.
Mosaic plot A graphical display for two categorical variables using
stacked rectangles. Stack width is proportional to
group size; rectangle height within a stack shows
relative frequency of each response category. Area of
each rectangle equals overall relative frequency of that
combination. Association visible as differing dividing-
line heights between stacks.
Association (in A relationship between two variables in which the value
graphs) of one provides information about the other. Positive
graphs) of one provides information about the other. Positive
(both increase together), negative (one increases as the
other decreases), or absent (no discernible
relationship).
Scatter plot A graphical display for two numerical variables, where
each observation is one point. X-axis = explanatory
variable; y-axis = response variable. Pattern of points
reveals direction, strength, and linearity of association.
The fundamental tool for bivariate numerical data.
Positive Points in a scatter plot trend from lower-left to upper-
association right — higher values of x accompany higher values of
y.
Negative Points trend from upper-left to lower-right — higher
association values of x accompany lower values of y.
Strip chart (dot A display for one numerical and one categorical
plot) variable in which each observation is a dot plotted at its
numerical value; group membership shown on the other
axis. Points jittered to reduce overlap. Ideal for small
samples (n < ~100 per group). Shows every individual
data point.
Violin plot A display for one numerical and one categorical
variable showing a smoothed, mirrored approximation
of the frequency distribution for each group. Width at
each value is proportional to frequency density. Best for
large samples where a strip chart would have severe
point overlap.
Multiple- Separate histograms for each group, stacked vertically
histogram and sharing the same x-axis scale for direct
method comparison. Shows distribution shape in detail. Best for
few groups with moderate-to-large n. Side-by-side (not
stacked) histograms are not recommended.
Line graph Uses dots connected by line segments to display a
summary measurement across an ordered series
(typically time). Connecting lines reveal trends, rates of
change, and temporal patterns. Baseline at zero makes
area under curve proportional to cumulative total.
Map (data Displays a numerical response variable at multiple
visualization) locations in physical space using color coding (gradient
or categorical). The spatial analogue of a line graph.
Can summarize enormous spatial datasets into single
interpretable images.
Display table A table designed to communicate a data pattern to a
general audience; appears in the main body of a paper.
Prioritizes clarity of pattern over completeness. Design
principles: compact, few significant figures,
rows/columns arranged to show pattern, categories
ordered by importance.
Data table A table storing raw data or detailed summaries for
reference, not communication. Usually too large for the
main body of a paper; placed in appendix or online
supplement. Prioritizes completeness and accuracy.
Inbreeding The probability that two gene copies in an individual
coefficient F are identical by descent (both trace back to the same
ancestor). F = 0: unrelated parents; F = 0.25: full
siblings (unrelated grandparents). Higher F → greater
inbreeding → increased homozygosity → inbreeding
depression.
CSV file format Comma-Separated Values: a plain text file format for
tabular data. Entries in the same row are separated by
commas; rows are separated by line breaks. Universal:
readable by all software now and in the future.
Recommended for storing biological data files.
Metadata / A separate file accompanying a data file, documenting:
auxiliary file data source, collection method, meaning of each
variable (with units), coding of categorical variables,
and any relevant notes. Ensures data remain
interpretable long after collection.
Tidy data Data organization principle: one row per
structure individual/observation; one column per variable.
Required by virtually all statistics and graphing
software. Avoids merged cells, subtotals, and
decorative formatting in data files.
Jitter The deliberate addition of a small random horizontal
displacement to data points in a strip chart, to reduce
overlap and make individual points visible. Does not
change the y-axis (measurement) values, only the visual
spread along the x-axis.
Data-ink ratio A principle from Tufte: maximize the proportion of "ink"
in a graph that encodes actual data. Remove everything
else (chartjunk). Every visual element should earn its
place by conveying information.
9. Quick-Reference: Which Graph for
Which Data?
Graph Selection Guide — Chapter 2 Summary
Recommended
Data Type / Goal Key Design Notes
Graph Type(s)
One categorical Bar graph; Frequency Y-axis starts at 0; bars
variable — show table separated; order nominal
frequency categories by frequency;
distribution report n
One categorical Bar graph (with Same rules as above; y-axis
variable — relative relative frequency on now 0–1 or 0–100%
frequencies y-axis)
One numerical Histogram; (Violin Bars contiguous; y-axis starts
variable — show plot or Strip chart as at 0; choose bin width
distribution alternatives) carefully; report n
Two categorical Mosaic plot; Grouped Try all three; mosaic often
variables — show bar graph; clearest for showing
association Contingency table association; grouped bar
shows absolute frequencies
Two numerical Scatter plot Explanatory variable on x-
variables — show axis; response on y-axis; note
association direction and strength of
association
One numerical + Strip chart (small n); Do NOT use bar graphs —
one categorical — Violin plot (large n); they hide the data; show
show differences Multiple histograms distribution shape within
between groups (few groups) groups
Time series — Line graph Connect time-ordered points
show trend over with line segments; y-axis at
time 0 makes area ∝ cumulative
total
Spatial data — Map (color gradient) Color encodes magnitude;
show variation "hotter" colors typically =
across locations higher values; include color
scale
CHAPTER 2 — CORE MESSAGE
Graph your data before doing anything else. Use graphs that
show the individual observations or the full frequency
distribution — not just averages. Eliminate chartjunk. Keep the
y-axis baseline at zero for bar graphs and histograms. Choose
the graph type that makes the main pattern in the data
immediately obvious to the eye. Organizing and storing data
properly ensures that these graphs can be created efficiently
and reliably, now and in the future.
Source: Whitlock, M.C., & Schluter, D. (2020). The Analysis of Biological Data (3rd ed.). Chapter
2: Displaying Data, pp. 115–182. MacMillan / W. H. Freeman.