Analyzing Lognormal Data: A Guide
Analyzing Lognormal Data: A Guide
Pharmacological Reviews
journal homepage: [Link]
REVIEW ARTICLE
Abstract 3
Significance Statement 3
I. Introduction 3
A. A motivating example 3
1. Incorrect analysis assuming sampling from normal distributions 3
2. Correct analysis assuming sampling from lognormal distributions 3
3. Why it can matter 4
B. History of lognormal distributions 4
C. Our goals in writing this review 4
II. Ratio scale variables 4
A. Definition of ratio scale variables 4
B. Examples of ratio variables and a counterexample 5
C. For ratio variables, experimental effects are best reported as ratios, not differences 5
III. Lognormal distributions 5
A. Multiplicative causes of variation lead to asymmetrical distributions 5
B. Review of logarithms 6
C. Logarithms convert a skewed distribution caused by multiplicative error to a symmetrical distribution 6
D. Relationships between normal and lognormal distributions 6
E. Lognormal distributions are common in biology and beyond 6
F. EC50, IC50, Kd, Km (and more) tend to be lognormal 7
1. Data demonstrating pharmacological parameters are lognormal 7
2. Simulations demonstrating pharmacological parameters are lognormal 7
G. Variables defined as the ratio of 2 lognormal distributions are lognormal 7
1. An interesting and impactful property of lognormal distributions 7
2. Examples in pharmacology where important parameters are the ratio of 2 lognormal variables 7
3. Why is the ratio of 2 lognormal variables lognormal? 8
IV. Descriptive statistics of lognormal distributions 8
A. The geometric mean 8
1. The GeoMean of an ideal lognormal distribution or population 8
2. How to calculate the GeoMean of a data set 8
3. Relationship between the GeoMean and the median 8
B. The geometric standard deviation 9
1. What is the GeoSD? 9
2. Calculating the GeoSD 9
3. A lognormal distribution with a small GeoSD is nearly identical to a normal distribution 9
4. How to write the GeoMean and GeoSD 9
C. Other ways to describe lognormal distributions 10
* Address correspondence to: Harvey J. Motulsky, GraphPad Software. E-mail: hmotulsky@[Link]; or Paul B.S. Clarke, Department of Pharmacology and Thera-
peutics, McGill University, 3655 Promenade Sir William Osler, Montreal, Quebec H3G 1Y6, Canada. E-mail: [Link]@[Link]
This article has supplemental material available at [Link].
[Link]
0031-6997/© 2025 The Authors. Published by Elsevier Inc. on behalf of American Society for Pharmacology and Experimental Therapeutics. This is an open access article
under the CC BY license ([Link]
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049
a r t i c l e i n f o a b s t r a c t
Associate Editor: Lynette Daws Lognormal distributions are pervasive in pharmacology and elsewhere in biomedical science, arising
naturally when biological effects multiply rather than add. Despite their ubiquity in pharmacological
parameters (eg, EC50, IC50, Kd, and Km), lognormal distributions are often overlooked or misunderstood,
leading to flawed data analysis. This largely nonmathematical review explains why lognormal distri-
butions are common, how to recognize them, and how to analyze them appropriately. We show that
many measured variables are lognormal. So are many derived parameters, particularly those defined as
ratios of lognormal variables. Through examples and simulations accessible to working scientists, we
demonstrate how misidentifying lognormal distributions as normal leads to reduced statistical power,
unnecessarily large sample sizes, false identification of outliers, and inappropriate reporting of effects
as differences rather than ratios. We challenge the common practice of using normality tests to decide
how to analyze data, showing that many data sets pass both normality and lognormality tests, espe-
cially with small sample sizes. Instead, we advocate for assuming lognormality based on the nature of
the variable. This review provides practical guidance on recognizing and presenting lognormal data,
and comparing data sets sampled from lognormal distributions. Based on Monte Carlo simulations, we
recommend the lognormal Welch’s t test or nonparametric Brunner-Munzel test for comparing 2
unpaired groups, the lognormal ratio paired t test for paired comparisons, and lognormal ANOVA for
3 groups. By recognizing and properly handling lognormal distributions, pharmacologists can design
more efficient experiments, obtain more reliable statistical inferences, and communicate their results
more effectively.
Significance Statement: Lognormal distributions are common in pharmacology and many scientific fields,
but they are often misunderstood or overlooked. This review provides a detailed guide to recognizing
and analyzing lognormal data, aiming to help pharmacologists perform more appropriate and more
powerful statistical analyses, draw more meaningful conclusions from their data, and communicate their
results more effectively.
© 2025 The Authors. Published by Elsevier Inc. on behalf of American Society for Pharmacology and
Experimental Therapeutics. This is an open access article under the CC BY license (http://
[Link]/licenses/by/4.0/).
I. Introduction of mean that everyone is familiar with. We will get to the geo-
metric mean in the next section.
A. A motivating example The P value (two-tailed) testing the null hypothesis that the sets
of data were sampled from identical normal distributions is
This motivating example demonstrates that analyzing 0.22. This is greater than the traditional threshold of 0.05, so the
lognormal data as if the values were sampled from a normal dis- null hypothesis of no difference would not be rejected.
tribution can lead to incorrect and misleading conclusions. Figure 1 Looking at the graph, the largest value in each group is much
compares EC50 values for control and treated conditions. larger than the rest. Indeed, Grubbs’ (1969) outlier test with a
set to 0.05 identified an outlier in each case.
1. Incorrect analysis assuming sampling from normal distributions 2. Correct analysis assuming sampling from lognormal distributions
Here are the results if the data were analyzed conventionally Now let us analyze correctly, assuming sampling from
with a 2-sample unpaired t test assuming sampling from normal lognormal distributions, using methods that will be explained in
distributions: detail below.
The means are 294 nM (control) and 575 nM (treated). The First, a quick reminder about samples and distributions.
difference is 282 nM (95% confidence interval [CI] of the Commonly used statistical tests (such as t tests) proceed by
difference: 175 to 739 nM). With such a wide CI), the data are assuming the null hypothesis, that is, that there is no real dif-
consistent with no difference, a moderate decrease, or a large ference between conditions (here, control vs treated). These
increase. In other words, no conclusion is possible. Note, here we tests then ask how frequently such an extreme (or even more
refer to the arithmetic mean (AMean)din other words, the type extreme) result would be obtained by taking random samples
3
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049
150 250
100
50
50
0 0
0 5 10 15 20 25 0 200 400 600
Sum of 4 dice Product of 4 dice
Fig. 2. Multiplicative factors lead to an asymmetrical distribution. The graphs show simulations of throwing 4 dice 1000 times. The graph on the left is a frequency distribution of
the sum of the values appearing on 4 dice. The graph on the right is a frequency distribution of the product of the 4 values. Our simulations were inspired by a blog by M.H. Nederlof
([Link]
It is not possible for either of those parameters to equal zero, but Consider the following example: a drug increases your measure
their values can be tiny and approach zero. from 5 to 10 U in first subject, from 10 to 20 U in a second subject,
Converting between units requires only multiplication or divi- and from 7 to 14 U in a third individual. These 3 results would be
sion, as is the case for weights, concentrations, durations, seen as identicaldthe drug has a doubling effect. In other words,
lengths, and EC50s. the underlying mechanism affecting changes in the variable is
It makes scientific sense to calculate a ratio. For example, 4 cm is multiplicative. The relative change is the same, regardless of the
twice as long as a distance of 2 cm; 6 L of water is 3 times as starting value. The unequal absolute differences (an increase of 5 U
much water as 2 L; and an EC50 of 5 mM is 5 times greater than vs 10 U vs 7 U) are most likely irrelevant.
an EC50 of 1 mM.
mechanisms that create distributions that are not lognormal but The 2 differences on the logarithmic scale are equidistant
look very similar. from log(1.0) ¼ 0.0, demonstrating that the logarithmic trans-
formation eliminates the asymmetry on the original scale, mak-
ing the effect of the multiplicative factor symmetrical on the
B. Review of logarithms
logarithmic scale.
The logarithm (base 10) of 1000 is the power of 10 that equals If an infinite population of values define a lognormal distri-
1000. The logarithm of 1000 is 3, because 103 ¼ 1000. The loga- bution, then the logarithms of those values define a normal
rithm of 10 is 1, because 101 ¼ 10. The logarithm of 0.001 is 3, distribution.
because 103 ¼ 1/103 ¼ 0.001. The logarithm of 3.162 is 0.5, If an infinite population of values define a normal distribution,
pffiffiffiffiffiffi
because 10 0.5 ¼ 10 ¼ 3.162. The logarithm of 1.0 is 0.0, because then the antilogarithms of those values define a lognormal dis-
10 0 ¼ 1.0. tribution. The antilogarithm is the inverse of the logarithm. The
The above logarithms are base 10 logarithms, also called antilogarithm of a common logarithm equals 10 to that power.
common logarithms, because the computations take 10 to some For example, the antilogarithm of 3 is 103 or 1000. The antilog-
power. They are sometimes written as “log10(x).” Mathemati- arithm of a natural logarithm equals e to that power. For example,
cians prefer natural logarithms using base e (2.718. . .), written as the antilogarithm of 6.908 is e6.908, often written as exp(6.908),
ln(x). Beware of the notation “log(x),” which can mean either which is 1000.
common or natural logarithm, depending on the field or
program. The term “lognormal,” often written “log-normal,” is potentially
The logarithms of values > 1.0 are positive. The logarithms of confusing as it can be incorrectly thought of as the “log of normal.”
values >0.0 and <1.0 are negative. The logarithms of zero and all But it is a mistake to think that values in a lognormal distribution
negative numbers are simply undefined, because there is no power are the logarithms of values from a normal distribution. We show it
of 10 that results in a negative number or zero. struck out because it is simply wrong. The term “antilognormal”
describes the distribution better than “lognormal” (Johnson et al,
C. Logarithms convert a skewed distribution caused by 1994), but that term is rarely (if ever) used.
multiplicative error to a symmetrical distribution The distribution of lognormal values is asymmetrical with a
positive skew, that is, with a long tail to the right (but as we will see,
Revisiting the earlier example where a factor randomly this asymmetry can be subtle in some cases). However, the distri-
doubles or halves a value of 10, resulting in 20 or 5, the loga- bution of their logarithms is symmetrical and forms a normal (ie,
rithms (base 10) of these values and the differences between Gaussian) distribution, as shown in Fig. 3.
them are:
E. Lognormal distributions are common in biology and beyond
logð20Þz1:30
There seems to be a prevailing sense that experimental data
logð10Þ ¼ 1
usually follow a normal distribution. But it has been realized for a
logð5Þz0:70 century that this is not true. A 75-year-old text states, “The normal
curve was, in fact, to the early statisticians what the circle was to
logð20Þ logð10Þz0:30 the Ptolemaic astronomers” (Yule and Kendall, 1950). We now
know that planets move in ellipses rather than circles, and that
logð10Þ logð5Þz0:30 lognormal distributions are pervasive.
Frequency
Fig. 3. Frequency distribution of a lognormal distribution. Left: A frequency distribution of a lognormal distribution with GeoMean ¼ 10 and GeoSD ¼ 2, defined later in this article.
Right: Frequency distribution of the common (base 10) logarithms of the values. Note that the log transformation converts the asymmetrical lognormal distribution to a sym-
metrical normal distribution.
6
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049
Lognormal distributions have been noted in data sets ranging the log-transformed data was closer to zero than the skewness of
from journal citation counts (Thelwall, 2016) to professorial sal- the raw data).
aries (Benzidia and Lubrano, 2020), and from pollution levels in We did similar simulations in more depth assessing the dis-
Los Angeles to the number of words in telephone conversations tribution of EC50, Kd, Koff, Kon, and Hill coefficient (Supplemental
(Limpert et al, 2001). Examples from biology and medicine Material). A lognormal distribution fitted all these parameter
include: neuronal firing rates in the central nervous system distributions better (larger R2) than a normal distribution did.
(Buzsa ki and Mizuseki, 2014); numerous blood analytes including The difference between the corrected Akaike information criteria
triglycerides (Carlson, 1960), free fatty acids (Heath, 1967), assesses how much better the data supports one model versus
ferritin (Custer et al, 1995), alkaline phosphatase, creatinine, another. If the difference is > 10, the worse-fitting model has
glucose, and iron (Flynn et al, 1974); abdominal fat (Stanforth essentially no support from the data (Burnham and Anderson,
et al, 2004), senile plaque size in Alzheimer disease (Hyman 2002; Portet, 2020). For our simulations, the smallest differ-
et al, 1995), blood pressure (Bodey and Michell, 1996), the num- ence in corrected Akaike information criteria was 30, demon-
ber of parasites per host (Shaw and Dobson, 1995), and cancer strating the simulated data for all the parameters support a
survival times (Qazi et al, 2007). For yet more examples, see lognormal distribution substantially better than a normal
Limpert et al (2001). distribution.
2. Simulations demonstrating pharmacological parameters are 2. Examples in pharmacology where important parameters are the
lognormal ratio of 2 lognormal variables
De Lean et al (1982) demonstrated that Kd values from simu- This relationship has ramifications in pharmacology because
lated repeated experiments are closer to lognormal than normal. many parameters are defined as the ratio of lognormal parameters.
Christopoulos (1998) used simulations to demonstrate that the Here are examples:
distributions of these parameters are lognormal: EC50, Hill coeffi-
cient, agonist efficacy t in the Black and Leff operational model of Bioequivalence. The comparison of drug formulations relies on
agonism (Black et al, 2010), the receptor/G-protein dissociation the ratio of areas under the curve between test and reference
equilibrium constant KG, and the cooperativity factor a. He found formulations. Area under the curve values tend to be lognor-
that the dissociation rate constant (Koff) was fit reasonably well by mally distributed because of the multiplicative nature of drug
both normal and lognormal distributions and concluded that Koff is absorption and elimination processes ([Link]
normal. But his data demonstrates that the lognormal distribution media/70958/download). Consequently, their ratio is
fits those values better than does the normal distribution (the sum- lognormal, which dictates the statistical approaches used in
of-squares was smaller for the lognormal fit, and the skewness of bioequivalence studies (Julious, 2004).
Table 1
Pharmacodynamic parameters reported to follow lognormal distributions
Tissue responses at fixed doses Tissue responses to fixed doses of norepinephrine or acetylcholine Fleming et al, 1972
EC50 Iris contraction induced by carbachol Patil, 1993
EC50 Cardiac muscle contraction induced by adrenaline and noradrenaline Kaumann et al, 1989
Ki and EC50 Ki and EC50 for atrial natriuretic peptide and norepinephrine Hancock et al, 1988
IC50 Cytotoxic effects of drugs on cancer cells in vitro Levasseur et al, 1998
Hill coefficient Electrophysiological response of rat olfactory receptor neurons to odorants Rospars et al, 2008
Hill coefficient Forceecalcium relationship in myocytes Walker et al, 2010
Kon, Koff, Kd Antibody-antigen binding Poulsen et al, 2011
Kon Microtubule binding to destabilizing kinesin-8 (Kif18B) Shrestha et al, 2023
Exponential time constants Single-molecule transport into liposomes, by a neurotransmitter: sodium symporter Fitzgerald et al, 2006
7
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049
Table 2
Pharmacokinetic parameters reported to follow lognormal distributions
Clearance, Vd, t1/2 Clearance, volume of distribution, and elimination half-life of theophylline after IV Steinijans et al, 1982; Julious and
infusion Debarnot, 2000
Cmax Either lognormal (Lacey et al) or not (Shen et al) Lacey et al, 1997; Shen et al, 2017
AUC Area under the curve (AUC) is close to lognormal Shen et al, 2017
Whole-body uptake Post-Chernobyl nuclear accident, whole-body counts of Cs-137 in women Wertelecki et al, 2016
Plasma levels Strontium plasma concentrations after oral administration to human volunteers Li et al, 2006
Drug concentration Intracellular drug concentration in tumors Zanotti-Fregonara and Hindie , 2011
Equilibrium dissociation constant (Kd). The Kd represents the arithmetic). As mentioned earlier, if a distribution is lognormal,
ratio of unbinding to binding rate constants (Koff/Kon). Both then the logarithm of that distribution is normal. Statisticians
rate constants arise from multiple molecular steps in protein- (and some biomedical scientists) sometimes report the mean
ligand interactions that combine multiplicatively, leading and SD of the log-transformed data. When interpreting such
to lognormal distributions. Their ratio, Kd, is therefore data, pay attention to whether common or natural logarithms
lognormal, impacting how we analyze drug-receptor binding were used.
data. Rather than reporting the summary statistics of the set of log-
The operational model. The transducer ratio t is defined as the arithms, we prefer to use the GeoMean and the geometric SD
total receptor density (termed R0 or Rt) divided by the (GeoSD).
coupling efficiency constant KE (Black and Leff, 1983; Kenakin
et al, 2012). Both parameters become lognormally distributed A. The geometric mean
through multiplicative cellular processesdRT through re-
ceptor expression and trafficking, KE through sequential steps 1. The GeoMean of an ideal lognormal distribution or population
in signal transduction. Consequently, their ratio t is The geometric mean, which we abbreviate GeoMean, is another
lognormal. name for the median of an ideal lognormal distribution. It is
Biased agonism. Ligand bias is quantified as the transduction expressed in the same units as the data. Figure 4 shows that the
coefficients (t/KA) for different signaling pathways. Because t/KA GeoMean, mode and AMean have different values for an ideal
represents a ratio of lognormally distributed parameters, it is lognormal distribution (left), but all are identical for an ideal
itself lognormal. When comparing 2 pathways (eg, G-protein normal distribution (right).
versus b-arrestin signaling), the standard approach calculates The concept of an AMean has been drilled into us since primary
the difference between log-transformed values for a test ligand school, and it may be hard to imagine how it is possible to have
[Dlog(t/KA)] and normalizes this to a reference ligand, yielding another kind of mean. Figure 5 shows one way to understand how
DDlog(t/KA) (Kenakin et al, 2012). Because the log-transformed there can be 2 distinct means.
transduction coefficients for individual pathways should be
normally distributed, their differences (Dlog and DDlog values)
2. How to calculate the GeoMean of a data set
will also be normal assuming that the pathways respond
To calculate a GeoMean, first calculate the common (ie, base 10)
independentlyda reasonable assumption given our under-
logarithms of all the values, compute their mean (call it m), and
standing of distinct signaling mechanisms. If the signaling
then calculate GeoMean ¼ 10m.
pathways are not independent, then a modified operational
Here, we use base 10 logarithms and the corresponding 10^
model has been proposed (Zhu et al, 2019).
antilog function, as this is what most biologists are familiar with.
Allosteric interactions. The cooperativity factor a is defined as a
Many statisticians, engineers, and physical scientists prefer the
ratio of binding constantsdspecifically, the ratio of a ligand’s
natural log (ln) and the corresponding exp() antilog function. The
binding affinity (Ki) when the allosteric site is empty to its
GeoMean (and the GeoSD) will be the same with either approach as
affinity when that site is occupied by a modulator (Ki(free)/
long as the same base is used for taking the logarithms and
Ki(co-bound)). Because both Ki values arise from binding
reversing that transform (antilog).
processes and are lognormally distributed, their ratio a is also
You will sometimes see the GeoMean defined as the n-th root of
lognormal (Christopoulos, 1998; Christopoulos and Kenakin,
the product of all values. As long as all values are positive, these 2
2002).
definitions are equivalent.
3. Why is the ratio of 2 lognormal variables lognormal?
Consider 2 lognormal variables X and Y. To understand the dis- 3. Relationship between the GeoMean and the median
tribution of their ratio Z ¼ X/Y, we can use a powerful mathematical As mentioned, for an ideal lognormal distribution, the GeoMean
tool: taking logarithms transforms multiplicative relationships into and median are identical.
additive ones. Take the logarithm of both sides: log(Z) ¼ log(X) For any particular data set randomly sampled from a
log(Y). Because X and Y are lognormal, both log(X) and log(Y) are lognormal distribution, the obtained GeoMean is equally likely to
normal by definition. When we subtract 2 independent normal be larger or smaller than the median. On average, the GeoMean is
variables, the result is also normal. The final step follows directly a more accurate estimate of the population median than is the
from the definition of a lognormal variable: because log(Z) is median computed from the same sample (Parkin, 1993; Vogel,
normal, Z itself must be lognormal. 2022). That is why the GeoMean is the standard way to quan-
tify the center of a lognormal distribution. However, the Geo-
IV. Descriptive statistics of lognormal distributions Mean can be a poor estimate of the median if data are sampled
from a skewed distribution that is not lognormal, or when data
A normal distribution is defined by 2 parameters: the mean are sampled from a (mostly) lognormal distribution with outliers
(ie, arithmetic mean or AMean) and the SD (ie, again, (Vogel, 2022).
8
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049
Fig. 4. The arithmetic mean, GeoMean, and median of an ideal lognormal distribution. Left: An ideal lognormal distribution showing distinct mode, arithmetic mean, and median
(GeoMean). The GeoMean is the median; half the values are larger, and half are smaller. The arithmetic mean is the center of gravity of the distribution. If you made the distribution
out of wood or plastic and included the tail that goes far beyond the right limit of the graph, it would balance at the arithmetic mean. Right: An ideal normal distribution showing
that the mode, arithmetic mean, and median are identical.
B. The geometric standard deviation The terms Geometric SD factor and multiplicative SD are
sometimes used as synonyms of GeoSD. Beware of the term s*,
1. What is the GeoSD? as it sometimes serves as an abbreviation for the SD of the nat-
The GeoSD (Kirkwood, 1979) quantifies both the spread and ural logarithms, and sometimes for the GeoSD (Limpert et al,
asymmetry of a lognormal distribution, as shown in Fig. 6. Unlike 2001).
the arithmetic (regular) SD, which has the same units as the data,
the GeoSD has no units. Multiplying the raw data by 1000, say, 3. A lognormal distribution with a small GeoSD is nearly identical
would make the arithmetic SD commensurately 1000 times larger, to a normal distribution
but would not change the GeoSD (Lewontin, 1966). The GeoSD al- A lognormal distribution with a small GeoSD closely resembles a
ways has a value 1.0. The GeoSD ¼ 1.0 only when all values are normal distribution (Fig. 7). How small does the GeoSD need to be?
identical. Any cutoff is somewhat arbitrary, but values of 1.2 (Limpert et al,
2001) and 1.3 (Elassaiss-Schaap and Duisters, 2020) have been
2. Calculating the GeoSD suggested.
The GeoSD is calculated as GeoSD ¼ 10s where s is the SD of the For variables that are always positive but are approximately
common logarithms of the values in the sample. Equivalently, normal, variation can be quantified as the coefficient of
GeoSD ¼ es , where s is the SD of the natural logarithms of the variation (CV) which equals SD/AMean. A lognormal distribution
values. Note that your choice to use natural or common logarithms with a small GeoSD looks very similar to a normal distribution
will alter the value of s but not GeoSD. with a CV ¼ GeoSD -1 (Haeckel and Wosniok, 2010). This
relationship makes sense because the minimum possible value
of the CV is 0.0, and the minimum possible value of the GeoSD is
1.0.
Figure 8 shows how similar normal and lognormal data can
appear. Each panel shows 10 simulated data sets (n ¼ 10 each). One
5 was sampled from a lognormal distribution with GeoSD ¼ 1.2 and
GeoMean ¼ 100, and the other was sampled from a normal dis-
tribution with a SD ¼ 20 and mean ¼ 100 (so the CV ¼ 0.2). It is
2 8 impossible to guess which is which (in fact, the left panel is
lognormal; the right panel is normal).
Figure 9 is identical, but with n ¼ 100 per group. You still cannot
tell which is normal and which is lognormal by inspection.
An example of a variable that fits both normal and lognormal
distribution is height. Slavskii et al (2021) reviewed the distri-
bution of height in multiple populations. Its CV is between 0.03
and 0.05, small enough so that normal and lognormal distribu-
tions both fit very well (but the lognormal distribution fits
4 slightly better, a finding that is only apparent with huge data
sets).
GeoSD = 1.25
GeoSD = 1.5
GeoSD = 2.0
0 1 2 3 4 0 1 2 3 4 0 1 2 3 4
GeoSD = 6 GeoSD = 10
GeoSD = 4
0 1 2 3 4 0 1 2 3 4 0 1 2 3 4
Fig. 6. The GeoSD quantifies asymmetry. All the graphs share the same area under the curves (1.0) and same GeoMean (1.0). Lognormal distributions with larger GeoSDs are more
skewed.
Enter the multiply superscripted (in the Insert Symbol dialog manuscript, it looks like an asterisk so does not communicate
of Word) followed by a slash: 103.0 / 4.53. This notation was the concept of multiply-or-divide very clearly. At the larger sizes
proposed by Limpert and Stahel (2011). used for presentations, the symbol is more understandable.
Enter a superscripted (and bold) period followed by a slash: If you want to avoid the phrase “multiplied or divided by,”
103.0 ./ 4.53. present a table with separate columns for GeoMean and GeoSD.
Insert a multiply sign and a superscripted ±1, for example
103.0 4.53±1. With the plus sign, this becomes 103.0 4.53.
With the minus sign, it becomes 103.0 4.531 ¼ 103.0 ÷ 4.53. C. Other ways to describe lognormal distributions
Use the Unicode “DIVISION TIMES” symbol (Uþ22C7) that su-
perimposes the multiply and divide ÷ symbols: 103.0 ⋇ 4.53. 1. The range that contains 68% or 95% of the values
To enter this character into Microsoft Word, enter “22C7” With a normal distribution, the range of values extending from
without the quotation marks, hold the Alt (Windows) or Option the [mean SD] to [mean þ SD] includes about two-thirds of the
(Macs) key, and tap X. Or use the Insert Symbol dialog. GraphPad population, more exactly 68.3%. With a lognormal distribution, the
Prism (starting with version 9) includes this symbol on the Math comparable range extends from [GeoMean/GeoSD] to [GeoMean
tab of its Insert Symbol dialog. At the font sizes used in a GeoSD] (Fig. 10).
Probability Density
Probability Density
10
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049
150 150
100 100
50 50
0 0
Fig. 8. Samples from lognormal vs normal populations. Each panel shows 10 simulated data sets (n ¼ 10 each). One was sampled from a lognormal distribution (GeoSD ¼ 1.2;
GeoMean ¼ 100), and the other from a normal distribution (SD ¼ 20; mean ¼ 100). You cannot tell which is which by looking at the graphs. (The graph on the left shows lognormal
data; the graph on the right shows normal data.)
Similarly, about 95% of values of a normal distribution are within sampling from a normal distribution would rarely create a data
the range [mean 2 SD] to [mean þ 2 SD]. The comparable range distribution that far (or further) from normal.
with a lognormal distribution is [GeoMean / GeoSD2] to
[GeoMean GeoSD2]. Note that the GeoSD is squared, not doubled. 2. How lognormality tests work
To test for lognormality, first transform all the data to their
2. Confidence interval of a GeoMean logarithms (it does not matter if you use common or natural log-
To calculate a CI of a GeoMeandperhaps better called a arithms). If the data were sampled from a lognormal distribution,
compatibility interval (Rafi and Greenland, 2020)dtransform the that set of logarithms would have been sampled from a normal
values to logarithms, compute the CI of the mean for the degree of distribution. Test this with one or more normality tests. If the P
confidence you want (usually 95%), and then reverse-transform the value is small, you will conclude that the distribution of the loga-
lower and upper confidence limits using the antilogarithm trans- rithms would be unlikely if the underlying distribution is normal,
form. The resulting interval will not be symmetrical around the so the distribution of the original values would be unlikely if the
GeoMean, but the asymmetry might be subtle. underlying distribution is lognormal.
200 200
150 150
100 100
50 50
0 0
Fig. 9. Even with n ¼ 100, a normal distribution with CV ¼ 20% is indistinguishable from a lognormal distribution with GeoSD ¼ 2.0. This matches Fig. 8, but with n ¼ 100 per data
set.
11
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049
0 GeoMean 2 4 6
Fig. 10. The middle 68.3% of a lognormal distribution. This lognormal distribution has a GeoMean ¼ 1 and a GeoSD ¼ 4. The shaded area extends from the GeoMean divided by the
GeoSD to the GeoMean multiplied by the GeoSD. If it looks to you like the shaded area contains > 68% of the total area, that is because the tail of the curve extends far to the right
beyond the limits of this graph and that tail has considerable area. In this example, the shaded range does not include the mode, the X-value at the peak of the curve. However, when
the GeoSD is much smaller, that range will include the mode.
GeoSD = 3
Probability Density
0 500 1000
Value
Fig. 11. Distributions used to sample data in Figs. 12 and 13. Both lognormal distributions have GeoSD ¼ 3. The GeoMeans differ by a factor of 3.2.
Figure 12 shows 4 simulated experiments of data from these Normality tests cannot reliably determine whether these dis-
distributions with n ¼ 10 per group. It is not always obvious by tributions are normal or lognormal. We ran 3 different normality
inspection that the data are not normal. tests (D’AgostinoePearson, ShapiroeWilk, and AndersoneDarling)
0 0 0 0
Control Treated Control Treated Control Treated Control Treated
Fig. 12. Random samples (n ¼ 10) from lognormal distributions shown in Fig. 11.
12
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049
0 0 0 0
Control Treated Control Treated Control Treated Control Treated
on 2500 simulated control data sets (n ¼ 10) similar to those in As mentioned earlier, only a ratio scale variable can be
Fig. 12. The simulated data sets failed the normality tests a bit more lognormal. When considering whether a variable can be
than half of the time (53%, 64%, and 66% for the 3 control sets; and lognormal, review this checklist of properties that define ratio
52%, 63%, and 65% for the treated data sets). In other words, nearly scale variables (discussed in the section Ratio scale variables):
half of the simulated lognormal data sets passed normality tests. ✓ Negative values must be inconceivable. Lognormal variables
Because the data are sampled from lognormal distributions, you can only be positive values.
would expect 5% of the simulated data sets to fail the lognormality ✓ Zero must either mean none of that variable or be the
tests with a set to 0.05, and indeed between 4.5% and 5.6% of the asymptotic value the variable approaches. However, no
simulated data sets failed each of the 3 tests. values can actually equal 0.0.
Normality tests can better detect the lack of normality of ✓ Converting between units must require only multiplication or
lognormal data when the sample size is larger. Figure 13 doubles division.
the sample size to 20 per group. Now, between 84% and 96% of the ✓ It must make scientific sense to calculate a ratio of 2 values.
simulated data sets (2 treatments; 3 normality tests; 10,000 sim-
ulations) fail the normality tests with P < .05. If a variable fails to meet any of these criteria, it is not a ratio
variable and so cannot be lognormal.
B. How to approach questions about lognormality
2. Consider the possibility that your data may follow a distribution
We have shown that normality tests can lead to inconsistent and that resembles lognormal
even incorrect conclusions, especially with small data sets. So how When working with positively skewed data, beware of other
should scientists decide when to assume lognormality? Here is the distributions that can resemble the lognormal distribution. Two
approach we recommend: such distributions are the g distribution and the distribution of the
ratio of 2 normally distributed variables.
Probability Density
Probability Density
0 5 10 15 0 1 2 3 4 5
Time Ratio
Fig. 14. A gamma distribution looks similar to a lognormal distribution. Equation Fig. 15. The distribution of the ratio of 2 normal distributions can look similar to a
where X is time (or duration), and k and theta are the 2 parameters that define the lognormal distribution. The graph shows the smoothed frequency distribution of
distribution (here, both are set to 2.0). Y ¼ (1/(gamma(k) * theta^k)) * X^(k-1) * exp(-X/ 10,000 ratios calculated with the numerator drawn from a normal distribution with
theta). mean ¼ 100 and SD ¼ 20, and the denominator drawn from a normal distribution with
mean ¼ 60 and SD ¼ 15.
400
100 100%
300 80 80%
EC50 (nM)
60 60%
200
SD
CV
40 40%
100
20 20%
0 0 0%
Control A B C 0 50 100 150 0 50 100 150
Mean Mean
Fig. 16. With lognormal distributions, the SDs are proportional to the mean. Left: Four data sets. Middle: The SDs are proportional to the mean. Right: The coefficient of variation (¼
SD/mean) is consistent for all 4 data sets. This is a clue that the data may be lognormal. In fact, the data were simulated from lognormal distributions with GeoSD ¼ 2.0 and various
GeoMeans.
where data are analyzed as if normal even though this rule of These problems, of course, are intrinsic to the difficulty in dis-
thumb is violated. tinguishing normal from lognormal distributions, not just running
normality tests. You will encounter the same issue by inspecting
frequency distributions or quantileequantile plots.
5. Do not base your decision only on normality and lognormality
tests D. Our recommendation: Choose to assume lognormality based on
Many scientists, we suspect, use the rule: Assume data are the nature of the variable without normality testing
normal until proven otherwise. Lognormality, for these scientists,
needs to be proven. In other words, “let the data decide.” Keene We strongly urge scientists to assume data are lognormal largely
(1995) presents 3 reasons why this is a bad idea (and we agree): based on the nature of the variable (or parameter) being compared,
and not to rely on normality or lognormality testing. This may
As we have seen in Figs. 8 and 9, the results of normality and sound like extreme advice far from the consensus, but others have
lognormality tests can be ambiguous. You may want to let the given the same advice:
data decide, but analyses of the data do not always lead to a clear
decision. Many data sets pass both normality and lognormality. “The theoretical justification for using this [the logarithmic]
Although it might seem logical to first run one test (here the transformation for most scientific observations is probably
normality test) and use that result to decide what to do next better than that for using no transformation at all… …If it were
(whether to log-transform the data), this kind of 2-stage the normal custom, when scientific observations show uncon-
approach to statistical testing can lead to misleading results so trolled variations large compared with the observations them-
is not recommended (Cartwright, 1991; Zimmerman, 1996, selves, to convert them to logarithms before estimating their
2004; Rochon et al, 2012; Gelman and Loken, 2014; Delacre et al, mean or variance, the usual result would be an increase in the
2017; Shamsudheen and Hennig, 2023). accuracy and scope of the conclusions drawn from them”
If you do multiple similar experiments, you might end up (Gaddum, 1945).
making different decisions for different experiments or even “I suggest that when an a priori decision about distribution has
different parts of the same experiment. It is better to analyze all to be made, the lognormal distribution should always be
similar data using the same assumptions. This is because the preferred over the normal distribution for data of this general
lognormality (or normality) assumption refers to the underlying type” (referring to variables, such as concentrations, that can
population, not to a particular sample. only have positive values) (Heath, 1967).
distribution that are negative
50%
% of values in a normal
40%
30%
20%
SD
10%
16%
0%
0.0 Mean 0 1 2 3 4
CV (=SD/Mean)
Fig. 17. A normal distribution with a high coefficient of variance (CV) contains a substantial number of negative values. Left: If the CV (which equals SD/mean) equals 1.0, then 16% of
the values in a normal distribution are negative. Right: The percentage of negative values in a normal distribution as a function of the CV calculated with this equation: 100 (1-
zdist (1/CV)). The blue dot shows that when CV ¼ 1, 16% of the values are negative. The green dot shows that when CV ¼ 0.6, only 5% of the values are negative. Equation: 100 (1-
zdist (1/CV)) (Prism format), 100 [Link] (1/CV, TRUE) (Excel format), or 100 pnorm (1/CV) (R format).
15
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049
“It is recommended that log transformed analyses should B. Outlier tests on the sample data
frequently be preferred to untransformed analyses, and that
careful consideration should be given to use of a log trans- When the raw data of Fig. 1 were analyzed using Grubbs’ outlier
formation at the protocol design stage. … If the use of a log test, the largest value in the treated group was identified as an
transformation is chosen on a case-by-case basis, then this will outlier as it is far from the rest of the data. But, as the next section
lead to inconsistencies and sometimes the wrong choice will be demonstrates, this test is very misleading with untransformed
made.” (Keene, 1995) lognormal data. After the sample data were log-transformed,
“It is proposed that all quantities should be considered to be Grubbs’ test did not identify an outlier in either group.
lognormal in clinical chemistry if the type of distribution is
unknown. Then, laboratories need not decide whether a distri- C. Outlier tests on lognormal data
bution is quasi-Gaussian or non-Gaussian.” (Haeckel and
Wosniok, 2010) Standard outlier tests fail dramatically with lognormal data.
“The lognormal distribution should be the first choice when Because lognormal distributions are naturally skewed, large values
modeling data taking (only) positive values. Its empirical as well that appear to be outliers are actually an expected feature of the
as theoretical justification is much stronger than for the normal distribution. Fig. 18 (left panel) demonstrates this problem using 20
distribution.” (Limpert and Stahel, 2017) simulated data sets (n ¼ 100 each, GeoMean ¼ 100, GeoSD ¼ 1.5).
“You should (usually) log transform your positive data. The Although many of these data sets look approximately normal,
reason for log transforming your data is not to deal with Grubbs’ outlier test (which assumes normality) incorrectly identified
skewness or to get closer to a normal distribution…The reason the largest value as an “outlier” (P < .05) in 12 of the 20 data sets.
for log transformation is in many settings it should make addi- The severity of this problem increases with both sample size and
tive and linear models make more sense. A multiplicative model GeoSD, as shown in the right panel of Fig. 18. Even with a very
on the original scale corresponds to an additive model on the log modest GeoSD of 1.25dwhere the distributions look nearly nor-
scale. For example, a treatment that increases prices by 2%, maldfar more than 5% of data sets have their largest value incor-
rather than a treatment that increases prices by $20. The log rectly flagged as an outlier. With larger sample sizes or larger
transformation is particularly relevant when the data vary a lot GeoSD values, nearly every data set has its largest value mis-
on the relative scale. Increasing prices by 2% has a much identified as an outlier.
different dollar effect for a $10 item than a $1000 item” (A. This problem is not limited to formal outlier tests. Visual in-
Gelman 2019; [Link] spection of data for outliers is equally unreliable with lognormal
21/you-should-usually-log-transform-your-positive-data/). data, as our eyes are naturally drawn to values that seem “too large”
when we expect a symmetric distribution. The key lesson is clear:
before testing for or removing outliers, you must first determine
VI. Do not use standard outlier tests with lognormal data whether your data might be lognormal. If the data are lognormal,
outlier detection should only be performed after logarithmic
A. Review of outlier tests transformation.
Most outlier tests evaluate whether extreme values are likely VII. Comparing 2 groups of lognormal data
to have come from a normal distribution. These tests compute
the probability that a value as extreme (or more extreme) as the A. Lognormal t test assuming sampling from lognormal
one observed would occur by chance if the data were sampled distributions with equal GeoSDs
from a normal distribution. If the P value is small (usually < .05),
that extreme value is identified as an “outlier,” and some in- 1. Calculating the lognormal t test and reporting the results
vestigators in some situations will remove that value from Figure 19 shows example data comparing 2 groups. With
further analyses (or might report results with and without the lognormal data, it is most common to compare the GeoMeans. For
outlier). this example, the GeoMean for the control values is 103 nM and the
In a set of many samples from a normal distribution, the most GeoMean for the treated samples is 302 nM. With lognormal data,
extreme value in 5% of those samples will be identified as an outlier. it rarely (if ever) makes sense to think about the absolute difference
Note that the 5% probability refers to the fraction of samples where between GeoMeans, but instead it makes scientific sense to think
the largest value is identified as an outlier, not the fraction of values about ratios (Wolfe and Carlin, 1999). The ratio is 302/103 ¼ 2.9. In
that are identified as outliers. other words, the treatment nearly tripled the EC50.
GeoSD=1.5
400 75% GeoSD=2.0
300
50%
200
25%
100
0 0%
N=5 N=20 N=100 N=1000
Fig. 18. Too many outliers identified with lognormal data. Left: Twenty simulated data sets with n ¼ 100, GeoMean ¼100, and GeoSD ¼ 1.5. The red dots on 12 of the data sets
denote outliers identified by Grubbs’ test (a ¼ 0.05). Right: Fraction of simulated lognormal data sets where Grubbs’ outlier test detected an outlier with P < .05. For each com-
bination of sample size and GeoSD, 1000 data sets were simulated. The horizontal line shows that Grubbs’ test identifies an outlier with P < .05 in 5% of data sets sampled from
normal distributions (with any sample size).
16
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049
4000 4
log(EC50, nM)
3000
EC50 (nM)
2000 2
1000 1
0 0
Control Treated Control Treated
Fig. 19. Example data for unpaired t test displayed on a linear axis (left) and after being transformed into common logarithms (right). These are the same data as in Fig. 1.
To calculate a CI and P value, we need to run a statistical test. 2. Graphing the results of a lognormal t test
This can easily be done by transforming all values to their loga- Figure 20 shows one way to plot the raw data and results of a
rithms to turn the lognormal distributions into normal distribu- lognormal t test. If you prefer plotting bar graphs with error bars,
tions, then running an unpaired t test on those logarithms. We call Figure 21 shows how. All 3 panels plot the GeoMeans. Lognormal
this the lognormal t test. The transformed values are shown on the distributions are asymmetrical, so the error bars in all 3 panels are
right panel of Fig. 19. Note that the distribution appears symmet- asymmetrical. The error bars on the left show the variation among
rical, as expected for data that are (before log transforming) the values, expressed as the GeoMean divided or multiplied by the
sampled from a lognormal distribution. The largest log- GeoSD. The error bars in the middle panel show how precisely the
transformed value in the treated group is just a bit larger than GeoMeans have been determined, expressed as the 95% CI of the
the rest, and Grubbs’ test on these data does not identify it as an GeoMeans. The error bars on the right show the GeoMean multi-
outlier. plied or divided by the geometric standard error (GeoSEM).
The transformed values were analyzed by a 2-sample t test Data are often presented as AMean with error bars representing
using GraphPad Prism 10.4, but any statistical software would the SD. If the values were sampled from a normal distribution, this
give the same result. The difference between the means of the would be straightforward to interpret. The range [mean SD] to
logEC50 values is 0.467. Recall that for any 2 positive values A [mean þ SD] would contain about two-thirds of the values. But this
and B, interpretation does not work when data are sampled from a
lognormal distribution (Fig. 22). When the GeoSD is reasonably
A A high so the data are noticeably asymmetrical, the symmetrical ± SD
logðAÞ logðBÞ ¼ log ; and so 10logðAÞlogðBÞ ¼ :
B B error bars do a poor job of displaying the distribution of the data. In
Therefore 100.467 ¼ 2.9 is the ratio of GeoMeans (as we already some cases, as shown in this example, the lower error bar can
determined). extend downward to a negative value, which makes no sense
The t test reports that the 95% CI for the difference between the because lognormal variables can never be negative.
means of the logarithms range from 0.09435 to 0.8409. Therefore,
the 95% CI for the ratio of GeoMeans ranges from 100.09435 to
100.8409, or 1.24 to 6.93. This gives a good sense of how precisely we 3. Terms to avoid when reporting ratio results
have determined the GeoMean ratio. When describing the relationship between 2 geometric means,
The above lognormal t test reports a 2-sided P value of .0154. it is essential to use clear and consistent language to avoid confu-
This P value tests the null hypothesis that the 2 sets of logEC50 sion. In the example above, the ratio of geometric means (treated/
values are sampled from normal distributions (or populations) control) is 2.9. Some scientists would report that value as a ratio,
that have identical means and SDs. Equivalently, if you refer and others would say that the treated response was 2.9 times the
instead to the EC50s (ie, not logged), then the null hypothesis is control response.
that the 2 sets of values are sampled from lognormal distributions The following terms are ambiguous, and we suggest simply
with identical GeoMeans and GeoSDs. If the null hypothesis were avoiding them:
true, then a ratio of GeoMeans of 2.9 or larger would occur only in
around 0.77% of experiments (one tail), and a ratio of 1/2.9 ¼ 0.34 Percent increase (or decrease). Some might state that the treated
or lower would occur in around 0.77% of experiments (the other GeoMean is 190% higher than that of the control (2.9 100%
tail). Why 0.77%? This is the P value (P ¼ .0154, ie, 1.54%), divided 100%). But it is also the case that the control GeoMean is 65.5%
equally between the lower and upper tails. The P value is less than lower than the treated GeoMean. The 2 are not symmetrical, so
the traditional cutoff (a) of .05. Therefore, if you chose that value we recommend avoiding both. Cole and Altman (2017) define a
of a and accept all the assumptions of a t test, you can reject that percentage difference that is symmetrical: 100 (difference/
null hypothesis. mean). In our example, the 2 GeoMeans were 103 nM (control)
Note the consistency of the CI and the P value. The P value is < and 302 nM (treated). The average is 202.5 nM, and the differ-
.05, and the 95% CI of the ratio (1.24e6.93, from above) does not ence is 199 nM. The difference/mean is 98% if you compute the
include the value that defines the null hypothesis (1.0). increase from control to treated, or 98% if you look at the
17
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049
10000
1000
EC50 (nM)
2.9x
100 5
.0 1
p=0
p = 0.015
10
95% CI
1
Control Treated 1 3 5 7 9
Ratio of GeoMeans
Fig. 20. Plotting the results of a 2-sample (unpaired) t test of lognormal data. Left: Raw data on logarithmic axis. Right: Ratio of GeoMeans with 95% CI.
decrease from treated to control. Although this definition is With the sample data (see Fig. 19), the lognormal Welch’s t test
symmetrical, it is used rarely. We do not recommend it. reports a P value of .016. The means of the logarithms differ by
Fold increase (or decrease). We agree with Small (2016) that the 0.4676 with a 95% CI ranging from 0.09349 to 0.8417. Take the
term fold increase should be avoided because it is used incon- antilog of all 3 values to obtain the ratio of GeoMeans (2.93) and its
sistently and so is ambiguous. Some would say there was “a 1.9- 95% CI (1.24e6.95). For this example, the lognormal t test (prior
fold increase”dthe difference between ratios of 1.0 (no change) section) and the lognormal Welch’s t test give nearly identical
and 2.9 (observed)dand others would say there was “a 2.9-fold results.
increase” (because the ratio is 2.9). Avoid this confusing term.
C. Comparing the lognormal t test with the lognormal Welch’s t test
B. The lognormal Welch’s t test
1. Power
In the previous example, we considered data sampled from 2 Figure 23 shows the results of simulations comparing the power
lognormal distributions with different geometric means but equal of the lognormal t test with the power of the lognormal Welch’s t
geometric SDs. Under these conditions, applying an unpaired t test test, when applied to lognormal data. They have nearly equal power
to the log-transformed datadreferred to as the lognormal t when the sample sizes are equal, even when the GeoSD values are
testdworks well because the 2 transformed distributions have unequal (left panel), but the lognormal Welch’s t test has more
similar variances. However, if the underlying distributions differ power when both sample sizes and GeoSD values differ (right
not only in their GeoMeans but also in their GeoSDs, the assump- panel).
tion of equal variances after log-transformation is violated. In this
scenario, the lognormal t test is no longer appropriate, and the log- 2. Type I error
transformed data should instead be analyzed using the Welch’s t The type I error rate is the probability of falsely rejecting the null
test. We refer to this approach as the lognormal Welch’s t test. hypothesis. With a set to 0.05, a well-behaved test should yield a
EC50 (nM)
0 0 0
Control Treated Control Treated Control Treated
GeoMean ⋇ GeoSD GeoMean with 95% CI GeoMean ⋇ GeoSEM
Fig. 21. Plotting the results of a 2-sample (unpaired) t test of lognormal data with error bars. Left: GeoMean, with error bars showing the GeoMean multiplied or divided by the
GeoSD. Middle: Same, but with 95% CI instead. Right: Same as the left panel, but replacing GeoSD with GeoSEM.
18
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049
100
Mean ± SD D. Nonparametric tests with lognormal data
1.0 1.0
0.8 0.8
0.6 0.6
Power
Power
0 0
0 10 20 30 40 0 10 20 30 40
n1 n1
Fig. 23. The power of the lognormal t test vs lognormal Welch’s t test, both applied to lognormal data. For each combination of GeoMean, GeoSD, and sample size, we simulated 2000 ex-
periments with values sampled from lognormal distributions, ran both statistical tests on each simulated data set, and tabulated the fraction of P values that are < .05 (our preset a). In both
panels, we set GeoMean 1 ¼1.6, GeoMean 2 ¼ 4.5, GeoSD 1 ¼1.2, and GeoSD 2 ¼ 6.0. However, in the right panel, one group is twice as larger as the other. The R code used for these simulations is
included in Supplemental Material. Left: The lognormal t test (blue) and lognormal Welch’s t test (red) have nearly equal power when GeoSD values are unequal, as long as sample sizes are
equal. Right: The lognormal Welch’s t test has more power when both sample sizes and GeoSD values differ, and when the group with the larger sample size also has a larger GeoSD.
19
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049
0.25 0.3
2 n1 = 0.5*n2
0.20 GeoSD 1 > GeoSD 2 GeoSD 1 > GeoSD 2
Type I error rate GeoMean 1 = GeoMean 2 GeoMean 1 = GeoMean 2
0.10
0.1
0.05
0 0
0 10 20 30 40 0 10 20 30 40
n1 n1
Fig. 24. The control of type I error by the lognormal t test and the lognormal Welch’s t test on lognormal data. For each combination of GeoMean, GeoSD, and sample size, we
simulated 2000 experiments sampled from lognormal distributions, ran both statistical tests on each simulated data set, and tabulated how frequently P < .05 (our preset a). The
GeoMeans are the same in both groups, that is, 1.6 (left panel) and 2.7 (right panel), whereas in both panels the GeoSD values differ (ie, GeoSD 1 ¼ 6.0 and GeoSD 2 ¼ 1.22). The R
code is included in the Supplemental Material. Left: Both lognormal t test (blue) and lognormal Welch’s t test (red) control the type I error reasonably well when GeoSD values are
unequal, as long as sample sizes are equal. Right: The lognormal Welch’s t test controls the type I error much better when both sample sizes and GeoSD values differ.
Mann-Whitney test (Karch, 2021). While not yet included in most When both sample sizes and GeoSDs are equal, the power of the
statistical software, the Brunner-Munzel test is available in R, Py- Mann-Whitney and Brunner-Munzel tests is almost identical, and is
thon, and Jamovi (Karch, 2023). close to the power of the lognormal t test and lognormal Welch’s t
test (Fig. 25, left panel). However, with unequal sample sizes and
2. The effect size reported by nonparametric tests GeoSDs, the power of the tests can vary considerably (see Fig. 25,
The Mann-Whitney test is often presented as a comparison of right panel). Under these conditions, the lognormal Welch’s t test
medians, which can seem appealing for analyzing data sampled has the most power. The nonparametric tests have a bit less power,
from lognormal distributions. Because the GeoMean of a lognormal with the Brunner-Munzel test performing better than the Mann-
sample estimates the median of its underlying distribution, one Whitney test. The lognormal t test has the least power of the 4 tests.
might be tempted to interpret the Mann-Whitney test as a
straightforward comparison of geometric means. Indeed, some
4. Type I error control
implementations of the Mann-Whitney test even report differences
Figure 26 compares the type I error rates of the statistical tests in
between medians with CIs. However, this interpretation is only
simulated lognormal data. In all these simulations, the GeoMeans
valid if the 2 distributions are identical in shape and differ only in
were set to the same value. When the distributions had identical
location (Stonehouse and Forrester, 1998; Fagerland and Sandvik,
GeoSDs and sample sizes (left panel), both nonparametric tests
2009).
maintained type I error rates near 0.05. However, when the dis-
This assumption that both distributions have the same shape
tributions have different GeoSDs (see Fig. 26, middle panel), the
does not hold with lognormal data (Divine et al, 2018). Although 2
type I error rate of the Mann-Whitney test is inflated, as previously
normal distributions will have the same shape when their SDs are
observed (Karch, 2021). The inflation of type I error rate for the
equal, 2 lognormal distributions with different GeoMeans will have
Mann-Whitney test is increased further when both sample sizes
different shapes even if they have the same GeoSD. Thus, the
and GeoSDs are set to unequal values (right panel). In contrast, the
assumption required to interpret the Mann-Whitney test as a me-
type I error rate of the Brunner-Munzel test remains near its
dian (or GeoMean) comparison rarely applies to lognormal data.
nominal value of 0.05 for all experimental designs, so long as the
Moreover, it has been shown that the Mann-Whitney test can yield
sample sizes in both groups are greater than or equal to about 10.
small P values when comparing lognormal distributions with the
same GeoMean but different GeoSDs, further complicating its
interpretation (Fagerland, 2012). 5. Conclusions about nonparametric tests for lognormal data
The results of a Brunner-Munzel test can be summarized as a For analyzing lognormal data, the lognormal Welch’s t test is the
probability of superiority (with a CI). This is the probability that a optimal choice for several reasons. First, it controls the type I error
randomly selected value from group A will be larger than a rate, that is, it provides acccurate P values when no effect exists.
randomly selected value from group B. For example, a probability of Second, it maximizes statistical power. Third, it provides easily
superiority of 0.8 indicates an 80% chance that a random value from interpretable effect sizes (ratios of GeoMeans), and directly ad-
group A exceeds a random value from group B. This is a less intu- dresses the data’s multiplicative nature. The lognormal t test has
itive way to summarize the result than the ratio of GeoMeans re- the same advantages except that when GeoSDs differ, it is less
ported by a lognormal t test. powerful than the lognormal Welch’s t test (see Fig. 25).
If you are not sure about the underlying distribution but suspect
3. The statistical power of nonparametric tests with lognormal data it might be lognormal, you may prefer to use a nonparametric test.
We ran Monte Carlo simulations to evaluate the power of If the experiment has equal sample sizes, both nonparametric tests
various statistical tests when analyzing lognormal data. For each perform similarly, but the Brunner-Munzel test has a bit more
experimental design, we specified the GeoMean, GeoSD, and power. If the sample sizes differ substantially, avoid the Mann-
sample size of each group. We simulated 2000 experiments, and Whitney test as it has lower power and larger type I errors than
tabulated the fraction of P values that are < .05 (our preset a). The R the Brunner-Munzel test. Thus, if you want to use a nonparametric
code used for these simulations is included in Supplemental test with lognormal data, choose the Brunner-Munzel test because
Material. it has more power and a smaller (more appropriate) type I error.
20
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049
1.0 1.0
0.8 0.8
0.6 0.6
Power
Power
0.4 0.4
n1 = n2 n1 = 0.5* n2
GeoSD 1 = GeoSD 2 GeoSD 1 < GeoSD 2
0.2 GeoMean 1 < GeoMean 2 0.2 GeoMean 1 < GeoMean 2
0.0 0.0
0 10 20 30 40 0 10 20 30 40
n1 n1
Fig. 25. Power of parametric and nonparametric tests applied to lognormal data. For each combination of GeoMean, GeoSD, and sample size, we simulated 2000 experiments,
sampling from lognormal distributions, ran 4 statistical tests on each simulated data set, and tabulated the fraction of P values that are < .05 (our preset a). The R code used for these
simulations is included in Supplemental Material. Left: The GeoSDs of the 2 groups were set to the same value (2.7), and the sample sizes were also made equal. The GeoMean of the
second group was 2.7 times the GeoMean of the first group (ie, 4.5 vs 1.6). The power of all the lognormal t test (blue, circles), lognormal Welch’s t test (red, circles), Mann-Whitney
test (blue, triangles), and Brunner-Munzel test (red, triangles) were nearly the same. Right: The sample sizes of the first group were set to half the sample size of the second, and the
GeoSD values were set to 1.2 and 6.0. Under these conditions, the power of the various tests differs.
E. Why the unpaired t test (without log transformation) should be When you run a t test assuming sampling from lognormal dis-
avoided with lognormal data tributions, the analogous effect size is reported as the ratio of the 2
GeoMeans. A ratio is a logical way to think about the size of an
1. Results of analyzing the sample data with an unpaired t test experimental effect. And it is helpful to report the ratio with its CI to
without log transformation give a sense of how precisely the ratio has been determined.
Let us return to our motivating example (see Fig. 1). The un-
paired t test (without log transformation) yielded P ¼ .22, so the
null hypothesis could not be rejected. The 95% CI for the difference b. Loss of statistical power (for a given sample size). Figure 27 pre-
between means ranged from 175 to 739. With such a wide CI, the sents additional simulations demonstrating that analyzing
data are consistent with no difference, a moderate decrease, or a lognormal data as normal results in a loss of statistical power. The
large increase. In other words, no conclusion is possible. simulated experiments had n ¼ 20 per group. All values were
sampled from lognormal distributions with GeoSD ¼ 4, and the
true effect size is 3 (the treated GeoMean is 3 times the control
2. Problems when analyzing lognormal data as normal GeoMean). The left side of the figure shows 1 of the 1000 simula-
tions. The right side shows P values for 1000 simulations.
a. Less useful effect size (difference, rather than ratio). When you All 1000 P values are graphed for all 4 analyses. For each test, the P
run a t test assuming sampling from normal distributions, the effect values vary over several orders of magnitude, more variability than
size is reported most simply as the difference between the 2 means. many scientists expect. Statistical power is defined as the probability
With ratio variables, the difference is rarely a useful way to view the that the P value will be < .05 (or any chosen cutoff), so is the fraction
effect, so many scientists do not bother reporting that difference or of the dots below the red line that defines P ¼ .05. Analyzing the data
its CI. correctly (assuming lognormal distributions) results in 69% power
n1 = n2 n1 = n2 n1 = 0.5*n2
0.25 0.25 0.25
GeoSD 1 = GeoSD 2 GeoSD 1 > GeoSD 2 GeoSD 1 > GeoSD 2
0.20
GeoMean 1 = GeoMean 2 0.20
GeoMean 1 = GeoMean 2 0.20
GeoMean 1 = GeoMean 2
Type I error rate
0 10 20 30 40 0 10 20 30 40 0 10 20 30 40
n1 n1 n1
Fig. 26. Type I error of parametric and nonparametric tests applied to lognormal data. The method matches that of Fig. 24. In all 3 panels, the GeoMeans of both groups were set to
1.6. Left: The GeoSDs of the 2 groups were set to the same value (6.0), and the sample sizes were also made equal. The type I error Mann-Whitney test (blue), and Brunner-Munzel
test (red) were nearly the same. Middle: The GeoSD values were set to 6.0 and 1.2, but the sample sizes were kept equal. Right: The GeoSDs were the same as in the middle panel,
but the sample sizes of the second group were set to half the sample size of the first. Under these conditions, the type I error of the Brunner-Munzel was always close to 0.05, but the
type I error of the Mann-Whitney test was several fold higher.
21
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049
0.1
0.05
5000
0.01
p-value
Response
0.001
2500
0.0001
GeoSD=4
0 0.00001 Effect = 3x
n = 20
Control Treated
0.000001
Fig. 27. Loss of statistical power when lognormal distributions are analyzed as normal. Left: A simulated experiment with n ¼ 20 per group, sampled from lognormal distributions
with GeoSD ¼ 4, and the true effect size is 3 (the treated GeoMean is 3 times the control GeoMean). Right: Results from 1000 simulated experiments with P values from 4 analyses
(from left to right): t test assuming lognormal data, Welch’s t tests assuming lognormal data, t test assuming normal data, and Welch’s t test assuming normal data. The powers (the
fraction of the P values < .05) are shown below each lane. The raw data are in Supplemental Material.
with the lognormal t test and 68% with the lognormal Welch’s t test. To summarize this contrived example: assuming a normal dis-
Analyzing the data incorrectly (assuming normal distributions) tribution requires a larger sample size: 64 per group versus 41 per
yields only 35% with the t test and 33% with the Welch’s t test. Similar group than if you assume the data are lognormal, a 56% increase.
simulations were reported by Fayers (2011). Figure 28 shows samples from the 2 hypothetical populations on
linear and logarithmic axes, and the right panel shows the needed
c. Increased sample size requirement (for constant power). sample size (per group) when computed correctly (assuming
Another way to assess how much it matters to identify lognormal lognormal) and incorrectly (assuming normal).
variables is to see how it impacts the calculation of necessary sample Limpert and Stahel (2011) used simulations with a variety of
size. GeoSD values (1.5e3.5) and a variety of intended effect sizes, in
Let us assume we are sampling from 2 lognormal distributions order to determine the degree to which sample size can be reduced
and running a 2-sample t test. The control GeoMean is 100 and we by recognizing lognormal distributions. They showed that incor-
are seeking sample size to detect a doubling to a GeoMean of 200. rectly assuming that lognormal data were sampled from normal
Assume GeoSD ¼ 3, and use conventional values for a (0.05, 2- distributions resulted in an increase of necessary sample size from
sided) and desired power (80%). somewhere between 20% and 300% depending on GeoSD and effect
Because the data are lognormal, a t test would be run on the size. One of their examples assumed GeoSD ¼ 2.4 (similar to our
logarithms of the values, so we need to convert to log scale before GEOSD ¼ 3), and they found that mistakenly analyzing the data as if
calculating necessary sample size. On a log scale, the hypothetical they were sampled from a normal distribution raised the required
means are log10(100) ¼ 2 and log10(200) ¼ 2.3, and the expected sample size from 10 to 16, a 60% increase (which matches Fig. 28).
SD ¼ log10(3) ¼ 0.477. For these parameters, sample size calculators
such as GraphPad Prism Cloud’s Power Analysis calculator or 3. Switching to the Welch’s t test does not solve the problem
G*Power (Faul et al, 2007) report the required sample size of 41 per The SDs in Fig. 1 differ considerably between the control and
group. Wolfe and Carlin (1999) presented equivalent calculations. treated groups. This is often the case for lognormal data, even when
But what if we wrongly assume the data are sampled from normal the GeoSDs are identical, because the SD of a lognormal distribution
distributions? This is a contrived situation, but let us do our best to do depends on both its GeoMean and GeoSD. These unequal SDs might
the corresponding sample size calculations assuming normal distri- tempt researchers to replace the usual unpaired t test with Welch’s
butions. Using the equation shown below, the corresponding AMeans t test, as this does not assume the variances (or SDs) of the groups
are 224.2 (control) and 448.4 (treated), and the corresponding SDs being compared are equal (Rasch et al, 2011; Delacre et al, 2017).
are 282.8 (control) and 565.7 (treated). The SDs differ because with With these sample data the results of the Welch’s t test (without
lognormal data with fixed GeoSD, the SD will be larger when the log transformation) are nearly identical to those of the regular t test.
GeoMean is larger. G*Power and Prism Cloud’s Power Analysis The P values (two-tailed) from both tests are .22. The CIs for the
calculator both allow you to specify different hypothetical SD values difference between means are also quite similar (t test, 175 to 739;
for the 2 populations. For this experimental design, both calculate Welch’s t test, 179 to 743).
that the resulting necessary sample size is 64 per group. Because the Welch’s t test does not assume the SDs are equal,
some might expect it to have more power with lognormal data. In
ln ðGeoSDÞ2
AMean ¼ GeoMean$e 2 fact, its power is a bit less than the t test with lognormal data
qffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi (Zimmerman and Zumbo, 1993; de Winter, 2016). Another reason
2 to avoid the Welch’s t test with lognormal data is that the observed
SD ¼ GeoMean2 $elnðGeoSDÞ 1
type I error rate of the Welch’s t test applied to data sampled from
22
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049
3000 10000 80
100 40
1000
10 20
0
1 0
Control Treated Control Treated Assume: Normal Lognormal
Fig. 28. Recognizing lognormality leads to smaller required sample size. The left and middle graphs show 100 values sampled from 2 hypothetical lognormal distributions with
GeoSD ¼ 3 with GeoMeans ¼ 100 and 200. The graph on the right shows the required sample sizes (per group) computed correctly assuming the data are lognormal or incorrectly
assuming sampling from normal distributions, setting a to 0.05 and desired power to 80%. The raw data are in Supplemental Material.
skewed distributions can be much higher than the preset value of a antilogarithm of all 3 values yields the results as ratios, which are far
(Ahad and Yahaya, 2014). easier to interpret (see Fig. 30, right panel). The GeoMean of the ratios
is 100.2075 ¼ 1.61, with a 95% CI ranging from 1.14 to 2.28. The 95% CI
VIII. Other comparisons of lognormal data does not include 1.0 (the value that denotes no change), which is
consistent with the P value being < .05. With GraphPad Prism (version
A. Paired t test of lognormal data 6.0 and later), choose the ratio paired t test to obtain the results
directly without needing to calculate logarithms and antilogarithms.
Figure 29 shows animal weight before and after an intervention Analyzed this way, the data suggest that the intervention
(left panel) and the absolute difference for each animal (right changes weight, but this is not super-convincing because the CI is
panel). The raw data are in Supplemental Material. Note that the set so wide, ranging from a 14% increase to a bit more than a doubling.
of differences is skewed, and appears to be lognormal (in fact, the
data were simulated so the differences are lognormal). B. One-way ANOVA with Dunnett’s test of lognormal data
1. Wrong analysis: Paired t test of untransformed data Figure 31 shows EC50 values collected in control conditions and
The paired t test looks at the set of differences and the null in the presence of 2 drugs in 7 experiments. In the presence of the
hypothesis that there is no effect of the intervention, in other words drugs, the EC50s tend to be larger so the pEC50s tend to be smaller.
that all the variation is due to normal random sampling. The P value
is .07 (two-tailed). The mean difference is 10.2 g with a 95% CI
ranging from 0.80 to 21.1 g. Because the P value is > .05, the 95% CI 1. Incorrect analysis assuming normal distributions
includes 0.0 (no difference). Interpretation of data always depends One-way ANOVA tests the null hypothesis that all 3 data sets
on the details. These data give a hint of a weight gain, but this large are sampled from normal distributions with the same mean and
a gain (or an equally large loss, because the P value is 2-sided) SD. One-way ANOVA of these EC50s results in P ¼ .07, high enough
would occur in 7% of experiments if the null hypothesis (no dif- that the null hypothesis is not rejected. Welch’s one-way ANOVA,
ference) is true. The CI (also called compatibility interval) is which does not assume equal SDs, results in a slightly higher P
consistent with a small decrease, no change, or a large increase. value (P ¼ .11).
Note that this analysis assumes the differences are sampled from a
normal distribution. 2. Analysis assuming lognormal distributions
Here are several reasons to assume these data in Fig. 31 are
2. Analysis of log-transformed data assuming lognormal sampled from lognormal distributions:
distribution of differences
There are 2 equivalent ways of thinking about running a EC50s are generally lognormal (Hancock et al, 1988; Kaumann
lognormal paired t test. et al, 1989; Christopoulos, 1998; Walker et al, 2010; Liang et al,
2015).
Transform all the values to logarithms and then run a paired t The SDs are very different for the 3 data sets (4, 50, and 150), but
test in order to test the null hypothesis that the means are equal. their CVs are more similar (63%, 77%, and 118%).
Compute the ratio of before/after for each animal, transform The CVs in all data sets are larger than 60%. As noted earlier, this
those ratios to logarithms, and run a 1-sample t test to test the (plus the fact that negative values are impossible) makes it
null hypothesis that the mean of those logarithms is zero exceedingly unlikely for the data to be sampled from a normal
(equivalently, the null hypothesis is that the GeoMean of the distribution.
ratios is 1.0).
To analyze the data assuming sampling from lognormal distri-
These 2 methods are totally equivalent. We prefer the second butions, the first step is to transform the values to logarithms. For
approach, because it focuses on before/after ratios. Figure 30 shows this example, we converted the data from nanomolar to molar,
the set of ratios and log (ratios), and the results of the ratio paired t transformed to log10 logarithms, and then multiplied those loga-
test. We used GraphPad Prism 10.4, but any statistical software rithms by negative 1 to compute the pEC50. Just as pH is the
would give the same result. The P value is .011. The mean log (ratio) negative logarithm of [Hþ], the pEC50 is the negative logarithm of
is 0.2075, with a 95% CI ranging from 0.0569 to 0.3580. Taking the EC50. For example, when the EC50 is 10 nM, which is 108 molar,
23
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049
80
the logEC50 is 8, and the pEC50 is 8. pEC50s are often used by For comparing 2 lognormal data sets, our simulations showed
pharmacologists to avoid negative numbers. the advantage of using the lognormal Welch’s t test (see section The
The middle panel (see Fig. 31) plots the pEC50 in the 3 groups. lognormal Welch’s t test). We suspect there are similar advantages to
One-way ANOVA of these values resulted in P ¼ .0003 (compare to using Welch’s ANOVA when analyzing the logarithms of lognormal
P ¼ .07 from ANOVA on untransformed data). Now focus on the data, but we have not run any simulations to test this idea.
right panel of Fig. 31. The follow-up Dunnett’s multiple-comparison
test compares the result of each drug to the control. Dunnett’s test C. Two-way ANOVA of lognormal data
reports the CI for the difference between logarithms. We trans-
formed the confidence limits to their antilogarithm to plot the ratio 1. Example and analysis assuming sampling from normal
of EC50s. For control versus drug A, the difference in logarithms is distributions
1.13 with a 95% CI ranging from 0.56 to 1.69. Transform all 3 values Suppose nicotine increases circulating levels of a certain hor-
to their antilogarithm (10 to those powers) and the ratio of EC50s is mone, and you wish to know whether the drug effect is different in
13.5, with the 95% CI ranging from 3.63 to 49.0. Because those CIs males and females. You measure hormone levels in control and
do not come close to 1.0 (the value that signifies no effect), the P drug-treated animals, in both sexes. The effect seems much larger
values are much < .05. in females (Fig. 32).
80 6 2.0
1.0
5
60 1.5 p = 0.011; n=12 pairs
log(weight, g)
4
Weight (g)
0.5
log(ratio)
Ratio
40 3 1.0
2 0.0
20 0.5
1.61
1
95% CI
0 0 0.0 –0.5
Before After Ratio Before After log(ratio) 0.5 1.0 1.5 2.0 2.5
Ratio
Fig. 30. Paired ratio t test. Left: Raw data and ratio. Middle: logarithms of raw data and ratio. Right: Summary of ratio t test showing: the GeoMean of the ratio, the 95% CI of that
ratio, and the P value testing the null hypothesis (ie, that the true ratio is 1.0).
24
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049
500 9
400
B / Control
EC50 (nM)
8 p = 0.0003
300
pEC50
200
7
A / Control
100
p = 0.0012
0 6
Fig. 31. One-way ANOVA example. Left: Raw data as EC50 in nanomolar. The values are in Supplemental Material. Middle: Transformed to pEC50 (convert nanomolar to molar,
transform to log10, then multiply by 1). Right: Results of Dunnett’s multiple comparisons test showing the 95% CI for the ratio of EC50s and the P values testing the null hypothesis
that the true ratio is 1.0. Both CIs and P values are corrected for multiple comparisons by the Dunnett calculations.
To demonstrate and quantify this result, run two-way ANOVA The largest value in the Male/Drug group is quite a bit larger
and focus on the results for interaction, which assesses whether the than the others, and is identified as an outlier by Grubbs’ outlier
drug effect differs between male and female animals. The interac- test (P < .01). But Grubbs’ test assumes the data (except for the
tion P value is tiny (<.0001). The drug effect is 25 ng/mL greater in possible outlier) are sampled from normal distributions. Values
females, with a 95% CI ranging from 18 to 32 ng/mL. This would be larger than the others are expected in data sampled from
convincing evidence of a SEX by DRUG interactiondif all the as- lognormal distributions.
sumptions behind the analysis are true.
For these reasons, especially the first, it makes sense to assume
lognormality, not normality. Can normality and lognormality tests help
ANOVA assumes all values are sampled from normal distribu-
decide? Not in this case. The data from the drug-treated males fail 3
tions. This is not obviously false, but there is a hint of
normality tests but pass all 3 lognormality tests, but the remaining 3 data
asymmetry.
sets pass 3 normality tests (P > .05) and also pass 3 lognormality tests.
ANOVA assumes that the underlying populations are not only
The right side of Fig. 32 shows the log-transformed data. Now
normal but that they all have the same SD. Here, the variability
the variation is similar across conditions, there are no longer any
(SD) clearly differs substantially between groups, violating a
obvious outliers, and the data look normally distributed.
major assumption underlying ANOVA.
Two-way ANOVA on the log-transformed data shows no evi-
2. Two-way ANOVA assuming sampling from lognormal dence of interaction (P ¼ .343), so it makes sense to look at the row
distributions (sex) and column (drug) effects. Tests of both null hypotheses (that
But… could the data be lognormal? There are reasons to think sex makes no difference, and that the treatment makes no differ-
so: ence) result in P < .0001. The effect of the drug is shown in the
ANOVA results as the difference between the mean of the log-
The measurement is concentration, a ratio variable which is transformed control and drug-treated values. It is reported as
often lognormal. 0.3301 (95% CI, 0.2852e0.3751). Use the 10^ (ie, 10 to the power of)
The variation is larger when the mean is larger, as expected for transform on all 3 values to express the drug effect as a ratio. The
lognormal data. drug-treated animals had a response 2.1 times that of the control
100 2.0
Log(Concentration, ng/ml)
Control
Concentration (ng/ml)
Drug
80
1.5
60
1.0
40
0.5 Control
20 Drug
0 0.0
Male Female Male Female
Fig. 32. Results of simulated experiment asking whether there is an interaction between drug response and sex. The values are in Supplemental Material. The solid horizontal lines
represent the arithmetic means.
25
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049
animals (95% CI, 1.9e2.4). Similar calculations show that the fe- Deleting such values would bias the results (leading to an
males had a response 3.1 times larger than the males (95% CI, increased GeoMean), because only the smallest value(s) would
2.8e3.4). be removed.
This example demonstrates: Replacing a “nondetect” value with the LOD would also bias the
results (increase the GeoMean) because the actual (unmeasur-
How one can be fooled by analyzing lognormal data as if the able) values are all less than the LOD.
values were sampled from normal distributions. Replacing nondetects with zero is not possible when assuming
Why it is important to differentiate additive effects from mul- lognormal distribution, because analyses of lognormal data first
tiplicative effects. Analyzing the raw data (left panel) suggested take the logarithm of all the values, and the logarithm of zero is
a substantial 2-way additive interaction, that is, with a larger not defined.
drug effect in females. However, analysis of the log-transformed
data (right panel) revealed no multiplicative interaction. Replacing nondetects with some value between zero and LOD is
That the decision of whether to assume sampling from the best solution. Verbosek (2011) used simulations of lognormal
lognormal distributions cannot depend entirely on normality data with different total sample size sizes and number of non-
and lognormality tests. detects, and recommends assigning a value of LOD/√2 to all values
that are less than the LOD.
In the example of Fig. 32, analyzing the data properly (after log The topic of how to deal with values too low to measure has
transformation) prevented falsely concluding there was an inter- been reviewed by Shoari and Dube (2018), but these authors do not
action. The converse can also happen: in some cases, analyzing the focus on lognormal data. Zhang et al (2009) recommend analyzing
data after log transformation reveals a multiplicative interaction data with too-small-to-measure values using nonparametric
that would have been missed had the data not been transformed. methods, and explain how to extend common nonparametric tests
to data with values below the detection limit. Zhou and Tu (1999)
devised a likelihood method for comparing data sets that are a
IX. Additional topics mixture of values from a lognormal distribution plus zeros.
A. How to handle values that are zero, negative, or below the limit B. Comparing the arithmetic means of lognormal distributions
of detection
Some researchers argue for reporting the AMean instead of the
1. If some values are zero or negative (occurs rarely) GeoMean in certain contexts. For example, Parkin and Robinson
By definition, an ideal lognormal distribution comprises only (1992) argue that the AMean provides a more meaningful com-
positive values. However, in the real world, some data sets that are parison than the GeoMean or median when comparing variables
close to lognormal nevertheless can contain values that are zero or such as pollutant concentrations across locations, where the
negative. This can occur in 3 ways: important consideration is the total mass of pollutant at a given
site.
Zeros can occur when the variable is a count, for example, Surprisingly, working with the AMean of lognormal data re-
number of immunopositive cells in a tissue section, or quires special methods. This is because the asymmetry of
number of days with rainfall. Lognormal distributions lognormal distributions causes random samples to often under-
describe continuous variables, so really are not appropriate represent large values. Thus, it is not appropriate to calculate the
for variables that are counted. The analysis of such data AMean by adding up all the values and dividing by the sample size,
should not be based on assuming sampling from a lognormal as this tends to underestimate the true population AMean. The
distribution. following papers describe appropriate procedures to compute the
Zeros can occur when the variable has a distribution that is not AMean and its CI (Zhou and Gau, 1997; Wu et al, 2003; Olsson,
entirely lognormal. One example is the Comet assay, used for 2005) and to compare AMeans of different groups (Zhou et al,
detecting DNA damage in eukaryotic cells (Bright et al, 2011). 1997).
The method uses gel electrophoresis to quantify the fraction of
the DNA that has been fragmented so appears in the tail of the C. The GeoMean as an average of ratios
“comet.” Combining data from many cells, the distribution is a
cluster of zeros plus a collection of values from a lognormal In this article, we explain the use of the GeoMean as a way to
distribution. Special methods are needed to analyze such data. summarize a set of values sampled from a lognormal distribution.
Subtracting a baseline or nonspecific value can lead to a differ- But the GeoMean is more widely applicable than that. It is the only
ence that is zero or negative. The true population value may be consistent way to average ratios (Fleming and Wallace, 1986),
larger than zero, but experimental (ie, random) error in total which makes it essential in areas such as physics and engineering
and/or baseline values can result in a zero or negative difference. (Mahajan, 2019), finance ([Link]
It is probably best to analyze such data without subtracting a investing/071113/[Link]), and other
baseline or nonspecific signal (and fit the baseline in the anal- domains ([Link]
ysis). Another approach is to add a positive constant to each statistics-for-data-visualizations-2619dbb3677a; Chargin, 2020).
value in the data set, so that all values become positive before This makes sense only when the ratios are unitless because they are
log transformation. the ratio of the same variable measured in 2 conditions, for
example, 2 treatments, 2 time points, or 2 genotypes.
This property makes geometric means essential in pharma-
2. If some values are below the detection limit (occurs rarely) cology whenever we need to average ratios, whether analyzing
With some experimental systems, a value may be too low to relative potencies or measuring fold-changes in receptor expres-
measure. You know the value cannot be zero (or negative), and that sion. The problem with using the AMean to summarize a set of
it is smaller than the limit of detection (LOD). Such values are called ratios is that the result depends on which group or treatment is
left-censored. How can such a value be accounted for? chosen as the baseline. For example, consider 2 experiments: in
26
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049
one, drug A is 3 times more potent than drug B, but in the other, it is 25
only one-third as potent. The AMean of these ratios is (3 þ 1/3)/2 ¼
1.67, implying that drug A is, on average, 1.67 times more potent
than drug B. If we instead measure the potency of B relative to A, we 20
arrive at the opposite conclusiondthat B is 1.67 times more potent
Skewness
than A. This inconsistency can lead to misleading interpretations.
15
In contrast, the geometric mean of 3.0 and one-third is 1.0,
correctly showing that neither drug is consistently more potent.
The GeoMean is a baseline-independent and consistent summary, 10
making it the appropriate method for averaging ratios.
5
D. Geometric Coefficient of Variation of lognormal data
ffi
p1 G. How much is lost when normal data are analyzed as if
GeoSEM ¼ GeoSD n
lognormal?
Note the similarity between the definitions of SEM and GeoSEM.
SEM equals SD multiplied by (1/√n), whereas the GeoSEM equals Analyzing lognormal data as if they were sampled from normal
GeoSD to the power of (1/√n). distributions can lead to major problems in data analysis. How bad
Beware of the earliest definition of the GeoSEM, which had the is the reverse issue: analyzing normal data as if they were sampled
same units as the data (Norris, 1940). This value was to be added to from lognormal distributions?
or subtracted from the GeoMean, which makes little sense for the This is a bit tricky to think through because all normal distri-
asymmetrical lognormal distribution. butions include negative values, which are impossible in lognormal
distributions. But if the CV is small enough, only a tiny fraction of
F. Why skewness is not a useful parameter with lognormal data samples will contain negative values. Therefore, it is only possible
to be unsure about whether data are sampled from normal versus
Skewnessdmore precisely Pearson’s moment coefficient of lognormal distributions when the CV is small. In this case, the
skewness, abbreviated G1dquantifies the asymmetry of a distri- lognormal distribution looks almost identical to a normal distri-
bution. A perfectly symmetrical distribution has a skewness of 0.0. bution, so the loss of power tends to be minimal (simulations not
Distributions with a long right tail, including lognormal distribu- shown).
tions, have positive skewness. Not surprisingly, there is a simple
relationship (Crow and Shimizu, 1988) between the GeoSD and the H. Performing lognormal comparisons with GraphPad Prism
skewness of a lognormal distribution (Fig. 33).
Although skewness of a lognormal distribution is related to Although data from lognormal distributions can be analyzed by
GeoSD, the left panel of Fig. 34 demonstrates 3 reasons why it is not programming languages such as R and Python, GraphPad Prism is
27
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049
30
11
9
Skewness
20
GeoSD
7
10 5
3
0 1
n= 5
50 0
n= 5
50 0
n= 0
n= 25
n= 00
n= 250
n= 0
n= 25
n= 00
n= 250
00
00
n= 00
n= 00
1
n=
1
n=
1
1
1
1
Fig. 34. Demonstration that skewness is not a useful measure of asymmetry of a lognormal data set. Each dot represents the analysis of one simulated data set drawn from a
lognormal distribution with GeoSD ¼ 3.0 and GeoMean ¼ 10. The left panel shows the values of Skewness from 100 simulated data sets of various sizes, and the right panel shows
the values of GeoSD from those same data sets. The horizontal black lines represent the medians. The horizontal red lines mark the population skewness (8.2) and GeoSD (3.0).
the only statistics program we know of that can directly perform t B. Lognormality in pharmacology
tests and ANOVA with lognormal data. Performing a paired t test
with lognormal data has been available since version 6. Choose the Measurements such as concentration, weight, and enzyme ac-
“ratio paired t test.” One- and 2-sample (unpaired) t tests and one- tivity are often lognormal.
way ANOVA can be done with lognormal data starting with version Key pharmacological parametersdincluding EC50, IC50, Kd, Km,
10.5. Simply select the option to assume sampling from lognormal Kon, Koff, clearance, and half-lifedfollow lognormal distributions.
distributions, and then choose between the lognormal t test and The ubiquity of lognormal distributions in pharmacology
the lognormal Welch’s t test. stems from the multiplicative nature of many chemical and
biological processes and also from the fact that a parameter
X. Summary formed as a ratio of 2 lognormal parameters will itself be
lognormal.
A. Properties of lognormal distributions
If a nonparametric approach is required, use the Brunner- To emphasize these points playfully, we conclude with a poem
Munzel test, which handles asymmetrical distributions better in the style of Dr Seuss.
than the Mann-Whitney test.
Oh the lognormal insights you will gain!
Do not rely on AMean ± SD error bars for lognormal data, as they
can be misleading when the true variation is asymmetrical. In When your data’s askew,
some cases, the lower error bar can even descend to an
And you don’t know what to do.
impossible negative value.
Express experimental effects on lognormal variables as ratios. A When your values spread wide,
75% decrease in EC50 represents the same effect size regardless of
All on the positive side.
the baseline EC50, just as a doubling in enzyme activity repre-
sents the same effect size regardless of the baseline activity. Avoid
For binding and clearance, EC50s galore,
the ambiguous terms fold change and percentage change.
For enzyme kinetics and so much more,
They multiply, multiply, that’s nature’s way!
E. Common misconceptions
Not adding like normal statistics would say.
Misconception: Lognormal distributions are rare special cases.
Reality: They are common. By welcoming lognormal, you’re thinking grows clear,
Misconception: Data should be considered normal until proven
Required sample size shrinks, no false outliers here.
lognormal.
Reality: For variables that must be positive, lognormal distribution Simple ratios illuminate the way,
is often more likely.
While absolute differences lead our insights astray.
Misconception: Lognormal distributions are always obviously
skewed.
When processes multiply rather than add,
Reality: With small GeoSD, they can be nearly symmetrical and
look very similar to normal distributions. Normal statistics can make results look bad.
Misconception: Effects should always be presented as absolute
Log-transform your data so analyses can thrive,
differences.
Reality: For lognormal variables, ratios are more meaningful Oh the insights you’ll gain, and the wisdom you’ll derive!
because they represent the same effect regardless of baseline.
Misconception: Log transformation is a form of p-hacking (invalid
data manipulation). Declaration of generative AI and AI-assisted technologies in
Reality: It is a valid statistical choice when justified by the nature of the writing process
the variable and should be prespecified in analysis plans.
Misconception: If a data set passes a normality test, the data During the preparation of this work the author(s) used [Link]
cannot be sampled from a lognormal distribution. 3 to enhance the manuscript’s clarity and conciseness, verify
Reality: Many data sets pass both normality and lognormality tests citation-reference consistency, and generate alternative versions of
(especially with small sample sizes). the concluding poem. After using this tool, the author(s) reviewed
Misconception: Standard outlier tests work for any distribution. and edited the content as needed and take(s) full responsibility for
Reality: These tests are invalid for untransformed lognormal data the content of the publication.
and can lead to inappropriate exclusion of legitimate high values.
Misconception: Reporting lognormal analyses requires that your Abbreviations
readers are facile with logarithms.
Reality: Results can be presented in original units using ratios and AMean, arithmetic mean; CI, confidence interval; GeoCV, geo-
GeoMeans without mentioning logarithms. metric coefficient of variation; GeoMean, geometric mean; GeoSD,
Misconception: CIs are always symmetrical. geometric standard deviation; GeoSEM, geometric standard error;
Reality: For lognormal data, the CIs of a GeoMean and the CI of a Kd, equilibrium dissociation constants; Km, Michaelis constant; Koff,
ratio of 2 GeoMeans are asymmetrical. dissociation rate constant; Kon, association rate constant; LOD, limit
of detection; LR, likelihood ratio; pEC50, negative logarithm (base
F. Perspective 10) of the EC50.
owns it. Trajen Head is the Senior Product Manager for Prism, and is de Winter J (2016) A case against the default use of Welch’s t-test. Int Rev Soc
Psychol 30:92e101.
a minority shareholder of the company that owns GraphPad Soft-
De Lean A, Hancock AA, and Lefkowitz RJ (1982) Validation and statistical analysis
ware. Paul B.S. Clarke declares no conflict of interest. of a computer modeling method for quantitative analysis of radioligand binding
data for mixtures of pharmacological receptor subtypes. Mol Pharmacol 21:
5e16.
Data availability Delacre M, Lakens Danie €l, and Leys C (2017) Why psychologists should by default
use Welch’s t-test instead of Student’s t-test. Int Rev Soc Psychol 30:92e101.
n AE, and Juarez-Colunga E (2018) The Wilcox-
The authors declare that all the data supporting the findings of Divine GW, Norton HJ, Baro
oneManneWhitney procedure fails as a test of medians. Am Stat 72:278e286.
this study are contained within the manuscript and Supplemental Elassaiss-Schaap J and Duisters K (2020) Variability in the log domain and limita-
Material. tions to its approximation by the normal distribution. CPT Pharmacometrics Syst
Pharmacol 9:245e257.
Fagerland MW (2012) T-tests, non-parametric tests, and large studiesda paradox of
Authorship contributions statistical practice? BMC Med Res Methodol 12:78.
Fagerland MW and Sandvik L (2009) The WilcoxoneManneWhitney test under
scrutiny. Stat Med 28:1487e1497.
Performed data analysis: Motulsky, Head.
Faul F, Erdfelder E, Lang AG, and Buchner A (2007) G*Power 3: a flexible statistical
Wrote or contributed to the writing of the manuscript: Motulsky, power analysis program for the social, behavioral, and biomedical sciences.
Head, Clarke. Behav Res Methods 39:175e191.
Fayers P (2011) Alphas, betas and skewy distributions: two ways of getting the
wrong answer. Adv Heal Sci Educ 16:291e296.
Supplemental material Fitzgerald JB, Schoeberl B, Nielsen UB, and Sorger PK (2006) Systems biology and
combination therapy in the quest for clinical efficacy. Nat Chem Biol 2:458e466.
Fleming PJ and Wallace J (1986) How not to lie with statistics: the correct way to
This article has supplemental material available at pharmrev. summarize benchmark results. Commun ACM 29:218e221.
[Link]. Fleming WW, Westfall DP, De la Lande IS, and Jellett LB (1972) Log-normal distri-
bution of equieffective doses of norepinephrine and acetylcholine in several
tissues. J Pharmacol Exp Ther 181:339e345.
References Flynn FV, Piper KAJ, Garcia-Webb P, McPherson K, and Healy MJR (1974) The fre-
quency distributions of commonly determined blood constituents in healthy
Ahad NA and Yahaya SSS (2014) Sensitivity analysis of Welch’s t-test. AIP Conf Proc blood donors. Clin Chim Acta 52:163e171.
1605:888e893. Gaddum JH (1945) Lognormal distributions. Nature 156:463e466.
Aitchison J and Brown JAC (1957) The lognormal distribution with special reference Galton F (1879) XII. The geometric mean in vital and social statistics. Proc R Soc Lond
to its uses in economics. J R Stat Soc Ser A (Gen) 120:481e482. 29:365e367.
Anderson TW and Darling DA (1954) A test of goodness of fit. J Am Stat Assoc 49: Gelman A and Loken E (2014) The statistical crisis in science. Am Sci 102:460.
765e769. Glantz S (2011), 7th ed Primer of Biostatistics, McGraw Hill, New York, NY.
Baldi B and Moore D (2017) Practice of Statistics in the Life Sciences, 4th ed, WH Glaser A (2018), 4th ed High-Yield Biostatistics, Epidemiology, and Public Health,
Freeman, New York, NY. Lippincott Williams & Wilkins, Philadelphia, PA.
Benzidia M and Lubrano M (2020) A Bayesian look at American academic wages: Grubbs FE (1969) Procedures for detecting outlying observations in samples.
from wage dispersion to wage compression. J Econ Inequal 18:213e238. Technometrics 11:1e21.
Black J and Leff P (1983) Operational models of pharmacological agonism. Proc Royal Haeckel R and Wosniok W (2010) Observed, unknown distributions of clinical
Society London B 220:141e162. chemical quantities should be considered to be log-normal: a proposal. Clin
Black JW, Leff P, Shankley NP, and Wood J (2010) An operational model of phar- Chem Lab Med 48:1393e1396.
macological agonism: the effect of E/[A] curve shape on agonist dissociation Hancock AA, Bush EN, Stanisic D, Kyncl JJ, and Lin CT (1988) Data normalization
constant estimation. Br J Pharm 160(Suppl 1):S54eS64. before statistical analysis: keeping the horse before the cart. Trends Pharmacol
Bland M (2015) An Introduction to Medical Statistics, 4th ed, Oxford University Press, Sci 9:29e32.
Oxford, UK. Havlicek LL and Peterson NL (1974) Robustness of the t test: a guide for researchers
Bodey AR and Michell AR (1996) Epidemiological study of blood pressure in do- on effect of violations of assumptions. Psychol Rep 34:1095e1114.
mestic dogs. J Small Anim Pr 37:116e125. Heath D (1967) Normal or log-normal: appropriate distributions. Nature 213:
Bright J, Aylott M, Bate S, Geys H, Jarvis P, Saul J, and Vonk R (2011) Recommen- 1159e1160.
dations on the statistical analysis of the Comet assay. Pharm Stat 10:485e493. Hyman BT, West HL, Rebeck GW, Buldyrev SV, Mantegna RN, Ukleja M, Havlin S, and
Brunner E and Munzel U (2000) The nonparametric Behrens-Fisher problem. Biom J Stanley HE (1995) Quantitative analysis of senile plaques in Alzheimer disease:
42:17e25. observation of log-normal size distribution and molecular epidemiology of
Burnham K and Anderson D (2002) Model Selection and Multimodel Inference: A differences associated with apolipoprotein E genotype and trisomy 21 (Down
Practical Information-Theoretic Approach, 2nd ed, Springer, New York, NY. syndrome). Proc Natl Acad Sci 92:3586e3590.
Buzsaki G and Mizuseki K (2014) The log-dynamic brain: how skewed distributions Irizarry R and Love M (2016), 1st ed Data Analysis for the Life Sciences with R,
affect network operations. Nat Rev Neurosci 15:264e278. Chapman and Hall/CRC, Boca Raton, FL.
Carlson LA (1960) Serum lipids in normal men. Acta Med Scand 167:377e397. Johnson NL, Kotz S, and Balakrishnan N (1994), 2nd ed Continuous Univariate
Cartwright AC (1991) International harmonization and consensus DIA meeting on Distributions, Wiley Interscience, Hoboken, NJ.
bioavailability and bioequivalence testing requirements and standards. Ther Julious SA (2004) Sample sizes for clinical trials with normal data. Stat Med 23:
Innov Regul Sci 25:471e482. 1921e1986.
Chargin W (2020) Why ratios want geometric means. Chargin blog. Viewed January Julious SA and Debarnot CAM (2000) Why are pharmacokinetic data summarized
11, 2025 from. [Link] by arithmetic means? J Biopharm Stat 10:55e71.
Christopoulos A (1998) Assessing the distribution of parameters in models of Karch JD (2021) Psychologists should use Brunner-Munzel’s instead of Mann-
ligandereceptor interaction: to log or not to log. Trends Pharmacol Sci 19: Whitney’s U test as the default nonparametric procedure. Adv Methods Pr
351e357. Psychol Sci 4:2515245921999602.
Christopoulos A and Kenakin T (2002) G protein-coupled receptor allosterism and Karch JD (2023) bmtest: a jamovi module for BrunnereMunzel’s testda robust
complexing. Pharmacol Rev 54:323e374. alternative to WilcoxoneManneWhitney’s test. Psych 5:386e395.
Cole TJ and Altman DG (2017) Statistics notes: what is a percentage difference? BMJ Kaumann AJ, Hall JA, Murray KJ, Wells FC, and Brown MJ (1989) A comparison of the
358:j3663. effects of adrenaline and noradrenaline on human heart: the role of 1- and 2-
Cox N (2010) Speaking Stata: the limits of sample skewness and kurtosis. Stata J 10: adrenoceptors in the stimulation of adenylate cyclase and contractile force. Eur
482e495. Hear J 10:29e37.
Crow E and Shimizu K (1988) Lognormal Distributions: Theory and Applications. Keene ON (1995) The log transformation is special. Stat Med 14:811e819.
Taylor and Francis, New York, NY. Kenakin T, Watson C, Muniz-Medina V, Christopoulos A, and Novick S (2012)
Curran-Everett D (2018) Explorations in statistics: the log transformation. Adv A simple method for quantifying functional selectivity and agonist bias. ACS
Physiol Educ 42:343e347. Chem Neurosci 3:193e203.
Custer EM, Finch CA, Sobel RE, and Zettner A (1995) Population norms for serum Kirby W (1974) Algebraic boundedness of sample statistics. Water Resour Res 10:
ferritin. J Lab Clin Med 126:88e94. 220e222.
D’Agostino RB, Belanger A, and D’Agostino RB Jr (1990) A suggestion for using Kirkwood T (1979) Geometric means and measures of dispersion. Biometrics 35:
powerful and informative tests of normality. Am Stat 44:316e321. 908e909.
Dancey C, Reidy J, and Rowe R (2012), 1st ed Statistics for the Health Sciences: A Koch AL (1966) The logarithm in biology 1. Mechanisms generating the log-normal
Non-Mathematical Introduction, SAGE Publications Ltd, London, UK. distribution exactly. J Theor Biol 12:276e290.
Daniels W and Cross C (2018), 11th ed Biostatistics: A Foundation for Analysis in the Koch AL (1969) The logarithm in biology II. Distributions simulating the log-normal.
Health Sciences, Wiley, Hoboken, NJ. J Theor Biol 23:251e268.
30
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049
Koopmans LH, Owen DB, and Rosenblatt JI (1964) Confidence intervals for the co- Shapiro SS and Wilk MB (1965) An analysis of variance test for normality (complete
efficient of variation for the normal and log normal distributions. Biometrika 51: samples). Biometrika 52:591e611.
25. Shaw DJ and Dobson AP (1995) Patterns of macroparasite abundance and aggre-
Lacey LF, Keene ON, Pritchard JF, and Bye A (1997) Common noncompartmental gation in wildlife populations: a quantitative review. Parasitology 111(Suppl):
pharmacokinetic variables: are they normally or log-normally distributed? S111eS133.
J Biopharm Stat 7:171e178. Shen M, Russek-Cohen E, and Slud EV (2017) Checking distributional assumptions
Levasseur LM, Faessel H, Slocum HK, and Greco WR (1998) Implications for clinical for pharmacokinetic summary statistics based on simulations with compart-
pharmacodynamic studies of the statistical characterization of an in vitro mental models. J Biopharm Stat 27:756e772.
antiproliferation assay. J Pharmacokinet Biopharm 26:717e733. Shoari N and Dube J (2018) Toward improved analysis of concentration data:
Lewontin R (1966) On the measurement of relative variability. Syst Zool 15: embracing nondetects. Environ Toxicol Chem 37:643e656.
141e142. Shrestha S, Ems-McClung SC, Hazelbaker MA, Yount AL, Shaw SL, and Walczak CE
Li WB, Ho€llriegl V, Roth P, and Oeh U (2006) Human biokinetics of strontium. Part I: (2023) Importin a/b promote Kif18B microtubule association and enhance
intestinal absorption rate and its impact on the dose coefficient of 90Sr after microtubule destabilization activity. Mol Biol Cell 34:ar30.
ingestion. Radiat Environ Biophys 45:115e124. Slavskii SA, Kuznetsov IA, Shashkova TI, Bazykin GA, Axenovich TI, Kondrashov FA,
Liang H, Li J, Di Y, Zhang A, and Zhu F (2015) Logarithmic transformation is essential and Aulchenko YS (2021) The limits of normal approximation for adult height.
for statistical analysis of fungicide EC50 values. J Phytopathol 163:456e464. Eur J Hum Genet 29:1082e1091.
Limpert E, Stahel W, and Abbt M (2001) Log-normal distributions across the sci- Small DS (2016) Let’s abolish fold higher and fold increase from our lexicon. Int J
ences: keys and clues. BioScience 51:341e352. Pharmacokinet 1:13e15.
Limpert E and Stahel WA (2011) Problems with using the normal distribution e and Stanforth PR, Jackson AS, Green JS, Gagnon J, Rankinen T, Despre s JP, Bouchard C,
ways to improve quality and efficiency of data analysis. PLoS One 6:e21403. Leon AS, Rao DC, Skinner JS, et al (2004) Generalized abdominal visceral fat
Limpert E and Stahel WA (2017) The log-normal distribution. Significance 14:8e9. prediction models for black and white adults aged 17e65 y: the HERITAGE
Mahajan S (2019) Don’t demean the geometric mean. Am J Phys 87:75e77. Family Study. Int J Obes 28:925e932.
Martinez MN and Bartholomew MJ (2017) What does it “mean”? A review of Steinijans VW, Eicke R, and Ahrens J (1982) Pharmacokinetics of theophylline in
interpreting and calculating different types of means and standard deviations. patients following short-term intravenous infusion. Eur J Clin Pharmacol 22:
Pharmaceutics 9:14. 417e422.
McAlister D (1879) XIII. The law of the geometric mean. Proc R Soc Lond 29: Stevens SS (1946) On the theory of scales of measurement. Science 103:677e680.
367e376. Stonehouse JM and Forrester GJ (1998) Robustness of the t and U tests under
Moser BK, Stevens GR, and Watts CL (1989) The two-sample t test versus sat- combined assumption violations. J Appl Stat 25:63e74.
terthwaite’s approximate f test. Commun Stat Theory Methods 18:3963e3975. Thelwall M (2016) Citation count distributions for large monodisciplinary journals.
Motulsky H (2017), 4th ed Intuitive Biostatistics, Oxford University Press, Oxford, J Inf 10:863e874.
UK. Thom H (1958) A note on the gamma distribution. Mon Weather Rev 86:
Norris N (1940) The standard errors of the geometric and harmonic means and 117e122.
their application to index numbers. Ann Math Statist 11:445e448. Verbosek T (2011) A comparison of parameters below the limit of detection in
Olsson U (2005) Confidence intervals for the mean of a log-normal distribution. geochemical analyses by substitution methods. RMZ M&G 5:393e404.
J Stat Educ 13. [Link] Vogel RM (2022) The geometric mean? Commun Stat Theory Methods 51:82e94.
Ott W (1995) Environmental Statistics and Data Analysis. CRC Press, Boca Raton, FL. Wahi M and Puzzullo J (2024), 2nd ed Biostatistics for Dummies, For Dummies,
Parkin TB (1993) Evaluation of statistical methods for determining differences be- Hoboken, NJ.
tween samples from lognormal populations. Agron J 85:747e753. Walker JS, Li X, and Buttrick PM (2010) Analyzing forceepCa curves. J Muscle Res Cell
Parkin TB and Robinson JA (1992) Analysis of lognormal data. Adv Soil Sci 20:193e235. Motil 31:59e69.
Patil PN (1993) Reactivity of human iris-sphincter to muscarinic drugs in vitro. Wertelecki W, Koerblein A, Ievtushok B, Zymak-Zakutnia N, Komov O, Kuznietsov I,
Naunyn Schmiedebergs Arch Pharmacol 347:568. Lapchenko S, and Sosyniuk Z (2016) Elevated congenital anomaly rates and
Portet S (2020) A primer on model selection using the Akaike information criterion. incorporated cesium-137 in the Polissia region of Ukraine. Birth Defects Res A
Infect Dis Model 5:111e128. Clin Mol Teratol 106:194e200.
Posten H, Yen H, and Owen D (1982) Robustness of the two-sample t-test under Wolfe R and Carlin JB (1999) Sample-size calculation for a log-transformed outcome
violations of the homogeneity of variance assumption. Commun Stat 11: measure. Control Clin Trials 20:547e554.
109e126. Wu J, Wong ACM, and Jiang G (2003) Likelihood-based confidence intervals for a
Poulsen TR, Jensen A, Haurum JS, and Andersen PS (2011) Limits for antibody af- log-normal mean. Stat Med 22:1849e1860.
finity maturation and repertoire diversification in hypervaccinated humans. Yule G and Kendall M (1950), 14th ed An Introduction to the Theory of Statistics,
J Immunol 187:4229e4235. Hafner, New York, NY.
Proost JH (2019) Calculation of the coefficient of variation of log-normally distrib- Zanotti-Fregonara P and Hindie E (2011) Lognormal distribution of cellular uptake
uted parameter values. Clin Pharmacokinet 58:1101e1102. of radiopharmaceuticals: implications for biologic response in cancer treat-
Qazi S, DuMez D, and Uckun F (2007) Meta analysis of advanced cancer survival ment. J Nucl Med 52:501e503.
data using lognormal parametric fitting: a statistical method to identify effec- Zar J (2009), 5th ed Biostatistical Analysis, Pearson, Upper Saddle River, NJ.
tive treatment protocols. Curr Pharm Des 13:1533e1544. Zhang D, Fan C, Zhang J, and Zhang C (2009) Nonparametric methods for mea-
Rafi Z and Greenland S (2020) Semantic and cognitive tools to aid statistical sci- surements below detection limit. Stat Med 28:700e715.
ence: replace confidence and significance by compatibility and surprise. BMC Zhou X and Gau S (1997) Confidence intervals for the log-normal mean. Stat Med
Med Res Methodol 20:244. 16:783e790.
Ramsey PH (1980) Exact type I error rates for robustness of Student’s t test with Zhou X and Tu W (1999) Comparison of several independent population means
unequal variances. J Educ Stat 5:337e349. when their samples contain log-normal and possibly zero observations. Bio-
Rasch D, Kubinger KD, and Moder K (2011) The two-sample t test: pre-testing its metrics 55:645e651.
assumptions does not pay off. Stat Pap 52:219e231. Zhou XH, Gao S, and Hui SL (1997) Methods for comparing the means of two in-
Rochon J, Gondan M, and Kieser M (2012) To test or not to test: preliminary dependent log-normal samples. Biometrics 53:1129e1135.
assessment of normality when comparing two independent samples. BMC Med Zhu X, Finlay DB, Glass M, and Duffull SB (2019) An intact model for quantifying
Res Methodol 12:81. functional selectivity. Sci Rep 9:2557.
Rospars JP, Lansky P, Chaput M, and Duchamp-Viret P (2008) Competitive and Zimmerman DW (1987) Comparative power of Student t test and Mann-
noncompetitive odorant interactions in the early neural coding of odorant Whitney U test for unequal sample sizes and variances. J Exp Educ 55:
mixtures. J Neurosci 28:2659e2666. 171e174.
Royston P (1992) Estimation, reference ranges and goodness of fit for the three- Zimmerman DW (1996) Some properties of preliminary tests of equality of variances
parameter log-normal distribution. Stat Med 11:897e912. in the two-sample location problem. J Gen Psychol 123:217e231.
Ruxton GD (2006) The unequal variance t-test is an underused alternative to Stu- Zimmerman DW (2004) A note on preliminary tests of equality of variances. Br J
dent’s t-test and the ManneWhitney U test. Behav Ecol 17:688e690. Math Stat Psychol 57:173e181.
Shamsudheen I and Hennig C (2023) Should we test the model assumptions before Zimmerman DW and Zumbo BD (1993) Rank transformations and the power of the
running a model-based test? J Data Sci Stat Vis 3. [Link] Student t test and Welch t’ test for non-normal populations with unequal
jdssv.v3i3.73. variances. Can J Exp Psychol 47:523e539.
31
With lognormal data, absolute differences between GeoMeans rarely provide meaningful insight because the data's multiplicative nature means that ratios better represent the relationship between groups. As described in Wolfe and Carlin (1999), the treatment in the example nearly tripled the EC50, reflected by a ratio of 2.9. This demonstrates how ratios encapsulate the proportional effect of treatments on lognormal data, offering a clearer scientific narrative .
Skewness in lognormal distributions necessitates larger sample sizes to achieve statistical significance, as skewed data increase variability, making it harder to detect genuine effects. The document highlights that analyzing such data as normal can drastically amplify the required sample size, sometimes by as much as 300% .
Outliers can significantly skew results, leading to false-positive findings, especially in lognormal data. The document suggests transforming data using logarithms, such as by employing a lognormal t test, which helps mitigate the impact of outliers and increases the robustness of the test outcomes .
The lognormal Welch’s t test is preferred over the lognormal t test when GeoSDs differ because it maintains control of the type I error rate and maximizes statistical power in these circumstances. This is in contrast to the lognormal t test, which becomes less powerful when GeoSDs are unequal .
Using Welch’s t test on untransformed lognormal data can lead to similar power and type I error rates compared to the regular t test, making it less effective. This indicates that log transformation is necessary to truly benefit from Welch’s test's usual ability to handle unequal variances across groups .
When dealing with lognormal data and unequal sample sizes, one should consider using the Brunner-Munzel test instead of the Mann-Whitney, as the latter shows lower power and higher type I error rates under these conditions. This choice ensures better statistical power and maintains an appropriate type I error rate .
Using untransformed data in a t test on lognormal distributions can lead to inaccurate results because such data violate the assumption of normality required by the t test. Skewed distributions increase the required sample size and can elevate type I error rates and reduce statistical power, leading to unreliable conclusions .
Transforming lognormal data into log-transformed values approximates a normal distribution, which satisfies many statistical test assumptions, such as those of t tests. This transformation allows for more accurate statistical modeling by providing symmetrical distributions that align with the normality assumptions of the test .
One should opt for the Brunner-Munzel test instead of the Mann-Whitney when sample sizes are unequal, as it better controls type I errors and offers more statistical power under these experimental designs .
Ratios of GeoMeans accurately represent the multiplicative nature of lognormal data, providing easily interpretable effect sizes and reliable comparisons across groups. This translates the complex interactions into straightforward ratios, which align better with the inherent data patterns and scientific interpretations .