0% found this document useful (0 votes)
27 views31 pages

Analyzing Lognormal Data: A Guide

This review article discusses the prevalence and significance of lognormal distributions in pharmacology and other scientific fields, emphasizing the importance of correctly identifying and analyzing such data. It provides practical guidance on recognizing lognormal data, appropriate statistical methods for analysis, and the implications of misidentifying lognormal distributions as normal. The authors advocate for assuming lognormality based on the nature of the variable rather than relying solely on normality tests, aiming to enhance the reliability of statistical inferences in pharmacological research.

Uploaded by

profjesg
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd
0% found this document useful (0 votes)
27 views31 pages

Analyzing Lognormal Data: A Guide

This review article discusses the prevalence and significance of lognormal distributions in pharmacology and other scientific fields, emphasizing the importance of correctly identifying and analyzing such data. It provides practical guidance on recognizing lognormal data, appropriate statistical methods for analysis, and the implications of misidentifying lognormal distributions as normal. The authors advocate for assuming lognormality based on the nature of the variable rather than relying solely on normality tests, aiming to enhance the reliability of statistical inferences in pharmacological research.

Uploaded by

profjesg
Copyright
© All Rights Reserved
We take content rights seriously. If you suspect this is your content, claim it here.
Available Formats
Download as PDF, TXT or read online on Scribd

Pharmacological Reviews 77 (2025) 100049

Pharmacological Reviews
journal homepage: [Link]

REVIEW ARTICLE

Analyzing lognormal data: A nonmathematical practical guide


Harvey J. Motulsky 1, * , Trajen Head 2 , Paul B.S. Clarke 3, *
1
GraphPad Software, Los Angeles, California
2
GraphPad Software, Boston, Massachusetts
3
Department of Pharmacology and Therapeutics, McGill University, Montreal, Quebec, Canada

Abstract 3
Significance Statement 3
I. Introduction 3
A. A motivating example 3
1. Incorrect analysis assuming sampling from normal distributions 3
2. Correct analysis assuming sampling from lognormal distributions 3
3. Why it can matter 4
B. History of lognormal distributions 4
C. Our goals in writing this review 4
II. Ratio scale variables 4
A. Definition of ratio scale variables 4
B. Examples of ratio variables and a counterexample 5
C. For ratio variables, experimental effects are best reported as ratios, not differences 5
III. Lognormal distributions 5
A. Multiplicative causes of variation lead to asymmetrical distributions 5
B. Review of logarithms 6
C. Logarithms convert a skewed distribution caused by multiplicative error to a symmetrical distribution 6
D. Relationships between normal and lognormal distributions 6
E. Lognormal distributions are common in biology and beyond 6
F. EC50, IC50, Kd, Km (and more) tend to be lognormal 7
1. Data demonstrating pharmacological parameters are lognormal 7
2. Simulations demonstrating pharmacological parameters are lognormal 7
G. Variables defined as the ratio of 2 lognormal distributions are lognormal 7
1. An interesting and impactful property of lognormal distributions 7
2. Examples in pharmacology where important parameters are the ratio of 2 lognormal variables 7
3. Why is the ratio of 2 lognormal variables lognormal? 8
IV. Descriptive statistics of lognormal distributions 8
A. The geometric mean 8
1. The GeoMean of an ideal lognormal distribution or population 8
2. How to calculate the GeoMean of a data set 8
3. Relationship between the GeoMean and the median 8
B. The geometric standard deviation 9
1. What is the GeoSD? 9
2. Calculating the GeoSD 9
3. A lognormal distribution with a small GeoSD is nearly identical to a normal distribution 9
4. How to write the GeoMean and GeoSD 9
C. Other ways to describe lognormal distributions 10

* Address correspondence to: Harvey J. Motulsky, GraphPad Software. E-mail: hmotulsky@[Link]; or Paul B.S. Clarke, Department of Pharmacology and Thera-
peutics, McGill University, 3655 Promenade Sir William Osler, Montreal, Quebec H3G 1Y6, Canada. E-mail: [Link]@[Link]
This article has supplemental material available at [Link].

[Link]
0031-6997/© 2025 The Authors. Published by Elsevier Inc. on behalf of American Society for Pharmacology and Experimental Therapeutics. This is an open access article
under the CC BY license ([Link]
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

1. The range that contains 68% or 95% of the values 10


2. Confidence interval of a GeoMean 11
V. How to decide if a variable is lognormal 11
A. Do not rely on tests of normality and lognormality 11
1. Review of normality tests 11
2. How lognormality tests work 11
3. Normality and lognormality tests are often inconclusive 11
B. How to approach questions about lognormality 13
1. Ask the right questions 13
2. Consider the possibility that your data may follow a distribution that resembles lognormal 13
a. The g distribution resembles the lognormal distribution 13
b. The distribution of the ratio of 2 normal distributions resembles a lognormal distribution 13
C. What to do when you are unsure about lognormality 13
1. Do not get fooled by variables that are already log transformed 13
2. Compare the consistency of the SD versus the consistency of the CV 14
3. Calculate the likelihood ratio of sampling from normal versus lognormal distributions 14
4. Consider the value of the CV 14
a. If the CV is small 14
b. If the CV is large 14
5. Do not base your decision only on normality and lognormality tests 15
D. Our recommendation: Choose to assume lognormality based on the nature of the variable without normality testing 15
VI. Do not use standard outlier tests with lognormal data 16
A. Review of outlier tests 16
B. Outlier tests on the sample data 16
C. Outlier tests on lognormal data 16
VII. Comparing 2 groups of lognormal data 16
A. Lognormal t test assuming sampling from lognormal distributions with equal GeoSDs 16
1. Calculating the lognormal t test and reporting the results 16
2. Graphing the results of a lognormal t test 17
3. Terms to avoid when reporting ratio results 17
B. The lognormal Welch’s t test 18
C. Comparing the lognormal t test with the lognormal Welch’s t test 18
1. Power 18
2. Type I error 18
3. When to use the lognormal Welch’s t test 19
D. Nonparametric tests with lognormal data 19
1. Understanding the Mann-Whitney and Brunner-Munzel tests 19
2. The effect size reported by nonparametric tests 20
3. The statistical power of nonparametric tests with lognormal data 20
4. Type I error control 20
5. Conclusions about nonparametric tests for lognormal data 20
E. Why the unpaired t test (without log transformation) should be avoided with lognormal data 21
1. Results of analyzing the sample data with an unpaired t test without log transformation 21
2. Problems when analyzing lognormal data as normal 21
a. Less useful effect size (difference, rather than ratio) 21
b. Loss of statistical power (for a given sample size) 21
c. Increased sample size requirement (for constant power) 22
3. Switching to the Welch’s t test does not solve the problem 22
VIII. Other comparisons of lognormal data 23
A. Paired t test of lognormal data 23
1. Wrong analysis: Paired t test of untransformed data 23
2. Analysis of log-transformed data assuming lognormal distribution of differences 23
B. One-way ANOVA with Dunnett’s test of lognormal data 23
1. Incorrect analysis assuming normal distributions 23
2. Analysis assuming lognormal distributions 23
C. Two-way ANOVA of lognormal data 24
1. Example and analysis assuming sampling from normal distributions 24
2. Two-way ANOVA assuming sampling from lognormal distributions 25
IX. Additional topics 26
A. How to handle values that are zero, negative, or below the limit of detection 26
1. If some values are zero or negative (occurs rarely) 26
2. If some values are below the detection limit (occurs rarely) 26
B. Comparing the arithmetic means of lognormal distributions 26
C. The GeoMean as an average of ratios 26
D. Geometric Coefficient of Variation of lognormal data 27
2
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

E. Geometric standard error of a geometric mean 27


F. Why skewness is not a useful parameter with lognormal data 27
G. How much is lost when normal data are analyzed as if lognormal? 27
H. Performing lognormal comparisons with GraphPad Prism 27
X. Summary 28
A. Properties of lognormal distributions 28
B. Lognormality in pharmacology 28
C. Recognizing lognormal data 28
D. Analyzing lognormal data 28
E. Common misconceptions 29
F. Perspective 29
I. Declaration of generative AI and AI-assisted technologies in the writing process 29
References 30

a r t i c l e i n f o a b s t r a c t

Associate Editor: Lynette Daws Lognormal distributions are pervasive in pharmacology and elsewhere in biomedical science, arising
naturally when biological effects multiply rather than add. Despite their ubiquity in pharmacological
parameters (eg, EC50, IC50, Kd, and Km), lognormal distributions are often overlooked or misunderstood,
leading to flawed data analysis. This largely nonmathematical review explains why lognormal distri-
butions are common, how to recognize them, and how to analyze them appropriately. We show that
many measured variables are lognormal. So are many derived parameters, particularly those defined as
ratios of lognormal variables. Through examples and simulations accessible to working scientists, we
demonstrate how misidentifying lognormal distributions as normal leads to reduced statistical power,
unnecessarily large sample sizes, false identification of outliers, and inappropriate reporting of effects
as differences rather than ratios. We challenge the common practice of using normality tests to decide
how to analyze data, showing that many data sets pass both normality and lognormality tests, espe-
cially with small sample sizes. Instead, we advocate for assuming lognormality based on the nature of
the variable. This review provides practical guidance on recognizing and presenting lognormal data,
and comparing data sets sampled from lognormal distributions. Based on Monte Carlo simulations, we
recommend the lognormal Welch’s t test or nonparametric Brunner-Munzel test for comparing 2
unpaired groups, the lognormal ratio paired t test for paired comparisons, and lognormal ANOVA for
3 groups. By recognizing and properly handling lognormal distributions, pharmacologists can design
more efficient experiments, obtain more reliable statistical inferences, and communicate their results
more effectively.

Significance Statement: Lognormal distributions are common in pharmacology and many scientific fields,
but they are often misunderstood or overlooked. This review provides a detailed guide to recognizing
and analyzing lognormal data, aiming to help pharmacologists perform more appropriate and more
powerful statistical analyses, draw more meaningful conclusions from their data, and communicate their
results more effectively.

© 2025 The Authors. Published by Elsevier Inc. on behalf of American Society for Pharmacology and
Experimental Therapeutics. This is an open access article under the CC BY license (http://
[Link]/licenses/by/4.0/).

I. Introduction of mean that everyone is familiar with. We will get to the geo-
metric mean in the next section.
A. A motivating example  The P value (two-tailed) testing the null hypothesis that the sets
of data were sampled from identical normal distributions is
This motivating example demonstrates that analyzing 0.22. This is greater than the traditional threshold of 0.05, so the
lognormal data as if the values were sampled from a normal dis- null hypothesis of no difference would not be rejected.
tribution can lead to incorrect and misleading conclusions. Figure 1  Looking at the graph, the largest value in each group is much
compares EC50 values for control and treated conditions. larger than the rest. Indeed, Grubbs’ (1969) outlier test with a
set to 0.05 identified an outlier in each case.
1. Incorrect analysis assuming sampling from normal distributions 2. Correct analysis assuming sampling from lognormal distributions
Here are the results if the data were analyzed conventionally Now let us analyze correctly, assuming sampling from
with a 2-sample unpaired t test assuming sampling from normal lognormal distributions, using methods that will be explained in
distributions: detail below.

 The means are 294 nM (control) and 575 nM (treated). The  First, a quick reminder about samples and distributions.
difference is 282 nM (95% confidence interval [CI] of the Commonly used statistical tests (such as t tests) proceed by
difference: 175 to 739 nM). With such a wide CI), the data are assuming the null hypothesis, that is, that there is no real dif-
consistent with no difference, a moderate decrease, or a large ference between conditions (here, control vs treated). These
increase. In other words, no conclusion is possible. Note, here we tests then ask how frequently such an extreme (or even more
refer to the arithmetic mean (AMean)din other words, the type extreme) result would be obtained by taking random samples
3
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

4000 B. History of lognormal distributions

Lognormal distributions, first described in 1879 (Galton, 1879;


McAlister, 1879) are asymmetrical distributions commonly
encountered in many fields of science. Despite their long history
3000 and widespread occurrence, they are often overlooked or misun-
derstood. Aitchison and Brown (1957) designated the lognormal
EC50 (nM)

distribution as the “Cinderella of distributions,” shunned


compared with its normal “sister,” and many scientists still
mistakenly consider lognormality to be an obscure topic that can
2000 generally be ignored. Many biostatistics texts do not even mention
lognormal distributions or geometric means (Zar, 2009; Glantz,
2011; Dancey et al, 2012; Irizarry and Love, 2016; Baldi and
Moore, 2017; Daniels and Cross, 2018; Glaser, 2018). We only
1000 know of a few biostatistics texts that devote more than a few
pages to this topic (Bland, 2015; Motulsky, 2017; Wahi and
Puzzullo, 2024).

C. Our goals in writing this review


0
Control Treated Prior reviews have urged scientists to realize how common
lognormal distributions are, and how much is gained by analyzing
lognormal data properly. These reviews, however, do not provide a
Fig. 1. Sample data to demonstrate why it is important to identify lognormal distri-
butions. These values were obtained by first randomly choosing values from lognormal
useful guide for working scientists to analyze lognormal data
distributions with GeoSD ¼ 4 and GeoMeans ¼ 100 and 400, and then rounding each because they are either too concise (Gaddum, 1945; Limpert et al,
value to the nearest integer (to make it easy for readers to analyze with their own 2001; Limpert and Stahel, 2011, 2017; Curran-Everett, 2018) or
software). The raw data are in Supplemental Material. too mathematical (Crow and Shimizu, 1988; Parkin and Robinson,
1992; Vogel, 2022; [Link]
distribution).
from an invisible population (distribution) comprising all anal-
In this review we aim to clarify the concept of lognormal dis-
ogous experiments that you could have performed if the null
tributions, explain their relevance in pharmacology, and provide
hypothesis were true.
guidance on how to work with these distributions. We limit our
 Geometric means (GeoMeans), explained below, are estimates
discussion to describing lognormal distributions and comparing
of the medians of the distributions that the data were sampled
data sets sampled from such distributions. We do not discuss the
from. The GeoMeans are 103 nM (control) and 302 nM
use of lognormal distributions in regression or in Bayesian ana-
(treated). The ratio is 2.9 with a 95% CI of the ratio ranging from
lyses. Neither do we discuss 3-parameter (ie, “shifted”) lognormal
1.24 to 6.93.
distributions, which can occur when there is a defined minimum
 The P value (two-tailed) testing the null hypothesis that both
value other than zero; although such variables are encountered
populations are identical and lognormal is 0.015. This is less
in medicine (Royston, 1992) they are rarely relevant to
than the traditional threshold of 0.05 so the null hypothesis
pharmacology.
would be rejected.
We point out that failing to recognize a lognormal distribution
 The section The lognormal Welch’s t test will explain an alter-
can lead to erroneous conclusions and misinterpretation of data.
native (or extension).
We also show that normality and lognormality tests are nowhere
 After log-transforming the raw data, Grubbs’ outlier test (a ¼
near as useful as many expect, and so we recommend that
0.05) did not identify an outlier in either group.
lognormal distributions be assumed (without testing) for many
of the variables measured by pharmacologists. Some readers
3. Why it can matter may prefer to skip ahead to our list of recommendations in the
For this example, analyzing the data correctly solved 3 Summary.
problems:
II. Ratio scale variables
 Reporting the experimental effect as a ratio is scientifically
sensible. The treatment approximately triples the EC50, and All lognormal variables are ratio scale variables, so let us start
the CI tells us the data are compatible with a ratio between a there.
1.2 and 6.9.
 With the correct analysis, the two-tailed P value is .015. Because A. Definition of ratio scale variables
this is less than the .05 threshold many scientists routinely use,
the null hypothesis of no treatment effect would be rejected. Many variables measured and compared in pharmacology are
Assuming lognormal distributions reversed the conclusion of ratio scale variables with the following properties (Stevens, 1946):
this experiment.
 Grubbs’ test detected no outliers when done properly on log-  Negative values are inconceivable. Only positive values are
transformed values. possible.
 Zero either means none of that variable or the asymptotic value
We will return to this example in the section Why the unpaired t the variable can approach. Examples of the former are weight or
test (without log transformation) should be avoided with lognormal length. A weight of zero means no weight. A height of zero
data. means no height. Examples of the latter are an EC50 or Km value.
4
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

150 250

Number of pairs of dice


Number of pairs of dice
200
100
150

100
50
50

0 0
0 5 10 15 20 25 0 200 400 600
Sum of 4 dice Product of 4 dice

Fig. 2. Multiplicative factors lead to an asymmetrical distribution. The graphs show simulations of throwing 4 dice 1000 times. The graph on the left is a frequency distribution of
the sum of the values appearing on 4 dice. The graph on the right is a frequency distribution of the product of the 4 values. Our simulations were inspired by a blog by M.H. Nederlof
([Link]

It is not possible for either of those parameters to equal zero, but Consider the following example: a drug increases your measure
their values can be tiny and approach zero. from 5 to 10 U in first subject, from 10 to 20 U in a second subject,
 Converting between units requires only multiplication or divi- and from 7 to 14 U in a third individual. These 3 results would be
sion, as is the case for weights, concentrations, durations, seen as identicaldthe drug has a doubling effect. In other words,
lengths, and EC50s. the underlying mechanism affecting changes in the variable is
 It makes scientific sense to calculate a ratio. For example, 4 cm is multiplicative. The relative change is the same, regardless of the
twice as long as a distance of 2 cm; 6 L of water is 3 times as starting value. The unequal absolute differences (an increase of 5 U
much water as 2 L; and an EC50 of 5 mM is 5 times greater than vs 10 U vs 7 U) are most likely irrelevant.
an EC50 of 1 mM.

III. Lognormal distributions


B. Examples of ratio variables and a counterexample
Variables on a ratio scale can follow various distributions,
including a Poisson distribution for counted variables, and an
 Enzyme velocities can only be positive, with zero indicating no
exponential or g distribution for waiting times. However, many
enzyme activity. If an enzyme velocity is reduced from 100 to 50
ratio scale variables are distributed as a lognormal distribution, as
units (U)/min, it is best quantified as a ratio. The velocity was cut
elaborated below.
in half. If the same inhibitor in a different system reduces ve-
locity from 50 to 25 U/min, the effect would be seen as the same:
the inhibitor cuts the enzyme velocity in half. The fact that the A. Multiplicative causes of variation lead to asymmetrical
differences are distinct (a change of 50 U/min vs a change of 25 distributions
U/min) would not usually be relevant.
 Equilibrium dissociation constants (Kd) can only have positive Lognormal distributions can arise when the effects of different
values. They cannot equal zero but can have any positive value. If factors on the variable are multiplicative. Figure 2 simulates
changing one amino acid in a receptor molecule alters the Kd of throwing 4 dice, computing their sum and product, repeating 1000
agonist binding from 5 to 1 mM, the finding is best summarized times, and tabulating the distribution of sums (left graph; sym-
as a ratiodthe Kd was reduced by a factor of 5 or (equivalently) metrical) and products (right graph; asymmetrical). This demon-
the affinity went up by a factor of 5. The difference (4 mM) is not strates that multiplicative variation leads to an asymmetrical
interesting and would not be reported. distribution.
 Temperature in either Celsius or Fahrenheit is a counterexample To understand why multiplicative factors lead to skewed
as it does not meet any of the 4 criteria defining a ratio variable. distributions, consider an example where a variable with an
Negative values are meaningful. A temperature of 0  C or 0  F average value of 10 is randomly doubled or halved. Doubling the
does not mean “no temperature.” Converting between the 2 value results in 20 (an increase of 10), whereas halving it results
temperature scales requires more than just multiplication or in 5 (a decrease of 5). The magnitude of the increase is greater
division. Ratios are not meaningful (200  C is not twice as hot as than that of the decrease, despite the symmetry of the multi-
100  C) so changes in temperature in  C or  F must be expressed plicative factor. This asymmetry in the magnitudes of change
as differences. For all 4 reasons, temperature in  C or  F is not a leads to a skewed distribution when compounded over many
ratio variable. In contrast, temperature in kelvin is a ratio vari- instances.
able because it has a true zero (absolute zero) and ratios are Another perspective is to consider the ratios between the
meaningful. values. A ratio of 1.0 represents no change from the average.
Doubling the average yields a ratio of 2.0, whereas halving it re-
sults in a ratio of 0.5. These ratios are not equidistant from 1.0 on a
C. For ratio variables, experimental effects are best reported as linear scaled2.0 is farther from 1.0 than is 0.5. This asymmetry in
ratios, not differences the ratios explains the skewed distribution caused by multipli-
cative errors.
When comparing the effects of different factors on a ratio var- Lognormal distributions can arise for reasons other than mul-
iable, it is usually more meaningful to focus on (or report) the tiplicative error. Koch (1966) reviews several mechanisms that lead
relative change (ratio) rather than the absolute change (difference). to lognormal distributions, and separately (Koch, 1969) reviews
5
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

mechanisms that create distributions that are not lognormal but The 2 differences on the logarithmic scale are equidistant
look very similar. from log(1.0) ¼ 0.0, demonstrating that the logarithmic trans-
formation eliminates the asymmetry on the original scale, mak-
ing the effect of the multiplicative factor symmetrical on the
B. Review of logarithms
logarithmic scale.

Logarithms have the unique ability to transform multiplicative


D. Relationships between normal and lognormal distributions
relationships into additive ones, because the logarithm of a product
equals the sum of the logarithms of its factors:
The relationships between normal and lognormal distributions
logða  bÞ ¼ logðaÞ þ logðbÞ: are simple.

The logarithm (base 10) of 1000 is the power of 10 that equals  If an infinite population of values define a lognormal distri-
1000. The logarithm of 1000 is 3, because 103 ¼ 1000. The loga- bution, then the logarithms of those values define a normal
rithm of 10 is 1, because 101 ¼ 10. The logarithm of 0.001 is 3, distribution.
because 103 ¼ 1/103 ¼ 0.001. The logarithm of 3.162 is 0.5,  If an infinite population of values define a normal distribution,
pffiffiffiffiffiffi
because 10 0.5 ¼ 10 ¼ 3.162. The logarithm of 1.0 is 0.0, because then the antilogarithms of those values define a lognormal dis-
10 0 ¼ 1.0. tribution. The antilogarithm is the inverse of the logarithm. The
The above logarithms are base 10 logarithms, also called antilogarithm of a common logarithm equals 10 to that power.
common logarithms, because the computations take 10 to some For example, the antilogarithm of 3 is 103 or 1000. The antilog-
power. They are sometimes written as “log10(x).” Mathemati- arithm of a natural logarithm equals e to that power. For example,
cians prefer natural logarithms using base e (2.718. . .), written as the antilogarithm of 6.908 is e6.908, often written as exp(6.908),
ln(x). Beware of the notation “log(x),” which can mean either which is 1000.
common or natural logarithm, depending on the field or
program. The term “lognormal,” often written “log-normal,” is potentially
The logarithms of values > 1.0 are positive. The logarithms of confusing as it can be incorrectly thought of as the “log of normal.”
values >0.0 and <1.0 are negative. The logarithms of zero and all But it is a mistake to think that values in a lognormal distribution
negative numbers are simply undefined, because there is no power are the logarithms of values from a normal distribution. We show it
of 10 that results in a negative number or zero. struck out because it is simply wrong. The term “antilognormal”
describes the distribution better than “lognormal” (Johnson et al,
C. Logarithms convert a skewed distribution caused by 1994), but that term is rarely (if ever) used.
multiplicative error to a symmetrical distribution The distribution of lognormal values is asymmetrical with a
positive skew, that is, with a long tail to the right (but as we will see,
Revisiting the earlier example where a factor randomly this asymmetry can be subtle in some cases). However, the distri-
doubles or halves a value of 10, resulting in 20 or 5, the loga- bution of their logarithms is symmetrical and forms a normal (ie,
rithms (base 10) of these values and the differences between Gaussian) distribution, as shown in Fig. 3.
them are:
E. Lognormal distributions are common in biology and beyond
logð20Þz1:30
There seems to be a prevailing sense that experimental data
logð10Þ ¼ 1
usually follow a normal distribution. But it has been realized for a
logð5Þz0:70 century that this is not true. A 75-year-old text states, “The normal
curve was, in fact, to the early statisticians what the circle was to
logð20Þ  logð10Þz0:30 the Ptolemaic astronomers” (Yule and Kendall, 1950). We now
know that planets move in ellipses rather than circles, and that
logð10Þ  logð5Þz0:30 lognormal distributions are pervasive.
Frequency

0 10 20 30 40 50 0.0 0.5 1.0 1.5 2.0


EC50, M log(EC50, M)

Fig. 3. Frequency distribution of a lognormal distribution. Left: A frequency distribution of a lognormal distribution with GeoMean ¼ 10 and GeoSD ¼ 2, defined later in this article.
Right: Frequency distribution of the common (base 10) logarithms of the values. Note that the log transformation converts the asymmetrical lognormal distribution to a sym-
metrical normal distribution.

6
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

Lognormal distributions have been noted in data sets ranging the log-transformed data was closer to zero than the skewness of
from journal citation counts (Thelwall, 2016) to professorial sal- the raw data).
aries (Benzidia and Lubrano, 2020), and from pollution levels in We did similar simulations in more depth assessing the dis-
Los Angeles to the number of words in telephone conversations tribution of EC50, Kd, Koff, Kon, and Hill coefficient (Supplemental
(Limpert et al, 2001). Examples from biology and medicine Material). A lognormal distribution fitted all these parameter
include: neuronal firing rates in the central nervous system distributions better (larger R2) than a normal distribution did.
(Buzsa ki and Mizuseki, 2014); numerous blood analytes including The difference between the corrected Akaike information criteria
triglycerides (Carlson, 1960), free fatty acids (Heath, 1967), assesses how much better the data supports one model versus
ferritin (Custer et al, 1995), alkaline phosphatase, creatinine, another. If the difference is > 10, the worse-fitting model has
glucose, and iron (Flynn et al, 1974); abdominal fat (Stanforth essentially no support from the data (Burnham and Anderson,
et al, 2004), senile plaque size in Alzheimer disease (Hyman 2002; Portet, 2020). For our simulations, the smallest differ-
et al, 1995), blood pressure (Bodey and Michell, 1996), the num- ence in corrected Akaike information criteria was 30, demon-
ber of parasites per host (Shaw and Dobson, 1995), and cancer strating the simulated data for all the parameters support a
survival times (Qazi et al, 2007). For yet more examples, see lognormal distribution substantially better than a normal
Limpert et al (2001). distribution.

G. Variables defined as the ratio of 2 lognormal distributions are


F. EC50, IC50, Kd, Km (and more) tend to be lognormal
lognormal
1. Data demonstrating pharmacological parameters are lognormal
1. An interesting and impactful property of lognormal distributions
Gaddum (1945) was probably the first to emphasize the
What is the distribution of a variable defined as the sum, dif-
importance of lognormal distributions in pharmacology. Since
ference, product, or ratio of 2 independent variables? It depends on
then, several pharmacological parameters have been shown to
whether those variables are normal or lognormal.
be distributed in a lognormal fashion. Table 1 lists pharmaco-
dynamic measures (notably EC50, IC50, Kd, Ki, and Hill coeffi-
 The difference or sum of independent normal variables is
cient). Table 2 lists pharmacokinetic measures such as area
normal.
under the curve, clearance, volume of distribution, elimination
 The ratio or product of normal variables is neither normal nor
half-life, plasma, and intracellular concentrations.
lognormal.
Although some papers (and our simulations; see below) report
 The difference or sum of lognormal variables is neither normal
that Hill Slopes are lognormal, Levasseur et al (1998) reported that
nor lognormal.
Hill slopes follow a normal distribution. We digitized their pub-
 The ratio or product of independent lognormal variables is
lished frequency distribution and fit it using Poisson weighting
lognormal (Limpert et al, 2001). As the next section will show,
(because all the values are counts). We found that a lognormal
this relationship has particular importance in pharmacology,
distribution fits somewhat better (pseudo R2 ¼ 0.84) than a normal
where many key parameters arise as ratios of lognormal
distribution (pseudo R2 ¼ 0.76).
variables.

2. Simulations demonstrating pharmacological parameters are 2. Examples in pharmacology where important parameters are the
lognormal ratio of 2 lognormal variables
De Lean et al (1982) demonstrated that Kd values from simu- This relationship has ramifications in pharmacology because
lated repeated experiments are closer to lognormal than normal. many parameters are defined as the ratio of lognormal parameters.
Christopoulos (1998) used simulations to demonstrate that the Here are examples:
distributions of these parameters are lognormal: EC50, Hill coeffi-
cient, agonist efficacy t in the Black and Leff operational model of  Bioequivalence. The comparison of drug formulations relies on
agonism (Black et al, 2010), the receptor/G-protein dissociation the ratio of areas under the curve between test and reference
equilibrium constant KG, and the cooperativity factor a. He found formulations. Area under the curve values tend to be lognor-
that the dissociation rate constant (Koff) was fit reasonably well by mally distributed because of the multiplicative nature of drug
both normal and lognormal distributions and concluded that Koff is absorption and elimination processes ([Link]
normal. But his data demonstrates that the lognormal distribution media/70958/download). Consequently, their ratio is
fits those values better than does the normal distribution (the sum- lognormal, which dictates the statistical approaches used in
of-squares was smaller for the lognormal fit, and the skewness of bioequivalence studies (Julious, 2004).

Table 1
Pharmacodynamic parameters reported to follow lognormal distributions

Parameter Parameter/Topic Reference

Tissue responses at fixed doses Tissue responses to fixed doses of norepinephrine or acetylcholine Fleming et al, 1972
EC50 Iris contraction induced by carbachol Patil, 1993
EC50 Cardiac muscle contraction induced by adrenaline and noradrenaline Kaumann et al, 1989
Ki and EC50 Ki and EC50 for atrial natriuretic peptide and norepinephrine Hancock et al, 1988
IC50 Cytotoxic effects of drugs on cancer cells in vitro Levasseur et al, 1998
Hill coefficient Electrophysiological response of rat olfactory receptor neurons to odorants Rospars et al, 2008
Hill coefficient Forceecalcium relationship in myocytes Walker et al, 2010
Kon, Koff, Kd Antibody-antigen binding Poulsen et al, 2011
Kon Microtubule binding to destabilizing kinesin-8 (Kif18B) Shrestha et al, 2023
Exponential time constants Single-molecule transport into liposomes, by a neurotransmitter: sodium symporter Fitzgerald et al, 2006

7
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

Table 2
Pharmacokinetic parameters reported to follow lognormal distributions

Parameter Parameter/Topic Reference

Clearance, Vd, t1/2 Clearance, volume of distribution, and elimination half-life of theophylline after IV Steinijans et al, 1982; Julious and
infusion Debarnot, 2000
Cmax Either lognormal (Lacey et al) or not (Shen et al) Lacey et al, 1997; Shen et al, 2017
AUC Area under the curve (AUC) is close to lognormal Shen et al, 2017
Whole-body uptake Post-Chernobyl nuclear accident, whole-body counts of Cs-137 in women Wertelecki et al, 2016
Plasma levels Strontium plasma concentrations after oral administration to human volunteers Li et al, 2006
Drug concentration Intracellular drug concentration in tumors Zanotti-Fregonara and Hindie , 2011

 Equilibrium dissociation constant (Kd). The Kd represents the arithmetic). As mentioned earlier, if a distribution is lognormal,
ratio of unbinding to binding rate constants (Koff/Kon). Both then the logarithm of that distribution is normal. Statisticians
rate constants arise from multiple molecular steps in protein- (and some biomedical scientists) sometimes report the mean
ligand interactions that combine multiplicatively, leading and SD of the log-transformed data. When interpreting such
to lognormal distributions. Their ratio, Kd, is therefore data, pay attention to whether common or natural logarithms
lognormal, impacting how we analyze drug-receptor binding were used.
data. Rather than reporting the summary statistics of the set of log-
 The operational model. The transducer ratio t is defined as the arithms, we prefer to use the GeoMean and the geometric SD
total receptor density (termed R0 or Rt) divided by the (GeoSD).
coupling efficiency constant KE (Black and Leff, 1983; Kenakin
et al, 2012). Both parameters become lognormally distributed A. The geometric mean
through multiplicative cellular processesdRT through re-
ceptor expression and trafficking, KE through sequential steps 1. The GeoMean of an ideal lognormal distribution or population
in signal transduction. Consequently, their ratio t is The geometric mean, which we abbreviate GeoMean, is another
lognormal. name for the median of an ideal lognormal distribution. It is
 Biased agonism. Ligand bias is quantified as the transduction expressed in the same units as the data. Figure 4 shows that the
coefficients (t/KA) for different signaling pathways. Because t/KA GeoMean, mode and AMean have different values for an ideal
represents a ratio of lognormally distributed parameters, it is lognormal distribution (left), but all are identical for an ideal
itself lognormal. When comparing 2 pathways (eg, G-protein normal distribution (right).
versus b-arrestin signaling), the standard approach calculates The concept of an AMean has been drilled into us since primary
the difference between log-transformed values for a test ligand school, and it may be hard to imagine how it is possible to have
[Dlog(t/KA)] and normalizes this to a reference ligand, yielding another kind of mean. Figure 5 shows one way to understand how
DDlog(t/KA) (Kenakin et al, 2012). Because the log-transformed there can be 2 distinct means.
transduction coefficients for individual pathways should be
normally distributed, their differences (Dlog and DDlog values)
2. How to calculate the GeoMean of a data set
will also be normal assuming that the pathways respond
To calculate a GeoMean, first calculate the common (ie, base 10)
independentlyda reasonable assumption given our under-
logarithms of all the values, compute their mean (call it m), and
standing of distinct signaling mechanisms. If the signaling
then calculate GeoMean ¼ 10m.
pathways are not independent, then a modified operational
Here, we use base 10 logarithms and the corresponding 10^
model has been proposed (Zhu et al, 2019).
antilog function, as this is what most biologists are familiar with.
 Allosteric interactions. The cooperativity factor a is defined as a
Many statisticians, engineers, and physical scientists prefer the
ratio of binding constantsdspecifically, the ratio of a ligand’s
natural log (ln) and the corresponding exp() antilog function. The
binding affinity (Ki) when the allosteric site is empty to its
GeoMean (and the GeoSD) will be the same with either approach as
affinity when that site is occupied by a modulator (Ki(free)/
long as the same base is used for taking the logarithms and
Ki(co-bound)). Because both Ki values arise from binding
reversing that transform (antilog).
processes and are lognormally distributed, their ratio a is also
You will sometimes see the GeoMean defined as the n-th root of
lognormal (Christopoulos, 1998; Christopoulos and Kenakin,
the product of all values. As long as all values are positive, these 2
2002).
definitions are equivalent.
3. Why is the ratio of 2 lognormal variables lognormal?
Consider 2 lognormal variables X and Y. To understand the dis- 3. Relationship between the GeoMean and the median
tribution of their ratio Z ¼ X/Y, we can use a powerful mathematical As mentioned, for an ideal lognormal distribution, the GeoMean
tool: taking logarithms transforms multiplicative relationships into and median are identical.
additive ones. Take the logarithm of both sides: log(Z) ¼ log(X)  For any particular data set randomly sampled from a
log(Y). Because X and Y are lognormal, both log(X) and log(Y) are lognormal distribution, the obtained GeoMean is equally likely to
normal by definition. When we subtract 2 independent normal be larger or smaller than the median. On average, the GeoMean is
variables, the result is also normal. The final step follows directly a more accurate estimate of the population median than is the
from the definition of a lognormal variable: because log(Z) is median computed from the same sample (Parkin, 1993; Vogel,
normal, Z itself must be lognormal. 2022). That is why the GeoMean is the standard way to quan-
tify the center of a lognormal distribution. However, the Geo-
IV. Descriptive statistics of lognormal distributions Mean can be a poor estimate of the median if data are sampled
from a skewed distribution that is not lognormal, or when data
A normal distribution is defined by 2 parameters: the mean are sampled from a (mostly) lognormal distribution with outliers
(ie, arithmetic mean or AMean) and the SD (ie, again, (Vogel, 2022).
8
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

Fig. 4. The arithmetic mean, GeoMean, and median of an ideal lognormal distribution. Left: An ideal lognormal distribution showing distinct mode, arithmetic mean, and median
(GeoMean). The GeoMean is the median; half the values are larger, and half are smaller. The arithmetic mean is the center of gravity of the distribution. If you made the distribution
out of wood or plastic and included the tail that goes far beyond the right limit of the graph, it would balance at the arithmetic mean. Right: An ideal normal distribution showing
that the mode, arithmetic mean, and median are identical.

B. The geometric standard deviation The terms Geometric SD factor and multiplicative SD are
sometimes used as synonyms of GeoSD. Beware of the term s*,
1. What is the GeoSD? as it sometimes serves as an abbreviation for the SD of the nat-
The GeoSD (Kirkwood, 1979) quantifies both the spread and ural logarithms, and sometimes for the GeoSD (Limpert et al,
asymmetry of a lognormal distribution, as shown in Fig. 6. Unlike 2001).
the arithmetic (regular) SD, which has the same units as the data,
the GeoSD has no units. Multiplying the raw data by 1000, say, 3. A lognormal distribution with a small GeoSD is nearly identical
would make the arithmetic SD commensurately 1000 times larger, to a normal distribution
but would not change the GeoSD (Lewontin, 1966). The GeoSD al- A lognormal distribution with a small GeoSD closely resembles a
ways has a value  1.0. The GeoSD ¼ 1.0 only when all values are normal distribution (Fig. 7). How small does the GeoSD need to be?
identical. Any cutoff is somewhat arbitrary, but values of 1.2 (Limpert et al,
2001) and 1.3 (Elassaiss-Schaap and Duisters, 2020) have been
2. Calculating the GeoSD suggested.
The GeoSD is calculated as GeoSD ¼ 10s where s is the SD of the For variables that are always positive but are approximately
common logarithms of the values in the sample. Equivalently, normal, variation can be quantified as the coefficient of
GeoSD ¼ es , where s is the SD of the natural logarithms of the variation (CV) which equals SD/AMean. A lognormal distribution
values. Note that your choice to use natural or common logarithms with a small GeoSD looks very similar to a normal distribution
will alter the value of s but not GeoSD. with a CV ¼ GeoSD -1 (Haeckel and Wosniok, 2010). This
relationship makes sense because the minimum possible value
of the CV is 0.0, and the minimum possible value of the GeoSD is
1.0.
Figure 8 shows how similar normal and lognormal data can
appear. Each panel shows 10 simulated data sets (n ¼ 10 each). One
5 was sampled from a lognormal distribution with GeoSD ¼ 1.2 and
GeoMean ¼ 100, and the other was sampled from a normal dis-
tribution with a SD ¼ 20 and mean ¼ 100 (so the CV ¼ 0.2). It is
2 8 impossible to guess which is which (in fact, the left panel is
lognormal; the right panel is normal).
Figure 9 is identical, but with n ¼ 100 per group. You still cannot
tell which is normal and which is lognormal by inspection.
An example of a variable that fits both normal and lognormal
distribution is height. Slavskii et al (2021) reviewed the distri-
bution of height in multiple populations. Its CV is between 0.03
and 0.05, small enough so that normal and lognormal distribu-
tions both fit very well (but the lognormal distribution fits
4 slightly better, a finding that is only apparent with huge data
sets).

0 2 8 4. How to write the GeoMean and GeoSD


If you assume sampling from a normal distribution, the mean
and SD are commonly written as 293.5 ± 515.4 (control data from
Fig. 5. The concept of different kinds of means. What is the mean of 2 and 8? The Fig. 1) and read as “293.5 plus or minus 515.4.” The plus-or-minus
upper panel shows that the arithmetic mean is 5 because the same delta takes you
from that geometric mean to each value (add 3, or subtract 3). The lower panel shows
symbol is standard. If you assume sampling from a lognormal
that the geometric mean is 4 because the same multiplicative factor takes you from distribution, the GeoMean and GeoSD would be stated as “103.0
that geometric mean to each value (multiply by 2, or divide by 2). multiplied or divided by 4.53.” This can be written in several ways:
9
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

GeoSD = 1.25
GeoSD = 1.5
GeoSD = 2.0

0 1 2 3 4 0 1 2 3 4 0 1 2 3 4

GeoSD = 6 GeoSD = 10
GeoSD = 4

0 1 2 3 4 0 1 2 3 4 0 1 2 3 4

Fig. 6. The GeoSD quantifies asymmetry. All the graphs share the same area under the curves (1.0) and same GeoMean (1.0). Lognormal distributions with larger GeoSDs are more
skewed.

 Enter the multiply  superscripted (in the Insert Symbol dialog manuscript, it looks like an asterisk so does not communicate
of Word) followed by a slash: 103.0 / 4.53. This notation was the concept of multiply-or-divide very clearly. At the larger sizes
proposed by Limpert and Stahel (2011). used for presentations, the symbol is more understandable.
 Enter a superscripted (and bold) period followed by a slash:  If you want to avoid the phrase “multiplied or divided by,”
103.0 ./ 4.53. present a table with separate columns for GeoMean and GeoSD.
 Insert a multiply sign and a superscripted ±1, for example
103.0  4.53±1. With the plus sign, this becomes 103.0  4.53.
With the minus sign, it becomes 103.0  4.531 ¼ 103.0 ÷ 4.53. C. Other ways to describe lognormal distributions
 Use the Unicode “DIVISION TIMES” symbol (Uþ22C7) that su-
perimposes the multiply  and divide ÷ symbols: 103.0 ⋇ 4.53. 1. The range that contains 68% or 95% of the values
To enter this character into Microsoft Word, enter “22C7” With a normal distribution, the range of values extending from
without the quotation marks, hold the Alt (Windows) or Option the [mean  SD] to [mean þ SD] includes about two-thirds of the
(Macs) key, and tap X. Or use the Insert Symbol dialog. GraphPad population, more exactly 68.3%. With a lognormal distribution, the
Prism (starting with version 9) includes this symbol on the Math comparable range extends from [GeoMean/GeoSD] to [GeoMean 
tab of its Insert Symbol dialog. At the font sizes used in a GeoSD] (Fig. 10).
Probability Density

Probability Density

0 50 100 150 200 0 50 100 150 200


Value Value
Fig. 7. When the CV is less than ~0.2, normal and lognormal distributions are similar. Left: A normal distribution with mean ¼ 100 and SD ¼ 20 (red), and a lognormal distribution
with GeoMean ¼ 100 and GeoSD ¼ 1.2 (blue). Note that the SD is 20% of the mean so the CV, is 20%. Adding 20% is the same as multiplying by 1.20, so the CV of 20% is roughly
equivalent to a GeoSD of 1.2. Right: A normal distribution with CV ¼ 0.1 and a lognormal distribution with GeoSD ¼ 1.1.

10
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

150 150

100 100

50 50

0 0

Fig. 8. Samples from lognormal vs normal populations. Each panel shows 10 simulated data sets (n ¼ 10 each). One was sampled from a lognormal distribution (GeoSD ¼ 1.2;
GeoMean ¼ 100), and the other from a normal distribution (SD ¼ 20; mean ¼ 100). You cannot tell which is which by looking at the graphs. (The graph on the left shows lognormal
data; the graph on the right shows normal data.)

Similarly, about 95% of values of a normal distribution are within sampling from a normal distribution would rarely create a data
the range [mean  2 SD] to [mean þ 2 SD]. The comparable range distribution that far (or further) from normal.
with a lognormal distribution is [GeoMean / GeoSD2] to
[GeoMean  GeoSD2]. Note that the GeoSD is squared, not doubled. 2. How lognormality tests work
To test for lognormality, first transform all the data to their
2. Confidence interval of a GeoMean logarithms (it does not matter if you use common or natural log-
To calculate a CI of a GeoMeandperhaps better called a arithms). If the data were sampled from a lognormal distribution,
compatibility interval (Rafi and Greenland, 2020)dtransform the that set of logarithms would have been sampled from a normal
values to logarithms, compute the CI of the mean for the degree of distribution. Test this with one or more normality tests. If the P
confidence you want (usually 95%), and then reverse-transform the value is small, you will conclude that the distribution of the loga-
lower and upper confidence limits using the antilogarithm trans- rithms would be unlikely if the underlying distribution is normal,
form. The resulting interval will not be symmetrical around the so the distribution of the original values would be unlikely if the
GeoMean, but the asymmetry might be subtle. underlying distribution is lognormal.

3. Normality and lognormality tests are often inconclusive


V. How to decide if a variable is lognormal Let us revisit the data sets of Fig. 8 where normal and
lognormal distributions appeared nearly identical. In the
Deciding whether to analyze data assuming sampling from left panel of lognormal data (GeoSD ¼ 1.2), 9 of the 10 data sets
lognormal distributions is not always a straightforward process. pass 3 different normality tests (AndersoneDarling, D’Ag-
ostinoePearson, and ShapiroeWilk) with P > .05, and also pass 3
A. Do not rely on tests of normality and lognormality different lognormality tests (the same tests run on the loga-
rithms of the values) with P > .05. The right panel of Fig. 8 shows
1. Review of normality tests normal data. All 10 data sets pass the 3 normality tests, and 9 pass
Normality tests use various criteria to test the null hypothesis the 3 lognormality tests.
that a data set was sampled from a normal distribution. Three What happens with a larger GeoSD? Figure 11 shows 2
commonly used normality tests are those named for D’Ag- lognormal distributions with different GeoMeans (100 and 322)
ostinoePearson (D’Agostino et al, 1990), ShapiroeWilk (Shapiro and identical GeoSD (3). With a GeoSD this large, the asym-
and Wilk, 1965), and AndersoneDarling (Anderson and Darling, metry of the distribution is not subtle. The 2 distributions are
1954). They use different methods to quantify the discrepancy quite distinct, but overlap considerably. Note that the samples
between the actual data distribution and an ideal normal distri- shown in Figs. 12 and 13 were randomly drawn from these
bution, so can produce different results. A small P value means that distributions.

200 200

150 150

100 100

50 50

0 0

Fig. 9. Even with n ¼ 100, a normal distribution with CV ¼ 20% is indistinguishable from a lognormal distribution with GeoSD ¼ 2.0. This matches Fig. 8, but with n ¼ 100 per data
set.

11
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

0 GeoMean 2 4 6
Fig. 10. The middle 68.3% of a lognormal distribution. This lognormal distribution has a GeoMean ¼ 1 and a GeoSD ¼ 4. The shaded area extends from the GeoMean divided by the
GeoSD to the GeoMean multiplied by the GeoSD. If it looks to you like the shaded area contains > 68% of the total area, that is because the tail of the curve extends far to the right
beyond the limits of this graph and that tail has considerable area. In this example, the shaded range does not include the mode, the X-value at the peak of the curve. However, when
the GeoSD is much smaller, that range will include the mode.

GeoMean = 100 322

GeoSD = 3
Probability Density

0 500 1000
Value
Fig. 11. Distributions used to sample data in Figs. 12 and 13. Both lognormal distributions have GeoSD ¼ 3. The GeoMeans differ by a factor of 3.2.

Figure 12 shows 4 simulated experiments of data from these Normality tests cannot reliably determine whether these dis-
distributions with n ¼ 10 per group. It is not always obvious by tributions are normal or lognormal. We ran 3 different normality
inspection that the data are not normal. tests (D’AgostinoePearson, ShapiroeWilk, and AndersoneDarling)

1000 1000 1000 1000

500 500 500 500

0 0 0 0
Control Treated Control Treated Control Treated Control Treated

Fig. 12. Random samples (n ¼ 10) from lognormal distributions shown in Fig. 11.

12
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

1000 1000 1000 1000

500 500 500 500

0 0 0 0
Control Treated Control Treated Control Treated Control Treated

Fig. 13. Same as Fig. 12, but with n ¼ 20 per sample.

on 2500 simulated control data sets (n ¼ 10) similar to those in  As mentioned earlier, only a ratio scale variable can be
Fig. 12. The simulated data sets failed the normality tests a bit more lognormal. When considering whether a variable can be
than half of the time (53%, 64%, and 66% for the 3 control sets; and lognormal, review this checklist of properties that define ratio
52%, 63%, and 65% for the treated data sets). In other words, nearly scale variables (discussed in the section Ratio scale variables):
half of the simulated lognormal data sets passed normality tests. ✓ Negative values must be inconceivable. Lognormal variables
Because the data are sampled from lognormal distributions, you can only be positive values.
would expect 5% of the simulated data sets to fail the lognormality ✓ Zero must either mean none of that variable or be the
tests with a set to 0.05, and indeed between 4.5% and 5.6% of the asymptotic value the variable approaches. However, no
simulated data sets failed each of the 3 tests. values can actually equal 0.0.
Normality tests can better detect the lack of normality of ✓ Converting between units must require only multiplication or
lognormal data when the sample size is larger. Figure 13 doubles division.
the sample size to 20 per group. Now, between 84% and 96% of the ✓ It must make scientific sense to calculate a ratio of 2 values.
simulated data sets (2 treatments; 3 normality tests; 10,000 sim-
ulations) fail the normality tests with P < .05. If a variable fails to meet any of these criteria, it is not a ratio
variable and so cannot be lognormal.
B. How to approach questions about lognormality
2. Consider the possibility that your data may follow a distribution
We have shown that normality tests can lead to inconsistent and that resembles lognormal
even incorrect conclusions, especially with small data sets. So how When working with positively skewed data, beware of other
should scientists decide when to assume lognormality? Here is the distributions that can resemble the lognormal distribution. Two
approach we recommend: such distributions are the g distribution and the distribution of the
ratio of 2 normally distributed variables.

1. Ask the right questions


a. The g distribution resembles the lognormal distribution. The g
When asking whether a variable is sampled from a lognormal
family of probability distributions (Thom, 1958; https://
distribution, be sure to frame your question appropriately.
[Link]/gamma-distribution-intuition-derivation-
and-examples-55f407423840) is commonly used to model waiting
 Recall that the normality assumption that forms the basis of t
times until a certain number of events occur. For example, it can be
tests, ANOVA, etc, does not refer to the data from a particular
used to model the time until a specific number of radioactive decay
experiment; rather it refers to the overall “invisible” distribution
events are observed as well as the time to onset of drug effects.
(or population) of data from which your experiment is just one
Similar to the lognormal distribution, the g distribution is
sample. Do not ask about the distribution of a particular data set.
defined only for positive values, is skewed with a long right tail, and
Any particular data set cannot be lognormal (or normal). It only
is defined by 2 parameters (the average event frequency and the
makes sense to ask whether it is reasonable to assume the data
number of events you are waiting for). With certain sets of pa-
were sampled from a lognormal (or normal) population or
rameters, a g distribution and a lognormal distribution can look
distribution.
nearly identical (Fig. 14). But it is uncommon for times or durations
 Ask about more than data collected in one particular experi-
to have a lognormal distribution, and it hard to imagine any vari-
ment. Decide how to analyze all experiments measuring a
able other than time or duration having a g distribution.
particular variable. The assumption about the distribution of
values will apply to all similar experiments with that variable.
 Do not ask if the underlying distribution is exactly lognormal (or b. The distribution of the ratio of 2 normal distributions resembles a
exactly normal). Ask about approximate agreement with the lognormal distribution. The distribution of the ratio of 2 normally
lognormal (or normal) ideal. Why approximate? One reason is distributed variables can be skewed (Fig. 15). This skewness is
that lognormal distributions extend up to infinity, but biological partly because values in the denominator can be close to zero.
variables have physical limits so their actual distribution cannot Although this skewed shape resembles a lognormal distribution, it
extend to infinity. This is also the case for normal distributions, is distinct and not lognormal.
which also extend down to negative infinity. Also note that if the
model that generates lognormal distributions is not exactly C. What to do when you are unsure about lognormality
correct, the distribution may not be exactly lognormal. For
example, a combination of additive and multiplicative factors 1. Do not get fooled by variables that are already log transformed
might lead to a distribution that is only approximately Beware of variables where the log transformation was already
lognormal. done during data collection. One example is acidity. Instruments do
13
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

Probability Density
Probability Density

0 5 10 15 0 1 2 3 4 5
Time Ratio
Fig. 14. A gamma distribution looks similar to a lognormal distribution. Equation Fig. 15. The distribution of the ratio of 2 normal distributions can look similar to a
where X is time (or duration), and k and theta are the 2 parameters that define the lognormal distribution. The graph shows the smoothed frequency distribution of
distribution (here, both are set to 2.0). Y ¼ (1/(gamma(k) * theta^k)) * X^(k-1) * exp(-X/ 10,000 ratios calculated with the numerator drawn from a normal distribution with
theta). mean ¼ 100 and SD ¼ 20, and the denominator drawn from a normal distribution with
mean ¼ 60 and SD ¼ 15.

not directly report the concentration of hydrogen ions, but rather


report the pH (which is 1 times the log10 of hydrogen ion con- distribution. For example, let us calculate the LR for the control data
centration in molar). Other examples are the amplitude of earth- of Fig. 1. AMean ¼ 293.5, GeoMean ¼ 103, GeoSD ¼ 4.53, CV ¼ 1.76,
quakes expressed on the Richter scale (the logarithm of amplitude) and n ¼ 20. Using the equation above, LR ¼ 3.9  1011, demon-
and drug potencies as pEC50 (1 times the logarithm of EC50). strating that the data are overwhelmingly more likely to have
With these data, the underlying variable (eg, concentration of [Hþ], been sampled from a lognormal distribution than from a normal
amplitude of earthquake, or EC50) is often lognormal, but the data distribution.
at hand are already log-transformed (to pH, decibels, or pEC50) so
are usually normal.
4. Consider the value of the CV
If a variable can only have positive values, you can learn a bit by
2. Compare the consistency of the SD versus the consistency of the thinking about its CV, which is the ratio of the SD divided by the
CV mean. Because the SD and mean are in the same units, the CV is a
When assessing if a variable is lognormal, do not just consider unitless ratio.
one data set at a time. Look for consistency among multiple data
sets of the same variable.
a. If the CV is small. We have already seen (see Figs. 8 and 9) that
The t test and ANOVA assume that the data sets are sampled
when the CV is lower than about 0.2, normal and lognormal dis-
from normal distributions with the same SD. In contrast, if all the
tributions are very similar.
data sets are sampled from lognormal distributions with the same
GeoSD, the SDs will differ and be approximately proportional to the
AMeans. In other words, when sampling from lognormal distribu- b. If the CV is large. All normal distributions span from negative
tions with the same GeoSD, you expect the CVs (defined as SD/ infinity to positive infinity, which means that some values are
AMean) to have similar values, but the SDs to have different values negative. However, many normal distributions are almost entirely
(Limpert et al, 2001). Figure 16 shows an example. This method has positive, with only a small negative tail that can be disregarded.
been used to demonstrate that height is lognormal, not normal After all, assumptions about the ideal distribution of data are, at
(Slavskii et al, 2021). best, only approximations.
The proportion of a normal distribution that is negative is a
function of the CV. The left panel of Fig. 17 displays a normal dis-
3. Calculate the likelihood ratio of sampling from normal versus
tribution with a CV ¼ 1.0, meaning that the mean equals the SD. In a
lognormal distributions
normal distribution, approximately 68% of the values fall within 1
What is the relative likelihood of a particular data set being
SD of the mean. Therefore, if the SD equals the mean, about 16% (¼
sampled from a normal distribution versus from a lognormal dis-
(100%  68%)/2) of the values are negative (see Fig. 17, left panel).
tribution? That likelihood ratio (LR) can be calculated by comparing
The right panel illustrates the relationship between the CV and the
the fit of the normal and lognormal distributions to a data set.
fraction of a normal distribution that is negative. The blue dot
Burnham and Anderson (2002) showed how to calculate the LR,
represents CV ¼ 1.0, indicating that 16% of the values are negative,
and we reduce the math to a simple equation derived in a
whereas the green dot represents CV ¼ 0.6, showing that 5% of the
Supplemental Material:
values are negative.
  If a variable can only take on positive values, a CV greater than
GeoMean lnðGeoSDÞ n
LR ¼ $ approximately 0.6 tells you that the data were probably not
AMean CV
sampled from a normal distribution, even as an approximation.
An LR > 1 means that the data are more likely to have been This is because in a typical sample, a nontrivial portion of the values
sampled from a normal distribution, whereas an LR < 1 means that would be negative if the distribution were normal. Limpert and
the data are more likely to have been sampled from a lognormal Stahel (2011) point out that it is common to see published papers
14
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

400
100 100%

300 80 80%
EC50 (nM)
60 60%
200

SD

CV
40 40%
100
20 20%

0 0 0%
Control A B C 0 50 100 150 0 50 100 150
Mean Mean

Fig. 16. With lognormal distributions, the SDs are proportional to the mean. Left: Four data sets. Middle: The SDs are proportional to the mean. Right: The coefficient of variation (¼
SD/mean) is consistent for all 4 data sets. This is a clue that the data may be lognormal. In fact, the data were simulated from lognormal distributions with GeoSD ¼ 2.0 and various
GeoMeans.

where data are analyzed as if normal even though this rule of These problems, of course, are intrinsic to the difficulty in dis-
thumb is violated. tinguishing normal from lognormal distributions, not just running
normality tests. You will encounter the same issue by inspecting
frequency distributions or quantileequantile plots.
5. Do not base your decision only on normality and lognormality
tests D. Our recommendation: Choose to assume lognormality based on
Many scientists, we suspect, use the rule: Assume data are the nature of the variable without normality testing
normal until proven otherwise. Lognormality, for these scientists,
needs to be proven. In other words, “let the data decide.” Keene We strongly urge scientists to assume data are lognormal largely
(1995) presents 3 reasons why this is a bad idea (and we agree): based on the nature of the variable (or parameter) being compared,
and not to rely on normality or lognormality testing. This may
 As we have seen in Figs. 8 and 9, the results of normality and sound like extreme advice far from the consensus, but others have
lognormality tests can be ambiguous. You may want to let the given the same advice:
data decide, but analyses of the data do not always lead to a clear
decision. Many data sets pass both normality and lognormality.  “The theoretical justification for using this [the logarithmic]
 Although it might seem logical to first run one test (here the transformation for most scientific observations is probably
normality test) and use that result to decide what to do next better than that for using no transformation at all… …If it were
(whether to log-transform the data), this kind of 2-stage the normal custom, when scientific observations show uncon-
approach to statistical testing can lead to misleading results so trolled variations large compared with the observations them-
is not recommended (Cartwright, 1991; Zimmerman, 1996, selves, to convert them to logarithms before estimating their
2004; Rochon et al, 2012; Gelman and Loken, 2014; Delacre et al, mean or variance, the usual result would be an increase in the
2017; Shamsudheen and Hennig, 2023). accuracy and scope of the conclusions drawn from them”
 If you do multiple similar experiments, you might end up (Gaddum, 1945).
making different decisions for different experiments or even  “I suggest that when an a priori decision about distribution has
different parts of the same experiment. It is better to analyze all to be made, the lognormal distribution should always be
similar data using the same assumptions. This is because the preferred over the normal distribution for data of this general
lognormality (or normality) assumption refers to the underlying type” (referring to variables, such as concentrations, that can
population, not to a particular sample. only have positive values) (Heath, 1967).
distribution that are negative

50%
% of values in a normal

40%

30%

20%
SD
10%
16%
0%
0.0 Mean 0 1 2 3 4
CV (=SD/Mean)

Fig. 17. A normal distribution with a high coefficient of variance (CV) contains a substantial number of negative values. Left: If the CV (which equals SD/mean) equals 1.0, then 16% of
the values in a normal distribution are negative. Right: The percentage of negative values in a normal distribution as a function of the CV calculated with this equation: 100  (1-
zdist (1/CV)). The blue dot shows that when CV ¼ 1, 16% of the values are negative. The green dot shows that when CV ¼ 0.6, only 5% of the values are negative. Equation: 100  (1-
zdist (1/CV)) (Prism format), 100  [Link] (1/CV, TRUE) (Excel format), or 100  pnorm (1/CV) (R format).

15
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

 “It is recommended that log transformed analyses should B. Outlier tests on the sample data
frequently be preferred to untransformed analyses, and that
careful consideration should be given to use of a log trans- When the raw data of Fig. 1 were analyzed using Grubbs’ outlier
formation at the protocol design stage. … If the use of a log test, the largest value in the treated group was identified as an
transformation is chosen on a case-by-case basis, then this will outlier as it is far from the rest of the data. But, as the next section
lead to inconsistencies and sometimes the wrong choice will be demonstrates, this test is very misleading with untransformed
made.” (Keene, 1995) lognormal data. After the sample data were log-transformed,
 “It is proposed that all quantities should be considered to be Grubbs’ test did not identify an outlier in either group.
lognormal in clinical chemistry if the type of distribution is
unknown. Then, laboratories need not decide whether a distri- C. Outlier tests on lognormal data
bution is quasi-Gaussian or non-Gaussian.” (Haeckel and
Wosniok, 2010) Standard outlier tests fail dramatically with lognormal data.
 “The lognormal distribution should be the first choice when Because lognormal distributions are naturally skewed, large values
modeling data taking (only) positive values. Its empirical as well that appear to be outliers are actually an expected feature of the
as theoretical justification is much stronger than for the normal distribution. Fig. 18 (left panel) demonstrates this problem using 20
distribution.” (Limpert and Stahel, 2017) simulated data sets (n ¼ 100 each, GeoMean ¼ 100, GeoSD ¼ 1.5).
 “You should (usually) log transform your positive data. The Although many of these data sets look approximately normal,
reason for log transforming your data is not to deal with Grubbs’ outlier test (which assumes normality) incorrectly identified
skewness or to get closer to a normal distribution…The reason the largest value as an “outlier” (P < .05) in 12 of the 20 data sets.
for log transformation is in many settings it should make addi- The severity of this problem increases with both sample size and
tive and linear models make more sense. A multiplicative model GeoSD, as shown in the right panel of Fig. 18. Even with a very
on the original scale corresponds to an additive model on the log modest GeoSD of 1.25dwhere the distributions look nearly nor-
scale. For example, a treatment that increases prices by 2%, maldfar more than 5% of data sets have their largest value incor-
rather than a treatment that increases prices by $20. The log rectly flagged as an outlier. With larger sample sizes or larger
transformation is particularly relevant when the data vary a lot GeoSD values, nearly every data set has its largest value mis-
on the relative scale. Increasing prices by 2% has a much identified as an outlier.
different dollar effect for a $10 item than a $1000 item” (A. This problem is not limited to formal outlier tests. Visual in-
Gelman 2019; [Link] spection of data for outliers is equally unreliable with lognormal
21/you-should-usually-log-transform-your-positive-data/). data, as our eyes are naturally drawn to values that seem “too large”
when we expect a symmetric distribution. The key lesson is clear:
before testing for or removing outliers, you must first determine
VI. Do not use standard outlier tests with lognormal data whether your data might be lognormal. If the data are lognormal,
outlier detection should only be performed after logarithmic
A. Review of outlier tests transformation.

Most outlier tests evaluate whether extreme values are likely VII. Comparing 2 groups of lognormal data
to have come from a normal distribution. These tests compute
the probability that a value as extreme (or more extreme) as the A. Lognormal t test assuming sampling from lognormal
one observed would occur by chance if the data were sampled distributions with equal GeoSDs
from a normal distribution. If the P value is small (usually < .05),
that extreme value is identified as an “outlier,” and some in- 1. Calculating the lognormal t test and reporting the results
vestigators in some situations will remove that value from Figure 19 shows example data comparing 2 groups. With
further analyses (or might report results with and without the lognormal data, it is most common to compare the GeoMeans. For
outlier). this example, the GeoMean for the control values is 103 nM and the
In a set of many samples from a normal distribution, the most GeoMean for the treated samples is 302 nM. With lognormal data,
extreme value in 5% of those samples will be identified as an outlier. it rarely (if ever) makes sense to think about the absolute difference
Note that the 5% probability refers to the fraction of samples where between GeoMeans, but instead it makes scientific sense to think
the largest value is identified as an outlier, not the fraction of values about ratios (Wolfe and Carlin, 1999). The ratio is 302/103 ¼ 2.9. In
that are identified as outliers. other words, the treatment nearly tripled the EC50.

500 100% GeoSD=1.25


with identified outlier
Fraction of data sets

GeoSD=1.5
400 75% GeoSD=2.0
300
50%
200
25%
100

0 0%
N=5 N=20 N=100 N=1000

Fig. 18. Too many outliers identified with lognormal data. Left: Twenty simulated data sets with n ¼ 100, GeoMean ¼100, and GeoSD ¼ 1.5. The red dots on 12 of the data sets
denote outliers identified by Grubbs’ test (a ¼ 0.05). Right: Fraction of simulated lognormal data sets where Grubbs’ outlier test detected an outlier with P < .05. For each com-
bination of sample size and GeoSD, 1000 data sets were simulated. The horizontal line shows that Grubbs’ test identifies an outlier with P < .05 in 5% of data sets sampled from
normal distributions (with any sample size).

16
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

4000 4

log(EC50, nM)
3000

EC50 (nM)
2000 2

1000 1

0 0
Control Treated Control Treated
Fig. 19. Example data for unpaired t test displayed on a linear axis (left) and after being transformed into common logarithms (right). These are the same data as in Fig. 1.

To calculate a CI and P value, we need to run a statistical test. 2. Graphing the results of a lognormal t test
This can easily be done by transforming all values to their loga- Figure 20 shows one way to plot the raw data and results of a
rithms to turn the lognormal distributions into normal distribu- lognormal t test. If you prefer plotting bar graphs with error bars,
tions, then running an unpaired t test on those logarithms. We call Figure 21 shows how. All 3 panels plot the GeoMeans. Lognormal
this the lognormal t test. The transformed values are shown on the distributions are asymmetrical, so the error bars in all 3 panels are
right panel of Fig. 19. Note that the distribution appears symmet- asymmetrical. The error bars on the left show the variation among
rical, as expected for data that are (before log transforming) the values, expressed as the GeoMean divided or multiplied by the
sampled from a lognormal distribution. The largest log- GeoSD. The error bars in the middle panel show how precisely the
transformed value in the treated group is just a bit larger than GeoMeans have been determined, expressed as the 95% CI of the
the rest, and Grubbs’ test on these data does not identify it as an GeoMeans. The error bars on the right show the GeoMean multi-
outlier. plied or divided by the geometric standard error (GeoSEM).
The transformed values were analyzed by a 2-sample t test Data are often presented as AMean with error bars representing
using GraphPad Prism 10.4, but any statistical software would the SD. If the values were sampled from a normal distribution, this
give the same result. The difference between the means of the would be straightforward to interpret. The range [mean  SD] to
logEC50 values is 0.467. Recall that for any 2 positive values A [mean þ SD] would contain about two-thirds of the values. But this
and B, interpretation does not work when data are sampled from a
  lognormal distribution (Fig. 22). When the GeoSD is reasonably
A A high so the data are noticeably asymmetrical, the symmetrical ± SD
logðAÞ  logðBÞ ¼ log ; and so 10logðAÞlogðBÞ ¼ :
B B error bars do a poor job of displaying the distribution of the data. In
Therefore 100.467 ¼ 2.9 is the ratio of GeoMeans (as we already some cases, as shown in this example, the lower error bar can
determined). extend downward to a negative value, which makes no sense
The t test reports that the 95% CI for the difference between the because lognormal variables can never be negative.
means of the logarithms range from 0.09435 to 0.8409. Therefore,
the 95% CI for the ratio of GeoMeans ranges from 100.09435 to
100.8409, or 1.24 to 6.93. This gives a good sense of how precisely we 3. Terms to avoid when reporting ratio results
have determined the GeoMean ratio. When describing the relationship between 2 geometric means,
The above lognormal t test reports a 2-sided P value of .0154. it is essential to use clear and consistent language to avoid confu-
This P value tests the null hypothesis that the 2 sets of logEC50 sion. In the example above, the ratio of geometric means (treated/
values are sampled from normal distributions (or populations) control) is 2.9. Some scientists would report that value as a ratio,
that have identical means and SDs. Equivalently, if you refer and others would say that the treated response was 2.9 times the
instead to the EC50s (ie, not logged), then the null hypothesis is control response.
that the 2 sets of values are sampled from lognormal distributions The following terms are ambiguous, and we suggest simply
with identical GeoMeans and GeoSDs. If the null hypothesis were avoiding them:
true, then a ratio of GeoMeans of 2.9 or larger would occur only in
around 0.77% of experiments (one tail), and a ratio of 1/2.9 ¼ 0.34  Percent increase (or decrease). Some might state that the treated
or lower would occur in around 0.77% of experiments (the other GeoMean is 190% higher than that of the control (2.9  100% 
tail). Why 0.77%? This is the P value (P ¼ .0154, ie, 1.54%), divided 100%). But it is also the case that the control GeoMean is 65.5%
equally between the lower and upper tails. The P value is less than lower than the treated GeoMean. The 2 are not symmetrical, so
the traditional cutoff (a) of .05. Therefore, if you chose that value we recommend avoiding both. Cole and Altman (2017) define a
of a and accept all the assumptions of a t test, you can reject that percentage difference that is symmetrical: 100  (difference/
null hypothesis. mean). In our example, the 2 GeoMeans were 103 nM (control)
Note the consistency of the CI and the P value. The P value is < and 302 nM (treated). The average is 202.5 nM, and the differ-
.05, and the 95% CI of the ratio (1.24e6.93, from above) does not ence is 199 nM. The difference/mean is 98% if you compute the
include the value that defines the null hypothesis (1.0). increase from control to treated, or 98% if you look at the
17
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

10000

1000
EC50 (nM)

2.9x
100 5
.0 1
p=0
p = 0.015
10
95% CI
1

Control Treated 1 3 5 7 9
Ratio of GeoMeans
Fig. 20. Plotting the results of a 2-sample (unpaired) t test of lognormal data. Left: Raw data on logarithmic axis. Right: Ratio of GeoMeans with 95% CI.

decrease from treated to control. Although this definition is With the sample data (see Fig. 19), the lognormal Welch’s t test
symmetrical, it is used rarely. We do not recommend it. reports a P value of .016. The means of the logarithms differ by
 Fold increase (or decrease). We agree with Small (2016) that the 0.4676 with a 95% CI ranging from 0.09349 to 0.8417. Take the
term fold increase should be avoided because it is used incon- antilog of all 3 values to obtain the ratio of GeoMeans (2.93) and its
sistently and so is ambiguous. Some would say there was “a 1.9- 95% CI (1.24e6.95). For this example, the lognormal t test (prior
fold increase”dthe difference between ratios of 1.0 (no change) section) and the lognormal Welch’s t test give nearly identical
and 2.9 (observed)dand others would say there was “a 2.9-fold results.
increase” (because the ratio is 2.9). Avoid this confusing term.
C. Comparing the lognormal t test with the lognormal Welch’s t test
B. The lognormal Welch’s t test
1. Power
In the previous example, we considered data sampled from 2 Figure 23 shows the results of simulations comparing the power
lognormal distributions with different geometric means but equal of the lognormal t test with the power of the lognormal Welch’s t
geometric SDs. Under these conditions, applying an unpaired t test test, when applied to lognormal data. They have nearly equal power
to the log-transformed datadreferred to as the lognormal t when the sample sizes are equal, even when the GeoSD values are
testdworks well because the 2 transformed distributions have unequal (left panel), but the lognormal Welch’s t test has more
similar variances. However, if the underlying distributions differ power when both sample sizes and GeoSD values differ (right
not only in their GeoMeans but also in their GeoSDs, the assump- panel).
tion of equal variances after log-transformation is violated. In this
scenario, the lognormal t test is no longer appropriate, and the log- 2. Type I error
transformed data should instead be analyzed using the Welch’s t The type I error rate is the probability of falsely rejecting the null
test. We refer to this approach as the lognormal Welch’s t test. hypothesis. With a set to 0.05, a well-behaved test should yield a

1000 1000 1000

800 800 800


EC50 (nM)
EC50 (nM)

EC50 (nM)

600 600 600

400 400 400

200 200 200

0 0 0
Control Treated Control Treated Control Treated
GeoMean ⋇ GeoSD GeoMean with 95% CI GeoMean ⋇ GeoSEM

Fig. 21. Plotting the results of a 2-sample (unpaired) t test of lognormal data with error bars. Left: GeoMean, with error bars showing the GeoMean multiplied or divided by the
GeoSD. Middle: Same, but with 95% CI instead. Right: Same as the left panel, but replacing GeoSD with GeoSEM.

18
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

200 data, we recommend routinely using the lognormal Welch’s t test


rather than regular lognormal t test. With unequal sample sizes,
this choice can matter a lot (see Figs. 23 and 24). With equal sample
sizes, the choice matters less, but if there is any possibility that the
150 2 populations have different GeoSD (that possibility can rarely be
ruled out), we recommend using the lognormal Welch’s t test.
EC50 (nM)

100
Mean ± SD D. Nonparametric tests with lognormal data

50 1. Understanding the Mann-Whitney and Brunner-Munzel tests


When analyzing data sampled from lognormal distributions,
researchers may notice that the data do not seem to be normal and
0 so consider using a nonparametric test that does not make any
assumption about sampling from any specified distribution. How-
ever, these nonparametric alternatives come with their own com-
–50 plexities and limitations.
Two nonparametric tests used to compare 2 groups of contin-
Fig. 22. Mean ± SD is not a useful way to quantify variability of lognormal data. Left:
uous data are the Brunner-Munzel test (Brunner and Munzel,
Simulated data (n ¼ 200, GeoMean ¼ 10, GeoSD ¼ 2.0). Right: Bar graph showing
arithmetic mean ± SD. Note that the error bar extends down to negative values, which 2000), and the better-known Mann-Whitney test, also called the
is impossible with lognormal data. The raw data are in Supplemental Material. Mann-Whitney U test. Wilcoxon independently derived an equiv-
alent test, so the names Wilcoxon rank sum test, Mann-Whitney-
Wilcoxon test, and Wilcoxon-Mann-Whitney test are also used
type I error rate near 5% when its assumptions are met and the null for this test. (Note: Do not confuse with the Wilcoxon signed-rank
hypothesis is true. test, which is for analysis of paired data.)
Although the t test is fairly robust to violations of the equal The Brunner-Munzel and Mann-Whitney tests share some
variance assumption when sample sizes are equal, its type I error fundamental properties. Both analyze the ranks of the values rather
rate can increase when there are pronounced differences in both than the values themselves and so make no assumptions about the
variance and sample size (Havlicek and Peterson, 1974; Ramsey, underlying distribution. Because a logarithmic transformation
1980; Posten et al, 1982; Zimmerman, 1987). We confirmed this preserves the rankings of positive values, each test gives identical
finding with lognormal data (Fig. 24). Both the lognormal t test and results on raw and log-transformed data. Both are tests of stochastic
the lognormal Welch’s t test do a reasonable job when the sample equality, testing whether values from one group are generally
sizes are equal, even when the GeoSD values are unequal (left larger (or smaller) than values from the other. More precisely, their
panel). However, the lognormal Welch’s t test did a much better job null hypothesis is that when randomly selecting one value from
of controlling type I errors when both sample sizes and GeoSD each group, there is a 50% chance the larger value came from either
values differed (right panel). group.
The Mann-Whitney test combines values from both groups,
3. When to use the lognormal Welch’s t test ranks them, and compares the average ranks of each group. The
Many researchers recommend using the Welch’s t test by Brunner-Munzel test takes a different approach, calculating for
default, rather than the traditional unpaired t test, regardless of each value in group A, the proportion of group B values that are
whether variances appear equal (Moser et al, 1989; Rasch et al, smaller, and vice versa. By combining these proportions, it extracts
2011; Delacre et al, 2017). Importantly, this choice should be more information from tied values than the Mann-Whitney test,
made without first performing a test for equal variances, as doing and it accurately detects differences between distributions even
so can inflate the overall type I error rate (Zimmerman, 2004; when they have different shapes. This versatility prompted Karch to
Ruxton, 2006; Delacre et al, 2017). Extending this idea to lognormal recommend using the Brunner-Munzel test routinely instead of the

1.0 1.0

0.8 0.8

0.6 0.6
Power

Power

0.4 2 0.4 n1 = 0.5*n2


GeoSD 1 < GeoSD 2 GeoSD 1 < GeoSD 2
0.2 GeoMean 1 < GeoMean 2 0.2 GeoMean 1 < GeoMean 2

0 0
0 10 20 30 40 0 10 20 30 40
n1 n1

Log t test Log Welch's test

Fig. 23. The power of the lognormal t test vs lognormal Welch’s t test, both applied to lognormal data. For each combination of GeoMean, GeoSD, and sample size, we simulated 2000 ex-
periments with values sampled from lognormal distributions, ran both statistical tests on each simulated data set, and tabulated the fraction of P values that are < .05 (our preset a). In both
panels, we set GeoMean 1 ¼1.6, GeoMean 2 ¼ 4.5, GeoSD 1 ¼1.2, and GeoSD 2 ¼ 6.0. However, in the right panel, one group is twice as larger as the other. The R code used for these simulations is
included in Supplemental Material. Left: The lognormal t test (blue) and lognormal Welch’s t test (red) have nearly equal power when GeoSD values are unequal, as long as sample sizes are
equal. Right: The lognormal Welch’s t test has more power when both sample sizes and GeoSD values differ, and when the group with the larger sample size also has a larger GeoSD.

19
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

0.25 0.3
2 n1 = 0.5*n2
0.20 GeoSD 1 > GeoSD 2 GeoSD 1 > GeoSD 2
Type I error rate GeoMean 1 = GeoMean 2 GeoMean 1 = GeoMean 2

Type I error rate


0.2
0.15

0.10
0.1
0.05

0 0
0 10 20 30 40 0 10 20 30 40
n1 n1

Log t test Log Welch's test

Fig. 24. The control of type I error by the lognormal t test and the lognormal Welch’s t test on lognormal data. For each combination of GeoMean, GeoSD, and sample size, we
simulated 2000 experiments sampled from lognormal distributions, ran both statistical tests on each simulated data set, and tabulated how frequently P < .05 (our preset a). The
GeoMeans are the same in both groups, that is, 1.6 (left panel) and 2.7 (right panel), whereas in both panels the GeoSD values differ (ie, GeoSD 1 ¼ 6.0 and GeoSD 2 ¼ 1.22). The R
code is included in the Supplemental Material. Left: Both lognormal t test (blue) and lognormal Welch’s t test (red) control the type I error reasonably well when GeoSD values are
unequal, as long as sample sizes are equal. Right: The lognormal Welch’s t test controls the type I error much better when both sample sizes and GeoSD values differ.

Mann-Whitney test (Karch, 2021). While not yet included in most When both sample sizes and GeoSDs are equal, the power of the
statistical software, the Brunner-Munzel test is available in R, Py- Mann-Whitney and Brunner-Munzel tests is almost identical, and is
thon, and Jamovi (Karch, 2023). close to the power of the lognormal t test and lognormal Welch’s t
test (Fig. 25, left panel). However, with unequal sample sizes and
2. The effect size reported by nonparametric tests GeoSDs, the power of the tests can vary considerably (see Fig. 25,
The Mann-Whitney test is often presented as a comparison of right panel). Under these conditions, the lognormal Welch’s t test
medians, which can seem appealing for analyzing data sampled has the most power. The nonparametric tests have a bit less power,
from lognormal distributions. Because the GeoMean of a lognormal with the Brunner-Munzel test performing better than the Mann-
sample estimates the median of its underlying distribution, one Whitney test. The lognormal t test has the least power of the 4 tests.
might be tempted to interpret the Mann-Whitney test as a
straightforward comparison of geometric means. Indeed, some
4. Type I error control
implementations of the Mann-Whitney test even report differences
Figure 26 compares the type I error rates of the statistical tests in
between medians with CIs. However, this interpretation is only
simulated lognormal data. In all these simulations, the GeoMeans
valid if the 2 distributions are identical in shape and differ only in
were set to the same value. When the distributions had identical
location (Stonehouse and Forrester, 1998; Fagerland and Sandvik,
GeoSDs and sample sizes (left panel), both nonparametric tests
2009).
maintained type I error rates near 0.05. However, when the dis-
This assumption that both distributions have the same shape
tributions have different GeoSDs (see Fig. 26, middle panel), the
does not hold with lognormal data (Divine et al, 2018). Although 2
type I error rate of the Mann-Whitney test is inflated, as previously
normal distributions will have the same shape when their SDs are
observed (Karch, 2021). The inflation of type I error rate for the
equal, 2 lognormal distributions with different GeoMeans will have
Mann-Whitney test is increased further when both sample sizes
different shapes even if they have the same GeoSD. Thus, the
and GeoSDs are set to unequal values (right panel). In contrast, the
assumption required to interpret the Mann-Whitney test as a me-
type I error rate of the Brunner-Munzel test remains near its
dian (or GeoMean) comparison rarely applies to lognormal data.
nominal value of 0.05 for all experimental designs, so long as the
Moreover, it has been shown that the Mann-Whitney test can yield
sample sizes in both groups are greater than or equal to about 10.
small P values when comparing lognormal distributions with the
same GeoMean but different GeoSDs, further complicating its
interpretation (Fagerland, 2012). 5. Conclusions about nonparametric tests for lognormal data
The results of a Brunner-Munzel test can be summarized as a For analyzing lognormal data, the lognormal Welch’s t test is the
probability of superiority (with a CI). This is the probability that a optimal choice for several reasons. First, it controls the type I error
randomly selected value from group A will be larger than a rate, that is, it provides acccurate P values when no effect exists.
randomly selected value from group B. For example, a probability of Second, it maximizes statistical power. Third, it provides easily
superiority of 0.8 indicates an 80% chance that a random value from interpretable effect sizes (ratios of GeoMeans), and directly ad-
group A exceeds a random value from group B. This is a less intu- dresses the data’s multiplicative nature. The lognormal t test has
itive way to summarize the result than the ratio of GeoMeans re- the same advantages except that when GeoSDs differ, it is less
ported by a lognormal t test. powerful than the lognormal Welch’s t test (see Fig. 25).
If you are not sure about the underlying distribution but suspect
3. The statistical power of nonparametric tests with lognormal data it might be lognormal, you may prefer to use a nonparametric test.
We ran Monte Carlo simulations to evaluate the power of If the experiment has equal sample sizes, both nonparametric tests
various statistical tests when analyzing lognormal data. For each perform similarly, but the Brunner-Munzel test has a bit more
experimental design, we specified the GeoMean, GeoSD, and power. If the sample sizes differ substantially, avoid the Mann-
sample size of each group. We simulated 2000 experiments, and Whitney test as it has lower power and larger type I errors than
tabulated the fraction of P values that are < .05 (our preset a). The R the Brunner-Munzel test. Thus, if you want to use a nonparametric
code used for these simulations is included in Supplemental test with lognormal data, choose the Brunner-Munzel test because
Material. it has more power and a smaller (more appropriate) type I error.
20
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

1.0 1.0

0.8 0.8

0.6 0.6
Power

Power
0.4 0.4
n1 = n2 n1 = 0.5* n2
GeoSD 1 = GeoSD 2 GeoSD 1 < GeoSD 2
0.2 GeoMean 1 < GeoMean 2 0.2 GeoMean 1 < GeoMean 2

0.0 0.0
0 10 20 30 40 0 10 20 30 40
n1 n1

Log t test Log Welch's test Mann-Whitney test Brunner-Munzel test

Fig. 25. Power of parametric and nonparametric tests applied to lognormal data. For each combination of GeoMean, GeoSD, and sample size, we simulated 2000 experiments,
sampling from lognormal distributions, ran 4 statistical tests on each simulated data set, and tabulated the fraction of P values that are < .05 (our preset a). The R code used for these
simulations is included in Supplemental Material. Left: The GeoSDs of the 2 groups were set to the same value (2.7), and the sample sizes were also made equal. The GeoMean of the
second group was 2.7 times the GeoMean of the first group (ie, 4.5 vs 1.6). The power of all the lognormal t test (blue, circles), lognormal Welch’s t test (red, circles), Mann-Whitney
test (blue, triangles), and Brunner-Munzel test (red, triangles) were nearly the same. Right: The sample sizes of the first group were set to half the sample size of the second, and the
GeoSD values were set to 1.2 and 6.0. Under these conditions, the power of the various tests differs.

E. Why the unpaired t test (without log transformation) should be When you run a t test assuming sampling from lognormal dis-
avoided with lognormal data tributions, the analogous effect size is reported as the ratio of the 2
GeoMeans. A ratio is a logical way to think about the size of an
1. Results of analyzing the sample data with an unpaired t test experimental effect. And it is helpful to report the ratio with its CI to
without log transformation give a sense of how precisely the ratio has been determined.
Let us return to our motivating example (see Fig. 1). The un-
paired t test (without log transformation) yielded P ¼ .22, so the
null hypothesis could not be rejected. The 95% CI for the difference b. Loss of statistical power (for a given sample size). Figure 27 pre-
between means ranged from 175 to 739. With such a wide CI, the sents additional simulations demonstrating that analyzing
data are consistent with no difference, a moderate decrease, or a lognormal data as normal results in a loss of statistical power. The
large increase. In other words, no conclusion is possible. simulated experiments had n ¼ 20 per group. All values were
sampled from lognormal distributions with GeoSD ¼ 4, and the
true effect size is 3 (the treated GeoMean is 3 times the control
2. Problems when analyzing lognormal data as normal GeoMean). The left side of the figure shows 1 of the 1000 simula-
tions. The right side shows P values for 1000 simulations.
a. Less useful effect size (difference, rather than ratio). When you All 1000 P values are graphed for all 4 analyses. For each test, the P
run a t test assuming sampling from normal distributions, the effect values vary over several orders of magnitude, more variability than
size is reported most simply as the difference between the 2 means. many scientists expect. Statistical power is defined as the probability
With ratio variables, the difference is rarely a useful way to view the that the P value will be < .05 (or any chosen cutoff), so is the fraction
effect, so many scientists do not bother reporting that difference or of the dots below the red line that defines P ¼ .05. Analyzing the data
its CI. correctly (assuming lognormal distributions) results in 69% power

n1 = n2 n1 = n2 n1 = 0.5*n2
0.25 0.25 0.25
GeoSD 1 = GeoSD 2 GeoSD 1 > GeoSD 2 GeoSD 1 > GeoSD 2
0.20
GeoMean 1 = GeoMean 2 0.20
GeoMean 1 = GeoMean 2 0.20
GeoMean 1 = GeoMean 2
Type I error rate

Type I error rate

Type I error rate

0.15 0.15 0.15

0.10 0.10 0.10

0.05 0.05 0.05

0 10 20 30 40 0 10 20 30 40 0 10 20 30 40
n1 n1 n1

Mann-Whitney test Brunner-Munzel test

Fig. 26. Type I error of parametric and nonparametric tests applied to lognormal data. The method matches that of Fig. 24. In all 3 panels, the GeoMeans of both groups were set to
1.6. Left: The GeoSDs of the 2 groups were set to the same value (6.0), and the sample sizes were also made equal. The type I error Mann-Whitney test (blue), and Brunner-Munzel
test (red) were nearly the same. Middle: The GeoSD values were set to 6.0 and 1.2, but the sample sizes were kept equal. Right: The GeoSDs were the same as in the middle panel,
but the sample sizes of the second group were set to half the sample size of the first. Under these conditions, the type I error of the Brunner-Munzel was always close to 0.05, but the
type I error of the Mann-Whitney test was several fold higher.

21
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

0.1
0.05
5000
0.01

p-value
Response

0.001
2500

0.0001

GeoSD=4
0 0.00001 Effect = 3x
n = 20

Control Treated
0.000001

Power: 69% 68% 35% 33%


Test: t-test Welch's t-test t-test Welch's t-test
Assume: Lognormal Lognormal Normal Normal

Fig. 27. Loss of statistical power when lognormal distributions are analyzed as normal. Left: A simulated experiment with n ¼ 20 per group, sampled from lognormal distributions
with GeoSD ¼ 4, and the true effect size is 3 (the treated GeoMean is 3 times the control GeoMean). Right: Results from 1000 simulated experiments with P values from 4 analyses
(from left to right): t test assuming lognormal data, Welch’s t tests assuming lognormal data, t test assuming normal data, and Welch’s t test assuming normal data. The powers (the
fraction of the P values < .05) are shown below each lane. The raw data are in Supplemental Material.

with the lognormal t test and 68% with the lognormal Welch’s t test. To summarize this contrived example: assuming a normal dis-
Analyzing the data incorrectly (assuming normal distributions) tribution requires a larger sample size: 64 per group versus 41 per
yields only 35% with the t test and 33% with the Welch’s t test. Similar group than if you assume the data are lognormal, a 56% increase.
simulations were reported by Fayers (2011). Figure 28 shows samples from the 2 hypothetical populations on
linear and logarithmic axes, and the right panel shows the needed
c. Increased sample size requirement (for constant power). sample size (per group) when computed correctly (assuming
Another way to assess how much it matters to identify lognormal lognormal) and incorrectly (assuming normal).
variables is to see how it impacts the calculation of necessary sample Limpert and Stahel (2011) used simulations with a variety of
size. GeoSD values (1.5e3.5) and a variety of intended effect sizes, in
Let us assume we are sampling from 2 lognormal distributions order to determine the degree to which sample size can be reduced
and running a 2-sample t test. The control GeoMean is 100 and we by recognizing lognormal distributions. They showed that incor-
are seeking sample size to detect a doubling to a GeoMean of 200. rectly assuming that lognormal data were sampled from normal
Assume GeoSD ¼ 3, and use conventional values for a (0.05, 2- distributions resulted in an increase of necessary sample size from
sided) and desired power (80%). somewhere between 20% and 300% depending on GeoSD and effect
Because the data are lognormal, a t test would be run on the size. One of their examples assumed GeoSD ¼ 2.4 (similar to our
logarithms of the values, so we need to convert to log scale before GEOSD ¼ 3), and they found that mistakenly analyzing the data as if
calculating necessary sample size. On a log scale, the hypothetical they were sampled from a normal distribution raised the required
means are log10(100) ¼ 2 and log10(200) ¼ 2.3, and the expected sample size from 10 to 16, a 60% increase (which matches Fig. 28).
SD ¼ log10(3) ¼ 0.477. For these parameters, sample size calculators
such as GraphPad Prism Cloud’s Power Analysis calculator or 3. Switching to the Welch’s t test does not solve the problem
G*Power (Faul et al, 2007) report the required sample size of 41 per The SDs in Fig. 1 differ considerably between the control and
group. Wolfe and Carlin (1999) presented equivalent calculations. treated groups. This is often the case for lognormal data, even when
But what if we wrongly assume the data are sampled from normal the GeoSDs are identical, because the SD of a lognormal distribution
distributions? This is a contrived situation, but let us do our best to do depends on both its GeoMean and GeoSD. These unequal SDs might
the corresponding sample size calculations assuming normal distri- tempt researchers to replace the usual unpaired t test with Welch’s
butions. Using the equation shown below, the corresponding AMeans t test, as this does not assume the variances (or SDs) of the groups
are 224.2 (control) and 448.4 (treated), and the corresponding SDs being compared are equal (Rasch et al, 2011; Delacre et al, 2017).
are 282.8 (control) and 565.7 (treated). The SDs differ because with With these sample data the results of the Welch’s t test (without
lognormal data with fixed GeoSD, the SD will be larger when the log transformation) are nearly identical to those of the regular t test.
GeoMean is larger. G*Power and Prism Cloud’s Power Analysis The P values (two-tailed) from both tests are .22. The CIs for the
calculator both allow you to specify different hypothetical SD values difference between means are also quite similar (t test, 175 to 739;
for the 2 populations. For this experimental design, both calculate Welch’s t test, 179 to 743).
that the resulting necessary sample size is 64 per group. Because the Welch’s t test does not assume the SDs are equal,
some might expect it to have more power with lognormal data. In
ln ðGeoSDÞ2
AMean ¼ GeoMean$e 2 fact, its power is a bit less than the t test with lognormal data
qffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi (Zimmerman and Zumbo, 1993; de Winter, 2016). Another reason
2 to avoid the Welch’s t test with lognormal data is that the observed
SD ¼ GeoMean2 $elnðGeoSDÞ  1
type I error rate of the Welch’s t test applied to data sampled from

22
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

3000 10000 80

Required sample size


1000 60
2000

100 40
1000
10 20
0
1 0
Control Treated Control Treated Assume: Normal Lognormal

Fig. 28. Recognizing lognormality leads to smaller required sample size. The left and middle graphs show 100 values sampled from 2 hypothetical lognormal distributions with
GeoSD ¼ 3 with GeoMeans ¼ 100 and 200. The graph on the right shows the required sample sizes (per group) computed correctly assuming the data are lognormal or incorrectly
assuming sampling from normal distributions, setting a to 0.05 and desired power to 80%. The raw data are in Supplemental Material.

skewed distributions can be much higher than the preset value of a antilogarithm of all 3 values yields the results as ratios, which are far
(Ahad and Yahaya, 2014). easier to interpret (see Fig. 30, right panel). The GeoMean of the ratios
is 100.2075 ¼ 1.61, with a 95% CI ranging from 1.14 to 2.28. The 95% CI
VIII. Other comparisons of lognormal data does not include 1.0 (the value that denotes no change), which is
consistent with the P value being < .05. With GraphPad Prism (version
A. Paired t test of lognormal data 6.0 and later), choose the ratio paired t test to obtain the results
directly without needing to calculate logarithms and antilogarithms.
Figure 29 shows animal weight before and after an intervention Analyzed this way, the data suggest that the intervention
(left panel) and the absolute difference for each animal (right changes weight, but this is not super-convincing because the CI is
panel). The raw data are in Supplemental Material. Note that the set so wide, ranging from a 14% increase to a bit more than a doubling.
of differences is skewed, and appears to be lognormal (in fact, the
data were simulated so the differences are lognormal). B. One-way ANOVA with Dunnett’s test of lognormal data

1. Wrong analysis: Paired t test of untransformed data Figure 31 shows EC50 values collected in control conditions and
The paired t test looks at the set of differences and the null in the presence of 2 drugs in 7 experiments. In the presence of the
hypothesis that there is no effect of the intervention, in other words drugs, the EC50s tend to be larger so the pEC50s tend to be smaller.
that all the variation is due to normal random sampling. The P value
is .07 (two-tailed). The mean difference is 10.2 g with a 95% CI
ranging from 0.80 to 21.1 g. Because the P value is > .05, the 95% CI 1. Incorrect analysis assuming normal distributions
includes 0.0 (no difference). Interpretation of data always depends One-way ANOVA tests the null hypothesis that all 3 data sets
on the details. These data give a hint of a weight gain, but this large are sampled from normal distributions with the same mean and
a gain (or an equally large loss, because the P value is 2-sided) SD. One-way ANOVA of these EC50s results in P ¼ .07, high enough
would occur in 7% of experiments if the null hypothesis (no dif- that the null hypothesis is not rejected. Welch’s one-way ANOVA,
ference) is true. The CI (also called compatibility interval) is which does not assume equal SDs, results in a slightly higher P
consistent with a small decrease, no change, or a large increase. value (P ¼ .11).
Note that this analysis assumes the differences are sampled from a
normal distribution. 2. Analysis assuming lognormal distributions
Here are several reasons to assume these data in Fig. 31 are
2. Analysis of log-transformed data assuming lognormal sampled from lognormal distributions:
distribution of differences
There are 2 equivalent ways of thinking about running a  EC50s are generally lognormal (Hancock et al, 1988; Kaumann
lognormal paired t test. et al, 1989; Christopoulos, 1998; Walker et al, 2010; Liang et al,
2015).
 Transform all the values to logarithms and then run a paired t  The SDs are very different for the 3 data sets (4, 50, and 150), but
test in order to test the null hypothesis that the means are equal. their CVs are more similar (63%, 77%, and 118%).
 Compute the ratio of before/after for each animal, transform  The CVs in all data sets are larger than 60%. As noted earlier, this
those ratios to logarithms, and run a 1-sample t test to test the (plus the fact that negative values are impossible) makes it
null hypothesis that the mean of those logarithms is zero exceedingly unlikely for the data to be sampled from a normal
(equivalently, the null hypothesis is that the GeoMean of the distribution.
ratios is 1.0).
To analyze the data assuming sampling from lognormal distri-
These 2 methods are totally equivalent. We prefer the second butions, the first step is to transform the values to logarithms. For
approach, because it focuses on before/after ratios. Figure 30 shows this example, we converted the data from nanomolar to molar,
the set of ratios and log (ratios), and the results of the ratio paired t transformed to log10 logarithms, and then multiplied those loga-
test. We used GraphPad Prism 10.4, but any statistical software rithms by negative 1 to compute the pEC50. Just as pH is the
would give the same result. The P value is .011. The mean log (ratio) negative logarithm of [Hþ], the pEC50 is the negative logarithm of
is 0.2075, with a 95% CI ranging from 0.0569 to 0.3580. Taking the EC50. For example, when the EC50 is 10 nM, which is 108 molar,
23
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

80

Weight difference (g)


60
60
Weight (g)
40
40
20
20
0
0
Before After After
-
Before
Fig. 29. Sample data for paired t test. Left: Raw data before and after an intervention; each line is a different animal. The values are in Supplemental Material. Right: The difference
(After e Before) for each animal.

the logEC50 is 8, and the pEC50 is 8. pEC50s are often used by For comparing 2 lognormal data sets, our simulations showed
pharmacologists to avoid negative numbers. the advantage of using the lognormal Welch’s t test (see section The
The middle panel (see Fig. 31) plots the pEC50 in the 3 groups. lognormal Welch’s t test). We suspect there are similar advantages to
One-way ANOVA of these values resulted in P ¼ .0003 (compare to using Welch’s ANOVA when analyzing the logarithms of lognormal
P ¼ .07 from ANOVA on untransformed data). Now focus on the data, but we have not run any simulations to test this idea.
right panel of Fig. 31. The follow-up Dunnett’s multiple-comparison
test compares the result of each drug to the control. Dunnett’s test C. Two-way ANOVA of lognormal data
reports the CI for the difference between logarithms. We trans-
formed the confidence limits to their antilogarithm to plot the ratio 1. Example and analysis assuming sampling from normal
of EC50s. For control versus drug A, the difference in logarithms is distributions
1.13 with a 95% CI ranging from 0.56 to 1.69. Transform all 3 values Suppose nicotine increases circulating levels of a certain hor-
to their antilogarithm (10 to those powers) and the ratio of EC50s is mone, and you wish to know whether the drug effect is different in
13.5, with the 95% CI ranging from 3.63 to 49.0. Because those CIs males and females. You measure hormone levels in control and
do not come close to 1.0 (the value that signifies no effect), the P drug-treated animals, in both sexes. The effect seems much larger
values are much < .05. in females (Fig. 32).

80 6 2.0
1.0
5
60 1.5 p = 0.011; n=12 pairs
log(weight, g)

4
Weight (g)

0.5
log(ratio)
Ratio

40 3 1.0

2 0.0
20 0.5
1.61
1
95% CI
0 0 0.0 –0.5
Before After Ratio Before After log(ratio) 0.5 1.0 1.5 2.0 2.5
Ratio
Fig. 30. Paired ratio t test. Left: Raw data and ratio. Middle: logarithms of raw data and ratio. Right: Summary of ratio t test showing: the GeoMean of the ratio, the 95% CI of that
ratio, and the P value testing the null hypothesis (ie, that the true ratio is 1.0).

24
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

500 9

400
B / Control
EC50 (nM)

8 p = 0.0003
300

pEC50
200
7
A / Control
100
p = 0.0012

0 6

Control Drug A Drug B Control Drug A Drug B 1 10 100


EC50 Ratio

Fig. 31. One-way ANOVA example. Left: Raw data as EC50 in nanomolar. The values are in Supplemental Material. Middle: Transformed to pEC50 (convert nanomolar to molar,
transform to log10, then multiply by 1). Right: Results of Dunnett’s multiple comparisons test showing the 95% CI for the ratio of EC50s and the P values testing the null hypothesis
that the true ratio is 1.0. Both CIs and P values are corrected for multiple comparisons by the Dunnett calculations.

To demonstrate and quantify this result, run two-way ANOVA  The largest value in the Male/Drug group is quite a bit larger
and focus on the results for interaction, which assesses whether the than the others, and is identified as an outlier by Grubbs’ outlier
drug effect differs between male and female animals. The interac- test (P < .01). But Grubbs’ test assumes the data (except for the
tion P value is tiny (<.0001). The drug effect is 25 ng/mL greater in possible outlier) are sampled from normal distributions. Values
females, with a 95% CI ranging from 18 to 32 ng/mL. This would be larger than the others are expected in data sampled from
convincing evidence of a SEX by DRUG interactiondif all the as- lognormal distributions.
sumptions behind the analysis are true.
For these reasons, especially the first, it makes sense to assume
lognormality, not normality. Can normality and lognormality tests help
 ANOVA assumes all values are sampled from normal distribu-
decide? Not in this case. The data from the drug-treated males fail 3
tions. This is not obviously false, but there is a hint of
normality tests but pass all 3 lognormality tests, but the remaining 3 data
asymmetry.
sets pass 3 normality tests (P > .05) and also pass 3 lognormality tests.
 ANOVA assumes that the underlying populations are not only
The right side of Fig. 32 shows the log-transformed data. Now
normal but that they all have the same SD. Here, the variability
the variation is similar across conditions, there are no longer any
(SD) clearly differs substantially between groups, violating a
obvious outliers, and the data look normally distributed.
major assumption underlying ANOVA.
Two-way ANOVA on the log-transformed data shows no evi-
2. Two-way ANOVA assuming sampling from lognormal dence of interaction (P ¼ .343), so it makes sense to look at the row
distributions (sex) and column (drug) effects. Tests of both null hypotheses (that
But… could the data be lognormal? There are reasons to think sex makes no difference, and that the treatment makes no differ-
so: ence) result in P < .0001. The effect of the drug is shown in the
ANOVA results as the difference between the mean of the log-
 The measurement is concentration, a ratio variable which is transformed control and drug-treated values. It is reported as
often lognormal. 0.3301 (95% CI, 0.2852e0.3751). Use the 10^ (ie, 10 to the power of)
 The variation is larger when the mean is larger, as expected for transform on all 3 values to express the drug effect as a ratio. The
lognormal data. drug-treated animals had a response 2.1 times that of the control

100 2.0
Log(Concentration, ng/ml)

Control
Concentration (ng/ml)

Drug
80
1.5
60
1.0
40
0.5 Control
20 Drug

0 0.0
Male Female Male Female
Fig. 32. Results of simulated experiment asking whether there is an interaction between drug response and sex. The values are in Supplemental Material. The solid horizontal lines
represent the arithmetic means.

25
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

animals (95% CI, 1.9e2.4). Similar calculations show that the fe-  Deleting such values would bias the results (leading to an
males had a response 3.1 times larger than the males (95% CI, increased GeoMean), because only the smallest value(s) would
2.8e3.4). be removed.
This example demonstrates:  Replacing a “nondetect” value with the LOD would also bias the
results (increase the GeoMean) because the actual (unmeasur-
 How one can be fooled by analyzing lognormal data as if the able) values are all less than the LOD.
values were sampled from normal distributions.  Replacing nondetects with zero is not possible when assuming
 Why it is important to differentiate additive effects from mul- lognormal distribution, because analyses of lognormal data first
tiplicative effects. Analyzing the raw data (left panel) suggested take the logarithm of all the values, and the logarithm of zero is
a substantial 2-way additive interaction, that is, with a larger not defined.
drug effect in females. However, analysis of the log-transformed
data (right panel) revealed no multiplicative interaction. Replacing nondetects with some value between zero and LOD is
 That the decision of whether to assume sampling from the best solution. Verbosek (2011) used simulations of lognormal
lognormal distributions cannot depend entirely on normality data with different total sample size sizes and number of non-
and lognormality tests. detects, and recommends assigning a value of LOD/√2 to all values
that are less than the LOD.
In the example of Fig. 32, analyzing the data properly (after log The topic of how to deal with values too low to measure has
transformation) prevented falsely concluding there was an inter- been reviewed by Shoari and Dube  (2018), but these authors do not
action. The converse can also happen: in some cases, analyzing the focus on lognormal data. Zhang et al (2009) recommend analyzing
data after log transformation reveals a multiplicative interaction data with too-small-to-measure values using nonparametric
that would have been missed had the data not been transformed. methods, and explain how to extend common nonparametric tests
to data with values below the detection limit. Zhou and Tu (1999)
devised a likelihood method for comparing data sets that are a
IX. Additional topics mixture of values from a lognormal distribution plus zeros.

A. How to handle values that are zero, negative, or below the limit B. Comparing the arithmetic means of lognormal distributions
of detection
Some researchers argue for reporting the AMean instead of the
1. If some values are zero or negative (occurs rarely) GeoMean in certain contexts. For example, Parkin and Robinson
By definition, an ideal lognormal distribution comprises only (1992) argue that the AMean provides a more meaningful com-
positive values. However, in the real world, some data sets that are parison than the GeoMean or median when comparing variables
close to lognormal nevertheless can contain values that are zero or such as pollutant concentrations across locations, where the
negative. This can occur in 3 ways: important consideration is the total mass of pollutant at a given
site.
 Zeros can occur when the variable is a count, for example, Surprisingly, working with the AMean of lognormal data re-
number of immunopositive cells in a tissue section, or quires special methods. This is because the asymmetry of
number of days with rainfall. Lognormal distributions lognormal distributions causes random samples to often under-
describe continuous variables, so really are not appropriate represent large values. Thus, it is not appropriate to calculate the
for variables that are counted. The analysis of such data AMean by adding up all the values and dividing by the sample size,
should not be based on assuming sampling from a lognormal as this tends to underestimate the true population AMean. The
distribution. following papers describe appropriate procedures to compute the
 Zeros can occur when the variable has a distribution that is not AMean and its CI (Zhou and Gau, 1997; Wu et al, 2003; Olsson,
entirely lognormal. One example is the Comet assay, used for 2005) and to compare AMeans of different groups (Zhou et al,
detecting DNA damage in eukaryotic cells (Bright et al, 2011). 1997).
The method uses gel electrophoresis to quantify the fraction of
the DNA that has been fragmented so appears in the tail of the C. The GeoMean as an average of ratios
“comet.” Combining data from many cells, the distribution is a
cluster of zeros plus a collection of values from a lognormal In this article, we explain the use of the GeoMean as a way to
distribution. Special methods are needed to analyze such data. summarize a set of values sampled from a lognormal distribution.
 Subtracting a baseline or nonspecific value can lead to a differ- But the GeoMean is more widely applicable than that. It is the only
ence that is zero or negative. The true population value may be consistent way to average ratios (Fleming and Wallace, 1986),
larger than zero, but experimental (ie, random) error in total which makes it essential in areas such as physics and engineering
and/or baseline values can result in a zero or negative difference. (Mahajan, 2019), finance ([Link]
It is probably best to analyze such data without subtracting a investing/071113/[Link]), and other
baseline or nonspecific signal (and fit the baseline in the anal- domains ([Link]
ysis). Another approach is to add a positive constant to each statistics-for-data-visualizations-2619dbb3677a; Chargin, 2020).
value in the data set, so that all values become positive before This makes sense only when the ratios are unitless because they are
log transformation. the ratio of the same variable measured in 2 conditions, for
example, 2 treatments, 2 time points, or 2 genotypes.
This property makes geometric means essential in pharma-
2. If some values are below the detection limit (occurs rarely) cology whenever we need to average ratios, whether analyzing
With some experimental systems, a value may be too low to relative potencies or measuring fold-changes in receptor expres-
measure. You know the value cannot be zero (or negative), and that sion. The problem with using the AMean to summarize a set of
it is smaller than the limit of detection (LOD). Such values are called ratios is that the result depends on which group or treatment is
left-censored. How can such a value be accounted for? chosen as the baseline. For example, consider 2 experiments: in
26
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

one, drug A is 3 times more potent than drug B, but in the other, it is 25
only one-third as potent. The AMean of these ratios is (3 þ 1/3)/2 ¼
1.67, implying that drug A is, on average, 1.67 times more potent
than drug B. If we instead measure the potency of B relative to A, we 20
arrive at the opposite conclusiondthat B is 1.67 times more potent

Skewness
than A. This inconsistency can lead to misleading interpretations.
15
In contrast, the geometric mean of 3.0 and one-third is 1.0,
correctly showing that neither drug is consistently more potent.
The GeoMean is a baseline-independent and consistent summary, 10
making it the appropriate method for averaging ratios.

5
D. Geometric Coefficient of Variation of lognormal data

The CV quantifies variability among values that can only be 0


positive. CV is defined as the SD/AMean. Because SD and AMean are 1 2 3 4
expressed in the same units, the CV is a unitless ratio. It is often
multiplied by 100 and reported as a percentage. The smallest GeoSD
possible value of CV is 0.0, which would only occur when all values
are identical (no variation). Fig. 33. Skewness of a lognormal distribution. The equation in this figure is equivalent
to Equation 4.8 in Crow and Shimizu (1988).
With lognormal data, variability and asymmetry are inter-
twined, and the GeoSD quantifies both. A lognormal distribution
with more variation is also more asymmetrical. The standard helpful to quantify the skewness of data sampled from lognormal
definition of the CV (SD/AMean) can be rewritten as a function of distributions.
only the GeoSD for data sampled from a lognormal distribution
(Koopmans et al, 1964; Ott, 1995; Elassaiss-Schaap and Duisters,  The figure shows the skewness of 100 simulated data sets
2020). sampled from a lognormal distribution with GeoSD ¼ 3.0, which
qffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffiffi corresponds to skewness ¼ 8.2. But most of the simulated
2 samples have skewness that are much smaller than that (Kirby,
GeoCV ¼ eln ðGeoSDÞ  1
1974; Cox, 2010).
We suggest not reporting the geometric CV (GeoCV) for 2 rea-  The maximum possible skewness is limited by sample size (n)
sons. First, it adds no information not already expressed by the according to this equation: (n  2)/√(n  1) (Kirby, 1974; Cox,
GeoSD (as GeoCV can be calculated from GeoSD). Second, alterna- 2010). Therefore, the skewness can be 8.2 or greater only
tive definitions of GeoCV are in use, leading to inconsistent results. when n is 71 or larger. But even huge samples rarely obtain
GeoCV has been defined as GeoSD 1 (Kirkwood, 1979), as skewness that large. With n ¼ 5000 (rightmost column), only
ln(GeoSD) (Proost, 2019), and as GeoSD/GeoMean (Martinez and 18% of the 100 simulated samples have skewness > 8.2.
Bartholomew, 2017).  Even with large samples, the skewness varies considerably from
sample to sample. Skewness is therefore not a reliable way to
E. Geometric standard error of a geometric mean assess the asymmetry of data sampled from a lognormal
distribution.
The GeoSEM of lognormal data is a unitless factor that can be
multiplied by or divided into the GeoMean (Kirkwood, 1979). This is In contrast, the right side of Fig. 34 shows that the GeoSD of
defined as the antilogarithm of the SEM of the log (values), which is simulated samples is well behaved. The values are centered on the
equivalent to: true GeoSD, and the variation decreases with larger samples.


p1 G. How much is lost when normal data are analyzed as if
GeoSEM ¼ GeoSD n

lognormal?
Note the similarity between the definitions of SEM and GeoSEM.
SEM equals SD multiplied by (1/√n), whereas the GeoSEM equals Analyzing lognormal data as if they were sampled from normal
GeoSD to the power of (1/√n). distributions can lead to major problems in data analysis. How bad
Beware of the earliest definition of the GeoSEM, which had the is the reverse issue: analyzing normal data as if they were sampled
same units as the data (Norris, 1940). This value was to be added to from lognormal distributions?
or subtracted from the GeoMean, which makes little sense for the This is a bit tricky to think through because all normal distri-
asymmetrical lognormal distribution. butions include negative values, which are impossible in lognormal
distributions. But if the CV is small enough, only a tiny fraction of
F. Why skewness is not a useful parameter with lognormal data samples will contain negative values. Therefore, it is only possible
to be unsure about whether data are sampled from normal versus
Skewnessdmore precisely Pearson’s moment coefficient of lognormal distributions when the CV is small. In this case, the
skewness, abbreviated G1dquantifies the asymmetry of a distri- lognormal distribution looks almost identical to a normal distri-
bution. A perfectly symmetrical distribution has a skewness of 0.0. bution, so the loss of power tends to be minimal (simulations not
Distributions with a long right tail, including lognormal distribu- shown).
tions, have positive skewness. Not surprisingly, there is a simple
relationship (Crow and Shimizu, 1988) between the GeoSD and the H. Performing lognormal comparisons with GraphPad Prism
skewness of a lognormal distribution (Fig. 33).
Although skewness of a lognormal distribution is related to Although data from lognormal distributions can be analyzed by
GeoSD, the left panel of Fig. 34 demonstrates 3 reasons why it is not programming languages such as R and Python, GraphPad Prism is
27
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

30
11

9
Skewness
20

GeoSD
7

10 5

3
0 1
n= 5

50 0

n= 5

50 0
n= 0
n= 25
n= 00
n= 250

n= 0
n= 25
n= 00
n= 250
00

00
n= 00

n= 00
1
n=

1
n=
1

1
1

1
Fig. 34. Demonstration that skewness is not a useful measure of asymmetry of a lognormal data set. Each dot represents the analysis of one simulated data set drawn from a
lognormal distribution with GeoSD ¼ 3.0 and GeoMean ¼ 10. The left panel shows the values of Skewness from 100 simulated data sets of various sizes, and the right panel shows
the values of GeoSD from those same data sets. The horizontal black lines represent the medians. The horizontal red lines mark the population skewness (8.2) and GeoSD (3.0).

the only statistics program we know of that can directly perform t B. Lognormality in pharmacology
tests and ANOVA with lognormal data. Performing a paired t test
with lognormal data has been available since version 6. Choose the  Measurements such as concentration, weight, and enzyme ac-
“ratio paired t test.” One- and 2-sample (unpaired) t tests and one- tivity are often lognormal.
way ANOVA can be done with lognormal data starting with version  Key pharmacological parametersdincluding EC50, IC50, Kd, Km,
10.5. Simply select the option to assume sampling from lognormal Kon, Koff, clearance, and half-lifedfollow lognormal distributions.
distributions, and then choose between the lognormal t test and  The ubiquity of lognormal distributions in pharmacology
the lognormal Welch’s t test. stems from the multiplicative nature of many chemical and
biological processes and also from the fact that a parameter
X. Summary formed as a ratio of 2 lognormal parameters will itself be
lognormal.
A. Properties of lognormal distributions

 Lognormal distributions arise naturally from multiplicative C. Recognizing lognormal data


biological and physical processes, whereas normal distributions
come from additive processes. Thus, lognormality often reflects  The decision to treat data as lognormal should be based pri-
fundamental biological processes and is not just a statistical marily on the nature of the variable and prior experience, and
“trick.” not on the result of normality or lognormality tests.
 All values in a lognormal distribution are positive. Zero and  For a variable that can only be positive: when the CV is greater
negative values cannot be part of a lognormal distribution. than about 0.6, the data cannot be sampled from a normal dis-
 Lognormal variables are ratio scale variables. Zero must repre- tribution because more than 5% of the values would be negative.
sent the absence of a quantity (eg, weight or concentration) or  With lognormal data, if the GeoSD remains consistent for all
an asymptotic limit that can be approached but never reached groups, expect the SD to vary proportionally with the group mean.
(eg, EC50 or Km).  Many data sets, especially with small sample sizes and small
 The geometric mean (GeoMean) of an ideal lognormal distri- GeoSD, pass both normality and lognormality tests, making
bution equals the median and is always smaller than the AMean. these tests often unreliable for distribution decisions.
The GeoMean of data randomly sampled from a lognormal
distribution provides the best estimate of the population
median. D. Analyzing lognormal data
 The GeoSD is a unitless factor that quantifies both spread and
asymmetry, and is always > 1.0.  Correctly recognizing lognormality can result in a smaller
 When the GeoSD is small (less than about 1.3), a lognormal dis- required sample size (with the same power). The statistical
tribution becomes nearly indistinguishable from a normal distri- power advantage of recognizing lognormality increases with
bution. Lognormal distributions with larger GeoSDs are skewed, larger GeoSD valuesdwhen GeoSD exceeds 2.0, analyzing data
but the skewness may not be obvious with small sample sizes. as if normal can reduce power by 50% or more.
 The product or ratio of 2 lognormal variables is also lognormal.  Outlier tests that assume a normal distribution are frequently
This partially explains why so many pharmacological parame- misleading when applied to untransformed lognormal data.
ters are lognormal.  Using the Welch’s t test on the log-transformed values is rec-
 When assessing lognormal data, treatment effects are best ommended. Using the Welch’s t test with untransformed
expressed as ratios rather than differences. This is because a lognormal data can lead to misleading results, as this approach
doubling, say, represents the same effect regardless of baseline. can result in an elevated type I error rate and decreased power.
28
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

 If a nonparametric approach is required, use the Brunner- To emphasize these points playfully, we conclude with a poem
Munzel test, which handles asymmetrical distributions better in the style of Dr Seuss.
than the Mann-Whitney test.
Oh the lognormal insights you will gain!
 Do not rely on AMean ± SD error bars for lognormal data, as they
can be misleading when the true variation is asymmetrical. In When your data’s askew,
some cases, the lower error bar can even descend to an
And you don’t know what to do.
impossible negative value.
 Express experimental effects on lognormal variables as ratios. A When your values spread wide,
75% decrease in EC50 represents the same effect size regardless of
All on the positive side.
the baseline EC50, just as a doubling in enzyme activity repre-
sents the same effect size regardless of the baseline activity. Avoid
For binding and clearance, EC50s galore,
the ambiguous terms fold change and percentage change.
For enzyme kinetics and so much more,
They multiply, multiply, that’s nature’s way!
E. Common misconceptions
Not adding like normal statistics would say.
Misconception: Lognormal distributions are rare special cases.
Reality: They are common. By welcoming lognormal, you’re thinking grows clear,
Misconception: Data should be considered normal until proven
Required sample size shrinks, no false outliers here.
lognormal.
Reality: For variables that must be positive, lognormal distribution Simple ratios illuminate the way,
is often more likely.
While absolute differences lead our insights astray.
Misconception: Lognormal distributions are always obviously
skewed.
When processes multiply rather than add,
Reality: With small GeoSD, they can be nearly symmetrical and
look very similar to normal distributions. Normal statistics can make results look bad.
Misconception: Effects should always be presented as absolute
Log-transform your data so analyses can thrive,
differences.
Reality: For lognormal variables, ratios are more meaningful Oh the insights you’ll gain, and the wisdom you’ll derive!
because they represent the same effect regardless of baseline.
Misconception: Log transformation is a form of p-hacking (invalid
data manipulation). Declaration of generative AI and AI-assisted technologies in
Reality: It is a valid statistical choice when justified by the nature of the writing process
the variable and should be prespecified in analysis plans.
Misconception: If a data set passes a normality test, the data During the preparation of this work the author(s) used [Link]
cannot be sampled from a lognormal distribution. 3 to enhance the manuscript’s clarity and conciseness, verify
Reality: Many data sets pass both normality and lognormality tests citation-reference consistency, and generate alternative versions of
(especially with small sample sizes). the concluding poem. After using this tool, the author(s) reviewed
Misconception: Standard outlier tests work for any distribution. and edited the content as needed and take(s) full responsibility for
Reality: These tests are invalid for untransformed lognormal data the content of the publication.
and can lead to inappropriate exclusion of legitimate high values.
Misconception: Reporting lognormal analyses requires that your Abbreviations
readers are facile with logarithms.
Reality: Results can be presented in original units using ratios and AMean, arithmetic mean; CI, confidence interval; GeoCV, geo-
GeoMeans without mentioning logarithms. metric coefficient of variation; GeoMean, geometric mean; GeoSD,
Misconception: CIs are always symmetrical. geometric standard deviation; GeoSEM, geometric standard error;
Reality: For lognormal data, the CIs of a GeoMean and the CI of a Kd, equilibrium dissociation constants; Km, Michaelis constant; Koff,
ratio of 2 GeoMeans are asymmetrical. dissociation rate constant; Kon, association rate constant; LOD, limit
of detection; LR, likelihood ratio; pEC50, negative logarithm (base
F. Perspective 10) of the EC50.

Lognormal distributions are not merely a statistical curi- Acknowledgments


ositydthey are fundamental to how biological and pharmacolog-
ical processes behave. When multiple factors influence a biological The authors thank Arthur Christopoulos, Martin Michel, and the
outcome through multiplication rather than addition, lognormal anonymous reviewers for helpful comments.
distributions naturally emerge.
Although the mathematical foundations may appear daunting, the Financial support
core concepts are intuitive, and the practical implications are pro-
found. Recognizing and properly accounting for lognormal distribu- Paul B.S. Clarke’s work was supported by the Canadian Institutes
tions results in smaller sample sizes, more reliable outlier detection, of Health Research [Grant 156045].
and more meaningful presentation of results. Most important,
assuming lognormal distributions shifts thinking about experimental Conflict of interest
effects from differences to ratios, and this often aligns better with
biological realityda doubling of enzyme activity (or a halving of drug Harvey J. Motulsky is the founder of GraphPad Software (creator
potency, say) represents the same effect regardless of baseline values. of Prism) and a minority shareholder of the company that now
29
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

owns it. Trajen Head is the Senior Product Manager for Prism, and is de Winter J (2016) A case against the default use of Welch’s t-test. Int Rev Soc
Psychol 30:92e101.
a minority shareholder of the company that owns GraphPad Soft-
De Lean A, Hancock AA, and Lefkowitz RJ (1982) Validation and statistical analysis
ware. Paul B.S. Clarke declares no conflict of interest. of a computer modeling method for quantitative analysis of radioligand binding
data for mixtures of pharmacological receptor subtypes. Mol Pharmacol 21:
5e16.
Data availability Delacre M, Lakens Danie €l, and Leys C (2017) Why psychologists should by default
use Welch’s t-test instead of Student’s t-test. Int Rev Soc Psychol 30:92e101.
 n AE, and Juarez-Colunga E (2018) The Wilcox-
The authors declare that all the data supporting the findings of Divine GW, Norton HJ, Baro
oneManneWhitney procedure fails as a test of medians. Am Stat 72:278e286.
this study are contained within the manuscript and Supplemental Elassaiss-Schaap J and Duisters K (2020) Variability in the log domain and limita-
Material. tions to its approximation by the normal distribution. CPT Pharmacometrics Syst
Pharmacol 9:245e257.
Fagerland MW (2012) T-tests, non-parametric tests, and large studiesda paradox of
Authorship contributions statistical practice? BMC Med Res Methodol 12:78.
Fagerland MW and Sandvik L (2009) The WilcoxoneManneWhitney test under
scrutiny. Stat Med 28:1487e1497.
Performed data analysis: Motulsky, Head.
Faul F, Erdfelder E, Lang AG, and Buchner A (2007) G*Power 3: a flexible statistical
Wrote or contributed to the writing of the manuscript: Motulsky, power analysis program for the social, behavioral, and biomedical sciences.
Head, Clarke. Behav Res Methods 39:175e191.
Fayers P (2011) Alphas, betas and skewy distributions: two ways of getting the
wrong answer. Adv Heal Sci Educ 16:291e296.
Supplemental material Fitzgerald JB, Schoeberl B, Nielsen UB, and Sorger PK (2006) Systems biology and
combination therapy in the quest for clinical efficacy. Nat Chem Biol 2:458e466.
Fleming PJ and Wallace J (1986) How not to lie with statistics: the correct way to
This article has supplemental material available at pharmrev. summarize benchmark results. Commun ACM 29:218e221.
[Link]. Fleming WW, Westfall DP, De la Lande IS, and Jellett LB (1972) Log-normal distri-
bution of equieffective doses of norepinephrine and acetylcholine in several
tissues. J Pharmacol Exp Ther 181:339e345.
References Flynn FV, Piper KAJ, Garcia-Webb P, McPherson K, and Healy MJR (1974) The fre-
quency distributions of commonly determined blood constituents in healthy
Ahad NA and Yahaya SSS (2014) Sensitivity analysis of Welch’s t-test. AIP Conf Proc blood donors. Clin Chim Acta 52:163e171.
1605:888e893. Gaddum JH (1945) Lognormal distributions. Nature 156:463e466.
Aitchison J and Brown JAC (1957) The lognormal distribution with special reference Galton F (1879) XII. The geometric mean in vital and social statistics. Proc R Soc Lond
to its uses in economics. J R Stat Soc Ser A (Gen) 120:481e482. 29:365e367.
Anderson TW and Darling DA (1954) A test of goodness of fit. J Am Stat Assoc 49: Gelman A and Loken E (2014) The statistical crisis in science. Am Sci 102:460.
765e769. Glantz S (2011), 7th ed Primer of Biostatistics, McGraw Hill, New York, NY.
Baldi B and Moore D (2017) Practice of Statistics in the Life Sciences, 4th ed, WH Glaser A (2018), 4th ed High-Yield Biostatistics, Epidemiology, and Public Health,
Freeman, New York, NY. Lippincott Williams & Wilkins, Philadelphia, PA.
Benzidia M and Lubrano M (2020) A Bayesian look at American academic wages: Grubbs FE (1969) Procedures for detecting outlying observations in samples.
from wage dispersion to wage compression. J Econ Inequal 18:213e238. Technometrics 11:1e21.
Black J and Leff P (1983) Operational models of pharmacological agonism. Proc Royal Haeckel R and Wosniok W (2010) Observed, unknown distributions of clinical
Society London B 220:141e162. chemical quantities should be considered to be log-normal: a proposal. Clin
Black JW, Leff P, Shankley NP, and Wood J (2010) An operational model of phar- Chem Lab Med 48:1393e1396.
macological agonism: the effect of E/[A] curve shape on agonist dissociation Hancock AA, Bush EN, Stanisic D, Kyncl JJ, and Lin CT (1988) Data normalization
constant estimation. Br J Pharm 160(Suppl 1):S54eS64. before statistical analysis: keeping the horse before the cart. Trends Pharmacol
Bland M (2015) An Introduction to Medical Statistics, 4th ed, Oxford University Press, Sci 9:29e32.
Oxford, UK. Havlicek LL and Peterson NL (1974) Robustness of the t test: a guide for researchers
Bodey AR and Michell AR (1996) Epidemiological study of blood pressure in do- on effect of violations of assumptions. Psychol Rep 34:1095e1114.
mestic dogs. J Small Anim Pr 37:116e125. Heath D (1967) Normal or log-normal: appropriate distributions. Nature 213:
Bright J, Aylott M, Bate S, Geys H, Jarvis P, Saul J, and Vonk R (2011) Recommen- 1159e1160.
dations on the statistical analysis of the Comet assay. Pharm Stat 10:485e493. Hyman BT, West HL, Rebeck GW, Buldyrev SV, Mantegna RN, Ukleja M, Havlin S, and
Brunner E and Munzel U (2000) The nonparametric Behrens-Fisher problem. Biom J Stanley HE (1995) Quantitative analysis of senile plaques in Alzheimer disease:
42:17e25. observation of log-normal size distribution and molecular epidemiology of
Burnham K and Anderson D (2002) Model Selection and Multimodel Inference: A differences associated with apolipoprotein E genotype and trisomy 21 (Down
Practical Information-Theoretic Approach, 2nd ed, Springer, New York, NY. syndrome). Proc Natl Acad Sci 92:3586e3590.
Buzsaki G and Mizuseki K (2014) The log-dynamic brain: how skewed distributions Irizarry R and Love M (2016), 1st ed Data Analysis for the Life Sciences with R,
affect network operations. Nat Rev Neurosci 15:264e278. Chapman and Hall/CRC, Boca Raton, FL.
Carlson LA (1960) Serum lipids in normal men. Acta Med Scand 167:377e397. Johnson NL, Kotz S, and Balakrishnan N (1994), 2nd ed Continuous Univariate
Cartwright AC (1991) International harmonization and consensus DIA meeting on Distributions, Wiley Interscience, Hoboken, NJ.
bioavailability and bioequivalence testing requirements and standards. Ther Julious SA (2004) Sample sizes for clinical trials with normal data. Stat Med 23:
Innov Regul Sci 25:471e482. 1921e1986.
Chargin W (2020) Why ratios want geometric means. Chargin blog. Viewed January Julious SA and Debarnot CAM (2000) Why are pharmacokinetic data summarized
11, 2025 from. [Link] by arithmetic means? J Biopharm Stat 10:55e71.
Christopoulos A (1998) Assessing the distribution of parameters in models of Karch JD (2021) Psychologists should use Brunner-Munzel’s instead of Mann-
ligandereceptor interaction: to log or not to log. Trends Pharmacol Sci 19: Whitney’s U test as the default nonparametric procedure. Adv Methods Pr
351e357. Psychol Sci 4:2515245921999602.
Christopoulos A and Kenakin T (2002) G protein-coupled receptor allosterism and Karch JD (2023) bmtest: a jamovi module for BrunnereMunzel’s testda robust
complexing. Pharmacol Rev 54:323e374. alternative to WilcoxoneManneWhitney’s test. Psych 5:386e395.
Cole TJ and Altman DG (2017) Statistics notes: what is a percentage difference? BMJ Kaumann AJ, Hall JA, Murray KJ, Wells FC, and Brown MJ (1989) A comparison of the
358:j3663. effects of adrenaline and noradrenaline on human heart: the role of 1- and 2-
Cox N (2010) Speaking Stata: the limits of sample skewness and kurtosis. Stata J 10: adrenoceptors in the stimulation of adenylate cyclase and contractile force. Eur
482e495. Hear J 10:29e37.
Crow E and Shimizu K (1988) Lognormal Distributions: Theory and Applications. Keene ON (1995) The log transformation is special. Stat Med 14:811e819.
Taylor and Francis, New York, NY. Kenakin T, Watson C, Muniz-Medina V, Christopoulos A, and Novick S (2012)
Curran-Everett D (2018) Explorations in statistics: the log transformation. Adv A simple method for quantifying functional selectivity and agonist bias. ACS
Physiol Educ 42:343e347. Chem Neurosci 3:193e203.
Custer EM, Finch CA, Sobel RE, and Zettner A (1995) Population norms for serum Kirby W (1974) Algebraic boundedness of sample statistics. Water Resour Res 10:
ferritin. J Lab Clin Med 126:88e94. 220e222.
D’Agostino RB, Belanger A, and D’Agostino RB Jr (1990) A suggestion for using Kirkwood T (1979) Geometric means and measures of dispersion. Biometrics 35:
powerful and informative tests of normality. Am Stat 44:316e321. 908e909.
Dancey C, Reidy J, and Rowe R (2012), 1st ed Statistics for the Health Sciences: A Koch AL (1966) The logarithm in biology 1. Mechanisms generating the log-normal
Non-Mathematical Introduction, SAGE Publications Ltd, London, UK. distribution exactly. J Theor Biol 12:276e290.
Daniels W and Cross C (2018), 11th ed Biostatistics: A Foundation for Analysis in the Koch AL (1969) The logarithm in biology II. Distributions simulating the log-normal.
Health Sciences, Wiley, Hoboken, NJ. J Theor Biol 23:251e268.

30
H.J. Motulsky, T. Head and P.B.S. Clarke Pharmacological Reviews 77 (2025) 100049

Koopmans LH, Owen DB, and Rosenblatt JI (1964) Confidence intervals for the co- Shapiro SS and Wilk MB (1965) An analysis of variance test for normality (complete
efficient of variation for the normal and log normal distributions. Biometrika 51: samples). Biometrika 52:591e611.
25. Shaw DJ and Dobson AP (1995) Patterns of macroparasite abundance and aggre-
Lacey LF, Keene ON, Pritchard JF, and Bye A (1997) Common noncompartmental gation in wildlife populations: a quantitative review. Parasitology 111(Suppl):
pharmacokinetic variables: are they normally or log-normally distributed? S111eS133.
J Biopharm Stat 7:171e178. Shen M, Russek-Cohen E, and Slud EV (2017) Checking distributional assumptions
Levasseur LM, Faessel H, Slocum HK, and Greco WR (1998) Implications for clinical for pharmacokinetic summary statistics based on simulations with compart-
pharmacodynamic studies of the statistical characterization of an in vitro mental models. J Biopharm Stat 27:756e772.
antiproliferation assay. J Pharmacokinet Biopharm 26:717e733. Shoari N and Dube  J (2018) Toward improved analysis of concentration data:
Lewontin R (1966) On the measurement of relative variability. Syst Zool 15: embracing nondetects. Environ Toxicol Chem 37:643e656.
141e142. Shrestha S, Ems-McClung SC, Hazelbaker MA, Yount AL, Shaw SL, and Walczak CE
Li WB, Ho€llriegl V, Roth P, and Oeh U (2006) Human biokinetics of strontium. Part I: (2023) Importin a/b promote Kif18B microtubule association and enhance
intestinal absorption rate and its impact on the dose coefficient of 90Sr after microtubule destabilization activity. Mol Biol Cell 34:ar30.
ingestion. Radiat Environ Biophys 45:115e124. Slavskii SA, Kuznetsov IA, Shashkova TI, Bazykin GA, Axenovich TI, Kondrashov FA,
Liang H, Li J, Di Y, Zhang A, and Zhu F (2015) Logarithmic transformation is essential and Aulchenko YS (2021) The limits of normal approximation for adult height.
for statistical analysis of fungicide EC50 values. J Phytopathol 163:456e464. Eur J Hum Genet 29:1082e1091.
Limpert E, Stahel W, and Abbt M (2001) Log-normal distributions across the sci- Small DS (2016) Let’s abolish fold higher and fold increase from our lexicon. Int J
ences: keys and clues. BioScience 51:341e352. Pharmacokinet 1:13e15.
Limpert E and Stahel WA (2011) Problems with using the normal distribution e and Stanforth PR, Jackson AS, Green JS, Gagnon J, Rankinen T, Despre s JP, Bouchard C,
ways to improve quality and efficiency of data analysis. PLoS One 6:e21403. Leon AS, Rao DC, Skinner JS, et al (2004) Generalized abdominal visceral fat
Limpert E and Stahel WA (2017) The log-normal distribution. Significance 14:8e9. prediction models for black and white adults aged 17e65 y: the HERITAGE
Mahajan S (2019) Don’t demean the geometric mean. Am J Phys 87:75e77. Family Study. Int J Obes 28:925e932.
Martinez MN and Bartholomew MJ (2017) What does it “mean”? A review of Steinijans VW, Eicke R, and Ahrens J (1982) Pharmacokinetics of theophylline in
interpreting and calculating different types of means and standard deviations. patients following short-term intravenous infusion. Eur J Clin Pharmacol 22:
Pharmaceutics 9:14. 417e422.
McAlister D (1879) XIII. The law of the geometric mean. Proc R Soc Lond 29: Stevens SS (1946) On the theory of scales of measurement. Science 103:677e680.
367e376. Stonehouse JM and Forrester GJ (1998) Robustness of the t and U tests under
Moser BK, Stevens GR, and Watts CL (1989) The two-sample t test versus sat- combined assumption violations. J Appl Stat 25:63e74.
terthwaite’s approximate f test. Commun Stat Theory Methods 18:3963e3975. Thelwall M (2016) Citation count distributions for large monodisciplinary journals.
Motulsky H (2017), 4th ed Intuitive Biostatistics, Oxford University Press, Oxford, J Inf 10:863e874.
UK. Thom H (1958) A note on the gamma distribution. Mon Weather Rev 86:
Norris N (1940) The standard errors of the geometric and harmonic means and 117e122.
their application to index numbers. Ann Math Statist 11:445e448. Verbosek T (2011) A comparison of parameters below the limit of detection in
Olsson U (2005) Confidence intervals for the mean of a log-normal distribution. geochemical analyses by substitution methods. RMZ M&G 5:393e404.
J Stat Educ 13. [Link] Vogel RM (2022) The geometric mean? Commun Stat Theory Methods 51:82e94.
Ott W (1995) Environmental Statistics and Data Analysis. CRC Press, Boca Raton, FL. Wahi M and Puzzullo J (2024), 2nd ed Biostatistics for Dummies, For Dummies,
Parkin TB (1993) Evaluation of statistical methods for determining differences be- Hoboken, NJ.
tween samples from lognormal populations. Agron J 85:747e753. Walker JS, Li X, and Buttrick PM (2010) Analyzing forceepCa curves. J Muscle Res Cell
Parkin TB and Robinson JA (1992) Analysis of lognormal data. Adv Soil Sci 20:193e235. Motil 31:59e69.
Patil PN (1993) Reactivity of human iris-sphincter to muscarinic drugs in vitro. Wertelecki W, Koerblein A, Ievtushok B, Zymak-Zakutnia N, Komov O, Kuznietsov I,
Naunyn Schmiedebergs Arch Pharmacol 347:568. Lapchenko S, and Sosyniuk Z (2016) Elevated congenital anomaly rates and
Portet S (2020) A primer on model selection using the Akaike information criterion. incorporated cesium-137 in the Polissia region of Ukraine. Birth Defects Res A
Infect Dis Model 5:111e128. Clin Mol Teratol 106:194e200.
Posten H, Yen H, and Owen D (1982) Robustness of the two-sample t-test under Wolfe R and Carlin JB (1999) Sample-size calculation for a log-transformed outcome
violations of the homogeneity of variance assumption. Commun Stat 11: measure. Control Clin Trials 20:547e554.
109e126. Wu J, Wong ACM, and Jiang G (2003) Likelihood-based confidence intervals for a
Poulsen TR, Jensen A, Haurum JS, and Andersen PS (2011) Limits for antibody af- log-normal mean. Stat Med 22:1849e1860.
finity maturation and repertoire diversification in hypervaccinated humans. Yule G and Kendall M (1950), 14th ed An Introduction to the Theory of Statistics,
J Immunol 187:4229e4235. Hafner, New York, NY.
Proost JH (2019) Calculation of the coefficient of variation of log-normally distrib- Zanotti-Fregonara P and Hindie  E (2011) Lognormal distribution of cellular uptake
uted parameter values. Clin Pharmacokinet 58:1101e1102. of radiopharmaceuticals: implications for biologic response in cancer treat-
Qazi S, DuMez D, and Uckun F (2007) Meta analysis of advanced cancer survival ment. J Nucl Med 52:501e503.
data using lognormal parametric fitting: a statistical method to identify effec- Zar J (2009), 5th ed Biostatistical Analysis, Pearson, Upper Saddle River, NJ.
tive treatment protocols. Curr Pharm Des 13:1533e1544. Zhang D, Fan C, Zhang J, and Zhang C (2009) Nonparametric methods for mea-
Rafi Z and Greenland S (2020) Semantic and cognitive tools to aid statistical sci- surements below detection limit. Stat Med 28:700e715.
ence: replace confidence and significance by compatibility and surprise. BMC Zhou X and Gau S (1997) Confidence intervals for the log-normal mean. Stat Med
Med Res Methodol 20:244. 16:783e790.
Ramsey PH (1980) Exact type I error rates for robustness of Student’s t test with Zhou X and Tu W (1999) Comparison of several independent population means
unequal variances. J Educ Stat 5:337e349. when their samples contain log-normal and possibly zero observations. Bio-
Rasch D, Kubinger KD, and Moder K (2011) The two-sample t test: pre-testing its metrics 55:645e651.
assumptions does not pay off. Stat Pap 52:219e231. Zhou XH, Gao S, and Hui SL (1997) Methods for comparing the means of two in-
Rochon J, Gondan M, and Kieser M (2012) To test or not to test: preliminary dependent log-normal samples. Biometrics 53:1129e1135.
assessment of normality when comparing two independent samples. BMC Med Zhu X, Finlay DB, Glass M, and Duffull SB (2019) An intact model for quantifying
Res Methodol 12:81. functional selectivity. Sci Rep 9:2557.
Rospars JP, Lansky P, Chaput M, and Duchamp-Viret P (2008) Competitive and Zimmerman DW (1987) Comparative power of Student t test and Mann-
noncompetitive odorant interactions in the early neural coding of odorant Whitney U test for unequal sample sizes and variances. J Exp Educ 55:
mixtures. J Neurosci 28:2659e2666. 171e174.
Royston P (1992) Estimation, reference ranges and goodness of fit for the three- Zimmerman DW (1996) Some properties of preliminary tests of equality of variances
parameter log-normal distribution. Stat Med 11:897e912. in the two-sample location problem. J Gen Psychol 123:217e231.
Ruxton GD (2006) The unequal variance t-test is an underused alternative to Stu- Zimmerman DW (2004) A note on preliminary tests of equality of variances. Br J
dent’s t-test and the ManneWhitney U test. Behav Ecol 17:688e690. Math Stat Psychol 57:173e181.
Shamsudheen I and Hennig C (2023) Should we test the model assumptions before Zimmerman DW and Zumbo BD (1993) Rank transformations and the power of the
running a model-based test? J Data Sci Stat Vis 3. [Link] Student t test and Welch t’ test for non-normal populations with unequal
jdssv.v3i3.73. variances. Can J Exp Psychol 47:523e539.

31

Common questions

Powered by AI

With lognormal data, absolute differences between GeoMeans rarely provide meaningful insight because the data's multiplicative nature means that ratios better represent the relationship between groups. As described in Wolfe and Carlin (1999), the treatment in the example nearly tripled the EC50, reflected by a ratio of 2.9. This demonstrates how ratios encapsulate the proportional effect of treatments on lognormal data, offering a clearer scientific narrative .

Skewness in lognormal distributions necessitates larger sample sizes to achieve statistical significance, as skewed data increase variability, making it harder to detect genuine effects. The document highlights that analyzing such data as normal can drastically amplify the required sample size, sometimes by as much as 300% .

Outliers can significantly skew results, leading to false-positive findings, especially in lognormal data. The document suggests transforming data using logarithms, such as by employing a lognormal t test, which helps mitigate the impact of outliers and increases the robustness of the test outcomes .

The lognormal Welch’s t test is preferred over the lognormal t test when GeoSDs differ because it maintains control of the type I error rate and maximizes statistical power in these circumstances. This is in contrast to the lognormal t test, which becomes less powerful when GeoSDs are unequal .

Using Welch’s t test on untransformed lognormal data can lead to similar power and type I error rates compared to the regular t test, making it less effective. This indicates that log transformation is necessary to truly benefit from Welch’s test's usual ability to handle unequal variances across groups .

When dealing with lognormal data and unequal sample sizes, one should consider using the Brunner-Munzel test instead of the Mann-Whitney, as the latter shows lower power and higher type I error rates under these conditions. This choice ensures better statistical power and maintains an appropriate type I error rate .

Using untransformed data in a t test on lognormal distributions can lead to inaccurate results because such data violate the assumption of normality required by the t test. Skewed distributions increase the required sample size and can elevate type I error rates and reduce statistical power, leading to unreliable conclusions .

Transforming lognormal data into log-transformed values approximates a normal distribution, which satisfies many statistical test assumptions, such as those of t tests. This transformation allows for more accurate statistical modeling by providing symmetrical distributions that align with the normality assumptions of the test .

One should opt for the Brunner-Munzel test instead of the Mann-Whitney when sample sizes are unequal, as it better controls type I errors and offers more statistical power under these experimental designs .

Ratios of GeoMeans accurately represent the multiplicative nature of lognormal data, providing easily interpretable effect sizes and reliable comparisons across groups. This translates the complex interactions into straightforward ratios, which align better with the inherent data patterns and scientific interpretations .

You might also like