What is A/B testing?
A/B testing (also known as split testing or bucket testing) is a method of
comparing two versions of a webpage or app against each other to determine which
one performs better. A/B testing is essentially an experiment where two or more
variants of a page are shown to users at random, and statistical analysis is used
to determine which variation performs better for a given conversion goal.
A/B testing (also known as bucket testing or split-run testing) is a user
experience research methodology. A/B tests consist of a randomized experiment with
two variants, A and B. It includes application of statistical hypothesis testing or
"two-sample hypothesis testing" as used in the field of statistics. A/B testing is
a way to compare two versions of a single variable, typically by testing a
subject's response to variant A against variant B, and determining which of the two
variants is more effective.
Like most fields, setting a date for the advent of a new method is difficult. The
first randomized double-blind trial, to assess the effectiveness of a homeopathic
drug, occurred in [Link] with advertising campaigns, which has been
compared to modern A/B testing, began in the early twentieth century.[17] The
advertising pioneer Claude Hopkins used promotional coupons to test the
effectiveness of his campaigns. However, this process, which Hopkins described in
his Scientific Advertising, did not incorporate concepts such as statistical
significance and the null hypothesis, which are used in statistical hypothesis
testing.[18] Modern statistical methods for assessing the significance of sample
data were developed separately in the same period. This work was done in 1908 by
William Sealy Gosset when he altered the Z-test to create Student's t-test.
With the growth of the internet, new ways to sample populations have become
available. Google engineers ran their first A/B test in the year 2000 in an attempt
to determine what the optimum number of results to display on its search engine
results page would be.[5] The first test was unsuccessful due to glitches that
resulted from slow loading times. Later A/B testing research would be more
advanced, but the foundation and underlying principles generally remain the same,
and in 2011, 11 years after Google's first test, Google ran over 7,000 different
A/B tests.
In 2012, a Microsoft employee working on the search engine Microsoft Bing created
an experiment to test different ways of displaying advertising headlines. Within
hours, the alternative format produced a revenue increase of 12% with no impact on
user-experience metrics.[4] Today, companies like Microsoft and Google each conduct
over 10,000 A/B tests annually.
Many companies now use the "designed experiment" approach to making marketing
decisions, with the expectation that relevant sample results can improve positive
conversion [Link] is an increasingly common practice as the tools and expertise
grow in this area.