Bischoff: Stet online hed
dek: Stet online dek
New
Math
How an AI proved dozens
of geometry theorems
from the International
Mathematical Olympiad
BY MANON BISCHOFF
© 2024 Scientific American
T HE INTERNATIONAL MATHEMATICAL OLYMPIAD (IMO) is probably the most presti-
gious competition for preuniversity students. Every year students from
around the world compete for its coveted bronze, silver and gold medals.
Soon artificial-intelligence programs could be competing with them, too.
In January a team led by Trieu H. Trinh lar to each other. By combining the prem-
of Google DeepMind and New York Uni- ises with the derived properties, the re-
versity unveiled a new AI program called searchers created a training data set con-
AlphaGeometry in the journal Nature. sisting of theorems and corresponding
The researchers reported that the pro- proofs. For example, a problem could in-
gram was able to solve 25 out of 30 geom- volve proving a certain characteristic of a
etry problems from past IMOs—a success triangle—say, that two of its angles were
with. This process can be repeated until the
AI and the deductive program reach the
desired conclusion. “The method sounds
plausible and in some ways similar to the
training of participants in the Internation-
al Mathematical Olympiad,” says Fields
Medalist Peter Scholze, who has won the
rate similar to that of human gold medal- equal. The corresponding solution would gold medal at the IMO three times.
ists. The AI also found a more general so- then consist of the steps that led the de- To test AlphaGeometry, the scientists
lution to a problem from the 2004 IMO ductive algorithm to it. selected 30 geometric problems that have
that had escaped the attention of experts. To solve problems at the level of an appeared in the IMO since 2000. The pro-
Over two days students competing in IMO, however, AlphaGeometry needed to gram previously used to solve geometric
the IMO must solve six problems from dif- go further. “The key missing piece is gen- problems, called Wu’s algorithm, man-
ferent mathematical domains. Some of erating new proof terms,” Trinh and his aged to solve only 10 correctly, and GPT-4
the problems are so complicated that even team wrote in their paper. For example, to failed on all of them, but AlphaGeometry
experts cannot solve them. They usually prove something about a triangle, you solved 25. According to the researchers,
have short, elegant solutions but require a might need to introduce new points and the AI outperformed most IMO partici-
lot of creativity, which makes them par- lines that weren’t mentioned in the prob- pants, who solved an average of 15.2 out of
ticularly interesting to AI researchers. lem—and that is something large language 30 problems. (Gold-medal winners solved
Translating a mathematical proof into models (LLMs) are well suited to do. an average of 25.9 problems correctly.)
a programming language that computers LLMs generate text by calculating the When the researchers looked through
know is a difficult task. There are formal probability of one word following another. the AI-generated proofs, they noticed that
programming languages specifically de- Trinh and his team were able to use their in the process of solving one problem, the
veloped for geometry, but they make little database to train AlphaGeometry on the- program hadn’t used all the information
use of methods from other areas of math- orems and proofs in a similar way. An provided—meaning that AlphaGeometry
ematics—so if a proof requires an inter- LLM does not learn the deductive steps set out on its own and found a solution to
mediate step that involves, say, complex involved in solving a problem; that work is a related but more general theorem. It was
numbers, programming languages spe- still done by other specialized algorithms. also apparent that complicated tasks—
cialized for geometry cannot be used. Instead the AI model concentrates on those in which IMO participants per-
To solve this problem, Trinh and his finding points, lines, and other useful formed poorly—generally required lon-
colleagues created a data set that doesn’t auxiliary objects. ger proofs from the AI. The machine, it
require the translation of human-gener- When AlphaGeometry is given a prob- seems, struggles with the same challenges
ated proofs into a formal language. They lem, a deductive algorithm first derives a as humans.
first had an algorithm generate a set of list of statements about it. If the statement AlphaGeometry can’t yet take part in
geometric “premises,” or starting points: to be proved is not included in that list, the the IMO, because geometry is only one
for example, a triangle with some of its AI gets involved. It might decide to add a third of the competition, but Trinh and his
measurements drawn in and additional fourth point X to a triangle ABC, for ex- colleagues have emphasized that their ap-
points marked along its sides. The re- ample, so that ABCX represents a paral- proach could be applied to other mathe-
searchers then used a deduc- lelogram—something that the matical subdisciplines, such as combina-
tive algorithm to infer further Manon Bischoff program learned to do from torics. Who knows—maybe in a few years
is a theoretical physicist
properties of the triangle, such and editor at Spektrum, previous training. In doing so, a nonhuman participant will take part in
as which angles matched and a partner publication the AI gives the deductive algo- the IMO for the first time. Maybe it will
which lines were perpendicu- of Scientific American. rithm new information to work even win gold.
A PR I L 2 02 4 SC I E N T I F IC A M ER IC A [Link] 41
© 2024 Scientific American
PHYSICS
STRANGE
METALS
Newly discovered materials bend the regular rules
of physics BY DOUGLAS NATELSON
ILLUSTRATION BY MARK ROSS
© 2024 Scientific American