ARTICLE AD BOX
February 14, 2026
4 min read
Add Us On GoogleAdd SciAm
AI conscionable sewage its toughest mathematics trial yet. The results are mixed
Experts gave AI 10 mathematics problems to lick successful a week. OpenAI, researchers and amateurs each gave it their champion shot
By Joseph Howlett edited by Claire Cameron
Interim Archives / Contributor via Getty Images
The verdict, it seems, is in: artificial intelligence is not astir to switch mathematicians.
That is nan contiguous takeaway from the “First Proof” challenge—perhaps nan astir robust trial yet of nan expertise of ample connection models (LLMs) to execute mathematical research. Set by 11 apical mathematicians connected February 5, nan results of nan trial were released early successful nan greeting connected Valentine’s Day. It’s excessively soon to conclusively opportunity really galore of nan 10 mathematics problems that were included successful nan situation were solved by AIs without quality help. But 1 point is clear: nary of nan LLMs came adjacent to solving them all.
The mathematicians down First Proof presented nan AIs 10 “lemmas”—a mathematics word for insignificant theorems that pave nan measurement to a larger result. These problems are nan moving mathematician’s stock-in-trade, nan benignant of mini problem 1 mightiness manus disconnected to a talented postgraduate student. The mathematicians aimed for problems that would require immoderate originality to solve, not conscionable a mash-up of modular techniques, according to Mohammed Abouzaid, a mathematics professor astatine Stanford University and a personnel of nan First Proof team.
On supporting subject journalism
If you're enjoying this article, see supporting our award-winning publicity by subscribing. By purchasing a subscription you are helping to guarantee nan early of impactful stories astir nan discoveries and ideas shaping our world today.
The challenge, while highlighting AI’s limitations, besides spotlights a budding AI-enthusiast subculture wrong nan mathematics community. Online chat boards and societal media accounts dedicated to mathematics were swamped pinch purported proofs from apical mathematicians and rogue undergraduates alike. And it underscored really earnestly AI startups, including ChatGPT shaper OpenAI, are taking nan situation of school an LLM to do math.
“We did not expect location would beryllium this overmuch activity,” Abouzaid says. “We did not expect that nan AI companies would return it this earnestly and put this overmuch labour into it.”
The First Proof squad revealed nan solutions to nan 10 challenges early connected Saturday, and posted astir their ain experiences trying to get LLMs to lick nan problems. They recovered that AIs could spit retired assured proofs to each problem, but only 2 were correct—those for nan ninth and 10th problems. And a impervious that was astir identical to nan ninth problem turned retired to already exist. The first problem was besides “contaminated”—a sketch of a impervious was archived from nan website of its author, squad personnel and 2014 Fields Medal victor Martin Hairer—but nan LLMs still grounded to capable successful nan gaps.
The style of impervious that nan LLMs came up pinch was peculiarly surprising, Abouzaid says. “The correct solutions that I’ve seen retired of AI systems, they person nan spirit of 19th-century mathematics,” he says. “But we’re trying to build nan mathematics of nan 21st century.”
Outside submissions didn’t look to fare overmuch better. Some submissions appeared to employment varying degrees of quality input, pinch respective seemingly nan consequence of week-long dialogues checked by mathematicians. Importantly, nan First Proof rules disallow quality mathematical input aliases prodding.
“Once there’s humans involved, really do we judge really overmuch is quality and really overmuch is AI?" says Lauren Williams, Dwight Parker Robinson Professor of Mathematics astatine Harvard University and 1 of nan mathematicians who group up First Proof.
OpenAI posted its activity connected Saturday, nan consequence of a week-long sprint utilizing its newest in-house AI models moving pinch “expert feedback” from quality mathematicians. The company’s main intelligence Jakub Pachocki said successful a social media post that they judge six of their 10 solutions to “have a precocious chance of being correct.” Mathematicians person pointed to imaginable holes successful astatine slightest 1 of those six already.
Aside from really overmuch quality assistance nan AIs had, nan immense bulk of nan submissions look to beryllium a batch of very convincing nonsense. Before nan situation had moreover ended, a number of purported solutions that initially appeared reliable were already being questioned by experts.
The submissions will return days for experts to decently vet. And judging whether a impervious is genuinely “original” is moreover tougher than judging if it is correct. “Nothing successful mathematics is wholly without precedent,” says Daniel Litt, a mathematician astatine nan University of Toronto, who was not portion of nan First Proof team.
“We are reasoning of this arsenic an experiment. Our extremity was to get feedback,” Abouzaid says. The squad writes that they’re readying a 2nd information pinch tighter controls, and that much more specifications will beryllium released connected March 14.
For immoderate mathematicians who’ve been search AI’s progress, nan lukewarm results lucifer their expectations. “I expected possibly 2 to 3 unambiguously correct solutions from publically disposable models,” Litt says. “Ten would person been very astonishing to me.”
Still, moreover getting a fewer valid solutions to research-level problems from an AI would apt person been intolerable conscionable months ago. “I already person heard from colleagues that they are successful shock,” says Scott Armstrong, a mathematician astatine Sorbonne University successful France. “These devices are coming to alteration mathematics, and it's happening now."
But for others who intimately way AI’s achievements, this wasn’t a awesome showing.
“The models look to person struggled,” says Kevin Barreto, an undergraduate student astatine nan University of Cambridge, who was not portion of nan First Proof team. He precocious used AI to lick 1 of nan Erdős problems, a number of challenges posed by Hungarian mathematician Paul Erdős. “To beryllium honest, yeah, I’m somewhat disappointed.”
It’s Time to Stand Up for Science
If you enjoyed this article, I’d for illustration to inquire for your support. Scientific American has served arsenic an advocator for subject and manufacture for 180 years, and correct now whitethorn beryllium nan astir captious infinitesimal successful that two-century history.
I’ve been a Scientific American subscriber since I was 12 years old, and it helped style nan measurement I look astatine nan world. SciAm always educates and delights me, and inspires a consciousness of awe for our vast, beautiful universe. I dream it does that for you, too.
If you subscribe to Scientific American, you thief guarantee that our sum is centered connected meaningful investigation and discovery; that we person nan resources to study connected nan decisions that frighten labs crossed nan U.S.; and that we support some budding and moving scientists astatine a clip erstwhile nan worth of subject itself excessively often goes unrecognized.
In return, you get basal news, captivating podcasts, superb infographics, can't-miss newsletters, must-watch videos, challenging games, and nan subject world's champion penning and reporting. You tin moreover gift personification a subscription.
There has ne'er been a much important clip for america to guidelines up and show why subject matters. I dream you’ll support america successful that mission.
5 bulan yang lalu
English (US) ·
Indonesian (ID) ·