AI Solves a Major Unsolved Math Problem. Not Everyone Is Happy

On 13 September, leading luminaries and up-and-coming talents in mathematics and computer science congregated at the annual Heidelberg Laureate Forum in Germany for a week of discussion, networking, and—naturally—wurst.

Having attended several of these events in the past, the mathematicians’ chatter is normally concerned with which famed researchers are attending, or what interesting problems they have been working on. But instead, every snippet of conversation caught in passing or any debate accidentally overheard was about how AI companies such as OpenAI, Anthropic, and Google are steamrollering their way through mathematics.

And there is good reason for these fervent discussions. Mathematics is the perfect testing ground for AI, involving step-by-step logical reasoning and answers that are automatically and objectively verifiable. This has led tech giants to develop their AI mathematical capabilities at a terrifying rate this year, leading to both new solutions to previously unsolved problems and a reckoning within the community as to what it means to do math in the age of AI.

Tech Giants Target the Millennium Problems

In short order, AI has gone from struggling with everyday research-level problems to solving a whole raft of teasers posed by prolific Hungarian mathematician Paul Erdős, then verifying the proof of Fermat’s Last Theorem, and—most recently and famously—OpenAI announcing they had solved the Navier–Stokes existence and smoothness problem. As attendee and young researcher Ailsa Robertson (University of Amsterdam, The Netherlands) put it: “AI and LLMs set the math community on fire over summer.”

One of seven extremely difficult Millennium Prize Problems posed in the year 2000 by the Clay Mathematics Institute (only one, the Poincaré conjecture, has been solved by humans so far), OpenAI’s claim to have solved the Navier–Stokes existence and smoothness problem represents a watershed moment for automated reasoning, if verified.

Fields Medalist Jacob Tsimerman (University of Toronto) was keen to recognize this achievement during a press conference at the Forum: “We’ve seen rapid increases in capabilities, even faster than many people, including myself, expected,” he said. “Is AI doing something new? I don’t know the details, but […] solving Navier–Stokes feels pretty definitive.”

Since this announcement, rumors have swirled about which of the remaining Millennium Problems will be next. One candidate is the Riemann hypothesis. In August, Anthropic quietly employed an unreleased version of Claude to tackle this problem, making important progress on a related problem but not on its main mission. And OpenAI is reportedly focusing on solving the Hodge conjecture. With these tech giants applying the full might of their most advanced unreleased models, many mathematicians see it as inevitable that at least some of these problems will be solved soon.

The tech giants are likely focusing on the Millennium Problems as a way of verifying the capabilities of their advanced AI systems, with the added bonus that solving these famous unsolved problems shows potential users and investors how powerful their technology is. “They’re really just solving these difficult mathematical problems as benchmarks, as some kind of PR stunt,” opined Fields Medalist Peter Scholze (University of Bonn, Germany), during a panel discussion at the event about AI in mathematical research.

Why OpenAI’s Navier–Stokes solution became a controversy

Alongside Scholze and Tsimerman at the panel discussion were thought-leading mathematicians Michael Harris (Columbia University, USA) and Geordie Williamson (University of Sydney, Australia). For them, these and many more problems researchers are now facing stem from the way in which tech giants fail to adhere to the norms and values deeply instilled in the mathematics community. And this was perfectly exemplified by OpenAI’s Navier–Stokes announcement.

“I think that OpenAI behaved extremely poorly, and that we should acknowledge that in the community,” said Williamson. Harris had a front-row seat to this alleged poor behavior, receiving what mathematician Tristan Buckmaster (New York University, USA)—who was making significant progress with Levent Alpöge of Anthropic on the Navier–Stokes existence and smoothness problem—claimed was correspondence between him and OpenAI that appeared coercive, censorious, and even threatening.

“I trusted Tristan’s account of this interaction,… and I guess I did my part in promoting his narrative,” Harris said. “But… on social media and traditional media, most reports are consistent with my takeaway; that is, they depict OpenAI as bullying and disrupting disciplinary norms.” OpenAI did not respond to requests for comment prior to publication.

Beyond OpenAI’s sportsmanship (or lack thereof), Williamson is concerned about what AI’s march across mathematics this summer does to the field, both for working mathematicians and for the knowledge that can become useful to the broader society.

“What we want as a mathematical community is understanding, but we measure this against unsolved problems, and the problem is that these two measurements are very, very quickly becoming uncorrelated,” he explained, referring to how AI solutions might give an answer but usually don’t develop new methodology that is understandable or useful. “So now,… we suddenly must re-evaluate things like how we assess people, who gets jobs, how do we educate people, etc. This is going to be a big challenge.”

Academia under pressure

While the tech giants tear up the rule book in their battle for supremacy, it is ordinary human mathematicians that are bearing the consequences. Unbridled access to powerful AI technologies is affecting how mathematicians work across the world.

Young mathematician Mita Ramabulana (University of Cape Town, South Africa) is a case in point. He said that two of the 10 advances in mathematics that OpenAI announced in August overlapped with his own work, but left him disappointed “because you spend some time thinking about these things, and these days you don’t know whether someone is just going to plug in a problem that you care about in some LLM and solve it.” He worries that unscrupulous researchers are using LLMs to scoop others or gain professional advantage.

For Robertson, currently studying for a PhD in quantum-safe cryptography, the problem is even more acute. She has seen all of her mathematics colleagues turn to using LLMs intensively in their research, with many maxing out their Pro subscriptions, and some even spending thousands of Euros on additional tokens: “And these are PhD students who don’t have thousands of Euros,” she added.

Robertson said that there are colleagues who feel coerced by the tech giants, with the likes of OpenAI announcing in July that it is giving away 100,000 free licenses to its frontier models for researchers in academia. And there are other colleagues who feel they simply have no choice: “If you don’t work at the rate at which you could work with LLMs then you will be behind your peers who will be applying for the same jobs as you”.

This is part of the reason why Robertson’s PhD is now a lot less mathematics-heavy and focused on the societal implications of transitioning to a quantum-safe ecosystem: “Because I don’t want to be in a career where you’re verifying LLM output”.

This article has been indexed from IEEE Spectrum

Read the original article: