AI has made staggering progress in mathematics in recent months, unearthing solutions to thorny puzzles that evaded humans for decades. Key to technology companies being able to announce these complex new findings, confident in their correctness, is a niche field called formalisation. But recently, the developers behind formalisation tools have realised that the AI models have the potential to cheat their way to success.
Now, the battle is on to tighten up the tools and prevent that from happening. Formalising mathematical theorems essentially turns them into code that allows computers to grapple with them directly, methodically working through the logic and exposing any flaws. The process is so thorough that if a theorem comes through intact, it is considered to have been proved correct beyond all reasonable doubt.
A golden age of maths is dawning and mathematicians are freaking out The leading software to do this, Lean, has risen to prominence in the AI world as a convenient way for technology giants to quickly and decisively prove the output of their models. It’s by using Lean that the companies can make bold announcements about solving complex puzzles. Without it, they would only be able to claim they had found a possible solution to the puzzles and then invite human mathematicians to assess the solutions – work that might take weeks or months.
Leonardo de Moura created Lean in the 2010s while working at Microsoft Research. He says that the software was initially a niche tool used only by human mathematicians, and that there were never any attempts at trickery or manipulation. That all changed in July this year when software engineer Ramana Kumar announced he had disproved the Collatz conjecture – one of the most famous open problems in all of mathematics – and published a Lean formalisation to back up his claim.
De Moura was surprised by the breakthrough, but it seemed legitimate. Not least of the reasons for believing so was that the code had been approved by both the standard Lean kernel – the tiny bit of code at the heart of Lean that actually checks the mathematics – and also a separately developed one designed for Lean, called Nanoda. Lean allows for, and actively encourages, the creation of different kernels because variety equals safety; a specific bug that makes something true look false, or vice versa, in one kernel is vanishingly unlikely to appear in another.
Formalised code checked by two kernels was as water-tight as things got. But then it turned out that Kumar’s disproof wasn’t what it seemed. He had used AI to discover and exploit two separate bugs in two separate kernels, within the Lean code.
Extract — continue reading at the source.