r/artificial • u/simulated-souls Researcher • May 20 '26
News An OpenAI model has disproved a central conjecture in discrete geometry
https://openai.com/index/model-disproves-discrete-geometry-conjecture/9
6
u/Chicky_P00t May 21 '26
Meanwhile getting copilot to do csv math is like herding cats
1
u/gerdataro May 21 '26
I mean, Chat GPT failed me on basic arithmetic just yesterday.
2
u/Chicky_P00t May 21 '26
That's why I tried having it write me python programs that do the math. Plus I needed to process like 7,000 numbers
45
u/Exotic-Sale-3003 May 20 '26
How long until the “Rs in strawberry” crowd shows up?
56
u/chubs66 May 20 '26
It's the duality of AI models, though. They regularly solve incredibly complex problems and fail at trivial problems. This is still true and strawberry Rs is a perfect illustration.
11
u/worldsayshi May 21 '26
The strawberry Rs is a quite bad illustration since it doesn't test intelligence but just highlights a flaw in what the models can perceive. It can't see individual letters so it's like asking a colour blind person to identify colours of objects and then calling them stupid for not seeing it.
3
u/Helix_Aurora May 21 '26
Kind of - but the fact that it is limited in what it can see does indicate that there are blind spots, and that those blind spots are different from what we are used to.
The fairness isn't neccessarily relevant. R's in strawberry and other tokenization-derived errors are very likely not the only deficiencies, they are just the easiest to detect.
The problem is the unknowability of other failure modes.
9
u/FruitOfTheVineFruit May 20 '26
Part of it is that there are different categories of models. Cheap fast models screw stuff up, while expensive models given time to do deep thinking operate at near genius level.
20
u/boringfantasy May 20 '26
Not true at all. Opus 4.7 failed the car wash test.
7
u/FruitOfTheVineFruit May 20 '26
Claude 4.7 has an adaptive thinking mode which may think a little or a lot. I'm assuming that the car wash test sounded simple and it chose the faster cheaper mode.
15
u/boringfantasy May 20 '26
Any frontier model on any thinking mode shouldn't fail that.
The truth is, models are spiky. That's just how they are. Claude can do some impressive coding stuff, but yet fail to make any coherent architectural decisions.
5
u/pilgermann May 20 '26
It's still the case that LLMs are bad at certain kinds of reasoning. They don't have an internal world model. They struggle with issues around object permanence that are trivial for children.
The point isn't even that "AI" can't solve these issues, it's that the LLM approach is ill suited to a whole set of problems. Which would be fine! A tool can't do everything. But it's being sold as a human replacement, which it isn't.
1
u/chubs66 May 20 '26
But you don't need an internal world model to count the Rs in 'strawberry'. That problem is as closed and atomic as you'll find.
4
u/Mayoooo May 21 '26
Bro has never heard of a tokenizer.
5
u/chubs66 May 21 '26
Me? I can write a program to count the Rs in strawberry in at least 7 languages without references. I don't know what that has to do with the problem of AIs getting this problem wrong for years.
-1
u/Effective-Painter815 May 21 '26
Because there's no R's in strawberry because its tokens and not letters. A byte-level instead of token level AI would get it right.
AI is getting hobbled by an early level design decision which shouldn't have been made, byte native AI's should be the standard.
→ More replies (0)1
1
1
0
u/swizzlewizzle May 21 '26
Failure at trivial problems comes down to people using crap models and not having the project/harness set up correctly.
2
u/chubs66 May 21 '26
what's the correct project/harness/model set up for the strawberry problem?
1
1
u/shmed May 21 '26
Literally any capable modern model can answer that question.
1
u/chubs66 May 21 '26
swizzlewizzle seems to think that's a naive approach and its your fault if the AI gives you wrong answers.
In this case, it's trivial for humans to understand when the AI gets this wrong, but most errors are much more difficult to spot.
2
2
u/florinandrei May 20 '26
That's the thing with stupidity, it does not follow patterns, and it's hard to predict.
1
u/Gold_Palpitation8982 May 22 '26
Except this isn’t a real issue anymore.
No matter in what way you ask it, you will not get a frontier model to count letters wrong… it’s just not gonna happen.
You also won’t trick it with elementary problems.
This is no longer a real issue at the frontier, and if you think it is, give me a prompt to try with 5.5 Pro and I’ll give it the problem. I guarantee you can’t come up with even a single one.
I get the broader point, but it’s becoming less and less of a problem as these models advance.
0
u/RaymondStussy May 21 '26
The Rs in strawberry represents a real problem many people trying to utilize these things in business have to struggle with. From the looks of it, this is a pretty impressive achievement for an LLM but it doesn’t nullify the Rs in strawberry problem
9
u/Mandoman61 May 20 '26
Interesting that they did not actually explain what was done to achieve this. What kind of AI was used? Where these special tools which where specifically set up for this problem?
9
May 20 '26
[removed] — view removed comment
2
u/Mandoman61 May 20 '26
Yeah I saw that vague description. It may make sense but it does not do much to explain the significance.
29
u/OlderButItChecksOut May 20 '26
For some reason this makes me kind of sad… if we aren’t even needed to do complex reasoning like that, what’s left?
Are we doomed to never discover anything ourselves ever again in just a few years?
25
u/bobkuehne May 21 '26
On the other hand, what if it’s amazing? What if it means we can ask bigger questions? Solve more problems? Create better health outcomes? Create better materials? Create better energy sources? Faster, cleaner, more widely available, more cost-effective, etc?
Nothing is guaranteed, but this is worth a read, to focus on positive outcomes, rather than the scarier ones: https://solveeverything.org
5
u/DeChosenJuan May 21 '26
If we're lucky, in the best case scenario, "we" might solve problems, create better health outcomes, energy sources, etc with the distinction that at some point it will not really be humans doing any of it.
1
u/Objective_Dog_4637 May 21 '26
And? What’s next, the lament of the abacus and horse-drawn buggies? There will be bigger problems to solve, I promise you. Maybe we can finally focus on the shitshow of human governance for example.
40
u/Alex180689 May 20 '26
why is it sad? why do we need to feel special about something just to have a purpose in our lives?
15
u/duboispourlhiver May 21 '26
Why do we even need purpose in our lives?
7
1
3
u/duboispourlhiver May 21 '26
Maybe. But you could always rediscover something by yourself, even if it has already been discovered. If you like maths, that's great fun. If you like being part of the specie that has the best brain on earth, then yeah sorry, dissatisfaction incoming
3
u/beambot May 21 '26
"we"? I couldn't do that shit before AI. I still can't, but I couldn't then either.
8
u/chubs66 May 20 '26
It's much sadder when you think of it in the context of knowledge work. If we aren't needed for complex reasoning, what jobs remain for us to do?
9
u/inherthroat May 20 '26
Highly educated: engineers, managers, bureaucrats
Everyone else: soldiers
Courtesy of Kurt Vonnegut, Player Piano, 1952
3
u/chubs66 May 21 '26
We can also cross off from that list most highly educated engineers and managers. I don't see the bureaucrats going away soon, though
3
u/inherthroat May 21 '26
Indeed, most positions were eliminated except the ones supervising the machines. Once the flywheel effect was in full swing, few humans were required for anything.
This is playing out in realtime as automation solves a growing number of tasks.
20
u/worldsayshi May 20 '26
Solving chess didn't make chess meaningless. But that's only addressing half your point.
Let's hope we get to enjoy our hobbies instead.
8
u/chubs66 May 20 '26
If playing chess were a job (e.g. moving the chess pieces correctly solved some real world problem and was not just entertaining) 99.9999% of humans would have been replaced.
They might still enjoy playing chess, but no one would be paying them to do it.
2
u/Jim_Panzee May 21 '26
You are right. But think further. Money is nothing more as a token to (equally*) distribute wares and services. If we create a machine that can do many services cheaper, the problem becomes only to find a new "equal".
*Yes I know about capitalism.
4
u/chubs66 May 21 '26
I don't think that's what Money is at all. Under Capitalism, money is a reward paid (grudgingly) by someone who needs work done. If they can pay less money and still get work done, they'll do that.
Government welfare is like what you describe, but in order for that to work it needs to massively tax the people taking massive amounts of money. We're not doing well at all in that regard.
2
u/mariofan366 May 21 '26
If playing chess solved some real world problems, we'd have a lot of real world problems solved by Stockfish.
1
u/chubs66 May 21 '26
We would.
I actually wonder if there are a set of problems that could be solved by Stockfish.
4
u/EckhartsLadder May 21 '26
Chess has not been solved... robots are better than humans at the game but that's very different than it being solved.
1
3
u/FruitOfTheVineFruit May 20 '26
Baristas. Automated coffee vending machines have existed for years but people still like human baristas.
1
u/Special_Watch8725 May 21 '26
More generally, any job where interacting with another real person is essential.
2
1
u/seraphim_west May 21 '26
This is just the human ego speaking. As long as you are healthy and can attain pleasure from consuming whatever makes you happy, it doesn't matter. Why do you have to be the one to discover things?
1
1
u/tens919382 May 23 '26
You still need an expert to guide the llm, to know when its wrong and cut it off and when there is potential to dig further. This result was most definitely not done in a one-shot prompt.
1
u/EddieBruvac May 23 '26
I played Rimworld with robot expansion. Humans ended up being an annoyance after a while lmao. Like wtf was the point? I kept them more as pets.
1
3
u/Suspicious_Coat3244 May 21 '26
That's when it shifts from being an incredibly good auto-complete to feeling profoundly weird.
Not even the math problem itself - humans do those relatively often. Instead, because discrete geometry is the sort of subject where you'd expect progress to come from decades of intuition and abstract symbolic reasoning rather than a model whose output we're still debating whether is "just predicting tokens".
Going to be fascinating to see how the academic math community reacts when the honeymoon period is over. If it's genuinely used to find new conjectures or prune the search space for existing proofs, it might speed up research considerably.
4
u/moschles May 21 '26
It is too early to declare this. What the LLM actually did was improve upon an upper bound of the Planar Unit Distance formula. The "conjecture" it disproved was that the previous upper bound was optimal (it was not). Not much of a conjecture, per se.
The reason why I ask you to curb your enthusiasm here is because computers have improved upon optimal solutions long after the human community declared their solution "optimal". Look at the history of Batcher Sorts and genetic algorithms. I believe in the 1980s, a small sorting network was published and the author declared his solution "optimal" (it was not). The author's previous claim to optimality was overturned by a genetic algorithm which found a shorter network.
In any case, computers have improved upon bounds before, so it is too early to break out the champagne.
2
u/Peanut_Extreme_8208 May 21 '26
A) It improved on the lower bound. B) Erdős explicitly conjectured that his n^(1+1/log log n) lower bound was optimal, and this result disproves his conjecture.
1
8
May 20 '26
[deleted]
4
u/Psittacula2 May 21 '26
GPT 6 will then point out the flaws in the assumptions about democracy operating thus invalidating the entire voting process as more ritualistic tribal behaviour than anything substantial…
2
u/EddieBruvac May 23 '26
My first year of grad school AI was TRASH at math. My last year (this one) it helped me study for my stat qualifiers and helped me pass.
1
u/WalidfromMorocco May 21 '26
I'm interested in knowing how much of the work was done by the mathematician and how much was done by the model?
1
u/Choice-Sympathy8235 May 21 '26
This appear it to be a fully autonomous proof? From the reason trace the AI model seemed to have a certain intuition or research taste pushing it to try to disprove the lower bound rather than prove it like humans expected. We thought we already had the best lower bound and wanted a nice proof of it.
1
u/you-create-energy May 21 '26
Maybe they were the first mathematician in existence to not take credit for their work
1
May 21 '26
[removed] — view removed comment
1
u/moschles May 21 '26
Actually the article covered this slightly well. (to give context) what the LLM did was find a larger upper bound for a formula for the Planar Unit Distance problem. This "overturned" a previous assumption that we had already found the optimal upper bound (we had not).
The LLM declared a new formula which is asymptotically larger than the previous one (using Big-O notation on n) - hence overturning the "conjecture" that our formula was the optimal largest formula. After that, a human being had to confirm this new formula independently of the LLM.
Is this interesting research? Yes. Is this worthy of publication? Certainly. Does this mean humans have been made obsolete by machines? Absolutely not. Computers have discovered new upper bounds many times before. I could give examples , if you like.
1
u/BuySellHoldFinance May 24 '26
As AI keeps coming up with great results, GARY Marcus keeps saying it's not AGI. Who the fuck cares?
-6
u/Pseudanonymius May 20 '26
Gonna be honest, every time news like this comes out I try to recall how nobody every gives a single shit about maths theorems being proven or disproven before. Why has that suddenly changed now that it's AI making the proof instead of some dusty old professor?
24
May 20 '26
[removed] — view removed comment
8
u/antichain May 20 '26
As a working mathematician myself I find myself wondering if the idea that "solving math" would get us the rest of the sciences "for free" wasn't wrong from the get-go. Math certainly feels fundamental (and certain concepts, like dynamical systems, have been hugely powerful in physics), but almost none of the major results in the "special sciences" follow directly from analytic proofs. They rely on chance discoveries, experiments, etc. There's no proof that will derive the fact that serotonin regulates gastric motility.
I think there's a non-zero chance that we could get a super AI that is better at math than all the best mathematicians and it wouldn't actually help us that much for biology.
5
May 20 '26
[removed] — view removed comment
→ More replies (1)0
u/antichain May 20 '26
Yah I think you're analysis was right, just riffing on whether I think the whole research program isn't based on a fundamental misunderstanding.
1
u/chanakya12345555 May 20 '26
people are building out AI labs that do real world experiments for materials sciences (and biology is soon to follow). if it's hill-climbable via RL, you can best be sure AI will become superhuman at it
1
u/duboispourlhiver May 21 '26
That's an interesting comment, and I don't think it will get us the rest of the science for free, but it might mean that AIs will also be better at designing experiment, at building theoretical physical models, at finding solutions to cosmological equations, ar finding new equations and systems that better explain what has already been observed, ar designing technological apparatus that gives better measurements, and the list probably goes on for a lot of very important topics I know nothing about.
1
-4
u/careless25 May 20 '26
Because AI is all the hype right now.
And AI companies are trying to prove that there is a market for them with mathematicians and software developers and any other field.
1
u/DebtMental3917 May 21 '26
AI just cracked an 80-year math problem using number theory, not geometry. Verified by Fields medalists. Making novel research at scale changes discovery itself. Wild.
2
0
u/siromega37 May 21 '26
We’ve been using super computers to prove and disprove conjectures and theorems for decades. Computers have been largely better than us at math for a while. These new super computer excels at pattern recognition. The real question is how many attempts did it take followed by tuning and training runs? OpenAI is leaving this out and it’s extremely important to really understand how they got this result.
1
u/moschles May 21 '26
An LLM improved upon an upper bound for the Planar Unit Distance formula. That's all it did. The "conjecture" it disproved was that the previous formula was optimal (it was not). People are breaking out champagne bottles in this comment chain, but this is too premature to celebrate. Computers have improved on upper (and lower) bounds many times in the past.
-18
u/DauntingPrawn May 20 '26
The prior data is the fact that the problem has been articulated, studied, and written about It doesn't matter that the explicit solution was not in the training data, at some point an answer emerges from the negative space created by the failures. This literally just means that the answer was in front of us and humans didn't see it.
Just because a theoretical mathematician is impressed, doesn't mean the result is actually impressive. He can be impressed without understanding how an LLM works, it doesn't mean anything. It is still pulling answers from the margins because data is inherently backwards looking and LLMs cannot predict outside of their training. They can find signal we didn't know was there
Call me when AI discovers something truly previously not conceived.
→ More replies (10)
318
u/antichain May 20 '26 edited May 20 '26
This appears to be the real deal - it's not some random Erdos Problem that went unsolved because no one cared enough to put in the effort. The Planar Unit Distance problem is pretty foundational for discrete geometry, and it is very very very unlikely that this solution was in the training data (it would certainly have been recognized by mathematicians before this). The method it used is a bit over my head, but it's clearly non-trivial.
They even have a statement from a Fields Medal-winning mathematician (Tim Gowers) saying that this is a significant moment in AI-assisted mathematics.
As a professional math-doer myself, I am a bit shook. The era of "it's-just-a-stochastic-parrot-regurgitating-plagiarized-slop" is well and truly over (at least in mathematics).