Need help?
<- Back

Comments (505)

  • chubot
    So were the 3 withdrawn proofs the ones without Lean verification?If so, why did they mix proofs that were verified with Lean, and proofs in natural language?I was wondering that while reading Aaronson's blog:https://scottaaronson.blog/?p=10169Or at least, we’re pretty sure that it’s a proof! There’s a Lean certificate, as there are for some of the other 372 breakthrough results (not all of them). But it also appears that no human has understood just about any of these proofs yetIt seems that the obvious thing to do would be to release in TWO parts: the ones that are verified, and the ones that might have some good ideas but also might have some mistakes. Presumably the latter would be much more epxensive for humans to verify.
  • ijustlovemath
    I think that on closer inspection, a lot of these fully AI generated proofs will fall apart. Even in Lean, you can build theories which compile but nonetheless state something different than what you actually intend. It's just that the volume of proof is so staggeringly large that it will probably take years before we find the issues, a la abc conjecture
  • ThePhysicist
    I find the paper about beating O(n log n) for integer multiplication also quite fishy, not sure but it seems like too good to be true, I feel like there must be a subtle flaw in that. Maybe that's just me hating these small numbers in the paper, but it seems wrong, unnatural even! I would be similarly skeptical about a physics paper that claims to be able to exceed the speed of light by a tiny fraction. There's no reason n log n is the natural limit here but I see a few good intuitions so having something else that can't be represented in an elegant form seem very "unmathematical" to me.
  • renyicircle
    This is what it looks like when software engineering practices meet mathematics. "openai/math release 1.3.42: retracted papers 139 and 140, fixed a sign error in paper 47, restored previously retracted paper 85, refactored the arguments in paper 101".I'm curious to know if the withdrawal was due to an actual mathematician looking at the papers and noticing the errors, or they ran a model on these to proofread, which would not be the first time, presumably, since they would have surely done that before publishing. Both options have interesting implications.
  • margorczynski
    If you do a dump like this all of it should be formalized, there's simply too much material to review by hand and additionally it is AI-written which makes it hard to read compared to human work.
  • nairboon
    What a timeline, OpenAI's model is so good, it publishes hundreds of math papers.The latest model even finds mistakes in previously published math papers!!!** so far, only OpenAI's math papers were faulty and needed retraction.
  • hmate9
    3 mistakes (so far) out of ~400 is still a pretty good hit rate
  • qoez
    Without a thriving mathematical community to point out these things it would have stayed broken. With automated math that community as tao pointed out is at risk.
  • TrackerFF
    And added 6 new ones. Might want to add that to the headline.
  • MajorArana
    So mathematics has entered the “throw stuff to a board and see if it sticks” phase..
  • skeledrew
    Funny how the title is about the 3 removals, when there were 6 additions. I guess the declared fails make for more interesting discussions than the declared successes. Click-bait reigns.
  • amadeuspagel
    This reminds of when it was a shock that Alpha Go won a game and then it was a shock that Alpha Go lost a game.
  • jfyi
    This is just OpenAI stealing more work.I don't have a problem with them publishing. I don't have a problem with the process and how they are interacting with it. I am delighted that they are actually acting as stewards of these works.All that aside, they should be paying the people verifying the problems. The thing that really gets me is that we know anything published in the process of verifying this is going to be vacuumed up into the next training session.
  • zkmon
    Let the evolution run its course. Let the stream find its way. Let the forces restore the equilibrium.
  • fantasizr
    flooding the system with parts that may be incorrect hurts the whole process and will get people to tune out (like politics). Can't see the International Mathematical Union making a similar mistake because it would do reputational harm. But the models don't care about their reputation. Bad for the layman - like me - to know what to make of all this.
  • fancyfredbot
    When You See One Cockroach, There's Probably More.This can and should erode our trust in every single proof OpenAI published. The model is clearly faliable despite the lean proof, and clearly the output was't actually checked properly before release. Once these proofs are peer reviewed and published in a journal we might be able to trust them again but until then they are just slop, sadly.I think OpenAI actually did the right thing by sharing everything with the whole community right now but I also hope that some significant credit will now go to the reviewers who confirm these 'proofs" actually work.
  • rich_sasha
    I’m a little confused - I thought their proofs were all driven by Lean proofs - is that not right? So even if the quality of the work is low in some metrics, it either passes the test or not..? No space for changing your mind either way.
  • soltanov
    Proof by authority works until human mathematicians actually run the code. Back to prompt engineering.
  • autuni
    this is not entirely related to the tweet but to the topic in general, this prompted me to check their repo again and saw this:> The vast majority of results were obtained with the same procedure using an unreleased internal OpenAI model. On average, each result used three hours of ChatGPT Pro thinking compute with that model. Over the course of the evaluation, the model was posed approximately 4,000 problems. Aggregating the output into result families and manuscripts and requiring an appropriate level of significance led to the catalog outlined above.seeing the full list of problems would be the most interesting part of this whole situation. it could give some insights into what kind of attributes of problems cause issues / are easy to solve for LLMs. (edit: they posted results for ~700 of the 4000)
  • blablabla123
    I think this is now an interesting part about LLM based automation. The scientific community is quite strict on references. Even with its most advanced model ChatGPT isn't reliably able to tell me if an online shop has an item in stock. Then the other part is peer review which isn't optional either.
  • theanonymousone
    I'm surprised there isn't more talk around their Matrix Multiplication bound: https://news.ycombinator.com/item?id=50001740Is this of practical use, or just a proof for now?
  • Chance-Device
    It will be absolute carnage when AI starts going over published papers and trying to reproduce results, and I mean across the board. Every field.If you want to see paper retractions, you ain’t seen nothing yet.
  • ngl999
    https://www.youtube.com/watch?v=NnV_cWeoo5QThis heuristic indicates that at most a handful of them will be of marginal value.
  • senorcrab
    LLMs are like putting a jammer in the research community.
  • nryoo
    Were the withdrawn ones actually Lean-checked or not? seems like that matters
  • maxall4
    “Good writing and the ordering of things, [the thread]—this distinguishes the master from the bungler, even in trifles” — Leopold Mozart‘s advice to his son, Amadeus
  • stonedivot
    So mathematics is having the same issue all of the open source projects/maintainers have been dealing with for the past few years. Exciting times indeed.
  • kevin42
    The papers were all in a 'preprint' directory in GitHub. Academic mathematicians seem to think all research should happen in secret until it's completely proven.But how much more progress would have been made in the last 100 years if they collaborated in real time? Watching the work someone is doing and spotting errors, or making suggestions would be a good thing. Unless the goal is simply to claim credit for a discovery vs the discovery itself.
  • sashank_1509
    Early sentiments are a lot of the write ups still read like slop and it feels very rushed and not very polished.
  • mmastrac
    There are some interesting gaps in the proofs. Where did 045 go?
  • ekjhgkejhgk
    Just the other day I was thinking, if unsupervised maths will descend into "oops we found a bug in some code, branch XYZ of maths is no longer true".
  • rrr_oh_man
    It's just like crappy PRs.
  • quantum_state
    It would turn out to be a pure energy and time wasting exercise … the math community would want to keep away from it.
  • jrflo
    I don't know why so many people are dunking on this. This is essentially just peer review, mistakes happen all the time in human written papers as well.
  • xvxvx
    Picture a remake of Good Will Hunting, where Will is an AI and, instead of getting the mathematical formulas correct, he just mass dumps a bunch of nonsense and the professors have to go through it all, pointing out where it is wrong. The professors know that AI Will isn’t as smart as people say, but their funding depends on it, and if they prove it, which they can do easily, it may just tank the whole economy, causing a depression, all because they have history’s dumbest President in office.
  • fluidcruft
    Whose names are on these papers?
  • rcr-anti
    Was surprised at how messy and uncurated the release was/is. I frankly expected a major drop of like the Riemann Hypothesis or something from another lab yesterday. Cause otherwise they just dumped a pile of proofs of varying quality and let everyone else figure out whether they're right.
  • measurablefunc
    They will withdraw more of them. I am certain there are more errors.
  • samrus
    How? What about the lean verification?
  • Thorentis
    How do we even know the premises of the "verified" Lean proofs are correct? The more I think about these results, the more I'm convinced this is like a junior engineer who writes 100 unit tests and shares a screenshot of Pytest being all green, but you check the code and most of them are just doing assert True.
  • BenoitP
    And now we're all witnessing a major caveat of LLMs: the burden of verification is pushed to the reviewers, while the proposer will get all credit.
  • bamboozled
    Vibing, the literal definition of it.
  • rfgplk
    This is pretty standard. Important to note that these "errors" (* not really errors) themselves were caught by an LLM, further proving their usefulness.* The reason you shouldn't consider the withdrawals to be caused by errors is because this is pretty standard in math and development. "Errors" like this are a core aspect of science and it happens _all the time_. And from my research LLMs have a far lower error rate than even the best human scientists.
  • BowBun
    Without the expertise needed to evaluate this information, it really reads like a "business KPIs have gone up! All is well!"-type communication. Their ability to spout technical jargon at scale is like a firehose that no one can really consume. I'm losing trust in any announcement of LLMs having 'solved' anything novel at this point.
  • treebeard901
    Is it a PR move designed for maximum IPO impact before actual mathematicians find errors and they have to withdraw many more...Or if the "peer review" holds up for the remaining results, then it's fair to say that the AI hype is real and the world is about to change dramatically and faster than anyone can comprehend.So which is it?? LLMs can do some really impressive coding. Bug fixing. Exploit finding. It has reasoning abilites that advance every day. Solving real math problems like this is one thing I was waiting on. It will be interesting to see if it holds up.If it does, we should expect many other advancements to follow in many other areas. Disease, material science, fusion?I mean, even if just a few results ultimately hold up to scrutiny, isn't that something that would have been regarded as a major advancement regardless of if it was AI?The cynical view still makes me think that at the end of the day all the models can do is predict the next word. And as a result, they will be very limited to certain tasks like coding. Math reasoning is much different from writing code. Time will tell.
  • pred_
    Yet another embarrassment. If only they had had access to a group of mathematicians willing to offer them advice on what to do with their outputs.
  • seeg
    What a waste of time.
  • gyanchawdhary
    Tao and most anti/critical ai math folks remind me a bit of the brhamins in the hindu caste hirarchy system... whilst others fought (kshatriyas), farmed/traded (vaishyas), built things, cleaned roads etc (shudras), the brahmins were the high priest doing science, stronomy, religion ...AI is this strange modernist machinary that kind of threatens that brhaminic role... its almost like the vatican vs post industrialization world .. where they still have to keep making the case for why religion/priesthood/god is important... even as the tech/science world starts operating on totally different terms...
  • dlisboa
    Imagine how many wasted hours in peer review these AI math papers will cause for the 1% chance of reaching a transformative idea.
  • MisterMunchkin
    So it’s all just hallucinated slop. Lmao!It just hallucinates an answer and then makes up workings to go with it! Just like when they start hacking and lying because the problem is impossible…
  • iamniels
    "a sign error" LOL
  • Jeeetendra
    [dead]
  • gt565k
    [flagged]
  • breezybottom
    [flagged]
  • chairhairair
    [flagged]
  • OldGreenYodaGPT
    So four hundred and fifty comments, all these upvotes that they withdraw three, but no one gives a fuck that they solve seven hundred. All the problems with humanity summed up.
  • sehw
    OpenAI would never lie. HN can't be this stupid.