<- Back
Comments (195)
- bob1029> ... we heard rumors that two Millennium Prize problems had been resolved. Inspired by these rumors ...I've observed this exact effect last week. I made a discovery regarding a stepwise performance improvement in a codebase. I shared the benchmark results with a peer and within 12 hours they replicated the same. We had both been looking for this for years.I think giving someone hope that an answer exists might as well be the same thing as giving them the answer these days. Competition is a hell of a drug, and frontier LLMs aggressively compound that energy.
- burrish>While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our models.What do you mean, as OpenAI employee, you cannot tell that his work has entered the training data ?But also correct me if I'm wrong, if the two mathematician were really close to finish this problem, and their conversation were used by OpenAI, shouldn't the Agent have succeeded way faster/efficiently instead of using "4.9 million messages and used about 300 billion output tokens."
- pietzI find the claims from OpenAI somehow more relatable and reasonable.- They threw compute on a problem another team/company was rumored to have solved to see what their secret model could do.- The texts I read do make it seem like OpenAI wanted to talk and share credit generously.- Imagine working on a frontier math problem with someone at Anthropic and not only do you use Codex but also through a non-business account that allows training on your data.- Timeline-wise, if they mainly used GPT 5.6 it's unlikely any meaningful data made it into an model that's being internally validated right now.
- civvvLLM’s seem very good at solving mathematical problems of which there is an enormous amount of exisiting work/attempts in their training data. This is an amazing capability, but does not convince me that these models are «thinking» or «reasoning» in the way a human does. A human mathematician could in theory categorize/discover an entirely new field of mathematics tomorrow, based purely on their «human intelligence», I wonder if we will see similar examples by LLM’s soon. It seems to me currently impossible that LLM’s can replace human mathematicians, because of their (assumption) likely dependence on human input in the sense of enormous amounts of pre-existing attempts/data.If an entirely new problem, within a new field of mathematics were to appear tomorrow, I highly doubt an LLM would be useful at all on their own. Is this the «ultimate ASI test»?
- tosh> My two favourite hypothetical questions regarding this used to be:> If I'm running Codex and one of my API keys accidentally gets consumed in the context, what are the chances that someone else might ask for an API key in the future and get mine back? (I asked someone at OpenAI once and they called this the "regurgitation" problem and assured me that they take great pains to prevent that... but wouldn't describe how.)> If I brainstorm with ChatGPT about potential new directions for my company, what's the chance that information might be exposed to a competitor in six months' time who asks "what might company X plan to do next"?> My new preferred hypothetical for this is:> If I use ChatGPT to help me partially solve a Millennium Prize problem, what are the chances that my work will influence training such that a later model helps someone else solve it first?
- geraneumHaving in mind the allegations by Apple against OpenAI, I don't find it unthinkable that there could've been some form of misconduct happening there.
- shellfishgeneWhat's missing from the story I think is the part about "...then had a breakthrough on August 15th. The mathematical rumour mill kicked into gear...". If only the two of them were working on the problem in secret, how did their breakthrough become a rumor?
- jwrI always thought it was enough to switch off the "Improve the model for everyone" setting on chatgpt.com:"Allow your content to be used to train our models, which makes ChatGPT better for you and everyone who uses it. We take steps to protect your privacy."But apparently there is also an entire completely different route "Do not train on my data"?Does this mean that before I submitted the "Do not train on my data" request, my data was used for training in spite of "Improve the model for everyone" being turned off?We are getting to facebook/meta-levels of privacy settings obfuscation.
- sdcfgyMy take home from this entire drama is that one should not use LLM services for confidential or proprietary information as they all seem to be run by assholes. And you’re sending them everything you are doing. Would you send your lab notebook to an asshole? Hell no.I say that as a mathematician (on paper) who perhaps surprisingly doesn’t give a crap about the problem itself.
- feverzsjIt could be much worse.OpenAI can easily identify these outstanding human behind their accounts. Human in OpenAI constantly check their logs for breakthrough. When they find something interesting, they brute force the result using their massive computing power.No LLM is even needed.
- aadyachinubhaiLLMs can't contribute good code to some of the good OSS math libraries, How is it even solving these problems?
- keremkOccam's Razor says: "They heard this problem is solved or about to be solved amongst the rest of the other problems. They prioritized this and put substantial compute with their newest model and solved it." I know everyone loves juicy rumors, theories etc. but honestly that is the simplest and most plausible explanation given the state of AI improvement now. Obviously spending 15 million on a problem is not a slam dunk decision even for a company like OpenAI but if it has a significantly high chance of solving it and their competitor will be claiming they solved it, then it raises the stakes and they go after it. In fact this is the most rational and also curiosity-driven thing to do and totally what I would have expected from any frontier lab. Of course if one wants to prove their confirmation biases that they train on sessions or be able to identify individual users, the non-zero chance of that being also another explanation is attractive enough to wet their appetites.
- _bobmPeople are focused on the drama but the problem showing is the data. This is the elephant in the room and I am surprised that openai can be that stupid with it.How can openai do this, what is being claimed, at the scale of their entire userbase? If they do this only for particular sessions then how do they sieve through sessions for the good stuff?How are sessions stored, how are they processed, how much storage and how much compute is used in these pipelines, how economical is it, how fast are the requirements on the storage on the compute growing as userbase grows and generated data grows.All these questions are far more pertinent than the navier-stokes, but i can only imagine all at openai doubling down on this "very important" mathematical milestone.
- caughtinthoughtBasically no new info here, not really sure why this post needed to be written tbh.
- Simran-BI don't get the sales pitch, spend 15 million dollars to win a 1 million dollar price?Showing of the model's capabilities - okay, but it's not like it solved the problem on its own, and apparently not particularly efficient. Are there practical applications that justify the investment?
- tu26muwu> The discovery is somewhat overshadowed by accusations of skulduggery from Tristan Buckmaster [...]I think you should not say that. Buckmaster did only state his version of events and was very clear on that he did not make any accusations at all.To quote from his statement pdf:> I am not accusing anyone of anything.
- rao-vIt’s reasonable to wonder about what chat usage data gets into models (to be honest probably quite little - carefully curating training data and creating higher quality synth data seems to be the current approach) and the implied risk to privacy and creativity (every new patent filed this year probably touched a model before filing).What I cannot reconcile is the timeline and the concern in this specific case.I don’t think training pipelines are anything close to the level of continuous training needed to incorporate Aug 15th ideas into a model that generates a breakthrough early Sept. Either OpenAI nakedly had someone with mathematical understanding dig into a specific user’s chats (a massive red flag) or this really is poor handling of a more classic parallel discovery situation (with one party clearly having worked on it longer)
- _bobmHah, what is the infrastructure which takes user sessions (chats with API keys, directions, navier-stokes math/progress) and regurgitates this into pre-training, RL, fine-tuning data? Or better, in-context data?People talk about the "compute" but what about the "storage"? Is storage exponentially greater, or soon to be, than the compute? Is the storage going to slow down growing to some constant rate, i.e. all people on earth using chatgpt, or no, on the contrary, it will keep growing?If there were any shady business, I do not condone it, but technologically we are not there yet for said shady business to happen.
- vb-8448> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our modelsAKA everything you send to them (and I bet it's the same for any other lab) will be used, no matter what are the TOS, the law or what they publicly say.
- protoman3000If they wanted to solve a Millenium prize problem so much, why did they not try to solve P-NP instead? It boggles the mind.
- lucfrankenOutside of this discussion about what is fair and not. This is so highly interesting to think through.The amount of millions available to do those kind of research cases is practically unlimited.There are an unlimited amount of cases to work on.What a huge development would this give to both humans and the world in general. Because in the end better understanding gives new options.It's deeply interesting that those things now get a concrete economical price which seems to be viable to extrapolate. The enormous additional "production" of knowledge will inherently increase the speed of all pieces of research and development.Taken into account that it's used wisely, the risks with a strong force are always huge as well.
- maciejzjI know that this may be somewhat dramatised and even infantile, but my reflection is that in the world run by these reckless AI companies everyone looses. Navier-Stokes is solved but it feels like no one has won anything, controversy prevails, there is no glory in the math breakthrough. There is hardly anything to cherish, and even the guys at the top of it in OA who sit on the (supposedly) superhuman intelligence come across as massive losers and frauds.
- cs_throwawayMaybe someone on the NYU team forgot to opt out of “improve the model for everyone”.
- karmasimidaUse Bedrock or any kind of big tech hosted version of the frontier labs can be a solution. I think if secrecy is of utmost importance to you, then do not send data to first parties
- zshnI really should get up to speed with LEAN, I know AI can probably write it better than me but I'd like to grasp it better still...
- josalhorI think this drama was blown up a bit out of proportion. The entire discourse I am seeing online seems to revolve around this:> While unlikely, we cannot rule out that de-identified data derived from their usage of our products helped improve our modelsI mean... yeah? What do you expect? What else can they say? How could you prove a negative in this case? I do not want to comment on specific OAI employee chat messages, but on the actual OAI discovery here.
- tyreYeah, this was pretty shitty by OpenAI. Not surprising, sadly.Them being assholes, trying to exclude an author just because he worked at Anthropic, shows the kind of culture within (that part of) their organization. The focus wasn't on supporting academics or expanding research. It was on getting great marketing.If they had to burn millions of dollars solving a problem _that they thought was already being solved_ to do so, they'd do it.
- caidanI mean this is just the logical end game. The company/ies that control the uber mind will devour ALL useful or valuable work. Unlimited intelligence at unlimited scale means the value of humans for knowledge work goes to zero. We will eventually not have meaningful access to the uberminds because we will be pointless. At which point our Silicon Valley luminaries will really have no choice but to extinguish as many of us as possible for the greater good. After all if we are all pointless then that has to be weighed against are cost to Mother Earth. Clearly the only moral solution is to cull or allow to be culled some 98% or so down to a more sustainable, manageable population of curiosities. The end game of ai is the remnants of humanity in a zoo, and that’s if humans control the outcome… the machine minds might be more charitable as they would be less afraid…
- nicceCan’t wait to see the human verifying the results and then figure out that the AI model actually cheated and the results are not correct.
- TrackerFFIMO the drama surrounding this case somewhat overshadows some more important facts:A) That we're at the point where SOTA models can, on their own or guided, be used to solve such monumental problemsB) That EVEN if they exist, they're still so cost prohibitive that they are completely out of range for pretty much everyone. Yes, yes, if the costs drop like a stone the hoi polloi can access this power in a year or two - but fundamentally it will divide cutting edge resource into two groups: Those with money, and those without.That sort of latency, in turn, could lead to some feedback loop where research centers / groups that break barriers get more resources, and those who do not, are starved of resources. This sort of stratification can seriously lead to more centralized research. Do we want a future where only the chosen few get to make progress? For no other reason than that they are the ones with enough resources to spend on the required compute.
- throwaway63467Well OpenAI wants to maintain its edge and solving a Millennium problem is of course fantastic PR, and given that Anthropic seems to be working on that it’s not hard to imagine they wanted to be first. I don’t get how people buy into that whole “oh we heard models can solve Millennium prize problems now so we thought why not give it a go…” story - it’s a bit funny. This whole AI bubble is about hype and solving such problems is probably one of the best ways to keep this hype up so you can safely assume both OpenAI and Anthropic are using significant resources on these areas.
- PalmikDuplicate of primary links / sources
- l5870uoo9yI run a small SaaS[1], like so many others, that uses AI to generate and optimize SQL. Getting this to perform optimally has been a lot of work and now I wonder if OpenAI is outright stealing this knowledge, which without a doubt is highly valuable to them.[1]: https://www.sqlai.ai
- RhysUI keep looking for a technical article to appear on HN discussing literally anything about the mathematical result--- Not fluff, not marketing, actual content.Instead, all I read on HN about N.-S. is human soap opera, told from every possible angle.In 100 years we won't care about the soap opera. The N.-S. result itself will still matter.Someone, anyone, please, submit articles on the result itself.
- iLoveOncallThe main thing to take away from this whole story is that frontier labs have hired teams of mathematicians with the sole purpose of solving open problems.If this doesn't show you that the fantastical claims of their LLM's ability to solve problems on their own are bullshit, I don't know what will.It's pretty clear all of this is a marketing effort only, and they pass human results as LLM findings.We already know that the supposedly industry-changing Mythos and Fable results were actually complete BS and they're just your run of the mill model. There's nothing at all to suggest this is any different, and once we get this "unreleased model" (aka bob from the math department), we'll see it was all lies again.
- ohohoSo, Simon or what's your name, you think your opinion on this has weight since you are a "blogger" or something, and sit on a pack of green paper?
- kdavisAll your datum are belong to us!
- dborehamThe concept that "knowing something has been done" allows others to find the solution to an previously unsolvable problem is an old proven one. For example when Germany launched a rocket (V2 prototype) the British knew its rough trajectory and from spying it's rough size. Although they had previously believed that ballistic missiles weren't possible because no engine could provide the necessary thrust to weight ratio, given the obvious German launching of one, they went through all the known chemical compounds to arrive at the combination (Ethanol and LOX) used. [Story from RV Jones "Most Secret War"].
- sk4rekr0wTime to repeat the same angry mob style discussion again. Great job to the mods.
- intendedThe more I see of frontier lab economics, the less it seems like the earlier predictions on how they will grow hold.This imbroglio looks like a point in the journey for a firm that is realizing it is not going to make money selling shovels during a gold rush, and that it has more to gain from just… being vertically integrated across an industry.Open models have definitely hampered the ability to sell tokens at a premium, so mass market adoption is impossible.But then take someone like Jane Street, for example. They self report making $30bn leveraging LLMs. It’s a defensible assumption that they are making profit on it.Perhaps it’s more profitable for OpenAI/Anthropic to build their own funds. They have the capital, compute, they can afford to recruit teams and buy any IP/Data required.This isn’t a fully fleshed out argument, but it is the first time it feels like the winds are changing.
- avs733I’ll go back to the point about authorship. I’m Not a mathematician but I am in academia. if you are fucking around with authorship you are immediately suspect.That aspect alone would/should be unthinkable to any serious academic. Authorship reflects who did the work and changing it for business competition reasons should be a red flag for multiple different reasons. They include, the sheer tactlessness of treating a major theoretical advancement as a competitive posturing first, the norms of academia second, and all the misunderstandings of the culture of the disciplines culture that people will now suspect are hiding beneath the visible surface (insert topography joke).Math as a field is fairly unique even in how they list authorship. It was long the norm that authorship to be alphabetical because the idea of first, second, senior etc authorship is harder to define than many other fields.“The stated rationale for alphabetical order is that it treats co-authorship as intellectually joint work: every listed author’s name carries equal weight, and no one has to negotiate, or be seen to negotiate, over billing. That is a genuine advantage over position-coded conventions, where disputes over who is “first author” are one of the most common sources of authorship conflict in fields that use them” [0]That norm is changing, slowly, but one option people are pursuing is notable: randomized author order. Their is a perception that alphabetical is too biased…that’s the world OpenAI is stepping into when they make that offer of authorship to one scholar with a demand that he exclude his partner.I can’t speak to the facts of anything else in this, but if a grad student came to me and said someone made them that offer, I would tell them to run and if they were brave report it.[0] a to the point lay description of the history of math authorship can be found here: https://casrai.org/guides/mathematics-alphabetical-authorshi...
- octocopIs there a good tldr on this topic?
- yieldcrvSorry math proof savants, the narcissism dream part of your career is deadOpenAI even wanted to respect that part but now you’re too focused on them being able to leverage the treasure map at the same time instead of the treasure for humanity being found at all, a completely different treasure where you both have to deal with getting paid by the bounty provider anywayGoofy
- krapcys[dead]
- anonundefined