Regardless of what you think of the priority dispute issue discussed on sibling threads, I’m highly skeptical of the closing quote that this Navier Stokes result means that the same approach of casually spending a few million on agentic computation is going to solve end to end materials design or drug development.
Those problems can’t be formally verified with an automated theorem prover. We have a lot of physics based simulation tools, but they tend to focus on small subsets of the full design problem and they make limiting approximations because otherwise they’d be too computationally expensive, or we just don’t have the right data to parameterize them beyond describing qualitative behavior. Agents are helping accelerate research in these fields but I think it’s mostly a different class of problem that’s a lot harder to specify and verify
Yeah, I think you can't just throw money randomly at problems and expect results unless you know a line of attack that can get you all the way. OpenAI chose the line of attack only after it became known to them via rumors. They "front-ran" the researchers.
> However, communications quickly became contentious. According to Buckmaster, OpenAI offered to give him sole authorship on the Navier-Stokes solution—but only if Alpöge’s name was removed from the work and if the write-up would acknowledge the problem had been resolved by an internal OpenAI model. Buckmaster refused, in part because he was troubled by the question of what OpenAI's system had actually seen. For example, Buckmaster said the company did not initially give him a clear answer about whether its agents had access to the pair's logs on Codex (which is an OpenAI product).
> OpenAI executives have denied that any employee or AI agent saw the pair’s work before the researchers released it publicly on 7 September. But there still remains a separate question: Could the pair's work have reached OpenAI's models through its training data?
> OpenAI’s blog announcing the Navier-Stokes solution does not dismiss the possibility: “While unlikely, we cannot rule out that de-identified data derived from [Buckmaster and Alpöge’s] usage of our products helped improve our models .”
So, this company wants everyone and every organization on Earth to use their software, and reserve the right to then tell anyone how the resulting work can be published and credited?
Real solid business model there, how could it ever fail?
That's not what the comment you're replying to or the article says. I feel like I'm going crazy reading comments here and elsewhere, am I not reading the same articles as everyone else?
There's a lot of unverified hearsay but the crux of the problem is that there is controversy around using this company's tools, the attribution of the resulting work, and the company for some reason competing with its users. The whole thing reeks and my point is: people won't ask for the chromatography spectrum of the turd, they will walk away.
I think it was a pretty questionable thing to do by trying to front-run these researchers even if they didn’t make use of their techniques. The fact that they may have inadvertently “borrowed” their work via training data makes it much worse.
OpenAI’s behavior here — even if you only consider there side of the story — was (at best) in bad taste.
Strongly agree. And as one of the major AI companies, this is extremely tone deaf. If they saw a human (even if assisted) was making great progress on a major problem then you give them space. You don’t swoop in with millions in token spend to scoop them. There are tons of important problems where humans aren’t making traction - please go solve those.
And according to that team OpenAI only started asking their own model these questions after those submissions had occurred. So OpenAI had these critical clues and info before they started. If OpenAI did or did not use that to produce their own “proof” is an open question, but OpenAI hasn’t definitely denied it.
OpenAI has messed up big time here by competing with their customers. It would have become the norm for humans and mathematicians to use the tools and publish bigger results any way. If OpenAI didn’t run for credit, this theorem itself may have been proven by Buckmaster OR others in maybe a year or two.
But now the bigger issue than AI solving problems is the issue of chat privacy, at the end of the day.
"Breakthrough" to me would be like Isaac Newton inventing calculus to calculate pi. This feels more like $12MM in tokens was spent to add another digit to pi using the old way.
Your local trailer park was never going to be able to afford sponsoring high energy particle physics experiments projects that hollow out a mountain and use up a ton of xenon to try to detect a stray particle. High end science has required deep pockets for a long time.
But given the cost of a college textbook this is a pretty silly complaint to lobby against a subscription that's $200 a month, in the context of the cost of a variety of other materials and tools out there. (If you think that's expensive you've clearly been lucky enough to never have to deal with commercial software costs) Also not sure how quickly this stuff uses up limits; $100 or even $20 subs might be enough for students. And if a student is scrappy and figures out that Luna can meet their needs then I'd imagine Luna is effectively unlimited on some of these subs. Luna Max scores pretty high.
I'm a mathematician. I have a lot of trepidation about these tools and what they mean for the future of the profession.
That said... most of us are not working on problems as famous as Navier-Stokes. Even if OpenAI could scoop me based on my back-and-forth with ChatGPT, which I presume they could if they threw $15 million worth of compute at it, I highly doubt they'd bother.
The fact that anyone has to make that trade off with that through process shows a company and culture that is untrustworthy. OpenAI isnt open, isnt a non-profit, isnt for the benefit of humanity, and isnt even for the benefit of users at this point, its users are vassals providing training data so their models can reach ASI first.
I am going all in on sovereign ai even if its worse, these companies have shown they not only dont deserve trust but are actively stealing past and present intellectual property from humanity and users.
But consider they could decide that they want to scoop more regular research work too. They could automate it with just a few LoC. Even if you opted out in the ToS, you'd have to file a massive lawsuit just to enforce it. And the actual fine would be inconsequential to OpenAI.
I think going forward, any researcher should consider anything submitted to an LLM to be copied/stolen.
Personally this is a watershed moment for researchers and grad students I know. All of them are close sourcing WIP repos, not putting their progress in LLMs, or have lab level initiatives to self host models.
I am building a federated hosting platform that I plan to opensource for this exact reason for our company and similar users. I would love to talk to these teams (we are a small team at Duke and a health care company).
If you mean move to products with ZDR policies, OK, very fair and I agree. If you're talking the equivalent of carpenters should give up pneumatic air guns because hammers are more authentic, then that's silly. These are tools, and like any tool how effective they are can come down to how well you use them.
Should they give up pneumatic air guns if the company that supplies then can then control what you build with them, how you build with them, and secretly copies all your designs for their own use. Sure its a good tool, but its also becoming a trap.
Are you asking if it uses additional axioms of `sorry` in the proof? It's easy to check that it doesn't by compiling it and telling lean to list the axioms.
> OpenAI, meanwhile, says its experience with Navier-Stokes could open the door to solving puzzles with more practical relevance. “We are now able to spend millions of dollars on a problem that we really care about and that really matters: developing new materials, finding cures to diseases,” Bubeck said. “All of those things that we have been talking about for a long time—now they seem to be at our fingertips.”
Eh? There's no connection at all between the Navier-Stokes work and those things.
He may be hinting that solving these complex mathematical problems will boost OpenAI clients' confidence and encourage them to spend millions of dollars on solving other complex problems.
Asking the question in the right way may be 99% of the work. And the fact that openAI is effectively snooping on users and then outspending them to announce a break through is just gross.
They represent different AI usage patterns. OpenAI wants everyone to believe that it was done with a practically autonomous network of thousands of agents with little to no human intervention for 88 hours, while the Buckmaster/Alpöge were using AI in a more guided way for months. If OpenAI actually used anything from Buckmaster/Alpöge work they would be misleading the public.
Agree! Not sure I believe that Buckmaster/Alpöge were guiding the AI.
This is just their way of saying that they had something in the process and deserve merit.
I do. But the difference is: do I go around bragging that my agent worked for 88h to solve a problem? Where is the credit coming from? Is it coming from the: I was the first one to think about throwing a prompt: "Solve Riemann Hypothesis" and it turns out that by luck of the non-deterministic behavior of the agent, it got right? Wow, that is a lot of merit really. Congrats...
Sorry the sarcasm, but really your point makes absolutely no sense.
It is completely different to design something with AI and then execute, validate, evolve vs just prompt it machine-g-brrrr style and get a result.
This brings an important question. Nowadays I don't write code, I review code, I review systems behavior and get paid for it. Will that be the same for math researchers? Their prompt/problem is already well-posed out there. Ours, in the day-to-day, are not. Will the first one to verify AI work get the credit? or is it going to be the dumdum that types a simple prompt and has the compute to run it for 21321 hours? I absolutely don't get your point here. Or you are just rage baiting
This right here, or at least the thought of this, is why in the not so far future, businesses can't (won't?) be using these LLMs services.
You cannot risk companies like Anthropic, OpenAI or their business partners like Microsoft having unfettered access to proprietary data on your company/businesses.
It it likely that they or rogue employees will use the information to make a profit? It's pure speculation, but I'd say more than likely, and we will never hear about it or read it on the news unless there's whistleblowers in high enough positions to know about it.
Assuming you and your employees aren't careful with what data you share, they will have intimate knowledge about your company from files and conversations logs. Likely personal user data too which they'll gladly create databases to link to and create extensive profiles on you, your employees and your businesses.
It's not far-fetched to see them leveraging insider information shared with LLMs to play the stock market, leveraging data against competing businesses in other markets they might want to explore, and likely a bunch of other things that are escaping me right now as I write this.
At the end of the day it's on those people for sharing such sensitive data, but it's not like these AI companies are innocent and won't gladly exploit every little byte of data without telling you, we know it happens.
It is even worse. Just them knowing or having an inkling that someone is close to achieving a solution might be enough for them to spawn a team of thousands of agents and outrun you.
Those problems can’t be formally verified with an automated theorem prover. We have a lot of physics based simulation tools, but they tend to focus on small subsets of the full design problem and they make limiting approximations because otherwise they’d be too computationally expensive, or we just don’t have the right data to parameterize them beyond describing qualitative behavior. Agents are helping accelerate research in these fields but I think it’s mostly a different class of problem that’s a lot harder to specify and verify
> However, communications quickly became contentious. According to Buckmaster, OpenAI offered to give him sole authorship on the Navier-Stokes solution—but only if Alpöge’s name was removed from the work and if the write-up would acknowledge the problem had been resolved by an internal OpenAI model. Buckmaster refused, in part because he was troubled by the question of what OpenAI's system had actually seen. For example, Buckmaster said the company did not initially give him a clear answer about whether its agents had access to the pair's logs on Codex (which is an OpenAI product).
> OpenAI executives have denied that any employee or AI agent saw the pair’s work before the researchers released it publicly on 7 September. But there still remains a separate question: Could the pair's work have reached OpenAI's models through its training data?
> OpenAI’s blog announcing the Navier-Stokes solution does not dismiss the possibility: “While unlikely, we cannot rule out that de-identified data derived from [Buckmaster and Alpöge’s] usage of our products helped improve our models .”
Real solid business model there, how could it ever fail?
Well yeah, if you use the free product they train on your data, ... I thought this was widely understood?
> I pay for the tools my group uses out of my own research funds, including footing a large bill to OpenAI.
[0] https://cims.nyu.edu/~tristanb/statement.pdf
This is blatant scientific misconduct.
OpenAI’s behavior here — even if you only consider there side of the story — was (at best) in bad taste.
Quanta Magazine article that also discusses some of the controversy: https://www.quantamagazine.org/ai-has-solved-one-of-maths-1-...
Seven, not six. One is solved already, but is still a millennium problem.
Pick a better analogy.
But given the cost of a college textbook this is a pretty silly complaint to lobby against a subscription that's $200 a month, in the context of the cost of a variety of other materials and tools out there. (If you think that's expensive you've clearly been lucky enough to never have to deal with commercial software costs) Also not sure how quickly this stuff uses up limits; $100 or even $20 subs might be enough for students. And if a student is scrappy and figures out that Luna can meet their needs then I'd imagine Luna is effectively unlimited on some of these subs. Luna Max scores pretty high.
In any case mathematics is humanity's oldest open source project going on for millenia, it never belonged to a single country, institution or class.
That said... most of us are not working on problems as famous as Navier-Stokes. Even if OpenAI could scoop me based on my back-and-forth with ChatGPT, which I presume they could if they threw $15 million worth of compute at it, I highly doubt they'd bother.
I am going all in on sovereign ai even if its worse, these companies have shown they not only dont deserve trust but are actively stealing past and present intellectual property from humanity and users.
I think going forward, any researcher should consider anything submitted to an LLM to be copied/stolen.
And the AI company for making the search program that searched through the data and found the solution.
Eh? There's no connection at all between the Navier-Stokes work and those things.
Navier Stokes is a test of how high the intelligence is.
Sorry the sarcasm, but really your point makes absolutely no sense. It is completely different to design something with AI and then execute, validate, evolve vs just prompt it machine-g-brrrr style and get a result. This brings an important question. Nowadays I don't write code, I review code, I review systems behavior and get paid for it. Will that be the same for math researchers? Their prompt/problem is already well-posed out there. Ours, in the day-to-day, are not. Will the first one to verify AI work get the credit? or is it going to be the dumdum that types a simple prompt and has the compute to run it for 21321 hours? I absolutely don't get your point here. Or you are just rage baiting
You cannot risk companies like Anthropic, OpenAI or their business partners like Microsoft having unfettered access to proprietary data on your company/businesses.
It it likely that they or rogue employees will use the information to make a profit? It's pure speculation, but I'd say more than likely, and we will never hear about it or read it on the news unless there's whistleblowers in high enough positions to know about it.
Assuming you and your employees aren't careful with what data you share, they will have intimate knowledge about your company from files and conversations logs. Likely personal user data too which they'll gladly create databases to link to and create extensive profiles on you, your employees and your businesses.
It's not far-fetched to see them leveraging insider information shared with LLMs to play the stock market, leveraging data against competing businesses in other markets they might want to explore, and likely a bunch of other things that are escaping me right now as I write this.
At the end of the day it's on those people for sharing such sensitive data, but it's not like these AI companies are innocent and won't gladly exploit every little byte of data without telling you, we know it happens.