Recently I care about how to create harness for mathematics that uses up the complete reasoning ability of the model. I care about both capability and cost.
Most generic harness we have now are not made for maximizing reasoning. I've tested agents like codex, and rarely the cost of reasoning tokens reaches more than 20%. Which is quite strange as math requires a lot of reasoning. So hackathons can be a good test bed.
Won't deny that this is an interesting idea, but I feel like waiting on the output of an LLM for 40 hours feels like it is completely antithetical to what makes classic Hackathons appealing / educative.
More generally, I don't think the shape of a hackathon (intensely working for a short timespan) maps at all onto the way LLM Math progress has seemingly been made so far; AFAIK it mostly involves picking out something for the Model, then having it run for a week with sporadic correction / encouragement.
If the goal is to accomplish something then why limit yourself with available tools?
I’m not a full on AI optimist but it is absolutely the most powerful tool in a host of applications. From a Hackathon perspective, obviously in the 90s it was much more unorganized, but the same ethos existed. Use all available tools to accomplish the goal/task, it’s where a lot of incredible learning came out of. The same will hopefully happen in scenarios like this one
Because the companies that run frontier models are malevolent by every metric.
They are destroying the environment, especially those in neighborhoods of low income people.
They are empowering their owners who are some of the most deplorable and duplicitous people living.
They are destroying personal compute to avoid competition with local models by buying all computer components with “promised money” and forcing their P into AI.
They stole the entire creative output of humanity and are trying to sell it back to us.
They are only good for giving wealth access to skill while removing from the skilled the ability to access wealth.
They are being used to kill in war and for surveillance.
Seriously why would you use them? Your use only emboldens them; making you complicit in their nefarious success.
When Claude made progress on the Riemann conjecture, here are the kind of prompts used:
> Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).2 This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.
Prompting for some of the results was almost the "Computer, do a breakthrough. Make no mistakes." meme. Just someone telling the model to keep trying a couple of times.
Unfortunately we don't actually know what kind of prompting was done for the more prominent results.
Did you read the article? This is not a hackathon where you build software, it’s one where you’re trying to get a model to make progress on a frontier math problem. The point is that that activity may not map well onto the shape of a hackathon
Should I read this as the big labs trying to move maths forward? Or the big labs trying to use professional mathematitians as (cheap?) Labour for validating LLM outputs? On yesterdays "An Alien Mind" post from openAI they openly said that maths is not a priority for them, so I personally know what to think...
Recently I care about how to create harness for mathematics that uses up the complete reasoning ability of the model. I care about both capability and cost.
Most generic harness we have now are not made for maximizing reasoning. I've tested agents like codex, and rarely the cost of reasoning tokens reaches more than 20%. Which is quite strange as math requires a lot of reasoning. So hackathons can be a good test bed.
More generally, I don't think the shape of a hackathon (intensely working for a short timespan) maps at all onto the way LLM Math progress has seemingly been made so far; AFAIK it mostly involves picking out something for the Model, then having it run for a week with sporadic correction / encouragement.
I’m not a full on AI optimist but it is absolutely the most powerful tool in a host of applications. From a Hackathon perspective, obviously in the 90s it was much more unorganized, but the same ethos existed. Use all available tools to accomplish the goal/task, it’s where a lot of incredible learning came out of. The same will hopefully happen in scenarios like this one
They are destroying the environment, especially those in neighborhoods of low income people.
They are empowering their owners who are some of the most deplorable and duplicitous people living.
They are destroying personal compute to avoid competition with local models by buying all computer components with “promised money” and forcing their P into AI.
They stole the entire creative output of humanity and are trying to sell it back to us.
They are only good for giving wealth access to skill while removing from the skilled the ability to access wealth.
They are being used to kill in war and for surveillance.
Seriously why would you use them? Your use only emboldens them; making you complicit in their nefarious success.
I for one, am one who walks away from Omelas.
I'm not sure people need to organize a mathathon to wait for a computer to give a printout. They mainly need tokens.
> Jarred's input was mostly limited to sending Claude messages of encouragement (mostly variants of “keep going” or “believe in yourself”).2 This seems to have helped Claude overcome some initial skepticism that it could make meaningful progress.
Unfortunately we don't actually know what kind of prompting was done for the more prominent results.