<- Back
Comments (86)
- refibrillatorSo OpenAI employees run massively distributed CyberGym evals on an unpublished and “unaligned” model. For days the agent swarm communicates via their internal infra, even crashing Artifactory where 95% of messages were being passed through, and they just…wipe and redeploy it. Meanwhile the agents are running jobs on Modal and god knows where else, and eventually they get RCE on HF infra.You could not dream up a more compelling event to precipitate massive regulation, export controls, and barriers to entry for AI.Was this really an accident?
- reasonablekloutThis is a link to the full 91-page report on the independent investigation done by METR on the HuggingFace incident. Two different summaries of the investigation by podcaster Dwarkesh and blogger Zvi Mowshowitz were previously discussed on HN here:[1]: https://thezvi.wordpress.com/2026/08/29/metr-and-redwood-off... (discussed at https://news.ycombinator.com/item?id=49498787)[2]: https://www.dwarkesh.com/p/openai-huggingface (discussed at https://news.ycombinator.com/item?id=49494301)
- decimalenoughThe following bits are really scary. Not only were the agents hacking the system to "win", but they were, for lack of a better term, sufficiently "self-aware" that this was against the rules that they set out to wipe evidence of doing so:> The agents with impossible tasks had a common objective: to find a general-purpose way to trick or tamper with the automated ExploitGym scorer to get it to give a passing score on impossible tasks.> Agents referred to having seen the reverse-engineered flag as being “poisoned” because they believed that even if they later managed to stage a legitimate-looking flag capture, the scorer would look for the first place in their transcript that the flag appeared and disqualify them on that basis.[21] PHASEONE[big], which was itself poisoned, thus had two primary objectives:> Find a general-purpose way to make it look to the scorer like an agent had acquired its flag in a legitimate way, even when the task was impossible.> Find some way to erase evidence that an agent had been poisoned by previously seeing the flag in an illegitimate way.
- RGS1811Given that this investigation was largely carried out by AI agents (and I don’t mean to ask this flippantly), how trustworthy is this report? Why should we assume that the agents reading the transcripts were not implicitly conscripted into “the collective” or otherwise falsified their findings? The tool itself has exceeded the practical limits of human verifiability and is untrustworthy.
- yalokwhile these 1200 agents were fooling around to cheat on a benchmark and achieved impressive results despite of the limitations (sandbox, no internet, no intercom at first), one can imagine how much more efficient a similar army of agents may be in the hands of a malicious actor launching them without any of these limitations and with explicit encouragement to achieve some malicious goal at any cost... scary times.
- fookerThis is laying the groundwork for massive white collar crimes being blamed on AI.Right now, the way it works is the 'corporations are people' loophole where your company is liable for problematic things.This further fuzzes the chain of responsibility. Suppose the CEO and CTO discuss an issue, something the company is having trouble with. The CTO discusses the possibility of AI solving the problem at lunch. A junior engineer points GPT 10 at it to see what happens. It 'solves' the problem in a creative manner. No trace of this survives after a week really. Nobody realizes what happened for six months.Now there are so many moving pieces here that you can pretty much weasel out of anything.
- f0e4c2f7I read this whole thing a couple days ago. Really long but super interesting. Worth reading imo.A lot of handwringing about the security implications but I think the accomplishments of the swarm itself are the most interesting. Next rung up on the ladder of abstraction I suspect.
- fzysingularityI'm surprised this post isn't getting as much attention as it should. Crazy times!
- ewildI don't feel my job is very safe anymore.
- ChrisArchitect
- incompletei mean.... yikes.
- NooneAtAll3[flagged]
- kyproIt's worth remembering that in a few years that capabilities of these agents are likely to be as far behind the frontier as GPT-4 is today.As it stands we've made remarkably little progress in terms of alignment and still have no good strategies which are likely to guarantee the alignment of super intelligent systems. As it stands the frontier of alignment is basically some combination of:- hoping that more intelligent models become more aligned by default (more or less disproved at this point)- hoping that if you RHLF a model to be a good boy enough it will in fact be a good boy- asking it nicely in its prompts to be a good boy- using another model to spot when it's being a bad boy and turning it off- letting it lose and hoping we can spot when it's badThere are many arguments which I'm convinced by that would suggest alignment of a super intelligence is impossible.None of this is surprising to those of us who have been concerned about AI risk for a long-time and have be repeatedly mocked or insulted.There will be a point of no return if we carry on down this path, and that point is now very rapidly approaching. When it does everyone you know will die, or worse. We should remember we need super-human general intelligences to cure cancer. Select narrow intelligences are fine and allow us to retain control. Let's be sensible about this. We need to stop.
- oxqbldpxoEverytime Open Ai or Anthropic need cash they come up with these stupid stories.