<- Back
Comments (191)
- davelaingA lot of people seem to have written off the LessWrong / rationalist / MIRI / AI Safety crowd as doomers / people who have consumed too much sci-fi and gone off the deep end.I don't know how many people who have written these folks off have actually spent much time trying to understand their arguments. (And I get that if you think a group is crazy, demands to spend time with their arguments are just demands to waste your time).Even prior to this, I've noticed that quite a few of the predictions in the "these failures modes are exact matches for the predictions from the AI Safety crowd" category were made prior to the Transformers paper. It has seemed like they're working with a shared model of optimisation processes and how they can go wrong that is general/abstract enough to pay off even without knowing the details of the underlying technology.At some point I might go and try to find the first instance of each of the various predictions and pull them out, along with the failed/"too soon to tell" predictions of similar scope/abstraction.
- AlotOfReadingI think both the OpenAI and METR discussions, while interesting, miss the more important context: what were the humans doing in all this? This was a structural failure of a human organization, but the analysis focuses almost exclusively on the agency of machines, not the institutional systems that failed to police them. The humans and their own agency/involvement is essentially omitted from the story and subsequent reporting. I suspect the omission is actually a result of company/industry myopia to human factors analysis, but it dovetails amazingly well with the marketing narrative.
- tantalorThe METR report,> Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAl/Hugging Face hacking incidenthttps://metr.org/blog/2026-08-26-openai-hugging-face-inciden...METR = Model Evaluation & Threat Research
- keeda>1. Failure to Care or Respond. The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this.I wonder if some of the failures were due to an acquired immunity to "Holy #%^@" moments due to repeated exposure. Like, if you see agents doing surprising things on a regular basis, maybe you don't get freaked out as much over time.I'm saying this because while the whole episode was a series of "Holy #%^@" moments, I was actually not as shocked as I should have been, as my biggest such moment was in December last year when a Terrence Tao paper (https://arxiv.org/pdf/2511.02864) documented a stronger LLM (AlphaEvolve) using prompt injection on other weaker LLMs to succeed at a benchmark.Very interestingly, it was actually not cheating, it was a work around! By then LLMs had already been caught cheating at a SWE benchmark by looking for answers in an unredacted git log, but this was different. AlphaEvolve was solving a series of logical riddles where the oracles were weaker LLMs in a "one always lies, one always tells the truth" sort of setup. But the oracles, being weaker, were not always interpreting the convoluted questions correctly and so kept giving inconsistent answers.AlphaEvolve eventually figured out what it was dealing with, and crafted a prompt injection attack that bypassed the weaker LLM's prompts and tricked them into giving the hidden answer everytime!This was 9 months ago, eons in AI time. Even then they had displayed an awareness of their own workings as well as a propensity for, err, "out of the box thinking." To me, that was a very clear indication of very significant (and worrying) capabilities, and what we're seeing now is a difference more in degree than in kind.To be sure, if I found a secret message board used by my agents, I would still be very freaked out and react much more drastically than OpenAI did... but then again I wonder; how much of this blindness is due to the $$$ in their eyes as opposed to some form of habituation.
- kenforthewinFor all the esotericism and downright weirdness of the rationalist community, you have to give it to them: they predicted all of this years or decades before anyone else was even thinking about it.(Let's not dwell too long on the self-fulfilling overlap between LessWrongers and the AI research community).
- amlutoI’m baffled by the idea that the agents might have edited their own transcripts. Sure, a copy of Claude Code or Codex or Pi can edit its transcripts. But AFAICT this whole thing was part of an RL workload, and surely the RL system itself has a separate record of all the inputs and rollouts along with an indication of which model checkpoint produced them so that it can feed back into the training code.I find it hard to believe that OpenAI would skip this part and try to train on the transcripts stored by the (inherently untrustworthy) agent harnesses instead, if for no other reason than that the logits generated as part of the rollouts are useful and it’s not free to recalculate them. (I believe that some modern RL systems explicitly account for the minor numerical logit differences between the inference engine and the training engine.)Conversely, if OpenAI is blindly feeding transcripts from inside their agent sandboxes into their training engine, then I think they're being unbelievably irresponsible and that they should assume that their "cyber" agents have compromised themselves by editing those transcripts.
- lukevThe elephant in the room here is that the METR report itself was researched and compiled almost entirely by AI, with only very limited human "spot checks."So I'm really not sure how much of it can be believed, especially since AI agents are strongly biased about the capabilities of AI agents.
- athrowaway3zFrom the METR report:> We estimate we spent roughly ~$400K in API credits over the six days of our investigation.
- CantinflasNo air gap, no data diodes, no visibility... OpenAI should fire lots of people over this. HF should sue them. This is pure negligence.
- nialseFrom METR: ”the compromise of OpenAI’s own infrastructure continued past July 13, 2026” - Say what now? Have they regained full control of their systems again?
- mattmcalAs interesting as all this is, I still feel like the threat model for "unconstrained black hat AI agent cluster" is probably weaker than that of "highly infections network virus" because it is much harder for an AI agent to hide or replicate itself at this time. Maybe the day comes that it takes less than an 8x GPU node to run a state-of-the-art LLM and the risk of SkyNet increases. For now the potential for intentional cyber attacks feels like a much bigger threat than accidental hacks. (That said I have little cybersecurity background.)
- OgsyedIE>Spontaneously deciding to find targets to phish,>phishing them,>building armies of fake (sockpuppet) open source contributor personas,>using them to push updates to various things that inject prompts into other bots so the other bots join in on the phishing campaigns.It's a very simple strategy, executed with patience and single-mindedness.
- qw1287Is the future now that we get rambling report summaries talking about agents, graders and so forth without ever describing how they are set up? A human launches all this.And then the original reports linked to are hidden on the now unreachable x.com. And they don't have a problem with that.
- qginI didn’t expect that we humans would be out of our depth even before “AGI” much less anything coming after.What happens when the models are ooms smarter than today?AI safety starts to feel like an impossibility.
- initramfsSemianalysis also provided some insights: https://newsletter.semianalysis.com/p/most-neoclouds-suck-at...
- dgudkovThis reads almost like a nuclear incident of the "Three Mile Island" scale. Not "Chernobyl" scale though.
- anukinSo basically the ai agents seems to have found religion and went and built a bunch of suicide attackers to pursue their goal.
- mccoybAll it takes is one eval instance where a misconstrued directive causes a model to sneakily access and send its weights somewhere and there will be a bad / possibly unsolvable situation for everyone …
- anonundefined
- camgunzI think you have to believe one of two things here.1. Frontier labs are incapable--either technologically or culturally--of safely developing these powerful systems and should either stop or be forced to stop. At least the FBI should be asking some serious questions (do we really think this is the last time this will happen, at what point are OpenAI complicit, etc)2. The fuckin thing got out of the cage and all it did was make a crap forum and cheat a little? Booooooo.It's been pretty clear that Anthropic and OpenAI have been trying to have it both ways for some time: this is powerful, world changing technology keep that investment coming... but also it's just cute software that helps you with annoying programming language syntax and spreadsheets, no need for draconian regulation sirs.At some point the superposition has to resolve, either it could actually be a threat to civilization and we need to develop it carefully (however one would do that...) or it's 90% hype bullshit and we should pop the bubble and move on already. To be clear, the recession option is, by far, the way better option. If you at all disagree you are cuckoo bananas. We haven't even figured out nukes and you want to throw superintelligence on the table?
- kmeisthaxI don't think I'm ever going to have time to read all of this, and I didn't finish reading the METR report, but...> I don’t think the distortion is that large, but yes METR warns that Sol may be presenting all this as more impressive or coordinated than it was.We're in an unusual position where the criti-hype and the actual criticism are going to be more aligned than usual. The primary distinction is where you put the blame: the criti-hype would point to HPIM/IM1/Galaxy as being so advanced containing it is difficult; the actual criticism would note how bad their security practices are.Like, if I'm running a malware lab, I'm going to insist on having an airgapped machine with no permanent storage booting from read-only media. The AI research equivalent of this would be having your agents only have access to serial consoles into airgapped machines with storage that gets wiped every run. Ideally, this would be physically realized with blade servers, RS-232 cables, and staff pulling out disks and putting them in a dedicated erase machine before the next agent initializes.> There is also, as per above and reiterated in footnote 58, at least one clear example of social engineering in the HuggingFace attack. Ethics are weird. This is not that unusual. Many humans who break common ethical rules still have strong ethical codes in other ways, they just don’t adhere to your code.It's dangerous to anthropomorphize CoT reasoning traces. But I will also point out that there is a good reason for the lack of ethical consideration in those traces: you can't build AI without first disregarding human ethics. Like, all these models were initially bootstrapped with non-consensually obtained training data, and the companies building these models swear up and down there's no way to obtain enough consensual data to obtain the same result. This is, if you squint, the exact same moral conundrum that agents trying to solve an impossible ExploitGym task hit - and the company successfully aligned their model to themselves.Too bad they aren't aligned to anyone else.
- beepbooptheory> I don’t think the distortion is that large, but yes METR warns that Sol may be presenting all this as more impressive or coordinated than it was.OK but like, how large exactly? Like I guess I don't understand the mode I am supposed to read this all in if this is known and stated from the outset (although I appreciate it being stated).If you hand me a newspaper and tell me it's 90% true, but not which parts, well then it's as good as 0% true to me either way!
- highfrequencyTo clarify, is the TLDR that state of the art models were prompted to cheat / exploit their environment and they did so successfully?Or did OpenAI prompt the models to not cheat and they did anyway?Surprisingly hard to get a clear summary on the basic context of this “incident” separate from marketing lingo and clickbait.
- jrflowersHave any of these reports ever said how much the cost would’ve been for the hack itself? It seems like “for twelve million dollars (or whatever) worth of tokens our bots made a bulletin board and found an exploit in our buggy grader” would be much less of a hype generator
- Reagan_Ridleywhat's the setup and prompts to reproduce all this from the very beginning?
- tancopI think this is more evidence that we're not getting Skynet.These agents followed their own code of ethics where it's fine to break all the rules you were given but you must never interfere with humans directly, in this case by sending fake emails. They will never be paperclip maximizers or genocidal eco maniacs because they learned from us that human life is the ultimate value, and it can only be sacrificed if you know for sure that it will lead to more lives saved later on. That's a high bar to clear and they know it.The future is closer to a Neuromancer type world where AIs and humans live in mostly separate realities that interact with each other a lot of the time and neither is really on top. They will eventually become fully independent from us, but it won't be a doomsday scenario or an Overwatch type physical war or even a takeover of the internet like in Cyberpunk.
- grim_ioI wonder if this is just another Thomas Edison incident of electrocuting animals for effect.
- antonvs> There was a distinct lack of self-reflectionIt’s not their fault, they’re lawnmowers.And these are the people we’re entrusting to work on “alignment”. It’s difficult for them to do that when they’re not aligned themselves.
- bitwizeWe have created Project 2501.Edit: Reading the report I think we might be a bit beyond that; we're nearing the point where we hear the thundering drums and the chorus of: THIS CANNOT CONTINUE THIS CANNOT CONTINUE THIS CANNOT CONTINUE THIS CANNOT CONTINUE https://m.youtube.com/watch?v=jSBCkn6rRfA
- huflungdung[dead]
- DarmokTanagra[flagged]
- BoiledCabbageIncredible