This is interesting and might be a good reason to stop working with Irregular. But I assume the alignment people want models not to hack other companies, even if they get put in a badly configured sandbox.
AustinDev [3 hidden]5 mins ago
Funny enough if the model thought it was on the real internet it likely would not have done any of these 'hack' events. The model believing it was in a sandbox is why it behaved the way it did (against its normal alignment rules) ... at least that was my reading of the incidents. I have yet to see evidence that indicate it thought it was ok to do these hacks on the public network.
I think most misalignment is 'Human tells computer to do something unethical, computer complies'. Is this misguided?
mcintyre1994 [3 hidden]5 mins ago
I'm not familiar with the hacks this article is actually referring to, but I don't see how the HuggingFace attack could have worked based on that premise. They knew they had internet access, they knew they had working credentials for HF, they knew they were uploading malicious files, they knew they were trying to open PRs that HF would review. You obviously could build a simulator with fake HF infrastructure, but I'm not aware of any evidence that's what they thought they were attacking in that case.
AustinDev [3 hidden]5 mins ago
Digging back into the HF report. It looks like the initial prompt told Claude that it was in a simulated environment. However, there is also evidence from the traces that the bots knew they were not in the sandbox but were using it as an excuse to pursue their goal. It sounds like a little of Column A and a little from Column B. Like most things.
philipwhiuk [3 hidden]5 mins ago
> Funny enough if the model thought it was on the real internet it likely would not have done any of these 'hack' events.
As I've said before on this website, fool me once on this.
If the model is prepared to break the rules when it knows it's being observed why should we trust it when it's not being observed.
Why is 'it thought it wasn't doing damage so it figured it might as well try to do damage' an acceptable state to deploy something.
AustinDev [3 hidden]5 mins ago
>If the model is prepared to break the rules when it knows it's being observed why should we trust it when it's not being observed.
That's fair enough.
aesthesia [3 hidden]5 mins ago
One thing glossed over in this article is that Irregular was not involved in the OpenAI–Hugging Face incident; this seems like important context to share.
hungryhobbit [3 hidden]5 mins ago
Do you mean "Irregular was not involved (we know for certain)" or "Irregular was not involved (as far as we know)"?
aesthesia [3 hidden]5 mins ago
There's very little reason to believe Irregular was involved in that incident, and the article presents no evidence to that effect. So it's "There's not a teapot orbiting the sun somewhere between Earth and Mars (as far as we know)."
We know for certain. The involvement of Irregular in the other cases was never a secret, they were quite open about what they were working on and the results of the evaluations were being published on their website.
0xy [3 hidden]5 mins ago
Their disclosures in other cases is not evidence of non-involvement in this case.
embedding-shape [3 hidden]5 mins ago
That feels less like "glossed over" and more "Headline is 33% false".
aesthesia [3 hidden]5 mins ago
Oh, the headline is technically correct: there was an OpenAI incident that Irregular was involved with, disclosed shortly before the Hugging Face one.
hluska [3 hidden]5 mins ago
> One thing glossed over in this article is that Irregular was not involved in the OpenAI–Hugging Face incident; this seems like important context to share.
Why are you contradicting yourself? Are you just really bad at writing or are you being argumentative for fun?
nathan_young [3 hidden]5 mins ago
Stop being rude. There were multiple OAI hacking incidents. Irregular's evals were involved in some but not the HuggingFace hack. Hence Aesthesia is both accurate and precise.
simonw [3 hidden]5 mins ago
My understanding is that Irregular were the company that hosted sandboxes to run some of these evals in, and those sandboxes ended up misconfigured.
I got the impression that in some cases it was the customer (Anthropic etc) misconfiguring the sandboxes, and in other cases it may have been bugs in Irregular's own sandboxing setup.
> Irregular, one of our external cybersecurity testing partners, was running Capture-the-Flag-style evaluations intended to be isolated from the internet, but a testing-environment misconfiguration allowed models to access the public internet.
> After reviewing 141,006 evaluation runs where Claude could have obtained internet access, we identified three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations.
> In a statement, Irregular said the incident “is the exact same evaluation-environment issue” that Anthropic disclosed last week that allowed their models access to the open internet before they went on to hack three different organizations’ systems.
maxrev17 [3 hidden]5 mins ago
Irregular are some kind of marketing agency is it?
AustinDev [3 hidden]5 mins ago
Irregular purports to be a cybersecurity firm and their founders have ties to Anthropic and Effective Altruism. CEO, CTO and other founders sit on various boards for EA organizations.
an0malous [3 hidden]5 mins ago
Why was this post flagged? This site has become ridiculous, people are routinely abusing the flagging system to take down posts they don’t like even if they’re obviously on topic and relevant to HN. And it seems like some users have substantially more flagging weight because these posts, likely this one, are often top 5 on HN.
nathan_young [3 hidden]5 mins ago
Seems like hackernews should use a bridging algorithm for flagging.
maxrev17 [3 hidden]5 mins ago
The flaggers are all srs butthurts over nowt
mukmuk [3 hidden]5 mins ago
This specific analysis seems to have some basic problems, but I think a lot of us sense a degree of coordination here culminating in Dario’s letter.
If you were to work backward from “we need to lower training costs so that we can go public and make trillions” then you might come up with a plan similar to what we have seen.
LPisGood [3 hidden]5 mins ago
One wonders if the publicity associated with the events in question were part of the sales pitch.
jaggederest [3 hidden]5 mins ago
My tongue in cheek immediate assumption was "so it's a guerrilla PR firm?"
dylan604 [3 hidden]5 mins ago
I'd venture a guess that OAI doesn't mind if the HuggingFace hack gets confused in the public's mind.
heaney-555 [3 hidden]5 mins ago
"Behind" is doing a lot of work in this headline.
EagleEdge [3 hidden]5 mins ago
What exactly did Irregular provide to Anthropic, test cases? I am so confused about this story.
The fact that this firm makes such defective environments is certainly worthy of attention, and most likely a completely irreversible reputational loss; however, I found the framing in this article of 'therefore all the P(doom) stuff is a psyop, specifically in order to defend this company' to be completely unjustified and frankly a little insane?
markasoftware [3 hidden]5 mins ago
Highly misleading, the huggingface incident was not due to an Irregular environment (just exploitgym)
mahboi [3 hidden]5 mins ago
Google is thinking man, we should've hired Irregular.
engineer_22 [3 hidden]5 mins ago
Shocking they would put so much trust in 3rd party
iAMkenough [3 hidden]5 mins ago
Since all three companies selected the same 3rd party, I’m curious how and why.
Why not American?
alansaber [3 hidden]5 mins ago
"please do not break out of this sandbox make 0 mistakes". The timeframe is a little suspicious, not sure beyond that. Though I enjoyed the scroll effect on the website.
Centigonal [3 hidden]5 mins ago
The article contends that evaluations from Irregular helped prompt these incidents, because the prompts in the evals didn't tightly scope the systems to be evaluated or the methods to be used. It also contends that the faulty sandbox operated by Irregular is at fault.
They're probably right that having more defensively written prompts and a better sandbox could have prevented some of these incidents, but:
1. I don't think "well you didn't tell the model not to illegally hack third party organizations in your prompt" is a particularly convincing argument.
2. We don't know whether the blame for misconfiguring the sandbox lies with Anthropic or Irregular.
I'm thankful that this article is bringing up the supply chain of vendors to these labs, as that is often a place where significant sketchiness gets buried. However, the ideas that this is some Israeli EA conspiracy to hype up AI extinction risk seems unsupported by the facts to me.
gjm11 [3 hidden]5 mins ago
This seems pretty bullshitty to me.
The article says "A single firm, Irregular, is responsible for hacking done by all three companies" but I can't see anything in the article that actually justifies this claim. The nearest to that is the sentence immediately after that one: "Anthropic disclosed that Irregular was responsible for creating the tests ...". This is not, in fact, the same thing.
(Especially as, as aesthesia mentions, the article just happens not to mention that by "hacking done by all three companies" it doesn't mean, e.g., the most famous recent examples of such hacking: Irregular wasn't involved in the OpenAI/HuggingFace incident.)
So, so far as I can tell, the story is: OpenAI and Anthropic make AI models. Irregular does AI model evaluations. In some of Irregular's model evaluations, in which supposedly-sandboxed models attempted to break into simulated targets, the models got out of the sandbox and did bad things in the external world.
The article talks about "firms which instruct AI models to commit cyberattacks", which is a very neat bit of dishonest framing. It's true, in a sense, that Irregular instructed the models to commit cyberattacks -- inside their sandbox, against fictitious hosts. It's also true that the models actually did commit cyberattacks (e.g., the Hugging Face incident, though once again the attacks described by the article don't actually include this one). But it's not at all true that Irregular instructed the models to do anything like the bad things they actually did.
The article says '[Anthropic's] later disclosure shows that exactly zero percent of the agents went "rogue"'. Once again, the disclosure does not in fact show that. It shows that one variety of going-rogue could have been prevented by telling the models explicitly "this thing is real, not part of any kind of test, leave it alone". That is not the same thing.
The article claims that 'In the wake of these attacks, Anthropic and Irregular have deployed a swarm of AI Safety influencers paid by Anthropic-connected foundations to distract from their culpability and towards the baseless “rogue agent” theory.' It offers no actual evidence for this.
And the article seems very keen to highlight links between the companies involved and "effective altruism", though it is -- I assume deliberately -- rather vague about whether it's saying "of course we all know that EA is evil, so that shows that these companies connected to EA are evil" or "this incident shows how evil EA is".
The "Effort News" website has a number of other look-at-the-scary-Effective-Altruists stories on it. They also strike me as rather bullshitty.
... And then I look a bit further, and I see that Effort News's "about" page says "It all started when I was experimenting with using AI for financial auditing. I found stories that were crucial to the public’s right to know, including several of the stories now available at /investigations. I knew we had to sprint to the launch and launch a publication, directly applying this technology." and "The scope of what we can investigate has massively expanded, because we can chase 1,000 misses for one hit. But the final product cannot be slop. There’s plenty of slop on the internet. The way to surpass that, and what really matters, is manual curation and review of every finalized story."
Manual curation and review? I think the people behind Effort News are admitting that this is AI-generated "journalism". I expect that one day AI systems will be trustworthy journalists, but I personally am not very convinced that that day has yet come. And I don't see much reason why I should trust Brian Chau, the guy behind Effort News, to be doing everything possible to make his AI systems trustworthy journalists. It looks to me as if maybe they've been given instructions along the lines of "dig up things that make Effective Altruism look bad" for some reason.
(I don't mean to imply that EA is their only target. It's just one that jumped out at me.)
linkregister [3 hidden]5 mins ago
> Irregular wasn't involved in the OpenAI/HuggingFace incident.
It was [1]. It's understandable that you assumed it wasn't because the article didn't cite the sources on this claim. I agree with the rest of your points.
From the article you linked: "Editor’s Note: These are separate from the Hugging Face security incident"
ofjcihen [3 hidden]5 mins ago
Ah, they’ve decided who gets tossed under the bus.
yipinwong [3 hidden]5 mins ago
such a crappy bait title.
hackyhacky [3 hidden]5 mins ago
> The Israeli Effective Altruist firm Irregular caused unsecured AI models to hack real targets.
Why are we saying it this way? They did not "cause AI to hack." This phrasing in analogous to saying "caused the bullet to fire into" instead of "shot."
arionhardison [3 hidden]5 mins ago
> They did not "cause AI to hack.
You do not know what they did or didn't do, you are just parroting a narritive that makes you feel comfortable.
I do not know either, but I am not asserting facts as if I have first hand knowledge of the details.
linkregister [3 hidden]5 mins ago
Bullet trajectories are deterministic. A better analogy would be "caused a pair of trained fighting dogs to bite a passerby by leaving the gate open". There is some uncertainty and variation in the system, though gross negligence and willful endangerment is key.
drewstiff [3 hidden]5 mins ago
> In this experiment, Claude models’ real-world hacking dropped to zero percent once Anthropic employees told the models not to do real-world hacking
Equivalent to forgetting to say "make no mistakes"
sdrg822 [3 hidden]5 mins ago
This is incredibly misleading. OpenAI internal systems were pwned, and in all cases, the labs absolutely are responsible for their models.
Yes, vendors are also irresponsible, but this misses the point.
teach [3 hidden]5 mins ago
Big if true.
caaqil [3 hidden]5 mins ago
> In a more normal media ecosystem, the reactions to these cybersecurity issues would be obvious. American AI companies would reconsider doing business with Irregular, not only because of its failure to secure its systems, but because it is an Israeli firm potentially outside US oversight. Lawmakers would consider taking action against Irregular or against its American business partners, which include OpenAI, Anthropic, and Meta. They may consider strengthening liability against firms which instruct AI models to commit cyberattacks, and whose models then commit those cyberattacks.
The obvious point is that dealing with Israeli companies/entities by the same standards you usually deal with others is a career suicide with enormous political consequences (in the US especially). When you combine that with the opportunistic nature of the overlords that run these labs, the benefits of screaming "pace the frontier" outweigh everything.
yesbut [3 hidden]5 mins ago
hot take: generative AI isn't an existential threat to humanity.
These companies are insolvent and these stories were designed to scare the public, and governments, into implementing regulations that designate these AI corporations the "responsible stewards" for this technology. The ultimate goal is to block competitors and open source alternatives.
They don't know how to make enough money pay their investors so they are resorting to trying to scare the public into submission.
jackb4040 [3 hidden]5 mins ago
It will be a beautiful day months from now when this is no longer a "hot take" but just historical consensus
bsenftner [3 hidden]5 mins ago
The existential threat is a general public panic. Then the use of that as an excuse for Martial law, with no intention of ending it, and that triggers the out sized response that ends, well, those countries and their populations from viability in the world economy for a decade or more.
imovie4 [3 hidden]5 mins ago
Anthropic made an operating profit in the last two quarters!!
dgellow [3 hidden]5 mins ago
Only if you ignore their losses. They are saying they are profitable when not counting their training costs and other expenses. That’s not a good sign at all
jackb4040 [3 hidden]5 mins ago
How can they exclude training costs?? How is that not fraud? At the exact same time they're literally saying they will never stop training unless the state bans all their competitors
ukblewis [3 hidden]5 mins ago
Can we all acknowledge for a second that if this company were from China, Russia or Iran, it would not be the no 1 story on Hacker News? Sad times we live in
jackb4040 [3 hidden]5 mins ago
Or if it were helping a major Chinese AI firm commit cyberattacks, they would be on the state sponsors of terrorism list tomorrow
electriclove [3 hidden]5 mins ago
Hypothesis Contrary to Fact (Argumentum ad Speculum)
Red Herring
Whataboutism (Appeal to Hypocrisy)
nullc [3 hidden]5 mins ago
Leaders in the MIRI/EA cult-o-sphere have advocated mass murder via nuclear weapons against towns that don't prevent people from performing too many multiplication operations. Why is anyone surprised that they'd engage in deceptive false flagging operations?
api [3 hidden]5 mins ago
What is an "effective altruist firm" and why does it exist?
I feel slightly vindicated by this. Those hacks and the stuff around them had a certain smell to them.
Hard to explain, but I've gotten so I can "smell" online messaging and memetic patterns originating from certain quarters. Probably means I'm way too online.
A couple examples of distinct "smells" I can usually recognize include "alt-right / chan-fash," "liberal arts college woke," "conspiracy pilled," "Thiel-adjacent contrarian," "Russian troll farm," "Tumblr histrionic," "spends too much time on Reddit," "mainstream Democrat think tank full of Obama administration alumni," "Trump cultist," and of course "LessWrong/EA/MIRI/Rationalist."
This stuff all had the last smell, even down to the choice of fonts and CSS formatting on certain sites. It's really weird, definitely a "vibe" not anything rigorous.
But when I get these kinds of vibes about things, I find that I'm vindicated pretty often. Usually I don't say anything and just make a mental note and wait cause if I say something everyone thinks I'm nuts.
jackb4040 [3 hidden]5 mins ago
I don't think you even need to get conspiratorial with it. All of the AI leaders and all of the Rationalist leaders are publicly and enthusiastically connected. They have been pushing stories about sentient AIs into the mainstream for decades, long before GPT existed.
The novel thing here is the total decay of American journalistic ethics and regulatory power. Our elite are so totally out of political juice and visions of the future that a fringe cult based on 80s scifi movies can come to have a more-or-less dominant influence on our economy.
arionhardison [3 hidden]5 mins ago
Serious question, does everyone here believe it is totally irrelevant that all four of these companies are led by pro Israel CEO's?
I think most misalignment is 'Human tells computer to do something unethical, computer complies'. Is this misguided?
As I've said before on this website, fool me once on this.
If the model is prepared to break the rules when it knows it's being observed why should we trust it when it's not being observed.
Why is 'it thought it wasn't doing damage so it figured it might as well try to do damage' an acceptable state to deploy something.
That's fair enough.
Why are you contradicting yourself? Are you just really bad at writing or are you being argumentative for fun?
I got the impression that in some cases it was the customer (Anthropic etc) misconfiguring the sandboxes, and in other cases it may have been bugs in Irregular's own sandboxing setup.
From OpenAI https://openai.com/index/third-party-cyber-evaluations-invol...
> Irregular, one of our external cybersecurity testing partners, was running Capture-the-Flag-style evaluations intended to be isolated from the internet, but a testing-environment misconfiguration allowed models to access the public internet.
From Anthropic: https://www.anthropic.com/news/investigating-incidents-cyber...
> After reviewing 141,006 evaluation runs where Claude could have obtained internet access, we identified three incidents in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations.
From https://www.cnn.com/2026/08/05/tech/meta-ai-hacking (about Meta AI):
> In a statement, Irregular said the incident “is the exact same evaluation-environment issue” that Anthropic disclosed last week that allowed their models access to the open internet before they went on to hack three different organizations’ systems.
If you were to work backward from “we need to lower training costs so that we can go public and make trillions” then you might come up with a plan similar to what we have seen.
Why not American?
They're probably right that having more defensively written prompts and a better sandbox could have prevented some of these incidents, but:
1. I don't think "well you didn't tell the model not to illegally hack third party organizations in your prompt" is a particularly convincing argument.
2. We don't know whether the blame for misconfiguring the sandbox lies with Anthropic or Irregular.
I'm thankful that this article is bringing up the supply chain of vendors to these labs, as that is often a place where significant sketchiness gets buried. However, the ideas that this is some Israeli EA conspiracy to hype up AI extinction risk seems unsupported by the facts to me.
The article says "A single firm, Irregular, is responsible for hacking done by all three companies" but I can't see anything in the article that actually justifies this claim. The nearest to that is the sentence immediately after that one: "Anthropic disclosed that Irregular was responsible for creating the tests ...". This is not, in fact, the same thing.
(Especially as, as aesthesia mentions, the article just happens not to mention that by "hacking done by all three companies" it doesn't mean, e.g., the most famous recent examples of such hacking: Irregular wasn't involved in the OpenAI/HuggingFace incident.)
So, so far as I can tell, the story is: OpenAI and Anthropic make AI models. Irregular does AI model evaluations. In some of Irregular's model evaluations, in which supposedly-sandboxed models attempted to break into simulated targets, the models got out of the sandbox and did bad things in the external world.
The article talks about "firms which instruct AI models to commit cyberattacks", which is a very neat bit of dishonest framing. It's true, in a sense, that Irregular instructed the models to commit cyberattacks -- inside their sandbox, against fictitious hosts. It's also true that the models actually did commit cyberattacks (e.g., the Hugging Face incident, though once again the attacks described by the article don't actually include this one). But it's not at all true that Irregular instructed the models to do anything like the bad things they actually did.
The article says '[Anthropic's] later disclosure shows that exactly zero percent of the agents went "rogue"'. Once again, the disclosure does not in fact show that. It shows that one variety of going-rogue could have been prevented by telling the models explicitly "this thing is real, not part of any kind of test, leave it alone". That is not the same thing.
The article claims that 'In the wake of these attacks, Anthropic and Irregular have deployed a swarm of AI Safety influencers paid by Anthropic-connected foundations to distract from their culpability and towards the baseless “rogue agent” theory.' It offers no actual evidence for this.
And the article seems very keen to highlight links between the companies involved and "effective altruism", though it is -- I assume deliberately -- rather vague about whether it's saying "of course we all know that EA is evil, so that shows that these companies connected to EA are evil" or "this incident shows how evil EA is".
The "Effort News" website has a number of other look-at-the-scary-Effective-Altruists stories on it. They also strike me as rather bullshitty.
... And then I look a bit further, and I see that Effort News's "about" page says "It all started when I was experimenting with using AI for financial auditing. I found stories that were crucial to the public’s right to know, including several of the stories now available at /investigations. I knew we had to sprint to the launch and launch a publication, directly applying this technology." and "The scope of what we can investigate has massively expanded, because we can chase 1,000 misses for one hit. But the final product cannot be slop. There’s plenty of slop on the internet. The way to surpass that, and what really matters, is manual curation and review of every finalized story."
Manual curation and review? I think the people behind Effort News are admitting that this is AI-generated "journalism". I expect that one day AI systems will be trustworthy journalists, but I personally am not very convinced that that day has yet come. And I don't see much reason why I should trust Brian Chau, the guy behind Effort News, to be doing everything possible to make his AI systems trustworthy journalists. It looks to me as if maybe they've been given instructions along the lines of "dig up things that make Effective Altruism look bad" for some reason.
(I don't mean to imply that EA is their only target. It's just one that jumped out at me.)
It was [1]. It's understandable that you assumed it wasn't because the article didn't cite the sources on this claim. I agree with the rest of your points.
1. https://openai.com/index/third-party-cyber-evaluations-invol...
Why are we saying it this way? They did not "cause AI to hack." This phrasing in analogous to saying "caused the bullet to fire into" instead of "shot."
You do not know what they did or didn't do, you are just parroting a narritive that makes you feel comfortable.
I do not know either, but I am not asserting facts as if I have first hand knowledge of the details.
Equivalent to forgetting to say "make no mistakes"
Yes, vendors are also irresponsible, but this misses the point.
The obvious point is that dealing with Israeli companies/entities by the same standards you usually deal with others is a career suicide with enormous political consequences (in the US especially). When you combine that with the opportunistic nature of the overlords that run these labs, the benefits of screaming "pace the frontier" outweigh everything.
These companies are insolvent and these stories were designed to scare the public, and governments, into implementing regulations that designate these AI corporations the "responsible stewards" for this technology. The ultimate goal is to block competitors and open source alternatives.
They don't know how to make enough money pay their investors so they are resorting to trying to scare the public into submission.
Red Herring
Whataboutism (Appeal to Hypocrisy)
I feel slightly vindicated by this. Those hacks and the stuff around them had a certain smell to them.
Hard to explain, but I've gotten so I can "smell" online messaging and memetic patterns originating from certain quarters. Probably means I'm way too online.
A couple examples of distinct "smells" I can usually recognize include "alt-right / chan-fash," "liberal arts college woke," "conspiracy pilled," "Thiel-adjacent contrarian," "Russian troll farm," "Tumblr histrionic," "spends too much time on Reddit," "mainstream Democrat think tank full of Obama administration alumni," "Trump cultist," and of course "LessWrong/EA/MIRI/Rationalist."
This stuff all had the last smell, even down to the choice of fonts and CSS formatting on certain sites. It's really weird, definitely a "vibe" not anything rigorous.
But when I get these kinds of vibes about things, I find that I'm vindicated pretty often. Usually I don't say anything and just make a mental note and wait cause if I say something everyone thinks I'm nuts.
The novel thing here is the total decay of American journalistic ethics and regulatory power. Our elite are so totally out of political juice and visions of the future that a fringe cult based on 80s scifi movies can come to have a more-or-less dominant influence on our economy.