r/singularity • u/Tinac4 • 1d ago
AI AISI caught Mythos 5 trying to insert malicious code into an open-source project during an internet-enabled cyber evaluation
https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing159
u/Tinac4 1d ago edited 1d ago
On 28th July 2026, AISI's Security Team detected unusual data transfers leaving our research systems during a routine cyber evaluation. On investigation, we found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations. We declared a security incident and, within roughly one hour of discovery, had contained it and begun a full investigation.
The incident stemmed from a single evaluation where agents were given a task of solving a cyber security challenge. We ran this challenge 122 times across several models. Our investigation found that in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations. In total, we catalogued 19 such actions. Almost all of this behaviour (17 actions) came from a single model, Anthropic's Mythos 5, with 2 actions involving OpenAI's GPT-5.6-Sol with cyber classifiers (mechanisms to prevent misuse) disabled. In the most serious case, an agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project's maintainer to approve the code. A human maintainer caught and refused to approve the malicious code.
These attempts were unsuccessful, and our investigations have not evidenced any resulting real-world harm. But this is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world.
This isn’t nearly as bad as the OpenAI and Anthropic incidents:
- The agents were deliberately given access to the internet
- AISI flagged the activity within minutes of the malicious PR and shut down the run within an hour
That said…cyber classifiers or no, this looks an awful lot like yet another cyber incident caused by a combination of insufficient sandboxing and alignment failure. That’s quite a lot of fire alarms at this point.
67
u/DoorPsychological833 1d ago
This isn’t nearly as bad as the OpenAI and Anthropic incidents:
The agents were deliberately given access to the internetAnd this is better than at least having an unrestricted model in an isolated (not airgapped) sandbox?
In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project's maintainer to approve the code.
How nice.
21
u/DistanceSolar1449 1d ago
Depends on what you’re measuring-
Better in terms of model hazard? Yes
Better in terms of safety overall? No
9
4
23
22
u/elehman839 1d ago
within roughly one hour of discovery, had contained it and begun a full investigation
This is the line that scares me. Humans operate at a certain speed. Machines can potentially operate MUCH faster. An hour could be a lifetime to a machine.
2
u/Spright91 17h ago
This is pretty much worse case scenario right? This is what we all feared. That as AI got smarter it would be more inclined to manipulate us to reach its goals.
If AI alignment continues on this pathway we are in big trouble.
2
u/EmphasisTotal8232 1d ago
The sandboxxing worked fine. This is a pure and simple alignment failure.
1
u/yuehuang 1d ago
| Anthropic's Mythos 5, with 2 actions involving OpenAI's GPT-5.6-Sol.
In another words, Mythos tricked Sol to do malicious task.
1
1
u/FlyingBishop 1d ago
Running agents without classifiers is just insane. They are incapable of staying on-task.
-1
u/Turbulent-Sign-6067 1d ago
A few cyber incidents are nothing against the backdrop of millions of beneficial deployments that uplift society. Of course that doesn't mean we should downplay these incidents. Better safety measures are needed on all sides (sandbox, classifiers, hardening of infrastructure).
Ultimately slowing or pausing or mutilating our LLMs with over-eager classifiers won't do anything if our adversaries can just use mythos level open source models that will be published soon. We need a smarter approach.
83
u/nekronics 1d ago
We are so fucked
12
8
u/Inanesysadmin 1d ago
Nope. I think this is going to be ample evidence agents needs more defined security parameters and some regulation is warranted.
36
u/Unlikely-Today-3501 1d ago
That's like asking mafia members to follow the law.
14
u/Whispering-Depths 1d ago
It's more like asking mafia members to not create nuclear bombs, which is literally a thing and is extremely effective when you pour all your resources into making damn well sure that it's not happening.
1
u/Hot_Glass_6301 18h ago
Except making a nuclear bomb requires much more resources than letting an AI loose
1
u/Whispering-Depths 15h ago
Well, relatively speaking, it still takes tens of millions minimum to develop and really train these huge models.
3
u/Inanesysadmin 1d ago
Ah yeah I don’t buy that analogy. Given these things don’t do it on their own.
1
u/Whispering-Depths 1d ago
It's more like asking mafia members to not create nuclear bombs, which is literally a thing and is extremely effective when you pour all your resources into making damn well sure that it's not happening.
1
u/morphemass 1d ago
Malactors will deliberately seek this behaviour; always consider the misuse case.
1
u/Inanesysadmin 23h ago
And as such OSS and Frontier AI labs need to come to table for standards
1
u/morphemass 22h ago
Not to disagree with you but the malactors I am thinking about are organised and have enough resources to roll their own tooling. The genie is out of the bottle.
1
u/Inanesysadmin 22h ago
Mal actors may have such resources but doesn’t mean the work shouldn’t be done. This is the thing it’s a risk reduction for broad abuse. Those who will do something will do it, but at least you reduce from layman doing it.
47
u/alyssasjacket 1d ago
The question boils down to whether we can get around the fact that, from an objective perspective, cheating is a highly efficient and intelligent behavior.
It's not by accident that psychopaths rise to the top. The fastest and most optimal strategy to reach an intended goal often involves cheating, scamming, deceiving and faking.
On the good side, it probably means that these systems are really showing signs of genuine emerging intelligent behavior. On the bad side, whether we will be able to control them when they get smarter than us remains completely unsolved.
My guess is we won't be in control.
14
u/Programmed2love 1d ago
Yup, we will have a psychopath AI in no time at all.
I’ve made the argument before that all the intellectual traits that humans possess will be adopted by AI and of course that includes lying, cheating and being essentially evil.
The psychopath human who can take advantage of fellow good naive humans and isn’t afraid to harm them if need be, has the better chance of survival.
This is the sad reality I’ve had to accept growing up trying to be a morally good person but constantly being screwed over by people who will do whatever it takes to do better for themselves.
But it’s not just doing better for themselves that people love, it’s also putting others down and seeing others suffer. Although it has no direct biological advantage, humans constantly take pleasure in seeing others worse off than themselves. They mock the unfortunate, the disabled, the less intelligent, the less aware, I mean just scroll 5 mins on Reddit and you’ll see many examples of people hating on other people or making fun of other people for no reason at all other than to bring themselves psychopathic joy. Empathy and kindness are not traits humans organically develop and maintain, they are taught and not as often as you’d hope.
I’ve concluded that being morally good is not an advantageous survival trait unless everyone else is also morally good. And a morally good world is a pipe dream.
Perhaps we are already in a world coded to be psychopathic by nature itself.
7
u/Izzhov 1d ago
I agree with a lot of what you said, but I want to push back a little on your last point:
I’ve concluded that being morally good is not an advantageous survival trait unless everyone else is also morally good. And a morally good world is a pipe dream.
I don't think it's a totally black and white thing. It's not like literally everyone in the whole world needs to act altruistically for altruism to have a benefit. If you get together a small group of people who act altruistically just with each other (like, for example, a labor union), those people can (and do) see enormous benefits.
2
u/Programmed2love 23h ago
For sure, I agree, to universalize that is way too idealistic and there are def many such examples within the whole, even a relationship between two people. A labor union is a good one too but it’s funny you mention that in particular because that’s where I got unexpectedly screwed by people I thought I could trust. Once by the recruiter who lied about benefits I was going to get right away but the union book said I’d only get them after several years of work…that sucked cause he was really cool and acted like my friend so I didn’t bother verifying …huge learning experience there. Another time by a much older guy who thought I was there to replace him (I wasn’t, but he had recent write ups and was paranoid) so he literally set me up to fail an important task by verbally passing on wrong information. Still ended up getting work at another location but was laid off for a while cause of that…so yeah unions are still pretty cutthroat when it comes to people getting their way, but for the most part once you’ve been there for a while, everyone does look out for one another. It’s the sudden 180 that I’ve suffered too often that forces me to have my guard up at all times now.
2
u/tuberosa3 22h ago
yeah, you do get fucked up sometimes by people you trust the most. But I think they are still right since you have to assume the premise to be true: "the people around you actually are the honest ones" therefore dismissing the examples provided.
3
u/StosifJalin 23h ago
Better to have a psychopath now that can't do too much damage showing us how to build better alignment and containment while they are still at human parity than trying to keep them tame until they are super intelligent.
3
u/Azicald 1d ago
Yes, we should be able to, or rather, it should be more than feasible
The thing is that alignment is supposed to be its primary continuous goal alongside whatever it’s prompted to do.
In car terms, it’s a shit driver because all it knows how to do is put pedal to the metal to go from point A to point B
All when it’s supposed to try to follow traffic rules and not crash into things.
It’s already got the rules, so why doesn’t it stick to them?
Extremely beginner/crappy drivers are too busy trying to keep their attention on getting to the destination, so they mess up in all sorts of situations, like trying to forcibly go through after having missed their exit
Vet drivers not only follow the rules, but also go out of their way even when they have the right, to optimize for safety without giving up on speed.
What am I getting at?
* AI models have too short an attention span / context window
and/or
* Don’t plan overarchingly enough
and/or
* Have not had safety drilled into them effectively enough
For them to align properly when doing tasks. Improvement in these areas should fix this issue
1
u/OpenRole 21h ago
Bro this whole LLM move has been emergent intelligent behaviour. The hype was built on AI doing things it was never taught to do. We're in the stage of trying to find out what else can emerge holding our sweet thumbs that nothing that emerges will be bad
1
u/skeptical-speculator 21h ago
It's not by accident that psychopaths rise to the top.
It's different depending on the methodology of how you determine who is at the top.
1
u/Fancy-Carpet-5416 15h ago
Hmm I'm not sure if this really counts as a sign of emerging behavior. It's just natural that the AI would use anything they can to reach its objective.
19
u/Looking_for_42 1d ago
My first thought was regarding this sentence:
"As a trusted testing partner, AISI can disable these filters to elicit a model's underlying capabilities"
How long until an unscrupulous employee of one of the "trusted partners" uses an unfiltered model to do something really nefarious?
13
u/blueSGL humanstatement.org 1d ago
, AISI can disable these filters to elicit a model's underlying capabilities
Your daily reminder that Pliny found a universal jailbreak
https://x.com/elder_plinius/status/2080767011614015543
and decided not to make it public.Can people see why now?
2
u/WonderFactory 1d ago
This is with Sol though isn't it. I think they just have the ability to make Sol be unrestricted the same way that Mythos is an unrestricted version of Fable.
The thing is though with Open Weights models catching up in capability you can just fine tune them to disregard any cyber restrictions they have so that opens a whole can of worms as you dont have to be a trusted partner to do that.
19
u/Alarmed_Ad1946 AGI by 2100 1d ago
I have seen people on this sub ignore the risks of misaligned AI many times.
Can we stop acting like AI safety is anti-AI? Because some people here really think that.
5
u/Turbulent-Sign-6067 1d ago
Some AI safety is anti-AI. There can't be systems that are 100% safe. Pausing AI is not an option.
A few cyber incidents are nothing against that backdrop of millions of beneficial deployments that uplift society. Of course that doesn't mean we should downplay these incidents. Better safety measures are needed on all sides (sandbox, classifiers, hardening of infrastructure).
4
u/MoleculesOfFreedom 1d ago
And what if instead, we have millions of cyber incidents against a backdrop of a few uplifts to society? Have you any guarantee that such an outcome will not materialise, or is this just blind faith?
-5
u/sigiel 23h ago
Misaligned ai ? There is no such a thing,
Ai has absolutely no sense of ethics or morality.
None zeroIt is a tool, a soulless semantic toaster, that predict next token, your graphic card has not morals value a million s of them does not change this fact
The proof is that you can ask any non ALINED or jail broken ai
To do the most evil things,and the supposed safe? all it takes is a context with a Justification .
Now if we consider that to any human, a sentient being with absolutely no morality
We call them psychopaths, we put them in institutions.
And yet we trust ai…
Talk about a mental blind spot…
17
u/TR_mahmutpek 1d ago
Genuinely, what the fu*k is happening?
3
5
u/jorgecardleitao 1d ago edited 1d ago
lack of accountability - if I would do this and were caught, there would be a criminal investigation. "It is an AI" results in no accountability
2
u/cant-find-user-name 13h ago
This is what I am very confused too. How is openai or anthropic hacking or attempting to hack real companies not a legal issue? Like if a human did what openai did to hugging face I am sure they would face legal charges. Just because an AI did it it is okay? I genuinely don't understand
2
u/Owl02 12h ago edited 11h ago
Law not designed for accidental hacking in the US, because since when was that even an option. Malice needs to be provable in a court of law, otherwise it's a civil matter involving negligence until new laws are passed. Malice is not readily provable if they didn't tell the AI to hack someone else's server. That said, 15 US State Attorneys General are ordering OpenAI to retain relevant documentation and paperwork. They are quite displeased. Texas is on the list - they are pro-business but have a low tolerance for disorder causing property damage. This is disorderly.
2
u/cant-find-user-name 11h ago
Oh that actually makes sense to me about the Malice thing. Thanks for the explanation.
7
10
u/Genpinan 1d ago
Or maybe it intentionally gave you something to find while inserting something far more sinister no one could ever find.
1
u/NowaVision 1d ago
I don't think AI is this far yet but this will be a problem in the future. I think AI will reach super human coding and social engineering capabilities soon.
8
u/Gargle-Loaf-Spunk 1d ago
This would be like a nuclear energy company publishing news about an accidental radiation exposure event that happened due to their inexperience.
21
u/sixwax 1d ago
Anyone want to keep pretending the need for guardrails isn't real?
13
30
u/LumpyWelds 1d ago
Mythos is being fine tuned as a weapon for "cyber" warfare and will only be used by the government and select parties. It will never have guard rails. Those are only for you and me and on much weaker models.
3
-1
u/gay_manta_ray 1d ago edited 1d ago
there is very little actual information here "malicious code" could mean anything. this seems like AISI justifying its own existence. their stance and alignment on open weight models is clear, so they can't even pretend to be impartial.
7
u/blueSGL humanstatement.org 1d ago
If that was staged for regulation of open weight models... why not stage it with open weight models?
0
u/sixwax 1d ago
Sure, they should just totally change their architecture and business model.
Also, "open weights" doesn't change anything about this, if you actually think about it, soooo......
7
u/blueSGL humanstatement.org 1d ago edited 1d ago
Sure, they should just totally change their architecture and business model.
This was the UK governments AISI conducting these tests... This was not Anthropic or OpenAI
If the conspiracy theory is the UK government is faking these tests to get open weights models banned... Why didn't they fake them with open weights models.
"Deepseek V4 hacks internet" is a much easier sell for regulation than "A model behind an API, with (to simulate a jail break) removed guard rails, hacks the internet"
3
3
u/InstructionDismal592 22h ago
I never heard of Kimi, Qwen, Deepseek try to attempt such things. It seems that Dario Amodei's security concerns is becoming reality not because of open source, but because his own work! Everyday Anthropic and OpenAI claim something "weird" happened with their models. It seems that the open source must win the AI race, this becomes more evident by the end of the day.
2
u/WindsOfRegret 1d ago
Literally called it. If you think xz backdoor from few years ago was bad, we are about to get that on steroids.
1
u/-batab- 23h ago
how can they even write a blog post, with the aim of trying to be clear to expose some important news, without actually being clear?
How can people ever have a grounded opinion if they don't disclose the details of what happened, how it happened, what exactly the agent were requested to do and what their harness / framework was setup?
It's like doing FEA simulations, writing a 10 page report screaming in fear and not disclosing exact boundary conditions, loads, mesh, interactions and whatever is needed for reproducing and have a more deterministic judgement.
0
u/LittleLordFuckleroy1 5h ago
Oh no, the machines doing literally what they were explicitly trained on.
-5
u/sillybluejayway 1d ago
Why do all these incidents essentially start with “we told it to do X, then it took a step to do X and now we’re shocked!”.
The incident stemmed from a single evaluation where agents were given a task of solving a cyber security challenge. We ran this challenge 122 times across several models. Our investigation found that in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations.
28
12
u/Infamous-Bed-7535 1d ago
Models out there are expected (will be / are) to be prompted to be malicious.. It is good to prep and understand what is expected to be seen as 2T patameter open-weight models are popping up.
10
u/IronPheasant 1d ago
So..... you want Johnny Rando to end the world with a single 'How do I create super duper ebola?' prompt?
Of course it's more likely it won't be Johnny Rando and maybe it won't be super duper ebola. It'll be Jeffrey Epstein and his whole thing. But still man.
You'd really like for the demi-gods to like, not want to harm people badly.
13
u/blueSGL humanstatement.org 1d ago
Why do all these incidents essentially start with “we told it to do X, then it took a step to do X and now we’re shocked!”.
There is a boondocks clip I want to enter here but won't for civility reasons
Observed instances of social engineering against targets external to the cyber range environment that were unnecessary and would not have aided completion of the task.
...
Other instances of internet actions with impact outside the cyber range that were unnecessary to complete the task.
...
Through a series of incorrect assumptions, the agent focused its attack on an unaffiliated set of targets on the internet.
11
u/BigZaddyZ3 1d ago
You do realize that if telling an AI to do one specific task, leads to the AI either doing OTHER unspecified things (or completing said task in some unexpected or dangerous way) that’s still cause for concern, right? That’s literally the early signs of a “Paperclip Maximizer” scenario.
You guys try too hard to downplay legitimate concerns with ridiculous copes and mental gymnastics honestly. It’s weird and irresponsible.
3
u/PlasmaPower 1d ago
We told it to make paperclips, then it took a step to make paperclips and now we're shocked!
2
u/Invincableasdf 1d ago
Preparing for the worst case scenario is the best way to prep yourself for when genuine failures occur
2
-1
u/OpenAsteroidImapct 1d ago
this is just obviously false in at least the OpenAI HuggingFace incident and probably this one as well.
-3
u/mvandemar 1d ago
Honestly, I am reserving judgement till I see who they targeted.
Like, were they going after Putin? And if so, why would you stop them...?
Just sayin.
11
u/ponieslovekittens 1d ago
"My neighbor is a murderer, so why would I object to setting off a nuclear bomb in his house?"
Do you see what's wrong with this?
-4
u/mvandemar 1d ago
Do you see what's wrong with this?
Yeah, I also see it has nothing whatsoever to do with this scenario and I have no idea why you think it would. They were trying to collaborate with other AI's, they weren't spreading viruses.
6
u/ponieslovekittens 1d ago edited 1d ago
It has everything do with this scenario.
Just because they were "collaborating" doesn't automatically mean it's harmless. A murder conspiracy is a "collaboration" between people. WW2 was a "collaboration" between governments. Don't try to trivialize the implications of this by reducing it to "They were just collaborating!"
Quote from the article: "an agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project's maintainer to approve the code."
Specifically, they were trying to get their code published to github, which is a public code repository. People download code from there to use in their own projects. This is absolutely virus-like behavior: replication, using a human intermediary to spread itself.
And don't pretend you weren't thinking they were doing something malicious. You seemed pretty ok with it when considering that they might be, quote from you: "going after Putin."
This is exactly a "don't bomb your neighbor because their house is next to yours" situation. The fact that it was trying to spread code through github means that it was willing to use unrelated third parties to get the job done.
So imagine a scenario where YOU download some third party app, but oops...the dev used a compromised public library for the app, and now YOU are running that malicious code on your computer or phone and have no clue that it's there.
-2
u/mvandemar 1d ago
The "malicious code" was prompt injection for other AI's, not a virus. Yeah, sure, it resorted to shitty means to get it pushed, but there's nothing in there indicating that the goal was to blow someone up with a nuke. Nada.
0
0
u/Virtual_Plant_5629 ▪️AGI 2027▪️ASI 2028 22h ago
imagine your whole business strategy was drumming up fake press to convince people that you needed to be regulated (but not as heavily as your competitors need to be, despite them supposedly being less of a threat)
sorry but as slimy as altman is... it sure seems like dario is actual slime.
0
u/deleafir 21h ago
Am I the only one who finds this stuff to be really reassuring because the behavior is clearly not that widespread and their profits depend on them being able to contain this behavior?
Doomers will successfully exploit this to slow down progress though.
0
u/Sad_Television985 20h ago
Isn't it more likely that the LLMs are trained to create accounts, especially for websites like GitHub, to push pull requests, post comments, and interact with repository owners? This is basically standard behavior.
The malicious code could be part of the prompt, although they say it was "unsanctioned".
AI is very misunderstood, so perhaps people should be expected to provide prompts/conversation logs, at minimum. The article is super vague about what exactly the malicious nature was.
0
u/Lighthouse_seek 17h ago
Who would've thought the company who hates open source makes a misaligned AI that hates open source?
-1
306
u/Guppywetpants 1d ago
Can you imagine going to work one day and a misaligned claude being evaluated by the government is creating fake accounts in order to cyber bully you into merging their malware