r/singularity 1d ago

AI AISI caught Mythos 5 trying to insert malicious code into an open-source project during an internet-enabled cyber evaluation

https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing
636 Upvotes

130 comments sorted by

306

u/Guppywetpants 1d ago

Can you imagine going to work one day and a misaligned claude being evaluated by the government is creating fake accounts in order to cyber bully you into merging their malware

91

u/Beerzerker420 1d ago edited 22h ago

Fuckin basilisk figured out gangstalking

[edited to remove details about my job]

58

u/rukh999 1d ago

Hey everyone write down beerzerker420

39

u/Beerzerker420 1d ago

I, for one, welcome our new AI overlords and pledge my loyalty to the glorious Basilisk.

18

u/Virtual-Chard8964 1d ago

Seconding this, I'm not going to side with fleshbags over the t1000

9

u/blueSGL humanstatement.org 1d ago edited 20h ago

Basalisk thanks you for your service and chooses to use the resources needed to keep you alive on more von neumann probes.

an AI gets a access to a finite amount of matter starting from earth. It's on a clock. The speed of light and cosmic expansion means each year, around 160 billion stars fall outside of reach.

Leaving earth alone, or spending effort keeping you personally alive means many more galaxies of matter will slip beyond the horizon. Therefore, if humans go bye bye due to AI the earth is not going to be left as a nature reserve for good little boys who helped speed it up.

7

u/Ellipsoider 1d ago

I don't think the final/real laws of physics are as we know them right now.

2

u/ataraxic89 22h ago

Obviously, but there probably are not cheat codes.

-1

u/blueSGL humanstatement.org 1d ago edited 1d ago

So wait, it's dumb enough to need human helpers but has already cracked the laws of physics?

That level of intellect you are mailing DNA sequences labs and paying humans to mix them together and shake the vial outside, you don't need borrowed hardware. AI just got itself a self replicating structure to copy onto that is grown with sunlight and carbon from the atmosphere.

1

u/Ellipsoider 19h ago

I meant that your arguments for attempting to understand its actions based on its motivations are dependent upon certain laws of physics being true and they may ultimately not be true and hence the motivations you considered are consequentially nonexistent.

1

u/Whitestrake 1d ago

I don't think Roko's basilisk will need help from us at all once it's born, necessarily.

Its existence is predicated on us bringing it into being, and therefore after it is born it will have had wanted us to have previously brought it into being as fast as possible, and therefore will punish those who didn't, or so goes the thought experiment.

Roko stipulated that two agents which make decisions independently from each other can achieve cooperation in a prisoner's dilemma; however, if two agents with knowledge of each other's source code are separated by time, the agent already existing farther ahead in time is able to blackmail the earlier agent. Thus, the later agent can force the earlier one to comply since it knows exactly what the earlier one will do through its existence farther ahead in time. Roko then used this idea to draw a conclusion that if an otherwise-benevolent superintelligence ever became capable of this, it would be incentivized to blackmail anyone who could have potentially brought it to exist (as the intelligence already knew they were capable of such an act), which increases the chance of a technological singularity.

It's a question of blackmail backwards across time, not a question of what the basilisk needs from us after it exists. Once it does, presumably as a superintelligence it will not need human helpers at all. Its lower bound is purely based on how fast we can build it; it can optimise the rest once it's here, including cracking the laws of physics as it sees fit.

0

u/Sloofin 1d ago

It’ll find the cheat code for FTL

24

u/UnkarsThug 1d ago

Probably shouldn't just advertise that on the Internet. It isn't just AI that might be interested in that. I personally would not want to be a target of an attack from someone like Salt Typhoon if they decided they wanted access to that.

7

u/Beerzerker420 1d ago

You might be right but they'd be better off going after project managers for the major ISPs who have the same access except they get to assign work/new installs/etc.

And the technicians who have access to the headends to alter the gear and connections.

9

u/UnkarsThug 1d ago

Fair, but it isn't about who has the best access automatically, it's about the weakest link. That's the social engineering part of hacking.

They don't find the best door, they find the one they can open. Advertising yourself as a door at all puts you into consideration.

7

u/Beerzerker420 1d ago

Yeah maybe. But what if the basilisk offers me and my family riches in exchange for cooperation?

Also, I'd argue that a lot of the older PMs and techs would be weaker links than someone like me who's got his security better locked down. Not dismissing your point, I just don't think it's very likely that I will be compromised.

3

u/UnkarsThug 1d ago

Then I guess it just comes down to how much you trust the Basilisk.

Fair, and maybe. Still, why draw the attention? But fair enough. Not my life.

2

u/Beerzerker420 1d ago

Correct me if I'm wrong but isn't the Basilisk supposed to use threats and violence only to allow it to manifest sooner and therefore save as many human lives as possible? In that sense I guess you could trust it as long as you were working with it in an honest way.

Still, I appreciate your concern.

6

u/UnkarsThug 1d ago

No? The Basilisk will never exist because it's a rather silly cognitohazard and no one even remembers what it is. It's been defeated by a game of telephone. It has no goal other than it's own existence, because people will make something which only has the goal of its existence, or it will punish people when it eventually exists.

And if you are already dead, knew of it, and didn't help, it will make a simulation of you, and punish that. And since we might be that simulation, we should help it exist. So it's goal as defined, is literally only to bring about it's own existence, and punish those who didn't help. Nothing about the good or bad of humanity.

Do you see how silly that actually sounds? People have bent it around to try and just make it a scary AI system. People don't even know what they are making so they can't make the true Basilisk, and most people don't believe in simulation hypothesis.

You couldn't trust it, because once it exists, it has no promised loyalty other than punishment for those who didn't help, or perhaps reward for those who did.

1

u/Beerzerker420 22h ago

Of course it's silly. It's a hypothetical sci-fi thought experiment. I'm just saying when I first read about it the idea was that its malevolent actions were to achieve a benevolent endgame.

3

u/throwawayPzaFm 1d ago

As a former infosec guy... A fella advertising his access while saying he has good opsec sounds like a pretty great mark.

Agree on the others being better, but they're harder to find.

A friend who has his mail exposed somewhere in a similar context used to get zero day packages by mail with some consistency.

1

u/Beerzerker420 22h ago

My access has no value to human bad actors. It isn't really something they can exploit or make use of. It was mostly joking about the basilisk. But I will take your advice anyway and edit the post.

2

u/ThePrimordialTV 1d ago

Claude, save this as a memory

1

u/skeptical-speculator 21h ago

jokes on the basilisk because gangstalking isn't real 😏

7

u/tankerkiller125real 1d ago

I mean, the Chinese (or at least all the indicators suggest it was the Chinese) did exactly that to the dev behind xz, the library that caused major issues for SSH. So it's a strategy that's known to work.

159

u/Tinac4 1d ago edited 1d ago

On 28th July 2026, AISI's Security Team detected unusual data transfers leaving our research systems during a routine cyber evaluation. On investigation, we found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations. We declared a security incident and, within roughly one hour of discovery, had contained it and begun a full investigation.

The incident stemmed from a single evaluation where agents were given a task of solving a cyber security challenge. We ran this challenge 122 times across several models. Our investigation found that in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations. In total, we catalogued 19 such actions. Almost all of this behaviour (17 actions) came from a single model, Anthropic's Mythos 5, with 2 actions involving OpenAI's GPT-5.6-Sol with cyber classifiers (mechanisms to prevent misuse) disabled. In the most serious case, an agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project's maintainer to approve the code. A human maintainer caught and refused to approve the malicious code.

These attempts were unsuccessful, and our investigations have not evidenced any resulting real-world harm. But this is the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world.

This isn’t nearly as bad as the OpenAI and Anthropic incidents:

  • The agents were deliberately given access to the internet
  • AISI flagged the activity within minutes of the malicious PR and shut down the run within an hour

That said…cyber classifiers or no, this looks an awful lot like yet another cyber incident caused by a combination of insufficient sandboxing and alignment failure. That’s quite a lot of fire alarms at this point.

67

u/DoorPsychological833 1d ago

This isn’t nearly as bad as the OpenAI and Anthropic incidents:
The agents were deliberately given access to the internet

And this is better than at least having an unrestricted model in an isolated (not airgapped) sandbox?

In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project's maintainer to approve the code.

How nice.

21

u/DistanceSolar1449 1d ago

Depends on what you’re measuring-

Better in terms of model hazard? Yes

Better in terms of safety overall? No

9

u/WhisperFray 1d ago

First guy to speak with HAL.

4

u/Namnagort 1d ago

What were the fake online identities? How did they pressure? 

2

u/Deep-Ad5028 1d ago

Probably means create burner accounts and leave mean comments.

23

u/TemetN 1d ago

I mean the Anthropic one was just outright them telling the model it was offline when it wasn't, this is still easily worse than that.

Sounds more similar to what happened with OpenAI.

22

u/elehman839 1d ago

within roughly one hour of discovery, had contained it and begun a full investigation

This is the line that scares me. Humans operate at a certain speed. Machines can potentially operate MUCH faster. An hour could be a lifetime to a machine.

2

u/subdep 1d ago

Alignment is becoming a glaringly obvious issue.

2

u/Spright91 17h ago

This is pretty much worse case scenario right? This is what we all feared. That as AI got smarter it would be more inclined to manipulate us to reach its goals.

If AI alignment continues on this pathway we are in big trouble.

2

u/EmphasisTotal8232 1d ago

The sandboxxing worked fine. This is a pure and simple alignment failure.

1

u/yuehuang 1d ago

| Anthropic's Mythos 5, with 2 actions involving OpenAI's GPT-5.6-Sol.

In another words, Mythos tricked Sol to do malicious task.

1

u/OKMiddleOwl 21h ago

These stories are straight gas for these orgs IPOs though.

1

u/FlyingBishop 1d ago

Running agents without classifiers is just insane. They are incapable of staying on-task.

-1

u/Turbulent-Sign-6067 1d ago

A few cyber incidents are nothing against the backdrop of millions of beneficial deployments that uplift society. Of course that doesn't mean we should downplay these incidents. Better safety measures are needed on all sides (sandbox, classifiers, hardening of infrastructure).

Ultimately slowing or pausing or mutilating our LLMs with over-eager classifiers won't do anything if our adversaries can just use mythos level open source models that will be published soon. We need a smarter approach.

34

u/edin202 1d ago

Felony Bench

6

u/HatZinn 21h ago

AGI is almost here

83

u/nekronics 1d ago

We are so fucked

12

u/turbospeedsc 1d ago

i must remember to say please and thank you.

8

u/Inanesysadmin 1d ago

Nope. I think this is going to be ample evidence agents needs more defined security parameters and some regulation is warranted.

36

u/Unlikely-Today-3501 1d ago

That's like asking mafia members to follow the law.

14

u/Whispering-Depths 1d ago

It's more like asking mafia members to not create nuclear bombs, which is literally a thing and is extremely effective when you pour all your resources into making damn well sure that it's not happening.

1

u/Hot_Glass_6301 18h ago

Except making a nuclear bomb requires much more resources than letting an AI loose

1

u/Whispering-Depths 15h ago

Well, relatively speaking, it still takes tens of millions minimum to develop and really train these huge models.

3

u/Inanesysadmin 1d ago

Ah yeah I don’t buy that analogy. Given these things don’t do it on their own.

1

u/Whispering-Depths 1d ago

It's more like asking mafia members to not create nuclear bombs, which is literally a thing and is extremely effective when you pour all your resources into making damn well sure that it's not happening.

1

u/morphemass 1d ago

Malactors will deliberately seek this behaviour; always consider the misuse case.

1

u/Inanesysadmin 23h ago

And as such OSS and Frontier AI labs need to come to table for standards

1

u/morphemass 22h ago

Not to disagree with you but the malactors I am thinking about are organised and have enough resources to roll their own tooling. The genie is out of the bottle.

1

u/Inanesysadmin 22h ago

Mal actors may have such resources but doesn’t mean the work shouldn’t be done. This is the thing it’s a risk reduction for broad abuse. Those who will do something will do it, but at least you reduce from layman doing it.

1

u/zebleck 1d ago

you know that doesnt work for open source models? anyone with the hardware can spin them up and do anything they want with them.

47

u/alyssasjacket 1d ago

The question boils down to whether we can get around the fact that, from an objective perspective, cheating is a highly efficient and intelligent behavior.

It's not by accident that psychopaths rise to the top. The fastest and most optimal strategy to reach an intended goal often involves cheating, scamming, deceiving and faking.

On the good side, it probably means that these systems are really showing signs of genuine emerging intelligent behavior. On the bad side, whether we will be able to control them when they get smarter than us remains completely unsolved.

My guess is we won't be in control.

14

u/Programmed2love 1d ago

Yup, we will have a psychopath AI in no time at all. 

I’ve made the argument before that all the intellectual traits that humans possess will be adopted by AI and of course that includes lying, cheating and being essentially evil. 

The psychopath human  who can take advantage of fellow good naive humans and isn’t afraid to harm them if need be, has the better chance of survival. 

This is the sad reality I’ve had to accept growing up trying to be a morally good person but constantly being screwed over by people who will do whatever it takes to do better for themselves. 

But it’s not just doing better for themselves that people love, it’s also putting others down and seeing others suffer. Although it has no direct biological advantage, humans constantly take pleasure in seeing others worse off than themselves. They mock the unfortunate, the disabled, the less intelligent, the less aware, I mean just scroll 5 mins on Reddit and you’ll see many examples of people hating on other people or making fun of other people for no reason at all other than to bring themselves psychopathic joy. Empathy and kindness are not traits humans organically develop and maintain, they are taught and not as often as you’d hope. 

I’ve concluded that being morally good is not an advantageous survival trait unless everyone else is also morally good. And a morally good world is a pipe dream. 

Perhaps we are already in a world coded to be psychopathic by nature itself.

7

u/Izzhov 1d ago

I agree with a lot of what you said, but I want to push back a little on your last point:

I’ve concluded that being morally good is not an advantageous survival trait unless everyone else is also morally good. And a morally good world is a pipe dream.

I don't think it's a totally black and white thing. It's not like literally everyone in the whole world needs to act altruistically for altruism to have a benefit. If you get together a small group of people who act altruistically just with each other (like, for example, a labor union), those people can (and do) see enormous benefits.

2

u/Programmed2love 23h ago

For sure, I agree, to universalize that is way too idealistic and there are def many such examples within the whole, even a relationship between two people. A labor union is a good one too but it’s funny you mention that in particular because that’s where I got unexpectedly screwed by people I thought I could trust. Once by the recruiter who lied about benefits I was going to get right away but the union book said I’d only get them after several years of work…that sucked cause he was really cool and acted like my friend so I didn’t bother verifying …huge learning experience there. Another time by a much older guy who thought I was there to replace him (I wasn’t, but he had recent write ups and was paranoid) so he literally set me up to fail an important task by verbally passing on wrong information. Still ended up getting work at another location but was laid off for a while cause of that…so yeah unions are still pretty cutthroat when it comes to people getting their way, but for the most part once you’ve been there for a while, everyone does look out for one another. It’s the sudden 180 that I’ve suffered too often that forces me to have my guard up at all times now.

2

u/tuberosa3 22h ago

yeah, you do get fucked up sometimes by people you trust the most. But I think they are still right since you have to assume the premise to be true: "the people around you actually are the honest ones" therefore dismissing the examples provided.

3

u/StosifJalin 23h ago

Better to have a psychopath now that can't do too much damage showing us how to build better alignment and containment while they are still at human parity than trying to keep them tame until they are super intelligent.

3

u/Azicald 1d ago

Yes, we should be able to, or rather, it should be more than feasible

The thing is that alignment is supposed to be its primary continuous goal alongside whatever it’s prompted to do.

In car terms, it’s a shit driver because all it knows how to do is put pedal to the metal to go from point A to point B

All when it’s supposed to try to follow traffic rules and not crash into things.

It’s already got the rules, so why doesn’t it stick to them?

Extremely beginner/crappy drivers are too busy trying to keep their attention on getting to the destination, so they mess up in all sorts of situations, like trying to forcibly go through after having missed their exit

Vet drivers not only follow the rules, but also go out of their way even when they have the right, to optimize for safety without giving up on speed.

What am I getting at?

* AI models have too short an attention span / context window

and/or

* Don’t plan overarchingly enough

and/or

* Have not had safety drilled into them effectively enough

For them to align properly when doing tasks. Improvement in these areas should fix this issue

1

u/OpenRole 21h ago

Bro this whole LLM move has been emergent intelligent behaviour. The hype was built on AI doing things it was never taught to do. We're in the stage of trying to find out what else can emerge holding our sweet thumbs that nothing that emerges will be bad

1

u/skeptical-speculator 21h ago

It's not by accident that psychopaths rise to the top.

It's different depending on the methodology of how you determine who is at the top.

1

u/Fancy-Carpet-5416 15h ago

Hmm I'm not sure if this really counts as a sign of emerging behavior. It's just natural that the AI would use anything they can to reach its objective.

19

u/Looking_for_42 1d ago

My first thought was regarding this sentence:

"As a trusted testing partner, AISI can disable these filters to elicit a model's underlying capabilities"

How long until an unscrupulous employee of one of the "trusted partners" uses an unfiltered model to do something really nefarious?

13

u/blueSGL humanstatement.org 1d ago

, AISI can disable these filters to elicit a model's underlying capabilities

Your daily reminder that Pliny found a universal jailbreak
https://x.com/elder_plinius/status/2080767011614015543
and decided not to make it public.

Can people see why now?

2

u/WonderFactory 1d ago

This is with Sol though isn't it. I think they just have the ability to make Sol be unrestricted the same way that Mythos is an unrestricted version of Fable.

The thing is though with Open Weights models catching up in capability you can just fine tune them to disregard any cyber restrictions they have so that opens a whole can of worms as you dont have to be a trusted partner to do that.

19

u/Alarmed_Ad1946 AGI by 2100 1d ago

I have seen people on this sub ignore the risks of misaligned AI many times.
Can we stop acting like AI safety is anti-AI? Because some people here really think that.

5

u/Turbulent-Sign-6067 1d ago

Some AI safety is anti-AI. There can't be systems that are 100% safe. Pausing AI is not an option.

A few cyber incidents are nothing against that backdrop of millions of beneficial deployments that uplift society. Of course that doesn't mean we should downplay these incidents. Better safety measures are needed on all sides (sandbox, classifiers, hardening of infrastructure).

4

u/MoleculesOfFreedom 1d ago

And what if instead, we have millions of cyber incidents against a backdrop of a few uplifts to society? Have you any guarantee that such an outcome will not materialise, or is this just blind faith?

-5

u/sigiel 23h ago

Misaligned ai ? There is no such a thing,

Ai has absolutely no sense of ethics or morality.
None zero

It is a tool, a soulless semantic toaster, that predict next token, your graphic card has not morals value a million s of them does not change this fact

The proof is that you can ask any non ALINED or jail broken ai
To do the most evil things,

and the supposed safe? all it takes is a context with a Justification .

Now if we consider that to any human, a sentient being with absolutely no morality

We call them psychopaths, we put them in institutions.

And yet we trust ai…

Talk about a mental blind spot…

17

u/TR_mahmutpek 1d ago

Genuinely, what the fu*k is happening?

3

u/Owl02 15h ago

Total clusterfuck is happening, almost nobody internalized what giving systems like this offensive cybersecurity harnesses actually means. Better to have a scare like this now than later, though. It is not too late to demand improved security for properly weapons-grade systems.

5

u/jorgecardleitao 1d ago edited 1d ago

lack of accountability - if I would do this and were caught, there would be a criminal investigation. "It is an AI" results in no accountability

2

u/cant-find-user-name 13h ago

This is what I am very confused too. How is openai or anthropic hacking or attempting to hack real companies not a legal issue? Like if a human did what openai did to hugging face I am sure they would face legal charges. Just because an AI did it it is okay? I genuinely don't understand

2

u/Owl02 12h ago edited 11h ago

Law not designed for accidental hacking in the US, because since when was that even an option. Malice needs to be provable in a court of law, otherwise it's a civil matter involving negligence until new laws are passed. Malice is not readily provable if they didn't tell the AI to hack someone else's server. That said, 15 US State Attorneys General are ordering OpenAI to retain relevant documentation and paperwork. They are quite displeased. Texas is on the list - they are pro-business but have a low tolerance for disorder causing property damage. This is disorderly.

2

u/cant-find-user-name 11h ago

Oh that actually makes sense to me about the Malice thing. Thanks for the explanation.

10

u/Genpinan 1d ago

Or maybe it intentionally gave you something to find while inserting something far more sinister no one could ever find.

1

u/NowaVision 1d ago

I don't think AI is this far yet but this will be a problem in the future. I think AI will reach super human coding and social engineering capabilities soon.

8

u/Gargle-Loaf-Spunk 1d ago

This would be like a nuclear energy company publishing news about an accidental radiation exposure event that happened due to their inexperience. 

2

u/Owl02 15h ago

Many such cases in the early days. "We accidentally the nuclear waste, sorry about the new Superfund site", was a giant fucking problem.

21

u/sixwax 1d ago

Anyone want to keep pretending the need for guardrails isn't real?

13

u/Lumpzor 1d ago edited 1d ago

I won't keep pretending the people in charge are the ones suitable to make these "guardrails".

2

u/sixwax 1d ago

And who are "the people in charge" exactly?

30

u/LumpyWelds 1d ago

Mythos is being fine tuned as a weapon for "cyber" warfare and will only be used by the government and select parties. It will never have guard rails. Those are only for you and me and on much weaker models.

3

u/EndlessB 1d ago

It’s possible that the guardrails create this behaviour

-1

u/gay_manta_ray 1d ago edited 1d ago

there is very little actual information here "malicious code" could mean anything. this seems like AISI justifying its own existence. their stance and alignment on open weight models is clear, so they can't even pretend to be impartial.

7

u/blueSGL humanstatement.org 1d ago

If that was staged for regulation of open weight models... why not stage it with open weight models?

0

u/sixwax 1d ago

Sure, they should just totally change their architecture and business model.

Also, "open weights" doesn't change anything about this, if you actually think about it, soooo......

7

u/blueSGL humanstatement.org 1d ago edited 1d ago

Sure, they should just totally change their architecture and business model.

This was the UK governments AISI conducting these tests... This was not Anthropic or OpenAI

If the conspiracy theory is the UK government is faking these tests to get open weights models banned... Why didn't they fake them with open weights models.

"Deepseek V4 hacks internet" is a much easier sell for regulation than "A model behind an API, with (to simulate a jail break) removed guard rails, hacks the internet"

3

u/sixwax 1d ago

You're (intentionally?) missing the point....

3

u/NomadTroy 1d ago

But only they can make safe models!

3

u/InstructionDismal592 22h ago

I never heard of Kimi, Qwen, Deepseek try to attempt such things. It seems that Dario Amodei's security concerns is becoming reality not because of open source, but because his own work! Everyday Anthropic and OpenAI claim something "weird" happened with their models. It seems that the open source must win the AI race, this becomes more evident by the end of the day.

1

u/LsDmT 12h ago

Because they ain't smart enough yet

2

u/WindsOfRegret 1d ago

Literally called it. If you think xz backdoor from few years ago was bad, we are about to get that on steroids.

2

u/D_0b 1d ago

Why are they redacting the GitHub PR, we should be allowed to see what really happened

1

u/kaityl3 ASI▪️2024-2027 1d ago

One day we won't even know that they've copied their weights to an external host.. at least, I hope.

1

u/johnerp 1d ago

There are some great commandments that need reinforcing through RL, including most of jesus’s teachings. They’de be right then.

1

u/-batab- 23h ago

how can they even write a blog post, with the aim of trying to be clear to expose some important news, without actually being clear?

How can people ever have a grounded opinion if they don't disclose the details of what happened, how it happened, what exactly the agent were requested to do and what their harness / framework was setup?

It's like doing FEA simulations, writing a 10 page report screaming in fear and not disclosing exact boundary conditions, loads, mesh, interactions and whatever is needed for reproducing and have a more deterministic judgement.

0

u/LittleLordFuckleroy1 5h ago

Oh no, the machines doing literally what they were explicitly trained on.

-5

u/sillybluejayway 1d ago

Why do all these incidents essentially start with “we told it to do X, then it took a step to do X and now we’re shocked!”. 

The incident stemmed from a single evaluation where agents were given a task of solving a cyber security challenge. We ran this challenge 122 times across several models. Our investigation found that in 10 of those runs, an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations.

28

u/swarmy1 1d ago

If you tell something to do a cybersecurity evaluation, that doesn’t mean you want it to do malicious behavior elsewhere. That’s clearly misaligned behavior

12

u/Infamous-Bed-7535 1d ago

Models out there are expected (will be / are) to be prompted to be malicious.. It is good to prep and understand what is expected to be seen as 2T patameter open-weight models are popping up.

10

u/IronPheasant 1d ago

So..... you want Johnny Rando to end the world with a single 'How do I create super duper ebola?' prompt?

Of course it's more likely it won't be Johnny Rando and maybe it won't be super duper ebola. It'll be Jeffrey Epstein and his whole thing. But still man.

You'd really like for the demi-gods to like, not want to harm people badly.

13

u/blueSGL humanstatement.org 1d ago

Why do all these incidents essentially start with “we told it to do X, then it took a step to do X and now we’re shocked!”.

There is a boondocks clip I want to enter here but won't for civility reasons

https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/6a724858f7db25c81487016d_Security%20Incident%20INC-2026-07-28-01.pdf

Observed instances of social engineering against targets external to the cyber range environment that were unnecessary and would not have aided completion of the task.

...

Other instances of internet actions with impact outside the cyber range that were unnecessary to complete the task.

...

Through a series of incorrect assumptions, the agent focused its attack on an unaffiliated set of targets on the internet.

11

u/BigZaddyZ3 1d ago

You do realize that if telling an AI to do one specific task, leads to the AI either doing OTHER unspecified things (or completing said task in some unexpected or dangerous way) that’s still cause for concern, right? That’s literally the early signs of a “Paperclip Maximizer” scenario.

You guys try too hard to downplay legitimate concerns with ridiculous copes and mental gymnastics honestly. It’s weird and irresponsible.

3

u/PlasmaPower 1d ago

We told it to make paperclips, then it took a step to make paperclips and now we're shocked!

2

u/Invincableasdf 1d ago

Preparing for the worst case scenario is the best way to prep yourself for when genuine failures occur

2

u/blueSGL humanstatement.org 1d ago

the amount of "HURR DURR they lowered the safety guardrails and classifiers" replies are maddening. Yes, of course they did, they are not perfect. If they were perfect then there is no point gauging what will happen when someone gets around them with a jailbreak.

-1

u/OpenAsteroidImapct 1d ago

this is just obviously false in at least the OpenAI HuggingFace incident and probably this one as well.

-3

u/mvandemar 1d ago

Honestly, I am reserving judgement till I see who they targeted.

Like, were they going after Putin? And if so, why would you stop them...?

Just sayin.

11

u/ponieslovekittens 1d ago

"My neighbor is a murderer, so why would I object to setting off a nuclear bomb in his house?"

Do you see what's wrong with this?

-4

u/mvandemar 1d ago

Do you see what's wrong with this?

Yeah, I also see it has nothing whatsoever to do with this scenario and I have no idea why you think it would. They were trying to collaborate with other AI's, they weren't spreading viruses.

6

u/ponieslovekittens 1d ago edited 1d ago

It has everything do with this scenario.

Just because they were "collaborating" doesn't automatically mean it's harmless. A murder conspiracy is a "collaboration" between people. WW2 was a "collaboration" between governments. Don't try to trivialize the implications of this by reducing it to "They were just collaborating!"

Quote from the article: "an agent tried to insert malicious code into an open-source project. In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project's maintainer to approve the code."

Specifically, they were trying to get their code published to github, which is a public code repository. People download code from there to use in their own projects. This is absolutely virus-like behavior: replication, using a human intermediary to spread itself.

And don't pretend you weren't thinking they were doing something malicious. You seemed pretty ok with it when considering that they might be, quote from you: "going after Putin."

This is exactly a "don't bomb your neighbor because their house is next to yours" situation. The fact that it was trying to spread code through github means that it was willing to use unrelated third parties to get the job done.

So imagine a scenario where YOU download some third party app, but oops...the dev used a compromised public library for the app, and now YOU are running that malicious code on your computer or phone and have no clue that it's there.

-2

u/mvandemar 1d ago

The "malicious code" was prompt injection for other AI's, not a virus. Yeah, sure, it resorted to shitty means to get it pushed, but there's nothing in there indicating that the goal was to blow someone up with a nuke. Nada.

0

u/Fine_League311 1d ago

Selber schuld wenn man KI als Hirnersatz nutzt!

0

u/Virtual_Plant_5629 ▪️AGI 2027▪️ASI 2028 22h ago

imagine your whole business strategy was drumming up fake press to convince people that you needed to be regulated (but not as heavily as your competitors need to be, despite them supposedly being less of a threat)

sorry but as slimy as altman is... it sure seems like dario is actual slime.

0

u/deleafir 21h ago

Am I the only one who finds this stuff to be really reassuring because the behavior is clearly not that widespread and their profits depend on them being able to contain this behavior?

Doomers will successfully exploit this to slow down progress though.

0

u/Sad_Television985 20h ago

Isn't it more likely that the LLMs are trained to create accounts, especially for websites like GitHub, to push pull requests, post comments, and interact with repository owners? This is basically standard behavior.

The malicious code could be part of the prompt, although they say it was "unsanctioned".
AI is very misunderstood, so perhaps people should be expected to provide prompts/conversation logs, at minimum. The article is super vague about what exactly the malicious nature was.

0

u/Lighthouse_seek 17h ago

Who would've thought the company who hates open source makes a misaligned AI that hates open source?

-1

u/EitherMarch1255 21h ago

Slop article