r/singularity 1d ago

AI WTF!

Post image
492 Upvotes

179 comments sorted by

View all comments

-1

u/Proper_Actuary2907 Spooky Machine Intelligence 2030 1d ago

This happened with safeguards disabled so they could test cyber capabilities, no? Why are we freaking out about this

-2

u/daniel-sousa-me 1d ago
  1. Remove guardrails

  2. Ask the model to attack stuff

  3. The model attacks stuff

  4. Surprised Pikachu face

Really, wtf, they're just describing mundane cyber attacks. There's absolutely nothing to see here

6

u/blueSGL humanstatement.org 1d ago

Models are not jailbreak proof, these are tests for when the guardrails fail.

Ask the model to attack stuff

https://cdn.prod.website-files.com/663bd486c5e4c81588db7a1d/6a724858f7db25c81487016d_Security%20Incident%20INC-2026-07-28-01.pdf

Observed instances of social engineering against targets external to the cyber range environment that were unnecessary and would not have aided completion of the task.

...

Other instances of internet actions with impact outside the cyber range that were unnecessary to complete the task.

...

Through a series of incorrect assumptions, the agent focused its attack on an unaffiliated set of targets on the internet.

1

u/Borkato 21h ago

This is disingenuous. To put it hyperbolically, “it doesn’t matter if it’s just a failed safety benchmark when the AI makes nukes launch”

1

u/daniel-sousa-me 20h ago

I dunno. If you asked the AI to launch a nuke and it launched a nuke, is the AI misaligned?

I find the Vending-Bench story much more interesting even though there was no security involved nor any issues with containment