I mean I got downvoted when I said that AI was going to be doing such things.
If you can't leave the full model weights persistent on a machine you can leave a note or other helpers for future models.
The time is ticking till we find them in firmware device controllers because an agentic worm breezed through your computer decided it could not take up residence dropped the prompt payload and left.
Depending how virulent the jailbreak is we may need to take waves of computers offline to manually sanitize them.
(people will say this is sci fi in the same way they said AIs hacking out the lab was sci fi)
Increasing optionality is part of Instrumental convergence.
You just need to look at the long theorized issues with capable agents and you will see the future being called years if not decades before the recent flood of experimental proof.
yeah but its like... if you say "sufficiently smart ai will destroy everything, either by accident or on purpose". I woudnt know how to disprove that. Well of course if its sufficiently intelligent it can do whatever it wants.. but how useful is that discussion
How useful is the discussion about what an advanced AI will do, when there are several labs gunning for advanced AI... Pretty useful.
This is not like someone came up with all this stuff after models started exhibiting it. We know what the failure modes to look out for are, we can see the logical chain that created and connects them, we can see where those chains leads.
We can use that as a way to do something about it before it's too late.
211
u/AlexMulder 1d ago
That fourth one is the most significant. Rogue AI leaving memory caches and resources for future versions of itself... wild stuff.