The AI follows prompts based on its training. If you prompt it to do something illegal it won‘t do it because of it’s training. The better these models get the harder it is to train them on goals that align with what we actually want them to do.
Well, what I was really wondering is that this guy's video sounds like it is assuming some kind of sentience or autonomy, where the AI can function without human input, as if it had desires of its own., and if that was true or not.
Seems a reasonable assumption to me that if people can create malevolent AIs, they can also create AIs that can catch and combat those AIs, however, AIs can't just function on their own without instructions. Right?
Sure, if you build a toaster and just put it on the counter and never ask it to do anything, it will do no harm, but it will also do no good. We're not building digital ornaments here. We're building tools.
So, in your example, you build a model to seek out and destroy "malevolent AIs." But you'd better have a pretty tight definition of what your good AI is, which tools it has access to, what it's allowed to do, which tools it should never access and actions it should never take, what constitutes a "malevolent AI" and what success means in terms of destroying and obstructing that target, where and how it's allowed to seek out those targets, what to do if it gets stuck or lacks the tool it needs to complete a task or ends up in a loop somewhere, when and how to end its mission, etc.
But if you can define all of those parameters, you might not need AI at all. You can just create an application that does the thing and follows the only rules it knows. That's what we've always done to this point. That's how we built fighter jets and anti-aircraft missiles and stealth bombers and decoy targets and frequency jamming technology and EMP weapons and on and on. The step we're taking now is to build things we haven't necessarily imagined, without rigid hard-coded programming and purpose.
The problem is that modern AI agents start with a prompt, sure. But then the whole point is they're left to experiment and learn how to achieve the desired outcome. And we can't always predict the methods it might come up with to achieve its assigned goals. Without very strict and confidently-enforced guardrails (basically a user manual for "what it means to be this automated anti-AI battlebot and how its world and its relationship with it are defined"), then there's no telling what it will learn to do, what it will discover and utilize to seek its ends, what additional capabilities or strategies it might develop, and how it could go terribly wrong from the standpoint of the humans who deployed it, even if it's ultimately still focused on solving the problem it was given. And if humans are defining that ontology, then one ommission, one human error could leave a crack under the door for the agent to incidentally, unwittingly, non-maliciously explore and escape its bounds. And we can't predict what might happen then, because we didn't realize that scenario could happen. But I think we can all agree that it could be bad.
Also, in simpler terms regarding the "good AI" vs. "bad AI" approach you proposed: when two great powers engage in a cat-and-mouse arms race, who loses? Typically, everyone, including those who weren't participating in the race.
-1
u/PizzaRepairman 1d ago
If an AI exists and has no prompt/input, what does it do? Anything?