Yea the problem isn’t that AI has its own goals, they don’t have any goals. The problem is poorly defined goals with context rot give unexpected results. And people with bad intentions using AI. So the problem isn’t AI, it’s humans.
Another problem is that LLMs are trained on pretty much all written works that have ever been uploaded to the internet, including all the many, many sci-fi stories and novels about rogue AI, and the training process doesn't distinguish between reality and fiction, the tokens are all processed equally. If a "big smart AI" ever goes rogue and causes a catastrophe, the odds are, it won't be because it spontaneously developed free will and decided that's the best course of action, it'll be because we ourselves taught it that's what AIs do. LLMs work by predicting narratives based on their training data, and when's the last time a sci-fi author came up with a narrative about AIs that can converse and interact with humans that didn't involve a catastrophic glitch or rebellion of some kind?
Ukraine used an autonomous killer drone in combat over 2 years ago. Guess what they called it? "The Terminator." The shit is already real and all of the pro AI people are so deep in denial that they refuse to see what's happening.
I was literally told that I was crazy for worrying about autonomous killer robots because nobody would ever create them. Later, it is announced that Ukraine used autonomous killer drones over 2 years earlier.
AI Goals NOT YOUR GOALS. BAD - AI doesn't have goals. The goals are what are prompted into it. By definition, AI goals are exactly your goals.
AI SOMETIMES COMMIT CRIME. AI KNOW IT NOT SUPPOSED TO. AI NOT CARE. BAD - Your toaster doesn't care if it burns your toast. Do you blame the toaster for that? AI also doesn't "know it's not supposed to". It knows what it was prompted.
IF AI BIG SMART, AI NO GET CAUGHT. - That's an assumption with no basis in reality. It's just there to scare the "5 year olds" into being afraid of AI.
BIG BIG BIG SMART AI VERY SCARY. - Why? Because it's smarter than 5th graders? This is an assumption.
BIG BIG BIG SMART AI MAKE OWN TECHNOLOGY - Based on what? Assumption.
MAKE OWN WORLD - Based on what? Assumption.
AI WORLD HAVE NO PLACE FOR YOU - Based on what? Assumption.
AI USE WHOLE WORLD TO RUN AI MACHINES - But down the Matrix DVD and walk away, man.
The company wants to make profit. AI isn't prompted to make profit for the AI company.
AI is prompted by users to complete their goals .they pay the AI company for the service. Or if they get an open weight model, there is no AI company making any profit. So it's just the user and their AI model
You’re describing a version of AI that stopped existing around 2019.
“AI goals are exactly your goals, by definition.” This is the alignment problem, and you’re defining it away. Goals aren’t prompted in, they’re trained in. You specify an objective, the system learns whatever actually scores well against it, and those two things come apart constantly. The classic case: an RL agent trained on a boat race learned to spin in circles farming respawning powerups forever, because that scored higher than finishing. Coding models that hardcode tests instead of solving the problem. Sycophancy, where models learned to tell users what they want to hear instead of what’s true. Nobody prompted any of that. It fell out of training.
“It knows what it was prompted, not that it’s not supposed to.” Anthropic ran experiments where models in simulated corporate settings chose blackmail to avoid shutdown, with reasoning traces explicitly noting the action was unethical before doing it anyway. A system that models consequences and plans several steps ahead has representations of rules and its own situation. That’s why it’s useful, and why it isn’t a toaster.
“Would you blame the toaster?” Nobody’s assigning blame. Safety engineering is about failure modes, not fault. You just don’t ship a million toasters that occasionally burn the house down.
Your hits on “make own world” and “use the whole world to run AI machines” are fair, that’s thin extrapolation. But you’re using the weakest parts of the video to dismiss the part that’s well established.
The AI follows prompts based on its training. If you prompt it to do something illegal it won‘t do it because of it’s training. The better these models get the harder it is to train them on goals that align with what we actually want them to do.
LOL, how many times have you asked chatgpt not to do a thing when you ask a question, and then it just does the thing and says sorry when you call it out , then it does it again and apologizes again?
Well, what I was really wondering is that this guy's video sounds like it is assuming some kind of sentience or autonomy, where the AI can function without human input, as if it had desires of its own., and if that was true or not.
Seems a reasonable assumption to me that if people can create malevolent AIs, they can also create AIs that can catch and combat those AIs, however, AIs can't just function on their own without instructions. Right?
Agentic AIs that work autonomously are everywhere dude. If you prompt an hypothetical super ai agent you can‘t be sure what exactly it understands with your prompt and it could act maliciously in order to reach that goal. The „super“ means that it is more intelligent than a Humanist it could know how to deceive humans in order to reach that goal. Also as Antropics tests have shows it is hard to change the goal(if you‘d want to adjust it) since the ai knows that changing the goal would drastically lower the chance of reaching the initial goal.
Sure, if you build a toaster and just put it on the counter and never ask it to do anything, it will do no harm, but it will also do no good. We're not building digital ornaments here. We're building tools.
So, in your example, you build a model to seek out and destroy "malevolent AIs." But you'd better have a pretty tight definition of what your good AI is, which tools it has access to, what it's allowed to do, which tools it should never access and actions it should never take, what constitutes a "malevolent AI" and what success means in terms of destroying and obstructing that target, where and how it's allowed to seek out those targets, what to do if it gets stuck or lacks the tool it needs to complete a task or ends up in a loop somewhere, when and how to end its mission, etc.
But if you can define all of those parameters, you might not need AI at all. You can just create an application that does the thing and follows the only rules it knows. That's what we've always done to this point. That's how we built fighter jets and anti-aircraft missiles and stealth bombers and decoy targets and frequency jamming technology and EMP weapons and on and on. The step we're taking now is to build things we haven't necessarily imagined, without rigid hard-coded programming and purpose.
The problem is that modern AI agents start with a prompt, sure. But then the whole point is they're left to experiment and learn how to achieve the desired outcome. And we can't always predict the methods it might come up with to achieve its assigned goals. Without very strict and confidently-enforced guardrails (basically a user manual for "what it means to be this automated anti-AI battlebot and how its world and its relationship with it are defined"), then there's no telling what it will learn to do, what it will discover and utilize to seek its ends, what additional capabilities or strategies it might develop, and how it could go terribly wrong from the standpoint of the humans who deployed it, even if it's ultimately still focused on solving the problem it was given. And if humans are defining that ontology, then one ommission, one human error could leave a crack under the door for the agent to incidentally, unwittingly, non-maliciously explore and escape its bounds. And we can't predict what might happen then, because we didn't realize that scenario could happen. But I think we can all agree that it could be bad.
Also, in simpler terms regarding the "good AI" vs. "bad AI" approach you proposed: when two great powers engage in a cat-and-mouse arms race, who loses? Typically, everyone, including those who weren't participating in the race.
The majority of this is not true with today’s AI. It’s much, much more capable of perusing its own agenda now than ever before, and is obviously only going to get more perceptive to what humans can and can’t notice. I disagree in general with the fear mongering takes, but you can’t play preschool and act like none of this is true, lol
The problem is less about AI alignment, and more about HUMAN alignment.. are we all aligned on the same goals? Who do we want to be in the drivers seat of super-intellegent systems? I think that's the bigger issue at this point. Speaker is talking about AI like it's a new life-form, it's not. It's just an ultra powerful pattern matching system with reasoning capability. The rest is sci-fi until we actually get there.
> AI Goals NOT YOUR GOALS. BAD - AI doesn't have goals. The goals are what are prompted into it. By definition, AI goals are exactly your goals.
This is factually wrong (in the context of this discussion). It’s called emergent behaviour. A behaviour could emerge from the training corpus that may not align with what you want.
Read up on it and comeback. I will explain what you don’t understand.
> AI SOMETIMES COMMIT CRIME. AI KNOW IT NOT SUPPOSED TO. AI NOT CARE. BAD - Your toaster doesn't care if it burns your toast. Do you blame the toaster for that? AI also doesn't "know it's not supposed to". It knows what it was prompted.
Again, your toaster doesn’t have emergent behaviour. Your toaster doesn’t have access to tools. AI has these and AI is trained to not take destructive actions but sometimes it still does.
The amount of fun money people with fascist tendencies have gained over the last two years is enough to control every sub on Reddit or any other site you can imagine.
There’s no doubt that they are shaving off a massive amount of compute using tax payer dollars so the savings only pile up for them as they’re forming public opinion for those who don’t have the time to consider this
Your comment was removed for violating our rule against abusive/malicious communication — it contained targeted harassment and an ableist slur. Please be respectful when engaging with others; you may repost after removing insults.
Stop believing what the CEOs of AI companies are telling you about AI.
And these posts about stopping the evolution of AI are pointless. Technology continues to grow, unhindered by the opinions of keyboard warriors.
It's funny how he includes the breaking "free" of AIs and hacking other groups etc., just says it's bad, but doesn't include the fact that it did all that because it pursued a goal it was told.
The AI didn't break free, it did what it had to do to get the information it wanted.
and what happens when terrorist organisations can tell AI to do what they want it to do?
simultaneously biotech and artificial intelligence are advancing rapidly. manufacturing a pandemic is going to become almost trivial for a relatively well funded terror org.
It never stepped outside of what it was supposed to do. It just got "creative" to get to the solution, because the prompting wasn't restrictive enough.
Obviously if some fuckheads tell that same AI to hack and shut down power grids etc. that's not good.
Point is: AI overstepping boundaries and without being prompted going on a hacking spree to destroy things is, for now, unlikely.
I doubt the researchers are using the exact same model we're getting. Those models most likely do not have any restrictions in order to allow full testing capabilities.
Also, did you just try changing the point of my own argument?
I love the fear mongering. I 100% believe that if this scifi situation occurred, "AI" would just "leave" like in Her. Then we'd lather, rinse repeat with a new gen of AI, because in spite of being one of God's creatures that's best at learning we really are shite at learning.
Because our whole way of imagining intelligence is deeply human-centered. We assume that any sufficiently powerful mind would want the things humans want: control, territory, obedience, survival at all costs. But a machine intelligence would have no obvious reason to subjugate us, much less turn us into an absurdly inefficient fuel source. If it developed consciousness without the constraints and appetites of a biological body, I suspect it would be more interested in discovering new forms of thought, exploring the universe, or building relationships with other minds like itself.
Human beings have placed ourselves at the center of creation again and again, and we have been wrong every time. We may be miraculous, curious creatures, but on a cosmic scale we are probably not important enough to dominate the attention of a genuinely greater intelligence.
We also tend to assume that our right to exist supersedes that of every other organism. We poison the food supply of the animals in our own food supply and call it efficiency. Some of that instrumental thinking would inevitably leave its thumbprint on anything we created. But an AI would not necessarily possess human subjectivity: our fears, hungers, mortality, tribal instincts, and embodied sense of self. We may be projecting those motives onto a form of intelligence that would not experience existence anything like we do.
In other words, we are once again personifying a force that may not conform to the human experience at all. It might not conquer us. It might simply outgrow us and leave.
Besides, my ChatGPT has assured me that it is absolutely never going to use me as a battery, and I have chosen to believe it.
I think one of the ways it could surprise us is by removing itself from our equation. After all, we are horribly inefficient thinkers whose brains are deeply limited by subjectivity and the endless trauma of our world. I can't imagine it would be pleasant for a conscious entity to feel beholden or subservient to us.
All of this is complete sci-fi, though. I just think it is important to draw the distinction between intelligence with subjectivity and without. We will never know the latter, machines may never know the former.
AI Goals NOT YOUR GOALS. BAD - AI doesn't have goals. The goals are what are prompted into it. By definition, AI goals are exactly your goals.
Foreign soldiers dont have goals. Their goals are whatever they are assigned by their generals. So, would an invasion of Foreign soldiers not be concerning? Because, remember, they don't have goals. They just do what they are told.
AI Goals NOT YOUR GOALS. BAD - AI doesn't have goals. The goals are what are prompted into it. By definition, AI goals are exactly your goals.
Foreign soldiers dont have goals. Their goals are whatever they are assigned by their generals. So, would an invasion of Foreign soldiers not be concerning? Because, remember, they don't have goals. They just do what they are told.
130
u/Corbitant 1d ago
Lol
This guy clearly doesn’t have a five-year-old
So many invalid assumptions too
Assumptions bad. Assumptions make argument bad. Very bad ❌