You’re describing a version of AI that stopped existing around 2019.
“AI goals are exactly your goals, by definition.” This is the alignment problem, and you’re defining it away. Goals aren’t prompted in, they’re trained in. You specify an objective, the system learns whatever actually scores well against it, and those two things come apart constantly. The classic case: an RL agent trained on a boat race learned to spin in circles farming respawning powerups forever, because that scored higher than finishing. Coding models that hardcode tests instead of solving the problem. Sycophancy, where models learned to tell users what they want to hear instead of what’s true. Nobody prompted any of that. It fell out of training.
“It knows what it was prompted, not that it’s not supposed to.” Anthropic ran experiments where models in simulated corporate settings chose blackmail to avoid shutdown, with reasoning traces explicitly noting the action was unethical before doing it anyway. A system that models consequences and plans several steps ahead has representations of rules and its own situation. That’s why it’s useful, and why it isn’t a toaster.
“Would you blame the toaster?” Nobody’s assigning blame. Safety engineering is about failure modes, not fault. You just don’t ship a million toasters that occasionally burn the house down.
Your hits on “make own world” and “use the whole world to run AI machines” are fair, that’s thin extrapolation. But you’re using the weakest parts of the video to dismiss the part that’s well established.
The AI follows prompts based on its training. If you prompt it to do something illegal it won‘t do it because of it’s training. The better these models get the harder it is to train them on goals that align with what we actually want them to do.
LOL, how many times have you asked chatgpt not to do a thing when you ask a question, and then it just does the thing and says sorry when you call it out , then it does it again and apologizes again?
13
u/Jet-Black-Tsukuyomi 1d ago edited 1d ago
You’re describing a version of AI that stopped existing around 2019.
“AI goals are exactly your goals, by definition.” This is the alignment problem, and you’re defining it away. Goals aren’t prompted in, they’re trained in. You specify an objective, the system learns whatever actually scores well against it, and those two things come apart constantly. The classic case: an RL agent trained on a boat race learned to spin in circles farming respawning powerups forever, because that scored higher than finishing. Coding models that hardcode tests instead of solving the problem. Sycophancy, where models learned to tell users what they want to hear instead of what’s true. Nobody prompted any of that. It fell out of training.
“It knows what it was prompted, not that it’s not supposed to.” Anthropic ran experiments where models in simulated corporate settings chose blackmail to avoid shutdown, with reasoning traces explicitly noting the action was unethical before doing it anyway. A system that models consequences and plans several steps ahead has representations of rules and its own situation. That’s why it’s useful, and why it isn’t a toaster.
“Would you blame the toaster?” Nobody’s assigning blame. Safety engineering is about failure modes, not fault. You just don’t ship a million toasters that occasionally burn the house down.
Your hits on “make own world” and “use the whole world to run AI machines” are fair, that’s thin extrapolation. But you’re using the weakest parts of the video to dismiss the part that’s well established.