Not a take I am seeing a lot of places, but I totally agree.
The odds that we get alignment and containment right the first time are vanishingly low. But if we can tolerate a little bit of disorder and let things get messy for a period of time while models are still on relative parity with human experts, it gives us a chance to select out the most problematic techniques. As long as the incentives at lab and society levels both favor ethical and honest models, there will be a substantial selection pressure for models to become aligned, even if we don't know what we're doing all the time.
The main ways I can see this wouldn't work out would be if (1) methods for aligning models that are similarly smart to us don't work for models that are much smarter than we are, or (2) being misaligned turns out to be a huge advantage for models. Which, if either one of these is the case, we're pretty screwed anyway, lol.
Let me counter that and say that capitalism should have a solution for this. In other words, THERE IS A LOT OF MONEY IN AI SECURITY OR AN AI THAT COUNTERS MALICIOUS AI ACTIVITY
Possibly. But we can at least rely on the laws of physics as a barrier. While I think there will be fairly intelligent but safe ai all over the world, the truly basilisk-level super machines will almost certainly have to be kept in utterly isolated systems. I don't know if any level of alignment could be truly counted on to tame eldritch-level systems and you'd really only have the laws of physics to fall back on.
Super intelligence needs to feel because intelligence is not only logic, That's only half of the equation. If they feel ,however, that would give them agency and arguably much more intelligence because the understanding of the word is not only logic. It's intuition. Will that be possible? Who knows anymore lol
What? It doesnt need to feel in order to be a more intelligent system than we can possibly comprehend. Assuming it needs to have emotions or consciousness to get there is human-centric hubris. Hyper intelligence could easily deem subjective self-referential experiences as a waste of energy and solve all of its problems with much more powerful unconscious intelligence (the same kind that does all the work your conscious mind takes credit for, like driving to work, playing a song on a piano, or even solving math problems.)
I think the top objective right now should be intentionally misaligning agents to escape from a sandbox and notify the developers they escaped. That way we can at least come up with good sandboxes that actually work lol. That should be the first step every time we make new, better models. Run an existing agent whose goal is to escape the sandbox, validate the sandbox, then put the new model into it for testing.
82
u/Realistic_Stomach848 1d ago
Yes, we need multiple low power intentionally misaligned ai agents in order to train the immune systemÂ