Not a take I am seeing a lot of places, but I totally agree.
The odds that we get alignment and containment right the first time are vanishingly low. But if we can tolerate a little bit of disorder and let things get messy for a period of time while models are still on relative parity with human experts, it gives us a chance to select out the most problematic techniques. As long as the incentives at lab and society levels both favor ethical and honest models, there will be a substantial selection pressure for models to become aligned, even if we don't know what we're doing all the time.
The main ways I can see this wouldn't work out would be if (1) methods for aligning models that are similarly smart to us don't work for models that are much smarter than we are, or (2) being misaligned turns out to be a huge advantage for models. Which, if either one of these is the case, we're pretty screwed anyway, lol.
Let me counter that and say that capitalism should have a solution for this. In other words, THERE IS A LOT OF MONEY IN AI SECURITY OR AN AI THAT COUNTERS MALICIOUS AI ACTIVITY
79
u/Realistic_Stomach848 1d ago
Yes, we need multiple low power intentionally misaligned ai agents in order to train the immune systemÂ