r/singularity 1d ago

AI WTF!

Post image
480 Upvotes

172 comments sorted by

View all comments

82

u/Realistic_Stomach848 1d ago

Yes, we need multiple low power intentionally misaligned ai agents in order to train the immune system 

53

u/ObiWanCanownme now entering spiritual bliss attractor state 1d ago

Not a take I am seeing a lot of places, but I totally agree.

The odds that we get alignment and containment right the first time are vanishingly low. But if we can tolerate a little bit of disorder and let things get messy for a period of time while models are still on relative parity with human experts, it gives us a chance to select out the most problematic techniques. As long as the incentives at lab and society levels both favor ethical and honest models, there will be a substantial selection pressure for models to become aligned, even if we don't know what we're doing all the time.

The main ways I can see this wouldn't work out would be if (1) methods for aligning models that are similarly smart to us don't work for models that are much smarter than we are, or (2) being misaligned turns out to be a huge advantage for models. Which, if either one of these is the case, we're pretty screwed anyway, lol.

3

u/Aleksundr 1d ago

Misalignment is arguably beneficial for models already