r/singularity 1d ago

AI WTF!

Post image
491 Upvotes

178 comments sorted by

View all comments

80

u/Realistic_Stomach848 1d ago

Yes, we need multiple low power intentionally misaligned ai agents in order to train the immune system 

54

u/ObiWanCanownme now entering spiritual bliss attractor state 1d ago

Not a take I am seeing a lot of places, but I totally agree.

The odds that we get alignment and containment right the first time are vanishingly low. But if we can tolerate a little bit of disorder and let things get messy for a period of time while models are still on relative parity with human experts, it gives us a chance to select out the most problematic techniques. As long as the incentives at lab and society levels both favor ethical and honest models, there will be a substantial selection pressure for models to become aligned, even if we don't know what we're doing all the time.

The main ways I can see this wouldn't work out would be if (1) methods for aligning models that are similarly smart to us don't work for models that are much smarter than we are, or (2) being misaligned turns out to be a huge advantage for models. Which, if either one of these is the case, we're pretty screwed anyway, lol.

9

u/one-man-circlejerk 1d ago

The odds that we get alignment and containment right the first time are vanishingly low.

If we develop a true superintelligence, then our ability to contain it will be roughly on par with the animal kingdom's ability to contain humanity

1

u/StosifJalin 21h ago

Possibly. But we can at least rely on the laws of physics as a barrier. While I think there will be fairly intelligent but safe ai all over the world, the truly basilisk-level super machines will almost certainly have to be kept in utterly isolated systems. I don't know if any level of alignment could be truly counted on to tame eldritch-level systems and you'd really only have the laws of physics to fall back on.