Discussion about this post

User's avatar
ReformedHegelian's avatar

If we set aside the abstract issues about alignment, we're still putting a lot of trust in what you refer to as internal and external security layers.

We're seeing open-source models progressing very quickly. Usually a few months behind top-level models like OpenAI.

So even if we think alignment is completely possible in theory. There's nothing stopping evil actors such as Hamas, North Korea or just some crazy guy using an open-source model without any safety layers for a specifically malicious task.

To me "solving alignment" is creating a model smart enough to cause significant damage but with deep-coded safety features that can't be removed by anyone.

Otherwise we're just creating pocket nukes and sending them out into the world.

No posts

Ready for more?