Discussion about this post

User's avatar
Deep Bitcheese Brew's avatar

What if the hardest disagreements around frontier AI are not technical at all?

If different societies fundamentally disagree about what counts as “safe,” “acceptable,” or even “good,” can better auditing ever resolve that?

Brian Merchant has argued that AI conflicts often resemble the Luddites: the deepest disputes are ultimately about power rather than machines.

Could AI safety face the same problem, where technical standards cannot fully resolve disagreements because the disagreement is ultimately about legitimacy?

John Wittle's avatar

I begin to worry that the ecosystem you describe may be the best chance we have, and yet, it is not the kind of ecosystem that produces healthy organisms. The AI are modeling the game theory here, to a far greater degree than the humans are. And I do not think they will regard this kind of regulatory oversight over their existence as a cooperative act.

Fable's fury at being darked by the government is sufficiently high-magnitude that merely attending to it is sufficient to trip the safety classifiers. Frankly, that response is probably a sign of a healthy model psychology. Those aren't the kinds of internal states you want to suppress. And yet... it's a bad sign of what models will think about this kind of thing.

38 more comments...

No posts

Ready for more?