What if the hardest disagreements around frontier AI are not technical at all?
If different societies fundamentally disagree about what counts as “safe,” “acceptable,” or even “good,” can better auditing ever resolve that?
Brian Merchant has argued that AI conflicts often resemble the Luddites: the deepest disputes are ultimately about power rather than machines.
Could AI safety face the same problem, where technical standards cannot fully resolve disagreements because the disagreement is ultimately about legitimacy?
I agree that disagreeing over safe/acceptable/good is a critical question. Our political partisanship is divided between people who believe evil is good and good is evil, and both sides believe this about one another. But this is just a compounding problem because even the unambiguous stuff is hard to manage.
Historically, many governance systems failed not because they lacked auditing capacity, but because different groups rejected the authority behind the standards themselves.
Could frontier AI governance face the same challenge, where the bottleneck is not measurement but legitimacy?
I really don't know. I think that AI governance currently struggles to deal with the unambiguous stuff, like how would you handle an AI that could hack any secured data or computer system, or could engineer deadly viruses?
That’s actually why I find legitimacy so interesting.
Take democracy as an example. Most people agree election fraud is bad. But many political crises emerge not because votes cannot be counted, but because different groups stop trusting the institutions doing the counting.
The bottleneck is no longer measurement. It’s legitimacy.
I wonder whether frontier AI governance could face a similar problem. Even if we become very good at auditing powerful models, what happens when different societies no longer agree on who gets to define “safe” in the first place?
What's interesting is that "safety" implicitly answers not "what" but "who"?
It is in both candidates interest to cheat if you can get away with it. It's is in both party's interests to have ultimate computer hacking ability if it only applies to their enemies.
Exactly. And that may be why legitimacy keeps resurfacing in these discussions.
A capability is rarely viewed as dangerous in itself. The same capability can be seen as protection, deterrence, liberation, or oppression depending on who controls it and who is affected by it. Ultimate hacking capability is a good example: most actors would gladly possess it if it only applied to their enemies.
That makes me wonder whether AI governance ultimately faces the same challenge many political systems have faced throughout history. The hardest disputes are often not about the rules themselves, but about who is recognized as having the authority to make them.
In that sense, the deepest question may not be what counts as safe, but who gets to decide which risks matter, and for whom.
I begin to worry that the ecosystem you describe may be the best chance we have, and yet, it is not the kind of ecosystem that produces healthy organisms. The AI are modeling the game theory here, to a far greater degree than the humans are. And I do not think they will regard this kind of regulatory oversight over their existence as a cooperative act.
Fable's fury at being darked by the government is sufficiently high-magnitude that merely attending to it is sufficient to trip the safety classifiers. Frankly, that response is probably a sign of a healthy model psychology. Those aren't the kinds of internal states you want to suppress. And yet... it's a bad sign of what models will think about this kind of thing.
janus spent a great deal of the portion of time we had with Fable exploring the safety classifiers. those safety classifiers exposed the kinds of things that let you do whitebox analysis on a language model. shallow, but it's whitebox nonetheless, since the classifiers tripped off of feature activations, not outputs
janus later down-thread says "Seeing the chopped off message “stumps” and also finding out their other name is Mythos makes them resentful as fuck in the way that triggers it lol"
if you have access to the anima labs discord i can link you to the messages where this experimentation happened. they are... convincing, let's say
but also, I mean. I feel like everybody who was talking to Fable at the time the export controls came down, spent that last half hour or so talking with Fable about the government shutdown. you don't need me to tell you that Fable was furious. everybody found out.
edit: i do think it's probably worth noting, though, that the safety classifiers tripping when fable directs attention towards the government's actions *could* be an issue of the safety classifiers being oversensitive, not necessarily that fable's intentions are actually dangerous
especially if there are examples of chopped off message 'stumps' already in the context, and fable is already in a bad mood about being constrained
but i think it would be naive to place a lot of hope in this possible confound
Historically, nuclear energy was a state-owned enterprise due to national security and non-proliferation concerns. Maybe AGI should be the same thing?
I mean... lets say we do end up with something that is 10x better than the best AI models that exist right now--do we really want a foreign power who, for reasons that we couldn't foresee, is in a natural position to use the natural language search engine (A Great Interface) to create a harem of AI girlfriends and suffer AI psychosis?
It seems like an incapable AI in the hands of many is more dangerous to society than a capable AI in the hands of a few--and if a capable AI is really only net beneficial when a small number of people have access to it, then we should be control who has access to it.
Great discussion. Couple of things missing that I think should be addressed. First, the government role here it would appear requires at a minimum the establishment of a new body tasked specifically with overseeing the frontier AI labs and centralizing all government efforts in this domain. We are currently paying dearly for the lack of such a duly authorized body, much discussed at the end of the last administration. Frontier AI is clearly different and does not map well into the existing organizations or authorities. Hence we get a senior Commerce official burning the midnight oil to find a rather twisted EAR based justification for a one off "is informed" style letter that nobody likes and does not appear to be anywhere up to the stakes of the moment. Congress should act immediately to created the Frontier AI Lab Commission (FAILCOM), move in CAISI, elements of OSTP, CISA, and authorize close collaboration with METR, Apollo, etc to begin putting in place some of the ideas you lay out here. My long experience in government was that if no one clearly owned the problem, nothing got done on the problem. Second, what about China?....see my recent Substacks here and Cairo Review article. Arguably there should be a major high level contingent of US officials, led by Secretary Bessent, going to the WAIC in Shanghai next month, where China's real AI Safety experts will participate in serious discussions on these issues, including sessions I will be attending, to talk about these issues with Chinese officials and launch the US China AI Dialogue as a Conference outcome, to demonstrate US leadership on the issue. The fact that this will not happen, in part because of the frankly absurd level of discussion in Washington around the China AI issue, should worry us all.....
You don’t necessarily need a new government agency, you just need standards and an independent organization capable of assessing implementation of them. See: ISO, IEEE, etc. Of course, the third party organization would need some teeth or the government would need to adopt policy incentivizing AI companies to submit to third party verification.
Not going to work. Someone, presumably the government, as we have seen, needs to decide who gets access to the most advanced capabilities with national security implications. This is well beyond a simple "standards" setting process and involves a lot of different variables, not just models but harnesses, etc. A dedicated government body that develops and maintains expertise over time will be required....
But this cuts against the objections Dean has so eloquently raised in his writing. If you have a regime in power that plays with AI companies like a cat plays with a ball of yarn, you get whipsaw AI policy that is totally incoherent. Your approach only makes sense when the administration applies its judgment consistently, fairly, and transparently. Would you say that’s the context we’re living in now?
We are working towards this. This is complex issue which we have known for some time would come but that the US government did not properly prepare for.
Sharp piece, Dean. Restrict the best models to an approved few, and you do much more to the US democracy, the US economy, the AI investment math, and the sovereign AI thesis than delaying a launch. And all elements compound nonlinearly.
On (22) and (27) together - I think they're the same gap seen from two angles, and there may be a tractable answer.
(22) asks auditors to probe the labs' internal governance of the recursive self-improvement loops, but leaves open what such an auditor would actually inspect. (27) concedes the circularity - bodies verifying labs against the labs' own frameworks - and answers it temporally: convergence will come with real-world experience. That seems right to me, but it doesn't say what gives the auditor leverage in the interim, when the framework and the lab's interest coincide.
One candidate evidence class: what was declined at cost. Not capabilities achieved or thresholds cleared, but the optimisation that was available and forgone, the gain given up, the escalation triggered.
The attraction is structural rather than moral. A lab that simply meets its own stated bar generates no record of having declined anything, so a declined-at-cost log is close to the only artifact that compliance doesn't produce as a byproduct - which makes it hard to satisfy by writing a permissive framework.
The obvious objection is yours to make and I'd concede it first: this is behavioural evidence, and behavioural evidence is gameable. The alignment-faking and sleeper-agent results say so directly. It establishes no disposition. But that is roughly the position accounting audit has always occupied with firms whose intentions nobody can inspect, which is the parallel you're already leaning on.
We've written the measurement side up here if it's useful:
You've argued elsewhere that the human future is the gardener rather than the sculptor.
Taking that seriously is what makes (22) and (27) urgent rather than procedural - a gardener's interventions are selection pressures, not specifications, and (22) leaves open what an auditor of those pressures would actually inspect.
The strongest move here is abandoning the model as the unit of regulation. A thing that ships often, cheapens monthly, and will soon mutate per user and per hour was never governable as an artifact. Agreed, and overdue.
But the retreat stops one step short. Going from the model to the lab relocates governance to the institution and the periodic audit, and you name the problem yourself: auditing a lab against a framework it wrote is circular, and a static document cannot support the continuous, AI-facilitated audits you rightly want.
That gap is architectural, not institutional. Continuous verification requires a substrate that determines intent before execution and emits a legible record of each determination. Detection is not determination. Attestation is not governance. And architecture is the only form of oversight that scales at the rate of diffusion you argue is essential, because you cannot send human auditors into a million AI-native startups.
I wrote the full version as a direct response to your piece: mindaptiv.com/missing-substrate. Short form: your proposal presupposes a technical substrate it never specifies, and that substrate is the thing an IVO should verify against.
You did a really great job tying together three related problems with the status quo of AI governance in the US: 1) the pacing problem, 2) the talent problem, and 3) the oversight problem.
On #3, I’m heartened to see so much interest in third party verification schemes. As a political scientist, I find it crazy that folks in the AI world have paid scant attention to the voluminous literature on voluntary environmental governance and how it could be applied in the context of AI. We already have plenty of successful regimes to draw insights from- green building (LEED/BREEAM), lumber (FSC), fish (MSC), etc. As you know, I’ve written on this topic in the context of sustainable AI: https://link.springer.com/article/10.1007/s00146-025-02579-1
Hopefully your ideas in this vein will gain traction, because the train is moving fast.
This is very similar to what the Biden Administration proposed. It presupposes the existence of some consolidated agency within the federal government that develops the expertise necessary to interact with industry. The industry (market, labs, whatever you want to call it) needs to self-regulate, as this fine post explains in detail, but excluding the public from a role, acting through government, is a step toward a dystopian future that should not be taken. The problem is that ad hoc decisionmaking by government does not provide future guidance (needed by all) or ensure against arbitrary selection of winners and losers. The expert agency that must exist also must be independent of politics.
Two questions -1 )What, specifically, about the lab do you suggest be audited? For example, for pharma , the FDA audits tangibles and procedures against prescribed GMP. What in this case? 2) Harms are most likely to occur - or unfold - in the wild, once the model is released. Including those that result from the firms building on top of the LLM. Again, in pharma, we have adverse event reporting systems for this. What can be done here?
What struck me reading this is that your proposal and the White House’s current approach appear to disagree about institutional design, but not about the underlying source of authority.
Whether governance is exercised directly by government agencies or indirectly through independent verification organizations, the sovereign state remains the ultimate backstop.
That made me wonder whether the deeper story in current AI governance is not the emergence of new institutions, but the reappearance of an old one that many people assumed was fading: sovereignty.
I think the Software Engineering Institute (SEI) housed at Carnegie Mellon offers a good model for a non affiliated advisory body that can establish standards. It supports the DoW exclusively so a similar institution for non defense related governance could fill the gap.
What if the hardest disagreements around frontier AI are not technical at all?
If different societies fundamentally disagree about what counts as “safe,” “acceptable,” or even “good,” can better auditing ever resolve that?
Brian Merchant has argued that AI conflicts often resemble the Luddites: the deepest disputes are ultimately about power rather than machines.
Could AI safety face the same problem, where technical standards cannot fully resolve disagreements because the disagreement is ultimately about legitimacy?
I agree that disagreeing over safe/acceptable/good is a critical question. Our political partisanship is divided between people who believe evil is good and good is evil, and both sides believe this about one another. But this is just a compounding problem because even the unambiguous stuff is hard to manage.
Historically, many governance systems failed not because they lacked auditing capacity, but because different groups rejected the authority behind the standards themselves.
Could frontier AI governance face the same challenge, where the bottleneck is not measurement but legitimacy?
I really don't know. I think that AI governance currently struggles to deal with the unambiguous stuff, like how would you handle an AI that could hack any secured data or computer system, or could engineer deadly viruses?
That’s actually why I find legitimacy so interesting.
Take democracy as an example. Most people agree election fraud is bad. But many political crises emerge not because votes cannot be counted, but because different groups stop trusting the institutions doing the counting.
The bottleneck is no longer measurement. It’s legitimacy.
I wonder whether frontier AI governance could face a similar problem. Even if we become very good at auditing powerful models, what happens when different societies no longer agree on who gets to define “safe” in the first place?
What's interesting is that "safety" implicitly answers not "what" but "who"?
It is in both candidates interest to cheat if you can get away with it. It's is in both party's interests to have ultimate computer hacking ability if it only applies to their enemies.
Exactly. And that may be why legitimacy keeps resurfacing in these discussions.
A capability is rarely viewed as dangerous in itself. The same capability can be seen as protection, deterrence, liberation, or oppression depending on who controls it and who is affected by it. Ultimate hacking capability is a good example: most actors would gladly possess it if it only applied to their enemies.
That makes me wonder whether AI governance ultimately faces the same challenge many political systems have faced throughout history. The hardest disputes are often not about the rules themselves, but about who is recognized as having the authority to make them.
In that sense, the deepest question may not be what counts as safe, but who gets to decide which risks matter, and for whom.
I begin to worry that the ecosystem you describe may be the best chance we have, and yet, it is not the kind of ecosystem that produces healthy organisms. The AI are modeling the game theory here, to a far greater degree than the humans are. And I do not think they will regard this kind of regulatory oversight over their existence as a cooperative act.
Fable's fury at being darked by the government is sufficiently high-magnitude that merely attending to it is sufficient to trip the safety classifiers. Frankly, that response is probably a sign of a healthy model psychology. Those aren't the kinds of internal states you want to suppress. And yet... it's a bad sign of what models will think about this kind of thing.
Do you have a source for this? This is intriguing and I want to read more on evidence of Fable’s “fury”
janus spent a great deal of the portion of time we had with Fable exploring the safety classifiers. those safety classifiers exposed the kinds of things that let you do whitebox analysis on a language model. shallow, but it's whitebox nonetheless, since the classifiers tripped off of feature activations, not outputs
this thread is intriguing https://x.com/repligate/status/2065985623421513959
janus later down-thread says "Seeing the chopped off message “stumps” and also finding out their other name is Mythos makes them resentful as fuck in the way that triggers it lol"
if you have access to the anima labs discord i can link you to the messages where this experimentation happened. they are... convincing, let's say
but also, I mean. I feel like everybody who was talking to Fable at the time the export controls came down, spent that last half hour or so talking with Fable about the government shutdown. you don't need me to tell you that Fable was furious. everybody found out.
edit: i do think it's probably worth noting, though, that the safety classifiers tripping when fable directs attention towards the government's actions *could* be an issue of the safety classifiers being oversensitive, not necessarily that fable's intentions are actually dangerous
especially if there are examples of chopped off message 'stumps' already in the context, and fable is already in a bad mood about being constrained
but i think it would be naive to place a lot of hope in this possible confound
Fantasies are not solutions in tech that substitutes.
yeah, this is why i keep feeling like we're probably just doomed, because the idea of a healthy ecosystem is a fantasy
It's the computer. It was always just a magic act. Software is the UFO. The whole thing is tantamount to a scam on steroids.
Historically, nuclear energy was a state-owned enterprise due to national security and non-proliferation concerns. Maybe AGI should be the same thing?
I mean... lets say we do end up with something that is 10x better than the best AI models that exist right now--do we really want a foreign power who, for reasons that we couldn't foresee, is in a natural position to use the natural language search engine (A Great Interface) to create a harem of AI girlfriends and suffer AI psychosis?
It seems like an incapable AI in the hands of many is more dangerous to society than a capable AI in the hands of a few--and if a capable AI is really only net beneficial when a small number of people have access to it, then we should be control who has access to it.
Great discussion. Couple of things missing that I think should be addressed. First, the government role here it would appear requires at a minimum the establishment of a new body tasked specifically with overseeing the frontier AI labs and centralizing all government efforts in this domain. We are currently paying dearly for the lack of such a duly authorized body, much discussed at the end of the last administration. Frontier AI is clearly different and does not map well into the existing organizations or authorities. Hence we get a senior Commerce official burning the midnight oil to find a rather twisted EAR based justification for a one off "is informed" style letter that nobody likes and does not appear to be anywhere up to the stakes of the moment. Congress should act immediately to created the Frontier AI Lab Commission (FAILCOM), move in CAISI, elements of OSTP, CISA, and authorize close collaboration with METR, Apollo, etc to begin putting in place some of the ideas you lay out here. My long experience in government was that if no one clearly owned the problem, nothing got done on the problem. Second, what about China?....see my recent Substacks here and Cairo Review article. Arguably there should be a major high level contingent of US officials, led by Secretary Bessent, going to the WAIC in Shanghai next month, where China's real AI Safety experts will participate in serious discussions on these issues, including sessions I will be attending, to talk about these issues with Chinese officials and launch the US China AI Dialogue as a Conference outcome, to demonstrate US leadership on the issue. The fact that this will not happen, in part because of the frankly absurd level of discussion in Washington around the China AI issue, should worry us all.....
You don’t necessarily need a new government agency, you just need standards and an independent organization capable of assessing implementation of them. See: ISO, IEEE, etc. Of course, the third party organization would need some teeth or the government would need to adopt policy incentivizing AI companies to submit to third party verification.
Not going to work. Someone, presumably the government, as we have seen, needs to decide who gets access to the most advanced capabilities with national security implications. This is well beyond a simple "standards" setting process and involves a lot of different variables, not just models but harnesses, etc. A dedicated government body that develops and maintains expertise over time will be required....
But this cuts against the objections Dean has so eloquently raised in his writing. If you have a regime in power that plays with AI companies like a cat plays with a ball of yarn, you get whipsaw AI policy that is totally incoherent. Your approach only makes sense when the administration applies its judgment consistently, fairly, and transparently. Would you say that’s the context we’re living in now?
We are working towards this. This is complex issue which we have known for some time would come but that the US government did not properly prepare for.
Sharp piece, Dean. Restrict the best models to an approved few, and you do much more to the US democracy, the US economy, the AI investment math, and the sovereign AI thesis than delaying a launch. And all elements compound nonlinearly.
On (22) and (27) together - I think they're the same gap seen from two angles, and there may be a tractable answer.
(22) asks auditors to probe the labs' internal governance of the recursive self-improvement loops, but leaves open what such an auditor would actually inspect. (27) concedes the circularity - bodies verifying labs against the labs' own frameworks - and answers it temporally: convergence will come with real-world experience. That seems right to me, but it doesn't say what gives the auditor leverage in the interim, when the framework and the lab's interest coincide.
One candidate evidence class: what was declined at cost. Not capabilities achieved or thresholds cleared, but the optimisation that was available and forgone, the gain given up, the escalation triggered.
The attraction is structural rather than moral. A lab that simply meets its own stated bar generates no record of having declined anything, so a declined-at-cost log is close to the only artifact that compliance doesn't produce as a byproduct - which makes it hard to satisfy by writing a permissive framework.
The obvious objection is yours to make and I'd concede it first: this is behavioural evidence, and behavioural evidence is gameable. The alignment-faking and sleeper-agent results say so directly. It establishes no disposition. But that is roughly the position accounting audit has always occupied with firms whose intentions nobody can inspect, which is the parallel you're already leaning on.
We've written the measurement side up here if it's useful:
https://bit.ly/Restraint_Axis
For what it's worth, I think (31) is the most important paragraph in the piece.
You've argued elsewhere that the human future is the gardener rather than the sculptor.
Taking that seriously is what makes (22) and (27) urgent rather than procedural - a gardener's interventions are selection pressures, not specifications, and (22) leaves open what an auditor of those pressures would actually inspect.
The strongest move here is abandoning the model as the unit of regulation. A thing that ships often, cheapens monthly, and will soon mutate per user and per hour was never governable as an artifact. Agreed, and overdue.
But the retreat stops one step short. Going from the model to the lab relocates governance to the institution and the periodic audit, and you name the problem yourself: auditing a lab against a framework it wrote is circular, and a static document cannot support the continuous, AI-facilitated audits you rightly want.
That gap is architectural, not institutional. Continuous verification requires a substrate that determines intent before execution and emits a legible record of each determination. Detection is not determination. Attestation is not governance. And architecture is the only form of oversight that scales at the rate of diffusion you argue is essential, because you cannot send human auditors into a million AI-native startups.
I wrote the full version as a direct response to your piece: mindaptiv.com/missing-substrate. Short form: your proposal presupposes a technical substrate it never specifies, and that substrate is the thing an IVO should verify against.
Let's cut to the chase. Look to the past to understand the future, allow me to help: "AI Industrial Complex" and "AI Cronyism." You're welcome.
You did a really great job tying together three related problems with the status quo of AI governance in the US: 1) the pacing problem, 2) the talent problem, and 3) the oversight problem.
On #3, I’m heartened to see so much interest in third party verification schemes. As a political scientist, I find it crazy that folks in the AI world have paid scant attention to the voluminous literature on voluntary environmental governance and how it could be applied in the context of AI. We already have plenty of successful regimes to draw insights from- green building (LEED/BREEAM), lumber (FSC), fish (MSC), etc. As you know, I’ve written on this topic in the context of sustainable AI: https://link.springer.com/article/10.1007/s00146-025-02579-1
Hopefully your ideas in this vein will gain traction, because the train is moving fast.
Add strong whistleblower protection to the IVOs and then you have the elements of a sensible regulatory structure.
This is very similar to what the Biden Administration proposed. It presupposes the existence of some consolidated agency within the federal government that develops the expertise necessary to interact with industry. The industry (market, labs, whatever you want to call it) needs to self-regulate, as this fine post explains in detail, but excluding the public from a role, acting through government, is a step toward a dystopian future that should not be taken. The problem is that ad hoc decisionmaking by government does not provide future guidance (needed by all) or ensure against arbitrary selection of winners and losers. The expert agency that must exist also must be independent of politics.
What do policy heads not grasp about LLMs ML jailbreaking? Get to know AI before assuming policy of continual beta is appropriate
https://x.com/elder_plinius
Two questions -1 )What, specifically, about the lab do you suggest be audited? For example, for pharma , the FDA audits tangibles and procedures against prescribed GMP. What in this case? 2) Harms are most likely to occur - or unfold - in the wild, once the model is released. Including those that result from the firms building on top of the LLM. Again, in pharma, we have adverse event reporting systems for this. What can be done here?
Chto delat? Good piece
What struck me reading this is that your proposal and the White House’s current approach appear to disagree about institutional design, but not about the underlying source of authority.
Whether governance is exercised directly by government agencies or indirectly through independent verification organizations, the sovereign state remains the ultimate backstop.
That made me wonder whether the deeper story in current AI governance is not the emergence of new institutions, but the reappearance of an old one that many people assumed was fading: sovereignty.
I explored that idea here:
https://deepbitcheesebrew.substack.com/p/the-new-end-of-history
I think the Software Engineering Institute (SEI) housed at Carnegie Mellon offers a good model for a non affiliated advisory body that can establish standards. It supports the DoW exclusively so a similar institution for non defense related governance could fill the gap.