Killer machines?

Over the past week, some prominent AI figures called for pacing frontier model development. They point, in part, to the recent OpenAI-HuggingFace security “incident”. For the longest time I privately disagree with the existential framing of model development and AI. While I don’t doubt the harm that superintelligence can do, I think that existential claims are only good servants insofar as they shed light on learned worries, but they are a bad master in that they distract from principled governance.

Perspectives on AI should graduate from arguing about categorical risks in isolation and look towards risk-proportionate evaluation frameworks. The question of whether a model is safe cannot remain context-agnostic. It is for this reason that a model is not a sufficient unit of governance. A model on its own doesn’t do harm. It is always a model plus something. I therefore consider AI systems instead of models on its own, even though the language is transferable.

Throughout, I define assurance to be justified confidence, supported by structured evidence, that a system remains within acceptable risk bounds for its intended use-case. Because much of what is acceptable is a normative judgment, I focus on the evidentiary rather than the normative aspects of this theme and so I decline to over-prescribe by defining what proportionality means. The common thread throughout my three points is that the conversation should be focused on the central question: “what makes us assured?”.

Three thoughts:

1. Ecosystem

It is more beneficial to look towards a responsible ecosystem rather than placing open-weight and closed-weight models at odds. I consider open-weight development to be crucial to a healthy safety ecosystem. Independently-operable models are important because researchers, organisations, and defenders cannot merely depend on the policies and alignments of a handful of model providers. I don’t propose that open-weight models are more responsible than others, but they certainly shape practice and thought in this regard. Given the current market state, I foresee that new and upcoming model developers will find trust and safety to be a differentiator even if speed-to-market has historically won. As a tidbit, HuggingFace investigators used the open-weight GLM-5.2 model in its investigations after Claude Opus 4.8 refused to perform forensic analysis. Placing open-weight and closed-weight models at odds therefore appear to further the capability asymmetry. However, open-weight harms are real. Bio and cyber concerns sharpen these harms, and the sharper regulation question concerns the capability asymmetry between regulated defenders and unregulated attackers. How does a coherent policy preserve defensive capabilities without carelessly externalising offensive capabilities? To this end, regulation and active (transnational) law enforcement should come hand-in-hand with responsible AI practice. We would be assured by regulations on open-weight models that are proportional to their demonstrated capabilities and foreseeable harms.

2. Institutional

I wholly disagree with light touch regulation towards AI, and I think they should continue to be placed within traditional safety paradigms. For instance, there is no reason to think why a security incident like HuggingFace should be shrugged off while conventional cybersecurity lapses result in audits and traces. Such breaches, as alluring as they are, are not merely an alignment problem. AI companies have a duty of care, for which they are accountable to society through containment, audit and accountability, and incident disclosure, all of which are traditional measures. There is also no surprise that the HuggingFace post-mortem advocates exactly for these measures. Though, such measures are sometimes seen to concede economic grounds to competitors, like in US-China dynamics. I don’t believe that frontier capability is the sole determiner of that digital arms race. Instead, integration and societal trust, with reliability and institutional legitimacy leading to proliferation, is an under-discussed framing. We would be assured by considering what systems are trustworthy enough to be embedded into everyday institutions like banks, and hospitals.

3. Epistemic

Much of AI governance has been built from the perspectives of data protection, provenance, and accountability. I think that the way forward for governance is to bridge all of these necessary, though ultimately insufficient, measures, and to look towards governing AI systems with a view towards explainability. Mechanistic interpretability is promising here, and I advocate for interpretability to become a first-class evaluation component for trust. The current reliance on chain-of-thought is misplaced, and I strongly believe that chain-of-thought should not be a privileged indication of a model’s transparency. Anecdotally, end-consumers intuitively equate chain-of-thought to transparency. Amodei’s post The Urgency of Interpretability from a year ago is relevant, and I echo the concerns that chain-of-thought is most capable to mislead because it doesn’t always reveal the true internal mechanism that an answer is arrived at.


Of course, an MRI for AI is years away, but current technologies can scaffold such a capability. Looking into the crystal ball, I consider this to be even more valuable for growing the field when this rationalisation is unintuitive (which arises, for instance, in reading attention pooling in terms of local and global features). I consider the north star to be that a system must behave acceptably in the environment it is intended for. The larger the scope, the more onerous this demand is. For high-risk applications, regulators should increasingly look to coherent evidence from a whole suite including mechanistic interpretability, behaviour evaluations, and adversarial testing. With transparency as a regulatory burden, we would be assured by a risk-proportionate framework that may reject a system for use-cases simply because it is uninterpretable.