
AI safety requires more than just slowing our pace
Safety requirements are non-negotiable. They depend on meeting concrete goals, not just adjusting a timeline
It has been a week of high drama in AI, precipitated by the resignation of the AI safety researcher Jacob Coxon from Anthropic. This followed several weeks of increasingly lurid and disturbing revelations about the OpenAI/Hugging Face incident.
My inbox yesterday included a message from Business Insider with the subject line: âÂÂAI doomsday debate reaches boiling pointâÂÂ.
Now, the Anthropic CEO, Dario Amodei, has written a 3,800-word, reassuringly phrased letter titled âÂÂWe Must Pace the Frontier,â describing his proposals for avoiding (or at least postponing) doomsday. Sam Altman of OpenAI, Elon Musk of xAI, Demis Hassabis of Google Deepmind, and Satya Nadella of Microsoft have all expressed support.
You may be forgiven for not immediately understanding what âÂÂpace the frontierâ means. (My first image was of Amodei walking deep in thought along the FinnishâÂÂRussian border.)
The phrase also appeared in JulyâÂÂs âÂÂPacing the Frontierâ open letter, signed by 1,386 employees of frontier AI labs, including Amodei himself.
While that letter may have upset the industryâÂÂs PR executives with its signatories noting âÂÂthe complete absence of credible plans for controlling superintelligent AI systemsâ and asserting that âÂÂbuilding things smarter than humans ⦠is, objectively, an insane and suicidal thing to doâÂÂ, AmodeiâÂÂs monograph goes out of its way to mollify investors.
The notion of pacing the frontier seems to come from Formula 1: when conditions become too dangerous for racing, a pace car comes onto the track and all the other cars have to follow it as a safe speed. Progress continues, without the danger.
Amodei writes: âÂÂTo be clear, pacing does not mean halting model training or technical progress.âÂÂ
AmodeiâÂÂs letter is prompted by his concern that âÂÂAI has been advancing drastically faster, driven primarily by ... recursive self-improvement.â ItâÂÂs as if he and Sam find themselves driving their F1 cars at 200mph neck-and-neck heading into the first corner, only to realize itâÂÂs covered in ice and they have no steering wheel. No wonder they want to slow down.
In brief, AmodeiâÂÂs proposal has three parts. The first is to have third-party AI system evaluators working inside each company, with full access to the systems; he commits Anthropic to this plan now, without waiting for the government to require it.
The second part of the plan asks all the frontier AI companies in âÂÂdemocratic countriesâ to âÂÂestablish common safety standards as well as limits on the rate of unchecked AI progressâÂÂ, with government regulation where needed. The third part would include âÂÂauthoritarian countriesâ in a broader compact.
Here, Amodei goes out of his way to reassure those in Washington who see AmericaâÂÂs lead in AI as its most important geopolitical asset.
On a casual reading, there are many reasons to believe that Amodei is calling for a general slowdown in the rate of progress. He talks about âÂÂlimits on the rate of unchecked AI progressâ and âÂÂsome kind of âÂÂspeed limitâ on the rate of recursive self- improvement (RSI)âÂÂ. He says: âÂÂProgress will still seem fast, and we must make wise use of the time we gain.âÂÂ
Slowing down would give companies a bit more time to work on safety; Amodei talks about one to two years of extra time for research on interpretability, alignment, and better testing methods.
Let me pause here to respond to AmodeiâÂÂs critics who say itâÂÂs just a bid to cement AnthropicâÂÂs lead with the help of government intervention. This is nonsense. In fact, the Wikipedia page on pacing in F1 races says it âÂÂeliminates any time and distance advantage that a leading driver may have had over the remaining field of competitorsâÂÂ.
Having said that, I think the pacing metaphor is completely misguided.
We cannot set a slower rate of progress for capabilities and then hope that provides enough time to get the safety right. The safety requirements are non-negotiable. We must set the safety requirements first, and further progress occurs only when they are met.
Imagine if Boeing said: âÂÂWeâÂÂre going to introduce a new plane every year, and we hope that provides enough time for some flight tests to be completed and for the results to be good.âÂÂ
We would say: âÂÂNo, you have that backwards; you can introduce a new plane only when it has passed all the tests and the government has issued an airworthiness certification. If that takes more than a year, so be it.âÂÂ
A more careful reading of the document suggests that Amodei agrees with this objection. For example, he says that rules should be of the form: âÂÂIf models have capability X, then they need to be accompanied by certifications of alignment properties Y and Z.âÂÂ
In other words, we set safety requirements, and developers have to show that they meet those requirements. This is in fact the âÂÂred linesâ approach that AI safety researchers have been calling for.
And it means that if developers canâÂÂt figure out how to meet the safety requirements, then they will have to halt. It would be, in F1 terminology, a red flag and not a pacing car.
There is no plausible alternative. Recursive self-improvement leading to superintelligent AI raises the risk of the irreversible loss of human control. The acceptable risk level for loss of control is perhaps one in 100m per year, not the one in 10 or one in five that the AI CEOs currently estimate.
And remember âÂÂthe complete absence of credible plans for controlling superintelligent AI systemsâÂÂ. At some point, progress along this technology path will halt, not because further progress is impossible, but because further progress is untenable when the technology is intrinsically unsafe. Humanity has a right to protect itself.
There is huge resistance to this conclusion. We have already sunk trillions of dollars into the current technology path and plan to sink trillions more. But the sunk cost fallacy is just that: if we double down on a mistake, itâÂÂs still a mistake.
The present level of attention to AI risk, the unanimity of the leading technology executives, and the forthcoming TrumpâÂÂXi summit give us a real opportunity to choose a different path. We must take it.
-
Stuart Russell is a distinguished professor of computer science at University of California, Berkeley, the president of the International Association for Safe and Ethical Artificial Intelligence and a Guardian US columnist
