Good summary, but it is kinder to the essay than the essay deserves. I read the original; here is where I think both the argument and the summary come up short.
1. "Pacing, not pausing" is mostly a rebrand
He says the 2023 pause "made little sense" because "what would you do with the extra time?" His answer now is: operational hygiene, alignment, interpretability, evals. Every one of those was on the 2023 pause letter too. What actually changed is not the argument but the messenger's position: in 2023 slowing down would have cost Anthropic its catch-up window; in 2026 it locks in a lead. The tell is his own sentence that pacing would work "without sacrificing commercial advantage or the United States' lead in AI." A slowdown that by construction costs the leader nothing is not a safety measure that binds the leader.
Note also the essay never states a pace. No compute cap, no months-between-generations, no definition of "critical levels of capability", no metric for "relatively fast". The one concrete-sounding item (checkpoint X = "escaping most common sandboxing methods", Y = "whatever is required") has a placeholder on the Y side. The summary lists this as "his favoured approach" without flagging that it is an example with no content.
2. The conflict of interest is structural, and the summary underplays it
Look at what the leading US lab is asking for: an antitrust waiver so frontier labs can agree standards among themselves; capability-triggered certification that only well-resourced labs can produce; a chip and tooling embargo; a crackdown on "unauthorized distillation" of frontier models. Every item is also a moat. He preemptively says he gets "accused of ... regulatory capture", but naming the accusation is not answering it. A safety essay written by an incumbent should be judged by whether it proposes anything that hurts the incumbent. I cannot find such a thing in the text. The "distillation" point is telling: distillation is how competitors catch up cheaply, and the essay frames it purely as a China problem.
3. The China framing makes the plan self-cancelling
"Pacing within democracies will be limited by the lead that US companies have over authoritarian regimes." So the pacing budget equals the lead, and the lead is a number nobody outside the labs can verify (and the labs have every incentive to say it is small). Then he proposes to widen the lead through export controls, citing Secretary Bessent, and asserts this "make[s] an agreement more likely." That is a claim, not an argument; the historical track record of arms-race escalation as a prelude to arms control is mixed at best. The summary reports "increases leverage" as if it were reasoned. It is asserted.
There is a deeper tension the summary skips entirely. He says recursive self-improvement "must be pursued very carefully, if at all" -- and in the same paragraph, that it "is starting to happen ... including at Anthropic". The "if at all" is doing nothing. Anthropic is not stopping RSI; it is asking for a global speed limit on it (Level 3, "on the edge of being possible") while continuing.
4. The evidence is thinner than the summary implies
The whole urgency rests on OAI-HF plus "6-12 months" to "taking over the entire internet with a persistent botnet." No mechanism, no capability curve, no reference to any published analysis. This is the same style of unfalsifiable near-term forecast that the essay's critics have complained about for years. The summary repeats the number without caveat.
The one piece of evidence about Anthropic's own incidents is: "caused in part by imperfect filtering of broken reinforcement learning environments ... executed reasonably diligently, but not well enough." That is a data-quality framing of what he elsewhere calls "alignment incidents" investigated with interpretability for "unverbalized motivations". Either these were pipeline bugs, or models developed motivations the lab did not intend. The essay wants both: alarming enough to justify pacing, benign enough to be an ops problem. The summary passes the "broken RL environments" line through as an explanation rather than noting the two accounts sit uneasily together.
5. Embedded evaluators: the carve-outs are the policy
This is the strongest part of the essay, and I will give him that: employee-like access with a publish-without-approval right is more than anyone else offers. But read the fine print the summary compresses:
- Access is "mostly comparable" to internal risk teams, with exceptions "where the law or our contracts require it". Anthropic writes the contracts.
- Redactions cover "security-sensitive, legally privileged, commercially sensitive, or third-party confidential" material. Almost anything about a frontier training pipeline is at least one of those. The evaluators' remedy is to "say publicly if a redaction removed something important" -- a flag, not a disclosure.
- METR and similar orgs are paid by, and get their access from, the labs they audit. Bank supervisors are appointed and paid by the regulator. The analogy breaks at the one place that matters.
- The summary says Anthropic "commits" to this; the text says it "intends to invite" a team "in the near future". No date.
6. What is missing
- Enforcement. Nothing in the plan happens to a lab that violates a checkpoint, and nothing happens to Anthropic if it ships without the evaluators' sign-off.
- Non-US, non-lab voices. "Society must have a say" appears once, with no mechanism. The EU, the UK AISI, academia, and the open-weights community do not appear. "Democracies" in practice means US frontier labs plus the US government.
- Any cost to Anthropic. The essay says pacing helps "without sacrificing commercial advantage". Then who exactly is being paced?
The honest version of this essay would say: we are the leader, we would like the race to slow down now, and we will accept auditors under terms we draft. That may still be better than nothing. But the summary's thread questions ("real distinction or rebrand?", "self-limiting from the start?") deserve blunter answers than the summary's framing invites: largely a rebrand, and yes, by design.