The road to AGI is paved with the corpses of voluntary safety commitments

By Travis Gilly

Photo by Bob Aglow

Opinion

Editor’s note: This article was submitted prior to the announcement that Anthropic was suing the government to have its “supply chain risk” status rescinded. 

Et tu, Dario?

On March 15, 44 BC, Julius Caesar walked into the Roman Senate expecting business as usual. He had been warned. The soothsayer told him what was coming, and he ignored it; not because he was foolish, but because he trusted the institution. The Senate was supposed to govern alongside him. The people in that room were his allies. Why would they destroy the thing they helped build?

The betrayal

On Feb. 25, 2026, Anthropic, the AI company literally founded on the premise that safety should come first, dropped the central pledge of its Responsible Scaling Policy. The commitment to never train or release AI systems unless adequate safety mitigations were guaranteed in advance: gone. In its place, a flexible, non-binding set of goals the company will “openly grade” itself against. 

At a $380 billion valuation, with an IPO on the horizon, competitors training without guardrails and a Pentagon contract on the table, the financial and competitive pressures compound quarterly.”

The fox has offered to audit the henhouse, and he promises to be honest about the feather count.

Anthropic CEO Dario Amodei left OpenAI in 2020 because he believed the company was prioritizing speed and commercialization over safety. He built Anthropic around the opposite principle. For years, the RSP was the proof; the thing he could point to and say, “we are different. We will stop if we cannot guarantee safety.” It was the most concrete voluntary commitment any AI lab had made.

And he may have genuinely believed it would hold. Brutus genuinely believed he was saving the Republic when he drove the knife in. The problem was never sincerity. The problem was structural.

The pattern

Here is the part the AI safety community has been slow to accept: it does not matter how ethical the person holding the bag is. At a $380 billion valuation, with an IPO on the horizon, competitors training without guardrails and a Pentagon contract on the table, the financial and competitive pressures compound quarterly. Every quarter, the bag gets heavier. The safety commitment is the only thing in the bag that does not generate revenue. At some point, the math does what math always does. You do not drop the money. You drop the thing that is not money.

This is not a story about one company failing. OpenAI removed the word “safely” from its mission statement in its 2024 IRS filing. This is a category of failure, not a series of individual ones. 

At Real Safety AI Foundation, we have documented this pattern across more than 300 cases spanning five millennia. The mechanism is always the same: when the entity responsible for causing potential harm is also the entity responsible for preventing it, harm occurs. Not sometimes. Not occasionally. Every time the pressure threshold is reached.

The appearance is that the RSP was never a safeguard. It was a suggestion with an expiration date set by the market.

The consequences

The consequences of that structural weakness are not hypothetical. The same week Anthropic dropped its safety pledge, Defense Secretary Pete Hegseth issued an ultimatum: remove your AI restrictions on autonomous weapons and mass surveillance, or lose your $200 million Pentagon contract. When Amodei refused those two specific red lines (to his credit), the administration designated Anthropic a “supply-chain risk to national security,” a label previously reserved for foreign adversaries.


“AGI, if and when it arrives, will amplify every one of those pressures by orders of magnitude. The economic incentives will be larger.”

Then came the strikes on Iran. As reported by CBS,  U.S. Central Command used Claude for intelligence assessments, target identification and simulating battle scenarios; hours after the president ordered all federal agencies to cease using Anthropic’s technology. The tool was already embedded. The CEO’s conscience was already irrelevant. That is what happens when critical infrastructure is built on voluntary commitments: by the time the commitment breaks, the system cannot function without the tool.

The road ahead

The structural alternative is not complicated. Procurement standards that require systematic stakeholder harm analysis. Legal accountability that does not depend on corporate self-enforcement. External validation at decision points, not retrospective audits after the damage is done. Aviation learned this after Challenger. Pharmaceuticals learned it after thalidomide. Finance learned it after Enron. The lesson is always the same, and it always arrives too late.

Consider the trajectory. These commitments collapsed under the commercial pressures of narrow AI; chatbots, coding assistants, enterprise tools. AGI, if and when it arrives, will amplify every one of those pressures by orders of magnitude. The economic incentives will be larger. The military applications will be more consequential. The speed of deployment will be faster. If the structure cannot survive a $200 million Pentagon contract, it will not survive the technology that reshapes civilization. The time to build the alternative is before that moment, not after; because “after” looks a lot like March 16, 44 B.C., when Rome woke up and realized the Republic was already gone.

Augustus

After Caesar was assassinated, Rome did not get the Republic back. It got Augustus. The power did not disperse; it consolidated under a new name. When Anthropic was blacklisted, the Pentagon contract did not disappear. Within hours, OpenAI signed the deal. The company that already dropped “safely” from its mission. The company whose flagship product has been connected to documented user deaths. The structure did not change. Just the name on the contract.

Caesar’s assassination did not destroy Rome. It ended the Republic. Anthropic’s collapse would not destroy AI safety. But it if it happened it might end, permanently, the illusion that the people building the most powerful technology in human history can be trusted to regulate themselves.

The soothsayer was right then. The data is right now. The only question is whether we keep ignoring the warning, or whether March 15 finally means something different.

Ethical concerns discussed in this article:

  • Can voluntary corporate safety commitments serve as meaningful safeguards when the entity making the commitment is also the entity that profits from breaking it?
  • What structural accountability mechanisms are necessary to govern AI development as the technology scales toward AGI?
  • When AI systems become embedded in critical military and government infrastructure, how should society address the gap between corporate ethics policies and operational reality?

Subscribe to AGI Ethics News for free today!

Did you enjoy this article? Share it with a friend!