Wire flash
Ex-CFIUS official: Trump AI safety pact lacks teeth, voluntary pledges won't slow race
Editorial responsibility
- No named human review is recorded for this page.
- Source reporting is collected, normalized, translated or condensed automatically when needed.
- Automatically published source-backed update
Connor Martin, former deputy director of CFIUS, analyzes the White House AI accord signed by President Trump and six leading AI companies. The one-page 'Joint Commitment on Frontier Responsibilities' outlines sensible safeguards including independent external evaluators and board-level oversight, but uses noncommittal language like 'believe' and 'should' rather than enforceable terms like 'will' or 'shall.' Martin argues the agreement does not alter commercial incentives for companies to prioritize capability over safety, noting that Anthropic and OpenAI face massive operating losses and competitive pressure from hyperscalers and open-weight models. The financial architecture built on continued rapid AI improvement creates systemic risk if capabilities plateau. However, the accord's proposals for independent external auditors and embedded evaluators offer a potential path forward if properly implemented with independence protections and full access privileges. Martin concludes that until the cost of safety failures exceeds the payoff of racing ahead, voluntary pledges will not slow companies down.
Source report
Connor Martin, former deputy director of the Committee on Foreign Investment in the United States (CFIUS) at the Treasury Department, previously focused on the national security implications of cross-border capital flows and the national security risks posed by artificial intelligence (AI) and other frontier technologies.
During his tenure at CFIUS, Martin's team negotiated agreements with private companies to compel specific actions or restrictions. The committee holds legal authority to mitigate national security risks arising from foreign investment transactions in the United States. By law, any such mitigation agreements must be: (1) effective, (2) verifiable, and (3) monitorable and enforceable.
This framework offers a useful lens for assessing the "White House Accord on Super Intelligence," signed Tuesday by President Donald Trump and leaders of six major U.S. AI companies—Anthropic, OpenAI, Google, Nvidia, Meta, and xAI. Trump commented that the agreement would involve "a tremendous self-policing aspect," is "morally binding," and that "we're going to have it be nice and safe."
Unfortunately, nothing in the document legally compels any signer to change their behavior. More importantly, nothing alters their commercial incentives to continue testing frontier capabilities. The document contains promising ideas but leaves to others the work of creating conditions that are effective, verifiable, monitorable, and enforceable.
A Conspicuously Noncommittal Agreement
The one-page agreement, titled "Joint Commitment on Frontier Responsibilities," does not actually compel the companies to take any action. It is closer in spirit to the "voluntary AI commitments" secured by the Biden administration in 2023 (though less detailed). It lacks the force of law or executive order, and the language is conspicuously noncommittal—with one revealing exception.
The document states: "[W]e believe each company should implement… four layers of controls and audits[.]" The words "believe" and "should"—terms any government lawyer would immediately strike from a CFIUS agreement because they are unfalsifiable—are used instead of "will" or "shall." The word "will" appears only at the end, when companies commit that they "will meet regularly to establish standards and best practices" on safety.
This is notable because one recommendation in Anthropic CEO Dario Amodei's "pacing the frontier" essay—which accelerated the AI safety debate in September—calls for the U.S. government to waive antitrust restrictions so labs can collaborate on safety. It remains unclear whether Trump's signature on this one-pager endorses that recommendation, but it could plausibly signal that federal antitrust enforcement would look the other way. Discussions are reportedly already ongoing between OpenAI, Anthropic, and Google DeepMind.
Commercial Incentives Unchanged
The nonbinding nature of the agreement also means it does nothing to alter the commercial incentives for companies to keep testing potentially dangerous models.
The "pacing the frontier" debate revolves around two related but distinct questions:
- Should the world's top AI companies slow down training of their best models to improve safety?
- Will they?
Even if leaders earnestly believe the answer to the first question is yes, the answer to the second depends on whether the cost of safety violations can be made to exceed the cost of sprinting ahead. This creates a business decision about where to direct finite resources—money, employees, and computing power.
To date, companies have grown by investing those resources in making their best models ("the frontier") even better at tasks consumers and enterprises want. "Pacing" means redirecting resources instead to operational improvements (such as data hygiene) and different kinds of model training and evaluation—focused not on effectiveness but on making actions more compliant (or aligned) and reasoning more transparent (or interpretable).
These are challenging engineering problems that will be expensive to solve. Reallocating resources toward them only makes business sense if the next dollar spent on safety earns as much as the next dollar spent on capability. According to Reuters, Anthropic's leaked IPO filing showed a 2025 operating loss of $8 billion, even as revenue grew by more than 1,000 percent to $4.6 billion. Additionally, huge amounts of Anthropic's cash are locked into infrastructure commitments. OpenAI has projected negative free cash flow of nearly $280 billion from 2026 through 2030.
Anthropic and OpenAI are squeezed between:
- Hyperscalers, which own the compute they currently rent and are developing their own capable models
- Open-weight models, which are free and good enough for most use cases
For these two companies, the promise of frontier capabilities is their product. Among other signers, Google and Meta already own significant compute and have other profitable business lines (xAI is cushioned by its merger with SpaceX). However, Google, Meta, and xAI are also reporting massive capital expenditures on AI infrastructure, and all three compete with OpenAI and Anthropic for the same enterprise relationships and consumer loyalty. If OpenAI and Anthropic don't stop testing at the frontier, why would they?
[Video: https://vimeo.com/1219226879]
The incentives extend further. The sector—and arguably the entire U.S. economy—is systemically exposed to the risk that model capabilities plateau. AI companies have swamped global credit markets, with over $500 billion in debt issued to the sector this year alone. They have also engaged in circular financing, where customers finance suppliers and vice versa. The Bank for International Settlements issued a bulletin on Thursday about the "exceptional" scale of these arrangements in the industry, "potentially amplifying contagion" in the event of a shock.
This financial architecture was built for a world in which AI model capabilities continue improving rapidly. To slow that rate of improvement without collapsing valuations, investors need to believe that safety is worth the delta.
Potential Ways Forward
Despite these concerns, the agreement's "four layers of controls and audits" provide the kernel of a path forward on AI safety—if companies can be incentivized to adhere to them.
- Proposals 1 and 2 call for "robust internal controls" and an "empower[ed]… internal team," respectively, each intended to "ensure" that models are not misused or behave in unintended ways. Presumably, all these companies already have such controls and teams, which have not prevented reported instances of AI breaking from its training and evaluation environments.
- Proposal 3 calls for companies to "[p]artner with an independent external auditor or evaluator to carry out independent assessments of whether the controls, monitoring, and detection are operating as intended." This independent evaluator approach is consistent with some CFIUS mitigation measures (e.g., "security officers" and "third-party monitors"). Anthropic and OpenAI have already committed to embedding such evaluators, and embedded evaluators have been endorsed by many leading AI scientists. However, the proposal omits crucial details for effectiveness—such as how evaluators would be kept meaningfully independent, protected from retaliation, and granted full access privileges.
- Proposal 4 calls for companies to "[d]esignate an independent committee of the board of directors to oversee and receive reports from the teams operating the controls and the internal and external auditors and evaluators, as well as to ensure any issues identified are remediated." This identifies specific individuals at the corporate governance level with visibility into AI safety issues and responsibility for fixing them. However, it fails to specify what would make the committee "independent," how it would be constituted, or how it would ensure issues are remediated.
Game theory dynamics in AI—not just among U.S. labs, but also between the United States and China—incentivize racing ahead to dominate the technology. Research by economists suggests that to truly "pace the frontier," all participants would have to share information transparently. The agreement's strongest commitment—that companies "will meet regularly to establish standards and best practices"—points toward this possibility. But, like the rest of the document, it lacks conditions that would make that commitment verifiable and monitorable, or any enforceable penalties for violating it.
Making Safety Stick
In CFIUS cases, the government is well-practiced at establishing clear requirements and clear lines of responsibility: Company X is required to take Action Y; if it doesn't, somebody is on the hook. The "Accord on Super Intelligence" lacks any such clarity and enforceability. It does not compel any company to take any action, and it doesn't specify who is on the hook for any violations.
This omission is striking because Treasury Secretary Scott Bessent—an essential figure in the Trump administration's approach to AI—has forcefully asserted that the government will not give AI companies a "liability shield." Although the agreement alludes to some possibility of "laws or regulations" in the future, there is no indication that anything meaningful is on the way—certainly not before the next Congress.
It should not come as a surprise, then, if the first real teeth for any slowdown come from the courts. Nothing prevents Hugging Face, for example, from suing OpenAI for being hacked by its agents this summer. Regardless of whether such a suit were ultimately successful, it would be reasonable to expect model testing to slow while the industry watched it play out.
For now, commercial pressures on leading AI companies continue to incentivize them to push the boundaries of model capabilities. Regardless of how it happens, the only way to change that is to raise the cost of safety violations. Establish those incentives, and tremendous self-policing really might follow.
This work represents the views solely of the author(s). The Council on Foreign Relations is an independent, nonpartisan membership organization, think tank, and publisher, and takes no institutional positions on matters of policy.
Source
Council on Foreign RelationsNeutral / independent