What could make AI self-regulation work?

What could make AI self-regulation work?

WASHINGTON—Earlier this week, the White House announced an accord on “Super Intelligence”—the Trump administration’s new preferred term for artificial intelligence (AI)—cosigned by the leaders of Google, Anthropic, OpenAI, Meta, xAI, and Nvidia. Many AI policy observers were left wondering whether and how a “morally binding” accord relying on voluntary internal controls would mitigate the doom-coded incidents of the last three months—from frontier models escaping test environments to agents probing government websites. To be clear, the accord does not suggest that companies are not accountable for the harm caused by their runaway agents. Taking the administration’s premise at face value, if existing law provides the backstop, the accord is an ex-ante layer meant to prevent harms before they occur. Both have to hold up. It is worth examining, then, how this collective corporate self-governance could work, and which pitfalls would undermine it. What makes collective self-governance work—and what lets it fail One strong precedent comes from the realm of nuclear power. After the nuclear accident at Three Mile Island, then US President Jimmy Carter established the Kemeny Commission, which found that “merely meeting the requirements of government regulation does not guarantee safety; therefore, the industry must also set and police its own standards of excellence[.]” Following the commission’s recommendations, utility CEOs formed the Institute of Nuclear Power Operations (INPO) in 1979. INPO relies mostly on peer pressure, but its inspections also inform underwriting by the industry’s mutual insurer, with each plant’s rating directly affecting premiums. Crucially, however, federal regulators remain an important check on the industry: INPO operates alongside the Nuclear Regulatory Commission, which oversees the plants themselves. The Bangladesh Accord shows the value of binding commitments even in the absence of a single federal regulator serving as oversight. After the Rana Plaza collapse in Bangladesh, which killed over 1,100 people working in sweatshop conditions, global clothing brands, under pressure from civil society, signed an agreement requiring independent inspections with public reports, mandatory repairs, brand underwriting of upgrade costs, and consequences for suppliers that refuse to improve. Disputes could be referred to the Permanent Court of Arbitration, potentially leading to punishments in companies’ home countries. Unions report that nearly 80 percent of hazards found in the original inspections were remediated. The reverse is also true: self-regulation fails when there are no explicit sanctions. Take, for instance, the Responsible Care initiative in the chemical industry. The initiative was launched in the aftermath of the Bhopal Gas Tragedy of 1984, a gas leak at a Union Carbide Corporation plant operating in India, which killed more than seven thousand people within days—and more than twenty thousand over the long term, by some estimates. Every firm faced the threat of financial losses and greater regulation after the accident, regardless of which company was at fault, creating a kind of commons. However, researchers have found that Responsible Care’s primary goals seem to be changing public opinion and lobbying against stronger regulation. Without third-party verification or explicit sanctions, Responsible Care became a cautionary tale. Four conditions for effective self-regulation Taken together, the examples above point to four conditions for effective self-regulation: First, shared exposure, where safety failures of one member can cause damage to all members. Second, consequences, such as substantial premiums or arbitration. Third, verification, paired with fourth, a credible federal backstop. The Joint Commitment on Frontier Responsibilities partially meets the first condition. Frontier AI has a collective reputation problem: a loss of trust in AI is likely to attach to the technology writ large rather than a specific company. However, the labs are also racing for market share, which creates a collective action problem. The accord also names no sanctions. It calls for independent external auditors but does not say who would carry out third-party evaluations. Additionally, with at least one accord signatory, xAI, creating an organizational culture that sidelines trust-and-safety teams, an audit without credible sanctions for noncompliance is unlikely to make a dent. The federal backstop is also uncertain. The text says codifying the measures into law “may make sense” in the future. The signatories also disagree among themselves about the nature of the risk, and even whether some kinds of risk are real at all. On the one hand, Anthropic CEO Dario Amodei argued that the industry must slow the pace at which it improves model capabilities and use that time to improve safety (a position broadly endorsed, to different degrees, by OpenAI‘s Sam Altman and Google DeepMind‘s Demis Hassabis). On the other hand, Nvidia CEO Jensen Huang has stated that the unique, catastrophic risks are overblown and existing law is sufficient. Meta’s Mark Zuckerberg has leaned toward self-regulation, framing trust and alignment as qualities that will differentiate products. President Donald Trump has asserted that the US risks losing ground to China if it slows down its own companies and called fears about an AI catastrophe a hoax. A coalition split over whether the risk itself is real will struggle to agree on common standards or sanctions, which pushes it toward the lowest common denominator. Putting power behind the pledge The accord is a fair starting point. Commitments to monitor models, staff internal teams to check that controls work, and bring in external auditors are solid building blocks. Over the coming months, the accord will have to pass concrete tests: Who selects and pays for the external auditors? Will the audit findings be published? Will there be any consequence attached to a failed audit? Will the circle widen beyond the six signatories to include smaller AI enterprises? In the longer term, the picture changes. If the race persists, whether with China or among the labs themselves, each firm has an incentive to free-ride on the others’ caution. In that case, government must step in with a “benign big gun”: an enforcement power credible enough that it rarely has to be deployed. Core safeguards must be mandatory, and not just reward companies for volunteering, since leniency offered in exchange for cooperation can make risky behavior more profitable. The role of a federal rule is not to replace industry’s own efforts but to make them stick.

Original Source

Read the full article at Atlanticcouncil →

KhanList aggregates and links to publicly available news content. We do not host full articles from third-party sources. Always verify important information with original sources.