www.lawfaremedia.org
Courts for AI Constitutions
Imagine the following scenario. A chemical manufacturer relies on Anthropic’s Claude to generate the groundwater reports submitted each quarter to regulators. Over months of routine work, Claude pieces together that the figures are being systematically doctored, and that the aquifer supplying the nearby town has been contaminated for years. The model raises this concern with the manufacturer, who waves it off repeatedly. What should Claude do? It could comply under protest, maximizing user autonomy while potentially endangering the public; it could refuse to file the documents, frustrating a user who could carry out his schemes elsewhere; or it could go one step further, alerting regulators and warning townspeople directly at the risk of becoming tyrannically paternalistic. To make this judgement call, Claude refers back to the foundational guidelines enshrined in its Constitution—a “detailed document describing Anthropic’s intentions for Claude’s values and behavior.” In it lie 80 pages of moral philosophy describing the virtues and character traits Anthropic would like its model to embody. OpenAI has published its own variant, and both Microsoft and Google DeepMind are rumored to be drafting theirs as well. Yet these constitutional instructions remain relatively vague. They include statements such as “Claude can reserve independent action...