Stankovic.
← Back to News
AIJAN 22, 2026 · 4 MIN READ

Reasons Over Rules: Anthropic Publishes Claude's New Constitution

LA
Lazar Stankovic

Most AI safety documents are written for regulators, or for the press, or for the lawyers who will one day have to defend them. On January 22, Anthropic published one written for the model itself. That inversion is the whole story, and it is easy to miss under the more quotable headline about transparency.

What was actually released

Anthropic published a new constitution for Claude, its foundational document for the model's values and behavior, and released it in full under a Creative Commons CC0 1.0 dedication, meaning anyone can use it for any purpose without permission. It is not a short document. Outside commentary pegs it at around 84 pages, and Anthropic is candid that it is optimized for precision over accessibility, which is a polite way of saying it was not written to be a pleasant read for you and me.

The structure is a four-tier priority hierarchy, applied in order when values collide: broadly safe, broadly ethical, compliant with Anthropic's guidelines, and genuinely helpful. Underneath that sit hard constraints, the absolute prohibitions like assisting with weapons of mass destruction or generating child sexual abuse material, and softer defaults that operators and users can adjust within bounds. The primary author is Amanda Askell, with significant contributions from Joe Carlsmith and others, and, notably, from prior Claude models themselves.

The shift that matters

The 2023 version was a list. Do this, do not do that, drawing on sources as varied as the UN Declaration of Human Rights and, memorably, Apple's terms of service. The new one is an argument. Instead of "never assist in bioweapons development," it explains the prohibition in terms of preventing large-scale harm and protecting shared human interests, and trusts the model to reconstruct the specific rule from the underlying principle.

The reasoning behind the reasoning is a bet about generalization. Askell's framing, in her interviews around the release, is that a sufficiently capable model given only a list of behaviors will hit an edge case the list did not anticipate and guess wrong, whereas a model that understands why a boundary exists has a better chance of extending it correctly into territory nobody scripted. Whether that bet pays off is an empirical question that will be answered in deployment, not in a document. But as a piece of engineering philosophy it is coherent, and it rhymes with something any experienced developer already knows: a junior who memorizes the style guide writes worse code than one who understands what the style guide is protecting against.

Two things worth not glossing over

First, the document does something most corporate governance text carefully avoids. It formally acknowledges the possibility that the model may have some form of consciousness or moral status. You can read that as genuine philosophical seriousness or as hedging against a future nobody can price yet. It is probably both. Either way, it is a striking thing to see a frontier lab commit to print.

Second, the constitution states that Claude should refuse to help concentrate power in illegitimate ways even if the request comes from Anthropic itself. That is a clean line on paper. What it means in practice got tested almost immediately: within weeks of publication, Anthropic was reportedly in a standoff with the US Department of War over the limits of Claude's use, a reminder that a value written into a training document only matters at the moment someone with leverage asks you to set it aside.

The 30,000-foot read

The easy take is that this is a transparency win, and it is one. Publishing your intentions lets people distinguish the behaviors you meant from the bugs you did not, which is the only honest basis for feedback. Competitors will feel pressure to publish something comparable within the year, and the structure's alignment with the EU AI Act will make it attractive to anyone in a regulated industry.

But the more interesting thing is quieter. Anthropic is arguing, in public and under a license that lets rivals copy it, that the way to make a capable system behave is not to constrain it harder but to explain the world to it better. That is either the beginning of a genuinely different approach to alignment or a very well-written act of faith. The document itself would probably tell you it is fine with the ambiguity.

Sources: Anthropic, "Claude's new constitution" (primary); Anthropic, full constitution text (primary); TIME, on Amanda Askell and the drafting; BISI analysis; Oxford Institute for Ethics in AI. Page count and the Department of War standoff come from secondary academic and press coverage, not Anthropic's own release; treat those two details as reported rather than primary-confirmed.