Constitutional AI does it differently. You write the principles down, hand them to the model, and let it critique its own work against them. Anthropic built the method to train Claude, and it’s the reason “what does this model actually value” is a question you can answer by reading a document instead of guessing.
What it means
Constitutional AI is a training method where a model checks its own answers against a written set of principles, called a constitution. The model drafts a response, critiques it against the constitution, then rewrites it. People still shape what counts as helpful. The harm-avoidance part gets handled by the model judging itself.
Why it matters
It cuts how much human labeling you need. Anthropic’s original research found the model came out both more helpful and less harmful with no human labels on harmlessness at all. Labeling is one of the biggest costs in training a model, so that math adds up fast.
The values become something you can read. Anthropic published Claude’s constitution in January 2026, roughly 23,000 words, released into the public domain (the 2023 version ran about 2,700 words, so it grew a bit). For scale, the US Constitution is around 7,500 words.
Order matters more than the rules do. Claude’s constitution ranks safety first, then ethics, then Anthropic’s own guidelines, then helpfulness. When a regulator asks how a model decides to refuse something, that ranking is the answer, which is probably why other labs end up publishing something similar.
Simple example
A head chef has thirty plates going out an hour. He can taste every one, catch the problems, and send them back, but that’s slow and it falls apart on a busy night. So he writes the standards down instead. How the sauce should look. What gets remade. Why a cold plate is worse than a late one.
The new cook reads all of it, checks his own plate against it, and fixes what’s wrong before it leaves the pass. The head chef still writes the standards. He just stopped tasting every plate.

