Constitutional AI

Constitutional AI


Standard harmlessness training needs people to label large volumes of unpleasant output. That is slow, expensive, and hard on the reviewers. Constitutional AI replaces most of that labour with an explicit document of principles and a model that applies them.

Training runs in two phases. In the supervised phase the model answers a prompt, then critiques its own answer against a principle from the constitution and rewrites it. Fine-tuning on the revised answers teaches the behaviour. In the reinforcement phase the model compares pairs of responses against the same principles, and those AI-generated preferences train the reward signal that RLHF would have taken from humans.

The property worth noting is legibility. The values the model is being trained toward exist as text a person can read, argue with, and version. Anthropic published the constitution behind Claude, so the principles can be inspected rather than inferred from behaviour.

Human judgement does not disappear. Someone still writes the principles, and the model's reading of an abstract principle can differ from the intent of whoever wrote it.

From generative AI strategy to custom agent development and retrieval architectures, we help you scale AI responsibly.
Discuss your AI project