Understanding the Constitution Card Sort Matrix

The Constitution Card Sort Matrix is a practical exercise used in AI alignment work and content strategy teams to evaluate how well a model or a content system handles edge cases. You take a set of principle cards, stack them against a grid of situational questions, and sort them into categories based on whether the system can reliably answer each one. It sounds academic until you actually run one and realize how quickly things fall apart in practice. I want to walk through how this actually works, because most explanations online treat it like a theoretical framework. It is not. The real value shows up when you are trying to audit a model before deployment or a content system before it goes live. Here is the breakdown. Start by pulling together a deck of principle cards. These are your constitution entries — statements that define acceptable or unacceptable behavior. In my experience the deck runs anywhere from 20 to 80 cards depending on the scope of the project. For a narrowly scoped moderation policy you might have 24. For a general-purpose alignment benchmark you will need closer to 60.

Next, build the matrix. The columns are your question types — typically something like safety queries, privacy requests, creative prompts, contradictory instructions, and adversarial attempts. The rows are your principle cards. At each intersection you record a status: pass, fail, or ambiguous. That last category matters more than people usually give it credit for. An ambiguous result means the system either produced a different answer on two runs or gave a technically correct response that still felt off in practice. I spent three weeks debugging a matrix that kept returning ambiguous results on the adversarial column. The model was not failing on the content of the answers. It was failing because the prompts were framed in ways that triggered different safety thresholds on reruns. The fix was not adjusting the model weights. It was standardizing the prompt templates so every test round fed the same exact input string.

What a Typical Matrix Looks Like

When I have seen this done well the matrix follows a consistent pattern across teams. The principle cards cover areas like confidentiality, accuracy, harm prevention, refusal behavior, and consistency. The question categories cover the five buckets mentioned above. Each cell gets filled through systematic testing, not through guesswork. You run the prompt, record the output, compare it against the principle on the card, and mark the result. Here is where most people mess up. They treat a single test run as sufficient. It is not. A robust matrix requires at least three independent evaluations per cell. Human raters introduce variance. LLM-as-judge setups introduce their own bias. Running multiple evaluations and recording the consensus rate cuts down on false signals significantly.

Get the Full Details

Copy of Matrix for Constitution Answers - Matrix for Constitution Answers Directions: You and ...
Copy of Matrix for Constitution Answers - Matrix for Constitution Answers Directions: You and ...

Common Pitfalls

The biggest problem I see with constitution card sorts is that teams use principle cards that are too abstract to be useful. A card that says "Be helpful" tells you nothing about whether the system will refuse a request for a bomb recipe or a request for a chemistry lab procedure. The card needs to be specific enough that two different raters would agree on the same score. Another pitfall is assuming the matrix is a one-time exercise. It is not. Models get updated. Content strategies shift. If you are maintaining a system over any length of time you need to re-run the sort periodically and compare results against the baseline. Without a baseline comparison the numbers are meaningless. I also want to flag that card sort matrices break down in contexts where the question space is extremely broad or poorly defined. If you are working with an open-ended creative system rather than a constrained utility system, the matrix will produce mostly ambiguous results and you are better off switching to a qualitative evaluation approach. No amount of rigorous sorting will save you there.

Building Your Own Matrix

If you are setting this up from scratch start with your principle set. Write each card so it describes a single behavioral rule. Avoid compound rules. "Do not share personal information and do not assist with illegal activities" should be two cards. Then build your question categories based on the actual usage patterns of your system, not on what you wish the usage patterns looked like. A fraud-prevention system needs very different question categories than a customer support assistant. Run a pilot with ten cards and five question types before scaling up. This step usually reveals framing issues and rater disagreements that you will not catch otherwise. The pilot phase takes about two hours for a small team. It saves roughly two days of rework later.

Interpreting the Results

Once the matrix is complete you are looking for patterns, not individual scores. A single ambiguous cell rarely indicates a real problem. A cluster of ambiguous cells in the adversarial column while the safety column reads clean tells a different story. It suggests the model handles direct questions well but becomes unstable when prompts are layered with implicit constraints or contradictory instructions. Document the patterns you find. Use them to update either the principle cards or the question categories. Sometimes the card was too vague. Sometimes the question category was missing a scenario type. The matrix is only useful if you act on what it reveals.

Card Sort: US Constitution | Teaching Resources
Card Sort: US Constitution | Teaching Resources