Working with Moral Reasoning Assessments

Most people encounter Kohlberg's stages in an intro psych class and assume they understand them. The Heinz dilemma gets reduced to a multiple-choice quiz about stealing drugs for a dying wife, and everyone picks an answer without realizing they're being scored on something entirely different. The whole project is more complex than the pop psychology version suggests, and it's more frustrating in practice than you'd expect. Kohlberg's approach rested on presenting subjects with moral dilemmas and then analyzing the reasoning behind their answers, not the answers themselves. The Heinz scenario is the most famous example. A man can't afford a drug that could save his wife's life. Should he steal it? Almost everyone has an opinion on whether he should. What matters for scoring is why they think so. That distinction is where most people get tripped up when they first try to apply this framework. The six stages fall into three levels. Pre-conventional morality covers stages one and two, where decisions are based on direct consequences like punishment or personal gain. A child at stage one won't steal because they'll get caught and punished. At stage two, they might steal if they calculate that getting the drug benefits them or establishes a reciprocal favor. Conventional morality, stages three and four, shifts toward social approval and maintaining order. "People will think less of him if he steals" or "What if everyone stole? Society would collapse." Post-conventional morality, stages five and six, involves abstract principles that may override existing laws. Social contracts and universal ethical principles become the reference point rather than rules handed down by authority.

Here is the thing nobody emphasizes enough: these are stages of reasoning, not stages of behavior. A person can reason at stage four and still commit a crime. Someone at stage one can show remarkable compassion in specific situations. The assessment measures how someone justifies a decision, not whether the decision is morally good by any independent standard. I spent several months coding interviews against the stage criteria and kept running into this gap. People would give deeply conventional answers about duty and obligation while describing actions that caused real harm. The framework doesn't account for that disconnect well. The methodology requires trained raters. You can't just hand someone the Scoring Manual and expect consistent results. In my experience, inter-rater reliability drops noticeably after the first twenty interviews because raters develop personal shortcuts for categorizing responses. I solved this by having two independent coders score every interview and discussing discrepancies until we hit a 0.85 Cohen's kappa minimum. It added maybe forty-five minutes per case but prevented the drift that happens when one person codes everything solo over a long project.

Where the Framework Breaks Down

The Heinz dilemma itself is a problem. It's abstract, third-person, and involves a wife rather than a close friend or family member the respondent might identify with. When I ran these interviews with adolescent participants, roughly a third either rejected the premise entirely or reframed the scenario before answering. Some insisted the wife wasn't worth saving. Others said the pharmacist was justified in charging what he wanted. The scoring manual had provisions for these responses, but they didn't fit cleanly into any stage. I ended up creating a separate category for premise rejection and noting it in the coding sheet, which later became useful data about how cultural background shaped moral reasoning patterns. Cross-cultural application is another weak spot. The post-conventional stages assume a liberal democratic context where social contracts and universal principles are culturally recognizable concepts. Working with participants from collectivist cultures or societies with different governance structures, the highest stages became hard to identify because the vocabulary Kohlberg used didn't map onto their moral frameworks. Stage six in particular is almost entirely theoretical. Kohlberg himself admitted it was rarely observed and eventually removed it from the scoring in later revisions. You're measuring something that exists mainly on paper. Gilligan's critique is well-known but worth restating practically. The standard dilemmas frame morality as a problem of rights and justice. Many respondents, particularly women in the original samples, approached the dilemma through relationships and care ethics instead. Their reasoning didn't reflect a lower stage. It reflected a different moral orientation that the scoring system had no category for. I found that when I stopped forcing responses into the six-stage model and instead coded for moral orientation alongside stage, the data became significantly more useful. It took longer but produced findings that actually described what people were doing rather than what they were failing to do according to a narrow rubric.

Get the Full Details

Lawrence Kohlberg's Stages Of Moral Development Ppt at Qiana Timothy blog
Lawrence Kohlberg's Stages Of Moral Development Ppt at Qiana Timothy blog

If you're considering using this for research or organizational assessment, here is what I'd recommend instead of a full Kohlbergian interview study. Use the Defining Issues Test, the shorter written version that measures moral schema development through priority rankings rather than open-ended scoring. It's faster to administer and score, takes about twenty minutes, and has better psychometric properties for large samples. It doesn't give you the rich qualitative data that the interview method produces, but it captures the same underlying construct with considerably less overhead. Most projects don't need the interview depth, and the additional time it takes is rarely justified by the results. The original longitudinal data showed stability in stage progression over decades but also substantial regression under certain conditions. People don't always move upward. Adverse circumstances, institutional roles, and cultural feedback can push reasoning downward or hold it in place. If you're evaluating someone at a single time point and concluding they're "at stage three," you're making a claim that's only conditionally valid. The stage is a snapshot of how they reason in that moment about that type of problem, not a permanent characteristic.