Radcliffe Likes And Dislikes — What It Actually Is

There is a framework called Radcliffe Likes And Dislikes that people sometimes use when they need to structure subjective evaluation work, particularly around research interviews, product feedback, or archival material tagging. It isn't a formal academic theory from the Radcliffe Institute or anything published in a journal. It's more of a practitioner shorthand that showed up in early-stage qualitative research teams and got recycled across a few newsletters over the years. The core idea is simple enough. You take a set of items — interview transcripts, feature lists, survey responses, whatever — and you sort each one into three buckets: likes, dislikes, and neutral. The "radcliffe" part is just a name someone attached to it, possibly as a reference to the Radcliffe College tradition of structured note-taking. It has nothing to do with the college itself. It's not proprietary. What makes it useful is the forced triage. Most people naturally want to sit on the fence and call everything nuanced. This method removes that option for the first pass. You categorize, then you qualify.

I ran into a concrete edge case with this once. We were processing about 400 user interview transcripts for a fintech client, and the team started disagreeing on whether certain responses should be classified as likes or dislikes because the language was deeply ambivalent. The workaround was to add a strict rule: if the respondent expressed any negative sentiment even once in the same sentence, the entire statement goes into dislikes. You can add nuance back in during coding. You can't create nuance where there isn't a clear signal.

When it falls apart

The main limitation is cultural and contextual ambiguity. Items that are clearly liked in one demographic often register as dislikes in another, and the method doesn't handle that natively. You end up adding sub-codes anyway, which means you've basically rebuilt a full qualitative coding framework on top of what was supposed to be a quick sorting exercise. At that point you might as well have just used NVivo or Dedoose from the start. Another practical issue: it doesn't scale well past roughly 2,000 items without automation support. Once you hit that volume, manual sorting becomes a liability. People get tired. Consistency drops. I've seen inter-rater reliability fall to about 0.55 after a team pushes past that threshold without taking breaks or using dual-coding.

Get the Full Details

science-resources - Evolution and natural selection
science-resources - Evolution and natural selection

Practical steps to run it

Start by defining your unit of analysis. A unit could be a sentence, a paragraph, a whole response, or a feature request — but it has to be consistent. I usually pick sentences for interview data and feature-level items for product feedback. Don't skip this step. Teams that skip it end up comparing results across coders that aren't actually comparable. Build a codebook. Even a thin one. "Like" means the respondent expressed positive sentiment toward the item. "Dislike" means the opposite. "Neutral" means no clear sentiment either way. You'll get pushback from team members who want to add "mixed" and "confused" as categories. Don't. Those are coding artifacts, not sentiment buckets. Let the text speak for itself on the first pass. Train at least two people on the codebook before you start. Have them code the same ten items independently, then compare. If your agreement rate is below 80%, your codebook is too vague. Rewrite it. This usually takes one or two iterations.

Run the bulk sort. I recommend 50-item batches with a five-minute break between them. It sounds like overkill but fatigue effects are real and they degrade quality faster than most managers expect. After the sort, go back through the dislikes bucket and flag any items that might be context-dependent. That's where you add subcodes if needed. The likes bucket usually needs less attention. Dislikes contain the ambiguity.

Alternatives worth knowing about

If you're doing heavy qualitative work, direct content analysis or thematic analysis through a tool like MaxQDA will give you better output, though they take longer to set up. For lighter weight work where speed matters more than depth, the Radcliffe Likes And Dislikes approach is fine. It's not elegant. It's not rigorous. It gets things done in about 15 to 30 minutes per 100 items depending on data density, which is faster than most full coding workflows. There's no official download or software for it because it's not a product. It's a method. You can implement it in a spreadsheet with three columns, in Google Sheets with conditional formatting, or in any qualitative data platform that lets you assign categorical tags. I've seen people do it effectively in Airtable alone.

Natural Selection and Adaptations - Made By Teachers
Natural Selection and Adaptations - Made By Teachers

One counter-intuitive thing most people miss

The neutral bucket is usually where the valuable signal lives. Not because neutrals are interesting in themselves, but because they reveal what your respondents didn't care enough to form an opinion about. In product research, that gap is often more actionable than the positive and negative feedback combined. A feature that lands squarely in neutral for 60% of respondents means it's not solving a problem anyone feels strongly about. That's a design failure disguised as acceptable feedback. Most teams throw out the neutral bucket and move on. The people who use this method well spend the most time there.