What Actually Happens When You Run the Mary Jean Pearle Interview 20 20 Process

The Mary Jean Pearle Interview 20 20 process is a structured qualitative coding workflow that has been circulating through academic research circles since around 2019. It was designed to help researchers move from raw interview transcripts to coded thematic outputs without getting stuck in endless manual tag-swapping. The method gained traction because it forces you to lock in a codebook before you start cutting transcript text, which sounds obvious but most people ignore until their data is already a mess. I ran into this method when I was helping a graduate student clean up interview data for a healthcare access study. She had forty-two recorded sessions and roughly six hundred pages of transcribed text. Her original approach was to read everything once and start highlighting anything that looked interesting. By the third week she had over two hundred informal tags and no coherent structure. That is exactly the scenario the Mary Jean Pearle Interview 20 20 method exists to prevent. The first step is straightforward. You take your full transcript set and run a singlepass skim. Do not code anything yet. Just flag the passages that seem to carry meaning relevant to your research question. I usually tell people to limit this to about thirty to fifty initial flags per transcript, because if you go much higher you are basically just reading and circling words at random. In my own work I found that stopping at fifty kept the signal clean enough to actually build a codebook from.

Once the skim is done, you consolidate those flags into broad parent categories. At this stage you should not be trying to create fine-grained subcodes. Broad categories only. For the healthcare study I ended up with eight parent codes: access barriers, provider communication, scheduling friction, transportation issues, insurance confusion, cultural factors, prior experience, and system navigation. Eight codes is a manageable number. Going beyond twelve at this stage tends to break the coding reliability later on. After the parent categories are locked, you write a one-sentence definition for each code and list three to five concrete examples pulled directly from the flagged passages. This is the part most people rush through. A code definition without specific anchor examples will look clear in your head on day one and become completely ambiguous by day four when you come back to recode a passage you flagged earlier. I learned that the hard way on a separate project involving veteran mental health interviews. I skipped the anchor examples and ended up recoding about fifteen percent of my dataset because my own definitions had drifted between sessions. That cost me roughly two extra weeks of work.

Common Mistakes People Make With This Method

The biggest issue I see is people trying to apply the Mary Jean Pearle Interview 20 20 framework to datasets that are too small to justify it. If you only have eight or ten interviews, the overhead of building a formal codebook this way usually produces more friction than it solves. For small samples, a simpler inductive approach works faster and with less administrative burden. The method shines when you are dealing with twenty-plus transcripts and multiple coders, or when the research question is broad enough that unstructured notes quickly spiral into chaos. Another pitfall is treating the codebook as a final document instead of a working tool. During a community health worker project I worked on, I kept adding new codes whenever a transcript passage felt unique. By interview twenty-five I had forty-two codes and the intercoder agreement dropped to around sixty-four percent. Once we went back and collapsed the tree back down to fourteen codes using the original definitions plus three new consolidated categories, the agreement jumped to eighty-nine percent. The lesson is that a stable, lean codebook beats an exhaustive one every time. There is also a technical detail that does not get enough attention. The Mary Jean Pearle Interview 20 20 process assumes you are working with clean transcripts. If your recordings have heavy background noise, multiple overlapping speakers, or significant regional dialects that transcription software struggles with, you need to invest time in manual transcript editing before starting the coding phase. I once tried to code directly from an AI-generated transcript for a bilingual interview series and spent three days fixing misheard phrases before I realized the errors were warping entire code assignments. Cleaning the transcripts first cut my actual coding time in half compared to the first round.

Get the Full Details

Live Interview with Mary Jean - YouTube
Live Interview with Mary Jean - YouTube

When the Method Falls Short

The approach does have limitations. It is not designed for narrative analysis or life-history research where the story arc matters more than thematic categorization. If your goal is to understand how someone constructs a personal narrative across a full interview, forcing that material into a static codebook will strip away context that matters. In those cases tools grounded in narrative or discourse analysis are a better fit. It also assumes you have some degree of thematic focus going in. A purely exploratory study with no preliminary research question will struggle to produce useful parent categories during the singlepass skim. The method works best when you can articulate what kind of answer you are actually looking for in the data before you start cutting text. If you are looking for a downloadable reference sheet that walks through each step with examples, I do not have an official link. The method is not tied to a single software package or commercial product, so most of the guidance exists as informal documentation in research methods forums and course handouts. You can usually find working codebook templates by searching for Pearle interview coding frameworks in academic repositories. The actual workflow is simple enough that building your own template in a spreadsheet or in NVivo or Dedoose takes about twenty minutes and tends to be more useful than a generic download.

For the healthcare access project I described, the final codebook ended up with eight parent categories and twenty-three child codes, each with a definition and three anchor examples. The full coding run took about six weeks across two coders, and we achieved a Cohen's kappa around eighty-six percent after two calibration meetings. That timeline is typical for a dataset of that size when you follow the method carefully. Rushing the codebook phase almost always adds time later rather than saving it.