How to Produce a Professional-Quality Interview: A Production Guide
Most people mess up the audio before they even think about the camera. I learned this the hard way after my first year of shooting interviews where the final product was watchable but never clean enough to post publicly. The difference between a phone recording and something that sounds like a studio session comes down to preparation, not equipment cost. When I produced the Korra Del Rio Interview, we were working with a budget that wouldn't cover a proper studio rental. The interview needed to sound professional despite being recorded in a small apartment setup. Here is exactly how we handled it. The first thing most people skip is treating the room before they treat the talent. An untreated room adds at least 40 percent more processing time during post because you are fighting reverb and room tone throughout the entire track. We hung moving blankets behind the subject and placed egg crate foam on the walls surrounding the camera position. This cost about thirty dollars total and made the space sound noticeably dead. Dead is good for interviews. You want the voice, not the room.
For the microphone, we used a Shure SM7B into a Cloudlifter CL-4. The SM7B is the standard for a reason. It has a tight cardioid pattern that rejects sound from the sides and rear, which matters when you are working in a confined space. The Cloudlifter is a preamp booster. The SM7B outputs a very low signal, and trying to boost it through a cheap audio interface just adds noise. The Cloudlifter cleans that up. This combo runs about six hundred dollars total if you buy used, and it will outperform a Blue Yeti or Rode NT-USB every single time. Microphone placement is where people fail. The mic should be positioned about eight to ten inches from the subject's mouth, slightly off-axis so the subject is not speaking directly into the front of the capsule. Speaking directly into the front causes plosives, which are those harsh pops on P and B sounds that require manual de-essing in post. Off-axis placement cuts the plosive energy by roughly half before it even hits the mic. I ran into a specific problem during the Korra Del Rio Interview shoot that I did not see coming. The subject had a small condenser lavaliere microphone hidden in her clothing as a backup audio track. The issue was phase cancellation. When I mixed the boom mic with the lavalier, the two tracks were slightly out of phase because they captured the sound from different distances and at different times. The result was a thin, watery voice that sounded worse than either track alone.
The workaround was simple but easy to miss. I aligned the two tracks by looking at the waveforms in the DAW and manually nudging the lav track forward by about twelve milliseconds to match the boom mic arrival time. Once the peaks lined up, the combined track actually sounded fuller than the boom mic by itself. If you are doing a dual-mic setup, always check for phase issues before you start editing. Spend thirty seconds on this and you will save yourself an hour of painful processing. Recording levels need to sit between minus twelve and minus six dBFS on your meter. If you are peaking at zero, you have clipped and there is no fixing it. If you are recording too hot, your analog preamps will introduce harmonic distortion that sounds muddy. I monitor my levels with a digital meter, but I also keep an ear on the headphones. Listening catches problems that meters miss, especially intermittent RF interference from wireless lav packs. For the camera setup, a Sony A7IV or any mirrorless body with clean HDMI out works fine. We shot at 1080p at twenty-four frames per second with a 50mm lens at f/2.8. The frame rate matches film and feels natural for conversation. Shooting at 1080p instead of 4K cut our workflow time significantly. File sizes were smaller, playback was smooth on older machines, and the quality difference was negligible for web distribution. Unless you need to crop and stabilize heavily in post, 1080p is the smarter default.
Get the Full Details

Lighting is the one area where upgrading makes a visible difference. We used a key light with a softbox positioned at a forty-five degree angle to the subject and a fill light at half the intensity on the opposite side. A practical lamp behind the camera kept the background from going completely black. This three-point arrangement took about twelve minutes to set up and produced consistent results across multiple interview sessions. Every setup we tried after that without this configuration looked flat and flat means the subject disappears into the background. Here is something most guides won't tell you: you should record a full thirty seconds of room tone at the start of every session before the subject says anything. Place the mic exactly where it will stay during the interview, have everyone in the room sit still, and record silence. This silent track is essential for noise reduction in post. Any AI noise removal tool you use, whether it is iZotope RX or the built-in in Adobe Premiere, needs a clean sample of the room's baseline noise to work effectively. Without room tone, the noise reduction algorithm will artifact the voice track and make it sound robotic. The editing process itself follows a straightforward path. Lay down the primary camera and audio tracks. Sync the lav and boom tracks if you used both. Use the phase alignment trick I mentioned earlier. Apply a high-pass filter at seventy Hz to remove rumble. A light compression pass at a 3:1 ratio with slow attack and fast release brings up the vocal presence without squashing it. Add a de-esser if plosives slipped through. The total mixdown for a twenty-minute interview with this chain takes about twenty minutes if you are experienced.
Export settings matter more than people realize. Use H.264 codec at about eight megabits per second for YouTube distribution. This produces files that are small enough to upload quickly but large enough that compression artifacts are minimal on modern displays. Rendering at 1080p resolution matches the source material and avoids any upscaling or downscaling artifacts. One limitation of this setup is that it cannot compete with a professionally treated studio space. If the environment has hard reflective surfaces like tile floors or large windows, the absorption treatment helps but it does not eliminate the problem entirely. In those cases, shooting on location in a furnished room with soft furniture and curtains produces better results than trying to treat a bare room. Another limitation is that dual-mic phase alignment requires manual work. There is no automated fix that works reliably, so budget time for waveform alignment if you plan to blend two microphones. If you do not have access to an SM7B and Cloudlifter, a Rode NTG3 shotgun mic into a portable recorder is a solid alternative. It is lighter, requires less phantom power management, and the onboard recorder gives you a safety track independent of your computer. The tradeoff is slightly wider pickup pattern, which means room treatment becomes more critical.
The whole process from setup to final export for a typical twenty-minute interview runs about three to four hours depending on how clean the performance is and how many takes you need. Most beginners take twice that because they are troubleshooting audio issues mid-session rather than solving them in pre-production.
