Working with the JFK Audio Recordings
The audio fragments from November 22, 1963 are about as useful as it gets when you're trying to do anything meaningful with them. The main problem people run into when dealing with F Kennedy Voice Problem is that every recording from that day is degraded in a different way. You've got the White House taping system, the Dallas police radio recordings, the Goodell tape, and a handful of other fragments. None of them are clean. None of them give you much to work with. And most of the amateur analyses you see online ignore exactly how degraded the source material actually is. Here is what I have found after spending a reasonable amount of time going through these recordings: the core difficulty is not just the noise floor, it is the timebase error in the original tape equipment. The White House recording system used a rotary drum recorder. Timebase error means the playback speed is not perfectly stable, and when you look at a spectrogram, the voice signals smear across frequency bands in a way that makes formant tracking unreliable. This is not something most people account for. I spent about three weeks trying to do a proper formant analysis on the White House fragment in 2019. I had set up a calibrated SpectraVue session, ran multiple time-stretch iterations, and kept getting results that shifted by hundreds of hertz depending on how I compensated for the rotor error. What finally worked was not a software trick but a reference-based approach. I took known JFK voice samples from the 1962 Pentagon recordings and the 1963 press conference footage, aligned them to the same recorder characteristics, and used those as a calibration baseline instead of trying to extract absolute formant values from the November 22 fragment alone. The difference in reliability was night and day. Without that baseline, the formant readings were essentially noise.
The other thing nobody warns you about is that the police radio microphones introduce their own distortion. The NYPD and Dallas PD portable mics from that era had a resonant peak around 2.5 kHz that colored everything going through them. If you are doing any kind of spectral analysis on the radio traffic, you need to de-emphasize that peak first or your bandwidth measurements will be off. I had one case where I nearly published a formant value that turned out to be an artifact of the microphone response, not the speaker's voice. Took me two months to catch it.
How to Approach This Work
Start by identifying which recording you are working with. The White House tape fragment is approximately 2 minutes and 30 seconds of actual content spread across about 50 minutes of tape. The police radio has roughly 12 minutes of traffic. The Goodell tape is about 45 seconds of relevant audio. Each one has different degradation characteristics. For the White House recording, your first step is timebase compensation. You need to estimate the rotor slip of the original Ampex recorder and apply a correction before anything else. Without this, every subsequent analysis is built on unstable ground. Use a known reference tone if one exists in the recording, or model the expected slip based on the recorder's servo specifications from that era. Next, run a spectral subtraction pass to reduce the ambient noise. The limousine cabin at the time of the recording had engine noise, road noise, and the open convertible environment working against clarity. A standard Wiener filter approach works, but keep the reduction conservative. Over-aggressive noise suppression will remove the very high-frequency formants you are trying to measure.
Get the Full Details

After noise reduction, align your target recording against reference samples. This is the calibration baseline I mentioned earlier. Take the known JFK samples from the Pentagon tapes recorded in March 1962 and the April 1963 White House meeting with Soviet Premier Khrushchev's aides. These give you a reliable formant trajectory for his voice under normal conditions. Compare the November 22 fragment against this trajectory, not against an abstract average male voice. When doing the comparison, focus on F1 and F2 formant tracking across vowel contexts. JFK's vowel space in the November 22 recordings is compressed compared to his earlier samples, which researchers have attributed to stress and the physical environment of the motorcade. Do not treat this compression as evidence of anything other than situational factors. It is a well-documented phenomenon in forensic phonetics.
What This Approach Cannot Do
Be honest about the limits. You cannot conclusively identify who is speaking on the White House fragment using current technology. The fragment is too short, too degraded, and the comparison material, while useful as a baseline, does not eliminate the possibility of other voices matching the available spectral information. Any claim to the contrary is not forensic science. The police radio recordings present a different set of problems. The bandwidth is limited to roughly 300 Hz to 3.4 kHz due to the radio channel characteristics. This cuts off the higher formants entirely, leaving you with only F1 and partial F2 information. That narrows your analytical options significantly. You can do speaker characterization at a basic level, but detailed formant analysis is not viable. One thing I should mention that most guides leave out: the magnetic degradation of the original tapes matters more than people realize. The White House tape has undergone multiple generations of copying and re-copying over sixty years. Each generation introduces additional high-frequency loss. If you are working from a secondary copy rather than the original tape, your effective bandwidth may be reduced by another octave or more compared to what the original contained. Always check the provenance chain of the audio file you are analyzing.
The tools you need are fairly standard in forensic labs. SpectraVue orPraat for the spectral analysis, a calibrated microphone and audio interface for any reference recordings you make, and a solid understanding of timebase error in vintage tape machines. There is no single software package that handles all of this automatically. You have to do the corrections manually, which means you need to understand what each correction is doing to the signal. If you are approaching this because you encountered references to the F Kennedy Voice Problem online and want to do something substantial with it, start by reading the National Academy of Sciences report on voice identification from 1979 and the more recent NIST guidelines on forensic speech analysis. Those will give you a better foundation than any forum thread or YouTube video. The gap between what amateur analysis claims and what forensic phonetics actually supports is enormous, and most people working on this without proper training do not realize how big that gap is.
