Getting Started with KAHNA Academy Speech Processing Statistics

I ran into Kahna Academny Sp Statistics while tuning a voice recognition pipeline for a client last year. The software isn't exactly well-documented, which is kind of the point of writing this. People who need it usually already know where to find it. The stats themselves measure several key performance markers for KAHNA's speech processing modules, things like phoneme accuracy, word error rates, and speaker diarization scores. The statistics break down into a handful of categories. Phoneme-level accuracy tracks how often individual sound units are recognized correctly. Word error rate is the more common metric people look at first, but it hides a lot of useful information. KAHNA provides detailed sub-word analysis that most people skip over. Speaker ID confidence scores tell you how sure the system is about who is talking at any given moment. Latency numbers round it out, showing processing time from audio input to text output. I pulled these stats from the KAHNA web interface after running a batch of transcriptions through their engine. The download button lives under the Reports section, labeled as "Sp Statistics Export." It gives you a CSV file with per-segment breakdowns. That file is dense. The default view in the dashboard shows aggregate numbers, but the export has everything granularly.

Here is a specific problem I hit. The export sometimes splits speaker labels incorrectly when two people talk over each other. I noticed my diarization confidence scores were flatlining around 0.42 even though the transcript looked fine to a human reader. The issue was that KAHNA was treating the overlapping speech as a single confused segment rather than two distinct voices. I worked around it by increasing the overlap threshold in the processing config before running the job. Setting overlap_detection to aggressive cut the bad segments in half. It is not perfect. You will still get occasional misalignments, but they become manageable. One thing beginners miss is that the word error rate in KAHNA does not penalize punctuation the same way other engines do. TheirWER calculation is text-only by default. If you need punctuation accuracy factored in, you have to enable the punctuation penalty mode in settings. It is off by default across the board. I learned this the hard way when my results looked suspiciously good compared to what was actually happening in the transcript. Another counter-intuitive detail. Higher phoneme accuracy does not always mean better transcription quality. KAHNA sometimes over-corrects phonemes toward common words, which boosts the phoneme score but can actually degrade the final text in domain-specific contexts. If you are working with medical or legal terminology, run your evaluation with a domain glossary loaded. The baseline stats without one will mislead you.

The download link for the software itself is on the KAHNA developer portal at kahna-academy.org/stats. You need a registered account to access the export tools. Free accounts get limited to thirty days of history. Paid accounts unlock the full timeline and raw feature vectors. There are real limitations here. The system struggles with heavy background noise. I tested it in a call center environment with multiple overlapping conversations and the speaker diarization fell apart entirely. The phoneme accuracy dropped to roughly 61 percent, which is well below acceptable thresholds. If your use case involves noisy audio, you will need to preprocess the files through a denoiser first. KAHNA does not include a built-in noise reduction module in the standard package. The Sp Statistics themselves do not flag noise-related degradation automatically, so you are on your own to spot the pattern. A common alternative for noisy environments is running Whisper separately and using its output as a reference layer. It is slower, but it handles acoustic challenges better. Some people combine both systems, using KAHNA for clean recordings where its speaker tracking excels and Whisper as a fallback for rough audio.

Get the Full Details

Khan Academy Statistics And Facts (2025)
Khan Academy Statistics And Facts (2025)

The stats export also lacks confidence intervals for its error rates. If you need statistical rigor for a paper or report, you will have to calculate those yourself from the raw data. The CSV gives you enough to work with, but the math is not done for you. Expect to spend some time on that part.