Getting started with Speech In Ohio Today
I first ran into this service when a client needed legal deposition transcripts turned around within 48 hours and couldn't rely on the usual three-day turnaround. The standard approach most people take is to sign up, upload a file, and wait. That works fine for short audio, but it breaks down quickly when you're dealing with multi-speaker recordings, poor audio quality, or thick regional accents. Here is what actually happens when you use Speech In Ohio Today and how to not waste your money on it. After you create an account, there is a settings page most people skip entirely. The default language model is set to general American English, which means if you are uploading recordings from rural Kentucky border counties or Cleveland neighborhoods with heavy Slavic influence, the accuracy drops by roughly 12 to 18 percent before you even hit submit. Go into Settings and switch the dialect model to Mid-Atlantic if your speakers lean that direction, or enable the custom vocabulary field. I had a client recording a series of zoning board meetings in Hamilton County where the transcriber kept rendering "zoning" as "soning" until I uploaded a custom word list with the actual ordinance numbers and street names they kept using. That one change brought the accuracy from about 78 percent up to 94 percent on the first pass. The upload process supports WAV, MP3, M4A, and FLAC files up to 2 gigabytes. Anything over 90 minutes gets chunked automatically by the system, but the auto-chunking often splits mid-sentence at paragraph breaks, which makes the exported text look fragmented. I now manually split any recording longer than 60 minutes before uploading using Audacity, saving each section as a separate WAV file and labeling them chronologically. The exports come back in order every time, and you avoid the awkward reassembly step later.
Understanding what the output actually looks like
The default export is a plain text file with no speaker identification. If you need speaker labels, you have to purchase the premium tier, which adds automated diarization at roughly $0.25 per audio minute. A full-length deposition that runs three hours will cost you about $45 extra just for the speaker tracking. That is not a terrible rate compared to manual annotation, but it is worth knowing upfront because the free tier gives you zero attribution. The platform also offers an SRT subtitle export, which most people do not realize is fully timestamped to the millisecond. I use that feature when I need to cross-reference a transcript against video evidence. The timestamps line up within a 200-millisecond margin of error, which is tight enough for legal review but loose enough that you should never rely on it for forensic audio analysis. If precision matters that much, you would be better off using a dedicated tool like Praat or Adobe Audition's waveform alignment features instead. There is a common misconception that the service improves with each upload in a session. It does not. Each file is processed independently on their servers, and the language model does not learn from your custom vocabulary across files unless you explicitly save that vocabulary to your account profile. I learned this the hard way when I spent six hours uploading trial exhibits one by one, only to discover the system had not remembered the specialized terminology I had entered on the first five files. Saving the vocabulary to your profile took about ten seconds and corrected the issue for the remaining twelve uploads.
When Speech In Ohio Today actually fails
The service struggles noticeably with background noise above 65 decibels. Restaurant recordings, crowded courtroom halls, and audio captured on smartphone microphones in windy conditions all produce garbage output. I tested this on a recording made at a public meeting in Dayton where the HVAC system was running and the mic was about eight feet from the speaker. The raw transcript had an error rate of nearly 34 percent. Running the same file through their noise reduction preprocessing step brought it down to about 19 percent, but even that baseline is unacceptable for professional use without manual correction. Another hard limitation is code-switching. If your speakers alternate between English and Spanish mid-sentence, the transcriber will output the English portions accurately and either drop the Spanish entirely or substitute phonetic approximations that are useless for translation. I encountered this with a client working on immigration hearings where the witnesses switched languages frequently. The workaround was to run the audio through a separate speech-to-text model trained on bilingual data first, then use the Speech In Ohio Today output only for the English segments and manually transcribe the Spanish parts. It added about forty-five minutes of work per hour of audio, but the result was reliable enough to submit. Pricing runs at $0.15 per minute for the basic plan and $0.40 per minute for the premium tier with diarization and custom vocabulary. There is a monthly subscription option at $29 that includes 120 minutes of processing, which works out to about $0.24 per minute if you hit the cap. If you go over, you are charged the per-minute rate on whatever remains. Most people do not track their usage and get surprised by the overage charges. I set a calendar reminder every month to check my dashboard and adjust my subscription tier accordingly, which has prevented any unexpected bills over the past fourteen months.
Get the Full Details

A practical workflow that actually saves time
Here is the sequence I follow now, and it takes me about fifteen minutes from receiving an audio file to having a clean, editable transcript ready for review. First, I listen to the entire file at 1.5x speed in Audacity to identify problem sections, loud background noise, and any speaker changes. I flag those timestamps and note them in a spreadsheet. Second, I preprocess the audio by applying a high-pass filter at 80 hertz to remove rumble and a noise gate set to -40 decibels to silence dead air. This step alone reduces the error rate by about 8 percent on average. Third, I upload the cleaned file with my saved custom vocabulary already loaded. Fourth, I download the transcript and open it in a text editor where I run a search for any flagged timestamps to manually verify the disputed sections. Fifth, I export the final version as a DOCX if the client needs formatting or as plain TXT if they just need the raw text. The whole pipeline replaces what used to take me about two hours per hour of audio. The preprocessing and manual verification steps are not optional if you want professional-grade output, but they are fast enough that most of the time cost goes away. The bottleneck is always the manual review of flagged sections, which varies wildly depending on how clean the original recording is. A well-mic'd studio recording needs almost no review. A phone call recorded in a moving car needs nearly line-by-line correction.
Alternatives worth considering
If your use case involves heavily accented speakers, continuous background noise, or multilingual content, Speech In Ohio Today is not the best tool. AssemblyAI handles accent variation significantly better and supports over thirty languages natively, though their pricing is steeper at $0.000218 per second, which translates to roughly $0.79 per minute. Descript is another option if you prefer a visual editor where you can edit the transcript and have the audio update in real time, but it requires a subscription rather than pay-per-use, which makes it expensive for occasional users. For one-off projects under thirty minutes with clear audio, the free tier of Speech In Ohio Today is perfectly adequate, but anything beyond that warrants a closer look at the paid options or the alternatives listed here.