Interview transcription
Upload the recording and get a clean, speaker-labeled transcript in minutes, synced to the audio so you can verify any quote in one click. Verbatim when you need it, cleaned up when you don't. The first 10 minutes of any file are free.
A real interview transcript synced to its audio. Press play, or click any word to jump straight to that moment. This is exactly what your recording becomes.
Researchers get timestamped citations for the appendix, journalists get quotes verified against the source audio before print, and HR gets an attributable record. A Word document never knew where minute 43 was; this does.
“It was almost an accident. I was working a completely different job, and I stumbled into a project that just clicked.”
“It was the first time the work didn't feel like work. I'd look up and three hours had gone by.”
Every line is timestamped and speaker-labeled, and the transcript is synced to the audio: click any quote to jump to that exact moment and hear it before you use it. The per-word confidence view shades the words the model was least sure about, so you check the few that matter instead of proofreading everything.
Can you record it, verbatim or cleaned, and what it will cost. Straight answers, right here on the page.
Recording law depends on where you (and they) are. Pick your region.
Typically: one-party consent
In most US states (federal law included), one party to the conversation may record it. If you're conducting the interview, that's you, so you can record. Telling the other person is still good practice and builds trust.
General guidance, not legal advice. Laws change and edge cases exist; when in doubt, ask and get the yes on tape.
Picking the wrong one costs you a redo. What's the transcript for?
Cleaned up, verify against audio
Trimming an 'um' from a quote is standard; changing words is not. Clean the transcript to read well, then click each quote to hear it before you print it.
FileToText keeps both: filler cleanup gives you the readable version, and the original stays one toggle away.
Automatic transcription plus your own review, versus typing it out yourself.
FileToText
$15
$3 / started audio hour (max $12 per file), ready in minutes
Typing it yourself
~20 hours
of your time (typing speech takes roughly 4x the audio length)
That is roughly 20 hours of your time back for $15, with transcripts the same morning, and your review goes into checking the quotes that matter instead of typing them. Hours round up: a 1 hour 2 minute recording bills as 2 hours. Not satisfied? Full refund, no questions asked.
The choices people skip until it's too late.
Speaker detection works by telling voices apart, and a classic interviewer-and-subject recording is close to the best case: two distinct voices, mostly taking turns. Upload, get "Speaker 1" and "Speaker 2," rename them once. For a panel or a group research session, set the speaker count at upload; guessing between five similar voices is genuinely harder, and the hint helps.
When your subject talks over you, expect to double-check a few attributions by ear and correct them in your export. That is true of every automatic system, ours included. The fix for next time costs nothing: let answers finish before you jump in. Your transcript improves and, conveniently, so do your interviews.
The most useful habit for anyone who quotes interviews: skim the transcript, mark the lines you might use, then click each one and listen. Ten seconds per quote. The per-word confidence view flags the words most likely to be wrong, which is exactly where a misquote would hide. If the interview was recorded on video, you can also export SRT or WebVTT. For the mechanics of getting audio off Zoom or a phone, our podcast page covers the same export steps.