Interview transcription

Interview transcription, every quote checkable against the audio

Upload the recording and get a clean, speaker-labeled transcript in minutes, synced to the audio so you can verify any quote in one click. Verbatim when you need it, cleaned up when you don't. The first 10 minutes of any file are free.

Transcribe your interviewFirst 10 minutes free. No credit card.

Hear it, read it, at the same time

A real interview transcript synced to its audio. Press play, or click any word to jump straight to that moment. This is exactly what your recording becomes.

Sample preview
play it · click any word to jump

Interviewer:Thanks for making the time. So to start, tell me a little about what you do and how you got here.

Guest:Sure. Um, so I run a small coffee roastery in Amsterdam, just east of the Oosterpark. We roast, uh, twice a week, and we supply about thirty cafes around the city. Before this, I was an accountant, which, you know, surprises people.

Interviewer:An accountant. That's quite a jump. What made you leave?

Guest:Honestly, it wasn't one big moment. I'd been doing the books for a cafe client, and I kept staying after the work was done, just asking about the machines, the beans, all of it. At some point, the owner said, "Well, if you're gonna hang around anyway, learn to roast." And that was it. I, I couldn't stop thinking about it after that.

Interviewer:So when did it become a real business?

Guest:About six years ago. We started with a tiny five-kilo machine in a shed. First year, we lost money. Second year, we broke even. And, um, by the third year, we could pay two salaries. Now there are seven of us.

Interviewer:Seven people. Is there a coffee you're especially proud of?

Guest:Yeah. There's a lot from Huehuetenango in Guatemala that we've bought from the same family for four seasons now. When it lands each spring, it's, uh, it's a bit like a harvest festival in here. Everyone stops to cup it together.

Interviewer:And what's been the hardest part of growing?

Guest:Hiring, without a doubt. I mean, roasting, you can learn. The machines are honest with you. But finding people who care about the boring parts, the cleaning, the logging, the consistency, that's rare. We got it wrong twice before we got it right.

Interviewer:What does a normal Tuesday look like for you these days?

Guest:Earlier than you'd think. The roaster gets switched on at seven because it needs a, a good forty minutes to warm up. Mornings are production. Afternoons are the unglamorous half: packing orders, answering wholesale emails, calibrating the grinders. And honestly, twice a week, I still do the deliveries myself on the cargo bike. You hear things from cafe owners that never make it into an email.

Interviewer:Is there a mistake from the early days that still makes you wince?

Guest:Oh, plenty. The worst one, we once shipped a whole batch that I'd roasted a full minute too short. It tasted like, um, like grass and peanuts. A customer called and asked very politely whether we were okay. We replaced everything, and I printed that invoice and hung it above the roaster. It's still there.

Interviewer:Last one: What would you tell someone who's sitting where you were in the office thinking about it?

Guest:Don't romanticize it. The first winter, I was loading sacks at six in the morning wondering what I'd done. But, like, if a thing keeps pulling at you for years, that's information. You can go back to spreadsheets. You can't go back in time.

Interviewer:That's a good place to end. Thank you.

Guest:Thanks for having me.

0:00
2:57

The point of an interview transcript: quotes you can stand behind

Researchers get timestamped citations for the appendix, journalists get quotes verified against the source audio before print, and HR gets an attributable record. A Word document never knew where minute 43 was; this does.

Citable quotes, verifiable in one click
It was almost an accident. I was working a completely different job, and I stumbled into a project that just clicked.
00:05·Guestclick to hear it in the transcript
It was the first time the work didn't feel like work. I'd look up and three hours had gone by.
00:21·Guestclick to hear it in the transcript

Every line is timestamped and speaker-labeled, and the transcript is synced to the audio: click any quote to jump to that exact moment and hear it before you use it. The per-word confidence view shades the words the model was least sure about, so you check the few that matter instead of proofreading everything.

The three decisions every interviewer makes

Can you record it, verbatim or cleaned, and what it will cost. Straight answers, right here on the page.

Can you record it? Check before you hit go

Recording law depends on where you (and they) are. Pick your region.

Typically: one-party consent

In most US states (federal law included), one party to the conversation may record it. If you're conducting the interview, that's you, so you can record. Telling the other person is still good practice and builds trust.

General guidance, not legal advice. Laws change and edge cases exist; when in doubt, ask and get the yes on tape.

Verbatim or cleaned up? Decide on purpose.

Picking the wrong one costs you a redo. What's the transcript for?

Cleaned up, verify against audio

Trimming an 'um' from a quote is standard; changing words is not. Clean the transcript to read well, then click each quote to hear it before you print it.

FileToText keeps both: filler cleanup gives you the readable version, and the original stays one toggle away.

What transcribing your interviews actually costs

Automatic transcription plus your own review, versus typing it out yourself.

5 hours

FileToText

$15

$3 / started audio hour (max $12 per file), ready in minutes

Typing it yourself

~20 hours

of your time (typing speech takes roughly 4x the audio length)

That is roughly 20 hours of your time back for $15, with transcripts the same morning, and your review goes into checking the quotes that matter instead of typing them. Hours round up: a 1 hour 2 minute recording bills as 2 hours. Not satisfied? Full refund, no questions asked.

Interview-specific things worth knowing

The choices people skip until it's too late.

Two-person interviews are the easy case

Speaker detection works by telling voices apart, and a classic interviewer-and-subject recording is close to the best case: two distinct voices, mostly taking turns. Upload, get "Speaker 1" and "Speaker 2," rename them once. For a panel or a group research session, set the speaker count at upload; guessing between five similar voices is genuinely harder, and the hint helps.

Overlapping speech is where you review

When your subject talks over you, expect to double-check a few attributions by ear and correct them in your export. That is true of every automatic system, ours included. The fix for next time costs nothing: let answers finish before you jump in. Your transcript improves and, conveniently, so do your interviews.

Never trust a line you haven't heard

The most useful habit for anyone who quotes interviews: skim the transcript, mark the lines you might use, then click each one and listen. Ten seconds per quote. The per-word confidence view flags the words most likely to be wrong, which is exactly where a misquote would hide. If the interview was recorded on video, you can also export SRT or WebVTT. For the mechanics of getting audio off Zoom or a phone, our podcast page covers the same export steps.

Interview transcription FAQ

Updated July 2026
Upload the recording (audio or video) and FileToText returns a speaker-labeled, timestamped transcript in minutes, synced to the audio so you can click any quote and hear it. Verbatim and cleaned-up versions are both available, and the per-word confidence view flags the words most likely to be wrong, which is exactly where a misquote hides. The first 10 minutes are free; beyond that it is $3 per audio hour, billed per started hour. Not satisfied? Full refund, no questions asked. A two-person interview separates cleanly; for a panel, set the speaker count at upload. Accuracy is 95%+ on clear audio, lower when people talk over each other. Export as TXT, DOCX, SRT, or WebVTT once you have verified the quotes you plan to use. See the live monthly benchmark →
It depends on where you and the other person are. Most US states (and federal law) are one-party consent, so as the interviewer you can record; a handful, like California and Illinois, need everyone's consent. The EU and UK treat a recording as personal data, so get informed consent and state the purpose. Whatever the rule, telling people you're recording and getting the yes on tape is the safe default. This is general guidance, not legal advice.
Decide on purpose. Discourse and conversation analysis need verbatim, because the hesitations are data. Thematic research, journalism, and podcast prep usually want the cleaned-up version so it reads well. HR records should stay verbatim as the official document. FileToText keeps both: filler-word cleanup gives you the readable version and the original stays one toggle away.
Never trust a transcript line you haven't heard. The transcript is synced to the audio word by word, so click any quote and playback jumps to that exact moment with the words highlighting as they play. The per-word confidence view shades the words the model was least sure about, which is exactly where a misquote would hide, so a ten-second listen per quote is enough.
A classic two-person interview is close to the best case: two distinct voices, mostly taking turns. It labels them automatically and you rename "Speaker 1" to a real name once. For panels or group sessions, set the speaker count at upload; guessing between several similar voices is genuinely harder and the hint helps. Overlapping speech is where every tool needs a quick human pass.
It is $3 per audio hour, billed per started hour, with a $3 minimum per file, and the transcript comes back in minutes with speaker labels and timestamps. Test the quality free on the first 10 minutes of any file. Not satisfied? Full refund, no questions asked.
Yes. Drop up to 10 recordings in one batch and each becomes its own transcript, so a day of fieldwork queues in one go. Every file keeps its own free first 10 minutes, and long sessions run up to 20 hours per file.