Learn how to properly transcribe an interview: consent and recording basics, verbatim vs cleaned, accurate speaker labels, and verifying quotes before you publish.

How to Transcribe an Interview Properly: A Practical Guide

By FileToTextPublished

A good interview transcript is more than a wall of text. It gets the words right, marks who said what, and holds up when you quote it back weeks later. Whether you are a journalist, a researcher running qualitative studies, a podcaster, or a hiring manager reviewing candidates, the process is the same. The fastest way to start is to upload your recording to FileToText's interview transcription tool, which turns an audio or video file into a labeled, timestamped draft you can then refine.

Automation does the bulk of the work in minutes. The rest, the part that separates a usable transcript from a sloppy one, is judgment: consent, formatting choices, speaker accuracy, and checking your quotes against the actual audio. Here is how to do all of it properly.

1. Sort out consent and recording law first

This is general guidance, not legal advice, but it matters before you hit record. In the United States, recording rules vary by state. Roughly a dozen states are "all-party" (sometimes called two-party) states, where everyone in the conversation must consent to being recorded. The rest are "one-party" states, where only one participant (which can be you) needs to consent. When people are in different states, the stricter rule usually wins, so default to getting consent from everyone.

In the EU and UK, personal data rules like GDPR apply to recordings of identifiable people. The safe, professional habit everywhere is simple: ask on the record at the start of the interview, get a clear yes, and keep that on the tape. For published journalism or research, you may also need a separate release for how the material will be used.

2. Get a clean recording

Accuracy starts at the microphone, not the software. Automatic transcription reaches 95%+ on clear audio, but drops with background noise, strong accents, overlapping speakers, and narrow-band phone audio. A few habits pay off later:

  • Record in a quiet room and cut background music or TV.
  • Give each speaker their own mic where you can, and ask people not to talk over each other.
  • For remote interviews, record locally rather than relying on a compressed phone line.
  • Do a 10-second test and play it back before the real conversation.

3. Turn the audio into a first draft

Upload your file and let the tool do the heavy lifting. FileToText accepts uploaded audio and video (MP3, WAV, M4A, MP4, MOV and more) and auto-extracts the audio from video. You get speaker labels, per-word timestamps, and a player that highlights each word as it plays so you can click any word to jump straight to that moment. The first 10 minutes of any file are free on a free account, so you can see the quality before committing.

One caveat: it works on files you upload, not on a YouTube page link. To transcribe a YouTube interview, download the video first, then upload it.

4. Choose verbatim or cleaned

Decide what the transcript is for before you edit it.

Verbatim

Every "um", false start, and stutter is captured. Use this for legal contexts, discourse analysis, or any research where how something was said matters as much as what was said.

Cleaned (or clean verbatim)

Filler words and stumbles are removed so the meaning reads clearly. Use this for articles, show notes, and quotes for publication. FileToText's filler-word cleanup helps here, and its AI summary gives you a prose recap of the conversation (not auto-chapters or bulleted action items) to orient yourself fast.

5. Fix the speaker labels

Automatic speaker separation is good, not perfect, especially when people interrupt each other. After the draft is generated, rename each speaker (from "Speaker 1" to a real name) and reassign any lines that were attributed to the wrong person. Do this pass early, because every later step depends on knowing who said what.

6. Verify every quote against the audio before you publish

This is the step most people skip, and it is where reputations are made or lost. Never publish a quote you have not heard yourself. Two features make this quick: the per-word confidence view flags words the model was unsure about (names, jargon, and numbers are common offenders), and the click-to-jump player lets you land on the exact second a quote was spoken and listen back. Check names, figures, dates, and anything you will put in quotation marks.

7. Cost: software versus a human service

Human transcription services are accurate and remain the right call when the stakes are legal or medical, but turnaround is typically a day or more. FileToText charges $3 per audio hour after the free first 10 minutes, billed per started hour with a $3 minimum, paid once per file, and returns the draft in minutes. If you transcribe regularly, subscriptions run from $10/mo for 600 minutes up to $50/mo for 6,000 minutes. There is a full refund if the output is not accurate.

For most interviews, the smart workflow is software for the draft plus your own verification pass. You capture the speed of automation while keeping a final human check on the quotes that matter. When you need to share, export to TXT, DOCX, SRT, or WebVTT, or send a read-only link, and translate into 90+ languages while keeping speaker labels and timestamps intact.

Common questions

Is it legal to record and transcribe an interview?

This is general guidance, not legal advice. In the US it depends on your state: one-party states need only your consent, while all-party (two-party) states require everyone to agree. In the EU and UK, GDPR applies to recordings of identifiable people. The safe habit everywhere is to ask for consent on the record at the start.

Should an interview transcript be verbatim or cleaned up?

It depends on the purpose. Use full verbatim (including "um" and false starts) for legal or research contexts where delivery matters. Use cleaned verbatim, with fillers removed, for articles, podcasts, and published quotes where readability matters more.

How accurate is automatic interview transcription?

FileToText reaches 95%+ on clear audio. Accuracy drops with background noise, strong accents, overlapping speakers, or narrow-band phone audio. A clean recording and a human verification pass close most of that gap.

How much does it cost to transcribe an interview?

The first 10 minutes of any file are free. After that FileToText charges $3 per audio hour, billed per started hour with a $3 minimum, paid once per file, so a one-hour interview costs about $3. Not satisfied? Full refund, no questions asked.

How do I get accurate speaker labels?

Start with a good recording (separate mics, minimal crosstalk), then refine in the tool. After the draft is generated you can rename each speaker to a real name and reassign any lines attributed to the wrong person.