Skip to content
Audio Editing & Production

Transcription and Audio Accessibility: The Decisions That Matter

Learn what accessibility rules require for audio transcripts, how to edit automatic drafts, label speakers, publish on the page, and when to get outside help.

Tomas Lindqvist Post-Production Lead 27 min read 22 views
Transcription and Audio Accessibility: The Decisions That Matter

Transcription and audio accessibility come down to one simple act: providing a text version of spoken audio so it can be read, searched and used by people who cannot listen or would rather not. A podcast episode, a recorded webinar, an interview, a training module and an audio guide all carry information that only reaches people who can hear it, at the pace it was spoken, in the order it was recorded. A good transcript removes all three limits. Deaf and hard-of-hearing readers can use it. So can someone on a train with no headphones, a non-native speaker who wants to check a word, a researcher looking for one quote, and a search engine that cannot listen at all.

This guide is for marketers, producers, communications teams and business owners who publish audio and want to get transcripts right without turning them into a second full-time job. It works through the questions people actually ask when they set up a transcription workflow: what the accessibility standards require, whether automatic transcription is good enough, how much editing is realistic, how to label speakers, where the transcript should live, and when it makes sense to hand the job to someone else. Each answer comes first, followed by the reasoning and the practical detail behind it.

The short version is that a transcript is not an afterthought to the audio. It is a second published asset, with its own quality bar, its own editing pass and its own place on the page. Teams that treat it that way get accessibility, discoverability and a reusable text record from every recording. Teams that paste in raw machine output get none of those things reliably.

What do accessibility standards actually require for audio content?

For prerecorded audio-only content, such as a podcast episode or a recorded talk with no video, the Web Content Accessibility Guidelines require a text alternative, in practice a transcript, at Level A, the most basic conformance level. That makes a transcript a baseline expectation, not a nice extra. The W3C Web Accessibility Initiative's guidance on transcripts and audio alternatives explains the requirement and the reasoning behind it: people who cannot hear the audio need an equivalent way to get the same information.

It helps to be precise about which requirement applies to which kind of media, because teams often mix them up:

  • Audio-only, prerecorded (podcasts, audio articles, recorded phone interviews): a transcript that presents the same information as the audio. This is WCAG success criterion 1.2.1, Level A.
  • Video with sound, prerecorded (webinars, explainer videos, recorded presentations): synchronized captions, under success criterion 1.2.2, also Level A. A transcript is a useful addition, but it does not replace captions.
  • Live audio and video: live video requires captions at Level AA under 1.2.4, and live audio-only content has its own, stricter criterion at AAA. Live content is a different operational problem, covered in our guide to getting live stream audio right.

"Equivalent" is the operative word. A transcript that summarizes the conversation, drops the tangents that happen to contain a key figure, or leaves out who said what does not give a reader the same information as a listener. Where meaningful non-speech sound carries information, such as a demonstration of a product's alert tone, laughter that changes the meaning of a remark, or a music cue introducing a new segment, the transcript should note it briefly in brackets.

Level AWCAG conformance level at which prerecorded audio-only content needs a text alternative
1.2.1WCAG success criterion covering prerecorded audio-only and video-only media
1.2.2WCAG success criterion requiring captions for prerecorded video with sound

Legal obligations vary by country and sector, and many public bodies, universities and regulated businesses reference WCAG directly in their procurement rules or their own policies. If your organization has signed a contract or policy that commits to a WCAG level, the transcript requirement comes with it. Even where no rule applies, the WCAG criteria are the clearest definition of "accessible" available, so they are the sensible benchmark to build to.

Is an automatic transcript good enough to publish as it is?

No. Automatic transcription is an excellent first draft, but its output needs human correction before it is published. Accuracy varies with accent, terminology and audio quality, and the errors it makes are concentrated in exactly the words that matter most.

Modern speech recognition handles clear, single-speaker, plain-language audio well. It struggles, predictably, in five situations:

  1. Proper nouns. Guest names, company names, product names and place names are often not in the model's vocabulary, so they come out as the nearest common words. A guest called Siobhan can become "Chevonne", and a software product can become an unrelated English phrase.
  2. Technical and industry terms. Jargon, acronyms and brand-specific vocabulary are misheard or split into several words. In a specialist episode, this can affect a noticeable share of the sentences that carry the actual expertise.
  3. Numbers. "Fifteen" and "fifty" are easy to confuse, and spoken figures such as "two point four million" can be rendered inconsistently. A wrong figure in a transcript is worse than a missing one, because readers quote it.
  4. Overlapping speech and crosstalk. When two people talk at once, the recognizer usually picks one and drops or garbles the other.
  5. Poor recordings. Room echo, background noise, distant microphones and compressed phone or video-call audio all reduce accuracy, sometimes sharply.

The practical consequence is that a transcript can look almost right at a glance and still be wrong in the places a reader relies on. A skim of the first paragraph tells you very little. The edit has to be a full pass against the audio, with particular attention to names, terms and numbers.

It is tempting to assume that automatic transcription is now accurate enough that editing is optional. It is not. Automatic output is often readable, but its errors cluster in names, technical terms and figures, which are the words readers search for and quote. Publishing it uncorrected puts those errors on the record under your name.

There is also a reputational side. An uncorrected transcript that misspells a guest's name or turns a client's product into nonsense is visible, permanent and searchable. For a business publishing audio as part of its marketing, that is a credibility problem as much as an accessibility one.

How accurate does a transcript need to be, and who should check it?

A published transcript should be accurate enough that a reader who never hears the audio gets the same facts, names and meaning as a listener, and it should be checked by someone who knows the subject, not just someone who can type. This is the hardest question in the whole process, because "accurate" means different things for different uses, and the person best placed to check it is often not the person with time to do it.

Choose a transcription style before you start

There are two broad styles, and mixing them in one transcript produces something that reads badly and is hard to check:

  • Clean verbatim removes filler words ("um", "you know"), false starts and repeated words, while keeping the speaker's meaning and wording. It is the right default for podcasts, interviews and marketing content, because it reads well and is still faithful.
  • Full verbatim keeps every hesitation, repetition and non-speech sound. It is needed for research interviews, legal and compliance records, and anywhere the manner of speaking is itself evidence.

Write the choice down in a short style sheet, along with how you handle numbers (numerals or words), units, brand capitalization, and bracketed notes such as [laughter] or [inaudible 12:34]. A one-page style sheet removes most of the inconsistency that creeps in when several people edit transcripts.

Split the check into two passes

The most reliable approach separates mechanical correction from subject-matter review. A transcript editor listens through the whole recording against the draft, fixes mishearings, applies the style sheet, labels speakers and marks anything uncertain with a timestamp. Then someone who knows the field, often the host, producer or the guest's own team, reviews only the flagged items plus every name, term and figure. The subject-matter pass is short because the editor has already done the heavy lifting and highlighted exactly where judgment is needed.

This matters because a skilled transcript editor can hear a word perfectly and still not know whether the guest said a product name or a common phrase that sounds the same. Only someone who knows the domain can resolve that. Sending the subject expert a full, unflagged transcript to "have a look at" rarely works; sending them a short list of questions with timestamps usually does. The same principle underpins good client audio review, described in our guide to reviewing audio with clients.

Insight: The accuracy of a transcript is decided less by the speech recognition engine than by what you feed it and who checks it. Three levers do most of the work:

  • A clean recording, with each speaker on their own close microphone, reduces errors before any software runs.
  • A custom vocabulary or glossary of names and terms, where the tool supports one, prevents the most damaging mistakes at the source.
  • A flagged, timestamped question list turns subject-matter review from an hour of reading into a few minutes of answers.

Decide what "done" means for each use

Set the bar by use. A transcript published as an accessibility alternative must be complete and faithful. A transcript used internally to find quotes can tolerate more rough edges. A transcript that forms part of a contract, a regulatory record or a translation source needs full verbatim, sign-off and version control. Put the expectation in writing so that "the transcript is ready" means the same thing to everyone.

Why do speaker labels matter so much in multi-speaker recordings?

Because without them, a reader cannot tell who said what, and in an interview or panel the speaker is often as important as the words. Speaker identification is necessary in any recording with more than one voice, and a transcript with no speaker labels fails the basic test of giving a reader the same information as a listener.

A listener hears the difference between the host and the guest instantly, through voice, accent and tone. A reader has only the text. If a claim, a recommendation or a disagreement is attributed to the wrong person, the transcript is not just unhelpful but misleading. In a panel discussion with four or five speakers, unlabeled text becomes almost unusable.

How to label speakers well

  • Use full names at first appearance, then a consistent short form. "Maya Chen:" at the start, then "Maya:" throughout, is easier to scan than repeating the full name, and clearer than initials.
  • Avoid generic labels such as "Speaker 1" and "Speaker 2" in anything you publish. They are acceptable in an internal draft, but a reader should not have to work out which number is the guest.
  • Start a new paragraph at every change of speaker, with the label in bold at the start. This is the single biggest readability improvement you can make.
  • Handle crosstalk explicitly. When two people speak at once, transcribe the main line and add the interjection in brackets, or break it into two short turns. Do not silently drop it if it changes the meaning.
  • Check diarization output carefully. Automatic speaker separation ("diarization") often switches labels when voices are similar, when someone laughs, or when a speaker interrupts. It is one of the most common sources of errors that survive a quick read.

Recording practice helps a great deal here. When each participant is recorded on a separate track, both automatic tools and human editors can attribute speech with far more confidence. That is one of the many reasons we recommend separate tracks in our guide to audio workflow and session organization. For hybrid meetings and recorded calls, the microphone setup in the room matters just as much, because a single distant microphone makes every voice sound alike to the software.

Which errors should a transcript editor prioritize?

Names, technical terms and figures, in that order, because they are what automatic systems get wrong most often and what readers rely on most. Everything else, such as punctuation, filler words and minor phrasing, matters, but a mistake there rarely changes the meaning.

A useful way to think about priority is to ask what harm each error causes if it reaches publication:

Error typeTypical exampleWhy it mattersHow to catch it
Misspelled or misheard nameA guest's surname rendered phoneticallyOffends the guest, breaks name searches, looks carelessConfirm spellings from the booking form or guest bio before editing
Wrong technical termAn acronym split into ordinary wordsMakes expert content read as nonsense and loses search relevanceKeep a running glossary per show or client
Wrong figure"Fifteen" transcribed as "fifty"Readers quote it; it can create a factual or compliance problemListen again to every number; check against source material
Wrong speaker attributionGuest's claim labeled as the host'sMisrepresents who said whatSeparate tracks; spot-check label changes
Dropped passageCrosstalk or a mumbled aside omittedTranscript is no longer equivalent to the audioMark gaps as [inaudible] with a timestamp rather than skipping
Punctuation and fillerRun-on sentences, stray "um"Reduces readability, rarely changes meaningApply the style sheet in the final read

Two habits pay off disproportionately. First, build a glossary before you transcribe, not after: collect guest names, company and product names, and recurring jargon from the episode notes, the guest's website and previous transcripts. Many transcription tools accept a custom vocabulary list, and a glossary also gives the human editor a reference. Second, treat every number as a checkpoint. Listen again to each one, and where the speaker cites a published figure, check it against the source. If the speaker simply misspoke, transcribe what was said and let the producer decide whether to add a correction note.

If you find that the same errors keep appearing, fix the cause upstream. Persistent problems with one speaker often trace back to a poor microphone or a noisy room. Our audio quality control checks include the recording issues that also degrade transcription.

Should the transcript be on the page or offered as a download?

On the page. A transcript published as HTML text alongside the audio player is accessible, readable on any device and indexable by search engines. A transcript hidden behind a download link, typically a PDF or a word-processor file, adds friction for every reader and often creates new accessibility problems of its own.

Consider what a download asks of the reader. They must notice the link, download a file, open it in another application, and then navigate a document that may not be tagged for screen readers or reflow on a phone screen. A screen reader user who could have read the text directly under the player now has an extra, error-prone task. A search engine may or may not index the file, and even if it does, the text is not associated with the episode page in the same way.

Good on-page implementations follow a few patterns:

  • Place the transcript directly below the player, or in a clearly labeled section on the same page. A collapsible "Show transcript" control is acceptable if it is a real, keyboard-accessible button and the text is in the page's HTML, not loaded only on click from somewhere else.
  • Use real text, not an image or embedded viewer. Text must be selectable, resizable and readable by assistive technology.
  • Keep timestamps light. A timestamp at each section heading or every few minutes helps readers jump to the audio; one on every line clutters the text.
  • Offer a download as an extra, not the only option. Some readers want a copy to print or annotate, so a download can sit alongside the on-page version.

Podcast platforms increasingly display transcripts inside their own apps, and Apple Podcasts for Creators documents how transcripts appear there and how creators can supply their own files. That is a welcome development, but it complements rather than replaces a corrected transcript on your own site, where you control accuracy, structure and links, and where the search value accrues to your domain.

  • Transcript is published as HTML text on the same page as the audio.
  • Every speaker is labeled by name, with a new paragraph at each change.
  • All names, technical terms and figures have been checked against the audio and source material.
  • Section headings break the transcript into readable parts.
  • Meaningful non-speech sounds are noted in brackets.
  • Any unresolved passage is marked [inaudible] with a timestamp, not silently omitted.
  • Any expand or collapse control works with a keyboard and a screen reader.

Is a transcript a substitute for captions on video?

No. For video with sound, captions are the accessibility requirement, and a transcript is a helpful addition, not a replacement. The two serve different needs, and treating one as the other is one of the most common mistakes in accessible media.

Captions are synchronized with the picture. A viewer who is deaf or hard of hearing needs to read what is said at the moment it is said, while watching the speaker's face, the slide on screen or the product being demonstrated. A transcript on the page below the video forces them to choose between reading and watching, and to keep their place in two places at once. For a talking-head video that might be tolerable; for a software demonstration, a training module or anything where the visuals and the speech depend on each other, it breaks the content.

Transcripts and captions are, however, closely related in production. A corrected transcript with timestamps is the ideal source for a caption file, and a finished caption file can be turned back into a readable transcript with some restructuring. The work is shared, but the deliverables are different:

Caption file
Timed text, broken into short lines sized for on-screen reading, synchronized to the video, with speaker identification and sound descriptions where needed. Usually delivered as a sidecar file for the player.
Transcript
Continuous, readable prose with speaker labels and headings, published on the page. Optimized for reading and searching rather than for timing.

For video that also needs descriptions of important visual information for blind and low-vision viewers, a third element, audio description or a descriptive transcript, comes into play. That goes beyond the scope of this article, but it is worth knowing that WCAG treats sound and visuals as separate accessibility problems, each with its own solution.

Should you use automatic transcription, human transcription or a hybrid?

For most organizations, a hybrid works best: an automatic first draft followed by a full human edit and a short subject-matter check. Pure automatic output is fast but not publishable, and fully manual transcription from scratch is accurate but slow for the kind of content most teams publish.

The right balance depends on volume, audio quality, subject matter and how much is riding on accuracy. The trade-offs look like this:

Pros

  • Automatic first drafts are produced in minutes, which makes transcribing every episode realistic even on a tight schedule.
  • Editing a draft is generally faster than typing from scratch, especially for clear, single-language recordings.
  • Timestamps and a rough speaker split come for free, which speeds up caption creation and navigation.
  • Custom vocabularies in many tools reduce repeated errors on names and terms across a series.
  • A human editing pass on top catches the errors that matter while keeping the cost of each episode predictable.

Cons

  • Automatic accuracy drops with heavy accents, specialist vocabulary, crosstalk and poor recordings, sometimes to the point where editing takes as long as typing.
  • Drafts that look fluent encourage superficial checking, so errors in names and figures survive.
  • Diarization errors can misattribute whole passages and are tedious to fix.
  • Uploading confidential recordings to a cloud service raises data-protection questions that need an answer before use.
  • Fully manual transcription remains the safer route for legal, medical and research material where every word is evidence.

The data-protection point deserves attention. Internal meetings, client calls, HR interviews and unreleased product discussions should not be sent to a transcription service until someone has checked its terms, where the data is processed and stored, and whether recordings are used to train models. The same care applies to any consent and release paperwork attached to the recording; our guide to audio rights and clearance records covers what to keep and why.

A good rule of thumb: if a quick test on a representative five-minute clip produces a draft that needs correction in most sentences, the audio or the subject matter is too difficult for an automatic-first workflow, and either the recording needs improving or a skilled human transcriber should work from scratch.

How should a transcript be structured so people can actually read it?

Structure it like an article: headings at each topic change, short paragraphs, speaker labels in bold, and a brief note at the top saying what the recording is. A transcript that is one unbroken wall of text is technically accessible but practically unreadable, and it wastes most of its search value.

Structuring a transcript well is a small editorial job with a large payoff. Most conversations already have a shape: an introduction, a series of topics and a close. Marking that shape with headings lets readers skim to the section they care about, lets screen reader users jump between headings, and gives search engines clear signals about what each part of the page covers.

  1. Add a short header block Title of the episode, date, the names and roles of the speakers, and the running time. One or two lines is enough.
  2. Identify topic changes While editing, mark where the conversation moves to a new subject. For a 40 to 60 minute interview, somewhere between five and ten sections is usually natural.
  3. Write descriptive headings Use headings that say what is discussed ("Why the pilot project stalled"), not "Part 3". Keep them factual, not promotional.
  4. Break long turns into paragraphs A single speaker talking for three minutes still needs paragraph breaks at natural pauses or shifts in thought.
  5. Add optional timestamps at each heading These let readers jump to the audio and help editors locate passages later.
  6. Link what speakers mention Where a guest mentions a resource you are allowed to link to, a link in the transcript is more useful than a vague reference, provided it does not alter what was said.
  7. Read the finished page top to bottom Check it on a phone screen, where long unbroken paragraphs are hardest to follow.

The same structure makes transcripts far more useful as raw material. A well-structured transcript can be turned into show notes, a blog post, social quotes, a newsletter summary or a training handout with much less effort than an unstructured one. Many teams find that the editing time they invest in structure is repaid by what they can reuse. Transcripts also belong in the project archive alongside the session files, as described in our practical guide to archiving audio sessions, because they are the fastest way to find content in a back catalog years later.

How do transcripts help people and search engines find audio?

They give search engines text to index where audio alone offers essentially none, which makes a transcript the only dependable way for spoken content to become searchable at all. Every question answered, every name mentioned and every topic discussed in an episode becomes something a page can be found for.

Search engines are built around text. Google Search Central's documentation is consistently clear that pages should present their important content in a form the crawler can read, and an audio file with a title and a two-sentence description gives the crawler very little to work with. A 45-minute interview may contain thousands of words of expert discussion, but without a transcript, almost none of it is visible to search.

A corrected, structured transcript on the page changes that in several ways:

  • Long-tail relevance. Conversations naturally cover specific questions and phrasings that a scripted article might not. Those become discoverable.
  • Accurate entities. Correctly spelled names of people, products and organizations help the page appear for searches on those names. Misspellings from raw automatic output do the opposite.
  • Clear topical structure. Headings in the transcript tell search engines what each section covers, in the same way they help human readers.
  • Internal search. Your own site search can find content inside episodes, which is invaluable once a back catalog grows beyond a few dozen recordings.

Two cautions keep this honest. First, a transcript's search value depends on its quality; a garbled raw transcript is thin, confusing content, not an asset. Second, publishing transcripts does not guarantee rankings. It makes the content eligible to be found, which is a precondition, not a promise. The accessibility case stands on its own regardless of search results.

What does a realistic transcription workflow look like for one episode?

A realistic workflow runs from a clean recording to an automatic draft, a full human edit, a short subject-matter check and on-page publication, and for a typical interview episode the editing is the part that takes the time. The worked example below is illustrative: it uses planning assumptions, not measured benchmarks, to show how the pieces fit together and where time goes.

Illustrative example: a 45-minute two-person interview

Imagine a B2B company publishing a fortnightly podcast. Each episode is a 45-minute conversation between a host and one guest, recorded remotely on separate tracks, covering a fairly technical subject with product names and figures. The team wants a corrected, structured transcript on each episode page on launch day.

A planning budget for one episode might look like this. The ranges are assumptions to adapt, and actual times depend heavily on audio quality, speaking speed, accents and how technical the conversation is.

StageWhoPlanning estimateWhat drives the range
Prepare glossary of names and termsProducer10 to 20 minutesNew guest or recurring topic; quality of the guest's bio
Automatic draft with custom vocabularySoftwareA few minutes of processingTool and file length
Full listen-through edit against audioTranscript editorAbout 1.5 to 3 times the running time, so roughly 70 to 135 minutesDraft accuracy, crosstalk, jargon density, editor experience
Speaker labels and structure with headingsTranscript editor15 to 30 minutesNumber of topic changes; diarization quality
Subject-matter review of flagged itemsHost or guest's team10 to 20 minutesNumber of flags; responsiveness
Publish on page and final read on mobileWeb editor15 to 20 minutesCMS and page template

On these assumptions, one episode needs somewhere around two and a half to four hours of people's time across several roles, with most of it in the editing stage. Over a fortnightly schedule, that is roughly a day of work a month. That figure is the one to plan and budget around, rather than the few minutes the automatic draft takes.

The example also shows where time can be saved without cutting corners. A better recording shortens the edit. A reused glossary shortens preparation for recurring guests and topics. A flagged question list shortens the subject-matter review. And a transcript template in the CMS shortens publication. None of these touches the accuracy bar.

  • Biggest time cost The full human edit against the audio, not the automatic draft
  • Biggest accuracy lever A clean, multitrack recording plus a glossary
  • Most damaging errors Names, technical terms and figures
  • Required for publication Speaker labels and a complete, faithful text
  • Best location HTML on the episode page, not a download
  • Video content Needs captions as well as, not instead of, a transcript

When should you bring in outside help with transcription?

Bring in help when a back catalog needs transcribing, when accuracy requirements are contractual, or when transcripts are needed in several languages. In each case, the work stops being a steady, manageable weekly task and becomes either a large one-off project or a job with consequences that a busy in-house team should not absorb informally.

A back catalog

Many organizations start transcribing new episodes and then realize that dozens or hundreds of older recordings have no transcript. Working through a back catalog in spare hours tends to stall. A structured project, with a shared style sheet, a combined glossary built from the whole series, batch processing and a consistent editing pass, finishes the job and produces transcripts that match each other. It is also a good moment to check that the original audio files are properly stored and documented, as covered in audio archiving and preservation.

Contractual or regulated accuracy

When a transcript is part of a contract, a compliance obligation, a research data set or a formal record, it needs a defined accuracy standard, full verbatim where appropriate, documented sign-off and version control. Those are process disciplines, and they are easier to guarantee with a team whose job is to follow them.

Several languages

Translating a transcript is a separate skill from transcribing one, and each language needs its own corrected source, its own glossary and a reviewer who knows the subject in that language. Translating an uncorrected transcript multiplies its errors. The right order is always to correct the original first, then translate, then review the translation.

Outside help is also worth considering when the underlying audio is the problem. If recordings are noisy, echoey or uneven in level, no amount of transcript editing fully compensates. Cleaning up dialogue before transcription through audio sweetening can make both the listening experience and the transcript noticeably better. Our wider audio editing and production service covers recording, editing and delivery as one workflow, so the transcript is planned from the start rather than bolted on at the end. Spoken content in public spaces raises similar questions; see how to get museum audio guides right for how text alternatives work in a visitor setting.

One point applies whoever does the work: the organization publishing the transcript owns its accuracy. Agree in writing on the transcription style, turnaround time, how uncertain passages are flagged, who signs off, and how confidential recordings are handled, before the first file changes hands.

Whatever route you choose, the standard stays the same: publish a corrected, structured transcript on the page with every piece of audio you release, label every speaker, check every name, term and figure, and add captions to anything with video. That is the standard that meets the accessibility requirement, serves readers and gives search engines something real to index.

The decisions that matter are not about which software to buy. They are about process: recording cleanly enough that automatic drafts are worth editing, choosing a transcription style and sticking to it, separating mechanical editing from subject-matter review, putting the transcript where people can use it, and knowing when a job has outgrown the in-house team. Get those right and a transcript stops being a compliance chore and becomes one of the most useful assets each recording produces.

Verdict Use automatic transcription for the first draft, never for the final one. A full human edit, clear speaker labels, careful checks on names, terms and figures, and on-page publication are what turn a machine draft into an accessible, searchable transcript. Treat captions for video as a separate requirement, and bring in help for back catalogs, contractual accuracy or multiple languages.

Where this comes from

The figures and practices above come from the sources listed.

Working on something like this?

We take on Audio Editing & Production work for teams who want it done once, properly. Tell us what you are building and we will tell you honestly whether we are the right studio for it. Start a project.

Where to go next

Spotted something wrong? Report an error on this page. We correct on the page and say what changed.

Frequently asked questions

It depends on your country, sector and any contracts or policies you have signed. The Web Content Accessibility Guidelines call for a transcript of prerecorded audio-only content at Level A, and many public bodies and regulated organizations adopt WCAG directly. Even where no rule applies, WCAG is the clearest benchmark for accessible audio.
You should not. Automatic tools often mishear names, specialist terms and numbers, and those errors are the ones readers quote and search for. Treat the automatic output as a draft and have someone edit it against the audio before it goes live.
Clean verbatim removes filler words, false starts and repetitions while keeping the speaker's meaning and wording, which suits podcasts and marketing content. Full verbatim keeps every hesitation and non-speech sound, which suits research, legal and compliance records. Pick one style per project and write it down.
A PDF on its own adds friction and may not work well with screen readers or on phones. Publish the transcript as text on the same page as the audio, and offer a download only as an optional extra for people who want a copy.
Video with sound needs synchronized captions, because viewers must read the words while watching the picture. A transcript is a helpful addition for reading and search, but it does not replace captions.
It varies with audio quality, accents, crosstalk and how technical the subject is. As a planning assumption, a full listen-through edit of an automatic draft often takes somewhere between one and a half and three times the running time, plus time for structure and review.
Transcripts give search engines text to index, which audio alone does not provide, so the topics and names discussed in an episode become findable. They make content eligible to be found but do not guarantee rankings, and a garbled raw transcript adds little value.
All services

The work behind this article, and what it costs.

Tomas Lindqvist

Picture and sound. Writes about editing, color, loudness and delivery specifications, including the ones that get deliveries rejected.

Keep reading

More in Audio Editing & Production