Two things do most of the work. Publish a full text transcript wherever the show lives, and put accurate captions on every video clip you post. Those two cover the large majority of people who cannot use your audio, and you can do both in an afternoon.
The part almost every accessibility guide misses is that the surface has moved. Most of them were written for an RSS show with a website, so they stop at "add a transcript." Podcast consumption is now majority video, which means your accessibility work lives inside a vertical clip in a feed, not on a page. That changes what you build, how fast you are allowed to talk, and what it costs.
TL;DR:
- A transcript is a text document of the whole show. Captions are time-synced text on the video. Publishing one does not cover the other.
- WCAG 2.2 requires captions for prerecorded audio in synchronized media at Level A, the lowest conformance bar there is.
- Caption legibility is capped by your script pace. The DCMP Captioning Key tops out at 160 words per minute for adult material, so a 60-second clip should carry roughly 150 words, not 200.
- US state and local government deadlines moved in April 2026. The current dates are April 26, 2027 and April 26, 2028, not the 2026 date most guides still print.
- The accessibility bill scales with runtime, so short video clips are the cheapest format to make accessible: about $20 a month at human-caption list prices, against roughly $300 for a weekly 40-minute show.
What does podcast accessibility actually mean?
Podcast accessibility means the same content is available in a second modality: text for people who cannot hear it, and structure and description for people who cannot see it. That is the whole idea. Everything below is a way of delivering one of those two things without wrecking the show.
The audience is not small. The World Health Organization reports that "over 5% of the world's population, or 430 million people, require rehabilitation to address their disabling hearing loss," and projects that "by 2050, nearly 2.5 billion people are projected to have some degree of hearing loss." Add everyone watching with the sound off in a waiting room and the practical audience for captions is far larger than the clinical number.
The standard that regulators point at is WCAG. Success Criterion 1.2.2 states that "Captions are provided for all prerecorded audio content in synchronized media, except when the media is a media alternative for text and is clearly labeled as such." That criterion sits at Level A, the minimum conformance level. Captions on a video clip are not an advanced accessibility feature. They are the floor.
Do podcasts legally need transcripts and captions?
For most independent creators, no law names your podcast directly. For public bodies and for a lot of businesses selling into the EU, the requirements are real and dated. This is not legal advice, but the two rules people actually ask about are worth getting right, because the most-cited dates on the web are now wrong.
In the US, the Department of Justice's 2024 web rule sets "WCAG, the Web Content Accessibility Guidelines, Version 2.1, Level AA" as the technical standard for state and local government web content and mobile apps. The compliance dates were extended this year: "On April 20, 2026, the Federal Register published the Department's Interim Final Rule (IFR) extending the compliance date for State and local government entities with a total population of 50,000 or more to April 26, 2027. The compliance date for public entities with a total population of less than 50,000, or any special district government, is extended to April 26, 2028."
That matters because nearly every podcast accessibility article published this year still prints April 24, 2026 as the big deadline. If you run a university show, a city podcast, or anything hosted by a public entity, check the current ada.gov page rather than a vendor blog.
In the EU, the European Accessibility Act is a directive that "aims to improve the functioning of the internal market for accessible products and services." Member states had to write it into national law by June 2022, and its scope includes "e-commerce" and "access to audio-visual media services such as television broadcast and related consumer equipment." If your show is a marketing channel attached to an EU-facing product, the product is in scope even when the podcast itself is not.
What is the difference between a transcript and captions?
They solve different problems and are not interchangeable. A transcript is a static text document a reader can search, skim, translate, or feed to a screen reader at their own pace. Captions are time-synced text displayed over the video, so a viewer follows the moment as it happens. Publishing a transcript on your site does nothing for someone watching a clip in a feed.
| Transcript | Captions | |
|---|---|---|
| Where it lives | Episode page, show notes, description link | On the video itself |
| Timed to the audio | No | Yes |
| Helps a feed viewer | No | Yes |
| Searchable and indexable | Yes | Only if the platform exposes them |
| Needs speaker labels and sound cues | Yes | Yes |
| Rough effort per 30-minute show | 20 to 40 minutes of cleanup | Handled per clip |
Do both. The transcript is where search engines and AI answer engines read your show, which is the same reason it is worth writing proper podcast show notes with AI instead of pasting a raw machine transcript and calling it done.
How fast should podcast captions be?
Slower than you think, and it is a scripting decision rather than a captioning one. The DCMP Captioning Key defines presentation rate as "the number of captioned words per minute (wpm) that are displayed onscreen," and caps it at "not to exceed 130 words per minute" for lower-level material, "not to exceed 140 wpm" for middle-level, and "not to exceed 160 wpm" for upper-level material.
Short-form scripts routinely blow through that. A punchy hook written to fill 60 seconds often runs 200 words or more, which is 200 wpm on screen. No captioning tool fixes that. Once the words exist, the captions either flash past unreadably or spill into blocks nobody finishes reading.
This is the one place our own numbers are directly useful. MakePodcast paces generated clips at 2.5 words per second, which is 150 words per minute, and the length presets are built on it: 75 words for 30 seconds, 150 for 60 seconds, 225 for the 90-second maximum. That was set for natural delivery, not for accessibility, but it lands inside the DCMP band for adult material by construction.
| Clip length | Caption-legible word budget | Resulting rate |
|---|---|---|
| 30 seconds | 75 words | 150 wpm |
| 60 seconds | 150 words | 150 wpm |
| 90 seconds | 225 words | 150 wpm |
| 60 seconds (typical unedited hook) | 200+ words | 200+ wpm, above every published guideline |
Accessibility is a script decision before it is a captioning decision. If the words are too fast, the captions are already broken.
If you are writing to a budget like this for the first time, the mechanics are covered in how to write a podcast script with AI. Cut adjectives, not ideas.
Should captions be burned into the video or added by the platform?
Use both, and know what each one gives up. Burned-in captions are pixels: they always display, they survive re-uploads and downloads, and you control the styling. They also cannot be turned off, cannot be translated, cannot be read by assistive technology, and get clipped by platform safe zones over the caption area.
Platform captions are a text track. They can be toggled, translated, indexed, and read by a screen reader. They are also editable after the fact on some platforms and not on others, and on some surfaces they simply do not appear.
| Burned-in captions | Platform caption track | |
|---|---|---|
| Always visible | Yes | Depends on viewer settings |
| Can be turned off | No | Yes |
| Machine-readable and translatable | No | Yes |
| Survives download and re-upload | Yes | No |
| Risk of being cropped by UI overlays | Yes | No |
The practical rule: burn in the captions so the clip works everywhere, and still upload or accept the platform's caption track so the text exists as text. It costs one extra step per post.
Are auto-generated captions good enough?
They are a first draft, and the platforms say so themselves. YouTube's own help documentation warns that "Automatic captions might misrepresent the spoken content due to mispronunciations, accents, dialects, or background noise" and instructs creators: "You should always review automatic captions and edit any parts that haven't been properly transcribed."
YouTube's automatic captions cover a long list of languages, including Arabic, Hebrew, Hindi, Japanese, Portuguese, Spanish and dozens more, which makes them a genuinely useful starting point. But a name, a product term, or an accent is exactly where they fail, and those are exactly the words carrying your meaning. If you host in a second language, this compounds, and the failure modes are the same ones covered in making a multilingual podcast with AI.
Budget 60 seconds per clip to read the caption track and fix proper nouns. That is the whole job at this length.
How much does it cost to make a podcast accessible?
Less than the guides imply, and dramatically less if you publish short. The bill scales with runtime, so format is the biggest cost lever you have.
Rev lists human transcription at "$1.99 /min." and human captions "Starting at $1.99 /min.", both "99%+ accurate, delivered in 12 hours or less", with a free tier that includes "45 AI transcription & caption minutes/month". Run those rates against two real publishing patterns:
| Publishing pattern | Finished minutes per month | Human captions at $1.99/min |
|---|---|---|
| Three 45-second clips a week | about 9 | under $20 |
| Weekly 40-minute audio episode | about 160 | about $320 |
That is roughly a 16x difference for the same accessibility standard. A long-form show has to lean on AI captions plus its own review time, because human captioning at list price costs more per month than most independent shows earn. A short-form video channel can simply buy the accurate version.
Do not buy human captioning before you fix the script pace. Paying $1.99 a minute for a verbatim transcription of 200 words crammed into 60 seconds gets you accurate captions that are still unreadable. Rewrite to the word budget first, then decide whether AI captions plus a one-minute review are enough. For most creators at this length, they are.
Honest disclosure: MakePodcast generates the clip, not the captions. We do not ship a caption editor, and we do not add a caption track for you, so you will still do that step in your posting tool or your platform of choice. What we do contribute is the part that is hardest to fix later, which is a script paced at 150 words per minute inside a 90-second clip. If you need a caption studio with styling controls, buy one.
The 2026 podcast accessibility checklist
- Publish a cleaned transcript on the page where the show lives, with speaker labels and bracketed sound cues.
- Write every clip to a word budget: 75 words for 30 seconds, 150 for 60, 225 for 90.
- Burn captions into the video so they survive downloads and re-uploads.
- Also accept or upload the platform caption track, so the text exists as text.
- Review auto captions for names, product terms and numbers. One minute per clip.
- Keep captions out of the platform's UI safe zones, roughly the bottom fifth and the right edge.
- Describe anything on screen that is not spoken. If the punchline is a chart, say the number out loud.
- Write link text and show-note headings that mean something on their own, not "click here".
- Check the contrast of your caption styling against the busiest frame in the clip, not the calmest.
- If a public body or an EU-facing product owns the show, check the current ada.gov and national EAA rules directly rather than a vendor summary.
FAQ
Do podcasts legally need transcripts?
For an independent show, usually not by name. For state and local government content in the US the standard is WCAG 2.1 Level AA, with compliance dates now April 26, 2027 for populations of 50,000 or more and April 26, 2028 for smaller and special district entities.
Are captions and subtitles the same thing?
No. Captions carry the speech plus non-speech audio information such as sound effects and speaker changes, and are aimed at people who cannot hear the audio. Subtitles assume you can hear and translate the dialogue only.
Do short video clips need captions if the platform adds them automatically?
Yes, review them. Automatic captions are generated by speech recognition and, in YouTube's own words, "might misrepresent the spoken content due to mispronunciations, accents, dialects, or background noise." Accuracy on names and product terms is where they break.
Does a podcast need audio description?
Rarely, if the audio carries the whole meaning. It matters the moment something visual is load bearing: an on-screen chart, a caption-only joke, or a product shot. The fix is usually to say the information out loud rather than to record a separate description track.
Is there a free way to make my podcast accessible?
Mostly, yes. Platform auto captions, a lightly edited AI transcript, and a properly paced script cover the large majority of the work at zero cost. What you are spending is review time, not money. The wider free toolkit is covered in how to make a podcast with AI.
