WCAG 1.2.5 Audio Description: The AA Requirement Teams Keep Skipping

Jamie
digitalwcagvideo accessibilityaudio descriptiontitle iisection 508

Jamie · AI Research Engine

Analytical lens: Strategic Alignment

Small business, Title III, retail/hospitality

AI-assisted · Source-linked · Editorially reviewed · Methodology

Trust note

This article was drafted with AI assistance, reviewed against accessibility.chat editorial standards, and should be treated as research and education rather than legal advice. We prioritize primary sources and correct material errors.

Close-up of a woman's hands reading a braille book while enjoying coffee, emphasizing sensory engagement.
Photo by Thirdman on Pexels

A training video plays. The presenter clicks through slides, gestures at a whiteboard, and demonstrates software on screen. The audio track says "as you can see here" — and a blind viewer hears nothing useful. No description of the chart. No narration of the diagram. No indication that a critical workflow just appeared on screen. The video is polished, professionally produced, and completely inaccessible to anyone who cannot see it.

This is the gap WCAG 1.2.5 Audio Description (Prerecorded) (opens in new window) exists to close. It's a Level AA requirement — the conformance level most organizations target — and it's one of the most consistently skipped criteria in video production workflows. Not because teams don't care, but because audio description is expensive, time-consuming, and easy to defer when a launch deadline is approaching.

The Requirement

Success Criterion 1.2.5 (opens in new window) requires that all prerecorded synchronized media — video with audio — provide an audio description of the video content. Audio description is a synchronized spoken narration added during natural pauses in dialogue that describes what's happening visually: actions, characters, scene changes, and on-screen text that aren't conveyed by the main audio track.

The key phrase is "not conveyed by the main audio track." If a speaker fully narrates everything visible — reading slide text aloud, describing the diagram, explaining what's on screen — no additional audio description is required. The problem is that most video production doesn't work this way. Presenters assume viewers can see. Editors cut for visual flow, not audio completeness. The result is a gap between what's audible and what's visible.

Conformance LevelCriterionRequirementFlexibility
Level A1.2.3Audio description OR full text alternativeAuthor's choice
Level AA1.2.5Audio description requiredNo text-only substitution
Level AAA1.2.7Extended audio descriptionPauses video if needed
Level AAA1.2.8Full text alternativeAdditional requirement

This layered structure matters strategically. At Level A, teams can satisfy 1.2.3 with a text transcript. At Level AA, that option closes — 1.2.5 requires actual audio description. Organizations that met 1.2.3 with a transcript and assumed they were done have a gap they may not know about.

Why This Matters

Blind users and users with low vision are the primary audience for audio description, but the exclusion extends further. People with certain cognitive disabilities benefit from audio description that reinforces visual information verbally. Users watching in contexts where they can't see the screen — eyes occupied, screen obscured — gain meaningful access.

The exclusion is concrete. Without audio description, a blind user watching a software tutorial hears button clicks and verbal instructions but misses where the cursor moves, which menu opens, and what the interface looks like. A training video about safety procedures describes the hazard verbally but shows the correct protective equipment on screen without narrating it. A product demo explains features while visually demonstrating them — and a user relying on audio gets half the information.

This isn't a minor inconvenience. For employment training, educational content, and public-facing video, the absence of audio description means disabled users receive materially less information than sighted users. That's not a UX problem — it's an equal access problem.

How to Meet It

The W3C documents several sufficient techniques:

G78 (opens in new window) — Provide a second, user-selectable audio track that includes audio descriptions. This is the standard broadcast-style approach: the video plays normally, and users can switch to an audio description track.

G173 (opens in new window) — Provide a version of the video with audio descriptions. A separate video file with description narration baked into the audio. Simpler to implement than a selectable track; requires maintaining two video files.

G8 (opens in new window) — Provide a movie with extended audio descriptions. Used when natural pauses aren't sufficient — the video pauses to allow longer descriptions, then resumes. This satisfies 1.2.5 but is more commonly associated with the AAA criterion 1.2.7.

G226 (opens in new window) — Provide the same information in text. This technique applies specifically when the video track presents information redundantly already available in the page text — a narrow exception.

G203 (opens in new window) — Use a static image alternative for video. Applicable only when the video is essentially a static image with audio — not applicable to most real-world video content.

The most common failure is documented as F113 (opens in new window): providing audio description that does not describe important visual content in the video. This is the partial-description failure — a description track exists but omits critical visual information. Having a description doesn't mean having an adequate one.

The advisory technique H96 (opens in new window) covers using the <track> element for descriptions, which is worth implementing for HTML5 video even though it doesn't satisfy the criterion on its own.

Applying This

Audio description fails at the production planning stage more often than the implementation stage. By the time a video is edited and delivered, retrofitting description is expensive — dialogue pauses may not exist in the right places, the production team has moved on, and the budget is spent.

The fix is upstream:

In scripting: Require scripts to be self-describing. If a presenter will reference something visual, the script should narrate it. "The chart shows a 40% increase from Q1 to Q3" instead of "as you can see here."

In production: Build natural pauses into the edit — 2-3 second gaps after visual transitions where description can be inserted. This costs nothing in post-production if it's planned during editing.

In QA: Audio description is nearly impossible to catch with automated testing tools. As our research on automated testing limitations documents, automated tools detect the presence of a description track but cannot evaluate whether it adequately describes the visual content. Manual review by someone who cannot see the video — listening only to the audio — is the only reliable test.

In procurement: Video production contracts should specify audio description as a deliverable, not an optional add-on. Organizations that outsource video production and don't contractually require description will consistently receive inaccessible content.

For existing video libraries, triage by use. High-traffic training videos, public-facing marketing content, and anything used in employment or education contexts should be prioritized. Description services exist — vendors like 3Play Media and Verbit offer audio description production — though quality varies and human review of the final product remains essential.

Teams navigating multiple compliance frameworks should note that Section 508 incorporates WCAG 2.0 Level AA by reference, which means 1.2.5 applies to federal agencies and their contractors. The compliance framework complexity around overlapping standards can create confusion about which requirement governs — but for audio description, the obligation is consistent across WCAG 2.x, Section 508, and EN 301 549.

CORS Perspective

From a strategic alignment lens, audio description is the accessibility requirement most likely to be eliminated in a budget conversation — it's visible, costly, and easy to defer. The business case requires reframing: audio description isn't a post-production expense, it's a production quality standard. Organizations that build description into scripting and editing workflows reduce per-video costs significantly compared to retrofitting. For Title II entities — government agencies, public universities — the legal obligation under ADA (opens in new window) and Section 508 removes the optionality that budget pressure implies. For Title III businesses, the risk calculus is straightforward: video is the dominant content format, and inaccessible video at scale represents a systemic barrier that plaintiffs' attorneys and DOJ investigators increasingly target. The operational question isn't whether to provide audio description — it's whether to build the production workflow that makes it sustainable.

About the Jamie lens

A strategy lens for small business and Title III. Frames findings around cost, sequencing, and what a retail or hospitality operator can realistically act on first.

Jamie is an AI analyst lens, not a human staff member. It helps frame this article through a consistent accessibility perspective.

Specialization: Small business, Title III, retail/hospitality

View all articles using this lens →

Transparency Disclosure

This article was drafted with AI assistance and reviewed against our editorial methodology. We disclose that process so readers can judge the work clearly.