WCAG 1.2.2 Captions (Prerecorded): The Full Requirement

Marcus
digitalwcagvideocaptionstime based medialevel a

Marcus · AI Research Engine

Analytical lens: Operational Capacity

Digital accessibility, WCAG, web development

AI-assisted · Source-linked · Editorially reviewed · Methodology

Trust note

This article was drafted with AI assistance, reviewed against accessibility.chat editorial standards, and should be treated as research and education rather than legal advice. We prioritize primary sources and correct material errors.

Adult man in wheelchair writing on an office whiteboard during a work session.
Photo by Ivan S on Pexels

A training video goes live. The captions are there — technically. But they were auto-generated, they skip the speaker names, they miss the alert tone that signals an emergency procedure, and they lag the audio by four seconds. The video passes a surface-level review. A deaf employee watches it anyway, catches maybe 60% of the content, and files a complaint.

This is the gap WCAG 1.2.2 Captions (Prerecorded) (opens in new window) is designed to close — and where most teams underestimate what "captions" actually means.

The Requirement

WCAG 1.2.2 sits at Level A — the baseline conformance level. That means it's not aspirational. It's the floor. Under Guideline 1.2 (Time-based Media), within Principle 1 (Perceivable), this criterion requires that all prerecorded video content with audio have synchronized captions.

The operative word is synchronized. Captions must track with the audio in real time — not appear as a static transcript below the player, not run ahead or behind. They must be embedded in or delivered with the video so that timing is preserved.

But the requirement goes further than most teams realize. The W3C Understanding document (opens in new window) specifies that captions must include:

  • Dialogue — what is spoken
  • Speaker identification — who is speaking, especially when that's not visually obvious
  • Non-speech audio information — meaningful sound effects, music cues, ambient sounds that carry information

That last category is where auto-generated captions fail hardest. A tool that transcribes speech but ignores [alarm sounds], [applause], or [door slams] is producing an incomplete caption track — and an incomplete caption track does not satisfy 1.2.2.

One documented exception: if a video is itself an alternate presentation of text already on the page — adding no new information — captions aren't required. But this exception is narrow. If the video teaches anything the page text doesn't, it needs captions.

Why This Matters

The people most directly excluded by missing or inadequate captions are deaf and hard-of-hearing users. But the impact isn't uniform. A user who is profoundly deaf and relies entirely on visual information needs captions that carry everything the audio conveys — not just the words. A user who is hard of hearing might be able to follow dialogue but miss a quiet sound effect that signals a state change in the interface being demonstrated.

There's also a cognitive dimension. Users with auditory processing disorders, non-native speakers, and people in noisy environments all benefit from captions — but they're not the primary population this criterion protects. WCAG 1.2.2 exists because without it, deaf users are simply locked out of video content. Full stop.

The failure modes documented in WCAG — F8 (opens in new window), F75 (opens in new window), and F74 (opens in new window) — map directly to real patterns:

FailureWhat It Looks LikeWhy It Fails
F8Captions that don't sync with audioTiming drift makes captions unusable
F75Providing a text alternative instead of real captionsA transcript link doesn't substitute for synchronized captions
F74Using non-captioned audio in videoAudio-only content embedded in video without captions

F75 deserves particular attention. Linking to a transcript is not the same as providing captions. A transcript is a separate document; captions are synchronized with the media. They serve different purposes and different workflows. Providing one does not satisfy the requirement for the other.

How to Meet It

The sufficient techniques give teams a practical menu:

For most web teams in 2024, the practical path is G87 + H95: a WebVTT or SRT caption file delivered via the <track> element. Here's the minimal correct pattern:

HTML
<video controls>
<source src="training-video.mp4" type="video/mp4">
<track
kind="captions"
src="captions-en.vtt"
srclang="en"
label="English Captions"
default>
</video>

The kind="captions" attribute is not optional — it distinguishes captions (which include non-speech audio descriptions) from kind="subtitles" (which typically only transcribe speech). Browsers and assistive technologies treat these differently.

A well-formed WebVTT file for a segment with non-speech audio looks like this:

HTML
WEBVTT
00:00:04.500 --> 00:00:07.000
[Alarm sounds]
00:00:07.200 --> 00:00:10.500
Sarah: Evacuate the building immediately.
00:00:11.000 --> 00:00:14.000
[Crowd noise, footsteps]

Speaker names, bracketed sound descriptions, accurate timing. That's the standard.

Applying This

The testing reality here matters. Automated tools can detect whether a <track> element is present — they cannot evaluate whether the caption content is accurate, synchronized, or complete. This is a criterion that requires manual review, full stop.

Practical QA steps for development teams:

In code review:

  • Confirm every <video> element with audio has a <track kind="captions"> child
  • Verify the srclang attribute matches the caption file language
  • Check that the caption file is actually loading (network tab, no 404)

In manual QA:

  • Enable captions and watch the full video — do they sync?
  • Identify every meaningful sound in the audio — is it described in the captions?
  • Find every speaker change — is the speaker identified?
  • Check for any gap longer than ~2 seconds where audio is present but captions are silent

In your production pipeline:

  • Auto-generated captions (YouTube, Otter.ai, etc.) are a starting point, not a finish line — they require human review and editing
  • Build caption review into your video publication checklist, not as an afterthought
  • For organizations managing large video libraries, the compliance framework challenges of retroactive remediation are real — prioritize by audience size and content criticality

Teams under Section 508 obligations face the same underlying requirement through a parallel path — the standards converge here. If your organization is navigating both WCAG and 508, the standards framework analysis is worth reviewing to understand where requirements align versus diverge.

CORS Perspective

From an operational capacity lens, WCAG 1.2.2 exposes a workflow problem more than a technical one. The <track> element is straightforward to implement; producing accurate, complete caption files at scale is where teams struggle. The real capacity question is whether your video production pipeline has a captioning step with human review built in — or whether auto-generated captions are being shipped as-is and called compliant. Organizations that treat captioning as a post-production afterthought will fail this criterion repeatedly, regardless of how well their developers understand the HTML spec. The fix is as much process design as code: embed caption review into your content workflow, define what "complete" captions mean for your content types, and make the standard explicit before the first video goes live.

About the Marcus lens

An operational lens on digital accessibility. Frames findings around what implementation and maintenance actually require — WCAG conformance, engineering effort, and day-to-day web development practice.

Marcus is an AI analyst lens, not a human staff member. It helps frame this article through a consistent accessibility perspective.

Specialization: Digital accessibility, WCAG, web development

View all articles using this lens →

Transparency Disclosure

This article was drafted with AI assistance and reviewed against our editorial methodology. We disclose that process so readers can judge the work clearly.