Auto-captions get blamed for wrong words. The failure nobody audits is the clock. Machine timing regularly pushes cues past what humans comfortably read — broadcast guidance puts comfort around 17 characters per second, and the auto-timed cues of a fast talker run 25–35. Your Deaf and hard-of-hearing viewers aren’t getting your words late; they’re getting them at a pace nobody can read.
Paste the SRT into the bench below. It measures every cue against a reading-speed limit you choose, marks the over-speed frames like an editor striking bad film, splices the timing where timing can fix it — and tells you honestly which cues only shorter text can fix, with the exact character budget a free chatbot needs to hit.
cues over the limit: — of — · worst — CPS at — · average — CPS
Set the limit and drift, then audit.
frame width = cue duration · diagonal strike = over the limit · dashed outline = timing can’t cure it, condense the text · white joint = a splice
select a frame to inspect its cue
Condensation layer — shrink the cues timing can’t cure (optional, free chatbot)
Contiguous fast speech can’t be fixed by timing — professional captioners edit the text down. Copy the flagged cues (with the character budget each one must fit), run the prompt from the prompt card in any free chatbot (Gemini, ChatGPT, Claude — no card), then paste the JSON back here. The bench re-audits the result: condensed text that still misses budget gets flagged, and any cue you cured earlier that grows past the limit again gets caught.
The free workflow
Export your auto-captions
In YouTube Studio (free, no card): Subtitles → your video → download the auto-captions. Any SRT or VTT from any tool works identically — the audit is source-agnostic.
Audit the reading speed
Every cue gets its chars-per-second measured against the limit you set — 13 for children’s content, ~15 for lectures, 17 for the broadcast comfort default, 20 as a hard ceiling. The worst cue auto-loads into the inspector so you can see exactly what a 30-CPS cue looks like.
Splice the timing
Over-speed cues pull their out-points into available gaps and may drift up to the lag you allow, re-syncing at natural pauses; cues that would sit on screen for 6½ seconds or more are cut into readable frames — the splice. The reel re-audits itself after every cut. What timing genuinely can’t cure, it says so.
Condense the remainder
Copy the flagged cues with their computed character budgets, run the strict-JSON prompt in a free chatbot, merge back, and watch the over-limit count fall. Your captions leave this tab only at this step, and only if you choose it.
What this replaces
| the old way | cost | the bench |
|---|---|---|
| Human caption editing and timing | ~$1–2 per video minute | $0, local |
| Pro caption editor seats | $10–30/mo | free chatbot + this tab |
| Manual re-timing of over-speed cues | 10–15 min per video | seconds, bounded by your drift setting |
| Unreadable captions left as-is | the audience they exist for | a measured, fixable reel |
The condensation prompt
You are a broadcast caption editor. I will paste auto-caption cues that exceed the reading-speed budget, each with its character budget in parentheses after the cue number. Return ONLY a valid JSON object — no prose, no markdown fence — matching exactly:
{"cues":[{"i":0,"text":"rewritten cue, within budget"}],"note":"one line, 20 words max, on what you cut"}
Rules: keep the meaning; drop fillers (uh, um, you know) and compress wordy phrasing; sentence case with punctuation; each text must fit its stated budget or be shorter; never merge, split, or reorder cues; "i" is the cue number shown before the budget.
Cues:
<PASTE THE FLAGGED CUES FROM THE BENCH HERE>
gemini · chatgpt · claude — free web tiers, no card · terms checked 2026-09-09 from training data (cold run — spot-check before republishing)
What breaks
- Contiguous fast speech can’t be cured by timing.If cues tile the timeline with no gaps, no re-timing math creates reading time — that’s what the condensation step is for, and the bench says so instead of pretending.
- Punctuation makes timing worse.Cleanup adds characters — a cured cue can drop back over the limit. The bench re-audits after every merge and names the ones that regressed.
- Overlaps and stacked speakers.One cue carrying two speakers needs a manual split first; the budget math assumes one voice per cue.
- Malformed pastes.No timing lines in the paste gets a named parse error and keeps the last reel — never a blank.
- Long files.Over 120 cues, the audit covers the prefix; the banner says so.
Keep the bench: the Caption Reel Audit app
The article audits one file and lets it go. The app keeps working:
- .srt and .vtt file exports — the re-timed reel downloads with timecodes intact (hand-copying text mangles them); ~10–15 minutes of manual re-timing per video saved.
- Channel audit history — every video’s counts, one glance; spot the episodes that regressed instead of re-pasting files.
- Printable compliance sheet — stats, flagged cues, and settings on one page for clients, teammates, or your future self.
$5 once · the free tier stays complete · no account
Open the Caption Reel Audit app