
Detect speech, not silence: building audio description for video that has none
Most video has no audio description because a human has to write it. I built a Fire TV app that writes it instead, using Amazon Transcribe and Polly, and two blind reviewers found the bugs my own testing never could.
Building audio description for the content that never got any
What I learned making Transcribe, Polly and a Fire TV write descriptions for
video nobody had ever described, and what two blind reviewers found that I
had missed.
video nobody had ever described, and what two blind reviewers found that I
had missed.
The seven seconds that explain the problem
Play any seven seconds of a film with your eyes shut. Someone stops walking.
The music changes. There is a small sound, and then quiet.
The music changes. There is a small sound, and then quiet.
You know something happened. You have no idea what.
That is the ordinary experience of watching video without audio description,
and audio description is not a new or exotic fix for it. A person watches the
film, writes down what is on screen, and a voice reads it out in the gaps
between the dialogue. It works. Blind and low-vision viewers rely on it. There
is a whole craft around doing it well.
and audio description is not a new or exotic fix for it. A person watches the
film, writes down what is on screen, and a voice reads it out in the gaps
between the dialogue. It works. Blind and low-vision viewers rely on it. There
is a whole craft around doing it well.
The catch is the economics. A human has to write it and someone has to pay
them, so it exists for a slice of what gets made and for everything else there
is nothing at all. Not a worse version. Nothing.
them, so it exists for a slice of what gets made and for everything else there
is nothing at all. Not a worse version. Nothing.
That is the pain point I wanted to attack: not "make description better", but
make it exist for the content that was never going to get any.
make it exist for the content that was never going to get any.
What I built
Sightline is a Fire TV app on Vega
OS that writes the description nobody was going to write, and speaks it while
you watch.
OS that writes the description nobody was going to write, and speaks it while
you watch.
It finds the gaps where no one is speaking, looks at what changed on screen
across each gap, writes only the part you would be lost without, measures the
line to check it fits the silence available, and speaks it there. One press
moves the description to a phone, so a blind viewer and a sighted viewer can
watch one screen and each get what they need. And you can ask it questions
mid-scene, which a pre-recorded description track structurally cannot answer,
because it was written before you had the question.
across each gap, writes only the part you would be lost without, measures the
line to check it fits the silence available, and speaks it there. One press
moves the description to a phone, so a blind viewer and a sighted viewer can
watch one screen and each get what they need. And you can ask it questions
mid-scene, which a pre-recorded description track structurally cannot answer,
because it was written before you had the question.
The AWS shape of it
1
2
3
4
5
6
7
8
9
video ──> Amazon S3 ──> Amazon Transcribe ──> word timings ──> gaps
│
frame pairs per gap ──> model ──> lines ──┤
│
Amazon Polly (generative) ──────┘
│
timeline + PCM ──> Fire TV app
│
questions ──> API Gateway ──> Lambda ──> answerBelow is what each of those actually taught me. I have tried to write the
notes I wanted to find before I started.
notes I wanted to find before I started.
Finding 1: detect speech, not silence. It was a 16x difference.
My first gap detector looked for silence, which is the obvious reading of
"description goes in the gaps."
"description goes in the gaps."
On a 52-second trailer, silence detection found 2.5 usable seconds. The
score never stops, so by the measure of "is it quiet", there is almost nowhere
to speak.
score never stops, so by the measure of "is it quiet", there is almost nowhere
to speak.
Detecting speech with Transcribe and treating everything else as available
found 39.8 seconds.
found 39.8 seconds.
That is a sixteenfold difference, and it is the single largest improvement in
the project. Music is not an obstacle to description. Dialogue is. Transcribe's
word-level timestamps are exactly the right primitive for this, and accuracy on
film dialogue over a score was excellent.
the project. Music is not an obstacle to description. Dialogue is. Transcribe's
word-level timestamps are exactly the right primitive for this, and accuracy on
film dialogue over a score was excellent.
A second benefit came free: feeding the transcript text to the model stops the
description repeating a line of dialogue that was just spoken.
description repeating a line of dialogue that was just spoken.
If you are timing anything against speech, ask for word timings and invert
them. Do not look for quiet.
them. Do not look for quiet.
Finding 2: Polly pads its output, and the speaking rate is not the rate you asked for
The generative engine sounds very good, and PCM output goes straight into the
Vega audio stream with no decoding on the device. Both were the right call.
Vega audio stream with no decoding on the device. Both were the right call.
Two things cost me a day.
Polly pads both ends of the audio with silence. If you are fitting speech
into a measured gap, that padding is charged against your budget. It has to be
trimmed before you measure anything.
into a measured gap, that padding is charged against your budget. It has to be
trimmed before you measure anything.
The effective speaking rate is not the rate you request. I assumed 170 wpm,
which is the rate I asked for, and every single line overflowed its gap. The
real figure, measured after trimming the padding, was 203 wpm.
which is the rate I asked for, and every single line overflowed its gap. The
real figure, measured after trimming the padding, was 203 wpm.
If you are doing timed speech, measure the rate of your specific voice and
engine before you build any budgeting logic on top of it. Do not trust the
number you passed in.
engine before you build any budgeting logic on top of it. Do not trust the
number you passed in.
Finding 3: S3 is fine, but it is most of the clock for short files
Transcribe takes input from a bucket rather than a direct upload, so every job
means writing an object and cleaning it up afterwards. The bucket itself was
unremarkable in the best possible way: private, one object per job, deleted
after, no surprises.
means writing an object and cleaning it up afterwards. The bucket itself was
unremarkable in the best possible way: private, one object per job, deleted
after, no surprises.
But for a 52-second clip, the round trip through S3 is most of the wall-clock
time of that stage. For batch work that is irrelevant. It mattered here because
users are waiting.
time of that stage. For batch work that is irrelevant. It mattered here because
users are waiting.
This also closed off a feature. I wanted spoken questions to work in-app on
iPhones, where Safari has no
recording audio and transcribing it server-side. Transcribe's batch path
through S3 is far slower than anyone will wait for an answer about the shot
they are looking at. Streaming transcription is the right tool and is on the
list; batch was never going to work for an interactive question.
iPhones, where Safari has no
SpeechRecognition at any origin, which meansrecording audio and transcribing it server-side. Transcribe's batch path
through S3 is far slower than anyone will wait for an answer about the shot
they are looking at. Streaming transcription is the right tool and is on the
list; batch was never going to work for an interactive question.
Finding 4: Bedrock refused at the account level, and the error said nothing useful
This is the one I most want to pass on, because I lost real time to it.
The Bedrock client is written and shipped in the repository. It never ran. Every
region, every model, bare model IDs and both inference-profile forms returned:
region, every model, bare model IDs and both inference-profile forms returned:
1
Error 002: Access to Bedrock models is not allowed for this accountMeanwhile S3, Polly and Transcribe all worked on the same credentials, in the
same session. The account is not in an Organization. And
same session. The account is not in an Organization. And
get-foundation-model-availability reported:1
2
3
4
authorizationStatus: AUTHORIZED
entitlementAvailability: AVAILABLE
regionAvailability: AVAILABLE
agreementAvailability: NOT_AVAILABLEThe last line is the actual problem, and the error message never mentions it.
Per-model toggles in the console do not clear it either.
Per-model toggles in the console do not clear it either.
"Access is not allowed for this account" gives a developer nothing to act on.
An error naming the missing agreement, or a console surface showing agreement
status next to the model, would turn a support ticket into a click. I shipped
against a direct model API instead, which is a worse outcome for everybody.
An error naming the missing agreement, or a console surface showing agreement
status next to the model, would turn a support ticket into a click. I shipped
against a direct model API instead, which is a worse outcome for everybody.
If you hit Error 002 with working credentials for other services, check
agreementAvailability before you spend an afternoon on IAM.What the blind reviewers changed, which is most of it
I did not want to guess at what would help, so I asked. Reviewers on the
ACB Audio Description Project mailing list answered
in detail, and almost none of the rules this thing follows are mine.
ACB Audio Description Project mailing list answered
in detail, and almost none of the rules this thing follows are mine.
Nothing during dialogue. Nothing that repeats what the sound already tells you.
Never "the camera pans", because the viewer is watching a story, not a shoot.
Never "the camera pans", because the viewer is watching a story, not a shoot.
One reviewer proposed a test I would not have thought of: attenuate the whole
audio file and check the classification does not move. It found a real bug. My
loader was reading 16-bit, this film decodes to samples above full scale, and
those were being clipped. Reading 32-bit float fixed it, and gain invariance
went to exactly 0.00 dB.
audio file and check the classification does not move. It found a real bug. My
loader was reading 16-bit, this film decodes to samples above full scale, and
those were being clipped. Reading 32-bit float fixed it, and gain invariance
went to exactly 0.00 dB.
A second tester, Dave Matters, used the player and found two faults in an
afternoon that all my local testing had missed. Description stopped a third of
the way through a clip, because a limit written to cap the number of
descriptions was truncating the film. And a scene where the software reported
someone watching a character sleep. Nobody was there. It had turned a camera
angle into a person.
afternoon that all my local testing had missed. Description stopped a third of
the way through a clip, because a limit written to cap the number of
descriptions was truncating the film. And a scene where the software reported
someone watching a character sleep. Nobody was there. It had turned a camera
angle into a person.
I did not catch either. He did, and he cannot see the screen.
The lesson I keep coming back to: "it looks fine" is not evidence
Three of the worst bugs in this project presented as working software.
A truncated film that played normally. A phone that tracked the television's
playhead perfectly while speaking nothing at all, because the local service
answered 501 to every HEAD request and the phone reads "not ok" as "the file
does not exist". And an app that went completely silent in the one state where
a blind user has nothing else to go on: the spoken offer to describe an
undescribed video was never audible, because that code path returned before the
audio stream was initialised.
playhead perfectly while speaking nothing at all, because the local service
answered 501 to every HEAD request and the phone reads "not ok" as "the file
does not exist". And an app that went completely silent in the one state where
a blind user has nothing else to go on: the spoken offer to describe an
undescribed video was never audible, because that code path returned before the
audio stream was initialised.
None of those threw an error. Two of them looked better than the working
version, because they were quieter.
version, because they were quieter.
If you are building accessibility features, the thing you have to test is the
thing you cannot see. Measure the audio. Count the utterances. Then give it to
somebody who actually relies on it.
thing you cannot see. Measure the audio. Count the utterances. Then give it to
somebody who actually relies on it.
Where it stops working, measured
A tool that claims to describe everything is lying about at least one case, so
here are three:
here are three:
| content | dialogue | descriptions produced |
|---|---|---|
| An animated film | 4% | 24 |
| A 1951 instructional film | 77% | 11 |
| A one-minute advert | 77% | none at all |
The advert result is correct behaviour, not a failure. It fills every second it
paid for, so there is nowhere to speak. That row is in the demo video, because
a system that knows where it stops is more trustworthy than one that claims it
never does.
paid for, so there is nowhere to speak. That row is in the demo video, because
a system that knows where it stops is more trustworthy than one that claims it
never does.
Links
- Demo video: https://youtu.be/vqJsFDnj7ko
- Live demo, no Fire TV needed: https://tushartechs.github.io/sightline/
- Source: https://github.com/TusharTechs/sightline
- The QA checks that catch invented characters and camera language, released
separately under MIT: https://github.com/TusharTechs/audio-description-qa
Built for Build, Ship, Shape: the Amazon Developer Hackathon 2026, on Vega OS.
Film in the demo is Sintel, (c) copyright Blender Foundation, www.sintel.org,
CC BY 3.0.
Film in the demo is Sintel, (c) copyright Blender Foundation, www.sintel.org,
CC BY 3.0.
Enjoyed reading this content? Let the author know!
Your likes, comments, shares, and saves help creators reach more builders.
Loading recommendations
Loading article