AWS Builder Center
Detect speech, not silence: building audio description for video that has none

Detect speech, not silence: building audio description for video that has none

Most video has no audio description because a human has to write it. I built a Fire TV app that writes it instead, using Amazon Transcribe and Polly, and two blind reviewers found the bugs my own testing never could.

Building audio description for the content that never got any

What I learned making Transcribe, Polly and a Fire TV write descriptions for
video nobody had ever described, and what two blind reviewers found that I
had missed.

The seven seconds that explain the problem

Play any seven seconds of a film with your eyes shut. Someone stops walking.
The music changes. There is a small sound, and then quiet.
You know something happened. You have no idea what.
That is the ordinary experience of watching video without audio description,
and audio description is not a new or exotic fix for it. A person watches the
film, writes down what is on screen, and a voice reads it out in the gaps
between the dialogue. It works. Blind and low-vision viewers rely on it. There
is a whole craft around doing it well.
The catch is the economics. A human has to write it and someone has to pay
them, so it exists for a slice of what gets made and for everything else there
is nothing at all. Not a worse version. Nothing.
That is the pain point I wanted to attack: not "make description better", but
make it exist for the content that was never going to get any.

What I built

Sightline  is a Fire TV app on Vega
OS that writes the description nobody was going to write, and speaks it while
you watch.
It finds the gaps where no one is speaking, looks at what changed on screen
across each gap, writes only the part you would be lost without, measures the
line to check it fits the silence available, and speaks it there. One press
moves the description to a phone, so a blind viewer and a sighted viewer can
watch one screen and each get what they need. And you can ask it questions
mid-scene, which a pre-recorded description track structurally cannot answer,
because it was written before you had the question.
There is a three-minute demo , and a
live version  that needs no Fire TV.

The AWS shape of it

1
2
3
4
5
6
7
8
9
video ──> Amazon S3 ──> Amazon Transcribe ──> word timings ──> gaps
│
frame pairs per gap ──> model ──> lines ──┤
│
Amazon Polly (generative) ──────┘
│
timeline + PCM ──> Fire TV app
│
questions ──> API Gateway ──> Lambda ──> answer
Below is what each of those actually taught me. I have tried to write the
notes I wanted to find before I started.

Finding 1: detect speech, not silence. It was a 16x difference.

My first gap detector looked for silence, which is the obvious reading of
"description goes in the gaps."
On a 52-second trailer, silence detection found 2.5 usable seconds. The
score never stops, so by the measure of "is it quiet", there is almost nowhere
to speak.
Detecting speech with Transcribe and treating everything else as available
found 39.8 seconds.
That is a sixteenfold difference, and it is the single largest improvement in
the project. Music is not an obstacle to description. Dialogue is. Transcribe's
word-level timestamps are exactly the right primitive for this, and accuracy on
film dialogue over a score was excellent.
A second benefit came free: feeding the transcript text to the model stops the
description repeating a line of dialogue that was just spoken.
If you are timing anything against speech, ask for word timings and invert
them. Do not look for quiet.

Finding 2: Polly pads its output, and the speaking rate is not the rate you asked for

The generative engine sounds very good, and PCM output goes straight into the
Vega audio stream with no decoding on the device. Both were the right call.
Two things cost me a day.
Polly pads both ends of the audio with silence. If you are fitting speech
into a measured gap, that padding is charged against your budget. It has to be
trimmed before you measure anything.
The effective speaking rate is not the rate you request. I assumed 170 wpm,
which is the rate I asked for, and every single line overflowed its gap. The
real figure, measured after trimming the padding, was 203 wpm.
If you are doing timed speech, measure the rate of your specific voice and
engine before you build any budgeting logic on top of it. Do not trust the
number you passed in.

Finding 3: S3 is fine, but it is most of the clock for short files

Transcribe takes input from a bucket rather than a direct upload, so every job
means writing an object and cleaning it up afterwards. The bucket itself was
unremarkable in the best possible way: private, one object per job, deleted
after, no surprises.
But for a 52-second clip, the round trip through S3 is most of the wall-clock
time of that stage. For batch work that is irrelevant. It mattered here because
users are waiting.
This also closed off a feature. I wanted spoken questions to work in-app on
iPhones, where Safari has no SpeechRecognition at any origin, which means
recording audio and transcribing it server-side. Transcribe's batch path
through S3 is far slower than anyone will wait for an answer about the shot
they are looking at. Streaming transcription is the right tool and is on the
list; batch was never going to work for an interactive question.

Finding 4: Bedrock refused at the account level, and the error said nothing useful

This is the one I most want to pass on, because I lost real time to it.
The Bedrock client is written and shipped in the repository. It never ran. Every
region, every model, bare model IDs and both inference-profile forms returned:
1
Error 002: Access to Bedrock models is not allowed for this account
Meanwhile S3, Polly and Transcribe all worked on the same credentials, in the
same session. The account is not in an Organization. And
get-foundation-model-availability reported:
1
2
3
4
authorizationStatus: AUTHORIZED
entitlementAvailability: AVAILABLE
regionAvailability: AVAILABLE
agreementAvailability: NOT_AVAILABLE
The last line is the actual problem, and the error message never mentions it.
Per-model toggles in the console do not clear it either.
"Access is not allowed for this account" gives a developer nothing to act on.
An error naming the missing agreement, or a console surface showing agreement
status next to the model, would turn a support ticket into a click. I shipped
against a direct model API instead, which is a worse outcome for everybody.
If you hit Error 002 with working credentials for other services, check
agreementAvailability before you spend an afternoon on IAM.

What the blind reviewers changed, which is most of it

I did not want to guess at what would help, so I asked. Reviewers on the
ACB Audio Description Project  mailing list answered
in detail, and almost none of the rules this thing follows are mine.
Nothing during dialogue. Nothing that repeats what the sound already tells you.
Never "the camera pans", because the viewer is watching a story, not a shoot.
One reviewer proposed a test I would not have thought of: attenuate the whole
audio file and check the classification does not move. It found a real bug. My
loader was reading 16-bit, this film decodes to samples above full scale, and
those were being clipped. Reading 32-bit float fixed it, and gain invariance
went to exactly 0.00 dB.
A second tester, Dave Matters, used the player and found two faults in an
afternoon that all my local testing had missed. Description stopped a third of
the way through a clip, because a limit written to cap the number of
descriptions was truncating the film. And a scene where the software reported
someone watching a character sleep. Nobody was there. It had turned a camera
angle into a person.
I did not catch either. He did, and he cannot see the screen.

The lesson I keep coming back to: "it looks fine" is not evidence

Three of the worst bugs in this project presented as working software.
A truncated film that played normally. A phone that tracked the television's
playhead perfectly while speaking nothing at all, because the local service
answered 501 to every HEAD request and the phone reads "not ok" as "the file
does not exist". And an app that went completely silent in the one state where
a blind user has nothing else to go on: the spoken offer to describe an
undescribed video was never audible, because that code path returned before the
audio stream was initialised.
None of those threw an error. Two of them looked better than the working
version, because they were quieter.
If you are building accessibility features, the thing you have to test is the
thing you cannot see. Measure the audio. Count the utterances. Then give it to
somebody who actually relies on it.

Where it stops working, measured

A tool that claims to describe everything is lying about at least one case, so
here are three:
contentdialoguedescriptions produced
An animated film4%24
A 1951 instructional film77%11
A one-minute advert77%none at all
The advert result is correct behaviour, not a failure. It fills every second it
paid for, so there is nowhere to speak. That row is in the demo video, because
a system that knows where it stops is more trustworthy than one that claims it
never does.

Links

  • Demo video: https://youtu.be/vqJsFDnj7ko
  • Live demo, no Fire TV needed: https://tushartechs.github.io/sightline/
  • Source: https://github.com/TusharTechs/sightline
  • The QA checks that catch invented characters and camera language, released
    separately under MIT: https://github.com/TusharTechs/audio-description-qa
Built for Build, Ship, Shape: the Amazon Developer Hackathon 2026, on Vega OS.
Film in the demo is Sintel, (c) copyright Blender Foundation, www.sintel.org,
CC BY 3.0.

Any opinions in this article are those of the individual author and may not reflect the opinions of AWS.
Enjoyed reading this content? Let the author know!

Your likes, comments, shares, and saves help creators reach more builders.

Loading recommendations

Loading article