A few weeks ago, I asked y'all where you got your subtitles. Someone recommended Bazarr, but I became intrigued by the idea of automated transcription. Subtitles are often summaries of the dialogue, and often use different vocabulary and construction, so I thought if I could get an automated transcription of any show, that would be the bee's knees.
So my new pipeline:
Show --> automatic transcription into SRT file --> manual correction & touch up --> compressed audio file of dialogue only
Automated transcription is... OK. There's plenty of stuff that it misses, so it's important to be listening the first time through for hallucinations and mistakes; otherwise you waste your time looking up irrelevant grammar and sabotage yourself listening for things that aren´t there. I don´t know if I'd recommend it to someone just starting out, but if you're phonetically aware to at least realize that the text does not match what you're hearing, I think automated transcription is ok.
And once you have a transcription keyed to the actual spoken dialogue, you can create a dialogue-only audio file, which means that instead of listening to a 100 minute movie for 20 minutes of dialogue, you can listen to 20 minutes of dialogue 5 times.
Anyone do similar? Tips? Tricks?
Details: faster-whisper with VOD enabled.
Yeah, there are two good reasons for that: lip sync, and adapting references to the local audience. The subtitles generally don't have to bother doing either of those, and furthermore they must limit their line size to make the subtitles readable at a normal speed. The result is that the sub and the dub diverge greatly - back when I still watched media in Spanish, I used to put subs and dubs on simultaneously to catch as much information from the original as possible
I am aware. It was short phrases that could easily have fit into subtitles unchanged that largely drove this suspiscion that there was some purpose, some consciousness, behind the divergence.