Speech Recognition WhisperX (Single Speaker)

Single-speaker local speech-to-text. Outputs plain text and SRT subtitles without diarization.

Speech Recognition WhisperX (Single Speaker) is optimized for recordings with one speaker — personal voice memos, vlog narration, solo podcasts, online courses, lecture recordings, and similar single-voice audio. Upload your audio file and the local WhisperX model produces an accurate transcript along with synchronized SRT subtitles. Skipping speaker diarization makes it noticeably faster and lighter than the multi-speaker variant. The whole pipeline runs on your on-prem GPU. Mixed Chinese / English and common Asian languages are supported; supplying a glossary in the notes helps with proper nouns. Outputs can be copied directly or downloaded as SRT for video editing. For recordings with multiple speakers where you need to know who said what, use "Speech Recognition WhisperX (Multi Speaker)" instead.