Speech Recognition gpt_transcribe

Cloud speech recognition with OpenAI's next-generation gpt-transcribe model — upload audio, get a high-accuracy plain-text transcript.

Speech Recognition gpt_transcribe uses OpenAI's next-generation gpt-transcribe model — the latest speech-recognition generation after whisper-1 and gpt-4o-transcribe, with further gains in noisy environments, mixed-language speech, proper nouns, and accents, while keeping sentence structure and punctuation intact for a highly readable transcript. Just upload an audio or video file; the system splits the audio into segments, sends them to the cloud, and merges the results into a single output. Common formats (mp3, wav, m4a, ogg, mp4, mov, webm) are supported with no pre-conversion needed. Best suited for meeting minutes, interviews, podcasts, and online-course recordings. This module outputs a plain-text transcript only and does not support SRT subtitles; for SRT use "Speech Recognition", for multiple speakers use "Multi-Speaker Speech Recognition", and for a live as-you-speak transcript use "Live Transcription gpt_live_transcribe". Maximum total upload size per request: 256MB (all files and message content combined). Output content formats: Text This module does not support continuous conversations.