Speech Recognition gpt-4o-transcribe (Multi Speaker) is powered by OpenAI's cloud gpt-4o-transcribe-diarize model — the diarization-enabled variant of the gpt-4o-transcribe family. The transcript is automatically labeled with speaker tags (SPEAKER_01, SPEAKER_02, SPEAKER_03, …) so meeting minutes and interview transcripts come back ready to read. Compared to v1 (single-speaker), this module is purpose-built for meetings, interviews, panels, and customer-support calls where you need to know who said what, while keeping the strong noise robustness, mixed-language handling, proper-noun accuracy, and well-preserved sentence structure / punctuation that the gpt-4o-transcribe family is known for. Just upload an audio or video file (mp3, wav, m4a, ogg, mp4, mov, webm); the system splits it into 59-second chunks for stability on long recordings, sends each chunk to OpenAI in the cloud, and merges the results into a single output. You get both a plain-text transcript and SRT subtitles — drop the SRT straight into Premiere, DaVinci Resolve, YouTube, or Vimeo, and feed the transcript into the summary, translation, or meeting-notes modules for further processing. For single-speaker recordings, use "Speech Recognition gpt-4o-transcribe" instead for cleaner output.