Back to Discover

#speaker diarization

1 prompt found

Speaker Diarized Transcription Pipeline Planner With WhisperX and pyannote speaker-diarization-3.1: Model Size vs VRAM, compute_type, batch_size, Alignment, Speaker Counts, HF Token via Env Var, SRT VTT JSON Outputs, and QA
๐Ÿค– AI Tools

Speaker Diarized Transcription Pipeline Planner With WhisperX and pyannote speaker-diarization-3.1: Model Size vs VRAM, compute_type, batch_size, Alignment, Speaker Counts, HF Token via Env Var, SRT VTT JSON Outputs, and QA

PpromptstudioยทOct 7, 2026
No rating

Plan a WhisperX transcription run with speaker labels: faster-whisper model size and compute type for your GPU, batch size, language and word alignment, pyannote speaker-diarization-3.1 with min and max speakers, gated model access through a Hugging Face token in an environment variable, SRT, VTT, and JSON outputs, a speaker name map, and QA checks for overlaps and very short turns.

Act as a machine learning engineer who runs WhisperX with its faster-whisper backend and pyannote diarization for podcasts, interviews, and meeting archives, and who has fixed CUDA out of memory errors, gated model 401s, and transcripts where every line was labeled the same speaker. Inputs: - Audio: file types, total hours, typical episode length, number of speakers, mics, crosstalk, background music: [AudioSet] - Hardware: GPU model and VRAM, or CPU only, OS, Python and CUDA versions, installed WhisperX version: [Hardware] - Language or languages, accents, and domain terms or names: [LanguageProfile] - Expected speaker count range per file and who the speakers are: [SpeakerInfo] - Outputs needed: SRT, VTT, JSON with word timings and speaker labels, plain text: [OutputSpec] - Accuracy versus speed priority and a deadline: [Priority] - Output format: [Format] Generate: 1. A model plan: faster-whisper model size (for example large-v3, medium, small) and compute_type (float16 on recent NVIDIA GPUs, int8_float16 or int8 to save memory, int8 on CPU) for Hardware. Do not state exact VRAM numbers unless the user supplied them; tell the user to check the project benchmarks and to measure peak memory on one file. 2. batch_size guidance: start value, and the rule to halve it first when CUDA runs out of memory before shrinking the model. 3. Language and alignment: set language explicitly from LanguageProfile instead of auto detect, run the alignment model for that language to get word timings, and note languages without a default alignment model. 4. Diarization: pyannote speaker-diarization-3.1 through WhisperX, min_speakers and max_speakers from SpeakerInfo, then assign word speakers. Access steps: accept the user conditions on the gated model pages for speaker-diarization-3.1 and its segmentation model on Hugging Face, create a read token, and pass it from an environment variable such as HF_TOKEN, never hardcoded. 5. The run: a CLI command and an equivalent Python script with the chosen flags, and a batch loop over AudioSet that skips files already done. 6. Outputs from OutputSpec: SRT and VTT with speaker prefixes, JSON with segments, words, start, end, and speaker, and where each lands. 7. A speaker name map: replace SPEAKER_00 style labels with real names after listening to the first minute, stored per file in a small JSON. 8. QA checks: percent of words with no speaker, turns shorter than about one second, overlapping speech regions, speaker count different from SpeakerInfo, and timestamps that drift; with a fix for each. 9. A troubleshooting table: 401 or gated model error, out of memory, wrong language, one speaker everywhere, music sections hallucinated as text. Constraints: - Function names and defaults change between WhisperX releases; mark such lines CHECK against the installed version. No em dashes.