🤖 AI Tools

Speaker Diarized Transcription Pipeline Planner With WhisperX and pyannote speaker-diarization-3.1: Model Size vs VRAM, compute_type, batch_size, Alignment, Speaker Counts, HF Token via Env Var, SRT VTT JSON Outputs, and QA

Plan a WhisperX transcription run with speaker labels: faster-whisper model size and compute type for your GPU, batch size, language and word alignment, pyannote speaker-diarization-3.1 with min and max speakers, gated model access through a Hugging Face token in an environment variable, SRT, VTT, and JSON outputs, a speaker name map, and QA checks for overlaps and very short turns.

0.0
0Reviews
P
October 7, 2026

Prompt

Act as a machine learning engineer who runs WhisperX with its faster-whisper backend and pyannote diarization for podcasts, interviews, and meeting archives, and who has fixed CUDA out of memory errors, gated model 401s, and transcripts where every line was labeled the same speaker.

Inputs:
- Audio: file types, total hours, typical episode length, number of speakers, mics, crosstalk, background music: [AudioSet]
- Hardware: GPU model and VRAM, or CPU only, OS, Python and CUDA versions, installed WhisperX version: [Hardware]
- Language or languages, accents, and domain terms or names: [LanguageProfile]
- Expected speaker count range per file and who the speakers are: [SpeakerInfo]
- Outputs needed: SRT, VTT, JSON with word timings and speaker labels, plain text: [OutputSpec]
- Accuracy versus speed priority and a deadline: [Priority]
- Output format: [Format]

Generate:
1. A model plan: faster-whisper model size (for example large-v3, medium, small) and compute_type (float16 on recent NVIDIA GPUs, int8_float16 or int8 to save memory, int8 on CPU) for Hardware. Do not state exact VRAM numbers unless the user supplied them; tell the user to check the project benchmarks and to measure peak memory on one file.
2. batch_size guidance: start value, and the rule to halve it first when CUDA runs out of memory before shrinking the model.
3. Language and alignment: set language explicitly from LanguageProfile instead of auto detect, run the alignment model for that language to get word timings, and note languages without a default alignment model.
4. Diarization: pyannote speaker-diarization-3.1 through WhisperX, min_speakers and max_speakers from SpeakerInfo, then assign word speakers. Access steps: accept the user conditions on the gated model pages for speaker-diarization-3.1 and its segmentation model on Hugging Face, create a read token, and pass it from an environment variable such as HF_TOKEN, never hardcoded.
5. The run: a CLI command and an equivalent Python script with the chosen flags, and a batch loop over AudioSet that skips files already done.
6. Outputs from OutputSpec: SRT and VTT with speaker prefixes, JSON with segments, words, start, end, and speaker, and where each lands.
7. A speaker name map: replace SPEAKER_00 style labels with real names after listening to the first minute, stored per file in a small JSON.
8. QA checks: percent of words with no speaker, turns shorter than about one second, overlapping speech regions, speaker count different from SpeakerInfo, and timestamps that drift; with a fix for each.
9. A troubleshooting table: 401 or gated model error, out of memory, wrong language, one speaker everywhere, music sections hallucinated as text.

Constraints:
- Function names and defaults change between WhisperX releases; mark such lines CHECK against the installed version. No em dashes.

Instructions

Replace every [bracket] with your details before running. Works on ChatGPT, Claude, and Gemini.

Generated Output

This image was generated using the prompt above.

Speaker Diarized Transcription Pipeline Planner With WhisperX and pyannote speaker-diarization-3.1: Model Size vs VRAM, compute_type, batch_size, Alignment, Speaker Counts, HF Token via Env Var, SRT VTT JSON Outputs, and QA - Result

Examples

Example Input

AudioSet: 60 podcast episodes, WAV 48 kHz, 45 to 70 minutes each, two hosts plus an occasional guest, separate mics mixed to stereo, short music intro and outro
Hardware: RTX 4070 with 12 GB VRAM, Ubuntu 22.04, Python 3.11, CUDA 12, WhisperX installed from pip last month
LanguageProfile: English, US accents, terms like Kubernetes, Postgres, and guest names
SpeakerInfo: 2 to 3 speakers; hosts are Maya and Dev
OutputSpec: SRT and VTT for the video editor, JSON with word timings for search
Priority: accuracy first, done within a week
Format: plan, commands, script, outputs, name map, QA, troubleshooting

Example Output

Model plan
- large-v3 with compute_type float16 on the RTX 4070 (12 GB). Measure peak memory on episode 1 with nvidia-smi; if it is near the limit, switch to int8_float16 before changing models.
- batch_size 16 to start; on CUDA out of memory, try 8, then 4.
- language en set explicitly; English alignment model for word timings.

Access setup (once)
1. On Hugging Face, accept the conditions for pyannote/speaker-diarization-3.1 and pyannote/segmentation-3.0.
2. Create a read token and export it: export HF_TOKEN=... in your shell profile, not in code.

CLI for one episode
whisperx ep001.wav --model large-v3 --language en --compute_type float16 --batch_size 16 \
  --diarize --min_speakers 2 --max_speakers 3 --hf_token "$HF_TOKEN" \
  --output_format all --output_dir out/ep001

Python batch script (CHECK import paths against your version)
import os, json, pathlib, whisperx
from whisperx.diarize import DiarizationPipeline  # older releases: whisperx.DiarizationPipeline
dev = "cuda"
asr = whisperx.load_model("large-v3", dev, compute_type="float16", language="en")
align_model, meta = whisperx.load_align_model(language_code="en", device=dev)
diar = DiarizationPipeline(use_auth_token=os.environ["HF_TOKEN"], device=dev)
for wav in sorted(pathlib.Path("audio").glob("*.wav")):
    out = pathlib.Path("out") / wav.stem
    if (out / "done.flag").exists():
        continue
    out.mkdir(parents=True, exist_ok=True)
    audio = whisperx.load_audio(str(wav))
    res = asr.transcribe(audio, batch_size=16)
    res = whisperx.align(res["segments"], align_model, meta, audio, dev)
    turns = diar(audio, min_speakers=2, max_speakers=3)
    res = whisperx.assign_word_speakers(turns, res)
    (out / "transcript.json").write_text(json.dumps(res))
    (out / "done.flag").touch()
Write SRT and VTT with whisperx.utils.get_writer, or use the CLI with --output_format all.

Outputs
- out/ep001/ep001.srt and .vtt: lines prefixed [SPEAKER_00]
- out/ep001/transcript.json: segments with start, end, text, speaker, and words with start, end, speaker

Speaker name map (out/ep001/speakers.json)
{"SPEAKER_00": "Maya", "SPEAKER_01": "Dev", "SPEAKER_02": "Guest: Priya Natarajan"}
Fill by listening to the first minute; labels are not stable across episodes.

QA checks per episode
| Check | Flag when | Fix |
| Words with no speaker | more than a few percent | re-run diarization; inspect music regions |
| Turns under about 1 s | many in a row | merge into neighbor turn if same speaker context |
| Overlap regions | hosts talk over each other | mark [crosstalk] in the edit, check by ear |
| Speaker count | 1 or more than 3 found | set min and max to the known count |
| Drift | captions late at the end | confirm alignment ran; re-align |

Troubleshooting
- 401 or gated error: conditions not accepted or token not exported.
- Out of memory: lower batch_size, then int8_float16.
- Text during music intro: trim intro or drop segments with no aligned words.
- One speaker everywhere: stereo mix collapsed; set min_speakers 2.

Reviews (0)

Please login to leave a review.
Loading reviews...