Speech & Audio
Transcribe speech, classify sounds or build voice interfaces.
Typical projects
Three ways to build it
Starter
Cloud speech-to-text API or Whisper (pretrained)State-of-the-art accuracy out of the box in ~100 languages.
Best for: No or little data, or new to ML
faster-whisper · Cloud STT APIs
Standard
Whisper (faster-whisper) self-hosted + LLM for summariesCheap at volume, private, and an LLM turns transcripts into insights.
Best for: Some labeled data and Python experience
faster-whisper · pyannote (speakers) · Anthropic SDK
Advanced
Fine-tuned Whisper / streaming ASR + speaker diarizationDomain vocabulary, real-time streaming and who-said-what.
Best for: Lots of data and an experienced team
Transformers · NVIDIA NeMo · pyannote · Triton
How success is measured
WER (Word Error Rate) for transcription; accuracy/F1 for sound classification
The data you'll need
- Collect real recordings: same microphones, accents, background noise as production.
- For custom vocabulary (product names), gather a list of terms.
- 1–10 hours of transcribed audio is enough to fine-tune for a domain.
Labeling
Correct auto-generated transcripts rather than typing from scratch (Label Studio supports audio).
Preparing the data
- Convert to 16 kHz mono WAV
- Split long recordings into segments (voice activity detection)
- Remove personal data from transcripts if required
Start with a baseline
Off-the-shelf Whisper or a cloud speech-to-text API on your recordings — measure WER.
Evaluating the model
- WER on a held-out set of real recordings
- Check names/numbers specifically
- Evaluate per accent / noise level
Monitoring in production
- Average confidence per recording
- Audio quality (clipping, silence)
- Processing time vs audio length
Common pitfalls
- Testing on clean audio but deploying in noisy environments
- Storing voice data without consent
Example code
CodeQuick start
# pip install faster-whisper
from faster_whisper import WhisperModel
model = WhisperModel("small", device="auto")
segments, info = model.transcribe("call.wav")
for s in segments:
print(f"[{s.start:.1f}s] {s.text}")CodeTrain your own model
# pip install faster-whisper anthropic
from faster_whisper import WhisperModel
import anthropic
asr = WhisperModel("large-v3", device="cuda", compute_type="float16")
segments, _ = asr.transcribe("call.wav", vad_filter=True)
transcript = " ".join(s.text for s in segments)
summary = anthropic.Anthropic().messages.create(
model="claude-haiku-4-5-20251001", max_tokens=400,
messages=[{"role": "user", "content":
f"Summarise this support call in 3 bullets and list action items:\n{transcript}"}])
print(summary.content[0].text)