Hello Model
← Model library

Speech & Audio

Transcribe speech, classify sounds or build voice interfaces.

Typical projects

Transcribe customer calls and summarise themVoice commands for a mobile appDetect machine sounds that indicate faults

Three ways to build it

Starter

Cloud speech-to-text API or Whisper (pretrained)

State-of-the-art accuracy out of the box in ~100 languages.

Best for: No or little data, or new to ML

faster-whisper · Cloud STT APIs

Standard

Whisper (faster-whisper) self-hosted + LLM for summaries

Cheap at volume, private, and an LLM turns transcripts into insights.

Best for: Some labeled data and Python experience

faster-whisper · pyannote (speakers) · Anthropic SDK

Advanced

Fine-tuned Whisper / streaming ASR + speaker diarization

Domain vocabulary, real-time streaming and who-said-what.

Best for: Lots of data and an experienced team

Transformers · NVIDIA NeMo · pyannote · Triton

How success is measured

WER (Word Error Rate) for transcription; accuracy/F1 for sound classification

The data you'll need

  • Collect real recordings: same microphones, accents, background noise as production.
  • For custom vocabulary (product names), gather a list of terms.
  • 1–10 hours of transcribed audio is enough to fine-tune for a domain.

Labeling

Correct auto-generated transcripts rather than typing from scratch (Label Studio supports audio).

Preparing the data

  • Convert to 16 kHz mono WAV
  • Split long recordings into segments (voice activity detection)
  • Remove personal data from transcripts if required

Start with a baseline

Off-the-shelf Whisper or a cloud speech-to-text API on your recordings — measure WER.

Evaluating the model

  • WER on a held-out set of real recordings
  • Check names/numbers specifically
  • Evaluate per accent / noise level

Monitoring in production

  • Average confidence per recording
  • Audio quality (clipping, silence)
  • Processing time vs audio length

Common pitfalls

  • Testing on clean audio but deploying in noisy environments
  • Storing voice data without consent

Example code

CodeQuick start
python
# pip install faster-whisper
from faster_whisper import WhisperModel
model = WhisperModel("small", device="auto")
segments, info = model.transcribe("call.wav")
for s in segments:
    print(f"[{s.start:.1f}s] {s.text}")
CodeTrain your own model
python
# pip install faster-whisper anthropic
from faster_whisper import WhisperModel
import anthropic

asr = WhisperModel("large-v3", device="cuda", compute_type="float16")
segments, _ = asr.transcribe("call.wav", vad_filter=True)
transcript = " ".join(s.text for s in segments)

summary = anthropic.Anthropic().messages.create(
    model="claude-haiku-4-5-20251001", max_tokens=400,
    messages=[{"role": "user", "content":
        f"Summarise this support call in 3 bullets and list action items:\n{transcript}"}])
print(summary.content[0].text)

Ready to build one? Get a personalised plan →

Or read about Tabular Classification next.