Hello Model
← Model library

Text Classification

Sort text into categories: sentiment, topic, intent, spam, priority.

Typical projects

Detect sentiment of product reviewsRoute support tickets to the right teamFilter spam emails

Three ways to build it

Starter

LLM zero-shot / few-shot classification via API

No training needed — describe the categories in a prompt. Great for prototyping or < 100 examples.

Best for: No or little data, or new to ML

Anthropic / OpenAI SDK · pydantic

Standard

Fine-tuned small transformer (DistilBERT / ModernBERT) or SetFit

Cheap to run, fast, and accurate once you have a few hundred labels.

Best for: Some labeled data and Python experience

Hugging Face Transformers · SetFit · datasets

Advanced

Fine-tuned larger encoder + active learning loop

Maximises accuracy on large, evolving datasets while keeping labeling cost low.

Best for: Lots of data and an experienced team

Transformers · Argilla · MLflow · ONNX Runtime

How success is measured

F1 score per class (macro-F1 when classes are imbalanced)

The data you'll need

  • Collect real text examples from the place the model will run (tickets, reviews).
  • 200–500 labeled examples per class is enough for fine-tuning a small transformer.
  • Write a one-page labeling guide so everyone labels the same way.

Labeling

Use Label Studio or Argilla. Or bootstrap: let an LLM pre-label and have humans correct it.

Preparing the data

  • Remove duplicates and boilerplate (signatures, quoted replies)
  • Keep raw text — modern models don't need stemming/stop-word removal
  • Stratified train/validation/test split

Start with a baseline

TF-IDF + logistic regression. Trains in seconds and is often 80–90% as good.

Evaluating the model

  • Per-class precision/recall
  • Read 50 misclassified examples — they reveal labeling issues
  • Test on very short and very long inputs

Monitoring in production

  • Class distribution over time
  • Low-confidence rate (send to human review)
  • New vocabulary / topics appearing

Common pitfalls

  • Inconsistent labels between annotators
  • Training on clean text but serving messy text

Example code

CodeQuick start
python
# pip install anthropic
import anthropic
client = anthropic.Anthropic()  # reads ANTHROPIC_API_KEY

LABELS = ["billing", "bug", "feature_request", "other"]
def classify(text: str) -> str:
    msg = client.messages.create(
        model="claude-haiku-4-5-20251001", max_tokens=10,
        messages=[{"role": "user", "content":
            f"Classify this support ticket into one of {LABELS}. "
            f"Answer with the label only.\n\n{text}"}])
    return msg.content[0].text.strip()

print(classify("I was charged twice this month"))
CodeTrain your own model
python
# pip install transformers datasets evaluate accelerate
from datasets import load_dataset
from transformers import (AutoTokenizer, AutoModelForSequenceClassification,
                          TrainingArguments, Trainer)

ds = load_dataset("csv", data_files={"train": "train.csv", "test": "test.csv"})
labels = sorted(set(ds["train"]["label"]))
l2id = {l: i for i, l in enumerate(labels)}
ds = ds.map(lambda x: {"labels": l2id[x["label"]]})

name = "distilbert-base-uncased"
tok = AutoTokenizer.from_pretrained(name)
ds = ds.map(lambda b: tok(b["text"], truncation=True, max_length=256), batched=True)
model = AutoModelForSequenceClassification.from_pretrained(name, num_labels=len(labels))

args = TrainingArguments("out", num_train_epochs=3, per_device_train_batch_size=16,
                         eval_strategy="epoch", learning_rate=2e-5)
trainer = Trainer(model=model, args=args, train_dataset=ds["train"],
                  eval_dataset=ds["test"], tokenizer=tok)
trainer.train()
trainer.save_model("model")

Ready to build one? Get a personalised plan →

Or read about Image Classification next.