LLM Assistant / RAG Chatbot
A chatbot or assistant that answers questions using your own documents and data.
Typical projects
Three ways to build it
Starter
Hosted LLM API (Claude) + managed RAG / long contextNo infrastructure to run; small corpora can even fit in the context window.
Best for: No or little data, or new to ML
Anthropic SDK · pypdf
Standard
RAG: hosted LLM + embeddings + vector DBScales to thousands of documents, answers cite sources, data stays updatable without retraining.
Best for: Some labeled data and Python experience
Anthropic SDK · sentence-transformers · pgvector / Qdrant / Chroma · FastAPI
Advanced
Hybrid search + reranker + agentic tool use, or self-hosted open-weights LLMHigher answer quality, and full data control when privacy requires on-prem.
Best for: Lots of data and an experienced team
vLLM · BM25 + embeddings · cross-encoder reranker · Ragas / promptfoo
How success is measured
Answer correctness and groundedness on a golden Q&A set; latency and cost per answer
The data you'll need
- Gather the documents the bot should know (PDF, HTML, Notion, Confluence…).
- Write 30–100 real questions with ideal answers — your golden set.
- Note which sources are authoritative and which are outdated.
Labeling
No training labels needed. Your golden Q&A set is the evaluation 'label'.
Preparing the data
- Extract text from documents (keep headings & page numbers)
- Split into chunks of ~300–800 tokens with overlap
- Create embeddings and store them in a vector database
- Attach metadata (source, date, permissions) to each chunk
Start with a baseline
Just put a few documents in the prompt of a hosted LLM and see how well it answers the golden set.
Evaluating the model
- Run the golden set after every change (prompt, chunking, model)
- Check retrieval hit-rate separately from answer quality
- Use an LLM-as-judge plus human spot checks
- Red-team: prompt injection, off-topic and unanswerable questions
Monitoring in production
- Thumbs up/down feedback
- Cost and tokens per conversation
- Questions with no good retrieval hits (content gaps)
- Latency p95
Common pitfalls
- Hallucination when retrieval misses — instruct the model to say 'I don't know'
- Leaking documents users shouldn't see — filter by permissions at retrieval time
Example code
CodeQuick start
# pip install anthropic pypdf
import anthropic
from pypdf import PdfReader
text = "\n".join(p.extract_text() for p in PdfReader("manual.pdf").pages)
client = anthropic.Anthropic()
resp = client.messages.create(
model="claude-sonnet-5-5", max_tokens=1024,
system="Answer only from the provided manual. If unsure, say you don't know.",
messages=[{"role": "user", "content": f"<manual>{text}</manual>\n\nHow do I reset the device?"}])
print(resp.content[0].text)CodeTrain your own model
# pip install anthropic chromadb sentence-transformers
import anthropic, chromadb
from chromadb.utils import embedding_functions
ef = embedding_functions.SentenceTransformerEmbeddingFunction("all-MiniLM-L6-v2")
db = chromadb.PersistentClient("./index").get_or_create_collection("docs", embedding_function=ef)
# 1) Index: chunks = list of (id, text, source)
# db.add(ids=[c[0] for c in chunks], documents=[c[1] for c in chunks],
# metadatas=[{"source": c[2]} for c in chunks])
# 2) Retrieve + generate
def answer(question: str) -> str:
hits = db.query(query_texts=[question], n_results=5)
context = "\n\n".join(f"[{m['source']}] {d}" for d, m in
zip(hits["documents"][0], hits["metadatas"][0]))
msg = anthropic.Anthropic().messages.create(
model="claude-sonnet-5-5", max_tokens=800,
system="Answer using only the context. Cite sources in [brackets].",
messages=[{"role": "user", "content": f"Context:\n{context}\n\nQuestion: {question}"}])
return msg.content[0].text
print(answer("What is the refund policy?"))