AI Guide
What Is RAG?
RAG lets AI answer questions using your own documents. It's how you build chatbots that know your company's data, your notes, or your product docs — without retraining the model.
Quick answer: RAG = Retrieval-Augmented Generation. When a user asks a question, you search your own documents for relevant text, then pass that text to the AI as context. The AI answers using the retrieved context — not just its training data.
💡 Why RAG Matters
- Grounds answers in facts. The AI stops making things up when it has your actual documents.
- Uses your data. Company docs, notes, PDFs, database entries — all can be searched.
- No retraining. Add new documents any time. No expensive model updates.
- Cheaper than fine-tuning. Fine-tuning costs thousands. RAG costs pennies.
- Cited answers. You can show users which source the answer came from.
🧠 How RAG Works (4 Steps)
1. Index your documents
Split each document into chunks (paragraphs). Convert each chunk to an embedding — a vector of ~1,500 numbers that captures meaning. Store the vector + original text.
2. User asks a question
Convert the question to an embedding using the same model.
3. Retrieve relevant chunks
Find the chunks whose embeddings are closest to the question's embedding. That's "similar meaning."
4. Generate with context
Send the retrieved chunks + the question to the AI. The AI answers using the context.
📊 RAG vs Fine-Tuning
- Setup time: RAG = hours. Fine-tuning = days or weeks.
- Cost: RAG = cents per query. Fine-tuning = thousands to train.
- Updates: RAG = add a new file. Fine-tuning = retrain the model.
- Accuracy: RAG grounds answers in retrieved facts. Fine-tuning bakes knowledge into the model.
- Best for: RAG for dynamic, factual data. Fine-tuning for style or behavior changes.
For 99% of use cases, RAG is the right answer.
🔢 What Is an Embedding?
An embedding is a list of numbers that represents the meaning of text. Similar meanings produce similar numbers.
"I love dogs" → [0.21, -0.55, 0.87, ...]
"I adore puppies" → [0.19, -0.52, 0.89, ...] ← very similar
"How to bake bread" → [-0.71, 0.44, 0.12, ...] ← very different
This is how computers find "similar meaning" without keyword matching.
🗄️ Vector Databases
A vector database stores embeddings and finds the closest matches fast.
- Pinecone — hosted, easy, generous free tier
- Supabase pgvector — if you already use Postgres
- Weaviate — open source, self-hostable
- Chroma — simplest for local projects
💡 Starting out? For under 500 documents, skip the vector DB. Store embeddings in a JSON file and compare in memory. Move to a real DB only when you need it.
🎯 When to Use RAG
- Customer support chatbot that knows your product docs
- Internal search for company wikis and PDFs
- Personal AI that answers questions about your notes
- Legal or medical AI that must cite sources
- Documentation Q&A for developer tools
🚫 Common Beginner Mistakes
- Chunks too large. Huge chunks dilute meaning. Aim for 200–500 words per chunk.
- Chunks too small. Tiny chunks lose context. Find the balance.
- Forgetting overlap. Adjacent chunks should overlap by ~10% so sentences aren't cut in half.
- Using different embedding models for indexing and querying. They must match.
- Not including the source. Show users which document the answer came from.
- Trusting output blindly. Even with RAG, models can hallucinate. Add a "I don't know" fallback.
❓ Frequently Asked Questions
What is RAG?
Retrieval-Augmented Generation. A technique that searches your own documents and passes the results to the AI before it answers.
Why use RAG instead of fine-tuning?
RAG is cheaper, faster to set up, and updates instantly. Fine-tuning is expensive and slow.
What is an embedding?
A vector of numbers representing text meaning. Similar meanings produce similar vectors.
What is a vector database?
A database optimised to store and query embeddings. Pinecone, Supabase pgvector, Weaviate, Chroma.
Do I need a vector DB for small projects?
No. For hundreds of documents, a JSON file + in-memory similarity is fine. Scale up when you need to.