Master Code
On The Go

Learn. Practice. Build.

AI Guide

What Is RAG?

RAG lets AI answer questions using your own documents. It's how you build chatbots that know your company's data, your notes, or your product docs — without retraining the model.

Quick answer: RAG = Retrieval-Augmented Generation. When a user asks a question, you search your own documents for relevant text, then pass that text to the AI as context. The AI answers using the retrieved context — not just its training data.

💡 Why RAG Matters

🧠 How RAG Works (4 Steps)

1. Index your documents

Split each document into chunks (paragraphs). Convert each chunk to an embedding — a vector of ~1,500 numbers that captures meaning. Store the vector + original text.

2. User asks a question

Convert the question to an embedding using the same model.

3. Retrieve relevant chunks

Find the chunks whose embeddings are closest to the question's embedding. That's "similar meaning."

4. Generate with context

Send the retrieved chunks + the question to the AI. The AI answers using the context.

📊 RAG vs Fine-Tuning

For 99% of use cases, RAG is the right answer.

🔢 What Is an Embedding?

An embedding is a list of numbers that represents the meaning of text. Similar meanings produce similar numbers.

"I love dogs"    → [0.21, -0.55, 0.87, ...]
"I adore puppies" → [0.19, -0.52, 0.89, ...]  ← very similar

"How to bake bread" → [-0.71, 0.44, 0.12, ...] ← very different

This is how computers find "similar meaning" without keyword matching.

🗄️ Vector Databases

A vector database stores embeddings and finds the closest matches fast.

💡 Starting out? For under 500 documents, skip the vector DB. Store embeddings in a JSON file and compare in memory. Move to a real DB only when you need it.

🎯 When to Use RAG

🚫 Common Beginner Mistakes

❓ Frequently Asked Questions

What is RAG?

Retrieval-Augmented Generation. A technique that searches your own documents and passes the results to the AI before it answers.

Why use RAG instead of fine-tuning?

RAG is cheaper, faster to set up, and updates instantly. Fine-tuning is expensive and slow.

What is an embedding?

A vector of numbers representing text meaning. Similar meanings produce similar vectors.

What is a vector database?

A database optimised to store and query embeddings. Pinecone, Supabase pgvector, Weaviate, Chroma.

Do I need a vector DB for small projects?

No. For hundreds of documents, a JSON file + in-memory similarity is fine. Scale up when you need to.

🎯 What's Next?

← All How-To Guides