RAG explained in 50 lines of Python (with a local model)

Retrieval-augmented generation in 49 lines: embed documents, retrieve the closest ones, and put them in the prompt. Without RAG, a local model invented a vacation policy; with RAG it answered correctly, and one rule stopped it guessing.
TL;DR — RAG (retrieval-augmented generation) means: find the pieces of your documents that match the question, paste them into the prompt, and ask the model to answer from them. Asked about a fictional company without RAG, a small local model invented answers (“10 vacation days”); with RAG, it answered correctly (“28 days”). When the answer wasn’t in the documents, it still invented one until we told it to say “I don’t know”.
Tested on 2026-09-28 with sentence-transformers 6.1.0 (all-MiniLM-L6-v2) and transformers 5.17.0 (Qwen2.5-0.5B-Instruct), both running locally with no API key. The handbook below is fictional, so the model can’t know it from training.
Key points
- RAG has three steps: index (embed your documents once), retrieve (find the closest chunks), generate (answer with those chunks in the prompt).
- It gives a model knowledge it was never trained on, such as your internal documents, without retraining.
- Without RAG, the model guessed confidently wrong answers for all three questions.
- RAG doesn’t stop guessing when the answer is missing. A low retrieval score is the warning sign.
- A one-line instruction (“If the answer is not in the context, say: I don’t know.”) fixed that case without breaking the others.
What does RAG look like in code?
The whole pipeline, including a fictional company handbook:
import numpy as np
from sentence_transformers import SentenceTransformer
from transformers import AutoModelForCausalLM, AutoTokenizer
# A made-up company handbook: facts the model cannot know from training
handbook = """
Nordwind Analytics was founded in 2019 and has offices in Leipzig and Porto.
Employees get 28 days of paid vacation per year, plus their birthday off.
Remote work is allowed up to three days per week; Wednesdays are office days for everyone.
The data team uses a shared GPU server called 'Kestrel' with four A100 cards.
GPU jobs longer than 12 hours must be booked in the #gpu-queue channel.
Expense reports are due by the 5th of the following month.
The coffee machine on the second floor is cleaned every Friday by the facilities team.
New hires get a 30-minute onboarding call with the CTO in their first week.
Production database access requires approval from two team leads.
The annual company retreat takes place in the Harz mountains in September.
""".strip().split("\n")
# 1) Index: embed every chunk once (here, one sentence = one chunk)
embedder = SentenceTransformer("sentence-transformers/all-MiniLM-L6-v2")
chunk_vecs = embedder.encode(handbook, normalize_embeddings=True)
# 2) Retrieve: find the chunks most similar to the question
def retrieve(question, k=2):
q = embedder.encode([question], normalize_embeddings=True)[0]
return [handbook[i] for i in np.argsort(chunk_vecs @ q)[::-1][:k]]
# 3) Generate: put the retrieved text into the prompt
name = "Qwen/Qwen2.5-0.5B-Instruct"
tok = AutoTokenizer.from_pretrained(name)
llm = AutoModelForCausalLM.from_pretrained(name).eval()
def ask(question, context=None, rule=""):
system = "Answer in one short sentence." + rule
if context:
system += " Use only this context:\n" + "\n".join(context)
msgs = [{"role": "system", "content": system}, {"role": "user", "content": question}]
ids = tok.apply_chat_template(msgs, add_generation_prompt=True, return_tensors="pt", return_dict=True)
out = llm.generate(**ids, max_new_tokens=40, do_sample=False, pad_token_id=tok.eos_token_id)
return tok.decode(out[0, ids["input_ids"].shape[1]:], skip_special_tokens=True).strip()
for question in ["How many vacation days do Nordwind employees get?",
"What is the name of the GPU server?",
"Which day must everyone be in the office?"]:
context = retrieve(question)
print(f"Q: {question}")
print(f" without RAG: {ask(question)}")
print(f" retrieved: {context[0]}")
print(f" with RAG: {ask(question, context)}\n")
Q: How many vacation days do Nordwind employees get?
without RAG: Nordwind employees are entitled to 10 vacation days per year.
retrieved: Employees get 28 days of paid vacation per year, plus their birthday off.
with RAG: Nordwind employees get 28 days of paid vacation per year, plus their birthday off.
Q: What is the name of the GPU server?
without RAG: The name of the GPU server is "GPU" (Graphics Processing Unit).
retrieved: The data team uses a shared GPU server called 'Kestrel' with four A100 cards.
with RAG: The data team uses a shared GPU server called 'Kestrel'.
Q: Which day must everyone be in the office?
without RAG: Everyone must be in the office on Monday.
retrieved: Remote work is allowed up to three days per week; Wednesdays are office days for everyone.
with RAG: Wednesdays.
| Question | Without RAG | With RAG | Truth (handbook) |
|---|---|---|---|
| Vacation days | 10 days | 28 days + birthday | 28 days + birthday |
| GPU server name | “GPU” | Kestrel | Kestrel |
| Office day | Monday | Wednesdays | Wednesdays |
Without RAG, the model didn’t say “I don’t know”. It produced plausible, specific and wrong answers. That is what “hallucination” looks like. do_sample=False makes the answers deterministic, so the same code gives the same output.
What happens when the answer isn’t in your documents?
The handbook never names the CEO. Retrieval still returns something, and the model builds an answer on it:
# continues from the code above (retrieve, ask, embedder, chunk_vecs)
question = "What is the name of the CEO?" # not in the handbook
q = embedder.encode([question], normalize_embeddings=True)[0]
print(f"Q: {question}")
print(f" best similarity: {(chunk_vecs @ q).max():.2f} -> {retrieve(question)[0]}")
print(f" with RAG: {ask(question, retrieve(question))}")
rule = " If the answer is not in the context, say: I don't know."
print(f" with RAG + 'say I don't know' rule: {ask(question, retrieve(question), rule)}")
Q: What is the name of the CEO?
best similarity: 0.22 -> New hires get a 30-minute onboarding call with the CTO in their first week.
with RAG: The CEO of Nordwind Analytics is Peter Schmid.
with RAG + 'say I don't know' rule: I don't know.
“Peter Schmid” appears nowhere in the handbook. The retrieved sentence was about the CTO, with a similarity of only 0.22:

With the “I don’t know” rule added, the model declined the CEO question. We also re-ran five answerable questions with the rule; all five still got correct answers (vacation, GPU server, office day, expense reports, company retreat).
How do you make a toy RAG production-ready?
| Toy version above | In practice |
|---|---|
| One sentence = one chunk | Split documents into chunks of a few hundred tokens, with some overlap |
| Top 2 chunks | Tune k; too few misses facts, too many adds noise and cost |
| Plain dot product over a matrix | A vector index or database for large collections |
| No threshold | Treat a low best similarity as “not found” before calling the model |
| One test question at a time | Keep a small set of questions with known answers and re-check them after every change |
Scores and the right threshold depend on the embedding model and your documents, so measure on your own data.
Sources
- Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks — Lewis et al., 2020, the paper that named RAG
- sentence-transformers/all-MiniLM-L6-v2, Qwen/Qwen2.5-0.5B-Instruct — models used, checked 2026-09-28
Related: What are embeddings? Semantic search with cosine similarity in Python · What does temperature do in an LLM?