Build RAG With Your Notes and Check What It Gets Wrong

By Sheng Pang · Published · 9 min read

AI basics rag embeddings ollama python

Our local model knew about Qwen3, but it could not name the laptop we tested it on.

Giving it the relevant paragraph worked. You can search a folder of text files and pass the matching paragraphs to Qwen3 with 62 lines of Python.

What you need before step one

You need Ollama running and two models downloaded. Our local model guide walks you through the setup.

ollama pull nomic-embed-text
ollama pull qwen3:8b

Nomic Embed Text turns text into vectors and its download was 274 MB. Qwen3 8B writes the answers and was 5.2 GB.

For the test, we exported 35 of our articles as plain text, 33,755 words. You can use your own folder of txt files.

The downloadable example and run records include that fixed corpus snapshot, the script, a reproduction runner and the full generation responses.

We ran it with Ollama 0.35.0 on an M2 Pro with 32 GB of memory, using Python's standard library throughout.

import json, sys, time, math, glob, os, urllib.request

OLLAMA = "http://localhost:11434"
EMBED_MODEL = "nomic-embed-text"
CHAT_MODEL = "qwen3:8b"
CHUNK_WORDS = 150
OVERLAP = 30

This function sends requests to Ollama's local HTTP API with a timeout so a stalled call cannot wait indefinitely.

def post(path, body):
    req = urllib.request.Request(OLLAMA + path, data=json.dumps(body).encode(), headers={"Content-Type": "application/json"})
    return json.load(urllib.request.urlopen(req, timeout=120))

Step 1, cut each file into 150 word chunks

Pasting all 33,755 words into every question would waste context and processing time. Part four explains those limits.

Instead, you cut each file into pieces and retrieve a few for each question. We started with 150 words and a 30 word overlap.

def chunk(text):
    words = text.split()
    step = CHUNK_WORDS - OVERLAP
    return [" ".join(words[i:i + CHUNK_WORDS]) for i in range(0, max(1, len(words) - OVERLAP), step)]

The overlap repeats words around each boundary so a short sentence crossing it can survive intact. It does not keep every paragraph together.

On the saved files, this produced 289 chunks. If you edit the files, your count and search results will change.

Step 2, turn every chunk into 768 numbers

Our embedding model returned 768 floats per chunk, which let you compare texts by meaning as described in part three.

For retrieval, Nomic requires task prefixes, search_document: for stored chunks and search_query: for questions. The Ollama model we downloaded did not add them for us.

def build_index(folder, out="index.json"):
    chunks = []
    for path in sorted(glob.glob(os.path.join(folder, "*.txt"))):
        for c in chunk(open(path, encoding="utf-8").read()):
            chunks.append({"source": os.path.basename(path), "text": c})
    t = time.time()
    vectors = post("/api/embed", {"model": EMBED_MODEL, "input": ["search_document: " + c["text"] for c in chunks], "truncate": False})["embeddings"]
    for c, v in zip(chunks, vectors):
        c["vector"] = v
    json.dump(chunks, open(out, "w"))
    print(f"{len(chunks)} chunks, {len(vectors[0])} numbers each, embedded in {time.time() - t:.1f}s")

The embed endpoint accepts a list, so we sent all 289 chunks together. Setting truncate to false makes oversized input fail rather than silently lose its ending.

Building this index took 5.4 seconds and wrote a 3.23 MB JSON file on our machine, though your timings will differ.

Step 3, that JSON file is your vector database

You can open the JSON file to see each chunk's text, source name and vector before using it to search your notes.

Scanning and sorting our 289 vectors took a median of 0.015 seconds across ten runs, excluding file loading and question embedding.

Linear extrapolation gives roughly 52 seconds and 11 GB of JSON at a million chunks. We did not benchmark that scale, and memory and parsing add costs too.

You can keep the JSON file while searches are fast enough for your use. Measure your own query time before choosing a database with an index.

Step 4, find the three chunks closest to the question

After you embed your question with the query prefix, cosine similarity compares the direction of that vector with each stored vector.

def cosine(a, b):
    dot = sum(x * y for x, y in zip(a, b))
    return dot / (math.sqrt(sum(x * x for x in a)) * math.sqrt(sum(y * y for y in b)))

Sort the chunks by their cosine scores, which range from minus one to one, and return the top three.

def search(question, index="index.json", k=3):
    chunks = json.load(open(index))
    q = post("/api/embed", {"model": EMBED_MODEL, "input": "search_query: " + question, "truncate": False})["embeddings"][0]
    scored = sorted(((cosine(q, c["vector"]), c) for c in chunks), key=lambda s: -s[0])
    return scored[:k]

For our five questions at this chunk size, full searches took 0.065 to 0.623 seconds including the HTTP call and file loading.

The winning score tells you which chunk ranked first. It does not tell you whether the chunk actually contains an answer.

Step 5, paste the chunks into the prompt and ask

Put your retrieved chunks between the instruction and the question. The model can still invent an answer, even when you ask it to stick to the notes.

def ask(question, context=None):
    if context:
        prompt = ("Answer using only the notes below. If the notes do not contain the answer, say you do not know.\n\n"
                  + "\n\n".join(f"[{i + 1}] {c['text']}" for i, c in enumerate(context))
                  + f"\n\nQuestion: {question}")
    else:
        prompt = question
    r = post("/api/generate", {"model": CHAT_MODEL, "prompt": prompt, "stream": False, "think": False, "options": {"num_predict": 250, "temperature": 0, "seed": 0}})
    return r["response"].strip(), r["prompt_eval_count"], r["eval_count"], r["total_duration"] / 1e9

if __name__ == "__main__":

With thinking off, temperature and seed set to zero, and a 250 token output cap, you can run the same comparison.

Ollama reports duration in nanoseconds, and the function converts it to seconds. These settings make runs easier to compare, but do not guarantee identical output everywhere.

The main block prints the retrieved text before showing the bare answer and the answer with notes.

    if sys.argv[1] == "index":
        build_index(sys.argv[2])
    else:
        question = sys.argv[1]
        t = time.time(); hits = search(question); st = time.time() - t
        print(f"search {st:.2f}s")
        for score, c in hits:
            print(f"  {score:.3f} {c['source']}: {c['text'][:90]}...")
        a, pi, po, d = ask(question)
        print(f"\nNO CONTEXT ({pi} in, {po} out, {d:.1f}s):\n{a}\n")
        a, pi, po, d = ask(question, [c for _, c in hits])
        print(f"WITH CONTEXT ({pi} in, {po} out, {d:.1f}s):\n{a}\n")
python3 rag.py index notes
python3 rag.py "Which laptop did we use to time Qwen3 8B, and how many tokens per second did it write?"

Save the pieces above as rag.py, or use the download, where you also get a runner that saves complete answers.

Here is what changed on three questions

We asked each question once without notes and once with the top three chunks. This is a small worked example, not an accuracy benchmark.

QuestionTop scoreBare promptWith chunks
Which laptop did we use to time Qwen3 8B, and how many tokens per second did it write?0.800Generic hardware advice, hit the 250 token cap, 13.4 sMatched the notes, 45 tokens, 3.7 s
What is the difference between UUID v4 and UUID v7?0.782Hit the cap before covering v7Covered both layouts, 149 tokens, 7.1 s
What is the refund policy for the Pro plan?0.556Suggested a 30 day window without a sourceSaid the notes contain no answer, 9 tokens, 2.3 s

For the laptop, the retrieved answer named a 2023 MacBook Pro, M2 Pro, 32 GB and about 24 tokens per second.

While the answer with notes matches our recorded laptop test, the bare answer discussed general hardware and stopped mid sentence.

The shorter answer finished sooner, but at 45 tokens against 250 you cannot infer a processing speed gain from those totals.

Using our v4 versus v7 article, the reply gave v4 122 random bits and v7 a 48 bit timestamp, consistent with RFC 9562.

The question your notes cannot answer needs a separate test

Our saved notes have no Pro refund policy, and the question leaves the service unnamed, which makes a specific answer suspect from the start.

The bare reply suggested a 30 day window. With notes, it answered: "The notes do not contain the answer."

We then removed only "If the notes do not contain the answer, say you do not know." The model still said the policy was not mentioned.

Both prompts refused this question, which leaves us without evidence that this sentence was needed in this run.

Despite having no answer in the corpus, the refund question scored 0.556, enough to pass a cutoff of 0.55.

If you use a score cutoff, test it against answerable and unanswerable questions from your own files. These three questions cannot establish a reliable threshold.

Chunk size changes what reaches the model

We rebuilt the same files with smaller and larger chunks. For the laptop question, the highest score was less useful than reading the winning text.

  • 50 words with 10 overlap: 853 chunks, 9.0 seconds to build. The top laptop chunk scored 0.786 and named the machine, but omitted its measured speed.
  • 150 words with 30 overlap: 289 chunks. The top chunk scored 0.800 and contained both the laptop and the speed.
  • 600 words with 60 overlap: 77 chunks, 3.9 seconds to build. The top chunk scored 0.627 and came from our token saving article, without the laptop answer.
  • The UUID question scored 0.770, 0.782 and 0.773 across those sizes. Its winning chunks all came from the UUID comparison.

Changing chunk size moved the refund score from 0.613 to 0.556 to 0.529, pushing this unanswerable question across that same 0.55 cutoff.

We did not generate answers for these two extra indexes. The comparison tests retrieval coverage, not whether Qwen3 could repair a missing or wrong chunk.

Where this version breaks

Your index is a snapshot, so editing a file does not update it. You need to rebuild, and this script rereads the whole index for each question.

The model retains its pretrained knowledge even when you supply notes and instruct it to answer only from them.

At 150 words, "how fast was the Mac" scored 0.713 for a generic hardware chunk that lacked our measured 24 tokens per second.

The right paragraph ranked third, so the model would still receive it. A wrong first match does not by itself prove the final answer will be wrong.

We chose three chunks without tuning that number. You might need more if an answer is spread across several files, at the cost of a longer prompt.

Ollama reported both models loaded, 9.9 GB for Qwen3 and 370 MB for the embedder. We did not test a machine with less memory.

Try it on a folder you actually care about

Index a few of your own documents and ask questions with answers you can check. Read the retrieved chunks before judging the generated reply.

If an answer is wrong, first check whether the needed text reached the prompt. A prompt edit cannot supply a paragraph that retrieval missed.

Then ask something your folder cannot answer and check whether the model says so.

← More AI basics articles  ·  All articles

↑ Top