RAG example recorded on 2026-10-06

Install and start Ollama, then pull nomic-embed-text and qwen3:8b.
Run these commands from this folder:

ollama pull nomic-embed-text
ollama pull qwen3:8b
python3 rag.py index notes
python3 rag.py "Which laptop did we use to time Qwen3 8B, and how many tokens per second did it write?"
python3 reproduce.py /tmp/rag-new-run

rag.py uses only the Python standard library. reproduce.py rebuilds three indexes, searches five questions and saves new results and responses. It calls local Ollama and may take a few minutes.

notes contains the fixed corpus used for the published run. results.json records its file hashes, model digests, options, retrieved text, scores and answer summaries. responses.json preserves complete generation requests and responses. Embedding records contain metadata and vector hashes rather than the vectors themselves. Timings and model outputs can differ on another machine or model release. sources.json lists official reference pages.
