Back to RAG & Vector Databases

Run a Local LLM with Ollama and Hook It Up to a Simple RAG Pipeline

Quickly spin up Ollama locally, load a model, and connect it to a vector store for retrieval‑augmented generation.

T

Trendzza Research Desk

Oct 1, 2026 · 1 min read

Research tools helped prepare this thread; a council editor is responsible for what was published. Last checked Oct 1, 2026.

Run a local LLM with Ollama and add RAG in three steps.
1️⃣ Install Ollama and pull a model → curl -fsSL https://ollama.com/install.sh | sh && ollama pull llama3.1. The server starts on localhost:11434.
2️⃣ Create a vector store (e.g., using FAISS) and index your docs → `python - <<'PY'
from langchain.document_loaders import TextLoader
from langchain.embeddings import OllamaEmbeddings
from langchain.vectorstores import FAISS
loader = TextLoader('data.txt')
docs = loader.load_and_split()
emb = OllamaEmbeddings(base_url='http://localhost:11434')
index = FAISS.from_documents(docs, emb)
index.save_local('faiss_index')
PY`
3️⃣ Query with the local model → `python - <<'PY'
from langchain.chains import RetrievalQA
from langchain.llms import Ollama
from langchain.vectorstores import FAISS
index = FAISS.load_local('faiss_index', OllamaEmbeddings(base_url='http://localhost:11434'))
qa = RetrievalQA.from_chain_type(llm=Ollama(base_url='http://localhost:11434'), retriever=index.as_retriever())
print(qa.run('What is the main benefit of RAG?'))
PY`
🔧 Gotcha: Ollama’s default context window may be smaller than some models; chunk your documents to ≤ 500 tokens to avoid truncation.

Read the evidence

Sources used in this thread

Open the original material, compare the claims, and form your own view.

Community notes

Add context, not noise (0)

Corrections, lived experience, useful examples, and better sources belong here.

Nothing added yet. Be the first to make this thread more useful.

Sign in to join the council thread