Run an LLM on Your Own Laptop with Ollama: A Practical 2026 Guide

By Sheng Pang · Published · 6 min read

For most of the AI boom, using a good model meant sending your text to someone else's server and paying per token. That is still true for the very best models. But open weight models, whose trained weights anyone can download, have improved to the point where a mid range laptop runs an assistant that is genuinely useful for everyday coding, writing and question answering. This guide gets you there in about fifteen minutes using Ollama, and is honest about what you give up.

Why bother

  • Privacy. Nothing leaves the machine. You can paste customer data, unreleased code or medical records without a data processing agreement.
  • Cost. Zero per token. Useful if you are building something that calls a model thousands of times, or just do not want another subscription.
  • Offline. Works on a plane or behind a firewall.
  • Control. The model does not change under you when a provider ships an update. You pin a version and it stays.
  • Learning. Running a model yourself demystifies it. You see the memory it takes, the speed it runs at and how it degrades when starved.

Hardware: what you actually need

The constraint is memory, not processor speed. A model's weights have to fit in fast memory, and roughly speaking a model needs about as many gigabytes as it has billions of parameters when compressed to 4 bits, which is how nearly everyone runs them locally.

Your machineModel size that runs wellWhat that gets you
8 GB RAM, no GPU3 to 4 billion parametersSummaries, simple questions, basic code completion. Noticeably limited.
16 GB RAM or Apple Silicon 16 GB7 to 12 billionA solid everyday assistant. Good at code, decent at reasoning.
32 GB Apple Silicon or 16 GB GPU14 to 30 billionClose to the hosted models of a year or two ago for most tasks.
64 GB or more Apple Silicon, or a 24 GB plus GPU30 to 70 billionCompetitive with mid tier hosted models. Slower to respond.

Apple Silicon Macs are unusually good at this because the CPU and GPU share one pool of memory, so a 32 GB MacBook has 32 GB available for the model. On a Windows or Linux machine the model needs to fit in the graphics card's memory to be fast, and spills to system RAM at a large speed penalty. A machine with no GPU still works, just slowly, at a few tokens per second.

Install and run

Download Ollama from its site or use your package manager, then:

ollama run qwen3:8b

That downloads the model, about five gigabytes, and drops you into a chat. Type a question. The first answer is slower while the model loads into memory; after that it streams at reading speed on a decent laptop. Type /bye to leave.

Some models worth trying as of autumn 2026, all available with a single pull:

  • Qwen3 in 8B, 14B and 30B sizes. The default recommendation for general use and code. Strong at multiple languages.
  • Gemma from Google, in sizes from 4B to 27B. Very good at writing and instruction following for its size.
  • Llama from Meta. Older but with the biggest ecosystem of fine tunes.
  • Coder variants of the above, tuned for code completion and editing. Pair one with an editor extension for a fully local Copilot.

Run ollama list to see what you have, ollama pull to fetch more, ollama rm to delete. Models are just files in a cache directory.

Using it from code

The part that matters for developers: Ollama runs a local server on port 11434 that speaks the same API shape as OpenAI's, so nearly every library and tool that talks to a hosted model can be pointed at it by changing a URL.

curl http://localhost:11434/v1/chat/completions -d '{
  "model": "qwen3:8b",
  "messages": [
    {"role": "system", "content": "You reply only with valid JSON."},
    {"role": "user", "content": "Give me three fruit names and their colours."}
  ],
  "response_format": {"type": "json_object"}
}'

The response is a JSON document with the model's reply inside it. Set the base URL in the OpenAI Python or JavaScript client to http://localhost:11434/v1 and your existing code works unchanged, with a throwaway API key. Ollama also has its own native endpoint with extra controls like context length and temperature, and a structured output mode where you pass a JSON Schema and the model is constrained to match it, which we cover in how to get reliable JSON from an LLM.

Coding agents and editor extensions increasingly let you point them at a local model too. Expect a real drop in quality compared to the frontier hosted models for agentic tasks, but for autocomplete and quick questions a 14B coder model is very usable.

Where local models fall short

Be realistic about the gap:

  • Reasoning depth. A 8B model is not going to debug a subtle concurrency issue or plan a multi step refactor reliably. The frontier hosted models are still a tier above.
  • Knowledge. Fewer parameters means fewer facts stored. Local models hallucinate more on specifics. Give them the documents they need rather than relying on recall.
  • Context length. Long contexts eat memory fast. A model that runs well with a 4,000 token window may not fit with 32,000.
  • Speed on big models. Anything above 30B on a laptop is slow enough to change how you work with it.
  • Battery and heat. Running a model pins the GPU. A laptop on battery will notice.

A sensible split that many developers have settled on: local model for private data, quick questions, autocomplete and anything you call in a loop; hosted frontier model for hard problems and agentic work.

Beyond Ollama

Ollama is the easy path. LM Studio offers a graphical interface with the same models. llama.cpp is the engine underneath most of these and gives you every knob if you want them. For serving many users at once, vLLM on a proper GPU server is the standard. But for one developer on one laptop, Ollama plus a Qwen or Gemma model is the answer in 2026, and the whole setup takes less time than reading this post did.

If you want to understand what those billions of parameters are and why the 4 bit compression works, our post on how neural networks learn explains what a weight is and the one on tokens explains why the context window is measured the way it is.

← Back to all articles