How it works
After installing Ollama on macOS, Windows or Linux, a command such as 'ollama run' followed by a model name downloads that model from its library and starts a chat. In the background it runs a server on your machine (port 11434) with its own REST API and an OpenAI-compatible one, so apps and editors can point at it instead of a cloud provider. Models are stored in compressed (quantised) form so that smaller ones fit in an ordinary laptop's memory.
Speed and quality depend on the hardware: small models run on a typical laptop, while large ones need a strong GPU or plenty of unified memory. Nothing leaves the machine and there is no per-token bill, which suits private data and offline use. For models too big to run at home, Ollama also sells a cloud service that runs larger open models with the same commands and API.
Ollama pros and cons
Pros
- Free, open source and quick to set up
- Private: prompts and data never leave your machine
- Works offline, with no per-token costs
- OpenAI-compatible local API that many tools already support
Cons
- Open models that fit on a laptop trail the largest cloud models
- Needs plenty of memory, and a good GPU for speed
- You manage updates, model choice and any server hosting yourself
When to use Ollama
Pick it when
- Private documents that must not go to a cloud provider
- Offline tools, local coding helpers and experiments
- Development and tests without paying for tokens
Skip it when
- Many users need fast answers from a large model (use a cloud API)
- Your users' devices are phones or low-powered laptops
Ollama pricing
Open source
Free (MIT) to run locally. Ollama's cloud has a free plan with starter credits, Pro at $20 a month and Max at $100 a month, each with monthly usage credits.
Ollama pricing page (opens in a new tab)Approximate, checked September 2026.What the other tools cost
Ollama vs the alternatives
Related terms
More in AI and LLMs
Run it yourself