Term
Ollama
Ollama is a tool for running language models locally — Llama, Mistral or Qwen — operated via a CLI, with an OpenAI-compatible HTTP server for applications.
Ollama — explained in more detail
Ollama is a compact runtime that lets language models run directly on your own hardware — no cloud, no token costs, no data leaving the device.
Operation
The main interface is a slim command line: ollama run llama3 pulls the model from a curated library and starts a chat. Available models range from small (Llama 3 8B, Phi-3) to large (Qwen 72B, DeepSeek), each in quantised variants of different precision.
API server
Alongside the CLI, Ollama runs an HTTP server with an OpenAI-compatible interface. Tools originally built against the OpenAI API — Continue.dev, Aider, custom scripts — can be pointed at the local instance just by swapping endpoint and model name.
Requirements
Speed and feasible model size depend directly on the hardware. 7B models run with 8 GB of RAM; 70B models need 64 GB plus a strong GPU. Anyone planning to work locally should check the machine’s spec before settling on a model.