Learning path Intermediate
Local LLMs: run models on your own hardware
Run language models on your own hardware, no API and no quotas: from Ollama to serving big models on a Mac, plus vector memory for RAG.
- 10 resources
- 7 views
- ~142 min
Run language models on your own machine — no API keys, no quotas and your data never leaves home. The path goes from installing Ollama to fine-tuning on an Apple Silicon Mac, and ends by giving your model memory with a vector database for RAG.
How to Install Ollama to Run LLMs on Your Computer
Ollama makes it trivial to run models like Llama 2 or Mistral on your own computer: one binary, one command, and quantised weights downloading to disk with no compilation required. Covers installation on macOS, Linux, and Windows with an honest look at what local inference can and cannot do compared to frontier models.
How to Install Ollama on macOS with Apple Silicon
Installing Ollama on an Apple Silicon Mac is as simple as running one Homebrew command. Then pick a model based on available RAM (Phi-3 for 8 GB, Llama 3.1 8B for 16 GB) and expose the local, OpenAI-compatible HTTP API on port 11434 to plug it into your own applications.
How to install llama.cpp on Linux, macOS and Docker
There are four ways to install llama.cpp: the llama.app script, which drops a 15 MB llama binary into ~/.local/bin; Homebrew; the server-v0.4.1 Docker image; or CMake. I tested them on 14 September 2026 on Linux arm64 with Gemma 4 E2B, in the terminal with llama cli and as an OpenAI-compatible API.
Gemma 4 locally, which size fits your GPU
Gemma 4 ships in five sizes under Apache 2.0. Quantised to 4 bits the weights run from 2.9 GB on the E2B to 17.5 GB on the 31B according to Google's own table, so your memory picks the size. And the 256K window adds another 10 GiB of cache on top.
Local multimodal models: the VRAM arithmetic on a 16 GB card
Qwen3-VL 8B at Q4_K_M is 4.68 GiB of weights plus a 1.08 GiB vision tower, and its KV cache costs 144 KiB per token. At 8,192 tokens of context the total is around 7 GiB and fits comfortably in 16 GB; at the native 256K it would ask for 36 GiB of cache alone and fits on no consumer card.
Maple-Preview, MiniCPM5-2B or Spark-X2.5, which CPU-only LLM to pick for an arm64 machine
Without a GPU, on an 18-core arm64 CPU with llama.cpp v0.4.1, Maple-Preview generated 172.84 tokens/s on 6 threads: twice MiniCPM5-2B and 3.6 times Spark-X2.5-4B. The trade-off is 5.79 GiB of RAM and reasoning it cannot switch off. MiniCPM5 fits in 3.27 GiB and Spark writes the best Spanish. All three made facts up.
What is oMLX and how it differs from MLX, Ollama and LM Studio
oMLX is a local inference server for Apple Silicon Macs that wraps Apple's MLX framework in a web process and exposes it through the OpenAI and Anthropic APIs. It adds continuous batching, on-disk KV caching and several models in memory at once, all driven from the menu bar.
How to install and tune oMLX on M5 Max 128 GB
Recipe for oMLX 0.6.4 on a Mac M5 Max with 128 GB: install, Claude Code and advanced settings with screenshots (Lightning MTP, VLM MTP and DFlash). Includes our own September 2026 benchmarks, with up to 1.88x more tokens per second.
How to Install PostgreSQL with pgvector Step by Step
This guide installs PostgreSQL 18.6 with pgvector 0.8.6 on Debian or Ubuntu from the official PGDG repository, or in Docker. It creates a dedicated role and database, tunes memory for production, and explains when HNSW beats IVFFlat, with recall and build times measured on 99,000 embeddings.
Local LLM Calculator
Will that model fit your GPU? Is self-hosting cheaper than the API? Estimate the VRAM and the cost break-even point.