Categories

How to Install

How to enable Paperless-ngx 3 local AI with Ollama

Paperless-ngx 3 suggests title, tags, correspondent and dates with an LLM and answers questions about your documents, all locally with Ollama. You upgrade from 2.20.15 by changing the image tag. On CPU, Gemma 4 E4B took 25.8 s per suggestion and used about 7 GiB, provided you dodge a bug with reasoning models.

Artificial Intelligence

How to import a Hugging Face model into Ollama 0.34

Since Ollama 0.34.1, ollama create no longer converts or quantizes safetensors weights. To import a Hugging Face model, download it with hf download, convert it to GGUF with llama.cpp's convert_hf_to_gguf.py, quantize it with llama-quantize and build it from a Modelfile. With MiniCPM5-2B on a CPU, the three stages took 35 s, 21 s and 2 s.

How to Install

How to install llama.cpp on Linux, macOS and Docker

There are four ways to install llama.cpp: the llama.app script, which drops a 15 MB llama binary into ~/.local/bin; Homebrew; the server-v0.4.1 Docker image; or CMake. I tested them on 14 September 2026 on Linux arm64 with Gemma 4 E2B, in the terminal with llama cli and as an OpenAI-compatible API.

Artificial Intelligence

Llama 3.2 at the edge: Meta bets on small

Meta released Llama 3.2 with models as small as 1B and 3B parameters, built specifically to run on the device rather than in a datacentre. A look at what they can genuinely do, where the small sizes fall short, and how they compare with the alternatives.

Artificial Intelligence

NPU in the PC: faster, cheaper local AI

Qualcomm, Intel and AMD Copilot+ processors have normalised the presence of an NPU in everyday PCs. A 40 TOPS NPU can run quantised Phi-3 Mini drawing just 5-10 W, versus 40-50 W for a laptop GPU doing the same task. What actually changes for running AI models locally, and when it is worth it.