Categories

AI Agents

How to use OpenCode with local models on llama.cpp and Ollama

OpenCode uses a local model when you declare an @ai-sdk/openai-compatible provider in opencode.json with the llama-server or Ollama URL and the model's context limit. Version 1.18.31 sends 7,516 tokens on its first turn, so the 4k context Ollama assigns by default without a large GPU is not enough: reserve 32k.

AI Agents

How to build an agent with Dify Agent and a local model

Dify Agent, the Linux-sandbox agent Dify introduced in 1.16, works with a local model once you install the Ollama plugin and switch on tool calling. With Dify 1.17.1 and Qwen3.5-4B on a CPU, the published agent returned correct figures in three out of three runs; Build mode finished none of its three sessions.

How to Install

How to enable Paperless-ngx 3 local AI with Ollama

Paperless-ngx 3 suggests title, tags, correspondent and dates with an LLM and answers questions about your documents, all locally with Ollama. You upgrade from 2.20.15 by changing the image tag. On CPU, Gemma 4 E4B took 25.8 s per suggestion and used about 7 GiB, provided you dodge a bug with reasoning models.

Artificial Intelligence

How to import a Hugging Face model into Ollama 0.34

Since Ollama 0.34.1, ollama create no longer converts or quantizes safetensors weights. To import a Hugging Face model, download it with hf download, convert it to GGUF with llama.cpp's convert_hf_to_gguf.py, quantize it with llama-quantize and build it from a Modelfile. With MiniCPM5-2B on a CPU, the three stages took 35 s, 21 s and 2 s.

Artificial Intelligence

Gemma 4 locally, which size fits your GPU

Gemma 4 ships in five sizes under Apache 2.0. Quantised to 4 bits the weights run from 2.9 GB on the E2B to 17.5 GB on the 31B according to Google's own table, so your memory picks the size. And the 256K window adds another 10 GiB of cache on top.

Artificial Intelligence

Function calling with Ollama on your own machine

Function calling lets a model you run with Ollama on your own machine ask your code to call a function (check the weather, query a database) and use the result to answer. Ollama has supported tools since July 2024; in 2026 models such as qwen3 and llama3.3 do it with reasonable reliability.

Artificial Intelligence

Ollama in 2024: Running LLMs Locally Without Pain

Ollama became the standard for running large language models locally in 2024. It wraps llama.cpp in a single binary with Docker-style CLI and an OpenAI-compatible API. Phi-3 Mini runs in 4 GB; Llama 3.1 8B Q4 needs 6 GB. For production traffic at scale, vLLM remains the correct choice.

Artificial Intelligence

How to Install Ollama on macOS with Apple Silicon

Installing Ollama on an Apple Silicon Mac is as simple as running one Homebrew command. Then pick a model based on available RAM (Phi-3 for 8 GB, Llama 3.1 8B for 16 GB) and expose the local, OpenAI-compatible HTTP API on port 11434 to plug it into your own applications.

Artificial Intelligence

LM Studio: Exploring AI Models from Your Desktop

LM Studio is a desktop app for Mac, Windows, and Linux that downloads and runs large language models on your own machine, with a polished chat interface and no terminal required. It includes an OpenAI-compatible API and RAG with your documents. For individual use it beats Ollama on user experience; for teams or production, OpenWebUI, vLLM, or TGI are the better fit.

Artificial Intelligence

How to Install Ollama to Run LLMs on Your Computer

Ollama makes it trivial to run models like Llama 2 or Mistral on your own computer: one binary, one command, and quantised weights downloading to disk with no compilation required. Covers installation on macOS, Linux, and Windows with an honest look at what local inference can and cannot do compared to frontier models.