Categories

AI Agents

How to use Shieldstral as a local guardrail for your agent

Shieldstral is the 3B safety classifier Mistral released in August 2026: it answers yes or no to a policy written in plain language. Converted to GGUF Q8_0 with llama.cpp and served on a CPU, it blocked no legitimate one among the 100 Spanish and English prompts I prepared, but let 8 of 17 injections through.

AI Agents

How to use OpenCode with local models on llama.cpp and Ollama

OpenCode uses a local model when you declare an @ai-sdk/openai-compatible provider in opencode.json with the llama-server or Ollama URL and the model's context limit. Version 1.18.31 sends 7,516 tokens on its first turn, so the 4k context Ollama assigns by default without a large GPU is not enough: reserve 32k.

AI Agents

How to use DeepSeek Harness with a local model

DeepSeek Harness (dsh) is the open-source agent harness DeepSeek released in August 2026. It works with a local model if you declare an openai-completions provider in settings.yaml that points at llama-server, with a placeholder key and the real context size. With Qwen3.5-4B on CPU it left the tests green in 4 of 5 attempts, but slowly.

How to Install

How to install OpenClaw with Docker and a local model

OpenClaw 2026.9.4 installs with Docker Compose from the official GHCR image, which ships for amd64 and arm64. To run it without a cloud API, connect it to a llama-server on the compose file's internal network. Publish port 18789 on 127.0.0.1 only, because the repository's compose file opens it on every interface.

Artificial Intelligence

How to import a Hugging Face model into Ollama 0.34

Since Ollama 0.34.1, ollama create no longer converts or quantizes safetensors weights. To import a Hugging Face model, download it with hf download, convert it to GGUF with llama.cpp's convert_hf_to_gguf.py, quantize it with llama-quantize and build it from a Modelfile. With MiniCPM5-2B on a CPU, the three stages took 35 s, 21 s and 2 s.

How to Install

How to install llama.cpp on Linux, macOS and Docker

There are four ways to install llama.cpp: the llama.app script, which drops a 15 MB llama binary into ~/.local/bin; Homebrew; the server-v0.4.1 Docker image; or CMake. I tested them on 14 September 2026 on Linux arm64 with Gemma 4 E2B, in the terminal with llama cli and as an OpenAI-compatible API.

Artificial Intelligence

llama.cpp: Optimisations That Keep Surprising

llama.cpp is the C++ library that powers Ollama and much of the local-LLM ecosystem. 2024 added speculative decoding with two- to three-fold speedups, an RPC server for sharding layers across machines, and a stable GGUF format. Ollama covers 90% of cases; going direct pays off with uncommon hardware or specific flags.

Artificial Intelligence

Ollama in 2024: Running LLMs Locally Without Pain

Ollama became the standard for running large language models locally in 2024. It wraps llama.cpp in a single binary with Docker-style CLI and an OpenAI-compatible API. Phi-3 Mini runs in 4 GB; Llama 3.1 8B Q4 needs 6 GB. For production traffic at scale, vLLM remains the correct choice.

Artificial Intelligence

Model Quantization and llama.cpp on Your Laptop

With quantization, model weights are stored with fewer bits (4, 5, or 8 instead of 16), so Llama 2 13B shrinks from 26 GB to about 7.5 GB. With llama.cpp it runs on an ordinary 16GB-RAM laptop with no dedicated GPU, and the quality loss is smaller than intuition suggests.