Tested with smolagents 1.26.0 · LiteLLM 1.101.0 · Ollama 0.34.1 · Qwen3-4B-Instruct-2507 · Python 3.12 · verified

Updated: 2026-09-16

smolagents is Hugging Face’s agent library, and its core idea: an agent reasons better when it writes its actions as Python code instead of filling in JSON objects. With its agent logic in a single file and support for almost any model, it is one of the most direct ways to stand up an agent that actually does things. This guide explains what it is, how its CodeAgent works and when to pick it over a bigger framework. We ran its examples on 16 September 2026 with smolagents 1.26.0 and a local model in Ollama, and the same explanation is available in Spanish.

Key takeaways

  • smolagents is an open-source library (Apache-2.0 license); version 1.26.0, released on 29 May 2026, was still the latest on 16 September 2026.
  • Its CodeAgent writes actions as executable Python code, which according to Hugging Face cuts steps (and therefore model calls) by around 30% compared with JSON-based tool calling.
  • It is model-agnostic: it works with Hugging Face inference providers, with OpenAI or Anthropic via LiteLLM, and with models you run on your own machine through Transformers or Ollama (a free engine that serves open-source language models on your own hardware, with no calls to an external API).
  • Tools can come from a @tool decorator, from LangChain, from an MCP server or from a Hub Space.
  • To run the generated code safely it offers sandboxed environments with E2B, Docker, Modal or Blaxel; the local interpreter it uses by default is not a security boundary.
  • With a 4-billion-parameter model in Ollama and no GPU, this guide’s search agent was right in 13 of 14 runs.

What is smolagents?

smolagents is a Python library built by Hugging Face to create and run AI agents with very little code. Its tagline is "Agents that think in code!", and its philosophy is to keep abstractions minimal above raw code. All the agent logic lives in one file, agents.py, instead of being spread across layers of classes as in other frameworks.

The documentation says that logic fits in about a thousand lines, but in version 1.26.0 agents.py has 1,813. Without blank lines, comments and docstrings, 1,288 remain, by our count.

The project is open (Apache-2.0) and had 29,350 GitHub stars on 16 September 2026. Its latest release, 1.26.0, removed the WebAssembly-based remote executor. If you come from wiring an agent loop by hand, as in the Anthropic SDK tutorial, smolagents spares you that boilerplate without hiding what happens inside.

Installing it is one line, with Python 3.10 or later:

pip install 'smolagents[toolkit,litellm]'

The toolkit extra adds default tools, such as web search, and litellm is the one LiteLLMModel needs: without it, creating the model fails with ModuleNotFoundError. OpenAIModel needs the openai extra, and MCP servers need mcp.

Agents that write code (CodeAgent)

The distinctive piece of smolagents is the CodeAgent. At each step of the reason-act-observe loop, the model does not return JSON with a tool name and its arguments; it returns a block of Python code. That block runs and its result flows back to the model as the next observation. Calling a tool is invoking a Python function.

The advantage is composability. In code you can nest calls, use loops and conditionals, store a result in a variable and reuse it, all in a single action. With JSON you would need one step and one model call per operation.

Hugging Face argues this approach yields around 30% fewer steps and better performance on hard tasks. The idea comes from the paper "Executable Code Actions Elicit Better LLM Agents" (CodeAct), which measured up to a 20% higher success rate when acting with code instead of structured text.

If you prefer the classic paradigm, there is also ToolCallingAgent, which uses the usual JSON tool calling. It suits cases where the model has well-tuned function calling or where your tools do one thing and need no chaining.

Tools and models (local or API)

smolagents is model-agnostic: it ties you to no provider. You pick the model class based on where you want to run inference (there are more, such as those for Azure OpenAI, Amazon Bedrock, vLLM and MLX):

  • InferenceClientModel: uses the inference providers on the Hugging Face Hub and is the quickstart option. It needs an HF_TOKEN; in 1.26.0 its default model is Qwen/Qwen3-Next-80B-A3B-Thinking.
  • LiteLLMModel: connects to OpenAI, Anthropic, Gemini and over a hundred models and providers through LiteLLM. It also serves Ollama with the ollama_chat/ prefix in the model name.
  • TransformersModel: loads an open model and runs it on your own machine with the transformers library.
  • OpenAIModel: points at any endpoint compatible with the OpenAI API, including a local server. It replaces OpenAIServerModel, which remains available as an alias.

For a model you run with Ollama on your own machine, point LiteLLMModel at http://localhost:11434, or OpenAIModel at http://localhost:11434/v1. With LiteLLM, set num_ctx=8192, because the smolagents guide warns that Ollama’s default context is too short. In our tests, the agent’s second step already sent over 4,200 tokens. To pick a model, read our comparison of open models with tool calling.

Tools are just as flexible: you define a function and add the @tool decorator, or import a collection from an MCP server with ToolCollection.from_mcp, from LangChain or from a Hub Space. MCP has a catch: on 16 September 2026, the mcp extra installed version 2.2.0 of the mcp package, and with it from_mcp fails with an ImportError. Pin the 1.x line:

pip install 'smolagents[mcp]' 'mcp[ws]<2'

With that, a local MCP server loaded its tools in our test. Also pass trust_remote_code=True: without it, from_mcp raises a ValueError.

A minimal example

This is a complete agent. It creates a CodeAgent with WebSearchTool, the search tool in the current quickstart, and a local model. First, pull Qwen3-4B-Instruct quantized to Q4_K_M (2.5 GB) into Ollama:

ollama pull hf.co/unsloth/Qwen3-4B-Instruct-2507-GGUF:Q4_K_M

You give it a task and the agent writes and runs the Python needed to solve it:

from smolagents import CodeAgent, LiteLLMModel, WebSearchTool

model = LiteLLMModel(
    model_id="ollama_chat/hf.co/unsloth/Qwen3-4B-Instruct-2507-GGUF:Q4_K_M",
    api_base="http://localhost:11434",
    num_ctx=8192,
)
agent = CodeAgent(tools=[WebSearchTool()], model=model)

result = agent.run(
    "Find Mount Teide's height and tell me how many times it fits in Everest."
)
print(result)

With a Hugging Face account, swap the model for InferenceClientModel() and export HF_TOKEN; we did not run that variant. In our local run, the first step called web_search once per mountain. In the second, the model copied the heights it had read and finished (log excerpt):

  teide_height_m = 3718  # in meters
  everest_height_m = 8848.86

  times_fit = everest_height_m / teide_height_m
  final_answer(times_fit)
Final answer: 2.3800053792361484
[Step 2: Duration 27.55 seconds| Input tokens: 6,297 | Output tokens: 252]

We repeated the task 14 times across Spanish and English, and 13 runs returned 2.38, with 3,715 or 3,718 m for Teide depending on the result the model read. In the other, the model treated the search text as a list and the agent returned 0 without any error.

With Ollama capped at 8 threads (PARAMETER num_thread 8) on a shared 18-core arm64 machine with no GPU and a load average of 8 to 10, each correct run took 33 to 58 s. With the default thread count and the load at 23, the first run took 16 minutes.

Defining your own tool is just as direct. smolagents turns the docstring into the description the model reads, and if an argument has no line under Args, the decorator raises a DocstringParsingException:

from smolagents import tool

@tool
def price_with_vat(price: float, rate: float = 21.0) -> float:
    """Compute the final price with VAT.

    Args:
        price: base price in euros.
        rate: VAT percentage, for example 21 for 21%.
    """
    return round(price * (1 + rate / 100), 2)

You pass the function in the tools list and the agent can call it inside its code, even three times in the same action:

agent = CodeAgent(tools=[price_with_vat], model=model)
print(agent.run(
    "Add up with VAT a 899-euro laptop and a 45.50-euro mouse at 21% "
    "and a 12-euro book at 4%."
))

In the 20 runs we made (10 per language), the model called the tool three times in a single step and returned 1155.32, in 5 to 12 s. With an earlier, vaguer prompt it passed 0.21 instead of 21 and returned 958.49 without any error, so check which arguments your tool receives.

smolagents versus bigger frameworks

The usual question is when to pick smolagents over LangGraph, LlamaIndex or OpenAI’s Agents SDK. The short answer: smolagents shines when you want an agent that acts by writing code and prefer little ceremony. Being so small, you can read it in full and debug it; there are no state graphs or orchestration layers to learn.

Bigger frameworks win when you need workflows with explicit state, complex conditional branches, long-running persistence or heavily structured multi-agent systems. smolagents also does multi-agent (one agent can manage others), but its natural ground is autonomous agents that solve a task by reasoning with code. Always run generated code in a sandbox, such as the E2B sandbox, never straight on your server.

Frequently asked questions

Why does an agent write code instead of JSON?

Because code is more expressive. In a single block it can chain one tool after another, use loops and store intermediate results in variables, which in JSON would demand one step and one model call per operation. According to Hugging Face, acting with code cuts steps by around 30% and improves results on complex tasks.

Do I need a paid API key to use smolagents?

Not necessarily. InferenceClientModel needs an HF_TOKEN, and a free Hugging Face account includes $0.10 of monthly credit for its inference providers. With TransformersModel or Ollama you run an open model on your own machine with no per-token cost. The LiteLLMModel and OpenAIModel classes let you move to a paid provider when you need to.

Is it safe for an agent to run the code it generates?

Only if you isolate it. smolagents integrates secure execution environments with E2B, Docker, Modal or Blaxel, which you pick with the executor_type parameter. Its default local interpreter, LocalPythonExecutor, is not a security boundary according to its own documentation; always use a sandbox in any real deployment.

Does smolagents work with a small local model?

Yes, with caveats. With Qwen3-4B-Instruct in Ollama 0.34.1 and no GPU, the search example was right in 13 of 14 runs and the VAT tool example in 20 of 20. The one failure was silent, a 0 with no error, so review the agent’s code and validate its answer before you use it.

Conclusion

smolagents bets on a concrete idea: an agent performs better when it expresses its actions as Python code. That decision, plus a core that fits in one file and support for any model, makes it a practical entry point into agent development. Start with the CodeAgent from the example, give it a useful tool, check its answers and, when you move to production, wrap execution in a sandbox. To back it with an open model, the natural next step is to learn how to install Ollama on your own machine.

Sources

  1. smolagents documentation
  2. smolagents on GitHub
  3. Introducing smolagents (Hugging Face blog)
  4. Executable Code Actions Elicit Better LLM Agents
  5. smolagents 1.26.0 release notes
  6. smolagents guided tour
  7. Secure code execution in smolagents
  8. Inference Providers pricing

Route: Frameworks to Build AI Agents