Langfuse: self-hosted agent observability
Table of contents
- Key takeaways
- What is Langfuse?
- Deploying it with Docker Compose
- Traces, spans and generations
- Instrumenting an agent with OpenTelemetry
- Evaluations and datasets
- Frequently asked questions
- Is Langfuse really free and open source?
- What is the difference between Langfuse and OpenTelemetry?
- Is Langfuse v3 still supported?
- Can I use Langfuse with local models?
- Conclusion
- Sources
Tested with Langfuse 4.37.0 · ClickHouse 25.12 · PostgreSQL 17 · Python SDK 4.15.4 · verified
Updated: 2026-09-16
Langfuse is an open-source platform to observe, debug and evaluate AI applications and agents. You can self-host it with Docker Compose on Postgres, ClickHouse, Redis and S3 storage, and its Python SDK, built on OpenTelemetry, captures traces, spans and generations with their cost and latency. This guide explains how to deploy it and instrument an agent.
Langfuse is an open-source platform that records everything your AI agent does: every model call, its cost, its latency and its output. You can debug and evaluate it with data instead of guesswork. The best part is that you can self-host the whole thing with Docker Compose, so your traces never leave your server.
In this guide you will see what Langfuse is and how to deploy version 4: v4.37.0, which we installed and tested on 16 September 2026. You will also see its data model (traces, spans and generations), how to instrument an agent with OpenTelemetry and how to set up evaluations with datasets. The same explanation is available in Spanish.
Key takeaways
- Langfuse is an open-source LLM engineering platform (MIT licence, except the
ee/folders of enterprise features) with over 34,600 stars on GitHub. Its stable branch is v4, released on 29 July 2026; the latest release, v4.37.0, is from 16 September 2026. - It is self-hostable: a single
docker compose upbrings up the web UI on port 3000 with four data stores behind it, so your traces stay in your own infrastructure. v4 requires ClickHouse 25.12 or later. - Its data model is observations of type span (a unit of work) and generation (a model call, with its tokens, cost and latency), grouped into traces; in v4 each trace is represented by its root observation.
- Instrumenting an agent takes one decorator: the Python SDK v4 (4.15.4 as of 16 September 2026) is built on OpenTelemetry, the open tracing standard, and ships direct integrations for OpenAI, LangChain, LiteLLM and LlamaIndex.
- It goes beyond logging: it manages prompts, stores test datasets and runs evaluations (including an LLM judge) to compare versions of your agent with objective metrics.
What is Langfuse?
Langfuse is an observability and evaluation platform for applications that use language models. Its own documentation defines it as "an open source LLM engineering platform". It goes beyond system metrics like CPU or memory and understands the concepts specific to an AI agent: the prompts, model responses, tokens consumed, dollar cost and the tools it calls.
The problem it solves is concrete. When an agent built with the Anthropic SDK or a LangGraph graph misbehaves, a print to the console is useless. You cannot see the exact prompt the model received, how many iterations the loop ran or how much each attempt cost you. Langfuse records that full execution tree and shows it to you on a navigable timeline.
The project started at the Y Combinator accelerator (Winter 2023 batch). In January 2026 it announced that ClickHouse had acquired it, with no immediate changes for users and a commitment to open source and self-hosting. Its code is free under the MIT licence, with the sole exception of the ee/ folders, which group enterprise features gated behind a licence key. Everything you need to observe and evaluate agents is in the open part.
Deploying it with Docker Compose
Langfuse’s big advantage over a cloud service is that you can run it entirely on your own machine, so neither your users’ prompts nor their responses ever leave your network. You need Git and Docker with Compose; for a virtual machine, the documentation recommends at least 4 cores and 16 GiB of memory. Clone the repository at the tag of the release you are deploying:
git clone --depth 1 --branch v4.37.0 \
https://github.com/langfuse/langfuse.git
cd langfuse
The docker-compose.yml marks every secret you must change before starting with # CHANGEME (grep -n CHANGEME docker-compose.yml lists them). All of them are read from environment variables, so you do not need to edit the YAML. Compose reads a .env file next to it, as we explain in environment variables and secrets in Docker Compose. This block generates those values with openssl:
PG=$(openssl rand -hex 16); CH=$(openssl rand -hex 16)
MINIO=$(openssl rand -hex 16); REDIS=$(openssl rand -hex 16)
cat > .env <<EOF
NEXTAUTH_SECRET=$(openssl rand -hex 32)
SALT=$(openssl rand -hex 32)
ENCRYPTION_KEY=$(openssl rand -hex 32)
POSTGRES_PASSWORD=$PG
DATABASE_URL=postgresql://postgres:$PG@postgres:5432/postgres
CLICKHOUSE_PASSWORD=$CH
MINIO_ROOT_PASSWORD=$MINIO
LANGFUSE_S3_EVENT_UPLOAD_SECRET_ACCESS_KEY=$MINIO
LANGFUSE_S3_MEDIA_UPLOAD_SECRET_ACCESS_KEY=$MINIO
LANGFUSE_S3_BATCH_EXPORT_SECRET_ACCESS_KEY=$MINIO
REDIS_AUTH=$REDIS
EOF
The images live on Docker Hub (langfuse/langfuse and langfuse/langfuse-worker), but the compose file requests them from docker.langfuse.com, a registry endpoint run by Reo.dev, Langfuse’s analytics provider, which counts each pull. Tags and digests are identical either way. On our test machine that domain resolved to 0.0.0.0 and docker pull failed, so we pulled from Docker Hub by dropping the prefix.
The compose file uses the floating :4 tag. On 16 September 2026, latest, 4 and 4.37.0 pointed to the same digest, and 3 to v3.225.8. To pin the exact release, add a docker-compose.override.yml, which Compose merges with the main file:
services:
langfuse-web:
image: langfuse/langfuse:4.37.0
langfuse-worker:
image: langfuse/langfuse-worker:4.37.0
The web UI (3000) and the MinIO S3 API (9090) listen on all interfaces, and the other ports (3030, 5432, 6379, 8123, 9000 and 9091) on 127.0.0.1 only. On the test machine another service held port 9000, so we remapped the ports with ports: !override lists in the override. If you change 3000 or 9090, also update NEXTAUTH_URL and LANGFUSE_S3_MEDIA_UPLOAD_ENDPOINT in the .env. Start the stack and check that the web UI answers:
docker compose up -d
curl -s http://localhost:3000/api/public/health
The response (on our remapped port) confirms the deployed release:
{"status":"OK","version":"4.37.0"}
Behind that single command, four data stores come up, each with its own job. Postgres holds transactional data (users, projects, prompts) and ClickHouse is the analytical database that absorbs the high volume of traces. Redis acts as a cache and event queue, and an S3-compatible store (by default a MinIO container built from the Chainguard image) keeps the raw events and attachments. v4 requires at least ClickHouse 25.12 (26.4 recommended), Postgres 15 and Redis 7.0, and the compose file ships ClickHouse 25.12, Postgres 17 and Redis 7.
With the images already downloaded, the web container ran the migrations and logged Ready 10 s after docker compose up -d, on an 18-core arm64 box with a load average of 2.9. A second clean install, with the load at 42, took 14 s (the documentation allows 2 to 3 minutes). Freshly started, the six containers used between 2.3 and 2.6 GiB of memory in total.
On that second install, the worker tried to load the model prices before the web container had finished the migrations, and generations arrived without a cost. Because it only tries at startup, check its log and restart it if you see Error upserting default model prices:
docker compose logs langfuse-worker | grep "model prices"
docker compose restart langfuse-worker
After the restart, the log showed Finished upserting default model prices and new traces carried their cost; earlier ones stayed without it.
The docker compose that ships with the repository is meant for testing and small deployments: it has no high availability, horizontal scaling or backups (that is what the Kubernetes chart is for). With the containers up, open http://localhost:3000, create the first account and a project, and note the API key pair (pk-lf-… and sk-lf-…) you will use to send traces. You can also create them at startup with the LANGFUSE_INIT_* variables of headless initialization, which is what we did in the test.
If you already run v3, pin it to the 3 tag and avoid latest, which now pulls v4. The migration requires upgrading ClickHouse first and backing up Postgres and ClickHouse, because there is no automatic way back; we cover it in how to migrate Langfuse from v3 to v4.
Traces, spans and generations
To get the most out of Langfuse it helps to be clear about its data model, which v4 has simplified. Every step of your application is an observation. A trace groups the observations that share the same trace_id and represents one complete request, for example "the user asked X and the agent replied Y". These are the three foundational observation types (Langfuse also ships more specialised ones, such as agent, tool, retriever or guardrail, for tracing specific pipeline steps):
- Span: a unit of work with a duration, such as a retrieval step against a vector database or a tool execution. Spans nest to form the execution tree.
- Generation: a special kind of span for model calls. It captures the model used, the input prompt, the response, the input and output tokens, the computed cost and the latency.
- Event: a single point in time, with no duration, to mark milestones.
In v4, Langfuse stores each observation in a wide ClickHouse table (events_full) that repeats its trace’s attributes, such as user, session or tags, on every row. A trace is no longer a separate entity: its root observation represents it. The UI replaces the traces table with a single Observations view, filtered to root observations by default so it shows one row per trace. On our fresh install, ClickHouse already had the events_full and events_core tables.
Above traces, sessions group the traces from one conversation, and scores attach a quality value, numeric, categorical or boolean, which may come from a user, a rule or a judge model. With those four pieces (traces, observations, sessions and scores) you can describe any AI application, from a simple chatbot to a multi-agent system in the style of CrewAI.
Instrumenting an agent with OpenTelemetry
Langfuse’s Python SDK is on v4 (4.15.4 was released on 16 September 2026) and is built on OpenTelemetry, the industry’s open standard for tracing. It does not lock you into a proprietary format. Against a v4 server, SDK v3 is deprecated, and real-time data needs 4.7.0 or later. One command installs it, together with the OpenAI client the example uses:
pip install langfuse==4.15.4 openai
The most direct way to instrument code is the @observe() decorator: wrap any function and Langfuse automatically creates a span with its input arguments and its return value. For model calls there is an even better shortcut: importing the OpenAI client from langfuse.openai turns every request into a full generation without touching anything else.
from langfuse import get_client, observe
from langfuse.openai import openai # already-instrumented OpenAI client
@observe()
def answer(question: str) -> str:
# This call logs itself as a generation,
# with model, tokens, cost and latency included.
response = openai.chat.completions.create(
model="gpt-5.6-luna",
messages=[{"role": "user", "content": question}],
)
return response.choices[0].message.content
print(answer("Summarise what agent observability is"))
get_client().flush()
Set LANGFUSE_PUBLIC_KEY, LANGFUSE_SECRET_KEY and LANGFUSE_BASE_URL (your http://localhost:3000), plus OPENAI_API_KEY. The older LANGFUSE_HOST name is marked as deprecated, although 4.15.4 still reads it. The final flush() sends pending traces before a short script exits.
We tested it against a local server that mimics the OpenAI API, without calling OpenAI. The Observations API v2 returned the answer span as the root and the generation with the gpt-5.6-luna model, its tokens and its cost, computed from the built-in price table. Old endpoints such as GET /api/public/traces returned 404 on this fresh install, so query v2:
curl -s -u "$LANGFUSE_PUBLIC_KEY:$LANGFUSE_SECRET_KEY" \
"$LANGFUSE_BASE_URL/api/public/v2/observations?fields=core,basic,model,usage"
If you already used OpenTelemetry, SDK v4 no longer exports every span: by default it sends Langfuse spans, spans with gen_ai.* attributes (the OpenTelemetry GenAI semantic conventions) and spans from known LLM libraries. To get HTTP or database spans back, create the client with Langfuse(should_export_span=lambda span: True).
Evaluations and datasets
Observing is the first step; the second is measuring whether your agent gets better or worse when you change a prompt or a model. Langfuse solves this with two tools that work together. A dataset is a collection of test cases, each with an input and, optionally, the expected output. You can build it by hand or promote real traces you care about into it from the UI.
Against that dataset you run an experiment: your agent processes each case and Langfuse saves each result as a trace linked to the dataset. Then you apply evaluators that assign scores. They can be deterministic checks (code evaluators in Python or TypeScript) or an LLM judge, that is, another model that rates correctness or tone against a rubric. In v4, evaluators work on observations: trace-level ones show up as Legacy and stop running once the install moves to the default write mode.
The UI compares two runs side by side, so you can see at a glance whether the new prompt version raises or lowers the average score. It is the same rigour you would apply to serving a model with vLLM in production, but applied to answer quality instead of server throughput.
Frequently asked questions
Is Langfuse really free and open source?
Yes. The bulk of Langfuse, including all of the observability, prompt management, datasets and evaluations, is free software under the MIT licence, and you can self-host it at no cost. The only exception is the repository’s ee/ folders (ee/, web/src/ee/ and worker/src/ee/), which group enterprise features such as project-level roles, audit logs or data retention policies, gated behind a paid licence key. You do not need it to observe and evaluate agents.
What is the difference between Langfuse and OpenTelemetry?
OpenTelemetry is an open standard and a set of libraries to generate and transport traces, but it includes neither storage nor a UI. Langfuse uses OpenTelemetry as the base of its SDK and adds what is missing: a specialised database (ClickHouse), an interface to explore the traces and the evaluation and prompt-management layers. In short, OpenTelemetry produces the traces and Langfuse receives them, stores them and lets you work with them.
Is Langfuse v3 still supported?
Yes: according to the official migration guide, v3 will receive security patches until the end of January 2027. The latest v3 release, v3.225.8, came out on 16 September 2026 with three fixes backported from v4, one of them a security fix. To migrate, first upgrade ClickHouse to 25.12, let the background migrations finish and back up Postgres and ClickHouse. Then start v4 in legacy or dual mode, which keeps the old read endpoints and trace-level evaluators working until you switch to events_only.
Can I use Langfuse with local models?
Yes. Because the SDK instruments your code, not the provider, it works with any model. That includes those you run on your own machine with Ollama (the tool that downloads and serves open-source language models on your own hardware) through its OpenAI-compatible API. By self-hosting Langfuse and serving the model locally you get a complete agent platform in which no data leaves your network, which is key in environments with strict privacy requirements.
If you are still choosing a backend, the full comparison is in AI agent observability tools, covering licences, self-hosting and where each one breaks.
Conclusion
Langfuse fills the gap between "my agent works on my laptop" and "my agent works in production and I know why". With a docker compose up you get a complete observability platform, which in v4 stores everything as observations in ClickHouse.
It records every trace, span and generation and computes the cost and latency of each call. It lets you evaluate changes with datasets and an LLM judge, all within your own network and under the MIT licence. The next step is to clone the repository at the v4.37.0 tag, generate the secrets and bring up the containers. Then install the SDK with pip install langfuse and add an @observe() to the first function of your agent to watch your first trace appear.
Sources
- Official Langfuse documentation
- Langfuse on GitHub
- Langfuse Python SDK releases
- OpenTelemetry, the tracing standard
- Langfuse v4.0.0 release notes
- Langfuse v4.37.0 release notes
- Official Langfuse v3 to v4 migration guide
- Deploying Langfuse with Docker Compose
- Langfuse application containers and images
- Python SDK v3 to v4 migration
- Langfuse joins ClickHouse
- Langfuse Enterprise licence features
Source code
Access all the source code for this post on GitHub.
View on GitHub