Langfuse: Open-Source LLM Observability You Can Self-Host
Langfuse brings full tracing, prompt versioning, dataset evaluation, and cost attribution to LLM apps - and you can run the entire stack on your own servers.
Mahmudul Haque Qudrati
CEO & ML Engineer
One AI engineering post, weekly
LLM benchmarks, prompt techniques, and token-cost breakdowns — not another AI news roundup.
Traditional APM tools (Datadog, New Relic) capture HTTP latency and error rates. They tell you a request took 4 seconds - but not that 3.8 seconds was an LLM call, that the prompt had 1,200 tokens, that the model used was gpt-4o, or that the user-visible answer was a hallucination. Langfuse fills this gap with LLM-native observability.
Core Hierarchy
Trace → Span → Generation
- A Trace represents one user interaction (e.g., one chat message)
- Spans are logical steps within a trace (retrieve documents, format prompt, call LLM)
- A Generation is a specific LLM call with input/output tokens, model name, latency, and cost
Team workspace
Ship faster with chat, meetings, and projects in one place — Zlyqor.
Python SDK Integration
pip install langfuse
Decorator-Based Tracing
from langfuse.decorators import observe, langfuse_context
from openai import OpenAI
client = OpenAI()
@observe()
def retrieve_docs(query: str) -> str:
# Simulate vector retrieval
return f"Documents for: {query}"
@observe()
def generate_answer(docs: str, question: str) -> str:
langfuse_context.update_current_observation(
input={"docs": docs, "question": question}
)
response = client.chat.completions.create(
model="gpt-4o-mini",
messages=[
{"role": "system", "content": f"Context: {docs}"},
{"role": "user", "content": question},
],
)
answer = response.choices[0].message.content
langfuse_context.update_current_observation(output=answer)
return answer
@observe()
def answer_question(question: str) -> str:
docs = retrieve_docs(question)
return generate_answer(docs, question)
result = answer_question("What is PagedAttention?")
Every @observe() call automatically creates a span in the trace. Langfuse captures function arguments as input and return values as output.
Prompt Management With Versioning
Store, version, and A/B test prompts in the Langfuse UI:
from langfuse import Langfuse
lf = Langfuse()
prompt = lf.get_prompt("answer-question", version=3)
compiled = prompt.compile(context="...", question="...")
Changing a prompt in production no longer requires a code deploy.
Dataset Creation From Production Traces
Mark any trace as a dataset item directly from the UI. Build ground-truth datasets from real user interactions, then run batch evaluations to compare prompt versions or models.
LLM-as-Judge Scoring
from langfuse import Langfuse
lf = Langfuse()
lf.score(
trace_id="trace-xyz",
name="faithfulness",
value=0.92,
comment="Answer matches source documents",
)
Automate this with an evaluator function that runs after each generation.
Self-Hosting With Docker
git clone https://github.com/langfuse/langfuse.git
cd langfuse
docker compose up -d
The stack includes PostgreSQL, Redis, and the Langfuse web server. Full instructions in the self-host guide.
User-Level Cost Attribution
Pass user_id to attribute token costs to individual users - essential for multi-tenant SaaS billing:
langfuse_context.update_current_trace(user_id="user-456", session_id="session-789")
The dashboard then shows cost-per-user histograms and per-session token usage.

Mahmudul Haque Qudrati
CEO & ML Engineer
Visionary technologist, software engineer, and machine learning specialist. Founder and CEO of Pristren, directing engineering teams that ship production-grade AI/ML pipelines, mission-critical full-stack applications, and developer tooling. Creator of Zlyqor, the unified team workspace platform. Author of 540+ technical guides and benchmark research reports on large language models, agentic workflows, Model Context Protocol (MCP), and modern web stacks.
More from Mahmudul
Related Articles
How to Use Claude to Make Videos Like Vox and Others
Claude can help you make Vox-style videos by generating scripts, editing with code, and automating animation. Here's a practical guide with real workflows and costs.
OpenAI Ends Cursor Model Access on Nov 12, 2026: What Developers Need to Know
OpenAI will terminate Cursor's access to its models on November 12, 2026, following SpaceX's acquisition. This guide explains the timeline, why it happened, and practical steps to migrate your workflow.
I Used Claude Code to Get a Second Opinion on My MRI: A Practical Overview
A developer fed his MRI scan to Claude Code and got a second opinion. Here's how the experiment worked, what it cost, and why you shouldn't rely on it for medical decisions.
// discussion
Comments