Key Takeaways
- Langfuse is MIT-licensed except for its
ee/,web/src/ee/, andworker/src/ee/directories, and self-hosting runs "the same infrastructure that powers Langfuse Cloud" with "no scalability limitations between the different versions" (Langfuse self-hosting docs)- Self-hosted LangSmith is "an add-on to the Enterprise plan designed for our largest, most security-conscious customers" — there is no free or open-source self-host path, so the deployment decision runs through a sales contract (LangChain docs)
- The billing units are not comparable: LangSmith charges per trace (one trace holds up to 25,000 runs), while Langfuse Cloud meters ingestion volume — so a 40-step agent run is one billable object on one platform and dozens on the other (LangChain docs)
- LangSmith base traces carry 14-day retention and extended traces 400-day, and Annotation Queues, Run Rules, and Feedback silently upgrade a base trace to the extended tier — which is where eval-heavy teams find their bill (LangChain docs)
- Langfuse gates data retention policies, audit logs, project-level RBAC, and server-side data masking behind an Enterprise licence key — the four things regulated buyers usually need on day one
Most teams pick an LLM observability platform by opening two pricing pages side by side and comparing the headline numbers. That comparison is close to meaningless, because the two platforms do not bill the same object. One charges for the request; the other charges for the steps inside it. An agent that fans out to twenty tool calls per turn lands in completely different cost brackets depending on which unit you are buying.
Langfuse vs LangSmith comes down to deployment rights and trace shape, not capability — both capture traces, spans, scores, prompts, and datasets. Langfuse is MIT-licensed and self-hostable on any tier, which makes it the default when trace payloads cannot leave your network. LangSmith wins when the application is built on LangChain or LangGraph and zero-config tracing plus a managed service is worth the Enterprise gate on self-hosting.
Self-Hosting Is the First Fork, and It Is a Licence Question
The self-hosting decision eliminates one of these platforms before you evaluate a single feature. Langfuse's repository is MIT-licensed with a narrow carve-out: its LICENSE scopes commercial terms to the ee/, web/src/ee/, and worker/src/ee/ directories and releases everything else under MIT Expat (langfuse/langfuse). Self-hosted LangSmith has no equivalent path — it is an Enterprise add-on reached through sales.
LangChain's own documentation is direct about this: self-hosted LangSmith is "an add-on to the Enterprise plan designed for our largest, most security-conscious customers," and "hybrid and self-hosted deployment options are available on Enterprise plan" (LangChain docs). Obtaining a licence key means going through sales and signing an enterprise contract. For a team that needs traces to stay inside a VPC next quarter, that is a procurement timeline, not a configuration change.
Langfuse's position is the inverse. Its self-hosting documentation states that "all core Langfuse features and APIs are available in Langfuse OSS (MIT licensed) without any limits," and that "when running Langfuse self-hosted, you use the same deployment infrastructure as Langfuse Cloud. There are no scalability limitations between the different versions" (Langfuse docs). The self-hosted build is not a throttled community edition — it is the production artefact.
Ownership changed without changing the licence. ClickHouse announced its acquisition of Langfuse on 16 January 2026, stating that Langfuse remains open source under its existing MIT licence for core features and that Langfuse Cloud continues as a standalone service (ClickHouse). The repository's LICENSE now carries a Copyright (c) 2023-2026 ClickHouse, Inc. line with the MIT terms intact — the licence text itself is the verification.
What Self-Hosting Langfuse Actually Costs You in Operations
Self-hosting Langfuse is free to licence and not free to run. The reference docker-compose.yml brings up six services: langfuse-web, langfuse-worker, ClickHouse, PostgreSQL, Redis, and MinIO for S3-compatible blob storage (langfuse/langfuse). That is an OLAP database, a relational database, a cache, and an object store to operate before you have logged a single trace.
Langfuse's own deployment matrix is honest about the gap between a demo and production. Docker Compose is scoped to "local use and testing" with the responsibility listed as a "single VM without high availability, scaling, or backups." Production self-hosting points to Kubernetes via Helm, or the maintained AWS, Azure, and GCP paths (Langfuse docs). Teams that evaluate on Compose and then budget for production on the same footing are the ones who get surprised.
The architecture is built for ingestion spikes rather than for minimal infrastructure. Traces arrive in batches at the web container, get written to S3 immediately, and only a reference is queued in Redis; the worker then ingests from S3 into ClickHouse. Because every tracing and evaluation event is persisted to blob storage before processing, a temporarily unavailable database delays ingestion instead of dropping it. That durability is why the component count is what it is — and why Prodinit treats a self-hosted Langfuse deployment as an infrastructure workstream with its own runbook, not a sidecar on an application cluster.
The Billing Unit Mismatch That Breaks Price Comparisons
LangSmith bills the trace; Langfuse Cloud meters ingestion volume. That single difference makes list-price comparison invalid for any application whose traces contain more than a handful of steps. LangSmith's documentation defines a run as "a single unit of work," comparable to an OpenTelemetry span, and caps a trace at 25,000 runs before further runs are rejected (LangChain docs).
Read that cap as a pricing statement, not a technical limit. A LangGraph agent that makes forty LLM and tool calls in one turn is one billable trace on LangSmith. The same workload metered per ingested event is roughly forty times larger. Deep-trace, agentic workloads therefore favour per-trace billing, while shallow, high-volume workloads — a single classification or RAG call per request — favour per-event metering, where you are not paying a trace-shaped minimum for a one-step operation.
So model your own trace shape before you compare anything. Pull a representative day of production traffic, count observations per trace at the median and the p95, then price that distribution against both vendors' current rates. List prices are deliberately not quoted here: both vendors reprice regularly, and the third-party comparison posts that do quote figures disagree materially on included allowances and overage rates. The authoritative numbers are on the Langfuse pricing page and the LangSmith pricing page — and your spans-per-trace ratio decides which column of either table applies to you.
There is a second-order cost on LangSmith worth knowing before you instrument. Its invoice carries two metrics, "LangSmith Traces (Base Charge)" and "LangSmith Traces (Extended Data Retention Upgrades)," where base traces retain for 14 days and extended traces for 400. Annotation Queues, Run Rules, and Feedback automatically upgrade a base trace to the extended tier, billed on top of the base charge (LangChain docs). An eval-heavy team posting feedback scores across sampled traffic is upgrading traces as a side effect of doing evals properly — exactly the practice our LLM evaluation rubric recommends.
Retention and Access Control: Where Both Platforms Charge
Retention is a compliance requirement for regulated buyers, and both platforms treat it as a paid capability. LangSmith's 14-day base retention is shorter than most audit windows, so extended retention is effectively mandatory in healthcare or finance — and because it is a per-trace upgrade, that compliance requirement scales with traffic rather than sitting in a flat line item.
Langfuse charges differently for the same outcome. Its Enterprise licence key gates nine features, and four of them are precisely what a regulated deployment needs on day one: data retention policies, audit logs, project-level RBAC roles, and server-side data masking. The others are protected prompt labels, UI customisation, organisation creators, the Org Management API and SCIM, and the Instance Management API (Langfuse docs). The key activates through a single LANGFUSE_EE_LICENSE_KEY environment variable on both containers, and removing it leaves core OSS features working while the Enterprise APIs return 403.
The practical read: "Langfuse is MIT-licensed" and "Langfuse is free for a HIPAA-scoped deployment" are different claims. Trace payloads containing PHI need server-side masking at ingestion and enforced retention policies, both of which sit behind the licence key. What self-hosting buys unconditionally is data residency — the traces never leave your network — which is usually the binding constraint anyway, as covered in our guide to private LLMs for regulated industries.
Langfuse vs LangSmith: Decision Table
The Langfuse vs LangSmith choice compresses into eight dimensions that decide real production deployments: licence, self-hosting rights, billing unit, trace shape, managed retention, framework coupling, operational burden, and team fit. No single row is decisive on its own — identify which rows are hard constraints for your workload and let those pick the platform.
| Dimension | Langfuse | LangSmith |
|---|---|---|
| Licence | MIT Expat, except ee/, web/src/ee/, worker/src/ee/ | Proprietary |
| Self-hosting | Available on every tier, same infrastructure as Cloud | Enterprise plan add-on, via sales contract |
| Billing unit | Metered ingestion volume | Per trace (up to 25,000 runs per trace) |
| Cost favours | Shallow, high-volume calls — one step per request | Deep agentic traces — many runs per request |
| Managed retention | Retention policies require an Enterprise licence key | 14 days base, 400 days extended (per-trace upgrade) |
| Framework coupling | Framework-agnostic SDKs | Zero-config tracing for LangChain and LangGraph |
| Self-host stack to operate | Web + worker, ClickHouse, PostgreSQL, Redis, S3/MinIO | Managed by default; Enterprise self-host otherwise |
| Best for | Data-residency constraints, non-LangChain stacks, cost control at volume | LangChain/LangGraph-native apps wanting zero instrumentation work |
What Prodinit Runs in Production
Prodinit standardises on self-hosted or cloud Langfuse across LLM engagements, and the deciding factor is almost always data control plus framework independence rather than price. Most client stacks are direct SDK calls to OpenAI, Azure OpenAI, or Bedrock rather than LangChain, so LangSmith's strongest advantage — zero-config tracing for LangChain and LangGraph — does not apply.
The clearest example is a voice AI platform running 10,000+ calls per day that Prodinit instrumented with Langfuse in weeks 0–2 of a model distillation engagement, documented in the model distillation case study. That single instrumentation pass carried three jobs: live A/B dashboards across a 10% → 25% → 50% → 75% → 90% progressive rollout, stage-gate evaluation data at each increment, and quarterly collection of 80,000–100,000 training examples for the next distillation run. Inference cost dropped 70% with zero rollbacks across all five stage transitions.
That third job is the reason the framework-agnostic choice mattered. The observability layer was also the training-data pipeline, and the traces had to be exportable into a fine-tuning workflow rather than locked to a retention tier. Our LLM observability in production guide covers the instrumentation patterns in detail, including async score posting that keeps quality checks off the request path.
Prodinit would still reach for LangSmith in one situation: a LangGraph-native application where the team wants tracing in an afternoon, has no data-residency constraint, and runs deep agentic traces that per-trace billing prices favourably. That is a real set of conditions, not a courtesy — it is just not the shape of most production stacks we are handed. Either way, instrument before you optimise: the most common failure mode is neither platform, but shipping without the hallucination detection the traces exist to support.
Get Prodinit's AI engineering guides in your inbox
Deep-dives on production LLMs, voice AI, and MLOps — published weekly. No sales emails.
Frequently Asked Questions
The licence is free — Langfuse is MIT Expat outside its ee/, web/src/ee/, and worker/src/ee/ directories, with core features available "without any limits." The infrastructure is not: a reference deployment runs ClickHouse, PostgreSQL, Redis, and S3-compatible storage alongside the web and worker containers. Budget engineering time for Kubernetes or Terraform, backups, and upgrades.
No. LangChain's documentation describes self-hosted LangSmith as an add-on to the Enterprise plan for its "largest, most security-conscious customers," and obtaining a licence key requires going through sales and signing an enterprise agreement. Evaluation keys can be requested from the sales team, but there is no open-source or free self-hosting path comparable to Langfuse's.
It depends entirely on observations per trace, which is why list-price comparison misleads. LangSmith bills per trace, holding up to 25,000 runs each, so a 40-step agent turn is one billable object. Metered-ingestion pricing charges for those steps individually. Deep agentic traces favour per-trace billing; shallow single-call requests favour per-event metering.
Yes — Langfuse is framework-agnostic, with SDKs that instrument direct API calls to OpenAI, Azure OpenAI, or Bedrock as readily as LangChain or LlamaIndex pipelines. This is the main reason Prodinit defaults to it: most production stacks we inherit call provider SDKs directly, so LangSmith's zero-config LangChain and LangGraph tracing advantage never comes into play.
Base traces retain for 14 days; extended traces retain for 400. Annotation Queues, Run Rules, and Feedback automatically upgrade a base trace to the extended tier, billed on top of the base charge. Teams running evals across sampled production traffic should expect extended-retention upgrades to appear on the invoice as a consequence of posting scores.