Skip to main content

Overview

Langfuse is an open-source LLM engineering platform: traces, token and cost analytics, prompt management, and evals. It pairs with GoModel in two complementary ways:
  • Gateway traces. GoModel’s OpenTelemetry exporter sends a trace for every request to Langfuse’s OTLP endpoint. Every client of the gateway is covered without touching application code, and each trace shows the model, the provider that served it, latency, token usage and cost, status, and every retry or failover. Prompts and completions are added only when you turn on content capture.
  • Application traces. The Langfuse SDK’s OpenAI drop-in wrapper, pointed at GoModel as its base URL. Langfuse records full prompts, completions, and token usage; GoModel supplies the routing, failover, caching, and its own audit log and cost tracking underneath.
Use either on its own, or both together — with W3C trace context propagation they merge into a single Langfuse trace per request (see One trace end to end). App (Langfuse SDK) -> GoModel -> OpenAI/Anthropic/Gemini/... with traces from both hops arriving in Langfuse.

1. Send gateway traces to Langfuse

Langfuse ingests OTLP traces at /api/public/otel, authenticated with a project’s API key pair. Base64-encode the keys:
Then enable GoModel’s exporter:
The quotes matter in a shell: the header value contains a space. In a Docker .env file, drop them — there they would become part of the value. Keep the default http/protobuf protocol: Langfuse’s OTLP endpoint does not speak gRPC. The API keys come from Project Settings -> API Keys in Langfuse; self-hosted deployments can also pre-provision them with the LANGFUSE_INIT_* variables. To run Langfuse itself, see its docker compose quickstart — git clone https://github.com/langfuse/langfuse.git && cd langfuse && docker compose up, UI on port 3000.
The Basic auth header carries your Langfuse secret key on every export. Use an https:// endpoint unless Langfuse runs on the same host; GoModel warns at startup when auth headers are configured with a plaintext non-loopback endpoint.

What appears in Langfuse

Each gateway request becomes one trace: a POST /v1/chat/completions span with a nested generation per provider call, named after the operation and model (chat gpt-5-mini). The generation carries the model the provider answered with, input and output token counts, and the finish reason, so Langfuse prices it from its model catalog and fills its cost and token dashboards. Its metadata also carries gen_ai.provider.name and gomodel.provider.name (the exact provider from your configuration), so a request that failed over shows one generation per attempt — which provider failed, with what error class, and which one answered, with the usage on the attempt that answered. Streamed chat completions, Responses API, and translated /v1/messages calls get a generation that lasts until the stream ends, with the same usage. /v1/messages requests that GoModel forwards natively to an Anthropic provider appear without usage. Langfuse computes cost from its own price list. GoModel’s cost tracking applies your pricing overrides and cache rates, so the two totals can differ.

Capture prompts and completions

Prompts and completions stay out of telemetry unless you turn on content capture:
Each generation then shows the request messages as its input and the response as its output, including tool calls and tool results. Images and files appear as their type only, and very long text is truncated; see Prompt and completion capture.
Content capture sends every sampled prompt and completion to Langfuse. Keep it off when that data must not leave the gateway; the audit log records it in your own storage instead.

2. Trace from the application with the Langfuse SDK

To trace from the application instead — for example to group calls into your own spans or attach Langfuse prompt versions — instrument it with Langfuse’s OpenAI wrapper and point it at GoModel. That takes two changed lines in an existing OpenAI-SDK app:
Langfuse records the messages, the completion, and token usage for every call, streaming included, and prices them from its model catalog. Because the base URL is GoModel, one client reaches every configured provider through provider-prefixed model ids (openai/gpt-5-mini, gemini/gemini-2.5-flash, virtual models), and every call still gets GoModel’s failover, caching, budgets, and audit trail. The JS/TS wrapper works the same way.

3. One trace end to end

With both paths enabled, an LLM call produces two separate Langfuse traces — the SDK’s and the gateway’s. Propagate W3C trace context to merge them: send a traceparent header with the request and GoModel parents its spans under the caller’s trace.
The resulting single trace nests the application span, the SDK generation with the prompt, GoModel’s server span, and the provider generation — what the user asked, what it cost, and how the gateway routed it, in one view. Both generations in that trace carry the call’s token usage, so Langfuse adds its cost twice to the trace total. When you rely on trace-level cost, filter cost views to one of the two generations, or send the gateway’s traces to a separate Langfuse project.
Verified with Langfuse v4.27.0 (self-hosted via docker compose) and Langfuse Python SDK 4.15.1: gateway OTLP traces for buffered and streaming chat across OpenAI, Anthropic, and Gemini, failover spans, SDK prompt capture through GoModel, and trace-context propagation merging both.

Troubleshooting

Notes

  • High-traffic gateways can sample gateway traces (OTEL_TRACES_SAMPLER=parentbased_traceidratio, OTEL_TRACES_SAMPLER_ARG=0.1) — with parentbased_*, requests whose caller sampled their trace keep their gateway spans, so propagated traces stay complete.
  • The full exporter reference — YAML config, per-signal endpoints, samplers, reload behavior — is in the OpenTelemetry guide. Everything there applies; Langfuse is just the OTLP backend.
  • For dashboards over request-rate and latency metrics, pair this with Prometheus metrics — Langfuse holds the traces, Prometheus the metrics.
Last modified on October 4, 2026