
Context
OVHcloud AI Endpoints provides an OpenAI-compatible API that gives applications access to a broad catalogue of open-weight models, including Qwen, Llama, Mistral, gpt-oss, embedding, guard and speech models, without requiring teams to manage inference infrastructure. It is the usage-based inference layer monitored by this architecture.
OVHcloud Managed Kubernetes Service (MKS) removes the operational burden of running the Kubernetes control plane: node pools, upgrades and high availability are handled by OVHcloud, and you keep full control over what runs on the worker nodes. It’s the natural place to run a stateless application like Langfuse’s web and worker processes, while the stateful pieces (databases, object storage) live in OVHcloud’s managed services next to it.
Langfuse provides the observability layer for applications using AI Endpoints. Because its OpenAI SDK integration is a drop-in replacement (from langfuse.openai import openai instead of import openai), pointing an existing OpenAI-compatible codebase at AI Endpoints and getting full tracing out of it is a two-line change, not a rewrite.
Together, these three pieces give a self-hosted, cost-attributed observability stack for any application calling AI Endpoints, without sending a single token of prompt or completion data outside infrastructure you control.

Langfuse on OVHcloud MKS for LLM observability and token consumption tracking
Architecture overview
Langfuse’s web and worker deployments run inside the MKS cluster. Ingress and certificate management run alongside them as separate, reusable cluster add-ons. Every stateful dependency – Postgres, Valkey, ClickHouse, object storage – is an OVHcloud managed service outside the cluster, reached over TLS.
1. Data flow
Putting the pieces together, a single traced request flows like this:

Data flow
1. The application calls AI Endpoints directly
The drop-in openai.OpenAI() client sends the request straight to the inference API – Langfuse sits outside this call entirely, so a Langfuse outage never affects whether the application can get a completion.
2. AI Endpoints streams the response back
The final chunk contains the token usage data.
3. The SDK exports what it just saw as an OpenTelemetry span batch
This is done asynchronously, via HTTPS, to /api/public/otel/v1/traces on the Langfuse instance.
This process runs in the background and adds no latency to the response the application has already received in step 2.
4. Langfuse web ingests the batch
It checks the request’s API key against its project in Postgres (a lookup cached in Valkey, not a fresh query on every call), writes the raw batch as-is to object storage, and pushes only a reference to it onto a Valkey queue.
5. The worker drains that queue on its own schedule, decoupled from any specific request
It pulls the reference from Valkey, fetches the full batch back from object storage, resolves the Prompt Management entry it’s linked to (if any) via Postgres, and persists the trace and generation records into ClickHouse (model, token counts, latency, prices)
6. The worker also handles anything scheduled or bulk
Batch exports and media uploads land in S3 independently of when the original request happened.
⚠️ Note: Nothing from step 3 onward can add latency to the application’s request in steps 1-2, that asynchronous handoff, and the fact that web never blocks on Postgres or ClickHouse to acknowledge a batch, is the entire reason tracing doesn’t cost anything on the critical path.
2. AI Endpoints APIs
Everything below calls one of two distinct OVHcloud AI Endpoints APIs:
- Inference API: https://oai.endpoints.kepler.ai.cloud.ovh.net/v1
This is the API your application actually talks to. It implements the OpenAI API surface, so any OpenAI-compatible SDK works against it by changing the base_url and the API key:
import openai
client = openai.OpenAI(
api_key="<your-ai-endpoints-token>",
base_url="https://oai.endpoints.kepler.ai.cloud.ovh.net/v1",
)
response = client.chat.completions.create(
model="Qwen3.5-397B-A17B",
messages=[{"role": "user", "content": "hello"}],
)If you retrieve the OpenAPI schema directly from the gateway (GET /openapi.json), you will see that the available interface goes well beyond chat suggestions.
It is useful to call the GET /v1/models