Skip to content
OVHcloud Blog
OVHcloud Blog
  • OVHcloud
  • Twitter
  • LinkedIn
  • Facebook
  • YouTube
  • Discord
  • Twitch
  • GitHub
  • RSS
  • Legal notice

vLLM

vLLM on OVHcloud MKS for high availability and full observability

Reference Architecture: Deploying a vision-language model with vLLM on OVHcloud MKS for high performance inference and full observability

Eléa Petton | 10/04/2026 | AI, GPU, Kubernetes, LLM, Open Source, OVHcloud, prometheus, Public Cloud, vLLM

Ensure complete digital sovereignty of your AI models with end-to-end control through open-source solutions on OVHcloud’s Managed Kubernetes Service. This reference architecture demonstrates […]

Reference Architecture: Deploying a vision-language model with vLLM on OVHcloud MKS for high performance inference and full observability Read More »

Serve LLM with vLLM and AI Deploy

How to serve LLMs with vLLM and OVHcloud AI Deploy

Mathieu Busquet | 29/05/2024 | AI, AI Deploy, AI Endpoints, Artificial Intelligence, Deep learning, GPU, LLaMA, LLaMA 3, LLM Serving, Mistral, Mixtral, vLLM

In this tutorial, we will learn how to serve Large Language Models (LLMs) using vLLM and the OVHcloud AI Products.

How to serve LLMs with vLLM and OVHcloud AI Deploy Read More »

Copyright © 2026 OVHcloud Blog
  • OVHcloud
  • Twitter
  • LinkedIn
  • Facebook
  • YouTube
  • Discord
  • Twitch
  • GitHub
  • RSS
  • Legal notice