
vLLM
vLLM is available in the Nebius Marketplace as a free, self-managed deployment — on a VM in minutes or on Managed Kubernetes — with an OpenAI-compatible API, PagedAttention, and continuous batching. Nebius also supports the vLLM open-source project itself, providing GPU clusters the project has used to test and optimize inference, including DeepSeek R1 work. Nebius publishes practical guides for serving open models with vLLM on Nebius AI Cloud.