Overview
vLLM is an inference and serving engine designed to run large language models with high throughput and efficient memory use. It gives teams a unified way to deploy a broad selection of open-source models across multiple hardware platforms.
The engine also provides an OpenAI-compatible API, making it suitable for applications that need a familiar interface for model integration.
Key Features
- Uses PagedAttention, advanced scheduling, and continuous batching to improve throughput and GPU utilization.
- Supports a unified API across NVIDIA CUDA, AMD ROCm, Intel XPU and Gaudi, Google Cloud TPU, AWS Neuron, CPUs, Apple Silicon, and other accelerators.
- Serves models from families including DeepSeek, Gemma, Llama, Mistral, MiniMax, Nemotron, Qwen, StepFun, and GLM.
- Offers stable builds for tested releases and nightly builds for access to newer changes.
Getting Started
vLLM requires Python 3.10 or later, with Python 3.12 or newer recommended. The project recommends the uv package manager for faster and more reliable installation, while Python packages and Docker are among the available setup routes.
Users select the build channel, target platform, package format, and relevant CUDA version before installing. Documentation covers additional platforms and troubleshooting for configurations beyond the quick-start options.
Resources and Community
The project publishes example notebooks and tutorials, performance benchmarks and comparisons, and a public roadmap. Community support is available through real-time discussions, a searchable forum, and issue tracking for bug reports and feature requests.
vLLM is a community project under the PyTorch Foundation, with development and testing resources supported by multiple organizations.