← All frameworks

vLLM

The serving engine most teams reach for when throughput matters.

Inference and serving

The serving engine most teams reach for when throughput matters. Built by UC Berkeley.

Others in inference and serving

Common questions

Is vLLM free to use?

vLLM is open source and free to run yourself. You still pay whichever model provider you point it at, and this tutorial is free with no signup.

Do I need to know Python to use vLLM?

Basic Python is enough. Functions, dictionaries and imports cover most of what vLLM asks of you.

When should I not use vLLM?

When something else in inference and serving fits the job better. The others in that group are listed below.

Lessons for vLLM are being written. Meanwhile the APIs for AI tutorial covers the same ground: state, tools, loops, memory and human approval. Most of it carries straight over.

Start the APIs for AI tutorial →