← All frameworks

TensorRT-LLM

NVIDIA's compiler and runtime for squeezing the most out of their GPUs.

Inference and serving

NVIDIA's compiler and runtime for squeezing the most out of their GPUs. Built by NVIDIA.

Others in inference and serving

Common questions

Is TensorRT-LLM free to use?

TensorRT-LLM is open source and free to run yourself. You still pay whichever model provider you point it at, and this tutorial is free with no signup.

Do I need to know Python to use TensorRT-LLM?

Basic Python is enough. Functions, dictionaries and imports cover most of what TensorRT-LLM asks of you.

When should I not use TensorRT-LLM?

When something else in inference and serving fits the job better. The others in that group are listed below.

Lessons for TensorRT-LLM are being written. Meanwhile the APIs for AI tutorial covers the same ground: state, tools, loops, memory and human approval. Most of it carries straight over.

Start the APIs for AI tutorial →