NVIDIA's compiler and runtime for squeezing the most out of their GPUs. Built by NVIDIA.
Others in inference and serving
Common questions
Is TensorRT-LLM free to use?
TensorRT-LLM is open source and free to run yourself. You still pay whichever model provider you point it at, and this tutorial is free with no signup.
Do I need to know Python to use TensorRT-LLM?
Basic Python is enough. Functions, dictionaries and imports cover most of what TensorRT-LLM asks of you.
When should I not use TensorRT-LLM?
When something else in inference and serving fits the job better. The others in that group are listed below.
Lessons for TensorRT-LLM are being written. Meanwhile the APIs for AI tutorial covers the same ground: state, tools, loops, memory and human approval. Most of it carries straight over.
Start the APIs for AI tutorial →