TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtim
Claim its page: indexed whatever its rank, translated into six languages, and enriched with what you write yourself.
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtim
Read the full guideTensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtim
TensorRT-LLM is an open-source project.
Yes. TensorRT-LLM is free and open source — you can use, modify and self-host it.
Add this live badge to your README — your GitHub stars and directory rank, refreshed daily.
[](https://opensourceai.tech/project/nvidia-tensorrt-llm.html)