T
TensorRT-LLM
NVIDIA/TensorRT-LLM
TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.
★14.3kstars
Python
NOASSERTION
Updated: Today
📋 Project at a Glance
Tap to expand
What's this?A blackwell/cuda tool in the Data & Infrastructure category, built with Python, open-source
Who made it?Maintained by NVIDIA team, 14.3K⭐ on GitHub, #136 out of 3133 in Data & Infrastructure
Why does it exist?The NVIDIA team recognized that existing blackwell tools in Data & Infrastructure were hard to use. TensorRT-LLM was designed to make cuda more accessible.
What can it do?Key use cases: llm-serving, moe, pytorch
How to install with AI?Use an AI coding assistant (Claude Code, Cursor, Copilot) to automatically set up pip dependencies and virtual env. Follow the README — the AI handles the rest.
🔗 github.com/NVIDIA/TensorRT-LLM | 官网 https://nvidia.github.io/TensorRT-LLM
🔗 github.com/NVIDIA/TensorRT-LLM | 官网 https://nvidia.github.io/TensorRT-LLM
Topics
blackwellcudallm-servingmoepytorch