# TensorRT-LLM

> NVIDIA's LLM inference framework for NVIDIA GPUs with TensorRT acceleration

- Category: [Inference & Compute](https://ailandscape.org/category/inference-compute) › Inference Optimization
- Homepage: https://nvidia.github.io/TensorRT-LLM/
- Repository: https://github.com/NVIDIA/TensorRT-LLM
- Added to the landscape: 2026-03-18

## Similar tools in Inference Optimization

- [llama.cpp](https://ailandscape.org/tool/llama-cpp): LLM inference in C/C++ for local deployment
- [ONNX Runtime](https://ailandscape.org/tool/onnx-runtime): Cross-platform ML model accelerator
- [SGLang](https://ailandscape.org/tool/sglang): Fast serving framework for LLMs and VLMs with structured generation support
- [Sol Engine](https://ailandscape.org/tool/sol-engine): NVIDIA's training-free acceleration framework for video diffusion inference
- [TensorRT](https://ailandscape.org/tool/tensorrt): NVIDIA's SDK for high-performance inference
- [vLLM](https://ailandscape.org/tool/vllm): High-throughput and memory-efficient LLM serving

---

Part of [AI Landscape](https://ailandscape.org), an open map of the AI ecosystem. Web page: https://ailandscape.org/tool/tensorrt-llm · Index for AI assistants: https://ailandscape.org/llms.txt
