📝 TensorRT
Description
< What is it? >
TensorRT is NVIDIA's SDK for optimizing deep-learning inference on NVIDIA GPUs. It takes a trained model and builds an optimized engine that an application can run for lower latency or higher throughput.
- Tensor refers to the multidimensional data used by neural networks
- RT means runtime.
It is a deployment optimizer and runtime, not a framework for training a model or improving its learned accuracy.
trained model → export / conversion → TensorRT engine build → GPU inference
Key points
< Model to engine >
- Train a model in a framework such as PyTorch.
- Export or convert the deployable graph, often through ONNX or a PyTorch integration.
- TensorRT builds a serialized engine for a selected GPU and configuration.
- The TensorRT runtime loads that engine and repeatedly runs inference on new inputs.
< Why it can be faster >
TensorRT can fuse compatible operations, choose efficient GPU kernels and tensor layouts, and use mixed or lower precision such as FP16 or INT8 when the model and accuracy requirements allow it. The real gain depends on the model, input shapes, GPU, batch size, and serving setup, so benchmark the deployed workload rather than assuming a fixed speedup.
< Precision and accuracy >
Lower precision can reduce memory use and improve throughput, but it can also change numerical results. Validate latency, throughput, and task quality on representative inputs before shipping an engine.
< What it is not >
- TensorRT normally runs a model after it has been trained; it does not replace PyTorch's training loop or backpropagation.
- An engine is an optimized deployment artifact, not a universally portable model checkpoint. Build and compatibility settings should match the intended GPU and runtime environment.
Related ideas
- PyTorch is commonly used to build and train the model before deployment.
- Inference Frameworks & Runtimes groups tools for serving models and executing inference.
Comparison
< PyTorch and TensorRT >
| Aspect | PyTorch | TensorRT |
|---|---|---|
| Main role | Build, train, evaluate, and run models | Optimize and run deployed inference on NVIDIA GPUs |
| Typical artifact | Model code and weights / checkpoint | Optimized engine |
| Main strength | Flexible experimentation and training | Low-latency, high-throughput GPU inference |
Video Tutorial
- Torch-TensorRT
- TensorFlow-TensorRT
- High Performance Deep Learning Inference