Welcome to trtutils

trtutils is a high-level Python interface for TensorRT inference, providing a simple and unified way to run arbitrary TensorRT engines. This library abstracts away the complexity of CUDA memory management, binding management, and engine execution.

Features

  • Simple, high-level interface for TensorRT inference

  • Automatic CUDA memory management, CUDA graphs, and pagelocked/unified memory

  • Support for arbitrary TensorRT engines on CUDA 11, 12, and 13

  • Engine building from ONNX, including strongly typed, DLA, and quantized builds

  • Built-in preprocessing and postprocessing capabilities

  • End-to-end image models for detection, classification, depth estimation, and hand-object interaction

  • Model download and ONNX export for many popular architectures

  • Comprehensive type hints and documentation

  • Performance benchmarking, profiling, and monitoring

Quick Start

from trtutils import TRTEngine

# Load your TensorRT engine
engine = TRTEngine("path_to_engine")

# Get input specifications
print(engine.input_shapes)  # Expected input shapes
print(engine.input_dtypes)  # Expected input data types

# Run inference
inputs = read_your_data()
outputs = engine.execute(inputs)

Documentation

Indices and Tables

Getting Help