trtutils.profiling package

Module contents

Submodule containing tools for profiling TensorRT engines.

Classes

LayerTiming

A dataclass to store per-layer profiling statistics.

ProfilerResult

A dataclass to store the complete profiling results.

LayerProfiler

A class that implements TensorRT’s IProfiler interface.

Functions

build_fused_layer_map()

Map individual ONNX layer names to their fused TRT layer.

identify_quantize_speedups_by_layer()

Identify which layers benefit most from INT8 quantization.

profile_engine()

Profile a TensorRT engine layer-by-layer.

resolve_fused_layer_value()

Look up a per-layer metric value, handling TensorRT layer fusion.

class trtutils.profiling.LayerProfiler(*args: Any, **kwargs: Any)[source]

Bases: IProfiler

A profiler that implements TensorRT’s IProfiler interface.

This class collects per-layer execution times across multiple inference iterations and can aggregate statistics for each layer.

report_layer_time(layer_name: str, ms: float) → None[source]

Record the execution time for a layer.

This method is called by TensorRT once per layer after inference, only if the profiler is not added to the context.

Parameters:
  • layer_name (str) – The name of the layer.

  • ms (float) – The execution time in milliseconds.

finalize_iteration() → None[source]

Finalize the current iteration by storing all layer timings.

This should be called after context.report_to_profiler() to commit the current iteration’s timings to the aggregate storage.

get_statistics() → list[LayerTiming][source]

Compute statistics for each layer across all iterations.

Returns:

A list of LayerTiming objects, one per layer, with aggregated statistics.

Return type:

list[LayerTiming]

reset() → None[source]

Reset all stored timings.

class trtutils.profiling.LayerTiming(name: str, mean: float, median: float, min: float, max: float, raw: list[float])[source]

Bases: object

A dataclass to store per-layer profiling statistics.

name

The name of the layer.

Type:

str

mean

The mean execution time in milliseconds.

Type:

float

median

The median execution time in milliseconds.

Type:

float

min

The minimum execution time in milliseconds.

Type:

float

max

The maximum execution time in milliseconds.

Type:

float

raw

The raw execution times in milliseconds across all iterations.

Type:

list[float]

name: str
mean: float
median: float
min: float
max: float
raw: list[float]
class trtutils.profiling.ProfilerResult(layers: Sequence[LayerTiming], total_time: LayerTiming, iterations: int)[source]

Bases: object

A dataclass to store the complete profiling results.

layers

The per-layer timing statistics.

Type:

list[LayerTiming]

total_time

The total execution time statistics across all layers.

Type:

LayerTiming

iterations

The number of profiling iterations performed.

Type:

int

layers: Sequence[LayerTiming]
total_time: LayerTiming
iterations: int
trtutils.profiling.build_fused_layer_map(profiled_layer_names: Sequence[str], onnx_layer_names: list[str] | None = None) → dict[str, tuple[str, int]][source]

Map individual ONNX layer names to their fused TRT layer.

TensorRT fuses layers in two ways:

  • GPU engines: "LayerA + LayerB + LayerC"

  • DLA engines: "{ForeignNode[FirstLayer...LastLayer]}"

This function parses both formats and maps each constituent ONNX layer name to the fused name and total constituent count.

Parameters:
  • profiled_layer_names (Sequence[str]) – Layer names from a profiled TensorRT engine (possibly fused).

  • onnx_layer_names (list[str] | None, optional) – Ordered list of all ONNX layer names. Required to resolve DLA ForeignNode ranges which only name the first and last layer.

Returns:

Maps individual layer name -> (fused_name, num_constituents).

Return type:

dict[str, tuple[str, int]]

trtutils.profiling.identify_quantize_speedups_by_layer(onnx: Path | str, *, build_func: BuildFunc | None = None, iterations: int = 100, warmup_iterations: int = 10, workspace: float = 4.0, ignore_mismatch_layers: bool = False, verbose: bool | None = None) → tuple[ProfilerResult, ProfilerResult, list[tuple[str, float]]][source]

Identify speedup by layer from INT8 quantization.

This function builds both FP16 and INT8 engines from an ONNX model (using either a user-provided build function or trtutils.builder.build_engine()), profiles them layer-by-layer, and computes the speedup percentage for each layer.

Profiled layers are grouped into chunks anchored by layers whose names correspond to ONNX nodes. Fused TRT-internal layers (e.g. __myl_Silu) between anchors are included in the preceding anchor’s chunk. When FP16 and INT8 fuse differently, chunks that appear in both profiles are compared directly.

Parameters:
  • onnx (Path | str) – The path to the ONNX model.

  • build_func (BuildFunc | None, optional) –

    A callable responsible for building engines. It must accept (onnx, output, *, fp16=None, int8=None) (or **kwargs including these) and build an engine at output with the requested precision flags.

    • When build_func is None (default), a thin wrapper over trtutils.builder.build_engine() is used with:

      • workspace forwarded from this function

      • profiling_verbosity=trt.ProfilingVerbosity.DETAILED

      • verbose forwarded from this function

    • To customize calibration, shapes, timing cache, hooks, etc., pass a functools.partial or custom function that captures those arguments.

  • iterations (int, optional) – The number of profiling iterations to run, by default 100.

  • warmup_iterations (int, optional) – The number of warmup iterations to run before profiling, by default 10.

  • workspace (float, optional) – The size of the workspace in gigabytes. Default is 4.0 GiB.

  • ignore_mismatch_layers (bool, optional) – Whether to ignore layers that cannot be resolved in both profiles. Default is False.

  • verbose (bool, optional) – Whether to output additional information to stdout. Default None/False.

Returns:

A tuple containing: - FP16 profiling results - INT8 profiling results - List of (layer_name, speedup_percent) tuples in execution order.

Positive values indicate INT8 is faster, negative values indicate INT8 is slower.

Return type:

tuple[ProfilerResult, ProfilerResult, list[tuple[str, float]]]

Raises:

ValueError – If a layer cannot be resolved in both profiles and mismatches are not ignored.

trtutils.profiling.profile_engine(engine: Path | str | TRTEngine, iterations: int = 100, warmup_iterations: int = 10, dla_core: int | None = None, device: int | None = None, *, warmup: bool | None = None, verbose: bool | None = None) → ProfilerResult[source]

Profile a TensorRT engine layer-by-layer.

This function runs inference multiple times and collects per-layer execution times using TensorRT’s IProfiler interface. It returns aggregated statistics (mean, median, min, max) for each layer across all iterations.

Notes

For best results, build the engine with profiling_verbosity set to DETAILED when calling build_engine. Otherwise, layer names may be numeric indices.

Parameters:
  • engine (Path | str | TRTEngine) – The engine to profile. Either a TRTEngine object or path to the engine file. If a path is given, then a TRTEngine will be created automatically.

  • iterations (int, optional) – The number of profiling iterations to run, by default 100.

  • warmup_iterations (int, optional) – The number of warmup iterations to run before profiling, by default 10.

  • dla_core (int, optional) – The DLA core to assign DLA layers of the engine to. Default is None. If None, any DLA layers will be assigned to DLA core 0.

  • device (int, optional) – The CUDA device index to use for the engine. Default is None, which uses the current device.

  • warmup (bool, optional) – Whether to do warmup iterations, by default None. If None, warmup will be set to True.

  • verbose (bool, optional) – Whether to output additional information to stdout. Default None/False.

Returns:

A dataclass containing per-layer timing statistics and total execution time.

Return type:

ProfilerResult

trtutils.profiling.resolve_fused_layer_value(layer_name: str, layer_values: dict[str, float], fused_map: dict[str, tuple[str, int]], *, verbose: bool | None = None, label: str = '') → float | None[source]

Look up a per-layer metric value, handling TensorRT layer fusion.

Resolution order:

  1. Exact match in layer_values

  2. Constituent of a fused layer (value split evenly among constituents)

  3. None — no data available

Parameters:
  • layer_name (str) – The ONNX layer name to look up.

  • layer_values (dict[str, float]) – Metric values keyed by TRT layer name (e.g., execution times).

  • fused_map (dict[str, tuple[str, int]]) – Mapping from constituent names to (fused_name, num_parts), as returned by build_fused_layer_map().

  • verbose (bool | None, optional) – Log warnings for unresolved layers.

  • label (str, optional) – Label for log messages (e.g., "FP16" or "INT8").

Returns:

The resolved value, or None if no data found.

Return type:

float | None