trtutils.profiling package¶
Module contents¶
Submodule containing tools for profiling TensorRT engines.
Classes¶
LayerTimingA dataclass to store per-layer profiling statistics.
ProfilerResultA dataclass to store the complete profiling results.
LayerProfilerA class that implements TensorRT’s IProfiler interface.
Functions¶
build_fused_layer_map()Map individual ONNX layer names to their fused TRT layer.
identify_quantize_speedups_by_layer()Identify which layers benefit most from INT8 quantization.
profile_engine()Profile a TensorRT engine layer-by-layer.
resolve_fused_layer_value()Look up a per-layer metric value, handling TensorRT layer fusion.
- class trtutils.profiling.LayerProfiler(*args: Any, **kwargs: Any)[source]¶
Bases:
IProfilerA profiler that implements TensorRT’s IProfiler interface.
This class collects per-layer execution times across multiple inference iterations and can aggregate statistics for each layer.
- report_layer_time(layer_name: str, ms: float) None[source]¶
Record the execution time for a layer.
This method is called by TensorRT once per layer after inference, only if the profiler is not added to the context.
- finalize_iteration() None[source]¶
Finalize the current iteration by storing all layer timings.
This should be called after context.report_to_profiler() to commit the current iteration’s timings to the aggregate storage.
- get_statistics() list[LayerTiming][source]¶
Compute statistics for each layer across all iterations.
- Returns:
A list of LayerTiming objects, one per layer, with aggregated statistics.
- Return type:
- class trtutils.profiling.LayerTiming(name: str, mean: float, median: float, min: float, max: float, raw: list[float])[source]¶
Bases:
objectA dataclass to store per-layer profiling statistics.
- class trtutils.profiling.ProfilerResult(layers: Sequence[LayerTiming], total_time: LayerTiming, iterations: int)[source]¶
Bases:
objectA dataclass to store the complete profiling results.
- layers¶
The per-layer timing statistics.
- Type:
- total_time¶
The total execution time statistics across all layers.
- Type:
- layers: Sequence[LayerTiming]¶
- total_time: LayerTiming¶
- trtutils.profiling.build_fused_layer_map(profiled_layer_names: Sequence[str], onnx_layer_names: list[str] | None = None) dict[str, tuple[str, int]][source]¶
Map individual ONNX layer names to their fused TRT layer.
TensorRT fuses layers in two ways:
GPU engines:
"LayerA + LayerB + LayerC"DLA engines:
"{ForeignNode[FirstLayer...LastLayer]}"
This function parses both formats and maps each constituent ONNX layer name to the fused name and total constituent count.
- Parameters:
- Returns:
Maps individual layer name ->
(fused_name, num_constituents).- Return type:
- trtutils.profiling.identify_quantize_speedups_by_layer(onnx: Path | str, *, build_func: BuildFunc | None = None, iterations: int = 100, warmup_iterations: int = 10, workspace: float = 4.0, ignore_mismatch_layers: bool = False, verbose: bool | None = None) tuple[ProfilerResult, ProfilerResult, list[tuple[str, float]]][source]¶
Identify speedup by layer from INT8 quantization.
This function builds both FP16 and INT8 engines from an ONNX model (using either a user-provided build function or
trtutils.builder.build_engine()), profiles them layer-by-layer, and computes the speedup percentage for each layer.Profiled layers are grouped into chunks anchored by layers whose names correspond to ONNX nodes. Fused TRT-internal layers (e.g.
__myl_Silu) between anchors are included in the preceding anchor’s chunk. When FP16 and INT8 fuse differently, chunks that appear in both profiles are compared directly.- Parameters:
onnx (Path | str) – The path to the ONNX model.
build_func (BuildFunc | None, optional) –
A callable responsible for building engines. It must accept
(onnx, output, *, fp16=None, int8=None)(or **kwargs including these) and build an engine atoutputwith the requested precision flags.When
build_funcisNone(default), a thin wrapper overtrtutils.builder.build_engine()is used with:workspaceforwarded from this functionprofiling_verbosity=trt.ProfilingVerbosity.DETAILEDverboseforwarded from this function
To customize calibration, shapes, timing cache, hooks, etc., pass a
functools.partialor custom function that captures those arguments.
iterations (int, optional) – The number of profiling iterations to run, by default 100.
warmup_iterations (int, optional) – The number of warmup iterations to run before profiling, by default 10.
workspace (float, optional) – The size of the workspace in gigabytes. Default is 4.0 GiB.
ignore_mismatch_layers (bool, optional) – Whether to ignore layers that cannot be resolved in both profiles. Default is False.
verbose (bool, optional) – Whether to output additional information to stdout. Default None/False.
- Returns:
A tuple containing: - FP16 profiling results - INT8 profiling results - List of (layer_name, speedup_percent) tuples in execution order.
Positive values indicate INT8 is faster, negative values indicate INT8 is slower.
- Return type:
tuple[ProfilerResult, ProfilerResult, list[tuple[str, float]]]
- Raises:
ValueError – If a layer cannot be resolved in both profiles and mismatches are not ignored.
- trtutils.profiling.profile_engine(engine: Path | str | TRTEngine, iterations: int = 100, warmup_iterations: int = 10, dla_core: int | None = None, device: int | None = None, *, warmup: bool | None = None, verbose: bool | None = None) ProfilerResult[source]¶
Profile a TensorRT engine layer-by-layer.
This function runs inference multiple times and collects per-layer execution times using TensorRT’s IProfiler interface. It returns aggregated statistics (mean, median, min, max) for each layer across all iterations.
Notes
For best results, build the engine with profiling_verbosity set to DETAILED when calling build_engine. Otherwise, layer names may be numeric indices.
- Parameters:
engine (Path | str | TRTEngine) – The engine to profile. Either a TRTEngine object or path to the engine file. If a path is given, then a TRTEngine will be created automatically.
iterations (int, optional) – The number of profiling iterations to run, by default 100.
warmup_iterations (int, optional) – The number of warmup iterations to run before profiling, by default 10.
dla_core (int, optional) – The DLA core to assign DLA layers of the engine to. Default is None. If None, any DLA layers will be assigned to DLA core 0.
device (int, optional) – The CUDA device index to use for the engine. Default is None, which uses the current device.
warmup (bool, optional) – Whether to do warmup iterations, by default None. If None, warmup will be set to True.
verbose (bool, optional) – Whether to output additional information to stdout. Default None/False.
- Returns:
A dataclass containing per-layer timing statistics and total execution time.
- Return type:
- trtutils.profiling.resolve_fused_layer_value(layer_name: str, layer_values: dict[str, float], fused_map: dict[str, tuple[str, int]], *, verbose: bool | None = None, label: str = '') float | None[source]¶
Look up a per-layer metric value, handling TensorRT layer fusion.
Resolution order:
Exact match in
layer_valuesConstituent of a fused layer (value split evenly among constituents)
None— no data available
- Parameters:
layer_name (str) – The ONNX layer name to look up.
layer_values (dict[str, float]) – Metric values keyed by TRT layer name (e.g., execution times).
fused_map (dict[str, tuple[str, int]]) – Mapping from constituent names to
(fused_name, num_parts), as returned bybuild_fused_layer_map().verbose (bool | None, optional) – Log warnings for unresolved layers.
label (str, optional) – Label for log messages (e.g.,
"FP16"or"INT8").
- Returns:
The resolved value, or
Noneif no data found.- Return type:
float | None