trtutils.image package

Subpackages

Submodules

Module contents

Utilities for using TensorRT on images.

Submodules

kernels

Kernels for image processing with TensorRT.

preprocessors

Preprocessors for images.

postprocessors

Postprocessors for images.

interfaces

Interfaces for image models.

onnx_models

Base ONNX models for creating ‘micro-engines’ for image processing.

sahi

SAHI (Slicing Aided Hyper Inference) for object detection.

Classes

Classifer

Wrapper around classification models.

DepthEstimator

Wrapper around depth estimation models.

Detector

Wrapper around detection models.

HandInteractionDetector

Wrapper around hand-object interaction models.

SAHI

SAHI wrapper for slicing aided inference.

ImageModel

Base class for models which process images.

class trtutils.image.SAHI(detector: DetectorInterface, slice_size: tuple[int, int] | None = None, slice_overlap: tuple[float, float] = (0.2, 0.2), iou_threshold: float = 0.5, *, agnostic_nms: bool = False, verbose: bool = False)[source]

Bases: object

Simple implementation of SAHI.

end2end(image: np.ndarray, conf_thres: float | None = None, nms_iou_thres: float | None = None, *, extra_nms: bool | None = None, agnostic_nms: bool | None = None, verbose: bool | None = None) → list[tuple[tuple[int, int, int, int], float, int]][source]

Perform end to end inference using detection model and SAHI.

Parameters:
  • image (np.ndarray) – The image to perform inference with.

  • conf_thres (float, optional) – The confidence threshold with which to retrieve bounding boxes. By default None

  • nms_iou_thres (float) – The IOU threshold to use during the optional/additional NMS operation. By default, None which will use value provided during initialization.

  • extra_nms (bool, optional) – Whether or not to perform an additional NMS operation. By default None, which will use value provided during initialization.

  • agnostic_nms (bool, optional) – Whether or not to perform class-agnostic NMS for the optional/additional operation. By default None, which will use value provided during initialization.

  • verbose (bool, optional) – Whether or not to log additional information.

Returns:

The detections where each entry is bbox, conf, class_id

Return type:

list[tuple[tuple[int, int, int, int], float, int]]

class trtutils.image.Classifier(engine_path: Path | str, warmup_iterations: int = 10, input_range: tuple[float, float] = (0.0, 1.0), preprocessor: str = 'trt', resize_method: str = 'linear', mean: tuple[float, float, float] | None = None, std: tuple[float, float, float] | None = None, dla_core: int | None = None, device: int | None = None, backend: str = 'auto', *, warmup: bool | None = None, pagelocked_mem: bool | None = None, unified_mem: bool | None = None, cuda_graph: bool | None = None, no_warn: bool | None = None, verbose: bool | None = None)[source]

Bases: ImageModel, ClassifierInterface

Implementation of image classifiers.

postprocess(outputs: list[np.ndarray], *, no_copy: bool | None = None, verbose: bool | None = None) → list[list[np.ndarray]][source]

Postprocess the outputs.

Parameters:
  • outputs (list[np.ndarray]) – The raw outputs from the engine to postprocess.

  • no_copy (bool, optional) – If True, do not copy the data from the allocated memory. If the data is not copied, it WILL BE OVERWRITTEN INPLACE once new data is generated.

  • verbose (bool, optional) – Whether or not to log additional information.

Returns:

The postprocessed outputs per image.

Return type:

list[list[np.ndarray]]

run(images: list[np.ndarray], *, preprocessed: bool | None = None, postprocess: Literal[False], no_copy: bool | None = None, verbose: bool | None = None) → list[np.ndarray][source]
run(images: list[np.ndarray], *, preprocessed: bool | None = None, postprocess: Literal[True] | None = None, no_copy: bool | None = None, verbose: bool | None = None) → list[list[np.ndarray]]
run(images: list[np.ndarray], *, preprocessed: bool | None = None, postprocess: bool | None = None, no_copy: bool | None = None, verbose: bool | None = None) → list[np.ndarray] | list[list[np.ndarray]]
run(images: np.ndarray, *, preprocessed: bool | None = None, postprocess: Literal[False], no_copy: bool | None = None, verbose: bool | None = None) → list[np.ndarray]
run(images: np.ndarray, *, preprocessed: bool | None = None, postprocess: Literal[True] | None = None, no_copy: bool | None = None, verbose: bool | None = None) → list[np.ndarray]
run(images: np.ndarray, *, preprocessed: bool | None = None, postprocess: bool | None = None, no_copy: bool | None = None, verbose: bool | None = None) → list[np.ndarray]

Run the model on input.

Parameters:
  • images (np.ndarray | list[np.ndarray]) – A single image (HWC format) or list of images to run the model on.

  • preprocessed (bool, optional) – Whether or not the inputs have been preprocessed. If None, will preprocess inputs.

  • postprocess (bool, optional) – Whether or not to postprocess the outputs. If None, will postprocess outputs.

  • no_copy (bool, optional) – If True, the outputs will not be copied out from the cuda allocated host memory. Instead, the host memory will be returned directly. This memory WILL BE OVERWRITTEN INPLACE by future inferences. In special case where, preprocessing and postprocessing will occur during run and no_copy was not passed (is None), then no_copy will be used for preprocessing and inference stages.

  • verbose (bool, optional) – Whether or not to log additional information.

Returns:

For single image with postprocess=True: list[np.ndarray] (single image outputs). For batch with postprocess=True: list[list[np.ndarray]] (per-image outputs). For postprocess=False: list[np.ndarray] (raw outputs).

Return type:

list[np.ndarray] | list[list[np.ndarray]]

Raises:

ValueError – If preprocessed inputs are not a single batch tensor.

get_classifications(outputs: list[np.ndarray], top_k: int = 5, *, verbose: bool | None = None) → list[tuple[int, float]][source]
get_classifications(outputs: list[list[np.ndarray]], top_k: int = 5, *, verbose: bool | None = None) → list[list[tuple[int, float]]]

Get the classifications from postprocessed outputs.

Parameters:
  • outputs (list[np.ndarray] | list[list[np.ndarray]]) – For single image: list[np.ndarray] (single image’s postprocessed outputs). For batch: list[list[np.ndarray]] (postprocessed outputs per image).

  • top_k (int, optional) – The number of top predictions to return. Default is 5.

  • verbose (bool, optional) – Whether or not to log additional information.

Returns:

For single image: list[tuple[int, float]] (classifications for single image). For batch: list[list[tuple[int, float]]] (classifications per image).

Return type:

list[tuple[int, float]] | list[list[tuple[int, float]]]

end2end(images: np.ndarray, top_k: int = 5, *, verbose: bool | None = None) → list[tuple[int, float]][source]
end2end(images: list[np.ndarray], top_k: int = 5, *, verbose: bool | None = None) → list[list[tuple[int, float]]]

Perform end to end inference for a batch of images.

Equivalent to running preprocess, run, postprocess, and get_classifications in that order. Makes some memory transfer optimizations under the hood to improve performance.

Parameters:
  • images (np.ndarray | list[np.ndarray]) – A single image (HWC format) or list of images to perform inference with.

  • top_k (int, optional) – The number of top predictions to return. Default is 5.

  • verbose (bool, optional) – Whether or not to log additional information.

Returns:

For single image: list[tuple[int, float]] (classifications). For batch: list[list[tuple[int, float]]] (classifications per image).

Return type:

list[tuple[int, float]] | list[list[tuple[int, float]]]

Raises:
  • RuntimeError – If end2end_graph is enabled and image dimensions change after first call.

  • RuntimeError – If end2end_graph is enabled and CUDA graph capture fails.

class trtutils.image.DepthEstimator(engine_path: Path | str, warmup_iterations: int = 10, input_range: tuple[float, float] = (0.0, 1.0), preprocessor: str = 'trt', resize_method: str = 'linear', mean: tuple[float, float, float] | None = (0.485, 0.456, 0.406), std: tuple[float, float, float] | None = (0.229, 0.224, 0.225), dla_core: int | None = None, device: int | None = None, backend: str = 'auto', *, warmup: bool | None = None, pagelocked_mem: bool | None = None, unified_mem: bool | None = None, cuda_graph: bool | None = None, no_warn: bool | None = None, verbose: bool | None = None)[source]

Bases: ImageModel, DepthEstimatorInterface

Implementation of depth estimators.

postprocess(outputs: list[np.ndarray], *, no_copy: bool | None = None, verbose: bool | None = None) → list[list[np.ndarray]][source]

Postprocess the outputs.

Parameters:
  • outputs (list[np.ndarray]) – The raw outputs from the engine to postprocess.

  • no_copy (bool, optional) – If True, do not copy the data from the allocated memory. If the data is not copied, it WILL BE OVERWRITTEN INPLACE once new data is generated.

  • verbose (bool, optional) – Whether or not to log additional information.

Returns:

The postprocessed depth maps per image.

Return type:

list[list[np.ndarray]]

run(images: list[np.ndarray], *, preprocessed: bool | None = None, postprocess: Literal[False], no_copy: bool | None = None, verbose: bool | None = None) → list[np.ndarray][source]
run(images: list[np.ndarray], *, preprocessed: bool | None = None, postprocess: Literal[True] | None = None, no_copy: bool | None = None, verbose: bool | None = None) → list[list[np.ndarray]]
run(images: list[np.ndarray], *, preprocessed: bool | None = None, postprocess: bool | None = None, no_copy: bool | None = None, verbose: bool | None = None) → list[np.ndarray] | list[list[np.ndarray]]
run(images: np.ndarray, *, preprocessed: bool | None = None, postprocess: Literal[False], no_copy: bool | None = None, verbose: bool | None = None) → list[np.ndarray]
run(images: np.ndarray, *, preprocessed: bool | None = None, postprocess: Literal[True] | None = None, no_copy: bool | None = None, verbose: bool | None = None) → list[np.ndarray]
run(images: np.ndarray, *, preprocessed: bool | None = None, postprocess: bool | None = None, no_copy: bool | None = None, verbose: bool | None = None) → list[np.ndarray]

Run the model on input.

Parameters:
  • images (np.ndarray | list[np.ndarray]) – A single image (HWC format) or list of images to run the model on.

  • preprocessed (bool, optional) – Whether or not the inputs have been preprocessed. If None, will preprocess inputs.

  • postprocess (bool, optional) – Whether or not to postprocess the outputs. If None, will postprocess outputs.

  • no_copy (bool, optional) – If True, the outputs will not be copied out from the cuda allocated host memory. Instead, the host memory will be returned directly. This memory WILL BE OVERWRITTEN INPLACE by future inferences. In special case where, preprocessing and postprocessing will occur during run and no_copy was not passed (is None), then no_copy will be used for preprocessing and inference stages.

  • verbose (bool, optional) – Whether or not to log additional information.

Returns:

For single image with postprocess=True: list[np.ndarray] (single image outputs). For batch with postprocess=True: list[list[np.ndarray]] (per-image outputs). For postprocess=False: list[np.ndarray] (raw outputs).

Return type:

list[np.ndarray] | list[list[np.ndarray]]

Raises:

ValueError – If preprocessed inputs are not a single batch tensor.

get_depth_maps(outputs: list[np.ndarray], *, verbose: bool | None = None) → np.ndarray[source]
get_depth_maps(outputs: list[list[np.ndarray]], *, verbose: bool | None = None) → list[np.ndarray]

Get the depth maps from postprocessed outputs.

Parameters:
  • outputs (list[np.ndarray] | list[list[np.ndarray]]) – For single image: list[np.ndarray] (single image’s postprocessed outputs). For batch: list[list[np.ndarray]] (postprocessed outputs per image).

  • verbose (bool, optional) – Whether or not to log additional information.

Returns:

For single image: np.ndarray (depth map of shape (1, H, W)). For batch: list[np.ndarray] (depth map per image).

Return type:

np.ndarray | list[np.ndarray]

end2end(images: np.ndarray, *, verbose: bool | None = None) → np.ndarray[source]
end2end(images: list[np.ndarray], *, verbose: bool | None = None) → list[np.ndarray]

Perform end to end inference for a batch of images.

Equivalent to running preprocess, run, postprocess, and get_depth_maps in that order. Makes some memory transfer optimizations under the hood to improve performance.

Parameters:
  • images (np.ndarray | list[np.ndarray]) – A single image (HWC format) or list of images to perform inference with.

  • verbose (bool, optional) – Whether or not to log additional information.

Returns:

For single image: np.ndarray (depth map of shape (1, H, W)). For batch: list[np.ndarray] (depth map per image).

Return type:

np.ndarray | list[np.ndarray]

Raises:
  • RuntimeError – If end2end_graph is enabled and image dimensions change after first call.

  • RuntimeError – If end2end_graph is enabled and CUDA graph capture fails.

class trtutils.image.Detector(engine_path: Path | str, warmup_iterations: int = 10, input_range: tuple[float, float] = (0.0, 1.0), preprocessor: str = 'trt', resize_method: str = 'letterbox', conf_thres: float = 0.1, nms_iou_thres: float = 0.5, mean: tuple[float, float, float] | None = None, std: tuple[float, float, float] | None = None, input_schema: InputSchema | str | None = None, output_schema: OutputSchema | str | None = None, dla_core: int | None = None, device: int | None = None, backend: str = 'auto', *, warmup: bool | None = None, pagelocked_mem: bool | None = None, unified_mem: bool | None = None, cuda_graph: bool | None = None, extra_nms: bool | None = None, agnostic_nms: bool | None = None, no_warn: bool | None = None, verbose: bool | None = None)[source]

Bases: ImageModel, DetectorInterface

Implementation of object detectors.

property input_schema: InputSchema

Get the input schema used by this detector.

property output_schema: OutputSchema

Get the output schema used by this detector.

postprocess(outputs: list[np.ndarray], ratios: list[tuple[float, float]], padding: list[tuple[float, float]], conf_thres: float | None = None, *, no_copy: bool | None = None, verbose: bool | None = None) → list[list[np.ndarray]][source]

Postprocess the outputs.

Parameters:
  • outputs (list[np.ndarray]) – The raw outputs from the engine to postprocess.

  • ratios (list[tuple[float, float]]) – The rescale ratios used during preprocessing for each image.

  • padding (list[tuple[float, float]]) – The padding values used during preprocessing for each image.

  • conf_thres (float, optional) – The confidence threshold to filter detections by. If not passed, will use value from constructor.

  • no_copy (bool, optional) – If True, do not copy the data from the allocated memory. If the data is not copied, it WILL BE OVERWRITTEN INPLACE once new data is generated.

  • verbose (bool, optional) – Whether or not to log additional information.

Returns:

The postprocessed outputs per image, each containing [bboxes, scores, class_ids].

Return type:

list[list[np.ndarray]]

run(images: list[ndarray], ratios: list[tuple[float, float]] | None = None, padding: list[tuple[float, float]] | None = None, conf_thres: float | None = None, *, preprocessed: bool | None = None, postprocess: Literal[False], no_copy: bool | None = None, verbose: bool | None = None) → list[ndarray][source]
run(images: list[ndarray], ratios: list[tuple[float, float]] | None = None, padding: list[tuple[float, float]] | None = None, conf_thres: float | None = None, *, preprocessed: bool | None = None, postprocess: Literal[True] | None = None, no_copy: bool | None = None, verbose: bool | None = None) → list[list[ndarray]]
run(images: list[ndarray], ratios: list[tuple[float, float]] | None = None, padding: list[tuple[float, float]] | None = None, conf_thres: float | None = None, *, preprocessed: bool | None = None, postprocess: bool | None = None, no_copy: bool | None = None, verbose: bool | None = None) → list[ndarray] | list[list[ndarray]]
run(images: ndarray, ratios: tuple[float, float] | None = None, padding: tuple[float, float] | None = None, conf_thres: float | None = None, *, preprocessed: bool | None = None, postprocess: Literal[False], no_copy: bool | None = None, verbose: bool | None = None) → list[ndarray]
run(images: ndarray, ratios: tuple[float, float] | None = None, padding: tuple[float, float] | None = None, conf_thres: float | None = None, *, preprocessed: bool | None = None, postprocess: Literal[True] | None = None, no_copy: bool | None = None, verbose: bool | None = None) → list[ndarray]
run(images: ndarray, ratios: tuple[float, float] | None = None, padding: tuple[float, float] | None = None, conf_thres: float | None = None, *, preprocessed: bool | None = None, postprocess: bool | None = None, no_copy: bool | None = None, verbose: bool | None = None) → list[ndarray]

Run the model on input.

Parameters:
  • images (np.ndarray | list[np.ndarray]) – A single image (HWC format) or list of images to run the model on.

  • ratios (tuple[float, float] | list[tuple[float, float]], optional) – The ratios generated during preprocessing. For single image, pass tuple. For batch, pass list.

  • padding (tuple[float, float] | list[tuple[float, float]], optional) – The padding values used during preprocessing. For single image, pass tuple. For batch, pass list.

  • conf_thres (float, optional) – Optional confidence threshold to filter detections via during postprocessing.

  • preprocessed (bool, optional) – Whether or not the inputs have been preprocessed. If None, will preprocess inputs.

  • postprocess (bool, optional) – Whether or not to postprocess the outputs. If None, will postprocess outputs. If postprocessing will occur and the inputs were passed already preprocessed, then the ratios and padding must be passed for postprocessing.

  • no_copy (bool, optional) – If True, the outputs will not be copied out from the cuda allocated host memory. Instead, the host memory will be returned directly. This memory WILL BE OVERWRITTEN INPLACE by future inferences. In special case where, preprocessing and postprocessing will occur during run and no_copy was not passed (is None), then no_copy will be used for preprocessing and inference stages.

  • verbose (bool, optional) – Whether or not to log additional information.

Returns:

For single image with postprocess=True: list[np.ndarray] (single image outputs). For batch with postprocess=True: list[list[np.ndarray]] (per-image outputs). For postprocess=False: list[np.ndarray] (raw outputs).

Return type:

list[np.ndarray] | list[list[np.ndarray]]

Raises:
  • RuntimeError – If postprocessing is running, but ratios/padding not found

  • ValueError – If preprocessed inputs are not a single batch tensor.

get_detections(outputs: list[ndarray], conf_thres: float | None = None, nms_iou_thres: float | None = None, *, extra_nms: bool | None = None, agnostic_nms: bool | None = None, verbose: bool | None = None) → list[tuple[tuple[int, int, int, int], float, int]][source]
get_detections(outputs: list[list[ndarray]], conf_thres: float | None = None, nms_iou_thres: float | None = None, *, extra_nms: bool | None = None, agnostic_nms: bool | None = None, verbose: bool | None = None) → list[list[tuple[tuple[int, int, int, int], float, int]]]

Get the bounding boxes from postprocessed outputs.

Parameters:
  • outputs (list[np.ndarray] | list[list[np.ndarray]]) – For single image: list[np.ndarray] (single image’s postprocessed outputs). For batch: list[list[np.ndarray]] (postprocessed outputs per image).

  • conf_thres (float, optional) – The confidence threshold with which to retrieve bounding boxes. By default None, which will use value passed during initialization.

  • nms_iou_thres (float) – The IOU threshold to use during the optional/additional NMS operation. By default, None which will use value provided during initialization.

  • extra_nms (bool, optional) – Whether or not to perform an additional NMS operation. By default None, which will use value provided during initialization.

  • agnostic_nms (bool, optional) – Whether or not to perform class-agnostic NMS for the optional/additional operation. By default None, which will use value provided during initialization.

  • verbose (bool, optional) – Whether or not to log additional information.

Returns:

For single image: list[tuple[…]] (detections for single image). For batch: list[list[tuple[…]]] (detections per image).

Return type:

list[tuple[…]] | list[list[tuple[…]]]

end2end(images: ndarray, conf_thres: float | None = None, nms_iou_thres: float | None = None, *, extra_nms: bool | None = None, agnostic_nms: bool | None = None, verbose: bool | None = None) → list[tuple[tuple[int, int, int, int], float, int]][source]
end2end(images: list[ndarray], conf_thres: float | None = None, nms_iou_thres: float | None = None, *, extra_nms: bool | None = None, agnostic_nms: bool | None = None, verbose: bool | None = None) → list[list[tuple[tuple[int, int, int, int], float, int]]]

Perform end to end inference for a batch of images.

Equivalent to running preprocess, run, postprocess, and get_detections in that order. Makes some memory transfer optimizations under the hood to improve performance.

Parameters:
  • images (np.ndarray | list[np.ndarray]) – A single image (HWC format) or list of images to perform inference with.

  • conf_thres (float, optional) – The confidence threshold with which to retrieve bounding boxes. By default None.

  • nms_iou_thres (float) – The IOU threshold to use during the optional/additional NMS operation. By default, None which will use value provided during initialization.

  • extra_nms (bool, optional) – Whether or not to perform an additional NMS operation. By default None, which will use value provided during initialization.

  • agnostic_nms (bool, optional) – Whether or not to perform class-agnostic NMS for the optional/additional operation. By default None, which will use value provided during initialization.

  • verbose (bool, optional) – Whether or not to log additional information.

Returns:

For single image: list[tuple[…]] (detections). For batch: list[list[tuple[…]]] (detections per image).

Return type:

list[tuple[…]] | list[list[tuple[…]]]

Raises:
  • RuntimeError – If the orig_image_size buffer is not valid

  • RuntimeError – If the scale_factor buffer is not valid

  • RuntimeError – If end2end_graph is enabled and image dimensions change after first call.

  • RuntimeError – If end2end_graph is enabled and CUDA graph capture fails.

class trtutils.image.HandInteractionDetector(engine_path: Path | str, warmup_iterations: int = 10, input_range: tuple[float, float] = (0.0, 1.0), preprocessor: str = 'trt', resize_method: str = 'letterbox', conf_thres: float = 0.3, pair_thres: float = 0.5, second_pair_thres: float | None = None, nms_iou_thres: float = 0.5, mean: tuple[float, float, float] | None = None, std: tuple[float, float, float] | None = None, dla_core: int | None = None, device: int | None = None, backend: str = 'auto', *, warmup: bool | None = None, pagelocked_mem: bool | None = None, unified_mem: bool | None = None, cuda_graph: bool | None = None, no_warn: bool | None = None, verbose: bool | None = None)[source]

Bases: ImageModel, HandInteractionDetectorInterface

Implementation of hand-object interaction detectors.

Wraps engines for hand-object interaction model families (e.g. HOI-DETR, Hands23) that follow a unified, fixed-K output contract: boxes (B,K,4) xyxy in network-input pixels, scores (B,K), labels (B,K) (0 hand, 1 object, 2 second object), pair_probs (B,K,K,C) where [...,0] is the no-interaction class, and an optional side (B,K) (0 left, 1 right, valid where label is hand). All model-specific heads (pairing MLP, attribute heads, touch gating) are expected to run in-graph so a single postprocessor handles every model in the family. Letterbox padding is centered and is undone by the postprocessor using the ratios/padding returned from preprocessing.

postprocess(outputs: list[np.ndarray], ratios: list[tuple[float, float]], padding: list[tuple[float, float]], conf_thres: float | None = None, *, no_copy: bool | None = None, verbose: bool | None = None) → list[list[np.ndarray]][source]

Postprocess the outputs.

Parameters:
  • outputs (list[np.ndarray]) – The raw outputs from the engine to postprocess.

  • ratios (list[tuple[float, float]]) – The rescale ratios used during preprocessing for each image.

  • padding (list[tuple[float, float]]) – The padding values used during preprocessing for each image.

  • conf_thres (float, optional) – The confidence threshold to filter candidates by. If not passed, will use the value from the constructor.

  • no_copy (bool, optional) – Kept for interface symmetry with the other postprocessors. NMS-based indexing always allocates new arrays.

  • verbose (bool, optional) – Whether or not to log additional information.

Returns:

The postprocessed outputs per image: [bboxes, scores, labels, pair_probs, side].

Return type:

list[list[np.ndarray]]

run(images: list[np.ndarray], ratios: list[tuple[float, float]] | None = None, padding: list[tuple[float, float]] | None = None, conf_thres: float | None = None, *, preprocessed: bool | None = None, postprocess: Literal[False], no_copy: bool | None = None, verbose: bool | None = None) → list[np.ndarray][source]
run(images: list[np.ndarray], ratios: list[tuple[float, float]] | None = None, padding: list[tuple[float, float]] | None = None, conf_thres: float | None = None, *, preprocessed: bool | None = None, postprocess: Literal[True] | None = None, no_copy: bool | None = None, verbose: bool | None = None) → list[list[np.ndarray]]
run(images: list[np.ndarray], ratios: list[tuple[float, float]] | None = None, padding: list[tuple[float, float]] | None = None, conf_thres: float | None = None, *, preprocessed: bool | None = None, postprocess: bool | None = None, no_copy: bool | None = None, verbose: bool | None = None) → list[np.ndarray] | list[list[np.ndarray]]
run(images: np.ndarray, ratios: tuple[float, float] | None = None, padding: tuple[float, float] | None = None, conf_thres: float | None = None, *, preprocessed: bool | None = None, postprocess: Literal[False], no_copy: bool | None = None, verbose: bool | None = None) → list[np.ndarray]
run(images: np.ndarray, ratios: tuple[float, float] | None = None, padding: tuple[float, float] | None = None, conf_thres: float | None = None, *, preprocessed: bool | None = None, postprocess: Literal[True] | None = None, no_copy: bool | None = None, verbose: bool | None = None) → list[np.ndarray]
run(images: np.ndarray, ratios: tuple[float, float] | None = None, padding: tuple[float, float] | None = None, conf_thres: float | None = None, *, preprocessed: bool | None = None, postprocess: bool | None = None, no_copy: bool | None = None, verbose: bool | None = None) → list[np.ndarray]

Run the model on input.

Parameters:
  • images (np.ndarray | list[np.ndarray]) – A single image (HWC format) or list of images to run the model on.

  • ratios (tuple[float, float] | list[tuple[float, float]], optional) – The ratios generated during preprocessing. For single image, pass tuple. For batch, pass list.

  • padding (tuple[float, float] | list[tuple[float, float]], optional) – The padding values used during preprocessing. For single image, pass tuple. For batch, pass list.

  • conf_thres (float, optional) – Optional confidence threshold to filter candidates by during postprocessing.

  • preprocessed (bool, optional) – Whether or not the inputs have been preprocessed. If None, will preprocess inputs.

  • postprocess (bool, optional) – Whether or not to postprocess the outputs. If None, will postprocess outputs. If postprocessing will occur and the inputs were passed already preprocessed, then the ratios and padding must be passed for postprocessing.

  • no_copy (bool, optional) – If True, the outputs will not be copied out from the cuda allocated host memory. Instead, the host memory will be returned directly. This memory WILL BE OVERWRITTEN INPLACE by future inferences. In special case where, preprocessing and postprocessing will occur during run and no_copy was not passed (is None), then no_copy will be used for preprocessing and inference stages.

  • verbose (bool, optional) – Whether or not to log additional information.

Returns:

For single image with postprocess=True: list[np.ndarray] (single image outputs). For batch with postprocess=True: list[list[np.ndarray]] (per-image outputs). For postprocess=False: list[np.ndarray] (raw outputs).

Return type:

list[np.ndarray] | list[list[np.ndarray]]

Raises:
  • ValueError – If preprocessed inputs are not a single batch tensor.

  • RuntimeError – If postprocessing is requested for already-preprocessed inputs without passing ratios/padding.

get_interactions(outputs: list[np.ndarray], pair_thres: float | None = None, second_pair_thres: float | None = None, *, verbose: bool | None = None) → list[HandInteraction][source]
get_interactions(outputs: list[list[np.ndarray]], pair_thres: float | None = None, second_pair_thres: float | None = None, *, verbose: bool | None = None) → list[list[HandInteraction]]

Get the hand-object interactions from postprocessed outputs.

Parameters:
  • outputs (list[np.ndarray] | list[list[np.ndarray]]) – For single image: list[np.ndarray] (single image’s postprocessed outputs). For batch: list[list[np.ndarray]] (postprocessed outputs per image).

  • pair_thres (float, optional) – The minimum link probability to link a hand to a first object. By default None, which uses the value from the constructor.

  • second_pair_thres (float, optional) – The minimum link probability to link a first object to a second object. By default None, which uses the value from the constructor.

  • verbose (bool, optional) – Whether or not to log additional information.

Returns:

For single image: list[HandInteraction] (interactions for single image). For batch: list[list[HandInteraction]] (interactions per image).

Return type:

list[HandInteraction] | list[list[HandInteraction]]

end2end(images: np.ndarray, *, conf_thres: float | None = None, pair_thres: float | None = None, second_pair_thres: float | None = None, verbose: bool | None = None) → list[HandInteraction][source]
end2end(images: list[np.ndarray], *, conf_thres: float | None = None, pair_thres: float | None = None, second_pair_thres: float | None = None, verbose: bool | None = None) → list[list[HandInteraction]]

Perform end to end inference for a batch of images.

Equivalent to running preprocess, run, postprocess, and get_interactions in that order. Makes some memory transfer optimizations under the hood to improve performance.

Parameters:
  • images (np.ndarray | list[np.ndarray]) – A single image (HWC format) or list of images to perform inference with.

  • conf_thres (float, optional) – The confidence threshold to filter candidates by. By default None, which uses the value from the constructor.

  • pair_thres (float, optional) – The minimum link probability to link a hand to a first object. By default None, which uses the value from the constructor.

  • second_pair_thres (float, optional) – The minimum link probability to link a first object to a second object. By default None, which uses the value from the constructor.

  • verbose (bool, optional) – Whether or not to log additional information.

Returns:

For single image: list[HandInteraction] (interactions). For batch: list[list[HandInteraction]] (interactions per image).

Return type:

list[HandInteraction] | list[list[HandInteraction]]

Raises:
  • RuntimeError – If end2end_graph is enabled and image dimensions change after first call.

  • RuntimeError – If end2end_graph is enabled and CUDA graph capture fails.

class trtutils.image.ImageModel(engine_path: Path | str, warmup_iterations: int = 10, input_range: tuple[float, float] = (0.0, 1.0), preprocessor: str = 'trt', resize_method: str = 'linear', mean: tuple[float, float, float] | None = None, std: tuple[float, float, float] | None = None, dla_core: int | None = None, device: int | None = None, backend: str = 'auto', *, warmup: bool | None = None, pagelocked_mem: bool | None = None, unified_mem: bool | None = None, cuda_graph: bool | None = None, no_warn: bool | None = None, verbose: bool | None = None)[source]

Bases: object

Abstract base class for image models.

property engine: TRTEngine

Get the underlying TRTEngine.

property name: str

Get the name of the engine.

property input_shape: tuple[int, int]

Get the width, height input shape.

property dtype: np.dtype

Get the dtype required by the model.

property batch_size: int

Get the batch size of the model.

property is_dynamic_batch: bool

Check if model has dynamic batch size.

update_input_range(input_range: tuple[float, float]) → None[source]

Update the input range of the model.

This will re-create all preprocessors with the new input range. Only preprocessors which have been created will be re-created.

Parameters:

input_range (tuple[float, float]) – The new input range.

update_mean_std(mean: tuple[float, float, float], std: tuple[float, float, float]) → None[source]

Update the mean and standard deviation of the model.

Parameters:
Raises:

ValueError – If the mean or std is not a tuple of 3 floats

get_random_input() → list[np.ndarray][source]

Generate random images for the model.

Returns:

A list containing one random image.

Return type:

list[np.ndarray]

mock_run(images: list[np.ndarray] | None = None) → list[np.ndarray][source]

Mock an execution of the model.

Parameters:

images (list[np.ndarray], optional) – Optional batch of images to use for execution. If None, random data will be generated.

Returns:

The raw outputs of the model.

Return type:

list[np.ndarray]

preprocess(images: ndarray, resize: str | None = None, method: str | None = None, *, no_copy: bool | None = None, verbose: bool | None = None) → tuple[ndarray, list[tuple[float, float]], list[tuple[float, float]]][source]
preprocess(images: list[ndarray], resize: str | None = None, method: str | None = None, *, no_copy: bool | None = None, verbose: bool | None = None) → tuple[ndarray, list[tuple[float, float]], list[tuple[float, float]]]

Preprocess the input images.

Parameters:
  • images (np.ndarray | list[np.ndarray]) – A single image (HWC format) or list of images to preprocess.

  • resize (str) – The method to resize the images with. Options are [letterbox, linear]. By default None, which will use the value passed during initialization.

  • method (str, optional) – The underlying preprocessor to use. Options are ‘cpu’, ‘cuda’, or ‘trt’. By default None, which will use the preprocessor stated in the constructor.

  • no_copy (bool, optional) – If True and using CUDA, do not copy the data from the allocated memory. If the data is not copied, it WILL BE OVERWRITTEN INPLACE once new data is generated.

  • verbose (bool, optional) – Whether or not to log additional information.

Returns:

The preprocessed batch tensor, list of ratios per image, and list of padding per image.

Return type:

tuple[np.ndarray, list[tuple[float, float]], list[tuple[float, float]]]