trtutils.image package¶
Subpackages¶
- trtutils.image.postprocessors package
- Module contents
- Functions
get_classifications()get_depth_maps()get_detections()get_interactions()postprocess_classifications()postprocess_depth()postprocess_detr()postprocess_detr_lbs()postprocess_efficient_nms()postprocess_hand_interactions()postprocess_rfdetr()postprocess_rtdetrv3()postprocess_yolov10()
- Module contents
- trtutils.image.preprocessors package
- trtutils.image.sahi package
Submodules¶
- trtutils.image.interfaces module
- Classes
ClassifierInterfaceDepthEstimatorInterfaceDepthEstimatorInterface.engineDepthEstimatorInterface.nameDepthEstimatorInterface.input_shapeDepthEstimatorInterface.dtypeDepthEstimatorInterface.preprocess()DepthEstimatorInterface.postprocess()DepthEstimatorInterface.run()DepthEstimatorInterface.get_depth_maps()DepthEstimatorInterface.end2end()
HandInteractionDetectorInterfaceHandInteractionDetectorInterface.engineHandInteractionDetectorInterface.nameHandInteractionDetectorInterface.input_shapeHandInteractionDetectorInterface.dtypeHandInteractionDetectorInterface.preprocess()HandInteractionDetectorInterface.postprocess()HandInteractionDetectorInterface.run()HandInteractionDetectorInterface.get_interactions()HandInteractionDetectorInterface.end2end()
DetectorInterfaceDetectorInterface.engineDetectorInterface.nameDetectorInterface.input_shapeDetectorInterface.dtypeDetectorInterface.input_schemaDetectorInterface.output_schemaDetectorInterface.preprocess()DetectorInterface.postprocess()DetectorInterface.run()DetectorInterface.get_detections()DetectorInterface.end2end()
- trtutils.image.kernels module
- trtutils.image.onnx_models module
Module contents¶
Utilities for using TensorRT on images.
Submodules¶
kernelsKernels for image processing with TensorRT.
preprocessorsPreprocessors for images.
postprocessorsPostprocessors for images.
interfacesInterfaces for image models.
onnx_modelsBase ONNX models for creating ‘micro-engines’ for image processing.
sahiSAHI (Slicing Aided Hyper Inference) for object detection.
Classes¶
ClassiferWrapper around classification models.
DepthEstimatorWrapper around depth estimation models.
DetectorWrapper around detection models.
HandInteractionDetectorWrapper around hand-object interaction models.
SAHISAHI wrapper for slicing aided inference.
ImageModelBase class for models which process images.
- class trtutils.image.SAHI(detector: DetectorInterface, slice_size: tuple[int, int] | None = None, slice_overlap: tuple[float, float] = (0.2, 0.2), iou_threshold: float = 0.5, *, agnostic_nms: bool = False, verbose: bool = False)[source]¶
Bases:
objectSimple implementation of SAHI.
- end2end(image: np.ndarray, conf_thres: float | None = None, nms_iou_thres: float | None = None, *, extra_nms: bool | None = None, agnostic_nms: bool | None = None, verbose: bool | None = None) list[tuple[tuple[int, int, int, int], float, int]][source]¶
Perform end to end inference using detection model and SAHI.
- Parameters:
image (np.ndarray) – The image to perform inference with.
conf_thres (float, optional) – The confidence threshold with which to retrieve bounding boxes. By default None
nms_iou_thres (float) – The IOU threshold to use during the optional/additional NMS operation. By default, None which will use value provided during initialization.
extra_nms (bool, optional) – Whether or not to perform an additional NMS operation. By default None, which will use value provided during initialization.
agnostic_nms (bool, optional) – Whether or not to perform class-agnostic NMS for the optional/additional operation. By default None, which will use value provided during initialization.
verbose (bool, optional) – Whether or not to log additional information.
- Returns:
The detections where each entry is bbox, conf, class_id
- Return type:
- class trtutils.image.Classifier(engine_path: Path | str, warmup_iterations: int = 10, input_range: tuple[float, float] = (0.0, 1.0), preprocessor: str = 'trt', resize_method: str = 'linear', mean: tuple[float, float, float] | None = None, std: tuple[float, float, float] | None = None, dla_core: int | None = None, device: int | None = None, backend: str = 'auto', *, warmup: bool | None = None, pagelocked_mem: bool | None = None, unified_mem: bool | None = None, cuda_graph: bool | None = None, no_warn: bool | None = None, verbose: bool | None = None)[source]¶
Bases:
ImageModel,ClassifierInterfaceImplementation of image classifiers.
- postprocess(outputs: list[np.ndarray], *, no_copy: bool | None = None, verbose: bool | None = None) list[list[np.ndarray]][source]¶
Postprocess the outputs.
- Parameters:
outputs (list[np.ndarray]) – The raw outputs from the engine to postprocess.
no_copy (bool, optional) – If True, do not copy the data from the allocated memory. If the data is not copied, it WILL BE OVERWRITTEN INPLACE once new data is generated.
verbose (bool, optional) – Whether or not to log additional information.
- Returns:
The postprocessed outputs per image.
- Return type:
- run(images: list[np.ndarray], *, preprocessed: bool | None = None, postprocess: Literal[False], no_copy: bool | None = None, verbose: bool | None = None) list[np.ndarray][source]¶
- run(images: list[np.ndarray], *, preprocessed: bool | None = None, postprocess: Literal[True] | None = None, no_copy: bool | None = None, verbose: bool | None = None) list[list[np.ndarray]]
- run(images: list[np.ndarray], *, preprocessed: bool | None = None, postprocess: bool | None = None, no_copy: bool | None = None, verbose: bool | None = None) list[np.ndarray] | list[list[np.ndarray]]
- run(images: np.ndarray, *, preprocessed: bool | None = None, postprocess: Literal[False], no_copy: bool | None = None, verbose: bool | None = None) list[np.ndarray]
- run(images: np.ndarray, *, preprocessed: bool | None = None, postprocess: Literal[True] | None = None, no_copy: bool | None = None, verbose: bool | None = None) list[np.ndarray]
- run(images: np.ndarray, *, preprocessed: bool | None = None, postprocess: bool | None = None, no_copy: bool | None = None, verbose: bool | None = None) list[np.ndarray]
Run the model on input.
- Parameters:
images (np.ndarray | list[np.ndarray]) – A single image (HWC format) or list of images to run the model on.
preprocessed (bool, optional) – Whether or not the inputs have been preprocessed. If None, will preprocess inputs.
postprocess (bool, optional) – Whether or not to postprocess the outputs. If None, will postprocess outputs.
no_copy (bool, optional) – If True, the outputs will not be copied out from the cuda allocated host memory. Instead, the host memory will be returned directly. This memory WILL BE OVERWRITTEN INPLACE by future inferences. In special case where, preprocessing and postprocessing will occur during run and no_copy was not passed (is None), then no_copy will be used for preprocessing and inference stages.
verbose (bool, optional) – Whether or not to log additional information.
- Returns:
For single image with postprocess=True: list[np.ndarray] (single image outputs). For batch with postprocess=True: list[list[np.ndarray]] (per-image outputs). For postprocess=False: list[np.ndarray] (raw outputs).
- Return type:
- Raises:
ValueError – If preprocessed inputs are not a single batch tensor.
- get_classifications(outputs: list[np.ndarray], top_k: int = 5, *, verbose: bool | None = None) list[tuple[int, float]][source]¶
- get_classifications(outputs: list[list[np.ndarray]], top_k: int = 5, *, verbose: bool | None = None) list[list[tuple[int, float]]]
Get the classifications from postprocessed outputs.
- Parameters:
outputs (list[np.ndarray] | list[list[np.ndarray]]) – For single image: list[np.ndarray] (single image’s postprocessed outputs). For batch: list[list[np.ndarray]] (postprocessed outputs per image).
top_k (int, optional) – The number of top predictions to return. Default is 5.
verbose (bool, optional) – Whether or not to log additional information.
- Returns:
For single image: list[tuple[int, float]] (classifications for single image). For batch: list[list[tuple[int, float]]] (classifications per image).
- Return type:
- end2end(images: np.ndarray, top_k: int = 5, *, verbose: bool | None = None) list[tuple[int, float]][source]¶
- end2end(images: list[np.ndarray], top_k: int = 5, *, verbose: bool | None = None) list[list[tuple[int, float]]]
Perform end to end inference for a batch of images.
Equivalent to running preprocess, run, postprocess, and get_classifications in that order. Makes some memory transfer optimizations under the hood to improve performance.
- Parameters:
- Returns:
For single image: list[tuple[int, float]] (classifications). For batch: list[list[tuple[int, float]]] (classifications per image).
- Return type:
- Raises:
RuntimeError – If end2end_graph is enabled and image dimensions change after first call.
RuntimeError – If end2end_graph is enabled and CUDA graph capture fails.
- class trtutils.image.DepthEstimator(engine_path: Path | str, warmup_iterations: int = 10, input_range: tuple[float, float] = (0.0, 1.0), preprocessor: str = 'trt', resize_method: str = 'linear', mean: tuple[float, float, float] | None = (0.485, 0.456, 0.406), std: tuple[float, float, float] | None = (0.229, 0.224, 0.225), dla_core: int | None = None, device: int | None = None, backend: str = 'auto', *, warmup: bool | None = None, pagelocked_mem: bool | None = None, unified_mem: bool | None = None, cuda_graph: bool | None = None, no_warn: bool | None = None, verbose: bool | None = None)[source]¶
Bases:
ImageModel,DepthEstimatorInterfaceImplementation of depth estimators.
- postprocess(outputs: list[np.ndarray], *, no_copy: bool | None = None, verbose: bool | None = None) list[list[np.ndarray]][source]¶
Postprocess the outputs.
- Parameters:
outputs (list[np.ndarray]) – The raw outputs from the engine to postprocess.
no_copy (bool, optional) – If True, do not copy the data from the allocated memory. If the data is not copied, it WILL BE OVERWRITTEN INPLACE once new data is generated.
verbose (bool, optional) – Whether or not to log additional information.
- Returns:
The postprocessed depth maps per image.
- Return type:
- run(images: list[np.ndarray], *, preprocessed: bool | None = None, postprocess: Literal[False], no_copy: bool | None = None, verbose: bool | None = None) list[np.ndarray][source]¶
- run(images: list[np.ndarray], *, preprocessed: bool | None = None, postprocess: Literal[True] | None = None, no_copy: bool | None = None, verbose: bool | None = None) list[list[np.ndarray]]
- run(images: list[np.ndarray], *, preprocessed: bool | None = None, postprocess: bool | None = None, no_copy: bool | None = None, verbose: bool | None = None) list[np.ndarray] | list[list[np.ndarray]]
- run(images: np.ndarray, *, preprocessed: bool | None = None, postprocess: Literal[False], no_copy: bool | None = None, verbose: bool | None = None) list[np.ndarray]
- run(images: np.ndarray, *, preprocessed: bool | None = None, postprocess: Literal[True] | None = None, no_copy: bool | None = None, verbose: bool | None = None) list[np.ndarray]
- run(images: np.ndarray, *, preprocessed: bool | None = None, postprocess: bool | None = None, no_copy: bool | None = None, verbose: bool | None = None) list[np.ndarray]
Run the model on input.
- Parameters:
images (np.ndarray | list[np.ndarray]) – A single image (HWC format) or list of images to run the model on.
preprocessed (bool, optional) – Whether or not the inputs have been preprocessed. If None, will preprocess inputs.
postprocess (bool, optional) – Whether or not to postprocess the outputs. If None, will postprocess outputs.
no_copy (bool, optional) – If True, the outputs will not be copied out from the cuda allocated host memory. Instead, the host memory will be returned directly. This memory WILL BE OVERWRITTEN INPLACE by future inferences. In special case where, preprocessing and postprocessing will occur during run and no_copy was not passed (is None), then no_copy will be used for preprocessing and inference stages.
verbose (bool, optional) – Whether or not to log additional information.
- Returns:
For single image with postprocess=True: list[np.ndarray] (single image outputs). For batch with postprocess=True: list[list[np.ndarray]] (per-image outputs). For postprocess=False: list[np.ndarray] (raw outputs).
- Return type:
- Raises:
ValueError – If preprocessed inputs are not a single batch tensor.
- get_depth_maps(outputs: list[np.ndarray], *, verbose: bool | None = None) np.ndarray[source]¶
- get_depth_maps(outputs: list[list[np.ndarray]], *, verbose: bool | None = None) list[np.ndarray]
Get the depth maps from postprocessed outputs.
- Parameters:
- Returns:
For single image: np.ndarray (depth map of shape (1, H, W)). For batch: list[np.ndarray] (depth map per image).
- Return type:
np.ndarray | list[np.ndarray]
- end2end(images: np.ndarray, *, verbose: bool | None = None) np.ndarray[source]¶
- end2end(images: list[np.ndarray], *, verbose: bool | None = None) list[np.ndarray]
Perform end to end inference for a batch of images.
Equivalent to running preprocess, run, postprocess, and get_depth_maps in that order. Makes some memory transfer optimizations under the hood to improve performance.
- Parameters:
- Returns:
For single image: np.ndarray (depth map of shape (1, H, W)). For batch: list[np.ndarray] (depth map per image).
- Return type:
np.ndarray | list[np.ndarray]
- Raises:
RuntimeError – If end2end_graph is enabled and image dimensions change after first call.
RuntimeError – If end2end_graph is enabled and CUDA graph capture fails.
- class trtutils.image.Detector(engine_path: Path | str, warmup_iterations: int = 10, input_range: tuple[float, float] = (0.0, 1.0), preprocessor: str = 'trt', resize_method: str = 'letterbox', conf_thres: float = 0.1, nms_iou_thres: float = 0.5, mean: tuple[float, float, float] | None = None, std: tuple[float, float, float] | None = None, input_schema: InputSchema | str | None = None, output_schema: OutputSchema | str | None = None, dla_core: int | None = None, device: int | None = None, backend: str = 'auto', *, warmup: bool | None = None, pagelocked_mem: bool | None = None, unified_mem: bool | None = None, cuda_graph: bool | None = None, extra_nms: bool | None = None, agnostic_nms: bool | None = None, no_warn: bool | None = None, verbose: bool | None = None)[source]¶
Bases:
ImageModel,DetectorInterfaceImplementation of object detectors.
- property input_schema: InputSchema¶
Get the input schema used by this detector.
- property output_schema: OutputSchema¶
Get the output schema used by this detector.
- postprocess(outputs: list[np.ndarray], ratios: list[tuple[float, float]], padding: list[tuple[float, float]], conf_thres: float | None = None, *, no_copy: bool | None = None, verbose: bool | None = None) list[list[np.ndarray]][source]¶
Postprocess the outputs.
- Parameters:
outputs (list[np.ndarray]) – The raw outputs from the engine to postprocess.
ratios (list[tuple[float, float]]) – The rescale ratios used during preprocessing for each image.
padding (list[tuple[float, float]]) – The padding values used during preprocessing for each image.
conf_thres (float, optional) – The confidence threshold to filter detections by. If not passed, will use value from constructor.
no_copy (bool, optional) – If True, do not copy the data from the allocated memory. If the data is not copied, it WILL BE OVERWRITTEN INPLACE once new data is generated.
verbose (bool, optional) – Whether or not to log additional information.
- Returns:
The postprocessed outputs per image, each containing [bboxes, scores, class_ids].
- Return type:
- run(images: list[ndarray], ratios: list[tuple[float, float]] | None = None, padding: list[tuple[float, float]] | None = None, conf_thres: float | None = None, *, preprocessed: bool | None = None, postprocess: Literal[False], no_copy: bool | None = None, verbose: bool | None = None) list[ndarray][source]¶
- run(images: list[ndarray], ratios: list[tuple[float, float]] | None = None, padding: list[tuple[float, float]] | None = None, conf_thres: float | None = None, *, preprocessed: bool | None = None, postprocess: Literal[True] | None = None, no_copy: bool | None = None, verbose: bool | None = None) list[list[ndarray]]
- run(images: list[ndarray], ratios: list[tuple[float, float]] | None = None, padding: list[tuple[float, float]] | None = None, conf_thres: float | None = None, *, preprocessed: bool | None = None, postprocess: bool | None = None, no_copy: bool | None = None, verbose: bool | None = None) list[ndarray] | list[list[ndarray]]
- run(images: ndarray, ratios: tuple[float, float] | None = None, padding: tuple[float, float] | None = None, conf_thres: float | None = None, *, preprocessed: bool | None = None, postprocess: Literal[False], no_copy: bool | None = None, verbose: bool | None = None) list[ndarray]
- run(images: ndarray, ratios: tuple[float, float] | None = None, padding: tuple[float, float] | None = None, conf_thres: float | None = None, *, preprocessed: bool | None = None, postprocess: Literal[True] | None = None, no_copy: bool | None = None, verbose: bool | None = None) list[ndarray]
- run(images: ndarray, ratios: tuple[float, float] | None = None, padding: tuple[float, float] | None = None, conf_thres: float | None = None, *, preprocessed: bool | None = None, postprocess: bool | None = None, no_copy: bool | None = None, verbose: bool | None = None) list[ndarray]
Run the model on input.
- Parameters:
images (np.ndarray | list[np.ndarray]) – A single image (HWC format) or list of images to run the model on.
ratios (tuple[float, float] | list[tuple[float, float]], optional) – The ratios generated during preprocessing. For single image, pass tuple. For batch, pass list.
padding (tuple[float, float] | list[tuple[float, float]], optional) – The padding values used during preprocessing. For single image, pass tuple. For batch, pass list.
conf_thres (float, optional) – Optional confidence threshold to filter detections via during postprocessing.
preprocessed (bool, optional) – Whether or not the inputs have been preprocessed. If None, will preprocess inputs.
postprocess (bool, optional) – Whether or not to postprocess the outputs. If None, will postprocess outputs. If postprocessing will occur and the inputs were passed already preprocessed, then the ratios and padding must be passed for postprocessing.
no_copy (bool, optional) – If True, the outputs will not be copied out from the cuda allocated host memory. Instead, the host memory will be returned directly. This memory WILL BE OVERWRITTEN INPLACE by future inferences. In special case where, preprocessing and postprocessing will occur during run and no_copy was not passed (is None), then no_copy will be used for preprocessing and inference stages.
verbose (bool, optional) – Whether or not to log additional information.
- Returns:
For single image with postprocess=True: list[np.ndarray] (single image outputs). For batch with postprocess=True: list[list[np.ndarray]] (per-image outputs). For postprocess=False: list[np.ndarray] (raw outputs).
- Return type:
- Raises:
RuntimeError – If postprocessing is running, but ratios/padding not found
ValueError – If preprocessed inputs are not a single batch tensor.
- get_detections(outputs: list[ndarray], conf_thres: float | None = None, nms_iou_thres: float | None = None, *, extra_nms: bool | None = None, agnostic_nms: bool | None = None, verbose: bool | None = None) list[tuple[tuple[int, int, int, int], float, int]][source]¶
- get_detections(outputs: list[list[ndarray]], conf_thres: float | None = None, nms_iou_thres: float | None = None, *, extra_nms: bool | None = None, agnostic_nms: bool | None = None, verbose: bool | None = None) list[list[tuple[tuple[int, int, int, int], float, int]]]
Get the bounding boxes from postprocessed outputs.
- Parameters:
outputs (list[np.ndarray] | list[list[np.ndarray]]) – For single image: list[np.ndarray] (single image’s postprocessed outputs). For batch: list[list[np.ndarray]] (postprocessed outputs per image).
conf_thres (float, optional) – The confidence threshold with which to retrieve bounding boxes. By default None, which will use value passed during initialization.
nms_iou_thres (float) – The IOU threshold to use during the optional/additional NMS operation. By default, None which will use value provided during initialization.
extra_nms (bool, optional) – Whether or not to perform an additional NMS operation. By default None, which will use value provided during initialization.
agnostic_nms (bool, optional) – Whether or not to perform class-agnostic NMS for the optional/additional operation. By default None, which will use value provided during initialization.
verbose (bool, optional) – Whether or not to log additional information.
- Returns:
For single image: list[tuple[…]] (detections for single image). For batch: list[list[tuple[…]]] (detections per image).
- Return type:
- end2end(images: ndarray, conf_thres: float | None = None, nms_iou_thres: float | None = None, *, extra_nms: bool | None = None, agnostic_nms: bool | None = None, verbose: bool | None = None) list[tuple[tuple[int, int, int, int], float, int]][source]¶
- end2end(images: list[ndarray], conf_thres: float | None = None, nms_iou_thres: float | None = None, *, extra_nms: bool | None = None, agnostic_nms: bool | None = None, verbose: bool | None = None) list[list[tuple[tuple[int, int, int, int], float, int]]]
Perform end to end inference for a batch of images.
Equivalent to running preprocess, run, postprocess, and get_detections in that order. Makes some memory transfer optimizations under the hood to improve performance.
- Parameters:
images (np.ndarray | list[np.ndarray]) – A single image (HWC format) or list of images to perform inference with.
conf_thres (float, optional) – The confidence threshold with which to retrieve bounding boxes. By default None.
nms_iou_thres (float) – The IOU threshold to use during the optional/additional NMS operation. By default, None which will use value provided during initialization.
extra_nms (bool, optional) – Whether or not to perform an additional NMS operation. By default None, which will use value provided during initialization.
agnostic_nms (bool, optional) – Whether or not to perform class-agnostic NMS for the optional/additional operation. By default None, which will use value provided during initialization.
verbose (bool, optional) – Whether or not to log additional information.
- Returns:
For single image: list[tuple[…]] (detections). For batch: list[list[tuple[…]]] (detections per image).
- Return type:
- Raises:
RuntimeError – If the orig_image_size buffer is not valid
RuntimeError – If the scale_factor buffer is not valid
RuntimeError – If end2end_graph is enabled and image dimensions change after first call.
RuntimeError – If end2end_graph is enabled and CUDA graph capture fails.
- class trtutils.image.HandInteractionDetector(engine_path: Path | str, warmup_iterations: int = 10, input_range: tuple[float, float] = (0.0, 1.0), preprocessor: str = 'trt', resize_method: str = 'letterbox', conf_thres: float = 0.3, pair_thres: float = 0.5, second_pair_thres: float | None = None, nms_iou_thres: float = 0.5, mean: tuple[float, float, float] | None = None, std: tuple[float, float, float] | None = None, dla_core: int | None = None, device: int | None = None, backend: str = 'auto', *, warmup: bool | None = None, pagelocked_mem: bool | None = None, unified_mem: bool | None = None, cuda_graph: bool | None = None, no_warn: bool | None = None, verbose: bool | None = None)[source]¶
Bases:
ImageModel,HandInteractionDetectorInterfaceImplementation of hand-object interaction detectors.
Wraps engines for hand-object interaction model families (e.g. HOI-DETR, Hands23) that follow a unified, fixed-K output contract:
boxes (B,K,4)xyxy in network-input pixels,scores (B,K),labels (B,K)(0 hand, 1 object, 2 second object),pair_probs (B,K,K,C)where[...,0]is the no-interaction class, and an optionalside (B,K)(0 left, 1 right, valid where label is hand). All model-specific heads (pairing MLP, attribute heads, touch gating) are expected to run in-graph so a single postprocessor handles every model in the family. Letterbox padding is centered and is undone by the postprocessor using the ratios/padding returned from preprocessing.- postprocess(outputs: list[np.ndarray], ratios: list[tuple[float, float]], padding: list[tuple[float, float]], conf_thres: float | None = None, *, no_copy: bool | None = None, verbose: bool | None = None) list[list[np.ndarray]][source]¶
Postprocess the outputs.
- Parameters:
outputs (list[np.ndarray]) – The raw outputs from the engine to postprocess.
ratios (list[tuple[float, float]]) – The rescale ratios used during preprocessing for each image.
padding (list[tuple[float, float]]) – The padding values used during preprocessing for each image.
conf_thres (float, optional) – The confidence threshold to filter candidates by. If not passed, will use the value from the constructor.
no_copy (bool, optional) – Kept for interface symmetry with the other postprocessors. NMS-based indexing always allocates new arrays.
verbose (bool, optional) – Whether or not to log additional information.
- Returns:
The postprocessed outputs per image: [bboxes, scores, labels, pair_probs, side].
- Return type:
- run(images: list[np.ndarray], ratios: list[tuple[float, float]] | None = None, padding: list[tuple[float, float]] | None = None, conf_thres: float | None = None, *, preprocessed: bool | None = None, postprocess: Literal[False], no_copy: bool | None = None, verbose: bool | None = None) list[np.ndarray][source]¶
- run(images: list[np.ndarray], ratios: list[tuple[float, float]] | None = None, padding: list[tuple[float, float]] | None = None, conf_thres: float | None = None, *, preprocessed: bool | None = None, postprocess: Literal[True] | None = None, no_copy: bool | None = None, verbose: bool | None = None) list[list[np.ndarray]]
- run(images: list[np.ndarray], ratios: list[tuple[float, float]] | None = None, padding: list[tuple[float, float]] | None = None, conf_thres: float | None = None, *, preprocessed: bool | None = None, postprocess: bool | None = None, no_copy: bool | None = None, verbose: bool | None = None) list[np.ndarray] | list[list[np.ndarray]]
- run(images: np.ndarray, ratios: tuple[float, float] | None = None, padding: tuple[float, float] | None = None, conf_thres: float | None = None, *, preprocessed: bool | None = None, postprocess: Literal[False], no_copy: bool | None = None, verbose: bool | None = None) list[np.ndarray]
- run(images: np.ndarray, ratios: tuple[float, float] | None = None, padding: tuple[float, float] | None = None, conf_thres: float | None = None, *, preprocessed: bool | None = None, postprocess: Literal[True] | None = None, no_copy: bool | None = None, verbose: bool | None = None) list[np.ndarray]
- run(images: np.ndarray, ratios: tuple[float, float] | None = None, padding: tuple[float, float] | None = None, conf_thres: float | None = None, *, preprocessed: bool | None = None, postprocess: bool | None = None, no_copy: bool | None = None, verbose: bool | None = None) list[np.ndarray]
Run the model on input.
- Parameters:
images (np.ndarray | list[np.ndarray]) – A single image (HWC format) or list of images to run the model on.
ratios (tuple[float, float] | list[tuple[float, float]], optional) – The ratios generated during preprocessing. For single image, pass tuple. For batch, pass list.
padding (tuple[float, float] | list[tuple[float, float]], optional) – The padding values used during preprocessing. For single image, pass tuple. For batch, pass list.
conf_thres (float, optional) – Optional confidence threshold to filter candidates by during postprocessing.
preprocessed (bool, optional) – Whether or not the inputs have been preprocessed. If None, will preprocess inputs.
postprocess (bool, optional) – Whether or not to postprocess the outputs. If None, will postprocess outputs. If postprocessing will occur and the inputs were passed already preprocessed, then the ratios and padding must be passed for postprocessing.
no_copy (bool, optional) – If True, the outputs will not be copied out from the cuda allocated host memory. Instead, the host memory will be returned directly. This memory WILL BE OVERWRITTEN INPLACE by future inferences. In special case where, preprocessing and postprocessing will occur during run and no_copy was not passed (is None), then no_copy will be used for preprocessing and inference stages.
verbose (bool, optional) – Whether or not to log additional information.
- Returns:
For single image with postprocess=True: list[np.ndarray] (single image outputs). For batch with postprocess=True: list[list[np.ndarray]] (per-image outputs). For postprocess=False: list[np.ndarray] (raw outputs).
- Return type:
- Raises:
ValueError – If preprocessed inputs are not a single batch tensor.
RuntimeError – If postprocessing is requested for already-preprocessed inputs without passing ratios/padding.
- get_interactions(outputs: list[np.ndarray], pair_thres: float | None = None, second_pair_thres: float | None = None, *, verbose: bool | None = None) list[HandInteraction][source]¶
- get_interactions(outputs: list[list[np.ndarray]], pair_thres: float | None = None, second_pair_thres: float | None = None, *, verbose: bool | None = None) list[list[HandInteraction]]
Get the hand-object interactions from postprocessed outputs.
- Parameters:
outputs (list[np.ndarray] | list[list[np.ndarray]]) – For single image: list[np.ndarray] (single image’s postprocessed outputs). For batch: list[list[np.ndarray]] (postprocessed outputs per image).
pair_thres (float, optional) – The minimum link probability to link a hand to a first object. By default None, which uses the value from the constructor.
second_pair_thres (float, optional) – The minimum link probability to link a first object to a second object. By default None, which uses the value from the constructor.
verbose (bool, optional) – Whether or not to log additional information.
- Returns:
For single image: list[HandInteraction] (interactions for single image). For batch: list[list[HandInteraction]] (interactions per image).
- Return type:
- end2end(images: np.ndarray, *, conf_thres: float | None = None, pair_thres: float | None = None, second_pair_thres: float | None = None, verbose: bool | None = None) list[HandInteraction][source]¶
- end2end(images: list[np.ndarray], *, conf_thres: float | None = None, pair_thres: float | None = None, second_pair_thres: float | None = None, verbose: bool | None = None) list[list[HandInteraction]]
Perform end to end inference for a batch of images.
Equivalent to running preprocess, run, postprocess, and get_interactions in that order. Makes some memory transfer optimizations under the hood to improve performance.
- Parameters:
images (np.ndarray | list[np.ndarray]) – A single image (HWC format) or list of images to perform inference with.
conf_thres (float, optional) – The confidence threshold to filter candidates by. By default None, which uses the value from the constructor.
pair_thres (float, optional) – The minimum link probability to link a hand to a first object. By default None, which uses the value from the constructor.
second_pair_thres (float, optional) – The minimum link probability to link a first object to a second object. By default None, which uses the value from the constructor.
verbose (bool, optional) – Whether or not to log additional information.
- Returns:
For single image: list[HandInteraction] (interactions). For batch: list[list[HandInteraction]] (interactions per image).
- Return type:
- Raises:
RuntimeError – If end2end_graph is enabled and image dimensions change after first call.
RuntimeError – If end2end_graph is enabled and CUDA graph capture fails.
- class trtutils.image.ImageModel(engine_path: Path | str, warmup_iterations: int = 10, input_range: tuple[float, float] = (0.0, 1.0), preprocessor: str = 'trt', resize_method: str = 'linear', mean: tuple[float, float, float] | None = None, std: tuple[float, float, float] | None = None, dla_core: int | None = None, device: int | None = None, backend: str = 'auto', *, warmup: bool | None = None, pagelocked_mem: bool | None = None, unified_mem: bool | None = None, cuda_graph: bool | None = None, no_warn: bool | None = None, verbose: bool | None = None)[source]¶
Bases:
objectAbstract base class for image models.
- property dtype: np.dtype¶
Get the dtype required by the model.
- update_input_range(input_range: tuple[float, float]) None[source]¶
Update the input range of the model.
This will re-create all preprocessors with the new input range. Only preprocessors which have been created will be re-created.
- update_mean_std(mean: tuple[float, float, float], std: tuple[float, float, float]) None[source]¶
Update the mean and standard deviation of the model.
- get_random_input() list[np.ndarray][source]¶
Generate random images for the model.
- Returns:
A list containing one random image.
- Return type:
list[np.ndarray]
- mock_run(images: list[np.ndarray] | None = None) list[np.ndarray][source]¶
Mock an execution of the model.
- preprocess(images: ndarray, resize: str | None = None, method: str | None = None, *, no_copy: bool | None = None, verbose: bool | None = None) tuple[ndarray, list[tuple[float, float]], list[tuple[float, float]]][source]¶
- preprocess(images: list[ndarray], resize: str | None = None, method: str | None = None, *, no_copy: bool | None = None, verbose: bool | None = None) tuple[ndarray, list[tuple[float, float]], list[tuple[float, float]]]
Preprocess the input images.
- Parameters:
images (np.ndarray | list[np.ndarray]) – A single image (HWC format) or list of images to preprocess.
resize (str) – The method to resize the images with. Options are [letterbox, linear]. By default None, which will use the value passed during initialization.
method (str, optional) – The underlying preprocessor to use. Options are ‘cpu’, ‘cuda’, or ‘trt’. By default None, which will use the preprocessor stated in the constructor.
no_copy (bool, optional) – If True and using CUDA, do not copy the data from the allocated memory. If the data is not copied, it WILL BE OVERWRITTEN INPLACE once new data is generated.
verbose (bool, optional) – Whether or not to log additional information.
- Returns:
The preprocessed batch tensor, list of ratios per image, and list of padding per image.
- Return type:
tuple[np.ndarray, list[tuple[float, float]], list[tuple[float, float]]]