trtutils.image.postprocessors package

Module contents

Postprocessors for images.

Functions

get_classifications()

Get the classifications from the output of a classification model.

get_depth_maps()

Get the depth maps from the output of a depth estimation model.

postprocess_classifications()

Postprocess the output of a classification model.

postprocess_depth()

Postprocess the output of a depth estimation model.

get_detections()

Get the detections from unified postprocessed outputs.

get_interactions()

Pair hands with objects from postprocessed hand-object interaction outputs.

postprocess_hand_interactions()

Postprocess the output of a hand-object interaction model.

postprocess_yolov10()

Postprocess the output of a YOLO-v10 model.

postprocess_rfdetr()

Postprocess the output of a RF-DETR model.

postprocess_detr()

Postprocess the output of a DETR-based model.

postprocess_detr_lbs()

Postprocess the output of a DETR-based model with LBS output order.

postprocess_rtdetrv3()

Postprocess the output of an RT-DETR v3 model.

postprocess_efficient_nms()

Postprocess the output of an EfficientNMS model.

trtutils.image.postprocessors.get_classifications(outputs: list[list[ndarray]], top_k: int = 5, *, verbose: bool | None = None) → list[list[tuple[int, float]]][source]

Get the classifications from the output of a classification network.

Parameters:
  • outputs (list[list[np.ndarray]]) – The postprocessed outputs per image from a classification network.

  • top_k (int, optional) – The number of top predictions to return. Default is 5.

  • verbose (bool, optional) – Whether or not to log additional information.

Returns:

The classifications per image, where each entry is (class_id, confidence).

Return type:

list[list[tuple[int, float]]]

trtutils.image.postprocessors.get_depth_maps(outputs: list[list[ndarray]], *, verbose: bool | None = None) → list[ndarray][source]

Get the depth maps from postprocessed depth estimation outputs.

Parameters:
  • outputs (list[list[np.ndarray]]) – The postprocessed outputs per image from a depth estimation network.

  • verbose (bool, optional) – Whether or not to log additional information.

Returns:

The depth map per image, each of shape (1, H, W) with values in [0, 1].

Return type:

list[np.ndarray]

trtutils.image.postprocessors.get_detections(outputs: list[list[ndarray]], conf_thres: float | None = None, nms_iou_thres: float = 0.5, *, extra_nms: bool | None = None, agnostic_nms: bool | None = None, verbose: bool | None = None) → list[list[tuple[tuple[int, int, int, int], float, int]]][source]

Convert postprocessed unified outputs to human-friendly detections.

Applies an optional confidence filter and optional CPU NMS. Input format is one list per image of [bboxes (N,4), scores (N,), class_ids (N,)].

Parameters:
  • outputs (list[list[np.ndarray]]) – Postprocessed outputs per image (unified format).

  • conf_thres (float, optional) – Confidence threshold; detections below are dropped.

  • nms_iou_thres (float) – IoU threshold for optional extra NMS.

  • extra_nms (bool, optional) – If True, run CPU NMS on the final detections.

  • agnostic_nms (bool, optional) – If True, use class-agnostic NMS when extra_nms is True.

  • verbose (bool, optional) – If True, log extra debug information.

Returns:

One list per image; each detection is ((x1, y1, x2, y2), score, class_id).

Return type:

list[list[tuple[tuple[int, int, int, int], float, int]]]

trtutils.image.postprocessors.get_interactions(outputs: list[list[ndarray]], pair_thres: float = 0.5, second_pair_thres: float | None = None, *, verbose: bool | None = None) → list[list[Tuple[Tuple[Tuple[int, int, int, int], float], Tuple[Tuple[int, int, int, int], float] | None, Tuple[Tuple[int, int, int, int], float] | None, int | None, int | None]]][source]

Pair hands with objects from postprocessed hand-object interaction outputs.

For each hand (label 0), links to the first object (label 1) with the highest link probability above pair_thres, then from that object to the second object (label 2) with the highest link probability above second_pair_thres. Objects may be shared across hands.

Parameters:
  • outputs (list[list[np.ndarray]]) – Postprocessed outputs per image, as returned by postprocess_hand_interactions.

  • pair_thres (float) – Minimum link probability to link a hand to a first object.

  • second_pair_thres (float, optional) – Minimum link probability to link a first object to a second object. Defaults to pair_thres.

  • verbose (bool, optional) – Whether or not to log additional information.

Returns:

One list of interactions per image, one entry per hand.

Return type:

list[list[HandInteraction]]

trtutils.image.postprocessors.postprocess_classifications(outputs: list[ndarray], *, no_copy: bool | None = None, verbose: bool | None = None) → list[list[ndarray]][source]

Postprocess outputs from a classification network.

Parameters:
  • outputs (list[np.ndarray]) – The outputs from a classification network with batch dimension.

  • no_copy (bool, optional) – If True, the outputs will not be copied out from the cuda allocated host memory. Instead, the host memory will be returned directly. This memory WILL BE OVERWRITTEN INPLACE by future inference calls.

  • verbose (bool, optional) – Whether or not to log additional information.

Returns:

The postprocessed outputs per image.

Return type:

list[list[np.ndarray]]

trtutils.image.postprocessors.postprocess_depth(outputs: list[ndarray], *, no_copy: bool | None = None, verbose: bool | None = None) → list[list[ndarray]][source]

Postprocess outputs from a depth estimation network.

Parameters:
  • outputs (list[np.ndarray]) – The outputs from a depth estimation network with batch dimension. Expected shape is (B, 1, H, W) or (B, H, W).

  • no_copy (bool, optional) – If True, the outputs will not be copied out from the cuda allocated host memory. Instead, the host memory will be returned directly. This memory WILL BE OVERWRITTEN INPLACE by future inference calls.

  • verbose (bool, optional) – Whether or not to log additional information.

Returns:

The postprocessed depth maps per image. Each image has a single depth map of shape (1, H, W) with values normalized to [0, 1].

Return type:

list[list[np.ndarray]]

trtutils.image.postprocessors.postprocess_detr(outputs: list[ndarray], ratios: list[tuple[float, float]], padding: list[tuple[float, float]], conf_thres: float | None = None, input_size: tuple[int, int] | None = None, *, no_copy: bool | None = None, verbose: bool | None = None) → list[list[ndarray]][source]

Postprocess DETR-based output (DEIM, RT-DETR, D-FINE).

Expects [scores, labels, boxes] with boxes (batch, num_queries, 4) as [x1, y1, x2, y2]. For models using orig_target_sizes, boxes are already in original image coords; no coordinate transform is applied.

Parameters:
  • outputs (list[np.ndarray]) – Raw DETR outputs [scores, labels, boxes].

  • ratios (list[tuple[float, float]]) – Preprocessing resize ratios per image.

  • padding (list[tuple[float, float]]) – Preprocessing padding per image.

  • conf_thres (float, optional) – Confidence threshold to filter detections.

  • input_size (tuple[int, int] | None) – Unused for standard DETR.

  • no_copy (bool, optional) – If True, return buffers without copying (overwritten by later preprocessing).

  • verbose (bool, optional) – If True, log extra debug information.

Returns:

One list per image, each [bboxes (N,4), scores (N,), class_ids (N,)].

Return type:

list[list[np.ndarray]]

trtutils.image.postprocessors.postprocess_detr_lbs(outputs: list[ndarray], ratios: list[tuple[float, float]], padding: list[tuple[float, float]], conf_thres: float | None = None, input_size: tuple[int, int] | None = None, *, no_copy: bool | None = None, verbose: bool | None = None) → list[list[ndarray]][source]

Postprocess DETR-style output when the engine returns (labels, boxes, scores).

Used for DEIM and D-FINE; reorders and delegates to postprocess_detr.

Parameters:
  • outputs (list[np.ndarray]) – Raw outputs in (labels, boxes, scores) order.

  • ratios (list[tuple[float, float]]) – Preprocessing resize ratios per image.

  • padding (list[tuple[float, float]]) – Preprocessing padding per image.

  • conf_thres (float, optional) – Confidence threshold to filter detections.

  • input_size (tuple[int, int] | None) – Unused.

  • no_copy (bool, optional) – If True, return buffers without copying (overwritten by later preprocessing).

  • verbose (bool, optional) – If True, log extra debug information.

Returns:

One list per image, each [bboxes (N,4), scores (N,), class_ids (N,)].

Return type:

list[list[np.ndarray]]

trtutils.image.postprocessors.postprocess_efficient_nms(outputs: list[ndarray], ratios: list[tuple[float, float]], padding: list[tuple[float, float]], conf_thres: float | None = None, input_size: tuple[int, int] | None = None, *, no_copy: bool | None = None, verbose: bool | None = None) → list[list[ndarray]][source]

Postprocess EfficientNMS plugin output.

Raw outputs are [num_dets, bboxes, scores, class_ids], each with batch dim at 0; unletterboxes and returns the unified format per image.

Parameters:
  • outputs (list[np.ndarray]) – Raw EfficientNMS outputs.

  • ratios (list[tuple[float, float]]) – Preprocessing resize ratios per image.

  • padding (list[tuple[float, float]]) – Preprocessing padding per image.

  • conf_thres (float, optional) – Optional extra confidence filter (use if plugin used a low cutoff).

  • input_size (tuple[int, int] | None) – Unused.

  • no_copy (bool, optional) – If True, return buffers without copying (overwritten by later preprocessing).

  • verbose (bool, optional) – If True, log extra debug information.

Returns:

One list per image, each [bboxes (N,4), scores (N,), class_ids (N,)] in original coords.

Return type:

list[list[np.ndarray]]

trtutils.image.postprocessors.postprocess_hand_interactions(outputs: list[ndarray], ratios: list[tuple[float, float]], padding: list[tuple[float, float]], conf_thres: float = 0.3, nms_iou_thres: float = 0.5, *, no_copy: bool | None = None, verbose: bool | None = None) → list[list[ndarray]][source]

Postprocess outputs from a hand-object interaction network.

Expects the unified engine output contract: [boxes (B,K,4), scores (B,K), labels (B,K), pair_probs (B,K,K,C), side (B,K)]; side is optional. Boxes are remapped out of letterboxed network-input coordinates, filtered by confidence, then deduplicated with class-aware NMS.

Parameters:
  • outputs (list[np.ndarray]) – Raw engine outputs [boxes, scores, labels, pair_probs, (side)].

  • ratios (list[tuple[float, float]]) – Preprocessing resize ratios per image.

  • padding (list[tuple[float, float]]) – Preprocessing padding per image.

  • conf_thres (float) – Confidence threshold used both to filter candidates and as the NMS score threshold.

  • nms_iou_thres (float) – IoU threshold for class-aware NMS.

  • no_copy (bool, optional) – Kept for interface symmetry with the other postprocessors. NMS-based indexing always allocates new arrays, so this has no effect.

  • verbose (bool, optional) – Whether or not to log additional information.

Returns:

One list per image: [bboxes (N,4), scores (N,), labels (N,), pair_probs (N,N,C), side (N,1) if present else (N,0)] in original image coordinates.

Return type:

list[list[np.ndarray]]

trtutils.image.postprocessors.postprocess_rfdetr(outputs: list[ndarray], ratios: list[tuple[float, float]], padding: list[tuple[float, float]], conf_thres: float | None = None, input_size: tuple[int, int] | None = None, *, no_copy: bool | None = None, verbose: bool | None = None) → list[list[ndarray]][source]

Postprocess RF-DETR output.

Expects [dets, labels]: dets (batch, num_queries, 4) in normalized [cx, cy, w, h]; labels (batch, num_queries, num_classes) as logits (class IDs 1-indexed). Converts to unified format in original coords.

Parameters:
  • outputs (list[np.ndarray]) – Raw RF-DETR outputs [dets, labels].

  • ratios (list[tuple[float, float]]) – Preprocessing resize ratios per image.

  • padding (list[tuple[float, float]]) – Preprocessing padding per image.

  • conf_thres (float, optional) – Confidence threshold to filter detections.

  • input_size (tuple[int, int] | None) – Model input (width, height) to denormalize bboxes; default 640x640.

  • no_copy (bool, optional) – If True, return buffers without copying (overwritten by later preprocessing).

  • verbose (bool, optional) – If True, log extra debug information.

Returns:

One list per image, each [bboxes (N,4), scores (N,), class_ids (N,)] in original coords.

Return type:

list[list[np.ndarray]]

trtutils.image.postprocessors.postprocess_rtdetrv3(outputs: list[ndarray], ratios: list[tuple[float, float]], padding: list[tuple[float, float]], conf_thres: float | None = None, input_size: tuple[int, int] | None = None, *, no_copy: bool | None = None, verbose: bool | None = None) → list[list[ndarray]][source]

Postprocess RT-DETR v3 (PaddlePaddle export) output.

Two tensors: combined_dets (total_N, 6) with rows (class_id, score, x1, y1, x2, y2) in original image coords, and num_dets_per_image [batch_size]. No coordinate transform is applied.

Parameters:
  • outputs (list[np.ndarray]) – [combined_dets, num_dets_per_image].

  • ratios (list[tuple[float, float]]) – Unused (boxes already in image coords).

  • padding (list[tuple[float, float]]) – Unused.

  • conf_thres (float, optional) – Confidence threshold to filter detections.

  • input_size (tuple[int, int] | None) – Unused.

  • no_copy (bool, optional) – If True, return buffers without copying (overwritten by later preprocessing).

  • verbose (bool, optional) – If True, log extra debug information.

Returns:

One list per image, each [bboxes (N,4), scores (N,), class_ids (N,)].

Return type:

list[list[np.ndarray]]

trtutils.image.postprocessors.postprocess_yolov10(outputs: list[ndarray], ratios: list[tuple[float, float]], padding: list[tuple[float, float]], conf_thres: float | None = None, input_size: tuple[int, int] | None = None, *, no_copy: bool | None = None, verbose: bool | None = None) → list[list[ndarray]][source]

Postprocess YOLO-v10 engine output.

Expects a single output of shape (batch, N, 6) with rows (x1, y1, x2, y2, score, class_id); unletterboxes and returns the unified format per image.

Parameters:
  • outputs (list[np.ndarray]) – Raw YOLO-v10 engine outputs.

  • ratios (list[tuple[float, float]]) – Preprocessing resize ratios per image.

  • padding (list[tuple[float, float]]) – Preprocessing padding per image.

  • conf_thres (float, optional) – Optional extra confidence filter (use if model used a low cutoff).

  • input_size (tuple[int, int] | None) – Unused.

  • no_copy (bool, optional) – If True, return buffers without copying (overwritten by later preprocessing).

  • verbose (bool, optional) – If True, log extra debug information.

Returns:

One list per image, each [bboxes (N,4), scores (N,), class_ids (N,)] in original coords.

Return type:

list[list[np.ndarray]]