YOLOv5：解读general.py_yolov5 general.py-CSDN博客

YOLOv5：解读general.py

前言
前提条件
相关介绍
general.py
参考

前言

记录一下自己阅读general.py代码的一些重要点，方便自己查阅。特别感谢，在参考里，列举的博文链接，写得很好，对本人阅读理解yolo.py代码，有很大帮助。
由于本人水平有限，难免出现错漏，敬请批评改正。
更多精彩内容，可点击进入YOLO系列专栏、自然语言处理
专栏或我的个人主页查看
基于DETR的人脸伪装检测
YOLOv7训练自己的数据集（口罩检测）
YOLOv8训练自己的数据集（足球检测）
YOLOv5：TensorRT加速YOLOv5模型推理
YOLOv5：IoU、GIoU、DIoU、CIoU、EIoU
玩转Jetson Nano（五）：TensorRT加速YOLOv5目标检测
YOLOv5：添加SE、CBAM、CoordAtt、ECA注意力机制
YOLOv5：yolov5s.yaml配置文件解读、增加小目标检测层
Python将COCO格式实例分割数据集转换为YOLO格式实例分割数据集
YOLOv5：使用7.0版本训练自己的实例分割模型（车辆、行人、路标、车道线等实例分割）
使用Kaggle GPU资源免费体验Stable Diffusion开源项目

前提条件

熟悉Python

general.py

clip_boxes

在 Python 中，torch.clamp_() 和numpy.clip()是一个用于限制数值范围的方法，该方法接受三个参数：最小值、最大值和需要被限制的数值。
方法的作用是将给定的数值限制在最小值和最大值之间，返回一个新的值。如果原始数值小于最小值，则返回最小值；如果原始数值大于最大值，则返回最大值；如果原始数值在最小值和最大值之间，则返回原始数值。

在这里插入图片描述

def clip_boxes(boxes, shape):
    # Clip boxes (xyxy) to image shape (height, width)
    if isinstance(boxes, torch.Tensor):  # faster individually
        boxes[..., 0].clamp_(0, shape[1])  # x1
        boxes[..., 1].clamp_(0, shape[0])  # y1
        boxes[..., 2].clamp_(0, shape[1])  # x2
        boxes[..., 3].clamp_(0, shape[0])  # y2
    else:  # np.array (faster grouped)
        boxes[..., [0, 2]] = boxes[..., [0, 2]].clip(0, shape[1])  # x1, x2
        boxes[..., [1, 3]] = boxes[..., [1, 3]].clip(0, shape[0])  # y1, y2

scale_boxes $\bigstar$

在这里插入图片描述

def scale_boxes(img1_shape, boxes, img0_shape, ratio_pad=None):
    # Rescale boxes (xyxy) from img1_shape to img0_shape
    if ratio_pad is None:  # calculate from img0_shape
        gain = min(img1_shape[0] / img0_shape[0], img1_shape[1] / img0_shape[1])  # gain  = old / new
        pad = (img1_shape[1] - img0_shape[1] * gain) / 2, (img1_shape[0] - img0_shape[0] * gain) / 2  # wh padding
    else:
        gain = ratio_pad[0][0]
        pad = ratio_pad[1]

    boxes[..., [0, 2]] -= pad[0]  # x padding
    boxes[..., [1, 3]] -= pad[1]  # y padding
    boxes[..., :4] /= gain
    clip_boxes(boxes, img0_shape)
    return boxes

xywh2xyxy

在这里插入图片描述

def xywh2xyxy(x):
    # Convert nx4 boxes from [x, y, w, h] to [x1, y1, x2, y2] where xy1=top-left, xy2=bottom-right
    y = x.clone() if isinstance(x, torch.Tensor) else np.copy(x)
    y[..., 0] = x[..., 0] - x[..., 2] / 2  # top left x
    y[..., 1] = x[..., 1] - x[..., 3] / 2  # top left y
    y[..., 2] = x[..., 0] + x[..., 2] / 2  # bottom right x
    y[..., 3] = x[..., 1] + x[..., 3] / 2  # bottom right y
    return y

non_max_suppression $\bigstar\bigstar\bigstar$

torchvision.ops.nms 是 PyTorch 的 torchvision 库中提供的一个函数，用于实现非极大值抑制（Non-Maximum Suppression，NMS）操作。这个函数对输入的候选框（bounding boxes）进行排序，并根据给定的 IoU 阈值去除重叠度较高的框。
函数的输入参数如下：
boxes：一个包含候选框的张量，每个框由一个或多个边界框坐标组成。每个边界框由四个元素表示，分别是左上角和右下角的坐标（x1, y1, x2, y2）。
scores：一个与 boxes 相同形状的张量，表示每个框的置信度分数。
iou_thres：一个阈值，用于控制哪些框被认为是重叠的。

函数的输出是一个张量，其中包含经过非极大值抑制处理后的结果。
传统的NMS算法，具体流程如下：
步骤一：将所有矩形框按照不同的类别标签分组，组内按照置信度高低得分进行排序；
步骤二：将步骤一中得分最高的矩形框拿出来，遍历剩余矩形框，计算与当前得分最高的矩形框的交并比，将剩余矩形框中大于设定的IOU阈值的框删除；
步骤三：将步骤二结果中，对剩余的矩形框重复步骤二操作，直到处理完所有矩形框；

在YOLOv5中，non_max_suppression函数，具体流程如下：

def non_max_suppression(
        prediction, 
        conf_thres=0.25,
        iou_thres=0.45,
        classes=None,
        agnostic=False,
        multi_label=False,
        labels=(),
        max_det=300,
        nm=0,  # number of masks
):
    """Non-Maximum Suppression (NMS) on inference results to reject overlapping detections
    Arguments:
            prediction : 1个 ，[bs, anchor_num*grid_w*grid_h, xywh+c+5classes] = [4,3*260*260+3*130*130+3*65*65*65,10] = [4, 266175, 10]
    
    Returns:
         list of detections, on (n,6) tensor per image [xyxy, conf, cls]
    """

    # Checks
    assert 0 <= conf_thres <= 1, f'Invalid Confidence threshold {conf_thres}, valid values are between 0.0 and 1.0'
    assert 0 <= iou_thres <= 1, f'Invalid IoU {iou_thres}, valid values are between 0.0 and 1.0'
    if isinstance(prediction, (list, tuple)):  # YOLOv5 model in validation model, output = (inference_out, loss_out)
        prediction = prediction[0]  # select only inference output

    device = prediction.device # 设置推理设备
    mps = 'mps' in device.type  # Apple MPS
    if mps:  # MPS not fully supported yet, convert tensors to CPU before NMS
        prediction = prediction.cpu()
    bs = prediction.shape[0]  # batch size
    nc = prediction.shape[2] - nm - 5  # number of classes # 这里的5表示[x,y,w,h,conf]这5个数值
    xc = prediction[..., 4] > conf_thres  # candidates # 预测框置信度> conf_thres阈值，选为候选框，xc = [True,False,....]

    # Settings
    # min_wh = 2  # (pixels) minimum box width and height # 预测物体宽度和高度的大小范围 [min_wh, max_wh]
    max_wh = 7680  # (pixels) maximum box width and height # 
    max_nms = 30000  # maximum number of boxes into torchvision.ops.nms()
    time_limit = 0.5 + 0.05 * bs  # seconds to quit after # 每个图像最多检测物体的个数 
    redundant = True  # require redundant detections # 是否需要冗余的detections
    multi_label &= nc > 1  # multiple labels per box (adds 0.5ms/img)
    merge = False  # use merge-NMS

    t = time.time()
    mi = 5 + nc  # mask start index
    # batch_size个output  存放最终筛选后的预测框结果
    output = [torch.zeros((0

YOLOv5：解读general.py

YOLOv5：解读general.py

前言

前提条件

相关介绍

general.py

clip_boxes

scale_boxes ★ \bigstar ★

xywh2xyxy

non_max_suppression ★ ★ ★ \bigstar\bigstar\bigstar ★★★

scale_boxes $\bigstar$

non_max_suppression $\bigstar\bigstar\bigstar$