hhzz
2026-10-11 08:34:02点赞:0阅读:0
关注
【YOLO 入门到精通 10】模型导出与部署:ONNX、TensorRT 与边缘设备实战
标签:ONNX TensorRT 模型部署 模型导出 边缘计算
难度:⭐⭐⭐⭐ | 阅读时长:约 65 分钟 | 系列:第 10/12 集
前置:09 目标跟踪 | 下一集:11 进阶任务

写在前面
.pt 文件适合训练调试,不适合直接上生产。上线需要 ONNX、TensorRT 等优化格式——速度可提升 2~5 倍,依赖也更轻。
本集手把手教你导出、推理、选型,覆盖服务器到树莓派的全链路部署。
本集学习目标
目录
10.1 为什么需要导出
| PyTorch .pt | 导出格式 |
|---|
| 依赖 | 重(PyTorch) | 轻(ONNX Runtime 等) |
| 速度 | 基准 | 优化后 2~5x |
| 跨平台 | 有限 | 广泛 |
| 生产 | ❌ | ✅ |
格式选型
| 格式 | 平台 | 加速 |
|---|
| ONNX | 跨平台 | 1.5~2x |
| TensorRT | NVIDIA GPU | 2~5x |
| OpenVINO | Intel | 2~3x |
| CoreML | Apple | 2~3x |
| TFLite | Android/嵌入式 | 2~4x |
10.2 基本导出操作
yolo export model=best.pt format=onnx
yolo export model=best.pt format=engine device=0 half=True
from ultralytics import YOLO
model = YOLO("runs/detect/train/weights/best.pt")
model.export(format="onnx", imgsz=640, simplify=True)
model.export(format="engine", device=0, half=True)
model.export(format="openvino")
model.export(format="coreml")
model.export(format="tflite", int8=True)
导出参数
| 参数 | 说明 |
|---|
format | 目标格式 |
imgsz | 固定输入尺寸 |
half | FP16 |
int8 | INT8 量化 |
dynamic | 动态 batch/尺寸 |
simplify | 简化 ONNX 图 |
workspace | TensorRT 工作空间 GB |
10.3 ONNX 导出与推理
model.export(format="onnx", dynamic=True, opset=17)
onnx_model = YOLO("best.onnx")
results = onnx_model("test.jpg")
ONNX Runtime 原生推理
import onnxruntime as ort
import numpy as np, cv2
session = ort.InferenceSession("best.onnx")
inp_name = session.get_inputs()[0].name
img = cv2.imread("test.jpg")
img = cv2.cvtColor(cv2.resize(img, (640,640)), cv2.COLOR_BGR2RGB)
img = img.transpose(2,0,1)[None].astype(np.float32) / 255.0
outputs = session.run(None, {inp_name: img})
10.4 TensorRT 加速
model.export(format="engine", device=0, half=True)
model.export(format="engine", device=0, int8=True, data="calibration.yaml")
trt_model = YOLO("best.engine")
results = trt_model("test.jpg")
print(results[0].speed)
| 精度 | 速度 | 精度损失 |
|---|
| FP32 | 1x | 无 |
| FP16 | 2~3x | 极小 |
| INT8 | 3~5x | 小 |
10.5 边缘设备部署
| 设备 | 格式 | 建议 |
|---|
| Jetson | TensorRT FP16 | imgsz=640, half=True |
| 树莓派 | TFLite INT8 | yolo26n, imgsz=320 |
| iOS | CoreML | nms=True |
| 服务器 GPU | TensorRT | batch 推理 |
10.6 量化技术
model.export(format="onnx", int8=True, data="calibration.yaml")
model.export(format="engine", int8=True, data="calibration.yaml", device=0)
10.7 基准测试
from ultralytics.utils.benchmarks import benchmark
benchmark(model="yolo26n.pt", data="coco8.yaml", imgsz=640, device=0)
yolo benchmark model=yolo26n.pt data=coco8.yaml
| 指标 | 含义 |
|---|
| Size | 文件大小 |
| mAP50-95 | 量化后精度 |
| Inference | 推理耗时 ms |
10.8 生产部署架构
FastAPI 服务
from fastapi import FastAPI, UploadFile
from ultralytics import YOLO
import cv2, numpy as np
app = FastAPI()
model = YOLO("best.engine")
@app.post("/detect")
async def detect(file: UploadFile):
data = await file.read()
img = cv2.imdecode(np.frombuffer(data, np.uint8), cv2.IMREAD_COLOR)
results = model(img, conf=0.5)
return {"detections": [
{"class": model.names[int(b.cls[0])], "conf": float(b.conf[0]), "bbox": b.xyxy[0].tolist()}
for b in results[0].boxes
]}
Docker
FROM ultralytics/ultralytics:latest-nvidia
WORKDIR /app
COPY best.engine api.py ./
RUN pip install fastapi uvicorn
EXPOSE 8000
CMD ["uvicorn", "api:app", "--host", "0.0.0.0", "--port", "8000"]
10.9 YOLO26 导出注意
YOLO26 默认端到端无 NMS:
model.export(format="onnx")
model.export(format="onnx", end2end=False)
10.10 踩坑指南
| 坑 | 解决 |
|---|
| TensorRT 导出失败 | 确认 CUDA/cuDNN 版本匹配 |
| ONNX 推理尺寸不对 | 导出时固定 imgsz |
| INT8 精度掉太多 | 增加校准数据量 |
| Mac 无法 TensorRT | 用 CoreML 或 ONNX |
| 动态尺寸导出后慢 | 生产用固定尺寸 |
10.11 本集小结
yolo export model=best.pt format=engine device=0 half=True
yolo export model=best.pt format=onnx
yolo export model=best.pt format=tflite int8=True
转载自 CSDN-专业IT技术社区
原文链接:https://blog.csdn.net/weixin_43025151/article/details/167370523