|
Hi folks! We are building a multi-class anomaly detection pipeline deployed on low-power ARM64 CPUs using ONNXRuntime + XNNPACK. To adhere to enterprise software licensing policies, we are using <12M parameter real-time transformer detectors: D-FINE (Nano/Small) and DEIMv2 (Pico/Nano). Few queries on top of my mind:
Any recommended integration patterns or guidance would be really helpful! |
Replies: 4 comments 1 reply
|
Can give you 1 and 2 with reasonable confidence — I'll leave 3 and 4 to someone who's actually deployed these on ARM. 1. Yep, that's the normal way round, not a workaround. 2. Main thing is to threshold the queries yourself before building the remain_inds = scores >= self.track_activation_threshold
inds_low = scores > 0.1
inds_high = scores < self.track_activation_threshold
inds_second = np.logical_and(inds_low, inds_high)Anything between |
|
Understood. Thanks for the detailed explanation! @capCafu |
|
Adding to the above, on Q3 and Q4 which are still open. Q4 (XNNPACK with transformer detectors): expect little gain, and measure it. If you test it, use the threading setup the docs recommend, otherwise the two threadpools fight each other: so = ort.SessionOptions()
so.intra_op_num_threads = 1
so.add_session_config_entry("session.intra_op.allow_spinning", "0")
sess = ort.InferenceSession(
"model.onnx", so,
providers=[("XnnpackExecutionProvider", {"intra_op_num_threads": 4}), "CPUExecutionProvider"],
)Then compare against plain Q3 (full high-res input vs InferenceSlicer): slice, for two reasons.
So keep the ONNX export at its native static shape and tile: slicer = sv.InferenceSlicer(
callback=run_onnx, # your ORT call returning sv.Detections
slice_wh=640,
overlap_wh=100,
overlap_filter=sv.OverlapFilter.NON_MAX_SUPPRESSION,
iou_threshold=0.5,
)Note that even though the model itself is NMS-free, you still need the overlap filter here: the same defect seen in two overlapping tiles gives two detections, and only the slicer can merge them. Latency scales with the number of tiles, so on a fixed-camera inspection setup crop to the ROI first to cut the tile count. Keep |
|
Thanks a lot for the detailed answer! @Y1D1R . |
Adding to the above, on Q3 and Q4 which are still open.
Q4 (XNNPACK with transformer detectors): expect little gain, and measure it.
The XNNPACK EP only takes a short list of ops: Conv, ConvTranspose, MaxPool, AveragePool, Softmax, Resize (bilinear), Gemm, MatMul (2D only) and their QLinear variants (docs). Everything else falls back to the CPU EP. In D-FINE / DEIMv2 that means the conv backbone can go to XNNPACK, but LayerNorm, the batched MatMuls of attention, GridSample in the deformable decoder, etc. stay on the CPU EP. You end up with a partitioned graph, and the default CPU EP (MLAS) already has NEON kernels on ARM64, so the difference is often small and can even be negative.
If you…