Skip to content
Discussion options

You must be logged in to vote

Adding to the above, on Q3 and Q4 which are still open.

Q4 (XNNPACK with transformer detectors): expect little gain, and measure it.
The XNNPACK EP only takes a short list of ops: Conv, ConvTranspose, MaxPool, AveragePool, Softmax, Resize (bilinear), Gemm, MatMul (2D only) and their QLinear variants (docs). Everything else falls back to the CPU EP. In D-FINE / DEIMv2 that means the conv backbone can go to XNNPACK, but LayerNorm, the batched MatMuls of attention, GridSample in the deformable decoder, etc. stay on the CPU EP. You end up with a partitioned graph, and the default CPU EP (MLAS) already has NEON kernels on ARM64, so the difference is often small and can even be negative.

If you…

Replies: 4 comments 1 reply

Comment options

You must be logged in to vote
0 replies
Comment options

You must be logged in to vote
0 replies
Comment options

You must be logged in to vote
0 replies
Answer selected by acode-x
Comment options

You must be logged in to vote
1 reply
@Y1D1R
Comment options

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment
Category
Q&A
Labels
None yet
3 participants