Add ARM fp8 dot kernels
I haven't been able to run this on real hardware to benchmark it, but the inner loop assembly looks reasonably nice. This is the 8x8 kernel:
```
.LBB7_9: // Parent Loop BB7_3 Depth=1
// Parent Loop BB7_5 Depth=2
// Parent Loop BB7_7 Depth=3
// => This Inner Loop Header: Depth=4
add x8, x10, x1
ldr d8, [x13, x5]
add x26, x29, x1
ldr d31, [x21, x5]
ldr d30, [x17, x5]
sub x9, x9, #8
ldr q9, [x8]
ldr q10, [x26]
ldr d29, [x15, x5]
add x1, x1, x14
ldr d28, [x16, x5]
add x5, x5, #8
fdot v23.4s, v9.16b, v8.4b[0]
fdot v23.4s, v10.16b, v8.4b[1]
fdot v22.4s, v9.16b, v31.4b[0]
fdot v22.4s, v10.16b, v31.4b[1]
fdot v25.4s, v9.16b, v30.4b[0]
fdot v25.4s, v10.16b, v30.4b[1]
fdot v26.4s, v9.16b, v29.4b[0]
fdot v26.4s, v10.16b, v29.4b[1]
fdot v27.4s, v9.16b, v28.4b[0]
fdot v27.4s, v10.16b, v28.4b[1]
ldr q9, [x8, #16]
ldr q10, [x26, #16]
fdot v18.4s, v9.16b, v8.4b[0]
fdot v18.4s, v10.16b, v8.4b[1]
fdot v17.4s, v9.16b, v31.4b[0]
fdot v17.4s, v10.16b, v31.4b[1]
fdot v20.4s, v9.16b, v30.4b[0]
fdot v20.4s, v10.16b, v30.4b[1]
fdot v21.4s, v9.16b, v29.4b[0]
fdot v21.4s, v10.16b, v29.4b[1]
fdot v24.4s, v9.16b, v28.4b[0]
fdot v24.4s, v10.16b, v28.4b[1]
ldr q9, [x8, #32]
ldr q10, [x26, #32]
fdot v2.4s, v9.16b, v8.4b[0]
fdot v2.4s, v10.16b, v8.4b[1]
fdot v6.4s, v9.16b, v31.4b[0]
fdot v6.4s, v10.16b, v31.4b[1]
fdot v7.4s, v9.16b, v30.4b[0]
fdot v7.4s, v10.16b, v30.4b[1]
fdot v16.4s, v9.16b, v29.4b[0]
fdot v16.4s, v10.16b, v29.4b[1]
fdot v19.4s, v9.16b, v28.4b[0]
fdot v19.4s, v10.16b, v28.4b[1]
ldr q9, [x8, #48]
ldr q10, [x26, #48]
add x8, x4, x9
fdot v0.4s, v9.16b, v8.4b[0]
fdot v0.4s, v10.16b, v8.4b[1]
fdot v3.4s, v9.16b, v31.4b[0]
fdot v3.4s, v10.16b, v31.4b[1]
fdot v4.4s, v9.16b, v30.4b[0]
fdot v4.4s, v10.16b, v30.4b[1]
fdot v5.4s, v9.16b, v29.4b[0]
fdot v5.4s, v10.16b, v29.4b[1]
fdot v1.4s, v9.16b, v28.4b[0]
fdot v1.4s, v10.16b, v28.4b[1]
cmp x8, #15
b.hi .LBB7_9
```
PiperOrigin-RevId: 933830198
XNNPACK is a highly optimized solution for neural network inference on ARM, x86, WebAssembly, and RISC-V platforms. XNNPACK is not intended for direct use by deep learning practitioners and researchers; instead it provides low-level performance primitives for accelerating high-level machine learning frameworks, such as TensorFlow Lite, TensorFlow.js, PyTorch, ONNX Runtime, ExecuTorch, and MediaPipe.
XNNPACK implements the following neural network operators:
All operators in XNNPACK support NHWC layout, but additionally allow custom stride along the Channel dimension. Thus, operators can consume a subset of channels in the input tensor, and produce a subset of channels in the output tensor, providing a zero-cost Channel Split and Channel Concatenation operations.
The table below presents single-threaded performance of XNNPACK library on three generations of MobileNet models and three generations of Pixel phones.
| Model | Pixel, ms | Pixel 2, ms | Pixel 3a, ms |
|---|---|---|---|
| FP32 MobileNet v1 1.0X | 82 | 86 | 88 |
| FP32 MobileNet v2 1.0X | 49 | 53 | 55 |
| FP32 MobileNet v3 Large | 39 | 42 | 44 |
| FP32 MobileNet v3 Small | 12 | 14 | 14 |
The following table presents multi-threaded (using as many threads as there are big cores) performance of XNNPACK library on three generations of MobileNet models and three generations of Pixel phones.
| Model | Pixel, ms | Pixel 2, ms | Pixel 3a, ms |
|---|---|---|---|
| FP32 MobileNet v1 1.0X | 43 | 27 | 46 |
| FP32 MobileNet v2 1.0X | 26 | 18 | 28 |
| FP32 MobileNet v3 Large | 22 | 16 | 24 |
| FP32 MobileNet v3 Small | 7 | 6 | 8 |
Benchmarked on March 27, 2020 with end2end_bench --benchmark_min_time=5 on an Android/ARM64 build with Android NDK r21 (bazel build -c opt --config android_arm64 :end2end_bench) and neural network models with randomized weights and inputs.
The table below presents multi-threaded performance of XNNPACK library on three generations of MobileNet models and three generations of Raspberry Pi boards.
| Model | RPi Zero W (BCM2835), ms | RPi 2 (BCM2836), ms | RPi 3+ (BCM2837B0), ms | RPi 4 (BCM2711), ms | RPi 4 (BCM2711, ARM64), ms |
|---|---|---|---|---|---|
| FP32 MobileNet v1 1.0X | 3919 | 302 | 114 | 72 | 77 |
| FP32 MobileNet v2 1.0X | 1987 | 191 | 79 | 41 | 46 |
| FP32 MobileNet v3 Large | 1658 | 161 | 67 | 38 | 40 |
| FP32 MobileNet v3 Small | 474 | 50 | 22 | 13 | 15 |
| INT8 MobileNet v1 1.0X | 2589 | 128 | 46 | 29 | 24 |
| INT8 MobileNet v2 1.0X | 1495 | 82 | 30 | 20 | 17 |
Benchmarked on Feb 8, 2022 with end2end-bench --benchmark_min_time=5 on a Raspbian Buster build with CMake (./scripts/build-local.sh) and neural network models with randomized weights and inputs. INT8 inference was evaluated on per-channel quantization schema.
XNNPACK is based on QNNPACK library. Over time its codebase diverged a lot, and XNNPACK API is no longer compatible with QNNPACK.