Over the last few months I have been working on Torx, an ExecuTorch backend in Mesa3D that delegates to the NPU drivers Mesa already contains. Initial Torx support has been merged already, with YOLOX support being currently under review.
Torx sits next to Teflon, the existing TensorFlow Lite delegate, and both use the same Gallium ML API. Today Torx works with the ethosu driver.Torx implements the backend API that ExecuTorch provides for running portions of a model's graph on specialized hardware such as NPUs or GPUs. It then uses the Gallium ML API to compile and execute those portions on the available hardware.
Unlike TensorFlow Lite, ExecuTorch compiles models in a step separate from execution. The model can therefore be compiled on a different machine from the target, where it is practical to apply more expensive optimizations than edge-class hardware could run.
I worked on this alongside Rob Herring from Arm, who supported the effort on the kernel side.
Below is a demo of object detection with YOLOX-Nano at 416x416, quantized to 8 bits integers, on an i.MX93 board with an Arm Ethos-U65 NPU. Inference takes 17 milliseconds per frame.
The convolutional layers of YOLOX run entirely on the U65. The final decoding steps need floating point operations, so they run on the CPU.
Once the other NPU drivers in Mesa support the ahead-of-time compilation extension to the Gallium ML API, they will be usable from ExecuTorch as well.
To try it yourself, you can follow the instructions in the documentation. If you prefer not to build ExecuTorch from sources, you can use this merge request from Meta's Anthony Shoumikhin.
.png)