2 comments

  • hsfzxjy 1 hour ago
    Author here. Feel free to ask any questions :)
    • __atx__ 18 minutes ago
      Nice work! I have been messing with doing some radio DSP on the NPU in my laptop and found the Intel graph compiler to be rather unpredictable. Weird random mis-lowering to fp16 ops instead of int8, unpredictable throughput depending on tensor sizes, reshapes being sometimes extremely expensive and sometimes basically free etc.

      Obviously I am mostly just driving the DPU, as the SHAVE cores are just not beefy enough, but maybe this will be useful at some point, at least to gain more insight into the architecture...

  • Gigachad 3 hours ago
    What is this useful for?
    • hsfzxjy 2 hours ago
      The main use is enabling developers to write their own NPU kernels according to their respective need, instead of picking workaround from what Intel provides in OpenVINO. I think the project would benefit three scenarios:

      1. Implement uncommon NN operators.

      2. Implement a high-precision (FP32, etc.) version of existing OpenVINO operators for numerically sensitive usage.

      3. Implement a mega kernel that fuse several small kernels together to reduce SHAVE invocations, and potentially improve performance.

      But most importantly, it gives you more control over the hardware you own. That control IMO should have been yours from day one you bought the machine.