Getting 50 GB/S Back from the Apple Neural Engine

(eiln.github.io)

27 points | by eiln 2 days ago

3 comments

  • VladVladikoff 14 minutes ago
    This website hijacked my back button during a simple page load. You should fix that, it’s not an acceptable way to behave.
  • Neywiny 38 minutes ago
    Just checking here- this systemverilog is a hypothetical telling of what you think is going on? Or do you have the actual source of the RTL?
  • eiln 2 days ago
    RTL performance erratum in the Apple M3 Neural Engine throttles DRAM weight streaming throughput down to 17–19 GB/s from the nominal 45–60 GB/s. Avoiding the problematic path in the kernel DMA engine's speculative prefetch ring increased Llama 3.2 1B token throughput from 10.0 to 24.3 tokens/s.