• Sources: primary, driver source, HN discussion
  • Summary: Eileen Yoon returns to an abandoned reverse-engineered M1 ANE driver to map the block, and reports that a task descriptor is not an instruction stream: each descriptor is a sequence of burst-write packets copying 32-bit words from IOMMU virtual memory into fixed hardware register blocks, with no load-weights opcode and no dynamic load or store from a virtual address. The compute side is 16 cores of 128 FP16 or 256 INT8 MAC lanes, 2048 lanes total, with a 32-bit Q16.16 accumulator that saturates at 2^15, and tanh implemented as a 33-entry piecewise-linear lookup table. The post, dated 2026-08-10, frames the design as assumptions about predictable CNN reuse that Apple committed to silicon in 2017 and that transformer decode broke, and reads the M5 folding ANE cores into the GPU as the end of the standalone NPU.
  • Why it matters: A fixed-function accelerator with no instruction set cannot be retargeted in software, which bounds what any third-party ANE backend can ever compile to.

send feedback on this story