- Sources: writeup, HN discussion
- Summary: The post, dated 2026-08-10, reports that transfers whose total weight size is an integer multiple of 1 MiB per core, measured as 1 MiB of static kernel data on each of 16 cores, hit a DMA erratum that drops streaming bandwidth from a nominal 45 to 60 GB/s to a floor of 17 to 19 GB/s. Multiples of 2048 are capped at the same floor, and the curves measured for one through six laps collapse onto a single profile. Splitting transfers off the boundary raised Llama 3.2 1B decode throughput from 10.0 to 24.3 tokens per second, and this is a separate post from the ANE register-interface writeup carried on 2026-09-12.
- Why it matters: The remedy is a transfer-sizing change in the compiler, so shipped M3 hardware recovers the lost bandwidth without a silicon revision.
send feedback on this story