- Sources: primary, discussion
- Summary: The authors measure the symbol distribution of 29 ternary models, find zeros reach 51.5 percent of weights, and replace five-trit packing, which rounds to 1.625 bits per weight, with a presence bitmap plus a compacted sign vector costing 2 minus the zero density. They report smaller storage on 26 of the 29 models, 1.485 bits per weight on the sparsest, and measured end-to-end decode gains up to 1.18 times on CPUs and 1.27 times on Xe2 GPUs. The preprint is not peer reviewed, and the optimized unpacking sequences target AVX-512, AVX2 and Intel Xe2 GPUs and are unreproduced.
- Why it matters: The deployment cost of a ternary model turns on the storage layout rather than on the information-theoretic bound the format is usually quoted against.
send feedback on this story