• Sources: primary, discussion
  • Summary: The post builds arithmetic coding and Shannon entropy from worked examples, up to the point where a better probability model is the only thing that lowers the size floor.
  • Why it matters: The post argues that the quantity the coder is bounded by and the quantity a language model is trained to reduce are the same measurement seen from different ends, which is its framing rather than a measured result.

send feedback on this story