• Sources: primary, discussion
  • Summary: Raschka explains looped transformers using the open-weight Nanbeige4.2-3B, which applies one stack of 22 blocks twice for 44 block applications with weights shared across passes. He sets out the tradeoff: effective depth doubles and transformer-block parameter memory roughly halves against 44 distinct blocks, while the forward pass and backpropagation still run through all 44 applications, so compute is not saved. He frames the claim that Astra uses recurrent depth as reporting by The Information rather than as a confirmed architecture, and treats the hidden chain-of-thought inference drawn from it as the part that does not follow.
  • Why it matters: Weight sharing trades parameter memory for compute, so a looped model is cheaper to hold and no cheaper to run.

send feedback on this story