• Sources: primary, technical post
  • Summary: NVIDIA announced at Hot Chips 2026 that Groq 3 LPX is in full production, and describes it as an extension of the Vera Rubin platform purpose-built to extend Vera Rubin's interactivity rather than its aggregate throughput. The announcement cites 3,400 output tokens per second on Gemma 4 31B at 100,000 tokens of context, a figure attributed to Artificial Analysis rather than published by NVIDIA with a method. A separate responsiveness claim of 4x carries no stated method. The release names Nebius as the first adopter.
  • Why it matters: Token generation rate sets how long an agent takes per step, and this part is sold on per-user interactivity at long context rather than on aggregate throughput.

send feedback on this story