- Sources: arXiv 2607.20723
- Summary: Sadegh Majidi, Niloofar Mireshghallah, and Kazem Taram submitted a preprint on 2026-07-22 describing a remote side channel that uses only per-token generation timing from an inference API. The method builds a timing model of how latency scales with model configuration and hardware parameters on current NVIDIA GPUs, then searches the architecture space against observed timings. The authors report that for Llama models a near-correct configuration of layer count, hidden dimension, and attention head count appears in the top ten candidates more than 90% of the time, and report evidence that Gemini Flash 2.5 runs speculative decoding with a draft context window near 128K tokens.
- Why it matters: Inference serving choices that providers treat as private are recoverable from ordinary API responses, with no access beyond a normal client.
- Follow-up: Watch for reproduction against other hosted endpoints and for serving-side timing mitigations.
send feedback on this story