• Sources: dri-devel message, r/linux discussion
  • Summary: Torvalds writes on dri-devel that repeatable job timeouts on Intel Xe hardware trace to a rounding error in the flat CCS reservation that has been in the driver for about two years. The effect is that part of the range the hardware has already taken is handed back to the allocator as usable VRAM, so any allocation placed there is silently corrupted. The linked Reddit submission title states that he used AI to debug the driver, and the message itself contains no mention of AI tooling of any kind, so that claim is not carried here.
  • Why it matters: A driver that publishes memory it has already stolen for hardware use corrupts whatever lands there, and when that is a GPU page table the failure surfaces as an unrelated engine reset.
  • Follow-up: Track which stable kernels take the fix.

send feedback on this story