Capturing data…
Capturing data…
FRI, 2 OCT · 93 ITEMS
"SparLeak: sparse-attention LLM inference leaks data on shared GPUs" — arXiv cs.LG · AI & Technology
This paper from researchers demonstrates a privacy vulnerability in large language model when using on . Sparse attention is a technique that reduces computation by focusing on only parts of the input, but it creates detectable memory access patterns that can leak information about the input data to co-located users on the same GPU hardware.
The claim originates from an arXiv preprint (2609.38830), a primary source where the researchers directly report their findings of a side-channel attack exploiting sparse attention on shared GPUs. As a preprint, it has not undergone peer review, but the description of the attack mechanism and results are the authors' own.
Genuinely new: the paper was posted to arXiv on October 1, 2026, with no earlier circulation.