Researchers presented FlashPrefill V2, an advancement to block-sparse attention that targets the computational bottleneck of processing long text sequences during prefill. The method adds a mean correction term that suppresses approximation error and redesigns the sparse attention operator to align with modern inference implementations and support quantized inference. On NVIDIA H20 GPUs, it achieved speedups of up to 47.26x and 27.19x over FlashAttention-2 at 128K context length under different precision settings, while remaining compatible with production inference frameworks.