https://openreview.net/forum?id=z1ph7lcmpv
SageAttention2: Efficient Attention with Smoothing Q and Per-thread Quantization | OpenReview
Although quantization for linear layers has been widely used, its application to accelerate the attention process remains limited. To further enhance the...
efficientattentionsmoothing