https://openreview.net/forum?id=KQ4pJFFqT1
LServe: Efficient Long-sequence LLM Serving with Unified Sparse Attention | OpenReview
Large language models (LLMs) have shown remarkable potential in processing long sequences, yet efficiently serving these long-context models remains...
sparse attentionefficientlongsequencellm
https://www.hashe.com/tech-news/deepseek-releases-sparse-attention-model-that-cuts-api-costs-in-half/
Deepseek Releases Sparse Attention Model That Cuts - DeepSee
Sep 30, 2025 - deepseek releases sparse attention model that cuts: DeepSeek has unveiled a groundbreaking experimental model aimed at significantly reducing inference costs
sparse attentiondeepseekreleasesmodelcuts
https://proceedings.iclr.cc/paper_files/paper/2025/hash/03645743ea35690f30d795d6bac149a5-Abstract-Conference.html
FlexPrefill: A Context-Aware Sparse Attention Mechanism for Efficient Long-Sequence Inference
sparse attentioncontextawaremechanismefficient
https://huggingface.co/papers/2406.16747
Paper page - Sparser is Faster and Less is More: Efficient Sparse Attention for Long-Range...
Join the discussion on this paper page
sparse attentionlong rangepaperfasterless
https://learnopoly.com/openbmb-publishes-minicpm4-ultra-effective-language-models-for-edge-devices-with-sparse-attention-and-rapid-inference/
OpenBMB Publishes Minicpm4: Ultra-effective Language Models For Edge Devices With Sparse Attention...
Jun 17, 2025 - Great language models have become integrated into AI systems, allowing tasks such as multilingual translation, virtual assistance and automated reasoning via
language modelsfor edgesparse attentionpublishesultra
https://deepai.org/publication/energon-towards-efficient-acceleration-of-transformers-using-dynamic-sparse-attention
Energon: Towards Efficient Acceleration of Transformers Using Dynamic Sparse Attention | DeepAI
Oct 18, 2021 - 10/18/21 - In recent years, transformer models have revolutionized Natural Language Processing (NLP) and also show promising performance on C...
sparse attentionenergontowardsefficientacceleration
https://deepwiki.com/deep-spin/adasplash/3.1-sparse-attention-with-block-masking
Sparse Attention with Block Masking | deep-spin/adasplash | DeepWiki
This document covers the sparse attention implementation with block masking provided by the `sparseattn` function in $1. This variant uses adaptive sparsity...
sparse attentionblockmaskingdeepspin
https://pure.nwpu.edu.cn/en/publications/predicting-and-understanding-student-learning-performance-using-m/
Predicting and Understanding Student Learning Performance Using Multi-Source Sparse Attention...
student learningsparse attentionunderstandingperformanceusing
https://proceedings.neurips.cc/paper/2020/hash/f0b76267fbe12b936bd65e203dc675c1-Abstract.html
Sparse and Continuous Attention Mechanisms
attention mechanismssparsecontinuous