https://openreview.net/forum?id=n77IeySrQl
Make Your LVLM KV Cache More Lightweight | OpenReview
Key-Value (KV) cache has become a de facto component of modern Large Vision-Language Models (LVLMs) for inference. While it enhances decoding efficiency in...
kv cachemakelightweightopenreview
https://openreview.net/forum?id=EQgEMAD4kv
CAKE: Cascading and Adaptive KV Cache Eviction with Layer Preferences | OpenReview
Large language models (LLMs) excel at processing long sequences, boosting demand for key-value (KV) caching. While recent efforts to evict KV cache have...
kv cachecakecascadingadaptive
https://www.techtarget.com/searchstorage/news/366637161/Nvidias-new-KV-cache-makes-waves-in-enterprise-storage
Nvidia's new KV cache makes waves in enterprise storage | TechTarget
Nvidia unveiled a KV cache system with its Vera Rubin and BlueField-4 chips that raises competitive and memory shortage concerns in the industry.
kv cacheenterprise storagenvidianew
https://openreview.net/forum?id=L057s2Rq8O
KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache | OpenReview
Efficiently serving large language models (LLMs) requires batching many requests together to reduce the cost per request. Yet, the key-value (KV) cache, which...
kv cachekivituningfreeasymmetric
https://openreview.net/forum?id=5t4ZAkPiJs
ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification |...
KV cache stores key and value states from previous tokens to avoid re-computation, yet it demands substantial storage space, especially for long sequences....
kv cacheaccurateefficient
https://openreview.net/forum?id=gkUyYcY1W9
SCBench: A KV Cache-Centric Analysis of Long-Context Methods | OpenReview
Long-context Large Language Models (LLMs) have enabled numerous downstream applications but also introduced significant challenges related to computational and...
kv cachelong contextcentric
https://openreview.net/forum?id=hHoK1kBPd9
PiKV: KV Cache Management System for MoE Architecture | OpenReview
As large-scale language models continue to scale up in both size and context length, the memory and communication cost of key-value (KV) cache storage has...
kv cachemanagement systemmoearchitectureopenreview
https://openreview.net/forum?id=8g9fs6mdEG
Streaming Video Question-Answering with In-context Video KV-Cache Retrieval | OpenReview
We propose ReKV, a novel training-free approach that enables efficient streaming video question-answering (StreamingVQA), by seamlessly integrating with...
streaming videoquestion answeringkv cache
https://www.micron.com/about/blog/company/insights/from-buzzword-to-bottom-line-understanding-the-why-behind-kv-cache-in-ai
Understanding KV cache in AI | Micron Technology Inc.
Find out more about KV cache, the secret behind how AI remembers your input and creates an almost instant output no matter how long the conversation has been.
kv cachemicron technologyunderstandingaiinc
https://llmsresearch.com/
LLMs Research — Independent Lab on LLM Inference, KV Cache & Agents
LLMs Research is an independent applied research lab working on large language model efficiency: inference, KV cache compression, adaptive compute for...
llm inferencekv cachellmsresearchindependent
https://lmcache.ai/en/
LMCache – Building the foundation of AI memory tensor with KV Cache Infrastructure
building the foundation
https://openreview.net/forum?id=FJFVmeXusW
Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and...
Key-Value (KV) caching is a common technique to enhance the computational efficiency of Large Language Models (LLMs), but its memory overhead grows rapidly...
https://cloud.google.com/blog/topics/developers-practitioners/boosting-llm-performance-with-tiered-kv-cache-on-google-kubernetes-engine/
Boosting LLM Performance with Tiered KV Cache on Google Kubernetes Engine | Google Cloud Blog
Boost LLM inference performance with LMCache on Google Kubernetes Engine. Discover how tiered KV cache expands NVIDIA GPU HBM with CPU RAM and local SSDs,...
https://arxiv.org/html/2510.16807v2
Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads
https://arxiv.org/abs/2511.00321
[2511.00321] Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache...
Abstract page for arXiv paper 2511.00321: Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits
https://openreview.net/forum?id=uNrFpDPMyo&referrer=%5Bthe%20profile%20of%20Suyu%20Ge%5D(%2Fprofile%3Fid%3D~Suyu_Ge1)
Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs | OpenReview
In this study, we introduce adaptive KV cache compression, a plug-and-play method that reduces the memory footprint of generative inference for Large Language...
https://huggingface.co/papers/2503.16257
Paper page - Plug-and-Play 1.x-Bit KV Cache Quantization for Video Large Language Models
Join the discussion on this paper page
https://arxiv.org/abs/2510.16807v2
[2510.16807v2] Improving Model Representation and Reducing KV Cache via Skip Connections with First...
Abstract page for arXiv paper 2510.16807v2: Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads
https://arxiv.org/abs/2510.16807
[2510.16807] Improving Model Representation and Reducing KV Cache via Skip Connections with First...
Abstract page for arXiv paper 2510.16807: Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads
https://dev.to/arshtechpro/turboquant-what-developers-need-to-know-about-googles-kv-cache-compression-eeg
TurboQuant: What Developers Need to Know About Google's KV Cache Compression - DEV Community
If you've ever run a large language model on your own hardware and watched your GPU memory vanish as... Tagged with ai, python, google.
https://cryptobriefing.com/reiner-pope-batch-size-dramatically-impacts-ai-latency-and-cost-kv-cache-is-key-for-autoregressive-models-and-efficient-inference-can-save-resources-dwarkesh/
Reiner Pope: Batch size dramatically impacts AI latency and cost, kv cache is key for...
Apr 29, 2026 - Efficient batching in AI models can slash costs and boost performance by up to a thousand times.
https://aclanthology.org/2024.acl-long.602/
Layer-Condensed KV Cache for Efficient Inference of Large Language Models - ACL Anthology
Haoyi Wu, Kewei Tu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.