Robuta

https://openreview.net/forum?id=n77IeySrQl Make Your LVLM KV Cache More Lightweight | OpenReview Key-Value (KV) cache has become a de facto component of modern Large Vision-Language Models (LVLMs) for inference. While it enhances decoding efficiency in... kv cachemakelightweightopenreview https://openreview.net/forum?id=EQgEMAD4kv CAKE: Cascading and Adaptive KV Cache Eviction with Layer Preferences | OpenReview Large language models (LLMs) excel at processing long sequences, boosting demand for key-value (KV) caching. While recent efforts to evict KV cache have... kv cachecakecascadingadaptive https://www.techtarget.com/searchstorage/news/366637161/Nvidias-new-KV-cache-makes-waves-in-enterprise-storage Nvidia's new KV cache makes waves in enterprise storage | TechTarget Nvidia unveiled a KV cache system with its Vera Rubin and BlueField-4 chips that raises competitive and memory shortage concerns in the industry. kv cacheenterprise storagenvidianew https://openreview.net/forum?id=L057s2Rq8O KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache | OpenReview Efficiently serving large language models (LLMs) requires batching many requests together to reduce the cost per request. Yet, the key-value (KV) cache, which... kv cachekivituningfreeasymmetric https://openreview.net/forum?id=5t4ZAkPiJs ZipCache: Accurate and Efficient KV Cache Quantization with Salient Token Identification |... KV cache stores key and value states from previous tokens to avoid re-computation, yet it demands substantial storage space, especially for long sequences.... kv cacheaccurateefficient https://openreview.net/forum?id=gkUyYcY1W9 SCBench: A KV Cache-Centric Analysis of Long-Context Methods | OpenReview Long-context Large Language Models (LLMs) have enabled numerous downstream applications but also introduced significant challenges related to computational and... kv cachelong contextcentric https://openreview.net/forum?id=hHoK1kBPd9 PiKV: KV Cache Management System for MoE Architecture | OpenReview As large-scale language models continue to scale up in both size and context length, the memory and communication cost of key-value (KV) cache storage has... kv cachemanagement systemmoearchitectureopenreview https://openreview.net/forum?id=8g9fs6mdEG Streaming Video Question-Answering with In-context Video KV-Cache Retrieval | OpenReview We propose ReKV, a novel training-free approach that enables efficient streaming video question-answering (StreamingVQA), by seamlessly integrating with... streaming videoquestion answeringkv cache https://www.micron.com/about/blog/company/insights/from-buzzword-to-bottom-line-understanding-the-why-behind-kv-cache-in-ai Understanding KV cache in AI | Micron Technology Inc. Find out more about KV cache, the secret behind how AI remembers your input and creates an almost instant output no matter how long the conversation has been. kv cachemicron technologyunderstandingaiinc https://llmsresearch.com/ LLMs Research — Independent Lab on LLM Inference, KV Cache & Agents LLMs Research is an independent applied research lab working on large language model efficiency: inference, KV cache compression, adaptive compute for... llm inferencekv cachellmsresearchindependent https://lmcache.ai/en/ LMCache – Building the foundation of AI memory tensor with KV Cache Infrastructure building the foundation https://openreview.net/forum?id=FJFVmeXusW Not All Heads Matter: A Head-Level KV Cache Compression Method with Integrated Retrieval and... Key-Value (KV) caching is a common technique to enhance the computational efficiency of Large Language Models (LLMs), but its memory overhead grows rapidly... https://cloud.google.com/blog/topics/developers-practitioners/boosting-llm-performance-with-tiered-kv-cache-on-google-kubernetes-engine/ Boosting LLM Performance with Tiered KV Cache on Google Kubernetes Engine | Google Cloud Blog Boost LLM inference performance with LMCache on Google Kubernetes Engine. Discover how tiered KV cache expands NVIDIA GPU HBM with CPU RAM and local SSDs,... https://arxiv.org/html/2510.16807v2 Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads https://arxiv.org/abs/2511.00321 [2511.00321] Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache... Abstract page for arXiv paper 2511.00321: Scalable Processing-Near-Memory for 1M-Token LLM Inference: CXL-Enabled KV-Cache Management Beyond GPU Limits https://openreview.net/forum?id=uNrFpDPMyo&referrer=%5Bthe%20profile%20of%20Suyu%20Ge%5D(%2Fprofile%3Fid%3D~Suyu_Ge1) Model Tells You What to Discard: Adaptive KV Cache Compression for LLMs | OpenReview In this study, we introduce adaptive KV cache compression, a plug-and-play method that reduces the memory footprint of generative inference for Large Language... https://huggingface.co/papers/2503.16257 Paper page - Plug-and-Play 1.x-Bit KV Cache Quantization for Video Large Language Models Join the discussion on this paper page https://arxiv.org/abs/2510.16807v2 [2510.16807v2] Improving Model Representation and Reducing KV Cache via Skip Connections with First... Abstract page for arXiv paper 2510.16807v2: Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads https://arxiv.org/abs/2510.16807 [2510.16807] Improving Model Representation and Reducing KV Cache via Skip Connections with First... Abstract page for arXiv paper 2510.16807: Improving Model Representation and Reducing KV Cache via Skip Connections with First Value Heads https://dev.to/arshtechpro/turboquant-what-developers-need-to-know-about-googles-kv-cache-compression-eeg TurboQuant: What Developers Need to Know About Google's KV Cache Compression - DEV Community If you've ever run a large language model on your own hardware and watched your GPU memory vanish as... Tagged with ai, python, google. https://cryptobriefing.com/reiner-pope-batch-size-dramatically-impacts-ai-latency-and-cost-kv-cache-is-key-for-autoregressive-models-and-efficient-inference-can-save-resources-dwarkesh/ Reiner Pope: Batch size dramatically impacts AI latency and cost, kv cache is key for... Apr 29, 2026 - Efficient batching in AI models can slash costs and boost performance by up to a thousand times. https://aclanthology.org/2024.acl-long.602/ Layer-Condensed KV Cache for Efficient Inference of Large Language Models - ACL Anthology Haoyi Wu, Kewei Tu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.