Robuta

https://arxiv.org/abs/2404.15420v3 [2404.15420v3] XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference Abstract page for arXiv paper 2404.15420v3: XC-Cache: Cross-Attending to Cached Context for Efficient LLM Inference