Robuta

https://arxiv.org/abs/2405.20985 [2405.20985] DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large... Abstract page for arXiv paper 2405.20985: DeCo: Decoupling Token Compression from Semantic Abstraction in Multimodal Large Language Models token compression https://thetokencompany.com/ Prompt Compression API | Cut LLM Token Costs | The Token Company Cut OpenAI, Anthropic, and Gemini API costs with accuracy held flat — or take a smaller cut and lift accuracy by several points instead. The bear-2 prompt... token costspromptcompressionapicut https://aclanthology.org/2023.findings-emnlp.655/ TCRA-LLM: Token Compression Retrieval Augmented Large Language Model for Inference Cost Reduction -... Junyi Liu, Liangzhi Li, Tong Xiang, Bowen Wang, Yiming Qian. Findings of the Association for Computational Linguistics: EMNLP 2023. 2023. large language model