Robuta

https://openreview.net/forum?id=HT0PSgajyN&referrer=%5Bthe%20profile%20of%20Qinsi%20Wang%5D(%2Fprofile%3Fid%3D~Qinsi_Wang2) CoreMatching: A Co-adaptive Sparse Inference Framework with Token and Neuron Pruning for... Vision-Language Models (VLMs) excel across diverse tasks but suffer from high inference costs in time and memory. Token sparsity mitigates inefficiencies in...