Robuta

https://arxiv.org/abs/2602.22678 [2602.22678] ViCLIP-OT: The First Foundation Vision-Language Model for Vietnamese Image-Text... Abstract page for arXiv paper 2602.22678: ViCLIP-OT: The First Foundation Vision-Language Model for Vietnamese Image-Text Retrieval with Optimal Transport