Robuta

https://huggingface.co/papers/2501.08326 Paper page - Omni-RGPT: Unifying Image and Video Region-level Understanding via Token Marks Join the discussion on this paper page