https://openreview.net/forum?id=wiBEFdAvl8L
GLIPv2: Unifying Localization and Vision-Language Understanding | OpenReview
We present a region-aware vision-language pre-trained model that serves both localization tasks (e.g., object detection, instance segmentation) and...
and visionlanguage understandingunifyinglocalizationopenreview