https://www.amazon.science/publications/yoro-lightweight-end-to-end-visual-grounding
YORO - Lightweight end to end visual grounding - Amazon Science
We present YORO - a multi-modal transformer encoder-only architecture for the Visual Grounding (VG) task. This task involves localizing, in an image, an object...
visual groundingyorolightweightendamazon
https://aclanthology.org/2020.lrec-1.527/
Visual Grounding Annotation of Recipe Flow Graph - ACL Anthology
Taichi Nishimura, Suzushi Tomori, Hayato Hashimoto, Atsushi Hashimoto, Yoko Yamakata, Jun Harashima, Yoshitaka Ushiku, Shinsuke Mori. Proceedings of the...
visual groundingrecipe flowannotationgraphacl
https://huggingface.co/papers/2307.08581
Paper page - BuboGPT: Enabling Visual Grounding in Multi-Modal LLMs
Join the discussion on this paper page
visual groundingmulti modalpaperenablingllms
https://openreview.net/forum?id=GlWzQhf2lV&referrer=%5Bthe%20profile%20of%20Ziqi%20Zhang%5D(%2Fprofile%3Fid%3D~Ziqi_Zhang5)
Exploiting Contextual Objects and Relations for 3D Visual Grounding | OpenReview
3D visual grounding, the task of identifying visual objects in 3D scenes based on natural language inputs, plays a critical role in enabling machines to...
visual groundingexploitingcontextualobjectsrelations
https://openreview.net/forum?id=dgQdvPZnH-t
LanguageRefer: Spatial-Language Model for 3D Visual Grounding | OpenReview
For robots to understand human instructions and perform meaningful tasks in the near future, it is important to develop learned models that comprehend...
language modelvisual groundingspatial3dopenreview
https://deepai.org/publication/learning-to-compose-and-reason-with-language-tree-structures-for-visual-grounding
Learning to Compose and Reason with Language Tree Structures for Visual Grounding | DeepAI
Jun 5, 2019 - 06/05/19 - Grounding natural language in images, such as localizing
https://huggingface.co/papers/2312.15043
Paper page - GroundVLP: Harnessing Zero-shot Visual Grounding from Vision-Language Pre-training and...
Join the discussion on this paper page
https://openreview.net/forum?id=zLWJR53KxC
Talk to Parallel LiDARs: A Human-LiDAR Interaction Method Based on 3D Visual Grounding | OpenReview
LiDAR sensors play a crucial role in various applications, especially in autonomous driving. Current research primarily focuses on optimizing perceptual models...
https://drive.google.com/file/d/1x3OOfIpW62DciVx6cFVX6JoZGJNVNZ2B/view?usp=sharing
The 5 Senses Grounding Technique Visual Cue Cards.pdf - Google Drive
the 5 sensesgrounding techniquecue cards
https://www.amazon.science/publications/mind-the-context-the-impact-of-contextualization-in-neural-module-networks-for-grounding-visual-referring-expression
Mind the context: The impact of contextualization in neural module networks for grounding visual...
Neural module networks (NMN) are a popular approach for grounding visual referring expressions. Prior implementations of NMN use pre-defined and fixed textual...
https://arxiv.org/html/2411.03405v1
Fine-Grained Spatial and Verbal Losses for 3D Visual Grounding
fine grainedspatialverballosses3d