https://robotics-transformer2.github.io/
RT-2: Vision-Language-Action Models
rt 2vision languageactionmodels
https://robotics-transformer.github.io/
RT-2 Vision-Language-Action
rt 2vision languageaction
https://www.anjiecheng.me/SpatialRGPT
SpatialRGPT: Grounded Spatial Reasoning in Vision-Language Models
SpatialRGPT: Grounded Spatial Reasoning in Vision-Language Models
spatial reasoningvision languagegroundedmodels
https://arxiv.org/html/2412.17417v2
Synth-Align: Improving Trustworthiness in Vision-Language Model with Synthetic Preference Data...
vision language model
https://huggingface.co/papers/2505.10610
Paper page - MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and...
Join the discussion on this paper page
vision language modelslong contextpaperbenchmarking
https://openreview.net/forum?id=sb7qHFYwBc
C-CLIP: Multimodal Continual Learning for Vision-Language Model | OpenReview
Multimodal pre-trained models like CLIP need large image-text pairs for training but often struggle with domain-specific tasks. Since retraining with...
vision language modelc clipcontinual learningmultimodalopenreview
https://openreview.net/forum?id=xkgfLXZ4e0
Correlating instruction-tuning (in multimodal models) with vision-language processing (in the...
Transformer-based language models, though not explicitly trained to mimic brain recordings, have demonstrated surprising alignment with brain activity....
instruction tuningmultimodal modelswith visionlanguage processing
https://www.thoughtworks.com/en-gb/radar/techniques/vision-language-models-for-end-to-end-document-parsing
Vision language models for end-to-end document parsing | Technology Radar | Thoughtworks United...
Document parsing often relies on multi-stage pipelines combining layout detection, traditional OCR and post-processing scripts. These approaches often struggle...
vision language models
https://www.datacamp.com/it/blog/vlms-ai-vision-language-models
Vision Language Models (VLMs) Explained | DataCamp
Vision language models (VLMs) are AI models that can understand and process both visual and textual data, enabling tasks like image captioning, visual question...
vision language modelsexplaineddatacamp
https://openreview.net/forum?id=mmpwZfXUk7&referrer=%5Bthe%20profile%20of%20Vladan%20Stojni%C4%87%5D(%2Fprofile%3Fid%3D~Vladan_Stojni%C4%871)
Label Propagation for Zero-shot Classification with Vision-Language Models | OpenReview
Vision-Language Models (VLMs) have demonstrated impressive performance on zero-shot classification, i.e. classification when provided merely with a list of...
vision language modelslabel propagationzero shot
https://www.thoughtworks.com/en-br/radar/techniques/vision-language-models-for-end-to-end-document-parsing
Vision language models for end-to-end document parsing | Technology Radar | Thoughtworks Brazil
Document parsing often relies on multi-stage pipelines combining layout detection, traditional OCR and post-processing scripts. These approaches often struggle...
vision language models
https://www.thoughtworks.com/es-ec/radar/techniques/vision-language-models-for-end-to-end-document-parsing
Vision language models for end-to-end document parsing | Technology Radar | Thoughtworks Ecuador
Document parsing often relies on multi-stage pipelines combining layout detection, traditional OCR and post-processing scripts. These approaches often struggle...
vision language models
https://openreview.net/forum?id=oSJqRF0Tkg&referrer=%5Bthe%20profile%20of%20Hongming%20Zhang%5D(%2Fprofile%3Fid%3D~Hongming_Zhang2)
Leopard: A Vision Language Model For Text-Rich Multi-Image Tasks | OpenReview
Text-rich images, where text serves as the central visual element guiding the overall understanding, are prevalent in real-world applications, such as...
vision language model
https://www.amazon.science/tag/vision-language-models-vlms
Vision-language models (VLMs) - Amazon Science
vision language modelsamazonscience
https://openreview.net/forum?id=Q5RYn6jagC
Understanding the Limits of Vision Language Models Through the Lens of the Binding Problem |...
Recent work has documented striking heterogeneity in the performance of state-of-the-art vision language models (VLMs), including both multimodal language...
the limits of visionlanguage modelsunderstanding
https://openreview.net/forum?id=eRleg6vy0Y&referrer=%5Bthe%20profile%20of%20Alyssa%20Unell%5D(%2Fprofile%3Fid%3D~Alyssa_Unell1)
Micro-Bench: A Microscopy Benchmark for Vision-Language Understanding | OpenReview
Recent advances in microscopy have enabled the rapid generation of terabytes of image data in cell biology and biomedical research. Vision-language models...
vision languagemicrobenchunderstandingopenreview
https://openreview.net/forum?id=JUwczEJY8I
Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning | OpenReview
Reinforcement learning (RL) requires either manually specifying a reward function, which is often infeasible, or learning a reward model from a large amount of...
vision language modelszero shotreinforcement learning
https://civitai.com/articles/27256/a-thousand-words-image-captioning-vision-language-model-interface
A Thousand Words - Image Captioning (Vision Language Model) interface | Civitai
I've spent a lot of time creating various "batch processing scripts" for various VLM's in the past ( Github repo search ). Instead, I decided to sp...
a thousand wordsvision language modelimage captioninginterfacecivitai
https://openreview.net/forum?id=zY37C8d6bS&referrer=%5Bthe%20profile%20of%20Ruifeng%20Chen%5D(%2Fprofile%3Fid%3D~Ruifeng_Chen1)
Semantic Temporal Abstraction via Vision-Language Model Guidance for Efficient Reinforcement...
Extracting temporally extended skills can significantly improve the efficiency of reinforcement learning (RL) by breaking down complex decision-making problems...
vision language modeltemporal abstractionsemanticvia
https://www.thoughtworks.com/en-cn/radar/techniques/vision-language-models-for-end-to-end-document-parsing
Vision language models for end-to-end document parsing | Technology Radar | Thoughtworks China
Document parsing often relies on multi-stage pipelines combining layout detection, traditional OCR and post-processing scripts. These approaches often struggle...
vision language models
https://arxiv.org/abs/2503.16538v1
[2503.16538v1] Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and...
Abstract page for arXiv paper 2503.16538v1: Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking
vision language models
https://www.thoughtworks.com/es-cl/radar/techniques/vision-language-models-for-end-to-end-document-parsing
Vision language models for end-to-end document parsing | Technology Radar | Thoughtworks Chile
Document parsing often relies on multi-stage pipelines combining layout detection, traditional OCR and post-processing scripts. These approaches often struggle...
vision language models
https://openreview.net/forum?id=Kt2VJrCKo4
VoLTA: Vision-Language Transformer with Weakly-Supervised Local-Feature Alignment | OpenReview
Vision-language pre-training (VLP) has recently proven highly effective for various uni- and multi-modal downstream applications. However, most existing...
vision languagevoltatransformer
https://www.autopilot-cvpr.net/
AUTOPILOT Workshop @ CVPR 2026 | Autonomous Driving & Vision-Language Models
Join the 3rd AUTOPILOT Workshop at CVPR 2026 in Denver, Colorado. Exploring autonomous understanding through open-world perception and integrated language...
autonomous drivingvision languageautopilotworkshopcvpr
https://www.qualcomm.com/developer/software/qualcomm-interactive-video-dataset-qivd
Interactive Video Dataset for Vision-Language Models | Qualcomm
Improve your AI's visual understanding and response with QIVD, featuring 2,900 video files and 13 categories of action attributes and object detection.
vision language modelsinteractive videodatasetqualcomm
https://openreview.net/forum?id=tZozeR3VV7
Backdooring Vision-Language Models with Out-Of-Distribution Data | OpenReview
The emergence of Vision-Language Models (VLMs) represents a significant advancement in integrating computer vision with Large Language Models (LLMs) to...
vision language modelswith outdistribution dataopenreview
https://vla-survey.github.io/
Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications
comprehensive survey of Vision-Language-Action (VLA) models for robotics. Explore the latest research on VLA architectures, learning paradigms, and real-world...
vision language
https://openreview.net/forum?id=ohJxgRLlLt
Large (Vision) Language Models are Unsupervised In-Context Learners | OpenReview
Recent advances in large language and vision-language models have enabled zero-shot inference, allowing models to solve new tasks without task-specific...
vision language modelsin contextlargeunsupervisedlearners
https://openreview.net/forum?id=hdYqGkSr9S&referrer=%5Bthe%20profile%20of%20Joseph%20Tighe%5D(%2Fprofile%3Fid%3D~Joseph_Tighe1)
Benchmarking Zero-Shot Recognition with Vision-Language Models: Challenges on Granularity and...
This paper introduces innovative benchmarks to evaluate Vision-Language Models (VLMs) in real-world zero-shot recognition tasks, focusing on the pivotal...
vision language modelszero shot
https://openreview.net/forum?id=trj2Jq8riA
Interpretable Vision-Language Survival Analysis with Ordinal Inductive Bias for Computational...
Histopathology Whole-Slide Images (WSIs) provide an important tool to assess cancer prognosis in computational pathology (CPATH). While existing survival...
vision languagesurvival analysisinductive bias
https://openreview.net/forum?id=7DpNGceUUV
Coordinated Robustness Evaluation Framework for Vision Language Models | OpenReview
Vision-language models, which integrate computer vision and natural language processing capabilities, have demonstrated significant advancements in tasks such...
vision language modelsevaluation frameworkcoordinatedrobustnessopenreview
https://openreview.net/forum?id=kcnIHUVeLMm&referrer=%5Bthe%20profile%20of%20Xingcheng%20Zhou%5D(%2Fprofile%3Fid%3D~Xingcheng_Zhou1)
Vision Language Models in Autonomous Driving and Intelligent Transportation Systems | OpenReview
The applications of Vision-Language Models (VLMs) in the fields of Autonomous Driving (AD) and Intelligent Transportation Systems (ITS) have attracted...
vision language modelsautonomous drivingintelligent transportation
https://arxiv.org/abs/2605.00583
[2605.00583] Jailbreaking Vision-Language Models Through the Visual Modality
Abstract page for arXiv paper 2605.00583: Jailbreaking Vision-Language Models Through the Visual Modality
vision language modelsthe visualjailbreakingmodality
https://www.utwente.nl/en/eemcs/dmb/assignments/open/master/Computer%20Vision%20and%20Biometrics/20260203-Hyperbolic%20Vision%20Language%20Models%20(VLMs)%20in%20medical%20imaging/
Hyperbolic Vision Language Models (VLMs) in medical imaging | Computer Vision and Biometrics |...
vision language modelsmedical imaginghyperbolic
https://arxiv.org/abs/2508.19652
[2508.19652] Self-Rewarding Vision-Language Model via Reasoning Decomposition
Abstract page for arXiv paper 2508.19652: Self-Rewarding Vision-Language Model via Reasoning Decomposition
vision language modelselfrewardingviareasoning
https://openreview.net/forum?id=3uVAA3ckxT
ADAPT: Adaptive Prompt Tuning for Vision-Language Models | OpenReview
Prompt tuning has emerged as an effective way for parameter-efficient fine-tuning. Conventional deep prompt tuning inserts continuous prompts of a fixed...
vision language modelsprompt tuningadaptopenreview
https://www.thoughtworks.com/en-ca/radar/techniques/vision-language-models-for-end-to-end-document-parsing
Vision language models for end-to-end document parsing | Technology Radar | Thoughtworks Canada
Document parsing often relies on multi-stage pipelines combining layout detection, traditional OCR and post-processing scripts. These approaches often struggle...
vision language models
https://arxiv.org/abs/2503.16538v2
[2503.16538v2] Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and...
Abstract page for arXiv paper 2503.16538v2: Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking
vision language models
https://www.wur.nl/en/vacancy/phd-position-vision-language-models-biodiversity-environmental-monitoring
PhD position: Vision-Language Models for Biodiversity & Environmental Monitoring | WUR
We are looking for a PhD to develop vision-language models for applications in biodiversity and environmental monitoring, using Earth observation, drone, and...
vision language modelsphd positionenvironmental monitoringbiodiversitywur
https://aclanthology.org/2024.lrec-main.973/
Medical Vision-Language Pre-Training for Brain Abnormalities - ACL Anthology
Masoud Monajatipoor, Zi-Yi Dou, Aichi Chien, Nanyun Peng, Kai-Wei Chang. Proceedings of the 2024 Joint International Conference on Computational Linguistics,...
medical visionpre trainingfor brainlanguageabnormalities
https://openreview.net/forum?id=U7aGAc0S5h&referrer=%5Bthe%20profile%20of%20Zhenlin%20Xu%5D(%2Fprofile%3Fid%3D~Zhenlin_Xu1)
Benchmarking Zero-Shot Recognition with Vision-Language Models: Challenges on Granularity and...
This paper presents novel benchmarks for evaluating vision-language models (VLMs) in zero-shot recognition, focusing on granularity and specificity. Although...
vision language modelszero shot
https://huggingface.co/papers/2305.10355
Paper page - Evaluating Object Hallucination in Large Vision-Language Models
Join the discussion on this paper page
large visionpaperevaluatingobjecthallucination
https://openreview.net/forum?id=ELR4ifkX3j
Efficient Open-set Test Time Adaptation of Vision Language Models | OpenReview
In dynamic real-world settings, models must adapt to changing data distributions, a challenge known as Test Time Adaptation (TTA). This becomes even more...
vision language modelsopen setefficienttesttime
https://www.amazon.science/publications/prompting-vision-language-models-for-aspect-controlled-generation-of-referring-expressions
Prompting vision-language models for aspect-controlled generation of referring expressions - Amazon...
Referring Expression Generation (REG) is the task of generating a description that unambiguously identifies a given target in the scene. Different from Image...
vision language models
https://www.datacamp.com/pl/blog/vlms-ai-vision-language-models
Vision Language Models (VLMs) Explained | DataCamp
Vision language models (VLMs) are AI models that can understand and process both visual and textual data, enabling tasks like image captioning, visual question...
vision language modelsexplaineddatacamp
https://arxiv.org/abs/2507.18053
[2507.18053] Resource Consumption Red-Teaming for Large Vision-Language Models
Abstract page for arXiv paper 2507.18053: Resource Consumption Red-Teaming for Large Vision-Language Models
resource consumptionred teaminglarge vision
https://openreview.net/forum?id=d5DJWgmMoX
Can Vision Language Models Learn from Visual Demonstrations of Ambiguous Spatial Reasoning? |...
Large vision-language models (VLMs) have become state-of-the-art for many computer vision tasks, with in-context learning (ICL) as a popular adaptation...
vision language models
https://vlm-jailbreaks.github.io/
Jailbreaking Vision-Language Models Through the Visual Modality | ICML 2026
Four visual jailbreak attacks exploiting the vision component of VLMs. ICML 2026.
vision language modelsthe visualjailbreakingmodalityicml
https://openreview.net/forum?id=N0I2RtD8je
Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning | OpenReview
Reinforcement learning (RL) requires either manually specifying a reward function, which is often infeasible, or learning a reward model from a large amount of...
vision language modelszero shotreinforcement learning
https://oecd.ai/en/catalogue/metric-use-cases/pali-3-vision-language-models-smaller,-faster,-stronger
PaLI-3 Vision Language Models: Smaller, Faster, Stronger - OECD.AI
The interest of the machine learning community in image synthesis has grown significantly in recent years, with the introduction of a wide range of deep...
vision language modelspali3smallerfaster
https://www.findbestmodel.app/
ModelMatch - Compare Vision-Language Models
Compare top open source vision-language models side-by-side, no coding needed
vision languagecomparemodels
https://arxiv.org/abs/2210.05335
[2210.05335] MAP: Multimodal Uncertainty-Aware Vision-Language Pre-training Model
Abstract page for arXiv paper 2210.05335: MAP: Multimodal Uncertainty-Aware Vision-Language Pre-training Model
vision languagepre training2210mapmultimodal
https://arxiv.org/abs/2405.10529?context=cs.CV
[2405.10529] Safeguarding Vision-Language Models Against Patched Visual Prompt Injectors
Abstract page for arXiv paper 2405.10529: Safeguarding Vision-Language Models Against Patched Visual Prompt Injectors
vision language models2405safeguarding
https://www.thoughtworks.com/radar/techniques/vision-language-models-for-end-to-end-document-parsing
Vision language models for end-to-end document parsing | Technology Radar | Thoughtworks
Document parsing often relies on multi-stage pipelines combining layout detection, traditional OCR and post-processing scripts. These approaches often struggle...
vision language modelsdocument parsingtechnology radarend
https://openreview.net/forum?id=iwp3H8uSeK
Distilling Out-of-Distribution Robustness from Vision-Language Foundation Models | OpenReview
We propose a conceptually simple and lightweight framework for improving the robustness of vision models through the combination of knowledge distillation and...
out offrom visionfoundation modelsdistillingdistribution
https://openreview.net/forum?id=0y3hGn1wOk
Benchmarking Vision Language Model Unlearning via Fictitious Facial Identity Dataset | OpenReview
Machine unlearning has emerged as an effective strategy for forgetting specific information in the training data. However, with the increasing integration of...
vision language modelbenchmarkingunlearning
https://www.frontiersin.org/research-topics/18532/identifying-analyzing-and-overcoming-challenges-in-vision-and-language-research
Frontiers | Identifying, Analyzing, and Overcoming Challenges in Vision and Language Research
Integration of computer vision and natural language processing is one of the holy grails in machine learning and artificial intelligence research with import...
overcoming challengesvision languagefrontiersidentifyinganalyzing
https://arxiv.org/abs/2405.19315
[2405.19315] Matryoshka Query Transformer for Large Vision-Language Models
Abstract page for arXiv paper 2405.19315: Matryoshka Query Transformer for Large Vision-Language Models
large vision2405matryoshkaquerytransformer
https://openreview.net/forum?id=gbc0m68fdw
Manipulate-Anything: Automating Real-World Robots using Vision-Language Models | OpenReview
Large-scale endeavors like RT-1 and widespread community efforts such as Open-X-Embodiment have contributed to growing the scale of robot demonstration data....
vision language modelsreal worldmanipulateanythingautomating
https://openreview.net/forum?id=C3PsmuhDq0&referrer=%5Bthe%20profile%20of%20Shayekh%20Bin%20Islam%5D(%2Fprofile%3Fid%3D~Shayekh_Bin_Islam1)
Behind Maya: Building a Multilingual Vision Language Model | OpenReview
In recent times, we have seen a rapid development of large Vision-Language Models (VLMs). They have shown impressive results on academic benchmarks, primarily...
vision language modelbehindmayabuildingmultilingual
https://openreview.net/forum?id=NEBa0bs5LR&referrer=%5Bthe%20profile%20of%20Yunhang%20Shen%5D(%2Fprofile%3Fid%3D~Yunhang_Shen1)
DS-VLM: Diffusion Supervision Vision Language Model | OpenReview
Vision-Language Models (VLMs) face two critical limitations in visual representation learning: degraded supervision due to information loss during gradient...
vision language modeldsvlmdiffusionsupervision
https://huggingface.co/papers/2603.03739
Paper page - PROSPECT: Unified Streaming Vision-Language Navigation via Semantic--Spatial Fusion...
Join the discussion on this paper page
vision language
https://openreview.net/forum?id=pQDzXVeBja
Effortless Vision-Language Model Specialization in Histopathology without Annotation | OpenReview
Recent advances in Vision-Language Models (VLMs) in histopathology, such as CONCH and QuiltNet, have demonstrated impressive zero-shot classification...
vision language modeleffortlessspecializationhistopathologywithout
https://www.thoughtworks.com/en-au/radar/techniques/vision-language-models-for-end-to-end-document-parsing
Vision language models for end-to-end document parsing | Technology Radar | Thoughtworks Australia
Document parsing often relies on multi-stage pipelines combining layout detection, traditional OCR and post-processing scripts. These approaches often struggle...
vision language models
https://openreview.net/forum?id=6zoKzrRn2h&referrer=%5Bthe%20profile%20of%20Zhihe%20Lu%5D(%2Fprofile%3Fid%3D~Zhihe_Lu1)
Beyond Sole Strength: Customized Ensembles for Generalized Vision-Language Models | OpenReview
vision language modelsbeyondsolestrengthcustomized
https://openreview.net/forum?id=WutpswD3ea
Assisted Few-Shot Learning for Vision-Language Models in Agricultural Stress Phenotype...
In the agricultural sector, labeled data for crop diseases and stresses are often scarce due to high annotation costs. We propose an Assisted Few-Shot Learning...
few shot learningvision language models
https://arxiv.org/abs/2505.13302?ref=disinfodocket.com
[2505.13302] Images Amplify Misinformation Sharing in Vision-Language Models
Abstract page for arXiv paper 2505.13302: Images Amplify Misinformation Sharing in Vision-Language Models
in visionimagesamplifymisinformationsharing
https://www.thoughtworks.com/en-ec/radar/techniques/vision-language-models-for-end-to-end-document-parsing
Vision language models for end-to-end document parsing | Technology Radar | Thoughtworks Ecuador
Document parsing often relies on multi-stage pipelines combining layout detection, traditional OCR and post-processing scripts. These approaches often struggle...
vision language models
https://openreview.net/forum?id=hssNWbMZHF&referrer=%5Bthe%20profile%20of%20Marzyeh%20Ghassemi%5D(%2Fprofile%3Fid%3D~Marzyeh_Ghassemi2)
Vision-Language Models Do Not Understand Negation | OpenReview
Many practical vision-language applications require models that understand negation, e.g., when using natural language to retrieve images which contain certain...
vision language modelsdo notunderstandnegationopenreview
https://openreview.net/forum?id=vQDKYYuqWA
Vision-Language Models Provide Promptable Representations for Reinforcement Learning | OpenReview
Humans can quickly learn new behaviors by leveraging background world knowledge. In contrast, agents trained with reinforcement learning (RL) typically learn...
vision language modelsreinforcement learningproviderepresentationsopenreview
https://arxiv.org/abs/2508.08604v1
[2508.08604v1] Transferable Model-agnostic Vision-Language Model Adaptation for Efficient...
Abstract page for arXiv paper 2508.08604v1: Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization
vision languagetransferablemodelagnosticadaptation
https://vislang.ai/
VisLang - Vision, Language and Learning Lab at Rice University
Research group at Rice University led by Vicente Ordonez, working at the intersection of computer vision, natural language processing, and machine learning.
vision languagelearning labriceuniversity
https://openreview.net/forum?id=XgYZT35N76&referrer=%5Bthe%20profile%20of%20Yiming%20Yang%5D(%2Fprofile%3Fid%3D~Yiming_Yang1)
Improve Vision Language Model Chain-of-thought Reasoning | OpenReview
Chain-of-thought (CoT) reasoning in vision language models (VLMs) is crucial for improving interpretability and trustworthiness. However, current training...
chain of thought reasoningvision language modelimproveopenreview
https://www.thoughtworks.com/es-es/radar/techniques/vision-language-models-for-end-to-end-document-parsing
Vision language models for end-to-end document parsing | Technology Radar | Thoughtworks Spain
Document parsing often relies on multi-stage pipelines combining layout detection, traditional OCR and post-processing scripts. These approaches often struggle...
vision language models
https://huggingface.co/papers/2210.09996
Paper page - Perceptual Grouping in Contrastive Vision-Language Models
Join the discussion on this paper page
vision languagepaperperceptualgroupingcontrastive
https://www.analyticsvidhya.com/blog/2024/03/introducing-moondream2-a-tiny-vision-language-model/
Introducing Moondream2: A tiny Vision-Language Model
Mar 30, 2024 - Explore Moondream2, a small vision-language model for efficient image-text tasks. Components, implementation, applications, and limitations.
vision languageintroducingtinymodel
https://arxiv.org/abs/2604.02318v1
[2604.02318v1] Stop Wandering: Efficient Vision-Language Navigation via Metacognitive Reasoning
Abstract page for arXiv paper 2604.02318v1: Stop Wandering: Efficient Vision-Language Navigation via Metacognitive Reasoning
vision languagestopwanderingefficient
https://openreview.net/forum?id=iVxxgZlXh6
LLaRA: Supercharging Robot Learning Data for Vision-Language Policy | OpenReview
Vision Language Models (VLMs) have recently been leveraged to generate robotic actions, forming Vision-Language-Action (VLA) models. However, directly adapting...
robot learningdata forvision languagesuperchargingpolicy
https://arxiv.org/html/2603.02972v1
TagaVLM: Topology-Aware Global Action Reasoning for Vision-Language Navigation
global actionvision languagetopologyawarereasoning
https://speakerdeck.com/keio_smilab/the-confluence-of-vision-language-and-robotics
The Confluence of Vision, Language, and Robotics - Speaker Deck
the confluencevision languageroboticsspeakerdeck
https://openreview.net/forum?id=vvoWPYqZJA
InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning | OpenReview
Large-scale pre-training and instruction tuning have been successful at creating general-purpose language models with broad competence. However, building...
vision language modelsgeneral purposeinstruction tuningtowards
https://huggingface.co/papers/2311.01477
Paper page - FAITHSCORE: Evaluating Hallucinations in Large Vision-Language Models
Join the discussion on this paper page
large visionpaperevaluatinghallucinationslanguage
https://www.datacamp.com/blog/top-vision-language-models
Top 10 Vision Language Models in 2026 | DataCamp
Discover the top open-source and proprietary vision-language models of 2026 for visual reasoning, image analysis, and computer vision.
vision language modelstop 102026datacamp
https://deepai.org/publication/illume-rationalizing-vision-language-models-by-interacting-with-their-jabber
ILLUME: Rationalizing Vision-Language Models by Interacting with their Jabber | DeepAI
Aug 17, 2022 - 08/17/22 - Bootstrapping from pre-trained language models has been proven to be an efficient approach for building foundation vision-language...
vision language modelsillumerationalizing
https://huggingface.co/papers/2401.12168
Paper page - SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Join the discussion on this paper page
vision language modelsspatial reasoningpaper
https://openreview.net/forum?id=SB2sWF3oCw
Understanding and Improving In-Context Learning on Vision-language Models | OpenReview
In-context learning (ICL) on large language models (LLMs) has received great attention, and this technique can also be applied to vision-language models (VLMs)...
in context learningvision language modelsunderstandingimproving
https://arxiv.org/html/2407.13588v1
Robust Calibration of Large Vision-Language Adapters
large visionrobustcalibrationlanguageadapters
https://arxiv.org/abs/2407.01449?ref=blog.mixpeek.com
[2407.01449] ColPali: Efficient Document Retrieval with Vision Language Models
Abstract page for arXiv paper 2407.01449: ColPali: Efficient Document Retrieval with Vision Language Models
document retrievalwith vision2407colpaliefficient
https://arxiv.org/abs/2306.11300
[2306.11300] RS5M and GeoRSCLIP: A Large Scale Vision-Language Dataset and A Large Vision-Language...
Abstract page for arXiv paper 2306.11300: RS5M and GeoRSCLIP: A Large Scale Vision-Language Dataset and A Large Vision-Language Model for Remote Sensing
a largevision language2306
https://www.thoughtworks.com/en-sg/radar/techniques/vision-language-models-for-end-to-end-document-parsing
Vision language models for end-to-end document parsing | Technology Radar | Thoughtworks Singapore
Document parsing often relies on multi-stage pipelines combining layout detection, traditional OCR and post-processing scripts. These approaches often struggle...
vision language models
https://openreview.net/forum?id=R2rasAEPVi
Leopard: A Vision Language Model for Text-Rich Multi- Image Tasks | OpenReview
Text-rich images, where text serves as the central visual element guiding the overall understanding, are prevalent in real-world applications, such as...
vision language model
https://openreview.net/forum?id=haJHr4UsQX
Causal Graphical Models for Vision-Language Compositional Understanding | OpenReview
Recent work has empirically shown that Vision-Language Models (VLMs) struggle to fully understand the compositional properties of the human language, usually...
graphical modelsvision languagecausalcompositionalunderstanding
https://www.esri.com/arcgis-blog/products/arcgis-pro/geoai/use-vision-language-models-to-optimize-object-classification
Use vision-language models to optimize object classification
Mar 10, 2026 - Rohit Singh demonstrates how to use vision-language models in scenarios where understanding both image and text content is crucial.
vision language modelsuseoptimizeobjectclassification
https://autopilot-cvpr.net/
AUTOPILOT Workshop @ CVPR 2026 | Autonomous Driving & Vision-Language Models
Join the 3rd AUTOPILOT Workshop at CVPR 2026 in Denver, Colorado. Exploring autonomous understanding through open-world perception and integrated language...
autonomous drivingvision languageautopilotworkshopcvpr
https://openreview.net/forum?id=iq0wUf6dUy
Reflective Planning: Vision-Language Models for Multi-Stage Long-Horizon Robotic Manipulation |...
Solving complex long-horizon robotic manipulation problems requires sophisticated high-level planning capabilities, the ability to reason about the physical...
vision language models
https://openreview.net/forum?id=friHAl5ofG
Vision Language Models are In-Context Value Learners | OpenReview
Predicting temporal progress from visual trajectories is important for intelligent robots that can learn, adapt, and improve. However, learning such progress...
vision language modelsin contextvaluelearnersopenreview
https://agentic-robot.github.io/
Agentic Robot: Vision-Language-Action Models in Embodied Agents
Agentic Robot: A Brain-Inspired Framework for Vision-Language-Action Models in Embodied Agents
robot visionagenticlanguageactionmodels
https://openreview.net/forum?id=XLiUcvHfzS
Post-hoc Probabilistic Vision-Language Models | OpenReview
Vision-language models (VLMs), such as CLIP and SigLIP, have found remarkable success in classification, retrieval, and generative tasks. For this, VLMs...
vision language modelspost hocprobabilisticopenreview
https://arxiv.org/abs/2505.05540
[2505.05540] Benchmarking Vision, Language, & Action Models in Procedurally Generated, Open Ended...
vision language
https://fastvlm.net/
FastVLM: Apple's Extremely Fast Vision Language Model
Runs directly on iPhone, first token output up to 85x faster!
vision languagefastvlmappleextremelymodel