Robuta

https://robotics-transformer2.github.io/ RT-2: Vision-Language-Action Models rt 2vision languageactionmodels https://robotics-transformer.github.io/ RT-2 Vision-Language-Action rt 2vision languageaction https://www.anjiecheng.me/SpatialRGPT SpatialRGPT: Grounded Spatial Reasoning in Vision-Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision-Language Models spatial reasoningvision languagegroundedmodels https://arxiv.org/html/2412.17417v2 Synth-Align: Improving Trustworthiness in Vision-Language Model with Synthetic Preference Data... vision language model https://huggingface.co/papers/2505.10610 Paper page - MMLongBench: Benchmarking Long-Context Vision-Language Models Effectively and... Join the discussion on this paper page vision language modelslong contextpaperbenchmarking https://openreview.net/forum?id=sb7qHFYwBc C-CLIP: Multimodal Continual Learning for Vision-Language Model | OpenReview Multimodal pre-trained models like CLIP need large image-text pairs for training but often struggle with domain-specific tasks. Since retraining with... vision language modelc clipcontinual learningmultimodalopenreview https://openreview.net/forum?id=xkgfLXZ4e0 Correlating instruction-tuning (in multimodal models) with vision-language processing (in the... Transformer-based language models, though not explicitly trained to mimic brain recordings, have demonstrated surprising alignment with brain activity.... instruction tuningmultimodal modelswith visionlanguage processing https://www.thoughtworks.com/en-gb/radar/techniques/vision-language-models-for-end-to-end-document-parsing Vision language models for end-to-end document parsing | Technology Radar | Thoughtworks United... Document parsing often relies on multi-stage pipelines combining layout detection, traditional OCR and post-processing scripts. These approaches often struggle... vision language models https://www.datacamp.com/it/blog/vlms-ai-vision-language-models Vision Language Models (VLMs) Explained | DataCamp Vision language models (VLMs) are AI models that can understand and process both visual and textual data, enabling tasks like image captioning, visual question... vision language modelsexplaineddatacamp https://openreview.net/forum?id=mmpwZfXUk7&referrer=%5Bthe%20profile%20of%20Vladan%20Stojni%C4%87%5D(%2Fprofile%3Fid%3D~Vladan_Stojni%C4%871) Label Propagation for Zero-shot Classification with Vision-Language Models | OpenReview Vision-Language Models (VLMs) have demonstrated impressive performance on zero-shot classification, i.e. classification when provided merely with a list of... vision language modelslabel propagationzero shot https://www.thoughtworks.com/en-br/radar/techniques/vision-language-models-for-end-to-end-document-parsing Vision language models for end-to-end document parsing | Technology Radar | Thoughtworks Brazil Document parsing often relies on multi-stage pipelines combining layout detection, traditional OCR and post-processing scripts. These approaches often struggle... vision language models https://www.thoughtworks.com/es-ec/radar/techniques/vision-language-models-for-end-to-end-document-parsing Vision language models for end-to-end document parsing | Technology Radar | Thoughtworks Ecuador Document parsing often relies on multi-stage pipelines combining layout detection, traditional OCR and post-processing scripts. These approaches often struggle... vision language models https://openreview.net/forum?id=oSJqRF0Tkg&referrer=%5Bthe%20profile%20of%20Hongming%20Zhang%5D(%2Fprofile%3Fid%3D~Hongming_Zhang2) Leopard: A Vision Language Model For Text-Rich Multi-Image Tasks | OpenReview Text-rich images, where text serves as the central visual element guiding the overall understanding, are prevalent in real-world applications, such as... vision language model https://www.amazon.science/tag/vision-language-models-vlms Vision-language models (VLMs) - Amazon Science vision language modelsamazonscience https://openreview.net/forum?id=Q5RYn6jagC Understanding the Limits of Vision Language Models Through the Lens of the Binding Problem |... Recent work has documented striking heterogeneity in the performance of state-of-the-art vision language models (VLMs), including both multimodal language... the limits of visionlanguage modelsunderstanding https://openreview.net/forum?id=eRleg6vy0Y&referrer=%5Bthe%20profile%20of%20Alyssa%20Unell%5D(%2Fprofile%3Fid%3D~Alyssa_Unell1) Micro-Bench: A Microscopy Benchmark for Vision-Language Understanding | OpenReview Recent advances in microscopy have enabled the rapid generation of terabytes of image data in cell biology and biomedical research. Vision-language models... vision languagemicrobenchunderstandingopenreview https://openreview.net/forum?id=JUwczEJY8I Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning | OpenReview Reinforcement learning (RL) requires either manually specifying a reward function, which is often infeasible, or learning a reward model from a large amount of... vision language modelszero shotreinforcement learning https://civitai.com/articles/27256/a-thousand-words-image-captioning-vision-language-model-interface A Thousand Words - Image Captioning (Vision Language Model) interface | Civitai I've spent a lot of time creating various "batch processing scripts" for various VLM's in the past ( Github repo search ). Instead, I decided to sp... a thousand wordsvision language modelimage captioninginterfacecivitai https://openreview.net/forum?id=zY37C8d6bS&referrer=%5Bthe%20profile%20of%20Ruifeng%20Chen%5D(%2Fprofile%3Fid%3D~Ruifeng_Chen1) Semantic Temporal Abstraction via Vision-Language Model Guidance for Efficient Reinforcement... Extracting temporally extended skills can significantly improve the efficiency of reinforcement learning (RL) by breaking down complex decision-making problems... vision language modeltemporal abstractionsemanticvia https://www.thoughtworks.com/en-cn/radar/techniques/vision-language-models-for-end-to-end-document-parsing Vision language models for end-to-end document parsing | Technology Radar | Thoughtworks China Document parsing often relies on multi-stage pipelines combining layout detection, traditional OCR and post-processing scripts. These approaches often struggle... vision language models https://arxiv.org/abs/2503.16538v1 [2503.16538v1] Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and... Abstract page for arXiv paper 2503.16538v1: Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking vision language models https://www.thoughtworks.com/es-cl/radar/techniques/vision-language-models-for-end-to-end-document-parsing Vision language models for end-to-end document parsing | Technology Radar | Thoughtworks Chile Document parsing often relies on multi-stage pipelines combining layout detection, traditional OCR and post-processing scripts. These approaches often struggle... vision language models https://openreview.net/forum?id=Kt2VJrCKo4 VoLTA: Vision-Language Transformer with Weakly-Supervised Local-Feature Alignment | OpenReview Vision-language pre-training (VLP) has recently proven highly effective for various uni- and multi-modal downstream applications. However, most existing... vision languagevoltatransformer https://www.autopilot-cvpr.net/ AUTOPILOT Workshop @ CVPR 2026 | Autonomous Driving & Vision-Language Models Join the 3rd AUTOPILOT Workshop at CVPR 2026 in Denver, Colorado. Exploring autonomous understanding through open-world perception and integrated language... autonomous drivingvision languageautopilotworkshopcvpr https://www.qualcomm.com/developer/software/qualcomm-interactive-video-dataset-qivd Interactive Video Dataset for Vision-Language Models | Qualcomm Improve your AI's visual understanding and response with QIVD, featuring 2,900 video files and 13 categories of action attributes and object detection. vision language modelsinteractive videodatasetqualcomm https://openreview.net/forum?id=tZozeR3VV7 Backdooring Vision-Language Models with Out-Of-Distribution Data | OpenReview The emergence of Vision-Language Models (VLMs) represents a significant advancement in integrating computer vision with Large Language Models (LLMs) to... vision language modelswith outdistribution dataopenreview https://vla-survey.github.io/ Vision-Language-Action Models for Robotics: A Review Towards Real-World Applications comprehensive survey of Vision-Language-Action (VLA) models for robotics. Explore the latest research on VLA architectures, learning paradigms, and real-world... vision language https://openreview.net/forum?id=ohJxgRLlLt Large (Vision) Language Models are Unsupervised In-Context Learners | OpenReview Recent advances in large language and vision-language models have enabled zero-shot inference, allowing models to solve new tasks without task-specific... vision language modelsin contextlargeunsupervisedlearners https://openreview.net/forum?id=hdYqGkSr9S&referrer=%5Bthe%20profile%20of%20Joseph%20Tighe%5D(%2Fprofile%3Fid%3D~Joseph_Tighe1) Benchmarking Zero-Shot Recognition with Vision-Language Models: Challenges on Granularity and... This paper introduces innovative benchmarks to evaluate Vision-Language Models (VLMs) in real-world zero-shot recognition tasks, focusing on the pivotal... vision language modelszero shot https://openreview.net/forum?id=trj2Jq8riA Interpretable Vision-Language Survival Analysis with Ordinal Inductive Bias for Computational... Histopathology Whole-Slide Images (WSIs) provide an important tool to assess cancer prognosis in computational pathology (CPATH). While existing survival... vision languagesurvival analysisinductive bias https://openreview.net/forum?id=7DpNGceUUV Coordinated Robustness Evaluation Framework for Vision Language Models | OpenReview Vision-language models, which integrate computer vision and natural language processing capabilities, have demonstrated significant advancements in tasks such... vision language modelsevaluation frameworkcoordinatedrobustnessopenreview https://openreview.net/forum?id=kcnIHUVeLMm&referrer=%5Bthe%20profile%20of%20Xingcheng%20Zhou%5D(%2Fprofile%3Fid%3D~Xingcheng_Zhou1) Vision Language Models in Autonomous Driving and Intelligent Transportation Systems | OpenReview The applications of Vision-Language Models (VLMs) in the fields of Autonomous Driving (AD) and Intelligent Transportation Systems (ITS) have attracted... vision language modelsautonomous drivingintelligent transportation https://arxiv.org/abs/2605.00583 [2605.00583] Jailbreaking Vision-Language Models Through the Visual Modality Abstract page for arXiv paper 2605.00583: Jailbreaking Vision-Language Models Through the Visual Modality vision language modelsthe visualjailbreakingmodality https://www.utwente.nl/en/eemcs/dmb/assignments/open/master/Computer%20Vision%20and%20Biometrics/20260203-Hyperbolic%20Vision%20Language%20Models%20(VLMs)%20in%20medical%20imaging/ Hyperbolic Vision Language Models (VLMs) in medical imaging | Computer Vision and Biometrics |... vision language modelsmedical imaginghyperbolic https://arxiv.org/abs/2508.19652 [2508.19652] Self-Rewarding Vision-Language Model via Reasoning Decomposition Abstract page for arXiv paper 2508.19652: Self-Rewarding Vision-Language Model via Reasoning Decomposition vision language modelselfrewardingviareasoning https://openreview.net/forum?id=3uVAA3ckxT ADAPT: Adaptive Prompt Tuning for Vision-Language Models | OpenReview Prompt tuning has emerged as an effective way for parameter-efficient fine-tuning. Conventional deep prompt tuning inserts continuous prompts of a fixed... vision language modelsprompt tuningadaptopenreview https://www.thoughtworks.com/en-ca/radar/techniques/vision-language-models-for-end-to-end-document-parsing Vision language models for end-to-end document parsing | Technology Radar | Thoughtworks Canada Document parsing often relies on multi-stage pipelines combining layout detection, traditional OCR and post-processing scripts. These approaches often struggle... vision language models https://arxiv.org/abs/2503.16538v2 [2503.16538v2] Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and... Abstract page for arXiv paper 2503.16538v2: Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking vision language models https://www.wur.nl/en/vacancy/phd-position-vision-language-models-biodiversity-environmental-monitoring PhD position: Vision-Language Models for Biodiversity & Environmental Monitoring | WUR We are looking for a PhD to develop vision-language models for applications in biodiversity and environmental monitoring, using Earth observation, drone, and... vision language modelsphd positionenvironmental monitoringbiodiversitywur https://aclanthology.org/2024.lrec-main.973/ Medical Vision-Language Pre-Training for Brain Abnormalities - ACL Anthology Masoud Monajatipoor, Zi-Yi Dou, Aichi Chien, Nanyun Peng, Kai-Wei Chang. Proceedings of the 2024 Joint International Conference on Computational Linguistics,... medical visionpre trainingfor brainlanguageabnormalities https://openreview.net/forum?id=U7aGAc0S5h&referrer=%5Bthe%20profile%20of%20Zhenlin%20Xu%5D(%2Fprofile%3Fid%3D~Zhenlin_Xu1) Benchmarking Zero-Shot Recognition with Vision-Language Models: Challenges on Granularity and... This paper presents novel benchmarks for evaluating vision-language models (VLMs) in zero-shot recognition, focusing on granularity and specificity. Although... vision language modelszero shot https://huggingface.co/papers/2305.10355 Paper page - Evaluating Object Hallucination in Large Vision-Language Models Join the discussion on this paper page large visionpaperevaluatingobjecthallucination https://openreview.net/forum?id=ELR4ifkX3j Efficient Open-set Test Time Adaptation of Vision Language Models | OpenReview In dynamic real-world settings, models must adapt to changing data distributions, a challenge known as Test Time Adaptation (TTA). This becomes even more... vision language modelsopen setefficienttesttime https://www.amazon.science/publications/prompting-vision-language-models-for-aspect-controlled-generation-of-referring-expressions Prompting vision-language models for aspect-controlled generation of referring expressions - Amazon... Referring Expression Generation (REG) is the task of generating a description that unambiguously identifies a given target in the scene. Different from Image... vision language models https://www.datacamp.com/pl/blog/vlms-ai-vision-language-models Vision Language Models (VLMs) Explained | DataCamp Vision language models (VLMs) are AI models that can understand and process both visual and textual data, enabling tasks like image captioning, visual question... vision language modelsexplaineddatacamp https://arxiv.org/abs/2507.18053 [2507.18053] Resource Consumption Red-Teaming for Large Vision-Language Models Abstract page for arXiv paper 2507.18053: Resource Consumption Red-Teaming for Large Vision-Language Models resource consumptionred teaminglarge vision https://openreview.net/forum?id=d5DJWgmMoX Can Vision Language Models Learn from Visual Demonstrations of Ambiguous Spatial Reasoning? |... Large vision-language models (VLMs) have become state-of-the-art for many computer vision tasks, with in-context learning (ICL) as a popular adaptation... vision language models https://vlm-jailbreaks.github.io/ Jailbreaking Vision-Language Models Through the Visual Modality | ICML 2026 Four visual jailbreak attacks exploiting the vision component of VLMs. ICML 2026. vision language modelsthe visualjailbreakingmodalityicml https://openreview.net/forum?id=N0I2RtD8je Vision-Language Models are Zero-Shot Reward Models for Reinforcement Learning | OpenReview Reinforcement learning (RL) requires either manually specifying a reward function, which is often infeasible, or learning a reward model from a large amount of... vision language modelszero shotreinforcement learning https://oecd.ai/en/catalogue/metric-use-cases/pali-3-vision-language-models-smaller,-faster,-stronger PaLI-3 Vision Language Models: Smaller, Faster, Stronger - OECD.AI The interest of the machine learning community in image synthesis has grown significantly in recent years, with the introduction of a wide range of deep... vision language modelspali3smallerfaster https://www.findbestmodel.app/ ModelMatch - Compare Vision-Language Models Compare top open source vision-language models side-by-side, no coding needed vision languagecomparemodels https://arxiv.org/abs/2210.05335 [2210.05335] MAP: Multimodal Uncertainty-Aware Vision-Language Pre-training Model Abstract page for arXiv paper 2210.05335: MAP: Multimodal Uncertainty-Aware Vision-Language Pre-training Model vision languagepre training2210mapmultimodal https://arxiv.org/abs/2405.10529?context=cs.CV [2405.10529] Safeguarding Vision-Language Models Against Patched Visual Prompt Injectors Abstract page for arXiv paper 2405.10529: Safeguarding Vision-Language Models Against Patched Visual Prompt Injectors vision language models2405safeguarding https://www.thoughtworks.com/radar/techniques/vision-language-models-for-end-to-end-document-parsing Vision language models for end-to-end document parsing | Technology Radar | Thoughtworks Document parsing often relies on multi-stage pipelines combining layout detection, traditional OCR and post-processing scripts. These approaches often struggle... vision language modelsdocument parsingtechnology radarend https://openreview.net/forum?id=iwp3H8uSeK Distilling Out-of-Distribution Robustness from Vision-Language Foundation Models | OpenReview We propose a conceptually simple and lightweight framework for improving the robustness of vision models through the combination of knowledge distillation and... out offrom visionfoundation modelsdistillingdistribution https://openreview.net/forum?id=0y3hGn1wOk Benchmarking Vision Language Model Unlearning via Fictitious Facial Identity Dataset | OpenReview Machine unlearning has emerged as an effective strategy for forgetting specific information in the training data. However, with the increasing integration of... vision language modelbenchmarkingunlearning https://www.frontiersin.org/research-topics/18532/identifying-analyzing-and-overcoming-challenges-in-vision-and-language-research Frontiers | Identifying, Analyzing, and Overcoming Challenges in Vision and Language Research Integration of computer vision and natural language processing is one of the holy grails in machine learning and artificial intelligence research with import... overcoming challengesvision languagefrontiersidentifyinganalyzing https://arxiv.org/abs/2405.19315 [2405.19315] Matryoshka Query Transformer for Large Vision-Language Models Abstract page for arXiv paper 2405.19315: Matryoshka Query Transformer for Large Vision-Language Models large vision2405matryoshkaquerytransformer https://openreview.net/forum?id=gbc0m68fdw Manipulate-Anything: Automating Real-World Robots using Vision-Language Models | OpenReview Large-scale endeavors like RT-1 and widespread community efforts such as Open-X-Embodiment have contributed to growing the scale of robot demonstration data.... vision language modelsreal worldmanipulateanythingautomating https://openreview.net/forum?id=C3PsmuhDq0&referrer=%5Bthe%20profile%20of%20Shayekh%20Bin%20Islam%5D(%2Fprofile%3Fid%3D~Shayekh_Bin_Islam1) Behind Maya: Building a Multilingual Vision Language Model | OpenReview In recent times, we have seen a rapid development of large Vision-Language Models (VLMs). They have shown impressive results on academic benchmarks, primarily... vision language modelbehindmayabuildingmultilingual https://openreview.net/forum?id=NEBa0bs5LR&referrer=%5Bthe%20profile%20of%20Yunhang%20Shen%5D(%2Fprofile%3Fid%3D~Yunhang_Shen1) DS-VLM: Diffusion Supervision Vision Language Model | OpenReview Vision-Language Models (VLMs) face two critical limitations in visual representation learning: degraded supervision due to information loss during gradient... vision language modeldsvlmdiffusionsupervision https://huggingface.co/papers/2603.03739 Paper page - PROSPECT: Unified Streaming Vision-Language Navigation via Semantic--Spatial Fusion... Join the discussion on this paper page vision language https://openreview.net/forum?id=pQDzXVeBja Effortless Vision-Language Model Specialization in Histopathology without Annotation | OpenReview Recent advances in Vision-Language Models (VLMs) in histopathology, such as CONCH and QuiltNet, have demonstrated impressive zero-shot classification... vision language modeleffortlessspecializationhistopathologywithout https://www.thoughtworks.com/en-au/radar/techniques/vision-language-models-for-end-to-end-document-parsing Vision language models for end-to-end document parsing | Technology Radar | Thoughtworks Australia Document parsing often relies on multi-stage pipelines combining layout detection, traditional OCR and post-processing scripts. These approaches often struggle... vision language models https://openreview.net/forum?id=6zoKzrRn2h&referrer=%5Bthe%20profile%20of%20Zhihe%20Lu%5D(%2Fprofile%3Fid%3D~Zhihe_Lu1) Beyond Sole Strength: Customized Ensembles for Generalized Vision-Language Models | OpenReview vision language modelsbeyondsolestrengthcustomized https://openreview.net/forum?id=WutpswD3ea Assisted Few-Shot Learning for Vision-Language Models in Agricultural Stress Phenotype... In the agricultural sector, labeled data for crop diseases and stresses are often scarce due to high annotation costs. We propose an Assisted Few-Shot Learning... few shot learningvision language models https://arxiv.org/abs/2505.13302?ref=disinfodocket.com [2505.13302] Images Amplify Misinformation Sharing in Vision-Language Models Abstract page for arXiv paper 2505.13302: Images Amplify Misinformation Sharing in Vision-Language Models in visionimagesamplifymisinformationsharing https://www.thoughtworks.com/en-ec/radar/techniques/vision-language-models-for-end-to-end-document-parsing Vision language models for end-to-end document parsing | Technology Radar | Thoughtworks Ecuador Document parsing often relies on multi-stage pipelines combining layout detection, traditional OCR and post-processing scripts. These approaches often struggle... vision language models https://openreview.net/forum?id=hssNWbMZHF&referrer=%5Bthe%20profile%20of%20Marzyeh%20Ghassemi%5D(%2Fprofile%3Fid%3D~Marzyeh_Ghassemi2) Vision-Language Models Do Not Understand Negation | OpenReview Many practical vision-language applications require models that understand negation, e.g., when using natural language to retrieve images which contain certain... vision language modelsdo notunderstandnegationopenreview https://openreview.net/forum?id=vQDKYYuqWA Vision-Language Models Provide Promptable Representations for Reinforcement Learning | OpenReview Humans can quickly learn new behaviors by leveraging background world knowledge. In contrast, agents trained with reinforcement learning (RL) typically learn... vision language modelsreinforcement learningproviderepresentationsopenreview https://arxiv.org/abs/2508.08604v1 [2508.08604v1] Transferable Model-agnostic Vision-Language Model Adaptation for Efficient... Abstract page for arXiv paper 2508.08604v1: Transferable Model-agnostic Vision-Language Model Adaptation for Efficient Weak-to-Strong Generalization vision languagetransferablemodelagnosticadaptation https://vislang.ai/ VisLang - Vision, Language and Learning Lab at Rice University Research group at Rice University led by Vicente Ordonez, working at the intersection of computer vision, natural language processing, and machine learning. vision languagelearning labriceuniversity https://openreview.net/forum?id=XgYZT35N76&referrer=%5Bthe%20profile%20of%20Yiming%20Yang%5D(%2Fprofile%3Fid%3D~Yiming_Yang1) Improve Vision Language Model Chain-of-thought Reasoning | OpenReview Chain-of-thought (CoT) reasoning in vision language models (VLMs) is crucial for improving interpretability and trustworthiness. However, current training... chain of thought reasoningvision language modelimproveopenreview https://www.thoughtworks.com/es-es/radar/techniques/vision-language-models-for-end-to-end-document-parsing Vision language models for end-to-end document parsing | Technology Radar | Thoughtworks Spain Document parsing often relies on multi-stage pipelines combining layout detection, traditional OCR and post-processing scripts. These approaches often struggle... vision language models https://huggingface.co/papers/2210.09996 Paper page - Perceptual Grouping in Contrastive Vision-Language Models Join the discussion on this paper page vision languagepaperperceptualgroupingcontrastive https://www.analyticsvidhya.com/blog/2024/03/introducing-moondream2-a-tiny-vision-language-model/ Introducing Moondream2: A tiny Vision-Language Model Mar 30, 2024 - Explore Moondream2, a small vision-language model for efficient image-text tasks. Components, implementation, applications, and limitations. vision languageintroducingtinymodel https://arxiv.org/abs/2604.02318v1 [2604.02318v1] Stop Wandering: Efficient Vision-Language Navigation via Metacognitive Reasoning Abstract page for arXiv paper 2604.02318v1: Stop Wandering: Efficient Vision-Language Navigation via Metacognitive Reasoning vision languagestopwanderingefficient https://openreview.net/forum?id=iVxxgZlXh6 LLaRA: Supercharging Robot Learning Data for Vision-Language Policy | OpenReview Vision Language Models (VLMs) have recently been leveraged to generate robotic actions, forming Vision-Language-Action (VLA) models. However, directly adapting... robot learningdata forvision languagesuperchargingpolicy https://arxiv.org/html/2603.02972v1 TagaVLM: Topology-Aware Global Action Reasoning for Vision-Language Navigation global actionvision languagetopologyawarereasoning https://speakerdeck.com/keio_smilab/the-confluence-of-vision-language-and-robotics The Confluence of Vision, Language, and Robotics - Speaker Deck the confluencevision languageroboticsspeakerdeck https://openreview.net/forum?id=vvoWPYqZJA InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning | OpenReview Large-scale pre-training and instruction tuning have been successful at creating general-purpose language models with broad competence. However, building... vision language modelsgeneral purposeinstruction tuningtowards https://huggingface.co/papers/2311.01477 Paper page - FAITHSCORE: Evaluating Hallucinations in Large Vision-Language Models Join the discussion on this paper page large visionpaperevaluatinghallucinationslanguage https://www.datacamp.com/blog/top-vision-language-models Top 10 Vision Language Models in 2026 | DataCamp Discover the top open-source and proprietary vision-language models of 2026 for visual reasoning, image analysis, and computer vision. vision language modelstop 102026datacamp https://deepai.org/publication/illume-rationalizing-vision-language-models-by-interacting-with-their-jabber ILLUME: Rationalizing Vision-Language Models by Interacting with their Jabber | DeepAI Aug 17, 2022 - 08/17/22 - Bootstrapping from pre-trained language models has been proven to be an efficient approach for building foundation vision-language... vision language modelsillumerationalizing https://huggingface.co/papers/2401.12168 Paper page - SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities Join the discussion on this paper page vision language modelsspatial reasoningpaper https://openreview.net/forum?id=SB2sWF3oCw Understanding and Improving In-Context Learning on Vision-language Models | OpenReview In-context learning (ICL) on large language models (LLMs) has received great attention, and this technique can also be applied to vision-language models (VLMs)... in context learningvision language modelsunderstandingimproving https://arxiv.org/html/2407.13588v1 Robust Calibration of Large Vision-Language Adapters large visionrobustcalibrationlanguageadapters https://arxiv.org/abs/2407.01449?ref=blog.mixpeek.com [2407.01449] ColPali: Efficient Document Retrieval with Vision Language Models Abstract page for arXiv paper 2407.01449: ColPali: Efficient Document Retrieval with Vision Language Models document retrievalwith vision2407colpaliefficient https://arxiv.org/abs/2306.11300 [2306.11300] RS5M and GeoRSCLIP: A Large Scale Vision-Language Dataset and A Large Vision-Language... Abstract page for arXiv paper 2306.11300: RS5M and GeoRSCLIP: A Large Scale Vision-Language Dataset and A Large Vision-Language Model for Remote Sensing a largevision language2306 https://www.thoughtworks.com/en-sg/radar/techniques/vision-language-models-for-end-to-end-document-parsing Vision language models for end-to-end document parsing | Technology Radar | Thoughtworks Singapore Document parsing often relies on multi-stage pipelines combining layout detection, traditional OCR and post-processing scripts. These approaches often struggle... vision language models https://openreview.net/forum?id=R2rasAEPVi Leopard: A Vision Language Model for Text-Rich Multi- Image Tasks | OpenReview Text-rich images, where text serves as the central visual element guiding the overall understanding, are prevalent in real-world applications, such as... vision language model https://openreview.net/forum?id=haJHr4UsQX Causal Graphical Models for Vision-Language Compositional Understanding | OpenReview Recent work has empirically shown that Vision-Language Models (VLMs) struggle to fully understand the compositional properties of the human language, usually... graphical modelsvision languagecausalcompositionalunderstanding https://www.esri.com/arcgis-blog/products/arcgis-pro/geoai/use-vision-language-models-to-optimize-object-classification Use vision-language models to optimize object classification Mar 10, 2026 - Rohit Singh demonstrates how to use vision-language models in scenarios where understanding both image and text content is crucial. vision language modelsuseoptimizeobjectclassification https://autopilot-cvpr.net/ AUTOPILOT Workshop @ CVPR 2026 | Autonomous Driving & Vision-Language Models Join the 3rd AUTOPILOT Workshop at CVPR 2026 in Denver, Colorado. Exploring autonomous understanding through open-world perception and integrated language... autonomous drivingvision languageautopilotworkshopcvpr https://openreview.net/forum?id=iq0wUf6dUy Reflective Planning: Vision-Language Models for Multi-Stage Long-Horizon Robotic Manipulation |... Solving complex long-horizon robotic manipulation problems requires sophisticated high-level planning capabilities, the ability to reason about the physical... vision language models https://openreview.net/forum?id=friHAl5ofG Vision Language Models are In-Context Value Learners | OpenReview Predicting temporal progress from visual trajectories is important for intelligent robots that can learn, adapt, and improve. However, learning such progress... vision language modelsin contextvaluelearnersopenreview https://agentic-robot.github.io/ Agentic Robot: Vision-Language-Action Models in Embodied Agents Agentic Robot: A Brain-Inspired Framework for Vision-Language-Action Models in Embodied Agents robot visionagenticlanguageactionmodels https://openreview.net/forum?id=XLiUcvHfzS Post-hoc Probabilistic Vision-Language Models | OpenReview Vision-language models (VLMs), such as CLIP and SigLIP, have found remarkable success in classification, retrieval, and generative tasks. For this, VLMs... vision language modelspost hocprobabilisticopenreview https://arxiv.org/abs/2505.05540 [2505.05540] Benchmarking Vision, Language, & Action Models in Procedurally Generated, Open Ended... vision language https://fastvlm.net/ FastVLM: Apple's Extremely Fast Vision Language Model Runs directly on iPhone, first token output up to 85x faster! vision languagefastvlmappleextremelymodel