https://ivi.fnwi.uva.nl/vislab/publication/yingjun-iclr-2026/
Prompt-Robust Vision-Language Models via Meta-Finetuning | VIS Lab
Apr 1, 2026 - ision-language models (VLMs) have demonstrated remarkable generalization across diverse tasks by leveraging large-scale image-text pretraining. However, their...
vision language modelspromptrobustviameta
https://www.utwente.nl/en/eemcs/ps/education/master%20theses/Adarsh-2/
Assignments: DRONE-BASED OBJECT DETECTION AND EXPLANATION USING VISION-LANGUAGE MODELS | Pervasive...
vision language modelsobject detection
https://research.nvidia.com/publication/2026-04_qcaleval-benchmarking-vision-language-models-quantum-calibration-plot
QCalEval: Benchmarking Vision-Language Models for Quantum Calibration Plot Understanding | Research
Quantum computing calibration depends on interpreting experimental data, and calibration plots provide the most universal human-readable representation for...
vision language modelsbenchmarking
https://arxiv.org/abs/2503.16538v1
[2503.16538v1] Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and...
Abstract page for arXiv paper 2503.16538v1: Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking
vision language models
https://developers.redhat.com/articles/2025/10/27/multimodal-ai-edge-deploy-vision-language-models-ramalama
Deploy vision language models with RamaLama | Red Hat Developer
Oct 27, 2025 - Learn how to deploy multimodal AI models on edge devices using the RamaLama CLI, from pulling your first vision language model (VLM) to serving it via an API.
vision language modelsred hatdeployramalamadeveloper
https://boris-portal.unibe.ch/entities/publication/6c36c151-d31b-44cf-878c-f80995408656
Benchmarking Zero-Shot Recognition with Vision-Language Models: Challenges on Granularity and...
vision language modelszero shot
https://seantrott.substack.com/p/vision-language-models-vlms-explained-24d
Vision-language models (VLMs), explained (pt. 2)
The view from Cognitive Science.
vision language modelsexplainedpt
https://arxiv.org/abs/2503.11609
[2503.11609] Rethinking Few-Shot Adaptation of Vision-Language Models in Two Stages
Abstract page for arXiv paper 2503.11609: Rethinking Few-Shot Adaptation of Vision-Language Models in Two Stages
vision language models
https://au.mathworks.com/help/vision/vision-language-models.html
Vision-Language Models - MATLAB & Simulink
Perform image classification, retrieval, captioning, and object detection tasks using vision-language models
vision language modelsmatlabsimulink
https://docs.nvidia.com/nim/vision-language-models/2.0.0/search.html
Search - NVIDIA NIM for Vision Language Models (VLMs)
vision language modelsnvidia nimsearch
https://research.nvidia.com/index.php/publication/2026-04_qcaleval-benchmarking-vision-language-models-quantum-calibration-plot
QCalEval: Benchmarking Vision-Language Models for Quantum Calibration Plot Understanding | Research
Quantum computing calibration depends on interpreting experimental data, and calibration plots provide the most universal human-readable representation for...
vision language modelsbenchmarking
https://www.n-ix.com/vision-language-models/
Vision language models: How they work and where to use them - N-iX
Jun 18, 2026 - Vision language models explained: how they work, real-world use cases, deployment tradeoffs, and the challenges to avoid based on N-iX engineering experience.
vision language modelshow they workwhere to use
https://hainet.substack.com/p/understanding-alignment-in-vision-language-model/comments
Comments - Understanding Alignment in Vision Language Models (VLMs)
At the core of VLMs lies the fundamental challenge of representing image feature vectors as LLM textual vectors while preserving the information. This blog...
vision language modelscommentsunderstandingalignment
https://www.nvidia.com/en-us/glossary/vision-language-models/
What are Vision-Language Models? | NVIDIA Glossary
Vision Language Models (VLMs) are multimodal generative AI models capable of reasoning over text, image and video prompts.
vision language modelsnvidiaglossary
https://tbd.ctr.utexas.edu/research-product/leveraging-vision-language-models-for-efficient-understanding-of-vulnerable-roadway-users-via-a-multimodal-traffic-sensing-approach/
Leveraging Vision-Language Models for Efficient Understanding of Vulnerable Roadway Users via a...
vision language models
https://arxiv.org/html/2603.08921v1
Vision-Language Models Encode Clinical Guidelines for Concept-Based Medical Reasoning
vision language modelsclinical guidelinesencode
https://gamma.umd.edu/researchdirections/crowdmultiagent/tgs/
VL-TGS: Trajectory Generation and Selection using Vision Language Models in Mapless Outdoor...
Abstract We present a multi-modal trajectory generation and selection algorithm for real-world mapless outdoor nav- igation in human-centered environments....
vision language models
https://experts.arizona.edu/en/publications/visually-grounded-planning-without-vision-language-models-infer-d-2/
Visually-grounded planning without vision: Language models infer detailed plans from high-level...
vision language models
https://researchconnect.buffalo.edu/en/publications/text-image-de-contextualization-detection-using-vision-language-m/
TEXT-IMAGE DE-CONTEXTUALIZATION DETECTION USING VISION-LANGUAGE MODELS - SUNY University at Buffalo
vision language models
https://research.google/pubs/benchmarking-vision-language-models-for-cultural-understanding/
Benchmarking Vision Language Models for Cultural Understanding
vision language modelsbenchmarkingculturalunderstanding
https://docs.nvidia.com/nim/vision-language-models/1.7.0/search.html
Search - NVIDIA NIM for Vision Language Models (VLMs)
vision language modelsnvidia nimsearch
https://docs.nvidia.com/nim/vision-language-models/1.5.0/search.html
Search - NVIDIA NIM for Vision Language Models (VLMs)
vision language modelsnvidia nimsearch
https://arxiv.org/html/2509.09731v1
Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning
vision language models
https://vlmgineer.github.io/release
VLMgineer: Vision Language Models as Robotic Toolsmiths
Vision Language Models as Robotic Toolsmiths
vision language modelsrobotic
https://docs.google.com/presentation/d/1zdfvmraJ_21zjJSpE7DMuRG54Bkze7hfKsAUnDG5u4Y/edit?usp=sharing
Fine-Grained Evaluation of Vision-Language Models through "Vision-Language Model as a Judge" -...
1 2024.02.16 Fine-Grained Evaluation of Vision-Language Models through "Vision-Language Model as a Judge"
vision language models
https://huggingface.co/papers/2401.12168
Paper page - SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities
Join the discussion on this paper page
vision language modelspaper pagespatial reasoning
https://arxiv.org/abs/2403.18715
[2403.18715] Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive...
Abstract page for arXiv paper 2403.18715: Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding
vision language models
https://arxiv.org/abs/2410.20971v1
[2410.20971v1] BlueSuffix: Reinforced Blue Teaming for Vision-Language Models Against Jailbreak...
Abstract page for arXiv paper 2410.20971v1: BlueSuffix: Reinforced Blue Teaming for Vision-Language Models Against Jailbreak Attacks
vision language modelsblue teaming
https://usbios.ai/
BiOS — Train it. Run it. Own it. | Fine-tune 250k+ open LLMs & vision-language models
Fine-tune 250k+ open models with 15+ methods including LoRA, QLoRA, and full fine-tune. 6 alignment algorithms (DPO, SimPO, ORPO, KTO), continued pre-training,...
https://www.anjiecheng.me/SpatialRGPT
SpatialRGPT: Grounded Spatial Reasoning in Vision-Language Models
SpatialRGPT: Grounded Spatial Reasoning in Vision-Language Models
spatial reasoninggroundedvisionlanguagemodels
https://robotics-transformer2.github.io/
RT-2: Vision-Language-Action Models
rtvisionlanguageactionmodels
https://arxiv.org/abs/2505.21061
[2505.21061] LPOI: Listwise Preference Optimization for Vision Language Models
Abstract page for arXiv paper 2505.21061: LPOI: Listwise Preference Optimization for Vision Language Models
preference optimizationvisionlanguagemodels
https://research.facebook.com/publications/taking-a-hint-leveraging-explanations-to-make-vision-and-language-models-more-grounded/
Taking a HINT: Leveraging Explanations to Make Vision and Language Models More Grounded - Meta...
In this work, we propose a generic approach called Human Importance-aware Network Tuning (HINT) that effectively leverages human demonstrations to improve...
https://publica.fraunhofer.de/entities/publication/b07a7aca-f1a7-42b5-9d7c-a1fba76a3663
Zero-Shot Open-Vocabulary OOD Object Detection and Grounding using Vision Language Models
Automated driving involves complex perception tasks that require a precise understanding of diverse traffic scenarios and confident navigation. Traditional...
https://arxiv.org/abs/2503.21817
[2503.21817] Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models via...
Abstract page for arXiv paper 2503.21817: Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models via Adaptive Token Skipping
https://arxiv.org/abs/2407.12366
[2407.12366] NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
Abstract page for arXiv paper 2407.12366: NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
https://is.mpg.de/en/publications/yeliuhe25
VERA: Explainable Video Anomaly Detection via Verbalized Learning of Vision-Language Models | MPI-IS
Our goal is to understand the principles of Perception, Action and Learning in autonomous systems that successfully interact with complex environments and to...
https://publications.ri.cmu.edu/toward-more-reliable-multimodal-systems-mitigating-hallucinations-in-large-vision-language-models
Toward More Reliable Multimodal Systems: Mitigating Hallucinations in Large Vision-Language Models...
more reliable
https://arxiv.org/abs/2508.01943
[2508.01943] ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied Tasks
Abstract page for arXiv paper 2508.01943: ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied Tasks