Robuta

https://ivi.fnwi.uva.nl/vislab/publication/yingjun-iclr-2026/ Prompt-Robust Vision-Language Models via Meta-Finetuning | VIS Lab Apr 1, 2026 - ision-language models (VLMs) have demonstrated remarkable generalization across diverse tasks by leveraging large-scale image-text pretraining. However, their... vision language modelspromptrobustviameta https://www.utwente.nl/en/eemcs/ps/education/master%20theses/Adarsh-2/ Assignments: DRONE-BASED OBJECT DETECTION AND EXPLANATION USING VISION-LANGUAGE MODELS | Pervasive... vision language modelsobject detection https://research.nvidia.com/publication/2026-04_qcaleval-benchmarking-vision-language-models-quantum-calibration-plot QCalEval: Benchmarking Vision-Language Models for Quantum Calibration Plot Understanding | Research Quantum computing calibration depends on interpreting experimental data, and calibration plots provide the most universal human-readable representation for... vision language modelsbenchmarking https://arxiv.org/abs/2503.16538v1 [2503.16538v1] Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and... Abstract page for arXiv paper 2503.16538v1: Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking vision language models https://developers.redhat.com/articles/2025/10/27/multimodal-ai-edge-deploy-vision-language-models-ramalama Deploy vision language models with RamaLama | Red Hat Developer Oct 27, 2025 - Learn how to deploy multimodal AI models on edge devices using the RamaLama CLI, from pulling your first vision language model (VLM) to serving it via an API. vision language modelsred hatdeployramalamadeveloper https://boris-portal.unibe.ch/entities/publication/6c36c151-d31b-44cf-878c-f80995408656 Benchmarking Zero-Shot Recognition with Vision-Language Models: Challenges on Granularity and... vision language modelszero shot https://seantrott.substack.com/p/vision-language-models-vlms-explained-24d Vision-language models (VLMs), explained (pt. 2) The view from Cognitive Science. vision language modelsexplainedpt https://arxiv.org/abs/2503.11609 [2503.11609] Rethinking Few-Shot Adaptation of Vision-Language Models in Two Stages Abstract page for arXiv paper 2503.11609: Rethinking Few-Shot Adaptation of Vision-Language Models in Two Stages vision language models https://au.mathworks.com/help/vision/vision-language-models.html Vision-Language Models - MATLAB & Simulink Perform image classification, retrieval, captioning, and object detection tasks using vision-language models vision language modelsmatlabsimulink https://docs.nvidia.com/nim/vision-language-models/2.0.0/search.html Search - NVIDIA NIM for Vision Language Models (VLMs) vision language modelsnvidia nimsearch https://research.nvidia.com/index.php/publication/2026-04_qcaleval-benchmarking-vision-language-models-quantum-calibration-plot QCalEval: Benchmarking Vision-Language Models for Quantum Calibration Plot Understanding | Research Quantum computing calibration depends on interpreting experimental data, and calibration plots provide the most universal human-readable representation for... vision language modelsbenchmarking https://www.n-ix.com/vision-language-models/ Vision language models: How they work and where to use them - N-iX Jun 18, 2026 - Vision language models explained: how they work, real-world use cases, deployment tradeoffs, and the challenges to avoid based on N-iX engineering experience. vision language modelshow they workwhere to use https://hainet.substack.com/p/understanding-alignment-in-vision-language-model/comments Comments - Understanding Alignment in Vision Language Models (VLMs) At the core of VLMs lies the fundamental challenge of representing image feature vectors as LLM textual vectors while preserving the information. This blog... vision language modelscommentsunderstandingalignment https://www.nvidia.com/en-us/glossary/vision-language-models/ What are Vision-Language Models? | NVIDIA Glossary Vision Language Models (VLMs) are multimodal generative AI models capable of reasoning over text, image and video prompts. vision language modelsnvidiaglossary https://tbd.ctr.utexas.edu/research-product/leveraging-vision-language-models-for-efficient-understanding-of-vulnerable-roadway-users-via-a-multimodal-traffic-sensing-approach/ Leveraging Vision-Language Models for Efficient Understanding of Vulnerable Roadway Users via a... vision language models https://arxiv.org/html/2603.08921v1 Vision-Language Models Encode Clinical Guidelines for Concept-Based Medical Reasoning vision language modelsclinical guidelinesencode https://gamma.umd.edu/researchdirections/crowdmultiagent/tgs/ VL-TGS: Trajectory Generation and Selection using Vision Language Models in Mapless Outdoor... Abstract We present a multi-modal trajectory generation and selection algorithm for real-world mapless outdoor nav- igation in human-centered environments.... vision language models https://experts.arizona.edu/en/publications/visually-grounded-planning-without-vision-language-models-infer-d-2/ Visually-grounded planning without vision: Language models infer detailed plans from high-level... vision language models https://researchconnect.buffalo.edu/en/publications/text-image-de-contextualization-detection-using-vision-language-m/ TEXT-IMAGE DE-CONTEXTUALIZATION DETECTION USING VISION-LANGUAGE MODELS - SUNY University at Buffalo vision language models https://research.google/pubs/benchmarking-vision-language-models-for-cultural-understanding/ Benchmarking Vision Language Models for Cultural Understanding vision language modelsbenchmarkingculturalunderstanding https://docs.nvidia.com/nim/vision-language-models/1.7.0/search.html Search - NVIDIA NIM for Vision Language Models (VLMs) vision language modelsnvidia nimsearch https://docs.nvidia.com/nim/vision-language-models/1.5.0/search.html Search - NVIDIA NIM for Vision Language Models (VLMs) vision language modelsnvidia nimsearch https://arxiv.org/html/2509.09731v1 Benchmarking Vision-Language Models on Chinese Ancient Documents: From OCR to Knowledge Reasoning vision language models https://vlmgineer.github.io/release VLMgineer: Vision Language Models as Robotic Toolsmiths Vision Language Models as Robotic Toolsmiths vision language modelsrobotic https://docs.google.com/presentation/d/1zdfvmraJ_21zjJSpE7DMuRG54Bkze7hfKsAUnDG5u4Y/edit?usp=sharing Fine-Grained Evaluation of Vision-Language Models through "Vision-Language Model as a Judge" -... 1 2024.02.16 Fine-Grained Evaluation of Vision-Language Models through "Vision-Language Model as a Judge" vision language models https://huggingface.co/papers/2401.12168 Paper page - SpatialVLM: Endowing Vision-Language Models with Spatial Reasoning Capabilities Join the discussion on this paper page vision language modelspaper pagespatial reasoning https://arxiv.org/abs/2403.18715 [2403.18715] Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive... Abstract page for arXiv paper 2403.18715: Mitigating Hallucinations in Large Vision-Language Models with Instruction Contrastive Decoding vision language models https://arxiv.org/abs/2410.20971v1 [2410.20971v1] BlueSuffix: Reinforced Blue Teaming for Vision-Language Models Against Jailbreak... Abstract page for arXiv paper 2410.20971v1: BlueSuffix: Reinforced Blue Teaming for Vision-Language Models Against Jailbreak Attacks vision language modelsblue teaming https://usbios.ai/ BiOS — Train it. Run it. Own it. | Fine-tune 250k+ open LLMs & vision-language models Fine-tune 250k+ open models with 15+ methods including LoRA, QLoRA, and full fine-tune. 6 alignment algorithms (DPO, SimPO, ORPO, KTO), continued pre-training,... https://www.anjiecheng.me/SpatialRGPT SpatialRGPT: Grounded Spatial Reasoning in Vision-Language Models SpatialRGPT: Grounded Spatial Reasoning in Vision-Language Models spatial reasoninggroundedvisionlanguagemodels https://robotics-transformer2.github.io/ RT-2: Vision-Language-Action Models rtvisionlanguageactionmodels https://arxiv.org/abs/2505.21061 [2505.21061] LPOI: Listwise Preference Optimization for Vision Language Models Abstract page for arXiv paper 2505.21061: LPOI: Listwise Preference Optimization for Vision Language Models preference optimizationvisionlanguagemodels https://research.facebook.com/publications/taking-a-hint-leveraging-explanations-to-make-vision-and-language-models-more-grounded/ Taking a HINT: Leveraging Explanations to Make Vision and Language Models More Grounded - Meta... In this work, we propose a generic approach called Human Importance-aware Network Tuning (HINT) that effectively leverages human demonstrations to improve... https://publica.fraunhofer.de/entities/publication/b07a7aca-f1a7-42b5-9d7c-a1fba76a3663 Zero-Shot Open-Vocabulary OOD Object Detection and Grounding using Vision Language Models Automated driving involves complex perception tasks that require a precise understanding of diverse traffic scenarios and confident navigation. Traditional... https://arxiv.org/abs/2503.21817 [2503.21817] Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models via... Abstract page for arXiv paper 2503.21817: Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models via Adaptive Token Skipping https://arxiv.org/abs/2407.12366 [2407.12366] NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models Abstract page for arXiv paper 2407.12366: NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models https://is.mpg.de/en/publications/yeliuhe25 VERA: Explainable Video Anomaly Detection via Verbalized Learning of Vision-Language Models | MPI-IS Our goal is to understand the principles of Perception, Action and Learning in autonomous systems that successfully interact with complex environments and to... https://publications.ri.cmu.edu/toward-more-reliable-multimodal-systems-mitigating-hallucinations-in-large-vision-language-models Toward More Reliable Multimodal Systems: Mitigating Hallucinations in Large Vision-Language Models... more reliable https://arxiv.org/abs/2508.01943 [2508.01943] ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied Tasks Abstract page for arXiv paper 2508.01943: ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied Tasks