https://vivo.tib.eu/fis/display/n30034
Patent Figure Classification using Large Vision-language Models
vision language modelspatentfigureclassificationusing
https://researchconnect.buffalo.edu/en/publications/text-image-de-contextualization-detection-using-vision-language-m/
TEXT-IMAGE DE-CONTEXTUALIZATION DETECTION USING VISION-LANGUAGE MODELS - SUNY University at Buffalo
vision language modelsuniversity at buffalotextimagedetection
https://zilliz.com/ai-faq/what-advancements-are-expected-in-visionlanguage-models-for-realtime-applications
What advancements are expected in Vision-Language Models for real-time applications? - Zilliz...
Vision-Language Models (VLMs) are expected to see significant advancements in real-time applications, mainly due to impr
vision language modelsfor realadvancementsexpectedtime
https://openreview.net/forum?id=1fpjV6xQ6Q&referrer=%5Bthe%20profile%20of%20Katia%20P.%20Sycara%5D(%2Fprofile%3Fid%3D~Katia_P._Sycara1)
Incorporating Generative Feedback for Mitigating Hallucinations in Large Vision-Language Models |...
While recent Large Vision-Language Models (LVLMs) have shown remarkable performance in multi-modal tasks, they are prone to generating hallucinatory text...
vision language modelsincorporatinggenerativefeedbackhallucinations
https://www.catalyzex.com/paper/replanning-human-robot-collaborative-tasks
Replanning Human-Robot Collaborative Tasks with Vision-Language Models via Semantic and Physical...
Replanning Human-Robot Collaborative Tasks with Vision-Language Models via Semantic and Physical Dual-Correction: Paper and Code. Human-Robot Collaboration...
vision language modelshumanrobotcollaborativetasks
https://lrec.elra.info/lrec2026-main-134
MEUR: A Benchmark for Evaluating Vision-Language Models on Multimodal Event Understanding and...
May 1, 2026 - Event understanding and reasoning play critical roles in thoroughly evaluating the capabilities of Vision-Language Models (VLMs); however, existing Visual Quest
vision language modelsbenchmarkevaluatingmultimodalevent
https://maestro-robot.github.io/
Maestro: Orchestrating Robotics Modules with Vision-Language Models for Zero-Shot Generalist Robots
Maestro: Orchestrating Robotics Modules with Vision-Language Models for Zero-Shot Generalist Robots
vision language modelszero shotmaestroroboticsmodules
https://papers.neurips.cc/paper_files/paper/2025/hash/12750d99d0faa73763108ff2bbeb54fd-Abstract-Datasets_and_Benchmarks_Track.html
DrVD-Bench: Do Vision-Language Models Reason Like Human Doctors in Medical Image Diagnosis?
vision language modelsmedical imagebenchreasonlike
https://datasciencedojo.com/blog/vision-language-models-moondream-2/
Vision Language Models: Introducing the new VLM Moondream 2
Explore Moondream 2, one of the cutting-edge vision language models, and its impact on AI applications in our detailed guide.
vision language modelsthe newintroducingvlm
https://docs.nvidia.com/nim/vision-language-models/1.3.0/search.html
Search - NVIDIA NIM for Vision Language Models (VLMs)
vision language modelssearchnvidianimvlms
https://uk.mathworks.com/help/vision/vision-language-models.html?s_tid=CRUX_lftnav
Vision-Language Models - MATLAB & Simulink
Perform image classification, retrieval, captioning, and object detection tasks using vision-language models
vision language modelsmatlabsimulink
https://deepai.org/publication/vlue-a-multi-task-benchmark-for-evaluating-vision-language-models
VLUE: A Multi-Task Benchmark for Evaluating Vision-Language Models | DeepAI
May 30, 2022 - 05/30/22 - Recent advances in vision-language pre-training (VLP) have demonstrated impressive performance in a range of vision-language (VL) ...
vision language modelsvluemultitaskbenchmark
https://aclanthology.org/2026.eacl-long.197/
Enhancing the Safety of Medical Vision-Language Models by Synthetic Demonstrations - ACL Anthology
Zhiyu Xue, Reza Abbasi-Asl, Ramtin Pedarsani. Proceedings of the 19th Conference of the European Chapter of the Association for Computational Linguistics...
vision language modelsthe safetyenhancingmedicalsynthetic
https://www.nec-labs.com/research/media-analytics/projects/foundational-vision-language-models/
Foundational Vision-Language Models | NEC Labs
Apr 9, 2025 - Our foundational models enable ubiquitous usage of computer vision across scenarios, applications and user preferences.
vision language modelsfoundationalneclabs
https://tldr.takara.ai/p/2510.22282
CityRiSE: Reasoning Urban Socio-Economic Status in Vision-Language Models via Reinforcement...
Harnessing publicly available, large-scale web data, such as street view and satellite imagery, urban socio-economic sensing is of paramount importance for a...
vision language modelseconomic statusreasoningurbansocio
https://www.datacamp.com/ko/blog/vlms-ai-vision-language-models
Vision Language Models (VLMs) Explained | DataCamp
Vision language models (VLMs) are AI models that can understand and process both visual and textual data, enabling tasks like image captioning, visual question...
vision language modelsvlmsexplaineddatacamp
https://aws.amazon.com/blogs/machine-learning/cost-effective-deployment-of-vision-language-models-for-pet-behavior-detection-on-aws-inferentia2/
Cost effective deployment of vision-language models for pet behavior detection on AWS Inferentia2 |...
May 6, 2026 - Tomofun, the Taiwan-headquartered pet-tech startup behind the Furbo Pet Camera, is redefining how pet owners interact with their pets remotely. To reduce costs...
vision language modelscost effectivepet behaviordeploymentdetection
https://jobs.accel.com/companies/facebook-2-07d2eb1b-d3ad-4084-93df-74a124ac9233/jobs/76244850-ai-research-scientist-vlm-vision-language-models
AI Research Scientist, VLM (vision language models) @ Facebook | Accel Job Board
Search job openings across the Accel network.
ai research scientistvision language modelsjob boardvlmfacebook
https://www.umassd.edu/events/cms/a-noise-based-defense-for-stealthy-backdoor-attacks-in-large-vision-language-models.php
Events: A Noise-Based Defense for Stealthy Backdoor Attacks in Large Vision-Language Models | UMass...
May 25, 2026 to May 25, 2026
vision language modelseventsnoisebaseddefense
https://ai.updf.com/paper-detail/unveiling-encoder-free-vision-language-models-diao-cui-11159e03ed50d72cd84f7949b09bf87b6d717c1a
Unveiling Encoder-Free Vision-Language Models
EVE is launched, an encoder-free vision-language model that can be trained and forwarded efficiently and can impressively rival the encoder-based VLMs of...
vision language modelsunveilingencoderfree
https://conf.researchr.org/details/fse-2026/fse-2026-research-papers/88/ViBR-Automated-Bug-Replay-from-Video-based-Reports-Using-Vision-Language-Models
ViBR: Automated Bug Replay from Video-based Reports Using Vision-Language Models (FSE 2026 -...
The ACM International Conference on the Foundations of Software Engineering (FSE) is an internationally renowned forum for researchers, practitioners, and...
vision language modelsautomatedbugreplayvideo
https://scholars.hkbu.edu.hk/en/publications/foundations-of-vision-language-models-concepts-and-roadmap/
Foundations of Vision-Language Models: Concepts and Roadmap - Hong Kong Baptist University
vision language modelshong kongfoundationsconceptsroadmap
https://papers.nips.cc/paper_files/paper/2025/hash/1f6af963e891e7efa229c24a1607fa7f-Abstract-Conference.html
Approximate Domain Unlearning for Vision-Language Models
vision language modelsapproximatedomainunlearning
https://www.thoughtworks.com/en-ec/radar/techniques/vision-language-models-for-end-to-end-document-parsing
Vision language models for end-to-end document parsing | Technology Radar | Thoughtworks Ecuador
Document parsing often relies on multi-stage pipelines combining layout detection, traditional OCR and post-processing scripts. These approaches often struggle...
vision language modelsdocument parsingtechnology radarendthoughtworks
https://the-decoder.com/vision-language-models-struggle-to-solve-simple-visual-puzzles-that-humans-find-intuitive/
Vision language models struggle to solve simple visual puzzles that humans find intuitive
Oct 27, 2024 - A new study from Germany's TU Darmstadt shows that even the most sophisticated AI image models fail at simple visual reasoning tasks.
vision language modelsto solvestrugglesimplevisual
https://paperium.net/article/en/17127/switch-kd-visual-switch-knowledge-distillation-for-vision-language-models
Switch-KD: Visual-Switch Knowledge Distillation for Vision-Language Models: Analysis, Review &...
Quick breakdown of the 'Switch-KD: Visual-Switch Knowledge Distillation for Vision-Language Models' paper. Methods, results, strengths/weaknesses expl
vision language modelsknowledge distillationswitchkdvisual
https://arxiv.org/abs/2505.21061
[2505.21061] LPOI: Listwise Preference Optimization for Vision Language Models
Abstract page for arXiv paper 2505.21061: LPOI: Listwise Preference Optimization for Vision Language Models
vision language modelspreferenceoptimization
https://research.ibm.com/publications/incorporating-structured-representations-into-pretrained-vision-and-language-models-using-scene-graphs
Incorporating Structured Representations into Pretrained Vision & Language Models Using Scene...
vision language modelsincorporatingstructuredrepresentationsusing
https://groovesquid.com/paper/summary-of-fake-news-detection-and-manipulation-reasoning-via-large-vision-language-models-by-ruihan-jin-et-al/
Summary of Fake News Detection and Manipulation Reasoning Via Large Vision-language Models, by...
Jul 13, 2025 - Fake News Detection and Manipulation Reasoning via Large Vision-Language Models by Ruihan Jin, Ruibo Fu, Zhengqi Wen, Shuai Zhang, Yukun Liu, Jianhua Tao First
vision language modelsfake newssummarydetectionmanipulation
https://www.nutrient.io/guides/python/extraction/extract-data-from-image-vlm/
Extracting data from images using vision language models | Nutrient Python SDK
Extract structured data from images using vision language models with Nutrient Python SDK.
vision language modelspython sdkdataimagesusing
https://repository.gatech.edu/entities/publication/944f2659-3269-42a0-91cd-55e3ca6492f3
Mutual exclusivity bias and spatial reasoning in Vision-Language Models
Despite rapid advancements in machine learning, enabling models to generalize beyond their training data, they still lag significantly behind the learning...
vision language modelsspatial reasoningmutualexclusivitybias
https://wandb.ai/byyoung3/ML_NEWS3/reports/LLaVA-o1-Advancing-structured-reasoning-in-vision-language-models--VmlldzoxMDMyMzc1Mg
LLaVA-o1: Advancing structured reasoning in vision-language models
Dec 3, 2024 - Discover how LLaVA-o1 tackles reasoning challenges in multimodal AI with structured problem-solving. Learn about its dataset, capabilities, and performance...
vision language modelsadvancingstructuredreasoning
https://huggingface.co/papers/2503.03278
Paper page - Enhancing Abnormality Grounding for Vision Language Models with Knowledge Descriptions
Join the discussion on this paper page
vision language modelspaperenhancingabnormalitygrounding
https://www.krcmic.com/tag/vision-language-models/
vision language models Archives | Krcmic.com - Online Marketing Professional - Personal Portfolio
vision language modelsonline marketingpersonal portfolioarchivesprofessional
https://liner.com/review/promptrobust-visionlanguage-models-via-metafinetuning
Prompt-Robust Vision-Language Models via Meta-Finetuning [Quick Review]
Regarding this ICLR 2026 paper, this review summarizes Promise, a meta-learning framework for prompt-robust vision-language models.
vision language modelsquick reviewpromptrobustvia
https://scholar.hit.edu.cn/en/publications/causal-tracing-of-object-representations-in-large-vision-language/
Causal Tracing of Object Representations in Large Vision Language Models: Mechanistic...
vision language modelscausal tracingobjectrepresentationslarge
https://cbirt.net/ai-and-biomedicine-enhancing-diagnosis-with-vision-language-models/
AI and Biomedicine: Enhancing Diagnosis with Vision-Language Models - CBIRT
Jun 30, 2024 - Advancing medical AI: Llama3-Med enhances biomedical visual question answering with new datasets and innovative image encoding.
vision language modelsaibiomedicineenhancingdiagnosis
https://www.qualcomm.com/developer/software/qualcomm-interactive-video-dataset-qivd
Interactive Video Dataset for Vision-Language Models | Qualcomm
Improve your AI's visual understanding and response with QIVD, featuring 2,900 video files and 13 categories of action attributes and object detection.
vision language modelsinteractive videodatasetqualcomm
https://www.proceedings.com/079017-4469.html
WikiDO: A New Benchmark Evaluating Cross-Modal Retrieval for Vision-Language Models -...
The world's premier source for conference proceedings, offering Print-on-Demand, DOI, and Content Hosting services.
vision language modelsnew benchmarkevaluatingcrossmodal
https://irep.mbzuai.ac.ae/items/7dce91fd-af23-4e11-b313-7123a17af18e
Towards Explainable and Controllable Vision-Language Models for Chest X-ray Imaging
While large-scale vision-language models (VLMs) show promise across various tasks, their application in safety-critical domains like medical imaging is...
vision language modelsx raytowardscontrollablechest
https://f4u.in/cost-effective-deployment-of-vision-language-models-for-pet-behavior-detection-on-aws-inferentia2/
Cost effective deployment of vision-language models for pet behavior detection on AWS Inferentia2 -...
May 6, 2026 - Tomofun, the Taiwan-headquartered pet-tech startup behind the Furbo Pet Camera, is redefining how pet owners interact with their pets remotely. Furbo combines
vision language modelscost effectivepet behaviordeploymentdetection
https://scipapermill.com/2025/11/30/vision-language-models-bridging-perception-reasoning-and-real-world-interaction/
Vision-Language Models: Bridging Perception, Reasoning, and Real-World Interaction
Dec 28, 2025 - Latest 50 papers on vision-language models: Nov. 30, 2025
vision language modelsreal worldbridgingperceptionreasoning