Sponsored by eBay
vision models | Shop on eBay
Buy and sell on the world's online marketplace. Top brands, low prices, and free shipping on many items.
https://docs.cloud.google.com/python/docs/reference/vertexai/latest/vertexai.preview.vision_models
Module vision_models (1.152.0) | Python client libraries | Google Cloud Documentation
module vision
https://research.google/blog/supercharge-your-computer-vision-models-with-the-tensorflow-object-detection-api/?ref=hackernoon.com
Supercharge your Computer Vision models with the TensorFlow Object Detection API
Posted by Jonathan Huang, Research Scientist and Vivek Rathod, Software Engineer (Cross-posted on the Google Open Source Blog) At Google, we develo...
computer vision modelsobject detectionsupercharge
https://aws.amazon.com/blogs/machine-learning/foundational-vision-models-and-visual-prompt-engineering-for-autonomous-driving-applications/
Foundational vision models and visual prompt engineering for autonomous driving applications |...
Nov 15, 2023 - Prompt engineering has become an essential skill for anyone working with large language models (LLMs) to generate high-quality and relevant texts. Although...
vision modelsprompt engineeringautonomous drivingfoundationalvisual
https://replit.com/guides/create-a-virtual-whiteboard-with-roboflow-vision-models
Create a virtual whiteboard with Roboflow vision models - Replit
Replit is an AI-driven software creation platform where everyone can build, share, and ship apps and websites, fast.
vision modelscreatevirtualwhiteboardroboflow
https://reactor.microsoft.com/en-us/reactor/events/26294/?ref=berberich.dev
Python + AI: Vision models | Microsoft Reactor
Learn new skills, meet new peers, and find career mentorship. Virtual events are running around the clock so join us anytime, anywhere!
python aivision modelsmicrosoftreactor
https://platform.kimi.ai/docs/guide/use-kimi-vision-model
Configure Kimi Vision Models - Kimi API Platform
Jul 29, 2026 - Kimi K3 is our flagship model for long-horizon coding and end-to-end knowledge work, with a 1M-token context window and industry-leading intelligence. The Kimi...
vision modelsconfigurekimiapiplatform
https://www.tensorflow.org/api_docs/python/tfm/vision/models/VideoClassificationModel?authuser=2
tfm.vision.models.VideoClassificationModel | TensorFlow v2.16.1
A video classification class builder.
vision modelstfmtensorflow
https://andrejusb.blogspot.com/2025/04/running-vision-models-on-apple-silicon.html
Andrej Baranovskij Blog: Running Vision Models on Apple Silicon with MLX-VLM
Blog about Oracle, Machine Learning and Cloud
vision models
https://usbios.ai/
BiOS — Train it. Run it. Own it. | Fine-tune 250k+ open LLMs & vision-language models
Fine-tune 250k+ open models with 15+ methods including LoRA, QLoRA, and full fine-tune. 6 alignment algorithms (DPO, SimPO, ORPO, KTO), continued pre-training,...
https://www.anjiecheng.me/SpatialRGPT
SpatialRGPT: Grounded Spatial Reasoning in Vision-Language Models
SpatialRGPT: Grounded Spatial Reasoning in Vision-Language Models
spatial reasoninggroundedvisionlanguagemodels
https://robotics-transformer2.github.io/
RT-2: Vision-Language-Action Models
rtvisionlanguageactionmodels
https://dev.to/aimodels-fyi/sora-a-review-on-background-technology-limitations-and-opportunities-of-large-vision-models-104h
Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models -...
Sora: A Review on Background, Technology, Limitations, and Opportunities of Large Vision Models. Tagged with machinelearning, ai, beginners, datascience.
https://arxiv.org/abs/2505.21061
[2505.21061] LPOI: Listwise Preference Optimization for Vision Language Models
Abstract page for arXiv paper 2505.21061: LPOI: Listwise Preference Optimization for Vision Language Models
preference optimizationvisionlanguagemodels
https://research.facebook.com/publications/taking-a-hint-leveraging-explanations-to-make-vision-and-language-models-more-grounded/
Taking a HINT: Leveraging Explanations to Make Vision and Language Models More Grounded - Meta...
In this work, we propose a generic approach called Human Importance-aware Network Tuning (HINT) that effectively leverages human demonstrations to improve...
https://ivi.fnwi.uva.nl/vislab/publication/yingjun-iclr-2026/
Prompt-Robust Vision-Language Models via Meta-Finetuning | VIS Lab
Apr 1, 2026 - ision-language models (VLMs) have demonstrated remarkable generalization across diverse tasks by leveraging large-scale image-text pretraining. However, their...
vision language modelspromptrobustviameta
https://www.utwente.nl/en/eemcs/ps/education/master%20theses/Adarsh-2/
Assignments: DRONE-BASED OBJECT DETECTION AND EXPLANATION USING VISION-LANGUAGE MODELS | Pervasive...
vision language modelsobject detection
https://zhangtemplar.github.io/battle-backbone/
Battle of the Backbones A Large-Scale Comparison of Pretrained Models across Computer Vision Tasks...
Oct 29, 2023 - This is my reading note for Battle of the Backbones: A Large-Scale Comparison of Pretrained Models across Computer Vision Tasks. This paper benchmarks...
https://arxiv.org/abs/2503.22081
[2503.22081] A Survey on Remote Sensing Foundation Models: From Vision to Multimodality
Abstract page for arXiv paper 2503.22081: A Survey on Remote Sensing Foundation Models: From Vision to Multimodality
https://research.nvidia.com/publication/2026-04_qcaleval-benchmarking-vision-language-models-quantum-calibration-plot
QCalEval: Benchmarking Vision-Language Models for Quantum Calibration Plot Understanding | Research
Quantum computing calibration depends on interpreting experimental data, and calibration plots provide the most universal human-readable representation for...
vision language modelsbenchmarking
https://arxiv.org/abs/2503.16538v1
[2503.16538v1] Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and...
Abstract page for arXiv paper 2503.16538v1: Leveraging Vision-Language Models for Open-Vocabulary Instance Segmentation and Tracking
vision language models
https://www.frontiersin.org/journals/neuroscience/articles/10.3389/fnins.2025.1615276/full
Frontiers | Editorial: Advances in computer vision: from deep learning models to practical...
Computer vision has emerged as one of the most transformative fields in artificial intelligence, with deep learning models driving unprecedented advancements...
deep learning modelscomputer vision
https://computervisionmodels.blogspot.com/
Computer Vision Models
I'm trying to write a new computer vision textbook. I'm going to post updated versions here as I do so. The plan is to first teach probability and machine...
computer visionmodels
https://developers.redhat.com/articles/2025/10/27/multimodal-ai-edge-deploy-vision-language-models-ramalama
Deploy vision language models with RamaLama | Red Hat Developer
Oct 27, 2025 - Learn how to deploy multimodal AI models on edge devices using the RamaLama CLI, from pulling your first vision language model (VLM) to serving it via an API.
vision language modelsred hatdeployramalamadeveloper
https://publica.fraunhofer.de/entities/publication/b07a7aca-f1a7-42b5-9d7c-a1fba76a3663
Zero-Shot Open-Vocabulary OOD Object Detection and Grounding using Vision Language Models
Automated driving involves complex perception tasks that require a precise understanding of diverse traffic scenarios and confident navigation. Traditional...
https://arxiv.org/abs/2503.21817
[2503.21817] Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models via...
Abstract page for arXiv paper 2503.21817: Skip-Vision: Efficient and Scalable Acceleration of Vision-Language Models via Adaptive Token Skipping
https://jobs.inria.fr/public/classic/fr/offres/2026-09994
2026-09994 - Computer Vision: PhD thesis on Simulatable Physics-aware World Models
Offre d'emploi Inria
computer visionphd thesis
https://arxiv.org/abs/2407.12366
[2407.12366] NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
Abstract page for arXiv paper 2407.12366: NavGPT-2: Unleashing Navigational Reasoning Capability for Large Vision-Language Models
https://is.mpg.de/en/publications/yeliuhe25
VERA: Explainable Video Anomaly Detection via Verbalized Learning of Vision-Language Models | MPI-IS
Our goal is to understand the principles of Perception, Action and Learning in autonomous systems that successfully interact with complex environments and to...
https://publications.ri.cmu.edu/toward-more-reliable-multimodal-systems-mitigating-hallucinations-in-large-vision-language-models
Toward More Reliable Multimodal Systems: Mitigating Hallucinations in Large Vision-Language Models...
more reliable
https://aws.amazon.com/blogs/machine-learning/vision-use-cases-with-llama-3-2-11b-and-90b-models-from-meta/
Vision use cases with Llama 3.2 11B and 90B models from Meta | Artificial Intelligence
Sep 25, 2024 - This is the first time that the Llama models from Meta have been released with vision capabilities. These new capabilities expand the usability of Llama models...
https://docs.cloud.google.com/vertex-ai/docs/reference/rpc/cloud.ai.large_models.vision
Package cloud.ai.large_models.vision | Vertex AI | Google Cloud Documentation
package cloudlarge modelsaivisionvertex
https://github.com/mit-han-lab/efficientvit
GitHub - mit-han-lab/efficientvit: Efficient vision foundation models for high-resolution...
Efficient vision foundation models for high-resolution generation and perception. - mit-han-lab/efficientvit
han lab
https://arxiv.org/abs/2508.01943
[2508.01943] ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied Tasks
Abstract page for arXiv paper 2508.01943: ROVER: Recursive Reasoning Over Videos with Vision-Language Models for Embodied Tasks
https://www.ideals.illinois.edu/items/132695
Models and data sources for economical computer vision in camera networks | IDEALS
models and data
https://umu.diva-portal.org/smash/record.jsf?pid=diva2:1801598
Evaluation of Tree Planting using Computer Vision models YOLO and U-Net
computer vision modelstree planting
https://github.com/codezakh/SelTDA
GitHub - codezakh/SelTDA: [CVPR 23] Q: How to Specialize Large Vision-Language Models to...
[CVPR 23] Q: How to Specialize Large Vision-Language Models to Data-Scarce VQA Tasks? A: Self-Train on Unlabeled Images! - codezakh/SelTDA
https://nim.nsc.liu.se/projects/8095/
Robotic Foundation Models: Vision-Language-Action Frameworks for Generalist Robotics -- extension...
foundation modelsroboticvisionlanguage