https://arxiv.org/abs/2603.25406
[2603.25406] MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal...
Abstract page for arXiv paper 2603.25406: MMaDA-VLA: Large Diffusion Vision-Language-Action Model with Unified Multi-Modal Instruction and Generation
vision language action model
https://arxiv.org/abs/2411.19650v1
[2411.19650v1] CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and...
Abstract page for arXiv paper 2411.19650v1: CogACT: A Foundational Vision-Language-Action Model for Synergizing Cognition and Action in Robotic Manipulation
vision language action model
https://arxiv.org/abs/2507.09160
[2507.09160] Tactile-VLA: Unlocking Vision-Language-Action Model's Physical Knowledge for Tactile...
Abstract page for arXiv paper 2507.09160: Tactile-VLA: Unlocking Vision-Language-Action Model's Physical Knowledge for Tactile Generalization
vision language action model
https://openreview.net/forum?id=UXdxYnkJtX
ShowUI: One Vision-Language-Action Model for Generalist GUI Agent | OpenReview
Graphical User Interface (GUI) automation holds significant promise for enhancing human productivity by assisting with digital tasks. While recent Large...
vision language action modelone
https://arxiv.org/html/2411.17465v1
ShowUI: One Vision-Language-Action Model for GUI Visual Agent
vision language action modeloneguivisualagent