English
Related papers

Related papers: OmniVaT: Single Domain Generalization for Multimod…

200 papers

To operate effectively in the real world, robots should integrate multimodal reasoning with precise action generation. However, existing vision-language-action (VLA) models often sacrifice one for the other, narrow their abilities to…

Robotics · Computer Science 2026-03-04 Shuai Yang , Hao Li , Bin Wang , Yilun Chen , Yang Tian , Tai Wang , Hanqing Wang , Feng Zhao , Yiyi Liao , Jiangmiao Pang

Videos convey richer information than images or text, capturing both spatial and temporal dynamics. However, most existing video customization methods rely on reference images or task-specific temporal priors, failing to fully exploit the…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Pengze Zhang , Yanze Wu , Mengtian Li , Xu Bai , Songtao Zhao , Fulong Ye , Chong Mou , Xinghui Li , Zhuowei Chen , Qian He , Mingyuan Gao

The goal of Audio-Visual Segmentation (AVS) is to localize and segment the sounding source objects from video frames. Research on AVS suffers from data scarcity due to the high cost of fine-grained manual annotations. Recent works attempt…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Kyungbok Lee , You Zhang , Zhiyao Duan

Reasoning over multiple modalities, e.g. in Visual Question Answering (VQA), requires an alignment of semantic concepts across domains. Despite the widespread success of end-to-end learning, today's multimodal pipelines by and large…

Computer Vision and Pattern Recognition · Computer Science 2021-09-10 Jan-Martin O. Steitz , Jonas Pfeiffer , Iryna Gurevych , Stefan Roth

Visual language tracking (VLT) has emerged as a cutting-edge research area, harnessing linguistic data to enhance algorithms with multi-modal inputs and broadening the scope of traditional single object tracking (SOT) to encompass video…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Xuchen Li , Shiyu Hu , Xiaokun Feng , Dailing Zhang , Meiqi Wu , Jing Zhang , Kaiqi Huang

Current visual representation learning remains bifurcated: vision-language models (e.g., CLIP) excel at global semantic alignment but lack spatial precision, while self-supervised methods (e.g., MAE, DINO) capture intricate local structures…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Shangzhe Di , Zhonghua Zhai , Weidi Xie

Visual-tactile fused sensing for object clustering has achieved significant progresses recently, since the involvement of tactile modality can effectively improve clustering performance. However, the missing data (i.e., partial data) issues…

Robotics · Computer Science 2021-02-16 Tao Zhang , Yang Cong , Gan Sun , Jiahua Dong , Yuyang Liu , Zhengming Ding

Single-domain generalization for object detection (S-DGOD) seeks to transfer learned representations from a single source domain to unseen target domains. While recent approaches have primarily focused on achieving feature invariance, they…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Zhenwei He , Hongsu Ni

Addressing the challenge of domain shift between datasets is vital in maintaining model performance. In the context of cross-domain object detection, the teacher-student framework, a widely-used semi-supervised model, has shown significant…

Computer Vision and Pattern Recognition · Computer Science 2024-09-19 Runou Yang , Tian Tian , Jinwen Tian

Multi-task dense prediction, which aims to jointly solve tasks like semantic segmentation and depth estimation, is crucial for robotics applications but suffers from domain shift when deploying models in new environments. While unsupervised…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Beomseok Kang , Niluthpol Chowdhury Mithun , Mikhail Sizintsev , Han-Pang Chiu , Supun Samarasekera

The paradigm of programmable diagram generation is evolving rapidly, playing a crucial role in structured visualization. However, most existing studies are confined to a narrow range of task formulations and language support, constraining…

Artificial Intelligence · Computer Science 2026-04-08 Haoyue Yang , Xuanle Zhao , Xuexin Liu , Feibang Jiang , Yao Zhu

A generalist robotic policy needs both semantic understanding for task planning and the ability to interact with the environment through predictive capabilities. To tackle this, we present MM-ACT, a unified Vision-Language-Action (VLA)…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Haotian Liang , Xinyi Chen , Bin Wang , Mingkang Chen , Yitian Liu , Yuhao Zhang , Zanxin Chen , Tianshuo Yang , Yilun Chen , Jiangmiao Pang , Dong Liu , Xiaokang Yang , Yao Mu , Wenqi Shao , Ping Luo

We present a single neural network architecture composed of task-agnostic components (ViTs, convolutions, and LSTMs) that achieves state-of-art results on both the ImageNav ("go to location in <this picture>") and ObjectNav ("find a chair")…

Computer Vision and Pattern Recognition · Computer Science 2023-03-15 Karmesh Yadav , Arjun Majumdar , Ram Ramrakhya , Naoki Yokoyama , Alexei Baevski , Zsolt Kira , Oleksandr Maksymets , Dhruv Batra

The widespread adoption of mobile devices and data collection technologies has led to an exponential increase in trajectory data, presenting significant challenges in spatio-temporal data mining, particularly for efficient and accurate…

Computer Vision and Pattern Recognition · Computer Science 2025-05-26 Yuanshao Zhu , James Jianqiao Yu , Xiangyu Zhao , Xiao Han , Qidong Liu , Xuetao Wei , Yuxuan Liang

Image-based Virtual Try-On (VTON) concerns the synthesis of realistic person imagery through garment re-rendering under human pose and body constraints. In practice, however, existing approaches are typically optimized for specific data…

Computer Vision and Pattern Recognition · Computer Science 2026-03-12 Zhaotong Yang , Yong Du , Shengfeng He , Yuhui Li , Xinzhe Li , Yangyang Xu , Junyu Dong , Jian Yang

Recent advances in Visual Question Answering (VQA) have demonstrated impressive performance in natural image domains, with models like LLaVA leveraging large language models (LLMs) for open-ended reasoning. However, their generalization…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Xinjin Li , Yulie Lu , Jinghan Cao , Yu Ma , Zhenglin Li , Yeyang Zhou

Multimodal representation learning models have demonstrated successful operation across complex tasks, and the integration of vision-language models (VLMs) has further enabled embedding models with instruction-following capabilities.…

Artificial Intelligence · Computer Science 2026-02-24 Wei-Yao Wang , Kazuya Tateishi , Qiyu Wu , Shusuke Takahashi , Yuki Mitsufuji

Recent works have proven that many relevant visual tasks are closely related one to another. Yet, this connection is seldom deployed in practice due to the lack of practical methodologies to transfer learned concepts across different…

Computer Vision and Pattern Recognition · Computer Science 2019-10-04 Pierluigi Zama Ramirez , Alessio Tonioni , Samuele Salti , Luigi Di Stefano

With the ever-increasing amount of data, the central challenge in multimodal learning involves limitations of labelled samples. For the task of classification, techniques such as meta-learning, zero-shot learning, and few-shot learning…

Computer Vision and Pattern Recognition · Computer Science 2021-06-29 Nihar Bendre , Kevin Desai , Peyman Najafirad

Visual grounding aims to align visual information of specific regions of images with corresponding natural language expressions. Current visual grounding methods leverage pre-trained visual and language backbones independently to obtain…

Computer Vision and Pattern Recognition · Computer Science 2024-07-09 Jiaxi Wang , Wenhui Hu , Xueyang Liu , Beihu Wu , Yuting Qiu , YingYing Cai