English
Related papers

Related papers: MOCHA: Multi-modal Objects-aware Cross-arcHitectur…

200 papers

In this paper, we reveal that most current efficient multimodal fine-tuning methods are hindered by a key limitation: they are directly borrowed from LLMs, often neglecting the intrinsic differences of multimodal scenarios and even…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Yake Wei , Yu Miao , Dongzhan Zhou , Di Hu

Transforming neutral, characterless input motions to embody the distinct style of a notable character in real time is highly compelling for character animation. This paper introduces MOCHA, a novel online motion characterization framework…

Graphics · Computer Science 2023-10-17 Deok-Kyeong Jang , Yuting Ye , Jungdam Won , Sung-Hee Lee

In many multi-objective reinforcement learning (MORL) applications, being able to systematically explore the Pareto-stationary solutions under multiple non-convex reward objectives with theoretical finite-time sample complexity guarantee is…

Machine Learning · Computer Science 2025-07-30 Fnu Hairi , Jiao Yang , Tianchen Zhou , Haibo Yang , Chaosheng Dong , Fan Yang , Michinari Momma , Yan Gao , Jia Liu

Multimodal object detection offers a promising prospect to facilitate robust detection in various visual conditions. However, existing two-stream backbone networks are challenged by complex fusion and substantial parameter increments. This…

Computer Vision and Pattern Recognition · Computer Science 2024-07-24 Weiying Xie , Yusi Zhang , Tianlin Hui , Jiaqing Zhang , Jie Lei , Yunsong Li

The fusion of language and vision in large vision-language models (LVLMs) has revolutionized deep learning-based object detection by enhancing adaptability, contextual reasoning, and generalization beyond traditional architectures. This…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Ranjan Sapkota , Manoj Karkee

We present LOWA, a novel method for localizing objects with attributes effectively in the wild. It aims to address the insufficiency of current open-vocabulary object detectors, which are limited by the lack of instance-level attribute…

Computer Vision and Pattern Recognition · Computer Science 2023-06-01 Xiaoyuan Guo , Kezhen Chen , Jinmeng Rao , Yawen Zhang , Baochen Sun , Jie Yang

In this paper, we for the first time explore helpful multi-modal contextual knowledge to understand novel categories for open-vocabulary object detection (OVD). The multi-modal contextual knowledge stands for the joint relationship across…

Computer Vision and Pattern Recognition · Computer Science 2023-08-31 Yifan Xu , Mengdan Zhang , Xiaoshan Yang , Changsheng Xu

Text-motion retrieval systems learn shared embedding spaces from motion-caption pairs via contrastive objectives. However, each caption is not a deterministic label but a sample from a distribution of valid descriptions: different…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Nikolai Warner , Cameron Ethan Taylor , Irfan Essa , Apaar Sadhwani

Vision-Language Models (VLMs) based on Mixture-of-Experts (MoE) architectures have emerged as a pivotal paradigm in multimodal understanding, offering a powerful framework for integrating visual and linguistic information. However, the…

Computer Vision and Pattern Recognition · Computer Science 2025-04-02 Xiaoda Yang , JunYu Lu , Hongshun Qiu , Sijing Li , Hao Li , Shengpeng Ji , Xudong Tang , Jiayang Xu , Jiaqi Duan , Ziyue Jiang , Cong Lin , Sihang Cai , Zejian Xie , Zhuoyang Song , Songxin Zhang

This work addresses the task of weakly-supervised object localization. The goal is to learn object localization using only image-level class labels, which are much easier to obtain compared to bounding box annotations. This task is…

Computer Vision and Pattern Recognition · Computer Science 2023-12-18 David Kim , Sinhae Cha , Byeongkeun Kang

The field of object detection and understanding is rapidly evolving, driven by advances in both traditional CNN-based models and emerging multi-modal large language models (LLMs). While CNNs like ResNet and YOLO remain highly effective for…

Computer Vision and Pattern Recognition · Computer Science 2025-10-13 Nirmal Elamon , Rouzbeh Davoudi

Collaborative game-based learning environments offer rich opportunities for small-group knowledge construction, yet automatically predicting student collaboration satisfaction remains challenging. A critical barrier is modality degradation:…

Machine Learning · Computer Science 2026-05-19 Wen-Hsin Tsai , Chia-Ming Lee , Yuk-Ying Tung

Recently, learning open-vocabulary semantic segmentation from text supervision has achieved promising downstream performance. Nevertheless, current approaches encounter an alignment granularity gap owing to the absence of dense annotations,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-07 Yajie Liu , Pu Ge , Qingjie Liu , Di Huang

Multiple object tracking has been a challenging field, mainly due to noisy detection sets and identity switch caused by occlusion and similar appearance among nearby targets. Previous works rely on appearance models built on individual or…

Computer Vision and Pattern Recognition · Computer Science 2019-03-08 Zheng Tang , Jenq-Neng Hwang

Single-modal object detection tasks often experience performance degradation when encountering diverse scenarios. In contrast, multimodal object detection tasks can offer more comprehensive information about object features by integrating…

Computer Vision and Pattern Recognition · Computer Science 2024-12-25 Chang Liu , Xin Ma , Xiaochen Yang , Yuxiang Zhang , Yanni Dong

Vision-Language Models (VLMs) have achieved impressive performance in cross-modal understanding across textual and visual inputs, yet existing benchmarks predominantly focus on pure-text queries. In real-world scenarios, language also…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Qing'an Liu , Juntong Feng , Yuhao Wang , Xinzhe Han , Yujie Cheng , Yue Zhu , Haiwen Diao , Yunzhi Zhuge , Huchuan Lu

LLM-based ASR overcomes multilingual data scarcity by projecting speech representations into the LLM space to leverage its robust semantic and reasoning capabilities. However, while previous approaches typically enhance performance by…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-05 Junjie Li , Jing Peng , Yangui Fang , Shuai Wang , Kai Yu

Advancements in cross-modal feature extraction and integration have significantly enhanced performance in few-shot learning tasks. However, current multi-modal object detection (MM-OD) methods often experience notable performance…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Zeyu Shangguan , Daniel Seita , Mohammad Rostami

The pre-trained vision and language (V\&L) models have substantially improved the performance of cross-modal image-text retrieval. In general, however, V\&L models have limited retrieval performance for small objects because of the rough…

Computer Vision and Pattern Recognition · Computer Science 2024-09-26 Naoya Sogi , Takashi Shibata , Makoto Terao

Conventional methods for object detection usually require substantial amounts of training data and annotated bounding boxes. If there are only a few training data and annotations, the object detectors easily overfit and fail to generalize.…

Computer Vision and Pattern Recognition · Computer Science 2020-08-31 Geonuk Kim , Hong-Gyu Jung , Seong-Whan Lee