English
Related papers

Related papers: Animation Needs Attention: A Holistic Approach to …

200 papers

Vision-Language-Action (VLA) models for autonomous driving show promise but falter in unstructured corner case scenarios, largely due to a scarcity of targeted benchmarks. To address this, we introduce Impromptu VLA. Our core contribution…

Computer Vision and Pattern Recognition · Computer Science 2025-05-30 Haohan Chi , Huan-ang Gao , Ziming Liu , Jianing Liu , Chenyu Liu , Jinwei Li , Kaisen Yang , Yangcheng Yu , Zeda Wang , Wenyi Li , Leichen Wang , Xingtao Hu , Hao Sun , Hang Zhao , Hao Zhao

Vision-language models (VLMs) have achieved impressive performance on multimodal reasoning tasks such as visual question answering, image captioning and so on, but their inference cost remains a significant challenge due to the large number…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Weichen Zhang , Zhui Zhu , Ningbo Li , Shilong Tao , Kebin Liu , Yunhao Liu

Multi-modal Large Language Models (MLLMs) are increasingly prominent in the field of artificial intelligence. Visual instruction fine-tuning (IFT) is a vital process for aligning MLLMs' output with user's intentions. High-quality and…

Computer Vision and Pattern Recognition · Computer Science 2024-01-18 Xiaotian Han , Yiqi Wang , Bohan Zhai , Quanzeng You , Hongxia Yang

The rapid advancements in vision-language models (VLMs), such as CLIP, have intensified the need to address distribution shifts between training and testing datasets. Although prior Test-Time Training (TTT) techniques for VLMs have…

Computer Vision and Pattern Recognition · Computer Science 2025-02-05 Yuto Kojima , Jiarui Xu , Xueyan Zou , Xiaolong Wang

Vision-Language Models (VLM) can support clinicians by analyzing medical images and engaging in natural language interactions to assist in diagnostic and treatment tasks. However, VLMs often exhibit "hallucinogenic" behavior, generating…

Artificial Intelligence · Computer Science 2024-10-11 Shenghuan Sun , Alexander Schubert , Gregory M. Goldgof , Zhiqing Sun , Thomas Hartvigsen , Atul J. Butte , Ahmed Alaa

Visual language models (VLMs) rapidly progressed with the recent success of large language models. There have been growing efforts on visual instruction tuning to extend the LLM with visual inputs, but lacks an in-depth study of the visual…

Computer Vision and Pattern Recognition · Computer Science 2024-05-20 Ji Lin , Hongxu Yin , Wei Ping , Yao Lu , Pavlo Molchanov , Andrew Tao , Huizi Mao , Jan Kautz , Mohammad Shoeybi , Song Han

Detecting temporal changes in geographical landscapes is critical for applications like environmental monitoring and urban planning. While remote sensing data is abundant, existing vision-language models (VLMs) often fail to capture…

Computer Vision and Pattern Recognition · Computer Science 2025-05-23 Hosam Elgendy , Ahmed Sharshar , Ahmed Aboeitta , Yasser Ashraf , Mohsen Guizani

Vision-Language-Action (VLA) models are advancing autonomous driving by replacing modular pipelines with unified end-to-end architectures. However, current VLAs face two expensive requirements: (1) massive dataset collection, and (2) dense…

Artificial Intelligence · Computer Science 2026-02-27 Ishaan Rawal , Shubh Gupta , Yihan Hu , Wei Zhan

Multiple instance learning (MIL)-based framework has become the mainstream for processing the whole slide image (WSI) with giga-pixel size and hierarchical image context in digital pathology. However, these methods heavily depend on a…

Computer Vision and Pattern Recognition · Computer Science 2025-02-13 Jiangbo Shi , Chen Li , Tieliang Gong , Yefeng Zheng , Huazhu Fu

Automatically generating data visualizations in response to human utterances on datasets necessitates a deep semantic understanding of the data utterance, including implicit and explicit references to data attributes, visualization tasks,…

Artificial Intelligence · Computer Science 2024-07-10 Hannah K. Bako , Arshnoor Bhutani , Xinyi Liu , Kwesi A. Cobbina , Zhicheng Liu

Manual slide creation is labor-intensive and requires expert prior knowledge. Existing natural language-based LLM generation methods struggle to capture the visual and structural nuances of slide designs. To address this, we formalize the…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Wenxin Tang , Jingyu Xiao , Wenxuan Jiang , Xi Xiao , Yuhang Wang , Xuxin Tang , Qing Li , Yuehe Ma , Junliang Liu , Shisong Tang , Michael R. Lyu

Thanks to the emerging of foundation models, the large language and vision models are integrated to acquire the multimodal ability of visual captioning, question answering, etc. Although existing multimodal models present impressive…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Bo Zhao , Boya Wu , Muyang He , Tiejun Huang

Videos are unique in their ability to capture actions which transcend multiple frames. Accordingly, for many years action recognition was the quintessential task for video understanding. Unfortunately, due to a lack of sufficiently diverse…

Computer Vision and Pattern Recognition · Computer Science 2026-05-06 Tanush Yadav , Mohammadreza Salehi , Jae Sung Park , Vivek Ramanujan , Hannaneh Hajishirzi , Yejin Choi , Ali Farhadi , Rohun Tripathi , Ranjay Krishna

Autonomous driving, particularly navigating complex and unanticipated scenarios, demands sophisticated reasoning and planning capabilities. While Multi-modal Large Language Models (MLLMs) offer a promising avenue for this, their use has…

Computer Vision and Pattern Recognition · Computer Science 2025-10-15 Hidehisa Arai , Keita Miwa , Kento Sasaki , Yu Yamaguchi , Kohei Watanabe , Shunsuke Aoki , Issei Yamamoto

Mobile app marketplaces require developers to disclose standardized content rating descriptors (CRDs) to inform users about potentially sensitive or restricted content. Ensuring the accuracy and consistency of these disclosures remains…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Dishanika Denipitiyage , Aruna Seneviratne , Suranga Seneviratne

The evolution of autonomous driving towards full automation demands robust interactive capabilities; however, the development of Vision-Language-Action (VLA) models is constrained by the sparsity of interactive scenarios and inadequate…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Haojie Feng , Peizhi Zhang , Mengjie Tian , Xinrui Zhang , Zhuoren Li , Junpeng Huang , Xiurong Wang , Junfan Zhu , Jianzhou Wang , Dongxiao Yin , Lu Xiong

State-of-the-art (SOTA) Automatic Speech Recognition (ASR) systems primarily rely on acoustic information while disregarding additional multi-modal context. However, visual information are essential in disambiguation and adaptation. While…

Artificial Intelligence · Computer Science 2025-10-17 Supriti Sinhamahapatra , Jan Niehues

We present Rodent-Bench, a novel benchmark designed to evaluate the ability of Multimodal Large Language Models (MLLMs) to annotate rodent behaviour footage. We evaluate state-of-the-art MLLMs, including Gemini-2.5-Pro, Gemini-2.5-Flash and…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Thomas Heap , Laurence Aitchison , Emma Cahill , Adriana Casado Rodriguez

Advertisement videos serve as a rich and valuable source of purpose-driven information, encompassing high-quality visual, textual, and contextual cues designed to engage viewers. They are often more complex than general videos of similar…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Zheyuan Zhang , Monica Dou , Linkai Peng , Hongyi Pan , Ulas Bagci , Boqing Gong

Video-language modeling has attracted much attention with the rapid growth of web videos. Most existing methods assume that the video frames and text description are semantically correlated, and focus on video-language modeling at video…

Computer Vision and Pattern Recognition · Computer Science 2022-12-06 Haoyu Lu , Mingyu Ding , Nanyi Fei , Yuqi Huo , Zhiwu Lu