English
Related papers

Related papers: Having Difficulty Understanding Manuals? Automatic…

200 papers

Multimodal Large Language Models (MLLMs) often struggle with fine-grained perception, such as identifying small objects in high-resolution images or detecting key moments in long videos. Existing methods typically rely on complex,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Sanghwan Kim , Rui Xiao , Stephan Alaniz , Yongqin Xian , Zeynep Akata

Building models that comprehends videos and responds specific user instructions is a practical and challenging topic, as it requires mastery of both vision understanding and knowledge reasoning. Compared to language and image modalities,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Ji Qi , Kaixuan Ji , Jifan Yu , Duokang Wang , Bin Xu , Lei Hou , Juanzi Li

Instructional videos are an important resource to learn procedural tasks from human demonstrations. However, the instruction steps in such videos are typically short and sparse, with most of the video being irrelevant to the procedure. This…

Computer Vision and Pattern Recognition · Computer Science 2023-04-27 Nikita Dvornik , Isma Hadji , Ran Zhang , Konstantinos G. Derpanis , Animesh Garg , Richard P. Wildes , Allan D. Jepson

Screen recordings of mobile applications are easy to obtain and capture a wealth of information pertinent to software developers (e.g., bugs or feature requests), making them a popular mechanism for crowdsourced app feedback. Thus, these…

Writing and maintaining UI tests for mobile apps is a time-consuming and tedious task. While decades of research have produced automated approaches for UI test generation, these approaches typically focus on testing for crashes or…

Software Engineering · Computer Science 2022-11-03 Yixue Zhao , Saghar Talebipour , Kesina Baral , Hyojae Park , Leon Yee , Safwat Ali Khan , Yuriy Brun , Nenad Medvidovic , Kevin Moran

In this era of videos, automatic video editing techniques attract more and more attention from industry and academia since they can reduce workloads and lower the requirements for human editors. Existing automatic editing systems are mainly…

Computer Vision and Pattern Recognition · Computer Science 2024-11-08 Panwen Hu , Nan Xiao , Feifei Li , Yongquan Chen , Rui Huang

Recent breakthroughs in Vision-Language (V&L) joint research have achieved remarkable results in various text-driven tasks. High-quality Text-to-video (T2V), a task that has been long considered mission-impossible, was proven feasible with…

Artificial Intelligence · Computer Science 2022-11-28 Yuxing Qiu , Feng Gao , Minchen Li , Govind Thattai , Yin Yang , Chenfanfu Jiang

Recent Video-Language Models (VLMs) achieve promising results on long-video understanding, but their performance still lags behind that achieved on tasks involving images or short videos. This has led to great interest in improving the long…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Lars Doorenbos , Federico Spurio , Juergen Gall

Despite the advances in text-to-image synthesis, particularly with diffusion models, generating visual instructions that require consistent representation and smooth state transitions of objects across sequential steps remains a formidable…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Quynh Phung , Songwei Ge , Jia-Bin Huang

In the paradigm of AI-generated content (AIGC), there has been increasing attention to transferring knowledge from pre-trained text-to-image (T2I) models to text-to-video (T2V) generation. Despite their effectiveness, these frameworks face…

Computer Vision and Pattern Recognition · Computer Science 2024-02-07 Susung Hong , Junyoung Seo , Heeseong Shin , Sunghwan Hong , Seungryong Kim

Visual instruction tuning has become the predominant technology in eliciting the multimodal task-solving capabilities of large vision-language models (LVLMs). Despite the success, as visual instructions require images as the input, it would…

Computation and Language · Computer Science 2025-02-18 Zikang Liu , Kun Zhou , Wayne Xin Zhao , Dawei Gao , Yaliang Li , Ji-Rong Wen

We introduce the novel task of Pano2Vid $-$ automatic cinematography in panoramic 360$^{\circ}$ videos. Given a 360$^{\circ}$ video, the goal is to direct an imaginary camera to virtually capture natural-looking normal field-of-view (NFOV)…

Computer Vision and Pattern Recognition · Computer Science 2016-12-08 Yu-Chuan Su , Dinesh Jayaraman , Kristen Grauman

The rapid advancement of diffusion models has greatly improved video synthesis, especially in controllable video generation, which is vital for applications like autonomous driving. Although DiT with 3D VAE has become a standard framework…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Ruiyuan Gao , Kai Chen , Bo Xiao , Lanqing Hong , Zhenguo Li , Qiang Xu

Recent advances in self-evolution video understanding frameworks have demonstrated the potential of autonomous learning without human annotations. However, existing methods often suffer from weakly controlled optimization and uncontrolled…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Guiyi Zeng , Junqing Yu , Yi-Ping Phoebe Chen , Xu Chen , Wei Yang , Zikai Song

We propose a step-by-step video-to-audio (V2A) generation method for finer controllability over the generation process and more realistic audio synthesis. Inspired by traditional Foley workflows, our approach aims to comprehensively capture…

Computer Vision and Pattern Recognition · Computer Science 2025-10-08 Akio Hayakawa , Masato Ishii , Takashi Shibuya , Yuki Mitsufuji

In this paper, we explore the capability of an agent to construct a logical sequence of action steps, thereby assembling a strategic procedural plan. This plan is crucial for navigating from an initial visual observation to a target visual…

Computer Vision and Pattern Recognition · Computer Science 2024-06-18 Kumaranage Ravindu Yasas Nagasinghe , Honglu Zhou , Malitha Gunawardhana , Martin Renqiang Min , Daniel Harari , Muhammad Haris Khan

Online courses have significantly lowered the barrier to accessing education, yet the varying content quality of these videos poses challenges. In this work, we focus on the task of automatically evaluating the quality of video course…

Multimedia · Computer Science 2025-01-07 Xiaoxuan Zhu , Zhouhong Gu , Sihang Jiang , Zhixu Li , Hongwei Feng , Yanghua Xiao

Zero-shot action recognition, which recognizes actions in videos without having received any training examples, is gaining wide attention considering it can save labor costs and training time. Nevertheless, the performance of zero-shot…

Computer Vision and Pattern Recognition · Computer Science 2023-06-13 Nan Wu , Hiroshi Kera , Kazuhiko Kawamoto

Recent advancements in instruction-based image editing and subject-driven generation have garnered significant attention, yet both tasks still face limitations in meeting practical user needs. Instruction-based editing relies solely on…

Computer Vision and Pattern Recognition · Computer Science 2025-10-09 Bin Xia , Bohao Peng , Yuechen Zhang , Junjia Huang , Jiyang Liu , Jingyao Li , Haoru Tan , Sitong Wu , Chengyao Wang , Yitong Wang , Xinglong Wu , Bei Yu , Jiaya Jia

We propose V2CNet, a new deep learning framework to automatically translate the demonstration videos to commands that can be directly used in robotic applications. Our V2CNet has two branches and aims at understanding the demonstration…

Computer Vision and Pattern Recognition · Computer Science 2019-03-27 Anh Nguyen , Thanh-Toan Do , Ian Reid , Darwin G. Caldwell , Nikos G. Tsagarakis
‹ Prev 1 8 9 10 Next ›