English
Related papers

Related papers: Zero-Shot Video Question Answering with Procedural…

200 papers

In this paper, we initiate an attempt of developing an end-to-end chat-centric video understanding system, coined as VideoChat. It integrates video foundation models and large language models via a learnable neural interface, excelling in…

Computer Vision and Pattern Recognition · Computer Science 2024-01-05 KunChang Li , Yinan He , Yi Wang , Yizhuo Li , Wenhai Wang , Ping Luo , Yali Wang , Limin Wang , Yu Qiao

Video action localization aims to find the timings of specific actions from a long video. Although existing learning-based approaches have been successful, they require annotating videos, which comes with a considerable labor cost. This…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Naoki Wake , Atsushi Kanehira , Kazuhiro Sasabuchi , Jun Takamatsu , Katsushi Ikeuchi

In traditional Visual Question Generation (VQG), most images have multiple concepts (e.g. objects and categories) for which a question could be generated, but models are trained to mimic an arbitrary choice of concept as given in their…

Machine Learning · Computer Science 2022-07-27 Nihir Vedd , Zixu Wang , Marek Rei , Yishu Miao , Lucia Specia

The increase in the availability of online videos has transformed the way we access information and knowledge. A growing number of individuals now prefer instructional videos as they offer a series of step-by-step procedures to accomplish…

Computation and Language · Computer Science 2023-09-22 Deepak Gupta , Kush Attal , Dina Demner-Fushman

We propose V-Doc, a question-answering tool using document images and PDF, mainly for researchers and general non-deep learning experts looking to generate, process, and understand the document visual question answering tasks. The V-Doc…

Artificial Intelligence · Computer Science 2022-06-01 Yihao Ding , Zhe Huang , Runlin Wang , Yanhang Zhang , Xianru Chen , Yuzhong Ma , Hyunsuk Chung , Soyeon Caren Han

Can we teach a robot to recognize and make predictions for activities that it has never seen before? We tackle this problem by learning models for video from text. This paper presents a hierarchical model that generalizes instructional…

Computer Vision and Pattern Recognition · Computer Science 2022-11-08 Fadime Sener , Rishabh Saraf , Angela Yao

We propose to perform video question answering (VideoQA) in a Contrastive manner via a Video Graph Transformer model (CoVGT). CoVGT's uniqueness and superiority are three-fold: 1) It proposes a dynamic graph transformer module which encodes…

Computer Vision and Pattern Recognition · Computer Science 2023-07-12 Junbin Xiao , Pan Zhou , Angela Yao , Yicong Li , Richang Hong , Shuicheng Yan , Tat-Seng Chua

Skilled human interviewers can extract valuable information from experts. This raises a fundamental question: what makes some questions more effective than others? To address this, a quantitative evaluation of question-generation models is…

Computer Vision and Pattern Recognition · Computer Science 2025-12-18 Huaying Zhang , Atsushi Hashimoto , Tosho Hirasawa

The Long-form Video Question-Answering task requires the comprehension and analysis of extended video content to respond accurately to questions by utilizing both temporal and contextual information. In this paper, we present…

Computer Vision and Pattern Recognition · Computer Science 2024-06-26 Yongliang Wu , Bozheng Li , Jiawang Cao , Wenbo Zhu , Yi Lu , Weiheng Chi , Chuyun Xie , Haolin Zheng , Ziyue Su , Jay Wu , Xu Yang

The next frontier for video generation lies in developing models capable of zero-shot reasoning, where understanding real-world scientific laws is crucial for accurate physical outcome modeling under diverse conditions. However, existing…

Computer Vision and Pattern Recognition · Computer Science 2025-12-03 Lanxiang Hu , Abhilash Shankarampeta , Yixin Huang , Zilin Dai , Haoyang Yu , Yujie Zhao , Haoqiang Kang , Daniel Zhao , Tajana Rosing , Hao Zhang

Existing methods for video question answering (VideoQA) often suffer from spurious correlations between different modalities, leading to a failure in identifying the dominant visual evidence and the intended question. Moreover, these…

Computer Vision and Pattern Recognition · Computer Science 2023-08-02 Yushen Wei , Yang Liu , Hong Yan , Guanbin Li , Liang Lin

Deep neural networks facilitate video question answering (VideoQA), but the real-world applications on video streams such as CCTV and live cast place higher demands on the solver. To address the challenges of VideoQA on long videos of…

Multimedia · Computer Science 2023-03-08 Weikai Kong , Shuhong Ye , Chenglin Yao , Jianfeng Ren

Video Question Answering (VideoQA) has made significant strides by leveraging multimodal learning to align visual and textual modalities. However, current benchmarks overwhelmingly focus on questions answerable through explicit visual…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Sirnam Swetha , Rohit Gupta , Parth Parag Kulkarni , David G Shatwell , Jeffrey A Chan Santiago , Nyle Siddiqui , Joseph Fioresi , Mubarak Shah

We propose a novel architecture design for video prediction in order to utilize procedural domain knowledge directly as part of the computational graph of data-driven models. On the basis of new challenging scenarios we show that…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Patrick Takenaka , Johannes Maucher , Marco F. Huber

When robots perform long action sequences, users will want to easily and reliably find out what they have done. We therefore demonstrate the task of learning to summarize and answer questions about a robot agent's past actions using natural…

Robotics · Computer Science 2023-06-19 Chad DeChant , Iretiayo Akinola , Daniel Bauer

Significant progress has been made in the field of video question answering (VideoQA) thanks to deep learning and large-scale pretraining. Despite the presence of sophisticated model structures and powerful video-text foundation models,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-03 Haopeng Li , Tom Drummond , Mingming Gong , Mohammed Bennamoun , Qiuhong Ke

Instructional videos provide detailed how-to guides for various tasks, with viewers often posing questions regarding the content. Addressing these questions is vital for comprehending the content, yet receiving immediate answers is…

Computer Vision and Pattern Recognition · Computer Science 2024-02-01 Saelyne Yang , Sunghyun Park , Yunseok Jang , Moontae Lee

Medical visual question answering (Med-VQA) is a machine learning task that aims to create a system that can answer natural language questions based on given medical images. Although there has been rapid progress on the general VQA task,…

Computer Vision and Pattern Recognition · Computer Science 2023-09-21 Louisa Canepa , Sonit Singh , Arcot Sowmya

Video Corpus Visual Answer Localization (VCVAL) includes question-related video retrieval and visual answer localization in the videos. Specifically, we use text-to-text retrieval to find relevant videos for a medical question based on the…

Computer Vision and Pattern Recognition · Computer Science 2024-12-23 Jiaxin Wu , Yiyang Jiang , Xiao-Yong Wei , Qing Li

Despite rapid advancements in video generation models, aligning their outputs with complex user intent remains challenging. Existing test-time optimization methods are typically either computationally expensive or require white-box access…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Yiwen Song , Tomas Pfister , Yale Song
‹ Prev 1 3 4 5 6 7 10 Next ›