中文
相关论文

相关论文: Constructing Hierarchical Q&A Datasets for Video S…

200 篇论文

Long-form video understanding remains fundamentally challenged by pervasive spatiotemporal redundancy and intricate narrative dependencies that span extended temporal horizons. While recent structured representations compress visual…

人工智能 · 计算机科学 2026-04-24 Yuehan Zhu , Jingqi Zhao , Jiawen Zhao , Xudong Mao , Baoquan Zhao

With the breakthrough of multi-modal large language models, answering complex visual questions that demand advanced reasoning abilities and world knowledge has become a much more important testbed for developing AI models than ever.…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Haibo Wang , Weifeng Ge

Multimodal counterfactual reasoning is a vital yet challenging ability for AI systems. It involves predicting the outcomes of hypothetical circumstances based on vision and language inputs, which enables AI models to learn from failures and…

计算机视觉与模式识别 · 计算机科学 2023-11-06 Te-Lin Wu , Zi-Yi Dou , Qingyuan Hu , Yu Hou , Nischal Reddy Chandra , Marjorie Freedman , Ralph M. Weischedel , Nanyun Peng

Hierarchies of concepts are useful in many applications from navigation to organization of objects. Usually, a hierarchy is created in a centralized manner by employing a group of domain experts, a time-consuming and expensive process. The…

人工智能 · 计算机科学 2015-08-04 Yuyin Sun , Adish Singla , Dieter Fox , Andreas Krause

Reasoning about causal and temporal event relations in videos is a new destination of Video Question Answering (VideoQA).The major stumbling block to achieve this purpose is the semantic gap between language and video since they are at…

计算机视觉与模式识别 · 计算机科学 2022-11-03 Shaoning Xiao , Long Chen , Kaifeng Gao , Zhao Wang , Yi Yang , Zhimeng Zhang , Jun Xiao

With the rapid advancement of video generation models such as Sora, video quality assessment (VQA) is becoming increasingly crucial for selecting high-quality videos from large-scale datasets used in pre-training. Traditional VQA methods,…

计算机视觉与模式识别 · 计算机科学 2025-09-16 Yanyun Pu , Kehan Li , Zeyi Huang , Zhijie Zhong , Kaixiang Yang

This paper proposes CQ-VQA, a novel 2-level hierarchical but end-to-end model to solve the task of visual question answering (VQA). The first level of CQ-VQA, referred to as question categorizer (QC), classifies questions to reduce the…

计算机视觉与模式识别 · 计算机科学 2020-02-18 Aakansha Mishra , Ashish Anand , Prithwijit Guha

Integrating vision and language has long been a dream in work on artificial intelligence (AI). In the past two years, we have witnessed an explosion of work that brings together vision and language from images to videos and beyond. The…

Visual interactivity understanding within visual scenes presents a significant challenge in computer vision. Existing methods focus on complex interactivities while leveraging a simple relationship model. These methods, however, struggle…

计算机视觉与模式识别 · 计算机科学 2024-03-18 Trong-Thuan Nguyen , Pha Nguyen , Khoa Luu

Visual Question Answering (VQA) is a challenging task that requires cross-modal understanding and reasoning of visual image and natural language question. To inspect the association of VQA models to human cognition, we designed a survey to…

计算机视觉与模式识别 · 计算机科学 2023-10-05 Liben Chen , Long Chen , Tian Ellison-Chen , Zhuoyuan Xu

Question-answering (QA) on video contents is a significant challenge for achieving human-level intelligence as it involves both vision and language in real-world settings. Here we demonstrate the possibility of an AI agent performing video…

计算机视觉与模式识别 · 计算机科学 2017-07-05 Kyung-Min Kim , Min-Oh Heo , Seong-Ho Choi , Byoung-Tak Zhang

Video understanding has achieved great success in representation learning, such as video caption, video object grounding, and video descriptive question-answer. However, current methods still struggle on video reasoning, including evidence…

计算机视觉与模式识别 · 计算机科学 2022-05-31 Jiangtong Li , Li Niu , Liqing Zhang

Inspired by the remarkable advances in video analytics, research teams are stepping towards a greater ambition -- movie understanding. However, compared to those activity videos in conventional datasets, movies are significantly different.…

计算机视觉与模式识别 · 计算机科学 2019-10-25 Yu Xiong , Qingqiu Huang , Lingfeng Guo , Hang Zhou , Bolei Zhou , Dahua Lin

Explanation and high-order reasoning capabilities are crucial for real-world visual question answering with diverse levels of inference complexity (e.g., what is the dog that is near the girl playing with?) and important for users to…

计算机视觉与模式识别 · 计算机科学 2019-09-24 Qingxing Cao , Bailin Li , Xiaodan Liang , Liang Lin

This work proposes TimeChat, a time-sensitive multimodal large language model specifically designed for long video understanding. Our model incorporates two key architectural contributions: (1) a timestamp-aware frame encoder that binds…

计算机视觉与模式识别 · 计算机科学 2024-03-29 Shuhuai Ren , Linli Yao , Shicheng Li , Xu Sun , Lu Hou

Learned visual representations often capture large amounts of semantic information for accurate downstream applications. Human understanding of the world is fundamentally grounded in hierarchy. To mimic this and further improve…

计算机视觉与模式识别 · 计算机科学 2023-11-27 Ethan Shen , Ali Farhadi , Aditya Kusupati

The feasibility of autonomous artificial thinking systems needs to compare the way the human beings acquire their information and develops the thought with the current capacities of the autonomous information systems. Our model uses four…

人工智能 · 计算机科学 2020-01-14 Joël Colloc

The recent development of artificial intelligence enables a machine to achieve a human level of intelligence. Problem-solving and decision-making are two mental abilities to measure human intelligence. Many scholars have proposed different…

人工智能 · 计算机科学 2023-02-23 Caesar Wu , Pascal Bouvry

Video question answering (VideoQA) is a challenging task that requires integrating spatial, temporal, and semantic information to capture the complex dynamics of video sequences. Although recent advances have introduced various approaches…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Zhongyu Yang , Zuhao Yang , Shuo Zhan , Tan Yue , Wei Pang , Yingfang Yuan

Video descriptions are crucial for blind and low vision (BLV) users to access visual content. However, current artificial intelligence models for generating descriptions often fall short due to limitations in the quality of human…

计算机视觉与模式识别 · 计算机科学 2025-03-03 Chaoyu Li , Sid Padmanabhuni , Maryam Cheema , Hasti Seifi , Pooyan Fazli