中文
相关论文

相关论文: Learning to Generate Long-term Future Narrations D…

200 篇论文

In this report, we describe the technical details of our approach for the Ego4D Long-Term Action Anticipation Challenge 2023. The aim of this task is to predict a sequence of future actions that will take place at an arbitrary time or…

计算机视觉与模式识别 · 计算机科学 2023-07-06 Tatsuya Ishibashi , Kosuke Ono , Noriyuki Kugo , Yuji Sato

While language models have become impactful in many real-world applications, video generation remains largely confined to entertainment. Motivated by video's inherent capacity to demonstrate physical-world information that is difficult to…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Junhao Cheng , Liang Hou , Xin Tao , Jing Liao

The rapid increase in video content production has resulted in enormous data volumes, creating significant challenges for efficient analysis and resource management. To address this, robust video analysis tools are essential. This paper…

计算机视觉与模式识别 · 计算机科学 2025-01-07 Ulindu De Silva , Leon Fernando , Billy Lau Pik Lik , Zann Koh , Sam Conrad Joyce , Belinda Yuen , Chau Yuen

We present "Narrative Weaver", a novel framework that addresses a fundamental challenge in generative AI: achieving multi-modal controllable, long-range, and consistent visual content generation. While existing models excel at generating…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Zhengjian Yao , Yongzhi Li , Xinyuan Gao , Quan Chen , Peng Jiang , Yanye Lu

In egocentric scenarios, anticipating both the next action and its visual outcome is essential for understanding human-object interactions and for enabling robotic planning. However, existing paradigms fall short of jointly modeling these…

计算机视觉与模式识别 · 计算机科学 2025-08-29 Binjie Zhang , Mike Zheng Shou

Research on text generation from multimodal inputs has largely focused on static images, and less on video data. In this paper, we propose a new task, narration generation, that is complementing videos with narration texts that are to be…

计算与语言 · 计算机科学 2021-01-19 Nikos Papasarantopoulos , Shay B. Cohen

Generalization in robot manipulation is essential for deploying robots in open-world environments and advancing toward artificial general intelligence. While recent Vision-Language-Action (VLA) models leverage large pre-trained…

机器人学 · 计算机科学 2025-12-09 Yichao Shen , Fangyun Wei , Zhiying Du , Yaobo Liang , Yan Lu , Jiaolong Yang , Nanning Zheng , Baining Guo

Automatically reasoning about future human behaviors is a difficult problem but has significant practical applications to assistive systems. Part of this difficulty stems from learning systems' inability to represent all kinds of behaviors.…

计算机视觉与模式识别 · 计算机科学 2020-04-07 Jiaqi Guan , Ye Yuan , Kris M. Kitani , Nicholas Rhinehart

Human activities generate various event sequences such as taxi trip records, bike-sharing pick-ups, crime occurrence, and infectious disease transmission. The point process is widely used in many applications to predict such events related…

Both text and video data are abundant on the internet and support large-scale self-supervised learning through next token or frame prediction. However, they have not been equally leveraged: language models have had significant real-world…

计算机视觉与模式识别 · 计算机科学 2024-02-28 Sherry Yang , Jacob Walker , Jack Parker-Holder , Yilun Du , Jake Bruce , Andre Barreto , Pieter Abbeel , Dale Schuurmans

Narrative visualization transforms data into engaging stories, making complex information accessible to a broad audience. Foundation models, with their advanced capabilities such as natural language processing, content generation, and…

人机交互 · 计算机科学 2025-02-14 Yi He , Ke Xu , Shixiong Cao , Yang Shi , Qing Chen , Nan Cao

In this work, we introduce (a) the new problem of anticipating object state changes in images and videos during procedural activities, (b) new curated annotation data for object state change classification based on the Ego4D dataset, and…

计算机视觉与模式识别 · 计算机科学 2024-12-03 Victoria Manousaki , Konstantinos Bacharidis , Filippos Gouidis , Konstantinos Papoutsakis , Dimitris Plexousakis , Antonis Argyros

Rapidly creating effective visualizations using expressive grammars is challenging for users who have limited time and limited skills in statistics and data visualization. Even high-level, dedicated visualization tools often require users…

人机交互 · 计算机科学 2018-11-06 Victor Dibia , Çağatay Demiralp

Short-term action anticipation (STA) in first-person videos is a challenging task that involves understanding the next active object interactions and predicting future actions. Existing action anticipation methods have primarily focused on…

计算机视觉与模式识别 · 计算机科学 2023-06-26 Sanket Thakur , Cigdem Beyan , Pietro Morerio , Vittorio Murino , Alessio Del Bue

Story generation is an important natural language processing task that aims to generate coherent stories automatically. While the use of neural networks has proven effective in improving story generation, how to learn to generate an…

计算与语言 · 计算机科学 2019-12-09 Gang Chen , Yang Liu , Huanbo Luan , Meng Zhang , Qun Liu , Maosong Sun

Visual Planning for Assistance (VPA) aims to predict a sequence of user actions required to achieve a specified goal based on a video showing the user's progress. Although recent advances in multimodal large language models (MLLMs) have…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Ce Zhang , Yale Song , Ruta Desai , Michael Louis Iuzzolino , Joseph Tighe , Gedas Bertasius , Satwik Kottur

Analyzing human actions in videos has gained increased attention recently. While most works focus on classifying and labeling observed video frames or anticipating the very recent future, making long-term predictions over more than just a…

计算机视觉与模式识别 · 计算机科学 2018-04-04 Yazan Abu Farha , Alexander Richard , Juergen Gall

Recent video generation models have shown promising results in producing high-quality video clips lasting several seconds. However, these models face challenges in generating long sequences that convey clear and informative events, limiting…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Junfei Xiao , Feng Cheng , Lu Qi , Liangke Gui , Jiepeng Cen , Zhibei Ma , Alan Yuille , Lu Jiang

Predictive foresight is important to intelligent embodied agents. Since the motor execution of a robot is intrinsically constrained by its visual perception of environmental geometry, effectively anticipating the future requires capturing…

机器人学 · 计算机科学 2026-03-12 Xiaoxu Xu , Hao Li , Jinhui Ye , Yilun Chen , Jia Zeng , Xinyi Chen , Linning Xu , Dahua Lin , Weixin Li , Jiangmiao Pang

We introduce LaViLa, a new approach to learning video-language representations by leveraging Large Language Models (LLMs). We repurpose pre-trained LLMs to be conditioned on visual input, and finetune them to create automatic video…

计算机视觉与模式识别 · 计算机科学 2022-12-09 Yue Zhao , Ishan Misra , Philipp Krähenbühl , Rohit Girdhar