中文
相关论文

相关论文: Learning to Generate Long-term Future Narrations D…

200 篇论文

Bridging vision and natural language is a longstanding goal in computer vision and multimedia research. While earlier works focus on generating a single-sentence description for visual content, recent works have studied paragraph…

多媒体 · 计算机科学 2020-05-15 Junnan Li , Yongkang Wong , Qi Zhao , Mohan S. Kankanhalli

Video storytelling is engaging multimedia content that utilizes video and its accompanying narration to attract the audience, where a key challenge is creating narrations for recorded visual scenes. Previous studies on dense video…

多媒体 · 计算机科学 2024-12-31 Dingyi Yang , Chunru Zhan , Ziheng Wang , Biao Wang , Tiezheng Ge , Bo Zheng , Qin Jin

Prospection, the act of predicting the consequences of many possible futures, is intrinsic to human planning and action, and may even be at the root of consciousness. Surprisingly, this idea has been explored comparatively little in…

机器人学 · 计算机科学 2018-04-03 Chris Paxton , Yotam Barnoy , Kapil Katyal , Raman Arora , Gregory D. Hager

Vision-language-action (VLA) models have advanced robot manipulation through large-scale pretraining, but real-world deployment remains challenging due to partial observability and delayed feedback. Reinforcement learning addresses this via…

Creating engaging narratives from visual data is crucial for automated digital media consumption, assistive technologies, and interactive entertainment. This survey covers methodologies used in the generation of these narratives, focusing…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Daniel A. P. Oliveira , Eugénio Ribeiro , David Martins de Matos

In recent years, a substantial body of work in visually grounded natural language processing has focused on real-life multimodal scenarios such as describing content depicted in images or videos. However, comparatively less attention has…

计算与语言 · 计算机科学 2025-08-21 Aditya K Surikuchi , Raquel Fernández , Sandro Pezzelle

Human daily activities can be concisely narrated as sequences of routine events (e.g., turning off an alarm) in video streams, forming an event vocabulary. Motivated by this, we introduce VLog, a novel video understanding framework that…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Kevin Qinghong Lin , Mike Zheng Shou

We present Palm, a solution to the Long-Term Action Anticipation (LTA) task utilizing vision-language and large language models. Given an input video with annotated action periods, the LTA task aims to predict possible future actions. We…

计算机视觉与模式识别 · 计算机科学 2023-06-30 Daoji Huang , Otmar Hilliges , Luc Van Gool , Xi Wang

Anticipating future activities in video is a task with many practical applications. While earlier approaches are limited to just a few seconds in the future, the prediction time horizon has just recently been extended to several minutes in…

计算机视觉与模式识别 · 计算机科学 2019-08-30 Yazan Abu Farha , Juergen Gall

With the rapid development of foundation video generation technologies, long video generation models have exhibited promising research potential thanks to expanded content creation space. Recent studies reveal that the goal of long video…

计算机视觉与模式识别 · 计算机科学 2026-03-09 X. Feng , H. Yu , M. Wu , S. Hu , J. Chen , C. Zhu , J. Wu , X. Chu , K. Huang

Physical world knowledge resides mainly in videos. Equipping Vision-Language-Action (VLA) models with such knowledge is fundamental for safe and generalizable planning. Predictive world modeling enables VLA to internalize physical dynamics…

In this report, we present a novel three-stage framework developed for the Ego4D Long-Term Action Anticipation (LTA) task. Inspired by recent advances in foundation models, our method consists of three stages: feature extraction, action…

计算机视觉与模式识别 · 计算机科学 2025-06-12 Qiaohui Chu , Haoyu Zhang , Yisen Feng , Meng Liu , Weili Guan , Yaowei Wang , Liqiang Nie

Foreseeing the future is one of the key factors of intelligence. It involves understanding of the past and current environment as well as decent experience of its possible dynamics. In this work, we address future prediction at the abstract…

计算机视觉与模式识别 · 计算机科学 2019-05-16 Mohammadreza Zolfaghari , Özgün Çiçek , Syed Mohsin Ali , Farzaneh Mahdisoltani , Can Zhang , Thomas Brox

Understanding human activity is a crucial yet intricate task in egocentric vision, a field that focuses on capturing visual perspectives from the camera wearer's viewpoint. Traditional methods heavily rely on representation learning that is…

计算机视觉与模式识别 · 计算机科学 2024-07-19 Sanghwan Kim , Daoji Huang , Yongqin Xian , Otmar Hilliges , Luc Van Gool , Xi Wang

Leveraging future observation modeling to facilitate action generation presents a promising avenue for enhancing the capabilities of Vision-Language-Action (VLA) models. However, existing approaches struggle to strike a balance between…

Inspired by human neurological structures for action anticipation, we present an action anticipation model that enables the prediction of plausible future actions by forecasting both the visual and temporal future. In contrast to current…

计算机视觉与模式识别 · 计算机科学 2019-12-17 Harshala Gammulle , Simon Denman , Sridha Sridharan , Clinton Fookes

Video-based long-term action anticipation is crucial for early risk detection in areas such as automated driving and robotics. Conventional approaches extract features from past actions using encoders and predict future events with…

计算机视觉与模式识别 · 计算机科学 2025-08-04 Yuji Sato , Yasunori Ishii , Takayoshi Yamashita

Recent advances in deep learning research, such as transformers, have bolstered the ability for automated agents to generate creative texts similar to those that a human would write. By default, transformer decoders can only generate new…

计算与语言 · 计算机科学 2022-12-21 Brian D. Zimmerman , Gaurav Sahu , Olga Vechtomova

Dense video captioning aims to identify the events of interest in an input video, and generate descriptive captions for each event. Previous approaches usually follow a two-stage generative process, which first proposes a segment for each…

计算机视觉与模式识别 · 计算机科学 2022-09-19 Wanrong Zhu , Bo Pang , Ashish V. Thapliyal , William Yang Wang , Radu Soricut

Predicting future events is an important activity with applications across multiple fields and domains. For example, the capacity to foresee stock market trends, natural disasters, business developments, or political events can facilitate…

计算与语言 · 计算机科学 2025-01-13 Petraq Nako , Adam Jatowt
‹ 上一页 1 2 3 10 下一页 ›