中文
相关论文

相关论文: Multimedia Generative Script Learning for Task Pla…

200 篇论文

Despite rapid advances in video generative models, robust metrics for evaluating visual and temporal correctness of complex human actions remain elusive. Critically, existing pure-vision encoders and Multimodal Large Language Models (MLLMs)…

计算机视觉与模式识别 · 计算机科学 2025-12-04 Xavier Thomas , Youngsun Lim , Ananya Srinivasan , Audrey Zheng , Deepti Ghadiyaram

People get informed of a daily task plan through diverse media involving both texts and images. However, most prior research only focuses on LLM's capability of textual plan generation. The potential of large-scale models in providing…

计算机视觉与模式识别 · 计算机科学 2025-06-16 Xiaoxin Lu , Ranran Haoran Zhang , Yusen Zhang , Rui Zhang

Recent advancements in video generation have enabled the development of ``world models'' capable of simulating potential futures for robotics and planning. However, specifying precise goals for these models remains a challenge; text…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Nate Gillman , Yinghua Zhou , Zitian Tang , Evan Luo , Arjan Chakravarthy , Daksh Aggarwal , Michael Freeman , Charles Herrmann , Chen Sun

Training multimodal large language models (MLLMs) for video understanding requires large-scale annotated data spanning diverse tasks such as object counting, question answering, and segmentation. However, collecting and annotating…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Tanzila Rahman , Renjie Liao , Leonid Sigal

We introduce Goal-Conditioned Visual Navigation Instruction Generation (GoViG), a new task that aims to generate contextually coherent navigation instructions solely from egocentric visual observations of initial and goal states. Unlike…

计算机视觉与模式识别 · 计算机科学 2026-04-30 Fengyi Wu , Yifei Dong , Yilong Dai , Guangyu Chen , Qifeng Wu , Huiting Huang , Hang Wang , Qi Dai , Alexander G. Hauptmann , Zhi-Qi Cheng

Text-driven human motion generation, as one of the vital tasks in computer-aided content creation, has recently attracted increasing attention. While pioneering research has largely focused on improving numerical performance metrics on…

计算机视觉与模式识别 · 计算机科学 2024-05-27 Yunyao Mao , Xiaoyang Liu , Wengang Zhou , Zhenbo Lu , Houqiang Li

Continual learning enables pre-trained generative vision-language models (VLMs) to incorporate knowledge from new tasks without retraining data from previous ones. Recent methods update a visual projector to translate visual information for…

计算机视觉与模式识别 · 计算机科学 2025-08-04 Hyundong Jin , Hyung Jin Chang , Eunwoo Kim

Generating coherent and useful image/video scenes from a free-form textual description is technically a very difficult problem to handle. Textual description of the same scene can vary greatly from person to person, or sometimes even for…

计算机视觉与模式识别 · 计算机科学 2020-12-01 Faria Huq , Nafees Ahmed , Anindya Iqbal

While recent advancements in generative models have achieved remarkable visual fidelity in video synthesis, creating coherent multi-shot narratives remains a significant challenge. To address this, keyframe-based approaches have emerged as…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Peixuan Zhang , Zijian Jia , Kaiqi Liu , Shuchen Weng , Si Li , Boxin Shi

Current work on image-based story generation suffers from the fact that the existing image sequence collections do not have coherent plots behind them. We improve visual story generation by producing a new image-grounded dataset, Visual…

计算与语言 · 计算机科学 2023-01-23 Xudong Hong , Asad Sayeed , Khushboo Mehra , Vera Demberg , Bernt Schiele

Access to high-quality education at scale is limited by the difficulty of providing student feedback on open-ended assignments in structured domains like computer programming, graphics, and short response questions. This problem has proven…

机器学习 · 计算机科学 2021-03-25 Ali Malik , Mike Wu , Vrinda Vasavada , Jinpeng Song , Madison Coots , John Mitchell , Noah Goodman , Chris Piech

The core of video understanding tasks, such as recognition, captioning, and tracking, is to automatically detect objects or actions in a video and analyze their temporal evolution. Despite sharing a common goal, different tasks often rely…

计算机视觉与模式识别 · 计算机科学 2024-03-28 Junke Wang , Dongdong Chen , Chong Luo , Bo He , Lu Yuan , Zuxuan Wu , Yu-Gang Jiang

Recent video and language pretraining frameworks lack the ability to generate sentences. We present Multimodal Video Generative Pretraining (MV-GPT), a new pretraining framework for learning from unlabelled videos which can be effectively…

计算机视觉与模式识别 · 计算机科学 2022-05-11 Paul Hongsuck Seo , Arsha Nagrani , Anurag Arnab , Cordelia Schmid

Emerging immersive display technologies efficiently utilize resources with perceptual graphics methods such as foveated rendering and denoising. Running multiple perceptual graphics methods challenges devices with limited power and…

计算机视觉与模式识别 · 计算机科学 2025-08-08 Doğa Yılmaz , He Wang , Towaki Takikawa , Duygu Ceylan , Kaan Akşit

Given the enormous number of instructional videos available online, learning a diverse array of multi-step task models from videos is an appealing goal. We introduce a new pre-trained video model, VideoTaskformer, focused on representing…

计算机视觉与模式识别 · 计算机科学 2023-03-24 Medhini Narasimhan , Licheng Yu , Sean Bell , Ning Zhang , Trevor Darrell

Creating meaningful visual narratives through human-AI collaboration requires understanding how text-image intertextuality emerges when textual intentions meet AI-generated visuals. We conducted a three-phase qualitative study with 15…

人机交互 · 计算机科学 2025-11-06 Mengyao Guo , Kexin Nie , Ze Gao , Black Sun , Xueyang Wang , Jinda Han , Xingting Wu

The performance of computer vision models in certain real-world applications (e.g., rare wildlife observation) is limited by the small number of available images. Expanding datasets using pre-trained generative models is an effective way to…

计算机视觉与模式识别 · 计算机科学 2024-12-25 Changjian Chen , Fei Lv , Yalong Guan , Pengcheng Wang , Shengjie Yu , Yifan Zhang , Zhuo Tang

Soft object manipulation tasks in domestic scenes pose a significant challenge for existing robotic skill learning techniques due to their complex dynamics and variable shape characteristics. Since learning new manipulation skills from…

机器人学 · 计算机科学 2023-09-06 Junjia Liu , Zhihao Li , Wanyu Lin , Sylvain Calinon , Kay Chen Tan , Fei Chen

We propose a hierarchically structured reinforcement learning approach to address the challenges of planning for generating coherent multi-sentence stories for the visual storytelling task. Within our framework, the task of generating a…

计算机视觉与模式识别 · 计算机科学 2019-01-21 Qiuyuan Huang , Zhe Gan , Asli Celikyilmaz , Dapeng Wu , Jianfeng Wang , Xiaodong He

Visual tracking is typically solved as a discriminative learning problem that usually requires high-quality samples for online model adaptation. It is a critical and challenging problem to evaluate the training samples collected from…

计算机视觉与模式识别 · 计算机科学 2020-04-02 Weichao Li , Xi Li , Omar Elfarouk Bourahla , Fuxian Huang , Fei Wu , Wei Liu , Zhiheng Wang , Hongmin Liu