中文
相关论文

相关论文: MiraData: A Large-Scale Video Dataset with Long Du…

200 篇论文

The December 2024 release of OpenAI's Sora, a powerful video generation model driven by natural language prompts, highlights a growing convergence between large language models (LLMs) and video synthesis. As these multimodal systems evolve…

计算机视觉与模式识别 · 计算机科学 2025-05-01 Misora Sugiyama , Hirokatsu Kataoka

Tracking dense 3D motion from monocular videos remains challenging, particularly when aiming for pixel-level precision over long sequences. We introduce DELTA, a novel method that efficiently tracks every pixel in 3D space, enabling…

计算机视觉与模式识别 · 计算机科学 2025-03-03 Tuan Duc Ngo , Peiye Zhuang , Chuang Gan , Evangelos Kalogerakis , Sergey Tulyakov , Hsin-Ying Lee , Chaoyang Wang

Recent video generation models have shown promising results in producing high-quality video clips lasting several seconds. However, these models face challenges in generating long sequences that convey clear and informative events, limiting…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Junfei Xiao , Feng Cheng , Lu Qi , Liangke Gui , Jiepeng Cen , Zhibei Ma , Alan Yuille , Lu Jiang

We study generative super-resolution (SR) in real-world scenarios where content and degradations vary across domains, genres, and segments. For example, images and videos may alternate between text overlays, fast motion, smooth cartoons,…

计算机视觉与模式识别 · 计算机科学 2026-04-30 Jiaqi Guo , Mingzhen Li , Haohong Wang , Aggelos K. Katsaggelos

Existing video generation models predominantly emphasize appearance fidelity while exhibiting limited ability to synthesize complex human motions, such as whole-body movements, long-range dynamics, and fine-grained human-environment…

计算机视觉与模式识别 · 计算机科学 2026-02-25 Haoyu Wang , Hao Tang , Donglin Di , Zhilu Zhang , Wangmeng Zuo , Feng Gao , Siwei Ma , Shiliang Zhang

Video description is the automatic generation of natural language sentences that describe the contents of a given video. It has applications in human-robot interaction, helping the visually impaired and video subtitling. The past few years…

计算机视觉与模式识别 · 计算机科学 2020-03-04 Nayyer Aafaq , Ajmal Mian , Wei Liu , Syed Zulqarnain Gilani , Mubarak Shah

We introduce \textbf{LongInsightBench}, the first benchmark designed to assess models' ability to understand long videos, with a focus on human language, viewpoints, actions, and other contextual elements, while integrating \textbf{visual,…

计算机视觉与模式识别 · 计算机科学 2025-10-22 ZhaoYang Han , Qihan Lin , Hao Liang , Bowen Chen , Zhou Liu , Wentao Zhang

Despite the considerable progress achieved in the long video generation problem, there is still significant room to improve the consistency of the generated videos, particularly in terms of their smoothness and transitions between scenes.…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Xingyao Li , Fengzhuo Zhang , Jiachun Pan , Yunlong Hou , Vincent Y. F. Tan , Zhuoran Yang

Action-conditioned world models are increasingly used as scalable simulators for robot learning, yet current evaluations provide limited evidence that their predictions are reliable under the actions they condition on. Existing benchmarks…

人工智能 · 计算机科学 2026-05-29 Tianzhuo Yang , Zihan Shen , Zirui Mi , Zhaoyi Zhang , Jiayi Zhou , Jiaming Ji , Juntao Dai , Jiawei Chen , Boyuan Chen , Yaodong Yang

The recent development of Sora leads to a new era in text-to-video (T2V) generation. Along with this comes the rising concern about its security risks. The generated videos may contain illegal or unethical content, and there is a lack of…

计算机视觉与模式识别 · 计算机科学 2024-09-10 Yibo Miao , Yifan Zhu , Yinpeng Dong , Lijia Yu , Jun Zhu , Xiao-Shan Gao

Omnidirectional or 360-degree video is being increasingly deployed, largely due to the latest advancements in immersive virtual reality (VR) and extended reality (XR) technology. However, the adoption of these videos in streaming encounters…

图像与视频处理 · 电气工程与系统科学 2024-03-08 Ahmed Telili , Ibrahim Farhat , Wassim Hamidouche , Hadi Amirpour

Significant advancements have been made in video generative models recently. Unlike image generation, video generation presents greater challenges, requiring not only generating high-quality frames but also ensuring temporal consistency…

计算机视觉与模式识别 · 计算机科学 2024-07-24 Jiahe Liu , Youran Qu , Qi Yan , Xiaohui Zeng , Lele Wang , Renjie Liao

Real-world user-generated short videos, especially those distributed on platforms such as WeChat Channel and TikTok, dominate the mobile internet. However, current large multimodal models lack essential temporally-structured, detailed, and…

计算机视觉与模式识别 · 计算机科学 2025-07-29 Yuying Ge , Yixiao Ge , Chen Li , Teng Wang , Junfu Pu , Yizhuo Li , Lu Qiu , Jin Ma , Lisheng Duan , Xinyu Zuo , Jinwen Luo , Weibo Gu , Zexuan Li , Xiaojing Zhang , Yangyu Tao , Han Hu , Di Wang , Ying Shan

Video storytelling is engaging multimedia content that utilizes video and its accompanying narration to attract the audience, where a key challenge is creating narrations for recorded visual scenes. Previous studies on dense video…

多媒体 · 计算机科学 2024-12-31 Dingyi Yang , Chunru Zhan , Ziheng Wang , Biao Wang , Tiezheng Ge , Bo Zheng , Qin Jin

Videos are more informative than images because they capture the dynamics of the scene. By representing motion in videos, we can capture dynamic activities. In this work, we introduce GPT-4 generated motion descriptions that capture…

计算机视觉与模式识别 · 计算机科学 2024-06-10 Chinmaya Devaraj , Cornelia Fermuller , Yiannis Aloimonos

Recent advances in text-to-video generation have achieved impressive performance on short clips, yet evaluating long-form generation under complex textual inputs remains a significant challenge. In response to this challenge, we present…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Xiangqing Zheng , Chengyue Wu , Kehai Chen , Min Zhang

The availability of high definition video content on the web has brought about a significant change in the characteristics of Internet video, but not many studies on characterizing video have been done after this change. Video…

多媒体 · 计算机科学 2014-08-26 Saba Ahsan , Varun Singh , Jörg Ott

Success in generative modeling across language, image, and video demonstrates that large, well-curated datasets are the key driver for building capable models. 3D Human motion, however, has lagged behind, constrained by an unsatisfying…

The advent of next-generation video generation models like \textit{Sora} poses challenges for AI-generated content (AIGC) video quality assessment (VQA). These models substantially mitigate flickering artifacts prevalent in prior models,…

计算机视觉与模式识别 · 计算机科学 2025-02-07 Shangkun Sun , Xiaoyu Liang , Bowen Qu , Wei Gao

Despite the promise of synthesizing high-fidelity videos, Diffusion Transformers (DiTs) with 3D full attention suffer from expensive inference due to the complexity of attention computation and numerous sampling steps. For example, the…

计算机视觉与模式识别 · 计算机科学 2025-02-18 Hangliang Ding , Dacheng Li , Runlong Su , Peiyuan Zhang , Zhijie Deng , Ion Stoica , Hao Zhang