中文
相关论文

相关论文: Tiger200K: Manually Curated High Visual Quality Vi…

200 篇论文

Demand for 2K video synthesis is rising with increasing consumer expectations for ultra-clear visuals. While diffusion transformers (DiTs) have demonstrated remarkable capabilities in high-quality video generation, scaling them to 2K…

计算机视觉与模式识别 · 计算机科学 2025-04-22 Jingjing Ren , Wenbo Li , Zhongdao Wang , Haoze Sun , Bangzhen Liu , Haoyu Chen , Jiaqi Xu , Aoxue Li , Shifeng Zhang , Bin Shao , Yong Guo , Lei Zhu

This paper presents an overview of the NTIRE 2025 Challenge on UGC Video Enhancement. The challenge constructed a set of 150 user-generated content videos without reference ground truth, which suffer from real-world degradations such as…

Encoding video content into compact latent tokens has become a fundamental step in video generation and understanding, driven by the need to address the inherent redundancy in pixel-level representations. Consequently, there is a growing…

计算机视觉与模式识别 · 计算机科学 2024-12-18 Anni Tang , Tianyu He , Junliang Guo , Xinle Cheng , Li Song , Jiang Bian

The scarcity of high-quality data remains a primary bottleneck in adapting multimodal generative models for medical image editing. Existing medical image editing datasets often suffer from limited diversity, neglect of medical image…

计算机视觉与模式识别 · 计算机科学 2026-05-13 Yongfan Lai , Wen Qian , Bo Liu , Hongyan Li , Hao Luo , Fan Wang , Bohan Zhuang , Shenda Hong

With the recent popularity of animated GIFs on social media, there is need for ways to index them with rich metadata. To advance research on animated GIF understanding, we collected a new dataset, Tumblr GIF (TGIF), with 100K animated GIFs…

计算机视觉与模式识别 · 计算机科学 2016-04-13 Yuncheng Li , Yale Song , Liangliang Cao , Joel Tetreault , Larry Goldberg , Alejandro Jaimes , Jiebo Luo

Controllable human video generation aims to produce realistic videos of humans with explicitly guided motions and appearances,serving as a foundation for digital humans, animation, and embodied AI.However, the scarcity of largescale,…

计算机视觉与模式识别 · 计算机科学 2026-04-24 Yuanchen Fei , Yude Zou , Zejian Kang , Ming Li , Jiaying Zhou , Xiangru Huang

Text-to-video generative models convert textual prompts into dynamic visual content, offering wide-ranging applications in film production, gaming, and education. However, their real-world performance often falls short of user expectations.…

计算机视觉与模式识别 · 计算机科学 2025-05-14 Wenhao Wang , Yi Yang

In recent years, with the vigorous development of the video game industry, the proportion of gaming videos on major video websites like YouTube has dramatically increased. However, relatively little research has been done on the automatic…

图像与视频处理 · 电气工程与系统科学 2022-04-15 Xiangxu Yu , Zhengzhong Tu , Neil Birkbeck , Yilin Wang , Balu Adsumilli , Alan C. Bovik

User-generated content platforms curate their vast repositories into thematic compilations that facilitate the discovery of high-quality material. Platforms that seek tight editorial control employ people to do this curation, but this…

人机交互 · 计算机科学 2020-08-13 Yan Chen , Andrés Monroy-Hernández , Ian Wehrman , Steve Oney , Walter S. Lasecki , Rajan Vaish

The success of modern text-to-image generation is largely attributed to massive, high-quality datasets. Currently, these datasets are curated through a filter-first paradigm that aggressively discards low-quality raw data based on the…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Zhiyang Liang , Ziyu Wan , Hongyu Liu , Dong Chen , Qiu Shen , Hao Zhu , Dongdong Chen

In an era where visual content generation is increasingly driven by machine learning, the integration of human feedback into generative models presents significant opportunities for enhancing user experience and output quality. This study…

计算机视觉与模式识别 · 计算机科学 2023-07-20 Dimitri von Rütte , Elisabetta Fedele , Jonathan Thomm , Lukas Wolf

Advances in generative artificial intelligence have altered multimedia creation, allowing for automatic cinematic video synthesis from text inputs. This work describes a method for creating 60-second cinematic movies incorporating Stable…

计算机视觉与模式识别 · 计算机科学 2025-06-13 Sridhar S , Nithin A , Shakeel Rifath , Vasantha Raj

In recent years, large-scale generative models for visual content (\textit{e.g.,} images, videos, and 3D objects/scenes) have made remarkable progress. However, training large-scale video generation models remains particularly challenging…

计算机视觉与模式识别 · 计算机科学 2025-10-23 Yongshun Zhang , Zhongyi Fan , Yonghang Zhang , Zhangzikang Li , Weifeng Chen , Zhongwei Feng , Chaoyue Wang , Peng Hou , Anxiang Zeng

We introduce Open-Sora Plan, an open-source project that aims to contribute a large generation model for generating desired high-resolution videos with long durations based on various user inputs. Our project comprises multiple components…

Text-to-video generation has shown promising results. However, by taking only natural languages as input, users often face difficulties in providing detailed information to precisely control the model's output. In this work, we propose…

计算机视觉与模式识别 · 计算机科学 2023-12-06 Hsin-Ping Huang , Yu-Chuan Su , Deqing Sun , Lu Jiang , Xuhui Jia , Yukun Zhu , Ming-Hsuan Yang

Text-to-image generation has achieved astonishing results, yet precise spatial controllability and prompt fidelity remain highly challenging. This limitation is typically addressed through cumbersome prompt engineering, scene layout…

计算机视觉与模式识别 · 计算机科学 2024-04-04 Petru-Daniel Tudosiu , Yongxin Yang , Shifeng Zhang , Fei Chen , Steven McDonagh , Gerasimos Lampouras , Ignacio Iacobacci , Sarah Parisot

The volume of User Generated Content (UGC) has increased in recent years. The challenge with this type of content is assessing its quality. So far, the state-of-the-art metrics are not exhibiting a very high correlation with perceptual…

计算机视觉与模式识别 · 计算机科学 2023-08-15 Xinyi Wang , Angeliki Katsenou , David Bull

Human video generation remains challenging due to the difficulty of jointly modeling human appearance, motion, and camera viewpoint under limited multi-view data. Existing methods often address these factors separately, resulting in limited…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Zhengwentai Sun , Keru Zheng , Chenghong Li , Hongjie Liao , Xihe Yang , Heyuan Li , Yihao Zhi , Shuliang Ning , Shuguang Cui , Xiaoguang Han

Directly editing ultra-high-resolution (UHR) images is valuable but underexplored, primarily due to the lack of high-quality data and the challenge in modeling high-frequency texture details. We introduce VINS-120K, the first large-scale…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Zhizhou Chen , Shanyan Guan , Zhanxin Gao , En Ci , Yanhao Ge , Wei Li , Zhenyu Zhang , Jian Yang , Ying Tai

Video Detailed Captioning (VDC) is a crucial task for vision-language bridging, enabling fine-grained descriptions of complex video content. In this paper, we first comprehensively benchmark current state-of-the-art approaches and…

计算机视觉与模式识别 · 计算机科学 2025-03-13 Luozheng Qin , Zhiyu Tan , Mengping Yang , Xiaomeng Yang , Hao Li