中文
相关论文

相关论文: TCR: Short Video Title Generation and Cover Select…

200 篇论文

This paper presents Video-P2P, a novel framework for real-world video editing with cross-attention control. While attention control has proven effective for image editing with pre-trained image generation models, there are currently no…

计算机视觉与模式识别 · 计算机科学 2023-03-09 Shaoteng Liu , Yuechen Zhang , Wenbo Li , Zhe Lin , Jiaya Jia

Video paragraph captioning aims to describe multiple events in untrimmed videos with descriptive paragraphs. Existing approaches mainly solve the problem in two steps: event detection and then event captioning. Such two-step manner makes…

计算机视觉与模式识别 · 计算机科学 2021-06-01 Yuqing Song , Shizhe Chen , Qin Jin

Recent advances in video generation have made it possible to produce visually compelling videos, with wide-ranging applications in content creation, entertainment, and virtual reality. However, most existing diffusion transformer based…

计算机视觉与模式识别 · 计算机科学 2025-10-22 Teng Hu , Jiangning Zhang , Zihan Su , Ran Yi

Video Temporal Grounding (VTG) aims to localize relevant temporal segments in videos given natural language queries. Despite recent progress with large vision-language models (LVLMs) and instruction-tuning, existing approaches often suffer…

计算机视觉与模式识别 · 计算机科学 2025-07-25 Ruizhe Chen , Zhiting Fan , Tianze Luo , Heqing Zou , Zhaopeng Feng , Guiyang Xie , Hansheng Zhang , Zhuochen Wang , Zuozhu Liu , Huaijian Zhang

The rapid proliferation of video content across various platforms has highlighted the urgent need for advanced video retrieval systems. Traditional methods, which primarily depend on directly matching textual queries with video metadata,…

信息检索 · 计算机科学 2025-10-10 Peyang Liu , Xi Wang , Ziqiang Cui , Wei Ye

This paper presents a review for the NTIRE 2025 Challenge on Short-form UGC Video Quality Assessment and Enhancement. The challenge comprises two tracks: (i) Efficient Video Quality Assessment (KVQ), and (ii) Diffusion-based Image…

图像与视频处理 · 电气工程与系统科学 2025-04-18 Xin Li , Kun Yuan , Bingchen Li , Fengbin Guan , Yizhen Shao , Zihao Yu , Xijun Wang , Yiting Lu , Wei Luo , Suhang Yao , Ming Sun , Chao Zhou , Zhibo Chen , Radu Timofte , Yabin Zhang , Ao-Xiang Zhang , Tianwu Zhi , Jianzhao Liu , Yang Li , Jingwen Xu , Yiting Liao , Yushen Zuo , Mingyang Wu , Renjie Li , Shengyun Zhong , Zhengzhong Tu , Yufan Liu , Xiangguang Chen , Zuowei Cao , Minhao Tang , Shan Liu , Kexin Zhang , Jingfen Xie , Yan Wang , Kai Chen , Shijie Zhao , Yunchen Zhang , Xiangkai Xu , Hong Gao , Ji Shi , Yiming Bao , Xiugang Dong , Xiangsheng Zhou , Yaofeng Tu , Ying Liang , Yiwen Wang , Xinning Chai , Yuxuan Zhang , Zhengxue Cheng , Yingsheng Qin , Yucai Yang , Rong Xie , Li Song , Wei Sun , Kang Fu , Linhan Cao , Dandan Zhu , Kaiwei Zhang , Yucheng Zhu , Zicheng Zhang , Menghan Hu , Xiongkuo Min , Guangtao Zhai , Zhi Jin , Jiawei Wu , Wei Wang , Wenjian Zhang , Yuhai Lan , Gaoxiong Yi , Hengyuan Na , Wang Luo , Di Wu , MingYin Bai , Jiawang Du , Zilong Lu , Zhenyu Jiang , Hui Zeng , Ziguan Cui , Zongliang Gan , Guijin Tang , Xinglin Xie , Kehuan Song , Xiaoqiang Lu , Licheng Jiao , Fang Liu , Xu Liu , Puhua Chen , Ha Thu Nguyen , Katrien De Moor , Seyed Ali Amirshahi , Mohamed-Chaker Larabi , Qi Tang , Linfeng He , Zhiyong Gao , Zixuan Gao , Guohua Zhang , Zhiye Huang , Yi Deng , Qingmiao Jiang , Lu Chen , Yi Yang , Xi Liao , Nourine Mohammed Nadir , Yuxuan Jiang , Qiang Zhu , Siyue Teng , Fan Zhang , Shuyuan Zhu , Bing Zeng , David Bull , Meiqin Liu , Chao Yao , Yao Zhao

Video Temporal Grounding (VTG) aims to localize temporal segments in long, untrimmed videos that align with a given natural language query. This task typically comprises two subtasks: Moment Retrieval (MR) and Highlight Detection (HD).…

计算机视觉与模式识别 · 计算机科学 2025-10-24 Minseok Kang , Minhyeok Lee , Minjung Kim , Donghyeong Kim , Sangyoun Lee

Video captioning, the task of describing the content of a video, has seen some promising improvements in recent years with sequence-to-sequence models, but accurately learning the temporal and logical dynamics involved in the task still…

计算与语言 · 计算机科学 2017-08-09 Ramakanth Pasunuru , Mohit Bansal

This paper presents a novel framework termed Cut-and-Paste for real-word semantic video editing under the guidance of text prompt and additional reference image. While the text-driven video editing has demonstrated remarkable ability to…

计算机视觉与模式识别 · 计算机科学 2023-11-21 Zhichao Zuo , Zhao Zhang , Yan Luo , Yang Zhao , Haijun Zhang , Yi Yang , Meng Wang

Reinforcement learning (RL) has shown strong potential for enhancing reasoning in multimodal large language models, yet existing video reasoning methods often rely on coarse sequence-level rewards or single-factor token selection,…

计算机视觉与模式识别 · 计算机科学 2026-01-28 Ziyue Wang , Sheng Jin , Zhongrong Zuo , Jiawei Wu , Han Qiu , Qi She , Hao Zhang , Xudong Jiang

Keyphrase generation (KG) aims to generate a set of keyphrases given a document, which is a fundamental task in natural language processing (NLP). Most previous methods solve this problem in an extractive manner, while recently, several…

计算与语言 · 计算机科学 2019-01-17 Wang Chen , Yifan Gao , Jiani Zhang , Irwin King , Michael R. Lyu

Fine-grained visual categorization is to recognize hundreds of subcategories belonging to the same basic-level category, which is a highly challenging task due to the quite subtle and local visual distinctions among similar subcategories.…

计算机视觉与模式识别 · 计算机科学 2019-02-21 Xiangteng He , Yuxin Peng

With the tremendous growth of videos over the Internet, video thumbnails, providing video content previews, are becoming increasingly crucial to influencing users' online searching experiences. Conventional video thumbnails are generated…

计算机视觉与模式识别 · 计算机科学 2019-10-17 Yitian Yuan , Lin Ma , Wenwu Zhu

To unfold the tremendous amount of multimedia data uploaded daily to social media platforms, effective topic modeling techniques are needed. Existing work tends to apply topic models on written text datasets. In this paper, we propose a…

计算与语言 · 计算机科学 2021-10-29 Lukas Stappen , Jason Thies , Gerhard Hagerer , Björn W. Schuller , Georg Groh

Temporal sentence grounding in videos (TSGV) faces challenges due to public TSGV datasets containing significant temporal biases, which are attributed to the uneven temporal distributions of target moments. Existing methods generate…

计算机视觉与模式识别 · 计算机科学 2025-01-15 Junlong Ren , Gangjian Zhang , Haifeng Sun , Hao Wang

The extraction of text information in videos serves as a critical step towards semantic understanding of videos. It usually involved in two steps: (1) text recognition and (2) text classification. To localize texts in videos, we can resort…

计算机视觉与模式识别 · 计算机科学 2022-06-07 Ye Liu , Changchong Lu , Chen Lin , Di Yin , Bo Ren

Attention mechanisms have attracted considerable interest in image captioning because of its powerful performance. Existing attention-based models use feedback information from the caption generator as guidance to determine which of the…

计算机视觉与模式识别 · 计算机科学 2018-07-11 Zhihao Zhu , Zhan Xue , Zejian Yuan

Contrastive Language-Image Pre-training (CLIP) has been widely studied and applied in numerous applications. However, the emphasis on brief summary texts during pre-training prevents CLIP from understanding long descriptions. This issue is…

计算与语言 · 计算机科学 2024-10-07 Jiapeng Wang , Chengyu Wang , Kunzhe Huang , Jun Huang , Lianwen Jin

Retrieval-Augmented Generation (RAG) is a powerful strategy for improving the factual accuracy of models by retrieving external knowledge relevant to queries and incorporating it into the generation process. However, existing approaches…

计算机视觉与模式识别 · 计算机科学 2025-05-30 Soyeong Jeong , Kangsan Kim , Jinheon Baek , Sung Ju Hwang

Video temporal grounding (VTG) is a fine-grained video understanding problem that aims to ground relevant clips in untrimmed videos given natural language queries. Most existing VTG models are built upon frame-wise final-layer CLIP…

计算机视觉与模式识别 · 计算机科学 2024-07-23 Ye Liu , Jixuan He , Wanhua Li , Junsik Kim , Donglai Wei , Hanspeter Pfister , Chang Wen Chen