中文
相关论文

相关论文: Phantom-Data : Towards a General Subject-Consisten…

200 篇论文

This work addresses the problem of vehicle identification through non-overlapping cameras. As our main contribution, we introduce a novel dataset for vehicle identification, called Vehicle-Rear, that contains more than three hours of…

计算机视觉与模式识别 · 计算机科学 2021-07-28 Icaro O. de Oliveira , Rayson Laroca , David Menotti , Keiko V. O. Fonseca , Rodrigo Minetto

Large-scale text-to-image diffusion models have achieved great success in synthesizing high-quality and diverse images given target text prompts. Despite the revolutionary image generation ability, current state-of-the-art models still…

计算机视觉与模式识别 · 计算机科学 2025-01-20 Jingyuan Zhu , Huimin Ma , Jiansheng Chen , Jian Yuan

The rapid development of generative models has made it increasingly crucial to develop detectors that can reliably detect synthetic images. Although most of the work has now focused on cross-generator generalization, we argue that this…

计算机视觉与模式识别 · 计算机科学 2025-10-08 Amirtaha Amanzadi , Zahra Dehghanian , Hamid Beigy , Hamid R. Rabiee

Generalizing video matting models to real-world videos remains a significant challenge due to the scarcity of labeled data. To address this, we present Video Mask-to-Matte Model (VideoMaMa) that converts coarse segmentation masks into pixel…

计算机视觉与模式识别 · 计算机科学 2026-01-21 Sangbeom Lim , Seoung Wug Oh , Jiahui Huang , Heeji Yoon , Seungryong Kim , Joon-Young Lee

Recent work on time-series models has leveraged self-supervised training to learn meaningful features and patterns in order to improve performance on downstream tasks and generalize to unseen modalities. While these pretraining methods have…

机器学习 · 计算机科学 2026-04-10 Paul Quinlan , Qingguo Li , Xiaodan Zhu

Recent advancements in text-to-image generation models have dramatically enhanced the generation of photorealistic images from textual prompts, leading to an increased interest in personalized text-to-image applications, particularly in…

计算机视觉与模式识别 · 计算机科学 2025-03-05 Xierui Wang , Siming Fu , Qihan Huang , Wanggui He , Hao Jiang

With growing concerns over data privacy, researchers have started using virtual data as an alternative to sensitive real-world images for training person re-identification (Re-ID) models. However, existing virtual datasets produced by game…

计算机视觉与模式识别 · 计算机科学 2025-11-10 Ruolin Li , Min Liu , Yuan Bian , Zhaoyang Li , Yuzhen Li , Xueping Wang , Yaonan Wang

Since its beginning visual recognition research has tried to capture the huge variability of the visual world in several image collections. The number of available datasets is still progressively growing together with the amount of samples…

计算机视觉与模式识别 · 计算机科学 2014-02-25 Tatiana Tommasi , Tinne Tuytelaars , Barbara Caputo

Image-to-Video generation (I2V) animates a static image into a temporally coherent video sequence following textual instructions, yet preserving fine-grained object identity under changing viewpoints remains a persistent challenge. Unlike…

计算机视觉与模式识别 · 计算机科学 2026-02-11 Mingyang Wu , Ashirbad Mishra , Soumik Dey , Shuo Xing , Naveen Ravipati , Hansi Wu , Binbin Li , Zhengzhong Tu

Multimodal summarization with multimodal output (MSMO) has emerged as a promising research direction. Nonetheless, numerous limitations exist within existing public MSMO datasets, including insufficient maintenance, data inaccessibility,…

计算机视觉与模式识别 · 计算机科学 2023-11-21 Jielin Qiu , Jiacheng Zhu , William Han , Aditesh Kumar , Karthik Mittal , Claire Jin , Zhengyuan Yang , Linjie Li , Jianfeng Wang , Ding Zhao , Bo Li , Lijuan Wang

Text-to-video retrieval enables users to find relevant video content using natural language queries, a task that has grown increasingly important with the rapid expansion of online video. Over the past six years, research has produced…

Recent advancements in audio-video joint generation models have demonstrated impressive capabilities in content creation. However, generating high-fidelity human-centric videos in complex, real-world physical scenes remains a significant…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Lei Zhu , Xing Cai , Yingjie Chen , Yiheng Li , Binxin Yang , Hao Liu , Jie Chen , Chen Li , Jing LYu

Customized text-to-video generation aims to produce high-quality videos that incorporate user-specified subject identities or motion patterns. However, existing methods mainly focus on personalizing a single concept, either subject identity…

计算机视觉与模式识别 · 计算机科学 2025-03-28 Chi-Pin Huang , Yen-Siang Wu , Hung-Kai Chung , Kai-Po Chang , Fu-En Yang , Yu-Chiang Frank Wang

To mitigate the risk of harmful outputs from large vision models (LVMs), we introduce the SafeSora dataset to promote research on aligning text-to-video generation with human values. This dataset encompasses human preferences in…

计算机视觉与模式识别 · 计算机科学 2024-06-21 Josef Dai , Tianle Chen , Xuyao Wang , Ziran Yang , Taiye Chen , Jiaming Ji , Yaodong Yang

Diffusion Transformer has shown remarkable abilities in generating high-fidelity videos, delivering visually coherent frames and rich details over extended durations. However, existing video generation models still fall short in…

计算机视觉与模式识别 · 计算机科学 2026-03-04 Zhaoyang Li , Dongjun Qian , Kai Su , Qishuai Diao , Xiangyang Xia , Chang Liu , Wenfei Yang , Tianzhu Zhang , Zehuan Yuan

This work concerns video-language pre-training and representation learning. In this now ubiquitous training scheme, a model first performs pre-training on paired videos and text (e.g., video clips and accompanied subtitles) from a large…

计算机视觉与模式识别 · 计算机科学 2021-04-14 Luowei Zhou , Jingjing Liu , Yu Cheng , Zhe Gan , Lei Zhang

One of the key shortcomings in current text-to-image (T2I) models is their inability to consistently generate images which faithfully follow the spatial relationships specified in the text prompt. In this paper, we offer a comprehensive…

Video face swapping is crucial in film and entertainment production, where achieving high fidelity and temporal consistency over long and complex video sequences remains a significant challenge. Inspired by recent advances in…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Zekai Luo , Zongze Du , Zhouhang Zhu , Hao Zhong , Muzhi Zhu , Wen Wang , Yuling Xi , Chenchen Jing , Hao Chen , Chunhua Shen

Multi-modal keyphrase generation aims to produce a set of keyphrases that represent the core points of the input text-image pair. In this regard, dominant methods mainly focus on multi-modal fusion for keyphrase generation. Nevertheless,…

计算机视觉与模式识别 · 计算机科学 2023-09-12 Yifan Dong , Suhang Wu , Fandong Meng , Jie Zhou , Xiaoli Wang , Jianxin Lin , Jinsong Su

Video-to-video synthesis (vid2vid) aims at converting an input semantic video, such as videos of human poses or segmentation masks, to an output photorealistic video. While the state-of-the-art of vid2vid has advanced significantly,…

计算机视觉与模式识别 · 计算机科学 2019-10-29 Ting-Chun Wang , Ming-Yu Liu , Andrew Tao , Guilin Liu , Jan Kautz , Bryan Catanzaro