中文
相关论文

相关论文: L-C4: Language-Based Video Colorization for Creati…

200 篇论文

Video captioning targets interpreting the complex visual contents as text descriptions, which requires the model to fully understand video scenes including objects and their interactions. Prevailing methods adopt off-the-shelf object…

计算机视觉与模式识别 · 计算机科学 2022-09-07 Hao Wang , Guosheng Lin , Steven C. H. Hoi , Chunyan Miao

In this paper, we study the importance of pre-training for the generalization capability in the color constancy problem. We propose two novel approaches based on convolutional autoencoders: an unsupervised pre-training algorithm using a…

计算机视觉与模式识别 · 计算机科学 2020-05-26 Firas Laakom , Jenni Raitoharju , Alexandros Iosifidis , Jarno Nikkanen , Moncef Gabbouj

In recent years, creative content generations like style transfer and neural photo editing have attracted more and more attention. Among these, cartoonization of real-world scenes has promising applications in entertainment and industry.…

计算机视觉与模式识别 · 计算机科学 2022-04-05 Zhenhuan Liu , Liang Li , Huajie Jiang , Xin Jin , Dandan Tu , Shuhui Wang , Zheng-Jun Zha

We propose a framework for automatic colorization that allows for iterative editing and modifications. The core of our framework lies in an imagination module: by understanding the content within a grayscale image, we utilize a pre-trained…

计算机视觉与模式识别 · 计算机科学 2024-04-09 Xiaoyan Cong , Yue Wu , Qifeng Chen , Chenyang Lei

Accurate color alignment in text-to-image (T2I) generation is critical for applications such as fashion, product visualization, and interior design, yet current diffusion models struggle with nuanced and compound color terms (e.g., Tiffany…

计算机视觉与模式识别 · 计算机科学 2025-09-15 Sung-Lin Tsai , Bo-Lun Huang , Yu Ting Shen , Cheng Yu Yeo , Chiang Tseng , Bo-Kai Ruan , Wen-Sheng Lien , Hong-Han Shuai

Video large language models (Video-LLMs) can temporally ground language queries and retrieve video moments. Yet, such temporal comprehension capabilities are neither well-studied nor understood. So we conduct a study on prediction…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Minjoon Jung , Junbin Xiao , Byoung-Tak Zhang , Angela Yao

Video captioning is a challenging task since it requires generating sentences describing various diverse and complex videos. Existing video captioning models lack adequate visual representation due to the neglect of the existence of gaps…

计算机视觉与模式识别 · 计算机科学 2021-10-14 Mingkang Tang , Zhanyu Wang , Zhenhua Liu , Fengyun Rao , Dian Li , Xiu Li

Recent advances in learned video codecs have demonstrated remarkable compression efficiency. Two fundamental design aspects are critical: the choice of inter-frame coding framework and the temporal information propagation strategy.…

图像与视频处理 · 电气工程与系统科学 2025-10-20 Kuan-Wei Ho , Yi-Hsin Chen , Martin Benjak , Jörn Ostermann , Wen-Hsiao Peng

Dense video captioning (DVC) aims to generate multi-sentence descriptions to elucidate the multiple events in the video, which is challenging and demands visual consistency, discoursal coherence, and linguistic diversity. Existing methods…

计算机视觉与模式识别 · 计算机科学 2021-11-22 Xu Yan , Zhengcong Fei , Shuhui Wang , Qingming Huang , Qi Tian

Layered image assets are widely used in real-world creative workflows, enabling non-destructive iteration and flexible re-composition. Recent advances in layered image generation and decomposition synthesize or recover layered…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Ryugo Morita , Stanislav Frolov , Brian Bernhard Moser , Ko Watanabe , Riku Takahashi , Issey Sukeda , Andreas Dengel

Video portraits relighting is critical in user-facing human photography, especially for immersive VR/AR experience. Recent advances still fail to recover consistent relit result under dynamic illuminations from monocular RGB stream,…

计算机视觉与模式识别 · 计算机科学 2021-04-02 Longwen Zhang , Qixuan Zhang , Minye Wu , Jingyi Yu , Lan Xu

Adapting text-to-image (T2I) latent diffusion models (LDMs) to video editing has shown strong visual fidelity and controllability, but challenges remain in maintaining causal relationships inherent to the video data generating process.…

True video understanding requires making sense of non-lambertian scenes where the color of light arriving at the camera sensor encodes information about not just the last object it collided with, but about multiple mediums -- colored…

计算机视觉与模式识别 · 计算机科学 2019-04-05 Jean-Baptiste Alayrac , João Carreira , Andrew Zisserman

In-context vision and language models like Flamingo support arbitrarily interleaved sequences of images and text as input. This format not only enables few-shot learning via interleaving independent supervised (image, text) examples, but…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Wanrong Zhu , Jack Hessel , Anas Awadalla , Samir Yitzhak Gadre , Jesse Dodge , Alex Fang , Youngjae Yu , Ludwig Schmidt , William Yang Wang , Yejin Choi

A fundamental characteristic common to both human vision and natural language is their compositional nature. Yet, despite the performance gains contributed by large vision and language pretraining, recent investigations find that most-if…

计算机视觉与模式识别 · 计算机科学 2024-04-18 Chenhao Zheng , Jieyu Zhang , Aniruddha Kembhavi , Ranjay Krishna

Multimodal ML models can process data in multiple modalities (e.g., video, images, audio, text) and are useful for video content analysis in a variety of problems (e.g., object detection, scene understanding). In this paper, we focus on the…

计算机视觉与模式识别 · 计算机科学 2020-06-09 Palash Goyal , Saurabh Sahu , Shalini Ghosh , Chul Lee

The dual-stream transformer architecture-based joint audio-video generation method has become the dominant paradigm in current research. By incorporating pre-trained video diffusion models and audio diffusion models, along with a…

计算机视觉与模式识别 · 计算机科学 2026-03-20 Bingqi Ma , Linlong Lang , Ming Zhang , Dailan He , Xingtong Ge , Yi Zhang , Guanglu Song , Yu Liu

Recent developments in multimodal methodologies have marked the beginning of an exciting era for models adept at processing diverse data types, encompassing text, audio, and visual content. Models like GPT-4V, which merge computer vision…

Recently, image-to-video (I2V) diffusion models have demonstrated impressive scene understanding and generative quality, incorporating image conditions to guide generation. However, these models primarily animate static images without…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Luis Denninger , Sina Mokhtarzadeh Azar , Juergen Gall

Unsupervised object-centric learning from videos is a promising approach to extract structured representations from large, unlabeled collections of videos. To support downstream tasks like autonomous control, these representations must be…

计算机视觉与模式识别 · 计算机科学 2025-03-19 Anna Manasyan , Maximilian Seitzer , Filip Radovic , Georg Martius , Andrii Zadaianchuk