中文
相关论文

相关论文: ControlCol: Controllability in Automatic Speaker V…

200 篇论文

Controllable image denoising aims to generate clean samples with human perceptual priors and balance sharpness and smoothness. In traditional filter-based denoising methods, this can be easily achieved by adjusting the filtering strength.…

计算机视觉与模式识别 · 计算机科学 2023-03-30 Zhaoyang Zhang , Yitong Jiang , Wenqi Shao , Xiaogang Wang , Ping Luo , Kaimo Lin , Jinwei Gu

Audio description (AD) makes video content accessible to millions of blind and low vision (BLV) users. However, creating high-quality AD involves a trade-off between the precision of human-crafted descriptions and the efficiency of…

人机交互 · 计算机科学 2025-08-05 Maryam Cheema , Sina Elahimanesh , Samuel Martin , Pooyan Fazli , Hasti Seifi

We propose a deep learning approach for user-guided image colorization. The system directly maps a grayscale image, along with sparse, local user "hints" to an output colorization with a Convolutional Neural Network (CNN). Rather than using…

计算机视觉与模式识别 · 计算机科学 2017-05-12 Richard Zhang , Jun-Yan Zhu , Phillip Isola , Xinyang Geng , Angela S. Lin , Tianhe Yu , Alexei A. Efros

Controllable image captioning models generate human-like image descriptions, enabling some kind of control over the generated captions. This paper focuses on controlling the caption length, i.e. a short and concise description or a long and…

计算机视觉与模式识别 · 计算机科学 2024-01-23 Elad Hirsch , Ayellet Tal

The proliferation of advanced tools for manipulating video has led to an arms race, pitting those who wish to sow disinformation against those who want to detect and expose it. Unfortunately, time favors the ill-intentioned in this race,…

图形学 · 计算机科学 2025-08-01 Peter F. Michael , Zekun Hao , Serge Belongie , Abe Davis

Self-supervised learning has driven significant progress in learning from single-subject, iconic images. However, there are still unanswered questions about the use of minimally-curated, naturalistic video data, which contain dense scenes…

计算机视觉与模式识别 · 计算机科学 2025-04-24 Alex N. Wang , Christopher Hoang , Yuwen Xiong , Yann LeCun , Mengye Ren

Text-to-Image (T2I) diffusion/flow models have recently achieved remarkable progress in visual fidelity and text alignment. However, they remain limited when users need to precisely control image layouts, something that natural language…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Amadou S. Sangare , Adrien Maglo , Mohamed Chaouch , Bertrand Luvison

Multi-modal learning, particularly among imaging and linguistic modalities, has made amazing strides in many high-level fundamental visual understanding problems, ranging from language grounding to dense event captioning. However, much of…

计算机视觉与模式识别 · 计算机科学 2019-10-28 Tanzila Rahman , Bicheng Xu , Leonid Sigal

Polarized color photography provides both visual textures and object surficial information in one single snapshot. However, the use of the directional polarizing filter array causes extremely lower photon count and SNR compared to…

图像与视频处理 · 电气工程与系统科学 2023-03-03 Zhuoxiao Li , Haiyang Jiang , Yinqiang Zheng

We present a novel approach to automatic image colorization by imitating the imagination process of human experts. Our imagination module is designed to generate color images that are context-correlated with black-and-white photos. Given a…

计算机视觉与模式识别 · 计算机科学 2021-08-23 Chenyang Lei , Yue Wu , Qifeng Chen

To enhance controllability in text-to-image generation, ControlNet introduces image-based control signals, while ControlNet++ improves pixel-level cycle consistency between generated images and the input control signal. To avoid the…

计算机视觉与模式识别 · 计算机科学 2025-11-10 Zonglin Lyu , Ming Li , Xinxin Liu , Chen Chen

Hyperspectral cameras face challenging spatial-spectral resolution trade-offs and are more affected by shot noise than RGB photos taken over the same total exposure time. Here, we present a colorization algorithm to reconstruct…

计算机视觉与模式识别 · 计算机科学 2024-03-19 M. Kerem Aydin , Qi Guo , Emma Alexander

This paper presents an innovative approach to enhance control over audio generation by emphasizing the alignment between audio and text representations during model training. In the context of language model-based audio generation, the…

Image captioning, a fundamental task in vision-language understanding, seeks to generate accurate natural language descriptions for provided images. Current image captioning approaches heavily rely on high-quality image-caption pairs, which…

计算机视觉与模式识别 · 计算机科学 2023-11-03 Chuanyang Jin

We present a method for harmonizing the lighting of a foreground video to match a target background scene, adjusting shadows, color tone, and illumination intensity (relightful harmonization). Unlike images, acquiring labeled data for…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Jun Myeong Choi , Jae Shin Yoon , Luchao Qi , Roni Sengupta , Joon-Young Lee

Blind and low-vision (BLV) people use audio descriptions (ADs) to access videos. However, current ADs are unalterable by end users, thus are incapable of supporting BLV individuals' potentially diverse needs and preferences. This research…

人机交互 · 计算机科学 2024-08-22 Rosiana Natalie , Ruei-Che Chang , Smitha Sheshadri , Anhong Guo , Kotaro Hara

Foley is a key element in video production, refers to the process of adding an audio signal to a silent video while ensuring semantic and temporal alignment. In recent years, the rise of personalized content creation and advancements in…

声音 · 计算机科学 2025-04-18 Roi Benita , Michael Finkelson , Tavi Halperin , Gleb Sterkin , Yossi Adi

Color vision is essential for human visual perception, but its impact on machine perception is still underexplored. There has been an intensified demand for understanding its role in machine perception for safety-critical tasks such as…

计算机视觉与模式识别 · 计算机科学 2025-02-14 Ming-Chang Chiu , Yingfei Wang , Derrick Eui Gyu Kim , Pin-Yu Chen , Xuezhe Ma

Video color grading is a critical post-production process that transforms flat, log-encoded raw footage into emotionally resonant cinematic visuals. Existing automated methods act as static, black-box executors that directly output edited…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Yuchen Guo , Junli Gong , Hongmin Cai , Yiu-ming Cheung , Weifeng Su

Color Constancy is the ability of the human visual system to perceive colors unchanged independently of the illumination. Giving a machine this feature will be beneficial in many fields where chromatic information is used. Particularly, it…

计算机视觉与模式识别 · 计算机科学 2019-12-05 Oleksii Sidorov