English
Related papers

Related papers: Sounding Video Generator: A Unified Framework for …

200 papers

Learning associations across modalities is critical for robust multimodal reasoning, especially when a modality may be missing during inference. In this paper, we study this problem in the context of audio-conditioned visual synthesis -- a…

Computer Vision and Pattern Recognition · Computer Science 2020-07-24 Anoop Cherian , Moitreya Chatterjee , Narendra Ahuja

Text-guided scalable vector graphics (SVG) synthesis has broad applications in icon and sketch generation. However, existing text-to-SVG methods often suffer from limited editability, suboptimal visual quality, and low sample diversity. To…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Ximing Xing , Haitao Zhou , Chuang Wang , Jing Zhang , Dong Xu , Qian Yu

Visual Speech Recognition (VSR) stands at the intersection of computer vision and speech recognition, aiming to interpret spoken content from visual cues. A prominent challenge in VSR is the presence of homophenes-visually similar lip…

Artificial Intelligence · Computer Science 2024-06-19 Young Jin Ahn , Jungwoo Park , Sangha Park , Jonghyun Choi , Kee-Eung Kim

Text-to-video generation has trailed behind text-to-image generation in terms of quality and diversity, primarily due to the inherent complexities of spatio-temporal modeling and the limited availability of video-text datasets. Recent…

Computer Vision and Pattern Recognition · Computer Science 2024-10-04 Xiefan Guo , Jinlin Liu , Miaomiao Cui , Liefeng Bo , Di Huang

With the rapid development of AI-generated content (AIGC), video generation has emerged as one of its most dynamic and impactful subfields. In particular, the advancement of video generation foundation models has led to growing demand for…

Video colorization aims to transform grayscale videos into vivid color representations while maintaining temporal consistency and structural integrity. Existing video colorization methods often suffer from color bleeding and lack…

Computer Vision and Pattern Recognition · Computer Science 2025-01-17 Zixun Fang , Zhiheng Liu , Kai Zhu , Yu Liu , Ka Leong Cheng , Wei Zhai , Yang Cao , Zheng-Jun Zha

Recent video and language pretraining frameworks lack the ability to generate sentences. We present Multimodal Video Generative Pretraining (MV-GPT), a new pretraining framework for learning from unlabelled videos which can be effectively…

Computer Vision and Pattern Recognition · Computer Science 2022-05-11 Paul Hongsuck Seo , Arsha Nagrani , Anurag Arnab , Cordelia Schmid

Text-to-Audio-Video (T2AV) generation aims to synthesize temporally coherent video and semantically synchronized audio from natural language, yet its evaluation remains fragmented, often relying on unimodal metrics or narrowly scoped…

Computer Vision and Pattern Recognition · Computer Science 2025-12-25 Zhe Cao , Tao Wang , Jiaming Wang , Yanghai Wang , Yuanxing Zhang , Jialu Chen , Miao Deng , Jiahao Wang , Yubin Guo , Chenxi Liao , Yize Zhang , Zhaoxiang Zhang , Jiaheng Liu

Visual grounding is a common vision task that involves grounding descriptive sentences to the corresponding regions of an image. Most existing methods use independent image-text encoding and apply complex hand-crafted modules or…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Ming Dai , Lingfeng Yang , Yihao Xu , Zhenhua Feng , Wankou Yang

We introduce and define a novel task-Scene-Aware Visually-Driven Speech Synthesis, aimed at addressing the limitations of existing speech generation models in creating immersive auditory experiences that align with the real physical world.…

Sound · Computer Science 2026-02-04 Chengyuan Ma , Jiawei Jin , Ruijie Xiong , Chunxiang Jin , Canxiang Yan , Wenming Yang

We propose a novel self-supervised approach for learning audio and visual representations from unlabeled videos, based on their correspondence. The approach uses an attention mechanism to learn the relative importance of convolutional…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Sudha Krishnamurthy

Video super-resolution (VSR) approaches have shown impressive temporal consistency in upsampled videos. However, these approaches tend to generate blurrier results than their image counterparts as they are limited in their generative…

Computer Vision and Pattern Recognition · Computer Science 2024-05-03 Yiran Xu , Taesung Park , Richard Zhang , Yang Zhou , Eli Shechtman , Feng Liu , Jia-Bin Huang , Difan Liu

We introduce the task of SVG extraction, which consists in translating specific visual inputs from an image into scalable vector graphics. Existing multimodal models achieve strong results when generating SVGs from clean renderings or…

Computer Vision and Pattern Recognition · Computer Science 2026-02-26 Marco Terral , Haotian Zhang , Tianyang Zhang , Meng Lin , Xiaoqing Xie , Haoran Dai , Darsh Kaushik , Pai Peng , Nicklas Scharpff , David Vazquez , Joan Rodriguez

Diffusion models have achieved impressive performance in video generation, but their iterative denoising process remains computationally expensive due to the large number of tokens processed at each timestep. Recently, progressive…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Shikang Zheng , Jingkai Huang , Jiacheng Liu , Guantao Chen , Lixuan , Yuqi Lin , Peiliang Cai , Linfeng Zhang

Singing voice synthesis (SVS) has advanced significantly, enabling models to generate vocals with accurate pitch and consistent style. As these capabilities improve, the need for reliable evaluation and optimization becomes increasingly…

Sound · Computer Science 2025-12-03 Xueyan Li , Yuxin Wang , Mengjie Jiang , Qingzi Zhu , Jiang Zhang , Zoey Kim , Yazhe Niu

Denoising in the sRGB image space is challenging due to large noise variability. Although end-to-end methods perform well, their effectiveness in real-world scenarios is limited by the scarcity of real noisy-clean image pairs, which are…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Jaekyun Ko , Dongjin Kim , Soomin Lee , Guanghui Wang , Tae Hyun Kim

The dominant paradigm for learning video-text representations -- noise contrastive learning -- increases the similarity of the representations of pairs of samples that are known to be related, such as text and video from the same sample,…

Computer Vision and Pattern Recognition · Computer Science 2021-01-15 Mandela Patrick , Po-Yao Huang , Yuki Asano , Florian Metze , Alexander Hauptmann , João Henriques , Andrea Vedaldi

Audio-visual question answering (AVQA) is a challenging task that requires multistep spatio-temporal reasoning over multimodal contexts. Recent works rely on elaborate target-agnostic parsing of audio-visual scenes for spatial grounding…

Computer Vision and Pattern Recognition · Computer Science 2023-12-11 Yuanyuan Jiang , Jianqin Yin

Transition videos play a crucial role in media production, enhancing the flow and coherence of visual narratives. Traditional methods like morphing often lack artistic appeal and require specialized skills, limiting their effectiveness.…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Rui Zhang , Yaosen Chen , Yuegen Liu , Wei Wang , Xuming Wen , Hongxia Wang

Text-to-Audio-Video (T2AV) generation is rapidly becoming a core interface for media creation, yet its evaluation remains fragmented. Existing benchmarks largely assess audio and video in isolation or rely on coarse embedding similarity,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Ziwei Zhou , Zeyuan Lai , Rui Wang , Yifan Yang , Zhen Xing , Yuqing Yang , Qi Dai , Lili Qiu , Chong Luo