English
Related papers

Related papers: Sounding that Object: Interactive Object-Aware Ima…

200 papers

Audio-driven talking face video generation has attracted increasing attention due to its huge industrial potential. Some previous methods focus on learning a direct mapping from audio to visual content. Despite progress, they often struggle…

Computer Vision and Pattern Recognition · Computer Science 2024-08-13 Weizhi Zhong , Junfan Lin , Peixin Chen , Liang Lin , Guanbin Li

We introduce the active audio-visual source separation problem, where an agent must move intelligently in order to better isolate the sounds coming from an object of interest in its environment. The agent hears multiple audio sources…

Computer Vision and Pattern Recognition · Computer Science 2021-08-27 Sagnik Majumder , Ziad Al-Halah , Kristen Grauman

We introduce SeeingSounds, a lightweight and modular framework for audio-to-image generation that leverages the interplay between audio, language, and vision-without requiring any paired audio-visual data or training on visual generative…

Gestures are essential for enhancing co-speech communication, offering visual emphasis and complementing verbal interactions. While prior work has concentrated on point-level motion or fully supervised data-driven methods, we focus on…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Jiahui Chen , Yang Huan , Runhua Shi , Chanfan Ding , Xiaoqi Mo , Siyu Xiong , Yinong He

We propose a framework to continuously learn object-centric representations for visual learning and understanding. Existing object-centric representations either rely on supervisions that individualize objects in the scene, or perform…

Computer Vision and Pattern Recognition · Computer Science 2022-04-05 Chuanyu Pan , Yanchao Yang , Kaichun Mo , Yueqi Duan , Leonidas Guibas

Modeling imaging sensor noise is a fundamental problem for image processing and computer vision applications. While most previous works adopt statistical noise models, real-world noise is far more complicated and beyond what these models…

Computer Vision and Pattern Recognition · Computer Science 2020-10-20 Ke-Chi Chang , Ren Wang , Hung-Jin Lin , Yu-Lun Liu , Chia-Ping Chen , Yu-Lin Chang , Hwann-Tzong Chen

We address the challenge of creating 3D assets for household articulated objects from a single image. Prior work on articulated object creation either requires multi-view multi-state input, or only allows coarse control over the generation…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Jiayi Liu , Denys Iliash , Angel X. Chang , Manolis Savva , Ali Mahdavi-Amiri

The generation of LiDAR scans is a growing topic with diverse applications to autonomous driving. However, scan generation remains challenging, especially when compared to the rapid advancement of image and 3D object generation. We consider…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Ellington Kirby , Mickael Chen , Renaud Marlet , Nermin Samet

Object-centric learning aims to represent visual data with a set of object entities (a.k.a. slots), providing structured representations that enable systematic generalization. Leveraging advanced architectures like Transformers, recent…

Computer Vision and Pattern Recognition · Computer Science 2023-09-25 Ziyi Wu , Jingyu Hu , Wuyue Lu , Igor Gilitschenski , Animesh Garg

Curating datasets for object segmentation is a difficult task. With the advent of large-scale pre-trained generative models, conditional image generation has been given a significant boost in result quality and ease of use. In this paper,…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Mischa Dombrowski , Hadrien Reynaud , Matthew Baugh , Bernhard Kainz

In multimedia applications such as films and video games, spatial audio techniques are widely employed to enhance user experiences by simulating 3D sound: transforming mono audio into binaural formats. However, this process is often complex…

Multimedia · Computer Science 2025-02-14 Xiaojing Liu , Ogulcan Gurelli , Yan Wang , Joshua Reiss

Generating realistic talking faces is a complex and widely discussed task with numerous applications. In this paper, we present DiffTalker, a novel model designed to generate lifelike talking faces through audio and landmark co-driving.…

Computer Vision and Pattern Recognition · Computer Science 2023-09-15 Zipeng Qi , Xulong Zhang , Ning Cheng , Jing Xiao , Jianzong Wang

Many applications of speech technology require more and more audio data. Automatic assessment of the quality of the collected recordings is important to ensure they meet the requirements of the related applications. However, effective and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-19 Qiang Huang , Thomas Hain

Efficient audio quality assessment is vital for streamlining audio codec development. Objective assessment tools have been developed over time to algorithmically predict quality ratings from subjective assessments, the gold standard for…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-28 Pablo M. Delgado , Jürgen Herre

Talking face generation aims to synthesize a sequence of face images that correspond to a clip of speech. This is a challenging task because face appearance variation and semantics of speech are coupled together in the subtle movements of…

Computer Vision and Pattern Recognition · Computer Science 2019-04-24 Hang Zhou , Yu Liu , Ziwei Liu , Ping Luo , Xiaogang Wang

General audio source separation is a key capability for multimodal AI systems that can perceive and reason about sound. Despite substantial progress in recent years, existing separation models are either domain-specific, designed for fixed…

Audio and Speech Processing · Electrical Eng. & Systems 2025-12-24 Bowen Shi , Andros Tjandra , John Hoffman , Helin Wang , Yi-Chiao Wu , Luya Gao , Julius Richter , Matt Le , Apoorv Vyas , Sanyuan Chen , Christoph Feichtenhofer , Piotr Dollár , Wei-Ning Hsu , Ann Lee

This study presents a novel method for generating music visualisers using diffusion models, combining audio input with user-selected artwork. The process involves two main stages: image generation and video creation. First, music captioning…

Multimedia · Computer Science 2024-12-10 Leonardo Pina , Yongmin Li

We propose Context Diffusion, a diffusion-based framework that enables image generation models to learn from visual examples presented in context. Recent work tackles such in-context learning for image generation, where a query image is…

Computer Vision and Pattern Recognition · Computer Science 2025-07-24 Ivona Najdenkoska , Animesh Sinha , Abhimanyu Dubey , Dhruv Mahajan , Vignesh Ramanathan , Filip Radenovic

Generating sound effects that humans want is an important topic. However, there are few studies in this area for sound generation. In this study, we investigate generating sound conditioned on a text prompt and propose a novel text-to-sound…

Sound · Computer Science 2023-05-01 Dongchao Yang , Jianwei Yu , Helin Wang , Wen Wang , Chao Weng , Yuexian Zou , Dong Yu

In a noisy conversation environment such as a dinner party, people often exhibit selective auditory attention, or the ability to focus on a particular speaker while tuning out others. Recognizing who somebody is listening to in a…

Computer Vision and Pattern Recognition · Computer Science 2023-03-29 Fiona Ryan , Hao Jiang , Abhinav Shukla , James M. Rehg , Vamsi Krishna Ithapu
‹ Prev 1 3 4 5 6 7 10 Next ›