English
Related papers

Related papers: Few-shot Acoustic Synthesis with Multimodal Flow M…

200 papers

Few-shot image generation, which aims to produce plausible and diverse images for one category given a few images from this category, has drawn extensive attention. Existing approaches either globally interpolate different images or fuse…

Computer Vision and Pattern Recognition · Computer Science 2023-08-31 Mengping Yang , Zhe Wang , Wenyi Feng , Qian Zhang , Ting Xiao

Zero-shot learning enables models to generalise to unseen classes by leveraging semantic information, bridging the gap between training and testing sets with non-overlapping classes. While much research has focused on zero-shot learning in…

Sound · Computer Science 2025-07-03 Ysobel Sims , Alexandre Mendes , Stephan Chalup

Despite Flow Matching and diffusion models having emerged as powerful generative paradigms for continuous variables such as images and videos, their application to high-dimensional discrete data, such as language, is still limited. In this…

Machine Learning · Computer Science 2024-11-06 Itai Gat , Tal Remez , Neta Shaul , Felix Kreuk , Ricky T. Q. Chen , Gabriel Synnaeve , Yossi Adi , Yaron Lipman

Achieving high-fidelity lip-speech synchronization in audio-driven talking portrait synthesis remains challenging. While multi-stage pipelines or diffusion models yield high-quality results, they suffer from high computational costs. Some…

Computer Vision and Pattern Recognition · Computer Science 2025-04-24 Ziqi Ni , Ao Fu , Yi Zhou

The zero-shot scenario for speech generation aims at synthesizing a novel unseen voice with only one utterance of the target speaker. Although the challenges of adapting new voices in zero-shot scenario exist in both stages -- acoustic…

Sound · Computer Science 2022-07-06 Yi Lei , Shan Yang , Jian Cong , Lei Xie , Dan Su

To generate new images for a given category, most deep generative models require abundant training images from this category, which are often too expensive to acquire. To achieve the goal of generation based on only a few images, we propose…

Computer Vision and Pattern Recognition · Computer Science 2020-03-26 Yan Hong , Li Niu , Jianfu Zhang , Liqing Zhang

Few-shot imitation learning relies on only a small amount of task-specific demonstrations to efficiently adapt a policy for a given downstream tasks. Retrieval-based methods come with a promise of retrieving relevant past experiences to…

Robotics · Computer Science 2024-10-14 Li-Heng Lin , Yuchen Cui , Amber Xie , Tianyu Hua , Dorsa Sadigh

Low-shot sketch-based image retrieval is an emerging task in computer vision, allowing to retrieve natural images relevant to hand-drawn sketch queries that are rarely seen during the training phase. Related prior works either require…

Computer Vision and Pattern Recognition · Computer Science 2020-06-23 Anjan Dutta , Zeynep Akata

The efficient mixing of fluid samples in miniaturized total analysis systems is essential for numerous applications including biological screening assays, chemical extraction, polymerization, cell analysis, and protein folding. Miniaturized…

Fluid Dynamics · Physics 2022-10-25 S. A. Iqrar , J. Park , M. Afzal , H. J. Sung

Rings like gold, thuds like wood! The sound we hear in a scene is shaped not only by the spatial layout of the environment but also by the materials of the objects and surfaces within it. For instance, a room with wooden walls will produce…

Computer Vision and Pattern Recognition · Computer Science 2026-04-24 Mahnoor Fatima Saad , Sagnik Majumder , Kristen Grauman , Ziad Al-Halah

Audio-driven bimanual piano motion generation requires precise modeling of complex musical structures and dynamic cross-hand coordination. However, existing methods often rely on acoustic-only representations lacking symbolic priors, employ…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Xuan Wang , Kai Ruan , Jiayi Han , Kaiyue Zhou , Gaoang Wang

Flow-based generative models have recently shown impressive performance for conditional generation tasks, such as text-to-image generation. However, current methods transform a general unimodal noise distribution to a specific mode of the…

Machine Learning · Computer Science 2025-02-14 Noam Issachar , Mohammad Salama , Raanan Fattal , Sagie Benaim

Diffusion-based generative models have achieved state-of-the-art performance for perceptual quality in speech enhancement (SE). However, their iterative nature requires numerous Neural Function Evaluations (NFEs), posing a challenge for…

Generating natural-sounding, multi-speaker dialogue is crucial for applications such as podcast creation, virtual agents, and multimedia content generation. However, existing systems struggle to maintain speaker consistency, model…

Few-shot image generation (FSIG) aims to learn to generate new and diverse samples given an extremely limited number of samples from a domain, e.g., 10 training samples. Recent work has addressed the problem using transfer learning…

Computer Vision and Pattern Recognition · Computer Science 2023-05-10 Yunqing Zhao , Keshigeyan Chandrasegaran , Milad Abdollahzadeh , Ngai-Man Cheung

Flow-matching models have enabled high-quality text-to-speech synthesis, but their iterative sampling process during inference incurs substantial computational cost. Although distillation is widely used to reduce the number of inference…

Sound · Computer Science 2026-02-11 Bin Lin , Peng Yang , Chao Yan , Xiaochen Liu , Wei Wang , Boyong Wu , Pengfei Tan , Xuerui Yang

Room Impulse Responses (RIRs) enable realistic acoustic simulation, with applications ranging from multimedia production to speech data augmentation. However, acquiring high-quality real-world RIRs is labor-intensive, and data scarcity…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-14 Kirak Kim , Sungyoung Kim

Flow instabilities, wave propagation phenomena, and structural interaction are current topics of the field "Flow acoustics" also named "Aeroacoustics". Assuming the theory of classical mechanics, aeroacoustic applications are modeled by the…

Fluid Dynamics · Physics 2024-01-23 Stefan Schoder

High-resolution remote sensing images (RSIs) are crucial for Earth observation applications, yet acquiring them is often limited by sensor constraints and costs. In recent years, generative super-resolution (SR) methods, particularly…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Jiangwei Mo , Xi Lu , Hanlin Wu

In real life, acoustic scenes and audio events are naturally correlated. Humans instinctively rely on fine-grained audio events as well as the overall sound characteristics to distinguish diverse acoustic scenes. Yet, most previous…

Sound · Computer Science 2022-05-03 Yuanbo Hou , Bo Kang , Wout Van Hauwermeiren , Dick Botteldooren
‹ Prev 1 4 5 6 7 8 10 Next ›