中文
相关论文

相关论文: A Unit Enhancement and Guidance Framework for Audi…

200 篇论文

This paper introduces the Procedural (audio) Variational autoEncoder (ProVE) framework as a general approach to learning Procedural Audio PA models of environmental sounds with an improvement to the realism of the synthesis while…

声音 · 计算机科学 2023-03-07 Danzel Serrano , Mark Cartwright

Retrieval-Augmented Generation (RAG) improves factual grounding by incorporating external knowledge into language model generation. However, when retrieved context is noisy, unreliable, or inconsistent with the model's parametric knowledge,…

计算与语言 · 计算机科学 2026-04-06 Jaemin Kim , Jong Chul Ye

Diffusion autoencoders (DAEs) are typically formulated as a noise prediction model and trained with a linear-$\beta$ noise schedule that spends much of its sampling steps at high noise levels. Because high noise levels are associated with…

计算机视觉与模式识别 · 计算机科学 2025-05-01 Pramook Khungurn , Sukit Seripanitkarn , Phonphrm Thawatdamrongkit , Supasorn Suwajanakorn

Character animation is a transformative field in computer graphics and vision, enabling dynamic and realistic video animations from static images. Despite advancements, maintaining appearance consistency in animations remains a challenge.…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Xiaoyu Jin , Zunnan Xu , Mingwen Ou , Wenming Yang

Despite the recent progress of audio-driven video generation, existing methods mostly focus on driving facial movements, leading to non-coherent head and body dynamics. Moving forward, it is desirable yet challenging to generate holistic…

Most of the prior studies in the spatial \ac{DoA} domain focus on a single modality. However, humans use auditory and visual senses to detect the presence of sound sources. With this motivation, we propose to use neural networks with audio…

声音 · 计算机科学 2021-05-14 Xinyuan Qian , Maulik Madhavi , Zexu Pan , Jiadong Wang , Haizhou Li

As artificial intelligence-generated content (AIGC) continues to evolve, video-to-audio (V2A) generation has emerged as a key area with promising applications in multimedia editing, augmented reality, and automated content creation. While…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Yuhuan You , Xihong Wu , Tianshu Qu

We propose PAV, Personalized Head Avatar for the synthesis of human faces under arbitrary viewpoints and facial expressions. PAV introduces a method that learns a dynamic deformable neural radiance field (NeRF), in particular from a…

计算机视觉与模式识别 · 计算机科学 2024-08-01 Akin Caliskan , Berkay Kicanaoglu , Hyeongwoo Kim

Real-time synthesis of high-fidelity 3D character motion from audio is a pivotal component for next-generation interactive avatars and virtual assistants. However, most existing approaches are limited to offline processing of complete audio…

计算机视觉与模式识别 · 计算机科学 2026-05-28 Bohong Chen , Yumeng Li , Yinglin Xu , Youyi Zheng , Yanlin Weng , Kun Zhou

Temporal alignment of fine-grained human actions in videos is important for numerous applications in computer vision, robotics, and mixed reality. State-of-the-art methods directly learn image-based embedding space by leveraging powerful…

计算机视觉与模式识别 · 计算机科学 2022-04-27 Taein Kwon , Bugra Tekin , Siyu Tang , Marc Pollefeys

We present READ Avatars, a 3D-based approach for generating 2D avatars that are driven by audio input with direct and granular control over the emotion. Previous methods are unable to achieve realistic animation due to the many-to-many…

计算机视觉与模式识别 · 计算机科学 2023-03-02 Jack Saunders , Vinay Namboodiri

Talking head generation is to synthesize a lip-synchronized talking head video by inputting an arbitrary face image and corresponding audio clips. Existing methods ignore not only the interaction and relationship of cross-modal information,…

计算机视觉与模式识别 · 计算机科学 2024-11-01 Sen Chen , Zhilei Liu , Jiaxing Liu , Longbiao Wang

Story visualization requires generating sequential imagery that aligns semantically with evolving narratives while maintaining rigorous consistency in character identity and visual style. However, existing methodologies often struggle with…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Jianzhang Zhang , Yijing Tian , Jiwang Qu , Chuang Liu

Traditional speech systems typically rely on separate, task-specific models for text-to-speech (TTS), automatic speech recognition (ASR), and voice conversion (VC), resulting in fragmented pipelines that limit scalability, efficiency, and…

声音 · 计算机科学 2026-01-19 Runyuan Cai , Yu Lin , Yiming Wang , Chunlin Fu , Xiaodong Zeng

Bias in speech emotion recognition (SER) systems often stems from spurious correlations between speaker characteristics and emotional labels, leading to unfair predictions across demographic groups. Many existing debiasing methods require…

音频与语音处理 · 电气工程与系统科学 2025-11-17 Yun-Shao Tsai , Yi-Cheng Lin , Huang-Cheng Chou , Hung-yi Lee

Principal component analysis (PCA) is a key tool in the field of data dimensionality reduction that is useful for various data science problems. However, many applications involve heterogeneous data that varies in quality due to noise…

机器学习 · 统计学 2023-11-14 Javier Salazar Cavazos , Jeffrey A. Fessler , Laura Balzano

Audio-visual embodied navigation, as a hot research topic, aims training a robot to reach an audio target using egocentric visual (from the sensors mounted on the robot) and audio (emitted from the target) input. The audio-visual…

声音 · 计算机科学 2022-10-06 Yinfeng Yu , Lele Cao , Fuchun Sun , Xiaohong Liu , Liejun Wang

Text-to-image diffusion models have achieved remarkable progress in recent years. However, training models for high-resolution image generation remains challenging, particularly when training data and computational resources are limited. In…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Ruonan Yu , Songhua Liu , Zhenxiong Tan , Xinchao Wang

Our goal is to create a realistic 3D facial avatar with hair and accessories using only a text description. While this challenge has attracted significant recent interest, existing methods either lack realism, produce unrealistic shapes, or…

计算机视觉与模式识别 · 计算机科学 2023-09-14 Hao Zhang , Yao Feng , Peter Kulits , Yandong Wen , Justus Thies , Michael J. Black

Audio-visual video segmentation (AVVS) aims to generate pixel-level maps of sound-producing objects that accurately align with the corresponding audio. However, existing methods often face temporal misalignment, where audio cues and…

计算机视觉与模式识别 · 计算机科学 2024-12-12 Kexin Li , Zongxin Yang , Yi Yang , Jun Xiao