中文
相关论文

相关论文: A Survey on Audio Synthesis and Audio-Visual Multi…

200 篇论文

Text-to-image generation has attracted significant interest from researchers and practitioners in recent years due to its widespread and diverse applications across various industries. Despite the progress made in the domain of vision and…

计算机视觉与模式识别 · 计算机科学 2023-05-25 Yutong Zhou , Nobutaka Shimada

Separating a song into vocal and accompaniment components is an active research topic, and recent years witnessed an increased performance from supervised training using deep learning techniques. We propose to apply the visual information…

声音 · 计算机科学 2021-07-02 Bochen Li , Yuxuan Wang , Zhiyao Duan

Audio and visual modalities are inherently connected in speech signals: lip movements and facial expressions are correlated with speech sounds. This motivates studies that incorporate the visual modality to enhance an acoustic speech signal…

声音 · 计算机科学 2023-06-02 Juan F. Montesinos , Daniel Michelsanti , Gloria Haro , Zheng-Hua Tan , Jesper Jensen

Deep Learning has implemented a wide range of applications and has become increasingly popular in recent years. The goal of multimodal deep learning (MMDL) is to create models that can process and link information using various modalities.…

机器学习 · 计算机科学 2022-02-21 Jabeen Summaira , Xi Li , Amin Muhammad Shoib , Jabbar Abdul

Audio-visual correlation learning aims to capture essential correspondences and understand natural phenomena between audio and video. With the rapid growth of deep learning, an increasing amount of attention has been paid to this emerging…

多媒体 · 计算机科学 2025-12-30 Luís Vilaça , Yi Yu , Paula Viana

Self-supervised learning has attracted plenty of recent research interest. However, most works for self-supervision in speech are typically unimodal and there has been limited work that studies the interaction between audio and visual…

音频与语音处理 · 电气工程与系统科学 2021-03-19 Abhinav Shukla , Stavros Petridis , Maja Pantic

Recent years have seen a significant increase in video content creation and consumption. Crafting engaging content requires the careful curation of both visual and audio elements. While visual cue curation, through techniques like optimal…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Chao Huang , Ruohan Gao , J. M. F. Tsang , Jan Kurcius , Cagdas Bilen , Chenliang Xu , Anurag Kumar , Sanjeel Parekh

This survey and application guide to multimodal large language models(MLLMs) explores the rapidly developing field of MLLMs, examining their architectures, applications, and impact on AI and Generative Models. Starting with foundational…

人工智能 · 计算机科学 2025-12-02 Chia Xin Liang , Pu Tian , Caitlyn Heqi Yin , Yao Yua , Wei An-Hou , Li Ming , Xinyuan Song , Tianyang Wang , Ziqian Bi , Ming Liu

Speech synthesis and music audio generation from symbolic input differ in many aspects but share some similarities. In this study, we investigate how text-to-speech synthesis techniques can be used for piano MIDI-to-audio synthesis tasks.…

声音 · 计算机科学 2022-02-25 Erica Cooper , Xin Wang , Junichi Yamagishi

Balancing dialogue, music, and sound effects with accompanying video is crucial for immersive storytelling, yet current audio mixing workflows remain largely manual and labor-intensive. While recent advancements have introduced the visually…

声音 · 计算机科学 2026-01-15 Junhua Huang , Chao Huang , Chenliang Xu

In today's tech-driven world, significant advancements in artificial intelligence and virtual reality have emerged. These developments drive research into exploring their intersection in the realm of soundscape. Not only do these…

声音 · 计算机科学 2025-04-11 Rima Ayoubi , Laurent Lescop , Sang Bum Park

The rapid advancement of deepfake technology poses a significant threat to digital media integrity. Deepfakes, synthetic media created using AI, can convincingly alter videos and audio to misrepresent reality. This creates risks of…

计算机视觉与模式识别 · 计算机科学 2024-10-21 Kashish Gandhi , Prutha Kulkarni , Taran Shah , Piyush Chaudhari , Meera Narvekar , Kranti Ghag

In an era defined by the explosive growth of data and rapid technological advancements, Multimodal Large Language Models (MLLMs) stand at the forefront of artificial intelligence (AI) systems. Designed to seamlessly integrate diverse data…

Automatic Singing Assessment and Singing Information Processing have evolved over the past three decades to support singing pedagogy, performance analysis, and vocal training. While the first approach objectively evaluates a singer's…

音频与语音处理 · 电气工程与系统科学 2026-01-21 Arthur N. dos Santos , Bruno S. Masiero

Multimodal learning has driven innovation across various industries, particularly in the field of music. By enabling more intuitive interaction experiences and enhancing immersion, it not only lowers the entry barriers to the music but also…

多媒体 · 计算机科学 2026-02-24 Sifei Li , Mining Tan , Feier Shen , Minyan Luo , Zijiao Yin , Fan Tang , Weiming Dong , Changsheng Xu

Thanks to advancements in deep learning, speech generation systems now power a variety of real-world applications, such as text-to-speech for individuals with speech disorders, voice chatbots in call centers, cross-linguistic speech…

Talking head synthesis, an advanced method for generating portrait videos from a still image driven by specific content, has garnered widespread attention in virtual reality, augmented reality and game production. Recently, significant…

计算机视觉与模式识别 · 计算机科学 2024-06-19 Ming Meng , Yufei Zhao , Bo Zhang , Yonggui Zhu , Weimin Shi , Maxwell Wen , Zhaoxin Fan

In the domain of Music Information Retrieval (MIR), Automatic Music Transcription (AMT) emerges as a central challenge, aiming to convert audio signals into symbolic notations like musical notes or sheet music. This systematic review…

声音 · 计算机科学 2024-06-24 Fatemeh Jamshidi , Gary Pike , Amit Das , Richard Chapman

The thud of a bouncing ball, the onset of speech as lips open -- when visual and audio events occur together, it suggests that there might be a common, underlying event that produced both signals. In this paper, we argue that the visual and…

计算机视觉与模式识别 · 计算机科学 2018-10-10 Andrew Owens , Alexei A. Efros

Self supervised representation learning has recently attracted a lot of research interest for both the audio and visual modalities. However, most works typically focus on a particular modality or feature alone and there has been very…

音频与语音处理 · 电气工程与系统科学 2020-02-21 Abhinav Shukla , Konstantinos Vougioukas , Pingchuan Ma , Stavros Petridis , Maja Pantic