中文
相关论文

相关论文: Audeo: Audio Generation for a Silent Performance V…

200 篇论文

We introduce an extensive new dataset of MIDI files, created by transcribing audio recordings of piano performances into their constituent notes. The data pipeline we use is multi-stage, employing a language model to autonomously crawl and…

声音 · 计算机科学 2025-07-01 Louis Bradshaw , Simon Colton

We focus on the task of generating sound from natural videos, and the sound should be both temporally and content-wise aligned with visual signals. This task is extremely challenging because some sounds generated \emph{outside} a camera can…

计算机视觉与模式识别 · 计算机科学 2023-07-19 Peihao Chen , Yang Zhang , Mingkui Tan , Hongdong Xiao , Deng Huang , Chuang Gan

The utilization of deep learning techniques in generating various contents (such as image, text, etc.) has become a trend. Especially music, the topic of this paper, has attracted widespread attention of countless researchers.The whole…

声音 · 计算机科学 2020-11-16 Shulei Ji , Jing Luo , Xinyu Yang

In movie productions, the Foley Artist is responsible for creating an overlay soundtrack that helps the movie come alive for the audience. This requires the artist to first identify the sounds that will enhance the experience for the…

声音 · 计算机科学 2020-06-29 Sanchita Ghose , John J. Prevost

Recent advances in diffusion models have showcased promising results in the text-to-video (T2V) synthesis task. However, as these T2V models solely employ text as the guidance, they tend to struggle in modeling detailed temporal dynamics.…

计算机视觉与模式识别 · 计算机科学 2023-05-24 Seungwoo Lee , Chaerin Kong , Donghyeon Jeon , Nojun Kwak

Human perceives rich auditory experience with distinct sound heard by ears. Videos recorded with binaural audio particular simulate how human receives ambient sound. However, a large number of videos are with monaural audio only, which…

声音 · 计算机科学 2021-05-04 Yan-Bo Lin , Yu-Chiang Frank Wang

Large-scale Text-to-Video (T2V) diffusion models have recently demonstrated unprecedented capability to transform natural language descriptions into stunning and photorealistic videos. Despite the promising results, a significant challenge…

计算机视觉与模式识别 · 计算机科学 2024-07-09 Xingyi Yang , Xinchao Wang

How does audio describe the world around us? In this paper, we propose a method for generating an image of a scene from sound. Our method addresses the challenges of dealing with the large gaps that often exist between sight and sound. We…

计算机视觉与模式识别 · 计算机科学 2023-03-31 Kim Sung-Bin , Arda Senocak , Hyunwoo Ha , Andrew Owens , Tae-Hyun Oh

Musicians delicately control their bodies to generate music. Sometimes, their motions are too subtle to be captured by the human eye. To analyze how they move to produce the music, we need to estimate precise 4D human pose (3D pose over…

计算机视觉与模式识别 · 计算机科学 2024-11-27 Seong Jong Yoo , Snehesh Shrestha , Irina Muresanu , Cornelia Fermüller

This paper presents <Dialogue in Resonance>, an interactive music piece for a human pianist and a computer-controlled piano that integrates real-time automatic music transcription into a score-driven framework. Unlike previous approaches…

声音 · 计算机科学 2025-05-23 Hayeon Bang , Taegyun Kwon , Juhan Nam

We demonstrate how conditional generation from diffusion models can be used to tackle a variety of realistic tasks in the production of music in 44.1kHz stereo audio with sampling-time guidance. The scenarios we consider include…

声音 · 计算机科学 2023-12-06 Mark Levy , Bruno Di Giorgi , Floris Weers , Angelos Katharopoulos , Tom Nickson

Generating combined visual and auditory sensory experiences is critical for the consumption of immersive content. Recent advances in neural generative models have enabled the creation of high-resolution content across multiple modalities…

计算机视觉与模式识别 · 计算机科学 2025-07-08 Rishit Dagli , Shivesh Prakash , Robert Wu , Houman Khosravani

We study the problem of making 3D scene reconstructions interactive by asking the following question: can we predict the sounds of human hands physically interacting with a scene? First, we record a video of a human manipulating objects…

计算机视觉与模式识别 · 计算机科学 2025-06-12 Yiming Dou , Wonseok Oh , Yuqing Luo , Antonio Loquercio , Andrew Owens

The recent success in StyleGAN demonstrates that pre-trained StyleGAN latent space is useful for realistic video generation. However, the generated motion in the video is usually not semantically meaningful due to the difficulty of…

计算机视觉与模式识别 · 计算机科学 2022-10-24 Seung Hyun Lee , Gyeongrok Oh , Wonmin Byeon , Chanyoung Kim , Won Jeong Ryoo , Sang Ho Yoon , Hyunjun Cho , Jihyun Bae , Jinkyu Kim , Sangpil Kim

The virtual world is being established in which digital humans are created indistinguishable from real humans. Producing their audio-related capabilities is crucial since voice conveys extensive personal characteristics. We aim to create a…

声音 · 计算机科学 2023-05-10 Wei Xue , Yiwen Wang , Qifeng Liu , Yike Guo

In recent years, Text-to-Audio Generation has achieved remarkable progress, offering sound creators powerful tools to transform textual inspirations into vivid audio. However, existing models predominantly operate directly in the acoustic…

音频与语音处理 · 电气工程与系统科学 2026-01-30 Zheqi Dai , Guangyan Zhang , Haolin He , Xiquan Li , Jingyu Li , Chunyat Wu , Yiwen Guo , Qiuqiang Kong

Multimodal generative models have shown remarkable progress in single-modality video and audio synthesis, yet truly joint audio-video generation remains an open challenge. In this paper, I explore four key contributions to advance this…

声音 · 计算机科学 2026-03-18 Alejandro Paredes La Torre

In this paper, we touch on the problem of markerless multi-modal human motion capture especially for string performance capture which involves inherently subtle hand-string contacts and intricate movements. To fulfill this goal, we first…

Emotions are fundamental to the creation and perception of music performances. However, achieving human-like expression and emotion through machine learning models for performance rendering remains a challenging task. In this work, we…

声音 · 计算机科学 2025-11-06 Ilya Borovik , Dmitrii Gavrilev , Vladimir Viro

We propose a method for adding sound-guided visual effects to specific regions of videos with a zero-shot setting. Animating the appearance of the visual effect is challenging because each frame of the edited video should have visual…

计算机视觉与模式识别 · 计算机科学 2023-04-17 Seung Hyun Lee , Sieun Kim , Innfarn Yoo , Feng Yang , Donghyeon Cho , Youngseo Kim , Huiwen Chang , Jinkyu Kim , Sangpil Kim