中文
相关论文

相关论文: Amanous: Distribution-Switching for Superhuman Pia…

200 篇论文

Motion-to-music and music-to-motion have been studied separately, each attracting substantial research interest within their respective domains. The interaction between human motion and music is a reflection of advanced human intelligence,…

声音 · 计算机科学 2024-11-05 Fuming You , Minghui Fang , Li Tang , Rongjie Huang , Yongqi Wang , Zhou Zhao

Thanks to the rapid development of diffusion models, unprecedented progress has been witnessed in image synthesis. Prior works mostly rely on pre-trained linguistic models, but a text is often too abstract to properly specify all the…

计算机视觉与模式识别 · 计算机科学 2023-08-15 Binbin Yang , Yi Luo , Ziliang Chen , Guangrun Wang , Xiaodan Liang , Liang Lin

We study and analyze the fundamental aspects of noise propagation in recurrent as well as deep, multi-layer networks. The main focus of our study are neural networks in analogue hardware, yet the methodology provides insight for networks in…

新兴技术 · 计算机科学 2020-06-29 Nadezhda Semenova , Xavier Porte , Louis Andreoli , Maxime Jacquot , Laurent Larger , Daniel Brunner

Timbre transfer aims to modify the timbral identity of a musical recording while preserving the original melody and rhythm. While single-instrument timbre transfer has made substantial progress, existing approaches to multi-instrument…

声音 · 计算机科学 2026-05-12 Leduo Chen , Junchuan Zhao , Shengchen Li

Structured metamaterials are at the core of extensive research, promising for acoustic and thermal engineering. Nevertheless, the computational cost required for correctly simulating large systems imposes to use a continuous model to…

软凝聚态物质 · 物理学 2022-03-14 Haoming Luo , Valentina M. Giordano , Anthony Gravouil , Anne Tanguy

Surgical instrument segmentation is a key component in developing context-aware operating rooms. Existing works on this task heavily rely on the supervision of a large amount of labeled data, which involve laborious and expensive human…

计算机视觉与模式识别 · 计算机科学 2020-08-28 Daochang Liu , Yuhui Wei , Tingting Jiang , Yizhou Wang , Rulin Miao , Fei Shan , Ziyu Li

Style transfer of polyphonic music recordings is a challenging task when considering the modeling of diverse, imaginative, and reasonable music pieces in the style different from their original one. To achieve this, learning stable…

声音 · 计算机科学 2018-11-30 Chien-Yu Lu , Min-Xin Xue , Chia-Che Chang , Che-Rung Lee , Li Su

Efficient text-to-image generation remains a challenging task due to the high computational costs associated with the multi-step sampling in diffusion models. Although distillation of pre-trained diffusion models has been successful in…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Jeeyung Kim , Ze Wang , Qiang Qiu

This paper presents a regularized sampling method for multiband signals, that makes it possible to approach the Landau limit, while keeping the sensitivity to noise at a low level. The method is based on band-limited windowing, followed by…

信息论 · 计算机科学 2015-05-18 J. Selva

Synthesizing realistic piano hand motions requires both precision and naturalness. Physics-based methods achieve precision but produce stiff motions; data-driven models learn natural dynamics but struggle with positional accuracy. Piano…

Regenerating singing voices with altered lyrics while preserving melody consistency remains challenging, as existing methods either offer limited controllability or require laborious manual alignment. We propose YingMusic-Singer-Plus, a…

音频与语音处理 · 电气工程与系统科学 2026-04-10 Chunbo Hao , Junjie Zheng , Guobin Ma , Yuepeng Jiang , Huakang Chen , Wenjie Tian , Gongyu Chen , Zihao Chen , Lei Xie

Virtual furniture synthesis, which seamlessly integrates reference objects into indoor scenes while maintaining geometric coherence and visual realism, holds substantial promise for home design and e-commerce applications. However, this…

计算机视觉与模式识别 · 计算机科学 2025-12-22 Qilong Wang , Xiaofan Ming , Zhenyi Lin , Jinwen Li , Dongwei Ren , Wangmeng Zuo , Qinghua Hu

We study timbre transfer as an inference-time editing problem for music audio. Starting from a strong pre-trained latent diffusion model, we introduce a lightweight procedure that requires no additional training: (i) a dimension-wise noise…

声音 · 计算机科学 2026-01-29 Ching Ho Lee , Javier Nistal , Stefan Lattner , Marco Pasini , George Fazekas

We propose {\it HumanDiffusion,} a diffusion model trained from humans' perceptual gradients to learn an acceptable range of data for humans (i.e., human-acceptable distribution). Conventional HumanGAN aims to model the human-acceptable…

人机交互 · 计算机科学 2023-06-22 Yota Ueda , Shinnosuke Takamichi , Yuki Saito , Norihiro Takamune , Hiroshi Saruwatari

Traditionally, music was treated as an analogue signal and was generated manually. In recent years, music is conspicuous to technology which can generate a suite of music automatically without any human intervention. To accomplish this…

声音 · 计算机科学 2019-08-06 Sanidhya Mangal , Rahul Modak , Poorva Joshi

Mastering dexterous manipulation with multi-fingered hands has been a grand challenge in robotics for decades. Despite its potential, the difficulty of collecting high-quality data remains a primary bottleneck for high-precision tasks.…

机器人学 · 计算机科学 2026-05-19 Amber Xie , Haozhi Qi , Dorsa Sadigh

Consistency models possess high capabilities for image generation, advancing sampling steps to a single step through their advanced techniques. Current advancements move one step forward consistency training techniques and eliminates the…

机器学习 · 计算机科学 2024-04-10 Mahmut S. Gokmen , Cody Bumgardner , Jie Zhang , Ge Wang , Jin Chen

Euclidean diffusion models have achieved remarkable success in generative modeling across diverse domains, and they have been extended to manifold cases in recent advances. Instead of explicitly utilizing the structure of special manifolds…

机器学习 · 计算机科学 2026-01-06 Zichen Liu , Wei Zhang , Tiejun Li

Segmentation of enhancement in LGE cardiac MRI is critical for diagnosing various ischemic and non-ischemic cardiomyopathies. However, creating pixel-level annotations for these images is challenging and labor-intensive, leading to limited…

人工智能 · 计算机科学 2026-03-20 Athira J. Jacob , Puneet Sharma , Daniel Rueckert

Automatic Music Transcription (AMT) is the task of recognizing notes in audio recordings of music. The State-of-the-Art (SotA) benchmarks have been dominated by deep learning systems. Due to the scarcity of high quality data, they are…

声音 · 计算机科学 2024-08-12 Lukáš Samuel Marták , Patricia Hu , Gerhard Widmer