中文
相关论文

相关论文: Live Music Diffusion Models: Efficient Fine-Tuning…

200 篇论文

Modern music retrieval systems often rely on fixed representations of user preferences, limiting their ability to capture users' diverse and uncertain retrieval needs. To address this limitation, we introduce Diff4Steer, a novel generative…

声音 · 计算机科学 2025-04-25 Xuchan Bao , Judith Yue Li , Zhong Yi Wan , Kun Su , Timo Denk , Joonseok Lee , Dima Kuzmin , Fei Sha

Text-to-music generation technology is progressing rapidly, creating new opportunities for musical composition and editing. However, existing music editing methods often fail to preserve the source music's temporal structure, including…

声音 · 计算机科学 2025-11-19 Yi Yang , Haowen Li , Tianxiang Li , Boyu Cao , Xiaohan Zhang , Liqun Chen , Qi Liu

We present LidarDM, a novel LiDAR generative model capable of producing realistic, layout-aware, physically plausible, and temporally coherent LiDAR videos. LidarDM stands out with two unprecedented capabilities in LiDAR generative…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Vlas Zyrianov , Henry Che , Zhijian Liu , Shenlong Wang

We introduce a novel approach for high-resolution talking head generation from a single image and audio input. Prior methods using explicit face models, like 3D morphable models (3DMM) and facial landmarks, often fall short in generating…

计算机视觉与模式识别 · 计算机科学 2024-12-11 Sejong Yang , Seoung Wug Oh , Yang Zhou , Seon Joo Kim

Instructional video editing applies edits to an input video using only text prompts, enabling intuitive natural-language control. Despite rapid progress, most methods still require fixed-length inputs and substantial compute. Meanwhile,…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Mohammadreza Salehi , Mehdi Noroozi , Luca Morreale , Ruchika Chavhan , Malcolm Chadwick , Alberto Gil Ramos , Abhinav Mehrotra

Building upon Diff-A-Riff, a latent diffusion model for musical instrument accompaniment generation, we present a series of improvements targeting quality, diversity, inference speed, and text-driven control. First, we upgrade the…

声音 · 计算机科学 2024-10-31 Javier Nistal , Marco Pasini , Stefan Lattner

Text-to-image (T2I) generation with Stable Diffusion models (SDMs) involves high computing demands due to billion-scale parameters. To enhance efficiency, recent studies have reduced sampling steps and applied network quantization while…

机器学习 · 计算机科学 2024-12-03 Bo-Kyeong Kim , Hyoung-Kyu Song , Thibault Castells , Shinkook Choi

With the development of artificial intelligence (AI) techniques, implementing AI-based techniques to improve wireless transceivers becomes an emerging research topic. Within this context, AI-based channel characterization and estimation…

信号处理 · 电气工程与系统科学 2025-10-29 Yuzhi Yang , Sen Yan , Weijie Zhou , Brahim Mefgouda , Ridong Li , Zhaoyang Zhang , Mérouane Debbah

Diffusion models (DMs) have recently emerged as SoTA tools for generative modeling in various domains. Standard DMs can be viewed as an instantiation of hierarchical variational autoencoders (VAEs) where the latent variables are inferred…

计算机视觉与模式识别 · 计算机科学 2022-10-12 Jiatao Gu , Shuangfei Zhai , Yizhe Zhang , Miguel Angel Bautista , Josh Susskind

Sounding Video Generation (SVG) is an audio-video joint generation task challenged by high-dimensional signal spaces, distinct data formats, and different patterns of content information. To address these issues, we introduce a novel…

计算机视觉与模式识别 · 计算机科学 2024-10-03 Mingzhen Sun , Weining Wang , Yanyuan Qiao , Jiahui Sun , Zihan Qin , Longteng Guo , Xinxin Zhu , Jing Liu

With the phenomenal success of diffusion models and ChatGPT, deep generation models (DGMs) have been experiencing explosive growth from 2022. Not limited to content generation, DGMs are also widely adopted in Internet of Things, Metaverse,…

网络与互联网体系结构 · 计算机科学 2023-03-31 Yinqiu Liu , Hongyang Du , Dusit Niyato , Jiawen Kang , Zehui Xiong , Dong In Kim , Abbas Jamalipour

Latent diffusion models (LDMs) achieve state-of-the-art performance across various tasks, including image generation and video synthesis. However, they generally lack robustness, a limitation that remains not fully explored in current…

机器学习 · 计算机科学 2025-06-10 Boris Martirosyan , Alexey Karmanov

Stable diffusion models represent the state-of-the-art in data synthesis across diverse domains and hold transformative potential for applications in science and engineering, e.g., by facilitating the discovery of novel solutions and…

机器学习 · 计算机科学 2025-10-23 Stefano Zampini , Jacob K. Christopher , Luca Oneto , Davide Anguita , Ferdinando Fioretto

Diffusion models, a powerful and universal generative AI technology, have achieved tremendous success in computer vision, audio, reinforcement learning, and computational biology. In these applications, diffusion models provide flexible…

机器学习 · 计算机科学 2024-04-12 Minshuo Chen , Song Mei , Jianqing Fan , Mengdi Wang

Diffusion models have emerged as powerful deep generative techniques, producing high-quality and diverse samples in applications in various domains including audio. While existing reviews provide overviews, there remains limited in-depth…

声音 · 计算机科学 2026-01-16 Ge Zhu , Yutong Wen , Zhiyao Duan

Autoregressive (AR) Large Language Models (LLMs) have demonstrated significant success across numerous tasks. However, the AR modeling paradigm presents certain limitations; for instance, contemporary autoregressive LLMs are trained to…

机器学习 · 计算机科学 2025-02-10 Justin Deschenaux , Caglar Gulcehre

Prevailing Dataset Distillation (DD) methods leveraging generative models confront two fundamental limitations. First, despite pioneering the use of diffusion models in DD and delivering impressive performance, the vast majority of…

计算机视觉与模式识别 · 计算机科学 2025-12-04 Letian Zhou , Songhua Liu , Xinchao Wang

Service-level mobile traffic prediction for individual users is essential for network efficiency and quality of service enhancement. However, current prediction methods are limited in their adaptability across different urban environments…

机器学习 · 计算机科学 2025-07-25 Shiyuan Zhang , Tong Li , Zhu Xiao , Hongyang Du , Kaibin Huang

Diffusion Language Models (DLMs) have recently achieved strong results in text generation. However, their multi-step sampling leads to slow inference, limiting practical use. To address this, we extend Inverse Distillation, a technique…

Along with the prosperity of generative artificial intelligence (AI), its potential for solving conventional challenges in wireless communications has also surfaced. Inspired by this trend, we investigate the application of the advanced…

信息论 · 计算机科学 2026-03-10 Xingyu Zhou , Le Liang , Jing Zhang , Peiwen Jiang , Yong Li , Shi Jin