中文
相关论文

相关论文: Detecting Musical Deepfakes

200 篇论文

As artificial intelligence becomes more and more ingrained in daily life, we present a novel system that uses deep learning for music recommendation and emotion-based detection. Through the use of facial recognition and the DeepFace…

计算机视觉与模式识别 · 计算机科学 2025-03-27 Swetha Kambham , Hubert Jhonson , Sai Prathap Reddy Kambham

Deepfakes are AI-synthesized multimedia data that may be abused for spreading misinformation. Deepfake generation involves both visual and audio manipulation. To detect audio-visual deepfakes, previous studies commonly employ two relatively…

声音 · 计算机科学 2025-06-10 Kuiyuan Zhang , Wenjie Pei , Rushi Lan , Yifang Guo , Zhongyun Hua

The detection and localization of deepfake content, particularly when small fake segments are seamlessly mixed with real videos, remains a significant challenge in the field of digital media security. Based on the recently released…

计算机视觉与模式识别 · 计算机科学 2024-09-12 Zhixi Cai , Abhinav Dhall , Shreya Ghosh , Munawar Hayat , Dimitrios Kollias , Kalin Stefanov , Usman Tariq

Cheapfake is a recently coined term that encompasses non-AI (``cheap'') manipulations of multimedia content. Cheapfakes are known to be more prevalent than deepfakes. Cheapfake media can be created using editing software for image/video…

The rapid advancement of AI-based music generation tools is revolutionizing the music industry but also posing challenges to artists, copyright holders, and providers alike. This necessitates reliable methods for detecting such AI-generated…

计算与语言 · 计算机科学 2025-07-01 Markus Frohmann , Gabriel Meseguer-Brocal , Markus Schedl , Elena V. Epure

All current benchmarks for multimodal deepfake detection manipulate entire frames using various generation techniques, resulting in oversaturated detection accuracies exceeding 94% at the video-level classification. However, these…

计算机视觉与模式识别 · 计算机科学 2024-08-07 Juho Jung , Sangyoun Lee , Jooeon Kang , Yunjin Na

Recent advances in Text-to-Speech (TTS) and Voice-Conversion (VC) using generative Artificial Intelligence (AI) technology have made it possible to generate high-quality and realistic human-like audio. This poses growing challenges in…

声音 · 计算机科学 2025-03-25 Xiang Li , Pin-Yu Chen , Wenqi Wei

In this paper, we propose a novel approach for generating music based on an artificial intelligence (AI) system. We analyze the features of music and use them to fit and predict the music. The fractional Fourier transform (FrFT) and the…

声音 · 计算机科学 2026-04-21 Li Ya , Chen Wei , Li Xiulai , Yu Lei , Deng Xinyi , Chen Chaofan

Music creation is typically composed of two parts: composing the musical score, and then performing the score with instruments to make sounds. While recent work has made much progress in automatic music generation in the symbolic domain,…

声音 · 计算机科学 2018-11-13 Bryan Wang , Yi-Hsuan Yang

The growing prominence of the field of audio deepfake detection is driven by its wide range of applications, notably in protecting the public from potential fraud and other malicious activities, prompting the need for greater attention and…

音频与语音处理 · 电气工程与系统科学 2024-12-12 Jiangyan Yi , Chu Yuan Zhang , Jianhua Tao , Chenglong Wang , Xinrui Yan , Yong Ren , Hao Gu , Junzuo Zhou

Generative models guided by text prompts are increasingly becoming more popular. However, no text-to-MIDI models currently exist due to the lack of a captioned MIDI dataset. This work aims to enable research that combines LLMs with symbolic…

音频与语音处理 · 电气工程与系统科学 2025-08-08 Jan Melechovsky , Abhinaba Roy , Dorien Herremans

Large-scale text-to-music generation models have significantly enhanced music creation capabilities, offering unprecedented creative freedom. However, their ability to collaborate effectively with human musicians remains limited. In this…

声音 · 计算机科学 2024-07-16 Yongyi Zang , Yixiao Zhang

In recent decades, neuroscientific and psychological research has traced direct relationships between taste and auditory perceptions. This article explores multimodal generative models capable of converting taste information into music,…

声音 · 计算机科学 2025-09-01 Matteo Spanio , Massimiliano Zampini , Antonio Rodà , Franco Pierucci

AI-generated synthetic media, also called Deepfakes, have significantly influenced so many domains, from entertainment to cybersecurity. Generative Adversarial Networks (GANs) and Diffusion Models (DMs) are the main frameworks used to…

We present Music Arena, an open platform for scalable human preference evaluation of text-to-music (TTM) models. Soliciting human preferences via listening studies is the gold standard for evaluation in TTM, but these studies are expensive…

Deepfakes are AI-generated media in which the original content is digitally altered to create convincing but manipulated images, videos, or audio. Among the various types of deepfakes, lip-syncing deepfakes are one of the most challenging…

计算机视觉与模式识别 · 计算机科学 2025-04-03 Soumyya Kanti Datta , Shan Jia , Siwei Lyu

Singing voice synthesis and singing voice conversion have significantly advanced, revolutionizing musical experiences. However, the rise of "Deepfake Songs" generated by these technologies raises concerns about authenticity. Unlike Audio…

声音 · 计算机科学 2023-09-07 Yuankun Xie , Jingjing Zhou , Xiaolin Lu , Zhenghao Jiang , Yuxin Yang , Haonan Cheng , Long Ye

The rapid advancement of generative adversarial networks (GANs) and diffusion models has enabled the creation of highly realistic deepfake content, posing significant threats to digital trust across audio-visual domains. While unimodal…

计算机视觉与模式识别 · 计算机科学 2025-11-12 Chende Zheng , Ruiqi Suo , Zhoulin Ji , Jingyi Deng , Fangbin Yi , Chenhao Lin , Chao Shen

Under the aegis of computer vision and deep learning technology, a new emerging techniques has introduced that anyone can make highly realistic but fake videos, images even can manipulates the voices. This technology is widely known as…

机器学习 · 计算机科学 2023-02-15 Bahar Uddin Mahmud , Afsana Sharmin

AI-empowered music processing is a diverse field that encompasses dozens of tasks, ranging from generation tasks (e.g., timbre synthesis) to comprehension tasks (e.g., music classification). For developers and amateurs, it is very difficult…

计算与语言 · 计算机科学 2023-10-26 Dingyao Yu , Kaitao Song , Peiling Lu , Tianyu He , Xu Tan , Wei Ye , Shikun Zhang , Jiang Bian