中文
相关论文

相关论文: Leveraging Mixture of Experts for Improved Speech …

200 篇论文

Deepfake is a generative deep learning algorithm that creates or changes facial features in a very realistic way making it hard to differentiate the real from the fake features It can be used to make movies look better as well as to spread…

计算机视觉与模式识别 · 计算机科学 2024-06-21 Nadeem Jabbar CH , Aqib Saghir , Ayaz Ahmad Meer , Salman Ahmad Sahi , Bilal Hassan , Siddiqui Muhammad Yasir

Generative deep learning algorithms have progressed to a point where it is difficult to tell the difference between what is real and what is fake. In 2018, it was discovered how easy it is to use this technology for unethical and malicious…

计算机视觉与模式识别 · 计算机科学 2020-09-23 Yisroel Mirsky , Wenke Lee

Recent advances in artificial speech and audio technologies have improved the abilities of deep-fake operators to falsify media and spread malicious misinformation. Anyone with limited coding skills can use freely available speech synthesis…

Deepfakes have recently raised significant trust issues and security concerns among the public. Compared to CNN face forgery detectors, ViT-based methods take advantage of the expressivity of transformers, achieving superior detection…

计算机视觉与模式识别 · 计算机科学 2025-08-21 Chenqi Kong , Anwei Luo , Peijun Bao , Yi Yu , Haoliang Li , Zengwei Zheng , Shiqi Wang , Alex C. Kot

Speech event detection is crucial for multimedia retrieval, involving the tagging of both semantic and acoustic events. Traditional ASR systems often overlook the interplay between these events, focusing solely on content, even though the…

计算与语言 · 计算机科学 2024-10-29 Jingqi Kang , Tongtong Wu , Jinming Zhao , Guitao Wang , Yinwei Wei , Hao Yang , Guilin Qi , Yuan-Fang Li , Gholamreza Haffari

Recent rapid advancements in deepfake technology have allowed the creation of highly realistic fake media, such as video, image, and audio. These materials pose significant challenges to human authentication, such as impersonation,…

密码学与安全 · 计算机科学 2023-09-12 Binh Le , Shahroz Tariq , Alsharif Abuadbba , Kristen Moore , Simon Woo

We propose a novel mixture of experts framework for field-of-view enhancement in binaural signal matching. Our approach enables dynamic spatial audio rendering that adapts to continuous talker motion, allowing users to emphasize or suppress…

Combining the strengths of many existing predictors to obtain a Mixture of Experts which is superior to its individual components is an effective way to improve the performance without having to develop new architectures or train a model…

计算机视觉与模式识别 · 计算机科学 2024-02-02 Kemal Oksuz , Selim Kuzucu , Tom Joy , Puneet K. Dokania

While the significant advancements have made in the generation of deepfakes using deep learning technologies, its misuse is a well-known issue now. Deepfakes can cause severe security and privacy issues as they can be used to impersonate a…

计算机视觉与模式识别 · 计算机科学 2022-03-02 Hasam Khalid , Shahroz Tariq , Minha Kim , Simon S. Woo

Audio deepfake detection (ADD) models are commonly evaluated using datasets that combine multiple synthesizers, with performance reported as a single Equal Error Rate (EER). However, this approach disproportionately weights synthesizers…

声音 · 计算机科学 2025-09-12 Chin Yuen Kwok , Jia Qi Yip , Zhen Qiu , Chi Hung Chi , Kwok Yan Lam

Advances in speech synthesis technologies, like text-to-speech (TTS) and voice conversion (VC), have made detecting deepfake speech increasingly challenging. Spoofing countermeasures often struggle to generalize effectively, particularly…

音频与语音处理 · 电气工程与系统科学 2025-01-27 Wen Huang , Yanmei Gu , Zhiming Wang , Huijia Zhu , Yanmin Qian

With the rapid advancement of generative audio models, distinguishing between human-composed and generated music is becoming increasingly challenging. As a response, models for detecting fake music have been proposed. In this work, we…

声音 · 计算机科学 2025-07-15 Tomasz Sroka , Tomasz Wężowicz , Dominik Sidorczuk , Mateusz Modrzejewski

This paper proposes an audio-visual deepfake detection approach that aims to capture fine-grained temporal inconsistencies between audio and visual modalities. To achieve this, both architectural and data synthesis strategies are…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Marcella Astrid , Enjie Ghorbel , Djamila Aouada

Solutions for defending against deepfake speech fall into two categories: proactive watermarking models and passive conventional deepfake detectors. While both address common threats, their differences in training, optimization, and…

声音 · 计算机科学 2025-06-18 Chia-Hua Wu , Wanying Ge , Xin Wang , Junichi Yamagishi , Yu Tsao , Hsin-Min Wang

This paper describes a submission to the Environment-Aware Speech and Sound Deepfake Detection Challenge (ESDD2) 2026, which addresses component-level deepfake detection using the CompSpoofV2 dataset, where speech and environmental sounds…

声音 · 计算机科学 2026-05-06 Khalid Zaman , Qixuan Huang , Muhammad Uzair , Masashi Unoki

The rapid advancement of artificial intelligence (AI) has enabled sophisticated audio generation and voice cloning technologies, posing significant security risks for applications reliant on voice authentication. While existing datasets and…

声音 · 计算机科学 2025-05-22 Kunyang Huang , Bin Hu

Detecting AI-generated images, particularly deepfakes, has become increasingly crucial, with the primary challenge being the generalization to previously unseen manipulation methods. This paper tackles this issue by leveraging the forgery…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Wentang Song , Zhiyuan Yan , Yuzhen Lin , Taiping Yao , Changsheng Chen , Shen Chen , Yandan Zhao , Shouhong Ding , Bin Li

The Mixture of Experts (MoE) architecture has excelled in Large Vision-Language Models (LVLMs), yet its potential in real-time open-vocabulary object detectors, which also leverage large-scale vision-language datasets but smaller models,…

计算机视觉与模式识别 · 计算机科学 2025-07-24 Yehao Lu , Minghe Weng , Zekang Xiao , Rui Jiang , Wei Su , Guangcong Zheng , Ping Lu , Xi Li

Despite the development of effective deepfake detectors in recent years, recent studies have demonstrated that biases in the data used to train these detectors can lead to disparities in detection accuracy across different races and…

计算机视觉与模式识别 · 计算机科学 2023-11-09 Yan Ju , Shu Hu , Shan Jia , George H. Chen , Siwei Lyu

Connectivity plays an ever-increasing role in modern society, with people all around the world having easy access to rapidly disseminated information. However, a more interconnected society enables the spread of intentionally false…

计算与语言 · 计算机科学 2023-02-03 João Vitorino , Tiago Dias , Tiago Fonseca , Nuno Oliveira , Isabel Praça
‹ 上一页 1 8 9 10 下一页 ›