中文
相关论文

相关论文: Identification of fake stereo audio

200 篇论文

Voice conversion (VC) aims at conversion of speaker characteristic without altering content. Due to training data limitations and modeling imperfections, it is difficult to achieve believable speaker mimicry without introducing processing…

音频与语音处理 · 电气工程与系统科学 2018-09-05 Tomi Kinnunen , Jaime Lorenzo-Trueba , Junichi Yamagishi , Tomoki Toda , Daisuke Saito , Fernando Villavicencio , Zhenhua Ling

Traditional stereo algorithms have focused their efforts on reconstruction quality and have largely avoided prioritizing for run time performance. Robots, on the other hand, require quick maneuverability and effective computation to observe…

机器人学 · 计算机科学 2016-02-18 Sudeep Pillai , Srikumar Ramalingam , John J. Leonard

Stereo matching algorithms usually consist of four steps, including matching cost calculation, matching cost aggregation, disparity calculation, and disparity refinement. Existing CNN-based methods only adopt CNN to solve parts of the four…

计算机视觉与模式识别 · 计算机科学 2018-03-29 Zhengfa Liang , Yiliu Feng , Yulan Guo , Hengzhu Liu , Wei Chen , Linbo Qiao , Li Zhou , Jianfeng Zhang

A naive approach for finding similar audio items would be to compare each entry from the feature vector of the test example with each feature vector of the candidates in a k-nearest neighbors fashion. There are already two problems with…

声音 · 计算机科学 2022-01-28 Kastriot Kadriu

Anti-spoofing is the task of speech authentication. That is, identifying genuine human speech compared to spoofed speech. The main focus of this paper is to suggest new representations for genuine and spoofed speech, based on the…

音频与语音处理 · 电气工程与系统科学 2022-10-28 Matan Karo , Arie Yeredor , Itshak Lapidot

For nearly a decade, deepfake detection has been framed as a classification task: given an audio or video clip, decide whether it is real or synthetic. Top detectors often report high accuracy on standard benchmarks; however, performance…

计算机与社会 · 计算机科学 2026-05-12 Jessee Ho , Shweta Khushu , Shaina Raza

Speaker Verification (SV) systems involve mainly two individual stages: feature extraction and classification. In this paper, we explore these two modules with the aim of improving the performance of a speaker verification system under…

音频与语音处理 · 电气工程与系统科学 2024-02-06 Kerlos Atia Abdalmalak , Ascensión Gallardo-Antol'in

Multi-taper estimators provide low-variance power spectrum estimates that can be used in place of the windowed discrete Fourier transform (DFT) to extract speech features such as mel-frequency cepstral coefficients (MFCCs). Even if past…

声音 · 计算机科学 2021-10-27 Xuechen Liu , Md Sahidullah , Tomi Kinnunen

Stereo Matching is one of the classical problems in computer vision for the extraction of 3D information but still controversial for accuracy and processing costs. The use of matching techniques and cost functions is crucial in the…

计算机视觉与模式识别 · 计算机科学 2022-10-31 Hamid Fsian , Vahid Mohammadi , Pierre Gouton , Saeid Minaei

Neural network-based speaker recognition has achieved significant improvement in recent years. A robust speaker representation learns meaningful knowledge from both hard and easy samples in the training set to achieve good performance.…

音频与语音处理 · 电气工程与系统科学 2022-10-31 Ruijie Tao , Kong Aik Lee , Zhan Shi , Haizhou Li

In this paper, we compare the performance of using binaural audio features in place of single-channel features for sound event detection. Three different binaural features are studied and evaluated on the publicly available TUT Sound Events…

声音 · 计算机科学 2017-10-10 Sharath Adavanne , Tuomas Virtanen

Fake audio detection is an emerging active topic. A growing number of literatures have aimed to detect fake utterance, which are mostly generated by Text-to-speech (TTS) or voice conversion (VC). However, countermeasures against…

声音 · 计算机科学 2024-09-02 Hao Gu , JiangYan Yi , Chenglong Wang , Yong Ren , Jianhua Tao , Xinrui Yan , Yujie Chen , Xiaohui Zhang

We propose Gated Stereo, a high-resolution and long-range depth estimation technique that operates on active gated stereo images. Using active and high dynamic range passive captures, Gated Stereo exploits multi-view cues alongside…

计算机视觉与模式识别 · 计算机科学 2023-05-23 Stefanie Walz , Mario Bijelic , Andrea Ramazzina , Amanpreet Walia , Fahim Mannan , Felix Heide

Deepfake audio presents a growing threat to digital security, due to its potential for social engineering, fraud, and identity misuse. However, existing detection models suffer from poor generalization across datasets, due to implicit…

声音 · 计算机科学 2025-05-13 Yasaman Ahmadiadli , Xiao-Ping Zhang , Naimul Khan

Speaker identification models are vulnerable to carefully designed adversarial perturbations of their input signals that induce misclassification. In this work, we propose a white-box steganography-inspired adversarial attack that generates…

We present VCNAC, a variable channel neural audio codec. Our approach features a single encoder and decoder parametrization that enables native inference for different channel setups, from mono speech to cinematic 5.1 channel surround…

We revisit the problem of visual depth estimation in the context of autonomous vehicles. Despite the progress on monocular depth estimation in recent years, we show that the gap between monocular and stereo depth accuracy remains large$-$a…

计算机视觉与模式识别 · 计算机科学 2020-07-09 Nikolai Smolyanskiy , Alexey Kamenev , Stan Birchfield

Speaker identification is a powerful, non-invasive and in-expensive biometric technique. The recognition accuracy, however, deteriorates when noise levels affect a specific band of frequency. In this paper, we present a sub-band based…

机器学习 · 计算机科学 2007-05-23 Unathi Mahola , Fulufhelo V. Nelwamondo , Tshilidzi Marwala

Recent advances in foundation models have enabled audio-generative models that produce high-fidelity sounds associated with music, events, and human actions. Despite the success achieved in modern audio-generative models, the conventional…

声音 · 计算机科学 2024-08-30 Tiantian Feng , Dimitrios Dimitriadis , Shrikanth Narayanan

Audio splicing is one of the most common manipulation techniques in the area of audio forensics. In this paper, the magnitudes of acoustic channel impulse response and ambient noise are proposed as the environmental signature. Specifically,…

密码学与安全 · 计算机科学 2014-11-27 Hong Zhao , Yifan Chen , Rui Wang , Hafiz Malik
‹ 上一页 1 8 9 10 下一页 ›