中文
相关论文

相关论文: Audio Splicing Detection and Localization Using En…

200 篇论文

Recent multi-modal audio-language models (ALMs) excel at text-audio retrieval but struggle with frame-wise audio understanding. Prior works use temporal-aware labels or unsupervised training to improve frame-wise capabilities, but they…

This paper describes a submission to the Environment-Aware Speech and Sound Deepfake Detection Challenge (ESDD2) 2026, which addresses component-level deepfake detection using the CompSpoofV2 dataset, where speech and environmental sounds…

声音 · 计算机科学 2026-05-06 Khalid Zaman , Qixuan Huang , Muhammad Uzair , Masashi Unoki

Segment Anything Model (SAM) has recently shown its powerful effectiveness in visual segmentation tasks. However, there is less exploration concerning how SAM works on audio-visual tasks, such as visual sound localization and segmentation.…

计算机视觉与模式识别 · 计算机科学 2023-05-04 Shentong Mo , Yapeng Tian

Realistic image forgeries involve a combination of splicing, resampling, cloning, region removal and other methods. While resampling detection algorithms are effective in detecting splicing and resampling, copy-move detection algorithms…

PRNU-based image processing is a key asset in digital multimedia forensics. It allows for reliable device identification and effective detection and localization of image forgeries, in very general conditions. However, performance impairs…

计算机视觉与模式识别 · 计算机科学 2020-01-20 Davide Cozzolino , Francesco Marra , Diego Gragnaniello , Giovanni Poggi , Luisa Verdoliva

Automatic speaker verification, like every other biometric system, is vulnerable to spoofing attacks. Using only a few minutes of recorded voice of a genuine client of a speaker verification system, attackers can develop a variety of…

声音 · 计算机科学 2019-06-20 Balamurali BT , Kin Wah Edward Lin , Simon Lui , Jer-Ming Chen , Dorien Herremans

Spin squeezing has been explored in atomic systems as a tool for quantum sensing, improving experimental sensitivity beyond the spin standard quantum limit for certain measurements. To optimize absolute metrological sensitivity, it is…

高能物理 - 唯象学 · 物理学 2025-02-21 Eric Boyers , Garry Goldstein , Alexander O. Sushkov

This work introduces a feature extracted from stereophonic/binaural audio signals aiming to represent a measure of perceived quality degradation in processed spatial auditory scenes. The feature extraction technique is based on a simplified…

音频与语音处理 · 电气工程与系统科学 2022-12-06 Pablo M. Delgado , Jürgen Herre

In this work, a multi-stage Machine Learning (ML) pipeline is proposed for pipe leakage detection in an industrial environment. As opposed to other industrial and urban environments, the environment under study includes many interfering…

机器学习 · 计算机科学 2022-05-06 Ibrahim Shaer , Abdallah Shami

In this paper, we address the problem of image splicing localization with a multi-stream network architecture that processes the raw RGB image in parallel with other handcrafted forensic signals. Unlike previous methods that either use only…

计算机视觉与模式识别 · 计算机科学 2022-12-05 Maria Siopi , Giorgos Kordopatis-Zilos , Polychronis Charitidis , Ioannis Kompatsiaris , Symeon Papadopoulos

Current text-to-speech algorithms produce realistic fakes of human voices, making deepfake detection a much-needed area of research. While researchers have presented various techniques for detecting audio spoofs, it is often unclear exactly…

Environmental Sound Classification (ESC) is an active research area in the audio domain and has seen a lot of progress in the past years. However, many of the existing approaches achieve high accuracy by relying on domain-specific features…

计算机视觉与模式识别 · 计算机科学 2020-04-17 Andrey Guzhov , Federico Raue , Jörn Hees , Andreas Dengel

This paper addresses the challenge of audio-visual single-microphone speech separation and enhancement in the presence of real-world environmental noise. Our approach is based on generative inverse sampling, where we model clean speech and…

音频与语音处理 · 电气工程与系统科学 2026-02-03 Yochai Yemini , Yoav Ellinson , Rami Ben-Ari , Sharon Gannot , Ethan Fetaya

Non-line-of-sight localization in signal-deprived environments is a challenging yet pertinent problem. Acoustic methods in such predominantly indoor scenarios encounter difficulty due to the reverberant nature. In this study, we aim to…

机器学习 · 计算机科学 2024-04-03 Yi Di Yuan , Swee Liang Wong , Jonathan Pan

Characterizing temporally correlated (``non-Markovian'') noise is a key prerequisite for achieving noise-tailored error mitigation and optimal device performance. Quantum noise spectroscopy can afford quantitative estimation of the noise…

量子物理 · 物理学 2024-09-12 Muhammad Qasim Khan , Wenzheng Dong , Leigh M. Norris , Lorenza Viola

Spatial audio quality is a highly multifaceted concept, with many interactions between environmental, geometrical, anatomical, psychological, and contextual considerations. Methods for characterization or evaluation of the geometrical…

音频与语音处理 · 电气工程与系统科学 2024-08-27 Karn N. Watcharasupat , Alexander Lerch

The iterative process of masking minimisation when mixing multitrack audio is a challenging optimisation problem, in part due to the complexity and non-linearity of auditory perception. In this article, we first propose a multitrack masking…

音频与语音处理 · 电气工程与系统科学 2021-01-07 David Ronan , Zheng Ma , Paul Mc Namara , Hatice Gunes , Joshua D. Reiss

Embedding audio signal segments into vectors with fixed dimensionality is attractive because all following processing will be easier and more efficient, for example modeling, classifying or indexing. Audio Word2Vec previously proposed was…

计算与语言 · 计算机科学 2018-11-08 Sung-Feng Huang , Yi-Chen Chen , Hung-yi Lee , Lin-shan Lee

With ever-increasing number of car-mounted electric devices and their complexity, audio classification is increasingly important for the automotive industry as a fundamental tool for human-device interactions. Existing approaches for audio…

声音 · 计算机科学 2018-04-11 Myounggyu Won , Haitham Alsaadan , Yongsoon Eun

Purpose: Surgical scene understanding is key to advancing computer-aided and intelligent surgical systems. Current approaches predominantly rely on visual data or end-to-end learning, which limits fine-grained contextual modeling. This work…