English
Related papers

Related papers: Asymmetric Phase Coding Audio Watermarking

200 papers

Automated audio captioning (AAC) is an audio-to-text task to describe audio contents in natural language. Recently, the advancements in large language models (LLMs), with improvements in training approaches for audio encoders, have opened…

Sound · Computer Science 2024-06-26 Jizhong Liu , Gang Li , Junbo Zhang , Heinrich Dinkel , Yongqing Wang , Zhiyong Yan , Yujun Wang , Bin Wang

Neural audio codecs discretize speech via residual vector quantization (RVQ), forming a coarse-to-fine hierarchy across quantizers. While codec models have been explored for representation learning, their discrete structure remains…

Sound · Computer Science 2026-03-19 Jinyang Wu , Zihan Pan , Qiquan Zhang , Sailor Hardik Bhupendra , Soumik Mondal

In recent years, neural networks (NNs) have been widely applied in acoustic echo cancellation (AEC). However, existing approaches struggle to meet real-world low-latency and computational requirements while maintaining performance. To…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-11 Xingchen Li , Boyi Kang , Ziqian Wang , Zihan Zhang , Mingshuai Liu , Zhonghua Fu , Lei Xie

Auditory attention decoding (AAD) identifies the attended speech stream in multi-speaker environments by decoding brain signals such as electroencephalography (EEG). This technology is essential for realizing smart hearing aids that address…

Signal Processing · Electrical Eng. & Systems 2026-01-26 Masahiro Yoshino , Haruki Yokota , Junya Hara , Yuichi Tanaka , Hiroshi Higashi

In this work, we propose novel decoding algorithms to enable streaming automatic speech recognition (ASR) on unsegmented long-form recordings without voice activity detection (VAD), based on monotonic chunkwise attention (MoChA) with an…

Audio and Speech Processing · Electrical Eng. & Systems 2021-07-16 Hirofumi Inaguma , Tatsuya Kawahara

Distributed computing is known as an emerging and efficient technique to support various intelligent services, such as large-scale machine learning. However, privacy leakage and random delays from straggling servers pose significant…

Information Theory · Computer Science 2023-10-31 Qicheng Zeng , Zhaojun Nan , Sheng Zhou

The advancement of artificial intelligence generated content (AIGC) has created a pressing need for robust image watermarking that can withstand both conventional signal processing and novel semantic editing attacks. Current deep…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Yichao Tang , Mingyang Li , Di Miao , Sheng Li , Zhenxing Qian , Xinpeng Zhang

While automated audio captioning (AAC) has made notable progress, traditional fully supervised AAC models still face two critical challenges: the need for expensive audio-text pair data for training and performance degradation when…

Sound · Computer Science 2025-01-07 Xiquan Li , Wenxi Chen , Ziyang Ma , Xuenan Xu , Yuzhe Liang , Zhisheng Zheng , Qiuqiang Kong , Xie Chen

Deep embedding based text-independent speaker verification has demonstrated superior performance to traditional methods in many challenging scenarios. Its loss functions can be generally categorized into two classes, i.e., verification and…

Machine Learning · Computer Science 2019-11-20 Zhongxin Bai , Xiao-Lei Zhang , Jingdong Chen

Generative AI advances rapidly, allowing the creation of very realistic manipulated video and audio. This progress presents a significant security and ethical threat, as malicious users can exploit DeepFake techniques to spread…

Multimedia · Computer Science 2025-06-09 Marcel Klemt , Carlotta Segna , Anna Rohrbach

Quantum machine learning has emerged as a promising tool for pattern recognition, yet many audio-focused approaches still treat spectrograms as generic images and do not explicitly exploit their time-frequency structure. We propose Q-Patch,…

Sound · Computer Science 2026-05-08 Lisan Al Amin , Rakib Hossain , Mahbubul Islam , Faisal Quader , Thanh Thi Nguyen

Generalizability, the capacity of a robust model to perform effectively on unseen data, is crucial for audio deepfake detection due to the rapid evolution of text-to-speech (TTS) and voice conversion (VC) technologies. A promising approach…

Sound · Computer Science 2025-04-16 Botao Zhao , Zuheng Kang , Yayun He , Xiaoyang Qu , Junqing Peng , Jing Xiao , Jianzong Wang

Error correction allows a quantum computer to preserve states long beyond the decoherence time of its physical qubits. Key to any scheme of error correction is the decoding algorithm, which estimates the error state of qubits from the…

Quantum Physics · Physics 2025-01-06 Stasiu Wolanski , Ben Barber

We consider an approach to fault tolerant quantum computing based on a simple error detecting code operating as the substrate for a conventional surface code. We develop a customised decoder to process the information about the likely…

Quantum Physics · Physics 2019-12-11 Xiaosi Xu , Qi Zhao , Xiao Yuan , Simon C. Benjamin

A new class of audio deepfakes-codecfakes (CFs)-has recently caught attention, synthesized by Audio Language Models that leverage neural audio codecs (NACs) in the backend. In response, the community has introduced dedicated benchmarks and…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-17 Orchid Chetia Phukan , Girish , Mohd Mujtaba Akhtar , Arun Balaji Buduru , Rajesh Sharma

In the audio modality, state-of-the-art watermarking methods leverage deep neural networks to allow the embedding of human-imperceptible signatures in generated audio. The ideal is to embed signatures that can be detected with high accuracy…

Sound · Computer Science 2025-04-16 Patrick O'Reilly , Zeyu Jin , Jiaqi Su , Bryan Pardo

Audio watermarking is increasingly used to verify the provenance of AI-generated content, enabling applications such as detecting AI-generated speech, protecting music IP, and defending against voice cloning. To be effective, audio…

Cryptography and Security · Computer Science 2025-03-28 Yizhu Wen , Ashwin Innuganti , Aaron Bien Ramos , Hanqing Guo , Qiben Yan

This work considers the design of short non-binary low-density parity-check (LDPC) codes over finite fields of order m, for channels with phase noise. In particular, m-ary differential phase-shift keying (DPSK) modulated code symbols are…

Information Theory · Computer Science 2019-08-09 Tudor Ninacs , Balázs Matuz , Gianluigi Liva , Giulio Colavolpe

The modern data compression is mainly based on two approaches to entropy coding: Huffman (HC) and arithmetic/range coding (AC). The former is much faster, but approximates probabilities with powers of 2, usually leading to relatively low…

Information Theory · Computer Science 2014-01-07 Jarek Duda

In this paper, we propose a neural-based coding scheme in which an artificial neural network is exploited to automatically compress and decompress speech signals by a trainable approach. Having a two-stage training phase, the system can be…

Sound · Computer Science 2016-01-25 Mahmood Yousefi-Azar , Farbod Razzazi