English
Related papers

Related papers: Wanilla: Sound Noninterference Analysis for WebAss…

200 papers

In recent years, developing a speech understanding system that classifies a waveform to structured data, such as intents and slots, without first transcribing the speech to text has emerged as an interesting research problem. This work…

Computer Vision and Pattern Recognition · Computer Science 2020-11-11 Mohamed Mhiri , Samuel Myer , Vikrant Singh Tomar

Speaker modeling is essential for many related tasks, such as speaker recognition and speaker diarization. The dominant modeling approach is fixed-dimensional vector representation, i.e., speaker embedding. This paper introduces a research…

Data hiding is essential for secure communication across digital media, and recent advances in Deep Neural Networks (DNNs) provide enhanced methods for embedding secret information effectively. However, previous audio hiding methods often…

Sound · Computer Science 2025-10-06 Wei Fan , Kejiang Chen , Xiangkun Wang , Weiming Zhang , Nenghai Yu

This letter investigates the problem of anti-jamming communications in dynamic and unknown environment through on-line learning. Different from existing studies which need to know (estimate) the jamming patterns and parameters, we use the…

Information Theory · Computer Science 2017-10-16 Xin Liu , Yuhua Xu , Luliang Jia , Qihui Wu , Alagan Anpalagan

This paper explores the application of artificial intelligence techniques in audio and voice processing, focusing on the integration of wake words and speaker recognition for secure access in embedded systems. With the growing prevalence of…

Non-intrusive load monitoring (NILM) is an advanced load monitoring technique that uses data-driven algorithms to disaggregate the total power consumption of a household into the consumption of individual appliances. However, real-world…

Machine Learning · Computer Science 2025-11-18 Sahar Moghimian Hoosh , Ilia Kamyshev , Henni Ouerdane

Existing weakly supervised semantic segmentation (WSSS) methods usually utilize the results of pre-trained saliency detection (SD) models without explicitly modeling the connections between the two tasks, which is not the most efficient…

Computer Vision and Pattern Recognition · Computer Science 2019-09-11 Yu Zeng , Yunzhi Zhuge , Huchuan Lu , Lihe Zhang

The optimization of a wavelet-based algorithm to improve speech intelligibility along with the full data set and results are reported. The discrete-time speech signal is split into frequency sub-bands via a multi-level discrete wavelet…

Sound · Computer Science 2022-07-25 Tianqu Kang , Anh-Dung Dinh , Binghong Wang , Tianyuan Du , Yijia Chen , Kevin Chau

Modern Large audio-language models (LALMs) power intelligent voice interactions by tightly integrating audio and text. This integration, however, expands the attack surface beyond text and introduces vulnerabilities in the continuous,…

Cryptography and Security · Computer Science 2026-04-17 Meng Chen , Kun Wang , Li Lu , Jiaheng Zhang , Tianwei Zhang

Matching-based methods, especially those based on space-time memory, are significantly ahead of other solutions in semi-supervised video object segmentation (VOS). However, continuously growing and redundant template features lead to an…

Computer Vision and Pattern Recognition · Computer Science 2022-08-23 Zhihui Lin , Tianyu Yang , Maomao Li , Ziyu Wang , Chun Yuan , Wenhao Jiang , Wei Liu

Automatic vulnerability detection on C/C++ source code has benefitted from the introduction of machine learning to the field, with many recent publications targeting this combination. In contrast, assembly language or machine code artifacts…

Cryptography and Security · Computer Science 2023-03-07 Clemens-Alexander Brust , Tim Sonnekalb , Bernd Gruner

Wireless sensor networks (WSN) have recently gained a great deal of attention as a topic of research, with a wide range of applications being explored such as communicating materials. Data dissemination and storage are very important issues…

Networking and Internet Architecture · Computer Science 2015-09-15 Kais Mekki , William Derigent , Eric Rondeau , Ahmed Zouinkhi , Mohamed Naceur Abdelkrim

Currently, the most widely used approach for speaker verification is the deep speaker embedding learning. In this approach, we obtain a speaker embedding vector by pooling single-scale features that are extracted from the last layer of a…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-09 Youngmoon Jung , Seong Min Kye , Yeunju Choi , Myunghun Jung , Hoirin Kim

Existing large language models (LLMs) evaluations use fixed-difficulty benchmarks that cannot adapt as models improve, and rarely isolate specific cognitive processes. We introduce Working Memory Fidelity-Active Manipulation (WMF-AM), a…

Artificial Intelligence · Computer Science 2026-05-05 Dengzhe Hou , Lingyu Jiang , Deng Li , Zirui Li , Fangzhou Lin , Kazunori D Yamada

Metasurfaces constitute effective media for manipulating and transforming impinging EM waves. Related studies have explored a series of impactful MS capabilities and applications in sectors such as wireless communications, medical imaging…

The intellectual property of deep neural network (DNN) models can be protected with DNN watermarking, which embeds copyright watermarks into model parameters (white-box), model behavior (black-box), or model outputs (box-free), and the…

Cryptography and Security · Computer Science 2025-07-25 Haonan An , Guang Hua , Yu Guo , Hangcheng Cao , Susanto Rahardja , Yuguang Fang

Speech-driven large language models (LLMs) are increasingly accessed through speech interfaces, introducing new security risks via open acoustic channels. We present Sirens' Whisper (SWhisper), the first practical framework for covert…

Cryptography and Security · Computer Science 2026-03-17 Zijian Ling , Pingyi Hu , Xiuyong Gao , Xiaojing Ma , Man Zhou , Jun Feng , Songfeng Lu , Dongmei Zhang , Bin Benjamin Zhu

Massive MIMO is one of the salient techniques for achieving high spectral efficiency in next generation wireless networks. Recently, a combined strategy of the massive MIMO and the artificial noise (AN), namely, {\it AN assisted massive…

Information Theory · Computer Science 2018-08-24 Changick Song

Automated speech intelligibility assessment is pivotal for hearing aid (HA) development. In this paper, we present three novel methods to improve intelligibility prediction accuracy and introduce MBI-Net+, an enhanced version of MBI-Net,…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-14 Ryandhimas E. Zezario , Fei Chen , Chiou-Shann Fuh , Hsin-Min Wang , Yu Tsao

Generative images have proliferated on Web platforms in social media and online copyright distribution scenarios, and semantic watermarking has increasingly been integrated into diffusion models to support reliable provenance tracking and…

Machine Learning · Computer Science 2026-02-26 Zheng Gao , Xiaoyu Li , Zhicheng Bao , Xiaoyan Feng , Jiaojiao Jiang
‹ Prev 1 8 9 10 Next ›