中文
相关论文

相关论文: GCI detection from raw speech using a fully-convol…

200 篇论文

In hyperspectral image (HSI) classification, spatial context has demonstrated its significance in achieving promising performance. However, conventional spatial context-based methods simply assume that spatially neighboring pixels should…

机器学习 · 计算机科学 2019-09-27 Sheng Wan , Chen Gong , Ping Zhong , Shirui Pan , Guangyu Li , Jian Yang

We propose an objective intelligibility measure (OIM), called the Gammachirp Envelope Similarity Index (GESI), that can predict speech intelligibility (SI) in older adults. GESI is a bottom-up model based on psychoacoustic knowledge from…

音频与语音处理 · 电气工程与系统科学 2025-10-30 Ayako Yamamoto , Fuki Miyazaki , Toshio Irino

This paper presents an innovative approach called BGTAI to simplify multimodal understanding by utilizing gloss-based annotation as an intermediate step in aligning Text and Audio with Images. While the dynamic temporal factors in textual…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Sen Fang , Sizhou Chen , Yalin Feng , Xiaofeng Zhang , Teik Toe Teoh

Gravitational wave detectors now under construction are sensitive to the phase of the incident gravitational waves. Correspondingly, the signals from the different detectors can be combined, in the analysis, to simulate a single detector of…

广义相对论与量子宇宙学 · 物理学 2009-12-31 Lee Samuel Finn

Several recent work on speech synthesis have employed generative adversarial networks (GANs) to produce raw waveforms. Although such methods improve the sampling efficiency and memory usage, their sample quality has not yet reached that of…

声音 · 计算机科学 2020-10-26 Jungil Kong , Jaehyeon Kim , Jaekyoung Bae

The collection of individually resolvable gravitational wave (GW) events makes up a tiny fraction of all GW signals which reach our detectors, while most lie below the confusion limit and go undetected. Like voices in a crowded room, the…

广义相对论与量子宇宙学 · 物理学 2022-03-02 Arianna I. Renzini , Boris Goncharov , Alexander C. Jenkins , Pat M. Meyers

The electroencephalography (EEG) signals recorded in parallel with speech are used to perform isolated and continuous speech recognition. During speaking process, one also hears his or her own speech and this speech perception is also…

音频与语音处理 · 电气工程与系统科学 2020-06-03 Gautam Krishna , Co Tran , Mason Carnahan , Ahmed Tewfik

Explainable AI (XAI) underwent a recent surge in research on concept extraction, focusing on extracting human-interpretable concepts from Deep Neural Networks. An important challenge facing concept extraction approaches is the difficulty of…

机器学习 · 计算机科学 2023-02-13 Dmitry Kazhdan , Botty Dimanov , Lucie Charlotte Magister , Pietro Barbiero , Mateja Jamnik , Pietro Lio

Language Identification, being an important aspect of Automatic Speaker Recognition has had many changes and new approaches to ameliorate performance over the last decade. We compare the performance of using audio spectrum in the log scale…

计算与语言 · 计算机科学 2017-05-19 Vrishabh Ajay Lakhani , Rohan Mahadev

We present an exploratory framework to test whether noise-like input can induce structured responses in language models. Instead of assuming that extraterrestrial signals must be decoded, we evaluate whether inputs can trigger linguistic…

天体物理仪器与方法 · 物理学 2025-06-04 Po-Chieh Yu

Prediction of epilepsy based on electroencephalogram (EEG) signals is a rapidly evolving field. Previous studies have traditionally applied 1D processing to the entire EEG signal. However, we have adopted the Gram Matrix method to transform…

机器学习 · 计算机科学 2025-12-16 Bihao You , Jiping Cui

Identifying speakers of quotations in narratives is an important task in literary analysis, with challenging scenarios including the out-of-domain inference for unseen speakers, and non-explicit cases where there are no speaker mentions in…

计算与语言 · 计算机科学 2024-02-20 Zhenlin Su , Liyan Xu , Jin Xu , Jiangnan Li , Mingdu Huangfu

This paper proposes a framework for modeling sound change that combines deep learning and iterative learning. Acquisition and transmission of speech is modeled by training generations of Generative Adversarial Networks (GANs) on unannotated…

计算与语言 · 计算机科学 2021-09-23 Gašper Beguš

Speech produced by human vocal apparatus conveys substantial non-semantic information including the gender of the speaker, voice quality, affective state, abnormalities in the vocal apparatus etc. Such information is attributed to the…

音频与语音处理 · 电气工程与系统科学 2022-08-16 Prathosh A. P. , Varun Srivastava , Mayank Mishra

This paper advances the design of CTC-based all-neural (or end-to-end) speech recognizers. We propose a novel symbol inventory, and a novel iterated-CTC method in which a second system is used to transform a noisy initial output into a…

计算与语言 · 计算机科学 2022-02-24 G. Zweig , C. Yu , J. Droppo , A. Stolcke

This paper introduces a novel method to separate noisy speech into low or high frequency frames, in order to improve fundamental frequency (F0) estimation accuracy. In this proposal, the target signal is analyzed by means of the ensemble…

音频与语音处理 · 电气工程与系统科学 2021-12-21 A. Queiroz , R. Coelho

In this study, we propose a new concept, the gammachirp envelope distortion index (GEDI), based on the signal-to-distortion ratio in the auditory envelope, SDRenv to predict the intelligibility of speech enhanced by nonlinear algorithms.…

声音 · 计算机科学 2020-07-21 Katsuhiko Yamamoto , Toshio Irino , Shoko Araki , Keisuke Kinoshita , Tomohiro Nakatani

For monaural speech enhancement, contextual information is important for accurate speech estimation. However, commonly used convolution neural networks (CNNs) are weak in capturing temporal contexts since they only build blocks that process…

音频与语音处理 · 电气工程与系统科学 2022-07-01 Xinmeng Xu , Yang Wang , Jie Jia , Binbin Chen , Jianjun Hao

Speech emotion recognition is a challenging task for three main reasons: 1) human emotion is abstract, which means it is hard to distinguish; 2) in general, human emotion can only be detected in some specific moments during a long…

声音 · 计算机科学 2019-05-03 Yuanyuan Zhang , Jun Du , Zirui Wang , Jianshu Zhang

Audio tagging aims to assign predefined tags to audio clips to indicate the class information of audio events. Sequential audio tagging (SAT) means detecting both the class information of audio events, and the order in which they occur…

声音 · 计算机科学 2022-10-25 Yuanbo Hou , Yun Wang , Wenwu Wang , Dick Botteldooren