中文
相关论文

相关论文: Causal-Anticausal Decomposition of Speech using Co…

200 篇论文

Topic classification systems on spoken documents usually consist of two modules: an automatic speech recognition (ASR) module to convert speech into text and a text topic classification (TTC) module to predict the topic class from the…

计算与语言 · 计算机科学 2021-06-17 Tan Liu , Wu Guo , Bin Gu

While generative text-to-speech (TTS) models approach human-level quality, monolithic metrics fail to diagnose fine-grained acoustic artifacts or explain perceptual collapse. To address this, we propose TTS-PRISM, a multi-dimensional…

计算与语言 · 计算机科学 2026-04-27 Xi Wang , Jie Wang , Xingchen Song , Baijun Song , Jingran Xie , Jiahe Shao , Zijian Lin , Di Wu , Meng Meng , Jian Luan , Zhiyong Wu

Counterfactual explanation is a form of interpretable machine learning that generates perturbations on a sample to achieve the desired outcome. The generated samples can act as instructions to guide end users on how to observe the desired…

机器学习 · 计算机科学 2023-03-28 Tri Dung Duong , Qian Li , Guandong Xu

In this paper, we propose a spatio-temporal contextual network, STC-Flow, for optical flow estimation. Unlike previous optical flow estimation approaches with local pyramid feature extraction and multi-level correlation, we propose a…

计算机视觉与模式识别 · 计算机科学 2020-11-04 Xiaolin Song , Yuyang Zhao , Jingyu Yang

Adapting to latent confounded shift remains a core challenge in modern AI. This setting is driven by hidden variables that induce spurious correlations between inputs and outputs during training, leading models to rely on non-causal…

机器学习 · 计算机科学 2026-05-14 Jialin Yu , Yuxiang Zhou , Haoxuan Li , Junchi Yu , Mengyue Yang , Yulan He , Nevin L. Zhang , Philip Torr , Ricardo Silva

This paper addresses the problem of under-determinded speech source separation from multichannel microphone singals, i.e. the convolutive mixtures of multiple sources. The time-domain signals are first transformed to the short-time Fourier…

声音 · 计算机科学 2019-04-11 Xiaofei Li , Laurent Girin , Radu Horaud

Dempster-Shafer Theory (DST) generalizes Bayesian probability theory, offering useful additional information, but suffers from a high computational burden. A lot of work has been done to reduce the complexity of computations used in…

人工智能 · 计算机科学 2021-07-15 Maxime Chaveroche , Franck Davoine , Véronique Cherfaoui

Dynamical zeta functions provide a powerful method to analyze low dimensional dynamical systems when the underlying symbolic dynamics is under control. On the other hand even simple one dimensional maps can show an intricate structure of…

混沌动力学 · 物理学 2007-05-23 G. Cristadoro

How can we learn unified representations for spoken utterances and their written text? Learning similar representations for semantically similar speech and text is important for speech translation. To this end, we propose ConST, a…

计算与语言 · 计算机科学 2022-05-06 Rong Ye , Mingxuan Wang , Lei Li

Previous studies demonstrated that a dynamic phone-informed compression of the input audio is beneficial for speech translation (ST). However, they required a dedicated model for phone recognition and did not test this solution for direct…

计算与语言 · 计算机科学 2021-10-15 Marco Gaido , Mauro Cettolo , Matteo Negri , Marco Turchi

The Aspect Sentiment Triplet Extraction (ASTE) task aims to extract aspect terms, opinion terms, and their corresponding sentiment polarity from a given sentence. It remains one of the most prominent subtasks in fine-grained sentiment…

计算与语言 · 计算机科学 2025-02-05 Qingling Li , Wushao Wen , Jinghui Qin

This paper introduces DiFlow-TTS, a novel zero-shot text-to-speech (TTS) system that employs discrete flow matching for generative speech modeling. We position this work as an entry point that may facilitate further advances in this…

End-to-end intent classification using speech has numerous advantages compared to the conventional pipeline approach using automatic speech recognition (ASR), followed by natural language processing modules. It attempts to predict intent…

计算与语言 · 计算机科学 2021-08-06 Yidi Jiang , Bidisha Sharma , Maulik Madhavi , Haizhou Li

Researches have shown accent classification can be improved by integrating semantic information into pure acoustic approach. In this work, we combine phonetic knowledge, such as vowels, with enhanced acoustic features to build an improved…

声音 · 计算机科学 2016-02-25 Zhenhao Ge

Lung sounds refer to the sound generated by air moving through the respiratory system. These sounds, as most biomedical signals, are non-linear and non-stationary. A vital part of using the lung sound for disease detection is discrimination…

声音 · 计算机科学 2020-12-23 Andrine Elsetrønning , Adil Rasheed , Jon Bekker , Omer San

Existing large-scale zero-shot text-to-speech (TTS) models deliver high speech quality but suffer from slow inference speeds due to massive parameters. To address this issue, this paper introduces ZipVoice, a high-quality…

音频与语音处理 · 电气工程与系统科学 2025-08-08 Han Zhu , Wei Kang , Zengwei Yao , Liyong Guo , Fangjun Kuang , Zhaoqing Li , Weiji Zhuang , Long Lin , Daniel Povey

Template-based segmentation, a widely used paradigm in medical imaging, propagates anatomical labels via deformable registration from a labeled atlas to a target image, and is often used to compute volumetric biomarkers for downstream…

图像与视频处理 · 电气工程与系统科学 2026-03-03 Matt Y. Cheung , Ashok Veeraraghavan , Guha Balakrishnan

A well-known but rarely used approach to text categorization uses conditional entropy estimates computed using data compression tools. Text affinity scores derived from compressed sizes can be used for classification and ranking tasks, but…

机器学习 · 计算机科学 2021-12-08 Nitya Kasturi , Igor L. Markov

The complete decomposition performed by blind source separation is computationally demanding and superfluous when only the speech of one specific target speaker is desired. In this paper, we propose a computationally efficient blind speech…

声音 · 计算机科学 2020-08-04 Lele Liao , Zhaoyi Gu , Jing Lu

We study effects of additive white noise on the cepstral representation of speech signals. Distribution of each individual cepstrum coefficient of speech is shown to depend strongly on noise and to overlap significantly with the cepstrum…

计算与语言 · 计算机科学 2007-05-23 Sergei Skorik , Frederic Berthommier