中文
相关论文

相关论文: GCI detection from raw speech using a fully-convol…

200 篇论文

Detecting synthetic from real speech is increasingly crucial due to the risks of misinformation and identity impersonation. While various datasets for synthetic speech analysis have been developed, they often focus on specific areas,…

声音 · 计算机科学 2025-07-18 Zhoulin Ji , Chenhao Lin , Hang Wang , Chao Shen

This thesis presents advancements in the detection of gravitational waves from compact binary coalescences, utilising the most sensitive observatories constructed to date. The research focuses on enhancing gravitational-wave signal searches…

广义相对论与量子宇宙学 · 物理学 2026-01-27 Arthur Tolley

Starting with a collection of traces generated by process executions, process discovery is the task of constructing a simple model that describes the process, where simplicity is often measured in terms of model size. The challenge of…

人工智能 · 计算机科学 2024-04-17 Hanan Alkhammash , Artem Polyvyanyy , Alistair Moffat

We propose a method to reduce false voice triggers of a speech-enabled personal assistant by post-processing the hypothesis lattice of a server-side large-vocabulary continuous speech recognizer (LVCSR) via a neural network. We first…

计算与语言 · 计算机科学 2020-03-03 Woojay Jeon , Leo Liu , Henry Mason

3D open-vocabulary scene understanding, crucial for advancing augmented reality and robotic applications, involves interpreting and locating specific regions within a 3D space as directed by natural language instructions. To this end, we…

计算机视觉与模式识别 · 计算机科学 2024-07-30 Yansong Qu , Shaohui Dai , Xinyang Li , Jianghang Lin , Liujuan Cao , Shengchuan Zhang , Rongrong Ji

Intrinsic image decomposition is a challenging, long-standing computer vision problem for which ground truth data is very difficult to acquire. We explore the use of synthetic data for training CNN-based intrinsic image decomposition…

计算机视觉与模式识别 · 计算机科学 2018-12-07 Zhengqi Li , Noah Snavely

Mask processing in the time-frequency (T-F) domain through the neural network has been one of the mainstreams for single-channel speech enhancement. However, it is hard for most models to handle the situation when harmonics are partially…

音频与语音处理 · 电气工程与系统科学 2022-03-17 Tianrui Wang , Weibin Zhu , Yingying Gao , Junlan Feng , Shilei Zhang

Attempts to develop speech enhancement algorithms with improved speech intelligibility for cochlear implant (CI) users have met with limited success. To improve speech enhancement methods for CI users, we propose to perform speech…

声音 · 计算机科学 2019-07-08 Nursadul Mamun , Soheil Khorram , John H. L. Hansen

Contemporary Human Computer Interaction (HCI) research relies primarily on neural network models for machine vision and speech understanding of a system user. Such models require extensively annotated training datasets for optimal…

人机交互 · 计算机科学 2023-11-14 Muhammad Ali Farooq , Dan Bigioi , Rishabh Jain , Wang Yao , Mariam Yiwere , Peter Corcoran

We present a novel approach to improve the performance of learning-based speech dereverberation using accurate synthetic datasets. Our approach is designed to recover the reverb-free signal from a reverberant speech signal. We show that…

音频与语音处理 · 电气工程与系统科学 2022-12-13 Rohith Aralikatti , Zhenyu Tang , Dinesh Manocha

Automatic emotion recognition for real-life appli-cations is a challenging task. Human emotion expressions aresubtle, and can be conveyed by a combination of several emo-tions. In most existing emotion recognition studies, each…

声音 · 计算机科学 2022-03-08 Jay Desai , Houwei Cao , Ravi Shah

Brain-Computer Interfaces (BCIs) rely on accurately decoding electroencephalography (EEG) motor imagery (MI) signals for effective device control. Graph Neural Networks (GNNs) outperform Convolutional Neural Networks (CNNs) in this regard,…

信号处理 · 电气工程与系统科学 2024-05-03 Htoo Wai Aung , Jiao Jiao Li , Yang An , Steven W. Su

Incremental text-to-speech (TTS) synthesis generates utterances in small linguistic units for the sake of real-time and low-latency applications. We previously proposed an incremental TTS method that leverages a large pre-trained language…

声音 · 计算机科学 2021-09-23 Takaaki Saeki , Shinnosuke Takamichi , Hiroshi Saruwatari

The intelligibility of natural speech is seriously degraded when exposed to adverse noisy environments. In this work, we propose a deep learning-based speech modification method to compensate for the intelligibility loss, with the…

音频与语音处理 · 电气工程与系统科学 2020-04-08 Haoyu Li , Szu-Wei Fu , Yu Tsao , Junichi Yamagishi

The recent developments in technology have re-warded us with amazing audio synthesis models like TACOTRON and WAVENETS. On the other side, it poses greater threats such as speech clones and deep fakes, that may go undetected. To tackle…

机器学习 · 计算机科学 2021-07-27 Arun Kumar Singh , Priyanka Singh , Karan Nathwani

To progress in the characterization of noise for current quantum computers, gate set tomography (GST) has emerged as a self-consistent tomographic protocol that can accurately estimate the complete set of noisy quantum gates, state…

量子物理 · 物理学 2025-09-25 Pablo Viñas , Alejandro Bermudez

Grammatical error correction (GEC) is one of the areas in natural language processing in which purely neural models have not yet superseded more traditional symbolic models. Hybrid systems combining phrase-based statistical machine…

计算与语言 · 计算机科学 2019-04-08 Felix Stahlberg , Christopher Bryant , Bill Byrne

A brain-computer interface (BCI) based on the motor imagery (MI) paradigm translates one's motor intention into a control signal by classifying the Electroencephalogram (EEG) signal of different tasks. However, most existing systems either…

数据结构与算法 · 计算机科学 2020-07-27 Eitan Netzer , Alex Frid , Dan Feldman

An important and difficult task in code-switched speech recognition is to recognize the language, as lots of words in two languages can sound similar, especially in some accents. We focus on improving performance of end-to-end Automatic…

计算与语言 · 计算机科学 2024-03-14 Yash Sharma , Basil Abraham , Preethi Jyothi

Language models (LMs) are increasingly extended with new learnable vocabulary tokens for domain-specific tasks, such as Semantic-ID tokens in generative recommendation. The standard practice initializes these new tokens as the mean of…