中文
相关论文

相关论文: CUSIDE: Chunking, Simulating Future Context and De…

200 篇论文

Dense video captioning is a task of localizing interesting events from an untrimmed video and producing textual description (captions) for each localized event. Most of the previous works in dense video captioning are solely based on visual…

计算机视觉与模式识别 · 计算机科学 2020-05-07 Vladimir Iashin , Esa Rahtu

Traditional speaker diarization systems have primarily focused on constrained scenarios such as meetings and interviews, where the number of speakers is limited and acoustic conditions are relatively clean. To explore open-world speaker…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Liangbin Huang , Xiaohua Liao , Chaoqun Cui , Shijing Wang , Zhaolong Huang , Yanlong Du , Wenji Mao

In this work we propose an inference technique, asynchronous revision, to unify streaming and non-streaming speech recognition models. Specifically, we achieve dynamic latency with only one model by using arbitrary right context during…

音频与语音处理 · 电气工程与系统科学 2020-11-04 Mingkun Huang , Meng Cai , Jun Zhang , Yang Zhang , Yongbin You , Yi He , Zejun Ma

Recently, multi-channel speech enhancement has drawn much interest due to the use of spatial information to distinguish target speech from interfering signal. To make full use of spatial information and neural network based masking…

音频与语音处理 · 电气工程与系统科学 2022-10-18 Shubo Lv , Yihui Fu , Yukai Jv , Lei Xie , Weixin Zhu , Wei Rao , Yannan Wang

Zero-shot streaming text-to-speech is an important research topic in human-computer interaction. Existing methods primarily use a lookahead mechanism, relying on future text to achieve natural streaming speech synthesis, which introduces…

机器学习 · 计算机科学 2025-06-03 Haiyang Sun , Shujie Hu , Shujie Liu , Lingwei Meng , Hui Wang , Bing Han , Yifan Yang , Yanqing Liu , Sheng Zhao , Yan Lu , Yanmin Qian

Speaker extraction seeks to extract the target speech in a multi-talker scenario given an auxiliary reference. Such reference can be auditory, i.e., a pre-recorded speech, visual, i.e., lip movements, or contextual, i.e., phonetic sequence.…

计算机视觉与模式识别 · 计算机科学 2022-10-13 Junjie Li , Meng Ge , Zexu Pan , Longbiao Wang , Jianwu Dang

The growing prevalence of online conferences and courses presents a new challenge in improving automatic speech recognition (ASR) with enriched textual information from video slides. In contrast to rare phrase lists, the slides within…

声音 · 计算机科学 2024-01-15 Fan Yu , Haoxu Wang , Xian Shi , Shiliang Zhang

This paper focuses on designing a noise-robust end-to-end Audio-Visual Speech Recognition (AVSR) system. To this end, we propose Visual Context-driven Audio Feature Enhancement module (V-CAFE) to enhance the input noisy audio speech with a…

声音 · 计算机科学 2022-07-14 Joanna Hong , Minsu Kim , Daehun Yoo , Yong Man Ro

This paper introduces a novel approach to speech restoration by integrating a context-related conditioning strategy. Specifically, we employ the diffusion-based generative restoration model, UNIVERSE++, as a backbone to evaluate the…

音频与语音处理 · 电气工程与系统科学 2025-08-13 Soo-Whan Chung , Min-Seok Choi

The representation learning of speech, without textual resources, is an area of significant interest for many low resource speech applications. In this paper, we describe an approach to self-supervised representation learning from raw audio…

音频与语音处理 · 电气工程与系统科学 2023-07-17 Varun Krishna , Tarun Sai , Sriram Ganapathy

Radio speech echo is a specific phenomenon in the air traffic control (ATC) domain, which degrades speech quality and further impacts automatic speech recognition (ASR) accuracy. In this work, a time-domain recognition-oriented speech…

声音 · 计算机科学 2024-07-31 Xincheng Yu , Dongyue Guo , Jianwei Zhang , Yi Lin

Recent progress in auditory intelligence has yielded high-performing systems for sound event detection (SED), acoustic scene classification (ASC), automated audio captioning (AAC), and audio question answering (AQA). Yet these tasks remain…

音频与语音处理 · 电气工程与系统科学 2025-08-12 Hyeonuk Nam

This paper presents our modeling and architecture approaches for building a highly accurate low-latency language identification system to support multilingual spoken queries for voice assistants. A common approach to solve multilingual…

音频与语音处理 · 电气工程与系统科学 2020-06-02 Chander Chandak , Zeynab Raeesy , Ariya Rastrow , Yuzong Liu , Xiangyang Huang , Siyu Wang , Dong Kwon Joo , Roland Maas

Applying speech super-resolution (SR) to recordings with severely low sampling rates is a critical challenge in digital archiving and investigative audio recovery. In these scenarios, the input lacks essential acoustic cues. Consequently,…

声音 · 计算机科学 2025-12-19 Jiajun Yuan , Xiaochen Wang , Yuhang Xiao , Yulin Wu , Chenhao Hu , Xueyang Lv

In this work, we propose a streaming speech recognition framework for Amdo Tibetan, built upon a hybrid CTC/Atten-tion architecture with a context-aware dynamic chunking mechanism. The proposed strategy adaptively adjusts chunk widths based…

计算与语言 · 计算机科学 2025-11-13 Chao Wang , Yuqing Cai , Renzeng Duojie , Jin Zhang , Yutong Liu , Nyima Tashi

Speech perception involves storing and integrating sequentially presented items. Recent work in cognitive neuroscience has identified temporal and contextual characteristics in humans' neural encoding of speech that may facilitate this…

计算与语言 · 计算机科学 2024-05-15 Oli Danyi Liu , Hao Tang , Naomi Feldman , Sharon Goldwater

Automatic Speech Recognition (ASR) models demonstrate outstanding performance on high-resource languages but face significant challenges when applied to low-resource languages due to limited training data and insufficient cross-lingual…

音频与语音处理 · 电气工程与系统科学 2025-06-17 Ming-Hao Hsu , Hung-yi Lee

We propose a novel approach to semi-supervised automatic speech recognition (ASR). We first exploit a large amount of unlabeled audio data via representation learning, where we reconstruct a temporal slice of filterbank features from past…

音频与语音处理 · 电气工程与系统科学 2020-05-15 Shaoshi Ling , Yuzong Liu , Julian Salazar , Katrin Kirchhoff

ASR systems have become increasingly widespread in recent years. However, their textual outputs often require post-processing tasks before they can be practically utilized. To address this issue, we draw inspiration from the multifaceted…

计算与语言 · 计算机科学 2023-09-22 Lei Zhang , Zhengkun Tian , Xiang Chen , Jiaming Sun , Hongyu Xiang , Ke Ding , Guanglu Wan

In a noisy environment, a lossy speech signal can be automatically restored by a listener if he/she knows the language well. That is, with the built-in knowledge of a "language model", a listener may effectively suppress noise interference…

机器学习 · 计算机科学 2019-07-03 Chien-Feng Liao , Yu Tsao , Xugang Lu , Hisashi Kawai