中文
相关论文

相关论文: Challenging the Boundaries of Speech Recognition: …

200 篇论文

In multilingual societies, social conversations often involve code-mixed speech. The current speech technology may not be well equipped to extract information from multi-lingual multi-speaker conversations. The DISPLACE challenge entails a…

Recognizing human non-speech vocalizations is an important task and has broad applications such as automatic sound transcription and health condition monitoring. However, existing datasets have a relatively small number of vocal sound…

声音 · 计算机科学 2022-06-22 Yuan Gong , Jin Yu , James Glass

While the last decade has witnessed significant advancements in Automatic Speech Recognition (ASR) systems, performance of these systems for individuals with speech disabilities remains inadequate, partly due to limited public training…

Despite impressive advancements in multilingual corpora collection and model training, developing large-scale deployments of multilingual models still presents a significant challenge. This is particularly true for language tasks that are…

Medical audio data is difficult to collect due to privacy regulations and high annotation costs arising from domain expertise. Thus, existing benchmarks tend to underrepresent complex medical audio scenarios. To address this challenge, we…

Generative AI tools, particularly those utilizing large language models (LLMs), are increasingly used in everyday contexts. While these tools enhance productivity and accessibility, little is known about how Deaf and Hard of Hearing (DHH)…

人机交互 · 计算机科学 2025-01-23 Shuxu Huffman , Si Chen , Kelly Avery Mack , Haotian Su , Qi Wang , Raja Kushalnagar

Whispering is a ubiquitous mode of communication that humans use daily. Despite this, whispered speech has been poorly served by existing speech technology due to a shortage of resources and processing methodology. To remedy this, this…

音频与语音处理 · 电气工程与系统科学 2023-03-15 Pablo Perez Zarazaga , Gustav Eje Henter , Zofia Malisz

Recent advancements in speech-language models have yielded significant improvements in speech tokenization and synthesis. However, effectively mapping the complex, multidimensional attributes of speech into discrete tokens remains…

Speech technology systems struggle with many downstream tasks for child speech due to small training corpora and the difficulties that child speech pose. We apply a novel dataset, SpeechMaturity, to state-of-the-art transformer models to…

计算与语言 · 计算机科学 2025-06-11 Theo Zhang , Madurya Suresh , Anne S. Warlaumont , Kasia Hitczenko , Alejandrina Cristia , Margaret Cychosz

Code-switching automatic speech recognition becomes one of the most challenging and the most valuable scenarios of automatic speech recognition, due to the code-switching phenomenon between multilingual language and the frequent occurrence…

Processing sequential multi-sensor data becomes important in many tasks due to the dramatic increase in the availability of sensors that can acquire sequential data over time. Human Activity Recognition (HAR) is one of the fields which are…

机器学习 · 计算机科学 2020-11-24 Zeyd Boukhers , Danniene Wete , Steffen Staab

Is chatbot able to completely replace the human agent? The short answer could be - "it depends...". For some challenging cases, e.g., dialogue's topical spectrum spreads beyond the training corpus coverage, the chatbot may malfunction and…

计算与语言 · 计算机科学 2020-12-15 Jiawei Liu , Zhe Gao , Yangyang Kang , Zhuoren Jiang , Guoxiu He , Changlong Sun , Xiaozhong Liu , Wei Lu

Hausa Natural Language Processing (NLP) has gained increasing attention in recent years, yet remains understudied as a low-resource language despite having over 120 million first-language (L1) and 80 million second-language (L2) speakers…

Metaphor is a fundamental cognitive mechanism that shapes scientific understanding, enabling the communication of complex concepts while potentially constraining paradigmatic thinking. Despite the prevalence of figurative language in…

计算与语言 · 计算机科学 2025-08-12 Anna Sofia Lippolis , Andrea Giovanni Nuzzolese , Aldo Gangemi

We present ZAEBUC-Spoken, a multilingual multidialectal Arabic-English speech corpus. The corpus comprises twelve hours of Zoom meetings involving multiple speakers role-playing a work situation where Students brainstorm ideas for a certain…

计算与语言 · 计算机科学 2024-03-28 Injy Hamed , Fadhl Eryani , David Palfreyman , Nizar Habash

Speech samples recorded in both indoor and outdoor environments are often contaminated with secondary audio sources. Most end-to-end monaural speech recognition systems either remove these background sounds using speech enhancement or train…

音频与语音处理 · 电气工程与系统科学 2022-02-04 Chaitanya Narisetty , Emiru Tsunoo , Xuankai Chang , Yosuke Kashiwagi , Michael Hentschel , Shinji Watanabe

In order to make spoken dialogue systems (such as Amazon Alexa or Google Assistant) more accessible and naturally interactive for people with cognitive impairments, appropriate data must be obtainable. Recordings of multi-modal spontaneous…

计算与语言 · 计算机科学 2020-10-01 Angus Addlesee , Pierre Albert

As language data and associated technologies proliferate and as the language resources community expands, it is becoming increasingly difficult to locate and reuse existing resources. Are there any lexical resources for such-and-such a…

计算与语言 · 计算机科学 2007-05-23 Steven Bird , Gary Simons

Meeting transcription is a field of high relevance and remarkable progress in recent years. Still, challenges remain that limit its performance. In this work, we extend a previously proposed framework for analyzing leakage in speech…

音频与语音处理 · 电气工程与系统科学 2025-09-15 Peter Vieting , Simon Berger , Thilo von Neumann , Christoph Boeddeker , Ralf Schlüter , Reinhold Haeb-Umbach

Large Audio-Language Models (LALMs), such as GPT-4o, have recently unlocked audio dialogue capabilities, enabling direct spoken exchanges with humans. The potential of LALMs broadens their applicability across a wide range of practical…

人工智能 · 计算机科学 2025-07-29 Kuofeng Gao , Shu-Tao Xia , Ke Xu , Philip Torr , Jindong Gu
‹ 上一页 1 8 9 10 下一页 ›