中文
相关论文

相关论文: CAVEMOVE: An Acoustic Database for the Study of Vo…

200 篇论文

General audio understanding is a fundamental goal for large audio-language models, with audio captioning serving as a cornerstone task for their development. However, progress in this domain is hindered by existing datasets, which lack the…

音频与语音处理 · 电气工程与系统科学 2026-03-26 Yadong Niu , Tianzi Wang , Heinrich Dinkel , Xingwei Sun , Jiahao Zhou , Gang Li , Jizhong Liu , Junbo Zhang , Jian Luan

With the recent advancements in AI, Intelligent Virtual Assistants (IVA) have become a ubiquitous part of every home. Going forward, we are witnessing a confluence of vision, speech and dialog system technologies that are enabling the IVAs…

计算与语言 · 计算机科学 2018-12-21 Shachi H Kumar , Eda Okur , Saurav Sahay , Juan Jose Alvarado Leanos , Jonathan Huang , Lama Nachman

Having knowledge of the environmental context of the user i.e. the knowledge of the users' indoor location and the semantics of their environment, can facilitate the development of many of location-aware applications. In this paper, we…

声音 · 计算机科学 2018-04-03 Muhammad A. Shah , Bhiksha Raj , Khaled A. Harras

Vocal entrainment is a social adaptation mechanism in human interaction, knowledge of which can offer useful insights to an individual's cognitive-behavioral characteristics. We propose a context-aware approach for measuring vocal…

音频与语音处理 · 电气工程与系统科学 2022-11-08 Rimita Lahiri , Md Nasir , Catherine Lord , So Hyun Kim , Shrikanth Narayanan

Voice activity detection is an essential pre-processing component for speech-related tasks such as automatic speech recognition (ASR). Traditional supervised VAD systems obtain frame-level labels from an ASR pipeline by using, e.g., a…

声音 · 计算机科学 2021-05-11 Heinrich Dinkel , Shuai Wang , Xuenan Xu , Mengyue Wu , Kai Yu

In recent years, a great emphasis has been put on engineering the acoustic signature of vehicles that represents the overall comfort level for passengers. Due to highly uncertain behavior of production cars, probabilistic metamodels or…

应用统计 · 统计学 2022-07-06 V. Prakash , O. Sauvage , J. Antoni , L. Gagliardini

Sound is a rich information medium that transmits through air; people communicate through speech and can even discern material through tapping and listening. To capture frequencies in the human hearing range, commercial microphones…

机器人学 · 计算机科学 2023-07-20 Monica S. Li , Hannah S. Stuart

Understanding and estimating driver trust and comfort are essential for the safety and widespread acceptance of autonomous vehicles. Existing works analyze user trust and comfort separately, with limited real-time assessment and…

人机交互 · 计算机科学 2025-08-26 Aditi Bhalla , Christian Hellert , Enkelejda Kasneci , Nastassja Becker

We present an open, reproducible reference for automotive sound quality that connects standardized psychoacoustic metrics with lightweight AI/ML baselines, with a specific focus on electric vehicles (EVs). We implement loudness (ISO…

音频与语音处理 · 电气工程与系统科学 2025-09-23 Mandip Goswami

This paper introduces a new database of voice recordings with the goal of supporting research on vulnerabilities and protection of voice-controlled systems (VCSs). In contrast to prior efforts, the proposed database contains both genuine…

密码学与安全 · 计算机科学 2019-07-03 Yuan Gong , Jian Yang , Jacob Huber , Mitchell MacKnight , Christian Poellabauer

We propose a methodology for training foundation models that enhances their in-context learning capabilities within the domain of bioacoustic signal processing. We use synthetically generated training data, introducing a…

Autonomous underwater vehicles (AUVs) are being tasked with increasingly complex missions. The acoustic communications required for AUVs are, by the nature of the medium, low bandwidth while adverse environmental conditions underwater often…

Sound is an information-rich medium that captures dynamic physical events. This work presents STReSSD, a framework that uses sound to bridge the simulation-to-reality gap for stochastic dynamics, demonstrated for the canonical case of a…

机器人学 · 计算机科学 2020-11-09 Carolyn Matl , Yashraj Narang , Dieter Fox , Ruzena Bajcsy , Fabio Ramos

The time and expense required to collect and label audio data has been a prohibitive factor in the availability of domain specific audio datasets. As the predictive specificity of a classifier depends on the specificity of the labels it is…

声音 · 计算机科学 2023-11-14 Blake Downward , Jon Nordby

Acoustic emotion recognition aims to categorize the affective state of the speaker and is still a difficult task for machine learning models. The difficulties come from the scarcity of training data, general subjectivity in emotion…

计算与语言 · 计算机科学 2018-04-02 Egor Lakomkin , Cornelius Weber , Sven Magg , Stefan Wermter

Silent speech interfaces have been recently proposed as a way to enable communication when the acoustic signal is not available. This introduces the need to build visual speech recognition systems for silent and whispered speech. However,…

计算机视觉与模式识别 · 计算机科学 2018-02-20 Stavros Petridis , Jie Shen , Doruk Cetin , Maja Pantic

We are witnessing a confluence of vision, speech and dialog system technologies that are enabling the IVAs to learn audio-visual groundings of utterances and have conversations with users about the objects, activities and events surrounding…

计算与语言 · 计算机科学 2019-12-27 Shachi H Kumar , Eda Okur , Saurav Sahay , Jonathan Huang , Lama Nachman

The imitation of percussive instruments via the human voice is a natural way for us to communicate rhythmic ideas and, for this reason, it attracts the interest of music makers. Specifically, the automatic mapping of these vocal imitations…

音频与语音处理 · 电气工程与系统科学 2020-09-25 Alejandro Delgado , SKoT McDonald , Ning Xu , Mark Sandler

Autonomous driving is a popular research area within the computer vision research community. Since autonomous vehicles are highly safety-critical, ensuring robustness is essential for real-world deployment. While several public multimodal…

Vision and voice are two vital keys for agents' interaction and learning. In this paper, we present a novel indoor navigation model called Memory Vision-Voice Indoor Navigation (MVV-IN), which receives voice commands and analyzes multimodal…

计算机视觉与模式识别 · 计算机科学 2020-09-02 Liqi Yan , Dongfang Liu , Yaoxian Song , Changbin Yu