English
Related papers

Related papers: HCU400: An Annotated Dataset for Exploring Aural P…

200 papers

Generative audio models are rapidly advancing in both capabilities and public utilization -- several powerful generative audio models have readily available open weights, and some tech companies have released high quality generative audio…

Storytelling is multi-modal in the real world. When one tells a story, one may use all of the visualizations and sounds along with the story itself. However, prior studies on storytelling datasets and tasks have paid little attention to…

Multimedia · Computer Science 2023-10-31 Jaeyeon Bae , Seokhoon Jeong , Seokun Kang , Namgi Han , Jae-Yon Lee , Hyounghun Kim , Taehwan Kim

This study investigates the aural and visual factors that influence appropriateness perception in soundscape evaluations in residential spaces, where people may spend most of their time in. Appropriateness in soundscape is derived from the…

Applied Physics · Physics 2023-12-29 Johann Kay Ann Tan , Siu-Kit Lau , Yoshimi Hasegawa

Voice signal classification based on human behaviours involves analyzing various aspects of speech patterns and delivery styles. In this study, a real-time dataset collection is performed where participants are instructed to speak twelve…

Sound · Computer Science 2024-07-08 Ali Raza , Faizan Younas

Human auditory perception is compositional in nature -- we identify auditory streams from auditory scenes with multiple sound events. However, such auditory scenes are typically represented using clip-level representations that do not…

Sound · Computer Science 2025-03-04 Sripathi Sridhar , Mark Cartwright

Despite extensive research on toxic speech detection in text, a critical gap remains in handling spoken Mandarin audio. The lack of annotated datasets that capture the unique prosodic cues and culturally specific expressions in Mandarin…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-22 Yu-Xiang Luo , Yi-Cheng Lin , Ming-To Chuang , Jia-Hung Chen , I-Ning Tsai , Pei Xing Kiew , Yueh-Hsuan Huang , Chien-Feng Liu , Yu-Chen Chen , Bo-Han Feng , Wenze Ren , Hung-yi Lee

The localization of sound sources by the human brain is computationally simulated from a neurobiological perspective. The simulation includes the neural representation of temporal differences in acoustic signals between the ipsilateral and…

Neurons and Cognition · Quantitative Biology 2008-10-31 Nikesh S. Dattani

When humans read or listen, they make implicit commonsense inferences that frame their understanding of what happened and why. As a step toward AI systems that can build similar mental models, we introduce GLUCOSE, a large-scale dataset of…

Computation and Language · Computer Science 2020-11-02 Nasrin Mostafazadeh , Aditya Kalyanpur , Lori Moon , David Buchanan , Lauren Berkowitz , Or Biran , Jennifer Chu-Carroll

We consider the problem of audio voice separation for binaural applications, such as earphones and hearing aids. While today's neural networks perform remarkably well (separating $4+$ sources with 2 microphones) they assume a known or fixed…

Sound · Computer Science 2022-07-18 Zhongweiyang Xu , Romit Roy Choudhury

Audio and sound generation has garnered significant attention in recent years, with a primary focus on improving the quality of generated audios. However, there has been limited research on enhancing the diversity of generated audio,…

Sound · Computer Science 2024-03-05 Zeyu Xie , Baihan Li , Xuenan Xu , Mengyue Wu , Kai Yu

Barriers to accessing mental health assessments including cost and stigma continues to be an impediment in mental health diagnosis and treatment. Machine learning approaches based on speech samples could help in this direction. In this…

Computation and Language · Computer Science 2023-12-27 Prabhat Agarwal , Akshat Jindal , Shreya Singh

Medical image segmentation supports clinical workflows by precisely delineating anatomical structures and lesions. However, medical image datasets medical image datasets suffer from acquisition noise and annotation ambiguity, causing…

Artificial Intelligence · Computer Science 2026-04-14 Ruiyang Li , Fang Liu , Licheng Jiao , Xinglin Xie , Jiayao Hao , Shuo Li , Xu Liu , Jingyi Yang , Lingling Li , Puhua Chen , Wenping Ma

Medical audio data is difficult to collect due to privacy regulations and high annotation costs arising from domain expertise. Thus, existing benchmarks tend to underrepresent complex medical audio scenarios. To address this challenge, we…

Automatic sound classification has a wide range of applications in machine listening, enabling context-aware sound processing and understanding. This paper explores methodologies for automatically classifying heterogeneous sounds…

Sound · Computer Science 2024-10-03 Panagiota Anastasopoulou , Jessica Torrey , Xavier Serra , Frederic Font

In this study, the notion of perceptual features is introduced for describing general music properties based on human perception. This is an attempt at rethinking the concept of features, in order to understand the underlying human…

Information Retrieval · Computer Science 2014-04-01 Anders Friberg , Erwin Schoonderwaldt , Anton Hedblad , Marco Fabiani , Anders Elowsson

During social interactions, understanding the intricacies of the context can be vital, particularly for socially anxious individuals. While previous research has found that the presence of a social interaction can be detected from ambient…

Human-Computer Interaction · Computer Science 2024-07-22 Varun Reddy , Zhiyuan Wang , Emma Toner , Max Larrazabal , Mehdi Boukhechba , Bethany A. Teachman , Laura E. Barnes

The neural mechanisms underlying the comprehension of meaningful sounds are yet to be fully understood. While previous research has shown that the auditory cortex can classify auditory stimuli into distinct semantic categories, the specific…

Neurons and Cognition · Quantitative Biology 2023-09-20 Kumar Neelabh , Vishnu Sreekumar

Datasets collected from the open world unavoidably suffer from various forms of randomness or noiseness, leading to the ubiquity of aleatoric (data) uncertainty. Quantifying such uncertainty is particularly pivotal for object detection,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-28 Peng Cui , Guande He , Dan Zhang , Zhijie Deng , Yinpeng Dong , Jun Zhu

We introduce the Free Universal Sound Separation (FUSS) dataset, a new corpus for experiments in separating mixtures of an unknown number of sounds from an open domain of sound types. The dataset consists of 23 hours of single-source audio…

Idioms are figurative expressions whose meanings often cannot be inferred from their individual words, making them difficult to process computationally and posing challenges for human experimental studies. This survey reviews datasets…

Computation and Language · Computer Science 2025-08-19 Michael Flor , Xinyi Liu , Anna Feldman