English
Related papers

Related papers: IMPACT: Industrial Machine Perception via Acoustic…

200 papers

Sounds carry an abundance of information about activities and events in our everyday environment, such as traffic noise, road works, music, or people talking. Recent machine learning methods, such as convolutional neural networks (CNNs),…

Sound · Computer Science 2023-05-31 Arshdeep Singh , Haohe Liu , Mark D. Plumbley

Audio pattern recognition is an important research topic in the machine learning area, and includes several tasks such as audio tagging, acoustic scene classification, music classification, speech emotion classification and sound event…

Sound · Computer Science 2020-08-25 Qiuqiang Kong , Yin Cao , Turab Iqbal , Yuxuan Wang , Wenwu Wang , Mark D. Plumbley

Discrete audio tokens are compact representations that aim to preserve perceptual quality, phonetic content, and speaker characteristics while enabling efficient storage and inference, as well as competitive performance across diverse…

Audio classifiers frequently face domain shift, when models trained on one dataset lose accuracy on data recorded in acoustically different conditions. Previous Test-Time Adaptation (TTA) research in speech and sound analysis often…

Sound · Computer Science 2025-11-25 Weichuang Shao , Iman Yi Liao , Tomas Henrique Bode Maul , Tissa Chandesa

Deep noise suppressors (DNS) have become an attractive solution to remove background noise, reverberation, and distortions from speech and are widely used in telephony/voice applications. They are also occasionally prone to introducing…

Sound · Computer Science 2022-04-15 Abu Zaher Md Faridee , Hannes Gamper

In recent years the automotive industry has been strongly promoting the development of smart cars, equipped with multi-modal sensors to gather information about the surroundings, in order to aid human drivers or make autonomous decisions.…

Audio and Speech Processing · Electrical Eng. & Systems 2023-01-31 Jun Yin , Stefano Damiano , Marian Verhelst , Toon van Waterschoot , Andre Guntoro

Analyzing real-world multimodal signals is an essential and challenging task for intelligent voice assistants (IVAs). Mainstream approaches have achieved remarkable performance on various downstream tasks of IVAs with pre-trained audio…

Sound · Computer Science 2024-09-17 Shengqiang Liu , Da Liu , Anna Wang , Zhiyu Zhang , Jie Gao , Yali Li

Anomalous audio in speech recordings is often caused by speaker voice distortion, external noise, or even electric interferences. These obstacles have become a serious problem in some fields, such as high-quality music mixing and speech…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-11 Qiang Huang , Thomas Hain

The Deep Noise Suppression (DNS) challenge is designed to foster innovation in the area of noise suppression to achieve superior perceptual speech quality. We recently organized a DNS challenge special session at INTERSPEECH and ICASSP…

Humans can robustly recognize and localize objects by integrating visual and auditory cues. While machines are able to do the same now with images, less work has been done with sounds. This work develops an approach for dense semantic…

Computer Vision and Pattern Recognition · Computer Science 2020-03-10 Arun Balajee Vasudevan , Dengxin Dai , Luc Van Gool

Distributed acoustic sensing (DAS) technology represents an innovative fiber-optic-based sensing methodology that enables real-time acoustic signal monitoring through the detection of minute perturbations along optical fibers. This sensing…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-27 Shuaikai Shi , Qijun Zong

Audio perception is a key to solving a variety of problems ranging from acoustic scene analysis, music meta-data extraction, recommendation, synthesis and analysis. It can potentially also augment computers in doing tasks that humans do…

Sound · Computer Science 2020-02-12 Prateek Verma , Kenneth Salisbury

Machine learning approaches often require training and evaluation datasets with a clear separation between positive and negative examples. This risks simplifying and even obscuring the inherent subjectivity present in many tasks. Preserving…

Human-Computer Interaction · Computer Science 2023-06-21 Lora Aroyo , Alex S. Taylor , Mark Diaz , Christopher M. Homan , Alicia Parrish , Greg Serapio-Garcia , Vinodkumar Prabhakaran , Ding Wang

This paper presents a multifunctional speech synthesis system that integrates voice cloning and emotion control speech synthesis within a unified framework. The goal of this work is to address longstanding challenges in achieving highly…

Computation and Language · Computer Science 2025-08-15 Fengping Tian , Chenyang Lyu , Xuanfan Ni , Haoqin Sun , Qingjuan Li , Zhiqiang Qian , Haijun Li , Longyue Wang , Zhao Xu , Weihua Luo , Kaifu Zhang

When detecting anomalous sounds in complex environments, one of the main difficulties is that trained models must be sensitive to subtle differences in monitored target signals, while many practical applications also require them to be…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-23 Kevin Wilkinghoff , Takuya Fujimura , Keisuke Imoto , Jonathan Le Roux , Zheng-Hua Tan , Tomoki Toda

Current state-of-the-art automatic speech recognition systems are trained to work in specific `domains', defined based on factors like application, sampling rate and codec. When such recognizers are used in conditions that do not match the…

Unlike human daily activities, existing publicly available sensor datasets for work activity recognition in industrial domains are limited by difficulties in collecting realistic data as close collaboration with industrial sites is…

Computer Vision and Pattern Recognition · Computer Science 2024-04-23 Naoya Yoshimura , Jaime Morales , Takuya Maekawa , Takahiro Hara

Rapid and efficient assessment of the future impact of research articles is a significant concern for both authors and reviewers. The most common standard for measuring the impact of academic papers is the number of citations. In recent…

Computation and Language · Computer Science 2025-03-28 Qichen Sun , Yuxing Lu , Kun Xia , Li Chen , He Sun , Jinzhuo Wang

Text-to-image diffusion models, such as Stable Diffusion, can produce high-quality and diverse images but often fail to achieve compositional alignment, particularly when prompts describe complex object relationships, attributes, or spatial…

Sound is a rich information medium that transmits through air; people communicate through speech and can even discern material through tapping and listening. To capture frequencies in the human hearing range, commercial microphones…

Robotics · Computer Science 2023-07-20 Monica S. Li , Hannah S. Stuart