中文
相关论文

相关论文: PlugSonic: a web- and mobile-based platform for bi…

200 篇论文

Objective: Three perceptually orthogonal auditory dimensions for multidimensional and multivariate data sonification are identified and experimentally validated. Background: Psychoacoustic investigations have shown that orthogonal…

声音 · 计算机科学 2020-01-22 Tim Ziemer , Holger Schultheis

Imagine being able to listen to the birds chirping in a park without hearing the chatter from other hikers, or being able to block out traffic noise on a busy street while still being able to hear emergency sirens and car honks. We…

声音 · 计算机科学 2023-11-02 Bandhav Veluri , Malek Itani , Justin Chan , Takuya Yoshioka , Shyamnath Gollakota

Background: The human mind is multimodal. Yet most behavioral studies rely on century-old measures such as task accuracy and latency. To create a better understanding of human behavior and brain functionality, we should introduce other…

机器学习 · 计算机科学 2021-11-30 Moein Razavi , Vahid Janfaza , Takashi Yamauchi , Anton Leontyev , Shanle Longmire-Monford , Joseph Orr

Human perceives rich auditory experience with distinct sound heard by ears. Videos recorded with binaural audio particular simulate how human receives ambient sound. However, a large number of videos are with monaural audio only, which…

声音 · 计算机科学 2021-05-04 Yan-Bo Lin , Yu-Chiang Frank Wang

Currently, artificial intelligence is profoundly transforming the audio domain; however, numerous advanced algorithms and tools remain fragmented, lacking a unified and efficient framework to unlock their full potential. Existing audio…

声音 · 计算机科学 2026-01-01 Cheng Zhu , Jing Han , Qianshuai Xue , Kehan Wang , Huan Zhao , Zixing Zhang

While being disturbed by environmental noises, the acoustic masking technique is a conventional way to reduce the annoyance in audio engineering that seeks to cover up the noises with other dominant yet less intrusive sounds. However,…

声音 · 计算机科学 2026-02-02 Chi Zuo , Martin B. Møller , Pablo Martínez-Nuevo , Huayang Huang , Yu Wu , Ye Zhu

Accurately estimating and simulating the physical properties of objects from real-world sound recordings is of great practical importance in the fields of vision, graphics, and robotics. However, the progress in these directions has been…

声音 · 计算机科学 2024-09-23 Xutong Jin , Chenxi Xu , Ruohan Gao , Jiajun Wu , Guoping Wang , Sheng Li

Imagine being in a crowded space where people speak a different language and having hearables that transform the auditory space into your native language, while preserving the spatial cues for all speakers. We introduce spatial speech…

计算与语言 · 计算机科学 2025-04-29 Tuochao Chen , Qirui Wang , Runlin He , Shyam Gollakota

This paper presents a sensory fusion neuromorphic dataset collected with precise temporal synchronization using a set of Address-Event-Representation sensors and tools. The target application is the lip reading of several keywords for…

We introduce SoundSpaces 2.0, a platform for on-the-fly geometry-based audio rendering for 3D environments. Given a 3D mesh of a real-world environment, SoundSpaces can generate highly realistic acoustics for arbitrary sounds captured from…

This work examines how head-mounted AR can be used to build an interactive sonic landscape to engage with a public sculpture. We describe a sonic artwork, "Listening To Listening", that has been designed to accompany a real-world sculpture…

人机交互 · 计算机科学 2020-12-07 Charles Patrick Martin , Zeruo Liu , Yichen Wang , Wennan He , Henry Gardner

Binaural audio gives the listener an immersive experience and can enhance augmented and virtual reality. However, recording binaural audio requires specialized setup with a dummy human head having microphones in left and right ears. Such a…

计算机视觉与模式识别 · 计算机科学 2021-11-17 Kranti Kumar Parida , Siddharth Srivastava , Gaurav Sharma

Sound is one of the most informative and abundant modalities in the real world while being robust to sense without contacts by small and cheap sensors that can be placed on mobile devices. Although deep learning is capable of extracting…

机器人学 · 计算机科学 2022-08-05 Xufeng Zhao , Cornelius Weber , Muhammad Burhan Hafez , Stefan Wermter

Rapid proliferation of mobile devices with various sensors have enabled Participatory Mobile Sensing (PMS). Several PMS platforms provide multiple functions for various sensing purposes, but they are suffering from the open issues: limited…

社会与信息网络 · 计算机科学 2022-08-24 Yuki Matsuda , Shogo Kawanaka , Hirohiko Suwa , Yutaka Arakawa , Keiichi Yasumoto

Modern autonomous systems are driving the critical need for next-generation adaptive materials and structures with embodied intelligence, i.e., the embodiment of memory, perception, learning, and decision-making within the mechanical…

应用物理 · 物理学 2025-11-18 Yuning Zhang , K. W. Wang

A crucial ability of mobile intelligent agents is to integrate the evidence from multiple sensory inputs in an environment and to make a sequence of actions to reach their goals. In this paper, we attempt to approach the problem of…

计算机视觉与模式识别 · 计算机科学 2020-03-10 Chuang Gan , Yiwei Zhang , Jiajun Wu , Boqing Gong , Joshua B. Tenenbaum

We introduce 3D-Speaker-Toolkit, an open-source toolkit for multimodal speaker verification and diarization, designed for meeting the needs of academic researchers and industrial practitioners. The 3D-Speaker-Toolkit adeptly leverages the…

音频与语音处理 · 电气工程与系统科学 2024-12-30 Yafeng Chen , Siqi Zheng , Hui Wang , Luyao Cheng , Tinglong Zhu , Rongjie Huang , Chong Deng , Qian Chen , Shiliang Zhang , Wen Wang , Xihao Li

Discrete audio tokens are compact representations that aim to preserve perceptual quality, phonetic content, and speaker characteristics while enabling efficient storage and inference, as well as competitive performance across diverse…

We study two foundational problems in audio language models: (1) how to design an audio tokenizer that can serve as an intermediate representation for both understanding and generation; and (2) how to build an audio foundation model that…

声音 · 计算机科学 2026-02-12 Dongchao Yang , Yuanyuan Wang , Dading Chong , Songxiang Liu , Xixin Wu , Helen Meng

In multimedia applications such as films and video games, spatial audio techniques are widely employed to enhance user experiences by simulating 3D sound: transforming mono audio into binaural formats. However, this process is often complex…

多媒体 · 计算机科学 2025-02-14 Xiaojing Liu , Ogulcan Gurelli , Yan Wang , Joshua Reiss