中文
相关论文

相关论文: SHroom: A Python Framework for Ambisonics Room Aco…

200 篇论文

An estimated 253 million people have visual impairments. These visual impairments affect everyday lives, and limit their understanding of the outside world. This can pose a risk to health from falling or collisions. We propose a solution to…

人机交互 · 计算机科学 2023-03-30 Alexander Mehta , Ritik Jalisatgi

We introduce a pioneering unified library that leverages depth anything, segment anything models to augment neural comprehension in language-vision model zero-shot understanding. This library synergizes the capabilities of the Depth…

计算机视觉与模式识别 · 计算机科学 2024-06-28 Mingxiao Huo , Pengliang Ji , Haotian Lin , Junchen Liu , Yixiao Wang , Yijun Chen

We introduce SpinPulse, an open-source python package for simulating spin qubit-based quantum computers at the pulse-level. SpinPulse models the specific physics of spin qubits, particularly through the inclusion of classical non-Markovian…

Audio-signal-processing and audio-machine-learning (ASP/AML) algorithms are ubiquitous in modern technology like smart devices, wearables, and entertainment systems. Development of such algorithms and models typically involves a formal…

音频与语音处理 · 电气工程与系统科学 2026-01-06 Georg Götz , Daniel Gert Nielsen , Steinar Guðjónsson , Finnur Pind

Polarized Resonant Soft X-ray scattering (P-RSoXS) has emerged as a powerful synchrotron-based tool that combines principles of X-ray scattering and X-ray spectroscopy. P-RSoXS provides unique sensitivity to molecular orientation and…

Existing image foundation models are not optimized for spherical images having been trained primarily on perspective images. PanoSAMic integrates the pre-trained Segment Anything (SAM) encoder to make use of its extensive training and…

计算机视觉与模式识别 · 计算机科学 2026-04-27 Mahdi Chamseddine , Didier Stricker , Jason Rambach

The practical deployment of Audio-Visual Speech Recognition (AVSR) systems is fundamentally challenged by significant performance degradation in real-world environments, characterized by unpredictable acoustic noise and visual interference.…

音频与语音处理 · 电气工程与系统科学 2025-12-17 Sungnyun Kim

Magnetron sputtering is an essential technique in combinatorial materials science, enabling the efficient synthesis of thin-film materials libraries with continuous compositional gradients. For exploring multidimensional search spaces,…

材料科学 · 物理学 2024-11-22 Felix Thelen , Rico Zehl , Jan Lukas Bürgel , Diederik Depla , Alfred Ludwig

Spatial filtering has been central in the development of large eddy simulation reduced order models (LES-ROMs) and regularized reduced order models (Reg-ROMs), In this paper, we perform a numerical investigation of spatial filtering. To…

数值分析 · 数学 2017-07-14 L. C. Berselli , D. Wells , X. Xie , T. Iliescu

Segment Anything Model 2 (SAM2) exhibits strong generalisation for promptable segmentation in video clips; however, its integration with the audio modality remains underexplored. Existing approaches either convert audio into visual prompts…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Yuyuan Liu , Yuanhong Chen , Chong Wang , Junlin Han , Junde Wu , Can Peng , Jingkun Chen , Yu Tian , Gustavo Carneiro

Ambisonics rendering has become an integral part of 3D audio for headphones. It works well with existing recording hardware, the processing cost is mostly independent of the number of sound sources, and it elegantly allows for rotating the…

音频与语音处理 · 电气工程与系统科学 2025-06-27 Or Berebi , Fabian Brinkmann , Stefan Weinzierl , Boaz Rafaely

Human and/or asset tracking using an attached sensor units helps understand their activities. Most common indoor localization methods for human tracking technologies require expensive infrastructures, deployment and maintenance. To overcome…

音频与语音处理 · 电气工程与系统科学 2024-03-27 Satoki Ogiso , Yoshiaki Bando , Takeshi Kurata , Takashi Okuma

This paper introduces a contribution made to one of the newest methods for simulating indirect lighting in dynamic scenes , the cascaded light propagation volumes . Our contribution consists on using Spherical Radial Basis Functions instead…

计算机视觉与模式识别 · 计算机科学 2024-07-25 Ludovic Silvestre , João Pereira

A microphone array can provide a mobile robot with the capability of localizing, tracking and separating distant sound sources in 2D, i.e., estimating their relative elevation and azimuth. To combine acoustic data with visual information in…

音频与语音处理 · 电气工程与系统科学 2020-07-23 Simon Michaud , Samuel Faucher , François Grondin , Jean-Samuel Lauzon , Mathieu Labbé , Dominic Létourneau , François Ferland , François Michaud

PySPH is an open-source, Python-based, framework for particle methods in general and Smoothed Particle Hydrodynamics (SPH) in particular. PySPH allows a user to define a complete SPH simulation using pure Python. High-performance code is…

Binaural audio delivers spatial cues essential for immersion, yet most consumer videos are monaural due to capture constraints. We introduce SIREN, a visually guided mono to binaural framework that explicitly predicts left and right…

声音 · 计算机科学 2026-04-01 Mingyeong Song , Seoyeon Ko , Junhyug Noh

Smart glasses are increasingly recognized as a key medium for augmented reality, offering a hands-free platform with integrated microphones and non-ear-occluding loudspeakers to seamlessly mix virtual sound sources into the real-world…

音频与语音处理 · 电气工程与系统科学 2024-09-24 Thomas Deppisch , Nils Meyer-Kahlen , Sebastià V. Amengual Garí

The hierarchical structure of 3D scene graphs shows a high relevance for representations purposes, as it fits common patterns from man-made environments. But, additionally, the semantic and geometric information in such hierarchical…

机器人学 · 计算机科学 2025-10-06 Hriday Bavle , Jose Luis Sanchez-Lopez , Muhammad Shaheer , Javier Civera , Holger Voos

In cluttered environments where visual sensors encounter heavy occlusion, such as in agricultural settings, tactile signals can provide crucial spatial information for the robot to locate rigid objects and maneuver around them. We introduce…

机器人学 · 计算机科学 2024-12-16 Moonyoung Lee , Uksang Yoo , Jean Oh , Jeffrey Ichnowski , George Kantor , Oliver Kroemer

While self-supervised learning (SSL) has revolutionized audio representation, the excessive parameterization and quadratic computational cost of standard Transformers limit their deployment on resource-constrained devices. To address this…

声音 · 计算机科学 2026-03-30 Harunori Kawano , Takeshi Sasaki