English
Related papers

Related papers: Towards Perception-Informed Latent HRTF Representa…

200 papers

Convolutional neural network (CNN) architectures have originated and revolutionized machine learning for images. In order to take advantage of CNNs in predictive modeling with audio data, standard FFT-based signal processing methods are…

Sound · Computer Science 2025-02-20 Pavol Harar , Roswitha Bammer , Anna Breger , Monika Dörfler , Zdenek Smekal

Audio-driven talking head synthesis strives to generate lifelike video portraits from provided audio. The diffusion model, recognized for its superior quality and robust generalization, has been explored for this task. However, establishing…

Multimedia · Computer Science 2024-09-17 Fa-Ting Hong , Yunfei Liu , Yu Li , Changyin Zhou , Fei Yu , Dan Xu

In recommender systems, models mostly use a combination of embedding layers and multilayer feedforward neural networks. The high-dimensional sparse original features are downscaled in the embedding layer and then fed into the fully…

Information Retrieval · Computer Science 2022-05-19 Mohan Hasama , Jing Li

Reinforcement Learning from Human Feedback (RLHF) has emerged as a popular paradigm for capturing human intent to alleviate the challenges of hand-crafting the reward values. Despite the increasing interest in RLHF, most works learn black…

Machine Learning · Computer Science 2024-10-14 Akansha Kalra , Daniel S. Brown

Hearable devices, equipped with one or more microphones, are commonly used for speech communication. Here, we consider the scenario where a hearable is used to capture the user's own voice in a noisy environment. In this scenario, own voice…

Audio and Speech Processing · Electrical Eng. & Systems 2025-08-20 Mattes Ohlenbusch , Christian Rollwage , Simon Doclo

Teleconferencing is becoming essential during the COVID-19 pandemic. However, in real-world applications, speech quality can deteriorate due to, for example, background interference, noise, or reverberation. To solve this problem, target…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-02 Yicheng Hsu , Yonghan Lee , Mingsian R. Bai

A learning-based method for estimating the magnitude distribution of sound fields from spatially sparse measurements is proposed. Estimating the magnitude distribution of acoustic transfer function (ATF) is useful when phase measurements…

Sound · Computer Science 2025-06-23 Shoichi Koyama , Kenji Ishizuka

Existing prescriptive compression strategies used in hearing aid fitting are designed based on gain averages from a group of users which are not necessarily optimal for a specific user. Nearly half of hearing aid users prefer settings that…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-02 Nasim Alamdari , Edward Lobarinas , Nasser Kehtarnavaz

Speech foundation models (SFMs) are increasingly hailed as powerful computational models of human speech perception. However, since their representations are inherently black-box, it remains unclear what drives their alignment with brain…

Neurons and Cognition · Quantitative Biology 2025-09-26 Riki Shimizu , Richard J. Antonello , Chandan Singh , Nima Mesgarani

Relative impulse responses between microphones are usually long and dense due to the reverberant acoustic environment. Estimating them from short and noisy recordings poses a long-standing challenge of audio signal processing. In this paper…

Sound · Computer Science 2016-11-17 Zbynek Koldovsky , Jiri Malek , Sharon Gannot

Deep neural networks have shown the ability to extract universal feature representations from data such as images and text that have been useful for a variety of learning tasks. However, the fruits of representation learning have yet to be…

Machine Learning · Computer Science 2023-03-28 Liam Collins , Hamed Hassani , Aryan Mokhtari , Sanjay Shakkottai

Learning representations for feature interactions to model user behaviors is critical for recommendation system and click-trough rate (CTR) predictions. Recent advances in this area are empowered by deep learning methods which could learn…

Information Retrieval · Computer Science 2019-11-26 Canran Xu , Ming Wu

A deep learning framework for dynamically rendering personal sound zones (PSZs) with head tracking is presented, utilizing a spatially adaptive neural network (SANN) that inputs listeners' head coordinates and outputs PSZ filter…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-04 Yue Qiao , Edgar Choueiri

Token representations in high-dimensional latent spaces often exhibit redundancy, limiting computational efficiency and reducing structural coherence across model layers. Hierarchical latent space folding introduces a structured…

Computation and Language · Computer Science 2025-08-11 Fenella Harcourt , Naderdel Piero , Gilbert Sutherland , Daphne Holloway , Harriet Bracknell , Julian Ormsby

Wearable devices such as smartwatches are becoming increasingly popular tools for objectively monitoring physical activity in free-living conditions. To date, research has primarily focused on the purely supervised task of human activity…

Signal Processing · Electrical Eng. & Systems 2021-05-26 Dimitris Spathis , Ignacio Perez-Pozuelo , Soren Brage , Nicholas J. Wareham , Cecilia Mascolo

Reinforcement learning from human feedback (RLHF) has emerged as a central framework for aligning large language models (LLMs) with human preferences. Despite its practical success, RLHF raises fundamental statistical questions because it…

Machine Learning · Statistics 2026-04-06 Pangpang Liu , Chengchun Shi , Will Wei Sun

An important aspect of a humanoid robot is audition. Previous work has presented robot systems capable of sound localization and source segregation based on microphone arrays with various configurations. However, no theoretical framework…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-09 Vladimir Tourbabin , Boaz Rafaely

The increasing accuracy of automatic chord estimation systems, the availability of vast amounts of heterogeneous reference annotations, and insights from annotator subjectivity research make chord label personalization increasingly…

Sound · Computer Science 2017-06-30 H. V. Koops , W. B. de Haas , J. Bransen , A. Volk

Machine hearing or listening represents an emerging area. Conventional approaches rely on the design of handcrafted features specialized to a specific audio task and that can hardly generalized to other audio fields. For example,…

Computer Vision and Pattern Recognition · Computer Science 2018-12-13 Imad Rida , Romain Hérault , Gilles Gasso

Recent advancements in speech-driven 3D talking head generation have made significant progress in lip synchronization. However, existing models still struggle to capture the perceptual alignment between varying speech characteristics and…

Graphics · Computer Science 2025-04-01 Lee Chae-Yeon , Oh Hyun-Bin , Han EunGi , Kim Sung-Bin , Suekyeong Nam , Tae-Hyun Oh
‹ Prev 1 4 5 6 7 8 10 Next ›