English
Related papers

Related papers: EasyCom: An Augmented Reality Dataset to Support A…

200 papers

Wearable devices such as AI glasses are transforming voice assistants into always-available, hands-free collaborators that integrate seamlessly with daily life, but they also introduce challenges like egocentric audio affected by motion and…

The scarcity of large-scale classroom speech data has hindered the development of AI-driven speech models for education. Classroom datasets remain limited and not publicly available, and the absence of dedicated classroom noise or Room…

Sound · Computer Science 2025-10-03 Ahmed Adel Attia , Jing Liu , Carol Espy Wilson

The goal of this paper is speech separation and enhancement in multi-speaker and noisy environments using a combination of different modalities. Previous works have shown good performance when conditioning on temporal or static visual…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-06 Akam Rahimi , Triantafyllos Afouras , Andrew Zisserman

We introduce and analyze a novel approach to the problem of speaker identification in multi-party recorded meetings. Given a speech segment and a set of available candidate profiles, we propose a novel data-driven way to model the distance…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-23 Nikolaos Flemotomos , Dimitrios Dimitriadis

The NeurIPS 2023 Machine Learning for Audio Workshop brings together machine learning (ML) experts from various audio domains. There are several valuable audio-driven ML tasks, from speech emotion recognition to audio event detection, but…

The scarcity of large-scale classroom speech data has hindered the development of AI-driven speech models for education. Public classroom datasets remain limited, and the lack of a dedicated classroom noise corpus prevents the use of…

Sound · Computer Science 2025-06-12 Ahmed Adel Attia , Jing Liu , Carl Espy-Wilson

We introduce EgoLife, a project to develop an egocentric life assistant that accompanies and enhances personal efficiency through AI-powered wearable glasses. To lay the foundation for this assistant, we conducted a comprehensive data…

One of the most challenging scenarios for smart speakers is multi-talker, when target speech from the desired speaker is mixed with interfering speech from one or more speakers. A smart assistant needs to determine which voice to recognize…

Audio and Speech Processing · Electrical Eng. & Systems 2022-05-19 Joe Caroselli , Arun Narayanan , Yiteng Huang

The rendering and display of text is a key use-case for augmented reality (AR). Here, we present the Read-AR, a dataset of reading in AR, for which we collected over 11,000 reading speeds and almost 6000 visual quality and comfort ratings…

Human-Computer Interaction · Computer Science 2026-05-14 Minjung Kim , Saeideh Ghahghaei Nezamabadi , Trisha Lian , Anand Singh

Interacting with a significant number of individuals on a daily basis is commonplace for many professionals, which can lead to challenges in recalling specific details: Who is this person? What did we talk about last time? The advant of…

Human-Computer Interaction · Computer Science 2025-05-20 Raphaël A. El Haddad , Zeyu Wang , Yeonsu Shin , Ranyi Liu , Yuntao Wang , Chun Yu

Augmented Reality (AR) solutions are providing tools that could improve applications in the medical and industrial fields. Augmentation can provide additional information in training, visualization, and work scenarios, to increase…

Human-Computer Interaction · Computer Science 2023-06-30 Akos Nagy , Thomas Lagkas , Panagiotis Sarigiannidis , Vasileios Argyriou

Group conversations are valuable for second language (L2) learners as they provide opportunities to practice listening and speaking, exercise complex turn-taking skills, and experience group social dynamics in a target language. However,…

Human-Computer Interaction · Computer Science 2025-06-02 Jad Bendarkawi , Ashley Ponce , Sean Mata , Aminah Aliu , Yuhan Liu , Lei Zhang , Amna Liaqat , Varun Nagaraj Rao , Andrés Monroy-Hernández

The SEAR Dataset is a novel multimodal resource designed to study the emerging threat of social engineering (SE) attacks orchestrated through augmented reality (AR) and multimodal large language models (LLMs). This dataset captures 180…

Artificial Intelligence · Computer Science 2025-06-02 Tianlong Yu , Chenghang Ye , Zheyu Yang , Ziyi Zhou , Cui Tang , Zui Tao , Jun Zhang , Kailong Wang , Liting Zhou , Yang Yang , Ting Bi

Multi-modal learning in the audio-language domain has seen significant advancements in recent years. However, audio-language learning faces challenges due to limited and lower-quality data compared to image-language tasks. Existing…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-10 David Xu

Data augmentation is a widely used technique for improving model performance in machine learning, particularly in computer vision and natural language processing. Recently, there has been increasing interest in applying augmentation…

Machine Learning · Computer Science 2023-05-05 Raad Khraishi , Ramin Okhrati

Recently, the AI community has made significant strides in developing powerful foundation models, driven by large-scale multimodal datasets. However, for audio representation learning, existing datasets suffer from limitations in the…

Sound · Computer Science 2024-09-10 Luoyi Sun , Xuenan Xu , Mengyue Wu , Weidi Xie

This paper addresses the challenge of speaker separation, which remains an active research topic despite the promising results achieved in recent years. These results, however, often degrade in real recording conditions due to the presence…

Sound · Computer Science 2024-11-14 Rawad Melhem , Assef Jafar , Oumayma Al Dakkak

Engaging in smooth conversations with others is a crucial social skill. However, differences in knowledge between conversation participants can sometimes hinder effective communication. To tackle this issue, this study proposes a real-time…

Human-Computer Interaction · Computer Science 2025-06-23 Yuichiro Fujimoto

Augmented reality (AR) games, particularly those designed for head-mounted displays, have grown increasingly prevalent. However, most existing systems depend on pre-scanned, static environments and rely heavily on continuous tracking or…

Human-Computer Interaction · Computer Science 2026-02-06 Liuchuan Yu , Ching-I Huang , Hsueh-Cheng Wang , Lap-Fai Yu

The rapid advances in audio analysis underscore its vast potential for humancomputer interaction, environmental monitoring, and public safety; yet, existing audioonly datasets often lack spatial context. To address this gap, we present two…

Sound · Computer Science 2025-12-10 Shuaihang Yuan , Congcong Wen , Muhammad Shafique , Anthony Tzes , Yi Fang