English
Related papers

Related papers: Real-World En Call Center Transcripts Dataset with…

200 papers

India is home to a multitude of languages of which 22 languages are recognised by the Indian Constitution as official. Building speech based applications for the Indian population is a difficult problem owing to limited data and the number…

Instant messaging with texts and stickers has become a widely adopted communication medium, enabling efficient expression of user semantics and emotions. With the increased use of stickers conveying information and feelings, sticker…

Information Retrieval · Computer Science 2025-07-11 Heng Er Metilda Chee , Jiayin Wang , Zhiqiang Guo , Weizhi Ma , Qinglang Guo , Min Zhang

Emergency Medical Services (EMS) are critical to patient survival in emergencies, but first responders often face intense cognitive demands in high-stakes situations. AI cognitive assistants, acting as virtual partners, have the potential…

We present Aria Everyday Activities (AEA) Dataset, an egocentric multimodal open dataset recorded using Project Aria glasses. AEA contains 143 daily activity sequences recorded by multiple wearers in five geographically diverse indoor…

Intonation is one of the important factors affecting the teaching language arts, so it is an urgent problem to be addressed by evaluating the teachers' intonation through artificial intelligence technology. However, the lack of an…

Sound · Computer Science 2023-12-15 Shuhua Liu , Chunyu Zhang , Binshuai Li , Niantong Qin , Huanting Cheng , Huayu Zhang

Privacy Masking is a critical concept under data privacy involving anonymization and de-anonymization of personally identifiable information (PII). Privacy masking techniques rely on Named Entity Recognition (NER) approaches under NLP…

Computation and Language · Computer Science 2025-04-18 Devansh Singh , Sundaraparipurnan Narayanan

Automated speech recognition (ASR) models have gained prominence for applications such as captioning, speech translation, and live transcription. This paper studies Whisper and two model variants: one optimized for live speech streaming and…

Sound · Computer Science 2025-03-14 Allison Andreyev

Deepfakes represent a growing concern across domains such as disinformation, fraud, and non-consensual media. In particular, the rise of video conference and identity-driven attacks in high-stakes scenarios--such as impostor hiring--demands…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Sarah Barrington , Maty Bohacek , Hany Farid

Speech synthesis systems can now produce highly realistic vocalisations that pose significant authenticity challenges. Despite substantial progress in deepfake detection models, their real-world effectiveness is often undermined by evolving…

Sound · Computer Science 2026-02-12 Qizhou Wang , Hanxun Huang , Guansong Pang , Sarah Erfani , Christopher Leckie

Research in emergency triage is restricted to structured electronic health records (EHR) due to regulatory constraints on nurse-patient interactions. We introduce TriageSim, a simulation framework for generating persona-conditioned triage…

Computation and Language · Computer Science 2026-04-02 Dipankar Srirag , Quoc Dung Nguyen , Aditya Joshi , Padmanesan Narasimhan , Salil Kanhere

We introduce TalkVerse, a large-scale, open corpus for single-person, audio-driven talking video generation designed to enable fair, reproducible comparison across methods. While current state-of-the-art systems rely on closed data or…

Computer Vision and Pattern Recognition · Computer Science 2025-12-18 Zhenzhi Wang , Jian Wang , Ke Ma , Dahua Lin , Bing Zhou

This paper presents a residential audio dataset to support sound event detection research for smart home applications aimed at promoting wellbeing for older adults. The dataset is constructed by deploying audio recording systems in the…

Sound · Computer Science 2024-10-07 Gabriel Bibbó , Thomas Deacon , Arshdeep Singh , Mark D. Plumbley

Speech recognition in highly-reverberant real environments remains a major challenge. An evaluation dataset for this task is needed. This report describes the generation of the Highly-Reverberant Real Environment database (HRRE). This…

Audio and Speech Processing · Electrical Eng. & Systems 2018-03-28 Juan Pablo Escudero , Victor Poblete , José Novoa , Jorge Wuth , Josué Fredes , Rodrigo Mahu , Richard Stern , Néstor Becerra Yoma

Systematic evaluation of speech separation and enhancement models under moving sound source conditions requires extensive and diverse data. However, real-world datasets often lack sufficient data for training and evaluation, and synthetic…

Sound · Computer Science 2025-03-07 Kai Li , Wendi Sang , Chang Zeng , Runxuan Yang , Guo Chen , Xiaolin Hu

The integration of various AI tools creates a complex socio-technical environment where employee-customer interactions form the core of work practices. This study investigates how customer service representatives (CSRs) at the power grid…

Human-Computer Interaction · Computer Science 2025-09-19 Kai Qin , Kexin Du , Yimeng Chen , Yueyan Liu , Jie Cai , Zhiqiang Nie , Nan Gao , Guohui Wei , Shengzhu Wang , Chun Yu

The lack of a publicly-available large-scale and diverse dataset has long been a significant bottleneck for singing voice applications like Singing Voice Synthesis (SVS) and Singing Voice Conversion (SVC). To tackle this problem, we present…

Sound · Computer Science 2025-05-15 Yicheng Gu , Chaoren Wang , Junan Zhang , Xueyao Zhang , Zihao Fang , Haorui He , Zhizheng Wu

By evaluating Large Language Models (LLMs) through uniform, text-only interfaces, current academic benchmarks obscure how the unique designs and affordances of distinct commercial platforms shape real-world user behavior and system…

Computation and Language · Computer Science 2026-05-19 Yueru Yan , Tuc Nguyen , Bo Su , Melissa Lieffers , Thai Le

The preservation of under-resourced languages requires digital tools and resources shaped by and for their speakers. We present the first dedicated ASR resources for Puno Quechua (ISO 639-3: qxp): (1) the largest speech corpus for any…

Computation and Language · Computer Science 2026-05-28 Elwin Huaman , Adrian Gamarra Lafuente , Johanna Cordova , Anna Korhonen

We introduce the first Natural Office Talkers in Settings of Far-field Audio Recordings (``NOTSOFAR-1'') Challenge alongside datasets and baseline system. The challenge focuses on distant speaker diarization and automatic speech recognition…

Healthcare AI holds the potential to increase patient safety, augment efficiency and improve patient outcomes, yet research is often limited by data access, cohort curation, and tooling for analysis. Collection and translation of electronic…

Software Engineering · Computer Science 2021-12-14 Raphael Y. Cohen , Vesela P. Kovacheva
‹ Prev 1 8 9 10 Next ›