English
Related papers

Related papers: Prediction of User Request and Complaint in Spoken…

200 papers

Speech-based analysis offers a scalable and non-invasive approach for detecting cognitive decline, yet progress has been constrained by the limited availability of clinically validated datasets collected under realistic conditions. We…

Recently, increasing research interests have focused on retrieval augmented generation (RAG) to mitigate hallucination for large language models (LLMs). Following this trend, we launch the FutureDial-RAG challenge at SLT 2024, which aims at…

Computation and Language · Computer Science 2024-09-17 Yucheng Cai , Si Chen , Yuxuan Wu , Yi Huang , Junlan Feng , Zhijian Ou

Recent findings show that pre-trained wav2vec 2.0 models are reliable feature extractors for various speaker characteristics classification tasks. We show that latent representations extracted at different layers of a pre-trained wav2vec…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-02 Ilja Baumann , Dominik Wagner , Franziska Braun , Sebastian P. Bayerl , Elmar Nöth , Korbinian Riedhammer , Tobias Bocklet

Audio Word2Vec offers vector representations of fixed dimensionality for variable-length audio segments using Sequence-to-sequence Autoencoder (SA). These vector representations are shown to describe the sequential phonetic structures of…

Computation and Language · Computer Science 2018-02-20 Chia-Hao Shen , Janet Y. Sung , Hung-Yi Lee

Over the years customers' expectation of getting information instantaneously has given rise to the increased usage of channels like virtual assistants. Typically, customers try to get their questions answered by low-touch channels like…

Computation and Language · Computer Science 2021-09-08 Ankush Chopra , Prateek Nagwanshi , Sohom Ghosh

This technical report proposes an audio captioning system for DCASE 2021 Task 6 audio captioning challenge. Our proposed model is based on an encoder-decoder architecture with bi-directional Gated Recurrent Units (BiGRU) using pretrained…

Sound · Computer Science 2021-10-08 Ayşegül Özkaya Eren , Mustafa Sert

This paper introduces Agentic-AI Healthcare, a privacy-aware, multilingual, and explainable research prototype developed as a single-investigator project. The system leverages the emerging Model Context Protocol (MCP) to orchestrate…

Cryptography and Security · Computer Science 2025-10-06 Mohammed A. Shehab

In this paper we introduce ClinQueryAgent, a system for translating natural language population health questions into executable database queries using agents with access to both local and external knowledge bases. Our novel architecture…

Information Retrieval · Computer Science 2026-05-20 Joseph S. Boyle , Anthony Dranfield , Mike O'Neil , Maria Liakata , Alison Q. Smithard

Achieving super-human performance in recognizing human speech has been a goal for several decades, as researchers have worked on increasingly challenging tasks. In the 1990's it was discovered, that conversational speech between two humans…

Computer Vision and Pattern Recognition · Computer Science 2021-07-28 Thai-Son Nguyen , Sebastian Stueker , Alex Waibel

Spoken question answering (SQA) systems are critical for digital assistants and other real-world use cases, but evaluating their performance is a challenge due to the importance of human-spoken questions. This study presents a new…

Computation and Language · Computer Science 2024-02-28 Yijing Wu , SaiKrishna Rallabandi , Ravisutha Srinivasamurthy , Parag Pravin Dakle , Alolika Gon , Preethi Raghavan

The dialogue experience with conversational agents can be greatly enhanced with multimodal and immersive interactions in virtual reality. In this work, we present an open-source architecture with the goal of simplifying the development of…

Artificial Intelligence · Computer Science 2023-08-08 Michele Yin , Gabriel Roccabruna , Abhinav Azad , Giuseppe Riccardi

What do deep neural speech models know about phonology? Existing work has examined the encoding of individual linguistic units such as phonemes in these models. Here we investigate interactions between units. Inspired by classic experiments…

Computation and Language · Computer Science 2024-07-04 Marianne de Heer Kloots , Willem Zuidema

Objective: This study aimed to evaluate which voice features can predict health deterioration in patients with chronic HF. Background: Heart failure (HF) is a chronic condition with progressive deterioration and acute decompensations, often…

We introduce the StatCan Dialogue Dataset consisting of 19,379 conversation turns between agents working at Statistics Canada and online users looking for published data tables. The conversations stem from genuine intents, are held in…

Computation and Language · Computer Science 2024-07-18 Xing Han Lu , Siva Reddy , Harm de Vries

Automated service agents require well-structured workflows to provide consistent and accurate responses to customer queries. However, these workflows are often undocumented, and their automatic extraction from conversations remains…

Computation and Language · Computer Science 2025-02-25 Prafulla Kumar Choubey , Xiangyu Peng , Shilpa Bhagavath , Caiming Xiong , Shiva Kumar Pentyala , Chien-Sheng Wu

Non-goal oriented dialog agents (i.e. chatbots) aim to produce varying and engaging conversations with a user; however, they typically exhibit either inconsistent personality across conversations or the average personality of all users.…

Computation and Language · Computer Science 2020-05-14 Alex Boyd , Raul Puri , Mohammad Shoeybi , Mostofa Patwary , Bryan Catanzaro

Listeners use short interjections, so-called backchannels, to signify attention or express agreement. The automatic analysis of this behavior is of key importance for human conversation analysis and interactive conversational agents.…

Computer Vision and Pattern Recognition · Computer Science 2023-06-05 Ahmed Amer , Chirag Bhuvaneshwara , Gowtham K. Addluri , Mohammed M. Shaik , Vedant Bonde , Philipp Müller

This work presents a practical solution to the problem of call center agent malpractice. A semi-supervised framework comprising of non-linear power transformation, neural feature learning and k-means clustering is outlined. We put these…

Machine Learning · Computer Science 2021-06-07 Şükrü Ozan , Leonardo Obinna Iheme

Current conversational AI systems often provide generic, one-size-fits-all interactions that overlook individual user characteristics and lack adaptive dialogue management. To address this gap, we introduce \textbf{HumAIne-chatbot}, an…

Human-Computer Interaction · Computer Science 2025-09-25 Georgios Makridis , George Fragiadakis , Jorge Oliveira , Tomaz Saraiva , Philip Mavrepis , Georgios Fatouros , Dimosthenis Kyriazis

With its crosslinguistic and cross-speaker diversity, the Mozilla Common Voice Corpus (CV) has been a valuable resource for multilingual speech technology and holds tremendous potential for research in crosslinguistic phonetics and speech…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-03 Miao Zhang , Aref Farhadipour , Annie Baker , Jiachen Ma , Bogdan Pricop , Eleanor Chodroff
‹ Prev 1 8 9 10 Next ›