English
Related papers

Related papers: DESAMO: A Device for Elder-Friendly Smart Homes Po…

200 papers

Simulating dementia patients with large language models (LLMs) is challenging due to the need to jointly model cognitive impairment, emotional dynamics, and nonverbal behaviors over long conversations. We present DemMA, an expert-guided…

Multiagent Systems · Computer Science 2026-01-13 Yutong Song , Jiang Wu , Kazi Sharif , Honghui Xu , Nikil Dutt , Amir Rahmani

With the development of IoT technologies in the past few years, a wide range of smart devices are deployed in a variety of environments aiming to improve the quality of human life in a cost efficient way. Due to the increasingly serious…

Cryptography and Security · Computer Science 2020-01-23 Ming-Chang Lee , Jia-Chun Lin , Olaf Owe

Emotion understanding is a critical yet challenging task. Most existing approaches rely heavily on identity-sensitive information, such as facial expressions and speech, which raises concerns about personal privacy. To address this, we…

Computer Vision and Pattern Recognition · Computer Science 2025-10-28 Deng Li , Bohao Xing , Xin Liu , Baiqiang Xia , Bihan Wen , Heikki Kälviäinen

In the era of Internet of Things (IoT) technologies the potential for privacy invasion is becoming a major concern especially in regards to healthcare data and Ambient Assisted Living (AAL) environments. Systems that offer AAL technologies…

Signal Processing · Electrical Eng. & Systems 2018-02-27 Ismini Psychoula , Erinc Merdivan , Deepika Singh , Liming Chen , Feng Chen , Sten Hanke , Johannes Kropf , Andreas Holzinger , Matthieu Geist

Connecting audio encoders with large language models (LLMs) allows the LLM to perform various audio understanding tasks, such as automatic speech recognition (ASR) and audio captioning (AC). Most research focuses on training an adapter…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-22 Weiqiao Shan , Yuang Li , Yuhao Zhang , Yingfeng Luo , Chen Xu , Xiaofeng Zhao , Long Meng , Yunfei Lu , Min Zhang , Hao Yang , Tong Xiao , Jingbo Zhu

Socially Assistive Robotics (SAR) has shown promise in supporting emotion regulation for neurodivergent children. Recently, there has been increasing interest in leveraging advanced technologies to assist parents in co-regulating emotions…

Human-Computer Interaction · Computer Science 2025-07-15 Jing Li , Felix Schijve , Sheng Li , Yuye Yang , Jun Hu , Emilia Barakova

While modern Text-to-Speech (TTS) systems achieve high fidelity for read-style speech, they struggle to generate Autonomous Sensory Meridian Response (ASMR), a specialized, low-intensity speech style essential for relaxation. The inherent…

Sound · Computer Science 2026-01-23 Leying Zhang , Tingxiao Zhou , Haiyang Sun , Mengxiao Bi , Yanmin Qian

Smart home assistants function best when user commands are direct and well-specified (e.g., "turn on the kitchen light"), or when a hard-coded routine specifies the response. In more natural communication, however, human speech is…

Human-Computer Interaction · Computer Science 2024-01-29 Evan King , Haoxiang Yu , Sangsu Lee , Christine Julien

We introduce DeSTA2.5-Audio, a general-purpose Large Audio Language Model (LALM) designed for robust auditory perception and instruction-following. Recent LALMs augment Large Language Models (LLMs) with auditory capabilities by training on…

The integration of pre-trained text-based large language models (LLM) with speech input has enabled instruction-following capabilities for diverse speech tasks. This integration requires the use of a speech encoder, a speech adapter, and an…

Computation and Language · Computer Science 2024-06-14 Suwon Shon , Kwangyoun Kim , Yi-Te Hsu , Prashant Sridhar , Shinji Watanabe , Karen Livescu

This paper investigates integrating large language models (LLMs) with advanced hardware, focusing on developing a general-purpose device designed for enhanced interaction with LLMs. Initially, we analyze the current landscape, where virtual…

Hardware Architecture · Computer Science 2024-08-21 Jiajun Xu , Qun Wang , Yuhang Cao , Baitao Zeng , Sicheng Liu

Interactions with virtual assistants typically start with a trigger phrase followed by a command. In this work, we explore the possibility of making these interactions more natural by eliminating the need for a trigger phrase. Our goal is…

Most automatic speech processing systems operate in ``open loop'' mode without user feedback about who said what, yet human-in-the-loop workflows can potentially enable higher accuracy. We propose an LLM-assisted in-meeting speaker…

Computation and Language · Computer Science 2026-05-29 Xinlu He , Yiwen Guan , Badrivishal Paurana , Pitipat Kongsomjit , Zilin Dai , Jacob Whitehill

The growing demand for home healthcare calls for tools that can support care delivery. In this study, we explore automatic health assessment from voice using real-world home care visit data, leveraging the diverse patient information it…

Audio and Speech Processing · Electrical Eng. & Systems 2025-10-22 Yu-Wen Chen , William Ho , Sasha M. Vergez , Grace Flaherty , Pallavi Gupta , Zhihong Zhang , Maryam Zolnoori , Margaret V. McDonald , Maxim Topaz , Zoran Kostic , Julia Hirschberg

Natural Language Processing (NLP) and Voice Recognition agents are rapidly evolving healthcare by enabling efficient, accessible, and professional patient support while automating grunt work. This report serves as my self project wherein…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-21 Kabir Kumar

With the rapid development of large language models (LLMs), which possess powerful natural language processing and generation capabilities, LLMs are poised to provide more natural and personalized user experiences. Their deployment on…

Artificial Intelligence · Computer Science 2026-03-03 Lianjun Liu , Hongli An , Pengxuan Chen , Longxiang Ye

Large language models (LLMs) have advanced in text and vision, but their reasoning on audio remains limited. Most existing methods rely on dense audio embeddings, which are difficult to interpret and often fail on structured reasoning…

Sound · Computer Science 2025-11-11 Termeh Taheri , Yinghao Ma , Emmanouil Benetos

Large language models (LLMs) have the potential to revolutionize smart home assistants by enhancing their ability to accurately understand user needs and respond appropriately, which is extremely beneficial for building a smarter home…

Computation and Language · Computer Science 2025-05-28 Silin Li , Yuhang Guo , Jiashu Yao , Zeming Liu , Haifeng Wang

Follow-up conversations with virtual assistants (VAs) enable a user to seamlessly interact with a VA without the need to repeatedly invoke it using a keyword (after the first query). Therefore, accurate Device-directed Speech Detection…

Audio and Speech Processing · Electrical Eng. & Systems 2024-11-06 Ognjen , Rudovic , Pranay Dighe , Yi Su , Vineet Garg , Sameer Dharur , Xiaochuan Niu , Ahmed H. Abdelaziz , Saurabh Adya , Ahmed Tewfik

The maturation of Large Audio Language Models (LALMs) has raised growing expectations for them to comprehend complex audio much like humans. Current efforts primarily replicate text-based reasoning by contextualizing audio content through a…