English
Related papers

Related papers: Efficient Multimodal Neural Networks for Trigger-l…

200 papers

Machine-generated speech is characterized by its limited or unnatural emotional variation. Current text to speech systems generates speech with either a flat emotion, emotion selected from a predefined set, average variation learned from…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-10 Sarath Sivaprasad , Saiteja Kosgi , Vineet Gandhi

Question answering (QA) systems are designed to answer natural language questions. Visual QA (VQA) and Spoken QA (SQA) systems extend the textual QA system to accept visual and spoken input respectively. This work aims to create a system…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-30 Nimrod Shabtay , Zvi Kons , Avihu Dekel , Hagai Aronowitz , Ron Hoory , Assaf Arbelle

Chatbots via large language models (LLMs) generate fluent responses but often struggle with when to speak, especially for brief, timely listener reactions during ongoing dialogue. We present a multimodal strategy for LLMs, which leverages…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Zikai Liao , Yi Ouyang , Yi-Lun Lee , Chen-Ping Yu , Yi-Hsuan Tsai , Zhaozheng Yin

Humans are capable of processing speech by making use of multiple sensory modalities. For example, the environment where a conversation takes place generally provides semantic and/or acoustic context that helps us to resolve ambiguities or…

Computation and Language · Computer Science 2019-02-21 Ozan Caglayan , Ramon Sanabria , Shruti Palaskar , Loïc Barrault , Florian Metze

Collaborative robots must effectively communicate their internal state to humans to enable a smooth interaction. Nonverbal communication is widely used to communicate information during human-robot interaction, however, such methods may…

Robotics · Computer Science 2024-05-01 Liam Roy , Dana Kulic , Elizabeth Croft

Intelligent conversational agents and virtual assistants, such as chatbots and voice assistants, have the potential of augmenting health service capacity to screen symptoms and deliver healthcare interventions. In this paper, we developed…

Human-Computer Interaction · Computer Science 2022-02-07 Abdalsalam Almzayyen , Angel Vela de la Garza Evia , Nick Coronato , Mehdi Boukhechba

In this paper, we present methods in deep multimodal learning for fusing speech and visual modalities for Audio-Visual Automatic Speech Recognition (AV-ASR). First, we study an approach where uni-modal deep networks are trained separately…

Computation and Language · Computer Science 2015-01-23 Youssef Mroueh , Etienne Marcheret , Vaibhava Goel

Vision and voice are two vital keys for agents' interaction and learning. In this paper, we present a novel indoor navigation model called Memory Vision-Voice Indoor Navigation (MVV-IN), which receives voice commands and analyzes multimodal…

Computer Vision and Pattern Recognition · Computer Science 2020-09-02 Liqi Yan , Dongfang Liu , Yaoxian Song , Changbin Yu

Acoustic sensing has proved effective as a foundation for numerous applications in health and human behavior analysis. In this work, we focus on the problem of detecting in-person social interactions in naturalistic settings from audio…

Sound · Computer Science 2022-03-23 Dawei Liang , Zifan Xu , Yinuo Chen , Rebecca Adaimi , David Harwath , Edison Thomaz

Voice activity detection (VAD) is the task of detecting speech in an audio stream, which is challenging due to numerous unseen noises and low signal-to-noise ratios in real environments. Recently, neural network-based VADs have alleviated…

Sound · Computer Science 2024-05-28 Jidong Jia , Pei Zhao , Di Wang

Zero-shot text-to-speech (TTS) aims to synthesize voices with unseen speech prompts, which significantly reduces the data and computation requirements for voice cloning by skipping the fine-tuning process. However, the prompting mechanisms…

Audio and Speech Processing · Electrical Eng. & Systems 2024-04-11 Ziyue Jiang , Jinglin Liu , Yi Ren , Jinzheng He , Zhenhui Ye , Shengpeng Ji , Qian Yang , Chen Zhang , Pengfei Wei , Chunfeng Wang , Xiang Yin , Zejun Ma , Zhou Zhao

Generating realistic human motions that naturally respond to both spoken language and physical objects is crucial for interactive digital experiences. Current methods, however, address speech-driven gestures or object interactions…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Sreehari Rajan , Kunal Bhosikar , Charu Sharma

Silent and whispered speech offer promise for always-available voice interaction with AI, yet existing methods struggle to balance vocabulary size, wearability, silence, and noise robustness. We present NasoVoce, a nose-bridge-mounted…

Human-Computer Interaction · Computer Science 2026-03-12 Jun Rekimoto , Yu Nishimura , Bojian Yang

Automatic speech recognition (ASR) for conversational code-switching speech remains challenging due to the scarcity of realistic, high-quality labeled speech data. This paper explores multilingual text-to-speech (TTS) models as an effective…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-06 Yue Heng Yeo , Yuchen Hu , Shreyas Gopal , Yizhou Peng , Hexin Liu , Eng Siong Chng

With the goal of more natural and human-like interaction with virtual voice assistants, recent research in the field has focused on full duplex interaction mode without relying on repeated wake-up words. This requires that in scenes with…

Sound · Computer Science 2024-09-17 Anna Wang , Da Liu , Zhiyu Zhang , Shengqiang Liu , Jie Gao , Yali Li

Human-machine interaction has been around for several decades now, with new applications emerging every day. One of the major goals that remain to be achieved is designing an interaction similar to how a human interacts with another human.…

Human-Computer Interaction · Computer Science 2022-12-27 Tauheed Khan Mohd , Nicole Nguyen , Ahmad Y Javaid

Wearable devices, such as smartwatches and head-mounted displays, are increasingly used for prolonged tasks like remote learning and work, but sustained interaction often leads to user fatigue, reducing efficiency and engagement. This study…

Machine Learning · Computer Science 2025-06-17 Yikan Wang

The dialogue experience with conversational agents can be greatly enhanced with multimodal and immersive interactions in virtual reality. In this work, we present an open-source architecture with the goal of simplifying the development of…

Artificial Intelligence · Computer Science 2023-08-08 Michele Yin , Gabriel Roccabruna , Abhinav Azad , Giuseppe Riccardi

Voice-based interfaces rely on a wake-up word mechanism to initiate communication with devices. However, achieving a robust, energy-efficient, and fast detection remains a challenge. This paper addresses these real production needs by…

Sound · Computer Science 2023-10-18 Fernando López , Jordi Luque , Carlos Segura , Pablo Gómez

Robustness against temporal variations is important for emotion recognition from speech audio, since emotion is ex-pressed through complex spectral patterns that can exhibit significant local dilation and compression on the time axis…

Audio and Speech Processing · Electrical Eng. & Systems 2020-03-10 Eric Guizzo , Tillman Weyde , Jack Barnett Leveson