English
Related papers

Related papers: Concurrent Crossmodal Feedback Assists Target-sear…

200 papers

Interactions with virtual assistants typically start with a trigger phrase followed by a command. In this work, we explore the possibility of making these interactions more natural by eliminating the need for a trigger phrase. Our goal is…

Conversational search is based on a user-system cooperation with the objective to solve an information-seeking task. In this report, we discuss the implication of such cooperation with the learning perspective from both user and system…

Artificial Intelligence · Computer Science 2020-01-10 Sharon Oviatt , Laure Soulier

Conversations contain a wide spectrum of multimodal information that gives us hints about the emotions and moods of the speaker. In this paper, we developed a system that supports humans to analyze conversations. Our main contribution is…

Human-Computer Interaction · Computer Science 2020-01-29 Joshua Y. Kim , Greyson Y. Kim , Kalina Yacef

Concurrent Speaker Detection (CSD), the task of identifying active speakers and their overlaps in an audio signal, is essential for various audio applications, including meeting transcription, speaker diarization, and speech separation.…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-16 Amit Eliav , Sharon Gannot

Efficient motion intent communication is necessary for safe and collaborative work environments with collocated humans and robots. Humans efficiently communicate their motion intent to other humans through gestures, gaze, and social cues.…

Visuo-tactile sensors aim to emulate human tactile perception, enabling robots to precisely understand and manipulate objects. Over time, numerous meticulously designed visuo-tactile sensors have been integrated into robotic systems, aiding…

Machine Learning · Computer Science 2025-04-02 Ruoxuan Feng , Jiangyu Hu , Wenke Xia , Tianci Gao , Ao Shen , Yuhao Sun , Bin Fang , Di Hu

Knowing who is in one's vicinity is key to managing privacy in everyday environments, but is challenging for people with visual impairments. Wearable cameras and other sensors may be able to detect such information, but how should this…

Human-Computer Interaction · Computer Science 2019-04-15 Tousif Ahmed , Rakibul Hasan , Kay Connelly , David Crandall , Apu Kapadia

The need for multimodal data integration arises naturally when multiple complementary sets of features are measured on the same sample. Under a dependent multifactor model, we develop a fully data-driven orchestrated approximate message…

Methodology · Statistics 2026-01-22 Sagnik Nandy , Zongming Ma

Individuals often align their speaking patterns with their interlocutors, a phenomenon linked to engagement and rapport. While well documented in task-oriented dialogues, less is known about entrainment in naturalistic, non-task and virtual…

Human-Computer Interaction · Computer Science 2026-04-20 Thanushi Withanage , Elizabeth Redcay , Carol Espy-Wilson

Information integration from different modalities is an active area of research. Human beings and, in general, biological neural systems are quite adept at using a multitude of signals from different sensory perceptive fields to interact…

Neural and Evolutionary Computing · Computer Science 2021-10-05 Shiv Shankar

This paper presents a novel manipulation strategy that uses keypoint correspondences extracted from visuo-tactile sensor images to facilitate precise object manipulation. Our approach uses the visuo-tactile feedback to guide the robot's…

Robotics · Computer Science 2024-05-24 Jeong-Jung Kim , Doo-Yeol Koh , Chang-Hyun Kim

Aligning machine learning systems with human expectations is mostly attempted by training with manually vetted human behavioral samples, typically explicit feedback. This is done on a population level since the context that is capturing the…

Artificial Intelligence · Computer Science 2025-06-23 Simon Werner , Katharina Christ , Laura Bernardy , Marion G. Müller , Achim Rettinger

In this paper, we propose a novel speech emotion recognition model called Cross Attention Network (CAN) that uses aligned audio and text signals as inputs. It is inspired by the fact that humans recognize speech as a combination of…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-27 Yoonhyung Lee , Seunghyun Yoon , Kyomin Jung

Gaze following estimates gaze targets of in-scene person by understanding human behavior and scene information. Existing methods usually analyze scene images for gaze following. However, compared with visual images, audio also provides…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Yuqi Hou , Zhongqun Zhang , Nora Horanyi , Jaewon Moon , Yihua Cheng , Hyung Jin Chang

Tactile sensing is critical for humans to perform everyday tasks. While significant progress has been made in analyzing object grasping from vision, it remains unclear how we can utilize tactile sensing to reason about and model the…

Recent studies of hearing aid benefits indicate that head movement behavior influences performance. To systematically assess these effects, movement behavior must be measured in realistic communication conditions. For this, the use of…

Medical Physics · Physics 2018-12-06 Maartje M. E. Hendrikse , Gerard Llorach , Giso Grimm , Volker Hohmann

We introduce a multimodal dataset where users express preferences through images. These images encompass a broad spectrum of visual expressions ranging from landscapes to artistic depictions. Users request recommendations for books or music…

Computer Vision and Pattern Recognition · Computer Science 2025-04-17 Se-eun Yoon , Hyunsik Jeon , Julian McAuley

In this study we describe a methodology to realize visual images cognition in the broader sense, by a cross-modal stimulation through the auditory channel. An original algorithm of conversion from bi-dimensional images to sounds has been…

Neurons and Cognition · Quantitative Biology 2017-05-16 Takahisa Kishino , Sun Zhe , Roberto Marchisio , Ruggero Micheletto

Data-driven approaches to tactile sensing aim to overcome the complexity of accurately modeling contact with soft materials. However, their widespread adoption is impaired by concerns about data efficiency and the capability to generalize…

Robotics · Computer Science 2020-03-06 Carmelo Sferrazza , Thomas Bi , Raffaello D'Andrea

Studies on emotion recognition (ER) show that combining lexical and acoustic information results in more robust and accurate models. The majority of the studies focus on settings where both modalities are available in training and…

Computation and Language · Computer Science 2019-06-26 Gustavo Aguilar , Viktor Rozgić , Weiran Wang , Chao Wang