English
Related papers

Related papers: A Bimodal Learning Approach to Assist Multi-sensor…

200 papers

Audiovisual data is everywhere in this digital age, which raises higher requirements for the deep learning models developed on them. To well handle the information of the multi-modal data is the key to a better audiovisual modal. We observe…

Sound · Computer Science 2023-09-27 Meng Liu , Ke Liang , Dayu Hu , Hao Yu , Yue Liu , Lingyuan Meng , Wenxuan Tu , Sihang Zhou , Xinwang Liu

As humans, we experience the world with all our senses or modalities (sound, sight, touch, smell, and taste). We use these modalities, particularly sight and touch, to convey and interpret specific meanings. Multimodal expressions are…

Machine Learning · Computer Science 2022-05-17 Anirudh Sundar , Larry Heck

Despite participants engaging in unimodal stimuli, such as watching images or silent videos, recent work has demonstrated that multi-modal Transformer models can predict visual brain activity impressively well, even with incongruent…

Neurons and Cognition · Quantitative Biology 2025-05-27 Subba Reddy Oota , Khushbu Pahwa , Mounika Marreddy , Maneesh Singh , Manish Gupta , Bapi S. Raju

In modern online learning, understanding and predicting student behavior is crucial for enhancing engagement and optimizing educational outcomes. This systematic review explores the integration of biosensors and Multimodal Learning…

Human-Computer Interaction · Computer Science 2025-09-10 Alvaro Becerra , Ruth Cobos , Charles Lang

In this paper, we introduce a novel audio-visual multi-modal bridging framework that can utilize both audio and visual information, even with uni-modal inputs. We exploit a memory network that stores source (i.e., visual) and target (i.e.,…

Computer Vision and Pattern Recognition · Computer Science 2022-04-05 Minsu Kim , Joanna Hong , Se Jin Park , Yong Man Ro

Word embeddings such as ELMo have recently been shown to model word semantics with greater efficacy through contextualized learning on large-scale language corpora, resulting in significant improvement in state of the art across many…

Computation and Language · Computer Science 2019-09-11 Shao-Yen Tseng , Panayiotis Georgiou , Shrikanth Narayanan

Sensor-aided beamforming reduces the overheads associated with beam training in millimeter-wave (mmWave) multi-input-multi-output (MIMO) communication systems. Most prior work, though, neglects the challenges associated with establishing…

Signal Processing · Electrical Eng. & Systems 2025-09-17 Kartik Patel , Robert W. Heath

Active speaker detection requires a solid integration of multi-modal cues. While individual modalities can approximate a solution, accurate predictions can only be achieved by explicitly fusing the audio and visual features and modeling…

Computer Vision and Pattern Recognition · Computer Science 2021-10-06 Juan León-Alcázar , Fabian Caba Heilbron , Ali Thabet , Bernard Ghanem

We are perceiving and communicating with the world in a multisensory manner, where different information sources are sophisticatedly processed and interpreted by separate parts of the human brain to constitute a complex, yet harmonious and…

Computer Vision and Pattern Recognition · Computer Science 2024-06-12 Ye Zhu , Yu Wu , Nicu Sebe , Yan Yan

Spoken language identification refers to the task of automatically predicting the spoken language in a given utterance. Conventionally, it is modeled as a speech-based language identification task. Prior techniques have been constrained to…

Music generation aims to create music segments that align with human aesthetics based on diverse conditional information. Despite advancements in generating music from specific textual descriptions (e.g., style, genre, instruments), the…

Sound · Computer Science 2025-04-21 Jiahao Song , Yuzhao Wang

Predicting future sensory states is crucial for learning agents such as robots, drones, and autonomous vehicles. In this paper, we couple multiple sensory modalities with exploratory actions and propose a predictive neural network…

Robotics · Computer Science 2021-09-17 Xiaohui Chen , Ramtin Hosseini , Karen Panetta , Jivko Sinapov

Noise pollution significantly affects our daily life and urban development. Urban Sound Tagging (UST) has attracted much attention recently, which aims to analyze and monitor urban noise pollution. One weakness of the previous UST studies…

Audio and Speech Processing · Electrical Eng. & Systems 2023-11-22 Jisheng Bai , Jianfeng Chen , Mou Wang

In order for robots to operate effectively in homes and workplaces, they must be able to manipulate the articulated objects common within environments built for and by humans. Previous work learns kinematic models that prescribe this…

Robotics · Computer Science 2016-07-04 Zhengyang Wu , Mohit Bansal , Matthew R. Walter

In recent years, multi-modal fusion has attracted a lot of research interest, both in academia, and in industry. Multimodal fusion entails the combination of information from a set of different types of sensors. Exploiting complementary…

Machine Learning · Computer Science 2020-08-27 Siddharth Roheda , Hamid Krim , Benjamin S. Riggan

Image-text matching is a key multimodal task that aims to model the semantic association between images and text as a matching relationship. With the advent of the multimedia information age, image, and text data show explosive growth, and…

Machine Learning · Computer Science 2024-06-24 Jinyin Wang , Haijing Zhang , Yihao Zhong , Yingbin Liang , Rongwei Ji , Yiru Cang

Event classification is inherently sequential and multimodal. Therefore, deep neural models need to dynamically focus on the most relevant time window and/or modality of a video. In this study, we propose the Multi-level Attention Fusion…

Computer Vision and Pattern Recognition · Computer Science 2021-06-15 Mathilde Brousmiche , Jean Rouat , Stéphane Dupont

Humans can robustly recognize and localize objects by integrating visual and auditory cues. While machines are able to do the same now with images, less work has been done with sounds. This work develops an approach for dense semantic…

Computer Vision and Pattern Recognition · Computer Science 2020-03-10 Arun Balajee Vasudevan , Dengxin Dai , Luc Van Gool

Attributes of sound inherent to objects can provide valuable cues to learn rich representations for object detection and tracking. Furthermore, the co-occurrence of audiovisual events in videos can be exploited to localize objects over the…

Computer Vision and Pattern Recognition · Computer Science 2021-11-05 Francisco Rivera Valverde , Juana Valeria Hurtado , Abhinav Valada

Our daily perceptual experience is driven by different neural mechanisms that yield multisensory interaction as the interplay between exogenous stimuli and endogenous expectations. While the interaction of multisensory cues according to…

Neurons and Cognition · Quantitative Biology 2018-07-17 German I. Parisi , Jonathan Tong , Pablo Barros , Brigitte Röder , Stefan Wermter