中文
相关论文

相关论文: Optimality and limitations of audio-visual integra…

200 篇论文

Current multi-modal models exhibit a notable misalignment with the human visual system when identifying objects that are visually assimilated into the background. Our observations reveal that these multi-modal models cannot distinguish…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Ruolin Shen , Xiaozhong Ji , Kai WU , Jiangning Zhang , Yijun He , HaiHua Yang , Xiaobin Hu , Xiaoyu Sun

In multimodal sentiment analysis (MSA), the performance of a model highly depends on the quality of synthesized embeddings. These embeddings are generated from the upstream process called multimodal fusion, which aims to extract and combine…

计算与语言 · 计算机科学 2021-09-17 Wei Han , Hui Chen , Soujanya Poria

Designers of digital solutions increasingly consult Large Language Models (LLMs) for their work. However, it remains unclear how this may affect the user experiences they produce and there are no established practices. We investigate how…

人机交互 · 计算机科学 2026-05-19 Eduard Kuric , Peter Demcak , Matus Krajcovic

In this study we describe a methodology to realize visual images cognition in the broader sense, by a cross-modal stimulation through the auditory channel. An original algorithm of conversion from bi-dimensional images to sounds has been…

神经元与认知 · 定量生物学 2017-05-16 Takahisa Kishino , Sun Zhe , Roberto Marchisio , Ruggero Micheletto

Educational multimedia has become increasingly important in modern learning environments because of its cost-effectiveness and ability to overcome the temporal and spatial limitations of traditional methods. However, the complex cognitive…

This paper investigates the inverse capabilities and broader utility of multimodal latent spaces within task-specific AI (Artificial Intelligence) models. While these models excel at their designed forward tasks (e.g., text-to-image…

机器学习 · 计算机科学 2025-08-01 Siwoo Park

In recent years, Visual Question Answering (VQA) has made significant strides, particularly with the advent of multimodal models that integrate vision and language understanding. However, existing VQA datasets often overlook the…

计算机视觉与模式识别 · 计算机科学 2024-12-12 Mohammadmostafa Rostamkhani , Baktash Ansari , Hoorieh Sabzevari , Farzan Rahmani , Sauleh Eetemadi

We propose a novel approach to multimodal sensor fusion for Ambient Assisted Living (AAL) which takes advantage of learning using privileged information (LUPI). We address two major shortcomings of standard multimodal approaches, limited…

计算机视觉与模式识别 · 计算机科学 2022-07-15 Alessandro Masullo , Toby Perrett , Tilo Burghardt , Ian Craddock , Dima Damen , Majid Mirmehdi

In recent years, there has been a significant increase in applications of multimodal signal processing and analysis, largely driven by the increased availability of multimodal datasets and the rapid progress in multimodal learning systems.…

图像与视频处理 · 电气工程与系统科学 2024-05-22 Hadi Hadizadeh , S. Faegheh Yeganli , Bahador Rashidi , Ivan V. Bajić

Numerical optimization of complex systems benefits from the technological development of computing platforms in the last twenty years. Unfortunately, this is still not enough, and a large computational time is still necessary when…

最优化与控制 · 数学 2022-09-07 Daniele Peri

Recently, multi-modality scene perception tasks, e.g., image fusion and scene understanding, have attracted widespread attention for intelligent vision systems. However, early efforts always consider boosting a single task unilaterally and…

计算机视觉与模式识别 · 计算机科学 2023-05-12 Zhu Liu , Jinyuan Liu , Guanyao Wu , Long Ma , Xin Fan , Risheng Liu

Binary program comprehension is critical for many use cases but is difficult, suffering from compounded uncertainty and lack of full automation. We seek methods to improve the effectiveness of the human-machine joint cognitive system…

人机交互 · 计算机科学 2024-09-23 Dennis Brown , Emily Mulder , Samuel Mulder

A chief goal of artificial intelligence is to build machines that think like people. Yet it has been argued that deep neural network architectures fail to accomplish this. Researchers have asserted these models' limitations in the domains…

机器学习 · 计算机科学 2024-08-09 Luca M. Schulze Buschoff , Elif Akata , Matthias Bethge , Eric Schulz

Multimodal sentiment analysis is a key technology in the fields of human-computer interaction and affective computing. Accurately recognizing human emotional states is crucial for facilitating smooth communication between humans and…

计算机视觉与模式识别 · 计算机科学 2026-01-07 Wangyuan Zhu , Jun Yu

Multimodal learning allows us to leverage information from multiple sources (visual, acoustic and text), similar to our experience of the real world. However, it is currently unclear to what extent auxiliary modalities improve performance…

计算与语言 · 计算机科学 2020-01-01 Tejas Srinivasan , Ramon Sanabria , Florian Metze

A single neuron is categorized as"multisensory" if there is a statistically significant difference between the response evoked by an audio-visual stimulus combination and that evoked by the most effective of its components individually.…

神经元与认知 · 定量生物学 2016-08-09 Hans Colonius , Adele Diederich

Audio-Video Emotion Recognition is now attacked with Deep Neural Network modeling tools. In published papers, as a rule, the authors show only cases of the superiority in multi-modality over audio-only or video-only modality. However, there…

信号处理 · 电气工程与系统科学 2021-08-02 Xin Chang , Władysław Skarbek

Robotic ultrasound systems can enhance medical diagnostics, but patient acceptance is a challenge. We propose a system combining an AI-powered conversational virtual agent with three mixed reality visualizations to improve trust and…

人机交互 · 计算机科学 2025-07-18 Tianyu Song , Felix Pabst , Ulrich Eck , Nassir Navab

Decoding human visual neural representations is a challenging task with great scientific significance in revealing vision-processing mechanisms and developing brain-like intelligent machines. Most existing methods are difficult to…

计算机视觉与模式识别 · 计算机科学 2023-03-31 Changde Du , Kaicheng Fu , Jinpeng Li , Huiguang He

This paper focuses on a highly practical scenario: how to continue benefiting from the advantages of multi-modal image fusion under harsh conditions when only visible imaging sensors are available. To achieve this goal, we propose a novel…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Hao Zhang , Yanping Zha , Zizhuo Li , Meiqi Gong , Jiayi Ma