English
Related papers

Related papers: Modality-Agnostic fMRI Decoding of Vision and Lang…

200 papers

Music information is often conveyed or recorded across multiple data modalities including but not limited to audio, images, text and scores. However, music information retrieval research has almost exclusively focused on single modality…

Sound · Computer Science 2021-06-03 Ho-Hsiang Wu , Magdalena Fuentes , Juan P. Bello

Reconstructing visual stimulus (image) only from human brain activity measured with functional Magnetic Resonance Imaging (fMRI) is a significant and meaningful task in Human-AI collaboration. However, the inconsistent distribution and…

Computer Vision and Pattern Recognition · Computer Science 2019-10-23 Ziqi Ren , Jie Li , Xuetong Xue , Xin Li , Fan Yang , Zhicheng Jiao , Xinbo Gao

Despite significant strides in visual quality assessment, the neural mechanisms underlying visual quality perception remain insufficiently explored. This study employed fMRI to examine brain activity during image quality assessment and…

Multimedia · Computer Science 2024-04-30 Yiming Zhang , Ying Hu , Xiongkuo Min , Yan Zhou , Guangtao Zhai

Clarifying the neural basis of speech intelligibility is critical for computational neuroscience and digital speech processing. Recent neuroimaging studies have shown that intelligibility modulates cortical activity beyond simple acoustics,…

Neurons and Cognition · Quantitative Biology 2025-11-05 Ching-Chih Sung , Shuntaro Suzuki , Francis Pingfan Chien , Komei Sugiura , Yu Tsao

Decoding visual stimuli from neural recordings is a critical challenge in the development of brain-computer interfaces (BCIs). Although recent EEG-based decoding approaches have made progress in tasks such as visual classification,…

Human-Computer Interaction · Computer Science 2024-12-31 Dongyang Li , Haoyang Qin , Mingyang Wu , Jiahua Tang , Yuang Cao , Chen Wei , Quanying Liu

Despite the impressive advancements achieved through vision-and-language pretraining, it remains unclear whether this joint learning paradigm can help understand each individual modality. In this work, we conduct a comparative analysis of…

Computer Vision and Pattern Recognition · Computer Science 2024-01-31 Zhuowan Li , Cihang Xie , Benjamin Van Durme , Alan Yuille

The human visual system is capable of processing continuous streams of visual information, but how the brain encodes and retrieves recent visual memories during continuous visual processing remains unexplored. This study investigates the…

Computation and Language · Computer Science 2024-10-01 Runze Xia , Congchi Yin , Piji Li

The intrication of brain signals drives research that leverages multimodal AI to align brain modalities with visual and textual data for explainable descriptions. However, most existing studies are limited to coarse interpretations, lacking…

Computer Vision and Pattern Recognition · Computer Science 2025-05-22 Weihao Xia , Cengiz Oztireli

Numerous studies have shown that multimodal LLMs process speech and images well but fail in non-intuitive ways rendering trivial tasks such as object counting unreliable. We investigate this behavior from an information-theoretic…

Computation and Language · Computer Science 2026-03-09 Jayadev Billa

In this study, we adopted visual motion imagery, which is a more intuitive brain-computer interface (BCI) paradigm, for decoding the intuitive user intention. We developed a 3-dimensional BCI training platform and applied it to assist the…

Signal Processing · Electrical Eng. & Systems 2020-05-19 Byoung-Hee Kwon , Ji-Hoon Jeong , Jeong-Hyun Cho , Seong-Whan Lee

While computer vision models have made incredible strides in static image recognition, they still do not match human performance in tasks that require the understanding of complex, dynamic motion. This is notably true for real-world…

Neurons and Cognition · Quantitative Biology 2025-04-09 Jacob Yeung , Andrew F. Luo , Gabriel Sarch , Margaret M. Henderson , Deva Ramanan , Michael J. Tarr

Decoding visual signals holds the tantalizing potential to unravel the complexities of cognition and perception. While recent studies have focused on reconstructing visual stimuli from neural recordings to bridge brain activity with visual…

Computational Engineering, Finance, and Science · Computer Science 2025-09-23 Zixiang Yin , Jiarui Li , Zhengming Ding

Objective. In this paper, we consider the problem of cross-subject decoding, where neural activity data collected from the prefrontal cortex of a given subject (destination) is used to decode motor intentions from the neural activity of a…

Neural and Evolutionary Computing · Computer Science 2022-02-23 Marko Angjelichinoski , Bijan Pesaran , Vahid Tarokh

This work introduces a novel approach to fMRI-based visual image reconstruction using a subject-agnostic common representation space. We show that the brain signals of the subjects can be aligned in this common space during training to form…

Image and Video Processing · Electrical Eng. & Systems 2025-10-10 Christos Zangos , Danish Ebadulla , Thomas Christopher Sprague , Ambuj Singh

We investigate optimal strategies for decoding perceived natural speech from fMRI data acquired from a limited number of participants. Leveraging Lebel et al. (2023)'s dataset of 8 participants, we first demonstrate the effectiveness of…

Neurons and Cognition · Quantitative Biology 2025-05-28 Louis Jalouzot , Alexis Thual , Yair Lakretz , Christophe Pallier , Bertrand Thirion

Integrating information from multiple modalities is arguably one of the essential prerequisites for grounding artificial intelligence systems with an understanding of the real world. Recent advances in video transformers that jointly learn…

Computer Vision and Pattern Recognition · Computer Science 2023-11-15 Dota Tianai Dong , Mariya Toneva

Camouflaged Object Detection (COD) aims to segment objects that blend seamlessly into complex backgrounds, with growing interest in exploiting additional visual modalities to enhance robustness through complementary information. However,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Hao Wang , Jiqing Zhang , Xin Yang , Baocai Yin , Lu Jiang , Zetian Mi , Huibing Wang

Vision-Language Models (VLMs) have shown solid ability for multimodal understanding of both visual and language contexts. However, existing VLMs often face severe challenges of hallucinations, meaning that VLMs tend to generate responses…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Jinjin Cao , Zhiyang Chen , Zijun Wang , Liyuan Ma , Weijian Luo , Guojun Qi

We propose NEURONA, a neuro-symbolic framework for fMRI decoding and concept grounding in neural activity. Leveraging image- and video-based fMRI question-answering datasets, NEURONA learns to decode interacting concepts from visual stimuli…

Neurons and Cognition · Quantitative Biology 2026-03-05 Yanchen Wang , Joy Hsu , Ehsan Adeli , Jiajun Wu

fMRI semantic category understanding using linguistic encoding models attempt to learn a forward mapping that relates stimuli to the corresponding brain activation. Classical encoding models use linear multi-variate methods to predict the…

Machine Learning · Computer Science 2018-12-04 Subba Reddy Oota , Adithya Avvaru , Naresh Manwani , Raju S. Bapi