English
Related papers

Related papers: TRIBE: TRImodal Brain Encoder for whole-brain fMRI…

200 papers

We present an exploration of machine learning architectures for predicting brain responses to realistic images on occasion of the Algonauts Challenge 2023. Our research involved extensive experimentation with various pretrained models.…

Neurons and Cognition · Quantitative Biology 2023-09-20 Riccardo Chimisso , Sathya Buršić , Paolo Marocco , Giuseppe Vizzari , Dimitri Ognibene

In this work, we present a network-specific approach for predicting brain responses to complex multimodal movies, leveraging the Yeo 7-network parcellation of the Schaefer atlas. Rather than treating the brain as a homogeneous system, we…

Neurons and Cognition · Quantitative Biology 2025-10-28 Andrea Corsico , Giorgia Rigamonti , Simone Zini , Luigi Celona , Paolo Napoletano

Understanding neural activity and information representation is crucial for advancing knowledge of brain function and cognition. Neural activity, measured through techniques like electrophysiology and neuroimaging, reflects various aspects…

Neurons and Cognition · Quantitative Biology 2024-07-22 Fengyu Yang , Chao Feng , Daniel Wang , Tianye Wang , Ziyao Zeng , Zhiyang Xu , Hyoungseob Park , Pengliang Ji , Hanbin Zhao , Yuanning Li , Alex Wong

Integrating information from multiple modalities is arguably one of the essential prerequisites for grounding artificial intelligence systems with an understanding of the real world. Recent advances in video transformers that jointly learn…

Computer Vision and Pattern Recognition · Computer Science 2023-11-15 Dota Tianai Dong , Mariya Toneva

Universal Multimodal Retrieval requires unified embedding models capable of interpreting diverse user intents, ranging from simple keywords to complex compositional instructions. While Multimodal Large Language Models (MLLMs) possess strong…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Xiangzhao Hao , Shijie Wang , Tianyu Yang , Tianyue Wang , Haiyun Guo , Jinqiao Wang

We present our submission to the Algonauts 2025 Challenge, where the goal is to predict fMRI brain responses to movie stimuli. Our approach integrates multimodal representations from large language models, video encoders, audio models, and…

Image and Video Processing · Electrical Eng. & Systems 2025-10-09 Robert Scholz , Kunal Bagga , Christine Ahrends , Carlo Alberto Barbano

Recent achievements in implantable brain-computer interfaces (iBCIs) have demonstrated the potential to decode cognitive and motor behaviors with intracranial brain recordings; however, individual physiological and electrode implantation…

Neurons and Cognition · Quantitative Biology 2025-06-17 Di Wu , Linghao Bu , Yifei Jia , Lu Cao , Siyuan Li , Siyu Chen , Yueqian Zhou , Sheng Fan , Wenjie Ren , Dengchang Wu , Kang Wang , Yue Zhang , Yuehui Ma , Jie Yang , Mohamad Sawan

Multimodal magnetic resonance imaging (MRI) constitutes the first line of investigation for clinicians in the care of brain tumors, providing crucial insights for surgery planning, treatment monitoring, and biomarker identification.…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Lucas Robinet , Ahmad Berjaoui , Elizabeth Cohen-Jonathan Moyal

fMRI semantic category understanding using linguistic encoding models attempts to learn a forward mapping that relates stimuli to the corresponding brain activation. State-of-the-art encoding models use a single global model (linear or…

Machine Learning · Computer Science 2020-06-02 Subba Reddy Oota , Naresh Manwani , Raju S. Bapi

This work presents our solutions to the Algonauts Project 2023 Challenge. The primary objective of the challenge revolves around employing computational models to anticipate brain responses captured during participants' observation of…

Computer Vision and Pattern Recognition · Computer Science 2023-08-02 Xuan-Bac Nguyen , Xudong Liu , Xin Li , Khoa Luu

Multimodal pretraining is an effective strategy for the trinity of goals of representation learning in autonomous robots: 1) extracting both local and global task progressions; 2) enforcing temporal consistency of visual representation; 3)…

Reasoning lies at the heart of intelligence, shaping the ability to make decisions, draw conclusions, and generalize across domains. In artificial intelligence, as systems increasingly operate in open, uncertain, and multimodal…

Extensive literature has drawn comparisons between recordings of biological neurons in the brain and deep neural networks. This comparative analysis aims to advance and interpret deep neural networks and enhance our understanding of…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Mai Gamal , Mohamed Rashad , Eman Ehab , Seif Eldawlatly , Mennatullah Siam

The Contrastive Language-Image Pre-training (CLIP) framework has become a widely used approach for multimodal representation learning, particularly in image-text retrieval and clustering. However, its efficacy is constrained by three key…

Computer Vision and Pattern Recognition · Computer Science 2025-12-09 Tiancheng Gu , Kaicheng Yang , Ziyong Feng , Xingjun Wang , Yanzhao Zhang , Dingkun Long , Yingda Chen , Weidong Cai , Jiankang Deng

Visual image reconstruction from functional Magnetic Resonance Imaging (fMRI) is a fundamental task in brain decoding, providing a crucial pathway for understanding human perceptual mechanisms and developing advanced brain-computer…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Yudan Ren , Pengcheng Shi , Zihan Ma , Xiaowei He , Xiao Li

We present MedARC's team solution to the Algonauts 2025 challenge. Our pipeline leveraged rich multimodal representations from various state-of-the-art pretrained models across video (V-JEPA2), speech (Whisper), text (Llama 3.2),…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Cesar Kadir Torrico Villanueva , Jiaxin Cindy Tu , Mihir Tripathy , Connor Lane , Rishab Iyer , Paul S. Scotti

Missing input sequences are common in medical imaging data, posing a challenge for deep learning models reliant on complete input data. In this work, inspired by MultiMAE [2], we develop a masked autoencoder (MAE) paradigm for multi-modal,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-04 Ayhan Can Erdur , Christian Beischl , Daniel Scholz , Jiazhen Pan , Benedikt Wiestler , Daniel Rueckert , Jan C Peeken

We use (multi)modal deep neural networks (DNNs) to probe for sites of multimodal integration in the human brain by predicting stereoencephalography (SEEG) recordings taken while human subjects watched movies. We operationalize sites of…

Machine Learning · Computer Science 2024-06-21 Vighnesh Subramaniam , Colin Conwell , Christopher Wang , Gabriel Kreiman , Boris Katz , Ignacio Cases , Andrei Barbu

Human intelligence is multimodal; we integrate visual, linguistic, and acoustic signals to maintain a holistic worldview. Most current pretraining methods, however, are limited to one or two modalities. We present i-Code, a self-supervised…

The exploration of brain activity and its decoding from fMRI data has been a longstanding pursuit, driven by its potential applications in brain-computer interfaces, medical diagnostics, and virtual reality. Previous approaches have…

Computer Vision and Pattern Recognition · Computer Science 2023-11-02 Xuelin Qian , Yun Wang , Jingyang Huo , Jianfeng Feng , Yanwei Fu