中文
相关论文

相关论文: MindAlign: Bridging EEG, Vision, and Language for …

200 篇论文

Recent studies demonstrate the use of a two-stage supervised framework to generate images that depict human perception to visual stimuli from EEG, referring to EEG-visual reconstruction. They are, however, unable to reproduce the exact…

多媒体 · 计算机科学 2022-08-19 Zesheng Ye , Lina Yao , Yu Zhang , Sylvia Gustin

Despite advancements in artificial intelligence, object recognition models still lag behind in emulating visual information processing in human brains. Recent studies have highlighted the potential of using neural data to mimic brain…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Zitong Lu , Yile Wang , Julie D. Golomb

EEG-based visual decoding aims to establish a mapping between neural signals and visual semantics. However, it remains constrained by the dual challenges of severe information granularity mismatch and the low signal-to-noise ratio (SNR) of…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Fan Yin , Chuhang Zheng , Peiliang Gong , Donghai Guan , Qi Zhu

Decoding imagined speech from non-invasive brain recordings is challenging because imagined datasets are scarce and difficult to align temporally across subjects and sessions In this work, we propose a new approach to the decoding of…

机器学习 · 计算机科学 2026-05-11 Maryam Maghsoudi , Shihab Shamma

Large Vision Language Models (LVLMs) exhibit strong visual understanding and reasoning abilities. However, whether their internal representations reflect human visual cognition is still under-explored. In this paper, we address this by…

人机交互 · 计算机科学 2026-03-10 Xin Xiao , Yang Lei , Haoyang Zeng , Xiao Sun , Xinyi Jiang , Yu Tian , Hao Wu , Kaiwen Wei , Jiang Zhong

Decoding visual representations from brain signals has attracted significant attention in both neuroscience and artificial intelligence. However, the degree to which brain signals truly encode visual information remains unclear. Current…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Jiawen Zheng , Haonan Jia , Ming Li , Yuhui Zheng , Yufeng Zeng , Yang Gao , Chen Liang

Electroencephalography (EEG)-based visual perception reconstruction has become an important area of research. Neuroscientific studies indicate that humans can decode imagined 3D objects by perceiving or imagining various visual information,…

计算机视觉与模式识别 · 计算机科学 2024-11-19 Xin Xiang , Wenhui Zhou , Guojun Dai

Advancements in non-invasive electroencephalogram (EEG)-based Brain-Computer Interface (BCI) technology have enabled communication through brain activity, offering significant potential for individuals with motor impairments. Existing…

信号处理 · 电气工程与系统科学 2024-09-26 Jingyuan Li , Yansen Wang , Nie Lin , Dongsheng Li

While capable of segregating visual data, humans take time to examine a single piece, let alone thousands or millions of samples. The deep learning models efficiently process sizeable information with the help of modern-day computing.…

计算机视觉与模式识别 · 计算机科学 2022-09-28 Alankrit Mishra , Nikhil Raj , Garima Bajwa

Deciphering the intricacies of the human brain has captivated curiosity for centuries. Recent strides in Brain-Computer Interface (BCI) technology, particularly using motor imagery, have restored motor functions such as reaching, grasping,…

计算与语言 · 计算机科学 2024-05-06 Hanwen Liu , Daniel Hajialigol , Benny Antony , Aiguo Han , Xuan Wang

The electroencephalogram (EEG) offers a non-invasive means by which a listener's auditory system may be monitored during continuous speech perception. Reliable auditory-EEG decoders could facilitate the objective diagnosis of hearing…

音频与语音处理 · 电气工程与系统科学 2023-12-18 Mike Thornton , Danilo Mandic , Tobias Reichenbach

We introduce a two-stage multitask learning framework for analyzing Electroencephalography (EEG) signals that integrates denoising, dynamical modeling, and representation learning. In the first stage, a denoising autoencoder is trained to…

机器学习 · 计算机科学 2026-02-24 Sucheta Ghosh , Felix Dietrich , Zahra Monfared

The intrication of brain signals drives research that leverages multimodal AI to align brain modalities with visual and textual data for explainable descriptions. However, most existing studies are limited to coarse interpretations, lacking…

计算机视觉与模式识别 · 计算机科学 2025-05-22 Weihao Xia , Cengiz Oztireli

Decoding visual images from brain activity has significant potential for advancing brain-computer interaction and enhancing the understanding of human perception. Recent approaches align the representation spaces of images and brain…

计算机视觉与模式识别 · 计算机科学 2025-06-11 Nona Rajabi , Antônio H. Ribeiro , Miguel Vasco , Farzaneh Taleb , Mårten Björkman , Danica Kragic

Unveiling visual semantics from neural signals such as EEG, MEG, and fMRI remains a fundamental challenge due to subject variability and the entangled nature of visual features. Existing approaches primarily align neural activity directly…

计算机视觉与模式识别 · 计算机科学 2025-11-07 Zehui Feng , Chenqi Zhang , Mingru Wang , Minuo Wei , Shiwei Cheng , Cuntai Guan , Ting Han

The translation of brain dynamics into natural language is pivotal for brain-computer interfaces (BCIs). With the swift advancement of large language models, such as ChatGPT, the need to bridge the gap between the brain and languages…

人机交互 · 计算机科学 2024-01-04 Yiqun Duan , Jinzhao Zhou , Zhen Wang , Yu-Kai Wang , Chin-Teng Lin

Leveraging the universal representations of pre-trained LLMs and MLLMs offers a promising path toward brain foundation models. However, visually-evoked EEG datasets remain scarce, leading existing methods to align neural signals mainly with…

人工智能 · 计算机科学 2026-05-26 Jun-Yu Pan , Yansen Wang , Enze Zhang , Bao-Liang Lu , Wei-Long Zheng , Dongsheng Li

In this work, we study the problem of cross-subject motor imagery (MI) decoding from electroencephalography (EEG) data. Multi-subject EEG datasets present several kinds of domain shifts due to various inter-individual differences (e.g.…

信号处理 · 电气工程与系统科学 2024-02-22 Georgios Zoumpourlis , Ioannis Patras

The ability to accurately interpret complex visual information is a crucial topic of multimodal large language models (MLLMs). Recent work indicates that enhanced visual perception significantly reduces hallucinations and improves…

Existing vision tokenization isolates the optimization of vision tokenizers from downstream training, implicitly assuming the visual tokens can generalize well across various tasks, e.g., image generation and visual question answering. The…

计算机视觉与模式识别 · 计算机科学 2025-05-16 Wenxuan Wang , Fan Zhang , Yufeng Cui , Haiwen Diao , Zhuoyan Luo , Huchuan Lu , Jing Liu , Xinlong Wang