中文
相关论文

相关论文: MIRAGE: Robust multi-modal architectures translate…

200 篇论文

Recent progress in task-optimized neural networks has established encoding models as a powerful tool for predicting brain responses to naturalistic stimuli, yet most existing approaches rely on unimodal representations. The emergence of…

机器学习 · 计算机科学 2026-05-29 Abdulkadir Gokce , Badr AlKhamissi , Martin Schrimpf

We release NSD-Imagery, a benchmark dataset of human fMRI activity paired with mental images, to complement the existing Natural Scenes Dataset (NSD), a large-scale dataset of fMRI activity paired with seen images that enabled unprecedented…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Reese Kneeland , Paul S. Scotti , Ghislain St-Yves , Jesse Breedlove , Kendrick Kay , Thomas Naselaris

Artificial intelligence (AI) has become a fundamental tool for assisting clinicians in analyzing ophthalmic images, such as optical coherence tomography (OCT). However, developing AI models often requires extensive annotation, and existing…

In the past five years, the use of generative and foundational AI systems has greatly improved the decoding of brain activity. Visual perception, in particular, can now be decoded from functional Magnetic Resonance Imaging (fMRI) with…

图像与视频处理 · 电气工程与系统科学 2024-03-15 Yohann Benchetrit , Hubert Banville , Jean-Rémi King

Spatial perception and reasoning are core components of human cognition, encompassing object recognition, spatial relational understanding, and dynamic reasoning. Despite progress in computer vision, existing benchmarks reveal significant…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Chonghan Liu , Haoran Wang , Felix Henry , Pu Miao , Yajie Zhang , Yu Zhao , Peiran Wu

The reconstruction of images observed by subjects from fMRI data collected during visual stimuli has made strong progress in the past decade, thanks to the availability of extensive fMRI datasets and advancements in generative models for…

Access to diverse, well-annotated medical images with interactive learning tools is fundamental for training practitioners in medicine and related fields to improve their diagnostic skills and understanding of anatomical structures. While…

计算机视觉与模式识别 · 计算机科学 2026-05-07 Miguel Diaz Benito , Cecilia Diana Albelda , Alvaro Garcia Martin , Jesus Bescos Cano , Marcos Escudero-Vinolo , Juan C. SanMiguel

Reconstructing visual stimuli from human brain activity (e.g., fMRI) bridges neuroscience and computer vision by decoding neural representations. However, existing methods often overlook critical brain structure-function relationships,…

计算机视觉与模式识别 · 计算机科学 2025-07-23 Sijin Yu , Zijiao Chen , Wenxuan Wu , Shengxian Chen , Zhongliang Liu , Jingxin Nie , Xiaofen Xing , Xiangmin Xu , Xin Zhang

Vision-language models (VLMs) excel at multimodal understanding, yet their text-only decoding forces them to verbalize visual reasoning, limiting performance on tasks that demand visual imagination. Recent attempts train VLMs to render…

计算机视觉与模式识别 · 计算机科学 2025-06-23 Zeyuan Yang , Xueyang Yu , Delin Chen , Maohao Shen , Chuang Gan

Degradation-agnostic image restoration aims to handle diverse corruptions with one unified model, but faces fundamental challenges in balancing efficiency and performance across different degradation types. Existing approaches either…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Bin Ren , Yawei Li , Xu Zheng , Yuqian Fu , Danda Pani Paudel , Hong Liu , Ming-Hsuan Yang , Luc Van Gool , Nicu Sebe

Deep learning models struggle with systematic compositional generalization, a hallmark of human cognition. We propose \textsc{Mirage}, a neuro-inspired dual-process model that offers a processing account for this ability. It combines a…

人工智能 · 计算机科学 2025-10-29 Alex Noviello , Claas Beger , Jacob Groner , Kevin Ellis , Weinan Sun

In daily life, we encounter diverse external stimuli, such as images, sounds, and videos. As research in multimodal stimuli and neuroscience advances, fMRI-based brain decoding has become a key tool for understanding brain perception and…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Pengyu Liu , Guohua Dong , Dan Guo , Kun Li , Fengling Li , Xun Yang , Meng Wang , Xiaomin Ying

Decoding visual experiences from fMRI offers a powerful avenue to understand human perception and develop advanced brain-computer interfaces. However, current progress often prioritizes maximizing reconstruction fidelity while overlooking…

机器学习 · 计算机科学 2025-10-09 Yuxiang Wei , Yanteng Zhang , Xi Xiao , Tianyang Wang , Xiao Wang , Vince D. Calhoun

To effectively leverage user-specific data, retrieval augmented generation (RAG) is employed in multimodal large language model (MLLM) applications. However, conventional retrieval approaches often suffer from limited retrieval accuracy.…

计算机视觉与模式识别 · 计算机科学 2026-03-04 Maoliang Li , Ke Li , Yaoyang Liu , Jiayu Chen , Zihao Zheng , Yinjun Wu , Chenchen Liu , Xiang Chen

Instruction-guided image editing has seen remarkable progress with models like FLUX.2 and Qwen-Image-Edit, yet they still struggle with complex scenarios with multiple similar instances each requiring individual edits. We observe that…

计算机视觉与模式识别 · 计算机科学 2026-04-08 Ziqian Liu , Stephan Alaniz

The integration of deep learning and neuroscience has been advancing rapidly, which has led to improvements in the analysis of brain activity and the understanding of deep learning models from a neuroscientific perspective. The…

神经元与认知 · 定量生物学 2023-06-21 Yu Takagi , Shinji Nishimoto

Every day, the human brain processes an immense volume of visual information, relying on intricate neural mechanisms to perceive and interpret these stimuli. Recent breakthroughs in functional magnetic resonance imaging (fMRI) have enabled…

计算机视觉与模式识别 · 计算机科学 2023-05-22 Matteo Ferrante , Furkan Ozcelik , Tommaso Boccato , Rufin VanRullen , Nicola Toschi

Reconstructing visual stimuli from measured functional magnetic resonance imaging (fMRI) has been a meaningful and challenging task. Previous studies have successfully achieved reconstructions with structures similar to the original images,…

计算机视觉与模式识别 · 计算机科学 2023-03-27 Yizhuo Lu , Changde Du , Dianpeng Wang , Huiguang He

Reconstructing human dynamic vision from brain activity is a challenging task with great scientific significance. Although prior video reconstruction methods have made substantial progress, they still suffer from several limitations,…

计算机视觉与模式识别 · 计算机科学 2025-02-20 Yizhuo Lu , Changde Du , Chong Wang , Xuanliu Zhu , Liuyun Jiang , Xujin Li , Huiguang He

In the pursuit to understand the intricacies of human brain's visual processing, reconstructing dynamic visual experiences from brain activities emerges as a challenging yet fascinating endeavor. While recent advancements have achieved…

计算机视觉与模式识别 · 计算机科学 2024-05-14 Jingyuan Sun , Mingxiao Li , Zijiao Chen , Marie-Francine Moens
‹ 上一页 1 2 3 10 下一页 ›