中文
相关论文

相关论文: SynMind: Reducing Semantic Hallucination in fMRI-B…

200 篇论文

Reconstructing visual stimuli from human brain activity (e.g., fMRI) bridges neuroscience and computer vision by decoding neural representations. However, existing methods often overlook critical brain structure-function relationships,…

计算机视觉与模式识别 · 计算机科学 2025-07-23 Sijin Yu , Zijiao Chen , Wenxuan Wu , Shengxian Chen , Zhongliang Liu , Jingxin Nie , Xiaofen Xing , Xiangmin Xu , Xin Zhang

Synthesizing medical images while preserving their structural information is crucial in medical research. In such scenarios, the preservation of anatomical content becomes especially important. Although recent advances have been made by…

图像与视频处理 · 电气工程与系统科学 2024-11-14 Ziqi Yu , Botao Zhao , Shengjie Zhang , Xiang Chen , Jianfeng Feng , Tingying Peng , Xiao-Yong Zhang

Understanding the hidden mechanisms behind human's visual perception is a fundamental question in neuroscience. To that end, investigating into the neural responses of human mind activities, such as functional Magnetic Resonance Imaging…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Yuankun Yang , Li Zhang , Ziyang Xie , Zhiyuan Yuan , Jianfeng Feng , Xiatian Zhu , Yu-Gang Jiang

In recent years, accelerated MRI reconstruction based on deep learning has led to significant improvements in image quality with impressive results for high acceleration factors. However, from a clinical perspective image quality is only…

图像与视频处理 · 电气工程与系统科学 2025-07-02 Jan Nikolas Morshuis , Christian Schlarmann , Thomas Küstner , Christian F. Baumgartner , Matthias Hein

The development of algorithms to accurately decode neural information has long been a research focus in the field of neuroscience. Brain decoding typically involves training machine learning models to map neural data onto a preestablished…

Decoding visual representations from brain signals has attracted significant attention in both neuroscience and artificial intelligence. However, the degree to which brain signals truly encode visual information remains unclear. Current…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Jiawen Zheng , Haonan Jia , Ming Li , Yuhui Zheng , Yufeng Zeng , Yang Gao , Chen Liang

Unveiling visual semantics from neural signals such as EEG, MEG, and fMRI remains a fundamental challenge due to subject variability and the entangled nature of visual features. Existing approaches primarily align neural activity directly…

计算机视觉与模式识别 · 计算机科学 2025-11-07 Zehui Feng , Chenqi Zhang , Mingru Wang , Minuo Wei , Shiwei Cheng , Cuntai Guan , Ting Han

Large Vision-Language Models (LVLMs) bridge the gap between visual and linguistic modalities, demonstrating strong potential across a variety of domains. However, despite significant progress, LVLMs still suffer from severe hallucination…

计算机视觉与模式识别 · 计算机科学 2025-12-23 Ruiqi Ma , Yu Yan , Chunhong Zhang , Minghao Yin , XinChao Liu , Zhihong Jin , Zheng Hu

Multimodal medical image fusion plays a crucial role in medical diagnosis by integrating complementary information from different modalities to enhance image readability and clinical applicability. However, existing methods mainly follow…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Haozhe Xiang , Han Zhang , Yu Cheng , Xiongwen Quan , Wanwan Huang

Hallucinations in large language model (LLM) outputs severely limit their reliability in knowledge-intensive tasks such as question answering. To address this challenge, we introduce REFIND (Retrieval-augmented Factuality hallucINation…

计算与语言 · 计算机科学 2025-04-09 DongGeon Lee , Hwanjo Yu

Decoding visual stimuli from brain recordings aims to deepen our understanding of the human visual system and build a solid foundation for bridging human and computer vision through the Brain-Computer Interface. However, reconstructing…

计算机视觉与模式识别 · 计算机科学 2023-03-30 Zijiao Chen , Jiaxin Qing , Tiange Xiang , Wan Lin Yue , Juan Helen Zhou

Multimodal Large Language Models (MLLMs) often struggle to accurately perceive fine-grained visual details, especially when targets are tiny or visually subtle. This challenge can be addressed through semantic-visual information fusion,…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Yuxiang Shen , Hailong Huang , Zhenkun Gao , Xueheng Li , Man Zhou , Chengjun Xie , Haoxuan Che , Xuanhua He , Jie Zhang

Visual image reconstruction from functional Magnetic Resonance Imaging (fMRI) is a fundamental task in brain decoding, providing a crucial pathway for understanding human perceptual mechanisms and developing advanced brain-computer…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Yudan Ren , Pengcheng Shi , Zihan Ma , Xiaowei He , Xiao Li

We propose SNI-SLAM, a semantic SLAM system utilizing neural implicit representation, that simultaneously performs accurate semantic mapping, high-quality surface reconstruction, and robust camera tracking. In this system, we introduce…

机器人学 · 计算机科学 2024-03-29 Siting Zhu , Guangming Wang , Hermann Blum , Jiuming Liu , Liang Song , Marc Pollefeys , Hesheng Wang

Speech language models align with human brain responses to natural language to an impressive degree. However, current models rely heavily on low-level speech features, indicating they lack brain-relevant semantics which limits their utility…

计算与语言 · 计算机科学 2025-03-05 Omer Moussa , Dietrich Klakow , Mariya Toneva

Reinforcement learning has recently improved the reasoning ability of Large Language Models and Multimodal LLMs, yet prevailing reward designs emphasise final-answer correctness and consequently tolerate process hallucinations--cases where…

计算机视觉与模式识别 · 计算机科学 2026-03-10 Yantao Li , Qiang Hui , Chenyang Yan , Kanzhi Cheng , Fang Zhao , Chao Tan , Huanling Gao , Jianbing Zhang , Kai Wang , Xinyu Dai , Shiguo Lian

Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in visual understanding and multimodal reasoning. However, LVLMs frequently exhibit hallucination phenomena, manifesting as the generated textual responses that…

计算机视觉与模式识别 · 计算机科学 2025-07-31 Ziyun Dai , Xiaoqiang Li , Shaohua Zhang , Yuanchen Wu , Jide Li

A system that enables blind or visually impaired users to access comics/manga would introduce a new medium of storytelling to this community. However, no such system currently exists. Generative vision-language models (VLMs) have shown…

Brain-to-image decoding has been recently propelled by the progress in generative AI models and the availability of large ultra-high field functional Magnetic Resonance Imaging (fMRI). However, current approaches depend on complicated…

计算机视觉与模式识别 · 计算机科学 2025-05-21 Marlène Careil , Yohann Benchetrit , Jean-Rémi King

To reduce scanning time and/or improve spatial/temporal resolution in some MRI applications, parallel MRI (pMRI) acquisition techniques with multiple coils acquisition have emerged since the early 1990s as powerful 3D imaging methods that…

最优化与控制 · 数学 2009-09-03 Lotfi Chaari , Jean-Christophe Pesquet , Philippe Ciuciu , Amel Benazza-Benyahia