中文
相关论文

相关论文: UniBrain: Unify Image Reconstruction and Captionin…

200 篇论文

Unified multimodal models often struggle with complex synthesis tasks that demand deep reasoning, and typically treat text-to-image generation and image editing as isolated capabilities rather than interconnected reasoning steps. To address…

计算机视觉与模式识别 · 计算机科学 2026-02-23 Dianyi Wang , Chaofan Ma , Feng Han , Size Wu , Wei Song , Yibin Wang , Zhixiong Zhang , Tianhang Wang , Siyuan Wang , Zhongyu Wei , Jiaqi Wang

Seeing is believing, however, the underlying mechanism of how human visual perceptions are intertwined with our cognitions is still a mystery. Thanks to the recent advances in both neuroscience and artificial intelligence, we have been able…

图像与视频处理 · 电气工程与系统科学 2023-08-17 Yu-Ting Lan , Kan Ren , Yansen Wang , Wei-Long Zheng , Dongsheng Li , Bao-Liang Lu , Lili Qiu

Text-Aware Image Restoration (TAIR) aims to recover high-quality images from low-quality inputs containing degraded textual content. While diffusion models provide strong generative priors for general image restoration, they often produce…

计算机视觉与模式识别 · 计算机科学 2025-12-10 Jin Hyeon Kim , Paul Hyunbin Cho , Claire Kim , Jaewon Min , Jaeeun Lee , Jihye Park , Yeji Choi , Seungryong Kim

Image recaptioning is widely used to generate training datasets with enhanced quality for various multimodal tasks. Existing recaptioning methods typically rely on powerful multimodal large language models (MLLMs) to enhance textual…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Yuchi Wang , Yishuo Cai , Shuhuai Ren , Sihan Yang , Linli Yao , Yuanxin Liu , Yuanxing Zhang , Pengfei Wan , Xu Sun

Unified Multimodal Models (UMMs) have demonstrated remarkable performance in text-to-image generation (T2I) and editing (TI2I), whether instantiated as assembled unified frameworks which couple powerful vision-language model (VLM) with…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Yuxin Song , Wenkai Dong , Shizun Wang , Qi Zhang , Song Xue , Tao Yuan , Hu Yang , Haocheng Feng , Hang Zhou , Xinyan Xiao , Jingdong Wang

Image fusion aims to integrate complementary information from multiple source images to produce a more informative and visually consistent representation, benefiting both human perception and downstream vision tasks. Despite recent…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Xingyuan Li , Songcheng Du , Yang Zou , HaoYuan Xu , Zhiying Jiang , Jinyuan Liu

Recent work has demonstrated that complex visual stimuli can be decoded from human brain activity using deep generative models, offering new ways to probe how the brain represents real-world scenes. However, many existing approaches first…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Pinyuan Feng , Hossein Adeli , Wenxuan Guo , Fan Cheng , Ethan Hwang , Nikolaus Kriegeskorte

Reconstructing human dynamic vision from brain activity is a challenging task with great scientific significance. Although prior video reconstruction methods have made substantial progress, they still suffer from several limitations,…

计算机视觉与模式识别 · 计算机科学 2025-02-20 Yizhuo Lu , Changde Du , Chong Wang , Xuanliu Zhu , Liuyun Jiang , Xujin Li , Huiguang He

Reconstructing human vision from brain activities has been an appealing task that helps to understand our cognitive process. Even though recent research has seen great success in reconstructing static images from non-invasive brain…

计算机视觉与模式识别 · 计算机科学 2023-05-22 Zijiao Chen , Jiaxin Qing , Juan Helen Zhou

Reconstructing visual stimulus (image) only from human brain activity measured with functional Magnetic Resonance Imaging (fMRI) is a significant and meaningful task in Human-AI collaboration. However, the inconsistent distribution and…

计算机视觉与模式识别 · 计算机科学 2019-10-23 Ziqi Ren , Jie Li , Xuetong Xue , Xin Li , Fan Yang , Zhicheng Jiao , Xinbo Gao

Understanding how the brain encodes visual information is a central challenge in neuroscience and machine learning. A promising approach is to reconstruct visual stimuli, essentially images, from functional Magnetic Resonance Imaging (fMRI)…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Zheng Huang , Enpei Zhang , Weikang Qiu , Yinghao Cai , Carl Yang , Elynn Chen , Xiang Zhang , Rex Ying , Dawei Zhou , Yujun Yan

Understanding how humans process visual information is one of the crucial steps for unraveling the underlying mechanism of brain activity. Recently, this curiosity has motivated the fMRI-to-image reconstruction task; given the fMRI data…

计算机视觉与模式识别 · 计算机科学 2024-09-19 Jaehoon Joo , Taejin Jeong , Seongjae Hwang

Reconstructing the viewed images from human brain activity bridges human and computer vision through the Brain-Computer Interface. The inherent variability in brain function between individuals leads existing literature to focus on…

计算机视觉与模式识别 · 计算机科学 2024-04-01 Ruijie Quan , Wenguan Wang , Zhibo Tian , Fan Ma , Yi Yang

Brain imaging analysis is fundamental in neuroscience, providing valuable insights into brain structure and function. Traditional workflows follow a sequential pipeline-brain extraction, registration, segmentation, parcellation, network…

图像与视频处理 · 电气工程与系统科学 2025-02-27 Yao Su , Keqi Han , Mingjie Zeng , Lichao Sun , Liang Zhan , Carl Yang , Lifang He , Xiangnan Kong

In the pursuit to understand the intricacies of human brain's visual processing, reconstructing dynamic visual experiences from brain activities emerges as a challenging yet fascinating endeavor. While recent advancements have achieved…

计算机视觉与模式识别 · 计算机科学 2024-05-14 Jingyuan Sun , Mingxiao Li , Zijiao Chen , Marie-Francine Moens

Visual captioning aims to generate textual descriptions given images or videos. Traditionally, image captioning models are trained on human annotated datasets such as Flickr30k and MS-COCO, which are limited in size and diversity. This…

计算机视觉与模式识别 · 计算机科学 2021-03-01 Marimuthu Kalimuthu , Aditya Mogadala , Marius Mosbach , Dietrich Klakow

Image captioning is a research hotspot where encoder-decoder models combining convolutional neural network (CNN) and long short-term memory (LSTM) achieve promising results. Despite significant progress, these models generate sentences…

计算机视觉与模式识别 · 计算机科学 2019-10-16 Hongwei Ge , Zehang Yan , Kai Zhang , Mingde Zhao , Liang Sun

Reconstructing visual stimuli from functional Magnetic Resonance Imaging fMRI enables fine-grained retrieval of brain activity. However, the accurate reconstruction of diverse details, including structure, background, texture, color, and…

神经与进化计算 · 计算机科学 2025-01-09 Haoyu Li , Hao Wu , Badong Chen

Image restoration aims to recover content from inputs degraded by various factors, such as adverse weather, blur, and noise. Perceptual Image Restoration (PIR) methods improve visual quality but often do not support downstream tasks…

图像与视频处理 · 电气工程与系统科学 2025-06-03 I-Hsiang Chen , Wei-Ting Chen , Yu-Wei Liu , Yuan-Chun Chiang , Sy-Yen Kuo , Ming-Hsuan Yang

Decoding visual stimuli from brain recordings aims to deepen our understanding of the human visual system and build a solid foundation for bridging human and computer vision through the Brain-Computer Interface. However, reconstructing…

计算机视觉与模式识别 · 计算机科学 2023-03-30 Zijiao Chen , Jiaxin Qing , Tiange Xiang , Wan Lin Yue , Juan Helen Zhou