中文
相关论文

相关论文: From Text to Image: Exploring GPT-4Vision's Potent…

200 篇论文

In this research, we introduce RefineNet, a novel architecture designed to address resolution limitations in text-to-image conversion systems. We explore the challenges of generating high-resolution images from textual descriptions,…

计算机视觉与模式识别 · 计算机科学 2024-01-01 Fan Shi

Existing large video-language models (LVLMs) struggle to comprehend long videos correctly due to limited context. To address this problem, fine-tuning long-context LVLMs and employing GPT-based agents have emerged as promising solutions.…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Yongdong Luo , Xiawu Zheng , Guilin Li , Shukang Yin , Haojia Lin , Chaoyou Fu , Jinfa Huang , Jiayi Ji , Fei Chao , Jiebo Luo , Rongrong Ji

Medical image interpretation is central to most clinical applications such as disease diagnosis, treatment planning, and prognostication. In clinical practice, radiologists examine medical images and manually compile their findings into…

计算机视觉与模式识别 · 计算机科学 2023-11-21 Nurbanu Aksoy , Nishant Ravikumar , Alejandro F Frangi

Artificial intelligence (AI) has significant potential to positively impact and advance medical imaging, including positron emission tomography (PET) imaging applications. AI has the ability to enhance and optimize all aspects of the PET…

计算机视觉与模式识别 · 计算机科学 2021-07-15 Arkadiusz Sitek , Sangtae Ahn , Evren Asma , Adam Chandler , Alvin Ihsani , Sven Prevrhal , Arman Rahmim , Babak Saboury , Kris Thielemans

Image to image translation aims to learn a mapping that transforms an image from one visual domain to another. Recent works assume that images descriptors can be disentangled into a domain-invariant content representation and a…

计算机视觉与模式识别 · 计算机科学 2020-08-13 Raul Gomez , Yahui Liu , Marco De Nadai , Dimosthenis Karatzas , Bruno Lepri , Nicu Sebe

The potential benefit of hybrid X-ray and MR imaging in the interventional environment is large due to the combination of fast imaging with high contrast variety. However, a vast amount of existing image enhancement methods requires the…

计算机视觉与模式识别 · 计算机科学 2019-05-09 Bernhard Stimpel , Christopher Syben , Tobias Würfl , Katharina Breininger , Katrin Mentl , Jonathan M. Lommen , Arnd Dörfler , Andreas Maier

The theranostic paradigm enables personalization of treatment by selecting patients with a diagnostic radiopharmaceutical and monitoring therapy using a matched therapeutic isotope. This strategy relies on accurate image reconstruction of…

We propose a Multifaceted Resilient Network(MRNet), a novel architecture developed for medical image-to-image translation that outperforms state-of-the-art methods in MRI-to-CT and MRI-to-MRI conversion. MRNet leverages the Segment Anything…

图像与视频处理 · 电气工程与系统科学 2024-12-05 Hyojeong Lee , Youngwan Jo , Inpyo Hong , Sanghyun Park

Computed tomography (CT) is a beneficial imaging tool for diagnostic purposes. CT scans provide detailed information concerning the internal anatomic structures of a patient, but present higher radiation dose and costs compared to X-ray…

图像与视频处理 · 电气工程与系统科学 2024-03-05 Benjamin Paulson , Joshua Goldshteyn , Sydney Balboni , John Cisler , Andrew Crisler , Natalia Bukowski , Julia Kalish , Theodore Colwell

This study extends previous research on spatial representations in multimodal AI systems. Although current models demonstrate a rich understanding of spatial information from images, this information is rooted in propositional…

人工智能 · 计算机科学 2024-09-24 Bridget Leonard , Kristin Woodard , Scott O. Murray

Vision Transformers (ViTs) are widely adopted in medical imaging tasks, and some existing efforts have been directed towards vision-language training for Chest X-rays (CXRs). However, we envision that there still exists a potential for…

计算机视觉与模式识别 · 计算机科学 2023-11-14 Umar Marikkar , Sara Atito , Muhammad Awais , Adam Mahdi

The rapid development of multimodal large language models (MLLMs), such as GPT-4V, has led to significant advancements. However, these models still face challenges in medical multimodal capabilities due to limitations in the quantity and…

计算机视觉与模式识别 · 计算机科学 2024-10-01 Junying Chen , Chi Gui , Ruyi Ouyang , Anningzhe Gao , Shunian Chen , Guiming Hardy Chen , Xidong Wang , Ruifei Zhang , Zhenyang Cai , Ke Ji , Guangjun Yu , Xiang Wan , Benyou Wang

Generative AI and large language models have the potential to drastically improve the landscape of computing education by automatically generating personalized feedback and content. Recent works have studied the capabilities of these models…

机器学习 · 计算机科学 2023-08-08 Adish Singla

Vision Transformers (ViT) have recently demonstrated the significant potential of transformer architectures for computer vision. To what extent can image-based deep reinforcement learning also benefit from ViT architectures, as compared to…

机器学习 · 计算机科学 2022-05-17 Tianxin Tao , Daniele Reda , Michiel van de Panne

Deep learning technology can be used as an assistive technology to help doctors quickly and accurately identify COVID-19 infections. Recently, Vision Transformer (ViT) has shown great potential towards image classification due to its global…

图像与视频处理 · 电气工程与系统科学 2022-07-06 Hongyan Xu , Xiu Su , Dadong Wang

Enhancement is an important step in post-processing digital images for personal use, in medical imaging, and for object recognition. Most existing manual techniques rely on region selection, similarity, and/or thresholding for editing,…

图形学 · 计算机科学 2019-09-05 Junyi Tu , Paul Rosen

Analyzing radiology reports is a time-consuming and error-prone task, which raises the need for an efficient automated radiology report analysis system to alleviate the workloads of radiologists and encourage precise diagnosis. In this…

计算与语言 · 计算机科学 2022-04-21 Song Wang , Mingquan Lin , Ying Ding , George Shih , Zhiyong Lu , Yifan Peng

This preliminary study explores the integration of GPT-4 Vision (GPT-4V) technology into teacher analytics, focusing on its applicability in observational assessment to enhance reflective teaching practice. This research is grounded in…

This is to present a text image classifier device that identifies textual content in images and then categorizes each image into one of four predefined categories, including Invoice, Form, Letter, or Report. The device supports a gallery…

计算机视觉与模式识别 · 计算机科学 2025-12-15 Aya Kaysan Bahjat

While the Transformer architecture has become the de-facto standard for natural language processing tasks, its applications to computer vision remain limited. In vision, attention is either applied in conjunction with convolutional…

‹ 上一页 1 8 9 10 下一页 ›