中文
相关论文

相关论文: From Text to Image: Exploring GPT-4Vision's Potent…

200 篇论文

This review surveys the state-of-the-art in text-to-image and image-to-image generation within the scope of generative AI. We provide a comparative analysis of three prominent architectures: Variational Autoencoders, Generative Adversarial…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Zineb Sordo , Eric Chagnon , Daniela Ushizima

The challenge of improving translation accuracy in GPT-4 is being addressed by harnessing a method known as in-context learning. This paper introduces a strategic approach to utilize in-context learning specifically for machine translation,…

计算与语言 · 计算机科学 2024-01-11 Yufeng Chen

Recent escalation in the field of computer vision underpins a huddle of algorithms with the magnificent potential to unravel the information contained within images. These computer vision algorithms are being practised in medical image…

图像与视频处理 · 电气工程与系统科学 2022-03-30 Arshi Parvaiz , Muhammad Anwaar Khalid , Rukhsana Zafar , Huma Ameer , Muhammad Ali , Muhammad Moazam Fraz

Large language models (LLMs) and vision-language models (VLMs) have been increasingly used in robotics for high-level cognition, but their use for low-level cognition, such as interpreting sensor information, remains underexplored. In…

机器人学 · 计算机科学 2024-08-09 Masashi Osada , Gustavo A. Garcia Ricardez , Yosuke Suzuki , Tadahiro Taniguchi

Clinicians spend a significant amount of time reviewing medical images and transcribing their findings regarding patient diagnosis, referral and treatment in text form. Vision-language models (VLMs), which automatically interpret images and…

Deep learning relies heavily on data augmentation to mitigate limited data, especially in medical imaging. Recent multimodal learning integrates text and images for segmentation, known as referring or text-guided image segmentation.…

计算机视觉与模式识别 · 计算机科学 2025-10-15 Shurong Chai , Rahul Kumar JAIN , Rui Xu , Shaocong Mo , Ruibo Hou , Shiyu Teng , Jiaqing Liu , Lanfen Lin , Yen-Wei Chen

The recent surge of foundation models in computer vision and natural language processing opens up perspectives in utilizing multi-modal clinical data to train large models with strong generalizability. Yet pathological image datasets often…

计算机视觉与模式识别 · 计算机科学 2023-07-28 Yunkun Zhang , Jin Gao , Mu Zhou , Xiaosong Wang , Yu Qiao , Shaoting Zhang , Dequan Wang

Our daily life is highly influenced by what we consume and see. Attracting and holding one's attention -- the definition of (visual) interestingness -- is essential. The rise of Large Multimodal Models (LMMs) trained on large-scale visual…

计算机视觉与模式识别 · 计算机科学 2025-10-16 Fitim Abdullahu , Helmut Grabner

In radiology, Artificial Intelligence (AI) has significantly advanced report generation, but automatic evaluation of these AI-produced reports remains challenging. Current metrics, such as Conventional Natural Language Generation (NLG) and…

Large multimodal models (LMMs) extend large language models (LLMs) with multi-sensory skills, such as visual understanding, to achieve stronger generic intelligence. In this paper, we analyze the latest model, GPT-4V(ision), to deepen the…

计算机视觉与模式识别 · 计算机科学 2023-10-12 Zhengyuan Yang , Linjie Li , Kevin Lin , Jianfeng Wang , Chung-Ching Lin , Zicheng Liu , Lijuan Wang

Diffusion models have revitalized the image generation domain, playing crucial roles in both academic research and artistic expression. With the emergence of new diffusion models, assessing the performance of text-to-image models has become…

计算机视觉与模式识别 · 计算机科学 2024-11-15 Chutian Meng , Fan Ma , Jiaxu Miao , Chi Zhang , Yi Yang , Yueting Zhuang

This study investigates the foundational characteristics of image-to-image translation networks, specifically examining their suitability and transferability within the context of routine clinical environments, despite achieving high levels…

图像与视频处理 · 电气工程与系统科学 2024-07-29 Mohammad R. Salmanpour , Amin Mousavi , Yixi Xu , William B Weeks , Ilker Hacihaliloglu

Image captioning as a multimodal task has drawn much interest in recent years. However, evaluation for this task remains a challenging problem. Existing evaluation metrics focus on surface similarity between a candidate caption and a set of…

计算与语言 · 计算机科学 2019-12-20 Huiyuan Xie , Tom Sherborne , Alexander Kuhnle , Ann Copestake

Determining the extent to which the perceptual world can be recovered from language is a longstanding problem in philosophy and cognitive science. We show that state-of-the-art large language models can unlock new insights into this problem…

计算与语言 · 计算机科学 2023-06-16 Raja Marjieh , Ilia Sucholutsky , Pol van Rijn , Nori Jacoby , Thomas L. Griffiths

Low-dose Positron Emission Tomography (PET) imaging presents a significant challenge due to increased noise and reduced image quality, which can compromise its diagnostic accuracy and clinical utility. Denoising diffusion probabilistic…

图像与视频处理 · 电气工程与系统科学 2025-03-03 Boxiao Yu , Savas Ozdemir , Jiong Wu , Yizhou Chen , Ruogu Fang , Kuangyu Shi , Kuang Gong

Large language models (LLMs) have demonstrated remarkable capabilities in natural language understanding and generation across various domains, including medicine. We present a comprehensive evaluation of GPT-4, a state-of-the-art LLM, on…

计算与语言 · 计算机科学 2023-04-13 Harsha Nori , Nicholas King , Scott Mayer McKinney , Dean Carignan , Eric Horvitz

Large language models (LLMs) have increased interest in vision language models (VLMs), which process image-text pairs as input. Studies investigating the visual understanding ability of VLMs have been proposed, but such studies are still…

计算与语言 · 计算机科学 2024-06-25 Jesse Atuhurra , Iqra Ali , Tatsuya Hiraoka , Hidetaka Kamigaito , Tomoya Iwakura , Taro Watanabe

The morphology of retinal blood vessels can indicate various diseases in the human body, and researchers have been working on automatic scanning and segmentation of retinal images to aid diagnosis. This project compares the performance of…

图像与视频处理 · 电气工程与系统科学 2023-03-20 Ifeyinwa Linda Anene , Yongmin Li

Medical images constitute a source of information essential for disease diagnosis, treatment and follow-up. In addition, due to its patient-specific nature, imaging information represents a critical component required for advancing…

计算机视觉与模式识别 · 计算机科学 2016-06-10 Dorin Comaniciu , Klaus Engel , Bogdan Georgescu , Tommaso Mansi

Text-to-image generation and text-guided image manipulation have received considerable attention in the field of image generation tasks. However, the mainstream evaluation methods for these tasks have difficulty in evaluating whether all…

计算机视觉与模式识别 · 计算机科学 2024-11-18 Mizuki Miyamoto , Ryugo Morita , Jinjia Zhou