中文
相关论文

相关论文: A Foundational Multimodal Vision Language AI Assis…

200 篇论文

Advances in markerless motion capture are expanding access to biomechanical movement analysis, making it feasible to obtain high-quality movement data from outpatient clinics, inpatient hospitals, therapy, and even home. Expanding access to…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Ruize Yang , Ann Kennedy , R. James Cotton

Generative models such as StyleGAN2 and Stable Diffusion have achieved state-of-the-art performance in computer vision tasks such as image synthesis, inpainting, and de-noising. However, current generative models for face inpainting often…

计算机视觉与模式识别 · 计算机科学 2023-04-14 Saman Motamed , Jianjin Xu , Chen Henry Wu , Fernando De la Torre

We introduce the task of Visual Dialog, which requires an AI agent to hold a meaningful dialog with humans in natural, conversational language about visual content. Specifically, given an image, a dialog history, and a question about the…

计算机视觉与模式识别 · 计算机科学 2017-08-03 Abhishek Das , Satwik Kottur , Khushi Gupta , Avi Singh , Deshraj Yadav , José M. F. Moura , Devi Parikh , Dhruv Batra

Vision Transformers (ViTs) have gained rapid adoption in computational pathology for their ability to model long-range dependencies through self-attention, addressing the limitations of convolutional neural networks that excel at local…

图像与视频处理 · 电气工程与系统科学 2026-01-15 Fuyao Chen , Yuexi Du , Elèonore V. Lieffrig , Nicha C. Dvornek , John A. Onofrey

Vision-threatening eye diseases pose a major global health burden, with timely diagnosis limited by workforce shortages and restricted access to specialized care. While multimodal large language models (MLLMs) show promise for medical image…

One of the key goals of artificial intelligence (AI) is the development of a multimodal system that facilitates communication with the visual world (image and video) using a natural language query. Earlier works on medical question…

计算机视觉与模式识别 · 计算机科学 2024-12-17 Deepak Gupta , Dina Demner-Fushman

The development of robust artificial intelligence models for histopathology diagnosis is severely constrained by the scarcity of expert-annotated lesion data, particularly for rare pathologies and underrepresented disease subtypes. While…

计算机视觉与模式识别 · 计算机科学 2026-01-14 Mohamad Koohi-Moghadam , Mohammad-Ali Nikouei Mahani , Kyongtae Tyler Bae

The use of artificial intelligence to enable precision medicine and decision support systems through the analysis of pathology images has the potential to revolutionize the diagnosis and treatment of cancer. Such applications will depend on…

There exist numerous diagnostic tasks in pathology. Conventional computational pathology formulates and tackles them as independent and individual image classification problems, thereby resulting in computational inefficiency and high…

图像与视频处理 · 电气工程与系统科学 2024-07-15 Anh Tien Nguyen , Keunho Byeon , Kyungeun Kim , Boram Song , Seoung Wan Chae , Jin Tae Kwak

Multimodal pathological image understanding has garnered widespread interest due to its potential to improve diagnostic accuracy and enable personalized treatment through integrated visual and textual data. However, existing methods exhibit…

计算机视觉与模式识别 · 计算机科学 2025-05-22 Zhe Xu , Cheng Jin , Yihui Wang , Ziyi Liu , Hao Chen

Recent advances in whole slide imaging (WSI) technology have led to the development of a myriad of computer vision and artificial intelligence (AI) based diagnostic, prognostic, and predictive algorithms. Computational Pathology (CPath)…

Recent advances in large vision-language models (LVLMs) have demonstrated strong performance on general-purpose medical tasks. However, their effectiveness in specialized domains such as dentistry remains underexplored. In particular,…

计算机视觉与模式识别 · 计算机科学 2025-09-12 Jing Hao , Yuxuan Fan , Yanpeng Sun , Kaixin Guo , Lizhuo Lin , Jinrong Yang , Qi Yong H. Ai , Lun M. Wong , Hao Tang , Kuo Feng Hung

In healthcare, AI techniques are widely used for tasks like risk assessment and anomaly detection. Despite AI's potential as a valuable assistant, its role in complex medical data analysis often oversimplifies human-AI collaboration…

人机交互 · 计算机科学 2024-07-23 Yang Ouyang , Chenyang Zhang , He Wang , Tianle Ma , Chang Jiang , Yuheng Yan , Zuoqin Yan , Xiaojuan Ma , Chuhan Shi , Quan Li

Multimodal large language models (LMMs) excel in world knowledge and problem-solving abilities. Through the use of a world-facing camera and contextual AI, emerging smart accessories aim to provide a seamless interface between humans and…

人机交互 · 计算机科学 2024-02-01 Robert Konrad , Nitish Padmanaban , J. Gabriel Buckmaster , Kevin C. Boyle , Gordon Wetzstein

Large language models (LLMs) have recently demonstrated their potential in clinical applications, providing valuable medical knowledge and advice. For example, a large dialog LLM like ChatGPT has successfully passed part of the US medical…

计算机视觉与模式识别 · 计算机科学 2023-02-15 Sheng Wang , Zihao Zhao , Xi Ouyang , Qian Wang , Dinggang Shen

In pathological research, education, and clinical practice, the decision-making process based on pathological images is critically important. This significance extends to digital pathology image analysis: its adequacy is demonstrated by the…

图像与视频处理 · 电气工程与系统科学 2024-08-19 Zhi-Bo Liu , Xiaobo Pang , Jizhao Wang , Shuai Liu , Chen Li

Despite strong performance of deep learning models in retinal disease detection, most systems produce static predictions without clinical reasoning or interactive explanation. Recent advances in multimodal large language models (MLLMs)…

计算机视觉与模式识别 · 计算机科学 2026-04-29 Ran Gu , Benjamin Hou , Mélanie Hébert , Asmita Indurkar , Yifan Yang , Emily Y. Chew , Tiarnán D. L. Keenan , Zhiyong Lu

Immunohistochemistry (IHC) provides information on protein expression in tissue sections and is commonly used to support pathology diagnosis and disease triage. While AI models for H\&E-stained slides show promise, their applicability to…

Generating accurate and clinically meaningful radiology reports from chest X-ray images remains a significant challenge in medical AI. While recent vision-language models achieve strong results in general radiology report generation, they…

计算机视觉与模式识别 · 计算机科学 2025-11-13 Nikolay Nechaev , Evgeniia Przhezdzetskaia , Dmitry Umerenkov , Dmitry V. Dylov

We investigate the multilingual and multimodal performance of a large language model-based artificial intelligence (AI) system, GPT-4o, using a diverse set of physics concept inventories spanning multiple languages and subject categories.…

物理教育 · 物理学 2025-07-14 Gerd Kortemeyer , Marina Babayeva , Giulia Polverini , Ralf Widenhorn , Bor Gregorcic
‹ 上一页 1 8 9 10 下一页 ›