中文
相关论文

相关论文: A Multimodal Benchmark Dataset and Model for Crop …

200 篇论文

Efficient and sustainable crop production process management is crucial to meet the growing global demand for food, fuel, and feed while minimizing environmental impacts. Traditional crop management practices, often developed through…

系统与控制 · 电气工程与系统科学 2024-10-15 Dong Chen , Yanbo Huang

The prevalence of vision-threatening eye diseases is a significant global burden, with many cases remaining undiagnosed or diagnosed too late for effective treatment. Large vision-language models (LVLMs) have the potential to assist in…

计算机视觉与模式识别 · 计算机科学 2025-02-06 Zhenyue Qin , Yu Yin , Dylan Campbell , Xuansheng Wu , Ke Zou , Yih-Chung Tham , Ninghao Liu , Xiuzhen Zhang , Qingyu Chen

Document Visual Question Answering (DocVQA) faces dual challenges in processing lengthy multimodal documents (text, images, tables) and performing cross-modal reasoning. Current document retrieval-augmented generation (DocRAG) methods…

信息检索 · 计算机科学 2025-11-10 Kuicai Dong , Yujing Chang , Shijie Huang , Yasheng Wang , Ruiming Tang , Yong Liu

Crops, fisheries and livestock form the backbone of global food production, essential to feed the ever-growing global population. However, these sectors face considerable challenges, including climate variability, resource limitations, and…

计算机视觉与模式识别 · 计算机科学 2026-05-06 Umair Nawaz , Muhammad Zaigham Zaheer , Ufaq Khan , Fahad Shahbaz Khan , Hisham Cholakkal , Salman Khan , Rao Muhammad Anwer

Early diagnosis of plant diseases is critical for global food safety, yet most AI solutions lack the generalization required for real-world agricultural diversity. These models are typically constrained to specific species, failing to…

计算机视觉与模式识别 · 计算机科学 2025-08-26 Saif Ur Rehman Khan , Muhammad Nabeel Asim , Sebastian Vollmer , Andreas Dengel

Multimodal/vision language models (VLMs) are increasingly being deployed in healthcare settings worldwide, necessitating robust benchmarks to ensure their safety, efficacy, and fairness. Multiple-choice question and answer (QA) datasets…

Unified vision large language models (VLLMs) have recently achieved impressive advancements in both multimodal understanding and generation, powering applications such as visual question answering and text-guided image synthesis. However,…

计算与语言 · 计算机科学 2025-09-19 Pengyu Wang , Shaojun Zhou , Chenkun Tan , Xinghao Wang , Wei Huang , Zhen Ye , Zhaowei Li , Botian Jiang , Dong Zhang , Xipeng Qiu

Most current medical vision language models struggle to jointly generate diagnostic text and pixel-level segmentation masks in response to complex visual questions. This represents a major limitation towards clinical application, as…

计算机视觉与模式识别 · 计算机科学 2025-10-22 Chengrun Li , Corentin Royer , Haozhe Luo , Bastian Wittmann , Xia Li , Ibrahim Hamamci , Sezgin Er , Anjany Sekuboyina , Bjoern Menze

While recent advancements in vision-language models have had a transformative impact on multi-modal comprehension, the extent to which these models possess the ability to comprehend generated images remains uncertain. Synthetic images, in…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Keqiang Sun , Junting Pan , Yuying Ge , Hao Li , Haodong Duan , Xiaoshi Wu , Renrui Zhang , Aojun Zhou , Zipeng Qin , Yi Wang , Jifeng Dai , Yu Qiao , Limin Wang , Hongsheng Li

Large Vision-Language Models (LVLMs) are capable of handling diverse data types such as imaging, text, and physiological signals, and can be applied in various fields. In the medical field, LVLMs have a high potential to offer substantial…

Differential diagnosis is an iterative process that integrates patient information with broader medical knowledge. Clinical case series such as the NEJM Clinicopathologic Conferences (CPCs), published continuously since 1923, feature expert…

Agriculture is a key sector of the economies of developing countries. It serves as a primary source of income and employment for rural populations. However, each year, a large portion of crops is wasted because of pests and diseases.…

计算机视觉与模式识别 · 计算机科学 2026-05-05 Muhammad Kaleem Ullah Khan

Recent advancements in Large Vision-Language Models (LVLMs) have demonstrated impressive capabilities across various multimodal tasks. They continue, however, to struggle with trivial scenarios such as reading values from Digital…

计算机视觉与模式识别 · 计算机科学 2025-09-01 João Valente , Atabak Dehban , Rodrigo Ventura

As a social being, we have an intimate bond with the environment. A plethora of things in human life, such as lifestyle, health, and food are dependent on the environment and agriculture. It comes under our responsibility to support the…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Semanto Mondal

Large language models (LLMs) have recently demonstrated their potential in clinical applications, providing valuable medical knowledge and advice. For example, a large dialog LLM like ChatGPT has successfully passed part of the US medical…

计算机视觉与模式识别 · 计算机科学 2023-02-15 Sheng Wang , Zihao Zhao , Xi Ouyang , Qian Wang , Dinggang Shen

Therapeutic art activities, such as expressive drawing and painting, require the synergy between creative visual production and interactive dialogue. Recent advancements in Multimodal Large Language Models (MLLMs) have expanded the capacity…

人机交互 · 计算机科学 2026-05-12 Le Lin , Zihao Zhu , Rainbow Tin Hung Ho , Jing Liao , Yuhan Luo

The task of visual dialog requires a multimodal chatbot to answer sequential questions from humans about image content. Prior work performs the standard likelihood training for answer generation on the positive instances (involving correct…

计算与语言 · 计算机科学 2022-11-28 Zihao Wang , Junli Wang , Changjun Jiang

Recent advances in generative medical models are constrained by modality-specific scenarios that hinder the integration of complementary evidence from imaging, pathology, and clinical notes. This fragmentation limits their evolution into…

计算机视觉与模式识别 · 计算机科学 2025-10-08 Jiawei Mao , Yuhan Wang , Lifeng Chen , Can Zhao , Yucheng Tang , Dong Yang , Liangqiong Qu , Daguang Xu , Yuyin Zhou

Plant disease diagnosis is critical for food security, yet training disease-recognition models that generalize across crops, pathogens, and field conditions remains challenging because labeled disease images are far less abundant and…

The classification of medical images is a pivotal aspect of disease diagnosis, often enhanced by deep learning techniques. However, traditional approaches typically focus on unimodal medical image data, neglecting the integration of diverse…

图像与视频处理 · 电气工程与系统科学 2025-11-11 Jun-En Ding , Chien-Chin Hsu , Chi-Hsiang Chu , Shuqiang Wang , Feng Liu