English
Related papers

Related papers: Evaluating Gemini LLM in Food Image-Based Recipe a…

200 papers

Clinical document classification is essential for converting unstructured medical texts into standardised ICD-10 diagnoses, yet it faces challenges due to complex medical language, privacy constraints, and limited annotated datasets. Large…

Computation and Language · Computer Science 2026-02-03 Akram Mustafa , Usman Naseem , Mostafa Rahimi Azghadi

Scientific discovery is an inherently creative and uncertain process, requiring reasoning beyond the recall of known knowledge. While many benchmarks have been proposed to evaluate large language model (LLM) performance on deep research…

Artificial Intelligence · Computer Science 2026-05-29 A. J. Lew , Y. Cao , M. J. Buehler

LLMs have demonstrated strong performance in data-rich domains such as programming, yet their reliability in engineering tasks remains limited. Circuit analysis--requiring multimodal understanding and precise mathematical…

Computers and Society · Computer Science 2026-04-17 Liangliang Chen , Weiyu Sun , Huiru Xie , Yongnuo Cai , Ying Zhang

Large language models (LLMs) are fundamentally transforming human-facing applications in the health and well-being domains: boosting patient engagement, accelerating clinical decision-making, and facilitating medical education. Although…

Computation and Language · Computer Science 2024-08-09 Iman Azimi , Mohan Qi , Li Wang , Amir M. Rahmani , Youlin Li

Recent breakthroughs in large multimodal models (LMMs) have significantly advanced both text-to-image (T2I) generation and image-to-text (I2T) interpretation. However, many generated images still suffer from issues related to perceptual…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Jiarui Wang , Huiyu Duan , Yu Zhao , Juntong Wang , Guangtao Zhai , Xiongkuo Min

Emotion recognition capabilities in multimodal AI systems are crucial for developing culturally responsive educational technologies, yet remain underexplored for Arabic language contexts where culturally appropriate learning tools are…

Computation and Language · Computer Science 2025-09-05 Bushra Asseri , Estabraq Abdelaziz , Maha Al Mogren , Tayef Alhefdhi , Areej Al-Wabil

Automating the classification of negative treatment in legal precedent is a critical yet nuanced NLP task where misclassification carries significant risk. To address the shortcomings of standard accuracy, this paper introduces a more…

Computation and Language · Computer Science 2026-05-19 M. Mikail Demir , M. Abdullah Canbaz

The rapid advancement of Large Language Models (LLMs) in the realm of mathematical reasoning necessitates comprehensive evaluations to gauge progress and inspire future directions. Existing assessments predominantly focus on problem-solving…

Computation and Language · Computer Science 2024-06-05 Xiaoyuan Li , Wenjie Wang , Moxin Li , Junrong Guo , Yang Zhang , Fuli Feng

Vision-language models (VLMs) are increasingly proposed as general-purpose solutions for visual recognition tasks, yet their reliability for agricultural decision support remains poorly understood. We benchmark a diverse set of open-source…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Earl Ranario , Mason J. Earles

The rapid evolution of Multi-modality Large Language Models (MLLMs) is driving significant advancements in visual understanding and generation. Nevertheless, a comprehensive assessment of their capabilities, concerning the fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Xiaorong Zhu , Ziheng Jia , Jiarui Wang , Xiangyu Zhao , Haodong Duan , Xiongkuo Min , Jia Wang , Zicheng Zhang , Guangtao Zhai

Communication network engineering in enterprise environments is traditionally a complex, time-consuming, and error-prone manual process. Most research on network engineering automation has concentrated on configuration synthesis, often…

Networking and Internet Architecture · Computer Science 2025-08-27 Beni Ifland , Elad Duani , Rubin Krief , Miro Ohana , Aviram Zilberman , Andres Murillo , Ofir Manor , Ortal Lavi , Hikichi Kenji , Asaf Shabtai , Yuval Elovici , Rami Puzis

Multi-modal Large Language Models (MLLMs) have shown impressive abilities in generating reasonable responses with respect to multi-modal contents. However, there is still a wide gap between the performance of recent MLLM-based applications…

The exponential growth of information presents a significant challenge for researchers and professionals seeking to remain at the forefront of their fields and this paper introduces an innovative framework for automatically generating…

Computational Engineering, Finance, and Science · Computer Science 2025-10-30 Andrei Lazarev , Dmitrii Sedov

Recent advancements in Vision-Language Models (VLMs) have revolutionized general visual understanding. However, their application in the food domain remains constrained by benchmarks that rely on coarse-grained categories, single-view…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Song Jin , Juntian Zhang , Xun Zhang , Zeying Tian , Fei Jiang , Guojun Yin , Wei Lin , Yong Liu , Rui Yan

Medical images and radiology reports are crucial for diagnosing medical conditions, highlighting the importance of quantitative analysis for clinical decision-making. However, the diversity and cross-source heterogeneity of these data…

Image and Video Processing · Electrical Eng. & Systems 2024-07-09 Yutong Zhang , Yi Pan , Tianyang Zhong , Peixin Dong , Kangni Xie , Yuxiao Liu , Hanqi Jiang , Zhengliang Liu , Shijie Zhao , Tuo Zhang , Xi Jiang , Dinggang Shen , Tianming Liu , Xin Zhang

The detection of Protected Health Information (PHI) in medical imaging is critical for safeguarding patient privacy and ensuring compliance with regulatory frameworks. Traditional detection methodologies predominantly utilize Optical…

Computer Vision and Pattern Recognition · Computer Science 2026-05-22 Tuan Truong , Guillermo Jimenez Perez , Pedro Osorio , Matthias Lenga

Food image segmentation is a critical and indispensible task for developing health-related applications such as estimating food calories and nutrients. Existing food image segmentation models are underperforming due to two reasons: (1)…

Computer Vision and Pattern Recognition · Computer Science 2021-05-13 Xiongwei Wu , Xin Fu , Ying Liu , Ee-Peng Lim , Steven C. H. Hoi , Qianru Sun

Recent advances in reasoning-focused large language models (LLMs) mark a shift from general LLMs toward models designed for complex decision-making, a crucial aspect in medicine. However, their performance in specialized domains like…