English
Related papers

Related papers: CMB: A Comprehensive Medical Benchmark in Chinese

200 papers

The recent progress of large language models (LLMs), including ChatGPT and GPT-4, in comprehending and responding to human instructions has been remarkable. Nevertheless, these models typically perform better in English and have not been…

Computation and Language · Computer Science 2023-04-18 Honglin Xiong , Sheng Wang , Yitao Zhu , Zihao Zhao , Yuxiao Liu , Linlin Huang , Qian Wang , Dinggang Shen

This paper explores the application of prompt engineering to enhance the performance of large language models (LLMs) in the domain of Traditional Chinese Medicine (TCM). We propose TCM-Prompt, a framework that integrates various pre-trained…

Computation and Language · Computer Science 2024-10-28 Yirui Chen , Qinyu Xiao , Jia Yi , Jing Chen , Mengyang Wang

The surge of large language models (LLMs) has driven significant progress in medical applications, including traditional Chinese medicine (TCM). However, current medical LLMs struggle with TCM diagnosis and syndrome differentiation due to…

Computation and Language · Computer Science 2025-10-01 Sibo Wei , Xueping Peng , Yi-Fei Wang , Tao Shen , Jiasheng Si , Weiyu Zhang , Fa Zhu , Athanasios V. Vasilakos , Wenpeng Lu , Xiaoming Wu , Yinglong Wang

Large Language Models (LLMs) are poised to transform healthcare under China's Healthy China 2030 initiative, yet they introduce new ethical and patient-safety challenges. We present a novel 12,000-item Q&A benchmark covering 11 ethics and 9…

Computation and Language · Computer Science 2025-05-13 Mouxiao Bian , Rongzhao Zhang , Chao Ding , Xinwei Peng , Jie Xu

Large language models (LLMs) have demonstrated remarkable capabilities across various applications, highlighting the urgent need for comprehensive safety evaluations. In particular, the enhanced Chinese language proficiency of LLMs,…

Computation and Language · Computer Science 2025-02-27 Shuyi Liu , Simiao Cui , Haoran Bu , Yuming Shang , Xi Zhang

Recent advances in large language models (LLMs) have led to substantial progress in domain-specific applications, particularly within the legal domain. However, general-purpose models such as GPT-4 often struggle with specialized subdomains…

Artificial Intelligence · Computer Science 2026-01-16 Zixun Lan , Maochun Xu , Yifan Ren , Rui Wu , Jianghui Zhou , Xueyang Cheng , Jianan Ding Ding , Xinheng Wang , Mingmin Chi , Fei Ma

Large language models (LLMs) are advancing rapidly in medical NLP, yet Traditional Chinese Medicine (TCM) with its distinctive ontology, terminology, and reasoning patterns requires domain-faithful evaluation. Existing TCM benchmarks are…

Generative large language models (LLMs) have shown great success in various applications, including question-answering (QA) and dialogue systems. However, in specialized domains like traditional Chinese medical QA, these models may perform…

Computation and Language · Computer Science 2023-09-06 Yang Tan , Mingchen Li , Zijie Huang , Huiqun Yu , Guisheng Fan

Despite the success of large language models (LLMs) in various domains, their potential in Traditional Chinese Medicine (TCM) remains largely underexplored due to two critical barriers: (1) the scarcity of high-quality TCM data and (2) the…

Computation and Language · Computer Science 2025-08-21 Junying Chen , Zhenyang Cai , Zhiheng Liu , Yunjin Yang , Rongsheng Wang , Qingying Xiao , Xiangyi Feng , Zhan Su , Jing Guo , Xiang Wan , Guangjun Yu , Haizhou Li , Benyou Wang

Open-ended medical LLM evaluation remains weakly grounded in physician-calibrated coverage of clinically relevant response criteria, especially in localized clinical settings. We introduce \textsc{ClinConsensus}, a Chinese medical benchmark…

Computation and Language · Computer Science 2026-05-28 Xiang Zheng , Han Li , Wenjie Luo , Weiqi Zhai , Yiyuan Li , Chuanmiao Yan , Xue Yang , Kailuan Wu , Ruyi Xu , Tianyun Lu , Tianyi Tang , Yubo Ma , Kexin Yang , Dayiheng Liu , Sen Yang , Lin Qu , Bing Zhao , Hu Wei

Natural medicines, particularly Traditional Chinese Medicine (TCM), are gaining global recognition for their therapeutic potential in addressing human symptoms and diseases. TCM, with its systematic theories and extensive practical…

Computation and Language · Computer Science 2025-05-20 Zhi Liu , Tao Yang , Jing Wang , Yexin Chen , Zhan Gao , Jiaxi Yang , Kui Chen , Bingji Lu , Xiaochen Li , Changyong Luo , Yan Li , Xiaohong Gu , Peng Cao

Large language models (LLMs) have demonstrated exceptional capabilities in general domains, yet their application in highly specialized and culturally-rich fields like Traditional Chinese Medicine (TCM) requires rigorous and nuanced…

Computation and Language · Computer Science 2025-11-18 Tianai Huang , Jiayuan Chen , Lu Lu , Pengcheng Chen , Tianbin Li , Bing Han , Wenchao Tang , Jie Xu , Ming Li

The advancement of large language models (LLMs) has enhanced the ability to generalize across a wide range of unseen natural language processing (NLP) tasks through instruction-following. Yet, their effectiveness often diminishes in…

As ChatGPT and GPT-4 spearhead the development of Large Language Models (LLMs), more researchers are investigating their performance across various tasks. But more research needs to be done on the interpretability capabilities of LLMs, that…

Computation and Language · Computer Science 2023-10-27 Dongfang Li , Jindi Yu , Baotian Hu , Zhenran Xu , Min Zhang

Evaluating large language models (LLMs) in medicine is crucial because medical applications require high accuracy with little room for error. Current medical benchmarks have three main types: medical exam-based, comprehensive medical, and…

The adoption of large language models (LLMs) to assist clinicians has attracted remarkable attention. Existing works mainly adopt the close-ended question-answering (QA) task with answer options for evaluation. However, many clinical…

Temporal reasoning is fundamental to human cognition and is crucial for various real-world applications. While recent advances in Large Language Models have demonstrated promising capabilities in temporal reasoning, existing benchmarks…

Computation and Language · Computer Science 2025-02-25 Zhenglin Wang , Jialong Wu , Pengfei LI , Yong Jiang , Deyu Zhou

Large language models (LLMs) are increasingly deployed in cost-sensitive and on-device scenarios, and safety guardrails have advanced mainly in English. However, real-world Chinese malicious queries typically conceal intent via homophones,…

Computation and Language · Computer Science 2026-01-06 Zhenhong Zhou , Shilinlu Yan , Chuanpu Liu , Qiankun Li , Kun Wang , Zhigang Zeng

Traditional Chinese Medicine (TCM), with a history spanning over two millennia, plays a role in global healthcare. However, applying large language models (LLMs) to TCM remains challenging due to its reliance on holistic reasoning, implicit…

Computation and Language · Computer Science 2025-10-21 Jiacheng Xie , Yang Yu , Yibo Chen , Hanyao Zhang , Lening Zhao , Jiaxuan He , Lei Jiang , Xiaoting Tang , Guanghui An , Dong Xu

Medical report interpretation plays a crucial role in healthcare, enabling both patient-facing explanations and effective information flow across clinical systems. While recent vision-language models (VLMs) and large language models (LLMs)…

Computer Vision and Pattern Recognition · Computer Science 2025-08-26 Fangxin Shang , Yuan Xia , Dalu Yang , Yahui Wang , Binglin Yang