English
Related papers

Related papers: DFIR-Metric: A Benchmark Dataset for Evaluating La…

200 papers

Large language models (LLMs) are incredible and versatile tools for text-based tasks that have enabled countless, previously unimaginable, applications. Retrieval models, in contrast, have not yet seen such capable general-purpose models…

Information Retrieval · Computer Science 2025-09-10 Julian Killingback , Hamed Zamani

Software vulnerability detection is generally supported by automated static analysis tools, which have recently been reinforced by deep learning (DL) models. However, despite the superior performance of DL-based approaches over rule-based…

Software Engineering · Computer Science 2024-05-03 Yanjing Yang , Xin Zhou , Runfeng Mao , Jinwei Xu , Lanxin Yang , Yu Zhangm , Haifeng Shen , He Zhang

The proliferation of Large Language Models (LLMs) has intensified concerns about manipulative or deceptive behaviors that can undermine user autonomy, trust, and well-being. Existing safety benchmarks predominantly rely on coarse binary…

Artificial Intelligence · Computer Science 2025-12-30 Sadia Asif , Israel Antonio Rosales Laguan , Haris Khan , Shumaila Asif , Muneeb Asif

Multimodal Large Language Models (MLLMs) have made substantial progress in recent years. However, their rigorous evaluation within specialized domains like finance is hindered by the absence of datasets characterized by professional-level…

Artificial Intelligence · Computer Science 2025-11-25 Shuangyan Deng , Haizhou Peng , Jiachen Xu , Rui Mao , Ciprian Doru Giurcăneanu , Jiamou Liu

Large Language Models (LLMs) have significantly advanced the field of Natural Language Processing (NLP), achieving remarkable performance across diverse tasks and enabling widespread real-world applications. However, LLMs are prone to…

Computation and Language · Computer Science 2024-06-12 Wen Luo , Tianshu Shen , Wei Li , Guangyue Peng , Richeng Xuan , Houfeng Wang , Xi Yang

The Test and Measurement domain, known for its strict requirements for accuracy and efficiency, is increasingly adopting Generative AI technologies to enhance the performance of data analysis, automation, and decision-making processes.…

Artificial Intelligence · Computer Science 2025-08-05 Emmanuel A. Olowe , Danial Chitnis

Large Vision-Language Models (LVLMs) suffer from hallucination issues, wherein the models generate plausible-sounding but factually incorrect outputs, undermining their reliability. A comprehensive quantitative evaluation is necessary to…

Computation and Language · Computer Science 2024-10-07 Haoyi Qiu , Wenbo Hu , Zi-Yi Dou , Nanyun Peng

Large Language Models (LLMs) are increasingly used across various domains, from software development to cyber threat intelligence. Understanding all the different fields of cybersecurity, which includes topics such as cryptography, reverse…

Artificial Intelligence · Computer Science 2024-06-04 Norbert Tihanyi , Mohamed Amine Ferrag , Ridhi Jain , Tamas Bisztray , Merouane Debbah

The rapid rise in popularity of Large Language Models (LLMs) with emerging capabilities has spurred public curiosity to evaluate and compare different LLMs, leading many researchers to propose their own LLM benchmarks. Noticing preliminary…

Artificial Intelligence · Computer Science 2025-05-15 Timothy R. McIntosh , Teo Susnjak , Nalin Arachchilage , Tong Liu , Paul Watters , Malka N. Halgamuge

Large Language Models (LLMs) have shown promise in various domains, including healthcare, with significant potential to transform mental health applications by enabling scalable and accessible solutions. This study aims to provide a…

Artificial Intelligence · Computer Science 2025-11-25 Abdelrahman Hanafi , Mohammed Saad , Noureldin Zahran , Radwa J. Hanafy , Mohammed E. Fouda

The advancement of large language models (LLMs) has enhanced the ability to generalize across a wide range of unseen natural language processing (NLP) tasks through instruction-following. Yet, their effectiveness often diminishes in…

A Large Language Model (LLM) as judge evaluates the quality of victim Machine Learning (ML) models, specifically LLMs, by analyzing their outputs. An LLM as judge is the combination of one model and one specifically engineered judge prompt…

Cryptography and Security · Computer Science 2026-03-24 Tom Biskupski , Stephan Kleber

Psychological support hotlines serve as critical lifelines for crisis intervention but encounter significant challenges due to rising demand and limited resources. Large language models (LLMs) offer potential support in crisis assessments,…

Computation and Language · Computer Science 2025-12-19 Guifeng Deng , Shuyin Rao , Tianyu Lin , Anlu Dai , Pan Wang , Junyi Xie , Haidong Song , Ke Zhao , Dongwu Xu , Zhengdong Cheng , Tao Li , Haiteng Jiang

As Large Language Models (LLMs) have risen in prominence over the past few years, there has been concern over the potential biases in LLMs inherited from the training data. Previous studies have examined how LLMs exhibit implicit bias, such…

Computation and Language · Computer Science 2025-12-30 Lake Yin , Fan Huang

Recent strides in Large Language Models (LLMs) have saturated many Natural Language Processing (NLP) benchmarks, emphasizing the need for more challenging ones to properly assess LLM capabilities. However, domain-specific and multilingual…

The evaluation of large language models (LLMs) via benchmarks is widespread, yet inconsistencies between different leaderboards and poor separability among top models raise concerns about their ability to accurately reflect authentic model…

Computation and Language · Computer Science 2026-01-19 Hongli Zhou , Hui Huang , Ziqing Zhao , Lvyuan Han , Huicheng Wang , Kehai Chen , Muyun Yang , Wei Bao , Jian Dong , Bing Xu , Conghui Zhu , Hailong Cao , Tiejun Zhao

Multi-modal Large Language Models (MLLMs) are increasingly prominent in the field of artificial intelligence. These models not only excel in traditional vision-language tasks but also demonstrate impressive performance in contemporary…

Computer Vision and Pattern Recognition · Computer Science 2023-12-06 Xiaotian Han , Quanzeng You , Yongfei Liu , Wentao Chen , Huangjie Zheng , Khalil Mrini , Xudong Lin , Yiqi Wang , Bohan Zhai , Jianbo Yuan , Heng Wang , Hongxia Yang

The exponential growth of digital content has generated massive textual datasets, necessitating the use of advanced analytical approaches. Large Language Models (LLMs) have emerged as tools that are capable of processing and extracting…

Computation and Language · Computer Science 2024-05-24 Benjamin M. Ampel , Chi-Heng Yang , James Hu , Hsinchun Chen

Recent advances in AIGC have exacerbated the misuse of malicious deepfake content, making the development of reliable deepfake detection methods an essential means to address this challenge. Although existing deepfake detection models…

Computer Vision and Pattern Recognition · Computer Science 2025-10-31 Changtao Miao , Yi Zhang , Weize Gao , Zhiya Tan , Weiwei Feng , Man Luo , Jianshu Li , Ajian Liu , Yunfeng Diao , Qi Chu , Tao Gong , Zhe Li , Weibin Yao , Joey Tianyi Zhou

Large Language Models (LLMs) are rapidly transforming the landscape of digital content creation. However, the prevalent black-box Application Programming Interface (API) access to many LLMs introduces significant challenges in…

Cryptography and Security · Computer Science 2026-01-21 Zhiyuan Fu , Junfan Chen , Lan Zhang , Ting Yang , Jun Niu , Hongyu Sun , Ruidong Li , Peng Liu , Jice Wang , Fannv He , Qiuling Yue , Yuqing Zhang