English
Related papers

Related papers: OpenLLM-RTL: Open Dataset and Benchmark for LLM-Ai…

200 papers

This study presents a comprehensive empirical evaluation of six state-of-the-art large language models (LLMs) for code generation, including both general-purpose and code-specialized models. Using a dataset of 944 real-world LeetCode…

Software Engineering · Computer Science 2025-12-23 Le Zhang , Suresh Kothari

Amid the expanding use of pre-training data, the phenomenon of benchmark dataset leakage has become increasingly prominent, exacerbated by opaque training processes and the often undisclosed inclusion of supervised data in contemporary…

Computation and Language · Computer Science 2024-04-30 Ruijie Xu , Zengzhi Wang , Run-Ze Fan , Pengfei Liu

Pre-trained code models rely heavily on high-quality pre-training data, particularly human-written reference comments that bridge code and natural language. However, these comments often become outdated as software evolves, degrading model…

Software Engineering · Computer Science 2025-04-29 Kang Yang , Xinjun Mao , Shangwen Wang , Yanlin Wang , Tanghaoran Zhang , Bo Lin , Yihao Qin , Zhang Zhang , Yao Lu , Kamal Al-Sabahi

Large language models (LLMs) are being increasingly integrated into practical hardware and firmware development pipelines for code generation. Existing studies have primarily focused on evaluating the functional correctness of LLM-generated…

Cryptography and Security · Computer Science 2026-01-21 Qirui Chen , Jingxian Shuai , Shuangwu Chen , Shenghao Ye , Zijian Wen , Xufei Su , Jie Jin , Jiangming Li , Jun Chen , Xiaobin Tan , Jian Yang

Recent Vision-based Large Language Models~(VisionLLMs) for autonomous driving have seen rapid advancements. However, such promotion is extremely dependent on large-scale high-quality annotated data, which is costly and labor-intensive. To…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Chaoqun Wang , Jie Yang , Xiaobin Hong , Ruimao Zhang

Current Large Language Models (LLMs) exhibit significant limitations, notably in structured, interpretable, and verifiable medical reasoning, alongside practical deployment challenges related to computational resources and data privacy.…

Computation and Language · Computer Science 2025-06-06 Boqin Zhuang , Chenxiao Song , Huitong Lu , Jiacheng Qiao , Mingqian Liu , Mingxing Yu , Ping Hong , Rui Li , Xiaoxia Song , Xiangjun Xu , Xu Chen , Yaoyao Ma , Yujie Gao

Large Language Model (LLM)-based systems present new opportunities for autonomous health monitoring in sensor-rich industrial environments. This study explores the potential of LLMs to detect and classify faults directly from sensor data,…

Artificial Intelligence · Computer Science 2025-09-30 Xian Yeow Lee , Lasitha Vidyaratne , Ahmed Farahat , Chetan Gupta

Large language models (LLMs) have been shown to be able to perform new tasks based on a few demonstrations or natural language instructions. While these capabilities have led to widespread adoption, most LLMs are developed by resource-rich…

Computation and Language · Computer Science 2023-06-28 BigScience Workshop , : , Teven Le Scao , Angela Fan , Christopher Akiki , Ellie Pavlick , Suzana Ilić , Daniel Hesslow , Roman Castagné , Alexandra Sasha Luccioni , François Yvon , Matthias Gallé , Jonathan Tow , Alexander M. Rush , Stella Biderman , Albert Webson , Pawan Sasanka Ammanamanchi , Thomas Wang , Benoît Sagot , Niklas Muennighoff , Albert Villanova del Moral , Olatunji Ruwase , Rachel Bawden , Stas Bekman , Angelina McMillan-Major , Iz Beltagy , Huu Nguyen , Lucile Saulnier , Samson Tan , Pedro Ortiz Suarez , Victor Sanh , Hugo Laurençon , Yacine Jernite , Julien Launay , Margaret Mitchell , Colin Raffel , Aaron Gokaslan , Adi Simhi , Aitor Soroa , Alham Fikri Aji , Amit Alfassy , Anna Rogers , Ariel Kreisberg Nitzav , Canwen Xu , Chenghao Mou , Chris Emezue , Christopher Klamm , Colin Leong , Daniel van Strien , David Ifeoluwa Adelani , Dragomir Radev , Eduardo González Ponferrada , Efrat Levkovizh , Ethan Kim , Eyal Bar Natan , Francesco De Toni , Gérard Dupont , Germán Kruszewski , Giada Pistilli , Hady Elsahar , Hamza Benyamina , Hieu Tran , Ian Yu , Idris Abdulmumin , Isaac Johnson , Itziar Gonzalez-Dios , Javier de la Rosa , Jenny Chim , Jesse Dodge , Jian Zhu , Jonathan Chang , Jörg Frohberg , Joseph Tobing , Joydeep Bhattacharjee , Khalid Almubarak , Kimbo Chen , Kyle Lo , Leandro Von Werra , Leon Weber , Long Phan , Loubna Ben allal , Ludovic Tanguy , Manan Dey , Manuel Romero Muñoz , Maraim Masoud , María Grandury , Mario Šaško , Max Huang , Maximin Coavoux , Mayank Singh , Mike Tian-Jian Jiang , Minh Chien Vu , Mohammad A. Jauhar , Mustafa Ghaleb , Nishant Subramani , Nora Kassner , Nurulaqilla Khamis , Olivier Nguyen , Omar Espejel , Ona de Gibert , Paulo Villegas , Peter Henderson , Pierre Colombo , Priscilla Amuok , Quentin Lhoest , Rheza Harliman , Rishi Bommasani , Roberto Luis López , Rui Ribeiro , Salomey Osei , Sampo Pyysalo , Sebastian Nagel , Shamik Bose , Shamsuddeen Hassan Muhammad , Shanya Sharma , Shayne Longpre , Somaieh Nikpoor , Stanislav Silberberg , Suhas Pai , Sydney Zink , Tiago Timponi Torrent , Timo Schick , Tristan Thrush , Valentin Danchev , Vassilina Nikoulina , Veronika Laippala , Violette Lepercq , Vrinda Prabhu , Zaid Alyafeai , Zeerak Talat , Arun Raja , Benjamin Heinzerling , Chenglei Si , Davut Emre Taşar , Elizabeth Salesky , Sabrina J. Mielke , Wilson Y. Lee , Abheesht Sharma , Andrea Santilli , Antoine Chaffin , Arnaud Stiegler , Debajyoti Datta , Eliza Szczechla , Gunjan Chhablani , Han Wang , Harshit Pandey , Hendrik Strobelt , Jason Alan Fries , Jos Rozen , Leo Gao , Lintang Sutawika , M Saiful Bari , Maged S. Al-shaibani , Matteo Manica , Nihal Nayak , Ryan Teehan , Samuel Albanie , Sheng Shen , Srulik Ben-David , Stephen H. Bach , Taewoon Kim , Tali Bers , Thibault Fevry , Trishala Neeraj , Urmish Thakker , Vikas Raunak , Xiangru Tang , Zheng-Xin Yong , Zhiqing Sun , Shaked Brody , Yallow Uri , Hadar Tojarieh , Adam Roberts , Hyung Won Chung , Jaesung Tae , Jason Phang , Ofir Press , Conglong Li , Deepak Narayanan , Hatim Bourfoune , Jared Casper , Jeff Rasley , Max Ryabinin , Mayank Mishra , Minjia Zhang , Mohammad Shoeybi , Myriam Peyrounette , Nicolas Patry , Nouamane Tazi , Omar Sanseviero , Patrick von Platen , Pierre Cornette , Pierre François Lavallée , Rémi Lacroix , Samyam Rajbhandari , Sanchit Gandhi , Shaden Smith , Stéphane Requena , Suraj Patil , Tim Dettmers , Ahmed Baruwa , Amanpreet Singh , Anastasia Cheveleva , Anne-Laure Ligozat , Arjun Subramonian , Aurélie Névéol , Charles Lovering , Dan Garrette , Deepak Tunuguntla , Ehud Reiter , Ekaterina Taktasheva , Ekaterina Voloshina , Eli Bogdanov , Genta Indra Winata , Hailey Schoelkopf , Jan-Christoph Kalo , Jekaterina Novikova , Jessica Zosa Forde , Jordan Clive , Jungo Kasai , Ken Kawamura , Liam Hazan , Marine Carpuat , Miruna Clinciu , Najoung Kim , Newton Cheng , Oleg Serikov , Omer Antverg , Oskar van der Wal , Rui Zhang , Ruochen Zhang , Sebastian Gehrmann , Shachar Mirkin , Shani Pais , Tatiana Shavrina , Thomas Scialom , Tian Yun , Tomasz Limisiewicz , Verena Rieser , Vitaly Protasov , Vladislav Mikhailov , Yada Pruksachatkun , Yonatan Belinkov , Zachary Bamberger , Zdeněk Kasner , Alice Rueda , Amanda Pestana , Amir Feizpour , Ammar Khan , Amy Faranak , Ana Santos , Anthony Hevia , Antigona Unldreaj , Arash Aghagol , Arezoo Abdollahi , Aycha Tammour , Azadeh HajiHosseini , Bahareh Behroozi , Benjamin Ajibade , Bharat Saxena , Carlos Muñoz Ferrandis , Daniel McDuff , Danish Contractor , David Lansky , Davis David , Douwe Kiela , Duong A. Nguyen , Edward Tan , Emi Baylor , Ezinwanne Ozoani , Fatima Mirza , Frankline Ononiwu , Habib Rezanejad , Hessie Jones , Indrani Bhattacharya , Irene Solaiman , Irina Sedenko , Isar Nejadgholi , Jesse Passmore , Josh Seltzer , Julio Bonis Sanz , Livia Dutra , Mairon Samagaio , Maraim Elbadri , Margot Mieskes , Marissa Gerchick , Martha Akinlolu , Michael McKenna , Mike Qiu , Muhammed Ghauri , Mykola Burynok , Nafis Abrar , Nazneen Rajani , Nour Elkott , Nour Fahmy , Olanrewaju Samuel , Ran An , Rasmus Kromann , Ryan Hao , Samira Alizadeh , Sarmad Shubber , Silas Wang , Sourav Roy , Sylvain Viguier , Thanh Le , Tobi Oyebade , Trieu Le , Yoyo Yang , Zach Nguyen , Abhinav Ramesh Kashyap , Alfredo Palasciano , Alison Callahan , Anima Shukla , Antonio Miranda-Escalada , Ayush Singh , Benjamin Beilharz , Bo Wang , Caio Brito , Chenxi Zhou , Chirag Jain , Chuxin Xu , Clémentine Fourrier , Daniel León Periñán , Daniel Molano , Dian Yu , Enrique Manjavacas , Fabio Barth , Florian Fuhrimann , Gabriel Altay , Giyaseddin Bayrak , Gully Burns , Helena U. Vrabec , Imane Bello , Ishani Dash , Jihyun Kang , John Giorgi , Jonas Golde , Jose David Posada , Karthik Rangasai Sivaraman , Lokesh Bulchandani , Lu Liu , Luisa Shinzato , Madeleine Hahn de Bykhovetz , Maiko Takeuchi , Marc Pàmies , Maria A Castillo , Marianna Nezhurina , Mario Sänger , Matthias Samwald , Michael Cullan , Michael Weinberg , Michiel De Wolf , Mina Mihaljcic , Minna Liu , Moritz Freidank , Myungsun Kang , Natasha Seelam , Nathan Dahlberg , Nicholas Michio Broad , Nikolaus Muellner , Pascale Fung , Patrick Haller , Ramya Chandrasekhar , Renata Eisenberg , Robert Martin , Rodrigo Canalli , Rosaline Su , Ruisi Su , Samuel Cahyawijaya , Samuele Garda , Shlok S Deshmukh , Shubhanshu Mishra , Sid Kiblawi , Simon Ott , Sinee Sang-aroonsiri , Srishti Kumar , Stefan Schweter , Sushil Bharati , Tanmay Laud , Théo Gigant , Tomoya Kainuma , Wojciech Kusa , Yanis Labrak , Yash Shailesh Bajaj , Yash Venkatraman , Yifan Xu , Yingxin Xu , Yu Xu , Zhe Tan , Zhongli Xie , Zifan Ye , Mathilde Bras , Younes Belkada , Thomas Wolf

Vision-Language Models (VLMs) have demonstrated impressive capabilities in code generation across various domains. However, their ability to replicate complex, multi-panel visualizations from real-world data remains largely unassessed. To…

We present LLMStructBench, a novel benchmark for evaluating Large Language Models (LLMs) on extracting structured data and generating valid JavaScript Object Notation (JSON) outputs from natural-language text. Our open dataset comprises…

Computation and Language · Computer Science 2026-02-17 Sönke Tenckhoff , Mario Koddenbrock , Erik Rodner

Natural language-driven no-code development allows users to specify software functionality using natural language (NL) instead of editing source code, promising increased productivity and democratized development. Large language models…

Software Engineering · Computer Science 2025-08-19 Le Deng , Zhonghao Jiang , Jialun Cao , Michael Pradel , Zhongxin Liu

Large Language Models (LLMs) have become a focal point of research across various domains, including software engineering, where their capabilities are increasingly leveraged. Recent studies have explored the integration of LLMs into…

Software Engineering · Computer Science 2024-10-14 Yi Wen Heng , Zeyang Ma , Zhenhao Li , Dong Jae Kim , Tse-Hsun , Chen

This paper embarks on an exploration into the Large Language Model (LLM) datasets, which play a crucial role in the remarkable advancements of LLMs. The datasets serve as the foundational infrastructure analogous to a root system that…

Computation and Language · Computer Science 2024-02-29 Yang Liu , Jiahuan Cao , Chongyu Liu , Kai Ding , Lianwen Jin

Benchmarks establish a standardized evaluation framework to systematically assess the performance of large language models (LLMs), facilitating objective comparisons and driving advancements in the field. However, existing benchmarks fail…

Computation and Language · Computer Science 2026-02-16 Ziqian Zhang , Xingjian Hu , Yue Huang , Kai Zhang , Ruoxi Chen , Yixin Liu , Qingsong Wen , Kaidi Xu , Xiangliang Zhang , Neil Zhenqiang Gong , Lichao Sun

Large Language Models (LLMs) are becoming integral to modern software development workflows, assisting developers with code generation, API explanation, and iterative problem-solving through natural language conversations. Despite…

Software Engineering · Computer Science 2025-09-15 Suzhen Zhong , Ying Zou , Bram Adams

The rapid advancement of large language models (LLMs) demands robust, unbiased, and scalable evaluation methods. However, human annotations are costly to scale, model-based evaluations are susceptible to stylistic biases, and…

Formal specification generation has recently drawn attention in software engineering as a way to improve program correctness without requiring manual annotations. Large Language Models (LLMs) have shown promise in this area, but early…

Software Engineering · Computer Science 2026-04-07 Ragib Shahariar Ayon , Shibbir Ahmed

Large Language Models (LLMs) have achieved remarkable success in natural language tasks, yet understanding their reasoning processes remains a significant challenge. We address this by introducing XplainLLM, a dataset accompanying an…

Computation and Language · Computer Science 2024-09-24 Zichen Chen , Jianda Chen , Ambuj Singh , Misha Sra

Large language models (LLMs) have shown potential in assisting scientific research, yet their ability to discover high-quality research hypotheses remains unexamined due to the lack of a dedicated benchmark. To address this gap, we…

Computation and Language · Computer Science 2026-04-21 Yujie Liu , Zonglin Yang , Tong Xie , Jinjie Ni , Ben Gao , Yuqiang Li , Shixiang Tang , Wanli Ouyang , Erik Cambria , Dongzhan Zhou

Program synthesis has been long studied with recent approaches focused on directly using the power of Large Language Models (LLMs) to generate code. Programming benchmarks, with curated synthesis problems and test-cases, are used to measure…

Software Engineering · Computer Science 2023-11-01 Jiawei Liu , Chunqiu Steven Xia , Yuyao Wang , Lingming Zhang