English
Related papers

Related papers: Designing an Evaluation Framework for Large Langua…

200 papers

To be included into chatbot systems, Large language models (LLMs) must be aligned with human conversational conventions. However, being trained mainly on web-scraped data gives existing LLMs a voice closer to informational text than actual…

Computation and Language · Computer Science 2024-07-30 Shaz Furniturewala , Kokil Jaidka , Yashvardhan Sharma

Large language models (LLMs) such as ChatGPT have received immense interest for their general-purpose language understanding and, in particular, their ability to generate high-quality text or computer code. For many professions, LLMs…

Computation and Language · Computer Science 2024-04-03 Simon Frieder , Julius Berner , Philipp Petersen , Thomas Lukasiewicz

Assessing the capabilities and limitations of large language models (LLMs) has garnered significant interest, yet the evaluation of multiple models in real-world scenarios remains rare. Multilingual evaluation often relies on translated…

Computation and Language · Computer Science 2025-12-02 Varun Gumma , Ananditha Raghunath , Mohit Jain , Sunayana Sitaram

The rapid advancement of Artificial Intelligence has resulted in the advent of Large Language Models (LLMs) with the capacity to produce text that closely resembles human communication. These models have been seamlessly integrated into…

Machine Learning · Computer Science 2025-02-04 Zenon Lamprou , Yashar Moshfeghi

Recommender systems have traditionally followed modular architectures comprising candidate generation, multi-stage ranking, and re-ranking, each trained separately with supervised objectives and hand-engineered features. While effective in…

Information Retrieval · Computer Science 2025-10-06 Rahul Raja , Anshaj Vats , Arpita Vats , Anirban Majumder

Modern large language models (LLMs) represent a paradigm shift in what can plausibly be expected of machine learning models. The fact that LLMs can effectively generate sensible answers to a diverse range of queries suggests that they would…

Computation and Language · Computer Science 2024-05-27 Dean Wyatte , Fatemeh Tahmasbi , Ming Li , Thomas Markovich

Evaluation of large language model (LLM) outputs requires users to make critical judgments about the best outputs across various configurations. This process is costly and takes time given the large amounts of data. LLMs are increasingly…

Conducting literature reviews for scientific papers is essential for understanding research, its limitations, and building on existing work. It is a tedious task which makes an automatic literature review generator appealing. Unfortunately,…

Lab results are often confusing and hard to understand. Large language models (LLMs) such as ChatGPT have opened a promising avenue for patients to get their questions answered. We aim to assess the feasibility of using LLMs to generate…

Computation and Language · Computer Science 2024-04-23 Zhe He , Balu Bhasuran , Qiao Jin , Shubo Tian , Karim Hanna , Cindy Shavor , Lisbeth Garcia Arguello , Patrick Murray , Zhiyong Lu

Large language models (LLMs) have shown potential in assisting scientific research, yet their ability to discover high-quality research hypotheses remains unexamined due to the lack of a dedicated benchmark. To address this gap, we…

Computation and Language · Computer Science 2026-04-21 Yujie Liu , Zonglin Yang , Tong Xie , Jinjie Ni , Ben Gao , Yuqiang Li , Shixiang Tang , Wanli Ouyang , Erik Cambria , Dongzhan Zhou

In the rapidly evolving field of artificial intelligence, large language models (LLMs) have demonstrated significant capabilities across numerous applications. However, the performance of these models in languages with fewer resources, such…

Computation and Language · Computer Science 2024-05-24 Birger Moell

Large language models (LLMs) are rapidly changing how researchers in materials science and chemistry discover, organize, and act on scientific knowledge. This paper analyzes a broad set of community-developed LLM applications in an effort…

Materials Science · Physics 2026-05-06 Aritra Roy , Kevin Shen , Andrew MacBride , Awwal Oladipupo , Mudassra Taskeen , Wojtek Treyde , Ruaa A. E. A. Abakar , Ahmad D. Abbas , Elsayed Abdelfatah , Abbas A. Abdullahi , Seham S. Abyah , Chahd Rahyl Adjmi , Fariha Agbere , Savyasanchi Aggarwal , Muhammad Ahmed , Tasnim Ahmed , Motasem Ajlouni , Mattias Akke , Hussein AlAdwan , Anwaar S. Alazani , Zahra A. Alharbi , Wajd A. Aljulyhi , Mohammed A. AlKubaish , Fatima A. Almahri , Sayed A. Almohri , David Obeh Alobo , Mohammed Alouni , Azizah S. Alqahtani , Omar Alsaigh , Husain Althagafi , Md. Aqib Aman , Lena Ara , Arifin , Ignacio Arretche , Abdulaziz Ashy , Syeda A. Asim , Amro Aswad , Adeel Atta , Sören Auer , Abdullah al Azmi , Toheeb Balogun , Suvo Banik , Viktoriia Baibakova , Shakira A. Baksh , Neus G. Bastús , Christina J. Bayard , Adib Bazgir , Louis Beal , Lejla Biberić , Wahid Billah , Ankita Biswas , Joshua Bocarsly , Montassar T. Bouzidi , Esma B. Boydas , Youssef Briki , Cailin Buchanan , Mauricio Cafiero , Damien Caliste , Yi Cao , Rafael E. Castañeda , Sruthy K. Chandy , Benjamin Charmes , Shayantan Chaudhuri , Yiming Chen , Alexander Chen , Jieneng Chen , Min-Hsueh Chiu , Defne Circi , Cinthya H. Contreras , Yoann Cure , Nathan Daelman , Roshini Dantuluri , Thomas Davy , William Dawson , Leonid Didukh , Rui Ding , Aminu R. Doguwa , Claudia Draxl , Sathya Edamadaka , Oulaya Elargab , Christina Ertural , Matthew L. Evans , Edvin Fako , Hossam Farag , Nur A. Fathurrahman , Merve Fedai , Rodrigo P. Ferreira , Giuseppe Fisicaro , Thomas Frank , Sasi K. Gaddipati , Abhijeet Gangan , Jennifer Garland , James Garrick , Luigi Genovese , Maryam Ghadrdran , Sandip Giri , Maxime Goulet , Jeremy Goumaz , Sara U. Gracia , Jacob Graham , Gabriel Graves , Kevin P. Greenman , Tim Greitemeier , Cameron Gruich , Sophie Gu , Salomé Guilbert , Hans Gundlach , Muriel F. Gusta , Mourad El Haddaoui , Alexander J. Haibel , Anubhab Haldar , Vehaan Handa , Hassan Harb , Nathan D. Harms , Abdullah Al Hasan , Abir Hassan , Qiyao He , Andrés Henao-Aristizábal , Bram Hoex , Sungil Hong , Alexander J. Horvath , Md. Shaib Hossain , Yanqi Huang , Yuqing Huang , Kostiantyn Hubaiev , Donald Intal , Katherine Inzani , Kevin Ishimwe , Tugba Isik , Gopal R. Iyer , Katharina Jager , Jan Janssen , Hyewon Jeong , Michael Jirasek , Tyler R. Josephson , Nisarg Joshi , Yassir Ben Kacem , Remya A. M. Kalapurakal , Rakesh R. Kamath , Sugan Kanagasenthinathan , Dohun Kang , Jason Kantorow , Kübra Kaygisiz , Murat Keceli , Farhana Keya , Muhammad U. Khan , Sartaaj Takrim Khan , Hyungjun Kim , Alexander Kister , Sascha Klawohn , Collin Kovacs , Pranav Krishnan , Maurycy Kryzanowski , Ritesh Kumar , Suman Kumari , Gourav Kumbhojkar , Ryo Kuroki , Shashank Kushwaha , Magdalena Lederbauer , Jaejun Lee , Seunghan Lee , Jeonghwan Lee , Bingcan Li , Calvin Li , Zhanzhao Li , Shi Li , Shicheng Li , Chengyan Liu , Hao Liu , Tung Yan Liu , Yutong Liu , Lucia Vina-Lopez , Chayaphol Lortaraparsert , Andre K. Y. Low , Saffron Luxford , Carlos Madariaga , Rishikesh Magar , Piyush R. Maharana , Rahul Mallela , Shoaib Mahmud , Natesan Mani , Umair Mansoor , Omar B. Mansour , Cassandra Masschelein , Kinga O. Mastej , Ankit Mathanker , Jeffrey Meng , Omran Mezghani , Yidong Ming , Rishav Mitra , Michail Mitsakis , Matthew Miyagishima , Ravikumar Mohan , Naveen R. Mohanraj , Trupti Mohanty , Bernadette Mohr , Francisco A. Molina-Bakhos , Jeremy Monat , Seyed Mohamad Moosavi , Shayan Mousavi , Arman Moussavi , Rubel Mozumber , Muhammad J. Mufti , Diyana Muhammed , Ram Munde , Mrigi Munjal , José A. Márquez , Shankha Nag , Giacomo Nagaro , Juno Nam , Jose M. Napoles-Duarte , Ry Nduma , Xuan-Vu Nguyen , Ebrahim Norouzi , Oluwatosin Ohiro , Ryotaro Okabe , Viejay Ordillo , Shuichiro Ozawa , Sebastian Pagel , Daniel Palmer , Angela Pan , Akash Pandey , Vivek Pandit , Prakul Pandit , Chiku Parida , Jaehee Park , Hyunsoo Park , Hemangi Patel , Shakul Pathak , Taradutt Pattnaik , Elena Patyukova , Noah Paulson , Deepak S. Pendyala , Erick S. Pepek , Martin H. Petersen , Thang D. Pham , Aniket Phutane , Sabila K. Pinky , Étienne Polack , Alison Polasik , Maria Politi , Tim Pongratz , Akhila Ponugoti , Fabio Priante , Thomas Michael Pruyn , Sai S. Puppala , Mohammad A. Qazi , Heike Quosdorf , Gollam Rabby , Mohammad J. Raei , Md. Habibur Rahman , A. B. M. Ashikur Rahman , Subhashree Rajasekaran , Tawfiqur Rakib , Hemanth N. Ramesh , Vrushali Ranadive , Karnamohit Ranka , Bojana Rankovic , Adwaith Ravichandran , Ilija Rašović , Sergei Rigin , Tatem Rios , Varun Rishi , Victor Naden Robinson , Lucas S. Rodrigues , Oswaldo Rodriguez , Mahule Roy , Diptendu Roy , Subhas Roy , Arokia Anto Royan M , Joseph F. Rudzinski , Muhammad Sabih , Subramanyam Sahoo , Srusti Bheem Sain , Thahira Saliya , Vignesh Sampath , Jesus Diaz Sanchez , Arthur S. S. Santos , Muliady Satria , Hasan M. Sayeed , Jörg Schaarschmidt , Philippe Schwaller , Nofit Segal , Abhishec Senthilvel , Sherjeel Shabih , Devanshu Shah , Faezeh Shahmoradi , Samiha Sharlin , Killian Sheriff , Qiuyu Shi , Abubakar D. Shuaibu , Ayesha Siddiqua , M. A. Shadab Siddiqui , Darian Smalley , Benjamin Smith , Taylor D. Sparks , Daniel T. Speckhard , Elena Stojanovska , Akshay Subramanian , Jiwon Sun , Yunkai Sun , Abdul W. Syed , Souvik Ta , Izumi Takahara , Kelly Tallau , Guannan Tang , Ans B. Tariq , Sui X. Tay , Nurlybek Temirbay , Surya P. Tiwari , Febin Tom , Tajah Trapier , Kasidet J. Trerayapiwat , Samanvya Tripathi , Hawra H. Tuhaifa , Mustafa Unal , Mohammad Uzair , Vallabh Vasudevan , Estefania Vazquez , Victor Venturi , Rahul Verma , Ashwini Verma , Alvaro Vazquez-Mayagoitia , Nicholas Wagner , Araki Wakiuchi , Hao Wan , Liaoyaqi Wang , Wolfgang Wenzel , Alexander Wieczorek , Sze H. Wong , Yue Wu , Tong Xie , Andrew Yi , Ziqi Yin , Jodie A. Yuwono , Nahed A. Zaid , Mohd Zaki , Shehtab Zaman , Maimuna U. Zarewa , Mahtab Zehtab , Baosen Zhang , Wenyu Zhang , Melody Zhang , Yangfan Zhang , Yuwen Zhang , Runze Zhang , Zongmin Zhang , Huanhuan Zhao , Yuanlong Bill Zheng , Ramzi Zidani , Xue Zong , Ian Foster , Ben Blaiszik

Large language models (LLMs) excel at generating empathic responses in text-based conversations. But, how reliably do they judge the nuances of empathic communication? We investigate this question by comparing how experts, crowdworkers, and…

Computation and Language · Computer Science 2025-10-06 Aakriti Kumar , Nalin Poungpeth , Diyi Yang , Erina Farrell , Bruce Lambert , Matthew Groh

Simulation powered by Large Language Models (LLMs) has become a promising method for exploring complex human social behaviors. However, the application of LLMs in simulations presents significant challenges, particularly regarding their…

Computers and Society · Computer Science 2025-02-26 Qian Wang , Zhenheng Tang , Bingsheng He

The development of large language models (LLMs) such as ChatGPT has brought a lot of attention recently. However, their evaluation in the benchmark academic datasets remains under-explored due to the difficulty of evaluating the generative…

Computation and Language · Computer Science 2023-07-07 Md Tahmid Rahman Laskar , M Saiful Bari , Mizanur Rahman , Md Amran Hossen Bhuiyan , Shafiq Joty , Jimmy Xiangji Huang

The rapid evolution of large language models (LLMs) has transformed conversational agents, enabling complex human-machine interactions. However, evaluation frameworks often focus on single tasks, failing to capture the dynamic nature of…

Computation and Language · Computer Science 2025-02-10 Pietro Alessandro Aluffi , Patrick Zietkiewicz , Marya Bazzi , Matt Arderne , Vladimirs Murevics

Large language models (LLMs) have moved far beyond their initial form as simple chatbots, now carrying out complex reasoning, planning, writing, coding, and research tasks. These skills overlap significantly with those that human scientists…

Cutting-edge Artificial Intelligence (AI) techniques keep reshaping our view of the world. For example, Large Language Models (LLMs) based applications such as ChatGPT have shown the capability of generating human-like conversation on…

The emergence of Large Language Models (LLMs) with increasingly sophisticated natural language understanding and generative capabilities has sparked interest in the Agent-based Modelling (ABM) community. With their ability to summarize,…

As Large Language Models (LLMs) transition from static tools to autonomous agents, traditional evaluation benchmarks that measure performance on downstream tasks are becoming insufficient. These methods fail to capture the emergent social…

Artificial Intelligence · Computer Science 2025-10-03 Zarreen Reza
‹ Prev 1 8 9 10 Next ›