English
Related papers

Related papers: Managing Uncertainty in LLM-Generated Procedural K…

200 papers

Trust is a crucial factor affecting the adoption of machine learning (ML) models. Qualitative studies have revealed that end-users, particularly in the medical domain, need models that can express their uncertainty in decision-making…

Machine Learning · Computer Science 2023-04-21 Andrew Houston , Georgina Cosma

Large language models (LLMs) have shown strong knowledge reserves and task-solving capabilities, but still face the challenge of severe hallucination, hindering their practical application. Though scientific theories and rules can…

Computation and Language · Computer Science 2026-04-09 Maotian Ma , Zheni Zeng , Zhenghao Liu , Yukun Yan

LLM-based coding agents are increasingly used to generate code, tests, and documentation. Still, their outputs can be plausible yet misaligned with developer intent and provide limited evidence for review in evolving projects. This limits…

Software Engineering · Computer Science 2026-04-14 Ragib Shahariar Ayon

Large Language Models (LLMs) have made significant progress in utilizing tools, but their ability is limited by API availability and the instability of implicit reasoning, particularly when both planning and execution are involved. To…

Computation and Language · Computer Science 2024-06-24 Cheng Qian , Chi Han , Yi R. Fung , Yujia Qin , Zhiyuan Liu , Heng Ji

This study introduces an innovative framework that employs large language models (LLMs) to automate the design and generation of curricula for reinforcement learning (RL). As mobile networks evolve towards the 6G era, managing their…

Machine Learning · Computer Science 2024-10-31 Omar Erak , Omar Alhussein , Shimaa Naser , Nouf Alabbasi , De Mi , Sami Muhaidat

Temporal logics are powerful tools that are widely used for the synthesis and verification of reactive systems. The recent progress on Large Language Models (LLMs) has the potential to make the process of writing such specifications more…

Machine Learning · Computer Science 2024-06-12 William Murphy , Nikolaus Holzer , Nathan Koenig , Leyi Cui , Raven Rothkopf , Feitong Qiao , Mark Santolucito

Automating activities through robots in unstructured environments, such as construction sites, has been a long-standing desire. However, the high degree of unpredictable events in these settings has resulted in far less adoption compared to…

Robotics · Computer Science 2024-07-23 Hossein Naderi , Alireza Shojaei , Lifu Huang

Large Language Models (LLMs) have demonstrated remarkable capabilities across diverse domains, including programming, planning, and decision-making. However, their performance often degrades when faced with highly complex problem instances…

Artificial Intelligence · Computer Science 2025-08-21 Yang Cheng , Zilai Wang , Weiyu Ma , Wenhui Zhu , Yue Deng , Jian Zhao

We introduce a modular prompting framework that supports safer and more adaptive use of large language models (LLMs) across dynamic, user-centered tasks. Grounded in human learning theory, particularly the Zone of Proximal Development…

Artificial Intelligence · Computer Science 2025-08-12 Vanessa Figueiredo

Despite the dramatic progress in Large Language Model (LLM) development, LLMs often provide seemingly plausible but not factual information, often referred to as hallucinations. Retrieval-augmented LLMs provide a non-parametric approach to…

Computation and Language · Computer Science 2023-11-09 Sai Munikoti , Anurag Acharya , Sridevi Wagle , Sameera Horawalavithana

The success of pretrained language models (PLMs) across a spate of use-cases has led to significant investment from the NLP community towards building domain-specific foundational models. On the other hand, in mission critical settings such…

Computation and Language · Computer Science 2024-07-18 Aman Sinha , Timothee Mickus , Marianne Clausel , Mathieu Constant , Xavier Coubez

Large language models (LLMs) are rapidly changing how researchers in materials science and chemistry discover, organize, and act on scientific knowledge. This paper analyzes a broad set of community-developed LLM applications in an effort…

Materials Science · Physics 2026-05-06 Aritra Roy , Kevin Shen , Andrew MacBride , Awwal Oladipupo , Mudassra Taskeen , Wojtek Treyde , Ruaa A. E. A. Abakar , Ahmad D. Abbas , Elsayed Abdelfatah , Abbas A. Abdullahi , Seham S. Abyah , Chahd Rahyl Adjmi , Fariha Agbere , Savyasanchi Aggarwal , Muhammad Ahmed , Tasnim Ahmed , Motasem Ajlouni , Mattias Akke , Hussein AlAdwan , Anwaar S. Alazani , Zahra A. Alharbi , Wajd A. Aljulyhi , Mohammed A. AlKubaish , Fatima A. Almahri , Sayed A. Almohri , David Obeh Alobo , Mohammed Alouni , Azizah S. Alqahtani , Omar Alsaigh , Husain Althagafi , Md. Aqib Aman , Lena Ara , Arifin , Ignacio Arretche , Abdulaziz Ashy , Syeda A. Asim , Amro Aswad , Adeel Atta , Sören Auer , Abdullah al Azmi , Toheeb Balogun , Suvo Banik , Viktoriia Baibakova , Shakira A. Baksh , Neus G. Bastús , Christina J. Bayard , Adib Bazgir , Louis Beal , Lejla Biberić , Wahid Billah , Ankita Biswas , Joshua Bocarsly , Montassar T. Bouzidi , Esma B. Boydas , Youssef Briki , Cailin Buchanan , Mauricio Cafiero , Damien Caliste , Yi Cao , Rafael E. Castañeda , Sruthy K. Chandy , Benjamin Charmes , Shayantan Chaudhuri , Yiming Chen , Alexander Chen , Jieneng Chen , Min-Hsueh Chiu , Defne Circi , Cinthya H. Contreras , Yoann Cure , Nathan Daelman , Roshini Dantuluri , Thomas Davy , William Dawson , Leonid Didukh , Rui Ding , Aminu R. Doguwa , Claudia Draxl , Sathya Edamadaka , Oulaya Elargab , Christina Ertural , Matthew L. Evans , Edvin Fako , Hossam Farag , Nur A. Fathurrahman , Merve Fedai , Rodrigo P. Ferreira , Giuseppe Fisicaro , Thomas Frank , Sasi K. Gaddipati , Abhijeet Gangan , Jennifer Garland , James Garrick , Luigi Genovese , Maryam Ghadrdran , Sandip Giri , Maxime Goulet , Jeremy Goumaz , Sara U. Gracia , Jacob Graham , Gabriel Graves , Kevin P. Greenman , Tim Greitemeier , Cameron Gruich , Sophie Gu , Salomé Guilbert , Hans Gundlach , Muriel F. Gusta , Mourad El Haddaoui , Alexander J. Haibel , Anubhab Haldar , Vehaan Handa , Hassan Harb , Nathan D. Harms , Abdullah Al Hasan , Abir Hassan , Qiyao He , Andrés Henao-Aristizábal , Bram Hoex , Sungil Hong , Alexander J. Horvath , Md. Shaib Hossain , Yanqi Huang , Yuqing Huang , Kostiantyn Hubaiev , Donald Intal , Katherine Inzani , Kevin Ishimwe , Tugba Isik , Gopal R. Iyer , Katharina Jager , Jan Janssen , Hyewon Jeong , Michael Jirasek , Tyler R. Josephson , Nisarg Joshi , Yassir Ben Kacem , Remya A. M. Kalapurakal , Rakesh R. Kamath , Sugan Kanagasenthinathan , Dohun Kang , Jason Kantorow , Kübra Kaygisiz , Murat Keceli , Farhana Keya , Muhammad U. Khan , Sartaaj Takrim Khan , Hyungjun Kim , Alexander Kister , Sascha Klawohn , Collin Kovacs , Pranav Krishnan , Maurycy Kryzanowski , Ritesh Kumar , Suman Kumari , Gourav Kumbhojkar , Ryo Kuroki , Shashank Kushwaha , Magdalena Lederbauer , Jaejun Lee , Seunghan Lee , Jeonghwan Lee , Bingcan Li , Calvin Li , Zhanzhao Li , Shi Li , Shicheng Li , Chengyan Liu , Hao Liu , Tung Yan Liu , Yutong Liu , Lucia Vina-Lopez , Chayaphol Lortaraparsert , Andre K. Y. Low , Saffron Luxford , Carlos Madariaga , Rishikesh Magar , Piyush R. Maharana , Rahul Mallela , Shoaib Mahmud , Natesan Mani , Umair Mansoor , Omar B. Mansour , Cassandra Masschelein , Kinga O. Mastej , Ankit Mathanker , Jeffrey Meng , Omran Mezghani , Yidong Ming , Rishav Mitra , Michail Mitsakis , Matthew Miyagishima , Ravikumar Mohan , Naveen R. Mohanraj , Trupti Mohanty , Bernadette Mohr , Francisco A. Molina-Bakhos , Jeremy Monat , Seyed Mohamad Moosavi , Shayan Mousavi , Arman Moussavi , Rubel Mozumber , Muhammad J. Mufti , Diyana Muhammed , Ram Munde , Mrigi Munjal , José A. Márquez , Shankha Nag , Giacomo Nagaro , Juno Nam , Jose M. Napoles-Duarte , Ry Nduma , Xuan-Vu Nguyen , Ebrahim Norouzi , Oluwatosin Ohiro , Ryotaro Okabe , Viejay Ordillo , Shuichiro Ozawa , Sebastian Pagel , Daniel Palmer , Angela Pan , Akash Pandey , Vivek Pandit , Prakul Pandit , Chiku Parida , Jaehee Park , Hyunsoo Park , Hemangi Patel , Shakul Pathak , Taradutt Pattnaik , Elena Patyukova , Noah Paulson , Deepak S. Pendyala , Erick S. Pepek , Martin H. Petersen , Thang D. Pham , Aniket Phutane , Sabila K. Pinky , Étienne Polack , Alison Polasik , Maria Politi , Tim Pongratz , Akhila Ponugoti , Fabio Priante , Thomas Michael Pruyn , Sai S. Puppala , Mohammad A. Qazi , Heike Quosdorf , Gollam Rabby , Mohammad J. Raei , Md. Habibur Rahman , A. B. M. Ashikur Rahman , Subhashree Rajasekaran , Tawfiqur Rakib , Hemanth N. Ramesh , Vrushali Ranadive , Karnamohit Ranka , Bojana Rankovic , Adwaith Ravichandran , Ilija Rašović , Sergei Rigin , Tatem Rios , Varun Rishi , Victor Naden Robinson , Lucas S. Rodrigues , Oswaldo Rodriguez , Mahule Roy , Diptendu Roy , Subhas Roy , Arokia Anto Royan M , Joseph F. Rudzinski , Muhammad Sabih , Subramanyam Sahoo , Srusti Bheem Sain , Thahira Saliya , Vignesh Sampath , Jesus Diaz Sanchez , Arthur S. S. Santos , Muliady Satria , Hasan M. Sayeed , Jörg Schaarschmidt , Philippe Schwaller , Nofit Segal , Abhishec Senthilvel , Sherjeel Shabih , Devanshu Shah , Faezeh Shahmoradi , Samiha Sharlin , Killian Sheriff , Qiuyu Shi , Abubakar D. Shuaibu , Ayesha Siddiqua , M. A. Shadab Siddiqui , Darian Smalley , Benjamin Smith , Taylor D. Sparks , Daniel T. Speckhard , Elena Stojanovska , Akshay Subramanian , Jiwon Sun , Yunkai Sun , Abdul W. Syed , Souvik Ta , Izumi Takahara , Kelly Tallau , Guannan Tang , Ans B. Tariq , Sui X. Tay , Nurlybek Temirbay , Surya P. Tiwari , Febin Tom , Tajah Trapier , Kasidet J. Trerayapiwat , Samanvya Tripathi , Hawra H. Tuhaifa , Mustafa Unal , Mohammad Uzair , Vallabh Vasudevan , Estefania Vazquez , Victor Venturi , Rahul Verma , Ashwini Verma , Alvaro Vazquez-Mayagoitia , Nicholas Wagner , Araki Wakiuchi , Hao Wan , Liaoyaqi Wang , Wolfgang Wenzel , Alexander Wieczorek , Sze H. Wong , Yue Wu , Tong Xie , Andrew Yi , Ziqi Yin , Jodie A. Yuwono , Nahed A. Zaid , Mohd Zaki , Shehtab Zaman , Maimuna U. Zarewa , Mahtab Zehtab , Baosen Zhang , Wenyu Zhang , Melody Zhang , Yangfan Zhang , Yuwen Zhang , Runze Zhang , Zongmin Zhang , Huanhuan Zhao , Yuanlong Bill Zheng , Ramzi Zidani , Xue Zong , Ian Foster , Ben Blaiszik

Scientific writing is an expert-domain task that demands deep domain knowledge, task-specific requirements and reasoning capabilities that leverage the domain knowledge to satisfy the task specifications. While scientific text generation…

Computation and Language · Computer Science 2026-04-20 Furkan Şahinuç , Subhabrata Dutta , Iryna Gurevych

Accurate uncertainty quantification remains a key challenge for standard LLMs, prompting the adoption of Bayesian and ensemble-based methods. However, such methods typically necessitate computationally expensive sampling, involving multiple…

Machine Learning · Computer Science 2025-07-25 Lakshmana Sri Harsha Nemani , P. K. Srijith , Tomasz Kuśmierczyk

Recent research on Vision-and-Language Navigation (VLN) indicates that agents suffer from poor generalization in unseen environments due to the lack of realistic training environments and high-quality path-instruction pairs. Most existing…

Robotics · Computer Science 2024-11-19 Yu Yan , Rongtao Xu , Jiazhao Zhang , Peiyang Li , Xiaodan Liang , Jianqin Yin

Despite the widespread adoption of large language models (LLMs) for recommendation, we demonstrate that LLMs often exhibit uncertainty in their recommendations. To ensure the trustworthy use of LLMs in generating recommendations, we…

Information Retrieval · Computer Science 2025-02-13 Wonbin Kweon , Sanghwan Jang , SeongKu Kang , Hwanjo Yu

LLM-based agents have emerged as promising tools, which are crafted to fulfill complex tasks by iterative planning and action. However, these agents are susceptible to undesired planning hallucinations when lacking specific knowledge for…

Computation and Language · Computer Science 2024-06-24 Ruixuan Xiao , Wentao Ma , Ke Wang , Yuchuan Wu , Junbo Zhao , Haobo Wang , Fei Huang , Yongbin Li

Large language models (LLMs) are promising tools for supporting security management tasks, such as incident response planning. However, their unreliability and tendency to hallucinate remain significant challenges. In this paper, we address…

Artificial Intelligence · Computer Science 2026-02-06 Kim Hammar , Tansu Alpcan , Emil Lupu

The rise of large language models (LLMs) has highlighted the importance of prompt engineering as a crucial technique for optimizing model outputs. While experimentation with various prompting methods, such as Few-shot, Chain-of-Thought, and…

Artificial Intelligence · Computer Science 2026-03-27 Michael Hewing , Vincent Leinhos

Large Language Models (LLMs) are employed across various high-stakes domains, where the reliability of their outputs is crucial. One commonly used method to assess the reliability of LLMs' responses is uncertainty estimation, which gauges…

‹ Prev 1 4 5 6 7 8 10 Next ›