English
Related papers

Related papers: RICA: Evaluating Robust Inference Capabilities Bas…

200 papers

To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we present Global PIQA, a participatory commonsense reasoning…

Computation and Language · Computer Science 2025-10-29 Tyler A. Chang , Catherine Arnett , Abdelrahman Eldesokey , Abdelrahman Sadallah , Abeer Kashar , Abolade Daud , Abosede Grace Olanihun , Adamu Labaran Mohammed , Adeyemi Praise , Adhikarinayum Meerajita Sharma , Aditi Gupta , Afitab Iyigun , Afonso Simplício , Ahmed Essouaied , Aicha Chorana , Akhil Eppa , Akintunde Oladipo , Akshay Ramesh , Aleksei Dorkin , Alfred Malengo Kondoro , Alham Fikri Aji , Ali Eren Çetintaş , Allan Hanbury , Alou Dembele , Alp Niksarli , Álvaro Arroyo , Amin Bajand , Amol Khanna , Ana Chkhaidze , Ana Condez , Andiswa Mkhonto , Andrew Hoblitzell , Andrew Tran , Angelos Poulis , Anirban Majumder , Anna Vacalopoulou , Annette Kuuipolani Kanahele Wong , Annika Simonsen , Anton Kovalev , Ashvanth. S , Ayodeji Joseph Lana , Barkin Kinay , Bashar Alhafni , Benedict Cibalinda Busole , Bernard Ghanem , Bharti Nathani , Biljana Stojanovska Đurić , Bola Agbonile , Bragi Bergsson , Bruce Torres Fischer , Burak Tutar , Burcu Alakuş Çınar , Cade J. Kanoniakapueo Kane , Can Udomcharoenchaikit , Catherine Arnett , Chadi Helwe , Chaithra Reddy Nerella , Chen Cecilia Liu , Chiamaka Glory Nwokolo , Cristina España-Bonet , Cynthia Amol , DaeYeop Lee , Dana Arad , Daniil Dzenhaliou , Daria Pugacheva , Dasol Choi , Daud Abolade , David Liu , David Semedo , Deborah Popoola , Deividas Mataciunas , Delphine Nyaboke , Dhyuthy Krishna Kumar , Diogo Glória-Silva , Diogo Tavares , Divyanshu Goyal , DongGeon Lee , Ebele Nwamaka Anajemba , Egonu Ngozi Grace , Elena Mickel , Elena Tutubalina , Elias Herranen , Emile Anand , Emmanuel Habumuremyi , Emuobonuvie Maria Ajiboye , Eryawan Presma Yulianrifat , Esther Adenuga , Ewa Rudnicka , Faith Olabisi Itiola , Faran Taimoor Butt , Fathima Thekkekara , Fatima Haouari , Filbert Aurelian Tjiaranata , Firas Laakom , Francesca Grasso , Francesco Orabona , Francesco Periti , Gbenga Kayode Solomon , Gia Nghia Ngo , Gloria Udhehdhe-oze , Gonçalo Martins , Gopi Naga Sai Ram Challagolla , Guijin Son , Gulnaz Abdykadyrova , Hafsteinn Einarsson , Hai Hu , Hamidreza Saffari , Hamza Zaidi , Haopeng Zhang , Harethah Abu Shairah , Harry Vuong , Hele-Andra Kuulmets , Houda Bouamor , Hwanjo Yu , Iben Nyholm Debess , İbrahim Ethem Deveci , Ikhlasul Akmal Hanif , Ikhyun Cho , Inês Calvo , Inês Vieira , Isaac Manzi , Ismail Daud , Itay Itzhak , Iuliia , Alekseenko , Ivan Belashkin , Ivan Spada , Ivan Zhelyazkov , Jacob Brinton , Jafar Isbarov , Jaka Čibej , Jan Čuhel , Jan Kocoń , Jauza Akbar Krito , Jebish Purbey , Jennifer Mickel , Jennifer Za , Jenny Kunz , Jihae Jeong , Jimena Tena Dávalos , Jinu Lee , João Magalhães , John Yi , Jongin Kim , Joseph Chataignon , Joseph Marvin Imperial , Jubeerathan Thevakumar , Judith Land , Junchen Jiang , Jungwhan Kim , Kairit Sirts , Kamesh R , Kamesh V , Kanda Patrick Tshinu , Kätriin Kukk , Kaustubh Ponkshe , Kavsar Huseynova , Ke He , Kelly Buchanan , Kengatharaiyer Sarveswaran , Kerem Zaman , Khalil Mrini , Kian Kyars , Krister Kruusmaa , Kusum Chouhan , Lainitha Krishnakumar , Laura Castro Sánchez , Laura Porrino Moscoso , Leshem Choshen , Levent Sencan , Lilja Øvrelid , Lisa Alazraki , Lovina Ehimen-Ugbede , Luheerathan Thevakumar , Luxshan Thavarasa , Mahnoor Malik , Mamadou K. Keita , Mansi Jangid , Marco De Santis , Marcos García , Marek Suppa , Mariam D'Ciofalo , Marii Ojastu , Maryam Sikander , Mausami Narayan , Maximos Skandalis , Mehak Mehak , Mehmet İlteriş Bozkurt , Melaku Bayu Workie , Menan Velayuthan , Michael Leventhal , Michał Marcińczuk , Mirna Potočnjak , Mohammadamin Shafiei , Mridul Sharma , Mrityunjaya Indoria , Muhammad Ravi Shulthan Habibi , Murat Kolić , Nada Galant , Naphat Permpredanun , Narada Maugin , Nicholas Kluge Corrêa , Nikola Ljubešić , Nirmal Thomas , Nisansa de Silva , Nisheeth Joshi , Nitish Ponkshe , Nizar Habash , Nneoma C. Udeze , Noel Thomas , Noémi Ligeti-Nagy , Nouhoum Coulibaly , Nsengiyumva Faustin , Odunayo Kareemat Buliaminu , Odunayo Ogundepo , Oghojafor Godswill Fejiro , Ogundipe Blessing Funmilola , Okechukwu God'spraise , Olanrewaju Samuel , Olaoye Deborah Oluwaseun , Olasoji Akindejoye , Olga Popova , Olga Snissarenko , Onyinye Anulika Chiemezie , Orkun Kinay , Osman Tursun , Owoeye Tobiloba Moses , Oyelade Oluwafemi Joshua , Oyesanmi Fiyinfoluwa , Pablo Gamallo , Pablo Rodríguez Fernández , Palak Arora , Pedro Valente , Peter Rupnik , Philip Oghenesuowho Ekiugbo , Pramit Sahoo , Prokopis Prokopidis , Pua Niau-Puhipau , Quadri Yahya , Rachele Mignone , Raghav Singhal , Ram Mohan Rao Kadiyala , Raphael Merx , Rapheal Afolayan , Ratnavel Rajalakshmi , Rishav Ghosh , Romina Oji , Ron Kekeha Solis , Rui Guerra , Rushikesh Zawar , Sa'ad Nasir Bashir , Saeed Alzaabi , Sahil Sandeep , Sai Pavan Batchu , SaiSandeep Kantareddy , Salsabila Zahirah Pranida , Sam Buchanan , Samuel Rutunda , Sander Land , Sarah Sulollari , Sardar Ali , Saroj Sapkota , Saulius Tautvaisas , Sayambhu Sen , Sayantani Banerjee , Sebastien Diarra , SenthilNathan. M , Sewoong Lee , Shaan Shah , Shankar Venkitachalam , Sharifa Djurabaeva , Sharon Ibejih , Shivanya Shomir Dutta , Siddhant Gupta , Silvia Paniagua Suárez , Sina Ahmadi , Sivasuthan Sukumar , Siyuan Song , Snegha A. , Sokratis Sofianopoulos , Sona Elza Simon , Sonja Benčina , Sophie Gvasalia , Sphurti Kirit More , Spyros Dragazis , Stephan P. Kaufhold , Suba. S , Sultan AlRashed , Surangika Ranathunga , Taiga Someya , Taja Kuzman Pungeršek , Tal Haklay , Tasi'u Jibril , Tatsuya Aoyama , Tea Abashidze , Terenz Jomar Dela Cruz , Terra Blevins , Themistoklis Nikas , Theresa Dora Idoko , Thu Mai Do , Tilek Chubakov , Tommaso Gargiani , Uma Rathore , Uni Johannesen , Uwuma Doris Ugwu , Vallerie Alexandra Putra , Vanya Bannihatti Kumar , Varsha Jeyarajalingam , Varvara Arzt , Vasudevan Nedumpozhimana , Viktoria Ondrejova , Viktoryia Horbik , Vishnu Vardhan Reddy Kummitha , Vuk Dinić , Walelign Tewabe Sewunetie , Winston Wu , Xiaojing Zhao , Yacouba Diarra , Yaniv Nikankin , Yash Mathur , Yixi Chen , Yiyuan Li , Yolanda Xavier , Yonatan Belinkov , Yusuf Ismail Abayomi , Zaid Alyafeai , Zhengyang Shan , Zhi Rui Tam , Zilu Tang , Zuzana Nadova , Baber Abbasi , Stella Biderman , David Stap , Duygu Ataman , Fabian Schmidt , Hila Gonen , Jiayi Wang , David Ifeoluwa Adelani

People judge interactions with large language models (LLMs) as successful when outputs match what they want, not what they type. Yet LLMs are trained to predict the next token solely from text input, not underlying intent. Because written…

Computation and Language · Computer Science 2026-03-13 Nadav Kunievsky , James A. Evans

More than one hundred benchmarks have been developed to test the commonsense knowledge and commonsense reasoning abilities of artificial intelligence (AI) systems. However, these benchmarks are often flawed and many aspects of common sense…

Artificial Intelligence · Computer Science 2023-02-24 Ernest Davis

Recent advancements in Chain-of-Thought (CoT) reasoning utilize complex modules but are hampered by high token consumption, limited applicability, and challenges in reproducibility. This paper conducts a critical evaluation of CoT…

Computation and Language · Computer Science 2024-06-12 Mengru Ding , Hanmeng Liu , Zhizhang Fu , Jian Song , Wenbo Xie , Yue Zhang

In the evolving landscape of conversational AI, generating concise, context-aware, and human-like dialogue using small and medium-sized language models (LLMs) remains a complex challenge. This study investigates the influence of LoRA rank,…

Computation and Language · Computer Science 2025-04-15 Chitranshu Harbola , Anupam Purwar

While Large Language Models (LLMs) have shown impressive capabilities in math problem-solving tasks, their robustness to noisy inputs is not well-studied. We propose ArithmAttack to examine how robust the LLMs are when they encounter noisy…

Computation and Language · Computer Science 2026-03-17 Zain Ul Abedin , Shahzeb Qamar , Lucie Flek , Akbar Karimi

Reasoning and inference are central to human and artificial intelligence. Modeling inference in human language is very challenging. With the availability of large annotated data (Bowman et al., 2015), it has recently become feasible to…

Computation and Language · Computer Science 2020-03-04 Qian Chen , Xiaodan Zhu , Zhenhua Ling , Si Wei , Hui Jiang , Diana Inkpen

Large Language Models (LLMs) are increasingly being deployed as the reasoning engines for agentic AI systems, yet they exhibit a critical flaw: a rigid adherence to explicit rules that leads to decisions misaligned with human common sense…

Artificial Intelligence · Computer Science 2025-10-16 Imran Khan

Recently, the community has achieved substantial progress on many commonsense reasoning benchmarks. However, it is still unclear what is learned from the training process: the knowledge, inference capability, or both? We argue that due to…

Computation and Language · Computer Science 2022-10-13 Hongming Zhang , Yintong Huo , Yanai Elazar , Yangqiu Song , Yoav Goldberg , Dan Roth

Humans spontaneously use increasingly efficient language as interactions progress, by adapting and forming ad-hoc conventions. This phenomenon has been studied extensively using reference games, showing properties of human language that go…

Computation and Language · Computer Science 2024-08-05 Yilun Hua , Yoav Artzi

Large Reasoning Models (LRMs) significantly improve the reasoning ability of Large Language Models (LLMs) by learning to reason, exhibiting promising performance in solving complex tasks. However, their deliberative reasoning process leads…

Computation and Language · Computer Science 2025-08-14 Yue Liu , Jiaying Wu , Yufei He , Ruihan Gong , Jun Xia , Liang Li , Hongcheng Gao , Hongyu Chen , Baolong Bi , Jiaheng Zhang , Zhiqi Huang , Bryan Hooi , Stan Z. Li , Keqin Li

Language models, characterized by their black-box nature, often hallucinate and display sensitivity to input perturbations, causing concerns about trust. To enhance trust, it is imperative to gain a comprehensive understanding of the…

Computation and Language · Computer Science 2025-01-03 Vatsal Gupta , Pranshu Pandya , Tushar Kataria , Vivek Gupta , Dan Roth

As language models accelerate scientific research by automating hypothesis generation and implementation, a new bottleneck emerges: evaluating and filtering hundreds of AI-generated ideas without exhaustive experimentation. We ask whether…

Machine Learning · Computer Science 2026-05-22 Srujan P Mule , Aniketh Garikaparthi , Manasi Patwardhan

Large Language Models (LLMs) are recruited in applications that span from clinical assistance and legal support to question answering and education. Their success in specialized tasks has led to the claim that they possess human-like…

Computation and Language · Computer Science 2024-07-10 Vittoria Dentella , Fritz Guenther , Elliot Murphy , Gary Marcus , Evelina Leivada

Retrieval-Augmented Language Models (RALMs) face significant challenges in reducing factual errors, particularly in document relevance evaluation and knowledge integration. We introduce a framework for structured relevance assessment that…

Artificial Intelligence · Computer Science 2025-07-30 Aryan Raj , Astitva Veer Garg , Anitha D

As large language models (LLMs) become integrated into everyday and high-stakes decision-making, they inherit the ambiguity and biases of human language. While they produce fluent and coherent outputs, they rely on statistical pattern…

Artificial Intelligence · Computer Science 2026-04-17 Rikard Rosenbacke , Carl Rosenbacke , Victor Rosenbacke , Martin McKee

Despite their advanced reasoning capabilities, state-of-the-art Multimodal Large Language Models (MLLMs) demonstrably lack a core component of human intelligence: the ability to `read the room' and assess deception in complex social…

Computer Vision and Pattern Recognition · Computer Science 2025-11-21 Caixin Kang , Yifei Huang , Liangyang Ouyang , Mingfang Zhang , Ruicong Liu , Yoichi Sato

The Retrieval-Augmented Language Model (RALM) has shown remarkable performance on knowledge-intensive tasks by incorporating external knowledge during inference, which mitigates the factual hallucinations inherited in large language models…

Computation and Language · Computer Science 2024-12-20 Yuan Xia , Jingbo Zhou , Zhenhui Shi , Jun Chen , Haifeng Huang

The rise of large language models (LLMs) has opened new opportunities in Recommender Systems (RSs) by enhancing user behavior modeling and content understanding. However, current approaches that integrate LLMs into RSs solely utilize either…

Information Retrieval · Computer Science 2024-03-26 Yunjia Xi , Weiwen Liu , Jianghao Lin , Chuhan Wu , Bo Chen , Ruiming Tang , Weinan Zhang , Yong Yu

The emergence of Large Language Models (LLMs) has driven rapid progress in multi-modal learning, particularly in the development of Large Vision-Language Models (LVLMs). However, existing LVLM training paradigms place excessive reliance on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Kaihua Tang , Jiaxin Qi , Jinli Ou , Yuhua Zheng , Jianqiang Huang
‹ Prev 1 4 5 6 7 8 10 Next ›