English
Related papers

Related papers: LoraxBench: A Multitask, Multilingual Benchmark Su…

200 papers

With nearly 1.5 billion people and more than 120 major languages, India represents one of the most diverse regions in the world. As multilingual Vision-Language Models (VLMs) gain prominence, robust evaluation methodologies are essential to…

Evaluating the multilingual and multicultural capabilities of Large Language Models (LLMs) is essential for their global utility. However, current benchmarks face three critical limitations: (1) fragmented evaluation dimensions that often…

Despite the remarkable advancements and widespread applications of deep neural networks, their ability to perform reasoning tasks remains limited, particularly in domains requiring structured, abstract thought. In this paper, we investigate…

Computation and Language · Computer Science 2025-09-16 Satyam Goyal , Soham Dan

Large Language Models (LLMs) achieve remarkable performance across various tasks, but their tendency to produce hallucinations limits reliable adoption. Benchmarks such as TruthfulQA have been developed to measure truthfulness, yet they are…

Computation and Language · Computer Science 2025-09-09 Lorenzo Alfred Nery , Ronald Dawson Catignas , Thomas James Tiam-Lee

To date, there exist almost no culturally-specific evaluation benchmarks for large language models (LLMs) that cover a large number of languages and cultures. In this paper, we present Global PIQA, a participatory commonsense reasoning…

Computation and Language · Computer Science 2025-10-29 Tyler A. Chang , Catherine Arnett , Abdelrahman Eldesokey , Abdelrahman Sadallah , Abeer Kashar , Abolade Daud , Abosede Grace Olanihun , Adamu Labaran Mohammed , Adeyemi Praise , Adhikarinayum Meerajita Sharma , Aditi Gupta , Afitab Iyigun , Afonso Simplício , Ahmed Essouaied , Aicha Chorana , Akhil Eppa , Akintunde Oladipo , Akshay Ramesh , Aleksei Dorkin , Alfred Malengo Kondoro , Alham Fikri Aji , Ali Eren Çetintaş , Allan Hanbury , Alou Dembele , Alp Niksarli , Álvaro Arroyo , Amin Bajand , Amol Khanna , Ana Chkhaidze , Ana Condez , Andiswa Mkhonto , Andrew Hoblitzell , Andrew Tran , Angelos Poulis , Anirban Majumder , Anna Vacalopoulou , Annette Kuuipolani Kanahele Wong , Annika Simonsen , Anton Kovalev , Ashvanth. S , Ayodeji Joseph Lana , Barkin Kinay , Bashar Alhafni , Benedict Cibalinda Busole , Bernard Ghanem , Bharti Nathani , Biljana Stojanovska Đurić , Bola Agbonile , Bragi Bergsson , Bruce Torres Fischer , Burak Tutar , Burcu Alakuş Çınar , Cade J. Kanoniakapueo Kane , Can Udomcharoenchaikit , Catherine Arnett , Chadi Helwe , Chaithra Reddy Nerella , Chen Cecilia Liu , Chiamaka Glory Nwokolo , Cristina España-Bonet , Cynthia Amol , DaeYeop Lee , Dana Arad , Daniil Dzenhaliou , Daria Pugacheva , Dasol Choi , Daud Abolade , David Liu , David Semedo , Deborah Popoola , Deividas Mataciunas , Delphine Nyaboke , Dhyuthy Krishna Kumar , Diogo Glória-Silva , Diogo Tavares , Divyanshu Goyal , DongGeon Lee , Ebele Nwamaka Anajemba , Egonu Ngozi Grace , Elena Mickel , Elena Tutubalina , Elias Herranen , Emile Anand , Emmanuel Habumuremyi , Emuobonuvie Maria Ajiboye , Eryawan Presma Yulianrifat , Esther Adenuga , Ewa Rudnicka , Faith Olabisi Itiola , Faran Taimoor Butt , Fathima Thekkekara , Fatima Haouari , Filbert Aurelian Tjiaranata , Firas Laakom , Francesca Grasso , Francesco Orabona , Francesco Periti , Gbenga Kayode Solomon , Gia Nghia Ngo , Gloria Udhehdhe-oze , Gonçalo Martins , Gopi Naga Sai Ram Challagolla , Guijin Son , Gulnaz Abdykadyrova , Hafsteinn Einarsson , Hai Hu , Hamidreza Saffari , Hamza Zaidi , Haopeng Zhang , Harethah Abu Shairah , Harry Vuong , Hele-Andra Kuulmets , Houda Bouamor , Hwanjo Yu , Iben Nyholm Debess , İbrahim Ethem Deveci , Ikhlasul Akmal Hanif , Ikhyun Cho , Inês Calvo , Inês Vieira , Isaac Manzi , Ismail Daud , Itay Itzhak , Iuliia , Alekseenko , Ivan Belashkin , Ivan Spada , Ivan Zhelyazkov , Jacob Brinton , Jafar Isbarov , Jaka Čibej , Jan Čuhel , Jan Kocoń , Jauza Akbar Krito , Jebish Purbey , Jennifer Mickel , Jennifer Za , Jenny Kunz , Jihae Jeong , Jimena Tena Dávalos , Jinu Lee , João Magalhães , John Yi , Jongin Kim , Joseph Chataignon , Joseph Marvin Imperial , Jubeerathan Thevakumar , Judith Land , Junchen Jiang , Jungwhan Kim , Kairit Sirts , Kamesh R , Kamesh V , Kanda Patrick Tshinu , Kätriin Kukk , Kaustubh Ponkshe , Kavsar Huseynova , Ke He , Kelly Buchanan , Kengatharaiyer Sarveswaran , Kerem Zaman , Khalil Mrini , Kian Kyars , Krister Kruusmaa , Kusum Chouhan , Lainitha Krishnakumar , Laura Castro Sánchez , Laura Porrino Moscoso , Leshem Choshen , Levent Sencan , Lilja Øvrelid , Lisa Alazraki , Lovina Ehimen-Ugbede , Luheerathan Thevakumar , Luxshan Thavarasa , Mahnoor Malik , Mamadou K. Keita , Mansi Jangid , Marco De Santis , Marcos García , Marek Suppa , Mariam D'Ciofalo , Marii Ojastu , Maryam Sikander , Mausami Narayan , Maximos Skandalis , Mehak Mehak , Mehmet İlteriş Bozkurt , Melaku Bayu Workie , Menan Velayuthan , Michael Leventhal , Michał Marcińczuk , Mirna Potočnjak , Mohammadamin Shafiei , Mridul Sharma , Mrityunjaya Indoria , Muhammad Ravi Shulthan Habibi , Murat Kolić , Nada Galant , Naphat Permpredanun , Narada Maugin , Nicholas Kluge Corrêa , Nikola Ljubešić , Nirmal Thomas , Nisansa de Silva , Nisheeth Joshi , Nitish Ponkshe , Nizar Habash , Nneoma C. Udeze , Noel Thomas , Noémi Ligeti-Nagy , Nouhoum Coulibaly , Nsengiyumva Faustin , Odunayo Kareemat Buliaminu , Odunayo Ogundepo , Oghojafor Godswill Fejiro , Ogundipe Blessing Funmilola , Okechukwu God'spraise , Olanrewaju Samuel , Olaoye Deborah Oluwaseun , Olasoji Akindejoye , Olga Popova , Olga Snissarenko , Onyinye Anulika Chiemezie , Orkun Kinay , Osman Tursun , Owoeye Tobiloba Moses , Oyelade Oluwafemi Joshua , Oyesanmi Fiyinfoluwa , Pablo Gamallo , Pablo Rodríguez Fernández , Palak Arora , Pedro Valente , Peter Rupnik , Philip Oghenesuowho Ekiugbo , Pramit Sahoo , Prokopis Prokopidis , Pua Niau-Puhipau , Quadri Yahya , Rachele Mignone , Raghav Singhal , Ram Mohan Rao Kadiyala , Raphael Merx , Rapheal Afolayan , Ratnavel Rajalakshmi , Rishav Ghosh , Romina Oji , Ron Kekeha Solis , Rui Guerra , Rushikesh Zawar , Sa'ad Nasir Bashir , Saeed Alzaabi , Sahil Sandeep , Sai Pavan Batchu , SaiSandeep Kantareddy , Salsabila Zahirah Pranida , Sam Buchanan , Samuel Rutunda , Sander Land , Sarah Sulollari , Sardar Ali , Saroj Sapkota , Saulius Tautvaisas , Sayambhu Sen , Sayantani Banerjee , Sebastien Diarra , SenthilNathan. M , Sewoong Lee , Shaan Shah , Shankar Venkitachalam , Sharifa Djurabaeva , Sharon Ibejih , Shivanya Shomir Dutta , Siddhant Gupta , Silvia Paniagua Suárez , Sina Ahmadi , Sivasuthan Sukumar , Siyuan Song , Snegha A. , Sokratis Sofianopoulos , Sona Elza Simon , Sonja Benčina , Sophie Gvasalia , Sphurti Kirit More , Spyros Dragazis , Stephan P. Kaufhold , Suba. S , Sultan AlRashed , Surangika Ranathunga , Taiga Someya , Taja Kuzman Pungeršek , Tal Haklay , Tasi'u Jibril , Tatsuya Aoyama , Tea Abashidze , Terenz Jomar Dela Cruz , Terra Blevins , Themistoklis Nikas , Theresa Dora Idoko , Thu Mai Do , Tilek Chubakov , Tommaso Gargiani , Uma Rathore , Uni Johannesen , Uwuma Doris Ugwu , Vallerie Alexandra Putra , Vanya Bannihatti Kumar , Varsha Jeyarajalingam , Varvara Arzt , Vasudevan Nedumpozhimana , Viktoria Ondrejova , Viktoryia Horbik , Vishnu Vardhan Reddy Kummitha , Vuk Dinić , Walelign Tewabe Sewunetie , Winston Wu , Xiaojing Zhao , Yacouba Diarra , Yaniv Nikankin , Yash Mathur , Yixi Chen , Yiyuan Li , Yolanda Xavier , Yonatan Belinkov , Yusuf Ismail Abayomi , Zaid Alyafeai , Zhengyang Shan , Zhi Rui Tam , Zilu Tang , Zuzana Nadova , Baber Abbasi , Stella Biderman , David Stap , Duygu Ataman , Fabian Schmidt , Hila Gonen , Jiayi Wang , David Ifeoluwa Adelani

Multimodal learning on video and text has seen significant progress, particularly in tasks like text-to-video retrieval, video-to-text retrieval, and video captioning. However, most existing methods and datasets focus exclusively on…

Multimedia · Computer Science 2025-07-15 Willy Fitra Hendria

Automatic assessment of cognitive impairment from spontaneous speech offers a promising, non-invasive avenue for early cognitive screening. However, current approaches often lack generalizability when deployed across different languages and…

Artificial Intelligence · Computer Science 2025-10-20 Rui Feng , Zhiyao Luo , Wei Wang , Yuting Song , Yong Liu , Tingting Zhu , Jianqing Li , Xingyao Wang

Recent advances in large language models (LLMs) have demonstrated remarkable capabilities on widely benchmarked high-resource languages. However, linguistic nuances of under-resourced languages remain unexplored. We introduce Batayan, a…

We present a benchmark suite of four datasets for evaluating the fairness of pre-trained language models and the techniques used to fine-tune them for downstream tasks. Our benchmarks cover four jurisdictions (European Council, USA,…

Computation and Language · Computer Science 2022-03-15 Ilias Chalkidis , Tommaso Pasini , Sheng Zhang , Letizia Tomada , Sebastian Felix Schwemer , Anders Søgaard

Building Natural Language Understanding (NLU) capabilities for Indic languages, which have a collective speaker base of more than one billion speakers is absolutely crucial. In this work, we aim to improve the NLU capabilities of Indic…

Computation and Language · Computer Science 2023-05-25 Sumanth Doddapaneni , Rahul Aralikatte , Gowtham Ramesh , Shreya Goyal , Mitesh M. Khapra , Anoop Kunchukuttan , Pratyush Kumar

Multilingual intent classification is central to customer-service systems on global logistics platforms, where models must process noisy user queries across languages and hierarchical label spaces. Yet most existing multilingual benchmarks…

Computation and Language · Computer Science 2026-03-25 Haoyu He , Jinyu Zhuang , Haoran Chu , Shuhang Yu , J , T AI Group , Hao Wang , Kunpeng Han

We review the recent literature (January 2022- October 2024) in South Asian languages on text-based language processing, multimodal models, and speech processing, and provide a spotlight analysis focused on 21 low-resource South Asian…

Computation and Language · Computer Science 2025-01-03 Pranav Gupta

Recent advances in large language models (LLMs) and medical LLMs (Med-LLMs) have demonstrated strong performance on general medical benchmarks. However, their capabilities in specialized medical fields, such as dentistry which require…

Computation and Language · Computer Science 2025-08-29 Hengchuan Zhu , Yihuan Xu , Yichen Li , Zijie Meng , Zuozhu Liu

Code generation benchmarks such as HumanEval are widely adopted to evaluate LLMs' capabilities. However, after consolidating the latest 24 benchmarks, we noticed three significant imbalances. First, imbalanced programming language. 95.8% of…

Machine Learning · Computer Science 2024-10-14 Jialun Cao , Zhiyong Chen , Jiarong Wu , Shing-chi Cheung , Chang Xu

The rapid advancement of large language models(LLMs) has intensified the need for domain and culture specific evaluation. Existing benchmarks are largely Anglocentric and domain-agnostic, limiting their applicability to India-centric…

Large Language Models (LLMs) have demonstrated remarkable performance across various disciplines and tasks. However, benchmarking their capabilities with multilingual spoken queries remains largely unexplored. In this study, we introduce…

Computation and Language · Computer Science 2025-05-27 Firoj Alam , Md Arid Hasan , Shammur Absar Chowdhury

Assessing the capabilities and limitations of large language models (LLMs) has garnered significant interest, yet the evaluation of multiple models in real-world scenarios remains rare. Multilingual evaluation often relies on translated…

Computation and Language · Computer Science 2025-12-02 Varun Gumma , Ananditha Raghunath , Mohit Jain , Sunayana Sitaram

Logging statements are central to debugging, failure diagnosis, and production observability, yet writing them requires developers to decide where to place a logging statement, which API and severity level to use, and what runtime…

Software Engineering · Computer Science 2026-04-21 Renyi Zhong , Yichen Li , Yulun Wu , Jinxi Kuang , Yintong Huo , Michael R. Lyu

Recent developments in Japanese large language models (LLMs) primarily focus on general domains, with fewer advancements in Japanese biomedical LLMs. One obstacle is the absence of a comprehensive, large-scale benchmark for comparison.…

Computation and Language · Computer Science 2024-09-23 Junfeng Jiang , Jiahao Huang , Akiko Aizawa

This paper presents a novel syllable-based tokenization approach for Indonesian large language models, inspired by the Gasing Literacy Learning System's pedagogical methodology. Drawing on information-theoretic principles, we develop a…

Computers and Society · Computer Science 2026-01-21 H. Situngkir , A. B. Lumbantobing , Y. Surya