English
Related papers

Related papers: Towards Analyzing N-language Polyglot Programs

200 papers

Underlying data distributions of natural language, programming code, and mathematical symbols vary vastly, presenting a complex challenge for large language models (LLMs) that strive to achieve high performance across all three domains…

Computation and Language · Computer Science 2024-03-27 Ning Ding , Yulin Chen , Ganqu Cui , Xingtai Lv , Weilin Zhao , Ruobing Xie , Bowen Zhou , Zhiyuan Liu , Maosong Sun

Code-mixing, the practice of switching between languages within a conversation, poses unique challenges for traditional NLP. Existing benchmarks are limited by their narrow language pairs and tasks, failing to adequately assess large…

Computation and Language · Computer Science 2025-09-09 Yilun Yang , Yekun Chai

Many software projects implement APIs and algorithms in multiple programming languages. Maintaining such projects is tiresome, as developers have to ensure that any change (e.g., a bug fix or a new feature) is being propagated, timely and…

Software Engineering · Computer Science 2023-09-13 Jiyang Zhang , Pengyu Nie , Junyi Jessy Li , Milos Gligoric

In this paper, we propose a model-agnostic cost-effective approach to developing bilingual base large language models (LLMs) to support English and any target language. The method includes vocabulary expansion, initialization of new…

Large language models (LLMs) are gaining increasing popularity in both academia and industry, owing to their unprecedented performance in various applications. As LLMs continue to play a vital role in both research and daily use, their…

Computation and Language · Computer Science 2024-01-01 Yupeng Chang , Xu Wang , Jindong Wang , Yuan Wu , Linyi Yang , Kaijie Zhu , Hao Chen , Xiaoyuan Yi , Cunxiang Wang , Yidong Wang , Wei Ye , Yue Zhang , Yi Chang , Philip S. Yu , Qiang Yang , Xing Xie

The rapid evolution of large language models (LLMs) has opened new possibilities for automating various tasks in software development. This paper evaluates the capabilities of the Llama 2-70B model in automating these tasks for scientific…

Software Engineering · Computer Science 2025-07-09 Patrick Diehl , Nojoud Nader , Maxim Moraru , Steven R. Brandt

As safety remains a crucial concern throughout the development lifecycle of Large Language Models (LLMs), researchers and industrial practitioners have increasingly focused on safeguarding and aligning LLM behaviors with human preferences…

Computation and Language · Computer Science 2024-07-11 Jiayang Song , Yuheng Huang , Zhehua Zhou , Lei Ma

Large Language Models (LLMs) are increasingly being integrated into various medical fields, including mental health support systems. However, there is a gap in research regarding the effectiveness of LLMs in non-English mental health…

Computation and Language · Computer Science 2026-02-10 Konstantinos Skianis , John Pavlopoulos , A. Seza Doğruöz

Large language model (LLM) agents have shown increasing promise for collaborative task completion. However, existing multi-agent frameworks often rely on static workflows, fixed roles, and limited inter-agent communication, reducing their…

Multiagent Systems · Computer Science 2026-02-13 Chengxuan Xia , Qianye Wu , Sixuan Tian , Yilun Hao

Large language models (LLMs) offer promise in generating educational content, providing instructor feedback, and reducing teacher workload on assessments. While prior studies have focused on studying LLM-powered learning analytics, limited…

Computation and Language · Computer Science 2024-11-08 Anand Syamkumar , Nora Tseng , Kaycie Barron , Shanglin Yang , Shamya Karumbaiah , Rheeya Uppal , Junjie Hu

LLM-based coding agents are increasingly used to generate code, tests, and documentation. Still, their outputs can be plausible yet misaligned with developer intent and provide limited evidence for review in evolving projects. This limits…

Software Engineering · Computer Science 2026-04-14 Ragib Shahariar Ayon

Large language models (LLMs) have been shown to be able to perform new tasks based on a few demonstrations or natural language instructions. While these capabilities have led to widespread adoption, most LLMs are developed by resource-rich…

Computation and Language · Computer Science 2023-06-28 BigScience Workshop , : , Teven Le Scao , Angela Fan , Christopher Akiki , Ellie Pavlick , Suzana Ilić , Daniel Hesslow , Roman Castagné , Alexandra Sasha Luccioni , François Yvon , Matthias Gallé , Jonathan Tow , Alexander M. Rush , Stella Biderman , Albert Webson , Pawan Sasanka Ammanamanchi , Thomas Wang , Benoît Sagot , Niklas Muennighoff , Albert Villanova del Moral , Olatunji Ruwase , Rachel Bawden , Stas Bekman , Angelina McMillan-Major , Iz Beltagy , Huu Nguyen , Lucile Saulnier , Samson Tan , Pedro Ortiz Suarez , Victor Sanh , Hugo Laurençon , Yacine Jernite , Julien Launay , Margaret Mitchell , Colin Raffel , Aaron Gokaslan , Adi Simhi , Aitor Soroa , Alham Fikri Aji , Amit Alfassy , Anna Rogers , Ariel Kreisberg Nitzav , Canwen Xu , Chenghao Mou , Chris Emezue , Christopher Klamm , Colin Leong , Daniel van Strien , David Ifeoluwa Adelani , Dragomir Radev , Eduardo González Ponferrada , Efrat Levkovizh , Ethan Kim , Eyal Bar Natan , Francesco De Toni , Gérard Dupont , Germán Kruszewski , Giada Pistilli , Hady Elsahar , Hamza Benyamina , Hieu Tran , Ian Yu , Idris Abdulmumin , Isaac Johnson , Itziar Gonzalez-Dios , Javier de la Rosa , Jenny Chim , Jesse Dodge , Jian Zhu , Jonathan Chang , Jörg Frohberg , Joseph Tobing , Joydeep Bhattacharjee , Khalid Almubarak , Kimbo Chen , Kyle Lo , Leandro Von Werra , Leon Weber , Long Phan , Loubna Ben allal , Ludovic Tanguy , Manan Dey , Manuel Romero Muñoz , Maraim Masoud , María Grandury , Mario Šaško , Max Huang , Maximin Coavoux , Mayank Singh , Mike Tian-Jian Jiang , Minh Chien Vu , Mohammad A. Jauhar , Mustafa Ghaleb , Nishant Subramani , Nora Kassner , Nurulaqilla Khamis , Olivier Nguyen , Omar Espejel , Ona de Gibert , Paulo Villegas , Peter Henderson , Pierre Colombo , Priscilla Amuok , Quentin Lhoest , Rheza Harliman , Rishi Bommasani , Roberto Luis López , Rui Ribeiro , Salomey Osei , Sampo Pyysalo , Sebastian Nagel , Shamik Bose , Shamsuddeen Hassan Muhammad , Shanya Sharma , Shayne Longpre , Somaieh Nikpoor , Stanislav Silberberg , Suhas Pai , Sydney Zink , Tiago Timponi Torrent , Timo Schick , Tristan Thrush , Valentin Danchev , Vassilina Nikoulina , Veronika Laippala , Violette Lepercq , Vrinda Prabhu , Zaid Alyafeai , Zeerak Talat , Arun Raja , Benjamin Heinzerling , Chenglei Si , Davut Emre Taşar , Elizabeth Salesky , Sabrina J. Mielke , Wilson Y. Lee , Abheesht Sharma , Andrea Santilli , Antoine Chaffin , Arnaud Stiegler , Debajyoti Datta , Eliza Szczechla , Gunjan Chhablani , Han Wang , Harshit Pandey , Hendrik Strobelt , Jason Alan Fries , Jos Rozen , Leo Gao , Lintang Sutawika , M Saiful Bari , Maged S. Al-shaibani , Matteo Manica , Nihal Nayak , Ryan Teehan , Samuel Albanie , Sheng Shen , Srulik Ben-David , Stephen H. Bach , Taewoon Kim , Tali Bers , Thibault Fevry , Trishala Neeraj , Urmish Thakker , Vikas Raunak , Xiangru Tang , Zheng-Xin Yong , Zhiqing Sun , Shaked Brody , Yallow Uri , Hadar Tojarieh , Adam Roberts , Hyung Won Chung , Jaesung Tae , Jason Phang , Ofir Press , Conglong Li , Deepak Narayanan , Hatim Bourfoune , Jared Casper , Jeff Rasley , Max Ryabinin , Mayank Mishra , Minjia Zhang , Mohammad Shoeybi , Myriam Peyrounette , Nicolas Patry , Nouamane Tazi , Omar Sanseviero , Patrick von Platen , Pierre Cornette , Pierre François Lavallée , Rémi Lacroix , Samyam Rajbhandari , Sanchit Gandhi , Shaden Smith , Stéphane Requena , Suraj Patil , Tim Dettmers , Ahmed Baruwa , Amanpreet Singh , Anastasia Cheveleva , Anne-Laure Ligozat , Arjun Subramonian , Aurélie Névéol , Charles Lovering , Dan Garrette , Deepak Tunuguntla , Ehud Reiter , Ekaterina Taktasheva , Ekaterina Voloshina , Eli Bogdanov , Genta Indra Winata , Hailey Schoelkopf , Jan-Christoph Kalo , Jekaterina Novikova , Jessica Zosa Forde , Jordan Clive , Jungo Kasai , Ken Kawamura , Liam Hazan , Marine Carpuat , Miruna Clinciu , Najoung Kim , Newton Cheng , Oleg Serikov , Omer Antverg , Oskar van der Wal , Rui Zhang , Ruochen Zhang , Sebastian Gehrmann , Shachar Mirkin , Shani Pais , Tatiana Shavrina , Thomas Scialom , Tian Yun , Tomasz Limisiewicz , Verena Rieser , Vitaly Protasov , Vladislav Mikhailov , Yada Pruksachatkun , Yonatan Belinkov , Zachary Bamberger , Zdeněk Kasner , Alice Rueda , Amanda Pestana , Amir Feizpour , Ammar Khan , Amy Faranak , Ana Santos , Anthony Hevia , Antigona Unldreaj , Arash Aghagol , Arezoo Abdollahi , Aycha Tammour , Azadeh HajiHosseini , Bahareh Behroozi , Benjamin Ajibade , Bharat Saxena , Carlos Muñoz Ferrandis , Daniel McDuff , Danish Contractor , David Lansky , Davis David , Douwe Kiela , Duong A. Nguyen , Edward Tan , Emi Baylor , Ezinwanne Ozoani , Fatima Mirza , Frankline Ononiwu , Habib Rezanejad , Hessie Jones , Indrani Bhattacharya , Irene Solaiman , Irina Sedenko , Isar Nejadgholi , Jesse Passmore , Josh Seltzer , Julio Bonis Sanz , Livia Dutra , Mairon Samagaio , Maraim Elbadri , Margot Mieskes , Marissa Gerchick , Martha Akinlolu , Michael McKenna , Mike Qiu , Muhammed Ghauri , Mykola Burynok , Nafis Abrar , Nazneen Rajani , Nour Elkott , Nour Fahmy , Olanrewaju Samuel , Ran An , Rasmus Kromann , Ryan Hao , Samira Alizadeh , Sarmad Shubber , Silas Wang , Sourav Roy , Sylvain Viguier , Thanh Le , Tobi Oyebade , Trieu Le , Yoyo Yang , Zach Nguyen , Abhinav Ramesh Kashyap , Alfredo Palasciano , Alison Callahan , Anima Shukla , Antonio Miranda-Escalada , Ayush Singh , Benjamin Beilharz , Bo Wang , Caio Brito , Chenxi Zhou , Chirag Jain , Chuxin Xu , Clémentine Fourrier , Daniel León Periñán , Daniel Molano , Dian Yu , Enrique Manjavacas , Fabio Barth , Florian Fuhrimann , Gabriel Altay , Giyaseddin Bayrak , Gully Burns , Helena U. Vrabec , Imane Bello , Ishani Dash , Jihyun Kang , John Giorgi , Jonas Golde , Jose David Posada , Karthik Rangasai Sivaraman , Lokesh Bulchandani , Lu Liu , Luisa Shinzato , Madeleine Hahn de Bykhovetz , Maiko Takeuchi , Marc Pàmies , Maria A Castillo , Marianna Nezhurina , Mario Sänger , Matthias Samwald , Michael Cullan , Michael Weinberg , Michiel De Wolf , Mina Mihaljcic , Minna Liu , Moritz Freidank , Myungsun Kang , Natasha Seelam , Nathan Dahlberg , Nicholas Michio Broad , Nikolaus Muellner , Pascale Fung , Patrick Haller , Ramya Chandrasekhar , Renata Eisenberg , Robert Martin , Rodrigo Canalli , Rosaline Su , Ruisi Su , Samuel Cahyawijaya , Samuele Garda , Shlok S Deshmukh , Shubhanshu Mishra , Sid Kiblawi , Simon Ott , Sinee Sang-aroonsiri , Srishti Kumar , Stefan Schweter , Sushil Bharati , Tanmay Laud , Théo Gigant , Tomoya Kainuma , Wojciech Kusa , Yanis Labrak , Yash Shailesh Bajaj , Yash Venkatraman , Yifan Xu , Yingxin Xu , Yu Xu , Zhe Tan , Zhongli Xie , Zifan Ye , Mathilde Bras , Younes Belkada , Thomas Wolf

Open source large language models (LLMs) have shown great improvements in recent times. However, many of these models are focused solely on popular spoken languages. We present a high quality dataset of more than 70k prompt-response pairs…

Computation and Language · Computer Science 2024-05-22 Peter Devine

Large language models (LLMs) showcase increasingly impressive English benchmark scores, however their performance profiles remain inconsistent across multilingual settings. To address this gap, we introduce PolyPrompt, a novel,…

Computation and Language · Computer Science 2025-06-04 Nathan Roll

Multimodal recommender systems (MRS) integrate heterogeneous user and item data, such as text, images, and structured information, to enhance recommendation performance. The emergence of large language models (LLMs) introduces new…

Information Retrieval · Computer Science 2025-05-16 Alejo Lopez-Avila , Jinhua Du

We are experiencing the rise of ChatGPT-like systems or LLMs in political turbulent times. We assume the need to regulate their use because of their bubble-shaping and polarizing potential. To regulate, we need a language that allows…

Computers and Society · Computer Science 2025-08-12 Aernout Schmidt , Kunbei Zhang

Retrieval-Augmented Generation (RAG) systems using Multimodal Large Language Models (MLLMs) show great promise for complex document understanding, yet their development is critically hampered by inadequate evaluation. Current benchmarks…

Computation and Language · Computer Science 2025-08-06 Wenxuan Shen , Mingjia Wang , Yaochen Wang , Dongping Chen , Junjie Yang , Yao Wan , Weiwei Lin

The quality of open-weight LLMs has seen significant improvement, yet they remain predominantly focused on English. In this paper, we introduce the EuroLLM project, aimed at developing a suite of open-weight multilingual LLMs capable of…

The performance of multilingual pretrained models is highly dependent on the availability of monolingual or parallel text present in a target language. Thus, the majority of the world's languages cannot benefit from recent progress in NLP…

Computation and Language · Computer Science 2022-04-07 Xinyi Wang , Sebastian Ruder , Graham Neubig

While new benchmarks for large language models (LLMs) are being developed continuously to catch up with the growing capabilities of new models and AI in general, using and evaluating LLMs in non-English languages remains a little-charted…

Computation and Language · Computer Science 2025-11-05 Špela Vintar , Taja Kuzman Pungeršek , Mojca Brglez , Nikola Ljubešić