English
Related papers

Related papers: Salamandra Technical Report

200 papers

We introduce Multilingual Document Parsing Benchmark, the first benchmark for multilingual digital and photographed document parsing. Document parsing has made remarkable strides, yet almost exclusively on clean, digital, well-formatted…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Zhang Li , Zhibo Lin , Qiang Liu , Ziyang Zhang , Shuo Zhang , Zidun Guo , Jiajun Song , Jiarui Zhang , Xiang Bai , Yuliang Liu

Open-source large language models are becoming increasingly available and popular among researchers and practitioners. While significant progress has been made on open-weight models, open training data is a practice yet to be adopted by the…

Computation and Language · Computer Science 2024-11-19 Catherine Arnett , Eliot Jones , Ivan P. Yamshchikov , Pierre-Carl Langlais

Continued pretraining and instruction tuning on large-scale multilingual data have proven to be effective in scaling large language models (LLMs) to low-resource languages. However, the unaligned nature of such data limits its ability to…

Computation and Language · Computer Science 2025-10-22 Yingli Shen , Wen Lai , Shuo Wang , Ge Gao , Kangyang Luo , Alexander Fraser , Maosong Sun

Program synthesis strives to generate a computer program as a solution to a given problem specification, expressed with input-output examples or natural language descriptions. The prevalence of large language models advances the…

Machine Learning · Computer Science 2023-03-01 Erik Nijkamp , Bo Pang , Hiroaki Hayashi , Lifu Tu , Huan Wang , Yingbo Zhou , Silvio Savarese , Caiming Xiong

Large Language Models (LLMs) play an ever-increasing role in the field of Artificial Intelligence (AI)--not only for natural language processing but also for code understanding and generation. To stimulate open and responsible research on…

Low-resource languages such as Sinhala are often overlooked by open-source Large Language Models (LLMs). In this research, we extend an existing multilingual LLM (Llama-3-8B) to better serve Sinhala. We enhance the LLM tokenizer with…

Computation and Language · Computer Science 2025-11-11 H. W. K. Aravinda , Rashad Sirajudeen , Samith Karunathilake , Nisansa de Silva , Surangika Ranathunga , Rishemjit Kaur

Code large language models (LLMs) have shown remarkable advances in code understanding, completion, and generation tasks. Programming benchmarks, comprised of a selection of code challenges and corresponding test cases, serve as a standard…

Multi-task and multilingual approaches benefit large models, yet speech processing for low-resource languages remains underexplored due to data scarcity. To address this, we present Granary, a large-scale collection of speech datasets for…

Mathematical reasoning is a cornerstone of human intelligence and a key benchmark for advanced capabilities in large language models (LLMs). However, the research community still lacks an open, large-scale, high-quality corpus tailored to…

Computation and Language · Computer Science 2025-04-04 Fan Zhou , Zengzhi Wang , Nikhil Ranjan , Zhoujun Cheng , Liping Tang , Guowei He , Zhengzhong Liu , Eric P. Xing

Since the release of ChatGPT, the field of Natural Language Processing has experienced rapid advancements, particularly in Large Language Models (LLMs) and their multimodal counterparts, Large Multimodal Models (LMMs). Despite their…

Computation and Language · Computer Science 2024-08-27 Florian Schneider , Sunayana Sitaram

As large language models continue to advance, their application in educational contexts remains underexplored and under-optimized. In this paper, we address this gap by introducing the first diverse benchmark tailored for educational…

Computation and Language · Computer Science 2026-01-07 Bin Xu , Yu Bai , Huashan Sun , Yiguan Lin , Siming Liu , Xinyue Liang , Yaolin Li , Zhuangzhi Dong , Jingren Zhang , Yufan Deng , Xinyu Zou , Yang Gao , Heyan Huang

In this report we describe the development of Command A, a powerful large language model purpose-built to excel at real-world enterprise use cases. Command A is an agent-optimised and multilingual-capable model, with support for 23…

Computation and Language · Computer Science 2025-04-15 Team Cohere , : , Aakanksha , Arash Ahmadian , Marwan Ahmed , Jay Alammar , Milad Alizadeh , Yazeed Alnumay , Sophia Althammer , Arkady Arkhangorodsky , Viraat Aryabumi , Dennis Aumiller , Raphaël Avalos , Zahara Aviv , Sammie Bae , Saurabh Baji , Alexandre Barbet , Max Bartolo , Björn Bebensee , Neeral Beladia , Walter Beller-Morales , Alexandre Bérard , Andrew Berneshawi , Anna Bialas , Phil Blunsom , Matt Bobkin , Adi Bongale , Sam Braun , Maxime Brunet , Samuel Cahyawijaya , David Cairuz , Jon Ander Campos , Cassie Cao , Kris Cao , Roman Castagné , Julián Cendrero , Leila Chan Currie , Yash Chandak , Diane Chang , Giannis Chatziveroglou , Hongyu Chen , Claire Cheng , Alexis Chevalier , Justin T. Chiu , Eugene Cho , Eugene Choi , Eujeong Choi , Tim Chung , Volkan Cirik , Ana Cismaru , Pierre Clavier , Henry Conklin , Lucas Crawhall-Stein , Devon Crouse , Andres Felipe Cruz-Salinas , Ben Cyrus , Daniel D'souza , Hugo Dalla-Torre , John Dang , William Darling , Omar Darwiche Domingues , Saurabh Dash , Antoine Debugne , Théo Dehaze , Shaan Desai , Joan Devassy , Rishit Dholakia , Kyle Duffy , Ali Edalati , Ace Eldeib , Abdullah Elkady , Sarah Elsharkawy , Irem Ergün , Beyza Ermis , Marzieh Fadaee , Boyu Fan , Lucas Fayoux , Yannis Flet-Berliac , Nick Frosst , Matthias Gallé , Wojciech Galuba , Utsav Garg , Matthieu Geist , Mohammad Gheshlaghi Azar , Ellen Gilsenan-McMahon , Seraphina Goldfarb-Tarrant , Tomas Goldsack , Aidan Gomez , Victor Machado Gonzaga , Nithya Govindarajan , Manoj Govindassamy , Nathan Grinsztajn , Nikolas Gritsch , Patrick Gu , Shangmin Guo , Kilian Haefeli , Rod Hajjar , Tim Hawes , Jingyi He , Sebastian Hofstätter , Sungjin Hong , Sara Hooker , Tom Hosking , Stephanie Howe , Eric Hu , Renjie Huang , Hemant Jain , Ritika Jain , Nick Jakobi , Madeline Jenkins , JJ Jordan , Dhruti Joshi , Jason Jung , Trushant Kalyanpur , Siddhartha Rao Kamalakara , Julia Kedrzycki , Gokce Keskin , Edward Kim , Joon Kim , Wei-Yin Ko , Tom Kocmi , Michael Kozakov , Wojciech Kryściński , Arnav Kumar Jain , Komal Kumar Teru , Sander Land , Michael Lasby , Olivia Lasche , Justin Lee , Patrick Lewis , Jeffrey Li , Jonathan Li , Hangyu Lin , Acyr Locatelli , Kevin Luong , Raymond Ma , Lukáš Mach , Marina Machado , Joanne Magbitang , Brenda Malacara Lopez , Aryan Mann , Kelly Marchisio , Olivia Markham , Alexandre Matton , Alex McKinney , Dominic McLoughlin , Jozef Mokry , Adrien Morisot , Autumn Moulder , Harry Moynehan , Maximilian Mozes , Vivek Muppalla , Lidiya Murakhovska , Hemangani Nagarajan , Alekhya Nandula , Hisham Nasir , Shauna Nehra , Josh Netto-Rosen , Daniel Ohashi , James Owers-Bardsley , Jason Ozuzu , Dennis Padilla , Gloria Park , Sam Passaglia , Jeremy Pekmez , Laura Penstone , Aleksandra Piktus , Case Ploeg , Andrew Poulton , Youran Qi , Shubha Raghvendra , Miguel Ramos , Ekagra Ranjan , Pierre Richemond , Cécile Robert-Michon , Aurélien Rodriguez , Sudip Roy , Sebastian Ruder , Laura Ruis , Louise Rust , Anubhav Sachan , Alejandro Salamanca , Kailash Karthik Saravanakumar , Isha Satyakam , Alice Schoenauer Sebag , Priyanka Sen , Sholeh Sepehri , Preethi Seshadri , Ye Shen , Tom Sherborne , Sylvie Shang Shi , Sanal Shivaprasad , Vladyslav Shmyhlo , Anirudh Shrinivason , Inna Shteinbuk , Amir Shukayev , Mathieu Simard , Ella Snyder , Ava Spataru , Victoria Spooner , Trisha Starostina , Florian Strub , Yixuan Su , Jimin Sun , Dwarak Talupuru , Eugene Tarassov , Elena Tommasone , Jennifer Tracey , Billy Trend , Evren Tumer , Ahmet Üstün , Bharat Venkitesh , David Venuto , Pat Verga , Maxime Voisin , Alex Wang , Donglu Wang , Shijian Wang , Edmond Wen , Naomi White , Jesse Willman , Marysia Winkels , Chen Xia , Jessica Xie , Minjie Xu , Bowen Yang , Tan Yi-Chern , Ivan Zhang , Zhenyu Zhao , Zhoujie Zhao

OpenDiLoCo is an open-source implementation and replication of the Distributed Low-Communication (DiLoCo) training method for large language models. We provide a reproducible implementation of the DiLoCo experiments, offering it within a…

Machine Learning · Computer Science 2024-07-11 Sami Jaghouar , Jack Min Ong , Johannes Hagemann

The large transformer-based language models demonstrate excellent performance in natural language processing. By considering the transferability of the knowledge gained by these models in one domain to other related domains, and the…

Cryptography and Security · Computer Science 2022-09-07 Chandra Thapa , Seung Ick Jang , Muhammad Ejaz Ahmed , Seyit Camtepe , Josef Pieprzyk , Surya Nepal

Practically all large language models have been pre-trained on data that is subject to global uncertainty related to copyright infringement and breach of contract. This creates potential risk for users and developers due to this uncertain…

Computation and Language · Computer Science 2025-04-11 Michael J Bommarito , Jillian Bommarito , Daniel Martin Katz

Foundational large language models (LLMs) can be instruction-tuned to perform open-domain question answering, facilitating applications like chat assistants. While such efforts are often carried out in a single language, we empirically…

Computation and Language · Computer Science 2024-02-01 Pinzhen Chen , Shaoxiong Ji , Nikolay Bogoychev , Andrey Kutuzov , Barry Haddow , Kenneth Heafield

Using representations provided by a large pre-trained model has become the primary strategy for achieving state-of-the-art results in a wide range of tasks. A recently proposed large pre-trained model, wav2vec 2.0, was seminal for several…

Computation and Language · Computer Science 2025-12-01 Jonatas Grosman , Cassio Almeida , Guilherme Schardong , Hélio Lopes

Financial LLMs hold promise for advancing financial tasks and domain-specific applications. However, they are limited by scarce corpora, weak multimodal capabilities, and narrow evaluations, making them less suited for real-world…

Pre-trained multilingual language models have become an important building block in multilingual natural language processing. In the present paper, we investigate a range of such models to find out how well they transfer discourse-level…

Computation and Language · Computer Science 2021-06-10 Murathan Kurfalı , Robert Östling

The use of language models for automatically evaluating long-form text (LLM-as-a-judge) is becoming increasingly common, yet most LLM judges are optimized exclusively for English, with strategies for enhancing their multilingual evaluation…

Computation and Language · Computer Science 2025-10-31 José Pombal , Dongkeun Yoon , Patrick Fernandes , Ian Wu , Seungone Kim , Ricardo Rei , Graham Neubig , André F. T. Martins
‹ Prev 1 8 9 10 Next ›