English
Related papers

Related papers: Sabi\'a-4 Technical Report

200 papers

Despite significant advances in speech processing, Portuguese remains under-resourced due to the scarcity of public, large-scale, and high-quality datasets. To address this gap, we present a new dataset, named TAGARELA, composed of over…

Large Language Models (LLMs) have significantly advanced the development of Legal Artificial Intelligence (Legal AI) in recent years, enhancing the efficiency and accuracy of legal tasks. To advance research and applications of LLM-based…

Computation and Language · Computer Science 2025-09-15 Zhitian Hou , Zihan Ye , Nanli Zeng , Tianyong Hao , Kun Zeng

We investigate transfer learning based on pre-trained neural machine translation models to translate between (low-resource) similar languages. This work is part of our contribution to the WMT 2021 Similar Languages Translation Shared Task…

Artificial Intelligence · Computer Science 2021-10-08 Ife Adebara , Muhammad Abdul-Mageed

This paper presents recent progress in the acoustic modelling of under-resourced code-switched (CS) speech in multiple South African languages. We consider two approaches. The first constructs separate bilingual acoustic models…

Computation and Language · Computer Science 2019-10-16 Astik Biswas , Emre Yılmaz , Febe de Wet , Ewald van der Westhuizen , Thomas Niesler

With the rapid integration of advanced reasoning capabilities into spoken dialogue models, the field urgently demands benchmarks that transcend simple interactions to address real-world complexity. However, current evaluations predominantly…

Computation and Language · Computer Science 2026-02-16 Yangzhuo Li , Shengpeng Ji , Yifu Chen , Tianle Liang , Haorong Ying , Yule Wang , Junbo Li , Jun Fang , Zhou Zhao

While Large Language Models demonstrate remarkable proficiency in high-level semantic planning, they remain limited in handling fine-grained, low-level web component manipulations. To address this limitation, extensive research has focused…

Artificial Intelligence · Computer Science 2026-01-22 Zhi Qiu , Jiazheng Sun , Chenxiao Xia , Jun Zheng , Xin Peng

Recent advances in Artificial Intelligence (AI) have leveraged promising results in solving complex problems in the area of Natural Language Processing (NLP), being an important tool to help in the expeditious resolution of judicial…

Artificial Intelligence · Computer Science 2023-05-12 Raphael Souza de Oliveira , Erick Giovani Sperandio Nascimento

Despite significant advancements and pervasive use of vision-language models, a paucity of studies has addressed their ethical implications. These models typically require extensive training data, often from hastily reviewed text and image…

Despite the recent advances in Large Language Models, benchmarks for evaluating legal writing remain scarce due to the inherent complexity of assessing open-ended responses in this domain. One of the key challenges in evaluating language…

Computation and Language · Computer Science 2025-05-01 Ramon Pires , Roseval Malaquias Junior , Rodrigo Nogueira

Binding precedents (s\'umulas vinculantes) constitute a juridical instrument unique to the Brazilian legal system and whose objectives include the protection of the Federal Supreme Court against repetitive demands. Studies of the…

Computation and Language · Computer Science 2025-05-29 Raphaël Tinarrage , Henrique Ennes , Lucas Resck , Lucas T. Gomes , Jean R. Ponciano , Jorge Poco

Recent advances in spoken language understanding benefited from Self-Supervised models trained on large speech corpora. For French, the LeBenchmark project has made such models available and has led to impressive progress on several tasks…

Computation and Language · Computer Science 2022-07-04 Marco Dinarelli , Marco Naguib , François Portet

Legal information retrieval in Portuguese remains difficult to evaluate systematically because available datasets differ widely in document type, query style, and relevance definition. We present JU\'A, a public benchmark for Brazilian…

Information Retrieval · Computer Science 2026-04-09 Jayr Pereira , Leandro Fernandes , Erick de Brito , Roberto Lotufo , Luiz Bonifacio

We present a benchmark suite of four datasets for evaluating the fairness of pre-trained language models and the techniques used to fine-tune them for downstream tasks. Our benchmarks cover four jurisdictions (European Council, USA,…

Computation and Language · Computer Science 2022-03-15 Ilias Chalkidis , Tommaso Pasini , Sheng Zhang , Letizia Tomada , Sebastian Felix Schwemer , Anders Søgaard

While the field of style transfer (ST) has been growing rapidly, it has been hampered by a lack of standardized practices for automatic evaluation. In this paper, we evaluate leading ST automatic metrics on the oft-researched task of…

Computation and Language · Computer Science 2021-10-22 Eleftheria Briakou , Sweta Agrawal , Joel Tetreault , Marine Carpuat

In this paper we describe the Portuguese-language podcast dataset we have released for academic research purposes. We give an overview of how the data was sampled, descriptive statistics over the collection, as well as information about the…

Computation and Language · Computer Science 2023-12-14 Ekaterina Garmash , Edgar Tanaka , Ann Clifton , Joana Correia , Sharmistha Jat , Winstead Zhu , Rosie Jones , Jussi Karlgren

In the last few years, three major topics received increased interest: deep learning, NLP and conversational agents. Bringing these three topics together to create an amazing digital customer experience and indeed deploy in production and…

Computation and Language · Computer Science 2021-07-27 Paulo Finardi , José Dié Viegas , Gustavo T. Ferreira , Alex F. Mansano , Vinicius F. Caridá

Significant strides have been made in natural language tasks, largely attributed to the emergence of powerful large language models (LLMs). These models, pre-trained on extensive and diverse corpora, have become increasingly capable of…

Computation and Language · Computer Science 2024-02-21 Ricardo Lopes , João Magalhães , David Semedo

The use of large language models (LLMs) for complex mathematical reasoning is an emergent area of research, with fast progress in methods, models, and benchmark datasets. However, most mathematical reasoning evaluations exhibit a…

This paper presents and makes publicly available the NILC-Metrix, a computational system comprising 200 metrics proposed in studies on discourse, psycholinguistics, cognitive and computational linguistics, to assess textual complexity in…