English
Related papers

Related papers: Sabi\'a-4 Technical Report

200 papers

Prompt engineering is crucial for unlocking the potential of Large Language Models (LLMs). Still, since manual prompt design is often complex, non-intuitive, and time-consuming, automatic prompt optimization has emerged as a research area.…

Computation and Language · Computer Science 2025-10-30 Sara Câmara , Eduardo Luz , Valéria Carvalho , Ivan Meneghini , Gladston Moreira

A new algorithm for voice automatic syllabic splitting in the Portuguese language is proposed, which is based on the envelope of the speech signal of the input audio file. A computational implementation in MatlabTM is presented and made…

Sound · Computer Science 2018-01-24 E. L. F. Da Silva , H. M. de Oliveira

In natural language processing (NLP), there is a need for more resources in Portuguese, since much of the data used in the state-of-the-art research is in other languages. In this paper, we pretrain a T5 model on the BrWac corpus, an…

Computation and Language · Computer Science 2020-10-12 Diedre Carmo , Marcos Piau , Israel Campiotti , Rodrigo Nogueira , Roberto Lotufo

We introduce LegalBench-BR, the first public benchmark for evaluating language models on Brazilian legal text classification. The dataset comprises 3,105 appellate proceedings from the Santa Catarina State Court (TJSC), collected via the…

Computation and Language · Computer Science 2026-04-23 Pedro Barbosa de Carvalho Neto

In this paper we present an efficient method for training models for speaker recognition using small or under-resourced datasets. This method requires less data than other SOTA (State-Of-The-Art) methods, e.g. the Angular Prototypical and…

The recent integration of visual capabilities into Large Language Models (LLMs) has the potential to play a pivotal role in science and technology education, where visual elements such as diagrams, charts, and tables are commonly used to…

Artificial Intelligence · Computer Science 2024-06-17 Nabor C. Mendonça

Leveraging research on the neural modelling of Portuguese, we contribute a collection of datasets for an array of language processing tasks and a corresponding collection of fine-tuned neural language models on these downstream tasks. To…

Computation and Language · Computer Science 2024-05-10 Tomás Osório , Bernardo Leite , Henrique Lopes Cardoso , Luís Gomes , João Rodrigues , Rodrigo Santos , António Branco

We investigate the effectiveness of GPT-3.5 and GPT-4, two large language models, as Grammatical Error Correction (GEC) tools for Brazilian Portuguese and compare their performance against Microsoft Word and Google Docs. We introduce a GEC…

Computation and Language · Computer Science 2023-07-19 Maria Carolina Penteado , Fábio Perez

The Brazilian Supreme Court receives tens of thousands of cases each semester. Court employees spend thousands of hours to execute the initial analysis and classification of those cases -- which takes effort away from posterior, more…

Large language models (LLMs) are increasingly deployed as task-oriented agents, where success depends on their ability to generate accurate function calls under realistic, multilingual conditions. However, existing agent evaluations largely…

Computation and Language · Computer Science 2025-09-19 Thales Sales Almeida , João Guilherme Alves Santos , Thiago Laitz , Giovana Kerche Bonás

Analyses of legislative behavior often rely on voting records, overlooking the rich semantic and rhetorical content of political speech. In this paper, we ask three complementary questions about parliamentary discourse: how things are said,…

Recent Speech-to-Text models often require a large amount of hardware resources and are mostly trained in English. This paper presents Speech-to-Text models for German, as well as for Spanish and French with special features: (a) They are…

Computation and Language · Computer Science 2021-10-18 Daniel Bermuth , Alexander Poeppel , Wolfgang Reif

This paper presents two studies on how Brazilian children (ages 9--11) use conversational agents (CAs) for schoolwork, discovery, and entertainment, and how structured scaffolds can enhance these interactions. In Study 1, a seven-week…

Human-Computer Interaction · Computer Science 2025-09-01 Vanessa Figueiredo

Recent advances in natural language processing have raised expectations for generative models to produce coherent text across diverse language varieties. In the particular case of the Portuguese language, the predominance of Brazilian…

Computation and Language · Computer Science 2025-02-21 Hugo Sousa , Rúben Almeida , Purificação Silvano , Inês Cantante , Ricardo Campos , Alípio Jorge

Context: Quantitative studies can identify statistical predictors of training quality, but they often fail to capture what professionals themselves consider genuinely useful learning experiences and why. Objective: This study qualitatively…

Software Engineering · Computer Science 2026-04-08 Rodrigo Siqueira , Antonio Oliveira , Breno Alves de Andrade , Lidiane C S Gomes , Danilo Monteiro Ribeiro

Significant advances have been made in natural language processing in recent years. However, our current deep learning approach to language modeling requires substantial resources in terms of data and computation. One of the side effects of…

Computation and Language · Computer Science 2025-07-25 Nicholas Kluge Corrêa , Aniket Sen , Sophia Falk , Shiza Fatimah

Large language models (LLMs) are increasingly used as sources of information, yet their reliability depends on the ability to search the web, select relevant evidence, and synthesize complete answers. While recent benchmarks evaluate…

Continued pretraining extends a language model's capabilities by further exposing it to additional data, often tailored to a specific linguistic or domain context. This strategy has emerged as an efficient alternative to full retraining…

Computation and Language · Computer Science 2025-12-16 Thales Sales Almeida , Rodrigo Nogueira , Hélio Pedrini

This paper reports on the development of a leaderboard of Open Large Language Models (LLM) for European Portuguese (PT-PT), and on its associated benchmarks. This leaderboard comes as a way to address a gap in the evaluation of LLM for…

Computation and Language · Computer Science 2026-03-16 João Silva , Luís Gomes , António Branco

The high compute cost associated with pretraining large language models limits their research. Two strategies have emerged to address this issue: domain specialization and pretraining with high-quality data. To explore these strategies, we…

Computation and Language · Computer Science 2025-07-29 Roseval Malaquias Junior , Ramon Pires , Roseli Romero , Rodrigo Nogueira