中文
相关论文

相关论文: Sabi\'a-4 Technical Report

200 篇论文

Despite significant advances in speech processing, Portuguese remains under-resourced due to the scarcity of public, large-scale, and high-quality datasets. To address this gap, we present a new dataset, named TAGARELA, composed of over…

Large Language Models (LLMs) have significantly advanced the development of Legal Artificial Intelligence (Legal AI) in recent years, enhancing the efficiency and accuracy of legal tasks. To advance research and applications of LLM-based…

计算与语言 · 计算机科学 2025-09-15 Zhitian Hou , Zihan Ye , Nanli Zeng , Tianyong Hao , Kun Zeng

We investigate transfer learning based on pre-trained neural machine translation models to translate between (low-resource) similar languages. This work is part of our contribution to the WMT 2021 Similar Languages Translation Shared Task…

人工智能 · 计算机科学 2021-10-08 Ife Adebara , Muhammad Abdul-Mageed

This paper presents recent progress in the acoustic modelling of under-resourced code-switched (CS) speech in multiple South African languages. We consider two approaches. The first constructs separate bilingual acoustic models…

计算与语言 · 计算机科学 2019-10-16 Astik Biswas , Emre Yılmaz , Febe de Wet , Ewald van der Westhuizen , Thomas Niesler

With the rapid integration of advanced reasoning capabilities into spoken dialogue models, the field urgently demands benchmarks that transcend simple interactions to address real-world complexity. However, current evaluations predominantly…

计算与语言 · 计算机科学 2026-02-16 Yangzhuo Li , Shengpeng Ji , Yifu Chen , Tianle Liang , Haorong Ying , Yule Wang , Junbo Li , Jun Fang , Zhou Zhao

While Large Language Models demonstrate remarkable proficiency in high-level semantic planning, they remain limited in handling fine-grained, low-level web component manipulations. To address this limitation, extensive research has focused…

人工智能 · 计算机科学 2026-01-22 Zhi Qiu , Jiazheng Sun , Chenxiao Xia , Jun Zheng , Xin Peng

Recent advances in Artificial Intelligence (AI) have leveraged promising results in solving complex problems in the area of Natural Language Processing (NLP), being an important tool to help in the expeditious resolution of judicial…

人工智能 · 计算机科学 2023-05-12 Raphael Souza de Oliveira , Erick Giovani Sperandio Nascimento

Despite significant advancements and pervasive use of vision-language models, a paucity of studies has addressed their ethical implications. These models typically require extensive training data, often from hastily reviewed text and image…

Despite the recent advances in Large Language Models, benchmarks for evaluating legal writing remain scarce due to the inherent complexity of assessing open-ended responses in this domain. One of the key challenges in evaluating language…

计算与语言 · 计算机科学 2025-05-01 Ramon Pires , Roseval Malaquias Junior , Rodrigo Nogueira

Binding precedents (s\'umulas vinculantes) constitute a juridical instrument unique to the Brazilian legal system and whose objectives include the protection of the Federal Supreme Court against repetitive demands. Studies of the…

计算与语言 · 计算机科学 2025-05-29 Raphaël Tinarrage , Henrique Ennes , Lucas Resck , Lucas T. Gomes , Jean R. Ponciano , Jorge Poco

Recent advances in spoken language understanding benefited from Self-Supervised models trained on large speech corpora. For French, the LeBenchmark project has made such models available and has led to impressive progress on several tasks…

计算与语言 · 计算机科学 2022-07-04 Marco Dinarelli , Marco Naguib , François Portet

Legal information retrieval in Portuguese remains difficult to evaluate systematically because available datasets differ widely in document type, query style, and relevance definition. We present JU\'A, a public benchmark for Brazilian…

信息检索 · 计算机科学 2026-04-09 Jayr Pereira , Leandro Fernandes , Erick de Brito , Roberto Lotufo , Luiz Bonifacio

We present a benchmark suite of four datasets for evaluating the fairness of pre-trained language models and the techniques used to fine-tune them for downstream tasks. Our benchmarks cover four jurisdictions (European Council, USA,…

计算与语言 · 计算机科学 2022-03-15 Ilias Chalkidis , Tommaso Pasini , Sheng Zhang , Letizia Tomada , Sebastian Felix Schwemer , Anders Søgaard

While the field of style transfer (ST) has been growing rapidly, it has been hampered by a lack of standardized practices for automatic evaluation. In this paper, we evaluate leading ST automatic metrics on the oft-researched task of…

计算与语言 · 计算机科学 2021-10-22 Eleftheria Briakou , Sweta Agrawal , Joel Tetreault , Marine Carpuat

In this paper we describe the Portuguese-language podcast dataset we have released for academic research purposes. We give an overview of how the data was sampled, descriptive statistics over the collection, as well as information about the…

In the last few years, three major topics received increased interest: deep learning, NLP and conversational agents. Bringing these three topics together to create an amazing digital customer experience and indeed deploy in production and…

计算与语言 · 计算机科学 2021-07-27 Paulo Finardi , José Dié Viegas , Gustavo T. Ferreira , Alex F. Mansano , Vinicius F. Caridá

Significant strides have been made in natural language tasks, largely attributed to the emergence of powerful large language models (LLMs). These models, pre-trained on extensive and diverse corpora, have become increasingly capable of…

计算与语言 · 计算机科学 2024-02-21 Ricardo Lopes , João Magalhães , David Semedo

The use of large language models (LLMs) for complex mathematical reasoning is an emergent area of research, with fast progress in methods, models, and benchmark datasets. However, most mathematical reasoning evaluations exhibit a…

This paper presents and makes publicly available the NILC-Metrix, a computational system comprising 200 metrics proposed in studies on discourse, psycholinguistics, cognitive and computational linguistics, to assess textual complexity in…