中文
相关论文

相关论文: Enhancing Portuguese Variety Identification with C…

200 篇论文

Language models have become foundational to many widely used systems. However, these seemingly advantageous models are double-edged swords. While they excel in tasks related to resource-rich languages like English, they often lose the fine…

计算与语言 · 计算机科学 2025-02-21 Hugo Sousa , Satya Almasian , Ricardo Campos , Alípio Jorge

The performance of large language models (LLMs) is deeply influenced by the quality and composition of their training data. While much of the existing work has centered on English, there remains a gap in understanding how to construct…

计算与语言 · 计算机科学 2025-09-11 Thales Sales Almeida , Rodrigo Nogueira , Helio Pedrini

As Large Language Models (LLMs) expand across multilingual domains, evaluating their performance in under-represented languages becomes increasingly important. European Portuguese (pt-PT) is particularly affected, as existing training data…

In natural language processing (NLP), there is a need for more resources in Portuguese, since much of the data used in the state-of-the-art research is in other languages. In this paper, we pretrain a T5 model on the BrWac corpus, an…

计算与语言 · 计算机科学 2020-10-12 Diedre Carmo , Marcos Piau , Israel Campiotti , Rodrigo Nogueira , Roberto Lotufo

Large Language Models (LLMs) exhibit significant performance variations depending on the linguistic and cultural context in which they are applied. This disparity signals the necessity of mature evaluation frameworks that can assess their…

计算与语言 · 计算机科学 2025-09-12 Thales Sales Almeida , Giovana Kerche Bonás , João Guilherme Alves Santos

Different of biases are reproduced in LLM-generated responses, including dialectal biases. A study based on prompt engineering was carried out to uncover how LLMs discriminate varieties of Brazilian Portuguese, specifically if…

计算与语言 · 计算机科学 2025-01-06 Raquel Meister Ko Freitag , Túlio Sousa de Gois

Language identification is an important first step in many IR and NLP applications. Most publicly available language identification datasets, however, are compiled under the assumption that the gold label of each instance is determined by…

计算与语言 · 计算机科学 2023-03-03 Marcos Zampieri , Kai North , Tommi Jauhiainen , Mariano Felice , Neha Kumari , Nishant Nair , Yash Bangera

To advance the neural encoding of Portuguese (PT), and a fortiori the technological preparation of this language for the digital age, we developed a Transformer-based foundation model that sets a new state of the art in this respect for two…

Brazilian Portuguese and European Portuguese are two varieties of the same language and, despite their close similarities, they exhibit several differences. However, there is a significant disproportion in the availability of resources…

计算与语言 · 计算机科学 2024-08-15 João Sanches , Rui Ribeiro , Luísa Coheur

At present, different deep learning models are presenting high accuracy on popular inference datasets such as SNLI, MNLI, and SciTail. However, there are different indicators that those datasets can be exploited by using some simple…

计算与语言 · 计算机科学 2019-10-25 Felipe Salvatore , Marcelo Finger , Roberto Hirata

This paper reports on the development of a leaderboard of Open Large Language Models (LLM) for European Portuguese (PT-PT), and on its associated benchmarks. This leaderboard comes as a way to address a gap in the evaluation of LLM for…

计算与语言 · 计算机科学 2026-03-16 João Silva , Luís Gomes , António Branco

Despite Portuguese being one of the most spoken languages in the world, there is a lack of high-quality information retrieval datasets in that language. We present Quati, a dataset specifically designed for the Brazilian Portuguese…

This paper investigates morphosyntactic covariation in Brazilian Portuguese (BP) to assess whether dialectal origin can be inferred from the combined behavior of linguistic variables. Focusing on four grammatical phenomena related to…

计算与语言 · 计算机科学 2026-04-28 Manoel Siqueira , Raquel Freitag

To advance the neural decoding of Portuguese, in this paper we present a fully open Transformer-based, instruction-tuned decoder model that sets a new state of the art in this respect. To develop this decoder, which we named Gerv\'asio PT*,…

计算与语言 · 计算机科学 2024-03-06 Rodrigo Santos , João Silva , Luís Gomes , João Rodrigues , António Branco

Image captioning (IC) refers to the automatic generation of natural language descriptions for images, with applications ranging from social media content generation to assisting individuals with visual impairments. While most research has…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Gabriel Bromonschenkel , Alessandro L. Koerich , Thiago M. Paixão , Hilário Tomaz Alves de Oliveira

Significant advances have been made in natural language processing in recent years. However, our current deep learning approach to language modeling requires substantial resources in terms of data and computation. One of the side effects of…

计算与语言 · 计算机科学 2025-07-25 Nicholas Kluge Corrêa , Aniket Sen , Sophia Falk , Shiza Fatimah

The Natural Language Processing task of determining "Who did what to whom" is called Semantic Role Labeling. For English, recent methods based on Transformer models have allowed for major improvements in this task over the previous state of…

计算与语言 · 计算机科学 2021-11-02 Sofia Oliveira , Daniel Loureiro , Alípio Jorge

This work addresses the cross-corpora generalization issue for the low-resourced spoken language identification (LID) problem. We have conducted the experiments in the context of Indian LID and identified strikingly poor cross-corpora…

音频与语音处理 · 电气工程与系统科学 2023-03-02 Spandan Dey , Md Sahidullah , Goutam Saha

Text-to-image generation has made significant advancements with the introduction of text-to-image diffusion models. These models typically consist of a language model that interprets user prompts and a vision model that generates…

计算机视觉与模式识别 · 计算机科学 2024-03-13 Shihao Zhao , Shaozhe Hao , Bojia Zi , Huaizhe Xu , Kwan-Yee K. Wong

Linguistic diversity is a human attribute which, with the advance of generative AIs, is coming under threat. This paper, based on the contributions of sociolinguistics, examines the consequences of the variety selection bias imposed by…

计算与语言 · 计算机科学 2026-03-24 Raquel Meister Ko Freitag
‹ 上一页 1 2 3 10 下一页 ›