中文
相关论文

相关论文: Cem Mil Podcasts: A Spoken Portuguese Document Cor…

200 篇论文

Podcasts provide highly diverse content to a massive listener base through a unique on-demand modality. However, limited data has prevented large-scale computational analysis of the podcast ecosystem. To fill this gap, we introduce a…

计算与语言 · 计算机科学 2025-12-22 Benjamin Litterer , David Jurgens , Dallas Card

This paper introduces PublicHearingBR, a Brazilian Portuguese dataset designed for summarizing long documents. The dataset consists of transcripts of public hearings held by the Brazilian Chamber of Deputies, paired with news articles and…

Speech provides a natural way for human-computer interaction. In particular, speech synthesis systems are popular in different applications, such as personal assistants, GPS applications, screen readers and accessibility tools. However, not…

Sentiment Analysis is one of the most classical and primarily studied natural language processing tasks. This problem had a notable advance with the proposition of more complex and scalable machine learning models. Despite this progress,…

计算与语言 · 计算机科学 2021-12-13 Frederico Souza , João Filho

Brazilian Portuguese and European Portuguese are two varieties of the same language and, despite their close similarities, they exhibit several differences. However, there is a significant disproportion in the availability of resources…

计算与语言 · 计算机科学 2024-08-15 João Sanches , Rui Ribeiro , Luísa Coheur

In this paper, we introduce How2, a multimodal collection of instructional videos with English subtitles and crowdsourced Portuguese translations. We also present integrated sequence-to-sequence baselines for machine translation, automatic…

计算与语言 · 计算机科学 2018-12-10 Ramon Sanabria , Ozan Caglayan , Shruti Palaskar , Desmond Elliott , Loïc Barrault , Lucia Specia , Florian Metze

Speech technologies rely on capturing a speaker's voice variability while obtaining comprehensive language information. Textual prompts and sentence selection methods have been proposed in the literature to comprise such adequate phonetic…

计算与语言 · 计算机科学 2024-02-09 Marcellus Amadeus , William Alberto Cruz Castañeda , Wilmer Lobato , Niasche Aquino

Despite Portuguese being one of the most spoken languages in the world, there is a lack of high-quality information retrieval datasets in that language. We present Quati, a dataset specifically designed for the Brazilian Portuguese…

While large language models (LLMs) show transformative potential in healthcare, their development remains focused on high-resource languages. This creates a critical barrier for other languages, as simple translation fails to capture unique…

Despite significant advances in speech processing, Portuguese remains under-resourced due to the scarcity of public, large-scale, and high-quality datasets. To address this gap, we present a new dataset, named TAGARELA, composed of over…

In Brazil, the governmental body responsible for overseeing and coordinating post-graduate programs, CAPES, keeps records of all theses and dissertations presented in the country. Information regarding such documents can be accessed online…

计算与语言 · 计算机科学 2019-05-07 Felipe Soares , Gabrielli Harumi Yamashita , Michel Jose Anzanello

This paper investigates morphosyntactic covariation in Brazilian Portuguese (BP) to assess whether dialectal origin can be inferred from the combined behavior of linguistic variables. Focusing on four grammatical phenomena related to…

计算与语言 · 计算机科学 2026-04-28 Manoel Siqueira , Raquel Freitag

This paper reports on the development of a leaderboard of Open Large Language Models (LLM) for European Portuguese (PT-PT), and on its associated benchmarks. This leaderboard comes as a way to address a gap in the evaluation of LLM for…

计算与语言 · 计算机科学 2026-03-16 João Silva , Luís Gomes , António Branco

Language models have become foundational to many widely used systems. However, these seemingly advantageous models are double-edged swords. While they excel in tasks related to resource-rich languages like English, they often lose the fine…

计算与语言 · 计算机科学 2025-02-21 Hugo Sousa , Satya Almasian , Ricardo Campos , Alípio Jorge

Analyses of legislative behavior often rely on voting records, overlooking the rich semantic and rhetorical content of political speech. In this paper, we ask three complementary questions about parliamentary discourse: how things are said,…

A new algorithm for voice automatic syllabic splitting in the Portuguese language is proposed, which is based on the envelope of the speech signal of the input audio file. A computational implementation in MatlabTM is presented and made…

声音 · 计算机科学 2018-01-24 E. L. F. Da Silva , H. M. de Oliveira

Multilingual discourse parsing is a very prominent research topic. The first stage for discourse parsing is discourse segmentation. The study reported in this article addresses a review of two on-line available discourse segmenters (for…

信息检索 · 计算机科学 2020-05-04 Iria da Cunha , Juan-Manuel Torres-Moreno

This paper presents and makes publicly available the NILC-Metrix, a computational system comprising 200 metrics proposed in studies on discourse, psycholinguistics, cognitive and computational linguistics, to assess textual complexity in…

Discourse parsing is an integral part of understanding information flow and argumentative structure in documents. Most previous research has focused on inducing and evaluating models from the English RST Discourse Treebank. However,…

计算与语言 · 计算机科学 2017-01-12 Chloé Braud , Maximin Coavoux , Anders Søgaard

Municipal meeting minutes are formal records documenting the discussions and decisions of local government, yet their content is often lengthy, dense, and difficult for citizens to navigate. Automatic summarization can help address this…

‹ 上一页 1 2 3 10 下一页 ›