中文
相关论文

相关论文: Text Information Retrieval in Tetun: A Preliminary…

200 篇论文

Text Simplification is a task that has been minimally explored for low-resource languages. Consequently, there are only a few manually curated datasets. In this paper, we present a human curated sentence-level text simplification dataset…

Open Information Extraction (Open IE) is the task of extracting structured information from textual documents, independent of domain. While traditional Open IE methods were based on unsupervised approaches, recently, with the emergence of…

计算与语言 · 计算机科学 2025-01-22 Marlo Souza , Bruno Cabral , Daniela Claro , Lais Salvador

Typhoon is a series of Thai large language models (LLMs) developed specifically for the Thai language. This technical report presents challenges and insights in developing Thai LLMs, including data preparation, pretraining,…

Sophisticated grammatical error detection/correction tools are available for a small set of languages such as English and Chinese. However, it is not straightforward -- if not impossible -- to adapt them to morphologically rich languages…

计算与语言 · 计算机科学 2024-10-17 Ali Gebeşçe , Gözde Gül Şahin

The proliferation of fake news has become a significant concern in recent times due to its potential to spread misinformation and manipulate public opinion. This paper presents a comprehensive study on detecting fake news in Brazilian…

Propinquity between Australian Indigenous communities' social structures and ICT purposed for cultural preservation is a modern area of research; historically hindered by the "digital divide" thus limiting plentiful literature and existing…

计算机与社会 · 计算机科学 2016-06-07 Sarah Van Der Meer , Stephen Smith , Vincent Pang

The digital exclusion of endangered languages remains a critical challenge in NLP, limiting both linguistic research and revitalization efforts. This study introduces the first computational investigation of Comanche, an Uto-Aztecan…

Tibetan, one of the major low-resource languages in Asia, presents unique linguistic and sociocultural characteristics that pose both challenges and opportunities for AI research. Despite increasing interest in developing AI systems for…

Multilingual text processing is useful because the information content found in different languages is complementary, both regarding facts and opinions. While Information Extraction and other text mining software can, in principle, be…

计算与语言 · 计算机科学 2014-01-14 Ralf Steinberger

English is the predominant language on the web, powering nearly half of the world's top ten million websites. Support for multilingual content is nevertheless growing, with many websites increasingly combining English with regional or…

计算与语言 · 计算机科学 2025-08-27 Masudul Hasan Masud Bhuiyan , Matteo Varvello , Yasir Zaki , Cristian-Alexandru Staicu

Collecting high-quality studio recordings of audio is challenging, which limits the language coverage of text-to-speech (TTS) systems. This paper proposes a framework for scaling a multilingual TTS model to 100+ languages using found data…

This paper introduces a high-quality open-source speech synthesis dataset for Kazakh, a low-resource language spoken by over 13 million people worldwide. The dataset consists of about 93 hours of transcribed audio recordings spoken by two…

音频与语音处理 · 电气工程与系统科学 2021-09-09 Saida Mussakhojayeva , Aigerim Janaliyeva , Almas Mirzakhmetov , Yerbolat Khassanov , Huseyin Atakan Varol

This work investigates the in-context learning abilities of pretrained large language models (LLMs) when instructed to translate text from a low-resource language into a high-resource language as part of an automated machine translation…

计算与语言 · 计算机科学 2024-10-28 Sara Court , Micha Elsner

Tip-of-the-tongue (ToT) known-item retrieval involves re-finding an item for which the searcher does not reliably recall an identifier. ToT information requests (or queries) are verbose and tend to include several complex phenomena, making…

信息检索 · 计算机科学 2026-01-29 Jaime Arguello , Fernando Diaz , Maik Fröebe , To Eun Kim , Bhaskar Mitra

In the era dominated by information overload and its facilitation with Large Language Models (LLMs), the prevalence of misinformation poses a significant threat to public discourse and societal well-being. A critical concern at present…

计算与语言 · 计算机科学 2024-11-05 Cem Üyük , Danica Rovó , Shaghayegh Kolli , Rabia Varol , Georg Groh , Daryna Dementieva

Traditional Text-to-Speech (TTS) systems rely on studio-quality speech recorded in controlled settings.a Recently, an effort known as noisy-TTS training has emerged, aiming to utilize in-the-wild data. However, the lack of dedicated…

Tunisians on social media tend to express themselves in their local dialect using Latin script (TUNIZI). This raises an additional challenge to the process of exploring and recognizing online opinions. To date, very little work has…

计算与语言 · 计算机科学 2020-10-15 Abir Messaoudi , Hatem Haddad , Moez Ben HajHmida , Chayma Fourati , Abderrazak Ben Hamida

Developing autonomous agents capable of performing complex, multi-step decision-making tasks specified in natural language remains a significant challenge, particularly in realistic settings where labeled data is scarce and real-time…

计算与语言 · 计算机科学 2025-06-09 Thomas Pouplin , Katarzyna Kobalczyk , Hao Sun , Mihaela van der Schaar

Etruscan is an ancient language spoken in Italy from the 7th century BC to the 1st century AD. There are no native speakers of the language at the present day, and its resources are scarce, as there exist only around 12,000 known…

计算与语言 · 计算机科学 2023-10-10 Gianluca Vico , Gerasimos Spanakis

Human computer conversation is regarded as one of the most difficult problems in artificial intelligence. In this paper, we address one of its key sub-problems, referred to as short text conversation, in which given a message from human,…

信息检索 · 计算机科学 2014-09-01 Zongcheng Ji , Zhengdong Lu , Hang Li