中文
相关论文

相关论文: Text Information Retrieval in Tetun: A Preliminary…

200 篇论文

Speech synthesis (text to speech, TTS) and recognition (automatic speech recognition, ASR) are important speech tasks, and require a large amount of text and speech pairs for model training. However, there are more than 6,000 languages in…

音频与语音处理 · 电气工程与系统科学 2020-08-11 Jin Xu , Xu Tan , Yi Ren , Tao Qin , Jian Li , Sheng Zhao , Tie-Yan Liu

In today's world anybody who wants to access any information the first choice is to use the web because it is the only source to provide easy and instant access to information. However web readers face many hurdles from web which includes…

人机交互 · 计算机科学 2011-06-09 Walayat Hussain , Osama Sohaib , Arif Ali

This paper presents techniques and findings for improving the performance of low-resource speech to text translation (ST). We conducted experiments on both simulated and real-low resource setups, on language pairs English - Portuguese, and…

计算与语言 · 计算机科学 2024-02-07 Santosh Kesiraju , Marek Sarvas , Tomas Pavlicek , Cecile Macaire , Alejandro Ciuba

Mapuzugun is the language of the Mapuche people. Due to political and historical reasons, its number of speakers has decreased and the language has been excluded from the educational system in Chile and Argentina. For this reason, it is…

计算与语言 · 计算机科学 2022-05-24 Cristian Ahumada , Claudio Gutierrez , Antonios Anastasopoulos

This study explores the use of large language models (LLMs) for translating English into Mambai, a low-resource Austronesian language spoken in Timor-Leste, with approximately 200,000 native speakers. Leveraging a novel corpus derived from…

计算与语言 · 计算机科学 2025-01-28 Raphaël Merx , Aso Mahmudi , Katrina Langford , Leo Alberto de Araujo , Ekaterina Vylomova

Despite Portuguese being one of the most spoken languages in the world, there is a lack of high-quality information retrieval datasets in that language. We present Quati, a dataset specifically designed for the Brazilian Portuguese…

Developing Automatic Speech Recognition (ASR) systems for Tunisian Arabic Dialect is challenging due to the dialect's linguistic complexity and the scarcity of annotated speech datasets. To address these challenges, we propose the LinTO…

计算与语言 · 计算机科学 2025-04-04 Hedi Naouara , Jean-Pierre Lorré , Jérôme Louradour

Large language models have made tremendous progress in recent years, but low-resource languages, like Tibetan, remain significantly underrepresented in their evaluation. Despite Tibetan being spoken by over seven million people, it has…

The complete freedom of expression in social media has its costs especially in spreading harmful and abusive content that may induce people to act accordingly. Therefore, the need of detecting automatically such a content becomes an urgent…

计算与语言 · 计算机科学 2021-10-12 Slim Gharbi , Heger Arfaoui , Hatem Haddad , Mayssa Kchaou

Social media enables data-driven analysis of public opinion on contested issues. Target-Stance Extraction (TSE) is the task of identifying the target discussed in a document and the document's stance towards that target. Many works classify…

计算与语言 · 计算机科学 2025-10-28 Ethan Mines , Bonnie Dorr

End-to-end text-to-speech (TTS) has shown great success on large quantities of paired text plus speech data. However, laborious data collection remains difficult for at least 95% of the languages over the world, which hinders the…

计算与语言 · 计算机科学 2019-07-03 Tao Tu , Yuan-Jui Chen , Cheng-chieh Yeh , Hung-yi Lee

This study introduces the continuous Educational Turkish Sign Language (E-TSL) dataset, collected from online Turkish language lessons for 5th, 6th, and 8th grades. The dataset comprises 1,410 videos totaling nearly 24 hours and includes…

计算与语言 · 计算机科学 2024-07-24 Şükrü Öztürk , Hacer Yalim Keles

We present a study of Tip-of-the-tongue (ToT) retrieval for music, where a searcher is trying to find an existing music entity, but is unable to succeed as they cannot accurately recall important identifying information. ToT information…

信息检索 · 计算机科学 2023-05-24 Samarth Bhargav , Anne Schuth , Claudia Hauff

This paper reports on the development of a text-to-speech (TTS) system for Mizo, a low-resource, tonal, and Tibeto-Burman language spoken primarily in the Indian state of Mizoram. The TTS was built with only 5.18 hours of data; however, in…

音频与语音处理 · 电气工程与系统科学 2026-01-06 Abhijit Mohanta , Remruatpuii , Priyankoo Sarmah , Rohit Sinha , Wendy Lalhminghlui

This paper reports some difficulties and some results when using dense retrievers on Amharic, one of the low-resource languages spoken by 120 millions populations. The efforts put and difficulties faced by University Addis Ababa toward…

信息检索 · 计算机科学 2025-03-25 Tilahun Yeshambel , Moncef Garouani , Serge Molina , Josiane Mothe

Micro-blogging through Twitter has made information short and to the point, and more importantly systematically searchable. This work is the first of a series in which quotidian observations about Tunisia are obtained using the…

社会与信息网络 · 计算机科学 2014-01-21 Meriem Ben-Salah Akin

While current information retrieval systems are effective for known-item retrieval where the searcher provides a precise name or identifier for the item being sought, systems tend to be much less effective for cases where the searcher is…

信息检索 · 计算机科学 2021-01-19 Jaime Arguello , Adam Ferguson , Emery Fine , Bhaskar Mitra , Hamed Zamani , Fernando Diaz

It is challenging to control the quality of online information due to the lack of supervision over all the information posted online. Manual checking is almost impossible given the vast number of posts made on online media and how quickly…

计算与语言 · 计算机科学 2022-03-16 Rini Anggrainingsih , Ghulam Mubashar Hassan , Amitava Datta

Dense retrieval is a basic building block of information retrieval applications. One of the main challenges of dense retrieval in real-world settings is the handling of queries containing misspelled words. A popular approach for handling…

The recent advances in natural language processing (NLP) are linked to training processes that require vast amounts of corpora. Access to this data is commonly not a trivial process due to resource dispersion and the need to maintain these…

计算与语言 · 计算机科学 2024-01-30 Rúben Almeida , Ricardo Campos , Alípio Jorge , Sérgio Nunes