中文
相关论文

相关论文: Text Information Retrieval in Tetun: A Preliminary…

200 篇论文

With the constant growth of the World Wide Web and the number of documents in different languages accordingly, the need for reliable language detection tools has increased as well. Platforms such as Twitter with predominantly short texts…

计算与语言 · 计算机科学 2016-08-31 Ivana Balazevic , Mikio Braun , Klaus-Robert Müller

Low-resource languages such as isiZulu and isiXhosa face persistent challenges in machine translation due to limited parallel data and linguistic resources. Recent advances in large language models suggest that self-reflection, prompting a…

计算与语言 · 计算机科学 2026-01-28 Nicholas Cheng

Pronunciation modeling is a key task for building speech technology in new languages, and while solid grapheme-to-phoneme (G2P) mapping systems exist, language coverage can stand to be improved. The information needed to build G2P models…

计算与语言 · 计算机科学 2021-01-28 Tania Chakraborty , Manasa Prasad , Theresa Breiner , Sandy Ritchie , Daan van Esch

Indigenous languages of the American continent are highly diverse. However, they have received little attention from the technological perspective. In this paper, we review the research, the digital resources and the available NLP systems…

计算与语言 · 计算机科学 2018-06-13 Manuel Mager , Ximena Gutierrez-Vasques , Gerardo Sierra , Ivan Meza

Among the pressing issues facing Australian and other First Nations peoples is the repatriation of the bodily remains of their ancestors, which are currently held in Western scientific institutions. The success of securing the return of…

计算与语言 · 计算机科学 2023-03-28 Md Abul Bashar , Richi Nayak , Gareth Knapman , Paul Turnbull , Cressida Fforde

There is growing interest in ASR systems that can recognize phones in a language-independent fashion. There is additionally interest in building language technologies for low-resource and endangered languages. However, there is a paucity of…

计算与语言 · 计算机科学 2021-04-05 David R. Mortensen , Jordan Picone , Xinjian Li , Kathleen Siminyu

Web search engines focus on serving highly relevant results within hundreds of milliseconds. Pre-trained language transformer models such as BERT are therefore hard to use in this scenario due to their high computational demands. We present…

信息检索 · 计算机科学 2021-12-06 Matěj Kocián , Jakub Náplava , Daniel Štancl , Vladimír Kadlec

With the emergence of automatic speech recognition (ASR) models, converting the spoken form text (from ASR) to the written form is in urgent need. This inverse text normalization (ITN) problem attracts the attention of researchers from…

计算与语言 · 计算机科学 2023-01-25 Szu-Jui Chen , Debjyoti Paul , Yutong Pang , Peng Su , Xuedong Zhang

Spoken Term Detection (STD) is the task of searching for words or phrases within audio, given either text or spoken input as a query. In this work, we use state-of-the-art Hindi, Tamil and Telugu ASR systems cross-lingually for lexical…

计算与语言 · 计算机科学 2020-11-13 Sanket Shah , Satarupa Guha , Simran Khanuja , Sunayana Sitaram

The internet is rife with unattributed, deliberately misleading, or otherwise untrustworthy content. Though large language models (LLMs) are often tasked with autonomous web browsing, the extent to which they have learned the simple…

计算与语言 · 计算机科学 2025-08-08 Gustaf Ahdritz , Anat Kleiman

Most data-to-text datasets are for English, so the difficulties of modelling data-to-text for low-resource languages are largely unexplored. In this paper we tackle data-to-text for isiXhosa, which is low-resource and agglutinative. We…

计算与语言 · 计算机科学 2024-03-13 Francois Meyer , Jan Buys

Nowadays, Information spreads at an unprecedented pace in social media and discerning truth from misinformation and fake news has become an acute societal challenge. Machine learning (ML) models have been employed to identify fake news but…

计算与语言 · 计算机科学 2024-05-08 Jasraj Singh , Fang Liu , Hong Xu , Bee Chin Ng , Wei Zhang

Direct speech-to-speech translation (S2ST) is among the most challenging problems in the translation paradigm due to the significant scarcity of S2ST data. While effort has been made to increase the data size from unlabeled speech by…

计算与语言 · 计算机科学 2022-10-27 Xuan-Phi Nguyen , Sravya Popuri , Changhan Wang , Yun Tang , Ilia Kulikov , Hongyu Gong

Institutions dependent on IT services and resources acknowledge the crucial significance of an IT help desk system, that act as a centralized hub connecting IT staff and users for service requests. Employing various Machine Learning models,…

信息检索 · 计算机科学 2025-08-11 Leonardo Santiago Benitez Pereira , Robinson Pizzio , Samir Bonho

We report findings of the TSAR-2022 shared task on multilingual lexical simplification, organized as part of the Workshop on Text Simplification, Accessibility, and Readability TSAR-2022 held in conjunction with EMNLP 2022. The task called…

计算与语言 · 计算机科学 2023-02-07 Horacio Saggion , Sanja Štajner , Daniel Ferrés , Kim Cheng Sheang , Matthew Shardlow , Kai North , Marcos Zampieri

Korea University Intelligent Signal Processing Lab. (KU-ISPL) developed speaker recognition system for SRE16 fixed training condition. Data for evaluation trials are collected from outside North America, spoken in Tagalog and Cantonese…

声音 · 计算机科学 2017-02-07 Suwon Shon , Hanseok Ko

The performance of automated speech recognition (ASR) systems is well known to differ for varied application domains. At the same time, vendors and research groups typically report ASR quality results either for limited use simplistic…

The Serbian language is a Slavic language spoken by over 12 million speakers and well understood by over 15 million people. In the area of natural language processing, it can be considered a low-resourced language. Also, Serbian is…

计算与语言 · 计算机科学 2023-04-13 Ulfeta A. Marovac , Aldina R. Avdić , Nikola Lj. Milošević

In our era of widespread false information, human fact-checkers often face the challenge of duplicating efforts when verifying claims that may have already been addressed in other countries or languages. As false information transcends…

计算与语言 · 计算机科学 2025-09-25 Ivan Vykopal , Matúš Pikuliak , Simon Ostermann , Tatiana Anikina , Michal Gregor , Marián Šimko

This research introduces a comprehensive Bahasa text-to-speech (TTS) dataset and a novel TTS model, EnGen-TTS, designed to enhance the quality and versatility of synthetic speech in the Bahasa language. The dataset, spanning…

声音 · 计算机科学 2024-10-10 Onkar Kishor Susladkar , Vishesh Tripathi , Biddwan Ahmed
‹ 上一页 1 8 9 10 下一页 ›